Nutanix Pushes Two-Tier Model for Agentic AI Costs

Nutanix will showcase its vision for enterprise agentic AI at AMD Advancing AI 2026, demonstrating how organisations can build scalable AI infrastructure while improving cost control, governance and data sovereignty as autonomous agents move from experimentation into production.

As enterprise AI moves to production-scale deployment, autonomous agents are increasingly retrieving context, interacting with enterprise tools and continuously executing complex workflows – a shift Nutanix says is driving new challenges around infrastructure scalability, governance and the economics of AI. With agents consuming and generating millions of tokens daily to perform routing, validation and execution tasks, the cost of rented infrastructure is becoming a bottleneck to scale.

Reserving frontier models for complex reasoning

Nutanix is proposing a two-tier intelligence strategy that reserves frontier AI models for complex reasoning while routing high-volume workloads to optimised models running on private infrastructure. Using an agent gateway, organisations can intelligently route requests, enforce policy, manage access and improve control over AI costs. At the AMD event, Nutanix will demonstrate the approach running on the Nutanix Cloud Platform and Nutanix Enterprise AI, combined with AMD EPYC processors and AMD Instinct MI355X accelerators.

  • Running long-lived AI agents while managing token costs
  • Supporting data sovereignty by keeping AI workloads within governed hybrid multicloud environments
  • Maximising infrastructure utilisation through shared inference infrastructure powered by AMD EPYC CPUs and AMD Instinct MI355X GPUs

“Owning your intelligence doesn’t mean completely abandoning frontier models; it means taking control of your routing, your volume, and your costs,” said Debo Dutta, Chief AI Officer, Nutanix. “This allows you to leverage expensive frontier models for a small fraction of tasks that require complex, edge-case reasoning, while the vast majority of your high-volume tasks are routed to highly optimised models running securely on your private infrastructure.”

The approach reflects a wider shift among infrastructure vendors towards helping enterprises manage the compounding costs of agentic AI at scale, as production deployments generate far higher and more continuous token volumes than earlier generative AI pilots.

Author


Discover more from techcoffeehouse.com

Subscribe to get the latest posts sent to your email.

Use promo code “TCH15” to get 15% off on checkout.

Share your thoughts

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from techcoffeehouse.com

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from techcoffeehouse.com

Subscribe now to keep reading and get access to the full archive.

Continue reading