Site icon techcoffeehouse.com

NetApp Unveils Novus Storage for AI Factories

Advertisements

NetApp has launched NetApp Novus, a storage architecture built for AI infrastructure providers running large GPU environments, designed to exceed 100 terabytes per second of aggregate throughput.

NetApp said neoclouds and GPU-as-a-service providers are building AI factories at a scale that traditional storage architectures were never designed to support. Under those older architectures, GPU utilisation can drop below 30 per cent when a data centre cannot feed its GPUs fast enough to generate a return on the investment.

“AI factories struggle and GPU economics collapse when data can’t keep up,” said Syam Nair, Chief Product Officer at NetApp. “HPC-era storage alone cannot serve AI Factory scale — the industry needs a true architecture built for it from the ground up. AI factories have long been searching for infrastructure that can scale at AI speed… We’ve just delivered the fastest storage on the planet, that is purpose built for the AI Era.”

Separating metadata from the data path

Novus is designed to eliminate scale limitations by separating metadata from the data path, so performance, capacity and concurrency can scale independently under a single namespace using industry-standard NFS. The initial release combines NetApp Novus Data Director, the metadata software, running on qualified Supermicro infrastructure, with NetApp ONTAP data services delivered through AFF A90 systems. NetApp said the architecture is built to exceed 100 TB/s throughput and support high utilisation across hundreds of thousands of GPUs, while preserving the enterprise data services customers already rely on from ONTAP. It’s orderable now, with a path to software-defined deployments over time.

Independent validation

Omdia said its testing found NetApp Novus demonstrated near-linear performance scaling as ONTAP clusters were added to a single global namespace. “Based on the rigorous testing we observed, our modeling projects that NetApp Novus disaggregated metadata and data architecture can scale up and out to 100TBps sequential read throughput with dozens of exabytes of effective capacity, and beyond,” said Tony Palmer, Chief Analyst with Omdia.

Mike Leone, Vice President and Principal Analyst at Moor Insights & Strategy, said storage architectures that can’t deliver data at cloud scale can become a limiting factor for training performance, tenant isolation and return on infrastructure investment. “NetApp Novus directly addresses that limit by combining disaggregated metadata management with high-performance data services, giving neocloud and GPU-as-a-service providers a more scalable foundation for AI factory operations,” he said.

Cisco is also involved: “Our longstanding collaboration with NetApp is about bringing compute, networking and data together to make complex infrastructure simpler for customers,” said Jeremy Foster, Senior Vice President and General Manager, Cisco Compute. “As AI pushes infrastructure to new levels of scale and speed, the real differentiation isn’t the GPU — it’s the data and how customers use it.”

Author

Exit mobile version