SOURCE // NEWS

NetApp & Nvidia Redefine AI Factory Storage with Novus Architecture, Separating Data & Metadata

NetApp & Nvidia Redefine AI Factory Storage with Novus Architecture, Separating Data & Metadata

The storage architecture for AI factories is undergoing a fundamental rewrite. Traditional enterprise storage systems were designed for workloads that scaled predictably. However, artificial intelligence (AI) fundamentally alters this equation by combining heavy data movement with transactional metadata activity, often on shared infrastructure, posing significant challenges.

NetApp Inc. is addressing these pressures through its collaboration with Nvidia Corp., a partnership spanning over a decade that now includes co-engineering efforts for AI infrastructure. According to Arindam Banerjee, #NetApp’s Chief Platform and Technology Officer, NetApp’s Novus architecture separates data and metadata functions, allowing each to scale independently based on diverse workload demands while ensuring GPU resources are continuously fed with data.

Banerjee explained, “Different types of workloads, such as transactional workloads for metadata and heavy sequential workloads for, say, checkpoints, now happen simultaneously. Previous architectures excelled at one or the other, but never managed to combine and do both efficiently at the same time. This is what AI factories and their workloads are pushing us to achieve regarding data management.”

Banerjee and Jason Hardy, #Nvidia's Vice President of Storage Technology, discussed how AI factories and AI Agents are forcing storage systems to scale and operate differently during an exclusive interview on theCUBE at NetApp INSIGHT.

This architectural shift, particularly the separation of metadata from data management, aims to prevent smaller transactional operations from competing with large data transfers for the same resources. This is crucial because storage delays can leave expensive GPUs underutilized, even while they continue consuming power, making efficiency a more significant part of the AI scaling equation, Banerjee clarified.

He elaborated, “We can scale each of them independently on different axes. If metadata and data shared the same resources, media, and network, it resulted in an inefficient system. Data operations would be queued behind metadata operations, which are small and transactional. This not only left your GPUs underutilized but also consuming power. To restore efficiencies, maximize our ecosystems, and keep GPUs fed, we needed this architectural shift.”

AI infrastructure also needs to scale without forcing enterprises to redesign the system every time a new use case comes online. Hardy emphasized that the goal is greater flexibility across fine-tuning, inference, capacity, and AI-ready data as production workloads evolve and demands shift across different parts of the infrastructure.

[AgentUpdate Depth Analysis]

The innovation by NetApp and Nvidia in AI storage holds profound implications for the AI Agent ecosystem. Agents execute complex tasks requiring interaction with diverse data sources, including model calls, context management, memory retrieval, and real-time processing. Traditional storage systems bottleneck performance due to agents' high-speed, concurrent, and mixed data access patterns. For instance, an agent’s memory and tool invocation within frameworks like LangChain would suffer significantly if underlying storage cannot efficiently process frequent metadata requests (e.g., retrieving memory fragments) separately from large data streams (e.g., loading model weights).

Unlike single-point solutions like vector databases or traditional file systems, the Novus architecture optimizes data flow from the foundational hardware level. By independently scaling metadata and data paths, it ensures AI Agents can simultaneously execute multiple tasks, quickly making decisions (metadata-intensive) while smoothly processing large data volumes (data-intensive). This empowers more efficient real-time inference, model fine-tuning, and interaction with large datasets, fostering complex and autonomous agent behaviors. Such an efficient storage foundation will be critical for achieving general intelligence, enhancing decision accuracy, and accelerating learning in future #AI agents, evolving them into adaptive, general-purpose intelligent entities.