The infrastructure for AI agents is rapidly evolving to support systems that continuously cycle through inference, feedback, and training. Cognition AI Inc.'s Devin now plays a significant role across the entire software development lifecycle, from planning and writing code to reviewing it and responding to production issues. This expanding role necessitates AI agent infrastructure capable of supporting continual learning at scale.
Silas Alberti, head of research and a founding team member at Cognition AI, emphasized the "always-on" nature of their operations. "While we still ship releases and package them up, the reality is we’re always training. We’re always trying to find the next data and the next reward signals to improve our models," he stated, highlighting the continuous improvement loop.
Alberti, alongside Chen Goldberg, executive vice president of product and engineering at CoreWeave Inc., discussed #continuous learning, distributed AI training, infrastructure reliability, and CoreWeave Forge during an exclusive interview with theCUBE Research’s Dave Vellante and John Furrier at the Fully Connected event.
This "always-on" training approach significantly raises the bar for infrastructure reliability. Alberti explained that training and inference are increasingly intertwined, particularly in reinforcement learning workloads where generated responses are used to refine models. Cognition AI has distributed its training across data centers in multiple countries and continents, making uptime across thousands of graphics processing units (GPUs) absolutely essential. He stressed, "If you run a big training run, you’re not just looking at the GPUs, but you also want uptime. If just one replica goes down, the whole training run goes down. Achieving 99.99% reliability is a very important metric."
New hardware also influences Cognition AI's research direction, according to Alberti. Early access to Nvidia Corp.'s Vera Rubin platform allows their researchers to study the system and adapt model architectures, aiming for better price-performance. "With each generation, price performance just goes up, so we can do more with the same amount of compute," Alberti noted. "Being early and actually being able to study the kernels and the dynamics of this new platform allows us to prioritize our research investments."
Announced during the event, CoreWeave Forge is a platform designed to connect the various stages of the AI loop: inference, observation, data curation, model improvement, and evaluation. It includes Agent Lens for tracing agent activity, along with model distillation and reinforcement learning capabilities. Its RL Rollouts service can hot-load updated model checkpoints into a live deployment without requiring a full system redeployment.
[AgentUpdate Depth Analysis] The collaboration between #Cognition AI and #CoreWeave, particularly their "always-on" training paradigm and deep integration with CoreWeave Forge, signifies a crucial shift in the AI Agent ecosystem from static model deployments to dynamic, continuously learning systems. Unlike frameworks such as LangChain or CrewAI that focus on agent orchestration and tool use at the application layer, Cognition AI's Devin and its underlying infrastructure emphasize the agent's inherent ability for continuous evolution. This addresses the challenge of models becoming "stale" in new environments by enabling real-time feedback and self-optimization. This paradigm shift holds transformative potential for applications requiring high autonomy and adaptability, such as complex system management or adaptive learning platforms. The emergence of specialized infrastructure providers like CoreWeave is pivotal in facilitating this next generation of AI Agents, transforming them from mere assistants into truly intelligent, evolving entities capable of independent thought and adaptation.



