According to a recent filing on the HKEX, Zhipu AI announced a major placement of new H shares and the issuance of zero-coupon convertible bonds worth RMB 20.14 billion, expected to raise net proceeds of approximately HK$39.3 billion (RMB 33.5 billion). This marks the third massive funding round for the company in less than nine months since its listing, bringing total accumulated funds to HK$75.5 billion.
The company disclosed that approximately 60% (about HK$23.5 billion) of the proceeds will be allocated to the R&D of the next-generation GLM foundation models and the "Fully Self-Training" architecture. #Zhipu defines this as a closed loop of Recursive Self-Improvement (RSI), where the next-gen model is trained within the environments built by its predecessors.
The "Fully Self-Training" paradigm focuses on three dimensions. The first is Data Self-Production, which reduces reliance on human labeling. By utilizing self-play, rule validation, execution checks, and joint model evaluations, Zhipu aims to automatically synthesize and filter high-quality data for pre-training and post-training stages.
The second dimension is Environment Self-Generation, allowing AI agents to harvest and transform real-world tasks. These agents will run test tasks, generate verification tools, and evaluate solvability, creating scalable training environments that feed directly into long-horizon tasks and reinforcement learning pipelines.
The third is Infrastructure Self-Optimization. Leveraging the model's advanced capabilities in coding and systems engineering, the system will autonomously optimize the kernels, operators, schedulers, and caching within its training and inference stack, facilitating self-guided infrastructure upgrades.
In addition to RSI, Zhipu will target deeper effective compute depth without escalating inference costs, aiming to bolster native multimodal capabilities and long-chain reasoning. Furthermore, reinforcement learning for long-horizon tasks will remain a priority to train models in real-world professional task decomposing, tool calling, and error recovery.
Regarding other funds, 15% will go toward strategic expansion and mergers, specifically focusing on AI agents and enterprise applications. Notably, the filing dropped clues about Zhipu’s A+H listing ambitions, specifying that bond terms would not be altered due to a prospective IPO on Shanghai’s STAR Market.
[AgentUpdate Depth Analysis] Zhipu's "Fully Self-Training" blueprint directly targets the ultimate bottleneck of AI Agent deployment: data scarcity and long-horizon planning. Compared to OpenAI's o1/o3 reasoning path and Anthropic's computer use paradigm, Zhipu's framework is unique in its emphasis on an entirely closed-loop, self-evolving system. By combining self-generating environments with infrastructure self-optimization, Zhipu is pioneering an "Agent Flywheel" that operates with minimal human intervention. This shift—handing the keys of system-level performance tuning over to the AI itself—represents a profound paradigm shift in software engineering. Ultimately, the future competitiveness of AI Agents will rely less on brute-force hardware scaling and more on the acceleration rate of their autonomous self-evolution in simulated environments.