SOURCE // NEWS

Memo and Suochen Tech Release Physical-WAM and RoboTwin-Phys Benchmark

Memo and Suochen Tech Release Physical-WAM and RoboTwin-Phys Benchmark

Currently, the commercialization of embodied AI is hitting a bottleneck. The BeTTER benchmark jointly released by Peking University and BeingBeyond highlights that state-of-the-art VLA models (Vision-Language-Action models) suffer a steep drop in success rates when facing spatial layout offsets. Additionally, a Stanford study identified two independent failure modes in contact-rich tasks: "accuracy failure" and "force failure." Essentially, current #VLA-driven robots are statistical imitators rather than physical actors, failing to comprehend real-world physical properties such as wall thickness, friction, or gravity shifts.

To address this challenge, at the 5th Global Digital Trade Expo, Memo, incubated by the AVS Research Center under Academician Gao Wen, together with its strategic investor Suochen Technology (688507.SH), launched two major achievements: Physical-WAM (Physical World-Action Model) and the RoboTwin-Phys physical drift evaluation benchmark. In August 2026, Suochen Technology, a leader in physical AI, completed a seed-round strategic investment in Memo to accelerate the joint development and industrial deployment of physical world models.

Robot training has long been restricted by data scarcity. Compared to billions of samples in NLP and CV, robotic physical interaction data remains at the million-level. Physical data collection is costly, hardware-intensive, and hard to transfer across different robot embodiments. In response, Memo introduced Physical-WAM, introducing a "Physical Token" representation layer between perceptual inputs and action outputs. This design allows the model not just to predict "the next frame," but to understand the physical mechanisms behind it.

Specifically, Physical-WAM consists of three modules: PhysLens translates multimodal inputs into unobservable physical properties (like friction, weight, and center of gravity) as physical tokens; PhysDream predicts physical risks like slippage or deformation; and PhysAct dynamically corrects actions based on these physical conditions. For instance, when a robot grasps a cup while filling it with water, the system senses the changing weight and automatically adjusts its grip force and position, demonstrating human-like physical intuition.

Beyond the model, Memo launched RoboTwin-Phys. Conventional benchmarks measure task completion rather than physical adaptation, showing poor robustness under parameter perturbations. Built on RoboTwin, RoboTwin-Phys introduces 13 types of physical perturbations across five dimensions: object properties, contact properties, damping, environment, and camera settings. By providing ground-truth data and paired evaluations, it accurately diagnoses failures in perception, prediction, or action execution, establishing a rigorous metric for embodied AI iteration.

[AgentUpdate Depth Analysis] Embodied AI agents have long been constrained by the "Sim-to-Real" gap. Traditional VLA models, such as Google's RT-2, rely heavily on end-to-end mapping from pixels to actions, lacking an explicit understanding of physical dynamics. The introduction of Physical-WAM by Memo and Suochen Technology represents a paradigm shift from "visual imitation" to "physical intuition." By decoupling key physical variables like friction and mass, the model enables zero-shot generalization in unfamiliar physical environments. Integrating classical physics priors with deep learning representations not only lowers the data collection bottleneck for #robotics but also paves the way for fully autonomous L4/L5 physical agents. This hybrid approach will significantly accelerate the evolution of the broader AI Agent ecosystem in physical-world applications.