Direct Perception Control Model real-robot asynchronous inference architecture Live onboard RGB and task language enter the perception side of a unified model to form shared scene context. A jointly optimized motion expert combines that context with live proprioceptive history and the previous action to continuously generate whole-body joint and hand targets. Vision refreshes the world context asynchronously while proprioception closes the fast physical feedback loop through a PD controller on a Unitree G1 humanoid. Symbiotic Attention denotes co-adaptation during training while inference remains causally ordered from context to motion targets. LIVE ROBOT RGB TASK LANGUAGE Walk forward, grab the cola and throw into the trash bin. ROBOT STATE DIRECT PERCEPTION CONTROL MODEL (DPC) LIVE PHYSICAL QUERY PERCEPTION EXPERT World Context Prediction Scene memory from vision and language MOTION EXPERT Motion Target Generation Live state correction along action trajectories SYMBIOTIC ATTN ACTION TRAJECTORY Joint + hand targets PD WHOLE-BODY EXECUTION