REAL-ROBOT INFERENCE
Direct Perception Control Model real-robot asynchronous inference architecture
Live onboard RGB and task language enter the perception side of a unified model to form shared scene context. A jointly optimized motion expert combines that context with live proprioceptive history and the previous action to continuously generate whole-body joint and hand targets. Vision refreshes the world context asynchronously while proprioception closes the fast physical feedback loop through a PD controller on a Unitree G1 humanoid. Symbiotic Attention denotes co-adaptation during training while inference remains causally ordered from context to motion targets.
INPUTS
MODEL
EXECUTE
ONBOARD VISION · ASYNCHRONOUS WORLD UPDATE
LIVE ROBOT RGB
TASK LANGUAGE
Walk forward, grab the cola
and throw into the trash bin.
ROBOT STATE
DIRECT PERCEPTION CONTROL MODEL (DPC)
LIVE PHYSICAL QUERY
PERCEPTION EXPERT
World Context Prediction
Scene memory from vision and language
MOTION EXPERT
Motion Target Generation
Live state correction along action trajectories
SYMBIOTIC ATTN
ACTION TRAJECTORY
Joint + hand targets
PD
WHOLE-BODY EXECUTION
LIVE PROPRIOCEPTION · CONTINUOUS PHYSICAL FEEDBACK