← Back to Top

Direct Perception Control Model

·SYMBIOSIS RESEARCH

Scaling perception helps robots understand the world, but understanding a task does not automatically make action reliable. We want to move the boundary of an intelligent model from generating a kinematic goal that must be tracked to generating physical action that depends on the robot’s current bodily state and can be executed directly.

TABLE OF CONTENTS

01 / PRELIMINARY

Revisiting Hierarchical Perception Control Paradigm

Most whole-body systems still use a modular architecture: System 1 understands the task and generates a high-level motion goal, while System 0 tracks that goal from the current body state.12 This makes existing controllers reusable, but introduces three problems across the complete Perception-to-Control execution path. First, an insufficient motion interface can discard task-relevant control information when System 1 compresses the complete context into Zt. Second, under separate training, coupled inference, System 1 and System 0 are optimized by different objectives even though they must act as one chain at deployment. Third, a low-level action boundary confines executable action to the control distribution learned by the frozen System 0.

Figure 01. The two-layer architecture maps complete task context through a Latent Policy and motion interface into a Whole-Body Tracker. Representation compression, fragmented optimization, and the low-level motion prior produce an insufficient motion interface, separate training with coupled inference, and a low-level action boundary.

Problem 1: Insufficient Motion Interface

A motion goal specifies what the body should do, while the tracker combines Zt with the current robot state to produce state-dependent Joint Targets. The limitation is not that the tracker lacks bodily state. It is that scaling only the goal generator improves task understanding and goal expression without expanding the tracker's physical action mapping across contact, balance, and dynamic conditions. Additional perceptual and reasoning capacity therefore still stops at the motion interface.

Figure 02 tracks the same samples from Latent Space to Joint Target Space. In Compare mode, latent-near pairs are matched by sample ID, and the distance panel measures their separation after state-conditioned action mapping. The larger separation in whole-body loco-manipulation is not a tracker error; it shows that decisive choices about contact, balance, and upper–lower-body coordination occur after the interface. Scaling only the upper-level goal generator leaves this mapping—and its physical capability boundary—unchanged.34

Figure 02. Pair Explorer tracks the same real samples from Latent Space to Joint Target Space, revealing how similar motion intents can correspond to different physical actions.

Problem 2: Separate Training, Coupled Inference

A hierarchical architecture appears as one continuous path during inference, but this continuity does not exist in its training objective. System 1's Latent Policy learns the motion interface, while System 0 independently learns tracking, stability, or control reward. They are connected through Zt only at deployment, so execution error and visual consequences cannot pass through the frozen interface to co-optimize System 1.5

Figure 03. The Latent Policy and Whole-Body Tracker receive latent-level and control-level supervision separately during training, then are chained through motion interface Zt at inference; the complete execution path therefore has no unified end-to-end action objective.

Problem 3: Low-Level Action Boundary

Final action is constrained not only by decoder form but also by the motion distribution covered by System 0's training data. If the low-level prior is learned primarily from locomotion-centric behavior, new loco-manipulation tasks at the high level must still project back into that execution distribution; joint-level actions outside it do not become reachable simply through stronger perception or reasoning.

Figure 04 uses Locomotion and Manipulation activation as its horizontal coordinates and relative sample density as height, forming a landscape of the two motion distributions. The locomotion-centric distribution remains concentrated around locomotion-heavy behavior, whereas loco-manipulation extends into mixed-activation regions. Their overlap marks the shared motion prior; the non-overlapping regions expose actions required by new tasks but absent from System 0's learned support.

Figure 04. Whole-body locomotion-centric data covers only a limited whole-body execution region, while whole-body loco-manipulation extends into areas where locomotion and manipulation are tightly coupled, exposing the distributional boundary of a frozen low-level prior.

02 / METHOD

DIRECT PERCEPTION CONTROL MODEL

The Direct Perception Control Model redesigns the execution path at the action interface, representation coupling, and closed-loop training. It removes the intermediate motion representation entirely: instead of producing a reference for a separate Whole-Body Tracker, one model maps vision, language, body state, action history, and execution feedback directly to executable joint and hand targets. Symbiotic Attention lets perceptual and control representations attend to each other under a shared action objective, so task understanding is constrained by action feasibility while action generation continuously draws on semantic context.

Figure 05. The Direct Perception Control Model maps perception and live body context directly to joint and hand targets without an intermediate motion interface. At its center, Symbiotic Attention couples perceptual and control representations during training, while asynchronous vision updates and continuous proprioceptive feedback sustain closed-loop execution.

For the Low-Level Action Boundary, DriftDistill exposes the Student to drifted states during closed-loop rollouts and uses Frozen Teacher corrections to learn recovery. Offline BC provides initialization; Stage 2 combines online rollout states with offline demonstrations, using repeated correction and distillation to progressively cover the states and recovery paths encountered at deployment. The intermediate interface is removed, perception and control are learned jointly, and the executable distribution expands through closed-loop training.

Figure 06. DriftDistill first initializes the Student with Offline BC. Stage 2 then jointly trains on drifted states visited by online rollouts and offline demonstrations, with the Frozen Teacher providing recovery targets so the Student can handle state shifts and recovery paths encountered at deployment.

03 / DATA

Boosting Loco-Manipulation Data

Direct physical-action learning requires supervision in the Joint Target Space where execution occurs. Existing human and robot corpora cannot be combined directly because their embodiments, action representations, and control interfaces differ.67 We therefore convert every source into G1-executable, time-aligned joint trajectories, then scale this common supervision to 15,010 hours.

Unifying Heterogeneous Motion Evidence

Conversion uses the most reliable motion signal available from each source. Teleoperated humanoid data are standardized across coordinate frames, joint definitions, and control rates. Head-mounted recordings are reconstructed and retargeted from recoverable whole-body motion. Egocentric video is lifted from hand, wrist, and end-effector trajectories into humanoid whole-body motion. All three paths produce the same joint-target supervision.

Figure 07. The heterogeneous data conversion pipeline uses the reliable motion information preserved in each source to produce unified, G1-executable, time-aligned joint trajectories.

Scaling the Unified Corpus

The resulting corpus contains 15,010 hours: 6,781 hours of Human Ego, 4,024 hours of Armed Robot, 3,660 hours of Wheeled Humanoid, and 545 hours of Bipedal Humanoid data. These sources span human and robot embodiments with different mobility and interaction capabilities8910; after conversion, all of them enter one execution space without tying training to a single platform.

Figure 08. The 15,010-hour embodied training corpus spans Human Ego data and multiple robot embodiments, providing a scalable foundation for whole-body loco-manipulation.

04 / PERFORMANCE

Performance

Mobile pick-and-place

The robot approaches the cone, bends and squats to grasp it with one hand, turns, walks two steps, and places it down—coupling grasping, loaded balance, turning, and transport.

Constrained whole-body loco-manipulation

Within a constrained posture and tight clearance, the torso and arms reconfigure while preserving stable support and precise object contact.

Hand–eye–foot coordination

The right foot controls the accelerator while the hands steer near the corner exit, combining visual timing, precise throttle control, and driving semantics.

05 / OUR VISION

Making Intelligence Physically Executable

Scaling perception has helped robots read the world and understand tasks. True Whole-Body Intelligence must also understand the task objective, bodily state, dynamic constraints, and the consequences of action. Direct Perception Control aims to shorten this Embodiment Gap. We are not training a stronger Motion Planner; we are training a system that learns continuously from understanding the world to using its own body to complete a task.

The next frontier of robot intelligence is not only better perception or better planning. It is learning how intelligence becomes physically executable action.

Citation

@article{
  symbiosis2026directperceptioncontrol,
  author = {Symbiosis Robotics Team},
  title = {Direct Perception Control Model},
  journal = {Symbiosis Robotics Blog},
  month = {August},
  year = {2026},
  url = {https://symbiosis-robotics.com/research/dpc},
}