Learning Holistic Whole-Body Loco-Manipulation with a Bipedal Mobile Manipulator

1ZJU-UIUC Institute 2LimX Dynamics
*Equal contribution   †Corresponding author
BipedalWBC performing board wiping, vertical placement, stabilization, and pick-and-place tasks

Given only a 6-DoF end-effector target, one learned controller coordinates reaching, postural adaptation, and stepping across diverse real-world manipulation tasks.

Abstract

Bipedal loco-manipulation enables robots to interact with objects beyond the nominal workspace of their arms by coordinating locomotion and manipulation. Realizing this capability requires a low-level whole-body controller that translates task-level manipulation goals into coordinated arm and leg motions while maintaining balance. We present a unified whole-body controller trained with reinforcement learning that directly maps 6-DoF end-effector targets to coordinated actions for the bipedal base and robotic arm.

Given only an end-effector target, the learned controller autonomously coordinates reaching, postural adaptation, and stepping without explicit base-velocity or footstep commands. A reward-gating strategy regulates the trade-offs among end-effector tracking, locomotion, and balance during training, while a temporal context estimator combines windowed Transformer encoding, recurrent GRU memory, and auxiliary dynamics prediction. Real-robot experiments demonstrate that the same controller supports VR teleoperation, a learned diffusion policy, and scripted trajectories through a common end-effector interface.

One Target, Whole-Body Coordination

The policy jointly controls all 14 arm and leg joints. Reward-gated training balances tracking, locomotion, and safety, while the Transformer-GRU estimator recovers dynamics-relevant context from observation history.

BipedalWBC training, temporal context estimator, actor, and deployment pipeline
Overview of the proposed controller and its shared task interface for simulation and real-world deployment.

Real-World Deployment

VR teleoperation, scripted trajectories, and learned policy interfaces

A common execution layer

VR teleoperation, learned diffusion policies, and scripted trajectories all provide the same end-effector command. The low-level policy decides when arm motion is sufficient and when the bipedal base must adapt or step.

Expanded reachability

Whole-body coordination expands the measured vertical end-effector range from approximately 38–163 cm with floating-base inverse kinematics to 3–191 cm with BipedalWBC.

188 cmvertical range
88.30%non-oracle simulation success
Vertical end-effector reachability comparison

Video Presentation

BibTeX

@misc{chen2027bipedalwbc,
  title  = {Learning Holistic Whole-Body Loco-Manipulation with a Bipedal Mobile Manipulator},
  author = {Zhongyu Chen and Yuxuan Nai and Qian Chen and Yidong Zhu and Chen Jing and Qihan Wang and Xudong Li and Zhizhan Li and Leixin Chang and Liangjing Yang and Hua Chen},
  year   = {2027},
  note   = {Under review},
  url    = {https://wholebodyrobotics.github.io/BipedalWBC/}
}