Dev48
Language
  • About
  • Services
  • Industries
  • Technologies
  • Articles
  • Contacts
Book a call
    Home/Articles/Bringing humanoids to lerobot
Dev48

© 2026 · All rights reserved.

Bringing Humanoids to LeRobot

Источник: wavymulder

Bringing Humanoids to LeRobot

Source: wavymulder

A Blog post by Martino Russi on Hugging Face

September 29, 2026•Updated: September 29, 2026

A π0.5 policy predicts SONIC latent motion tokens from language prompt, robot state, and camera observations. A fast controller running on the G1 decodes those tokens into stable 29-DoF whole-body motion.

This post shows how we connect open-source hardware, teleoperation, datasets, and policy learning in LeRobot to build that workflow on the Unitree G1. You’ll learn how demonstrations become a π0.5 policy that predicts SONIC motion tokens, and how a fast whole-body controller turns those tokens into robot movement. Finally you'll learn about the larger humanoid ecosystem at LeRobot.

Why humanoids need a different control stack

LeRobot is a full-stack Python library for open-source robot learning. Early integrations focused on arms and mobile manipulators, where a policy can send action chunks directly to the actuators.

A biped, by contrast, must balance continuously while walking, reaching, and responding to disturbances. VLA inference alone is too slow to control those dynamics directly. We therefore put a fast controller between the learned policy and the robot: the policy predicts compact motion tokens, and the controller turns them into dynamically feasible joint targets.

Starting from OpenHLM

Our starting point is OpenHLM: Open Humanoid Loco-Manipulation.

In the OpenHLM paper, the authors train OpenPI's π0.5 on whole-body demonstrations of the Unitree G1. The policy takes language, camera observations, and robot state and predicts chunks of SONIC latent motion tokens. SONIC's decoder turns those tokens into joint targets, which the robot's low-level PD controller tracks.

In other words, π0.5 learns the task-level behavior while SONIC handles fast whole-body execution.

The OpenHLM paper also shows that useful loco-manipulation behaviors can be learned from relatively modest amounts of task-specific data: performance across its benchmark largely saturates at around 40 demonstrations per task.

We follow the same basic recipe, while deviating from it slightly:

  • Collect whole-body teleoperation demonstrations.
  • Record the 64-D latent action representation from SONIC's encoder instead of the robot's joint targets.
  • Train the VLA to predict this latent representation.
  • Let SONIC's decoder execute it.

The physical G1 setup

We equip a Unitree G1 with grippers, cameras, and additional CAN hardware, much of it adapted from the Unfolding Robotics project.

We first mount OpenArm grippers onto the G1 wrists using custom 3D-printed adapters.

OpenArm grippers with custom wrist adapters

An additional camera is mounted on the G1's head using a printed holder, giving the policy a stable egocentric view independent of the wrist cameras. A printed enclosure on the G1's back houses the CAN-FD adapter and additional gripper and camera cabling. Its removable faceplate keeps the interface accessible without dismantling the robot.

G1 backpack enclosure closed (center) and open (right).

The printable parts are available in lerobot/unitree-g1-hardware-modifications.

Turning human motion into robot actions

During teleoperation, the operator provides a kinematic reference through a VR headset: head and wrist poses, together with ankle poses or navigation commands depending on the teleoperation mode, which we toggle between during the episode depending on need.

This reference motion is converted into SMPL format and passed through SONIC's encoder.

VR teleoperation. The operator supplies head and wrist motion while SONIC fills in the rest of the body. The resulting latent motion stream becomes the action target used for imitation learning.

SONIC is a relatively small autoencoder (42 million parameters), trained on more than 100 million motion frames from BONES-SEED, retargeted to the Unitree G1. Its compact size lets it run directly on the robot, where low latency matters: delays between observations and actions make the executed trajectory diverge from the desired one.

In this architecture, different motion sources such as VR poses, reference joints, human motion, video-derived motion, or learned-policy output get encoded into a shared 64-D quantized motion token. The decoder combines that token with a short history of proprioceptive state and produces executable joint targets.

SONIC's universal control-policy architecture. The LeRobot integration focuses on the robot-side decode path: the caller supplies a motion token, and the controller turns it into executable G1 motion.

Separating the VLA from the controller means that the learned policy does not have to rediscover balance and whole-body coordination from a relatively small task dataset.

From demonstrations to a VLA

To showcase the pipeline we trained a can pick-and-place policy using nepyope/can_clean_final.

The dataset contains roughly 100 episodes at 50 fps, or about 71 minutes of whole-body teleoperation. Each episode uses the language instruction "Bring the can to the white table."

Three camera streams are recorded for every timestep, all at 480×640.

The state is 31-dimensional: 29 G1 joint positions + 2 gripper states. The action is 66-dimensional: 64 for the SONIC latent + 2 gripper commands.

We then fine-tune lerobot/pi05_base for 12,000 steps on 4×H100 GPUs. The resulting policy is published as nepyope/pi05-can-to-martino-12k.

The complete training and deployment commands live in the documentation rather than this post. See G1 training and inference for the reproducible workflow.

ViBE and Depth Dodge

What if you didn't need a VLA? In this experiment we attempt to replicate the results from ViBE and ORCS: we train a LoRA adapter on top of SONIC in mjlab using PPO, with the goal to avoid thrown balls using depth data from its head camera, together with its own body state.

Architecture overview: head-camera depth passes through frozen Theia-Tiny and a 16-frame temporal adapter to form a 128-D visual summary. Seven LoRA adapters feed it, together with the G1's body state, into SONIC to produce joint targets.

The vision pathway uses a frozen Theia-Tiny, a small vision foundation model for robot learning distilled from multiple off-the-shelf vision foundation models. A compact learned temporal adapter summarizes 16 depth-camera frames (about 0.6 seconds) into a 128-dimensional vector. Seven low-rank adapters let SONIC combine that vector with the robot's current state to produce joint targets. During training, an auxiliary objective asks the percept to predict the ball's position and velocity. In simulation, this visual signal was key: PPO with reward alone did not learn to dodge reliably. During rollout, the G1 uses its head-camera depth directly, as shown above.

In a three-seed simulation evaluation, the released depth policy dodged 79.1% of resolved throws. The same policy with blank depth dodged 0%, while a privileged oracle given the simulator's true ball state reached 97.4%. This is a simulation result, not a measured hardware success rate. The project also includes a physical G1 deployment clip. It is a hardware demonstration rather than a controlled evaluation, and the clip does not identify the exact checkpoint or camera settings. Find implementation, deployment notes, and released weights here.

More than one whole-body controller

It's also possible to control the G1 with NVIDIA GR00T-WholeBodyControl and Amazon FAR Holosoma in LeRobot. To compare them, check out this space.

More broadly, we support:

  • Dataset visualization. G1 URDF playback in the LeRobot visualizer, so robot motion can also be inspected visually.
  • Full proprioceptive state. The interface exposes joint positions and velocities, IMU data, estimated joint torques, and motor temperatures.
  • Arm IK. G1 arm kinematics and inverse kinematics can map end-effector targets into joint targets.
  • Gravity compensation. The interface supports it for upper-body joints.
  • MuJoCo simulation. The same Unitree G1 interface can connect to the packaged MuJoCo environment, allowing controller and teleoperation code to be exercised in simulation before running on the physical robot.

The Unitree G1 documentation contains the setup and commands for these workflows.

Open-source LeRobot humanoid hardware

We have also open-sourced several pieces of hardware for humanoids over the years.

The Homunculus glove directly measures finger-joint motion for recording dex-manip datasets and teleoperation. The latest open hardware release includes a custom ESP32-C3 board that reads 16 SS49E Hall-effect sensors through a multiplexer, along with firmware, a KiCad PCB design, printable parts, a URDF model, and an interactive viewer. The materials cost is under $20, and the design includes an interactive 3D assembly guide.

The glove is also integrated into LeRobot as the HomunculusGlove teleoperator: it reads the serial sensor stream, guides per-finger range-of-motion calibration, and maps the glove’s motion to the HOPE-Jr hand. That gives builders a low-cost route from a wearable input device to a robot hand and a LeRobot teleoperation workflow.

We later extended the idea into the Homunculus exoskeleton for arm teleoperation. Rather than reconstructing the operator's motion from only a few tracked points, the exoskeleton measures joint motion directly and retargets it to the robot.

A public example is nepyope/unitree_box_move_blue_full, containing 550 episodes and roughly 475k frames of G1 teleoperation data.

The hardware is open source. Firmware, PCB files, STEP/STL geometry, and URDF descriptions are available across:

For the LeRobot calibration and data-collection workflow, see the Homunculus loco-manipulation documentation.

HopeJr demonstration at SeedStudio

These devices were originally built to drive HOPEJr, the open-source DIY humanoid arm originally designed by TheRobotStudio. HOPEJr is a full-size humanoid with dexterous, tendon-driven hands, built from 3D-printed parts and off-the-shelf servos. The LeRobot-integrated arm and hand alone cost around €500.

LeRobot Humanoid

On the locomotion side, Virgile Batto built LeRobot Humanoid, a fully 3D-printed 12-DoF biped (6 per leg) driven by RobStride CAN-FD actuators and a Raspberry Pi 5. Its BOM is around $2,500. The project also publishes public Onshape CAD, BOM, wiring, assembly docs, motor commissioning tools, a MuJoCo simulation, calibration, safety tooling, and more as open-source resources.

Humanoid datasets on the Hub

HIW-500 is one of the largest examples: 500+ hours and 23K episodes of Unitree G1 teleoperation across 12 real homes.

Other whole-body collections include Unitree's UnifoLM-WBT datasets and OpenHLM's HLM-12 benchmark.

eidon-ai/tracker-pov contains 1,274 hours of egocentric video paired with 7-point IMU arm tracking, recorded during ordinary household work.

EgoHumanoid, for example, combines G1 demonstrations with robot-free egocentric human data collected from a VR headset and head-mounted camera, aligning both into a common training interface.

BONES-SEED, the corpus behind SONIC, contains more than 142K annotated human motions (roughly 288 hours) including Unitree G1 retargets.

Other examples include OMG-Data, with roughly 798K G1 motion episodes, and NVIDIA's GRAIL, and many more.

Where humanoids in LeRobot go next

Our goal is to connect all the open-source tools, models and data for humanoids that are out there so that they can be used together, thus fostering the community and making it so that researchers and AIs can spend more time working on robot learning instead of rebuilding the surrounding stack.

Next, we will support ASIMOV, the open-source humanoid from Menlo Research.

Get started with LeRobot.

– The LeRobot team ❤️

More info: Unitree G1 docs · SONIC whole-body control · Homunculus loco-manipulation · G1 hardware modifications · can_clean_final dataset · pi05-can-to-martino-12k policy · HIW-500-LeRobot ·

← All articles

More in AI & Machine Learning

All →
OpenAI DevDay live updates: Altman faces safety questions as company unveils new featuresПресса
OpenAI

OpenAI DevDay live updates: Altman faces safety questions as company unveils new features

Mistral CEO says U.S. AI safety debate masks competitors’ 'negligence'Пресса
Mistral

Mistral CEO says U.S. AI safety debate masks competitors’ 'negligence'

Ex-Tesla team raises $12.5M to put supply chains on autopilot
Пресса
Tesla

Ex-Tesla team raises $12.5M to put supply chains on autopilot

OpenAI reportedly ditches model over safety concernsПресса
OpenAI

OpenAI reportedly ditches model over safety concerns

OpenAI abandons plan to release upcoming model as safety concerns escalateПресса
OpenAI

OpenAI abandons plan to release upcoming model as safety concerns escalate

Source: Inference provider Modal Labs closing in on $750M round at $15.75B valuationПресса
Modal

Source: Inference provider Modal Labs closing in on $750M round at $15.75B valuation

More from wavymulder

What Is Jev AI? A Practical Guide to System One and Executable Decisions
wavymulder

What Is Jev AI? A Practical Guide to System One and Executable Decisions

YODAS v3: A 1 Million Hour Dataset for the Next Generation of Open Voice AI Research
wavymulder

YODAS v3: A 1 Million Hour Dataset for the Next Generation of Open Voice AI Research

How to Use the Jev AI Model: A Step-by-Step Developer Guide
wavymulder

How to Use the Jev AI Model: A Step-by-Step Developer Guide

Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization
wavymulder

Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization