FAME: Force-Adaptive RL for Expanding the Manipulation Envelope of a Full-Scale Humanoid
Under review · 2026
Bimanual manipulation loads the hands, and those forces travel down the kinematic chain into the stance. The same push produces a different moment at the base depending on how the arms are posed, so neither the force nor the arm configuration alone tells the legs what they are resisting. Can a standing policy that represents the two together keep a humanoid upright, and its hands where the task put them?
We present FAME, a force-adaptive reinforcement learning framework for the Unitree H1-2. An upper-body context encoder compresses torso and arm joint positions together with both hand forces into a latent that conditions the lower-body standing policy. At deployment the hand forces are reconstructed online from joint torques through rigid-body inverse dynamics, so no wrist force/torque sensor is needed.
How FAME works
FAME has three parts: an upper-body context encoder, a base standing policy conditioned on its latent, and a model-based estimator that supplies the hand forces at deployment. Training exposes the policy to isotropically sampled 3D forces at each hand while a pose curriculum widens the range of arm configurations as standing improves.
Encode force and arm configuration jointly
The encoder μθ, a small MLP, maps the 15 torso and arm joint positions and the two 3D hand forces to an 8-D latent, zt = μθ([qub, FL, FR]). It is trained end-to-end with the policy through PPO, with no force labels or auxiliary regression loss. The policy sees a three-step proprioceptive history together with the three most recent latents and outputs position targets for the 12 leg joints, tracked by a PD controller at 50 Hz.
Estimate hand force from joint torques
The ground reaction is unknown and discontinuous at contact, which blocks estimating the hand force from the full-body dynamics. FAME restricts the equation of motion to the seven rows of one arm. Moving an arm joint does not displace either foot, so the foot-contact Jacobians vanish on those rows and the ground reaction drops out without being measured or modelled:
Seven equations for three unknowns are solved by least squares, with the Jacobian and inverse dynamics evaluated online in Pinocchio. The estimator follows from the robot model rather than from data collected under a particular controller, so it is independent of the policy it serves. Every result below runs with this estimate in the loop.
Standing under swept hand forces
We sample 100 upper-body configurations from the training pose distribution and apply 25 force draws to each: 2,500 loaded trials per policy. Each trial is run twice, with and without the force, and the two hand trajectories are differenced. A trial succeeds only if the robot stays upright, does not step more than 0.15 m, and keeps peak hand deviation below a tolerance T.
Task success
Share of 2,500 loaded trials, estimated force in the loop
| Variant | Success, T = 150 mm | Success, T = 250 mm | Hand deviation ΔEE (mm) |
|---|---|---|---|
| FAME | 38.9 | 48.5 | 179.6 |
| ALMI | 24.7 | 39.3 | 202.3 |
| +Curr+Fraw | 16.6 | 21.7 | 220.3 |
| FAME-pose | 4.3 | 5.6 | 242.4 |
Adding force to the context lifts success from 4.3% to 38.9%. Giving the policy the force without the joint encoding reaches only 16.6%, so most of the gain comes from representing force and arm configuration together, not from having the force signal. ALMI scores 24.7%: it stays upright but recovers by stepping, which relocates the hands.
Left FAME · Right FAME-pose. Same curriculum and latent width, but FAME-pose never sees the force, and falls under the load.
Left FAME · Right +Curr+Fraw. Both receive the estimated force; only FAME encodes it with the arm configuration.
Left ALMI · Right FAME, asymmetric arm pose under sampled wrist forces.
Kitchen tasks and hardware
Forces produced by a real task are neither isotropic nor independent of the robot's own motion. We test whether the result holds when the load comes from the task itself, and when the policy runs on the physical robot.
Single-arm pick in a MuJoCo kitchen
On the GOLEM platform, FAME and ALMI sit behind a shared velocity-command interface, with the upper body driven by the same inverse-kinematics stack so that the lower body is the only variable. The task is a single-arm pick of a 3 kg object from the counter, over 10 seeds.
- Falls
- 1/10ALMI 5/10
- Median peak pelvis displacement
- 0.065 mALMI 0.641 m
- Median peak tilt
- 7.1°ALMI 30.5°
Loads on the Unitree H1-2
The policy runs at 50 Hz using only the IMU and joint encoders; no wrist force/torque sensor is used. We compare FAME against FAME-pose under two loads: RE1, an asymmetric single-arm load that creates an unbalanced moment, and RE2, a symmetric load shared across both arms. FAME held its stance in 3/3 trials of each; FAME-pose fell in 3/3 of each.
Played at 1.75×. FAME-pose appears as “Base+Curr” in the full video.
Scope and limitations
FAME is a standing controller: success requires the feet to stay put, and evaluation poses are drawn from the training distribution, so the results measure behaviour inside the trained envelope rather than extrapolation beyond it. On hardware the force estimate is noisier than in simulation. Extending the joint encoding to load-carrying locomotion, where the feet may move but the hands must still arrive, is the natural next step.
Cite this work
BibTeX
@article{pudasaini2026fame,
title = {{FAME}: Force-Adaptive {RL} for Expanding the Manipulation
Envelope of a Full-Scale Humanoid},
author = {Pudasaini, Niraj and Zhang, Yutong and Lavering, Jensen
and Conway, Max and Roncone, Alessandro and Correll, Nikolaus},
journal = {arXiv preprint arXiv:2603.08961},
year = {2026}
}