FAME
Paper Video Code (Releasing Soon)

FAME: Force-Adaptive RL for Expanding the Manipulation Envelope of a Full-Scale Humanoid

  • Niraj Pudasaini1
  • Yutong Zhang2
  • Jensen Lavering2
  • Max Conway2
  • Alessandro Roncone2
  • Nikolaus Correll1

1 University of Notre Dame 2 University of Colorado Boulder

Under review · 2026

Bimanual manipulation loads the hands, and those forces travel down the kinematic chain into the stance. The same push produces a different moment at the base depending on how the arms are posed, so neither the force nor the arm configuration alone tells the legs what they are resisting. Can a standing policy that represents the two together keep a humanoid upright, and its hands where the task put them?

We present FAME, a force-adaptive reinforcement learning framework for the Unitree H1-2. An upper-body context encoder compresses torso and arm joint positions together with both hand forces into a latent that conditions the lower-body standing policy. At deployment the hand forces are reconstructed online from joint torques through rigid-body inverse dynamics, so no wrist force/torque sensor is needed.

Left: the H1-2 stands under a 30 N load while an upper-body context encoder maps arm joints and hand forces to a latent that conditions a lower-body RL policy. Right: the policy in simulation during bowl pick-and-place and refrigerator door opening.
Overview. Left: arm and torso configuration (ℝ15) and hand forces [FL, FR] (ℝ6) are encoded into a latent context zt for the standing policy; the H1-2 stands under a 30 N load. Right: the same policy in the GOLEM kitchen simulation during bowl pick-and-place and refrigerator door opening, where the forces come from the task.
01

How FAME works

FAME has three parts: an upper-body context encoder, a base standing policy conditioned on its latent, and a model-based estimator that supplies the hand forces at deployment. Training exposes the policy to isotropically sampled 3D forces at each hand while a pose curriculum widens the range of arm configurations as standing improves.

Training (top): a pose curriculum and a 3D hand-force sampler perturb H1-2 robots in IsaacGym; the upper-body context encoder and standing policy are trained with PPO. Deployment (bottom): measured robot state feeds a Pinocchio-based force estimator, and the frozen encoder and policy output 12-D lower-body joint targets at 50 Hz.
Training and deployment. In simulation the encoder receives the sampled hand forces exactly. On the robot, the same frozen encoder receives forces estimated from measured joint states and torques.

Encode force and arm configuration jointly

The encoder μθ, a small MLP, maps the 15 torso and arm joint positions and the two 3D hand forces to an 8-D latent, zt = μθ([qub, FL, FR]). It is trained end-to-end with the policy through PPO, with no force labels or auxiliary regression loss. The policy sees a three-step proprioceptive history together with the three most recent latents and outputs position targets for the 12 leg joints, tracked by a PD controller at 50 Hz.

Estimate hand force from joint torques

The ground reaction is unknown and discontinuous at contact, which blocks estimating the hand force from the full-body dynamics. FAME restricts the equation of motion to the seven rows of one arm. Moving an arm joint does not displace either foot, so the foot-contact Jacobians vanish on those rows and the ground reaction drops out without being measured or modelled:

[Jh⊤F]A = RNEA(q, q̇, q̈)A + BAq̈A + DAq̇A − τA

Seven equations for three unknowns are solved by least squares, with the Jacobian and inverse dynamics evaluated online in Pinocchio. The estimator follows from the robot model rather than from data collected under a particular controller, so it is independent of the policy it serves. Every result below runs with this estimate in the loop.

02

Standing under swept hand forces

We sample 100 upper-body configurations from the training pose distribution and apply 25 force draws to each: 2,500 loaded trials per policy. Each trial is run twice, with and without the force, and the two hand trajectories are differenced. A trial succeeds only if the robot stays upright, does not step more than 0.15 m, and keeps peak hand deviation below a tolerance T.

Task success

Share of 2,500 loaded trials, estimated force in the loop

Task success (%) and force-induced hand deviation by variant
VariantSuccess, T = 150 mmSuccess, T = 250 mmHand deviation ΔEE (mm)
FAME38.948.5179.6
ALMI24.739.3202.3
+Curr+Fraw16.621.7220.3
FAME-pose4.35.6242.4
FAME-pose encodes arm configuration only. +Curr+Fraw receives the same estimated force, concatenated rather than encoded. ALMI is an adversarially trained locomotion policy with no force input.

Adding force to the context lifts success from 4.3% to 38.9%. Giving the policy the force without the joint encoding reaches only 16.6%, so most of the gain comes from representing force and arm configuration together, not from having the force signal. ALMI scores 24.7%: it stays upright but recovers by stepping, which relocates the hands.

Left FAME · Right FAME-pose. Same curriculum and latent width, but FAME-pose never sees the force, and falls under the load.

03

Kitchen tasks and hardware

Forces produced by a real task are neither isotropic nor independent of the robot's own motion. We test whether the result holds when the load comes from the task itself, and when the policy runs on the physical robot.

Single-arm pick in a MuJoCo kitchen

On the GOLEM platform, FAME and ALMI sit behind a shared velocity-command interface, with the upper body driven by the same inverse-kinematics stack so that the lower body is the only variable. The task is a single-arm pick of a 3 kg object from the counter, over 10 seeds.

Left FAME · Right ALMI, same seed.
Falls
1/10ALMI 5/10
Median peak pelvis displacement
0.065 mALMI 0.641 m
Median peak tilt
7.1°ALMI 30.5°

Loads on the Unitree H1-2

The policy runs at 50 Hz using only the IMU and joint encoders; no wrist force/torque sensor is used. We compare FAME against FAME-pose under two loads: RE1, an asymmetric single-arm load that creates an unbalanced moment, and RE2, a symmetric load shared across both arms. FAME held its stance in 3/3 trials of each; FAME-pose fell in 3/3 of each.

Held FAME
Fell FAME-pose

Played at 1.75×. FAME-pose appears as “Base+Curr” in the full video.

Hip pitch, ankle pitch and elbow trajectories and torques for RE1 and RE2. With FAME the joints stay near their standing configuration while loaded; with FAME-pose they drift until the robot falls. Snapshot sequences below show each case.
Joint trajectories on hardware. With FAME, hip- and ankle-pitch stay near the nominal standing posture while loaded (green). With FAME-pose they drift until balance is lost (red).

Scope and limitations

FAME is a standing controller: success requires the feet to stay put, and evaluation poses are drawn from the training distribution, so the results measure behaviour inside the trained envelope rather than extrapolation beyond it. On hardware the force estimate is noisier than in simulation. Extending the joint encoding to load-carrying locomotion, where the feet may move but the hands must still arrive, is the natural next step.

Cite this work

BibTeX

@article{pudasaini2026fame,
  title   = {{FAME}: Force-Adaptive {RL} for Expanding the Manipulation
             Envelope of a Full-Scale Humanoid},
  author  = {Pudasaini, Niraj and Zhang, Yutong and Lavering, Jensen
             and Conway, Max and Roncone, Alessandro and Correll, Nikolaus},
  journal = {arXiv preprint arXiv:2603.08961},
  year    = {2026}
}