Offline adaptation · No target-light interaction
World-model relighting, source-retentive replay, and frozen policy references adapt a trained HIL-RL policy to changed illumination without additional target-light interaction.
Real-robot rollout shown at 3× speed
Abstract
Despite strong performance in contact-rich manipulation, human-in-the-loop reinforcement learning (HIL-RL) remains vulnerable to visual distribution shifts. A policy that achieves near-perfect success at its training workstation can fail at a workstation only a few meters away when changes in lamp placement and window light alter its visual observations. Repeated demonstrations and HIL limit deployment scalability, while naive shifted-light fine-tuning can forget the source policy. We present RoHIL for offline fine-tuning without new real-robot interaction.
RoHIL combines (i) a world-model-based image relighter that resynthesizes source trajectories under virtual HDRI environments while retaining real actions and rewards; (ii) Illumination-Retention Replay (IRR), a data-level anti-forgetting mechanism that interleaves relit adaptation transitions with original-light retention transitions to preserve source-workstation Bellman coverage; and (iii) an anchored Bellman–actor regularizer that constrains critic-value and deterministic-action drift from the source policy. Across four real-robot manipulation tasks under significant cross-workstation illumination variations, RoHIL substantially improves shifted-light performance where standard HIL-RL collapses while preserving source-workstation performance, eliminating the need to re-collect data and retrain for every new workstation and environment.
Motivation
A HIL-SERL policy trained at one workstation reaches near-perfect success there, but moving the same robot a few metres away — under different lamp positions and daylight — breaks it. Below are rollouts of the source-workstation policy on four tasks after an illumination shift, while task geometry, camera poses, dynamics, action semantics, and rewards remain fixed. Generic image perturbations do not reproduce temporally consistent shadows, highlights, and reflections. RoHIL adapts the recorded experience without target-light interaction.
All rollouts on this page are played at 3× real-time speed.
Framework Overview
Starting from an online-best HIL-SERL policy that degrades under lighting shifts, RoHIL runs an offline robust fine-tune in three coupled stages. The visual encoder is exposed to illumination-diverse evidence via world-model relighting, while replay-buffer balancing and frozen-source output references jointly suppress source-domain forgetting in Bellman support, critic values, and deterministic actions.
Stage 01
Cosmos-Transfer1-DiffusionRenderer processes contiguous recorded camera streams under four HDRI conditions: bright, dim, cool, and natural light. Proprioception, actions, rewards, termination flags, view identity, and transition indices are retained; only the synchronized RGB fields change.
Stage 02
IRR retains the RLPD 50/50 replay/demonstration split. Within the replay half, α controls the probability of sampling source-policy replay; the demonstration half draws from original and relit expert anchors. The USB development sweep selects α = 0.75, which is fixed for the other three tasks.
Stage 03
On source and relit demonstration states, LQref preserves the frozen source critics' values for recorded expert actions, while LAref preserves normalized deterministic actions. Their shared weight decays linearly from 1 to 0.33 during adaptation.
Experimental Results
We evaluate RoHIL against five baselines (BC, ACT, HG-DAgger, HIL-SERL, IBRL) on four real-robot manipulation tasks under a controlled lighting-shift gradient. At physical light intensity 60, RoHIL records the highest success rate on every task. Across intensities 10–100, its fixed-checkpoint success curve remains above the other evaluated checkpoints. USB is the development task; the selected recipe is held fixed for the other tasks.
Experimental Demos
Below are real-robot rollouts of the RoHIL-fine-tuned policy on each task, executed under physical lighting shifts. The adaptation horizon and replay ratio are selected on USB insertion at intensity 60, then held fixed for RAM insertion, table wiping, and circuit-breaker actuation. Target-light rollouts are excluded from replay and learner updates.
All rollouts on this page are played at 3× real-time speed.
Experimental Setup
All experiments use a Franka Emika Panda controlled at 10 Hz through six-dimensional end-effector delta-twist commands. The first three tasks use two wrist-mounted Intel RealSense RGB streams; table wiping adds a side view. Physical evaluation varies task-light intensity while task geometry, camera poses, dynamics, action semantics, and rewards remain fixed.
Hardware setup. For each of the four tasks, we annotate the wrist cameras (red), workspace lights (yellow), and side cameras when present. Camera baselines and light placement differ per task. The same adaptation recipe is used across all four tasks after selection on USB insertion.