ICML 2026 Embodied AIReinforcement LearningVLMs

LAGEA: Language Guided Embodied Agents for Robotic Manipulation

Abdul Monaf Chowdhury1, Akm Moshiur Rahman Mazumder2, Safaeid Hossain Arib1, Rabeya Akter1

1University of Dhaka   2Independent University, Bangladesh

Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)

TL;DR

A similarity score summarizes the outcome, not the cause. LAGEA asks a VLM to diagnose why an episode failed using a fixed error schema, anchors that diagnosis to the key frames where it happened, and turns it into a shaping reward that fades as the robot gets competent.

  • 80.0%MT10 fixed goals (+4.0 over FuRL)
  • 70.4%MT10 random goals (+5.8 over FuRL)
  • 51.7%Gymnasium-Robotics Fetch (+7.5 over FuRL)
  • 98.9%with structured feedback vs. 78.9% free-form
Overview figure for LAGEA: Language Guided Embodied Agents for Robotic Manipulation
The LAGEA loop. Key frames from each trajectory are sent to a frozen VLM with an error taxonomy. The structured reflection is embedded, aligned with visual states, and fused with goal progress into a shaping reward for the policy.

Abstract

Robotic manipulation benefits from foundation models that describe goals, but today’s agents still lack a principled way to learn from their own mistakes. We ask whether natural language can serve as feedback, an error-reasoning signal that helps embodied agents diagnose what went wrong and correct course. We introduce LaGEA (Language Guided Embodied Agents), a framework that turns episodic, schema-constrained reflections from a vision language model (VLM) into temporally grounded guidance for reinforcement learning. LaGEA summarizes each attempt in concise language, localizes the decisive moments in the trajectory, aligns feedback with visual state in a shared representation, and converts goal progress and feedback agreement into bounded, step-wise shaping rewards whose influence is modulated by an adaptive, failure-aware coefficient. This design yields dense signals early when exploration needs direction and gracefully recedes as competence grows. On the Meta-World MT10 and Robotic Fetch embodied manipulation benchmark, LaGEA improves average success over the state-of-the-art (SOTA) methods by 9.0% on random goals, 5.3% on fixed goals, and 17% on fetch tasks, while converging faster. These results support our hypothesis: language, when structured and grounded in time, is an effective mechanism for teaching robots to self-reflect on mistakes and make better choices.

A score is not a diagnosis

Motivation

Recent methods use vision–language models as reward models. They compare the current camera frame with a goal description and return a similarity score. That score says how far the robot is from success, but not why it is failing. It cannot tell grasping the wrong object from approaching at the wrong angle or pushing too gently.

LAGEA makes the VLM produce that diagnosis explicitly, then solves three problems that stop language from being a usable training signal. Small VLMs drift and hallucinate, so feedback must be structured. One sentence per episode gives poor credit assignment, so it must be anchored in time. And text is not a reward, so it must be aligned with what the robot sees.

What a reflection looks like

After every episode the VLM must answer in a fixed JSON schema with one primary error code from a small taxonomy, a one-sentence explanation, the key frames, a suggested fix and a confidence.

Error codeMeaning
wrong_objectInteracted with the wrong object
bad_approach_directionApproached the object from the wrong angle or direction
failed_graspContact without a stable grasp, slipped or never closed
insufficient_forceTouched the right object but without enough motion or force
drift_from_goalDrifted away from the goal with no course correction

Failure on button-press-topdown

{
  "task": "button-press-topdown-v2-goal-observable",
  "outcome": "failure",
  "primary_error": {
    "code": "bad_approach_direction",
    "explanation": "The gripper came from the side, sliding off the button instead of a vertical press."
  },
  "secondary_factors": [],
  "key_frame_indices": [18, 22],
  "suggested_fix": "Approach from directly above the button; align gripper normal to the button surface, then press straight down.",
  "confidence": 0.85,
  "summary": "The robot failed to press the button correctly because it approached from the side instead of a vertical press."
}

The reflection is produced by a frozen Qwen2.5-VL-3B and embedded with a frozen GPT-2 into a 768-dimensional feedback vector.

Feedback generation pipeline from key frames to VLM reflection to feedback embedding
Feedback generation. Key frames selected from an episode are analysed by the VLM, which returns structured feedback that is then encoded into a feedback embedding.

Method

Find the moments, align, reward progress, fade the guidance

1. Find the moments that mattered

Broadcasting one feedback vector to every step ignores when the outcome was decided. LAGEA scores each frame by how close it is to the goal and how sharply that closeness is changing, using image embeddings \(x_t\) and the goal embedding \(g\).

\[ s_t = \cos(x_t, g), \qquad v_t = s_t - s_{t-1}, \qquad a_t = v_t - v_{t-1} \]
\[ p_t = \omega_s\,[z(s_t)]_+ + \omega_v\, z(|v_t|) + \omega_a\, z(|a_t|) \]

The highest-saliency frames, spaced apart in time, become key frames \(\mathcal{K}\). A triangular kernel spreads their influence to nearby steps, giving per-step weights that concentrate learning on approach, contact and reversal.

\[ \tilde w_t = \max_{k \in \mathcal{K}} \Big(1 - \tfrac{|t-k|}{h+1}\Big)_+, \qquad w_t = \beta + (1-\beta)\, \tilde w_t \]

2. Put feedback and pixels in one space

Small MLP projectors map image states \(z_t\), the feedback \(z_f\) and the goal \(z_g\) onto a shared unit sphere. A weighted binary loss calibrates absolute agreement (successful steps pull image and feedback together, failures push them apart), and an InfoNCE loss shapes relative geometry across the batch.

\[ \mathcal{L}_{\mathrm{align}} = \lambda_{\mathrm{bce}}\, \mathcal{L}_{\mathrm{bce}} + \lambda_{\mathrm{nce}}\, \mathcal{L}_{\mathrm{nce}}, \qquad \psi_t = \langle z_t, z_f \rangle \]

3. Reward progress, not position

Goal potential and feedback potential used to compute delta rewards
Delta-based rewards. (a) A goal potential aligns the current state with the goal image and instruction. (b) A feedback potential aligns the state with the VLM feedback. Their temporal differences form the shaping reward.

LAGEA rewards the change in goal agreement and in feedback agreement, so the signal is positive when the robot moves the right way. Feedback reward is gated by the key-frame weights, and the two are mixed according to how well the instruction and the feedback agree.

\[ r_t^{\mathrm{goal}} = \tanh\!\Big(\tfrac{\gamma\,\phi_{t+1} - \phi_t}{\tau_{\mathrm{goal}}}\Big), \qquad r_t^{\mathrm{fb}} = \hat w_t \tanh\!\Big(\tfrac{\gamma\,\psi_{t+1} - \psi_t}{\tau_f}\Big) \]
\[ \tilde r_t = (1-\alpha)\, r_t^{\mathrm{goal}} + \alpha\, r_t^{\mathrm{fb}}, \qquad \alpha = \mathrm{clip}\big(\alpha_{\mathrm{base}} \cdot \tfrac12 (1 + \langle z_y, z_f \rangle)\big) \]

4. Let the guidance fade

Dense language rewards can overpower the sparse task reward. LAGEA applies shaping only on failures and scales it by a coefficient that shrinks as estimated progress \(P\) grows, so guidance is strong early and recedes as the policy becomes competent.

\[ \rho_t = \rho_{\min} + (\rho_{\max} - \rho_{\min})(1 - P), \qquad r_t = r_t^{\mathrm{task}} + \mathbf{1}[r_t^{\mathrm{task}} < 0]\, \rho_t\, \tilde r_t \]

Results

Success rate (%) · higher is better

Meta-World MT10, fixed goals

TaskSACLIVLIV-ProjRelayFuRL w/o goal imgFuRLLAGEA
button-press-topdown0006080100100
door-open500080100100100
drawer-close100100100100100100100
drawer-open2000408080100
peg-insert-side0000000
pick-place0000000
push00004080100
reach608080100100100100
window-close60604080100100100
window-open80402080100100100
Average37.028.024.054.070.076.080.0

Average over five seeds.

Meta-World MT10, random goals

TaskSACRelayFuRLLAGEA
button-press-topdown16.0 (32.0)56.0 (38.3)64.0 (32.6)96 (8)
door-open78.0 (39.2)80.0 (30.3)96.0 (8.0)100 (0)
drawer-close100.0 (0.0)100.0 (0.0)100.0 (0.0)100 (0)
drawer-open40.0 (49.0)50.0 (42.0)84.0 (27.3)92 (9.8)
pick-place0.0 (0.0)0.0 (0.0)0.0 (0.0)4 (4.9)
peg-insert-side0.0 (0.0)0.0 (0.0)0.0 (0.0)0
push0.0 (0.0)0.0 (0.0)6.0 (8.0)12 (4)
reach100.0 (0.0)100.0 (0.0)100.0 (0.0)100 (0)
window-close86.0 (28.0)96.0 (4.9)100.0 (0.0)100 (0)
window-open78.0 (39.2)92.0 (7.5)96.0 (4.9)100 (0)
Average49.8 (7.9)57.4 (7.0)64.6 (5.0)70.4 (1.85)

Mean (standard deviation) over five seeds.

Gymnasium-Robotics Fetch

The four Fetch tasks: Reach, Push, PickAndPlace and Slide
Fetch benchmark. A 7-DoF arm on Reach, Push, PickAndPlace and Slide.
TaskSACRelayFuRLLAGEA
Reach100 (0)100 (0)100 (0)100 (0)
Push26.67 (4.71)30 (8.16)40 (8.16)53.33 (4.71)
PickAndPlace10 (8.16)20 (0)33.33 (9.43)43.33 (4.71)
Slide0 (0)0 (0)3.33 (4.71)10 (8.16)
Average34.1737.544.1751.67

Mean (standard deviation) over three seeds.

Faster convergence

Success-rate learning curves for SAC, FuRL and LAGEA on eight MT10 tasks
Learning curves on eight Meta-World tasks. LAGEA reaches high success in far fewer environment steps than FuRL and SAC, which plateau late or stall.

Averaged over nine MT10 settings, LAGEA converges in 92.4 minutes of wall-clock time versus 94.9 for FuRL, despite querying a VLM.

What matters

Ablations

Reward ablation: 79% without goal delta, 80% without feedback delta, 80% without adaptive rho, 99% full LaGEA
Every reward term matters. Removing the goal delta, the feedback delta or the adaptive schedule drops success to 79–80%, versus 99% for the full model.
Feedback reward on drawer-open with and without key frames
Key frames sharpen credit. On drawer-open, key-frame gating focuses the feedback reward on the decisive moments.
</tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tbody> </table> </div>
  • Structure is the biggest single lever. Constraining the VLM to a schema lifts success by 20 points over free-form reflections.
  • Timing matters as much as content. Saliency-based key frames beat random or uniform frames by about 12 points.
  • The policy transfers across viewpoints. Trained from one camera, LAGEA loses at most 2.7 points from three unseen ones.

Limitations

LAGEA is evaluated in simulation, and contact-rich tasks such as peg insertion and pick-and-place remain largely unsolved. The approach is robust to viewpoint shifts but does not claim to handle severe self-occlusion or cluttered scenes, and VLM feedback can still be wrong. Real-robot and sim-to-real experiments are natural next steps.

VariantSuccess (%)
Feedback format (6 tasks)
Free-form text78.89
Schema-constrained JSON98.89
Key-frame selection
Random frames68.0
Uniform frames67.3
LAGEA key frames80.0
VLM for reflection
SmolVLM256.0
OpenQwen2VL66.7
InternVL268.0
Qwen2.5-VL-3B80.0
Text encoder for feedback
MPNet70.7
BGE71.3
LIV72.7
GPT-280.0
Camera viewpoint (zero-shot)
Training view80.0
Directly overhead79.3
Front-left diagonal77.3
Behind-left diagonal79.3

Citation

@inproceedings{chowdhury2026lagea,
  title     = {{LAGEA}: Language Guided Embodied Agents for Robotic Manipulation},
  author    = {Chowdhury, Abdul Monaf and Mazumder, Akm Moshiur Rahman and
               Arib, Safaeid Hossain and Akter, Rabeya},
  booktitle = {Proceedings of the 43rd International Conference on Machine Learning},
  series    = {Proceedings of Machine Learning Research},
  volume    = {306},
  year      = {2026}
}