LAGEA: Language Guided Embodied Agents for Robotic Manipulation
1University of Dhaka 2Independent University, Bangladesh
Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)
TL;DR
A similarity score summarizes the outcome, not the cause. LAGEA asks a VLM to diagnose why an episode failed using a fixed error schema, anchors that diagnosis to the key frames where it happened, and turns it into a shaping reward that fades as the robot gets competent.
- 80.0%MT10 fixed goals (+4.0 over FuRL)
- 70.4%MT10 random goals (+5.8 over FuRL)
- 51.7%Gymnasium-Robotics Fetch (+7.5 over FuRL)
- 98.9%with structured feedback vs. 78.9% free-form

Abstract
Robotic manipulation benefits from foundation models that describe goals, but today’s agents still lack a principled way to learn from their own mistakes. We ask whether natural language can serve as feedback, an error-reasoning signal that helps embodied agents diagnose what went wrong and correct course. We introduce LaGEA (Language Guided Embodied Agents), a framework that turns episodic, schema-constrained reflections from a vision language model (VLM) into temporally grounded guidance for reinforcement learning. LaGEA summarizes each attempt in concise language, localizes the decisive moments in the trajectory, aligns feedback with visual state in a shared representation, and converts goal progress and feedback agreement into bounded, step-wise shaping rewards whose influence is modulated by an adaptive, failure-aware coefficient. This design yields dense signals early when exploration needs direction and gracefully recedes as competence grows. On the Meta-World MT10 and Robotic Fetch embodied manipulation benchmark, LaGEA improves average success over the state-of-the-art (SOTA) methods by 9.0% on random goals, 5.3% on fixed goals, and 17% on fetch tasks, while converging faster. These results support our hypothesis: language, when structured and grounded in time, is an effective mechanism for teaching robots to self-reflect on mistakes and make better choices.
A score is not a diagnosis
Motivation
Recent methods use vision–language models as reward models. They compare the current camera frame with a goal description and return a similarity score. That score says how far the robot is from success, but not why it is failing. It cannot tell grasping the wrong object from approaching at the wrong angle or pushing too gently.
LAGEA makes the VLM produce that diagnosis explicitly, then solves three problems that stop language from being a usable training signal. Small VLMs drift and hallucinate, so feedback must be structured. One sentence per episode gives poor credit assignment, so it must be anchored in time. And text is not a reward, so it must be aligned with what the robot sees.
What a reflection looks like
After every episode the VLM must answer in a fixed JSON schema with one primary error code from a small taxonomy, a one-sentence explanation, the key frames, a suggested fix and a confidence.
| Error code | Meaning |
|---|---|
wrong_object | Interacted with the wrong object |
bad_approach_direction | Approached the object from the wrong angle or direction |
failed_grasp | Contact without a stable grasp, slipped or never closed |
insufficient_force | Touched the right object but without enough motion or force |
drift_from_goal | Drifted away from the goal with no course correction |
Failure on button-press-topdown
{
"task": "button-press-topdown-v2-goal-observable",
"outcome": "failure",
"primary_error": {
"code": "bad_approach_direction",
"explanation": "The gripper came from the side, sliding off the button instead of a vertical press."
},
"secondary_factors": [],
"key_frame_indices": [18, 22],
"suggested_fix": "Approach from directly above the button; align gripper normal to the button surface, then press straight down.",
"confidence": 0.85,
"summary": "The robot failed to press the button correctly because it approached from the side instead of a vertical press."
}
The reflection is produced by a frozen Qwen2.5-VL-3B and embedded with a frozen GPT-2 into a 768-dimensional feedback vector.

Method
Find the moments, align, reward progress, fade the guidance
1. Find the moments that mattered
Broadcasting one feedback vector to every step ignores when the outcome was decided. LAGEA scores each frame by how close it is to the goal and how sharply that closeness is changing, using image embeddings \(x_t\) and the goal embedding \(g\).
The highest-saliency frames, spaced apart in time, become key frames \(\mathcal{K}\). A triangular kernel spreads their influence to nearby steps, giving per-step weights that concentrate learning on approach, contact and reversal.
2. Put feedback and pixels in one space
Small MLP projectors map image states \(z_t\), the feedback \(z_f\) and the goal \(z_g\) onto a shared unit sphere. A weighted binary loss calibrates absolute agreement (successful steps pull image and feedback together, failures push them apart), and an InfoNCE loss shapes relative geometry across the batch.
3. Reward progress, not position

LAGEA rewards the change in goal agreement and in feedback agreement, so the signal is positive when the robot moves the right way. Feedback reward is gated by the key-frame weights, and the two are mixed according to how well the instruction and the feedback agree.
4. Let the guidance fade
Dense language rewards can overpower the sparse task reward. LAGEA applies shaping only on failures and scales it by a coefficient that shrinks as estimated progress \(P\) grows, so guidance is strong early and recedes as the policy becomes competent.
Results
Success rate (%) · higher is better
Meta-World MT10, fixed goals
| Task | SAC | LIV | LIV-Proj | Relay | FuRL w/o goal img | FuRL | LAGEA |
|---|---|---|---|---|---|---|---|
| button-press-topdown | 0 | 0 | 0 | 60 | 80 | 100 | 100 |
| door-open | 50 | 0 | 0 | 80 | 100 | 100 | 100 |
| drawer-close | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| drawer-open | 20 | 0 | 0 | 40 | 80 | 80 | 100 |
| peg-insert-side | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| pick-place | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| push | 0 | 0 | 0 | 0 | 40 | 80 | 100 |
| reach | 60 | 80 | 80 | 100 | 100 | 100 | 100 |
| window-close | 60 | 60 | 40 | 80 | 100 | 100 | 100 |
| window-open | 80 | 40 | 20 | 80 | 100 | 100 | 100 |
| Average | 37.0 | 28.0 | 24.0 | 54.0 | 70.0 | 76.0 | 80.0 |
Average over five seeds.
Meta-World MT10, random goals
| Task | SAC | Relay | FuRL | LAGEA |
|---|---|---|---|---|
| button-press-topdown | 16.0 (32.0) | 56.0 (38.3) | 64.0 (32.6) | 96 (8) |
| door-open | 78.0 (39.2) | 80.0 (30.3) | 96.0 (8.0) | 100 (0) |
| drawer-close | 100.0 (0.0) | 100.0 (0.0) | 100.0 (0.0) | 100 (0) |
| drawer-open | 40.0 (49.0) | 50.0 (42.0) | 84.0 (27.3) | 92 (9.8) |
| pick-place | 0.0 (0.0) | 0.0 (0.0) | 0.0 (0.0) | 4 (4.9) |
| peg-insert-side | 0.0 (0.0) | 0.0 (0.0) | 0.0 (0.0) | 0 |
| push | 0.0 (0.0) | 0.0 (0.0) | 6.0 (8.0) | 12 (4) |
| reach | 100.0 (0.0) | 100.0 (0.0) | 100.0 (0.0) | 100 (0) |
| window-close | 86.0 (28.0) | 96.0 (4.9) | 100.0 (0.0) | 100 (0) |
| window-open | 78.0 (39.2) | 92.0 (7.5) | 96.0 (4.9) | 100 (0) |
| Average | 49.8 (7.9) | 57.4 (7.0) | 64.6 (5.0) | 70.4 (1.85) |
Mean (standard deviation) over five seeds.
Gymnasium-Robotics Fetch

| Task | SAC | Relay | FuRL | LAGEA |
|---|---|---|---|---|
| Reach | 100 (0) | 100 (0) | 100 (0) | 100 (0) |
| Push | 26.67 (4.71) | 30 (8.16) | 40 (8.16) | 53.33 (4.71) |
| PickAndPlace | 10 (8.16) | 20 (0) | 33.33 (9.43) | 43.33 (4.71) |
| Slide | 0 (0) | 0 (0) | 3.33 (4.71) | 10 (8.16) |
| Average | 34.17 | 37.5 | 44.17 | 51.67 |
Mean (standard deviation) over three seeds.
Faster convergence

Averaged over nine MT10 settings, LAGEA converges in 92.4 minutes of wall-clock time versus 94.9 for FuRL, despite querying a VLM.
What matters
Ablations


| Variant | Success (%) | |
|---|---|---|
| Feedback format (6 tasks) | ||
| Free-form text | 78.89 | </tr>|
| Schema-constrained JSON | 98.89 | </tr>|
| Key-frame selection | ||
| Random frames | 68.0 | </tr>|
| Uniform frames | 67.3 | </tr>|
| LAGEA key frames | 80.0 | </tr>|
| VLM for reflection | ||
| SmolVLM2 | 56.0 | </tr>|
| OpenQwen2VL | 66.7 | </tr>|
| InternVL2 | 68.0 | </tr>|
| Qwen2.5-VL-3B | 80.0 | </tr>|
| Text encoder for feedback | ||
| MPNet | 70.7 | </tr>|
| BGE | 71.3 | </tr>|
| LIV | 72.7 | </tr>|
| GPT-2 | 80.0 | </tr>|
| Camera viewpoint (zero-shot) | ||
| Training view | 80.0 | </tr>|
| Directly overhead | 79.3 | </tr>|
| Front-left diagonal | 77.3 | </tr>|
| Behind-left diagonal | 79.3 | </tr> </tbody> </table> </div>|
Citation
@inproceedings{chowdhury2026lagea,
title = {{LAGEA}: Language Guided Embodied Agents for Robotic Manipulation},
author = {Chowdhury, Abdul Monaf and Mazumder, Akm Moshiur Rahman and
Arib, Safaeid Hossain and Akter, Rabeya},
booktitle = {Proceedings of the 43rd International Conference on Machine Learning},
series = {Proceedings of Machine Learning Research},
volume = {306},
year = {2026}
}