LAGEA: Language Guided Embodied Agents for Robotic Manipulation
1University of Dhaka 2Independent University, Bangladesh
Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)
TL;DR
A similarity score summarizes the outcome, not the cause. LAGEA asks a VLM to diagnose why an episode failed using a fixed error schema, anchors that diagnosis to the key frames where it happened, and turns it into a shaping reward that fades as the robot gets competent.
- 80.0%MT10 fixed goals (+4.0 over FuRL)
- 70.4%MT10 random goals (+5.8 over FuRL)
- 51.7%Gymnasium-Robotics Fetch (+7.5 over FuRL)

Abstract
Robotic manipulation benefits from foundation models that describe goals, but today’s agents still lack a principled way to learn from their own mistakes. We ask whether natural language can serve as feedback, an error-reasoning signal that helps embodied agents diagnose what went wrong and correct course. We introduce LaGEA (Language Guided Embodied Agents), a framework that turns episodic, schema-constrained reflections from a vision language model (VLM) into temporally grounded guidance for reinforcement learning. LaGEA summarizes each attempt in concise language, localizes the decisive moments in the trajectory, aligns feedback with visual state in a shared representation, and converts goal progress and feedback agreement into bounded, step-wise shaping rewards whose influence is modulated by an adaptive, failure-aware coefficient. This design yields dense signals early when exploration needs direction and gracefully recedes as competence grows. On the Meta-World MT10 and Robotic Fetch embodied manipulation benchmark, LaGEA improves average success over the state-of-the-art (SOTA) methods by 9.0% on random goals, 5.3% on fixed goals, and 17% on fetch tasks, while converging faster. These results support our hypothesis: language, when structured and grounded in time, is an effective mechanism for teaching robots to self-reflect on mistakes and make better choices.
From score to diagnosis
Motivation
Prior VLM-as-reward methods compare the current observation with a goal description and return a similarity score. That score says how far the robot is from success, but not why it failed: whether it grasped the wrong object, came in from the wrong direction, or pushed with too little force. LAGEA makes the VLM produce that diagnosis explicitly and uses it as a learning signal.

bad_approach_direction, failed_grasp, insufficient_force) and a short explanation, which keeps feedback consistent and machine-usable.Method
Four ideas
- Find the moments that matterKey frames are chosen by combining goal similarity with motion cues (velocity and acceleration), so the VLM looks at the decisive part of the episode rather than at uniformly sampled frames.
- Put feedback and pixels in one spaceSmall projectors, trained with a binary alignment loss and InfoNCE, embed feedback, goal, and visual states in a shared space.
- Reward progress, not positionPotential-based shaping rewards the change in goal and feedback agreement, weighted by how well the instruction and the reflection agree.
- Let the guidance fadeAn adaptive schedule driven by a moving average of success reduces VLM influence as the policy becomes competent.
Results
Average success rate (%) · higher is better
| Benchmark | SAC | Relay | FuRL | LAGEA |
|---|---|---|---|---|
| Meta-World MT10 · fixed goals (5 seeds) | 37.0 | 54.0 | 76.0 | 80.0 |
| Meta-World MT10 · random goals (5 seeds) | 49.8 | 57.4 | 64.6 | 70.4 |
| Gymnasium-Robotics Fetch · 4 tasks (3 seeds) | 34.17 | 37.5 | 44.17 | 51.67 |

What matters
Ablations on MT10 (fixed goals)
- Key frames: LAGEA's key-frame selection (80.0%) outperforms random (68.0%) and uniform (67.3%) frame choices.
- VLM choice: Qwen2.5-VL-3B (80.0%) beats InternVL2 (68.0%), OpenQwen2VL (66.7%), and SmolVLM2 (56.0%).
- Viewpoint robustness: performance holds across unseen camera viewpoints (77.3–79.3% vs. 80.0% on the training view).
Citation
@inproceedings{chowdhury2026lagea,
title = {{LAGEA}: Language Guided Embodied Agents for Robotic Manipulation},
author = {Chowdhury, Abdul Monaf and Mazumder, Akm Moshiur Rahman and
Arib, Safaeid Hossain and Akter, Rabeya},
booktitle = {Proceedings of the 43rd International Conference on Machine Learning},
series = {Proceedings of Machine Learning Research},
volume = {306},
year = {2026}
}