ICML 2026 Embodied AIReinforcement LearningVLMs

LAGEA: Language Guided Embodied Agents for Robotic Manipulation

Abdul Monaf Chowdhury1, Akm Moshiur Rahman Mazumder2, Safaeid Hossain Arib1, Rabeya Akter1

1University of Dhaka   2Independent University, Bangladesh

Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)

TL;DR

A similarity score summarizes the outcome, not the cause. LAGEA asks a VLM to diagnose why an episode failed using a fixed error schema, anchors that diagnosis to the key frames where it happened, and turns it into a shaping reward that fades as the robot gets competent.

  • 80.0%MT10 fixed goals (+4.0 over FuRL)
  • 70.4%MT10 random goals (+5.8 over FuRL)
  • 51.7%Gymnasium-Robotics Fetch (+7.5 over FuRL)
Overview figure for LAGEA: Language Guided Embodied Agents for Robotic Manipulation
The LAGEA loop. Key frames from each trajectory are sent to a frozen VLM with an error taxonomy; the structured reflection is embedded, aligned with visual states, and fused with goal progress into a shaping reward for the policy.

Abstract

Robotic manipulation benefits from foundation models that describe goals, but today’s agents still lack a principled way to learn from their own mistakes. We ask whether natural language can serve as feedback, an error-reasoning signal that helps embodied agents diagnose what went wrong and correct course. We introduce LaGEA (Language Guided Embodied Agents), a framework that turns episodic, schema-constrained reflections from a vision language model (VLM) into temporally grounded guidance for reinforcement learning. LaGEA summarizes each attempt in concise language, localizes the decisive moments in the trajectory, aligns feedback with visual state in a shared representation, and converts goal progress and feedback agreement into bounded, step-wise shaping rewards whose influence is modulated by an adaptive, failure-aware coefficient. This design yields dense signals early when exploration needs direction and gracefully recedes as competence grows. On the Meta-World MT10 and Robotic Fetch embodied manipulation benchmark, LaGEA improves average success over the state-of-the-art (SOTA) methods by 9.0% on random goals, 5.3% on fixed goals, and 17% on fetch tasks, while converging faster. These results support our hypothesis: language, when structured and grounded in time, is an effective mechanism for teaching robots to self-reflect on mistakes and make better choices.

From score to diagnosis

Motivation

Prior VLM-as-reward methods compare the current observation with a goal description and return a similarity score. That score says how far the robot is from success, but not why it failed: whether it grasped the wrong object, came in from the wrong direction, or pushed with too little force. LAGEA makes the VLM produce that diagnosis explicitly and uses it as a learning signal.

Examples of structured VLM feedback with error codes and explanations
Structured reflections. The VLM answers in a fixed JSON schema with an error code (e.g. bad_approach_direction, failed_grasp, insufficient_force) and a short explanation, which keeps feedback consistent and machine-usable.

Method

Four ideas

  1. Find the moments that matterKey frames are chosen by combining goal similarity with motion cues (velocity and acceleration), so the VLM looks at the decisive part of the episode rather than at uniformly sampled frames.
  2. Put feedback and pixels in one spaceSmall projectors, trained with a binary alignment loss and InfoNCE, embed feedback, goal, and visual states in a shared space.
  3. Reward progress, not positionPotential-based shaping rewards the change in goal and feedback agreement, weighted by how well the instruction and the reflection agree.
  4. Let the guidance fadeAn adaptive schedule driven by a moving average of success reduces VLM influence as the policy becomes competent.

Results

Average success rate (%) · higher is better

BenchmarkSACRelayFuRLLAGEA
Meta-World MT10 · fixed goals (5 seeds)37.054.076.080.0
Meta-World MT10 · random goals (5 seeds)49.857.464.670.4
Gymnasium-Robotics Fetch · 4 tasks (3 seeds)34.1737.544.1751.67
Success-rate learning curves for SAC, FuRL and LAGEA on eight MT10 tasks
Faster convergence. Success over 1M environment steps on eight MT10 tasks. LAGEA solves several tasks substantially earlier than FuRL.

What matters

Ablations on MT10 (fixed goals)

  • Key frames: LAGEA's key-frame selection (80.0%) outperforms random (68.0%) and uniform (67.3%) frame choices.
  • VLM choice: Qwen2.5-VL-3B (80.0%) beats InternVL2 (68.0%), OpenQwen2VL (66.7%), and SmolVLM2 (56.0%).
  • Viewpoint robustness: performance holds across unseen camera viewpoints (77.3–79.3% vs. 80.0% on the training view).

Citation

@inproceedings{chowdhury2026lagea,
  title     = {{LAGEA}: Language Guided Embodied Agents for Robotic Manipulation},
  author    = {Chowdhury, Abdul Monaf and Mazumder, Akm Moshiur Rahman and
               Arib, Safaeid Hossain and Akter, Rabeya},
  booktitle = {Proceedings of the 43rd International Conference on Machine Learning},
  series    = {Proceedings of Machine Learning Research},
  volume    = {306},
  year      = {2026}
}