RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents

1Nanyang Technological University 2Beihang University 3National University of Singapore 4Cloud Butterfly Technology 5Shanghai Jiao Tong University
*Equal Contribution, ‡Corresponding Author

Abstract

A foundation model should not act in isolation as an embodied agent. Yet, existing methods often optimize individual components of the agent stack, such as memory, context, skills, or action interfaces, rather than treating the supporting system itself as a unified policy. Moreover, interaction alone does not yield self-improvement unless execution experience is converted into persistent, validated system changes. We therefore propose RoboFoundry, the first embodied agentic framework that formulates this process as Self-Evolving System-as-Policy. RoboFoundry diagnoses capability gaps in decision-making and memory management, converts execution traces into validated task-specific system updates, and promotes recurring improvements to the general system. Evolution operates over two complementary surfaces: a context system that manages active internal context and persistent file-system memory, and a hierarchical skill system that organizes atomic skills, reusable compositions, and failure-conditioned recovery. A shared semantic interface separates embodiment-invariant decisions from embodiment-specific execution, allowing evolved system capabilities to transfer across heterogeneous robots. On EmbodiedBench, RoboFoundry achieves state-of-the-art performance, notably improving GPT-5.5 by 27.8%. It also brings Qwen3.7-Plus to near parity with GPT-5.5 (70.3% vs. 72.7%), showing consistent gains from system-as-policy evolution across foundation models. For long-horizon memory, RoboFoundry outperforms all baselines on RoboMemArena by at least 39.0%, even against methods assisted by external foundation models. On LIBERO-PRO, it further outperforms Cap-Agent0 by 243.8–679.7% across all perturbation types. In real-world deployments, RoboFoundry demonstrates zero-shot transfer and online evolution across robots and tasks, highlighting its potential for fully autonomous embodied agents.

RoboFoundry overview: the Act–Reflect–Repair–Promote loop across embodiments, with results on EmbodiedBench, RoboMemArena, and LIBERO-PRO.
From Code-as-Policy execution to System-as-Policy evolution: The central Act–Reflect–Repair–Promote loop executes across embodiments, attributes failures to capabilities, revises the task-level system, and promotes validated changes to the general system.

Stronger systems, more capable agents: Explore RoboFoundry across foundation models and EmbodiedBench’s four embodied task suites.

View all results · Success rate (%)
Overall comparison on EmbodiedBench. Avg. is the arithmetic mean of success rates across the four suites. All scores are percentages. Superscript gains are reported in percentage points (+pp) over the corresponding foundation model; bold scores indicate the best result in each column.
MethodAvg.EB-ALFREDEB-HabitatEB-NavigationEB-Manipulation
RoboFoundry (GPT-6 Astra)78.0+6.190.0+2.786.7+11.480.2+3.254.9+6.9
RoboFoundry (GPT-5.5)72.7+15.884.0+7.388.7+24.772.0+17.345.9+13.8
RoboFoundry (Qwen3.7-Plus)70.3+10.481.3+8.080.3+15.371.9+7.647.7+10.5
RoboFoundry (GLM5.3-Flash)69.3+7.480.0+8.375.0+14.372.7+1.749.5+5.1
RoboFoundry (Qwen3.8-27B)66.5+14.278.0+12.377.0+19.068.3+16.042.8+9.5
GPT-6 Astra71.987.375.377.048.0
GLM5.3-Flash62.071.760.771.044.4
Qwen3.7-Plus60.073.365.064.337.2
GPT-5.556.976.764.054.732.1
Qwen3.8-27B52.365.758.052.333.3
Claude-3.5-Sonnet50.564.068.044.725.4
GPT-4o50.456.359.057.728.5
Claude-3.7-Sonnet49.967.758.745.028.3

Better long-term memory : Explore RoboFoundry leading performance across RoboMemArena’s four task categories.

View all results · TSR / CSR (%)
Long-horizon memory results on RoboMemArena. RoboFoundry achieves the best TSR and CSR overall and across all four categories. Each entry reports TSR / CSR (%).
MethodOverallTransferOcclusionCountingSequence
RoboFoundry53.5 / 72.870.5 / 78.138.3 / 57.655.2 / 82.575.2 / 92.1
PrediMem38.5 / 55.222.5 / 45.227.3 / 38.445.7 / 69.372.5 / 89.5
MemER27.3 / 49.120.0 / 36.116.4 / 33.227.1 / 65.165.0 / 79.1
HiF-VLA16.9 / 39.817.5 / 38.912.7 / 27.18.6 / 45.942.5 / 70.2
π0.521.5 / 38.720.0 / 42.812.7 / 17.214.3 / 50.960.0 / 71.6
MemoryVLA15.0 / 35.315.0 / 37.27.3 / 13.114.3 / 55.137.5 / 65.2

Adapting to perturbations: Explore how RoboFoundry improves position and task success rates under object, goal, and spatial perturbations on LIBERO-PRO.

View all results · Position / task success (%)
LIBERO-PRO results under perturbations. Each entry reports position / task success rate (%). † denotes privileged simulator object poses.
MethodObjectGoalSpatial
OpenVLA0.0 / 0.00.0 / 0.00.0 / 0.0
π00.0 / 0.00.0 / 0.00.0 / 0.0
π0.517.0 / 1.038.0 / 0.020.0 / 1.0
CaP-Agent021.8 / 18.225.6 / 16.811.8 / 14.0
ASPIRE98.0 / 95.081.0 / 45.051.0 / 60.0
Harness VLA (Codex)81.0 / 69.094.0 / 91.075.0 / 66.0
Harness VLA (CC)94.0 / 80.088.0 / 90.087.0 / 87.5
RoboFoundry96.0 / 98.088.0 / 86.092.0 / 91.5
RoboFoundry†99.0 / 100.096.0 / 100.098.0 / 99.2

Better memory across agent configurations: Explore how RoboFoundry improves long-term memory in standalone π0.5 and PrediMem through context evolution.

View all results · TSR / CSR (%)
Context evolution across agent configurations. RoboFoundry improves both standalone π0.5 and PrediMem across all four categories. Each entry reports TSR / CSR (%); averages are reproduced as reported.
Memory categoryπ0.5PrediMem
Base+RoboFoundryBase+RoboFoundry
Transferring20.0 / 42.862.5 / 71.222.5 / 45.267.5 / 76.2
Occlusion12.7 / 17.233.6 / 47.227.3 / 38.436.4 / 56.4
Counting14.3 / 50.958.8 / 79.645.7 / 69.353.7 / 80.3
Sequence60.0 / 71.670.0 / 90.872.5 / 89.573.0 / 90.5
Avg.26.8 / 45.656.2 / 72.242.0 / 60.657.7 / 75.9

Task-specific commits, general-system improvements: Follow RoboFoundry as it validates task-specific updates and promotes them to the general system, raising held-in success from 21.4% to 78.0% on the EB-Habitat spatial subset.

View the full evolution trace
EB-Habitat spatial subset, RoboFoundry (GPT-5.5). Scores are held-in success rates (%); held-out checks are reported separately. The animation replays the reported trace with illustrative timing.
EvaluationCandidate (%)Retained (%)Held-out checkDecision
021.421.4SkippedInitial system
135.735.7PassCommit
242.942.9PassCommit
350.050.0PassCommit
442.950.0FailDiscard
542.950.0SkippedDiscard
678.078.0PassTag & promote

See the system in action: Explore RoboFoundry across different real-world experiments. Videos are shown at 2× speed.

BibTeX

@misc{liang2026robofoundry,
  title={RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents},
  author={Jingsong Liang and Shuhao Liao and Shizhe Zhang and Diyuan Hou and Yuxin Cai and Xinjian Deng and Chengyang He and Wenhui Huang and Runjia Tan and Zhidong Wang and Lan Yu and Xuesong Tian and Guillaume Sartoretti and Jie Luo and Yao Mu and Wenjun Wu and Wanhua Li and Chen Lv},
  year={2026},
  eprint={2609.32862},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2609.32862}
}