RoboHarn-Evo

Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation

Turn physical experience into reusable knowledge.

Shifeng Bao1,3* Fanding Huang2* Yihan Lin1,3 Youhe Feng1,3 Guanlin Li1,3Chen Zhao1,3 Yang Li1,3 Jiawei He5 Cheng Chi1,4‡ Jing Zhang1,4†

1 Renmin University of China2 Tsinghua University

Full affiliations

1 School of Information, Renmin University of China
2 Tsinghua University
3 Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China
4 Engineering Research Center of Database and Business Intelligence, Beijing, China
5 XYZ Embodied AI, Beijing, China

* Equal contribution† Corresponding author‡ Project leader

TASK 3 · UNCOVER, COUNT & PRESS ACTIONS 5× · WAITING 50× SCROLL TO EXPLORE
+24.2pp

RMBench success gain

Full HPK vs. no HPK · GPT-5.5
48.3 → 75.0%

Improvement through interaction

0 to 80 learning rollouts · GPT-5.5
41.7 → 62.5%

Real-world success

Simulation knowledge → after real-world RSI
THE IDEA

Experience should change
what a robot knows.

Good task reasoning only works when physical actions produce the intended effects. A successful grasp, for example, does not mean the object has been placed correctly.

RoboHarn-Evo turns interaction into Hierarchical Physical Knowledge (HPK): Task Knowledge guides which subtask to execute and when it is complete; Action Knowledge guides object-relative geometry and physical effects. A dual-loop harness retrieves this knowledge during execution, then revises it using evidence across episodes.

Read the abstract

Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop harness that evolves Hierarchical Physical Knowledge (HPK) from physical experience. HPK couples two levels of reusable knowledge: Task Knowledge captures which subtask should be executed and when it is complete, while Action Knowledge captures object-relative geometric strategies and their physical effects. During execution, the agent retrieves knowledge at the corresponding decision level and grounds it in the current scene under the task goal. Across episodes, physical feedback is used to revise historical knowledge, update its applicability, and organize reusable entries for subsequent retrieval. Experiments on RMBench show that HPK improves average success by up to 24.2 percentage points across different agent models. With 80 interaction rollouts, held-out success rises from 48.3% to 75.0% for GPT-5.5 and from 70.0% to 88.3% for GPT-6. RoboHarn-Evo also resolves over 83% of historical knowledge errors while retaining 95.8% of valid knowledge, and transfers zero-shot from RMBench to RoboDojo with gains of 35.0 and 25.0 percentage points. These results demonstrate that physical interaction can be accumulated into reusable knowledge for improving subsequent manipulation.

ROBOHARN-EVO / CROSS-TASK DEMONSTRATIONDownload video ↗
01 / THE METHOD

Two levels of knowledge.
Two loops of improvement.

The VLM and low-level executor stay fixed.
Experience changes the knowledge they use.

T
TASK KNOWLEDGE

What to do. When it is done.

Subtask selection, goals, ordering, and completion conditions.

Finish the required press count before switching to the next color.
A
ACTION KNOWLEDGE

How to make it happen.

Object-relative geometry, interaction strategies, and physical effects.

Approach from above, press once, release, and verify the effect.
The dual-loop harness. Local physical effects and subtask completion are checked separately.
Skill-routed retrieval & evidence-driven maintenance

Skill summaries route queries to related entries. The agent selects applicable atomic knowledge, while task goals constrain action grounding. Across episodes, observed outcomes support, oppose, or leave a claim unverified. Maintenance merges equivalent entries and revises their applicability using the physical evidence.

Skill-routed retrieval and evidence-driven maintenance from the paper.
02 / THE EVIDENCE

Better knowledge.
Better manipulation.

Matched within-model comparisons, frozen held-out
evaluations, and transfer to new environments.

RMBENCH · SIX TASKS

HPK improves different backbones.

Compare the same harness and executor with and without Hierarchical Physical Knowledge.

Without HPKFull HPKSUCCESS RATE (%)
View values as a table
RMBench task success rates (%) · GPT-5.5
TaskWithout HPKFull HPKGain (pp)
CONTINUED INTERACTION

Learning without weight updates.

GPT-5.5GPT-6

Mean held-out success over three learning histories. Knowledge snapshots are frozen at evaluation; evaluation feedback is excluded from updates.

View checkpoints & variability
Mean success ± sample standard deviation (%)
K rolloutsGPT-5.5GPT-6
KNOWLEDGE MAINTENANCE

Revise errors. Retain what works.

95.8%

of initially correct knowledge retained
23 / 24 entries · both models

Initial errors resolved83.3% GPT-5.591.7% GPT-6

After 80 source rollouts, 10/12 and 11/12 initial errors are repaired or deactivated. Knowledge quality is assessed by an external GPT-6 judge using recorded observations and execution evidence.

SIMULATION → REAL WORLD

Physical experience transfers.

On three real-world tasks, further knowledge updates raise mean success from 41.7% to 62.5%, and normalized task progress from 70.5% to 84.2%.

Paper evaluation: eight scenes per task. Full task completion counts as success. These aggregate results are separate from the three demonstrations below.

RMBench → RoboDojo
Zero-shot success gain with frozen source knowledge
+35.0pp · GPT-5.5+25.0pp · GPT-6
Real-world evaluation reported in the paper.
03 / REAL ROBOT DEMONSTRATIONS

Knowledge carried forward.

From simulation priors to real-world experience,
then to a composed manipulation task.

3× SPEED
Grasp the cupSOURCE 03:30
04 / RESOURCES

Build on RoboHarn-Evo.

Read the paper, explore the implementation,
or cite this work.

BIBTEX
@misc{bao2026roboharnevo,
  title = {RoboHarn-Evo: Evolving Hierarchical
    Physical Knowledge for Self-Improving
    Robotic Manipulation},
  author = {Shifeng Bao and Fanding Huang and
    Yihan Lin and Youhe Feng and Guanlin Li and
    Chen Zhao and Yang Li and Jiawei He and
    Cheng Chi and Jing Zhang},
  year = {2026},
  eprint = {2609.37583},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url = {https://arxiv.org/abs/2609.37583}
}

Paper figure