What to do. When it is done.
Subtask selection, goals, ordering, and completion conditions.
Finish the required press count before switching to the next color.
Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation
Turn physical experience into reusable knowledge.
RMBench success gain
Full HPK vs. no HPK · GPT-5.5Improvement through interaction
0 to 80 learning rollouts · GPT-5.5Real-world success
Simulation knowledge → after real-world RSIGood task reasoning only works when physical actions produce the intended effects. A successful grasp, for example, does not mean the object has been placed correctly.
RoboHarn-Evo turns interaction into Hierarchical Physical Knowledge (HPK): Task Knowledge guides which subtask to execute and when it is complete; Action Knowledge guides object-relative geometry and physical effects. A dual-loop harness retrieves this knowledge during execution, then revises it using evidence across episodes.
Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop harness that evolves Hierarchical Physical Knowledge (HPK) from physical experience. HPK couples two levels of reusable knowledge: Task Knowledge captures which subtask should be executed and when it is complete, while Action Knowledge captures object-relative geometric strategies and their physical effects. During execution, the agent retrieves knowledge at the corresponding decision level and grounds it in the current scene under the task goal. Across episodes, physical feedback is used to revise historical knowledge, update its applicability, and organize reusable entries for subsequent retrieval. Experiments on RMBench show that HPK improves average success by up to 24.2 percentage points across different agent models. With 80 interaction rollouts, held-out success rises from 48.3% to 75.0% for GPT-5.5 and from 70.0% to 88.3% for GPT-6. RoboHarn-Evo also resolves over 83% of historical knowledge errors while retaining 95.8% of valid knowledge, and transfers zero-shot from RMBench to RoboDojo with gains of 35.0 and 25.0 percentage points. These results demonstrate that physical interaction can be accumulated into reusable knowledge for improving subsequent manipulation.
The VLM and low-level executor stay fixed.
Experience changes the knowledge they use.
Subtask selection, goals, ordering, and completion conditions.
Finish the required press count before switching to the next color.
Object-relative geometry, interaction strategies, and physical effects.
Approach from above, press once, release, and verify the effect.
Skill summaries route queries to related entries. The agent selects applicable atomic knowledge, while task goals constrain action grounding. Across episodes, observed outcomes support, oppose, or leave a claim unverified. Maintenance merges equivalent entries and revises their applicability using the physical evidence.

Matched within-model comparisons, frozen held-out
evaluations, and transfer to new environments.
Compare the same harness and executor with and without Hierarchical Physical Knowledge.
| Task | Without HPK | Full HPK | Gain (pp) |
|---|
Mean held-out success over three learning histories. Knowledge snapshots are frozen at evaluation; evaluation feedback is excluded from updates.
| K rollouts | GPT-5.5 | GPT-6 |
|---|
of initially correct knowledge retained
23 / 24 entries · both models
After 80 source rollouts, 10/12 and 11/12 initial errors are repaired or deactivated. Knowledge quality is assessed by an external GPT-6 judge using recorded observations and execution evidence.
On three real-world tasks, further knowledge updates raise mean success from 41.7% to 62.5%, and normalized task progress from 70.5% to 84.2%.
Paper evaluation: eight scenes per task. Full task completion counts as success. These aggregate results are separate from the three demonstrations below.
From simulation priors to real-world experience,
then to a composed manipulation task.
Read the paper, explore the implementation,
or cite this work.
@misc{bao2026roboharnevo,
title = {RoboHarn-Evo: Evolving Hierarchical
Physical Knowledge for Self-Improving
Robotic Manipulation},
author = {Shifeng Bao and Fanding Huang and
Yihan Lin and Youhe Feng and Guanlin Li and
Chen Zhao and Yang Li and Jiawei He and
Cheng Chi and Jing Zhang},
year = {2026},
eprint = {2609.37583},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.37583}
}