AtomTree
Accepted at IROS

AtomTreeA Hierarchical Framework for State-Aware Embodied Instruction Following with LLMs

Build the plan as the environment unfolds: decompose, execute, remember, and replan from the latest state.

Haotian Zhai University of Minnesota haoti002@umn.edu
Tianming Sha Stony Brook University tianming.sha@stonybrook.edu
Junnan Li University of Wisconsin-Madison jli2786@wisc.edu
Paper · Awaiting arXiv Code · Preparing release
01 · Method

Plan one decision at a time.

AtomTree incrementally constructs an explicit goal tree. At each active node, the LLM sees the current subgoal, symbolic state, and memory, then chooses whether to decompose the goal or execute one grounded action.

AtomTree method overview comparing supervised, few-shot, and state-aware hierarchical planning
Figure 2 from the paper. Unlike one-shot action sequences, AtomTree progressively refines the task into a state-aware tree whose leaves are executable actions.
01

Decompose or Execute

Every active subgoal is either refined into ordered children or grounded as one atomic action.

02

State-Aware Memory

Seen objects and action outcomes preserve long-horizon context without replaying the full trace.

03

Hierarchical Replanning

A failed or stalled action sends planning back to its parent, where a revised subtree is generated.

DetectAction fails or stalls
→
BacktrackReturn to the parent goal
→
RegenerateBuild a new subtree from updated state
Evaluation infrastructure

Reconstructing hidden test goals.

ALFRED's public test release omits the hidden task_type and pddl_params, preventing direct local evaluation on Test Seen and Test Unseen. We use an LLM through codex exec to recover these goal specifications from the public language instructions.

View reconstructed test labels ↗
Public
Instructions
→ codex exec → task_type
+ pddl_params
96.16% Exact match on the official validation entries
1,578 / 1,641 reconstructed correctly

Method development and parameter selection use only the public validation splits. These reconstructed goals are used only for AtomTree's final local evaluation on Test Seen and Test Unseen; the test splits are not used for tuning. They are unofficial inferred labels, not official ALFRED test ground truth. Many remaining mismatches involve ambiguity or noise in the human-written annotations and canonical target names.

02 · Results

Zero-shot. State-aware. Strong across splits.

On ALFRED, AtomTree achieves the highest reported Success Rate among the compared zero-shot methods on all four splits, while using only the high-level goal instruction.

40.98Valid SeenSuccess Rate (%)
42.39Valid UnseenSuccess Rate (%)
39.92Test SeenSuccess Rate (%)
41.20Test UnseenSuccess Rate (%)

Against a strong zero-shot baseline

Test Success Rate (%) · Socratic-Planner and AtomTree both operate zero-shot.

Test Seen18.66 → 39.92
Test Unseen13.47 → 41.20
Socratic-Planner · zero-shotAtomTree · zero-shot
03 · Demos

Five tasks, one evolving tree.

The featured tomato rollout above shows replanning. These four additional rollouts show memory, grounded object interaction, multi-object progress, and another long-horizon recovery.

Clean lettuce, then place it in the fridge

Recorded outcomes keep the planner aligned after cleaning so it can advance to placement.

Clean a plate and place it on the counter

The hierarchical goal is progressively refined into executable navigation and interaction actions.

Move two spray bottles to the toilet

Separate subgoals preserve progress across two object instances until both are placed.

Cook an egg, then return it to the fridge

Updated state and memory support a revised subtree when execution can no longer progress.

BibTeX

@inproceedings{zhai2026atomtree,
  title  = {AtomTree: A Hierarchical Framework for State-Aware
            Embodied Instruction Following with LLMs},
  author = {Zhai, Haotian and Sha, Tianming and Li, Junnan},
  year   = {2026}
}