Two arms share one coordinate frame, and the agent decides how to use them. It watches an overhead demonstration video and reproduces the folds: first the near long edge onto the far long edge, to make a doubled strip, and then both short ends inward. Two configurations were evaluated with five trials each. In sequential folding the demonstration folds the right short end first and then the left. In simultaneous folding the demonstration brings both short ends inward together. The agent chooses the grasp points, which arm does what and when, and may stage a placement to keep the arms clear of each other. The towel changes shape during grasping and folding.
Success criterionSuccess requires both inward folds with reasonably aligned edges, verified in an overhead image after both grippers open and both arms withdraw. Small wrinkles and creases are allowed, and the final position on the table is not restricted.
Sequential 5 of 5, at a mean 50.8 minutes and USD 24.14, among the highest in the suite. Simultaneous 3 of 5, at a mean 21.4 minutes and USD 7.46, the lowest success rate of the eight configurations.
The agent watches one human demonstration video and then assembles eight 3D-printed parts into four pairs in the demonstrated order: a ridged sleeve over a broad round base, a hexagonal nut over a short post, a smooth sleeve over a stepped cylinder, and a long pin into a hollow tube. Nothing tells the agent which part mates with which; it has to read the relationships and the insertion motions out of the video and out of the scene in front of it. The parts are printed from AutoMate meshes, with a nominal radial clearance as small as 0.74 mm, requiring AGP to align and insert the parts precisely.
Success criterionSuccess requires all four pairs to be assembled at the end, with pairs that were finished earlier still intact, and everything stable after the gripper opens and the arm withdraws.
8 of 10 trials succeeded. Successful trials took a mean of 37.2 minutes, ranging from 16.0 to 59.1, with mean usage of 13.02 million tokens and mean inference cost of USD 16.62.
The four pairs
The four pairs
Pair
Hole diameter (mm)
Shaft diameter (mm)
Radial clearance (mm)
Ridged sleeve and broad round base
37.41
35.86
0.78
Hexagonal nut and short post
28.18
25.80
1.19
Smooth sleeve and stepped cylinder
42.86
37.20
2.83
Long pin and hollow tube
18.01
16.54
0.74
Scroll the table sideways for the rest of the columns.
Nominal print dimensions as recorded by the authors, in demonstration order. Radial clearance is half the diametral clearance, computed from the two columns beside it and rounded once, and gives the transverse offset available when the two axes are parallel. What the parts actually tolerate also depends on printing accuracy and on angular alignment.
Ridged sleeve and broad round base. Radial clearance 0.78 mm.Hexagonal nut and short post. Radial clearance 1.19 mm.Smooth sleeve and stepped cylinder. Radial clearance 2.83 mm.Long pin and hollow tube. Radial clearance 0.74 mm, the tightest of the four.An assembly example after 47.2 minutes. All four pairs are assembled.
The agent gets six cubes on a table and a single photograph of the structure it should build. It has to work out which cube is which, how they support each other, and an order of placements that keeps the partial structure standing. Three goal structures were evaluated with ten trials each: a pyramid with three blue cubes in the base, two grey cubes bridging the base seams and one grey cube at the peak; two towers whose colour orders run blue, grey, blue and grey, blue, grey from the bottom; and a six block tower alternating grey and blue above a grey base that stays where it is.
Success criterionSuccess requires the complete target structure, still standing after the gripper opens and the arm withdraws.
Pyramid succeeded in 10 of 10 trials at a mean of 21.6 minutes, two towers in 10 of 10 at 20.6 minutes, and the six block tower in 9 of 10 at 28.2 minutes. Together, 29 of 30 block construction trials succeeded.
The goal image for the pyramid.Pyramid example after the arm withdrew.
The goal image for the two towers.Two towers example in its final state.
The goal image for the six block tower.Six block tower example in its final state.
Six dice lie on the table with different faces upward. The agent receives a language instruction to turn them so that the upward face matches the goal, together with a goal photograph that fixes which face that is. The initial and target faces are balanced across predefined conditions. Nothing restricts where the dice end up or how they are rotated in the table plane, so the agent has to choose a grasp for each die that makes the required rotation reachable, and then release it so that it settles on the intended face.
Success criterionSuccess requires all six dice to show the requested number after placement. The task prompt leaves the dice positions and in-plane rotations unspecified.
10 of 10 trials succeeded. Successful trials took a mean of 37.9 minutes, ranging from 22.6 to 56.0, with mean usage of 17.08 million tokens and mean inference cost of USD 21.07.
The goal photograph supplied with the instruction: three black and three light blue dice, six pips upward.Dice flipping example after the arm withdrew.
The agent receives a text instruction and a goal photograph specifying the requested face.
A language instruction asks the agent to pick up a potato-shaped object and throw it into a stationary bowl. The required motion is a circular swing driven by one joint, with the other five joint targets held fixed and the gripper opening during the swing, so the outcome depends on the relationship between arm motion and release timing. This configuration uses a separate runtime with timed motion programs: the agent specifies joint offsets at time knots and a schedule for the gripper, and the executor interpolates the commands on a 50 Hz grid.
Success criterionSuccess requires the object to separate from the moving hand, travel through the air, and come to rest inside the target region, which here is the bowl.
Two recorded sessions, both successful, at a mean 22.8 minutes, from 22.2 to 23.3, at an average of 6.26 million tokens and USD 8.94. The throw itself completes an average of 13.9 minutes after the task is delivered, so about 40% of that time is the analysis and verification that follow it.
The two sessions were selected rather than drawn from a planned batch. They document successful throws but do not establish a success rate. In both sessions, the object was released during the swing and came to rest in the bowl. Occlusion leaves the rising part of the flight and the apex unresolved.
An image the agent made during a throwing experiment, showing eight wrist frames around the release labelled with the time since the motion program was dispatched.
Another agent-made image from the same trial: the wrist view with the fitted bowl rim drawn on top, used to place the target before the swing.A throwing example after 22.2 minutes with a single arm. The object is at rest inside the bowl.
Which agent, which model, how much thinking
Eight configurations, five trials each. Three of them do badly, and they stay in the table.
The comparison uses a smaller assembly task: four parts forming two pairs, a hexagonal ring over a short post and a circular sleeve over a stepped cylinder, learned from one human demonstration. Every configuration receives the same task references, the same robot interface and the same time and request budgets. Success requires both pairs to remain assembled after release and arm withdrawal.
Agent, model and thinking effort on two pair assembly. GPT-6 Astra appears at three thinking efforts; the other models use their high setting.
Scroll the table sideways for the rest of the columns.
Reading the table
GPT-6 Astra succeeds in every trial at all three thinking efforts, in about nine to ten minutes and USD 4.09 to 4.79 per success, so on this task more thinking effort buys little.
GPT-5.6 Sol also succeeds in every trial, more slowly and more cheaply than Astra.
Within Claude Code, Fable 5.1 succeeds less often than Opus 5 and uses fewer tokens on its successful trials, yet takes more time and costs more per success.
Comparisons across Codex CLI and Claude Code are comparisons of whole systems, not of models alone.
The single Terra figure is one trial, and the time, token and cost cells for that row have no spread because there is only one successful trial to average.
Table view
Model comparison, per trial: the data plotted in the figure above.
Means, minima and maxima are over successful trials only. GPT-5.6 Luna has no successful trial, and seven failed trials have no usable resource record.
Source: the pinned per-trial two pair assembly records used for model comparison.
One mark per trial for each of the eight configurations, successes filled and failures hollow. Each trial uses a single arm.
GPT-6 Astra, high effort
Two pair assembly example (5/5). GPT-6 Astra at high effort in Codex CLI, single arm. 7.9 min real at 8x. The other arm runs a separate trial outside this crop.
Task
Two pair assembly
Outcome
success
Model
GPT-6 Astra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
7.9 min real
Clip
58 s
GPT-5.6 Sol
Two pair assembly example (5/5). GPT-5.6 Sol at high effort in Codex CLI, single arm. 9.1 min real at 8x. Nothing else is running: four parts, one arm.
Task
Two pair assembly
Outcome
success
Model
GPT-5.6 Sol (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
9.1 min real
Clip
64 s
Claude Opus 5
Two pair assembly example (5/5). Claude Opus 5 at high effort in Claude Code, single arm. 14.4 min real at 8x. No failed commands; its report notes that the steep overhead view made the upright posts look like they were lying down, and that it checked with wrist depth.
Task
Two pair assembly
Outcome
success
Model
Claude Opus 5 (high)
Harness
Claude Code
Arm
single
Speed
8x
Real time
14.4 min real, 12.8 min shown
Clip
96 s
The other arm on the rail sometimes runs a separate trial at the same time. Where that is visible in a clip, the caption says so and the frame is cropped to the arm being discussed.
What these numbers rest on
Time, token and cost cells are computed from the per-trial usage records. The Success column is different. The resource rows can be recomputed from data files; the denominators cannot all be. All five GPT-5.6 Luna sessions and all five Claude Fable 5.1 sessions are now on record, but three of the Luna sessions ended without a final report, through a timeout or an abort, and one Fable session has no report either. Those outcomes are recorded by the authors.
Reusing experience
Every trial so far started from nothing. These three studies let the agent keep what it learned. Experience is ordinary files: measurements, procedures, corrections and working scripts, written at the end of one execution and read at the start of the next. Each execution still begins with a fresh agent context, so the only thing that carries over is what was written down.
Two comparisons on this page are two separate clips, not one composed video: ring cycle 1 beside ring cycle 4, and the same model with and without the transferred store. Each pair starts at the same task second and runs at the same speed, so the shorter run simply finishes first. Play them together; each still states its own trial and real elapsed time.
One arm takes a ridged ring off its base and puts it back, five times over, in a single continuous session with GPT-6 Astra at high thinking effort. Each cycle starts a new agent process that reads the previous checkpoint and updates the object poses from what it can see now. The first cycle has nothing to read: its first insertion stalls, the agent lifts the part safely, measures an offset of about 9 mm between the centre of the grasped ring and the tool, and corrects for it. Later cycles reuse that corrected alignment procedure and check it against current observations.
Ring, cycle 1: 870 s
Cycle 1 of 5, with nothing in the experience store. The first insertion stalls; the agent measures a 9 mm grasp offset and retries.
Task
Ring disassembly and reassembly, cycle 1
Outcome
success after one retry
Model
GPT-6 Astra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
14.5 min real
Clip
109 s
Ring, cycle 4: 372 s
The fourth cycle of the same continuous session, run by a fresh agent that reads what the earlier cycles wrote down. No retries.
Task
Ring disassembly and reassembly, cycle 4
Outcome
success
Model
GPT-6 Astra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
6.2 min real
Clip
47 s
Seconds per cycle
Seconds per cycle
Cycle
Total runtime
Reasoning and programming
Tool execution
Experience writing
Retries
Calls
1
989.7
702.9
167.4
119.4
1
37 / 37
2
522.9
282.1
126.7
114.1
0
32 / 33
3
489.9
245.0
127.3
117.6
0
32 / 33
4
484.4
247.7
124.6
112.2
0
32 / 33
5
461.7
268.3
125.0
68.5
0
32 / 33
Scroll the table sideways for the rest of the columns.
Total runtime falls from 989.7 to 461.7 seconds. The largest reduction follows the first cycle, after the agent establishes an alignment procedure and recovers from a failed insertion. Most of the savings come from reasoning and programming.
Timing definitions
Task execution includes reasoning and programming and tool execution. Tool execution covers recorded time running programs, processing images and operating the robot. Overlapping activities count once, and tool time takes precedence when tools run during a model response. The remaining task time estimates reasoning and programming, including model response and tool startup delays. Experience writing covers the subsequent saving of updated experience. Components sum to total runtime up to rounding. Calls lists motion and gripper requests followed by image capture requests.
These are five dependent repetitions in one session, with no matched control in which every cycle starts from empty experience. They cannot separate the effect of the saved experience from anything else that changes between cycles.
Table view
Reusing experience: ring disassembly and reassembly — seconds per execution.
#
Reasoning and programming
Tool execution
Experience writing
Total
1
702.9
167.4
119.4
989.7
2
282.1
126.7
114.1
522.9
3
245.0
127.3
117.6
489.9
4
247.7
124.6
112.2
484.4
5
268.3
125.0
68.5
461.7
All values in seconds. Components sum to total runtime up to rounding.
Source: Appendix Table 7, Execution breakdown over repeated tasks.
Ring task timing over five cycles. Most of the reduction comes from reasoning and programming.
Five executions use GPT-6 Astra at high thinking effort with a single arm. Each starts with a fresh agent context and reads saved experience and programs from earlier executions. Total runtime decreases from 621.4 to 462.4 seconds. Reasoning and programming falls from 351.9 to 186.6 seconds, while tool execution rises slightly from 160.5 to 175.5 seconds.
Successful assembly is assessed by the agent from recorded images. Task execution includes the return to the observation posture, followed by experience writing.
Timing definitions
Task execution includes reasoning and programming and tool execution. Tool execution covers recorded time running programs, processing images and operating the robot. Overlapping activities count once, and tool time takes precedence when tools run during a model response. The remaining task time estimates reasoning and programming, including model response and tool startup delays. Experience writing covers the subsequent saving of updated experience. Components sum to total runtime up to rounding. Calls lists motion and gripper requests followed by image capture requests.
Table view
Reusing experience: two pair assembly — seconds per execution.
#
Reasoning and programming
Tool execution
Experience writing
Total
1
351.9
160.5
108.9
621.4
2
281.7
186.3
100.1
568.1
3
237.1
173.8
96.6
507.5
4
221.0
174.1
96.6
491.8
5
186.6
175.5
100.3
462.4
All values in seconds. Components sum to total runtime up to rounding.
Source: Appendix Table 7, Execution breakdown over repeated tasks.
Two pair assembly timing over five executions. Reasoning and programming decreases while tool execution takes slightly longer.
GPT-6 Astra ran the two pair assembly task five times and saved geometry, procedures, corrections and scripts. GPT-5.6 Terra, at high thinking effort, then received that collection with read access only, frozen, unchanged between trials. The comparison is Terra with nothing against Terra reading the Astra collection, with the same task references, the same interface and the same budgets on both sides.
Handing experience to a smaller model
Condition
Success
Time (min)
Tokens (M)
Cost (USD)
GPT-5.6 Terra, empty experience
1/5
19.0
6.28
1.98
GPT-5.6 Terra, reading the Astra collection
4/5
12.5 [10.1, 16.9]
5.56 [4.81, 5.87]
1.64 [1.56, 1.74]
GPT-6 Astra, the five runs that built the collection
5/5
7.6 [6.5, 9.1]
2.50 [2.24, 2.83]
3.79 [3.45, 4.39]
Scroll the table sideways for the rest of the columns.
Means over successful trials, with the observed minimum and maximum beneath, the same accounting as the model comparison above. The empty-experience figure is the one successful Terra trial from that table, so it has no spread.
The 5 of 5 in the third row is one of the four Success counts on this page that rest on the authors' own record. All five runs are on file with full time, token and cost accounting, but the operator verdict column of that batch record was left blank, so the count comes from the single decision of 11 September 2026, in which Yang Lin, the author who ran the trials, confirmed all four of these counts, rather than from the record.
No experience: 19.0 min
Two pair assembly, GPT-5.6 Terra at high effort in Codex CLI, single arm, empty experience store: the only success in that row (1/5). 19.0 min real at 8x, with several settle-miss recoveries. Shown from task second 90 at the same speed as the clip beside it.
Task
Two pair assembly, empty experience
Outcome
success
Model
GPT-5.6 Terra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
19.0 min real, 17.7 min shown
Clip
132 s
Reading a frozen experience store: 11.2 min
The same task and the same model, now reading a frozen experience store written by GPT-6 Astra and never written back to: 4/5 successful, 11.2 min real. Same 8x and the same task second 90 start, so the clip simply ends sooner.
Task
Two pair assembly, frozen experience from GPT-6 Astra
Outcome
success
Model
GPT-5.6 Terra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
11.2 min real, 9.5 min shown
Clip
71 s
With the collection in front of it, Terra succeeds in four trials instead of one, finishes faster, and uses fewer tokens. It does not reach the model that wrote the collection: Astra still finished in 7.6 minutes on average. Cost moves the other way, because Terra is the cheaper model per token.
Five trials per condition. The conclusion rests on very few successful runs.
The collection is frozen. Terra can read it but cannot add to it, so this measures transfer, not further accumulation.
Two runs of this study were discarded because of infrastructure problems and are excluded from the reported results.
Table view
Experience transfer, per trial: the data plotted in the figure above.
GPT-5.6 Terra, reading a frozen GPT-6 Astra experience store
4/5
12.5 [10.1, 16.9]
5.56 [4.81, 5.87]
1.64 [1.56, 1.74]
10.1, 11.2, 11.7, 16.9, 19.6 (failed)
GPT-6 Astra, building the experience store
5/5
7.6 [6.5, 9.1]
2.50 [2.24, 2.83]
3.79 [3.45, 4.39]
6.5, 7.0, 7.3, 8.2, 9.1
Means, minima and maxima are over successful trials only. The third row is a different model, shown for reference.
Source: the pinned per-trial records of the two pair assembly experience transfer runs.
Completion time and token use for GPT-5.6 Terra with and without the Astra collection, over successful trials.
In the agent's own words
One failed assembly attempt preserved (cap30 failure): aligning TCP directly to receiver led to visible lateral offset and SETTLE_MISS at commanded z .028 (actual .0308). Fresh depth cap29 showed effective ring +.009 x from TCP; lifted safely .065, closed failure, started retry, corrected x then seated successfully. No hardware fault. Initial high downward observation IK candidates rejected without motion; successful observation is (.23,0,.15). Ordinary joint tolerance can allow 5–8 mm TCP error; final approaches use Cartesian. Repeated deltas accumulated x drift; use absolute measured targets. Ring moved/effective pose changed after release; remeasure. No force measurements.
What a fresh agent actually reads: a sanitized excerpt of the saved experience, in the agent's own words. Saved after cycle 1 of the ring study, from the corrections field of the experience store. Verbatim: the field carries no path, no name and no identifier, so nothing was cut.
Where the time goes
The timelines of all 62 main result trials are reconstructed by matching robot session records with timestamped agent traces, covering 7,756 paired tool calls. The means below include successful and unsuccessful trials and cover the full recorded execution, including final reporting.
Where the time goes Mean minutes per trial.
Task
Trials
Total
Visual loop
Remaining response
Local tools
Action service
Check
Report
Four pair assembly
10
48.5
25.5
9.9
0.3
11.6
0.5
0.8
Pyramid
10
22.3
8.7
3.4
0.1
8.8
0.6
0.7
Two towers
10
21.4
8.5
4.7
0.1
7.1
0.2
0.9
Six block tower
10
29.8
14.1
5.0
0.1
9.3
0.5
0.8
Dice flipping
10
38.6
16.2
7.7
0.2
12.9
0.8
0.8
Targeted throwing, potato
2
23.7
7.2
12.7
0.7
1.9
0.2
0.9
Towel folding, sequential
5
52.7
25.4
10.0
0.5
14.7
0.3
1.9
Towel folding, simultaneous
5
23.2
12.2
1.6
0.1
7.1
0.3
2.0
Scroll the table sideways for the rest of the columns.
Mean minutes per trial.
Table view
Mean minutes per trial by category, over all evaluated trials.
Configuration
Trials
Visual loop
Remaining response
Local tools
Action service
Check
Report
Total
Four pair assembly
10
25.5
9.9
0.3
11.6
0.5
0.8
48.5
Block construction · pyramid
10
8.7
3.4
0.1
8.8
0.6
0.7
22.3
Block construction · two towers
10
8.5
4.7
0.1
7.1
0.2
0.9
21.4
Block construction · six block tower
10
14.1
5.0
0.1
9.3
0.5
0.8
29.8
Dice flipping
10
16.2
7.7
0.2
12.9
0.8
0.8
38.6
Targeted throwing · potato
2
7.2
12.7
0.7
1.9
0.2
0.9
23.7
Towel folding · sequential
5
25.4
10.0
0.5
14.7
0.3
1.9
52.7
Towel folding · simultaneous
5
12.2
1.6
0.1
7.1
0.3
2.0
23.2
All values in minutes. Rows sum to the total.
Source: the execution time allocation records, rebuilt for all 62 main result trials.
Mean minutes per trial for each configuration, stacked by category.
What each column counts
Visual loop: acquiring images, geometric queries, visual tool runs, and the model response that comes straight after an image. That response often carries perception, planning and code generation together, so the split between this column and the next depends on where the images fall.
Remaining response: the same agent's model and client time outside the other categories, including reasoning, generation, queueing and transport.
Local tools: shell commands, reading the interface documentation, listing files, local Python, and process polling.
Action service: controller planning, movement, settling, and requests that were refused. It includes stretches when the robot is standing still.
Check: from the last action to the reporting boundary. Report: writing the final result.
Overlapping intervals are counted once. For towel folding the two arms' request logs are merged and concurrent intervals counted once.
The visual loop accounts for 25.5 of the 48.5 mean minutes in four pair assembly and 14.1 minutes in the six block tower. Throwing uses 1.9 minutes of action service and 12.7 minutes of remaining response, reflecting the analysis performed after the object lands.
Where it fails, and what the agent said about it
Five of the 62 main result trials were unsuccessful, as were 11 of the 40 model comparison trials.
The pin that was never found
Three pairs went together. The hollow tube was knocked over and stood back up twice. On the final insertion the gripper was already partly open when the arm moved, the long pin left the jaws, and the agent could not locate it again.
A four pair assembly failure
Three pairs are assembled; on the last insertion the gripper opens too early, the long pin leaves the jaws and is never found.
Full caption
An unsuccessful four pair assembly example. Three pairs are assembled; on the last insertion the gripper opens too early, the long pin leaves the jaws and is never found. 83:30-86:30 at 8x, then 86:30-89:00 at 8x. The agent stops and reports the failure itself.
Task
Four pair assembly
Outcome
failure
Model
GPT-6 Astra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
89.2 min real, 5.5 min shown
Clip
41 s
Verdict
Agent report: failure. Human review: failure. They agree.
Six cubes standing, then not
The fourth cube went down out of alignment, and the agent regrasped and corrected it while the stack was still intact. The sixth level was beyond a straight-down reach, so it built the top two cubes on the table and carried them up as a pair. All six stood after release. The tower came down during the withdrawal that followed.
All six stand, then fall
All six cubes stand, then the tower collapses as the arm withdraws and one cube lands out of reach. The collapse plays at 8x.
Full caption
An unsuccessful six block tower example. 35.6 min. The agent builds the top two cubes as a pair on the table and transfers them; all six stand, then the tower collapses as the arm withdraws and one cube lands out of reach. Side view: 27:20-31:24 at 8x, then the collapse at 8x; 41 s and 63 s of the run are skipped between the three pieces.
Task
Six block tower
Outcome
failure
Model
GPT-6 Astra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
35.6 min real, 4.8 min shown
Clip
38 s
Verdict
Agent report: failure. Human review: failure. They agree.
One corner that stayed unfolded
The agent carried both short ends inward at the same time and staged the placements, which is what the demonstration asks for.
A corner left unfolded
The concurrent inward carry works, but earlier corner slips and a curled long edge leave a loose triangular corner.
Full caption
An unsuccessful simultaneous towel folding example. 27.6 min. The concurrent inward carry works, but earlier corner slips and a curled long edge leave a loose triangular corner. 8x. The agent called it an approximate match; the human review counted it a failure.
Task
Towel folding, simultaneous
Outcome
failure
Model
GPT-6 Astra (high)
Harness
Codex CLI
Arm
bimanual
Speed
8x
Real time
27.6 min real, 6.2 min shown
Clip
46 s
Verdict
Agent report: an approximate or partial match. Human review: failure. This is a difference of degree rather than a contradiction, and the trial is counted as a failure.
A smaller model that knows when to stop
The circular pair went together. On the second pair the hexagonal ring slipped and knocked its post out of reach near the base of the arm. The agent tried to recover, could not, and reported partial success rather than success.
GPT-5.6 Terra, one of four failures
The round pair goes together; then the hex ring slips and pushes its post into the zone the arm cannot reach.
Full caption
Two pair assembly example (1/5 for this model). GPT-5.6 Terra at high effort in Codex CLI, single arm. The round pair goes together; then the hex ring slips and pushes its post into the zone the arm cannot reach. 22.1 min to its own report, which calls it partial success, not task success. 8x.
Task
Two pair assembly
Outcome
failure
Model
GPT-5.6 Terra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
22.1 min real, 5.7 min shown
Clip
46 s
Verdict
Agent report: partial success, not task success. Human review: failure.
Across all 112 counted slots
1trial where the agent stated an outcome the review contradicted
10reports the fixed parser could not classify, which is a limit of the parser and not a disagreement
Every trial starts a fresh agent session with the interface documentation, the task instructions, the reference media and an empty workspace. Solutions and logs from earlier trials are not supplied.
Default configuration
GPT-6 Astra at high thinking effort, inside Codex CLI, for every main result configuration.
Fixed start
Each trial begins from a documented physical reset and an initial scene drawn from a predefined set of placements. Model and experience comparisons use matched scenes with counterbalanced order.
Budgets
Each task has a fixed wall-clock budget and a fixed budget for observation and action requests, shared across configurations. Recovery inside those budgets is allowed.
What ends a trial
The agent reporting completion, the budget running out, or a reset or human intervention. A reset or intervention ends the autonomous trial. Physical failure, budget exhaustion, intervention and infrastructure interruption are recorded separately.
What counts as success
The physical outcome and the execution video, judged against criteria fixed before evaluation. The required assembly, support, orientation and fold relationships have to survive gripper release and arm withdrawal.
Completion time
Measured from the delivery of the task and its references to the final physical verification. It includes reference processing, model calls, local computation, sensing, motion and recovery. Time is reported over successful trials only, so a configuration with no successful trial has no time.
Tokens and cost
Tokens are the sum of input and output usage across every model request in the trial, including reasoning and retries, reported in millions. Repeatedly submitted context is counted on each request, and cached input counts once inside that request's input total. Cost uses standard API rates with cached input priced as cached.
† Throwing time includes the trajectory analysis and verification that follow the throw. The throw itself completes an average of 13.9 minutes after the task is delivered, so roughly 40% of the reported 22.8 minutes comes after the object has landed.
What the capability column means
Long horizon reasoning (Long horizon)
the trial runs many steps and earlier progress has to survive later ones
Action diversity (Action variety)
the motion itself has to change, through grasp-dependent reorientation or a timed release
Object diversity (Object variety)
the object deforms, so the shape grasped is not the shape planned for
Reading the table
Success is the number of completely successful trials over the number of evaluated trials. Partial completion does not count.
Time, tokens and cost summarize successful trials within each configuration.
Taken together, the eight rows hold 57 successes in 62 trials.
Table view
Main results, per trial: the data plotted in the figure above.
Means, minima and maxima are over successful trials only.
Source: the per-trial completion time, token usage and cost table for all 62 main result trials.
One mark per trial for each configuration, successes filled and failures hollow, with the trial index beneath. The strips carry the same 62 trials as the table.
How AGP works
The loop
Observe
The agent asks the robot for an overhead image, a wrist image with aligned depth, or the current joint and end effector state. Captures arrive with calibration and timestamps and are saved into the agent's own workspace.
Write a program
The agent writes Python during the run to read those files: measure a part from depth, fit a surface, intersect a camera ray with a plane, triangulate a point seen from two views, or compose several robot calls into one step.
Command motion
The agent submits a target: either joint angles or an end effector position and orientation, plus a separate gripper opening. The robot runtime turns the target into joint trajectories under workspace and motion limits, using Mink for inverse kinematics.
Check
Each motion returns the measured state and the arrival error. The agent takes a new image when it wants one, and decides for itself whether the physical result matches its estimate.
Revise
A failed insertion can prompt a closer look, a new alignment estimate, or a different approach direction. The agent keeps going until it reports completion or ends the attempt inside the trial budget.
What the agent gets
A task specification: a human demonstration video, a goal photograph, a language instruction, or a combination, plus reference materials and completion criteria.
A documented robot interface with seven commands: state, status, help, frames, move_ee, move_joints and gripper.
A writable workspace where it can save images, scripts and notes for the duration of the trial.
Fixed budgets for elapsed time and for observation and action requests.
What the agent does not get
No simulation rehearsal.
No learned motion policy or skill library underneath it. The agent computes its own targets.
No solutions or logs from earlier trials, except in the experience studies, which are recorded separately.
No teleoperation. A reset or a human intervention ends the autonomous trial.
The platform
The robot is an I2RT YAM arm with six revolute joints and a parallel gripper over a tabletop workspace. A wrist-mounted Intel RealSense D405 supplies aligned colour and depth at 640 by 360 pixels, and a fixed overhead Logitech BRIO supplies rectified colour at 1920 by 1080. Bimanual towel folding adds a second arm expressed in the first arm's coordinate frame. Targeted throwing uses a separate runtime with timed motion programs, so that the swing and the gripper release can be scheduled together.
The tabletop station: two I2RT YAM arms on a shared rail, each with a parallel gripper and a wrist camera, under one fixed overhead camera.
What the agent sees from overhead: the first 1920 by 1080 capture of the two pair assembly trial replayed below, second by second.What the agent sees from the wrist: the 640 by 360 colour frame returned by the same request, with depth aligned to it. The agent asks for both views at once and reads them together.
Agent self-report and human review
Most trials end with a report written by the agent, and every trial has a separate review of the final state and the execution video by a person. Five of the 112 counted runs ended with no report at all, so on those there is nothing on the agent's side to compare. Where both exist they can differ. On this page a disagreement means the agent stated an outcome that the review contradicted: the agent called a trial successful and the review called it a failure, or the other way round. Trials whose report could not be classified automatically are counted apart, under a label of their own, because that is a limit of the parser and not a judgement by the agent.
The pin that was never found
Four pair assembly example. GPT-6 Astra at high thinking effort, Codex CLI, single arm. 89.2 minutes, ended at the timeout.
Three pairs went together. The hollow tube was knocked over and stood back up twice. On the final insertion the gripper was already partly open when the arm moved, the long pin left the jaws, and the agent could not locate it again. It stopped and reported the failure rather than claiming the task was done.
Agent report: failure. Human review: failure. They agree.
Six cubes standing, then not
Six block tower example. GPT-6 Astra high, Codex CLI, single arm. 35.6 minutes.
The fourth cube went down out of alignment, and the agent regrasped and corrected it while the stack was still intact. The sixth level was beyond a straight-down reach, so it built the top two cubes on the table and carried them up as a pair. All six stood after release. The tower came down during the withdrawal that followed. The agent then measured where the cubes had landed, found one grey cube at about 0.70 m in the robot base frame against a 0.65 m radial grasp limit, and ended the attempt as unsuccessful with time and command budget still left.
All six cubes standing after release.The same scene after the arm withdrew.
Agent report: failure. Human review: failure. They agree.
This is why withdrawal counts as part of the task, and it is also where recovery stops: reobservation and regrasping fixed the earlier alignment error, but they could not reach a cube that had left the workspace.
The agent carried both short ends inward at the same time and staged the placements, which is what the demonstration asks for. Repeated corner slips earlier in the run, and a curled long-edge fold, left a loose triangular corner that is plain in the final overhead image.
An example, counted as a failure.An example, counted as a success, for comparison.
Agent report: an approximate or partial match. Human review: failure. This is a difference of degree rather than a contradiction, and the trial is counted as a failure.
A smaller model that knows when to stop
Two pair assembly, GPT-5.6 Terra at high thinking effort, Codex CLI, single arm, an example. 22.1 minutes.
The circular pair went together. On the second pair the hexagonal ring slipped and knocked its post out of reach near the base of the arm. The agent tried to recover, could not, and reported partial success rather than success.
Agent report: partial success, not task success. Human review: failure.
Across all 112 counted slots
1trial where the agent stated an outcome the review contradicted
10reports the fixed parser could not classify, which is a limit of the parser and not a disagreement
5runs that ended without a final report
In the agent's own words
The main fold order was attempted correctly, but the final flaps are skewed, unevenly overlapped, and not reasonably aligned. Success: no.
Towel folding, sequential. The review counted this example a success; the agent counted it a failure.
The page states two counts, how many trials carry an explicit disagreement between the agent's stated outcome and the human review, and how many reports the automatic parser could not classify. The second group is much the larger of the two, and it is not evidence of disagreement.
Programs the agent wrote while the task was running
Across 65 recorded executions, 18 sessions retained 37 Python working files. These comprise 29 files for perception and geometry, 7 for composing robot calls, and 1 for drawing reference points into an image. The logs hold a further 135 shell events with inline Python across 48 sessions. One case can be followed all the way from a camera image to a physical motion.
Four pair assembly
Observe the parts
The agent compared wrist camera views and depth measurements. Bright white surfaces sometimes gave unreliable depth readings.
Compute a position from two views
An executed Python command converted selected pixels in captures 1 and 16 into viewing rays using their camera calibration. It computed the closest points on the two rays.
Revise the alignment
The hexagonal ring caught on the peg. Contact stops led to lifts and revised alignment using the visible ring hole and peg. The ring passed over the peg and settled near its foot after release.
Check the final state
The trial record reports success. The final photographs show all four assemblies after the gripper released the parts and the arm withdrew.
In the agent's own words
Executed 3D estimation code from the assembly run, with comments and line breaks added. The ray function uses the calibration for each capture to return the camera position and viewing direction.
# Camera positions p, q and viewing directions a, b
p,a=ray(1,450,175)
q,b=ray(16,293,189)
dot=lambda a,b:sum(x*y for x,y in zip(a,b))
# Geometry between the two viewing rays
d=[q[i]-p[i] for i in range(3)]
aa=dot(a,a); bb=dot(b,b); ab=dot(a,b)
ad=dot(a,d); bd=dot(b,d)
# Solve for the closest point on each ray
t=(ad*bb-ab*bd)/(aa*bb-ab*ab)
s=(ab*ad-aa*bd)/(aa*bb-ab*ab)
# Print the two estimated 3D positions as [x, y, z]
print([p[i]+t*a[i] for i in range(3)],
[q[i]+s*b[i] for i in range(3)])
What the programs contribute to performance is not established. Settling that would need comparisons that hold the interface capabilities and the inherited experience fixed.
Limitations
Study limitations
Time and money. Mean completion time per successful trial runs from 20.6 to 50.8 minutes and mean inference cost from USD 7.46 to USD 24.14. Both are real barriers to deployment.
Coordinated deformable manipulation is the weakest case. Simultaneous towel folding is the lowest configuration at 3 of 5, and sequential folding costs 50.8 minutes and USD 24.14 on average.
Cost and time grow with task complexity, as the six block tower shows against the pyramid and the two towers, and a large share of the overhead is reviewing and verifying outcomes. About 40% of the throwing time comes after the throw.
Image-conditioned decision cycles take a large share of elapsed time, with perception, planning and code generation coupled inside single model responses. There is substantial room to cut observation and decision overhead and the cost of writing and transferring experience.
Model choice matters. On the same two pair assembly task, GPT-5.6 Terra reaches 1 of 5, GPT-5.6 Luna 0 of 5 and Claude Fable 5.1 3 of 5.
Cite this work
@article{jia2026agentaspolicy,
title = {Agent as Policy for Robotic Manipulation},
author = {Jia, Mengzhao and Lin, Yang and Zhang, Xixin and Zhang, Zhihan and Liu, Xiaobai and Jiang, Meng},
journal = {arXiv preprint arXiv:2609.12541},
year = {2026}
}
Credits
Assembly part geometry
Adapted from the AutoMate dataset. Tang et al., AutoMate: Specialist and Generalist Assembly Policies over Diverse Geometries, 2024, arXiv:2407.08028. The parts on this page are 3D prints of those meshes, scaled with their aspect ratios preserved.
Robot arm
I2RT YAM arm and the I2RT robot software.
Inverse kinematics
Mink, by Kevin Zakka: Python inverse kinematics based on MuJoCo.