Skip to the content

Every number

The tables, studies and protocol that the main page states in one line each.

Two arms share one coordinate frame, and the agent decides how to use them. It watches an overhead demonstration video and reproduces the folds: first the near long edge onto the far long edge, to make a doubled strip, and then both short ends inward. Two configurations were evaluated with five trials each. In sequential folding the demonstration folds the right short end first and then the left. In simultaneous folding the demonstration brings both short ends inward together. The agent chooses the grasp points, which arm does what and when, and may stage a placement to keep the arms clear of each other. The towel changes shape during grasping and folding.

Success criterionSuccess requires both inward folds with reasonably aligned edges, verified in an overhead image after both grippers open and both arms withdraw. Small wrinkles and creases are allowed, and the final position on the table is not restricted.

Sequential 5 of 5, at a mean 50.8 minutes and USD 24.14, among the highest in the suite. Simultaneous 3 of 5, at a mean 21.4 minutes and USD 7.46, the lowest success rate of the eight configurations.

Overhead photograph of the folded towel.
Overhead photograph of the folded towel.

The agent watches one human demonstration video and then assembles eight 3D-printed parts into four pairs in the demonstrated order: a ridged sleeve over a broad round base, a hexagonal nut over a short post, a smooth sleeve over a stepped cylinder, and a long pin into a hollow tube. Nothing tells the agent which part mates with which; it has to read the relationships and the insertion motions out of the video and out of the scene in front of it. The parts are printed from AutoMate meshes, with a nominal radial clearance as small as 0.74 mm, requiring AGP to align and insert the parts precisely.

Success criterionSuccess requires all four pairs to be assembled at the end, with pairs that were finished earlier still intact, and everything stable after the gripper opens and the arm withdraws.

8 of 10 trials succeeded. Successful trials took a mean of 37.2 minutes, ranging from 16.0 to 59.1, with mean usage of 13.02 million tokens and mean inference cost of USD 16.62.

The four pairs

The four pairs
PairHole diameter (mm)Shaft diameter (mm)Radial clearance (mm)
Ridged sleeve and broad round base37.4135.860.78
Hexagonal nut and short post28.1825.801.19
Smooth sleeve and stepped cylinder42.8637.202.83
Long pin and hollow tube18.0116.540.74

Scroll the table sideways for the rest of the columns.

Nominal print dimensions as recorded by the authors, in demonstration order. Radial clearance is half the diametral clearance, computed from the two columns beside it and rounded once, and gives the transverse offset available when the two axes are parallel. What the parts actually tolerate also depends on printing accuracy and on angular alignment.

Render of two unassembled parts side by side on an empty ground: a blue ring with a ridged inner wall on the left, a grey round base with a raised hub on the right.
Ridged sleeve and broad round base. Radial clearance 0.78 mm.
Render of two unassembled parts side by side on an empty ground: a blue hexagonal nut with a round hole on the left, a grey cylindrical post standing on a hexagonal flange on the right.
Hexagonal nut and short post. Radial clearance 1.19 mm.
Render of two unassembled parts side by side on an empty ground: a blue open sleeve with a smooth bore on the left, a grey stepped cylinder on the right.
Smooth sleeve and stepped cylinder. Radial clearance 2.83 mm.
Render of two unassembled parts side by side on an empty ground: a blue pin with a flat lobed head on the left, a grey hollow tube standing on a flange on the right.
Long pin and hollow tube. Radial clearance 0.74 mm, the tightest of the four.
Overhead photograph of four assembled part pairs standing upright on a dark table.
An assembly example after 47.2 minutes. All four pairs are assembled.

The agent gets six cubes on a table and a single photograph of the structure it should build. It has to work out which cube is which, how they support each other, and an order of placements that keeps the partial structure standing. Three goal structures were evaluated with ten trials each: a pyramid with three blue cubes in the base, two grey cubes bridging the base seams and one grey cube at the peak; two towers whose colour orders run blue, grey, blue and grey, blue, grey from the bottom; and a six block tower alternating grey and blue above a grey base that stays where it is.

Success criterionSuccess requires the complete target structure, still standing after the gripper opens and the arm withdraws.

Pyramid succeeded in 10 of 10 trials at a mean of 21.6 minutes, two towers in 10 of 10 at 20.6 minutes, and the six block tower in 9 of 10 at 28.2 minutes. Together, 29 of 30 block construction trials succeeded.

Photograph of six cubes built into a pyramid, three blue in the base, two grey above, one grey on top.
The goal image for the pyramid.
Overhead photograph of a completed six-cube pyramid.
Pyramid example after the arm withdrew.
Photograph of two three-cube towers with different colour orders.
The goal image for the two towers.
Overhead photograph of two completed three-cube towers.
Two towers example in its final state.
Photograph of a six-cube tower alternating grey and blue.
The goal image for the six block tower.
Overhead photograph of a completed six-cube tower.
Six block tower example in its final state.

Six dice lie on the table with different faces upward. The agent receives a language instruction to turn them so that the upward face matches the goal, together with a goal photograph that fixes which face that is. The initial and target faces are balanced across predefined conditions. Nothing restricts where the dice end up or how they are rotated in the table plane, so the agent has to choose a grasp for each die that makes the required rotation reachable, and then release it so that it settles on the intended face.

Success criterionSuccess requires all six dice to show the requested number after placement. The task prompt leaves the dice positions and in-plane rotations unspecified.

10 of 10 trials succeeded. Successful trials took a mean of 37.9 minutes, ranging from 22.6 to 56.0, with mean usage of 17.08 million tokens and mean inference cost of USD 21.07.

Photograph of six dice on a table, all showing six pips on the upward face.
The goal photograph supplied with the instruction: three black and three light blue dice, six pips upward.
Overhead photograph of six dice all showing six pips upward.
Dice flipping example after the arm withdrew.

The agent receives a text instruction and a goal photograph specifying the requested face.

A language instruction asks the agent to pick up a potato-shaped object and throw it into a stationary bowl. The required motion is a circular swing driven by one joint, with the other five joint targets held fixed and the gripper opening during the swing, so the outcome depends on the relationship between arm motion and release timing. This configuration uses a separate runtime with timed motion programs: the agent specifies joint offsets at time knots and a schedule for the gripper, and the executor interpolates the commands on a 50 Hz grid.

Success criterionSuccess requires the object to separate from the moving hand, travel through the air, and come to rest inside the target region, which here is the bowl.

Two recorded sessions, both successful, at a mean 22.8 minutes, from 22.2 to 23.3, at an average of 6.26 million tokens and USD 8.94. The throw itself completes an average of 13.9 minutes after the task is delivered, so about 40% of that time is the analysis and verification that follow it.

The two sessions were selected rather than drawn from a planned batch. They document successful throws but do not establish a success rate. In both sessions, the object was released during the swing and came to rest in the bowl. Occlusion leaves the rising part of the flight and the apex unresolved.

A four by two grid of eight wrist-camera frames, read left to right then top to bottom: a pale potato held in the dark gripper jaws under a teal fitting, then the jaws opening and the potato falling out of the frame. Each frame carries a red label giving its frame number and a time from 0.905 to 1.138 seconds.
An image the agent made during a throwing experiment, showing eight wrist frames around the release labelled with the time since the motion program was dispatched.
Wrist-camera view of a bowl with an ellipse drawn along its rim.
Another agent-made image from the same trial: the wrist view with the fitted bowl rim drawn on top, used to place the target before the swing.
Photograph of a potato-shaped object resting inside a bowl on a table.
A throwing example after 22.2 minutes with a single arm. The object is at rest inside the bowl.

Which agent, which model, how much thinking

Eight configurations, five trials each. Three of them do badly, and they stay in the table.

The comparison uses a smaller assembly task: four parts forming two pairs, a hexagonal ring over a short post and a circular sleeve over a stepped cylinder, learned from one human demonstration. Every configuration receives the same task references, the same robot interface and the same time and request budgets. Success requires both pairs to remain assembled after release and arm withdrawal.

Agent, model and thinking effort on two pair assembly. GPT-6 Astra appears at three thinking efforts; the other models use their high setting.
ModelSuccess (n/N)Time (min)Tokens (M)Cost (USD)
Codex CLI
GPT-6 Astra, low effort5/5GPT-6 Astra, low effort example: successGPT-6 Astra, low effort example: successGPT-6 Astra, low effort example: successGPT-6 Astra, low effort example: successGPT-6 Astra, low effort example: success9.9[6.0, 19.0]3.69[1.44, 7.72]4.79[2.15, 9.56]
GPT-6 Astra, medium effort5/5GPT-6 Astra, medium effort example: successGPT-6 Astra, medium effort example: successGPT-6 Astra, medium effort example: successGPT-6 Astra, medium effort example: successGPT-6 Astra, medium effort example: success9.2[6.3, 15.4]3.02[1.84, 5.80]4.09[2.76, 7.48]
GPT-6 Astra, high effort5/5GPT-6 Astra, high effort example: successGPT-6 Astra, high effort example: successGPT-6 Astra, high effort example: successGPT-6 Astra, high effort example: successGPT-6 Astra, high effort example: success9.2[7.9, 10.6]3.30[2.58, 4.02]4.47[3.72, 5.19]
GPT-5.6 Sol5/5GPT-5.6 Sol example: successGPT-5.6 Sol example: successGPT-5.6 Sol example: successGPT-5.6 Sol example: successGPT-5.6 Sol example: success14.2[9.1, 19.3]7.13[4.12, 9.72]3.94[2.40, 5.30]
GPT-5.6 Terra1/5GPT-5.6 Terra example: failureGPT-5.6 Terra example: failureGPT-5.6 Terra example: failureGPT-5.6 Terra example: successGPT-5.6 Terra example: failure19.06.281.98
GPT-5.6 Luna0/5GPT-5.6 Luna example: failureGPT-5.6 Luna example: failureGPT-5.6 Luna example: failureGPT-5.6 Luna example: failureGPT-5.6 Luna example: failureno successful trialno successful trialno successful trial
Claude Code
Claude Opus 55/5Claude Opus 5 example: successClaude Opus 5 example: successClaude Opus 5 example: successClaude Opus 5 example: successClaude Opus 5 example: success22.2[14.4, 31.9]12.40[9.05, 16.73]9.75[6.94, 13.03]
Claude Fable 5.13/5Claude Fable 5.1 example: successClaude Fable 5.1 example: successClaude Fable 5.1 example: successClaude Fable 5.1 example: failureClaude Fable 5.1 example: failure27.7[18.3, 39.8]9.27[6.87, 11.10]11.30[8.33, 14.36]

Scroll the table sideways for the rest of the columns.

Reading the table

  • GPT-6 Astra succeeds in every trial at all three thinking efforts, in about nine to ten minutes and USD 4.09 to 4.79 per success, so on this task more thinking effort buys little.
  • GPT-5.6 Sol also succeeds in every trial, more slowly and more cheaply than Astra.
  • Within Claude Code, Fable 5.1 succeeds less often than Opus 5 and uses fewer tokens on its successful trials, yet takes more time and costs more per success.
  • Comparisons across Codex CLI and Claude Code are comparisons of whole systems, not of models alone.
  • The single Terra figure is one trial, and the time, token and cost cells for that row have no spread because there is only one successful trial to average.
Two pair assembly across agents, models and thinking effortStrip plot with an outcome column. Each row is one agent and model configuration evaluated on the same two pair assembly task, five trials each. The outcome column shows every evaluated trial: a filled mark is a success, a hollow crossed mark is a failure. The three scales to the right plot every trial that has a complete resource record, for completion time in minutes, tokens in millions and inference cost in US dollars. The vertical rule is the mean over successful trials. The same numbers are given in the table below the figure.Two pair assembly across agents, models and thinking effortThe same task, the same references and the same budgets in every row; five trials each.Outcome shows all five trials. Filled = success, hollow and crossed = failure.Outcome5 trialsCodex CLIClaude CodeGPT-6 Astra · low effortGPT-6 Astra · medium effortGPT-6 Astra · high effortGPT-5.6 SolGPT-5.6 TerraGPT-5.6 LunaClaude Opus 5Claude Fable 5.1Elapsed timeminutes010203040no trial with a resource recordTokensmillions, per trial061218no trial with a resource recordInference costUS dollars051015no trial with a resource recordsuccessful trialfailed trialmean over successful trialsThe mean and the bracketed range are over successful trials only, so a failed trial's mark can sit outside that range.7 failed trials have no usable time, token or cost record and appear in the outcome column only: all five GPT-5.6 Luna trials and two Claude Fable 5.1 trials.Those two success counts, and the GPT-5.6 Terra count, rest in part on the authors' own record of trials whose session files are incomplete.Source: the pinned per-trial two pair assembly records used for model comparison.
Table view
Model comparison, per trial: the data plotted in the figure above.
ConfigurationAgentSuccessTime (min), mean [min, max]Tokens (M), mean [min, max]Cost (USD), mean [min, max]Per-trial completion time (min)
GPT-6 Astra · low effortCodex CLI5/59.9 [6.0, 19.0]3.69 [1.44, 7.72]4.79 [2.15, 9.56]6.0, 6.3, 7.4, 10.9, 19.0
GPT-6 Astra · medium effortCodex CLI5/59.2 [6.3, 15.4]3.02 [1.84, 5.80]4.09 [2.76, 7.48]6.3, 7.6, 7.6, 8.8, 15.4
GPT-6 Astra · high effortCodex CLI5/59.2 [7.9, 10.6]3.30 [2.58, 4.02]4.47 [3.72, 5.19]7.9, 8.9, 9.0, 9.7, 10.6
GPT-5.6 SolCodex CLI5/514.2 [9.1, 19.3]7.13 [4.12, 9.72]3.94 [2.40, 5.30]9.1, 12.3, 15.2, 15.5, 19.3
GPT-5.6 TerraCodex CLI1/519.0 [19.0, 19.0]6.28 [6.28, 6.28]1.98 [1.98, 1.98]14.4 (failed), 19.0, 21.4 (failed), 22.1 (failed), 30.4 (failed)
GPT-5.6 LunaCodex CLI0/5
Claude Opus 5Claude Code5/522.2 [14.4, 31.9]12.40 [9.05, 16.72]9.75 [6.94, 13.03]14.4, 19.2, 20.0, 25.5, 31.9
Claude Fable 5.1Claude Code3/527.7 [18.3, 39.8]9.27 [6.87, 11.10]11.30 [8.33, 14.36]18.3, 25.2, 39.8

Means, minima and maxima are over successful trials only. GPT-5.6 Luna has no successful trial, and seven failed trials have no usable resource record.

Source: the pinned per-trial two pair assembly records used for model comparison.

One mark per trial for each of the eight configurations, successes filled and failures hollow. Each trial uses a single arm.

GPT-6 Astra, high effort

Two pair assembly example (5/5). GPT-6 Astra at high effort in Codex CLI, single arm. 7.9 min real at 8x. The other arm runs a separate trial outside this crop.

Task
Two pair assembly
Outcome
success
Model
GPT-6 Astra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
7.9 min real
Clip
58 s

GPT-5.6 Sol

Two pair assembly example (5/5). GPT-5.6 Sol at high effort in Codex CLI, single arm. 9.1 min real at 8x. Nothing else is running: four parts, one arm.

Task
Two pair assembly
Outcome
success
Model
GPT-5.6 Sol (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
9.1 min real
Clip
64 s

Claude Opus 5

Two pair assembly example (5/5). Claude Opus 5 at high effort in Claude Code, single arm. 14.4 min real at 8x. No failed commands; its report notes that the steep overhead view made the upright posts look like they were lying down, and that it checked with wrist depth.

Task
Two pair assembly
Outcome
success
Model
Claude Opus 5 (high)
Harness
Claude Code
Arm
single
Speed
8x
Real time
14.4 min real, 12.8 min shown
Clip
96 s

The other arm on the rail sometimes runs a separate trial at the same time. Where that is visible in a clip, the caption says so and the frame is cropped to the arm being discussed.

What these numbers rest on

Time, token and cost cells are computed from the per-trial usage records. The Success column is different. The resource rows can be recomputed from data files; the denominators cannot all be. All five GPT-5.6 Luna sessions and all five Claude Fable 5.1 sessions are now on record, but three of the Luna sessions ended without a final report, through a timeout or an abort, and one Fable session has no report either. Those outcomes are recorded by the authors.

Reusing experience

Every trial so far started from nothing. These three studies let the agent keep what it learned. Experience is ordinary files: measurements, procedures, corrections and working scripts, written at the end of one execution and read at the start of the next. Each execution still begins with a fresh agent context, so the only thing that carries over is what was written down.

Two comparisons on this page are two separate clips, not one composed video: ring cycle 1 beside ring cycle 4, and the same model with and without the transferred store. Each pair starts at the same task second and runs at the same speed, so the shorter run simply finishes first. Play them together; each still states its own trial and real elapsed time.

One arm takes a ridged ring off its base and puts it back, five times over, in a single continuous session with GPT-6 Astra at high thinking effort. Each cycle starts a new agent process that reads the previous checkpoint and updates the object poses from what it can see now. The first cycle has nothing to read: its first insertion stalls, the agent lifts the part safely, measures an offset of about 9 mm between the centre of the grasped ring and the tool, and corrects for it. Later cycles reuse that corrected alignment procedure and check it against current observations.

Ring, cycle 1: 870 s

Cycle 1 of 5, with nothing in the experience store. The first insertion stalls; the agent measures a 9 mm grasp offset and retries.

Task
Ring disassembly and reassembly, cycle 1
Outcome
success after one retry
Model
GPT-6 Astra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
14.5 min real
Clip
109 s

Ring, cycle 4: 372 s

The fourth cycle of the same continuous session, run by a fresh agent that reads what the earlier cycles wrote down. No retries.

Task
Ring disassembly and reassembly, cycle 4
Outcome
success
Model
GPT-6 Astra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
6.2 min real
Clip
47 s

Seconds per cycle

Seconds per cycle
CycleTotal runtimeReasoning and programmingTool executionExperience writingRetriesCalls
1989.7702.9167.4119.4137 / 37
2522.9282.1126.7114.1032 / 33
3489.9245.0127.3117.6032 / 33
4484.4247.7124.6112.2032 / 33
5461.7268.3125.068.5032 / 33

Scroll the table sideways for the rest of the columns.

Total runtime falls from 989.7 to 461.7 seconds. The largest reduction follows the first cycle, after the agent establishes an alignment procedure and recovers from a failed insertion. Most of the savings come from reasoning and programming.

Timing definitions

Task execution includes reasoning and programming and tool execution. Tool execution covers recorded time running programs, processing images and operating the robot. Overlapping activities count once, and tool time takes precedence when tools run during a model response. The remaining task time estimates reasoning and programming, including model response and tool startup delays. Experience writing covers the subsequent saving of updated experience. Components sum to total runtime up to rounding. Calls lists motion and gripper requests followed by image capture requests.

These are five dependent repetitions in one session, with no matched control in which every cycle starts from empty experience. They cannot separate the effect of the saved experience from anything else that changes between cycles.

Reusing experience: ring disassembly and reassemblyFive stacked columns show reasoning and programming, tool execution and experience writing. Total runtime falls from 989.7 to 461.7 seconds.Reusing experience: ring disassembly and reassemblyFive executions using saved measurements, programs and lessons.The largest drop in reasoning and programming follows the first cycle.02505007501 000seconds989.71522.92489.93484.44461.75cycleReasoning and programmingTool executionExperience writingTask execution combines reasoning and programming with tool execution.Total runtime also includes experience writing. Values follow Appendix Table 7.
Table view
Reusing experience: ring disassembly and reassembly — seconds per execution.
#Reasoning and programmingTool executionExperience writingTotal
1702.9167.4119.4989.7
2282.1126.7114.1522.9
3245.0127.3117.6489.9
4247.7124.6112.2484.4
5268.3125.068.5461.7

All values in seconds. Components sum to total runtime up to rounding.

Source: Appendix Table 7, Execution breakdown over repeated tasks.

Ring task timing over five cycles. Most of the reduction comes from reasoning and programming.

Five executions use GPT-6 Astra at high thinking effort with a single arm. Each starts with a fresh agent context and reads saved experience and programs from earlier executions. Total runtime decreases from 621.4 to 462.4 seconds. Reasoning and programming falls from 351.9 to 186.6 seconds, while tool execution rises slightly from 160.5 to 175.5 seconds.

Successful assembly is assessed by the agent from recorded images. Task execution includes the return to the observation posture, followed by experience writing.

Timing definitions

Task execution includes reasoning and programming and tool execution. Tool execution covers recorded time running programs, processing images and operating the robot. Overlapping activities count once, and tool time takes precedence when tools run during a model response. The remaining task time estimates reasoning and programming, including model response and tool startup delays. Experience writing covers the subsequent saving of updated experience. Components sum to total runtime up to rounding. Calls lists motion and gripper requests followed by image capture requests.

Reusing experience: two pair assemblyFive stacked columns show reasoning and programming, tool execution and experience writing. Total runtime falls from 621.4 to 462.4 seconds.Reusing experience: two pair assemblyFive executions, each a fresh agent context reading saved experience.Reasoning and programming decreases from 351.9 to 186.6 seconds.0200400600seconds621.41568.12507.53491.84462.45executionReasoning and programmingTool executionExperience writingMost of the time savings come from reasoning and programming.Tool execution takes slightly longer. Values follow Appendix Table 7.
Table view
Reusing experience: two pair assembly — seconds per execution.
#Reasoning and programmingTool executionExperience writingTotal
1351.9160.5108.9621.4
2281.7186.3100.1568.1
3237.1173.896.6507.5
4221.0174.196.6491.8
5186.6175.5100.3462.4

All values in seconds. Components sum to total runtime up to rounding.

Source: Appendix Table 7, Execution breakdown over repeated tasks.

Two pair assembly timing over five executions. Reasoning and programming decreases while tool execution takes slightly longer.

GPT-6 Astra ran the two pair assembly task five times and saved geometry, procedures, corrections and scripts. GPT-5.6 Terra, at high thinking effort, then received that collection with read access only, frozen, unchanged between trials. The comparison is Terra with nothing against Terra reading the Astra collection, with the same task references, the same interface and the same budgets on both sides.

Handing experience to a smaller model
ConditionSuccessTime (min)Tokens (M)Cost (USD)
GPT-5.6 Terra, empty experience1/519.06.281.98
GPT-5.6 Terra, reading the Astra collection4/512.5 [10.1, 16.9]5.56 [4.81, 5.87]1.64 [1.56, 1.74]
GPT-6 Astra, the five runs that built the collection5/57.6 [6.5, 9.1]2.50 [2.24, 2.83]3.79 [3.45, 4.39]

Scroll the table sideways for the rest of the columns.

Means over successful trials, with the observed minimum and maximum beneath, the same accounting as the model comparison above. The empty-experience figure is the one successful Terra trial from that table, so it has no spread.

The 5 of 5 in the third row is one of the four Success counts on this page that rest on the authors' own record. All five runs are on file with full time, token and cost accounting, but the operator verdict column of that batch record was left blank, so the count comes from the single decision of 11 September 2026, in which Yang Lin, the author who ran the trials, confirmed all four of these counts, rather than from the record.

No experience: 19.0 min

Two pair assembly, GPT-5.6 Terra at high effort in Codex CLI, single arm, empty experience store: the only success in that row (1/5). 19.0 min real at 8x, with several settle-miss recoveries. Shown from task second 90 at the same speed as the clip beside it.

Task
Two pair assembly, empty experience
Outcome
success
Model
GPT-5.6 Terra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
19.0 min real, 17.7 min shown
Clip
132 s

Reading a frozen experience store: 11.2 min

The same task and the same model, now reading a frozen experience store written by GPT-6 Astra and never written back to: 4/5 successful, 11.2 min real. Same 8x and the same task second 90 start, so the clip simply ends sooner.

Task
Two pair assembly, frozen experience from GPT-6 Astra
Outcome
success
Model
GPT-5.6 Terra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
11.2 min real, 9.5 min shown
Clip
71 s

With the collection in front of it, Terra succeeds in four trials instead of one, finishes faster, and uses fewer tokens. It does not reach the model that wrote the collection: Astra still finished in 7.6 minutes on average. Cost moves the other way, because Terra is the cheaper model per token.

  • Five trials per condition. The conclusion rests on very few successful runs.
  • The collection is frozen. Terra can read it but cannot add to it, so this measures transfer, not further accumulation.
  • Two runs of this study were discarded because of infrastructure problems and are excluded from the reported results.
Handing one model's experience to a smaller modelStrip plot with an outcome column. Three rows: GPT-5.6 Terra with no prior experience, GPT-5.6 Terra reading a frozen GPT-6 Astra experience store, and the GPT-6 Astra run that wrote that store, shown for reference. The outcome column shows all five trials of each condition. The three scales plot completion time in minutes, tokens in millions and inference cost in US dollars, for every trial; failures are hollow and crossed. The vertical rule is the mean over successful trials. The same numbers are given in the table below the figure.Handing one model's experience to a smaller modelTwo pair assembly. The smaller model succeeds in 1 of 5 trials on its own, and in 4 of 5 when it may read a frozen experience store.Outcome5 trialsGPT-5.6 Terrano prior experience · 1/5 successfulGPT-5.6 Terrareading a frozen GPT-6 Astra store · 4/5 successfulGPT-6 Astrathe run that wrote the store · reference · 5/5 successfulElapsed timeminutes08162432Tokensmillions, per trial0369Inference costUS dollars01.534.5successful trialfailed trialmean over successful trialsThe store is written once by the larger model and then frozen: the smaller model may read it but never changes it, and it is identical in every trial.The third row is the larger model's own run, shown only to place the two conditions in context. It is a different model, not a third arm of the comparison.Four successful trials is a small sample, and the two conditions were collected on different days. Two runs discarded for infrastructure faults are excluded.Source: the pinned per-trial records of the two pair assembly experience transfer runs, using the model comparison accounting rules.
Table view
Experience transfer, per trial: the data plotted in the figure above.
ConditionSuccessTime (min), mean [min, max]Tokens (M), mean [min, max]Cost (USD), mean [min, max]Per-trial completion time (min)
GPT-5.6 Terra, empty experience1/519.0 [19.0, 19.0]6.28 [6.28, 6.28]1.98 [1.98, 1.98]14.4 (failed), 19.0, 21.4 (failed), 22.1 (failed), 30.4 (failed)
GPT-5.6 Terra, reading a frozen GPT-6 Astra experience store4/512.5 [10.1, 16.9]5.56 [4.81, 5.87]1.64 [1.56, 1.74]10.1, 11.2, 11.7, 16.9, 19.6 (failed)
GPT-6 Astra, building the experience store5/57.6 [6.5, 9.1]2.50 [2.24, 2.83]3.79 [3.45, 4.39]6.5, 7.0, 7.3, 8.2, 9.1

Means, minima and maxima are over successful trials only. The third row is a different model, shown for reference.

Source: the pinned per-trial records of the two pair assembly experience transfer runs.

Completion time and token use for GPT-5.6 Terra with and without the Astra collection, over successful trials.

In the agent's own words

One failed assembly attempt preserved (cap30 failure): aligning TCP directly to receiver led to visible lateral offset and SETTLE_MISS at commanded z .028 (actual .0308). Fresh depth cap29 showed effective ring +.009 x from TCP; lifted safely .065, closed failure, started retry, corrected x then seated successfully. No hardware fault. Initial high downward observation IK candidates rejected without motion; successful observation is (.23,0,.15). Ordinary joint tolerance can allow 5–8 mm TCP error; final approaches use Cartesian. Repeated deltas accumulated x drift; use absolute measured targets. Ring moved/effective pose changed after release; remeasure. No force measurements.

What a fresh agent actually reads: a sanitized excerpt of the saved experience, in the agent's own words. Saved after cycle 1 of the ring study, from the corrections field of the experience store. Verbatim: the field carries no path, no name and no identifier, so nothing was cut.

Where the time goes

The timelines of all 62 main result trials are reconstructed by matching robot session records with timestamped agent traces, covering 7,756 paired tool calls. The means below include successful and unsuccessful trials and cover the full recorded execution, including final reporting.

Where the time goes Mean minutes per trial.
TaskTrialsTotalVisual loopRemaining responseLocal toolsAction serviceCheckReport
Four pair assembly1048.525.59.90.311.60.50.8
Pyramid1022.38.73.40.18.80.60.7
Two towers1021.48.54.70.17.10.20.9
Six block tower1029.814.15.00.19.30.50.8
Dice flipping1038.616.27.70.212.90.80.8
Targeted throwing, potato223.77.212.70.71.90.20.9
Towel folding, sequential552.725.410.00.514.70.31.9
Towel folding, simultaneous523.212.21.60.17.10.32.0

Scroll the table sideways for the rest of the columns.

Mean minutes per trial.

Where the execution time goesStacked bar chart. One bar per task configuration, showing mean minutes per trial split into six categories: visual loop, remaining model response, local tools, action service, check and report. Averages are over all evaluated trials in the configuration, successes and failures together. The total is printed at the end of each bar. The same numbers are given in the table below the figure.Where the execution time goesMean minutes per trial, averaged over all evaluated trials in the configuration — failures included.Rebuilt from the timestamped agent traces paired with the robot session records.0204060minutesFour pair assemblymean over 10 trials48.5 min25.59.911.6Block construction · pyramidmean over 10 trials22.3 min8.78.8Block construction · two towersmean over 10 trials21.4 min8.54.77.1Block construction · six block towermean over 10 trials29.8 min14.15.09.3Dice flippingmean over 10 trials38.6 min16.27.712.9Targeted throwing · potatomean over 2 trials23.7 min7.212.7Towel folding · sequentialmean over 5 trials52.7 min25.410.014.7Towel folding · simultaneousmean over 5 trials23.2 min12.27.1Visual loopRemaining responseLocal toolsAction serviceCheckReportVisual loop covers image acquisition, geometric queries, visual tools and the model response that immediately follows an image. Action service is the time the robotinterface spends carrying out a motion or gripper request. Inline numbers are minutes; a segment is labelled only where the label fits inside it.Source: the execution time allocation records, rebuilt for all 62 main result trials.
Table view
Mean minutes per trial by category, over all evaluated trials.
ConfigurationTrialsVisual loopRemaining responseLocal toolsAction serviceCheckReportTotal
Four pair assembly1025.59.90.311.60.50.848.5
Block construction · pyramid108.73.40.18.80.60.722.3
Block construction · two towers108.54.70.17.10.20.921.4
Block construction · six block tower1014.15.00.19.30.50.829.8
Dice flipping1016.27.70.212.90.80.838.6
Targeted throwing · potato27.212.70.71.90.20.923.7
Towel folding · sequential525.410.00.514.70.31.952.7
Towel folding · simultaneous512.21.60.17.10.32.023.2

All values in minutes. Rows sum to the total.

Source: the execution time allocation records, rebuilt for all 62 main result trials.

Mean minutes per trial for each configuration, stacked by category.

What each column counts

  • Visual loop: acquiring images, geometric queries, visual tool runs, and the model response that comes straight after an image. That response often carries perception, planning and code generation together, so the split between this column and the next depends on where the images fall.
  • Remaining response: the same agent's model and client time outside the other categories, including reasoning, generation, queueing and transport.
  • Local tools: shell commands, reading the interface documentation, listing files, local Python, and process polling.
  • Action service: controller planning, movement, settling, and requests that were refused. It includes stretches when the robot is standing still.
  • Check: from the last action to the reporting boundary. Report: writing the final result.
  • Overlapping intervals are counted once. For towel folding the two arms' request logs are merged and concurrent intervals counted once.

The visual loop accounts for 25.5 of the 48.5 mean minutes in four pair assembly and 14.1 minutes in the six block tower. Throwing uses 1.9 minutes of action service and 12.7 minutes of remaining response, reflecting the analysis performed after the object lands.

Where it fails, and what the agent said about it

Five of the 62 main result trials were unsuccessful, as were 11 of the 40 model comparison trials.

The pin that was never found

Three pairs went together. The hollow tube was knocked over and stood back up twice. On the final insertion the gripper was already partly open when the arm moved, the long pin left the jaws, and the agent could not locate it again.

A four pair assembly failure

Three pairs are assembled; on the last insertion the gripper opens too early, the long pin leaves the jaws and is never found.

Full caption

An unsuccessful four pair assembly example. Three pairs are assembled; on the last insertion the gripper opens too early, the long pin leaves the jaws and is never found. 83:30-86:30 at 8x, then 86:30-89:00 at 8x. The agent stops and reports the failure itself.

Task
Four pair assembly
Outcome
failure
Model
GPT-6 Astra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
89.2 min real, 5.5 min shown
Clip
41 s
Verdict
Agent report: failure. Human review: failure. They agree.

Six cubes standing, then not

The fourth cube went down out of alignment, and the agent regrasped and corrected it while the stack was still intact. The sixth level was beyond a straight-down reach, so it built the top two cubes on the table and carried them up as a pair. All six stood after release. The tower came down during the withdrawal that followed.

All six stand, then fall

All six cubes stand, then the tower collapses as the arm withdraws and one cube lands out of reach. The collapse plays at 8x.

Full caption

An unsuccessful six block tower example. 35.6 min. The agent builds the top two cubes as a pair on the table and transfers them; all six stand, then the tower collapses as the arm withdraws and one cube lands out of reach. Side view: 27:20-31:24 at 8x, then the collapse at 8x; 41 s and 63 s of the run are skipped between the three pieces.

Task
Six block tower
Outcome
failure
Model
GPT-6 Astra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
35.6 min real, 4.8 min shown
Clip
38 s
Verdict
Agent report: failure. Human review: failure. They agree.

One corner that stayed unfolded

The agent carried both short ends inward at the same time and staged the placements, which is what the demonstration asks for.

A corner left unfolded

The concurrent inward carry works, but earlier corner slips and a curled long edge leave a loose triangular corner.

Full caption

An unsuccessful simultaneous towel folding example. 27.6 min. The concurrent inward carry works, but earlier corner slips and a curled long edge leave a loose triangular corner. 8x. The agent called it an approximate match; the human review counted it a failure.

Task
Towel folding, simultaneous
Outcome
failure
Model
GPT-6 Astra (high)
Harness
Codex CLI
Arm
bimanual
Speed
8x
Real time
27.6 min real, 6.2 min shown
Clip
46 s
Verdict
Agent report: an approximate or partial match. Human review: failure. This is a difference of degree rather than a contradiction, and the trial is counted as a failure.

A smaller model that knows when to stop

The circular pair went together. On the second pair the hexagonal ring slipped and knocked its post out of reach near the base of the arm. The agent tried to recover, could not, and reported partial success rather than success.

GPT-5.6 Terra, one of four failures

The round pair goes together; then the hex ring slips and pushes its post into the zone the arm cannot reach.

Full caption

Two pair assembly example (1/5 for this model). GPT-5.6 Terra at high effort in Codex CLI, single arm. The round pair goes together; then the hex ring slips and pushes its post into the zone the arm cannot reach. 22.1 min to its own report, which calls it partial success, not task success. 8x.

Task
Two pair assembly
Outcome
failure
Model
GPT-5.6 Terra (high)
Harness
Codex CLI
Arm
single
Speed
8x
Real time
22.1 min real, 5.7 min shown
Clip
46 s
Verdict
Agent report: partial success, not task success. Human review: failure.

Across all 112 counted slots

  • 1trial where the agent stated an outcome the review contradicted
  • 10reports the fixed parser could not classify, which is a limit of the parser and not a disagreement
  • 5runs that ended without a final report

How a trial was run and judged

How a trial was run and judged

One agent, one session

Every trial starts a fresh agent session with the interface documentation, the task instructions, the reference media and an empty workspace. Solutions and logs from earlier trials are not supplied.

Default configuration

GPT-6 Astra at high thinking effort, inside Codex CLI, for every main result configuration.

Fixed start

Each trial begins from a documented physical reset and an initial scene drawn from a predefined set of placements. Model and experience comparisons use matched scenes with counterbalanced order.

Budgets

Each task has a fixed wall-clock budget and a fixed budget for observation and action requests, shared across configurations. Recovery inside those budgets is allowed.

What ends a trial

The agent reporting completion, the budget running out, or a reset or human intervention. A reset or intervention ends the autonomous trial. Physical failure, budget exhaustion, intervention and infrastructure interruption are recorded separately.

What counts as success

The physical outcome and the execution video, judged against criteria fixed before evaluation. The required assembly, support, orientation and fold relationships have to survive gripper release and arm withdrawal.

Completion time

Measured from the delivery of the task and its references to the final physical verification. It includes reference processing, model calls, local computation, sensing, motion and recovery. Time is reported over successful trials only, so a configuration with no successful trial has no time.

Tokens and cost

Tokens are the sum of input and output usage across every model request in the trial, including reasoning and retries, reported in millions. Repeatedly submitted context is counted on each request, and cached input counts once inside that request's input total. Cost uses standard API rates with cached input priced as cached.

Throwing time includes the trajectory analysis and verification that follow the throw. The throw itself completes an average of 13.9 minutes after the task is delivered, so roughly 40% of the reported 22.8 minutes comes after the object has landed.

What the capability column means

Long horizon reasoning (Long horizon)
the trial runs many steps and earlier progress has to survive later ones
Action diversity (Action variety)
the motion itself has to change, through grasp-dependent reorientation or a timed release
Object diversity (Object variety)
the object deforms, so the shape grasped is not the shape planned for

Reading the table

  • Success is the number of completely successful trials over the number of evaluated trials. Partial completion does not count.
  • Time, tokens and cost summarize successful trials within each configuration.
  • Taken together, the eight rows hold 57 successes in 62 trials.
Every trial in the main evaluationStrip plot. One mark per evaluated trial for each of the eight task configurations, on three scales: completion time in minutes, total tokens in millions, and inference cost in US dollars. Filled marks are successful trials; hollow crossed marks are failed trials. The vertical rule on each row is the mean over successful trials. The same numbers are given in the table below the figure.Every trial in the main evaluation62 trials across eight task configurations. Filled = success, hollow and crossed = failure.The vertical rule is the mean over successful trials.Four pair assembly8/10 successfulBlock construction · pyramid10/10 successfulBlock construction · two towers10/10 successfulBlock construction · six block tower9/10 successfulDice flipping10/10 successfulTargeted throwing · potato2/2 successfulTowel folding · sequential5/5 successfulTowel folding · simultaneous3/5 successfulElapsed timeminutes0255075100Tokensmillions, per trial0153045Inference costUS dollars02550successful trialfailed trialmean over successful trialsAssembly, blocks and dice have ten trials per configuration, towels five, and throwing two. Means, minima and maxima are computed over successful trials only, so a failed trial'smark can sit outside the bracketed range over successful trials. Cost is the model inference cost at the recorded list rates.Failed-trial times end at the final report boundary or termination. Source: the 62 per-trial timing and usage records.
Table view
Main results, per trial: the data plotted in the figure above.
ConfigurationSuccessTime (min), mean [min, max]Tokens (M), mean [min, max]Cost (USD), mean [min, max]Per-trial completion time (min)
Four pair assembly8/1037.2 [16.0, 59.1]13.02 [3.96, 27.12]16.62 [5.55, 31.50]38.6, 53.0, 89.2 (failed), 47.2, 90.2 (failed), 16.3, 30.3, 16.0, 37.4, 59.1
Block construction · pyramid10/1021.6 [14.9, 32.2]9.24 [6.27, 12.94]11.69 [8.01, 15.28]16.8, 23.2, 26.9, 14.9, 32.2, 20.0, 17.0, 17.9, 24.1, 22.4
Block construction · two towers10/1020.6 [14.8, 29.2]7.94 [4.94, 14.39]9.92 [6.27, 16.85]21.0, 19.4, 14.8, 22.6, 29.2, 24.4, 17.6, 17.8, 21.9, 17.0
Block construction · six block tower9/1028.2 [11.6, 55.0]11.57 [4.50, 19.85]14.93 [6.50, 26.78]46.4, 35.6 (failed), 20.1, 17.3, 55.0, 25.4, 31.0, 29.0, 17.9, 11.6
Dice flipping10/1037.9 [22.6, 56.0]17.08 [10.16, 25.00]21.07 [12.57, 31.20]36.4, 56.0, 47.9, 51.4, 22.6, 37.3, 33.6, 32.1, 26.4, 34.9
Targeted throwing · potato2/222.8 [22.2, 23.3]6.26 [6.25, 6.26]8.94 [8.71, 9.18]22.2, 23.3
Towel folding · sequential5/550.8 [33.8, 63.3]18.12 [10.13, 25.26]24.14 [14.16, 30.90]61.6, 50.7, 44.6, 63.3, 33.8
Towel folding · simultaneous3/521.4 [15.7, 31.2]5.06 [4.27, 5.68]7.46 [6.16, 8.67]14.0 (failed), 15.7, 27.6 (failed), 31.2, 17.4

Means, minima and maxima are over successful trials only.

Source: the per-trial completion time, token usage and cost table for all 62 main result trials.

One mark per trial for each configuration, successes filled and failures hollow, with the trial index beneath. The strips carry the same 62 trials as the table.

How AGP works

The loop

  1. Observe

    The agent asks the robot for an overhead image, a wrist image with aligned depth, or the current joint and end effector state. Captures arrive with calibration and timestamps and are saved into the agent's own workspace.

  2. Write a program

    The agent writes Python during the run to read those files: measure a part from depth, fit a surface, intersect a camera ray with a plane, triangulate a point seen from two views, or compose several robot calls into one step.

  3. Command motion

    The agent submits a target: either joint angles or an end effector position and orientation, plus a separate gripper opening. The robot runtime turns the target into joint trajectories under workspace and motion limits, using Mink for inverse kinematics.

  4. Check

    Each motion returns the measured state and the arrival error. The agent takes a new image when it wants one, and decides for itself whether the physical result matches its estimate.

  5. Revise

    A failed insertion can prompt a closer look, a new alignment estimate, or a different approach direction. The agent keeps going until it reports completion or ends the attempt inside the trial budget.

What the agent gets

  • A task specification: a human demonstration video, a goal photograph, a language instruction, or a combination, plus reference materials and completion criteria.
  • A documented robot interface with seven commands: state, status, help, frames, move_ee, move_joints and gripper.
  • A writable workspace where it can save images, scripts and notes for the duration of the trial.
  • Fixed budgets for elapsed time and for observation and action requests.

What the agent does not get

  • No simulation rehearsal.
  • No learned motion policy or skill library underneath it. The agent computes its own targets.
  • No solutions or logs from earlier trials, except in the experience studies, which are recorded separately.
  • No teleoperation. A reset or a human intervention ends the autonomous trial.

The platform

The robot is an I2RT YAM arm with six revolute joints and a parallel gripper over a tabletop workspace. A wrist-mounted Intel RealSense D405 supplies aligned colour and depth at 640 by 360 pixels, and a fixed overhead Logitech BRIO supplies rectified colour at 1920 by 1080. Bimanual towel folding adds a second arm expressed in the first arm's coordinate frame. Targeted throwing uses a separate runtime with timed motion programs, so that the swing and the gripper release can be scheduled together.

Photograph of two robot arms bolted to an aluminium rail that lies on a wooden tabletop, facing each other, with an overhead camera on a mast between them and cabling running down past the table edge.
The tabletop station: two I2RT YAM arms on a shared rail, each with a parallel gripper and a wrist camera, under one fixed overhead camera.
Overhead camera frame of a wooden tabletop with four 3D-printed parts lying apart on it, a white cylinder, a teal hexagonal nut, a teal ring and a white hexagonal nut, one robot arm reaching over the table and the base of a second arm visible.
What the agent sees from overhead: the first 1920 by 1080 capture of the two pair assembly trial replayed below, second by second.
Wrist-camera view looking down at the wooden tabletop, the two dark gripper jaws at the bottom edge of the frame, a white cylindrical post and a teal hexagonal nut lying on the table ahead of them.
What the agent sees from the wrist: the 640 by 360 colour frame returned by the same request, with depth aligned to it. The agent asks for both views at once and reads them together.

Agent self-report and human review

Most trials end with a report written by the agent, and every trial has a separate review of the final state and the execution video by a person. Five of the 112 counted runs ended with no report at all, so on those there is nothing on the agent's side to compare. Where both exist they can differ. On this page a disagreement means the agent stated an outcome that the review contradicted: the agent called a trial successful and the review called it a failure, or the other way round. Trials whose report could not be classified automatically are counted apart, under a label of their own, because that is a limit of the parser and not a judgement by the agent.

The pin that was never found

Four pair assembly example. GPT-6 Astra at high thinking effort, Codex CLI, single arm. 89.2 minutes, ended at the timeout.

Three pairs went together. The hollow tube was knocked over and stood back up twice. On the final insertion the gripper was already partly open when the arm moved, the long pin left the jaws, and the agent could not locate it again. It stopped and reported the failure rather than claiming the task was done.

Agent report: failure. Human review: failure. They agree.

Six cubes standing, then not

Six block tower example. GPT-6 Astra high, Codex CLI, single arm. 35.6 minutes.

The fourth cube went down out of alignment, and the agent regrasped and corrected it while the stack was still intact. The sixth level was beyond a straight-down reach, so it built the top two cubes on the table and carried them up as a pair. All six stood after release. The tower came down during the withdrawal that followed. The agent then measured where the cubes had landed, found one grey cube at about 0.70 m in the robot base frame against a 0.65 m radial grasp limit, and ended the attempt as unsuccessful with time and command budget still left.

Overhead view of a completed six-cube tower with the gripper open beside it.
All six cubes standing after release.
Overhead view of six cubes scattered across the table.
The same scene after the arm withdrew.

Agent report: failure. Human review: failure. They agree.

This is why withdrawal counts as part of the task, and it is also where recovery stops: reobservation and regrasping fixed the earlier alignment error, but they could not reach a cube that had left the workspace.

One corner that stayed unfolded

Simultaneous towel folding example. GPT-6 Astra high, Codex CLI, bimanual. 27.6 minutes.

The agent carried both short ends inward at the same time and staged the placements, which is what the demonstration asks for. Repeated corner slips earlier in the run, and a curled long-edge fold, left a loose triangular corner that is plain in the final overhead image.

Overhead photograph of a folded towel with an unfolded triangular corner.
An example, counted as a failure.
Overhead photograph of a folded towel with both end flaps in place.
An example, counted as a success, for comparison.

Agent report: an approximate or partial match. Human review: failure. This is a difference of degree rather than a contradiction, and the trial is counted as a failure.

A smaller model that knows when to stop

Two pair assembly, GPT-5.6 Terra at high thinking effort, Codex CLI, single arm, an example. 22.1 minutes.

The circular pair went together. On the second pair the hexagonal ring slipped and knocked its post out of reach near the base of the arm. The agent tried to recover, could not, and reported partial success rather than success.

Agent report: partial success, not task success. Human review: failure.

Across all 112 counted slots

  • 1trial where the agent stated an outcome the review contradicted
  • 10reports the fixed parser could not classify, which is a limit of the parser and not a disagreement
  • 5runs that ended without a final report
In the agent's own words

The main fold order was attempted correctly, but the final flaps are skewed, unevenly overlapped, and not reasonably aligned. Success: no.

Towel folding, sequential. The review counted this example a success; the agent counted it a failure.

The page states two counts, how many trials carry an explicit disagreement between the agent's stated outcome and the human review, and how many reports the automatic parser could not classify. The second group is much the larger of the two, and it is not evidence of disagreement.

Programs the agent wrote while the task was running

Across 65 recorded executions, 18 sessions retained 37 Python working files. These comprise 29 files for perception and geometry, 7 for composing robot calls, and 1 for drawing reference points into an image. The logs hold a further 135 shell events with inline Python across 48 sessions. One case can be followed all the way from a camera image to a physical motion.

Four pair assembly

  1. Observe the parts

    The agent compared wrist camera views and depth measurements. Bright white surfaces sometimes gave unreliable depth readings.

  2. Compute a position from two views

    An executed Python command converted selected pixels in captures 1 and 16 into viewing rays using their camera calibration. It computed the closest points on the two rays.

  3. Revise the alignment

    The hexagonal ring caught on the peg. Contact stops led to lifts and revised alignment using the visible ring hole and peg. The ring passed over the peg and settled near its foot after release.

  4. Check the final state

    The trial record reports success. The final photographs show all four assemblies after the gripper released the parts and the arm withdrew.

In the agent's own words

Executed 3D estimation code from the assembly run, with comments and line breaks added. The ray function uses the calibration for each capture to return the camera position and viewing direction.

# Camera positions p, q and viewing directions a, b
p,a=ray(1,450,175)
q,b=ray(16,293,189)
dot=lambda a,b:sum(x*y for x,y in zip(a,b))

# Geometry between the two viewing rays
d=[q[i]-p[i] for i in range(3)]
aa=dot(a,a); bb=dot(b,b); ab=dot(a,b)
ad=dot(a,d); bd=dot(b,d)

# Solve for the closest point on each ray
t=(ad*bb-ab*bd)/(aa*bb-ab*ab)
s=(ab*ad-aa*bd)/(aa*bb-ab*ab)

# Print the two estimated 3D positions as [x, y, z]
print([p[i]+t*a[i] for i in range(3)],
      [q[i]+s*b[i] for i in range(3)])

View the assembly run and its final photograph

What the programs contribute to performance is not established. Settling that would need comparisons that hold the interface capabilities and the inherited experience fixed.

Limitations

Study limitations

  • Time and money. Mean completion time per successful trial runs from 20.6 to 50.8 minutes and mean inference cost from USD 7.46 to USD 24.14. Both are real barriers to deployment.
  • Coordinated deformable manipulation is the weakest case. Simultaneous towel folding is the lowest configuration at 3 of 5, and sequential folding costs 50.8 minutes and USD 24.14 on average.
  • Cost and time grow with task complexity, as the six block tower shows against the pyramid and the two towers, and a large share of the overhead is reviewing and verifying outcomes. About 40% of the throwing time comes after the throw.
  • Image-conditioned decision cycles take a large share of elapsed time, with perception, planning and code generation coupled inside single model responses. There is substantial room to cut observation and decision overhead and the cost of writing and transferring experience.
  • Model choice matters. On the same two pair assembly task, GPT-5.6 Terra reaches 1 of 5, GPT-5.6 Luna 0 of 5 and Claude Fable 5.1 3 of 5.

Cite this work

@article{jia2026agentaspolicy,
  title   = {Agent as Policy for Robotic Manipulation},
  author  = {Jia, Mengzhao and Lin, Yang and Zhang, Xixin and Zhang, Zhihan and Liu, Xiaobai and Jiang, Meng},
  journal = {arXiv preprint arXiv:2609.12541},
  year    = {2026}
}

Credits

Assembly part geometry
Adapted from the AutoMate dataset. Tang et al., AutoMate: Specialist and Generalist Assembly Policies over Diverse Geometries, 2024, arXiv:2407.08028. The parts on this page are 3D prints of those meshes, scaled with their aspect ratios preserved.
Robot arm
I2RT YAM arm and the I2RT robot software.
Inverse kinematics
Mink, by Kevin Zakka: Python inverse kinematics based on MuJoCo.
Coding agent
Codex CLI, by OpenAI. Claude Code, by Anthropic.