← Back to The Signal
Techne · Physical AI · Oikos

The House Changed.
The Checkpoint Did Not.

Figure's Helix 2.5 transferred three learned chores across 30 unfamiliar homes. The result is not general domestic labor. It is evidence that the robotics bottleneck is moving from motion to transfer, reliability, and operating economics.

Physical AIHumanoidsZero-shot transferCapital allocation

Evidence cutoff: September 19, 2026. Company evaluations and operating claims remain attributed. A demonstration is not a deployment, and a deployment is not yet a profitable robot-hour.

An original unbranded humanoid robot carefully makes a bed in an unfamiliar home
Original As Above conceptual artwork. Generalization becomes visible when a policy can act carefully inside a room it has never seen.

Architecture readout

  • Transfer: Figure says one fixed Helix 2.5 checkpoint per chore was evaluated across 30 homes where no training data had been collected. The homes and objects were unseen. The chores were not.
  • Reliability: Figure reports 56% aggregate full-task success after Index pretraining, versus 9% for an otherwise comparable policy trained from scratch. That is a meaningful research result and a 44% failure rate under the company's own rubric.
  • Market: Industrial arms and mobile robots still own the mature work. Humanoids are entering bounded industrial cells, developer fleets, and supervised home trials.
  • Capital: The investable bridge is not a walking body. It is lower adaptation cost, more accepted work per scheduled hour, less assistance, and service economics that improve at the second site.
  • Counter-signal: Another company can report a higher success rate on a narrower task without disproving Figure. Robotics metrics are inseparable from task scope, horizon, intervention rules, and evaluation independence.

A robot entered a house it had never seen. The furniture had moved. The towels were different. The toys were scattered in new places. No engineer collected a new demonstration after arrival. No model update followed.

The house changed. The checkpoint did not.

That is the essential fact inside Figure AI's September 17 Helix 2.5 announcement. The company evaluated three long-horizon chores across 30 Bay Area rental homes: putting scattered toys into a basket, folding towels, and making a bed. Figure says it used a fixed checkpoint for each chore, collected no training data in the evaluation homes, and applied no adaptation to the unfamiliar rooms or manipulated objects.

The result deserves attention. It also requires discipline.

The chores were specified with task data collected elsewhere. The evaluation was company-run. Figure reported 56% aggregate success under a full-completion rubric, with safety intervention counted as failure. A policy trained from scratch reached 9% under the same setup. The leap is not from ignorance to arbitrary housework. It is from a learned behavior in one distribution to that behavior surviving a change of place and object.

That is narrower than a general household worker. It is more important than another choreographed gait video.

This desk has followed physical AI through identity, joints, verification, production, and the invoice. Helix 2.5 moves the argument to the next layer: can a trained behavior travel?

What zero-shot means when matter pushes back

In language models, zero-shot often means answering a class of question without examples in the prompt. In robotics, the phrase must carry more weight because a mistake can break an object, stop a line, or injure a person.

There are at least four different claims that can hide inside the same term:

  1. New objects: the robot handles items it did not encounter during task training.
  2. New environments: the same behavior works under a different layout, lighting condition, geometry, or background.
  3. New embodiments: the policy transfers to a different hand, arm, sensor suite, or body.
  4. New tasks: the robot performs a behavior that was not represented in its task-specific training.

Helix 2.5 is evidence for the first two, not the fourth. Figure is unusually explicit about this boundary. The evaluation homes and objects were unseen, while the three chores were defined through fine-tuning data gathered elsewhere.

That qualifier does not diminish the result. It locates it.

A commercial robot rarely fails because it has never seen the concept of a towel. It fails because this towel is heavy, partly occluded, folded into itself, lying under different light, on a surface that changes the grasp. Real work is a long tail of small differences. Environmental transfer is how a useful policy begins to escape the laboratory.

What Figure actually measured

Figure used an unforgiving scoring rule. The whole chore had to finish. No partial credit. A safety intervention failed the attempt.

Figure reports that Index pretraining lifted aggregate full-task success from 9% to 56%. Index is the company's human-behavior pretraining system. The company says it now produces roughly 35 minutes of new human experience per second and that the largest pretraining run improved downstream action prediction in a manner consistent with its scaling forecast.

All of those statements come from Figure. They are not independent replications. Brett Adcock is the founder, chief executive, and the most interested possible narrator of the Figure thesis. The right response is neither dismissal nor surrender. Attribute the claim, inspect the evaluation, find the denominator, and ask what evidence must appear next.

The decisive number is not 56% by itself. It is 56% under a defined full-task rubric, on three learned chores, across unfamiliar homes, without in-home adaptation. Change the task horizon or intervention rule and the percentage changes meaning.

Figure makes a second claim with direct commercial implications. It says Helix 2.5 used half as much task-specific data as a representative Helix 02 behavior while expanding evaluation from one training house to 30 unfamiliar homes. If that relationship survives external use, behavior specification becomes cheaper and the addressable deployment field becomes wider.

That is the real product hypothesis. Not a robot that knows every chore, but a pretraining system that lowers the marginal cost of teaching the next useful behavior.

What the result does not prove

Helix 2.5 does not establish that Figure 03 is ready for unsupervised domestic service.

A 56% completion rate means the company's system failed 44% of evaluated tasks under its own protocol. Some failures will be harmless. Others will require intervention, reset, cleanup, or repair. In a research lab, a failure is an observation. In a home, it is a service call, a frustrated customer, or a safety event.

The test also does not establish:

Figure itself calls the result first evidence, not the completion of household robotics. That is the correct register.

The goal is not to make the breakthrough smaller. It is to prevent a legitimate breakthrough from being asked to prove a market it has not yet reached.

The benchmark problem: 99% of what?

Sunday Robotics reported a 99.1% zero-shot success rate for ACT-2 across 785 garment-folding attempts in unfamiliar homes. The company counted 778 successful folds across nine garment types and said it performed no per-home adaptation.

That number is stronger than 56% and less comparable than it looks.

Sunday's unit of evaluation was one garment fold and stack. Figure's unit was a multi-step room-level chore with many objects or a whole bed. Sunday excluded categories such as socks and bras from its preview. Figure required every object in the task to reach an acceptable final state. Both evaluations were produced by the companies building the systems.

The lesson is not that one robot is better. The lesson is that robotics lacks a shared grammar for reliability.

MeasureWhat must travel with itWhy it matters
Success rateTask definition, horizon, object count, tolerance, retriesA short fold and a room-level cleanup cannot share a denominator.
Zero-shotWhat was unseen: object, room, body, or taskThe phrase otherwise hides the actual transfer.
AutonomyInterventions, remote assistance, reset protocolHuman rescue can turn a failure into a polished clip.
DeploymentCustomer, paid status, hours, acceptance criteriaA pilot site is not recurring commercial demand.
ReliabilityScheduled hours, accepted work, downtime, service costThe operating burden determines the invoice.

A shared benchmark will eventually matter. A shared accounting language matters now.

The 2026 deployment ladder

Humanoid robotics is not replacing industrial automation as one coherent wave. Different forms are clearing different gates.

1. Industrial arms and mobile robots: the cash-clearing base

Industrial arms, collaborative robots, and autonomous mobile robots remain the mature business. Their environments can be engineered. Their safety cases are established. Their tasks are narrow enough to measure cycle time, uptime, and payback.

If a machine only needs to move totes from one station to another, wheels remain cheaper and simpler than legs. If a fixed arm can tend the machine, a humanoid is usually the wrong first purchase. Morphology is an economic decision, not a cultural preference.

Humanoids become interesting where the facility was designed around people and the task demands a changing combination of walking, reaching, manipulation, and tool use. The target is the awkward remainder, not every automated task.

2. Supervised industrial humanoids: the operating frontier

Figure, Agility, Boston Dynamics, and Apptronik are each moving through bounded industrial work, but their evidence differs.

Figure says its earlier Figure 02 deployment at BMW logged more than 1,250 runtime hours and handled more than 90,000 parts over an 11-month program. Those numbers are company-reported, but they describe work in a real plant rather than a stage.

Agility says Digit has accumulated more than 65,000 operating hours and passed 100,000 totes in GXO operations. Digit 5 is planned for early access in the first half of 2027 and wider availability later that year. The product page warns that specifications remain preliminary. The operational hours matter. They do not settle field-service economics or unit profitability.

Boston Dynamics introduced its production Atlas in January, with 2026 deployments committed to Hyundai's robotics center and Google DeepMind. It is no longer accurate to describe Atlas only as a research mobility benchmark. It is now an enterprise product entering an industrial evidence phase.

Apptronik's Apollo program is paired with manufacturing and logistics partners, while Google DeepMind reports that one Gemini Robotics 2 checkpoint controlled multiple embodiments, including Apollo 2 with different hands and a dual-arm Franka system. That is another kind of transfer claim: a model crossing bodies rather than houses.

3. Training fleets: the factory as classroom

Tesla's July investor update said Fremont production lines for Optimus were under installation and that initial builds would support internal data collection and development through the Optimus Academy. The language is important. Lines under installation are not a shipped fleet, and a training fleet is not an external product.

Tesla's advantage thesis remains powerful: use manufacturing scale, vehicle AI infrastructure, and internal factories as a learning field. The missing register is the same one this desk applies elsewhere: accepted work, assistance burden, external customer evidence, and cost.

4. Home-first systems: autonomy with an operating backstop

1X offers NEO as a home-first robot, with a $20,000 ownership option and a later $499 monthly plan. Its order materials are candid that early users receive basic autonomy and that complex tasks can use scheduled Expert Mode remote supervision.

That operating model is not a disqualification. It is a disclosure. Teleoperation can produce customer value, collect edge-case data, and bridge the autonomy gap. The economics depend on how often the bridge is used, how much it costs, and whether assistance declines with experience.

5. Developer bodies: price compression before labor substitution

Unitree lists the G1 from $13,500, while its store lists an R1 configuration from $4,900. These prices change access for laboratories, developers, and educators. Unitree also warns that the category remains in an early exploration phase and that buyers of basic configurations should not assume unrestricted secondary development.

Price compression matters even before autonomous work arrives. Cheaper bodies expand the population of experiments. They do not establish that the same machines can complete useful work reliably enough to replace labor.

The brain layer is becoming a starting point

The architectural shift underneath these bodies is the vision-language-action model.

A VLA takes perception and instruction and produces actions. The category now includes Helix, Physical Intelligence's policies, Google DeepMind's Gemini Robotics lineage, NVIDIA's GR00T stack, Skild's S1, and open projects such as OpenVLA.

The change is not that one model can control every robot. The change is that teams increasingly begin with a pretrained physical representation rather than learning a narrow behavior from zero.

That changes the economics of development in three ways:

  1. Fewer task-specific demonstrations may be required. Figure's 2x claim is one company's evidence of this direction.
  2. More value moves into data mixture and evaluation. The source, diversity, and failure coverage of pretraining data become strategic assets.
  3. Embodiment becomes a compatibility problem. A model that can cross hands, arms, or bodies may lower integration cost, but the physical system still imposes torque, latency, reach, thermal, and safety limits.

The robot foundation model is becoming a default starting point. It is not yet a universal operating system.

The Oikos read: generalization is a cost curve

For capital, the most important part of Helix 2.5 is not the bed. It is the implied cost curve.

A robot vendor must pay to specify a behavior, adapt it to a site, maintain it as the environment changes, and support failures after deployment. If pretraining lets the same policy survive more variation with less local data, the marginal cost of a new site falls. If it does not, every customer becomes a custom engineering project.

That distinction separates a platform from a services business wearing a robot costume.

Figure's capital structure shows the scale of the wager. Its Nscale agreement contemplates access to as many as 100,000 Vera Rubin GPUs, beginning with an initial $3.5 billion compute commitment and a first targeted deployment in the second half of 2027. The commitment is not current robot revenue. It is a claim that more compute and human-behavior data will lower the cost of useful physical generalization.

The evidence should eventually appear below the model:

A scaling law is scientifically interesting. A falling deployment cost is economically decisive.

Where the value may accrue

LayerWhy it could winWhat to verifyFailure mode
Data and computePretraining improves transfer across tasks, sites, and bodiesPerformance per unit of compute; data rights; utilizationSpend rises faster than accepted work
Actuators, joints, sensorsEvery body needs reliable motion and perceptionYield, durability, qualified capacity, architecture dependenceSupply response removes scarcity or designs change
Robot platformsIntegrated hardware and model can compound fleet learningUptime, intervention, gross margin, fleet retentionSupport cost scales with deployment
Integrators and incumbentsThey own workflow, safety, installed base, and customer trustCommissioning time, recurring service, customer paybackPlatform vendors internalize the layer
OperatorsAutomation can raise throughput and resilienceAccepted hours, quality, labor redeployment, total paybackPilots remain theater

The glamorous layer does not automatically own the margin. Compute can be paid before a robot ships. Component suppliers can benefit before autonomy matures. Integrators can earn cash while platforms absorb field-learning costs. The end user can capture the largest surplus if competition pushes robot prices down faster than service costs rise.

This is why market size is not the thesis. The thesis is the mechanism by which a solved constraint becomes cash.

Four paths through 2028

1. Transfer compounds inside bounded work

Pretrained policies require fewer demonstrations, and supervised industrial tasks spread across similar sites. Humanoids become useful in human-designed facilities before they become general household workers.

Confirmation: lower site-specific data requirements, higher held-out success, falling intervention, repeat customer expansion.
Disconfirmation: every site still requires a bespoke engineering team.

2. The model improves faster than the service burden

Evaluation scores rise, but maintenance, teleoperation, resets, and customer support remain expensive. Robots work more often without becoming attractive businesses.

Confirmation: autonomy metrics improve while gross margin and utilization stall.
Disconfirmation: accepted hours grow faster than support headcount and depreciation.

3. Specialized forms keep most of the work

Arms, cobots, and wheeled systems absorb the valuable tasks while humanoids remain useful only where human morphology is economically necessary.

Confirmation: buyers redesign processes around simpler forms rather than pay for general bodies.
Disconfirmation: humanoids repeatedly win whole-workflow contracts on total cost and adaptability.

4. The household becomes a data market before it becomes a labor market

Home-first robots rely on remote assistance, collect edge cases, and improve slowly. Early buyers fund the learning field. The dominant asset is the dataset and operating system, not near-term replacement of domestic labor.

Confirmation: supervised fleets expand while unsupervised task reliability remains below consumer expectations.
Disconfirmation: independent tests show near-continuous autonomy across varied homes and long task horizons.

As Above, so below

Correspondence is useful here as a method, not evidence.

Above is the checkpoint: a compressed model of action, intention, and possibility.

Below is the house: a twisted towel, a hidden toy, an unfamiliar bed, a human who expects the work to finish.

The model is general only to the degree that the levels correspond.

Mentalism begins the movement in representation. Cause and Effect follows it into matter. The robot acts, the room changes, and responsibility must remain attached to a principal. The more intelligence leaves the screen, the less acceptable it becomes for identity, authority, and failure history to disappear behind the word autonomous.

The Great Work is not a perfect demonstration. It is repeated refinement under resistance. One successful fold is possibility. A stable distribution of accepted work across unfamiliar places is transformation.

The operating conclusion

Helix 2.5 is a real signal because the environment changed while the checkpoint stayed fixed. It is not proof of a general household employee, a consumer product, or a profitable robot-hour.

The next frontier is now visible:

  1. transfer a learned behavior without local retraining;
  2. complete the whole task at a commercial reliability threshold;
  3. recover safely when the long tail arrives;
  4. lower assistance and service cost as the fleet grows;
  5. bind every consequential action to a named principal;
  6. turn accepted work into a second customer and a second invoice.

Figure called extended household footage boring on purpose. That is the right aspiration, even if the current reliability is not yet boring enough.

In robotics, boring is the compliment. A useful machine is the one you stop watching.

Above: a policy that generalizes. Below: a machine that finishes. The future belongs to the system that can do both, repeatedly, under burden, with an accountable principal.

What to watch next

Primary sources

Evidence cutoff: September 19, 2026. Figure, Sunday Robotics, Agility Robotics, Boston Dynamics, Tesla, 1X, Unitree, and Google DeepMind performance, product, deployment, capacity, and benchmark statements remain attributed to their sources and have not been independently audited by As Above. Success rates from different companies are not directly comparable unless task scope, horizon, intervention rules, objects, and evaluation methods match. This publication is educational analysis, not individualized investment, procurement, legal, or safety advice. AI tools assisted source discovery and production; Marc Theiler retains editorial responsibility.

The weekly Signal

Follow the breakthrough until it becomes accepted work.

One source-linked intelligence letter on agents, capital, markets, and the body that has to live with them.

Share this Signal