Where you are. You can place any robotics system on four axes. This lesson adds the question you should ask of everything you read from here on: how much of this was demonstrated, and how much was announced?
Two tabs
In one tab, a humanoid empties a dishwasher. It picks up a plate, turns, walks to the cupboard, puts the plate away, comes back for the next. Four minutes, one continuous take, no cuts, no hand entering frame to reset anything. You watch it twice. It is good.
In the other tab, a different company’s news page. A humanoid crosses a factory floor carrying a parts bin. A customer’s logo sits in the corner. The word on the page is deployed. It is also good. It is also thirty seconds long.
Ask one question of each: what would I have to see to conclude I was wrong?
For the first, that is easy. A cut. A reset. A human hand. The video is one uninterrupted take of one task, and the claim it supports is exactly as large as the video: this machine did this thing, once, all the way through, in its builder’s own building.
For the second there is nothing to check. How many robots are on that floor? For what fraction of a shift? Autonomous, or is somebody wearing a headset in another building? The clip does not say, and no amount of rewatching will make it say.
Same production quality, same impressive machine. One of them is evidence. The other is a photograph of a claim.
The idea in one paragraph
As of August 2026, robotics has more capital, more talent and more polished video than at any point in its history, and the distance between what has been demonstrated and what has been announced has never been wider. The useful skill is not knowing who is ahead; that ranking will change before you finish this course. The skill is grading claims. Four questions do most of the work: who ran the evaluation, on whose hardware, was a human in the loop, and is a named third party on the record with money. Apply them and the field sorts into a small set of things that are real, a larger set that are early, and a very large set that are a video.
The four questions
- Who ran the evaluation? The builder, on a benchmark it designed, or somebody else?
- On whose hardware, in whose building? Home turf, or a customer’s floor?
- Was a human in the loop? Full teleoperation, shared autonomy with a remote takeover, or nothing?
- Is a named third party on the record? A named customer, a signed contract, a disclosed hour count, a filing.
Answering all four sorts the claim onto a ladder.
Wider than the screen; scroll it sideways.
One subtlety saves you from misusing this. A single uncut take is strong evidence about a capability and almost no evidence about a deployment. A named customer and a five-figure hour count are the reverse: strong evidence of deployment, silent on how much of the work the robot actually does. Both mistakes are common, and they run in opposite directions.
Four companies to practise on
Run the questions across the field as it stood in mid-2026.
| Company | Strongest thing you can check | Still not disclosed |
|---|---|---|
| Agility Robotics (Digit) | Named customers including GXO, Schaeffler and Mercado Libre; more than 65,000 cumulative operating hours across nine customer facilities; Toyota’s Canadian manufacturing arm signed in February 2026 for seven Digits | Revenue. A $2.5 billion SPAC merger was announced in June 2026 and, per reporting on the initial filings, full revenue figures were not in them |
| Figure (F.03) | A four-minute unedited dishwasher run published in January 2026, 61 sequential actions with no reset and no human intervention; robots at BMW’s Spartanburg plant in South Carolina from 30 June 2026, following an earlier pilot BMW says ran through production of more than 30,000 cars | How many robots, autonomous versus teleoperated, throughput, uptime. None of it is on the announcement page |
| 1X (NEO) | A published consumer price, $20,000 outright or $499 a month; pre-orders open since October 2025, with first deliveries stated for 2026 | Verified customer deliveries. As of mid-2026 no publicly verified delivery had surfaced, and 1X’s own materials still described shipping in the future tense |
| Tesla (Optimus) | Prototypes exist and are shown publicly; on the Q4 2025 earnings call in January 2026 Musk said the robots were still in R&D and not doing useful work | A unit count. Tesla has published none, reporting puts deployed units somewhere between a few hundred and about a thousand against an earlier target of a thousand doing productive work by end of 2025, and by most accounts volume production had not started as of mid-2026 |
The rest, briefly. Boston Dynamics unveiled a production Atlas at CES 2026 and committed its 2026 units to its parent Hyundai and to Google DeepMind rather than to outside customers; Hyundai’s published plan puts Atlas on real plant processes, starting with parts sequencing, from 2028. Apptronik raised a further $520 million, and its Apollo 2 is the platform Google DeepMind demonstrates on. Unitree shipped more than 5,500 humanoids in 2025 and claims the largest unit share, though at least one outside analyst puts AgiBot ahead; either way those are development hardware rather than deployed labour. And the boring half of the industry, the Fanucs and ABBs and KUKAs, kept quietly making money running the paradigms from lesson 0.5 at enormous scale.
The best published numbers sit below the videos
Google DeepMind announced Gemini Robotics 2 on 30 July 2026. In its own publication, on hardware and tasks it chose, it reported whole-body picking on an Apptronik Apollo at 76.3% from a shelf, 68.4% from a table and 45.7% from the floor. Two-arm gripper work on a Franka setup ran 74.2% on general pick-and-place up to 89.6% on precise insertion. Multi-finger dexterity is where the floor is: 92% to unscrew a bulb, but 44% to tie a trash bag, 40% on a ziplock, 36% to screw the bulb back in. DeepMind also names movement speed as a remaining weakness.
The other result worth carrying is a shape rather than a value. Physical Intelligence reported in late 2025 that adding reinforcement learning on the robot’s own autonomous attempts, plus expert corrections at the moments it went wrong, on top of a model trained by imitation, more than doubled throughput and cut failures by half or better on its hardest tasks. Those figures are self-reported and have not been independently reproduced. The shape is what matters: imitation learning alone hits a reliability ceiling, and the current move is reinforcement learning on the robot’s own mistakes. You will meet both halves in Module 3.
One more warning before you compare anything. There is no trustworthy cross-model comparison in this field. Every vendor benchmarks on hardware it picked, tasks it wrote and scene resets it defined. Two success rates from two labs are not on the same scale and cannot be ranked.
What you can actually touch
The commercially interesting robots are mostly unavailable. The research stack mostly is not.
| Thing | Status, August 2026 |
|---|---|
| π₀, π₀-FAST, π₀.₅ (Physical Intelligence) | Open weights, Apache-2.0, in the openpi repository |
| GR00T N1.7 (NVIDIA) | Open weights under NVIDIA’s own model licence; the code is Apache-2.0 |
| SmolVLA (Hugging Face) | Open, 450M parameters, the one you can realistically train end to end at home |
| Gemini Robotics ER 2 (DeepMind) | Closed weights, but callable through an API today |
| Gemini Robotics 2, Helix (Figure), π*₀.₆ | Announced; weights not released |
LeRobot is the centre of gravity for the open half: around 26,500 stars in August 2026, an active pull-request queue, and coverage of the whole loop from teleoperation through recording and training to deployment. Two practical notes. Its dataset format reached v3, so older tutorials and published datasets may no longer load; and it moves fast enough that you should pin a release tag in requirements.txt rather than installing whatever is current.
Hardware follows the same discipline. The SO-101’s official bill of materials was $229.88 for a leader and follower pair in mid-2026, or $121.94 for the follower alone, and that excludes 3D printing, cameras and a power supply. A pre-assembled motor kit from one common vendor was $288.99, printed parts $35 more.
For simulation, MuJoCo remains the correct default: it installs with pip, runs natively on Apple Silicon on the CPU, and needs no GPU. It ships several point releases a year, so write 3.x and move on.
Where the leverage is
Wider than the screen; scroll it sideways.
Job postings are primary sources; salary aggregators are not, so trust the first and discount the second. Postings consistently ask for deep PyTorch or JAX rather than tutorial familiarity, experience training models that take images and language together, named simulators, sim-to-real transfer, ROS 2 as a baseline, and large-scale demonstration-data pipelines. That last one is the unglamorous data-engineering half of the job, and it is the part your existing background maps onto directly.
Apply the rubric to your own offer letter too. Headline compensation at these companies is dominated by equity denominated in a last private-round valuation. That is an option on the humanoid thesis being right, not the same asset as cash.
The rare profile, and the reason Module 6 exists, is fluency in both the agentic layer and the policy layer. Very few people have both.
Wider than the screen; scroll it sideways.
Check yourself
1. A four-minute uncut dishwasher video, and a thirty-second clip captioned “deployed at a customer site”. Which supports the larger claim, and a claim about what?
They support different claims and neither dominates. The uncut take is strong evidence of a capability: this system completed this task, once, without intervention, in the builder’s own space. It says nothing about a second task, a different kitchen, or the hundredth attempt. The captioned clip gestures at a deployment but supports almost nothing, because it discloses no count, no duration and no autonomy level. If you had to bet on which company is closer to shipping, the uncut take is worth more; if you had to bet on which has revenue, neither clip helps and you should go looking for a contract.
2. A company announces more than $300 million in committed multi-year orders. What do you still not know?
Whether any of it has been recognised as revenue, what milestones release it, how long it is spread over, and who the customers are. In the actual 2026 case most of the figure came from one three-year contract with an unnamed customer. An unnamed counterparty is the tell: a named customer can be contacted and can confirm or deny, which is exactly what makes naming them costly. Question four of the rubric is doing all the work here.
3. DeepMind published a 45.7% success rate for picking an object off the floor with a frontier model. Why is that more useful than any demo video, and why is it probably an upper bound?
More useful because it is a rate over many trials rather than a selected trial, and because it is on the record from the group with the most to lose from being wrong. An upper bound because the lab chose the tasks, the hardware, the scene resets and which numbers to publish. Every incentive pushes a published set of results toward the favourable end of what was measured, and none of it was run by an outside party.
4. Two labs report success rates for their vision-language-action models on picking things up. Why can you not order them?
Because they are not measuring the same thing. Different robots, different grippers, different task definitions, different tolerance for what counts as success, different scene resets between trials, and usually different definitions of a trial. Absent a shared benchmark run by a third party on shared hardware, the two numbers share a percent sign and nothing else. The honest move is to compare each model against a baseline you ran yourself on your own setup.
5. Someone tells you the SO-101 costs $100. What is wrong with that, and how do you check?
It repeats a 2025 launch-coverage figure that does not describe a purchasable 2026 kit, and it silently excludes printed parts, a camera and a power supply. Check by pricing the official bill of materials at source and then adding the things it excludes. The mid-2026 answer was $229.88 for a leader-follower pair from the official parts list, and realistically $300 to $400 all-in without a 3D printer. This is question one of the rubric applied to a price: who ran the number, and what did they leave out?
6. Why is an equity package at a large private valuation not equivalent to cash compensation?
Because the valuation is a price set in a negotiated private round, not a market clearing price, and the shares are usually illiquid for years. The value is contingent on the humanoid thesis being right on a timeline that, by the evidence in this lesson, has slipped repeatedly for nearly everyone. It is a position, and it should be sized like one.
Do this
Start notes/00-claim-ledger.md and give it five columns: claim, source, grade, date checked, re-check by.
Put three rows in it today. Pick one humanoid company’s deployment claim, one model release, and one hardware price. For each, run the four questions and write the grade as a rung on the ladder rather than a feeling: edited reel, uncut take, partner floor, named customer, disclosed hours. Set the re-check date from the half-life figure, so version numbers get weeks and prices get months.
Then do the part that actually teaches. Take one row and chase it to the primary artifact: the repository, the docs page, the vendor’s product listing, the filing. Write one line on the difference between what the artifact says and what the coverage said. In my own research for this lesson that difference was the single largest source of error, in both directions - confident numbers with no source behind them, and real caveats that were dropped somewhere between the filing and the headline.
What you can now do
You can grade any robotics claim by evidence rather than production values: name who ran the evaluation, on whose hardware, whether a human was in the loop, and whether a third party is on the record. You can place the major humanoid companies on that ladder as of August 2026, name which models and tools are actually open and usable, quote realistic hardware prices instead of headline ones, and tell which parts of everything above will expire first.