Pantheon, a robotics startup building general-purpose models for robots, says it has collected more than one million unique tasks of human demonstration video at an average of $10 for each hour of footage. The company published the account on 6 October. Every number in it comes from Pantheon about Pantheon, and the post shows no results from a robot model trained on the data.

Robots need humans to produce their training examples because no internet-scale archive of machines doing chores exists. Language models learned from text that already sat online. Robot models need someone to perform tasks while cameras record, and the quality of that recording decides what can be learned.

Pantheon collects with UMI, short for Universal Manipulation Interface: a handheld gripper fitted with cameras that a person uses like a tool. According to the post, it sits between head-mounted video of people (easy to gather, but unlike a robot arm) and remote-controlled robots (accurate, but slow and costly). The company’s target is dexterous manipulation, meaning fine hand work such as fitting a part into a socket, as opposed to simply picking something up and putting it down.

The figures need careful reading. Pantheon says it grew from a core operations team of five to 90 operators in eight weeks, and that it now records over 16,000 unique tasks a day with 45 operators per shift. It says supplier prices for ready-made UMI data run around $60 an hour, and up to $150 for custom work. That $60 is the market price Pantheon says it avoided. Its own cost, shown in a chart in the post, started at $71.92 per recorded hour in early August, wandered between roughly $33 and $57 for weeks, and reached $10 only in the week of 21 September. The text calls $10 an average, which the chart does not support. The post also does not say what the $10 includes, though it says operators are paid triple the going local living wage.

The collection method is the real idea. Standard UMI work records a short task, stops, resets the scene, and repeats, so a lot of time is lost to overhead. Pantheon calls its alternative freeform collection: an operator gets about 15 objects on a table and does whatever they like, with no instructions and no labels. That data was unusable when collected. The company bet that vision-language models, which can describe video, would soon be good enough to label it, and says that took five months, using its own open-source tool, Argus.

Scripted collection covers what freeform misses. Freeform sessions drift toward picking and placing, so a task generator, Nomos, combines a catalogue of objects with a list of actions to assign specific jobs, and Pantheon says it can shift the mix within minutes. It also records failures and recoveries on purpose, arguing that robots trained only on clean successes cannot recover from a dropped object.

Pantheon describes the result as, as far as it knows, the broadest UMI collection anywhere. No outside party has checked that. It is releasing a 100-hour annotated sample, which gives researchers a way to test the quality claim themselves.

Until a model trained on this data beats one trained on cheaper or smaller sets, the post documents a cheaper input, not a better robot. Teams buying robot data should ask for that comparison, and for what the $10 covers, before treating the price as a benchmark.

Pantheon (pantheon.inc), research post on its UMI data collection operation, published 6 October 2026.