Ryan Greenblatt, Redwood Research’s chief scientist, told podcast host Dwarkesh Patel this month that most of what has improved AI pretraining data over the past several years has little to do with human expertise. Shuchao Bi, who helped start YouTube Shorts, later led multimodal post-training work at OpenAI, and now builds models at Meta Superintelligence Labs, made a closely related argument in a talk he gave more than a year earlier. Neither appearance was framed as investment commentary, but the overlap between them lands on a question with direct stakes for any company whose value depends on an exclusive data pile.
Greenblatt’s framing starts with a thought experiment: hold compute and data fixed, then ask how much a model still improves on technique alone. He calls that residual algorithmic progress and argues it explains a large share of the field’s advancement since GPT-3. Recreate GPT-3 today at its original compute budget, he told Patel, and current technique alone would produce something well past GPT-4’s capability, with no new data and no additional hardware.
Bi’s talk did not use Greenblatt’s exact term, but the substance lines up. He argued that raw scraped text is not the most useful arrangement for training, and that further scaling gains depend on reordering and reweighting that data rather than collecting more of it, a process he described as “equalizing intelligence per token.” He also broke down how humans build knowledge, through proposing problems, absorbing existing work, testing ideas against feedback, and folding results back in, then argued AI systems can now speed up nearly every step of that loop, including deciding what to work on in the first place.
Greenblatt is more direct about human expert data specifically. He told Patel that most pretraining data gains trace back to research on which datasets work and to what he called “schleppy labor” in filtering, not to expert-curated content. He pointed to the shift from OpenWebText, an older open training corpus, to FineWeb, a newer one, as an example: reproducible with computing hardware and no human expert data. He acknowledged a richer internet and a larger pool of people posting online as separate effects but said both are smaller than the gains from better scraping and curation technique.
The overlap prompted MBI Deep Dives, the investment research newsletter that flagged both appearances in an August 18 note, to reconsider its own position. The newsletter had previously pointed to Alphabet’s reported bid for Spirit Airlines’ customer data during the carrier’s bankruptcy proceedings as evidence that proprietary data still carries real value. After encountering Greenblatt and Bi, MBI Deep Dives wrote that its prior view had shifted, while stopping short of discarding the data-moat thesis and describing its own read of the AI landscape as something to hold loosely given how fast it changes.
The commercial question neither speaker addressed directly is sharper than the research framing suggests. If algorithmic technique and better data distributions really can substitute for volume, the moat weakens for any company pricing itself on an exclusive corpus: social platforms sitting on years of posts, publishers negotiating licensing deals, specialty vendors selling annotated domain data. Their pitch depends on scale being durable and hard to replicate. What would settle the argument is a lab matching a data-rich competitor’s benchmark results using a smaller, differently weighted dataset and materially less compute, replicated by someone outside the lab that first reported it.
One caveat belongs here that neither Greenblatt nor Bi supplies. Both work at organizations whose edge depends on algorithmic and training technique rather than exclusive data access, which is exactly the factor their argument elevates. The claim also remains untested where it counts commercially: no lab or platform has walked away from a data-exclusivity deal on the strength of it, and Alphabet’s pursuit of Spirit Airlines’ records, reported the same week these ideas resurfaced, is a company still betting the other way. Anyone raising capital on a proprietary dataset should be ready to name the specific benchmark result that would prove technique, not data, is doing the work.
Adapted from reporting and commentary in MBI Deep Dives (mbi-deepdives.com), published August 18, 2026.