Index Ventures led the financing behind Intelligence, the startup that built Design Arena, TechCrunch reported Monday. The seed round totals $7.9 million. Conviction’s Sarah Guo and Mike Vernal also joined, alongside A* and Valkyrie. The financing is a wager that AI’s hardest remaining problem is not getting an answer right but deciding whether the answer is any good.

Design Arena started as a side effect. Co-founder Grace Li and a group of college friends were building an AI game engine a few weeks before their 2025 graduation, and the models could produce games that ran without producing games anyone wanted to play. That gap, between functional and fun, had no obvious scoring method. Li’s team concluded that only real people could settle the question, and built a system to collect that judgment at scale.

For individual users, Design Arena works something like a model router with a personality. A prompt window accepts a request, a dropdown sets the format (websites, images, and roughly a dozen other visual categories), and the platform generates several candidate outputs. The user then works through a series of head-to-head choices until every output is ranked from best to worst.

The consumer product is the visible layer. The business is enterprise: frontier labs plug into Design Arena’s ranking pipeline for continuous human feedback on their media-generating models, without having to recruit and manage a testing panel of their own. Li told TechCrunch that user indifference is the point. People rarely know or care which model produced a given output, so their preferences reveal what audiences actually want rather than which brand they favor. That pipeline now generates $60 million in annualized revenue, Li said, and the platform counts 5.3 million users worldwide.

Turning taste into a metric requires deciding whose taste counts, and Design Arena’s public answer is thinner than its revenue figure. The company requires users to log in, which lets it track how preferences shift by region and over time; Li pointed to web dashboards in Asia trending toward a more maximalist style than their Western counterparts. Beyond that geographic detail, TechCrunch’s account does not explain how Design Arena weights votes, screens for bots, or corrects for a rater pool that may skew toward people who enjoy ranking AI outputs rather than the broader population a lab ultimately ships to. A preference leaderboard is only as trustworthy as the crowd casting the votes, and the reporting leaves that crowd mostly undescribed.

Elsewhere in today’s issue are benchmarks built around a single correct answer to chase. Design Arena is a bet that the bigger opportunity now sits in the categories that got skipped, precisely because no one could agree on what counts as right.

Human-feedback platforms are not an automatic business. Chris Dixon’s a16z crypto fund backed Yupp, a similarly styled human-feedback startup, with $33 million. Yupp signed frontier labs as customers too and counted more than 1.3 million users, then closed down earlier this year regardless. LM Arena, a comparable ranking platform for text output, closed its Series A four months after the paid product launched, pulling in $150 million during January. Design Arena’s pitch rests on the claim that automated benchmarks are gameable, a point TechCrunch tied to last week’s breach at Hugging Face, and that human rankings are harder to fake.

For a lab shipping image, video, or interface-generating models, the math is straightforward. Building an equivalent panel of engaged testers across multiple continents costs more than paying for Design Arena’s data feed, at least until a rival undercuts the price or a lab decides evaluation data is too strategic to outsource. Teams evaluating multimodal or design-facing models should treat an undisclosed rater methodology as a real limitation, not a footnote, before wiring someone else’s taste into a product roadmap.

TechCrunch’s Russell Brandom first reported the Design Arena funding round and product details on August 3, 2026.