Luxobench, a new benchmark site published by Christian Keil, tests whether AI models can execute a full hardware project rather than answer questions about one. The task: start from a CAD model and a concept render of an interactive desk lamp and deliver a shoppable parts basket, build instructions, and any needed firmware, engineered so a hobbyist with a soldering iron can assemble it from mail-order parts.

Luxobench’s rubric scores each attempt out of 100: 30 points for basic function (does it power on), 20 for advanced function (movement, expressions, tracking a person nearby), 20 for cost (a target under $250), 20 for ease of manufacturing and assembly, and 10 for lead time from ordering to delivery. Firm requirements: a camera, a screen, movement, a light, and a footprint that fits comfortably on a desk.

This is a self-published benchmark from its creator, with design and implementation choices left to each model. Luxobench does not name which models it has tested or publish scores in the material reviewed here.

The format matters more than any single result: it shifts AI evaluation from text and code toward physical execution, including sourcing, cost control, and manufacturability.

Per Luxobench (luxobench.com), published by Christian Keil.