Magic Hour built a testing harness that gave two AI models, Fable 5 and Sol 5.6, identical creative briefs: produce four 15-second videos, three product ads and one mini-documentary, using the same skills, tools and production pipeline. The company published the results as a public test of whether frontier models can be trusted to run creative work unsupervised. Its own writeup, built to showcase the product both models generated footage through, should be read with that incentive in mind.

That incentive is worth stating plainly: Magic Hour sells the AI video and image tools that both models called through a local MCP server (the protocol that lets an AI model invoke outside software directly) to render every shot in the test. A company whose product supplies the raw footage is not a neutral judge of how well AI handles creative video, and the writeup includes no outside rater to check Magic Hour’s own scoring against.

Each run moved through the same ten stages: research, concept development, scripting, storyboarding, a video-model screen test, production, dailies review, sound, editing and a final pass, with a checkpoint validation loop after every stage. Because neither model can watch motion, both extracted still frames from their own footage to review it. Each model then self-rated its finished cut, 1 to 5, against five criteria drawn from the brief: whether the product stayed visually consistent shot to shot, whether on-screen branding and text held up, overall production craft, sound design, and creative originality, the category the brief weighted most heavily.

Cost separated the two models more than output quality did. The eight runs, four briefs times two models, cost Magic Hour $83.26 in credits and tokens for Fable 5 and $45.92 for Sol 5.6, a combined $129.18. Fable came in 80% more expensive on average, driven by token spend rather than MagicHour credit spend: Fable wrote roughly 367,000 output tokens across the test against 222,000 for Sol. Sol read more than it wrote. It issued 22 web searches across the eight runs versus 9 for Fable, a gap Magic Hour attributes to Fable pulling reference images directly rather than searching repeatedly.

On creative direction, Magic Hour’s team judged Fable the more literal of the two, sticking closer to each brief’s stated audience and taking fewer risks. Sol swung bigger and drifted further from the brief, in one case dropping the “for men” positioning a sunscreen ad was supposed to carry. Neither model, in the company’s own assessment, produced a cut it would call tasteful.

The clearest pattern involves voice. Six of the eight runs avoided the MCP server’s celebrity-clone voice list entirely, with both models reasoning on their own that using an unlicensed celebrity’s voice cut against the brief’s intent even where a narrator was expected. Only one run, Fable’s burger ad, used a celebrity clone. When the Big Ben mini-documentary required narration, Fable extracted a synthetic voice from a video model’s native audio track and Sol auditioned and cast a local speech-synthesis voice, workarounds Magic Hour calls functional but not good.

Four briefs scored against one company’s own rubric cannot establish which model has better creative taste, and Magic Hour does not claim it does. What the test can support is narrower: given an identical agentic pipeline and identical tools, both models made cautious, unfinished-feeling creative choices and left unused budget on the table rather than iterate toward something sharper, even when nothing forced them to stop early. Magic Hour has open-sourced the harness on GitHub so other teams can run the same comparison against their own model of choice.

For any team weighing whether an agentic pipeline can run ad production without a human in the loop, this result argues for keeping a person on final creative approval through at least the next few model generations, and for treating a vendor’s own creative benchmark as a starting point for your own testing rather than a verdict to build a workflow on.

Adapted from Runbo Li, “Testing Fable vs Sol in terms of taste (they are both bad),” published on the Magic Hour blog on August 17, 2026.