Inherent says an AI agent it calls Faraday reached the same conclusions as existing peer-reviewed studies more accurately than systems built on Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5, using a base model roughly an order of magnitude smaller than either. The London startup ran the comparison itself. There is no independent lab or third-party benchmark cited to confirm the result.
That distinction matters. The claim comes from Inherent’s own research post and a company post on X, not from a peer-reviewed evaluation or a neutral leaderboard. Readers should treat the outperformance figure as a company-reported result until an outside group replicates it, which is a bit ironic given that replication is the exact task being scored.
The setup itself is specific. Faraday was handed published papers and asked to work out their conclusions on its own, with the true outcome kept hidden until it finished, a task cofounder and chief scientist Edward Hughes compared to how PhD students often begin their research training. Inherent also graded Faraday on what Hughes called research taste: whether the agent chose sound experiments and designed them well, not just whether it landed on the correct number.
The model behind Faraday is Qwen 3.6, a 27 billion parameter system, small next to the frontier-scale models it was measured against. Inherent trained it using reinforcement learning, a method that rewards good outcomes rather than encoding fixed rules, betting that approach generalizes better across scientific fields than training on descriptions of how research gets done.
If a 27 billion parameter model can match or beat far larger systems on a bounded, checkable task like reproducing a paper’s results, that is a data point for task-specific training beating raw scale, at least within narrow, verifiable domains. It is the same logic behind smaller coding and math models that have closed gaps with generalist frontier systems on their own turf. It does not yet say anything about Faraday’s ability to generate genuinely new scientific findings, which is the harder problem Inherent says it actually wants to solve.
Notably, Inherent did not build its own coding tool for Faraday. The agent uses OpenAI’s GPT-5.5 Codex to write and run code, a choice the company frames as similar to scientists relying on existing lab software instead of building every instrument themselves.
The company is small: about a dozen staff working in person from an office in London’s King’s Cross, with plans to grow to 20 to 25 people by year end. Hughes, a Google DeepMind alumnus along with three of his cofounders, has also called publicly for the UK to drop “garden leave,” a UK rule that keeps people who quit a company from starting or joining a competitor for months. He said the policy affected him personally before he left to start Inherent. Inherent raised a $50 million seed round earlier this year; this result is a separate development from that raise.
For operators evaluating AI research tools, the near-term takeaway is narrower than the headline suggests. Faraday’s edge is demonstrated on paper replication, a bounded task with a known correct answer, not on open-ended discovery. Anyone piloting an AI research agent for the next quarter should ask vendors for the same kind of held-out, answer-known test before trusting claims about novel hypothesis generation.
Reported by Anna Heim for TechCrunch, published August 22, 2026.