Melanie Mitchell, who studies cognition and computation at the Santa Fe Institute, argues that large language models think in ways fundamentally unlike human cognition, and that most current testing methods are not built to notice the difference. She laid out the case in a Quanta Magazine podcast interview published Thursday.

Mitchell’s central claim is that despite being trained on human-generated text, these systems process and produce language through mechanisms that diverge sharply from human reasoning. She calls this an “alien intelligence,” a framing she says other researchers, including Stanford’s Michael Frank, have also used to argue that AI cognition should be studied the way developmental psychologists study infants or comparative psychologists study animals: as an unfamiliar mind, not a smaller version of an adult human one.

To do that rigorously, Mitchell proposes six principles for assessing machine cognition, drawn from a paper she references in the interview. First, researchers should recognize their own tendency to project human-like understanding onto systems that produce fluent English. Second, they should treat their own hypotheses skeptically and build control experiments rather than confirming what they expect to find. Third, benchmark items need novel variations to test whether a system generalizes or has simply memorized a pattern. Fourth, models should be probed rather than treated as unopenable black boxes. Fifth, researchers should separate what a system can be shown to do (performance) from what it is actually capable of (competence). Sixth, failure cases deserve as much publication and analysis as successes, since negative results often expose the real mechanism at work.

Mitchell illustrates the risk of skipping this rigor with Clever Hans, a horse in early-1900s Germany that appeared to solve arithmetic by tapping its hoof the correct number of times. Scientists at the time were convinced the horse could count. A psychologist, Oskar Pfungst, ran a controlled test: when the questioner did not know the answer, or was hidden from the horse’s view, Clever Hans failed. He was reading involuntary facial cues, not doing math. Mitchell draws a direct parallel to an AI system that appeared to answer questions about scientific diagrams correctly, until a control experiment removed the diagrams entirely and the system kept answering just as well, exploiting a spurious link between the wording of the question and the answer rather than reasoning about the image.

She also points to a gap in how AI research operates compared with other experimental sciences: independent replication, the practice of a second lab confirming a first lab’s result before it is trusted, is not common practice in AI research the way it is in psychology or biology.

None of this is presented as settled science. Mitchell frames it as a research agenda, not a verdict on whether any given model reasons or does not.

For anyone buying AI systems on the strength of benchmark scores, Mitchell’s argument has a direct consequence. If those benchmarks were designed around assumptions about how humans solve problems, and the system being scored solves problems through a different mechanism entirely, a high score does not necessarily mean the capability it claims to measure. Procurement teams evaluating vendors on leaderboard rankings should ask whether the benchmark’s control conditions, in Mitchell’s sense, have ever been tested, not just whether the score is high.

Reported by Quanta Magazine, published August 20, 2026, based on an interview with Melanie Mitchell for the podcast The Joy of Why.