Harvey, the legal AI company, says an agent that studies its own graded work between jobs now clears every check on 15.7 percent of legal tasks, up from 2.9 percent without that habit. That is a fivefold jump. It also means the agent still falls short somewhere on more than 84 of every 100 tasks.

The numbers come from a post on X by Niko Grupen, published 8 October. The post never states his employer outright, but it describes Harvey’s Legal Agent Benchmark and says “we evaluated our own variant of wake-sleep”, so this article treats every figure as Harvey’s own.

Wake-sleep is an old idea. A classic learning algorithm alternates between two phases: in the wake phase the system works on real data, and in the sleep phase it goes offline to improve how it will behave next time. Later systems such as DreamCoder used the same rhythm to build up reusable knowledge. What Harvey adds is a version for agents that spend hours on a single task. The agent does real work, a review model then reads the graded record of that work, and the lessons it writes are handed to the agent on its next job.

Nothing here retrains the model. The lessons are plain text. Some work as checklists for what a finished memo or draft needs, and others as tips on method. All of them sit in a memory store and get placed in the agent’s context before the next task. Grupen’s post says the approach is mostly non-parametric and could later extend to actual training. For a builder, that means no fine-tuning bill, and a memory a person can read, edit, and delete.

“All-pass” means every rubric criterion on a task is met, a far harder bar than partial credit. On the softer measure the gain is modest: against the vanilla agent, the memory-equipped one met over 10 percent more rubric items on each familiar task and nearly 5 percent more on each new one. A task passes only when every criterion does, which is why small per-criterion gains can move the all-pass rate so much while most tasks still fall short.

The sample is small and the evidence is internal. Harvey ran 10 cycles over 196 tasks, using synthetic client matters and GPT-6 Luna as the agent model. The tasks come from two practice areas, Corporate M&A and Capital Markets. Of the 196, 110 were for training and 86 for validation. The post does not say how many tasks fall in the familiar and new groups, and it offers no replication from outside Harvey. It also never says what the headline rates are averaged over: 15.7 percent and 2.9 percent of 86 tasks are not whole numbers, so they cannot be a simple count over the validation set. Judges are two LLMs scoring against expert-written rubrics.

The first cycle added 122 lessons, and later cycles refined them. One rule about stating deadlines as durations grew, over nine cycles, to cover weekends and conflicting windows. Harvey put a commit gate, built on a model it calls Jev, in front of the memory to reject any lesson that names a client or draws on just one training example. The agent also changed how it worked, making about 2.5 times as many tool calls to run code and check its output, which made runs slower and costlier. Retrieving only relevant lessons halved the cost per task at similar quality.

The grader carries the whole design. The agent never sees its scores, but the review model does, and without a trustworthy way to tell good attempts from bad ones the loop would write down confident nonsense. Legal tasks with expert rubrics are an unusually friendly case. The technique should travel to work with checkable output, such as code with tests or reconciliations against known totals, and travel badly to work where quality is a matter of taste.

A team wanting to try it needs four things: logged trajectories, a grader it trusts, a gate that blocks overfit lessons, and a held-out set that the memory never sees. Harvey reports the largest gains in the earliest cycles, so a few cycles should show whether your grader is good enough to build on. Budget for the extra tool calls first, because the retrieval savings arrive only after the memory has grown.

Based on a post on X by Niko Grupen describing Harvey’s wake-sleep research, published 8 October 2026.