Harvey, the legal AI company AI Insiders has covered before for its cross-matter memory feature and its own Legal Agent Benchmark, shifted its research effort in a new direction: post-training a base model rather than shipping another product layer. The company published results for Harvey Tenet, a legal-agent model built on top of Kimi K3, the open-weight model from China’s Moonshot AI. A US legal-AI vendor choosing a Chinese base model to post-train, rather than a domestic frontier model, is itself notable, and Harvey did not explain the choice beyond describing Kimi K3 as its starting point.

Harvey, working with Fireworks research, trained Tenet using asynchronous reinforcement learning inside simulated legal work. An agent gets a partner’s task instruction, a client matter’s documents, and a rubric describing what a passing work product must contain, then drafts deliverables and gets graded by an LLM judge against that rubric. On the held-out LAB set, Harvey says the untrained Kimi K3 baseline clears roughly half of what the post-trained checkpoint finishes, a gap the company rounds to nearly double; the edge narrows to 20 percent on the LAB Contracts variant. Harvey also says the same checkpoint generalizes to two outside benchmarks it never trained on: Mercor’s APEX Agents and Crosby’s Redline Bench.

Every one of those figures comes from Harvey’s own evaluation, on tasks Harvey substantially controls, including a benchmark it authored (LAB) and rubrics graded by another AI model. A vendor’s internal benchmark run can show its post-training changed model behavior in the intended direction. It cannot establish how the model performs against a legal team’s actual matters, under adversarial conditions, or against a rival vendor’s model graded by a neutral third party, none of which Harvey’s disclosure appears to include.

The more consequential claim sits underneath the benchmark table. Harvey post-trained three narrower models, for M&A diligence, high-volume document review, and firm-specific knowledge search. Each is trained separately and deployed as a tool or subagent that Tenet routes to, rather than a capability baked into one generalist model. On its diligence benchmark, a task that can require reading up to 80 million tokens of documents, Harvey says no baseline model cleared 43.8 percent of rubric criteria, while its self-distilled, harness-trained version reached 60.1 percent. On its document review product, Harvey says the post-trained model cut cost per cell by roughly a factor of ten while improving citation quality.

Harvey’s framing is explicit: these gains came from training inside the production harness and task structure, not from feeding the model more legal text. If that holds up outside Harvey’s own tests, it reframes how any vertical AI company should spend its research budget. A startup building for accounting, healthcare intake, or insurance claims may not need a bigger proprietary corpus so much as a realistic simulated harness and a rubric tight enough to reward the right behavior. That lowers the barrier to post-training relative to raw data acquisition, and it makes the choice of base model, now demonstrably including non-US open-weight options, a live strategic decision rather than a default.

For legal teams evaluating Harvey next, the open question is whether Tenet’s gains hold on real client matters and against a benchmark Harvey did not write. Ask for a third-party or customer-run evaluation before treating Harvey’s own numbers as a performance ceiling.

Harvey published these results on its company blog in the “Harvey Tenet Research Preview” post.