Perplexity has released two search models that look at a document page the way a person does, as an image, and skip the text-extraction step most companies still run first. The family is called pplx-embed-v2-late, and Perplexity published it on its engineering blog on 7 October.
An embedding model turns text or an image into numbers a computer can compare, which is how a search system decides that a question and a page are about the same thing. Most of them squash an entire document into one set of numbers. These keep a set for every word-sized piece instead, an approach called late interaction because the matching of question to document happens at the end, piece by piece.
Picture a 200-page contract boiled down to one index card, then asked to find the clause on subcontractor liability. Perplexity describes this as an information bottleneck that tightens as documents get longer and cover more topics. Keeping a set per word lets each word of your question find its best match anywhere on the page, and the scores add up. The cost is storage and comparison work that grows with document length, which Perplexity says plainly.
The practical headline is the images. The usual pipeline for PDFs, slide decks and scans runs OCR (software that reads text off a picture), then searches the output. Perplexity says that step introduces errors and throws away layout, table structure and figures. These models take the rendered page straight in, so a typed question can find a chart or a table without anyone building and tuning an extraction stage first.
The second detail is the one a buyer should read twice. Perplexity trained a large model and distilled it into two sizes, 9B and 0.6B parameters, that share the same number space. A team can index its documents once with the 9B model, which is slow and expensive but runs only offline, then answer live queries with the 0.6B model, which is cheap and fast. Perplexity measured the gain: indexing with the 9B model added 1.6 percentage points on its domain-specific text tests and lifted a visual-document score from 62.3% to 63.5%, with query cost unchanged. By its own account, that recovers roughly half the quality gap between the two models on text.
The company also points to storage. Its models store 128 numbers per word-sized piece, against 2,048 for Tencent’s EVIE-4.5B and 4,096 for EVIE-8B and NVIDIA’s Nemotron ColEmbed V2 8B. Since late interaction multiplies storage, that gap decides which corpora are affordable to index.
Every result here comes from Perplexity’s own testing, and “state of the art” is the company’s phrase. On ViDoRe V3, a public benchmark for finding the right page among document images, Perplexity says its 0.6B model keeps pace with rivals roughly five times its working size, a count that includes only the parts that run when an image is encoded, about 340 million of its 594 million total. Its 9B model scores 65.2% there, trailing only EVIE. On MADQA, where an AI agent searches 800 PDFs to answer 500 questions, Perplexity reports 92.4% accuracy for the 9B model and 90.1% for the 0.6B, against 88.9% for a Mixedbread retriever. Mixedbread’s agentic search system scored higher at 93.4%, a gap Perplexity says falls within its confidence interval. Two of the other benchmarks, Q2D-Web and PPLX-Q2I, are Perplexity’s own creations built from its production traffic.
Both models are available on Hugging Face and work with the sentence-transformers and transformers libraries. Perplexity says it will roll out late-interaction support on its API Platform progressively, with no dates given, and that a technical report will follow later this year.
A team with a PDF-heavy search project should test the 0.6B model against its own pages, and look at the size of the index before the accuracy numbers, because per-word storage is where this approach gets expensive.
Perplexity, “Multimodal embeddings beyond a single vector,” Perplexity engineering blog, 7 October 2026.