Two companies released small, open AI models on 7 October 2026 that share an unfashionable idea: not every AI job needs a chatbot that writes. Liquid AI’s new d1 “decision models” pick an answer in a single step instead of generating text word by word [1]. Perplexity’s new pplx-embed-v2-late models turn text, images and document pages into searchable fingerprints, and let a small model search an index built by a large one [2][3].
Both releases are confirmed by the companies. Speed and benchmark figures below are the companies’ own, and we label them that way.
Why “one pass” matters
Most AI models people meet are generative. They write their answer one token (a small chunk of text) at a time, and each token is another trip through the model. That is flexible, but it is slow and costly when all you need is a short, structured answer: yes or no, which of these five options, is this message toxic, which document matches.
Both releases skip the writing step:
- A decision model reads the input and outputs its choice directly, in one trip through the network.
- An embedding model reads a query or a document and outputs a set of numbers that describe it. Search then becomes a matter of comparing numbers, not generating prose.
That is why these models can be small and fast enough to run on edge hardware, such as a robot’s onboard computer, or cheaply at very large scale.
Liquid AI d1: what was released?
Liquid AI released two open-weight models in its d1 family [1]:
- d1-3B, trained from Liquid’s LFM2.5-VL-3B vision-language model. It takes text and images.
- d1-omni-600M, which Liquid calls its first experimental checkpoint. It is built from a 350-million-parameter encoder with added vision and audio encoders, and takes either text and image or text and audio.
Liquid says that, unlike its generative models, “our d1 decision models don’t produce tokens. Instead, they produce an answer in a single forward pass” [1]. Both are on Hugging Face, with day-one support for llama.cpp, a popular tool for running models locally [1]. Liquid describes them as open-weight models you can download, fine-tune and deploy “without restrictions” [1]; we did not review the licence text for this draft.
How fast does Liquid say d1 is?
These are Liquid’s figures, measured by Liquid (on Jetson hardware, in collaboration with NVIDIA), one request at a time [1]. For a single question, d1-3B answers in:
- About 8 ms on an NVIDIA GeForce RTX 4090
- About 9 ms on an AMD MI325X
- About 16 ms on an NVIDIA Jetson AGX Thor
- About 26 ms on a Jetson AGX Orin 64 GB
- About 30 ms on an Apple M5 Pro
- About 50 ms on a Jetson Orin Nano
Liquid also reports that three questions about the same input take only about 1.3 times as long as one: on the Jetson AGX Thor, 16 ms becomes 20 ms [1]. Longer inputs take longer; Liquid’s own table shows a 3,400-token input taking 220 ms on the AGX Thor and 1,640 ms on the Orin Nano [1].
Because decision models do not generate output tokens, Liquid measures end-to-end time from input to answer [1]. That is a sensible metric, but it means these numbers are not directly comparable with “tokens per second” figures for chatbots.
How good does Liquid say it is?
Again, Liquid’s figures [1]:
- On its Decision Index v0.2.1 (public split), d1-3B scores 48.57. Liquid says that is ahead of every model under 10 billion parameters and on par with a decision model called Decider 35B-A3B, which is 12 times its size. d1-omni-600M scores 15.95 on the same index.
- On seven public text benchmarks covering reading comprehension, toxicity detection, intent classification, medical questions and cross-lingual understanding, d1-3B averages 82.9, against 81.1 for Decider 4B. d1-omni-600M averages 78.4, against 77.1 for Decider 2B.
Liquid is open about gaps. It does not report results on the private vision split of its index, and says dedicated audio decision benchmarks are “currently an open problem” [1]. It calls d1-omni-600M an early research release under active development [1].
What would you use d1 for?
Liquid’s examples are real-time: ten demos that run d1-3B in a loop over live camera input, from gesture-controlled games to live content moderation, reading one answer per frame [1]. With NVIDIA, it also shows d1-3B steering through a simulated environment in Isaac Sim, served on a Jetson device [1].
The common thread is a fixed set of choices that must be made quickly and repeatedly: is this frame safe, which way to turn, which intent did the user express. That is a narrower job than open conversation, and a much cheaper one.
Perplexity pplx-embed-v2-late: what was released?
Perplexity released two retrieval models on Hugging Face [2][3]:
| Model | Active parameters | Vector size | ViDoRe v3 (image) | ViDoRe v3 (markdown) |
|---|---|---|---|---|
| pplx-embed-v2-late-0.6b | 340M | 128 | 62.3% | 61.2% |
| pplx-embed-v2-late-9b | 7.4B | 128 | 65.2% | 64.7% |
The scores are Perplexity’s, on ViDoRe v3, a public benchmark for retrieving visual documents (measured as nDCG@10, a standard ranking-quality metric) [2].
Perplexity describes them as multimodal “late-interaction” retrievers for text, images and visual documents, built on Qwen3.5 [2]. Both were distilled from an internal 18-billion-parameter teacher model that Perplexity has not released [2].
What does “late interaction” mean, in plain words?
Many search systems squash a whole document into one list of numbers, a single “vector”. That is compact, but detail gets lost.
Late-interaction models keep one small vector for each token instead. Perplexity’s models produce one 128-number vector per token, and score a match by letting each part of the query find its best-matching part of the document (a method called MaxSim) [2]. Different parts of a question can match different parts of a page.
The practical upshot is search across text, images and visual documents [2][3], including rendered pages such as PDFs, without OCR, the step that normally converts a page image into text first (see source note). Tables, charts and layouts do not have to survive a text conversion to be found.
What is the “small model queries a big index” trick?
This is the most useful part for anyone running search at scale. The 0.6B and 9B models share one embedding space, so “the 0.6B model can query an index built with the 9B model” [2].
Indexing happens once and can be slow and expensive. Queries happen constantly and need to be fast and cheap. Splitting the job lets you build a higher-quality index with the big model, then answer every user query with the small one, without rebuilding the index.
There are practical limits on the model page. Text-only and image-only inputs must be encoded in separate batches; mixed text-and-image inputs are not supported [2]. The models need recent versions of the sentence-transformers and transformers libraries [2].
What this does not prove
- That d1 is as fast on your hardware. The latency numbers are Liquid’s, measured one request at a time on specific devices [1]. Real systems add pre-processing, networking and batching.
- That d1 beats larger models generally. The headline comparison is on Liquid’s own Decision Index, and d1 only does structured decisions. It is not a chatbot replacement.
- That d1-omni-600M is production-ready. Liquid calls it experimental [1].
- That Perplexity’s models are the best retrievers. The scores are Perplexity’s on one benchmark family. We could not open Perplexity’s full blog post (it timed out and returned 403), so we have not checked its wider benchmark claims.
- That “open” means fully reproducible. Perplexity has released the models, not its 18B teacher [2].
The Bottom Line
Liquid AI’s d1 and Perplexity’s pplx-embed-v2-late are small, open models built for jobs where writing is wasted effort. d1 picks a structured answer in one pass, in about 8 ms on an RTX 4090 by Liquid’s own figures [1]. Perplexity’s models search text, images and PDF pages without OCR, and let a cheap small model query an index built by a large one [2][3]. Neither is a headline-grabbing chatbot. Both are the kind of plumbing that makes AI fast and cheap enough to put everywhere.
Sources
- Liquid AI, “Open d1: Edge decision models for text, vision, and audio” (blog; benchmarks and latency are Liquid’s figures), 7 October 2026. https://www.liquid.ai/blog/d1-open
- Perplexity, pplx-embed-v2-late-9b model card, Hugging Face (model sizes, shared embedding space, ViDoRe v3 scores, training note, usage limits). https://huggingface.co/perplexity-ai/pplx-embed-v2-late-9b
- Perplexity, “We’re releasing pplx-embed-v2-late, two late-interaction embedding models that retrieve text, images, and pages with a shared embedding space for cross-model querying,” Perplexity Community (Announcements), 7 October 2026. https://community.perplexity.ai/t/were-releasing-pplx-embed-v2-late-two-late-interaction-embedding-models-that-retrieve-text-images-and-pages-with-a-shared-embedding-space-for-cross-model-querying/6296
Source note: Perplexity’s blog post “Multimodal embeddings beyond a single vector” (https://www.perplexity.ai/hub/blog/multimodal-embeddings-beyond-a-single-vector) timed out via our fetch tool and returned 403 via curl. The “without OCR” description follows the research brief’s confirmed wording and Perplexity’s announcement of retrieval over “text, images, and pages”.

