Liquid AI open d1: when a fast decision beats an AI answer
Liquid AI's new open-weight d1-3B and d1-omni-600M models return decisions without generating long answers. See verified benchmarks, on-device options, and how to test routing workflows.
Liquid AI released two open-weight d1 decision models on October 7, 2026: d1-3B and the smaller experimental d1-omni-600M. These are not general-purpose assistants built to write long explanations. Their core purpose is to read a state, evaluate a fixed question or set of candidate outcomes, and return probabilities without generating an answer token by token. That design is relevant to teams that repeatedly classify support tickets, route leads, inspect images, or make on-device decisions with strict latency limits.
What makes a decision model different?
A typical language model may respond to a classification task with a paragraph or a JSON object that must be parsed and checked. Liquid's d1 architecture instead handles the decision in a single forward pass, returning probabilities across the allowed outcomes. For a workflow where the only valid answers are, for example, 'sales', 'support', 'spam' or 'human review', the model does not need to spend time composing a written explanation. The trade-off is that the outcome set and the business logic need to be specified well in advance.
Choosing the right model type
| Use a decision model | Use a generative assistant |
|---|---|
| Classify a support request into fixed categories | Draft a helpful, contextual customer response |
| Route a lead or document to a known team | Research an ambiguous situation and explain options |
| Score many repeated inputs with low latency | Synthesize a detailed report or recommendation |
| Make constrained on-device decisions | Handle evolving, open-ended conversation |
The two releases and their limits
Liquid d1 open-weight models
| Model | Primary design and caveat |
|---|---|
| d1-3B | Approximately three billion parameters; text and image input; stronger reported decision quality and broader device support |
| d1-omni-600M | Experimental smaller checkpoint; text with image or text with audio; lower memory footprint but weaker reported decision scores |
Both checkpoints are available as open weights through Hugging Face, with local inference support including llama.cpp. Liquid's release describes support across Apple, AMD, Qualcomm and NVIDIA environments. These are model distribution and implementation claims, not a guarantee that every device configuration achieves the same throughput or that deployment has no licensing, integration or maintenance obligations.
What the published benchmarks actually say
Liquid reports a Decision Index v0.2.1 public-split score of 48.57 for d1-3B and 15.95 for the d1-omni-600M experimental checkpoint. Across seven text benchmarks presented in the announcement, their reported averages are 82.9 and 78.4 respectively. The company also reports an 8 ms single-question response for d1-3B on an RTX 4090, 16 ms on Jetson AGX Thor, 26 ms on Jetson AGX Orin and 50 ms on an Orin Nano. These are vendor-reported measurements on specified setups, not independently reproduced end-to-end application benchmarks.
Single-question latency excludes many real deployment costs, including data capture, preprocessing, queues, network hops, business rules and escalation. A model that responds in milliseconds can still be the wrong choice if confidence is poorly calibrated or wrong decisions have expensive consequences.
A real workflow: lead and support routing
- Define the closed set of outcomes: qualified sales inquiry, customer support, spam and human review.
- Build a labeled evaluation set from actual inquiries, removing private identifiers and sensitive fields where feasible.
- Run the decision model and compare its predictions with a simple rules-based baseline and an existing generative model.
- Measure false positives, false negatives, calibration and time per decision; assess whether different errors have different business costs.
- Choose a conservative confidence threshold. Send uncertain or high-impact cases to a person instead of forcing an automatic decision.
- Only automate routing after monitoring, logging, versioning and rollback are established. Reevaluate the system when inquiry patterns change.
For operations teams, the opportunity is not to replace a general-purpose AI agent with a smaller model everywhere. It is to separate recurring micro-decisions from the occasional task requiring reasoning and written explanation. A cheap, fast classifier may route 90% of predictable cases, while a human or generative system handles ambiguous exceptions. The 90% figure here is an illustrative system design goal, not a reported d1 accuracy measurement. Whether that split is economical depends on the real traffic, error costs, infrastructure and review burden.
Where the open-weight release matters
Open weights make it possible to examine, adapt and deploy a checkpoint closer to the data source, including on devices where sending every input to a hosted model is undesirable. For robots, cameras, field devices and local business systems, network reliability and responsiveness can be as important as benchmark scores. The smaller omni model is particularly interesting as an experimental route to combined text-and-image or text-and-audio decisions, but Liquid explicitly labels it an early checkpoint. The release says dedicated audio decision benchmarks are not yet mature.
- Test the exact input format: text, text plus image, or text plus audio rather than assuming one model supports all combinations equally.
- Confirm hardware latency with representative payloads, not only a short single-question benchmark.
- Review deployment and model license terms before commercial rollout, and document what version is deployed.
- Keep a rule-based or human fallback when the model is uncertain, unavailable or sees unfamiliar data.
- Protect sensitive input and audit decisions that affect people, access or money.
Liquid AI's open d1 release is a useful reminder that many AI applications need a dependable decision, not an essay. Its architecture and local distribution are worth testing for classification, routing and on-device work. The right evaluation is application-specific: accuracy, latency, hardware requirements, review workload and the cost of being wrong.
Explore related Make Better resources
- Claude Haiku 5.5 and high-volume workflow costs— Related inference cost trade-offs
- What should you automate first?— Determine whether recurring work needs a structured process
Sources & useful resources
- Liquid AI: Open d1— Primary launch article, October 7, 2026
- Liquid AI: Model catalog— Official model lineup and deployment options