Liquid AI's Open d1: 3B and 600M Decision Models That Run at 16 ms on Jetson AGX Thor and 8 ms on RTX 4090
Liquid AI shipped d1-3B and d1-omni-600M as open weights on Oct 7 — single-token decision models that hit 16 ms on Jetson AGX Thor and 8 ms on RTX 4090. Day-zero llama.cpp support, GGUF builds, packed-batch throughput up to 1,106/sec on AMD MI325X. License is lfm1.0 (not OSI-approved) and the vision/audio benchmarks are limited.
Liquid AI released d1-3B and d1-omni-600M as open weights on October 7, 2026 — the first open-weight decision-model pair small enough to live on a Jetson Orin Nano and fast enough to make a routing decision before the user notices the click. Single-question latency is 16 ms on Jetson AGX Thor, 8 ms on RTX 4090, 9 ms on AMD MI325X, and packed-batch throughput reaches 1,106 decisions per second on a single MI325X. Day-zero llama.cpp support, GGUF builds the same day, and Hugging Face checkpoints that load in LM Studio the morning after the launch. (Liquid AI blog, HF blog)
A “decision model” is a specific category that matters: it produces a single yes/no/score token from a closed-set question — intent classification, PII detection, safety routing, guardrail decisions, RAG relevance scoring. Zero output tokens. The compute budget is one forward pass. That’s the budget d1-3B was built for.
The benchmark claim, framed accurately
Liquid AI ran d1-3B across seven public text benchmarks — SQuAD 2.0, Civil Comments, MASSIVE intent, PubMedQA, BoolQ, XNLI, PAWS-X — and reports 82.9 mean, ahead of the open Decider 4B baseline at 81.1 and ahead of the larger Decider 35B-A3B on Decision Index 0.2.1 (48.57 vs 47.11). The d1-3B is the smaller model. (HF blog, explainx.ai)
Framing this matters: a 3B-parameter model beating a 4B baseline is impressive. A 3B model beating a 35B-active MoE on a decision index is the headline. The MoE loses because per-token compute for a single forward pass is high; the d1-3B wins because a dense 3B is cheap enough to run on every request and still beat the sparse model on the closed-set questions a decision layer actually sees.
The d1-omni-600M is multimodal: text + image OR text + audio, built on LFM2.5-Encoder-350M with bolted-on vision and audio encoders. It’s the first small-multimodal decision model with open weights. (MarkTechPost)
Single-question latency across six hardware targets
This is the table that matters for edge deployment. All six rows from Liquid’s own benchmark, single closed-set question on d1-3B:
| Hardware | Single-question latency | Notes |
|---|---|---|
| Jetson AGX Thor | 16 ms | Edge AI module, 100W class |
| Jetson AGX Orin (64 GB) | 26 ms | Edge AI module, 60W class |
| Jetson Orin Nano | 50 ms | Edge AI module, 15W class |
| Apple M5 Pro | 30 ms | Laptop-class Apple Silicon |
| RTX 4090 | 8 ms | Desktop GPU |
| AMD MI325X | 9 ms | Datacenter GPU |
The Thor number is the unlock: a 16 ms routing decision fits inside a 60 Hz frame budget with margin to spare. The Orin Nano at 50 ms is the floor — that’s a real-time budget on a $249 dev kit. (Liquid AI blog, explainx.ai, Jetson AI Lab)
Packed-batch throughput scales linearly with hardware budget:
| Hardware | Decisions/sec (packed batch) |
|---|---|
| AMD MI325X | 1,106/sec |
| RTX 4090 | 475/sec (64 states per batch) |
| Jetson AGX Thor | 262/sec |
| Jetson AGX Orin (64 GB) | 110/sec |
| Jetson Orin Nano | 38/sec |
At 1,106 decisions per second on a single MI325X, a serving tier can route an entire cluster’s traffic through one GPU. At 38/sec on an Orin Nano, a single edge box handles its own requests without calling home.
What a decision model actually does
The “decision model” category is heating up for a reason. The pattern: take a generative model’s output (or input), ask a closed-set question about it, route on the answer. Is this email PII? Is this query on-topic? Should this agent’s proposed tool call be allowed? Should this RAG retrieval be expanded or stopped?
The category has consolidated around a handful of entrants: Perplexity’s pplx-decider API, OpenAI’s Decisions API (shipped at DevDay alongside GPT-6 Luna), llama.cpp’s day-zero decision-model support, and now GEV-26B-Decide (quantized to 17 GiB). d1-3B is the first small-multimodal entrant with open weights. (AI Weekly)
The practical unlock: a generative model is overkill for routing. A decision model is one forward pass to a single token. Cost is one to two orders of magnitude lower per call. Latency is 10x lower. Accuracy on closed-set questions can match or beat the larger model because the question is narrow.
Loading d1-3B in LM Studio
LM Studio exposes the same llama.cpp runtime that the GGUF checkpoints target. Walkthrough for a workstation with a single RTX 4090:
- Download the GGUF. From the d1-3B model card, pull the GGUF file — Q4_K_M is the right default (~2.0 GB on disk), Q8_0 if you have the VRAM headroom (~3.4 GB on disk). The repo also has a
-omni-600M-GGUFfor the multimodal variant (~0.5 GB at Q4_K_M). - Open LM Studio → Search → “LiquidAI/d1-3B” (or use “My Models” → Import → point at the downloaded
.gguf). Pick the Embedding or Completion runtime — d1-3B is a single-token forward pass, not a chat model. - GPU offload config. The GPU Layers slider controls how many transformer layers go to VRAM. For a 3B model on a 24 GB RTX 4090, set the slider to the full model depth (all layers) — d1-3B fits with full offload plus headroom for K/V cache and activations. The system RAM requirement is still load-bearing even with full GPU offload: VRAM holds the K/V cache and intermediate activations, RAM holds the model weights. Budget roughly 2x the GGUF file size for runtime headroom (~4 GB RAM for the Q4_K_M build).
- Context length. Set to 512 or 1024 — these are single-question forward passes, not long-context chat. The model is trained on short prompts.
- Test it. A request like “Is this email PII? yes/no” returns a single token in ~8 ms on the 4090. A batch of 64 closed-set questions returns in ~135 ms wall-clock at 475 decisions/sec.
For the d1-omni-600M, the same steps apply; budget ~0.5 GB VRAM and ~1 GB RAM, and use the multimodal runtime in LM Studio.
llama.cpp from the command line
If you skip the GUI:
# Clone and build llama.cpp (CUDA build for NVIDIA, Metal for Apple Silicon)
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp && cmake -B build -DGGML_CUDA=ON && cmake --build build --config Release
# Download the GGUF
huggingface-cli download LiquidAI/d1-3B-GGUF \
d1-3B-Q4_K_M.gguf --local-dir ./models
# Run a packed-batch decision pass on RTX 4090
./build/bin/llama-batched \
-m ./models/d1-3B-Q4_K_M.gguf \
--n-gpu-layers 32 \
--batch-size 64 \
--ctx-size 512 \
-p "Is this email PII? yes/no"
The --n-gpu-layers 32 flag is the offload control — set it to the model depth to put every transformer layer on GPU. For the d1-omni-600M, lower the layer count to match the smaller depth. The --batch-size 64 flag is what unlocks the 475 decisions/sec throughput on RTX 4090 — a single-state batch is the 8 ms number; packed batches amortize weight loads across requests.
Caveats you need to know
License. d1-3B and d1-omni-600M ship under Liquid’s lfm1.0 license — not OSI-approved. Commercial use conditions apply. Read the license text before shipping a product on top of these weights. (Liquid AI license)
Loading requires trust_remote_code=True. The d1 architecture is not yet in transformers’ mainline model registry, so from_pretrained() will execute code from the Liquid repo. Pin a revision (revision="<commit-sha>") before loading in production — without a pin, a future commit to the repo changes what code runs on your machine. Required dependencies: transformers>=5.14, torch, torchvision, pillow (for the omni variant). (Liquid AI blog)
Vision and audio benchmarks are limited. The vision split on Decision Index v0.3 is a private benchmark — not reproducible outside Liquid. Liquid’s own docs call the audio decision benchmarks “currently an open problem.” If you’re evaluating d1-omni-600M for a vision or audio routing task, you’ll need to build your own eval set. (HF blog)
The 8 ms RTX 4090 number is one warm call, not throughput. AI Weekly’s digest noted the measurement methodology: bf16 + CUDA graphs + median of 20 warm calls. That’s single warm requests only, not token throughput. Burst traffic, cold-cache requests, or first-call compilation may be materially slower. If you’re sizing for production, measure your own p50/p95/p99 cold-cache latency — Liquid’s table is the best-case number. (AI Weekly)
What this means in practice
If you’re building a routing/guardrail layer today, you have three options: call a hosted decision API (Perplexity, OpenAI), run a quantized larger model on your own GPU (GEV-26B-Decide at 17 GiB), or run d1-3B locally with full GPU offload on commodity hardware.
The d1-3B path is the first one that fits on a Jetson Orin Nano at 50 ms. The decision layer is no longer a datacenter concern. It’s an edge concern.
The d1-omni-600M is the first multimodal decision model with open weights. If you’re routing on image or audio inputs, the only open-weight option at sub-1B parameters is now Liquid’s. Whether the multimodal quality holds in third-party reproduction is the open question.
What to verify when you load it:
- The 82.9 mean benchmark. Reproduce on your own eval set before shipping.
- The latency numbers. Measure on your hardware with your prompt distribution — single warm calls vs cold cache, single state vs packed batch.
- The license. lfm1.0 is not OSI-approved; commercial conditions apply.
- The repo code.
trust_remote_code=Trueexecutes remote code; pin a revision before production.
Sources
- Liquid AI — d1 open release
- Hugging Face — d1-3B and d1-omni-600M blog
- Jetson AI Lab — d1-3B model page
- Hugging Face — d1-3B model card
- MarkTechPost — Liquid AI releases open-weight d1-3B and d1-omni-600M
- explainx.ai — Liquid AI Open d1-3B / omni-600M
- BenchLM — d1-3B benchmarks
- Liquid AI on X — d1 launch thread
- AI Weekly — AI News Today, October 7: Top Stories