Liquid AI's Open d1: 3B and 600M Decision Models That Run at 16 ms on Jetson AGX Thor and 8 ms on RTX 4090

Liquid AI shipped d1-3B and d1-omni-600M as open weights on Oct 7 — single-token decision models that hit 16 ms on Jetson AGX Thor and 8 ms on RTX 4090. Day-zero llama.cpp support, GGUF builds, packed-batch throughput up to 1,106/sec on AMD MI325X. License is lfm1.0 (not OSI-approved) and the vision/audio benchmarks are limited.

Liquid AI released d1-3B and d1-omni-600M as open weights on October 7, 2026 — the first open-weight decision-model pair small enough to live on a Jetson Orin Nano and fast enough to make a routing decision before the user notices the click. Single-question latency is 16 ms on Jetson AGX Thor, 8 ms on RTX 4090, 9 ms on AMD MI325X, and packed-batch throughput reaches 1,106 decisions per second on a single MI325X. Day-zero llama.cpp support, GGUF builds the same day, and Hugging Face checkpoints that load in LM Studio the morning after the launch. (Liquid AI blog, HF blog)

A “decision model” is a specific category that matters: it produces a single yes/no/score token from a closed-set question — intent classification, PII detection, safety routing, guardrail decisions, RAG relevance scoring. Zero output tokens. The compute budget is one forward pass. That’s the budget d1-3B was built for.

The benchmark claim, framed accurately

Liquid AI ran d1-3B across seven public text benchmarks — SQuAD 2.0, Civil Comments, MASSIVE intent, PubMedQA, BoolQ, XNLI, PAWS-X — and reports 82.9 mean, ahead of the open Decider 4B baseline at 81.1 and ahead of the larger Decider 35B-A3B on Decision Index 0.2.1 (48.57 vs 47.11). The d1-3B is the smaller model. (HF blog, explainx.ai)

Framing this matters: a 3B-parameter model beating a 4B baseline is impressive. A 3B model beating a 35B-active MoE on a decision index is the headline. The MoE loses because per-token compute for a single forward pass is high; the d1-3B wins because a dense 3B is cheap enough to run on every request and still beat the sparse model on the closed-set questions a decision layer actually sees.

The d1-omni-600M is multimodal: text + image OR text + audio, built on LFM2.5-Encoder-350M with bolted-on vision and audio encoders. It’s the first small-multimodal decision model with open weights. (MarkTechPost)

Single-question latency across six hardware targets

This is the table that matters for edge deployment. All six rows from Liquid’s own benchmark, single closed-set question on d1-3B:

Hardware Single-question latency Notes
Jetson AGX Thor 16 ms Edge AI module, 100W class
Jetson AGX Orin (64 GB) 26 ms Edge AI module, 60W class
Jetson Orin Nano 50 ms Edge AI module, 15W class
Apple M5 Pro 30 ms Laptop-class Apple Silicon
RTX 4090 8 ms Desktop GPU
AMD MI325X 9 ms Datacenter GPU

The Thor number is the unlock: a 16 ms routing decision fits inside a 60 Hz frame budget with margin to spare. The Orin Nano at 50 ms is the floor — that’s a real-time budget on a $249 dev kit. (Liquid AI blog, explainx.ai, Jetson AI Lab)

Packed-batch throughput scales linearly with hardware budget:

Hardware Decisions/sec (packed batch)
AMD MI325X 1,106/sec
RTX 4090 475/sec (64 states per batch)
Jetson AGX Thor 262/sec
Jetson AGX Orin (64 GB) 110/sec
Jetson Orin Nano 38/sec

At 1,106 decisions per second on a single MI325X, a serving tier can route an entire cluster’s traffic through one GPU. At 38/sec on an Orin Nano, a single edge box handles its own requests without calling home.

What a decision model actually does

The “decision model” category is heating up for a reason. The pattern: take a generative model’s output (or input), ask a closed-set question about it, route on the answer. Is this email PII? Is this query on-topic? Should this agent’s proposed tool call be allowed? Should this RAG retrieval be expanded or stopped?

The category has consolidated around a handful of entrants: Perplexity’s pplx-decider API, OpenAI’s Decisions API (shipped at DevDay alongside GPT-6 Luna), llama.cpp’s day-zero decision-model support, and now GEV-26B-Decide (quantized to 17 GiB). d1-3B is the first small-multimodal entrant with open weights. (AI Weekly)

The practical unlock: a generative model is overkill for routing. A decision model is one forward pass to a single token. Cost is one to two orders of magnitude lower per call. Latency is 10x lower. Accuracy on closed-set questions can match or beat the larger model because the question is narrow.

Loading d1-3B in LM Studio

LM Studio exposes the same llama.cpp runtime that the GGUF checkpoints target. Walkthrough for a workstation with a single RTX 4090:

  1. Download the GGUF. From the d1-3B model card, pull the GGUF file — Q4_K_M is the right default (~2.0 GB on disk), Q8_0 if you have the VRAM headroom (~3.4 GB on disk). The repo also has a -omni-600M-GGUF for the multimodal variant (~0.5 GB at Q4_K_M).
  2. Open LM Studio → Search → “LiquidAI/d1-3B” (or use “My Models” → Import → point at the downloaded .gguf). Pick the Embedding or Completion runtime — d1-3B is a single-token forward pass, not a chat model.
  3. GPU offload config. The GPU Layers slider controls how many transformer layers go to VRAM. For a 3B model on a 24 GB RTX 4090, set the slider to the full model depth (all layers) — d1-3B fits with full offload plus headroom for K/V cache and activations. The system RAM requirement is still load-bearing even with full GPU offload: VRAM holds the K/V cache and intermediate activations, RAM holds the model weights. Budget roughly 2x the GGUF file size for runtime headroom (~4 GB RAM for the Q4_K_M build).
  4. Context length. Set to 512 or 1024 — these are single-question forward passes, not long-context chat. The model is trained on short prompts.
  5. Test it. A request like “Is this email PII? yes/no” returns a single token in ~8 ms on the 4090. A batch of 64 closed-set questions returns in ~135 ms wall-clock at 475 decisions/sec.

For the d1-omni-600M, the same steps apply; budget ~0.5 GB VRAM and ~1 GB RAM, and use the multimodal runtime in LM Studio.

llama.cpp from the command line

If you skip the GUI:

# Clone and build llama.cpp (CUDA build for NVIDIA, Metal for Apple Silicon)
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp && cmake -B build -DGGML_CUDA=ON && cmake --build build --config Release

# Download the GGUF
huggingface-cli download LiquidAI/d1-3B-GGUF \
  d1-3B-Q4_K_M.gguf --local-dir ./models

# Run a packed-batch decision pass on RTX 4090
./build/bin/llama-batched \
  -m ./models/d1-3B-Q4_K_M.gguf \
  --n-gpu-layers 32 \
  --batch-size 64 \
  --ctx-size 512 \
  -p "Is this email PII? yes/no"

The --n-gpu-layers 32 flag is the offload control — set it to the model depth to put every transformer layer on GPU. For the d1-omni-600M, lower the layer count to match the smaller depth. The --batch-size 64 flag is what unlocks the 475 decisions/sec throughput on RTX 4090 — a single-state batch is the 8 ms number; packed batches amortize weight loads across requests.

Caveats you need to know

License. d1-3B and d1-omni-600M ship under Liquid’s lfm1.0 license — not OSI-approved. Commercial use conditions apply. Read the license text before shipping a product on top of these weights. (Liquid AI license)

Loading requires trust_remote_code=True. The d1 architecture is not yet in transformers’ mainline model registry, so from_pretrained() will execute code from the Liquid repo. Pin a revision (revision="<commit-sha>") before loading in production — without a pin, a future commit to the repo changes what code runs on your machine. Required dependencies: transformers>=5.14, torch, torchvision, pillow (for the omni variant). (Liquid AI blog)

Vision and audio benchmarks are limited. The vision split on Decision Index v0.3 is a private benchmark — not reproducible outside Liquid. Liquid’s own docs call the audio decision benchmarks “currently an open problem.” If you’re evaluating d1-omni-600M for a vision or audio routing task, you’ll need to build your own eval set. (HF blog)

The 8 ms RTX 4090 number is one warm call, not throughput. AI Weekly’s digest noted the measurement methodology: bf16 + CUDA graphs + median of 20 warm calls. That’s single warm requests only, not token throughput. Burst traffic, cold-cache requests, or first-call compilation may be materially slower. If you’re sizing for production, measure your own p50/p95/p99 cold-cache latency — Liquid’s table is the best-case number. (AI Weekly)

What this means in practice

If you’re building a routing/guardrail layer today, you have three options: call a hosted decision API (Perplexity, OpenAI), run a quantized larger model on your own GPU (GEV-26B-Decide at 17 GiB), or run d1-3B locally with full GPU offload on commodity hardware.

The d1-3B path is the first one that fits on a Jetson Orin Nano at 50 ms. The decision layer is no longer a datacenter concern. It’s an edge concern.

The d1-omni-600M is the first multimodal decision model with open weights. If you’re routing on image or audio inputs, the only open-weight option at sub-1B parameters is now Liquid’s. Whether the multimodal quality holds in third-party reproduction is the open question.

What to verify when you load it:

  • The 82.9 mean benchmark. Reproduce on your own eval set before shipping.
  • The latency numbers. Measure on your hardware with your prompt distribution — single warm calls vs cold cache, single state vs packed batch.
  • The license. lfm1.0 is not OSI-approved; commercial conditions apply.
  • The repo code. trust_remote_code=True executes remote code; pin a revision before production.

Sources