• Liquid AI
  • Decision Models
  • Agent Routing
  • Open Weights

Liquid AI d1 decision models: small choices inside a larger workflow

Liquid AI's open-weight d1-3B targets structured decisions in one forward pass. Learn where a small decision model fits and what the latency means.

Real frame from atomic.chat's original Snake decision demo shared by Liquid AI; the scoreboard is one creator run, not an owner benchmark.
Original atomic.chat demonstration frame. Vendor demo, not our benchmark.

TL;DR

  • Liquid AI released open-weight d1-3B on October 7, 2026, for structured decisions in one forward pass with zero output tokens.
  • Liquid reports 8 ms for one question on an RTX 4090, bf16, median of 20 runs, with compilation and CUDA graphs.
  • That figure is a model benchmark, not end-to-end agent latency.
  • The 86-versus-3 Snake result belongs to atomic.chat's particular RTX 3090 demo. It is not my benchmark and does not show NVIDIA funding.

Sometimes the useful AI answer is simply up, down, left or right. The complete Snake demo by atomic.chat, shared by Liquid AI, makes that idea tangible. The video ends at 86 apples versus 3 in that particular run. The creator's setup is an RTX 3090, and the result belongs to that demo.

Liquid AI's d1 release describes d1-3B as an open-weight decision model. The model card reports an 8 ms result for one question or request at a time on an NVIDIA RTX 4090, in bf16, as the median of 20 runs with compilation and CUDA graphs. Those are useful conditions to record. They also define what the number means.

It is not a universal response time. It is not the time for a browser, network, tool call, game loop or full agent workflow. NVIDIA's Jetson guide publishes separate measurements. Liquid's collaboration on those measurements does not establish NVIDIA funding.

Put the small model at a decision boundary

A larger model can plan a multi-step task. A small decision model can choose among a known set of next actions. The boundary needs to be typed:

  • one input shape;
  • one allowed decision set;
  • one fallback when the answer is invalid or uncertain;
  • one record of the decision and the evidence it used.

In an agent workflow, that could mean selecting a map, trend or table; choosing a tool; deciding whether a result needs a deeper review; or sending a bounded handoff to another worker. The model should not invent a new action outside the list. If the decision is unclear, the system should escalate instead of converting uncertainty into a random route.

This is where retained context matters. The decision model needs the current task state, not the entire history. The state should say what the user asked, what has already been approved, which outputs exist and what check must run next. A route change then has a clear contract.

Measure the complete path

I recently moved background workloads to an M3 Ultra Mac Studio so projects can continue while I manage them from my MacBook. With multiple agents and parallel projects, I am interested in the small decisions inside each workflow.

The first experiment should not ask whether d1 is “faster.” It should ask whether a bounded decision improves the completed task. Record:

  1. time to receive the decision;
  2. time to execute the selected action;
  3. invalid or escalated decisions;
  4. retries and human review;
  5. the final acceptance result.

Keep the model latency separate from network, tool and rendering time. Keep a creator demo separate from an owner run. A one-question median can guide a design, but it cannot certify a production route.

Decision model, lead model, and evidence

The larger model should keep the reasoning that needs broad context. The decision model should make the narrow choice. A deterministic function should calculate values when the task has a known formula. A reviewer or validator should check the result.

That division is similar to the question-shaped interface I explored in Tiersel. Jev can help choose a view; deterministic code calculates the numbers; a person can inspect the records. The components are separate, and each leaves a different kind of evidence.

Agent's receipt model is useful for the same reason. The task output, checks, route, time and qualified cost can be sealed together, while runs, artifacts and decisions keep their own source records. If the proof is missing or stale, the result stays Unverified. A fast decision without that trail is only a fast guess.

A modest next test

I would start with tool selection or view selection, where the legal choices are small and the acceptance check is clear. Run the same task with a fixed policy, a larger model and d1. Count complete workflows, not only model calls. Include the creator's demo conditions in the report, then replace them with the conditions of the machine and route that matter to you.

The useful promise of d1 is not that every agent should become tiny. It is that some repeated choices may deserve a small, inspectable component so the larger model can spend its context on judgment.

The source event is my October 8 LinkedIn post. The attached video is the complete original atomic.chat media with its branding and scoreboard intact.

See the Agent receipt example when you want to define the evidence a routed task should leave behind.

Turn the numbers into shipped work.

Agent runs these choices for you: a persistent AI worker with memory and rules, on your Claude and Codex subscriptions.