How To Deploy Multimodal Open D1 Decision Models At The Edge
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How To Deploy Multimodal Open D1 Decision Models At The Edge on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Liquid AI released two open-weight models designed to return structured decisions in one forward pass: d1-3B and experimental d1-omni-600M. The company reports strong scores on seven public datasets and response times below 50 milliseconds for d1-3B on three Jetson devices, but independent replication and detailed vision and audio results are not provided.

Liquid AI has released two open-weight models, d1-3B and d1-omni-600M, built to return structured decisions in a single forward pass rather than generate a conventional sequence of text, as described in the original analysis. The company is positioning them for tasks such as classifying requests and assessing urgency, including on edge devices; its performance figures are company-reported and have not been independently replicated in the supplied release information.

The models are intended for tasks where a system needs to produce a decision or structured response, rather than a long-form answer, a focus also explored in recent coverage of open AI models. Liquid AI gives examples including routing a customer request to a team, judging its urgency and answering a question about an image. Both sets of weights are available on Hugging Face, according to the company.

d1-3B is based on Liquid AI’s LFM2.5-VL-3B vision-language model and accepts text and images, a multimodal approach also seen in other open-source AI models. The company reports a score of 48.57 on Decision Index 0.2.1. Its smaller model, d1-omni-600M, is based on the LFM2.5-Encoder-350M bidirectional encoder with added vision and audio encoders. It can process text paired with an image or text paired with audio. Liquid AI describes it as an early research release still under development.

On three NVIDIA Jetson devices, Liquid AI reports that d1-3B answered one question in 16 milliseconds on Jetson AGX Thor, 26 milliseconds on Jetson AGX Orin 64 GB and 50 milliseconds on Jetson Orin Nano. The company also reports 8 milliseconds per question on an NVIDIA RTX 4090 and 9 milliseconds on an AMD MI325X. These figures describe the company’s tests, not guaranteed performance in other applications or configurations.

At a glance
announcementWhen: Released; date not specified in the sup…
The developmentLiquid AI has made two open-weight decision models available, including a text-and-image model and a smaller experimental model with image and audio inputs.
At a glance
announcementWhen: Released in 2026; available on Hugging…
The developmentLiquid AI released d1-3B and experimental d1-omni-600M, two open-weight models designed for fast, structured decisions from text and visual or audio inputs.

Where Single-Pass Decisions Could Fit

The release targets a practical deployment trade-off: some products need a model to make a bounded decision quickly, but do not need a general-purpose text generator. A structured output for routing, classification or urgency scoring could suit systems that process requests close to where data is collected, where network delay or hardware limits affect product design.

Liquid AI’s reported Jetson timings make the edge-use case concrete, but they are not a complete measure of deployment readiness. Response time alone does not establish whether a model makes the right decision, handles ambiguous input reliably or meets a product’s safety requirements. Developers would need to test the models on their own workloads, devices and acceptable error thresholds before relying on them.

The smaller model may also interest teams working under memory or compute constraints. Liquid AI reports that d1-omni-600M scored a mean of 78.4 across its selected datasets, compared with 77.1 for the listed Decider 2B result. That company-run comparison is limited to the reported evaluation and does not establish that the smaller model will outperform alternatives on other tasks.

Amazon

NVIDIA Jetson AGX Thor developer kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmarks Behind the Release

Liquid AI reports results across seven public datasets covering reading comprehension, toxicity detection, intent classification, medical question answering and cross-lingual understanding. The named datasets include SQuAD 2.0, Civil Comments, MASSIVE intent, PubMedQA, BoolQ, XNLI and PAWS-X. The company gives mean scores of 82.9 for d1-3B and 78.4 for d1-omni-600M, compared with listed scores of 81.1 for Decider 4B and 77.1 for Decider 2B.

The dataset-level results are not uniformly higher: the source material says d1-3B scores below Decider 4B on BoolQ, MASSIVE intent and XNLI. A mean across seven selected datasets summarizes that test set; it does not show performance on every kind of decision task. Liquid AI also says Decision Index version 0.3 includes only a private vision split, while audio decision benchmarks remain an open problem.

The models build on Liquid AI’s Liquid Foundation Models. The release also points to demonstrations in the company’s System One Arcade Hugging Face Space. Its instructions specify Transformers version 5.14 or later and loading the models with supplied code enabled.

““Best decision model under 10B on the Decision Index 0.2.1””

— Liquid AI

Amazon

edge AI decision model hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evidence Still Missing for Deployment

The release information provided does not include independent evaluations, confidence intervals or enough methodological detail to establish how closely the benchmark conditions match a particular deployment. The seven public datasets cover selected capabilities, not every decision task, and their average scores do not establish reliability or safety in production.

Liquid AI says it checked that d1-3B retained vision capabilities from its underlying vision-language model and that d1-omni-600M handled its supported modalities. However, the release provides no vision or audio benchmark scores. It also reports no speed measurements for d1-omni-600M, which the company identifies as experimental and still under development.

Other operational questions remain unanswered in the supplied material: how the models respond to ambiguous inputs, how often a human should review their decisions, and how performance changes with different input sizes, software setups or production workloads. The reported timings and scores should not be treated as answers to those questions.

Amazon

multimodal AI decision models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing Models on Local Workloads

The next practical step for developers is to evaluate the open weights against their own tasks and hardware. Liquid AI points users to its Hugging Face model releases and demos, and specifies the software requirements for loading them. Local testing can establish whether reported latency and task performance carry over to a particular application; the release does not give a timeline for further evaluations or for changes to the experimental omni model.

More evidence would be needed to compare performance across independent test setups, especially for image and audio decisions. Until such results are available, d1-3B’s figures remain company-reported measurements, while d1-omni-600M should be treated as an early research model rather than a validated production option.

Amazon

vision-language AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did Liquid AI release?

Liquid AI released d1-3B and d1-omni-600M, open-weight models designed to return structured decisions in a single forward pass.

What inputs can the models process?

d1-3B accepts text and images. Liquid AI says d1-omni-600M can process text paired with an image or text paired with audio, and describes that model as an early research release.

How fast is d1-3B on edge devices?

Liquid AI reports one-question response times of 16 milliseconds on Jetson AGX Thor, 26 milliseconds on Jetson AGX Orin 64 GB and 50 milliseconds on Jetson Orin Nano. These are company-reported test results, not performance guarantees.

Have the benchmark results been independently verified?

The supplied release information does not provide independent replication. Its mean scores come from Liquid AI’s evaluation of seven public datasets and should be read within that limited scope.

Where can developers get the models?

Liquid AI says both models are available as open weights on Hugging Face and points to demos in its System One Arcade Hugging Face Space. Its instructions specify Transformers version 5.14 or later and require the supplied code to be enabled when loading the models.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Atari Surges In Global Coverage

Search interest in Atari spikes with 13 mentions this week, indicating rising global attention amid unconfirmed developments, prompting industry speculation.

The Microduck Toy Masks A Robust Open Stack AI Engine

Hugging Face introduces Microduck, a small, affordable robot with open reinforcement learning platform, signaling a shift in accessible embodied AI.

Is NeoMME The Best Multimodal-native And Multilingual AI Encoder Available?

Hugging Face releases NeoMME, a multimodal and multilingual encoder promising efficiency and performance, but independent validation is pending.

OpenAI’s Jalapeño Chip: Leading Or Lagging In The AI Race?

OpenAI releases initial results for Jalapeño, its custom inference chip, showing strong efficiency against NVIDIA but with limitations on deployment and scope.