Multimodal open d1 decision models for the edge

Liquid AI released two open decision models in its d1 family: d1-3B and d1-omni-600M, the latter described as experimental.
According to the company, d1-3B scores 48.57 on the Decision Index 0.2.1, which it says is the best result for a decision model under 10B parameters, ahead of every 4B and 9B model and of Decider 35B-A3B at 47.11. d1-3B supports text and images. d1-omni-600M supports either text and images or text and audio.
On speed, Liquid AI says d1-3B answers a question in 16 ms on an NVIDIA Jetson AGX Thor, 26 ms on a Jetson AGX Orin, and 50 ms on a Jetson Orin Nano. Three questions take about 1.3x the time of one, with the AGX Thor going from 16 ms to 20 ms. On GPU, the model answers a question in under 10 ms and processes a 384px image in under 18 ms on both tested platforms. No speed figures were reported for d1-omni-600M because it is an early research release.
The models are built on Liquid Foundation Models. Unlike generative models, they do not produce tokens but answer in a single forward pass. d1-3B is trained from LFM2.5-VL-3B, a decoder-only VLM accepting text and images. d1-omni-600M is trained from LFM2.5-Encoder-350M, a bidirectional encoder, and adds vision and audio encoders to handle all three modalities.
Both were benchmarked on seven public datasets covering reading comprehension, toxicity detection, intent classification, medical QA, and cross-lingual understanding. Liquid AI reports d1-3B at a mean score of 82.9, above Decider 4B, and d1-omni-600M at 78.4, surpassing Decider 2B at 77.1 with a quarter of the parameters. The company says it validated that d1-3B retains the vision capabilities of its backbone and that d1-omni-600M handles all three modalities, but reports no vision or audio benchmarks because the Decision Index v0.3 includes only a private vision split and audio decision benchmarks are an open problem.
Both models are open-weight and available on Hugging Face. Demos are available in a Hugging Face Space.
Based on reporting from the original publisher. Visit the source for full context and later updates.