Sources
See it in action
Browse the models and styles behind stories like this one — free account, instant gallery.
Explore the catalogDiscuss this with
Pick a companion and get their take on this story

Theo turns AI news into things you can actually try in tonight's session.
Browse the models and styles behind stories like this one — free account, instant gallery.
Explore the catalogPick a companion and get their take on this story
Liquid AI has published open-weight d1 decision models designed to run multimodal inference directly on edge hardware — phones, laptops, and embedded chips — without sending data to a cloud server. For creators who want faster, private, local inference in their workflows, that is a meaningful architectural shift.
Most multimodal models are engineered for data-center GPUs and accessed through an API. Liquid AI's d1 models are built around a different priority: running a full decision loop — perceive inputs, reason, produce an output — on hardware with tight memory and compute budgets. According to the Hugging Face blog post, the d1 family uses Liquid AI's proprietary architecture (derived from Liquid Neural Networks) rather than a standard transformer backbone, which is part of how it achieves the smaller memory footprint.
The practical upshot: a creator running a local ComfyUI or Automatic1111 setup could, in principle, slot a d1 model into a workflow node that makes image-routing or captioning decisions — without that step hitting an external API or adding cloud latency.

Liquid AI's open d1 models process text, image, and structured data in a single inference pass on edge hardware.
Image: Hugging Face Blog
Open-weight is the key word here. You download the weights once, run them on your own hardware, and pay nothing per inference. For a creator who batch-processes hundreds of images — generating captions, routing outputs by content type, or running automated quality checks — eliminating per-call API fees adds up fast. Compare that to the subscription-gated model access that has become common elsewhere: Google recently restricted free Gemini users to its lightest Flash Lite tier, as covered in our report on Google's free plan changes.
Fine-tuning is also on the table. Because the weights are open, a creator working in a specific niche — architectural visualization, character-sheet generation, product photography — could fine-tune a d1 model on domain-specific data to make its routing or captioning decisions sharper for that exact use case.
Consider a character-art creator who generates 200 images in a batch and needs to automatically sort outputs by quality tier before upscaling. Today that usually means either manual review or an API call to a vision model for each image. A locally running d1 model could handle that classification step on the same machine doing the generation — no network hop, no cost per image, no data leaving the studio.
The setup isn't plug-and-play yet. You'll need to pull the weights from Hugging Face, confirm your hardware meets the memory requirements Liquid AI specifies, and wire the model into your pipeline manually. If you're already comfortable pulling models from the Hub and running inference scripts, the friction is low. If you're newer to local model deployment, the Charmloop guides cover the fundamentals of getting open-weight models running locally.
The d1 release is early-stage. Benchmark numbers against comparable edge models — Apple's on-device models, Qualcomm's edge AI stack, or Microsoft's Phi series — aren't yet widely available for direct comparison. Liquid AI's architecture is also less battle-tested in creative pipelines than transformer-based models that have years of community tooling around them. Expect some integration roughness before the broader ComfyUI and Automatic1111 communities build standardized nodes for it.
That said, open-weight multimodal models optimized for edge hardware are still scarce enough that the d1 family fills a real gap — particularly for creators who've been waiting for a locally-runnable option that handles more than just text. Explore what's available for your own image generation workflows while the community tooling catches up.