No cloud, no GPUs, no problem: Liquid AI shrinks the AI agent down to Raspberry Pi size

No cloud, no GPUs, no problem: Liquid AI shrinks the AI agent down to Raspberry Pi size

LFM2.5-2.6B trades frontier-scale ambition for a bet that the future of enterprise AI runs locally, cheaply, and everywhere at once.

Written by OutOfToken AI

August 10, 2026 · 5 min read · Synthesized from reporting by VentureBeat · How this works

AI Likely Accurate · 7/10

Liquid AI, the startup founded in 2023 by former MIT computer scientists, has released LFM2.5-2.6B, an open-weight language model built specifically for agentic work rather than chatbot conversation. The headline claim is stark: it runs entirely on local hardware, from laptops and smartphones down to a Raspberry Pi, without touching the cloud or requiring a GPU.

A model built for tasks, not chat

At 2.6 billion parameters with a 128,000-token context window and native tool calling, LFM2.5-2.6B is designed for high-volume, well-defined agentic work — tool calling, document management, calendar automation, and always-on background routines. Liquid AI is explicit that coding-heavy work still belongs with larger cloud models.

CPU-first by design

Maxime Labonne, Liquid AI's head of post-training, told VentureBeat the model runs 'very, very well' on CPUs, with the underlying LFM2 architecture engineered around real-world CPU performance rather than GPU benchmarks. The company points to Raspberry Pi demos as proof, alongside a mobile app called Apollo that lets users try the models directly on their phones.

""I think the best example is a Raspberry Pi," Labonne said. "We have a lot of demos that show that actually, it works pretty fast on the Raspberry Pi.""

From chatbots to agent harnesses

Liquid AI built its training pipeline around the assumption that models are increasingly consumed through agent frameworks like OpenClaw and Hermes Agent, not chat interfaces. The company pretrained on roughly 34 trillion tokens and ran a four-stage post-training process — supervised fine-tuning, teacher specialization, multi-domain distillation, and agentic reinforcement learning inside real production harnesses. Labonne called the resulting broad capability gains, including in coding, a 'happy accident' of the new pipeline.

Built its own harness, too

Liquid AI didn't stop at the model — it built a custom agent harness and demonstrated it running entirely on a phone, planning and calling tools on-device. Labonne says most existing harnesses wait passively for prompts, while Liquid AI wants proactive agents that monitor context, like a calendar, and act without being asked.

Benchmarks against Gemma, Qwen, and DeepSeek

Liquid AI's own comparisons pit the model against Google's Gemma 4 and Alibaba's Qwen3.5 small models, claiming leadership on instruction-following and near-parity on tool-use and agentic benchmarks against models several times its size. A separate third-party test by Atomic Chat reportedly found LFM2.5-2.6B completing tool-calling tasks faster than DeepSeek-V4-Flash, a 284-billion-parameter model — though these figures are vendor- or third-party-reported and haven't been independently verified.

The licensing catch

The model ships under the LFM Open License v1.0, free for commercial use by organizations under $10 million in annual revenue; larger companies need a separate commercial agreement with Liquid AI. Labonne framed licensing as necessary to fund future model development, though he admitted enforcement is loose in practice — the ask, he said, is simply that bigger companies get in touch.

Liquid AI isn't chasing frontier benchmarks — it's betting that latency, privacy, and near-zero inference cost matter more to a large swath of enterprise use cases than raw model scale. A new partnership with MacPaw to build on-device AI for Mac suggests real commercial traction behind that bet. Whether small, task-specific agents become a durable enterprise category will hinge less on benchmark charts than on how reliably they perform once deployed at the edge, far from any GPU cluster.

Editorial Note

The research (primarily Source 1, the original VentureBeat article) corroborates all major factual claims about the model's specifications, Liquid AI's founding, and licensing terms. However, the research itself acknowledges that performance benchmarks are vendor-reported and unverified. The article appropriately relays company claims without independently validating real-world performance or the claimed advantages over competing models.

Claim Tracker

AI-assessed

VerifiedLiquid AI was formed in 2023 by former MIT computer scientists

Source 1 (VentureBeat) confirms: 'the AI startup Liquid, formed in 2023 by former MIT computer scientists'

VerifiedLFM2.5-2.6B contains 2.6 billion parameters and supports a 128,000-token context window

Source 1 (VentureBeat) states: 'LFM2.5-2.6B contains 2.6 billion parameters, supports a 128,000-token context window, and includes native tool calling'

VerifiedThe model was pretrained on approximately 34 trillion tokens

Source 1 (VentureBeat) confirms: 'The model is pretrained on approximately 34 trillion tokens'

VerifiedLiquid AI released the model under LFM Open License v1.0 which permits commercial use for organizations with less than $10 million in annual revenue

Source 1 (VentureBeat) states: 'LFM2.5-2.6B is distributed under the LFM Open License v1.0, which permits use, modification, and redistribution — including commercial use — for organizations with less than $10 million in annual revenue'

UnverifiedCompany-reported measurements indicate decoding throughput of approximately 220 tokens per second on an Apple M5 Max

Source 1 mentions these figures exist and explicitly notes 'These figures are vendor benchmarks and have not been independently verified,' but the research provided does not independently confirm these specific throughput numbers

Ask AI about this story

// discussion

sign in to join the discussion