Native
Transformers, vLLM, SGLang, or fine-tuning
Use the exact LiquidAI/LFM2.5-2.6B ID.
Liquid AI model record · Source checked August 5, 2026
Liquid AI publishes a 2.69B dense, text-only checkpoint with 131,072-token context and native tool calling. It ships in native, GGUF, MLX, and ONNX formats. That record supports local evaluation—not a promise that every phone, harness, task, or Flowith account is ready for it.
LiquidAI/LFM2.5-2.6B, post-trained for agentic workloads
2.69B-parameter dense, text-only LFM2.5 model
131,072 tokens in the official model card
Native weights, GGUF, MLX, and ONNX
Tool use, data extraction, RAG, and long-context workflows
Liquid AI does not recommend it for agentic coding or knowledge-heavy tasks
Not established by the provider release; verify the live workspace separately
Format decision
Transformers, vLLM, SGLang, or fine-tuning
Use the exact LiquidAI/LFM2.5-2.6B ID.
llama.cpp and compatible local applications
Choose a quantization only after measuring quality and memory.
Apple Silicon local inference
Measure unified-memory pressure with the intended context.
Cross-platform edge and hardware-accelerated deployment
Validate the exported graph and target execution provider.
Choose a checkpoint, runtime, device budget, tool boundary, and evaluation gate.
Start with Transformers, GGUF, or MLX and verify a reproducible prompt contract.
Use the official checkpoint names and pick by adaptation versus ready-to-test agent behavior.
Map local CPU, Apple Silicon, cross-platform edge, or GPU serving to the published format.
Budget weights, KV cache, runtime overhead, tools, and safety margin before claiming fit.
Validate schemas, permissions, idempotency, traces, and human approval in the target harness.