Liquid AI LFM2.5 Tutorial: A 1.2B Edge AI Model That Runs on Your Phone
Liquid AI LFM2.5 is a 1.2B lightweight model for edge devices. Runs on phones, supports multimodal, with Ollama & vLLM deployment guide.
Published 264 days ago. Content may be outdated.
What is LFM2.5?
If you’ve been looking for an AI model that runs smoothly on phones or edge devices, Liquid AI’s LFM2.5 might be exactly what you need.
It’s a model family specifically optimized for local deployment. The smallest version has only 1.2B parameters, but it’s surprisingly capable. Built on the LFM2 architecture with extended pre-training and reinforcement learning, it’s both fast and resource-efficient.
What Models Are Available?
LFM2.5 isn’t just one model—it’s a whole family covering text, vision, and audio:
| Model | Parameters | Best For |
|---|---|---|
| LFM2.5-1.2B-Instruct | 1.2B | Daily chat, instruction following—first choice for most use cases |
| LFM2.5-1.2B-Base | 1.2B | Pre-trained base model, pick this if you want to fine-tune |
| LFM2.5-1.2B-JP | 1.2B | Japanese-optimized model |
| LFM2.5-VL-1.6B | 1.6B | Vision-language model that can see images |
| LFM2.5-Audio-1.5B | 1.5B | Supports both speech input and output |
In short: use Instruct for chatting, Base for fine-tuning, and pick the multimodal versions if you need vision or audio.
Flexible Deployment Formats
This is pretty thoughtful—LFM2.5 comes in several formats that work with most inference frameworks:
- HuggingFace Native: Works directly with Transformers or vLLM
- GGUF Quantized: Great for llama.cpp, LM Studio, Ollama users—runs on CPU too
- ONNX Format: Use this for phones and embedded devices
- MLX Format: Mac exclusive, runs smoothly on M-series chips
How to Run It?
Option 1: Using Transformers (Most Universal)
from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer
model_id = "LiquidAI/LFM2.5-1.2B-Instruct"
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
dtype="bfloat16",
# attn_implementation="flash_attention_2" # Enable if your GPU supports it
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)
prompt = "What is edge computing?"
input_ids = tokenizer.apply_chat_template(
[{"role": "user", "content": prompt}],
add_generation_prompt=True,
return_tensors="pt",
tokenize=True,
).to(model.device)
output = model.generate(
input_ids,
do_sample=True,
temperature=0.3,
min_p=0.15,
repetition_penalty=1.05,
max_new_tokens=512,
streamer=streamer,
)
Option 2: Using Ollama (Easiest)
One command and you’re done—perfect for quick testing:
ollama run hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF
Option 3: Using vLLM (Recommended for Production)
If you need high throughput, vLLM is the better choice:
vllm serve LiquidAI/LFM2.5-1.2B-Instruct
Want to Fine-tune? It’s Supported
LFM2.5 is fine-tuning friendly, with several official approaches:
- SFT + Unsloth: Supervised fine-tuning with LoRA, fast
- SFT + TRL: HuggingFace’s official library, great ecosystem
- DPO + TRL: Try this for preference alignment
The Base model is especially suited for:
- Training language-specific versions (Chinese, Japanese, etc.)
- Vertical domain adaptation (medical, legal, finance)
- Training on your private data
- Experimenting with new training methods
Task-Specific Nano Models
Liquid AI also released a series of task-specific models called “Nanos”—smaller parameters, more focused:
| Model | What It Does |
|---|---|
| LFM2-1.2B-Extract | Extracts structured info from documents, outputs JSON |
| LFM2-350M-ENJP-MT | Japanese-English bidirectional translation, very fast |
| LFM2-1.2B-RAG | Q&A model specifically for RAG systems |
| LFM2-1.2B-Tool | Optimized for tool calling, more accurate Function Calling |
| LFM2-350M-Math | Small but capable math reasoning model |
| LFM2-ColBERT-350M | For retrieval and reranking |
| LFM2-2.6B-Transcript | Meeting transcript summarization |
These small models have fewer parameters but perform well on specific tasks, and they’re faster and more resource-efficient.
Summary
LFM2.5 has a clear positioning: it’s for developers who want to run AI locally or on edge devices. The 1.2B parameter size can run on phones, and multimodal support is fairly comprehensive. If you’re building edge AI applications or looking for a lightweight model to fine-tune, give it a try.
Related Links
More Articles