StableLearn Logo

Search Content

News 3 min read

Liquid AI LFM2.5 Tutorial: A 1.2B Edge AI Model That Runs on Your Phone

Liquid AI LFM2.5 is a 1.2B lightweight model for edge devices. Runs on phones, supports multimodal, with Ollama & vLLM deployment guide.

Cover image for Liquid AI LFM2.5 Tutorial: A 1.2B Edge AI Model That Runs on Your Phone

Published 264 days ago. Content may be outdated.

What is LFM2.5?

If you’ve been looking for an AI model that runs smoothly on phones or edge devices, Liquid AI’s LFM2.5 might be exactly what you need.

It’s a model family specifically optimized for local deployment. The smallest version has only 1.2B parameters, but it’s surprisingly capable. Built on the LFM2 architecture with extended pre-training and reinforcement learning, it’s both fast and resource-efficient.

What Models Are Available?

LFM2.5 isn’t just one model—it’s a whole family covering text, vision, and audio:

ModelParametersBest For
LFM2.5-1.2B-Instruct1.2BDaily chat, instruction following—first choice for most use cases
LFM2.5-1.2B-Base1.2BPre-trained base model, pick this if you want to fine-tune
LFM2.5-1.2B-JP1.2BJapanese-optimized model
LFM2.5-VL-1.6B1.6BVision-language model that can see images
LFM2.5-Audio-1.5B1.5BSupports both speech input and output

In short: use Instruct for chatting, Base for fine-tuning, and pick the multimodal versions if you need vision or audio.

Flexible Deployment Formats

This is pretty thoughtful—LFM2.5 comes in several formats that work with most inference frameworks:

  • HuggingFace Native: Works directly with Transformers or vLLM
  • GGUF Quantized: Great for llama.cpp, LM Studio, Ollama users—runs on CPU too
  • ONNX Format: Use this for phones and embedded devices
  • MLX Format: Mac exclusive, runs smoothly on M-series chips

How to Run It?

Option 1: Using Transformers (Most Universal)

   from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer

model_id = "LiquidAI/LFM2.5-1.2B-Instruct"
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    dtype="bfloat16",
    # attn_implementation="flash_attention_2"  # Enable if your GPU supports it
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)

prompt = "What is edge computing?"

input_ids = tokenizer.apply_chat_template(
    [{"role": "user", "content": prompt}],
    add_generation_prompt=True,
    return_tensors="pt",
    tokenize=True,
).to(model.device)

output = model.generate(
    input_ids,
    do_sample=True,
    temperature=0.3,
    min_p=0.15,
    repetition_penalty=1.05,
    max_new_tokens=512,
    streamer=streamer,
)

Option 2: Using Ollama (Easiest)

One command and you’re done—perfect for quick testing:

   ollama run hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF

If you need high throughput, vLLM is the better choice:

   vllm serve LiquidAI/LFM2.5-1.2B-Instruct

Want to Fine-tune? It’s Supported

LFM2.5 is fine-tuning friendly, with several official approaches:

  • SFT + Unsloth: Supervised fine-tuning with LoRA, fast
  • SFT + TRL: HuggingFace’s official library, great ecosystem
  • DPO + TRL: Try this for preference alignment

The Base model is especially suited for:

  • Training language-specific versions (Chinese, Japanese, etc.)
  • Vertical domain adaptation (medical, legal, finance)
  • Training on your private data
  • Experimenting with new training methods

Task-Specific Nano Models

Liquid AI also released a series of task-specific models called “Nanos”—smaller parameters, more focused:

ModelWhat It Does
LFM2-1.2B-ExtractExtracts structured info from documents, outputs JSON
LFM2-350M-ENJP-MTJapanese-English bidirectional translation, very fast
LFM2-1.2B-RAGQ&A model specifically for RAG systems
LFM2-1.2B-ToolOptimized for tool calling, more accurate Function Calling
LFM2-350M-MathSmall but capable math reasoning model
LFM2-ColBERT-350MFor retrieval and reranking
LFM2-2.6B-TranscriptMeeting transcript summarization

These small models have fewer parameters but perform well on specific tasks, and they’re faster and more resource-efficient.

Summary

LFM2.5 has a clear positioning: it’s for developers who want to run AI locally or on edge devices. The 1.2B parameter size can run on phones, and multimodal support is fairly comprehensive. If you’re building edge AI applications or looking for a lightweight model to fine-tune, give it a try.

Share Article

More Articles