A smartphone-compatible visual language model, 'LFM2.5-VL-3B,' has been released, supporting Japanese and usable for UI recognition and OCR.



Liquid AI, an AI company based in the United States, has announced its visual language model, ' LFM2.5-VL-3B .' LFM2.5-VL-3B is a lightweight VLM that has been confirmed to work on smartphones and is capable of document OCR and UI recognition. It is released as an open model and can be downloaded free of charge.

LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge — Blog — Liquid AI

https://www.liquid.ai/blog/lfm2-5-vl-3b

LFM2.5-VL-3B is a VLM built on the same core model as LFM2.5-2.6B . It has 3.1 billion parameters and uses less than 3.3GB of memory.



The following are the results of measuring the inference performance and visual recognition performance of 'LFM2.5-VL-3B(3.1B)', 'LFM2-VL-3B(3.1B)', 'gemma-4-E2B-it(5.1B)', 'gemma-4-E4B-it(8B)', 'InternVL 3.5 2B(2.4B)', 'InternVL 3.5 4B(4.7B)', 'Qwen3.5-2B(2.3B)', and 'Qwen3.5-4B(4.7B)'. LFM2.5-VL-3 recorded a higher score than relatively large models such as gemma-4-E4B-it and Qwen3.5-4B.



The LFM2.5-VL-3B supports NVIDIA, AMD, Apple, and Qualcomm processors and can run on laptops and smartphones. The following diagram shows the 'delay to output the first token (milliseconds)', 'decode speed (tok/s)', and 'memory usage (MB)' when inputting images and text into various models: a PC with an AMD Ryzen AI Max+ 395, a Mac with an Apple M5 Max, and a Galaxy S26 Ultra with a Qualcomm SoC. The LFM2.5-VL-3B can decode 20 tokens per second on a Galaxy S26 Ultra.



The graph below compares the time it takes for various models to output the first token when inputting images and text using NVIDIA's H100. The LFM2.5-VL-3B can process a task that inputs '256x256 pixel, 5-frame video' with a low delay of 34 milliseconds.



The supported languages are English, Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Spanish, Vietnamese, Thai, Indonesian, Hindi, Russian, and Polish. The following video shows how the LFM2.5-VL-3B performs OCR on documents.

LFM2.5-VL-3B: Document OCR and Layout Understanding in WebGPU - YouTube


It can also accurately recognize the screen UI of PCs and smartphones.

LFM2.5-VL-3B: Screen and UI Understanding in WebGPU - YouTube


LFM2.5-VL-3B is available at the following link. It already supports operation with 'llama.cpp', 'MLX', 'vLLM', 'SGLang', and 'ONNX', and can support everything from on-device processing to large-scale service deployments.

LiquidAI/LFM2.5-VL-3B · Hugging Face
https://huggingface.co/LiquidAI/LFM2.5-VL-3B

The documentation can also be found at the following link.

LFM2.5-VL-3B - Liquid Docs
https://docs.liquid.ai/lfm/models/lfm25-vl-3b

in AI, Posted by log1o_hf