A high-speed and high-precision visual language model, 'Zamba2-VL,' has been released, developed with an architecture faster than Transformer.



AI development company

Zyphra has released ' Zamba2-VL ,' a visual language model (VLM). Zamba2-VL enables faster image recognition processing compared to models of similar size.

Zamba2-VL: Hybrid SSM Vision-Language Models
https://www.zyphra.com/our-work/zamba2-vl




Zamba2-VL is a VLM built on the 'SSM-Transformer' hybrid architecture, which combines the mainstream AI architecture 'Transformer' with the AI architecture ' Mamba2 ' announced in 2024. By adopting SSM-Transformer, it is said to be able to perform high-speed processing with the same quality as Transformer-based models of the same scale.




Zamba2-VL is available in three versions: 'Zamba2-VL-1.2B' with 2 billion parameters, 'Zamba2-VL-2.7B' with 2.7 billion parameters, and 'Zamba2-VL-7B' with 7 billion parameters. The graph below shows the time to output the first token on the horizontal axis and the average benchmark score on the vertical axis, demonstrating that Zamba2-VL can perform image recognition processing with higher accuracy compared to models with comparable speeds.



'Zamba2-VL-1.2B', 'Zamba2-VL-2.7B', and 'Zamba2-VL-7B' are released as open models and can be downloaded from the following link. The license is the Apache License 2.0 .

Zyphra/Zamba2-VL-1.2B · Hugging Face
https://huggingface.co/Zyphra/Zamba2-VL-1.2B

Zyphra/Zamba2-VL-2.7B · Hugging Face
https://huggingface.co/Zyphra/Zamba2-VL-2.7B

Zyphra/Zamba2-VL-7B · Hugging Face
https://huggingface.co/Zyphra/Zamba2-VL-7B

in AI, Posted by log1o_hf