DeepSeek has released 'DeepSeek-V4-Flash-Vision-Exp,' an open model with image recognition capabilities, achieving performance equivalent to Claude Opus 4.8 in tasks including image recognition.



Chinese company DeepSeek has released ' DeepSeek-V4-Flash-Vision-Exp ' as an open model. DeepSeek-V4-Flash-Vision-Exp is the first multimodal model in the DeepSeek-V4 series to support image input and has achieved benchmark scores equivalent to Claude Opus 4.8 in tasks including image recognition.




DeepSeek-V4-Flash-Vision-Exp is a multimodal model with 305 billion parameters. It is based on the DeepSeek-V4-Flash architecture, with an image processing module incorporated and image recognition capabilities added through post-training.

The benchmark scores for 'DeepSeek-V4-Flash-Vision-Exp', 'DeepSeek-V4-Flash-0731', and 'Claude Opus 4.8' are as follows. DeepSeek-V4-Flash-Vision-Exp successfully added image recognition capabilities while maintaining text processing performance equal to or better than DeepSeek-V4-Flash-0731, and recorded a score equivalent to Claude Opus 4.8 in tests that included image recognition tasks.



DeepSeek-V4-Flash-Vision-Exp is released as an open model and can be downloaded from Hugging Face and ModelScope. It is licensed under the MIT License.

deepseek-ai/DeepSeek-V4-Flash-Vision-Exp · Hugging Face
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp



DeepSeek-V4-Flash-Vision-Exp · Model Center
https://modelscope.cn/models/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp



DeepSeek-V4-Flash-Vision-Exp is also available via API. The supported image formats are JPEG, PNG, GIF, and WebP. The maximum resolution for input images is 8192 pixels on the longest side; if you input more than 15 images, the longest side is limited to 4096 pixels. Also, large images are resized to approximately 800x800 pixels while maintaining their aspect ratio, so even with large images, the number of tokens per image will be limited to around 384. The API documentation is available at the following link.

Vision | DeepSeek API Docs
https://api-docs.deepseek.com/guides/vision



DeepSeek's API fees vary between peak and off-peak hours. Peak hours are 'Monday to Friday, 1am to 4am and 6am to 10am UTC,' during which the DeepSeek-V4-Flash-Vision-Exp fee per million tokens is $0.44 (approximately ¥70.55) for input, $0.014 (approximately ¥2.24) for cached input, and $1.32 (approximately ¥211.65) for output. Off-peak fees are half of peak fees.

in AI, Posted by log1o_hf