Microsoft has released 'MAI-Image-2.5-Pro,' an AI-powered image generation tool, integrating its proprietary AI into Office products and other services to reduce reliance on third-party solutions.



Microsoft has released ' MAI-Image-2.5-Pro, ' an AI-powered image generator. It supports image generation from text and image editing.

Introducing MAI-Image-2.5-Pro and MAI-Voice-2-Flash | Microsoft AI

https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/

MAI-Image-2.5-Pro | AI Model Catalog | Microsoft Foundry Models
https://ai.azure.com/catalog/models/MAI-Image-2.5-Pro

MAI-Image-2.5-Pro is positioned as a quality-focused version of MAI-Image-2.5 , which was released in May 2026. Examples of images generated by MAI-Image-2.5-Pro are shown below. Its strengths include 'high-definition portraits,' 'accurate text rendering,' 'composing high-quality scenes from vague instructions,' and 'reproducing the reflection and weight of materials.'



MAI-Image-2.5-Pro supports text input of up to 32,000 tokens, as well as JPEG and PNG image input. The maximum output image resolution is 1,048,576 pixels (equivalent to 1024 x 1024 pixels). The input prompts are optimized for English, and the text rendering is also limited to English only.

The API charges per million tokens are $5 (approximately 820 yen) for text input, $8 (approximately 1310 yen) for image input, and $106 (approximately 17,360 yen) for image output.

Microsoft is highlighting that it has already integrated MAI-Image-2.5, which was released in 2026, into its products. For example, its image generation service Bing Image Creator initially used OpenAI's DALL-E series, but as of the time of writing, MAI-Image-2.5 is the default model.



PowerPoint also allows for image generation and editing using MAI-Image-2.5. According to Microsoft, this can reduce costs by up to 84% compared to OpenAI's GPT-Image-2.



Microsoft also released a public preview version of its speech synthesis AI, 'MAI-Voice-2 Flash,' which it had announced it would be developing in June 2026. MAI-Voice-2-Flash can generate up to 45 seconds of audio from text input. Supported languages include English, Italian, Spanish (Mexico), Hindi, English (Australia), French, German, Portuguese (Brazil), Korean, Portuguese (Portugal), Spanish (Spain), Simplified Chinese, Turkish, Russian, Thai, Dutch, Romanian, and Hungarian.

MAI-Voice-2-Flash | AI Model Catalog | Microsoft Foundry Models
https://ai.azure.com/catalog/models/MAI-Voice-2-Flash

in AI, Posted by log1o_hf