Here are the benchmark results for Gemini 3.6 Flash.



On July 21, 2026, Google announced '

Gemini 3.6 Flash, ' a new model in its 'Gemini Flash' series of AI models that prioritize high-speed and low-cost processing. Numerous benchmark results of Gemini 3.6 Flash from various third-party organizations have been reported.

Gemini 3.6 Flash - Intelligence, Performance & Price Analysis
https://artificialanalysis.ai/models/gemini-3-6-flash

On July 21, 2026, Google announced three new models in its 'Gemini Flash' series: 'Gemini 3.6 Flash,' 'Gemini 3.5 Flash Lite,' and 'Gemini 3.5 Flash Cyber.' Gemini 3.6 Flash is positioned as a general-purpose model that improves coding and knowledge work, Gemini 3.5 Flash Lite is a lightweight model that performs large amounts of processing at low cost, and Gemini 3.5 Flash Cyber is a cybersecurity model specifically designed for vulnerability detection and remediation.

Google announces 'Gemini 3.6 Flash,' 'Gemini 3.5 Flash Lite,' and 'Gemini 3.5 Flash Cyber,' and is also developing Gemini 4 - GIGAZINE



Gemini 3.6 Flash is a model that enhances coding, knowledge work, and multimodal processing based on feedback from developers and users of Gemini 3.5 Flash, which was announced in May 2026. According to Google, Gemini 3.6 Flash reduces the number of output tokens by 17% compared to Gemini 3.5 Flash in the Artificial Analysis Index, and also requires fewer inference steps and tool calls when performing multi-stage tasks, enabling more efficient execution of long-running agent processes.

Artificial Analysis, a website that compares and analyzes the performance and characteristics of AI models, has recorded benchmark results for Gemini 3.6 Flash across various metrics. The following graph highlights its intelligence, and while Gemini 3.6 Flash, shown in green, does not quite reach the level of top-of-the-line models such as Anthropic's ' Claude Fable 5 ' or OpenAI's ' GPT-5.6 Sol ,' it achieves scores comparable to other cutting-edge models.



The following are intelligence ratings independently measured by Artificial Analysis. The top left shows the agent performing real-world tasks, the top right shows tasks in the financial sector, the bottom left shows the agent coding and terminal use, and the bottom right shows coding. All of these scores are comparable to those of Claude's higher-end models, the GPT-5.6 series and GPT-5.5 series, as well as Chinese open models such as the 'Kimi K3' and 'Grok 4.5 high.'



Furthermore, the following is a comparison of the Gemini 3.6 Flash's score on the 'AA-Briefcase' agent-based knowledge work benchmark developed by Artificial Analysis with other top-performing models. A higher number indicates better performance, and the Gemini 3.6 Flash scored '961'.



The AA-Omniscience Index, which measures the reliability of knowledge and the degree of hallucination, assigns positive scores to correct answers and negative scores to hallucinations . Scores range from -100 to 100, with 0 meaning an equal number of correct and incorrect answers. The Gemini 3.6 Flash scored '24' on the AA-Omniscience Index, which is comparable to many higher-end models.



The following graph shows the number of tokens output per second, with the Gemini 3.6 Flash recording a particularly excellent score.



In tests that recorded the time required to achieve a certain intelligence score, a lower number indicates better performance, and the Gemini 3.6 Flash scored '1.3,' placing it second only to the GPT-5.6 series.



Furthermore, in tests comparing the response of

HutchDB , a structured data storage and MCP service designed for AI agents, with Gemini 3.6 Flash, Gemini 3.5 Flash Lite, Claude Opus 4.8, and GPT-5.6 Terra, Gemini 3.6 Flash was shown to be 6 times faster and cost-effective than Claude Opus 4.8 when simply comparing speed and cost.



The latency, recorded in seconds until the first response token is received, is as follows, demonstrating that this model has excellent response speed.



The following is the weighted average cost per task, with the Gemini 3.6 Flash costing $0.50 (approximately 81.5 yen), placing it somewhere in the middle among the cutting-edge models.



Overall, Gemini 3.6 Flash scored '50 points' on the Artificial Analysis index, surpassing the median score of '31 points' for other inference models in the same price range. It excels particularly in token output speed, with Gemini 3.6 Flash outputting '303.6 tokens per second' compared to the median of '78.5 tokens per second' for other inference models in the same price range.

In addition, the results of a benchmark devised by engineer Simon Willison, 'Pelican on a Bicycle,' have also been

reported . The 'Pelican on a Bicycle' created by Gemini 3.6 Flash is as follows: While the shape of the bicycle and the figure riding it are visible, the pelican's beak is strangely distorted, and the pelican's rear end is not resting on the saddle but rather fused together, so it's not a very good result. However, since the previous model, Gemini 3.5 Flash Lite, could not generate either a pelican or a bicycle, it can be said that the performance has improved significantly. However, Willison points out that there are cases where the older model can output a pelican on a bicycle better, so a simple comparison as a benchmark is not possible.



Below is a pelican riding a bicycle, generated by the AI app builder

Playcode using Gemini 3.6 Flash. As in Mr. Wilson's test, the bicycle is shaped correctly and the pelican is shown pedaling, but the beak is oddly shaped, the fish it's holding is poorly done, and the saddle and rear end are fused together. However, according to Playcode, this output is 'only slightly different' from Claude Fable 5, meaning it achieved a similar level of output at about one-fifth the cost.



In

a benchmark comparing Gemini 3.6 Flash (high), Gemini 3.5 Flash (high), Gemini 3.6 Flash (medium), and Gemini 3.5 Flash (medium), Gemini 3.6 Flash scored slightly higher, but was also shown to be more expensive. Google claims that '3.6 Flash is more token-efficient,' but this test suggests that the token efficiency may not actually be that much better compared to 3.5 Flash.



in AI, Posted by log1e_dh