Artificial Analysis has published the results of its measurements of the performance of various AIs on iPhones, measuring performance in a realistic environment by limiting context length and response time.

Some locally executable AI models are small enough to run on smartphones. Artificial Analysis, a company that analyzes AI performance, has now released benchmark results focusing on 'how they perform on the iPhone 17 Pro.'
Intelligence at pocket scale: Benchmarking small models and mobile phones | Artificial Analysis
Introducing Pipette: A benchmarking suite for on-device intelligence — Blog — Liquid AI
https://www.liquid.ai/blog/pipette-on-device-ai-benchmarking-by-liquid-ai
Among the open models that can run locally, some are mega-models that require high-performance data center-level hardware, while others are small models that run on smartphones. However, many benchmark tests are optimized for evaluating the performance of mega-models, and small models may fail to complete the tests and receive a score of '0'. In addition, inference stacks that can run on smartphones are less sophisticated than those for data centers, so comparing mega-models and small models on the same infrastructure and benchmark tests will yield results that do not reflect reality.
Therefore, Artificial Analysis collaborated with AI development company Liquid AI to build a system to measure the performance of small models when run on smartphones. Of the multiple tests, Artificial Analysis developed the system for intelligence performance testing, while Liquid AI developed the system for inference performance testing. Five benchmark tests were used: BFCL , IFBench , AA-Omniscience , GPQA Diamond , and MATH-500 .
Artificial Analysis runs multiple AI models on an iPhone 17 Pro and compiles the results. The models tested are limited to 'AI models that fit within 8GB in size when quantized to 4 bits,' such as 'Gemma 4 E2B' and 'LFM2-2.6B-Exp.'

The average benchmark scores when the context length was limited to 16K tokens are as follows. The highest scores were achieved by 'Nanbeige4.2-3B' and 'LFM2.5-2.6B'.

The average scores when response times were limited to a maximum of 1 minute are as follows: 'LFM2.5-8B-A1B' came out on top, followed by 'LFM2-2.6B-Exp', 'Gemma 4 E4B (Non-reasoning)', and 'Granite 4.1 8B'.

The graph below shows the 'seconds taken to output 256 tokens' on the horizontal axis and the 'average benchmark score when the context length is limited to 16K tokens' on the vertical axis. 'Nanbeige4.2-3B' had a benchmark score equivalent to 'LFM2.5-2.6B,' but it can be seen that the response time was longer.

If we plot peak memory usage on the horizontal axis, it looks something like this.

When the number of output tokens is plotted on the horizontal axis, it can be seen that the token efficiency of 'Qwen3.5 9B (Reasoning)' and 'Qwen3.5 4B (Reasoning)' is relatively low.

Artificial Analysis's future update plans include 'the continuous addition of new models' and 'model comparisons based on memory usage rather than a single quantization accuracy.'
Liquid AI has named the system used for performance measurement 'Pipette,' and has made its source code publicly available on GitHub. They have also released iOS and Android apps that allow users to run benchmark tests using Pipette.
Pipette by Liquid app - App Store
https://apps.apple.com/jp/app/pipette-by-liquid/id6772314671
Pipette - App on Google Play
https://play.google.com/store/apps/details?id=ai.liquid.pipette
- Continued
I tried out 'Pipette,' a benchmark app that can measure the AI performance of your smartphone - GIGAZINE

Related Posts:
in AI, Posted by log1o_hf






