Maple-Preview, an AI that runs on iPhones and boasts performance equivalent to Bonsai 27B while being 13 times faster, represents another step forward for local AI.

AI research firm
Maple-Preview Model Card — DeepGrove
https://deepgrove.ai/maple-preview
Maple-Preview is a MoE model with a total of 20.2 billion parameters and 1.49 billion active parameters.
Many AI researchers are attempting to run AI models on devices with limited memory, but most use the method of 'quantizing existing AI models.' Quantization is a technique that can reduce the size of an AI model by lowering its computational precision, but DeepGrove points out that 'the approach of lowering the computational precision of existing models is fundamentally wrong. It unnecessarily limits both performance and efficiency,' and 'the accuracy of a model at runtime should match the accuracy at training.' Maple-Preview is designed to perform processing in ternary form from the design stage, resulting in a small and high-performance model.
The graph below shows the number of tokens output per second during decoding on the horizontal axis and the average benchmark score on the vertical axis. It is clear that Maple-Preview is faster and more powerful than other models designed for smartphones such as the Ternary Bonsai 27B and Gemma 4 E4B.

The graph below compares the benchmark scores of 'Maple-Preview', 'Qwen3.5 35B-A3B', 'Ternary Bonsai 27B', 'gpt-oss-20b', and 'GLM-4.7-Flash'.

Maple-Preview was reportedly able to accurately solve problems from the International Mathematical Olympiad.

The graph below shows the context length on the horizontal axis and the additional memory usage required to maintain the context on the vertical axis. Maple-Preview can keep memory usage low even when handling a large number of tokens.

The graph below compares the memory usage of each model with the memory usage including the context of 131,000 tokens. Maple-Preview consumes only 7.69GB of memory even when handling 131,000 tokens.

When the same prompt was entered into both Maple-Preview and Binary Bonsai 27B on an iPhone, Maple-Preview was found to be 13 times faster.

Maple-Preview is developed as an open model and is distributed at the following link. It is licensed under the MIT License.
deepgrove/maple-preview · Hugging Face

A demo site where you can actually use Maple-Preview is also available.
DeepGrove
https://chat.deepgrove.ai/

The following is the result of entering the question 'Should I put a screen protector on my smartphone? Is it okay not to? In the case of an iPhone 17?' on the demo site. It responded in Japanese, but the content was quite questionable, resulting in sentences that were both understandable and not.

As the name suggests, Maple-Preview is in the preview stage and will continue to be improved. DeepGrove also aims to build a system that 'tunes the AI model based on conversation content to optimize it for individual users.'
Related Posts:
in AI, Posted by log1o_hf







