Thinking Machines Lab has released 'Inkling-Small,' which delivers performance equivalent to Inkling but in almost one-quarter the size.



AI startup Thinking Machines Lab has released '

Inkling- Small,' an open-weight model that achieves performance equivalent to 'Inkling,' an AI model with 975 billion parameters released just two weeks ago, but at approximately one-third to one-fourth the size.

Introducing Inkling-Small - Thinking Machines Lab
https://thinkingmachines.ai/news/inkling-small/



thinkingmachines/Inkling-Small · Hugging Face

https://huggingface.co/thinkingmachines/Inkling-Small

Thinking Machines debuts Inkling Small open source AI model nearing performance of predecessor at about 1/4 size | VentureBeat
https://venturebeat.com/technology/thinking-machines-debuts-inkling-small-open-source-ai-model-nearing-performance-of-predecessor-at-about-1-4-size

Inkling-Small is a Mixture-of-Experts (MoE) Transformer model with a total of 276 billion parameters and 12 billion active parameters, trained on an NVIDIA GB300 NVL72 .

Like Inkling, it features native inference capabilities for audio and images, variable thinking load, and a context window of up to 1 million tokens, and is said to deliver balanced performance across a wide range of benchmarks.

Compared to its predecessor, Inkling, Inklink-Small delivers comparable performance with significantly fewer computing resources.

The graph below shows the scores from 10 different benchmarks performed on five models: Inkling-Small, Inkling, DeepSeek V4 Flash, Gemini 3.5 Flash-Lite, and GPT 5.6 Luna. Inkling-Small and Inkling generally recorded similar scores, except for the SimpleQA Verified benchmark, where Inkling-Small lagged significantly behind the others.



According to Thinking Machines Lab, Inkling-Small performs at or above the level of Inkling in inference and agent tasks, and offers a good balance of performance and token count compared to open weight models in the same weight class.

Furthermore, just like Inkling, Inkling-Small has been developed specifically for voice processing, making it an ideal candidate for voice applications. On the other hand, visual processing is not neglected, and it exhibits performance that approaches that of some closed models.



In terms of safety, it inherits the same post-training safety measures as Inkling, incorporating measures that meet Thinking Machines Lab's internal safety standards. It demonstrates performance on par with Inkling and is fully competitive with other models in 'FORTRESS,' which measures safety related to the risk of crime and violent acts, and 'StrongREJECT,' which measures whether the AI model rejects requests that are clearly considered harmful.



Inkling-Small, along with Inkling, is available through Thinking Machines Lab's ' Tinker ' service. Additionally, the BF16 and NVFP4 models are available on Hugging Face.

in AI, Posted by logc_nt