Thinking Machines Lab Inkling Performance Analysis
- •Thinking Machines Lab’s Inkling scores 836 Elo on the AA-Briefcase agentic knowledge work benchmark.
- •Inkling achieves a 19.3% rubric score, surpassing DeepSeek V4 Flash max and Gemini 3.5 Flash-Lite.
- •The model shows a presentation Elo of 863 and analytical quality Elo of 764 across tasks.
Thinking Machines Lab’s Inkling model achieved an Elo score of 836 on the AA-Briefcase benchmark, a testing platform for agentic knowledge work involving thousands of input files. The model currently ranks behind leading open-weights models like Nemotron 3 Ultra and GLM-5.2 but maintains a higher position than DeepSeek V4 Flash. The benchmark measures performance through binary rubric checks, pairwise analytical quality grading, and pairwise presentation quality grading.
Regarding the AA-Briefcase rubric score, Inkling reached 19.3%. This result sits below MiMo-V2.5-Pro at 21.4%, yet outperforms DeepSeek V4 Flash max at 18.7% and Gemini 3.5 Flash-Lite at 14.8%. Performance is lowest when the model processes non-standard file types outside of Excel, PowerPoint, PDF, or Word, despite its native multimodal (capable of processing multiple data formats) capabilities. Inkling displays superior presentation quality with an Elo of 863 compared to an analytical quality Elo of 764.
In terms of resource usage, Inkling consumes an average of 52K output tokens per task and 5M tokens for the full benchmark suite. Excel deliverables require the highest token usage, followed by Word, PDF, PowerPoint, and Other file types. The model exhibits a mean of 81 turns per task with a median of 49, indicating a wide distribution in interaction length. Despite maintaining one of the higher average turns per task, Inkling utilizes an average of 0.5 tool calls (functions allowing models to interact with external software) per turn, which is relatively low compared to other models.