Inkling Small、Inklingに肉薄
- •Inkling Smallは276B総パラメータでArtificial Analysis Intelligence Indexのスコア40を記録
- •Thinking Machinesのモデルは12B active MoE parametersながら、Inklingとの差を1ポイントに抑えた
- •Inkling SmallはHumanity's Last Exam、GPQA Diamond、CritPt、SciCodeでInklingを上回った
Thinking Machines LabはJuly 30, 2026、Inklingの2週間後となる2番目のモデルとしてInkling Smallを公開し、Artificial AnalysisはArtificial Analysis Intelligence Indexで40点を付けた。元OpenAI CTOのミラ・ムラティ(Mira Murati)が創業したサンフランシスコ拠点のAI研究所は、Inkling SmallをApache 2.0 licenseのopen weights reasoning modelとして提供した。Inkling Smallは276B total parameters、12B active MoE(タスクごとにモデルの一部だけを経路選択する仕組み)、text・image・speech入力、text出力、256K token context windowを備え、Inklingは1M tokensに対応する。
Inkling Smallは41点のInklingに1ポイント差まで迫りながら、Inklingの975B total parametersと41B active parametersの3分の1未満で動作した。Artificial Analysisによると、同サイズ以下のopen weights modelでIntelligence Indexのスコアがこれを上回るものはない。DeepSeek V4 Flash (max)は284B total、13B active parametersで40点、MiniMax-M3は23B activeで44点、GLM-5.2 (max)は40B activeで51点、MiniMax-M2.7は230B total、10B active parametersで2ポイント低い。
Inkling Smallは複数のcodingとfrontier reasoning評価でInklingに並ぶか上回った。Humanity's Last Examは32%対30%、GPQA Diamondは89%対87%、CritPtは8%対5%、SciCodeは49%対46%、Terminal Bench v2.1は同じ55%だった。長期的なagentic knowledge-work benchmarkであるAA-Briefcaseでは917 Eloを記録し、Inklingの839を上回った。rubric pass ratesは20%対19%でほぼ同じだった一方、Inkling Smallは平均34 turnsでタスクを終え、Inklingは81だった。
Inkling Smallは一部のagentic tasksとfactual knowledge measuresではInklingを下回った。τ³-Bankingは15%対24%、GDPval-AA v2は1269対1237 Eloでわずかに先行し、AA-Omniscience IndexはInklingの2に対して-9だった。AA-Omniscience Accuracyも31%対40%と低く、Hallucination Rateは57%対63%だった。
Inkling SmallのIntelligence Index taskあたり平均出力は~24K output tokensで、Inklingの~25Kをわずかに下回り、同程度の知能水準にあるDeepSeek V4 Flash (max)の~45K、GPT-5.4 mini (xhigh)の~78Kも下回った。Intelligence Index全体の実行では、Inkling Smallが~131M output tokensを使用し、Inklingの~128Mをわずかに上回った。