Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

Token Savings Stack Fails Pilot

Token Savings Stack Fails Pilot

DEV.to·Friday, August 28, 2026
  • •Token optimization stack ended after pilot showed 97% token savings came from a failed run
  • •Full benchmark plan required about 4,800 agent runs and more than $1,200 on Haiku
  • •Author recommends cost per solved task and turn count instead of token-only dashboards
  • •Token optimization stack ended after pilot showed 97% token savings came from a failed run
  • •Full benchmark plan required about 4,800 agent runs and more than $1,200 on Haiku
  • •Author recommends cost per solved task and turn count instead of token-only dashboards
  • •Token optimization stack ended after pilot showed 97% token savings came from a failed run
  • •Full benchmark plan required about 4,800 agent runs and more than $1,200 on Haiku
  • •Author recommends cost per solved task and turn count instead of token-only dashboards
  • •Token optimization stack ended after pilot showed 97% token savings came from a failed run
  • •Full benchmark plan required about 4,800 agent runs and more than $1,200 on Haiku
  • •Author recommends cost per solved task and turn count instead of token-only dashboards

Shreyash ended a public side project called token-optimization-stack after a few weeks because a pilot benchmark showed token-savings metrics could hide failed coding-agent work and a full test would cost too much to run properly. The project began after he kept hitting token limits while using Claude Code for engineering work and wanted to reduce token spend without lowering output quality. He also built token-stack-benchmarks to test whether the stack worked against real tasks.

The first stack combined 5 tools: Graphify for querying codebases as knowledge graphs (maps code relationships), Serena for symbol-based navigation and edits, Headroom for advertised context compression, LiteLLM for model routing, and Caveman for compressing agent output. Headroom was removed because `headroom init claude` only registered an on-demand MCP tool, `headroom doctor` showed nothing was routed through it without a separate proxy process and `ANTHROPIC_BASE_URL`, and `mcp serve` crashed against a current MCP SDK unless pinned to `mcp<2`. LiteLLM was also removed because its usage-based routing balanced load across provider endpoints rather than routing each task by complexity, while Claude Code used one fixed model for an entire session.

The tested stack kept Graphify, Serena, LeanCTX, and Caveman. The rigorous benchmark plan used real tasks from SWE-bench Verified and Multi-SWE-bench, 16 repos, 5 stack versions, and 3 repeats each, totaling about 4,800 agent runs. Shreyash ran only a cheaper pilot: 31 tasks, 2 stack versions, 1 repeat, `claude-haiku-4-5` at medium effort, and only 11 of 31 task pairs partly finished. That pilot already cost about $5.60 in raw API spend, excluding EC2 costs, Docker builds, and setup time across EC2 and an Apple Silicon Mac.

Scaling the pilot's per-run cost to the full benchmark would exceed $1,200 in API spend on the cheapest model before producing trustworthy results. Sonnet cost 2x Haiku on both input and output tokens, listed as $2/$10 per million tokens versus Haiku's $1/$5, so the same test would pass $2,400 on Sonnet. Haiku also failed as an evaluator of the idea: it produced zero correct fixes on Java tasks and broke 2 of the 3 Python tasks that the plain baseline had solved.

The partial pilot showed why token reduction alone was misleading. On 3 Python tasks that the baseline solved correctly, the full stack failed one by asking a one-line clarification on the first turn and stopping with `num_turns: 1`; the baseline took 30 turns on the same prompt and fixed the bug. The dashboard made that failed run look best because it used 97% fewer tokens. A second stacked run produced a patch that applied cleanly but left the target test failing. A third remained correct but used more tokens than the baseline.

Shreyash argued that reports should measure cost per solved task, not cost per task, because the 97%-savings run solved zero tasks and should count as infinitely expensive. He also proposed checking turn count as a cheap warning signal, since a 1-turn run where the baseline needed 30 turns can reveal silent failure before full correctness scoring. He left both repos public and named 3 follow-up paths: investigate the wrong-but-plausible patch, run the full 5-arm test to check whether Caveman caused the 97% failure, and rerun the pilot on Sonnet or Opus to test whether the Haiku weakness prediction holds.

Shreyash ended a public side project called token-optimization-stack after a few weeks because a pilot benchmark showed token-savings metrics could hide failed coding-agent work and a full test would cost too much to run properly. The project began after he kept hitting token limits while using Claude Code for engineering work and wanted to reduce token spend without lowering output quality. He also built token-stack-benchmarks to test whether the stack worked against real tasks.

The first stack combined 5 tools: Graphify for querying codebases as knowledge graphs (maps code relationships), Serena for symbol-based navigation and edits, Headroom for advertised context compression, LiteLLM for model routing, and Caveman for compressing agent output. Headroom was removed because `headroom init claude` only registered an on-demand MCP tool, `headroom doctor` showed nothing was routed through it without a separate proxy process and `ANTHROPIC_BASE_URL`, and `mcp serve` crashed against a current MCP SDK unless pinned to `mcp<2`. LiteLLM was also removed because its usage-based routing balanced load across provider endpoints rather than routing each task by complexity, while Claude Code used one fixed model for an entire session.

The tested stack kept Graphify, Serena, LeanCTX, and Caveman. The rigorous benchmark plan used real tasks from SWE-bench Verified and Multi-SWE-bench, 16 repos, 5 stack versions, and 3 repeats each, totaling about 4,800 agent runs. Shreyash ran only a cheaper pilot: 31 tasks, 2 stack versions, 1 repeat, `claude-haiku-4-5` at medium effort, and only 11 of 31 task pairs partly finished. That pilot already cost about $5.60 in raw API spend, excluding EC2 costs, Docker builds, and setup time across EC2 and an Apple Silicon Mac.

Scaling the pilot's per-run cost to the full benchmark would exceed $1,200 in API spend on the cheapest model before producing trustworthy results. Sonnet cost 2x Haiku on both input and output tokens, listed as $2/$10 per million tokens versus Haiku's $1/$5, so the same test would pass $2,400 on Sonnet. Haiku also failed as an evaluator of the idea: it produced zero correct fixes on Java tasks and broke 2 of the 3 Python tasks that the plain baseline had solved.

The partial pilot showed why token reduction alone was misleading. On 3 Python tasks that the baseline solved correctly, the full stack failed one by asking a one-line clarification on the first turn and stopping with `num_turns: 1`; the baseline took 30 turns on the same prompt and fixed the bug. The dashboard made that failed run look best because it used 97% fewer tokens. A second stacked run produced a patch that applied cleanly but left the target test failing. A third remained correct but used more tokens than the baseline.

Shreyash argued that reports should measure cost per solved task, not cost per task, because the 97%-savings run solved zero tasks and should count as infinitely expensive. He also proposed checking turn count as a cheap warning signal, since a 1-turn run where the baseline needed 30 turns can reveal silent failure before full correctness scoring. He left both repos public and named 3 follow-up paths: investigate the wrong-but-plausible patch, run the full 5-arm test to check whether Caveman caused the 97% failure, and rerun the pilot on Sonnet or Opus to test whether the Haiku weakness prediction holds.

Read original (English)·Aug 25, 2026
Coding#coding agents#token optimization#claude code#swe bench#multi swe bench#graphify#serena#leanctx#caveman#litellm