AI 파이프라인 비용 25% 절감
- •AI 파이프라인은 provider, prompt 내용, 저가 모델 변경 없이 비용을 25% 줄였다
- •`tool_choice`와 더 엄격한 schema로 Sonnet이 약 4분의 1 비용에 Opus와 같은 pass-parity에 도달했다
- •`max_tokens`를 16000에서 4000으로 낮추자 일부 Anthropic 실패 경로에서 출력 비용이 줄었다
Marc는 August 1 Dev.to에 올린 글에서 `ai.generate`, `ai.extract`, `ai.classify`라는 3가지 단계가 각각 run time에 provider, model, prompt revision을 독립적으로 결정하는 AI 파이프라인을 설명했다. 이 파이프라인은 Sonnet을 기본 모델로 사용했고, Opus는 더 비싸지만 더 안정적인 선택지로 취급됐다.
비용 절감은 provider 변경이나 prompt 재작성 없이 configuration 조정에서 나왔다. `tool_choice`(모델의 도구 사용 경로를 강제하는 설정)와 더 엄격한 output schema를 함께 적용하자, 팀의 eval에서 Sonnet은 Opus의 기존 출력과 같은 pass-parity에 도달하면서 비용은 약 4분의 1 수준이었다. Anthropic이 agentic task에 권장한 `effort: "xhigh"`는 같은 pass rate에서 `"high"`보다 thinking tokens를 약 2배 만들었고, 명시적 override로 Opus를 쓸 때 accuracy gain은 0인데 비용은 2x가 됐다. 또한 실제 단계 필요량보다 여전히 높은 `max_tokens`를 16000에서 4000으로 낮추자, Anthropic이 일부 failure path에서 cap 기준으로 과금할 수 있어 effective output cost가 다시 줄었다.
팀은 각 단계 뒤에 prompt revision, provider, model, token usage를 포함한 immutable provenance records를 저장해 이 3가지 절감 지점을 찾았다. 저자는 이 패턴을 LLM call에 대한 event sourcing(상태 변화를 event로 기록하는 방식)에 비유했고, user content를 저장하지 않고도 6개월 동안 tenant-level prompt auditing이 가능했다고 말했다.
글은 adaptive thinking과 forced tool use를 함께 켜면 Anthropic이 “Thinking may not be enabled when tool_choice forces tool use.”라는 400 error로 request를 거부한다고도 경고했다. 해결책은 request 전송 전에 forced `tool_choice`를 자동 감지해 adaptive thinking을 비활성화하는 방식이었다. Prompt caching은 `cache_control`을 마지막 static block에 두고 tools와 system prompt를 하나의 cacheable prefix로 구성할 때만 작동했으며, 그 설정 없이 roughly 16k tokens를 넘으면 충분한 streamed response가 돌아오기 전에 request가 SDK의 HTTP timeout에 걸렸다.