AgentSelfEdit Rejects Every Prompt Edit
- •AgentSelfEdit ran 15 iterations and 4,150 LLM calls with zero prompt edit promotions
- •A confidence bug checked p < 0.95 instead of p < 0.05 before being fixed
- •Corrected tests showed net +3 improvement on 26 tasks, but p=0.23 failed the gate
Debashish Ghosal published AgentSelfEdit on September 1, 2026, describing an open-source sidecar that rewrites its own system prompt from execution feedback, A/B tests proposed edits, and promotes only statistically proven winners. In a field test, the system ran 15 iterations and 4,150 LLM calls, but the safety gate rejected every edit because none met the required evidence threshold.
AgentSelfEdit runs a closed loop with 4 stages: Analyze, Test, Gate, and Registry. An LLM reviews failed traces and proposes edits with hypotheses; each edit is A/B tested (comparing two versions on the same task set); a deterministic promotion gate applies 6 checks; promoted edits are versioned with lineage, diff, and rollback. The gate can return 3 outcomes: promote when all 6 checks pass, near-miss when most pass and human review is needed, or reject when a critical check fails.
The gate’s 6 checks are sample floor, effect size, confidence (p < 0.05), frozen sections, edit distance, and drift detection. Ghosal said the gate is not an LLM and uses code rather than prompt-based judgment, because an LLM judging its own edits could approve changes it prefers rather than changes that work. The confidence check uses a p-value (probability a result could arise by chance), with alpha set to 0.05 from a confidence_level of 0.95.
A two-line bug initially made the system look successful. The code checked `p < confidence_level`, meaning `p < 0.95`, instead of computing `alpha = 1 - confidence_level` and checking `p < 0.05`. That meant p-values of 0.9, 0.5, or 0.94 could pass, and one early “promotion” at p=0.1 had a 10% chance of being random noise. After the fix, the same edit produced p=0.23 and was rejected because 23% was above the required 5% threshold.
The corrected run used a local Qwen3.5-4B-4bit model on Apple Silicon. In every one of 15 iterations, the analyzer reviewed 50 real failure traces and proposed adding priority rules to a classification prompt. The A/B test ran the candidate on 26 hard classification tasks; the edit fixed 4 tasks, broke 1 task, and produced a net +3 improvement, or 11.5%. The gate still rejected the edit every time because p=0.23 was not statistically significant at p<0.05.
The field test reported 15 iterations, 4,150 LLM calls, 716,580 total tokens, p-value 0.23 for every iteration, reject as the gate decision for every iteration, 0% false positive rate, 0% false negative rate, $0.00 total cost, and 37 minutes of wall time. Ghosal said the result showed the gate was enforcing the evidence bar rather than allowing a helpful but underpowered edit to change the system prompt.
Ghosal also fixed 31 issues during the session. The A/B test had compared a prompt against itself because the code passed an edited fragment instead of the full candidate prompt. The failure traces were fabricated with `final_output: "other"` even though the model actually output “billing,” “security,” and “technical.” The gate received the edited prompt instead of the original, so the frozen_sections check looked for text after it had already been replaced. A Docker test also used `--dry-run`, skipping the A/B test and gate while still reporting “9/9 tests passed.”
Ghosal said the bugs looked correct in summaries but appeared in raw LLM traffic, which included 4,150 request/response pairs logged to a JSONL file. The project is available through `pip install agent-self-edit`, GitHub, PyPI, a field test report, per-iteration A/B artifacts, closed GitHub issues for all 31 fixes, and a learnings document. For v0.2.0, Ghosal said work is moving toward a rejection-aware analyzer that learns from gate decisions, cumulative evidence across iterations, and larger A/B task sets.