Execution Trees Debug AI Agents
- •Raju Dandigam proposes execution trees as a debugging model for AI-agent runs and causal paths
- •AgentInspect TypeScript examples record planning, sibling tool calls, model-facing steps, and local trace inspection
- •Four tree shapes expose nested ownership, fallbacks, retries, and parallel siblings without replacing raw trace data
Raju Dandigam argued on September 2 that AI-agent debugging should start with execution trees rather than flat logs, because agent behavior depends on the path a run took. The article uses AgentInspect, an open-source TypeScript toolkit maintained by Dandigam, and synthetic fixtures verified against agent-inspect@6.17.4 to show how local traces can make causal structure visible to developers.
A flat timeline can show that at 09:00:00.000 a plan started, at 09:00:00.020 an inventory request started, at 09:00:00.060 it failed with 503, at 09:00:00.061 another inventory request started, at 09:00:00.120 it succeeded, and at 09:00:00.150 the answer completed. Dandigam says that sequence still forces the developer to reconstruct causality mentally, while an execution tree can show a support-agent run with plan, failed fetch-inventory, successful fetch-inventory, and draft-answer as explicit steps.
AgentInspect records developer-selected boundaries through TypeScript wrappers such as inspectRun and step. In the sample travel-planner run, the code records a plan step, two sibling tool calls for search-flights and search-hotels using Promise.all, and a final rank-options model-facing step, then stores the trace under ./.agent-inspect for local inspection with npx agent-inspect view travel-planner --dir .agent-inspect --summary.
The article identifies 4 tree shapes that reveal different bug classes: nested work, fallbacks, repeated siblings, and parallel siblings. A 3-level fixture shows outer at 120ms, middle at 80ms, and inner at 50ms, making ownership visible when inner work fails. A fallback fixture shows primary-search failing after 100ms, fallback-search succeeding after 200ms, and handle-recovered-result finishing after 50ms, preserving recovery behavior even when the final answer succeeds.
Repeated sibling operations show retries that a final success status could hide. One fixture records fetch-inventory failing after 40ms, failing again after 45ms, succeeding after 60ms, and then handle-recovered-result taking 30ms; Dandigam says the tree provides evidence that a retry policy was exercised, though it does not prove the policy was correct. A parallel fixture lists search-hotels at 300ms, search-flights at 200ms, and search-cars at 100ms as sibling operations, warning that those durations should not simply be added because the steps may overlap.
Dandigam says execution trees are a view, not the full evidence model. The underlying trace may include identifiers, timestamps, status, inputs or outputs, observations, and metadata, while different projections answer different questions: tree for what path happened, check for whether an invariant held, diff for what changed between runs, report for reviewer-facing reading, and bundle for shareable evidence.
The workflow turns suspicious shapes into deterministic checks. A CLI trajectory check can require search-flights and fail on recorded failed observations, while the experimental TraceContract API can express tool requirements, forbidden tools, maximum calls, ordering, run status, duration, model allowlists, and token ceilings. Because TraceContract is beta in the referenced release, the article says teams should pin the version and test exact semantics before using it as a CI gate.
Execution trees do not prove answer correctness. A retrieval step can return irrelevant documents, a model call can produce unsupported claims, and a tool can succeed while returning stale data, so Dandigam recommends semantic evaluators, domain tests, and human review for content quality. The article’s closing design principle is to debug the path, not only the final answer, by preserving causal structure, exposing unsuccessful work after recovery, and converting suspicious patterns into repeatable checks.