RCA · Gemini 2.5 Pro debugger
11.1%→19.4%
Tag+Step Exact
L2: 29.9→36.8%; exact root step: 20.8→29.2%.
EMNLP 2026
1University of Illinois Urbana-Champaign, Urbana, IL, USA 2Stanford University, Stanford, CA, USA 3Yale University, New Haven, CT, USA §Leading authors †Co-senior authors
CUADebug localizes the earliest causal error, grounds it in trajectory evidence, and guides re-execution.
Interactive evidence
Examples span perception, grounding and interaction, reasoning and control, and external or system failures.
Results
Headline RCA and recovery use N=144 Claude failures; Overall Acc. uses all 361 tasks.
RCA · Gemini 2.5 Pro debugger
L2: 29.9→36.8%; exact root step: 20.8→29.2%.
Single re-rollout
Recovery: 13.89→29.86%. Human oracle: 68.98% overall.
Continual re-rollout
Recovery: 12.50→25.69%. Human oracle: 67.87% overall.
Re-execution
Recovery uses N=144 Claude failures; Overall Acc. uses all 361 tasks.
Single: restart before the machine-predicted root step for Machine RCA and before the human-labeled step otherwise. Continual: complete turn 50, then continue from step 51.
Claude 4.5 Sonnet · N=144 failures
| Method | RCA source | Avg. turns | Input tok. | Output tok. | Task Acc. (%) | Overall Acc. (%) | Root error New + Fixed |
|---|---|---|---|---|---|---|---|
| Baseline | none | 28.01 | 272.7K | 3.5K | 13.89 | 61.77 | 20 + 11 |
| Self-debug | self | 31.53 | 324.4K | 3.9K | 15.28 | 62.33 | 22 + 14 |
| Machine RCA | machine RCA | 23.03 | 290.8K | 2.9K | 28.47 | 67.59 | 40 + 25 |
| Human oracle | human annotation | 22.37 | 296.9K | 3.7K | 31.94 | 68.98 | 46 + 29 |
| Our method | RCA + memory | 28.28 | 319.1K | 2.9K | 29.86 | 68.14 | 43 + 24 |
Claude 4.5 Sonnet · N=144 failures
| Method | Original turn | Avg. turns | Input tok. | Output tok. | Acc. (%) | Overall Acc. (%) | Root error New + Fixed |
|---|---|---|---|---|---|---|---|
| Baseline | 50 | 38.92 | 390.3K | 3.8K | 12.50 | 61.22 | 18 + 15 |
| Self-debug | 50 | 43.76 | 401.6K | 4.4K | 13.89 | 61.77 | 20 + 12 |
| Machine RCA | 50 | 47.13 | 405.9K | 5.1K | 21.53 | 64.82 | 31 + 19 |
| Our method | 50 | 47.44 | 426.2K | 5.4K | 25.69 | 66.48 | 37 + 23 |
| Human oracle | 50 | 45.75 | 419.1K | 4.8K | 29.17 | 67.87 | 42 + 25 |
Single re-rollout compares complete packages because restart sources and prompt formats differ.
CUA error taxonomy
The paper taxonomy separates what the agent perceived, where it acted, how it reasoned, what the system did, and whether the task was feasible.
Did the agent misunderstand what was visible in the observation?
Did the agent know the intended operation but execute it on the wrong target or with the wrong mechanics?
Did the agent choose, maintain, or revise the wrong plan?
Did the environment, tool, or benchmark setup prevent otherwise valid progress?
Does the failure fall outside the four execution modules above?
Aligned examples
Examples span each P/G/R/S family and open in the trajectory viewer.