EMNLP 2026

CUADebug: Diagnosing and Repairing
Computer-Use Agent Failures

Weijia ZhangKunlun ZhuZeyi Liu1 Yinting Chen1Tianyi Ma1Jiateng Liu1 Jiaxun Zhang1Bingxuan Li1Pan Lu2Xiangru Tang3 Heng Ji1†Jiaxuan You1†

1University of Illinois Urbana-Champaign, Urbana, IL, USA 2Stanford University, Stanford, CA, USA 3Yale University, New Haven, CT, USA §Leading authors Co-senior authors

CUADebug localizes the earliest causal error, grounds it in trajectory evidence, and guides re-execution.

204
human-annotated failures
11.1 19.4%
Tag+Step Exact Gemini 2.5 Pro debugger · Claude N=144
61.77 68.14%
single Overall Acc. Recovery: 13.89 → 29.86%
61.22 66.48%
continual Overall Acc. Recovery: 12.50 → 25.69%
Pipeline. CUADebugger inspects failed trajectories, produces structured RCA, retrieves relevant memories, and guides re-rollout.

Why root-cause diagnosis

The visible failure may follow an earlier mistake.

Motivating example. Why CUA failures require root-cause diagnosis rather than terminal-state inspection.

Interactive evidence

Explore selected aligned trajectories.

Examples span perception, grounding and interaction, reasoning and control, and external or system failures.

--

Loading · --------

Loading trajectory...

Task instruction

Loading task...

Step -- / --
Post-action state
Loading visual evidence...
Trajectory timelineRed marks the root cause; amber marks downstream drift.

Results

Diagnosis and re-execution.

Headline RCA and recovery use N=144 Claude failures; Overall Acc. uses all 361 tasks.

Single re-rollout

61.77%68.14%

Overall Acc.

Baseline 61.77CUADebug 68.14

Recovery: 13.89→29.86%. Human oracle: 68.98% overall.

Continual re-rollout

61.22%66.48%

Overall Acc.

Baseline 61.22CUADebug 66.48

Recovery: 12.50→25.69%. Human oracle: 67.87% overall.

Re-execution

Re-rollout comparison.

Recovery uses N=144 Claude failures; Overall Acc. uses all 361 tasks.

Single: restart before the machine-predicted root step for Machine RCA and before the human-labeled step otherwise. Continual: complete turn 50, then continue from step 51.

Table 3Single re-rollout

Claude 4.5 Sonnet · N=144 failures

Method RCA source Avg. turns Input tok. Output tok. Task Acc. (%) Overall Acc. (%) Root error New + Fixed
Baselinenone28.01272.7K3.5K13.8961.7720 + 11
Self-debugself31.53324.4K3.9K15.2862.3322 + 14
Machine RCAmachine RCA23.03290.8K2.9K28.4767.5940 + 25
Human oraclehuman annotation22.37296.9K3.7K31.9468.9846 + 29
Our methodRCA + memory28.28319.1K2.9K29.8668.1443 + 24
Table 4Continual re-rollout

Claude 4.5 Sonnet · N=144 failures

Method Original turn Avg. turns Input tok. Output tok. Acc. (%) Overall Acc. (%) Root error New + Fixed
Baseline5038.92390.3K3.8K12.5061.2218 + 15
Self-debug5043.76401.6K4.4K13.8961.7720 + 12
Machine RCA5047.13405.9K5.1K21.5364.8231 + 19
Our method5047.44426.2K5.4K25.6966.4837 + 23
Human oracle5045.75419.1K4.8K29.1767.8742 + 25

Single re-rollout compares complete packages because restart sources and prompt formats differ.

Benchmark distribution. CUAErrorBench annotation distribution by trajectory source.
RCA by category. Category-level RCA F1 for the CUADebugger trial using Gemini 2.5 Pro as the debugger on Claude 4.5 Sonnet trajectories (144-task split), including the Others (O) category.

CUA error taxonomy

Five families. Thirty defined subtypes.

The paper taxonomy separates what the agent perceived, where it acted, how it reasoned, what the system did, and whether the task was feasible.

PPerception5 subtypes · 36 annotations+

Did the agent misunderstand what was visible in the observation?

  • P1 Visual hallucination
  • P2 Misrecognition / OCR error
  • P3 Cross-modal misbinding
  • P4 Observation omission
  • P5 Semantic misunderstanding
GGrounding and Interaction4 subtypes · 25 annotations+

Did the agent know the intended operation but execute it on the wrong target or with the wrong mechanics?

  • G1 Coordinate / element grounding error
  • G2 Visibility / accessibility error
  • G3 Interaction mechanics error
  • G4 Distraction / adversarial misdirection
RTask Reasoning and Control13 subtypes · 110 annotations+

Did the agent choose, maintain, or revise the wrong plan?

  • R1 Constraint violation
  • R2 Impossible plan / impossible action
  • R3 Decomposition failure
  • R4 Inefficient / redundant strategy
  • R5 Action-intent misalignment
  • R6 Invalid / malformed action
  • R7 Parameter / argument error
  • R8 Context loss / over-simplification
  • R9 Memory hallucination
  • R10 Progress misjudgment
  • R11 Outcome misinterpretation
  • R12 Failed self-correction
  • R13 Causal misattribution
SExternal/System7 subtypes · 13 annotations+

Did the environment, tool, or benchmark setup prevent otherwise valid progress?

  • S1 Rendering / layout failure
  • S2 Timing / race condition
  • S3 Unexpected system behavior
  • S4 Step / resource limit
  • S5 Tool / API failure
  • S6 Environment instability
  • S7 Benchmark / evaluation artifact
OOthers1 subtype · 20 annotations+

Does the failure fall outside the four execution modules above?

  • O1 Infeasible task

Aligned examples

Selected human–debugger aligned cases.

Examples span each P/G/R/S family and open in the trajectory viewer.

Paper figure