中
← Back to projects
Project notes

CriticGUI

Can a multimodal LLM be a reliable critic for GUI agents?

Multimodal UnderstandingCritic EvaluationState AwarenessFailure Analysis

CriticGUI evaluates whether a multimodal LLM can judge a GUI action correctly and explain its outcome. The focus is state-aware, instruction-grounded criticism—not action generation.

Click to inspect the full diagram.

Why a critic needs to observe state

A plan can describe the intended steps, but it cannot guarantee what the interface will do. Opening a dialog may change the screen substantially; focusing a field may barely change it. The next instruction needs to be grounded in the state actually reached.

What exactly is being judged?

The critic receives the instruction hierarchy, action code and visual evidence. It returns a success/failure label and an explanation. A semantic atomic action may contain several physical inputs; it is not necessarily a single click.

01 / Goal

Add a watermark saying “Draft”.

02 / Subtask

Enter “Draft” and apply the watermark.

03 / Atomic action

Focus the watermark text field.

Judge the current instruction. Focusing the field can be successful even though the watermark has not yet been created.

Progress and completion are different questions. Focusing the field may complete the atomic instruction while leaving the watermark subtask unfinished. The evaluation must name the level being judged. In my notes, separating helpful from complete became a design question—not an additional validated metric in the results below.

What annotation revealed

Annotating trajectories exposed three ways a seemingly plausible judgment can miss the actual task.

Small changes can carry the decisive evidence

Selecting the link-editing tool may leave most of the document unchanged. Tool mode, cursor appearance and local interface cues matter more than the overall visual similarity of the two screens.

A correct label can come with the wrong reason

In my link-region annotation notes, a failed selection was attributed to incomplete coverage of the target text. The model also predicted failure, but argued that no rectangle was visible in the final screenshot—even though the interface had already opened the Create Link dialog. The label agreed with the annotation, while the explanation relied on a different reading of the evidence. The drawing process matters here: its geometry may no longer be visible in the final screenshot.

Label agreement alone does not establish a faithful explanation. This is a qualitative annotation finding, not a separate quantitative explanation evaluation.

A completed operation can violate the user’s goal

Another task asked the agent to reorder PDF pages. The interface had instead reached a deletion confirmation. Clicking OK successfully deleted a page, reducing the page count from three to two, but moved away from the user’s goal. The appropriate response was to cancel the mistaken operation. A critic must detect this conflict rather than treating local execution success as task success.

Build a sample around one atomic action

Each sample aligns one atomic instruction with its action code, before/after screenshots and a matching video clip. The illustrated sample checks whether “GUICritic” was entered into a signature field. Its label and explanation refer to that local result, not the completion of the full signature workflow.

Why add human demonstrations?

Agent exploration is a useful source of natural successes and failures, but it inherits the capabilities of the executing MLLM. On a hard task, the agent may fail before reaching the very state we want a critic to assess. Collecting only what that agent can execute therefore narrows the task distribution and can couple the evaluation too closely to the collection model.

Human demonstrations extend the states and interactions the benchmark can cover. A person can complete difficult procedures, expose subtle intermediate changes, and introduce controlled mistakes. The two sources complement each other: agent exploration provides naturally occurring behavior, while human collection makes difficult cases accessible and repeatable. The agent branch follows WorldGUI; the CriticGUI repository provides the human-demonstration toolkit.

The difficulty of executing an action is not the same as the difficulty of judging its outcome. Agent collection can stall at one state and produce repeated, weakly related samples. Human demonstrations help reach later states and construct subtle contrasts, including interactions that are easy to perform but hard to evaluate.

From a reviewed plan to critic examples

  1. Draft the plan. An MLLM proposes a procedure from the user query and initial screen, optionally using a tutorial video or transcript.
  2. Refine and freeze it. A human corrects the workflow and splits non-atomic operations. The plan has three levels: high-level milestone → low-level subtask → atomic interaction. Only the reviewed version is treated as the ground-truth plan.
  3. Record one atomic interaction at a time. The demonstrator sees all three levels, executes the finest instruction, and records the input events and visual evidence. An atomic interaction can contain multiple physical key presses or mouse events.
  4. Create controlled negative cases. Restore the relevant starting state, deliberately change the execution, and record a separate variant under the same intended instruction. Document the deviation rather than simply relabeling the successful recording.
  5. Review the evidence. Judge the observed outcome, explain it, and retain the plan, action representation, screenshots and video together. An attempted negative is not automatically a failure: the label must follow what actually happened.

The released tool validates plan structure, supports step-wise recording and separate variants, and exports human-reviewed examples. Plan correction, state restoration and perturbation remain human decisions.

Designing informative contrasts

My construction notes explored two main places to introduce difficulty. These are research design options; the released recorder supports separate attempts and human review, rather than automatically generating every augmentation.

What changes?ExampleWhat must the critic understand?
Starting stateA prerequisite is missing, the step is partly done, or the goal is already satisfied.Whether the action is appropriate or needed in this state.
Action or outcomeA nearby wrong target, incomplete text selection, or imprecise placement.Whether the result actually meets the instruction.
Instruction–action pairingRe-pair neighboring steps with different instructions.Whether the evidence matches the intended step; labels require review.

State is relative to the instruction: an unchanged screen can mean an ineffective action, an already-completed task, or a subtle change missed by the observation. Repetition and re-pairing do not automatically produce negatives.

Recovery starts from the state reached after the error. Copying the original correct action is only appropriate when its preconditions still hold. Likewise, an atomic interaction is a semantic unit: a drag may contain move, press, move and release events.

What the experiment found

The initial dataset contains 540 examples from 56 tasks (266 successful and 274 failed actions), covering Acrobat, PowerPoint, Word, Windows Settings and web tasks. GPT-4o achieved the following results:

MetricReported result
Accuracy73.3%
Balanced accuracy73.9%
True positive rate84.9%
True negative rate63.0%

Recognizing failures was harder than recognizing successes. Error analysis highlights overreliance on the final screenshot, missed state changes and confusion between atomic progress and full-task completion.

What remains open

This is a compact, offline benchmark with an initial GPT-4o evaluation. It does not establish performance across current models or prove that a critic improves an online agent.

BERTScore is proposed for comparing explanations, but the reported classification table does not establish explanation faithfulness. Broader model coverage and direct checks of whether reasons match the visual evidence remain important next steps.

How the design evolved

Early notes tried to divide the plan around screen updates. I later separated the concerns: the plan should preserve semantic goals and subtasks, while execution decides when to observe the interface again. A fixed plan cannot replace state-aware execution.

I also explored a unified step-check/action-critic model, iterative self-correction and video-based atomic-action recognition. These remain research directions in the notes, not trained models or demonstrated online improvements reported here.

← Back to projects