Visual Abstention

in Unified Multimodal Models

Chufan Shi1,*, Cheng Yang2,*, Tiannuo Yang1, Isadora White2

Yiwei Chen2, Taylor Berg-Kirkpatrick2, Xuezhe Ma1

1University of Southern California 2University of California San Diego

*Equal contribution

chufansh@usc.edu · chy085@ucsd.edu

tl;dr

When a requested edit is impossible under the task's rules, a model should say so and draw nothing. We call this visual abstention. On our Draw-or-Decline (DoD) benchmark, unified multimodal models (UMMs) almost never abstain: the strongest editor refuses only 0.4% of infeasible requests. Our paired training method VisTA teaches a model to reason first and then draw or decline: VisTA-BAGEL refuses 93.0% of infeasible requests without any reminder, while completing 74.3% of feasible edits.

Visual abstention in a maze task: draw a valid path when one exists, and decline when no path exists.
Visual abstention in a maze task. (a) A valid path exists, so the model should draw it. (b) No path exists, yet the model still draws through the wall. (c) With visual abstention, the model states that no solution exists and ends its response without generating an image.
01 — Overview

Editing ability does not bring abstention

UMMs integrate understanding and generation, yet their generative behavior is rarely governed by what they understand about the task. We formalize visual abstention, measure it with a paired benchmark, and show that it can be learned.

1Visual Abstention

The task: if no valid output exists under the stated constraints, recognize it, explain why, and decline to generate an image.

2Draw-or-Decline

The benchmark: 1,050 feasible–infeasible pairs across 7 task categories that measure editing success and refusal of infeasible requests together.

3VisTA

The method: paired supervision that teaches a model to judge feasibility before deciding whether to draw. We apply it to BAGEL to obtain VisTA-BAGEL.

02 — Benchmark

Draw-or-Decline: minimal feasible–infeasible pairs

Each feasible request has an infeasible counterpart that shares its input image (six real-world categories built from UniREditBench) or its instruction (programmatically generated mazes), so the two differ mainly in whether they can be completed. A model must check the request against the input to handle both correctly.

One feasible-infeasible pair from each of the seven DoD categories.
One pair from each of the 7 categories: material modification, medium interaction, motion state change, pose adjustment, spatial arrangement, temporal evolution, and maze solving (150 pairs each).
03 — Evaluation

UMMs rarely notice the conflict

We evaluate 8 UMMs and 4 decoupled pipelines (a VLM rewriting the request for an image editor), with Kimi-K3 as the judge (98.5% agreement with the human majority on refusals, 97.0% on edits).

01

Strong editors do not refuse

Without a reminder, no UMM refuses more than 1.7% of infeasible requests; UniREdit-BAGEL edits 68.4% of feasible requests but refuses only 0.4%.

02

They plan the edit as if it were possible

At least 96.8% of their reasoning traces go ahead without acknowledging the conflict: they describe objects that are not there, or quietly rewrite the request into a feasible one.

03

Prompting trades editing for refusals

An explicit infeasibility reminder raises refusals, but editing success falls for every model. Decoupling understanding from generation notices more conflicts, yet still does not refuse reliably and edits worse.

How the reasoning of five UMMs handles infeasible requests.
How the reasoning of 5 UMMs handles infeasible requests, grouped into Refusal, Acknowledge then Comply, Premise Rewriting, and Premise Acceptance.
04 — VisTA

Reason, then draw or decline

VisTA (Visual Transformation and Abstention) pairs every feasible training example with a closely related infeasible one. The response always starts with reasoning: a feasible request continues to the edited image; an infeasible one explains the conflict, ends with [ABSTAIN], and generates no image. We train VisTA-BAGEL from UniREdit-BAGEL on 38,224 pairs.

D93.0% refusal

Refuses 93.0% of infeasible requests without any reminder, up from 0.4% for its starting point.

E74.3% editing

Completes 74.3% of feasible edits, more than any of the 8 evaluated UMMs; learning to refuse does not cost editing accuracy.

F0.8% false refusal

Falsely refuses only 0.8% of feasible requests, so the refusals stay selective.

Editing success versus refusal success on DoD without and with the reminder.
Editing success versus refusal success on DoD, (a) without and (b) with the reminder. Circles are existing UMMs, triangles are decoupled pipelines, and the star is VisTA-BAGEL.
Citation

BibTeX

If you find this work useful, please consider citing:

@article{shi2026visualabstention,
  title   = {Visual Abstention in Unified Multimodal Models},
  author  = {Shi, Chufan and Yang, Cheng and Yang, Tiannuo and White, Isadora and
             Chen, Yiwei and Berg-Kirkpatrick, Taylor and Ma, Xuezhe},
  journal = {arXiv preprint},
  year    = {2026}
}