SolveEdit

Benchmarking Visual Problem Solving in Generative Models

Wenjie Shu1,†, Yexin Liu2,†, Harold Haodong Chen2, Xuerui Qiu3, Zehan Wang4,
Yidi Zhang1, Yizhan Chen1, Zunwei Wang1, Minghao Liu5, Qi Chen1,*,
Harry Yang2,*, Xiaogang Xu4

1 ZODA · 2 HKUST · 3 UCAS · 4 ZJU · 5 UTokyo

†Equal contribution. *Corresponding authors.

Leaderboard

11 models · 2,728 cases

Direct model evaluation on all 2,728 SolveEdit cases. Video models are scored from their final frames.

RankModelTypeRequired R ↑Damage D ↓SolveScore ↑
1GPT-Image-2Commercial image67.623.957.0
2Seedream 5.0 ProCommercial image64.317.656.6
3Gemini 3.1 Flash ImageCommercial image66.023.155.5
4Qwen Image 2.0 ProCommercial image50.728.840.6
5Kling V3Image to video39.134.027.5
6FLUX.2 [dev]Open source image31.732.322.8
7Qwen-Image-Edit-2509Open source image32.747.520.8
8FLUX.1 Kontext [dev]Open source image24.222.518.7
9BAGELOpen source image22.445.512.1
10OmniGen2Open source image18.145.99.7
11HunyuanVideo-1.5Image to video17.355.57.4

Scores are percentages. Higher is better for R and SolveScore; lower is better for D.

SolveEdit benchmark overview

Abstract

Generative visual systems are often asked to transform an existing scene while preserving what should remain. SolveEdit evaluates this ability as visual problem solving: given an image and a goal, a model must infer a valid transition from the request and visual evidence, execute it, and protect unrelated content. The benchmark contains 2,728 cases with atomic required and protected conditions, enabling reference-free completion and preservation scores.

The problem: an edit is a transition

Most image-editing evaluations state the desired transformation explicitly. Real visual tasks are less direct: the request may leave the target, destination, or final state to be inferred from the scene, or may refer to a rule shown in the image. A successful system therefore has to answer two linked questions: what transition is valid? and can it render that transition without changing unrelated content?

1. Determine
Ground the request and recover the scene-dependent transition.
2. Execute
Ask an image or video generator to realize the transition.
3. Preserve
Check required completion and protected content separately.

What the benchmark measures

SolveEdit application breadth

SolveEdit spans ten application domains, from semantic grouping and numerical correction to route tracing, structural assembly, and constrained action.

SolveEdit formulation and evaluation overview

SolveEdit separates transition determination from image rendering. Each case records required conditions for the intended change and protected conditions for content that should remain intact.

SolveEdit benchmark statistics

Benchmark at a glance

2,728
visual problem cases
29,460
atomic contract criteria
57.0%
best direct SolveScore

Cases are organized by the information that fixes the transition: instruction-specified (IS), state-dependent (SD), and rule-dependent (RD). SolveScore reports required completion together with collateral changes to protected content. This separates a model choosing the wrong solution from a model rendering the right solution poorly.

Planning before generation

SolveEdit-Plan uses Inspect and Resolve to recover unresolved transition variables before the final editing call. It improves GPT-Image-2 from 57.0% to 71.6% under a matched one-generation budget and transfers to the tested open-source image and video generators.

Planning results across controls and generators

Representative cases

The three examples below show how the information source changes across IS, SD, and RD cases. Each pair shows the input and a generated output.

IS: Selective cleanup

IS input
Input
IS output
Output

SD: Road bridge repair

SD input
Input
SD output
Output

RD: Storyboard connection repair

RD input
Input
RD output
Output