θvectoreditgymtheta labs · svg benchmark

theta labs / svg editing benchmark

VectorEditGym

Human visual repair instructions, hidden SVG targets, and an auditable specification: repair every visible defect closely enough while preserving the rest of the scene semantically.

authors

YAYug Aditi Guptayug@thetalab.tech
PHPrannay Hebbar

Shared material and equal contribution by both authors.

leaderboard

A repair can be approximate. Its side effects cannot.

34 model endpoints, one scored outcome per task, no fallback routing. A full pass requires every requested repair within calibrated perceptual tolerance, a valid SVG, and semantic preservation outside the requested fields. Invalid outputs receive zero repair progress; UCR is conditional on valid outputs.

Updated 2026-07-20Evaluator semantic-perceptual-binary-2026-07-21Corpus 2a62410b5de7
Submit your run
top ten / specification gates
full table34 entries
#SolverProviderFull passNearProgressCleanSourceValid UCRValidTargetTruncatedErrorsCostTasksDate
#1
Claude Sonnet 5
Anthropic15.0%2.5%43.7%42.5%42.5%0.8%62.5%0.0%40.0%0.0%$2.032402026-07-20
2
KAT Coder Air V2.5
KwaiPilot12.5%0.0%27.3%25.0%25.0%0.3%42.5%0.0%57.5%0.0%$0.128402026-07-20
3
MiniMax M3
MiniMax10.0%0.0%33.0%25.0%25.0%1.0%45.0%0.0%55.0%0.0%$0.240402026-07-20
4
HY 3
Tencent7.5%2.5%56.8%37.5%37.5%1.8%100.0%0.0%0.0%0.0%$0.092402026-07-20
5
Gemma 4 31B IT
Google5.0%5.0%54.8%47.5%47.5%0.8%87.5%2.5%5.0%0.0%$0.084402026-07-20
6
Qwen3.5 397B A17B
Qwen5.0%0.0%39.8%32.5%32.5%1.4%65.0%0.0%32.5%0.0%$0.618402026-07-20
7
DeepSeek V4 Pro
DeepSeek5.0%2.5%30.6%20.0%20.0%1.2%47.5%0.0%52.5%0.0%$0.508402026-07-20
8
Kimi K2.6
Moonshot AI5.0%5.0%16.0%15.0%15.0%0.3%25.0%0.0%72.5%2.5%$1.077402026-07-20
9
Qwen3 Coder
Qwen2.5%0.0%54.7%45.0%45.0%1.4%100.0%0.0%0.0%0.0%$0.295402026-07-20
10
Gemini 3.1 Flash Lite
Google2.5%0.0%54.5%25.0%25.0%2.5%100.0%0.0%0.0%0.0%$0.269402026-07-20
11
Devstral 2512
Mistral2.5%5.0%46.0%57.5%57.5%0.6%97.5%0.0%0.0%0.0%$0.358402026-07-20
12
Inkling
Thinking Machines2.5%0.0%26.6%27.5%27.5%0.8%45.0%0.0%55.0%0.0%$0.806402026-07-20
13
GPT-5
OpenAI2.5%0.0%9.1%7.5%7.5%0.3%15.0%0.0%85.0%0.0%$1.945402026-07-20
14
Nemotron 3 Ultra
NVIDIA2.5%0.0%4.9%5.0%5.0%0.5%7.5%0.0%92.5%0.0%$0.622402026-07-20
15
Claude Haiku 4.5
Anthropic0.0%0.0%54.0%27.5%27.5%2.0%100.0%0.0%0.0%0.0%$0.789402026-07-20
16
Seed 2.0 Mini
ByteDance0.0%2.5%51.2%35.0%35.0%2.0%92.5%0.0%2.5%0.0%$0.174402026-07-20
17
GPT-OSS 120B
OpenAI0.0%5.0%49.9%42.5%42.5%1.1%95.0%0.0%5.0%0.0%$0.057402026-07-20
18
Llama 4 Maverick
Meta0.0%0.0%46.9%20.0%20.0%2.5%92.5%0.0%0.0%0.0%$0.143402026-07-20
19
Mistral Small 2603
Mistral0.0%0.0%39.0%5.0%5.0%4.7%95.0%0.0%0.0%0.0%$0.119402026-07-20
20
GPT-5.4 Nano
OpenAI0.0%0.0%38.0%22.5%22.5%2.2%97.5%0.0%0.0%0.0%$0.177402026-07-20
21
GPT-OSS 20B
OpenAI0.0%2.5%36.5%27.5%27.5%1.6%72.5%0.0%27.5%0.0%$0.030402026-07-20
22
Phi-4
Microsoft0.0%0.0%30.7%12.5%12.5%3.8%92.5%0.0%0.0%0.0%$0.026402026-07-20
23
Laguna M.1
Poolside0.0%2.5%23.3%25.0%22.5%1.4%42.5%0.0%55.0%0.0%$0.101402026-07-20
24
Ling 2.6 Flash
InclusionAI0.0%0.0%19.4%15.0%15.0%4.8%87.5%0.0%5.0%0.0%$0.006402026-07-20
25
Granite 4.1 8B
IBM0.0%0.0%18.5%10.0%10.0%14.3%77.5%0.0%2.5%0.0%$0.019402026-07-20
26
GLM 4.7
Z.ai0.0%0.0%15.1%12.5%12.5%0.9%30.0%0.0%70.0%0.0%$0.549402026-07-20
27
MiMo V2.5 Pro
Xiaomi0.0%0.0%6.4%5.0%5.0%0.8%10.0%0.0%87.5%0.0%$0.384402026-07-20
28
Qwen3.6 35B A3B
Qwen0.0%0.0%4.9%2.5%2.5%1.4%7.5%0.0%92.5%0.0%$0.236402026-07-20
29
Trinity Large Thinking
Arcee AI0.0%0.0%1.8%0.0%0.0%0.0%2.5%0.0%97.5%0.0%$0.176402026-07-20
30
Gemini 3.1 Pro Preview
Google0.0%0.0%0.0%0.0%0.0%n/a0.0%0.0%100.0%0.0%$2.494402026-07-20
31
Nemotron 3 Nano
NVIDIA0.0%0.0%0.0%0.0%0.0%n/a0.0%0.0%80.0%0.0%$0.051402026-07-20
32
OLMo 3 32B Think
AI20.0%0.0%0.0%0.0%0.0%n/a0.0%0.0%0.0%100.0%$0.000402026-07-20
33
Reka Flash 3
Reka0.0%0.0%0.0%0.0%0.0%n/a0.0%0.0%12.5%0.0%$0.040402026-07-20
34
Step 3.7 Flash
StepFun0.0%0.0%0.0%0.0%0.0%n/a0.0%0.0%100.0%0.0%$0.234402026-07-20

paper analysis

The binary score separates four distinct outcomes.

Full paper
Mutually exclusive decomposition of benchmark outcomes by evaluation gate
Each output is assigned once: full pass, completed repair with side effects, valid incomplete repair, or invalid artifact.
Scatter plot of repair progress against valid-output unintended change rate
Repair progress and collateral change on valid outputs are measured independently.
Scatter plot of model repair progress against benchmark run cost
Provider cost does not fully predict repair quality.