Speech · Cursor · Edit history

Omni2Web

Benchmarking Audiovisual
Website Development

“Change this.”
Which element? What edit?

Minghao Han1,2,*Zhenghao Xing3,*Xize Cheng2Yuxuan Wang2Junming Lin2,4Ling Wang2,5Yinsong Yan2,5Yunfei Chu2Qize Yang2Jin Xu2,†

1 FDU2 Alibaba Token Hub, Alibaba Group3 CUHK4 THU5 PolyU

* Equal contribution.† Corresponding author.

Instances
918
Edit steps
13,907
Operation types
12
Models evaluated
23

01 / The task

Follow the cue.
Make the right edit.

A screen recording carries more than words. To edit the right part of a webpage, a model must connect spoken requests with cursor movements, page context, and earlier instructions—including requests the user later undoes.

The orange curve follows the mouse trajectory through the recording. The model returns a complete edited HTML page. Click to enlarge.
AEFS

Direct Editing

Can the model turn a recording and the source HTML into the requested final webpage?

Recording + HTML Edited HTML

BIRS

Instruction Recovery

Can the model resolve weak references into explicit targets and editing actions?

Recording Explicit instructions

CEFS / NIU

Instruction Utility

How useful are recovered instructions when passed to a fixed DeepSeek-v4-pro executor?

Instructions + HTML Fixed executor

EFS: Edit Fidelity Score. IRS: Instruction Recovery Score. NIU: Normalized Instruction Utility. The tracks provide complementary diagnostic views of intent recovery and code execution.

02 / The benchmark

Small references.
Long editing histories.

918 Chinese and English recordings cover 129 source webpages. Each instance contains 10–20 editing steps, with requests that span text, style, layout, content, and revisions.

Manually written transcripts and step-level rubrics connect each request to its intended target and action. Final-page evaluation excludes requirements canceled by whole-step reversions, while still checking the resulting requested state.

Dataset anatomy of Omni2Web. Chinese: 465 instances; English: 453. Median recording duration: 128.04 seconds.

Source webpages come from WebVR and are obtained through its original distribution.

03 / Main results

Finding the target
is only the beginning.

Our leaderboard covers 23 open- and closed-source models. The best direct-editing EFS is 56.95, and the best instruction-recovery IRS is 57.01. Both leave substantial room for improvement.

The leaderboard includes the 17 models reported in the paper and 6 additional models.

Main results on 918 instances. All scores are on a 0–100 scale; higher is better.
ModelInput

Omni denotes models with native audiovisual input; VL denotes vision-language models evaluated with visual input and a manual transcript. “—” indicates an unavailable Track A result. NIU uses the dataset-level oracle EFS as its denominator.

01

Grounding is not enough.

Step-level analyses show that a model can locate the correct target and still fail to implement the requested edit.

02

Timing matters.

Controlled input and temporal-alignment experiments support the value of synchronized speech and visual context for recovering user intent.

03

Recovered intent can travel further.

For Qwen3.5-Omni-Plus, passing recovered instructions to the fixed executor raises EFS from 29.56 in direct editing to 51.08.

Room to improve

A gap remains even with
a fixed executor.

Oracle instructions are explicit reference instructions supplied to the same coding model.

Best recovered instructions60.24
Oracle instructions89.69

EFS · DeepSeek-v4-pro executor · 918 instances

04 / Citation

Working with Omni2Web?

Please cite the paper if you use the benchmark or evaluation code.

@misc{han2026omni2web,
  title={Omni2Web: Benchmarking Audiovisual Website Development},
  author={Han, Minghao and Xing, Zhenghao and Cheng, Xize and Wang, Yuxuan and Lin, Junming and Wang, Ling and Yan, Yinsong and Chu, Yunfei and Yang, Qize and Xu, Jin},
  year={2026}
}