Direct Editing
Can the model turn a recording and the source HTML into the requested final webpage?
Recording + HTML Edited HTML
Speech · Cursor · Edit history
Benchmarking Audiovisual
Website Development
“Change this.”
Which element? What edit?
1 FDU2 Alibaba Token Hub, Alibaba Group3 CUHK4 THU5 PolyU
01 / The task
A screen recording carries more than words. To edit the right part of a webpage, a model must connect spoken requests with cursor movements, page context, and earlier instructions—including requests the user later undoes.
Can the model turn a recording and the source HTML into the requested final webpage?
Recording + HTML Edited HTML
Can the model resolve weak references into explicit targets and editing actions?
Recording Explicit instructions
How useful are recovered instructions when passed to a fixed DeepSeek-v4-pro executor?
Instructions + HTML Fixed executor
EFS: Edit Fidelity Score. IRS: Instruction Recovery Score. NIU: Normalized Instruction Utility. The tracks provide complementary diagnostic views of intent recovery and code execution.
02 / The benchmark
918 Chinese and English recordings cover 129 source webpages. Each instance contains 10–20 editing steps, with requests that span text, style, layout, content, and revisions.
Manually written transcripts and step-level rubrics connect each request to its intended target and action. Final-page evaluation excludes requirements canceled by whole-step reversions, while still checking the resulting requested state.
Source webpages come from WebVR and are obtained through its original distribution.
03 / Main results
Our leaderboard covers 23 open- and closed-source models. The best direct-editing EFS is 56.95, and the best instruction-recovery IRS is 57.01. Both leave substantial room for improvement.
The leaderboard includes the 17 models reported in the paper and 6 additional models.
| Model | Input |
|---|
Omni denotes models with native audiovisual input; VL denotes vision-language models evaluated with visual input and a manual transcript. “—” indicates an unavailable Track A result. NIU uses the dataset-level oracle EFS as its denominator.
Step-level analyses show that a model can locate the correct target and still fail to implement the requested edit.
Controlled input and temporal-alignment experiments support the value of synchronized speech and visual context for recovering user intent.
For Qwen3.5-Omni-Plus, passing recovered instructions to the fixed executor raises EFS from 29.56 in direct editing to 51.08.
Room to improve
Oracle instructions are explicit reference instructions supplied to the same coding model.
04 / Citation
Please cite the paper if you use the benchmark or evaluation code.
@misc{han2026omni2web,
title={Omni2Web: Benchmarking Audiovisual Website Development},
author={Han, Minghao and Xing, Zhenghao and Cheng, Xize and Wang, Yuxuan and Lin, Junming and Wang, Ling and Yan, Yinsong and Chu, Yunfei and Yang, Qize and Xu, Jin},
year={2026}
}