Animation Bench
| Rank | Model | Score | Visual | Motion | Layout | Task wins | $ / task |
|---|---|---|---|---|---|---|---|
| 1 | GPT-6.1 Sol | 0.606 ±0.044 | 0.704 | 0.512 | 0.624 | 14 | $0.46 |
| 2 | GPT-6 Astra | 0.594 ±0.045 | 0.710 | 0.473 | 0.628 | 17 | $3.04 |
| 3 | Claude Fable 5.1 | 0.548 ±0.039 | 0.662 | 0.430 | 0.565 | 6 | $3.89 |
| 4 | GPT-6 Sol | 0.516 ±0.032 | 0.640 | 0.381 | 0.539 | 7 | $0.45 |
| 5 | Claude Opus 5.5 | 0.507 ±0.038 | 0.631 | 0.383 | 0.517 | 4 | $1.12 |
- 1GPT-6.1 Sol0.606±0.044
- Visual
- 0.704
- Motion
- 0.512
- Layout
- 0.624
- Task wins
- 14
- $ / task
- $0.46
- 2GPT-6 Astra0.594±0.045
- Visual
- 0.710
- Motion
- 0.473
- Layout
- 0.628
- Task wins
- 17
- $ / task
- $3.04
- 3Claude Fable 5.10.548±0.039
- Visual
- 0.662
- Motion
- 0.430
- Layout
- 0.565
- Task wins
- 6
- $ / task
- $3.89
- 4GPT-6 Sol0.516±0.032
- Visual
- 0.640
- Motion
- 0.381
- Layout
- 0.539
- Task wins
- 7
- $ / task
- $0.45
- 5Claude Opus 5.50.507±0.038
- Visual
- 0.631
- Motion
- 0.383
- Layout
- 0.517
- Task wins
- 4
- $ / task
- $1.12
One selected run per model and task. Scores are reproduction scores on a 0–1 scale, not success rates. GPT-6.1 Sol was added on 1 October 2026 with the same tasks, agent and scorer, recorded in a later scoring run; recording conditions alone can move a 48-task mean by about 0.01, so its lead over GPT-6 Astra is within the uncertainty.
Reference vs. 4 agent reconstructions
Frame-locked to the same moments. Hover a model to overlay the reference. For the full task set, see the appendix.
Background
Frontend generation is a common commercial use case for coding agents. Current agents can reproduce much of a page’s visual palette, typography, and static layout. A production page, however, is not a single frame. It also contains transitions and interactions triggered by scrolling, hovering, clicking, dragging, and cursor movement. Can frontier models reconstruct those behaviours as well as the static appearance?
We built Animation Bench to answer that question and tested four frontier models on 48 tasks from real websites. Each model ran in its own sandbox with the same task inputs and had to deliver a self-contained HTML file. Across the four models, mean visual similarity ranges from 0.63 to 0.71, while mean motion consistency ranges from 0.38 to 0.47.
As we enter the era of recursive self-improvement (RSI), we need to pay more attention to the verifiers used to measure progress. Left unchecked, model performance may saturate on tasks in domains with easily verifiable outcomes. Animation Bench applies this scrutiny to web animation.
Design philosophy
We set out to measure end-to-end animation reconstruction. The benchmark follows five design principles:
- Reproduction > Generation: Score of each task is measured directly and objectively against the real animation.
- Frames in, motion out: Each model received between 12 and 24 frames, depending on the task, and network capture of the page. It did not receive the site’s source code or a description of the timing.
- Real sites selected for animation coverage:
- We sourced animations from commercial sites spanning editorial, portfolio, e-commerce, product, and brand work.
- The 48 animations cover the common trigger types: autoplay, scroll, hover, cursor-follow, click, drag, and state change.
- Separate scoring axes: Three axes measure complementary aspects of reconstruction quality: visual similarity, motion consistency, and layout correctness. A reconstruction can match the layout while missing the motion, which a single undifferentiated score would obscure.
- Test the running artifact: We open each reconstruction in a real browser, drive it as a user would, and record it frame by frame.
Methodology
We used Computer-1 using the Harbor framework to run 192 evaluations: four models on the same 48 tasks, all at maximum reasoning effort. GPT-6.1 Sol was added later in the same setup, for 240 in total. Each model ran in its own sandbox with a 1280×720 desktop, a shell, the task's reference frames, and the page's network capture. The required output was one self-contained HTML file. The reported results contain one selected run for each model and task pair.
- 01Input12–24 frames+ the page’s network capture
- 02AgentComputer-1run through Harbor · max reasoning
- 03Sandbox1280×720desktop + bash shell
- 04Outputindex.htmlone self-contained file
- 05Scored3 axesvisual · motion · layout
Tasks
Each task is one precise animation on a production website with a fixed start and end state.
Input: The model receives 12 to 24 still frames sampled across the animation, with capture timestamps where available, and a HAR capture containing the HTML, CSS, JavaScript, fonts, and images loaded by the live page. It does not receive the site’s source project, a video, or any timing description beyond the frames and their timestamps.
Output: The required output is one self-contained index.html file with inline CSS and JavaScript and no external requests.
Evaluation: We open the file in a headless browser at 1280×720, drive it with the task’s trigger (wait, scroll, hover, drag, click), capture it frame by frame the same way we captured the original, and compare the two captures.
The 48 tasks come from 32 sites and were selected to cover different triggers and animation properties. Each task is tagged by trigger, animated property, timing pattern, spatial extent, and site category.
| Trigger | Tasks | What the model has to recover |
|---|---|---|
| Scroll | 21 tasks | Progress tied to scroll position: pinned sequences, scroll-driven text, scroll-linked transforms |
| Plays by itself | 12 tasks | Load-in entrances, ambient loops, auto-cycling scenes |
| Hover | 6 tasks | Reveals and state changes on pointer enter, and their reversal on leave |
| Click / drag / key | 4 tasks | Drag carousels, dial scrubbing, spring physics |
| Opens or changes | 4 tasks | Page transitions, view switches, expand-collapse |
| Cursor-follow | 1 task | A contextual cursor that changes over specific elements |
Scoring
Each result is compared with the original recording frame by frame. We score three axes, each composed of sub-scores that measure different failure modes.
Visual similarity: does it look right at a given moment?
Each reconstruction frame is compared with the reference frame at the same moment. Five sub-scores are averaged with fixed weights:
| Sub-score | Weight | How it is computed | What it catches |
|---|---|---|---|
| MS-SSIM | 0.32 | Structural similarity of each frame pair, at several scales | Shapes in the wrong places; layout drift |
| LPIPS | 0.32 | 1 − perceptual distance between each frame pair (a learned metric) | Whether a person would say the frames look alike |
| Colour | 0.13 | Overlap of the two foreground colour histograms | Wrong palette |
| Edges | 0.08 | F1 of the two frames’ edge maps | Right colour, wrong geometry |
| Coverage | 0.15 | Ratio of foreground fill: min(fR, fC) / max(fR, fC) | Missing or extra blocks; blank or letterboxed pages |
Motion consistency: does it move right over time?
The motion score compares how the two recordings change over time. Three sub-scores describe the pattern of motion; two penalties then scale the result down when the amount or the placement of that motion is wrong:
| Sub-score | Weight | How it is computed | What it catches |
|---|---|---|---|
| Energy | 0.50 | ½ timing + ½ burstiness of the frame-to-frame pixel-change curve. Timing is the correlation of the two normalised curves; burstiness is the ratio of their coefficients of variation | Motion at the wrong moments; a smooth fade where the original snaps, or the reverse |
| Flow | 0.20 | The same two terms on the optical-flow magnitude curve | Real movement vs fades and flicker |
| Trajectory | 0.30 | 1 − RMS distance between the normalised cumulative paths of the moving region’s centroid | Wrong direction or order of movement |
| G_amount | penalty | min(r, 1/r)^0.35, where r is the ratio of the reconstruction's total motion to the reference's | Too little or too much motion overall |
| G_placement | penalty | (fill ratio)^0.35: the reconstruction's occupied screen area relative to the original | Content in a corner, letterboxed, or missing |
Layout correctness: is it built like the original?
The layout score runs OCR on both recordings and compares the words it finds, frame by frame:
| Sub-score | Weight | How it is computed | What it catches |
|---|---|---|---|
| Presence | 0.35 | F1 of OCR words matched between the two frames (a match allows up to 30% character error) | Words missing or invented |
| Accuracy | 0.25 | 1 − character error rate over the matched words | Misspelt copy |
| Order | 0.15 | 1 − 2 · inversions / n(n − 1) of the matched words, top to bottom | Wrong reading order |
| Alignment | 0.25 | Σ IoU of matched word boxes / (matched + unmatched) | Words in the wrong positions |
Overall score
The overall score is a weighted mean of the three axes. The weights depend on what triggers the animation. Interaction is scored separately and held out, so the three weights are renormalised. Canvas-heavy tasks set the layout weight to 0.05, since OCR cannot see into a canvas.
| Trigger | Visual | Motion | Layout | Interaction (held out) |
|---|---|---|---|---|
| autoplay | 0.35 | 0.35 | 0.20 | 0.10 |
| scroll | 0.30 | 0.35 | 0.20 | 0.15 |
| hover | 0.30 | 0.20 | 0.20 | 0.30 |
| cursor | 0.30 | 0.20 | 0.15 | 0.35 |
| gesture | 0.30 | 0.20 | 0.20 | 0.30 |
| state-change | 0.30 | 0.25 | 0.20 | 0.25 |
All scores are reproduction scores on a 0–1 scale. The code also applies a nominal 0.99 ceiling to the overall score; it never binds (the highest score in the set is 0.890).
Scoring example: oxigen-voxel-palm-pinned
The following reconstruction from oxigen.sa shows how the scoring system works end to end. In the original, a palm tree made of glowing voxels grows over a voxel landscape while the section remains pinned and the copy changes during scrolling. Opus 5.5 produced a recognisable but substantially different result.
Visual similarity: 0.453
Results
Dimensions
Across all five models, mean visual similarity exceeds mean motion consistency. At the task level, visual exceeds motion in 219 of 240 reconstructions. The average gap is 0.23, and the two are only moderately correlated (r = 0.47). A page with high visual similarity is only somewhat more likely to have high motion consistency.
Within this sample, cost has little association with score. Mean spend per task ranges from $0.45 for GPT-6 Sol to $3.89 for Fable 5.1, an 8.6-fold difference, while mean overall score spans 0.099. Per-model Spearman correlations between spend and score range from −0.24 to +0.24.
Spending 1.5–1.9× more on a task moved the score by at most 0.09, and not in one direction.
Where animation reconstructions fail
The final results indicate that motion is the gap. We deeply investigated all 192 generated pages and replayed a subset side by side with the reference. Each failure below is observable in the output, countable across the set, and has a named example to follow along.
Motion timing and sequencing break down
Models do well at capturing motion location (the location gate averages 0.88). However, across the 183 reconstructions in which motion was captured, the timing term averages 0.57, compared with 0.50 when each reconstruction is paired with the reference from a different task. The models often identify which parts of the page should move but reproduce the timing only weakly.
Instead, the motion tends to arrive all at once. In a typical reconstruction the single biggest change between two frames accounts for 29% of all its movement; in the original it is 19%.
Nine of the 32 intros that should play once were implemented as loops that restart indefinitely. On one eight-second sequence, three of the four models completed the full sequence within about one second.
Reconstructing a timed UI sequence
On wisprflow.ai, a two-option pill (Dictation | Notetaker) runs one sequence:
- The white thumb sits on “Dictation”.
- Sliding to “Notetaker” stretches to the width of the label.
- As it lands, the letters of “Notetaker” ripple: each lifts and drops in turn, left to right.
- Easing back to “Dictation”, the letters stay still.
Every model recognised the component and reproduced its appearance (visual 0.90–0.97, layout ≈ 0.95 for all four). Two models also recognised the ripple and implemented a suitable mechanism: Claude Opus 5.5 used a per-letter @keyframes wave with a 70 ms stagger, while Claude Fable 5.1 used a per-character transform sequence. Each model reconstructed a different part of the interaction:
| Slide on time (f1) | Ripple after slide | Returns on time (f10) | |
|---|---|---|---|
| GPT-6.1 Sol | early, on hover before the click | on hover, before the slide | |
| Claude Opus 5.5 | fires immediately | never returns | |
| Claude Fable 5.1 | starts already switched | short faint | |
| GPT-6 Sol | 2 frames late | none | 1 frame late |
| GPT-6 Astra | 7 frames late | barely | never returns |
What makes this case so hard? The pill is small, and each letter of the ripple lifts by only a few pixels. The stills are unevenly spaced: the first is taken almost seven seconds into the recording, a full second passes before the slide, eight quick frames about 150 ms apart catch the ripple, and a second and a half passes before the return. To rebuild it, a model has to read the timestamps as well as the pictures, and turn a few pixels of difference into a sequence with an order, a pause and a return. No single frame shows any of that.
GPT-6 Sol reproduces the component, but not its timing. Its page slides the thumb 0.9 seconds after load, then flips it back and forth every 3.3 seconds, indefinitely. There is no ripple at all: the letters are never split apart, so they cannot move one at a time. Every frame of it looks right (visual 0.94, layout 0.95), and it scores 0.36 on motion.
Motion is incomplete or missing
Reconstructions more often move too little than too much. The median reconstruction carries 0.63 of the reference’s motion, and 38% carry less than half; 10 of 192 overshoot by more than double. GPT-6 Sol is the most restrained, at a median of 0.56. Combined with the timeline finding above, the typical rebuild is a correct-looking page that moves less, and in fewer, larger steps.
Staggered sequences collapse into simultaneous motion
Staggered choreography, in which elements enter one after another, is the most common timing pattern in the set (32 tasks). In 34 of 128 reconstructions of staggered tasks, we found no delay or stagger construct: every element starts together. GPT-6 Sol accounts for 16 of the 34. Under the current scoring, the within-task penalty is small (0.02 on motion), so this is more evident in the code than in the aggregate score.
Complex visual assets are approximated
The most expensive element of a commercial animation is often its hero asset: a WebGL scene, a 3D product, or a photographic sequence. Fifteen of the 16 canvas tasks come from sites that use WebGL or three.js. GPT-6 Astra used WebGL on 5 of the 16 tasks by inlining the site’s captured three.js code; the other models rebuilt these scenes with Canvas2D.
None of the four reconstructions reproduces the full sweep. Claude Opus 5.5 has the highest overall score at 0.537, followed by Fable 5.1 at 0.498, GPT-6 Sol at 0.486, and GPT-6 Astra at 0.338. Sol has the highest visual score (0.708), while Opus has the highest motion score (0.408).
Page structure, behaviour, and text drift from the reference
Eleven of the 52 reconstructions we inspected visually added dark side bars that do not appear in the reference. These pages are letterboxed into a fixed-aspect column instead of filling the viewport (Sol 6, Opus 5.5 3, Fable 5.1 2). At pudding.cool (opens in a new tab) Astra and Fable 5.1 are close to the pinned word cloud reference; Opus 5.5 adds fixed black bars on both sides, while GPT-6 Sol runs ahead of the scroll inside a letterbox.
Secondary interactions and states are omitted
Several reconstructions implement the headline behaviour and drop what surrounds it:
- Pinned sections that do not pin. The content scrolls past instead of holding while the animation plays (3 of 52 inspected). On the slowdown footer, Astra holds the block while the icons rotate; Sol and Fable 5.1 let it scroll away.
- Incomplete dragging effect. On kaviengcreative.com (opens in a new tab), it shows the original interaction: dragging the cards should fly them into a grid. Astra assembles the grid; Fable 5.1 brings the cards forward but never settles them into it; GPT-6 Sol fades the title but the cards never assemble; Opus 5.5 does nothing on drag.

Text is present but misplaced
The models did not invent new copy, and none of the reconstructions contains placeholder text. The visible text comes from the captured page, but it is often displayed at the wrong time or position. Across the sampled frames, roughly 41% of the reference text labels never appear on screen in the reconstruction, and 13% are misspelled. Of the text that a reconstruction does display, 38% has no counterpart in the corresponding reference frame. Bounding-box alignment, which measures whether words occupy the same positions, averages 0.21, the lowest term on any axis.
The visual axis shows the same split between palette and geometry. Colour agreement averages 0.85; edge agreement (whether outlines and borders line up) averages 0.26 and is the weakest visual term in 184 of 192 reconstructions. The models get the palette but don’t get shapes.
Screenshot replay replaces real reconstruction
Fourteen reconstructions solved the task by embedding the reference frames themselves as images and stepping through them on a timer, on scroll, or on hover (Sol 10, Astra 3, Opus 5.5 1). It is the purest form of screenshot mimicry: correct at twelve instants by construction, and wrong everywhere between them. Within the same task, flipbooks score 0.08 lower on motion than reconstructions that rebuild the animation, and slightly lower overall.
Conclusion
Implications
For model labs. The bottleneck to full webpage reproduction is neither perception nor code generation. Frontier models frequently select an appropriate implementation technique but fail to reproduce the complete animation sequence. They misjudge timing, overlap, and duration.
For benchmarks & RL environments. Current benchmarks often grade reconstructed webpages primarily by static visual similarity. Animation benchmarks should also evaluate whether models preserve temporal structure, including duration, overlap, pauses, and event order. A mechanical three-axis scoring system could support a training environment for these behaviours.
Final thoughts
Can frontier models rebuild a web animation rather than only its first frame? Not yet. They reproduce much of its palette, typography, and layout, and they often identify an appropriate implementation technique. What they fail to recover is time: event order, pauses, movement duration, and whether the page returns to its initial state. Every model scored lower on motion than on visual similarity. Screenshot-based evaluation does not capture this temporal gap, but users notice it within seconds.
Appendix: Tasks
All 48 tasks with each model’s overall score, site, trigger and difficulty. One selected generation per task and model, scored against one reference capture; the best score on each task is in bold.
| Task | Sol 6.1 | Astra | Fable | Opus | Sol | Site | Genre | Trigger | Difficulty |
|---|---|---|---|---|---|---|---|---|---|
| 0.570 | 0.385 | 0.618 | 0.497 | 0.465 | adcker.com | portfolio | hover | medium | |
| 0.524 | 0.541 | 0.499 | 0.518 | 0.510 | altitude101.com | portfolio | scroll | hard | |
| 0.651 | 0.635 | 0.606 | 0.451 | 0.681 | altitude101.com | portfolio | scroll | hard | |
| 0.645 | 0.697 | 0.613 | 0.575 | 0.611 | ausify.com.au | saas | drag / click | hard | |
| 0.584 | 0.478 | 0.513 | 0.570 | 0.540 | basement.studio | portfolio | plays by itself | hard | |
| 0.642 | 0.616 | 0.749 | 0.608 | 0.703 | benxrun.com | portfolio | scroll | hard | |
| 0.691 | 0.674 | 0.650 | 0.559 | 0.635 | berd.xyz | saas | plays by itself | hard | |
| 0.599 | 0.604 | 0.595 | 0.448 | 0.433 | charmling.app | ecommerce | scroll | hard | |
| 0.489 | 0.338 | 0.498 | 0.537 | 0.486 | ciaoenergy.com | ecommerce | plays by itself | hard | |
| 0.475 | 0.546 | 0.462 | 0.489 | 0.521 | ciaoenergy.com | ecommerce | scroll | hard | |
| 0.527 | 0.557 | 0.562 | 0.518 | 0.582 | ciaoenergy.com | ecommerce | scroll | hard | |
| 0.768 | 0.799 | 0.673 | 0.528 | 0.730 | cipher.tv | portfolio | plays by itself | hard | |
| 0.776 | 0.745 | 0.805 | 0.619 | 0.574 | dialkit.dev | app-ui | drag / click | hard | |
| 0.526 | 0.609 | 0.569 | 0.552 | 0.467 | 2025.driftime.com | brand | scroll | hard | |
| 0.890 | 0.888 | 0.844 | 0.719 | 0.673 | gufram.it | ecommerce | plays by itself | hard | |
| 0.725 | 0.572 | 0.465 | 0.384 | 0.392 | kaviengcreative.com | portfolio | drag / click | hard | |
| 0.631 | 0.536 | 0.554 | 0.552 | 0.528 | maximatherapy.com | brand | plays by itself | hard | |
| 0.654 | 0.669 | 0.495 | 0.604 | 0.743 | maximatherapy.com | brand | plays by itself | medium | |
| 0.732 | 0.633 | 0.576 | 0.546 | 0.469 | monopo.london | portfolio | hover | hard | |
| 0.645 | 0.577 | 0.406 | 0.565 | 0.550 | examples.motion.dev | app-ui | drag / click | easy | |
| 0.702 | 0.529 | 0.578 | 0.627 | 0.522 | neutomni.com | portfolio | plays by itself | hard | |
| 0.841 | 0.766 | 0.758 | 0.742 | 0.501 | neutomni.com | portfolio | scroll | hard | |
| 0.606 | 0.575 | 0.353 | 0.448 | 0.575 | otsuka-air.jp | ecommerce | plays by itself | hard | |
| 0.475 | 0.427 | 0.434 | 0.395 | 0.479 | oxigen.sa | brand | scroll | hard | |
| 0.642 | 0.566 | 0.549 | 0.537 | 0.553 | palmo.co.in | ecommerce | scroll | hard | |
| 0.583 | 0.634 | 0.524 | 0.467 | 0.452 | palmo.co.in | ecommerce | scroll | medium | |
| 0.548 | 0.559 | 0.558 | 0.498 | 0.490 | papertiger.com | portfolio | scroll | hard | |
| 0.525 | 0.582 | 0.623 | 0.252 | 0.528 | papumba.com | saas | opens / changes | medium | |
| 0.518 | 0.517 | 0.458 | 0.472 | 0.440 | pixel.melbourne | portfolio | scroll | medium | |
| 0.400 | 0.402 | 0.314 | 0.300 | 0.409 | pixel.melbourne | portfolio | hover | hard | |
| 0.797 | 0.828 | 0.806 | 0.537 | 0.512 | pudding.cool | editorial | scroll | hard | |
| 0.674 | 0.715 | 0.411 | 0.473 | 0.461 | rapidkert.com | brand | scroll | hard | |
| 0.534 | 0.546 | 0.455 | 0.488 | 0.429 | raycast.com | saas | plays by itself | hard | |
| 0.832 | 0.851 | 0.649 | 0.641 | 0.587 | rebelliously-optimistic.com | brand | scroll | hard | |
| 0.850 | 0.887 | 0.601 | 0.636 | 0.617 | rebelliously-optimistic.com | brand | scroll | hard | |
| 0.638 | 0.609 | 0.544 | 0.494 | 0.633 | slowdowncreative.com | portfolio | cursor-follow | medium | |
| 0.290 | 0.297 | 0.322 | 0.347 | 0.326 | slowdowncreative.com | portfolio | scroll | medium | |
| 0.269 | 0.288 | 0.264 | 0.278 | 0.264 | slowdowncreative.com | portfolio | scroll | medium | |
| 0.872 | 0.870 | 0.702 | 0.882 | 0.485 | slowdowncreative.com | portfolio | hover | easy | |
| 0.341 | 0.340 | 0.343 | 0.342 | 0.338 | slowdowncreative.com | portfolio | hover | easy | |
| 0.643 | 0.702 | 0.552 | 0.504 | 0.593 | brand.squarespace.com | brand | hover | medium | |
| 0.586 | 0.594 | 0.574 | 0.527 | 0.427 | truus.co | portfolio | scroll | hard | |
| 0.399 | 0.481 | 0.433 | 0.431 | 0.481 | victorfuruya.com | portfolio | scroll | medium | |
| 0.458 | 0.408 | 0.398 | 0.408 | 0.470 | victorfuruya.com | portfolio | opens / changes | medium | |
| 0.321 | 0.364 | 0.386 | 0.382 | 0.360 | victorfuruya.com | portfolio | plays by itself | medium | |
| 0.429 | 0.482 | 0.452 | 0.391 | 0.299 | victorfuruya.com | portfolio | opens / changes | medium | |
| 0.731 | 0.783 | 0.804 | 0.815 | 0.750 | wisprflow.ai | saas | opens / changes | medium | |
| 0.835 | 0.831 | 0.688 | 0.194 | 0.487 | wisprflow.ai | saas | plays by itself | hard |
Citation
If you use Animation Bench, cite this post as:
@misc{physera2026animationbench, title = {Animation Bench: Evaluating Frontier Models on Web Animation Reconstruction}, author = {Ashwarya Maratha and Tim Cvetko and Himanshu Dubey and Soham Parekh}, year = {2026}, month = sep, howpublished = {Physera}, url = {https://www.physera.ai/research/animation}}Partner with us
Animation Bench is our first research work aimed at closing this gap for the next generation of frontier models, helping them improve on the qualities users actually perceive. If you are working on frontend generation, or environments for agents that build software, we’d love to hear from you. Reach us at hello@physera.ai.