Physera AI

Updated: 01 October 2026v1.0

Animation Bench

Frontier multimodal coding agents can already recreate visually plausible web animations, but current evaluation methods fail to discriminate between screenshot parity and shippable frontend reconstruction. Animation Bench tests whether coding agents can reconstruct a web page's behavior, not just its appearance.

Frontier Models
5
Real Web Animations
48
Reconstructions
240
Live Commercial Websites
32
Leaderboard

Animation Bench

5 models · 48 tasks · 240 reconstructions
RankModelScoreVisualMotionLayoutTask wins$ / task
1GPT-6.1 Sol0.606 ±0.0440.7040.5120.62414$0.46
2GPT-6 Astra0.594 ±0.0450.7100.4730.62817$3.04
3Claude Fable 5.10.548 ±0.0390.6620.4300.5656$3.89
4GPT-6 Sol0.516 ±0.0320.6400.3810.5397$0.45
5Claude Opus 5.50.507 ±0.0380.6310.3830.5174$1.12
  1. 1GPT-6.1 Sol0.606±0.044
    Visual
    0.704
    Motion
    0.512
    Layout
    0.624
    Task wins
    14
    $ / task
    $0.46
  2. 2GPT-6 Astra0.594±0.045
    Visual
    0.710
    Motion
    0.473
    Layout
    0.628
    Task wins
    17
    $ / task
    $3.04
  3. 3Claude Fable 5.10.548±0.039
    Visual
    0.662
    Motion
    0.430
    Layout
    0.565
    Task wins
    6
    $ / task
    $3.89
  4. 4GPT-6 Sol0.516±0.032
    Visual
    0.640
    Motion
    0.381
    Layout
    0.539
    Task wins
    7
    $ / task
    $0.45
  5. 5Claude Opus 5.50.507±0.038
    Visual
    0.631
    Motion
    0.383
    Layout
    0.517
    Task wins
    4
    $ / task
    $1.12

One selected run per model and task. Scores are reproduction scores on a 0–1 scale, not success rates. GPT-6.1 Sol was added on 1 October 2026 with the same tasks, agent and scorer, recorded in a later scoring run; recording conditions alone can move a 48-task mean by about 0.01, so its lead over GPT-6 Astra is within the uncertainty.

Reference vs. 4 agent reconstructions

Frame-locked to the same moments. Hover a model to overlay the reference. For the full task set, see the appendix.

1 / 13
Referenceneutomni.com · scroll
GPT-6.1 Soloverall 0.841motion 0.76
GPT-6 Astraoverall 0.766motion 0.56
Claude Fable 5.1overall 0.758motion 0.54
Claude Opus 5.5overall 0.742motion 0.57
GPT-6 Soloverall 0.501motion 0.39
Neutomni rolling shapeneutomni.com · scroll · 14 frames · 9.2 s
frame 1/14 · t = 0.0 s

Background

Frontend generation is a common commercial use case for coding agents. Current agents can reproduce much of a page’s visual palette, typography, and static layout. A production page, however, is not a single frame. It also contains transitions and interactions triggered by scrolling, hovering, clicking, dragging, and cursor movement. Can frontier models reconstruct those behaviours as well as the static appearance?

We built Animation Bench to answer that question and tested four frontier models on 48 tasks from real websites. Each model ran in its own sandbox with the same task inputs and had to deliver a self-contained HTML file. Across the four models, mean visual similarity ranges from 0.63 to 0.71, while mean motion consistency ranges from 0.38 to 0.47.

As we enter the era of recursive self-improvement (RSI), we need to pay more attention to the verifiers used to measure progress. Left unchecked, model performance may saturate on tasks in domains with easily verifiable outcomes. Animation Bench applies this scrutiny to web animation.

Design philosophy

We set out to measure end-to-end animation reconstruction. The benchmark follows five design principles:

  1. Reproduction > Generation: Score of each task is measured directly and objectively against the real animation.
  2. Frames in, motion out: Each model received between 12 and 24 frames, depending on the task, and network capture of the page. It did not receive the site’s source code or a description of the timing.
  3. Real sites selected for animation coverage:
    1. We sourced animations from commercial sites spanning editorial, portfolio, e-commerce, product, and brand work.
    2. The 48 animations cover the common trigger types: autoplay, scroll, hover, cursor-follow, click, drag, and state change.
  4. Separate scoring axes: Three axes measure complementary aspects of reconstruction quality: visual similarity, motion consistency, and layout correctness. A reconstruction can match the layout while missing the motion, which a single undifferentiated score would obscure.
  5. Test the running artifact: We open each reconstruction in a real browser, drive it as a user would, and record it frame by frame.

Methodology

We used Computer-1 using the Harbor framework to run 192 evaluations: four models on the same 48 tasks, all at maximum reasoning effort. GPT-6.1 Sol was added later in the same setup, for 240 in total. Each model ran in its own sandbox with a 1280×720 desktop, a shell, the task's reference frames, and the page's network capture. The required output was one self-contained HTML file. The reported results contain one selected run for each model and task pair.

  1. 01Input12–24 frames+ the page’s network capture
  2. 02AgentComputer-1run through Harbor · max reasoning
  3. 03Sandbox1280×720desktop + bash shell
  4. 04Outputindex.htmlone self-contained file
  5. 05Scored3 axesvisual · motion · layout
5 models × 48 tasks = 240 runs · one selected run per pair

Tasks

Each task is one precise animation on a production website with a fixed start and end state.

Input: The model receives 12 to 24 still frames sampled across the animation, with capture timestamps where available, and a HAR capture containing the HTML, CSS, JavaScript, fonts, and images loaded by the live page. It does not receive the site’s source project, a video, or any timing description beyond the frames and their timestamps.

Output: The required output is one self-contained index.html file with inline CSS and JavaScript and no external requests.

Evaluation: We open the file in a headless browser at 1280×720, drive it with the task’s trigger (wait, scroll, hover, drag, click), capture it frame by frame the same way we captured the original, and compare the two captures.

The 48 tasks come from 32 sites and were selected to cover different triggers and animation properties. Each task is tagged by trigger, animated property, timing pattern, spatial extent, and site category.

TriggerTasksWhat the model has to recover
Scroll21 tasksProgress tied to scroll position: pinned sequences, scroll-driven text, scroll-linked transforms
Plays by itself12 tasksLoad-in entrances, ambient loops, auto-cycling scenes
Hover6 tasksReveals and state changes on pointer enter, and their reversal on leave
Click / drag / key4 tasksDrag carousels, dial scrubbing, spring physics
Opens or changes4 tasksPage transitions, view switches, expand-collapse
Cursor-follow1 taskA contextual cursor that changes over specific elements

Scoring

Each result is compared with the original recording frame by frame. We score three axes, each composed of sub-scores that measure different failure modes.

Visual similarity: does it look right at a given moment?

Each reconstruction frame is compared with the reference frame at the same moment. Five sub-scores are averaged with fixed weights:

Sub-scoreWeightHow it is computedWhat it catches
MS-SSIM0.32Structural similarity of each frame pair, at several scalesShapes in the wrong places; layout drift
LPIPS0.321 − perceptual distance between each frame pair (a learned metric)Whether a person would say the frames look alike
Colour0.13Overlap of the two foreground colour histogramsWrong palette
Edges0.08F1 of the two frames’ edge mapsRight colour, wrong geometry
Coverage0.15Ratio of foreground fill: min(fR, fC) / max(fR, fC)Missing or extra blocks; blank or letterboxed pages

Motion consistency: does it move right over time?

The motion score compares how the two recordings change over time. Three sub-scores describe the pattern of motion; two penalties then scale the result down when the amount or the placement of that motion is wrong:

Sub-scoreWeightHow it is computedWhat it catches
Energy0.50½ timing + ½ burstiness of the frame-to-frame pixel-change curve. Timing is the correlation of the two normalised curves; burstiness is the ratio of their coefficients of variationMotion at the wrong moments; a smooth fade where the original snaps, or the reverse
Flow0.20The same two terms on the optical-flow magnitude curveReal movement vs fades and flicker
Trajectory0.301 − RMS distance between the normalised cumulative paths of the moving region’s centroidWrong direction or order of movement
G_amountpenaltymin(r, 1/r)^0.35, where r is the ratio of the reconstruction's total motion to the reference'sToo little or too much motion overall
G_placementpenalty(fill ratio)^0.35: the reconstruction's occupied screen area relative to the originalContent in a corner, letterboxed, or missing

Layout correctness: is it built like the original?

The layout score runs OCR on both recordings and compares the words it finds, frame by frame:

Sub-scoreWeightHow it is computedWhat it catches
Presence0.35F1 of OCR words matched between the two frames (a match allows up to 30% character error)Words missing or invented
Accuracy0.251 − character error rate over the matched wordsMisspelt copy
Order0.151 − 2 · inversions / n(n − 1) of the matched words, top to bottomWrong reading order
Alignment0.25Σ IoU of matched word boxes / (matched + unmatched)Words in the wrong positions

Overall score

The overall score is a weighted mean of the three axes. The weights depend on what triggers the animation. Interaction is scored separately and held out, so the three weights are renormalised. Canvas-heavy tasks set the layout weight to 0.05, since OCR cannot see into a canvas.

TriggerVisualMotionLayoutInteraction (held out)
autoplay0.350.350.200.10
scroll0.300.350.200.15
hover0.300.200.200.30
cursor0.300.200.150.35
gesture0.300.200.200.30
state-change0.300.250.200.25

All scores are reproduction scores on a 0–1 scale. The code also applies a nominal 0.99 ceiling to the overall score; it never binds (the highest score in the set is 0.890).

Scoring example: oxigen-voxel-palm-pinned

The following reconstruction from oxigen.sa shows how the scoring system works end to end. In the original, a palm tree made of glowing voxels grows over a voxel landscape while the section remains pinned and the copy changes during scrolling. Opus 5.5 produced a recognisable but substantially different result.

The original's palm assembles, grows and fills the frame as you scroll. Opus 5.5 draws a cyan fountain that barely changes, and its copy scrolls up under the logo instead of staying pinned.

Visual similarity: 0.453

Sub-scoreWeightScoreWeight × scorecontribution
MS-SSIM0.320.420.134
LPIPS0.320.340.109
Colour0.130.770.100
Edges0.080.290.023
Coverage0.150.580.087
Visual similarity0.4530.453

Results

Dimensions

0.20.40.60.81.0Visual similarity0–1Motion consistency0–1Layout correctness0–1
Visual similarity
Sol 6.10.704
Astra0.710
Fable 5.10.662
Sol0.640
Opus 5.50.631
Motion consistency
Sol 6.10.512
Astra0.473
Fable 5.10.430
Sol0.381
Opus 5.50.383
Layout correctness
Sol 6.10.624
Astra0.628
Fable 5.10.565
Sol0.539
Opus 5.50.517

Across all five models, mean visual similarity exceeds mean motion consistency. At the task level, visual exceeds motion in 219 of 240 reconstructions. The average gap is 0.23, and the two are only moderately correlated (r = 0.47). A page with high visual similarity is only somewhat more likely to have high motion consistency.

GPT-6.1 Sol42 / 48 tasks score higher on visual than motion
0.000.000.250.250.500.500.750.751.001.00visual = motionadcker-menu-services-hover: visual 0.621, motion 0.544altitude101-glass-ring-word-swap: visual 0.561, motion 0.453altitude101-words-scroll-rates: visual 0.705, motion 0.609ausify-vibe-canvas-carousel: visual 0.768, motion 0.473basement-studio-graffiti-hero: visual 0.553, motion 0.557benxrun-skyline-chapter-scroll: visual 0.847, motion 0.423berd-window-morphs-into-app: visual 0.880, motion 0.606charmling-99-charms-flythrough: visual 0.653, motion 0.547ciaoenergy-cans-fan-scroll-spin: visual 0.685, motion 0.247ciaoenergy-cans-sideways-selection: visual 0.602, motion 0.361ciaoenergy-text-dancing-scroll: visual 0.723, motion 0.373cipher-loader-stills-ring: visual 0.761, motion 0.761dialkit-dials-shape-headline: visual 0.889, motion 0.566driftime-2025-pinned-scroll-morph: visual 0.612, motion 0.468gufram-zero-gravity-collage-hero: visual 0.884, motion 0.868kavieng-cards-fly-to-grid-drag: visual 0.754, motion 0.673maxima-splash-curtain-whale-part2: visual 0.680, motion 0.564maxima-splash-curtain-whale-scene: visual 0.767, motion 0.593monopo-london-webgl-sections: visual 0.762, motion 0.675motion-dev-animation: visual 0.764, motion 0.801neutomni-preloader-cut-along-line: visual 0.834, motion 0.616neutomni-process-rolling-shape: visual 0.901, motion 0.764otsuka-zeroz-intro-reveal: visual 0.757, motion 0.458oxigen-voxel-palm-pinned: visual 0.614, motion 0.340palmo-coconut-crack-scroll: visual 0.769, motion 0.554palmo-pure-fresh-clean-words: visual 0.679, motion 0.460papertiger-card-stack-to-fullbleed: visual 0.599, motion 0.542papumba-play-explore-ipad-transition: visual 0.557, motion 0.481pixel-melbourne-crafty-bunch-scroll: visual 0.558, motion 0.446pixel-melbourne-menu-directors-hover: visual 0.397, motion 0.424pudding-essential-words-pinned-cloud: visual 0.883, motion 0.636rapidkert-soil-dive-pinned: visual 0.662, motion 0.671raycast-animation: visual 0.750, motion 0.347rebelliously-optimistic-four-commitments: visual 0.915, motion 0.744rebelliously-optimistic-pinned-hero: visual 0.910, motion 0.881slowdown-featured-work-view-work-cursor: visual 0.535, motion 0.791slowdown-footer-services-rolling-labels: visual 0.450, motion 0.047slowdown-footer-slow-down-reveal: visual 0.393, motion 0.191slowdown-nav-hover-bullets: visual 0.912, motion 0.744slowdown-process-experience-reveal: visual 0.531, motion 0.206squarespace-brand-logo-hover-reveal: visual 0.690, motion 0.675truus-letters-scatter-along-path: visual 0.799, motion 0.303victor-furuya-core-values-scroll: visual 0.595, motion 0.234victor-furuya-make-it-matter-collapse: visual 0.690, motion 0.417victor-furuya-manifesto-text: visual 0.567, motion 0.000victor-furuya-work-index-transition: visual 0.439, motion 0.314wisprflow-dictation-notetaker-toggle: visual 0.968, motion 0.271wisprflow-hero-text-ribbons: visual 0.942, motion 0.836visual similaritymotion consistency
GPT-6 Astra42 / 48 tasks score higher on visual than motion
0.000.000.250.250.500.500.750.751.001.00visual = motionadcker-menu-services-hover: visual 0.571, motion 0.000altitude101-glass-ring-word-swap: visual 0.594, motion 0.504altitude101-words-scroll-rates: visual 0.750, motion 0.536ausify-vibe-canvas-carousel: visual 0.804, motion 0.508basement-studio-graffiti-hero: visual 0.481, motion 0.405benxrun-skyline-chapter-scroll: visual 0.848, motion 0.362berd-window-morphs-into-app: visual 0.899, motion 0.544charmling-99-charms-flythrough: visual 0.661, motion 0.550ciaoenergy-cans-fan-scroll-spin: visual 0.534, motion 0.101ciaoenergy-cans-sideways-selection: visual 0.730, motion 0.356ciaoenergy-text-dancing-scroll: visual 0.777, motion 0.379cipher-loader-stills-ring: visual 0.766, motion 0.808dialkit-dials-shape-headline: visual 0.930, motion 0.298driftime-2025-pinned-scroll-morph: visual 0.687, motion 0.554gufram-zero-gravity-collage-hero: visual 0.902, motion 0.845kavieng-cards-fly-to-grid-drag: visual 0.653, motion 0.324maxima-splash-curtain-whale-part2: visual 0.677, motion 0.380maxima-splash-curtain-whale-scene: visual 0.814, motion 0.491monopo-london-webgl-sections: visual 0.730, motion 0.460motion-dev-animation: visual 0.764, motion 0.724neutomni-preloader-cut-along-line: visual 0.673, motion 0.371neutomni-process-rolling-shape: visual 0.937, motion 0.555otsuka-zeroz-intro-reveal: visual 0.656, motion 0.438oxigen-voxel-palm-pinned: visual 0.440, motion 0.431palmo-coconut-crack-scroll: visual 0.750, motion 0.422palmo-pure-fresh-clean-words: visual 0.747, motion 0.495papertiger-card-stack-to-fullbleed: visual 0.605, motion 0.557papumba-play-explore-ipad-transition: visual 0.623, motion 0.574pixel-melbourne-crafty-bunch-scroll: visual 0.557, motion 0.443pixel-melbourne-menu-directors-hover: visual 0.397, motion 0.430pudding-essential-words-pinned-cloud: visual 0.913, motion 0.687rapidkert-soil-dive-pinned: visual 0.698, motion 0.703raycast-animation: visual 0.765, motion 0.368rebelliously-optimistic-four-commitments: visual 0.937, motion 0.847rebelliously-optimistic-pinned-hero: visual 0.973, motion 0.897slowdown-featured-work-view-work-cursor: visual 0.532, motion 0.705slowdown-footer-services-rolling-labels: visual 0.438, motion 0.076slowdown-footer-slow-down-reveal: visual 0.400, motion 0.232slowdown-nav-hover-bullets: visual 0.943, motion 0.688slowdown-process-experience-reveal: visual 0.548, motion 0.178squarespace-brand-logo-hover-reveal: visual 0.752, motion 0.780truus-letters-scatter-along-path: visual 0.840, motion 0.316victor-furuya-core-values-scroll: visual 0.604, motion 0.455victor-furuya-make-it-matter-collapse: visual 0.759, motion 0.195victor-furuya-manifesto-text: visual 0.676, motion 0.000victor-furuya-work-index-transition: visual 0.427, motion 0.456wisprflow-dictation-notetaker-toggle: visual 0.970, motion 0.424wisprflow-hero-text-ribbons: visual 0.963, motion 0.834visual similaritymotion consistency
Claude Fable 5.144 / 48 tasks score higher on visual than motion
0.000.000.250.250.500.500.750.751.001.00visual = motionadcker-menu-services-hover: visual 0.704, motion 0.425altitude101-glass-ring-word-swap: visual 0.565, motion 0.423altitude101-words-scroll-rates: visual 0.633, motion 0.591ausify-vibe-canvas-carousel: visual 0.712, motion 0.507basement-studio-graffiti-hero: visual 0.491, motion 0.519benxrun-skyline-chapter-scroll: visual 0.942, motion 0.548berd-window-morphs-into-app: visual 0.879, motion 0.501charmling-99-charms-flythrough: visual 0.602, motion 0.583ciaoenergy-cans-fan-scroll-spin: visual 0.675, motion 0.291ciaoenergy-cans-sideways-selection: visual 0.590, motion 0.333ciaoenergy-text-dancing-scroll: visual 0.675, motion 0.490cipher-loader-stills-ring: visual 0.748, motion 0.552dialkit-dials-shape-headline: visual 0.898, motion 0.562driftime-2025-pinned-scroll-morph: visual 0.681, motion 0.497gufram-zero-gravity-collage-hero: visual 0.841, motion 0.793kavieng-cards-fly-to-grid-drag: visual 0.586, motion 0.349maxima-splash-curtain-whale-part2: visual 0.626, motion 0.443maxima-splash-curtain-whale-scene: visual 0.660, motion 0.335monopo-london-webgl-sections: visual 0.694, motion 0.340motion-dev-animation: visual 0.739, motion 0.000neutomni-preloader-cut-along-line: visual 0.754, motion 0.428neutomni-process-rolling-shape: visual 0.923, motion 0.538otsuka-zeroz-intro-reveal: visual 0.461, motion 0.187oxigen-voxel-palm-pinned: visual 0.485, motion 0.416palmo-coconut-crack-scroll: visual 0.680, motion 0.468palmo-pure-fresh-clean-words: visual 0.566, motion 0.526papertiger-card-stack-to-fullbleed: visual 0.490, motion 0.615papumba-play-explore-ipad-transition: visual 0.694, motion 0.611pixel-melbourne-crafty-bunch-scroll: visual 0.539, motion 0.352pixel-melbourne-menu-directors-hover: visual 0.381, motion 0.350pudding-essential-words-pinned-cloud: visual 0.901, motion 0.648rapidkert-soil-dive-pinned: visual 0.519, motion 0.286raycast-animation: visual 0.649, motion 0.272rebelliously-optimistic-four-commitments: visual 0.823, motion 0.576rebelliously-optimistic-pinned-hero: visual 0.889, motion 0.444slowdown-featured-work-view-work-cursor: visual 0.444, motion 0.660slowdown-footer-services-rolling-labels: visual 0.458, motion 0.117slowdown-footer-slow-down-reveal: visual 0.393, motion 0.181slowdown-nav-hover-bullets: visual 0.722, motion 0.650slowdown-process-experience-reveal: visual 0.545, motion 0.196squarespace-brand-logo-hover-reveal: visual 0.649, motion 0.487truus-letters-scatter-along-path: visual 0.748, motion 0.365victor-furuya-core-values-scroll: visual 0.586, motion 0.294victor-furuya-make-it-matter-collapse: visual 0.652, motion 0.284victor-furuya-manifesto-text: visual 0.734, motion 0.000victor-furuya-work-index-transition: visual 0.365, motion 0.477wisprflow-dictation-notetaker-toggle: visual 0.953, motion 0.513wisprflow-hero-text-ribbons: visual 0.851, motion 0.597visual similaritymotion consistency
Claude Opus 5.546 / 48 tasks score higher on visual than motion
0.000.000.250.250.500.500.750.751.001.00visual = motionadcker-menu-services-hover: visual 0.610, motion 0.306altitude101-glass-ring-word-swap: visual 0.532, motion 0.502altitude101-words-scroll-rates: visual 0.512, motion 0.388ausify-vibe-canvas-carousel: visual 0.707, motion 0.367basement-studio-graffiti-hero: visual 0.554, motion 0.534benxrun-skyline-chapter-scroll: visual 0.799, motion 0.426berd-window-morphs-into-app: visual 0.786, motion 0.361charmling-99-charms-flythrough: visual 0.544, motion 0.357ciaoenergy-cans-fan-scroll-spin: visual 0.633, motion 0.408ciaoenergy-cans-sideways-selection: visual 0.610, motion 0.390ciaoenergy-text-dancing-scroll: visual 0.718, motion 0.349cipher-loader-stills-ring: visual 0.622, motion 0.404dialkit-dials-shape-headline: visual 0.716, motion 0.475driftime-2025-pinned-scroll-morph: visual 0.621, motion 0.509gufram-zero-gravity-collage-hero: visual 0.801, motion 0.514kavieng-cards-fly-to-grid-drag: visual 0.538, motion 0.173maxima-splash-curtain-whale-part2: visual 0.594, motion 0.503maxima-splash-curtain-whale-scene: visual 0.724, motion 0.464monopo-london-webgl-sections: visual 0.682, motion 0.291motion-dev-animation: visual 0.769, motion 0.646neutomni-preloader-cut-along-line: visual 0.702, motion 0.569neutomni-process-rolling-shape: visual 0.899, motion 0.567otsuka-zeroz-intro-reveal: visual 0.563, motion 0.299oxigen-voxel-palm-pinned: visual 0.453, motion 0.362palmo-coconut-crack-scroll: visual 0.698, motion 0.418palmo-pure-fresh-clean-words: visual 0.578, motion 0.496papertiger-card-stack-to-fullbleed: visual 0.431, motion 0.515papumba-play-explore-ipad-transition: visual 0.531, motion 0.000pixel-melbourne-crafty-bunch-scroll: visual 0.546, motion 0.382pixel-melbourne-menu-directors-hover: visual 0.552, motion 0.045pudding-essential-words-pinned-cloud: visual 0.525, motion 0.441rapidkert-soil-dive-pinned: visual 0.526, motion 0.427raycast-animation: visual 0.649, motion 0.341rebelliously-optimistic-four-commitments: visual 0.789, motion 0.615rebelliously-optimistic-pinned-hero: visual 0.865, motion 0.553slowdown-featured-work-view-work-cursor: visual 0.426, motion 0.499slowdown-footer-services-rolling-labels: visual 0.497, motion 0.133slowdown-footer-slow-down-reveal: visual 0.385, motion 0.218slowdown-nav-hover-bullets: visual 0.917, motion 0.776slowdown-process-experience-reveal: visual 0.539, motion 0.199squarespace-brand-logo-hover-reveal: visual 0.615, motion 0.388truus-letters-scatter-along-path: visual 0.742, motion 0.409victor-furuya-core-values-scroll: visual 0.598, motion 0.317victor-furuya-make-it-matter-collapse: visual 0.759, motion 0.195victor-furuya-manifesto-text: visual 0.724, motion 0.000victor-furuya-work-index-transition: visual 0.394, motion 0.258wisprflow-dictation-notetaker-toggle: visual 0.897, motion 0.615wisprflow-hero-text-ribbons: visual 0.414, motion 0.000visual similaritymotion consistency
GPT-6 Sol45 / 48 tasks score higher on visual than motion
0.000.000.250.250.500.500.750.751.001.00visual = motionadcker-menu-services-hover: visual 0.597, motion 0.244altitude101-glass-ring-word-swap: visual 0.593, motion 0.415altitude101-words-scroll-rates: visual 0.775, motion 0.609ausify-vibe-canvas-carousel: visual 0.701, motion 0.493basement-studio-graffiti-hero: visual 0.612, motion 0.405benxrun-skyline-chapter-scroll: visual 0.869, motion 0.569berd-window-morphs-into-app: visual 0.873, motion 0.469charmling-99-charms-flythrough: visual 0.554, motion 0.319ciaoenergy-cans-fan-scroll-spin: visual 0.708, motion 0.247ciaoenergy-cans-sideways-selection: visual 0.610, motion 0.451ciaoenergy-text-dancing-scroll: visual 0.824, motion 0.385cipher-loader-stills-ring: visual 0.790, motion 0.636dialkit-dials-shape-headline: visual 0.701, motion 0.339driftime-2025-pinned-scroll-morph: visual 0.477, motion 0.463gufram-zero-gravity-collage-hero: visual 0.839, motion 0.368kavieng-cards-fly-to-grid-drag: visual 0.578, motion 0.141maxima-splash-curtain-whale-part2: visual 0.645, motion 0.355maxima-splash-curtain-whale-scene: visual 0.765, motion 0.810monopo-london-webgl-sections: visual 0.636, motion 0.140motion-dev-animation: visual 0.766, motion 0.475neutomni-preloader-cut-along-line: visual 0.678, motion 0.489neutomni-process-rolling-shape: visual 0.522, motion 0.387otsuka-zeroz-intro-reveal: visual 0.729, motion 0.396oxigen-voxel-palm-pinned: visual 0.634, motion 0.344palmo-coconut-crack-scroll: visual 0.757, motion 0.388palmo-pure-fresh-clean-words: visual 0.544, motion 0.462papertiger-card-stack-to-fullbleed: visual 0.497, motion 0.486papumba-play-explore-ipad-transition: visual 0.554, motion 0.495pixel-melbourne-crafty-bunch-scroll: visual 0.545, motion 0.346pixel-melbourne-menu-directors-hover: visual 0.407, motion 0.438pudding-essential-words-pinned-cloud: visual 0.525, motion 0.410rapidkert-soil-dive-pinned: visual 0.524, motion 0.393raycast-animation: visual 0.623, motion 0.246rebelliously-optimistic-four-commitments: visual 0.738, motion 0.461rebelliously-optimistic-pinned-hero: visual 0.873, motion 0.502slowdown-featured-work-view-work-cursor: visual 0.537, motion 0.771slowdown-footer-services-rolling-labels: visual 0.457, motion 0.135slowdown-footer-slow-down-reveal: visual 0.392, motion 0.182slowdown-nav-hover-bullets: visual 0.575, motion 0.127slowdown-process-experience-reveal: visual 0.533, motion 0.192squarespace-brand-logo-hover-reveal: visual 0.695, motion 0.480truus-letters-scatter-along-path: visual 0.569, motion 0.243victor-furuya-core-values-scroll: visual 0.555, motion 0.476victor-furuya-make-it-matter-collapse: visual 0.758, motion 0.372victor-furuya-manifesto-text: visual 0.739, motion 0.000victor-furuya-work-index-transition: visual 0.377, motion 0.000wisprflow-dictation-notetaker-toggle: visual 0.943, motion 0.358wisprflow-hero-text-ribbons: visual 0.551, motion 0.388visual similaritymotion consistency
Each dot is one page a model built. Dots below the diagonal look better than they move: 219 of 240 do.

Within this sample, cost has little association with score. Mean spend per task ranges from $0.45 for GPT-6 Sol to $3.89 for Fable 5.1, an 8.6-fold difference, while mean overall score spans 0.099. Per-model Spearman correlations between spend and score range from −0.24 to +0.24.

Same model, cheap tasks vs expensive tasks

Spending 1.5–1.9× more on a task moved the score by at most 0.09, and not in one direction.

ModelΔ score
GPT-6.1 Sol$0.35 → $0.57 per task · ρ +0.05−0.002
GPT-6 Astra$2.20 → $3.88 per task · ρ +0.10+0.032
Claude Fable 5.1$2.66 → $5.12 per task · ρ +0.24+0.085
Claude Opus 5.5$0.84 → $1.40 per task · ρ −0.12−0.061
GPT-6 Sol$0.37 → $0.54 per task · ρ −0.24−0.063
cheaper half of the model's 48 taskspricier halfaxis: mean overall score
Each model’s 48 tasks split at its median spend. The hollow dot is the mean overall score of the cheaper half, the filled dot the pricier half. Two models did slightly better on the tasks they spent more on, three did slightly worse (GPT-6.1 Sol by 0.002); none moved by more than 0.09.

Where animation reconstructions fail

The final results indicate that motion is the gap. We deeply investigated all 192 generated pages and replayed a subset side by side with the reference. Each failure below is observable in the output, countable across the set, and has a named example to follow along.

Motion timing and sequencing break down

Models do well at capturing motion location (the location gate averages 0.88). However, across the 183 reconstructions in which motion was captured, the timing term averages 0.57, compared with 0.50 when each reconstruction is paired with the reference from a different task. The models often identify which parts of the page should move but reproduce the timing only weakly.

Instead, the motion tends to arrive all at once. In a typical reconstruction the single biggest change between two frames accounts for 29% of all its movement; in the original it is 19%.

Nine of the 32 intros that should play once were implemented as loops that restart indefinitely. On one eight-second sequence, three of the four models completed the full sequence within about one second.

Reference
GPT-6.1 Soloverall 0.841motion 0.76
GPT-6 Astraoverall 0.766motion 0.56
neutomni.com (opens in a new tab) As you scroll, a white outline rolls along a track above four red cards, turning from a square into a pentagon and then a circle. GPT-6.1 Sol fills the full width and turns the shape at the same scroll positions as the original, reaching each card in step (motion 0.76). GPT-6 Astra keeps pace with the reference (motion 0.56).
Case study · Wispr Flow

Reconstructing a timed UI sequence

On wisprflow.ai, a two-option pill (Dictation | Notetaker) runs one sequence:

  1. The white thumb sits on “Dictation”.
  2. Sliding to “Notetaker” stretches to the width of the label.
  3. As it lands, the letters of “Notetaker” ripple: each lifts and drops in turn, left to right.
  4. Easing back to “Dictation”, the letters stay still.

Every model recognised the component and reproduced its appearance (visual 0.90–0.97, layout ≈ 0.95 for all four). Two models also recognised the ripple and implemented a suitable mechanism: Claude Opus 5.5 used a per-letter @keyframes wave with a 70 ms stagger, while Claude Fable 5.1 used a per-character transform sequence. Each model reconstructed a different part of the interaction:

Slide on time (f1)Ripple after slideReturns on time (f10)
GPT-6.1 Solearly, on hover before the clickon hover, before the slide
Claude Opus 5.5fires immediatelynever returns
Claude Fable 5.1starts already switchedshort faint
GPT-6 Sol2 frames latenone1 frame late
GPT-6 Astra7 frames latebarelynever returns

What makes this case so hard? The pill is small, and each letter of the ripple lifts by only a few pixels. The stills are unevenly spaced: the first is taken almost seven seconds into the recording, a full second passes before the slide, eight quick frames about 150 ms apart catch the ripple, and a second and a half passes before the return. To rebuild it, a model has to read the timestamps as well as the pictures, and turn a few pixels of difference into a sequence with an order, a pause and a return. No single frame shows any of that.

GPT-6 Sol reproduces the component, but not its timing. Its page slides the thumb 0.9 seconds after load, then flips it back and forth every 3.3 seconds, indefinitely. There is no ripple at all: the letters are never split apart, so they cannot move one at a time. Every frame of it looks right (visual 0.94, layout 0.95), and it scores 0.36 on motion.

Reference
GPT-6 Astraoverall 0.783motion 0.42
Claude Fable 5.1overall 0.804motion 0.51
Claude Opus 5.5overall 0.815motion 0.61
GPT-6 Soloverall 0.750motion 0.36
Wispr Flow toggle, the reference and all four models. Opus 5.5 slides on time and never returns (motion 0.62); Fable 5.1 starts already switched and returns on time (0.51); Sol and Astra slide late, and neither ripples as the original does (0.36, 0.42).

Motion is incomplete or missing

Reconstructions more often move too little than too much. The median reconstruction carries 0.63 of the reference’s motion, and 38% carry less than half; 10 of 192 overshoot by more than double. GPT-6 Sol is the most restrained, at a median of 0.56. Combined with the timeline finding above, the typical rebuild is a correct-looking page that moves less, and in fewer, larger steps.

Reference
GPT-6 Soloverall 0.521motion 0.45
ciaoenergy.com (opens in a new tab) As the page scrolls, the reference steps the can and its copy through five flavours and ends on a full-screen “ZERO BULLSHIT” frame. GPT-6 Sol reproduces the flavour sequence but lags behind the reference and never reaches the final frame (overall 0.521, visual 0.610, motion 0.451).

Staggered sequences collapse into simultaneous motion

Staggered choreography, in which elements enter one after another, is the most common timing pattern in the set (32 tasks). In 34 of 128 reconstructions of staggered tasks, we found no delay or stagger construct: every element starts together. GPT-6 Sol accounts for 16 of the 34. Under the current scoring, the within-task penalty is small (0.02 on motion), so this is more evident in the code than in the aggregate score.

The same Wispr toggle, rebuilt by GPT-6 Sol. The thumb slides, but the letters of Notetaker change as one block, with no ripple (motion 0.36).

Complex visual assets are approximated

The most expensive element of a commercial animation is often its hero asset: a WebGL scene, a 3D product, or a photographic sequence. Fifteen of the 16 canvas tasks come from sites that use WebGL or three.js. GPT-6 Astra used WebGL on 5 of the 16 tasks by inlining the site’s captured three.js code; the other models rebuilt these scenes with Canvas2D.

Reference
GPT-6 Astraoverall 0.338motion 0.10
Claude Fable 5.1overall 0.498motion 0.29
ciaoenergy.com (opens in a new tab) The reference forms a diagonal sweep of cans. GPT-6 Astra preserves some details of the original renders but scores 0.338 overall (visual 0.534, motion 0.101). Claude Fable 5.1 shows six cans that do not form the reference arrangement and scores 0.498 overall (visual 0.675, motion 0.291).

None of the four reconstructions reproduces the full sweep. Claude Opus 5.5 has the highest overall score at 0.537, followed by Fable 5.1 at 0.498, GPT-6 Sol at 0.486, and GPT-6 Astra at 0.338. Sol has the highest visual score (0.708), while Opus has the highest motion score (0.408).

Page structure, behaviour, and text drift from the reference

Eleven of the 52 reconstructions we inspected visually added dark side bars that do not appear in the reference. These pages are letterboxed into a fixed-aspect column instead of filling the viewport (Sol 6, Opus 5.5 3, Fable 5.1 2). At pudding.cool (opens in a new tab) Astra and Fable 5.1 are close to the pinned word cloud reference; Opus 5.5 adds fixed black bars on both sides, while GPT-6 Sol runs ahead of the scroll inside a letterbox.

pudding.cool (opens in a new tab) Claude Fable 5.1 fills the page like the original (visual 0.90); GPT-6 Sol squeezes it into a column between dark side bars (visual 0.53, coverage 0.30).

Secondary interactions and states are omitted

Several reconstructions implement the headline behaviour and drop what surrounds it:

  • Pinned sections that do not pin. The content scrolls past instead of holding while the animation plays (3 of 52 inspected). On the slowdown footer, Astra holds the block while the icons rotate; Sol and Fable 5.1 let it scroll away.
  • Incomplete dragging effect. On kaviengcreative.com (opens in a new tab), it shows the original interaction: dragging the cards should fly them into a grid. Astra assembles the grid; Fable 5.1 brings the cards forward but never settles them into it; GPT-6 Sol fades the title but the cards never assemble; Opus 5.5 does nothing on drag.
Slowdown footer: reference row and four model rows; pinned block with rotating icons.
Slowdown footer: pinned block with rotating icons
kaviengcreative.com (opens in a new tab) Dragging should fly the cards into a grid. GPT-6 Astra assembles it (overall 0.57); Claude Opus 5.5 leaves the page still under the drag (overall 0.38).

Text is present but misplaced

The models did not invent new copy, and none of the reconstructions contains placeholder text. The visible text comes from the captured page, but it is often displayed at the wrong time or position. Across the sampled frames, roughly 41% of the reference text labels never appear on screen in the reconstruction, and 13% are misspelled. Of the text that a reconstruction does display, 38% has no counterpart in the corresponding reference frame. Bounding-box alignment, which measures whether words occupy the same positions, averages 0.21, the lowest term on any axis.

The visual axis shows the same split between palette and geometry. Colour agreement averages 0.85; edge agreement (whether outlines and borders line up) averages 0.26 and is the weakest visual term in 184 of 192 reconstructions. The models get the palette but don’t get shapes.

dialkit.dev (opens in a new tab) Claude Fable 5.1 sets the headline at the original’s size and position (layout 0.91); Claude Opus 5.5 has the same words, smaller and lighter, so they no longer sit where the original’s do (layout 0.62, box alignment 0.00).

Screenshot replay replaces real reconstruction

Fourteen reconstructions solved the task by embedding the reference frames themselves as images and stepping through them on a timer, on scroll, or on hover (Sol 10, Astra 3, Opus 5.5 1). It is the purest form of screenshot mimicry: correct at twelve instants by construction, and wrong everywhere between them. Within the same task, flipbooks score 0.08 lower on motion than reconstructions that rebuild the animation, and slightly lower overall.

brand.squarespace.com (opens in a new tab) GPT-6 Sol’s page is sixteen stored screenshots swapped on a timer (motion 0.48); GPT-6 Astra animates the reveal itself (motion 0.78).

Conclusion

Implications

For model labs. The bottleneck to full webpage reproduction is neither perception nor code generation. Frontier models frequently select an appropriate implementation technique but fail to reproduce the complete animation sequence. They misjudge timing, overlap, and duration.

For benchmarks & RL environments. Current benchmarks often grade reconstructed webpages primarily by static visual similarity. Animation benchmarks should also evaluate whether models preserve temporal structure, including duration, overlap, pauses, and event order. A mechanical three-axis scoring system could support a training environment for these behaviours.

Final thoughts

Can frontier models rebuild a web animation rather than only its first frame? Not yet. They reproduce much of its palette, typography, and layout, and they often identify an appropriate implementation technique. What they fail to recover is time: event order, pauses, movement duration, and whether the page returns to its initial state. Every model scored lower on motion than on visual similarity. Screenshot-based evaluation does not capture this temporal gap, but users notice it within seconds.

Appendix: Tasks

All 48 tasks with each model’s overall score, site, trigger and difficulty. One selected generation per task and model, scored against one reference capture; the best score on each task is in bold.

TaskSol 6.1AstraFableOpusSolSiteGenreTriggerDifficulty
0.5700.3850.6180.4970.465adcker.comportfoliohovermedium
0.5240.5410.4990.5180.510altitude101.comportfolioscrollhard
0.6510.6350.6060.4510.681altitude101.comportfolioscrollhard
0.6450.6970.6130.5750.611ausify.com.ausaasdrag / clickhard
0.5840.4780.5130.5700.540basement.studioportfolioplays by itselfhard
0.6420.6160.7490.6080.703benxrun.comportfolioscrollhard
0.6910.6740.6500.5590.635berd.xyzsaasplays by itselfhard
0.5990.6040.5950.4480.433charmling.appecommercescrollhard
0.4890.3380.4980.5370.486ciaoenergy.comecommerceplays by itselfhard
0.4750.5460.4620.4890.521ciaoenergy.comecommercescrollhard
0.5270.5570.5620.5180.582ciaoenergy.comecommercescrollhard
0.7680.7990.6730.5280.730cipher.tvportfolioplays by itselfhard
0.7760.7450.8050.6190.574dialkit.devapp-uidrag / clickhard
0.5260.6090.5690.5520.4672025.driftime.combrandscrollhard
0.8900.8880.8440.7190.673gufram.itecommerceplays by itselfhard
0.7250.5720.4650.3840.392kaviengcreative.comportfoliodrag / clickhard
0.6310.5360.5540.5520.528maximatherapy.combrandplays by itselfhard
0.6540.6690.4950.6040.743maximatherapy.combrandplays by itselfmedium
0.7320.6330.5760.5460.469monopo.londonportfoliohoverhard
0.6450.5770.4060.5650.550examples.motion.devapp-uidrag / clickeasy
0.7020.5290.5780.6270.522neutomni.comportfolioplays by itselfhard
0.8410.7660.7580.7420.501neutomni.comportfolioscrollhard
0.6060.5750.3530.4480.575otsuka-air.jpecommerceplays by itselfhard
0.4750.4270.4340.3950.479oxigen.sabrandscrollhard
0.6420.5660.5490.5370.553palmo.co.inecommercescrollhard
0.5830.6340.5240.4670.452palmo.co.inecommercescrollmedium
0.5480.5590.5580.4980.490papertiger.comportfolioscrollhard
0.5250.5820.6230.2520.528papumba.comsaasopens / changesmedium
0.5180.5170.4580.4720.440pixel.melbourneportfolioscrollmedium
0.4000.4020.3140.3000.409pixel.melbourneportfoliohoverhard
0.7970.8280.8060.5370.512pudding.cooleditorialscrollhard
0.6740.7150.4110.4730.461rapidkert.combrandscrollhard
0.5340.5460.4550.4880.429raycast.comsaasplays by itselfhard
0.8320.8510.6490.6410.587rebelliously-optimistic.combrandscrollhard
0.8500.8870.6010.6360.617rebelliously-optimistic.combrandscrollhard
0.6380.6090.5440.4940.633slowdowncreative.comportfoliocursor-followmedium
0.2900.2970.3220.3470.326slowdowncreative.comportfolioscrollmedium
0.2690.2880.2640.2780.264slowdowncreative.comportfolioscrollmedium
0.8720.8700.7020.8820.485slowdowncreative.comportfoliohovereasy
0.3410.3400.3430.3420.338slowdowncreative.comportfoliohovereasy
0.6430.7020.5520.5040.593brand.squarespace.combrandhovermedium
0.5860.5940.5740.5270.427truus.coportfolioscrollhard
0.3990.4810.4330.4310.481victorfuruya.comportfolioscrollmedium
0.4580.4080.3980.4080.470victorfuruya.comportfolioopens / changesmedium
0.3210.3640.3860.3820.360victorfuruya.comportfolioplays by itselfmedium
0.4290.4820.4520.3910.299victorfuruya.comportfolioopens / changesmedium
0.7310.7830.8040.8150.750wisprflow.aisaasopens / changesmedium
0.8350.8310.6880.1940.487wisprflow.aisaasplays by itselfhard

Citation

If you use Animation Bench, cite this post as:

BibTeXphysera2026animationbench
@misc{physera2026animationbench,  title        = {Animation Bench: Evaluating Frontier Models on Web Animation Reconstruction},  author       = {Ashwarya Maratha and Tim Cvetko and Himanshu Dubey and Soham Parekh},  year         = {2026},  month        = sep,  howpublished = {Physera},  url          = {https://www.physera.ai/research/animation}}

Partner with us

Animation Bench is our first research work aimed at closing this gap for the next generation of frontier models, helping them improve on the qualities users actually perceive. If you are working on frontend generation, or environments for agents that build software, we’d love to hear from you. Reach us at hello@physera.ai.

Get our latest research

New benchmarks and findings from Physera, straight to your inbox.