Algorithmic Motion: The Other AI Video Revolution

When people say AI video, the image that comes to mind is usually a diffusion model dreaming directly in pixels: an impossible camera move, a photorealistic creature, a street that changes historical periods as someone walks through it. This was the first great shock of AI moving images. The machine appeared to have acquired a camera, a studio, actors, weather, lighting, and special effects all at once.
But a second route has arrived more quietly. Instead of generating the frames, an AI writes a program that produces them. Typography moves along a path. A diagram assembles itself. Data becomes shape, rhythm, and color. Photographs, interfaces, illustrations, footage, and sound can all be placed on the same designed canvas. A browser renders the program frame by frame and exports the result as video.
This is AI-driven algorithmic motion: motion graphics generated by code.
It came later than diffusion video and initially looked less revolutionary. It does not conjure a dragon or an imaginary city. Yet I suspect it may become the more pervasive force. Diffusion video expands what a camera can appear to capture. Algorithmic motion expands what can be made visible at all.
The Lineage of Algorithmic Motion
Motion graphics made with code are not new. Artists and designers have long worked with creative coding, procedural animation, and browser graphics. Processing, first released in 2001, was conceived as a software sketchbook that would bring programming into the practice of visual artists and designers. It has been used for animation, data visualization, installations, stage design, and film.
The family of approaches is much broader than any one tool. Manim uses Python to construct precise mathematical and technical animations from objects, transformations, and scenes. SVG makes vector geometry scriptable inside the browser. The HTML Canvas element provides a different kind of browser surface: rather than arranging HTML inside it, JavaScript draws lines, shapes, images, and pixels onto a bitmap. p5.js builds a friendlier creative-coding environment around this kind of drawing, and libraries such as p5.brush give its generated images brushes, hatching, natural fills, and a more illustrated or hand-rendered character. Lottie carries authored vector animation as structured data between platforms. GSAP supplies a rich timeline and motion vocabulary; Three.js extends the browser into 3D; and the Blender Python API opens a complete 3D creation suite to procedural control. These systems differ greatly, but they share one decisive property: their production logic can be expressed as code (objects, parameters, relationships, and time), even when the final Canvas image is a bitmap.
These approaches are not exclusive. A single shot can place DOM typography over a Canvas particle field, animate an SVG diagram with GSAP, bring in a Lottie character, use a p5.brush rendering as an illustrated texture, and set a Three.js object into genuine depth. Photographs, generated images, recorded footage, and sound can enter the same composition. The useful analogy is collage: different materials retain their own properties, but acquire new meaning through how they are cut, layered, synchronized, and placed in relation to one another.
Remotion and HyperFrames are two prominent frameworks for programmatic video. They occupy a somewhat different level. DOM, SVG, Canvas, p5.js, Lottie, and Three.js answer the question: What is this image made from, and how is it drawn? An animation runtime such as GSAP answers: How do those things change over time? Remotion and HyperFrames answer the larger video question: What is the composition, which frame or moment are we rendering, when do its scenes and media appear, and how does the result become a playable preview or exported file? In that sense, they sit one level above the visual materials as composition-and-rendering frameworks.
Their internal models differ. Remotion is React-first and frame-driven. It gives a React component the current frame number; the component calculates what should appear at that frame, while the composition supplies its duration, frame rate, width, and height. HyperFrames is HTML-first and timeline-driven. HTML describes clips, media, and nested compositions; scripts provide seekable animation timelines, commonly through GSAP; and the renderer moves those timelines to the requested moment before capturing the frame.
Neither framework dictates a visual style or requires a single drawing method. A Remotion component or HyperFrames composition can contain DOM text, SVG, Canvas, Lottie, Three.js, photographs, generated images, and video. They are closer to the stage, clock, and conductor for the collage than to any one material within it. The distinction is not absolute. Manim and Blender, for example, combine drawing, animation, composition, and rendering inside more self-contained systems. But it is a useful way to understand browser-based video.
Regardless of the approach, the result can look like something made in After Effects; the difference is that every frame is the output of a program rather than the result of manipulating a timeline by hand.
However, the genuinely new part of this algorithmic motion is not the code. It is that the code no longer has to be written by a programmer or motion designer.
Code is the door through which AI enters the medium. Language models have become exceptionally capable at reading, writing, and revising software. When a creative system exposes a stable code interface, an agent does not need the years of training required to operate its graphical interface. It can assemble objects, calculate positions, search a codebase, change parameters, render the result, inspect it, and try again at software speed. What once required patient manual construction can suddenly feel instantaneous, to the point of magic.
A person can describe an idea, optionally provide references and source material, and wait for the result to come back in a few minutes. The agent translates visual intention into geometry, type, timing, easing, camera movement, and compositing. It can take screenshots of its own creation, a significant step forward in Anthropic’s Opus 5.5. When something is wrong, it can alter the program: move this label, slow that transition, hold the image two seconds longer, replace one asset without rebuilding the scene. Because the animation is deterministic, the same frame can be returned to, inspected, and changed in a much more affordable way.
The person asking for an awesome video about whatever topic does not need to know anything about SVG, GSAP, FFmpeg, or how an LLM makes videos at all. The specialist craft that once set the price of entry is largely absorbed by the agent. Nor does the request need design direction to come back presentable. The default design sense that arrived with Opus 5.5 is surprisingly good by everyday standards, and that is a big reason the approach has had such a huge impact. Presentable is not the same as significant.
Moving Images in the Context of Visual Culture
There is no doubt that twentieth-century visual culture is deeply indebted to cinema. Cinema taught mass audiences to read the close-up, the cut, the moving camera, parallel action, and the transformation of time through montage. Television, advertising, music video, and eventually online video inherited much of that grammar.
Generative video models are no exception. Trained on footage, a diffusion model learns what footage looks like and reproduces its grammar. It is a neural extension of cinema, and its dominant uses show it: a character, a place, a mood, a few seconds of story. In China, this has already become an industry. According to China Network Audio-Video Association data reported by Yicai Global, more than 95 percent of the 128,000 mini-dramas released in the first quarter of 2026 were AI-generated. They are serialized as short, phone-first episodes and released at a volume that conventional production could not sustain.
That lineage also explains its weakness. A diffusion model does not assemble a scene. It draws a sample. Practitioners describe the process as a lottery: you pull the lever and hope. Reference images, keyframes, LoRAs, and motion controls narrow the odds, but they do not turn the result into a completely specified scene.
Cinema can live with this. A director never got exactly what was on the page either. She shot takes and chose one, and the surprising take is often the better one. Diffusion video inherits that economy: generate ten, keep the best. The lottery turns into a liability when the work stops being a story and becomes a statement. Preserve this face, move that object, fix one word, repeat the action with exactly the same timing: each small correction may mean another pull of the lever, or a separate repair workflow. What feels miraculous during exploration becomes frustrating and wasteful when the work demands continuity and exact revision.
But the contemporary screen is not simply a small movie theatre. Count what moves on yours in a single day: a weather map, a loading bar, a lower third naming a speaker, a lyric video, a year-in-review, a twenty-second explainer on how a tax rule works. This is a second tradition running beside cinema, descended from Saul Bass’s title sequences, broadcast graphics, and interface design. In After Effects, or the Velvet Revolution, Lev Manovich described a new hybrid visual language that emerged in the 1990s through film titles and television graphics and gradually came to dominate visual culture. Its grammar is not only the cut or the close-up. It is the label beside the number, the bar that grows when the figure does, the object that becomes a diagram, the beat that triggers a transition. These forms are made of relationships, not scenes.
Code is exceptionally good at relationships.
A diffusion clip is a photograph of a spreadsheet. Every number is visible and none is connected to anything. Code is the spreadsheet: change one cell and the rest recalculates. Code can keep text accurate, align objects, bind a graphic to data, repeat a design language across scenes, and produce versions for another language or screen size. A color can be changed everywhere. A number can be updated without redrawing a shot. The same visual idea can become a horizontal film, a vertical post, an interactive explanation, or a live graphic. Once motion is code, it inherits the powers of software: variables, components, versions, automation, and exact revision.
Nor is this a rival to diffusion. Instead of immutable clips, code creates a visual system, and that system can host generated images or footage, often to good effect: a diffusion shot of a harbor can sit behind code-driven type and a chart bound to live data. It does not surrender the entire frame to generation. The origin of an asset, its visual representation, and the way it moves are separate decisions. A photograph can sit beside an animated diagram. A real product interface can open inside an abstract field of type. A line can be driven by musical pitch while particles react to loudness. Code gives the generator a place to work and an editor above it. Motion graphics have always been a place where media meet. AI makes that compositional power available through language.
This is why algorithmic motion may become a major form of AI-made moving image. It won’t replace cinema, animation, or photography, but it fits the working middle of visual culture: everything that needs to communicate, persuade, teach, identify, or make information felt, along with things that never had a picture at all, like an argument or a piece of music.
This is no longer a forecast. In September 2026, almost immediately after Opus 5.5 became public, X (formerly Twitter) filled with short motion reels, product films, explainers, interface animations, music videos, and 3D experiments written by coding agents. One community index catalogued more than a thousand Opus 5.5 video examples, while an open collection grouped hundreds of works across motion graphics, UI, explainers, games, and 3D scenes.
On September 18, HeyGen published a Code2Video benchmark built around 168 human-designed motion-graphics references. A medium that needed no benchmark a few months earlier is now being measured against professional work.
The volume is not proof of quality. It is evidence that the capability has crossed a threshold.
Explaining Concepts through Motion
For me, this approach was never a novelty. It was always there and now it just does things much better with little intervention.
I spend a great deal of time explaining abstract concepts: how an AI agent is organized, how a workflow changes, where responsibility moves inside a system. Prose can state the relationships, but a moving diagram can give them space, sequence, and memory. Once AI made that kind of construction affordable, I wanted to see whether an argument could acquire a visual world of its own.
My first experiments were explainers. This is the most obvious use because explanation already has an internal structure that can be made spatial: parts, layers, sequences, contrasts, causes, transformations.
In The Anatomy of Agentic AI, the subject becomes a stack of concentric systems: the agent core, teams, triggers, interfaces, environments, and finally the world outside. In Who Carries the Work?, different relationships with AI become a relay race. The runner, baton, lanes, and destination persist while the argument changes around them.

These are modest experiments, and you can probably find stronger work circulating elsewhere. What these quick tests taught me was more useful than their polish. The interesting part is no longer making these explainers. It is understanding what they can do, what they cannot, and, as a result, how to steer AI to produce them.
An explainer becomes convincing when its visual world thinks with the narration. Objects must recur with meaning. A transition should develop an idea rather than merely replace one scene with another. The sequence needs rhythm, contrast, anticipation, and release. AI can now perform a remarkable amount of the construction, but direction still determines whether the construction adds up to thought.
Visualizing Music
Motion graphics do not need to explain an argument. They can also interpret something that already unfolds in time. A few days later another question occurred to me: could this newly acquired ability make music visible?
For a quick test I chose a short and sweet piece: Grieg’s Arietta.
The first attempt placed a circle-of-fifths dial in the sky above a snowed-in Norwegian valley. It was already more than a generic audio visualizer. An automatic transcription supplied 439 performance-timed notes, separated into melody, inner voice, and bass. A note’s angle showed its place on the circle of fifths; radius showed its octave; duration and loudness shaped its glow. Melody notes became connected stars, inner notes became short twinkles, and bass notes lit arcs on their octave rings. A ghost of the opening constellation remained on the dial and brightened when the theme returned. Behind it, the valley changed at a slower speed: dusk became night, snow followed the pianist’s local tempo, the aurora followed larger sections and dynamics, and a second farmhouse window lit at the return.
The mapping was musically detailed, but I found its governing idea too pedantic. It could tell the viewer where a chord belonged without necessarily conveying why the returning phrase felt tender or why a chromatic note carried emotional weight. Teaching compositional technique could remain one layer of the visualization, but it is not my primary purpose. The image needed to transmit some of the music’s emotional power.
The next experiments kept the same performance data but loosened the metaphor. One turned the melody into an aurora; another poured it into a reproducible ink-fluid simulation; a murmuration sent musical energy through 1,500 birds; and the version below treated the performance as a landscape of moving traces. Pitch determines height, loudness becomes heat, the accompaniment appears as an inner voice, and the bass rings on the water. Harmonic tension changes the color. One turn of the luminous path, measured in musical time, runs from the opening to the return, so the theme eventually passes over its own trail.

The result is not a conventional audio visualizer. It is closer to a reading of the piece through motion and space, using what the visual medium inherently offers: shape and color. The musical data supplies exact events, but the mapping from sound to space is interpretive. A different listener could build an entirely different world from the same performance.
A History of Cinema, as Understood by AI
As a trained film scholar, cinema is the subject in which I am most able to test the machine rather than merely be impressed by it. I was curious about what AI knew beyond obvious facts. Could it know how a film looks? More importantly, could it identify why an image or sequence remains striking, through composition, mise-en-scène, camera movement, or montage?

The answer surprised me. Opus 5.5 could identify a well-known film for its most memorable visual element. Nosferatu became a clawed shadow climbing a staircase: architecture and silhouette doing the work of horror. Lawrence of Arabia became a tiny rider emerging from an almost empty horizon: scale and duration made visible as composition. Seven Samurai became bodies charging through diagonal rain: movement understood through the forces surrounding it. Psycho became shower lines, a drain, a spiral, and an eye: not one famous still but an editorial chain.
This does not establish that the model understands a film as a viewer or scholar does. Its “memory” may be assembled from recurring stills, descriptions, criticism, and other traces in its training. Has it really seen those films? I doubt it. But operationally, it can do something more interesting than retrieve a title or regurgitate a film review. It can select a salient visual motif from a film and rebuild a scene around it.
Opus 5.5 was also unexpectedly clever about the transitions between films. The rain of Seven Samurai straightens into shower water; the drain spiral opens into an eye; the bone from 2001 becomes a satellite and then the outline of a shark’s fin; the fire trails from Back to the Future turn upright and fall as the green code of The Matrix. The transitions are not decorative glue. They make a history of cinema into one continuously mutating visual thought.

Prevalence vs. Significance
I began with a large claim: algorithmic motion will be a significant force in shaping visual culture. Its reach and ease of production make that plausible. But prevalence is a matter of counting, and significance is a matter of looking. Nothing guarantees that the culture it produces will be worth looking at.
The more AI motion graphics circulate, the more they begin to resemble one another. Code avoids the strange fingers and ugly faces of diffusion video, and its text stays legible. But like every AI-generated artifact, it cannot escape cliché.
A model has absorbed an immense repertoire of existing design, and asked simply to “make it impressive,” it moves toward the statistical center of impressiveness. We have already seen the first wave in AI video: trailers, advertisements, and social clips that are technically astonishing but interchangeable in rhythm, framing, and feeling.
If algorithmic motion becomes cheap enough to fill every screen, that default could homogenize visual culture faster than cinema did. It can make visual slop with perfect edges.
Slop is not incompetence. A slop sequence can be flawless: the type is crisp, the easing is smooth, the colors agree. What it lacks is taste, an answer to the question of why things look the way they do. Where taste is present, form and content, medium and message, arrive at a synergy: the motion of the type is the argument, and the color is the mood. In slop, a zoom, a transition, and a field of particles appear because the model can make them, and nothing has decided what the viewer should notice now.
My first Arietta visualization shows how close competence can come. Every note sat correctly on the circle of fifths, and the image still could not say why the returning theme felt tender. Nothing was wrong with it. What it lacked was an answer to what the image was for, and that answer came from my rejecting the mapping, not from the model.
Which raises a hard question: does AI have taste?
In a recent essay about AI and taste, I made a distinction that matters here. AI has become extraordinarily capable at executing a chosen direction. Its ability to select the direction has not advanced at the same rate. Taste begins when several options are plausible, no rule can settle the choice, and someone must still rank them: this direction has life; that one is competent but inert; a third will become expensive without becoming interesting.
The Opus 5.5 videos complicate any comfortable claim that taste belongs permanently to humans. The model sometimes chooses the right metaphor, movement, or stillness without being told. In the history of cinema, it picked the image that would stand for each film, and found that the rain of Seven Samurai could straighten into the shower water of Psycho. That suggests visual taste is at least partly learnable. It also suggests where taste lives: in the weights (the parameters fixed during training), as a tendency, not as instructions you can use to steer the model. Prompts and skills (packaged instructions the model reads before it works) could not get earlier models there. The model produces good work without being told how because the judgment was absorbed in training.
But a taste absorbed into the weights is a shared taste, and everyone who calls the model gets the same one. That is the statistical center again, now at its most refined. It can make a sequence tasteful. It cannot give it a reason to exist. Good taste gets cheap once everyone has it.
For now, cheaper execution makes human selection more valuable rather than less. The work has moved from making the frames to directing what the frames are for. We own the purpose, progression, source material, recurring visual ideas, and the standard by which the result is judged. Neither technical validity nor one polished screenshot establishes that a sequence is compelling.
Direction also means control. Greater power creates a greater need to inspect what the machine did and to stop it from silently changing the things that matter. My work on Hyperframe Studio grew out of both problems. A useful AI motion workflow cannot be only prompt, generation, and export. It needs visible decisions: an art direction, real assets, a shooting script synchronized to sound, static design review before expensive motion work, scene-level revision, and final human judgment of the sequence. The machine removes mechanical labor. It does not remove the need to see.
Software Made for Agents
Immediately before the Opus 5.5 motion-graphics wave, OpenAI’s Astra produced a similar shock in 3D. In one OpenAI case study, Astra built an editable house in Blender through the bpy Python API, inspected and repaired its own renders, directed a camera tour, and constructed a pipeline that transferred the scene into Unreal Engine. Other experiments used the same model to build game assets and procedural worlds. The striking improvement was not a new 3D generation system. It was the agent’s ability to operate the tools already there much better.
This suggests a new way to recognize progress in AI. A model does not need to become visibly more philosophical or score higher on an abstract measure of intelligence. It may simply become much better at using Blender, writing a motion-graphics timeline, or coordinating an existing production system.
Software is increasingly designed not only for a human moving a pointer through menus, but for an agent calling tools. Epic’s experimental Unreal MCP, for example, exposes editor operations such as spawning actors, configuring lights, creating material instances, inspecting widgets, and running tests to any compatible AI agent. The graphical interface remains, but it is no longer the only entrance.
The pattern will recur. A major authoring platform exposes a reliable API, command line, or MCP server; an agent gains access to the medium’s objects and operations; and work that once required prolonged manual execution accelerates dramatically. Algorithmic motion is not an isolated case. It is one instance of a larger shift in which software is made for agents.
AI is the King Midas of our time. Whatever it touches turns to gold, and the things it touches first are the ones that can be reached through code.
The myth is usually told as a warning, and its details fit our case closely. Gold is valuable because it is rare. In a world where everything has been touched and turned into gold, what value remains?

Comments