How Animation, Backgrounds, and Effects Enter the Game
Making Games with AI · Part 9

This article revisits development in June–July 2026 and early August, using records through August 2, 2026. References to “current” behavior, values, images, candidates, and validation describe that historical snapshot, not today's released game.
Previous: Why AI Art Isn't Ready to Use
The previous article followed static art from candidate to usable component. Animation adds another problem: a character has more than one picture, buildings change after taking damage, the battlefield moves with the camera, and skill effects must appear at the right location.
All of these are visuals, but they entered Kibble Street TD through different paths. Characters needed an action state, buildings needed a health state, backgrounds needed layers, and effects needed a trigger and an ending. Images were only the starting point. Integration gave them executable rules.
Character animation is a state library, not a video
The dog ranged unit had six states: idle, run, attack, stun, death, and victory. They shared one identity but behaved differently. Idle and run looped, attack played once, death held its final frame, and victory looped during results.

Production did not replace every action simply to make them uniform. Existing idle and attack sequences remained usable; run, stun, death, and victory were rebuilt separately. That kept rework focused and made each state's rhythm, ground contact, and proportions easier to judge.
The difficult part was making every state feel like the same character. Ears, backpack, clothing, weapon, proportions, and major colors needed to stay consistent. Running feet needed one ground line. Falling could change the pose's height without changing the character's build. The human fixed the standard appearance and action intent; AI produced state-specific candidates, frames, and previews around that reference. Six unrelated text prompts were liable to produce six similar but different characters.
An atlas needs a frame-layout manifest
The running sequence became an 858×143 atlas containing six 143×143 frames. The image alone was insufficient. A manifest recorded each frame's x and y position, width, height, and state.

Runtime read the manifest and checked frame count, bounds, and playback mode instead of guessing a grid. Run looped at 0.11 seconds per frame; stun looped at 0.15; death advanced at 0.20 and held its final frame. Missing layout information, inconsistent counts, or out-of-bounds rectangles caused an explicit loading error rather than a silent fallback to a still image.
Those checks exposed problems early. An atlas incorrectly divided into five frames could otherwise flicker only near the end of an action, leaving a much harder investigation across unit behavior, image dimensions, and animation code.
Buildings animate state changes
Character actions followed movement and attacks. The dog main base instead used four sections: a normal loop, transition into damage, damaged loop, and transition into collapse.

Normal and damaged states each used four looping frames to keep flags, lighting, and details moving. Entering damage played a six-frame transition once. Reaching zero health played eight collapse frames and left the final rubble frame visible. That separated what the building was now from what had just happened to it.
A single long loop would not let battle logic choose when damage or collapse occurred. It would also complicate pausing, acceleration, and reentering the view. With states and transitions separated, the game could select the correct sequence.
The battlefield is more than one full-screen image
The battlefield used a 1080×1920 portrait reference within a world 4320 units wide. At this development stage it combined four layers, from back to front: distant scenery, road, middle-distance buildings, and foreground.

The distant layer supplied sky and mountains. The road established the ground beneath units. Buildings defined the neighborhood, while foreground elements covered the nearest parts of the view. Separating them allowed changes to buildings without repainting the sky, or to foreground overlap without changing the road. Transparent areas allowed runtime composition instead of permanently baking every relationship into one picture.
More layers were not automatically better. Each added dimensions, placement, transparent edges, and scroll boundaries to maintain. These four existed because each had a distinct responsibility. Decoration that did not need independent movement or overlap did not need another layer.
Distant scenery moves more slowly
As the camera moved across the battlefield, road, buildings, and foreground traveled the full distance. Mountains traveled five-sixths as far. The difference suggested depth: nearby streets passed quickly, while distant scenery felt farther away.

The difficult part was the boundary, not the formula. The maximum main-scene offset was 3240, and the shorter distant layer allowed 2700. Left, middle, and right positions needed checking for exposed empty space and layers drifting out of alignment.
AI could repeatedly read dimensions, calculate those positions, and produce comparison previews. The human still judged whether parallax was useful, which layer should move more slowly, and whether the difference looked distracting.
Effects stay independent of characters and backgrounds
The freeze item began with a five-frame impact, then played a four-frame cold mist beneath affected units. Neither sequence belonged to the ranged unit's action atlas or to the background. They were temporary visual elements positioned near each affected unit.

Effect size followed the unit's ground shadow, so tanks, melee, ranged, and siege units did not each need a separate freeze asset. The animation communicated freezing; it did not decide damage or when movement stopped and resumed. Those decisions belonged to battle rules. Temporary visuals removed themselves after their assigned playback, while the character retained its own state animation.
This avoided copying the same skill into every character atlas. Revising the mist did not change a character's action. Backgrounds, characters, effects, and rules could evolve separately while preserving their trigger and placement contracts.
Who is responsible for this part?
The human fixes character identity, action meaning, ground contact, layer order, timing, and the final look. They also decide whether a change belongs to characters, buildings, backgrounds, or effects.
AI prepares state-specific candidates around the approved identity, organizes frames and manifests, produces boundary and composition previews, and checks dimensions, counts, transparency, and references. Runtime validation keeps the manifest and image in agreement.
The aim is not one-shot generation of a whole animated game. It is making each state clear enough to inspect and improve.
Where to start
Start with one easily judged running state, not every action. Prepare a full-body reference pose, make a short six-frame loop, fix the ground line, and record each frame rectangle. Run it in the real scene and watch loop joins, scale, and ground contact.
Once that works, add attacks, hit reactions, and death. Add independently moving background layers when depth requires them. A small working version exposes problems sooner than dozens of unfinished atlases.
Animation, scenery, and effects make the picture richer, but watching them still cannot establish whether the game is enjoyable.
