Back to notes
4K launch videos written as code
28 Sep 2026 · 2 min read
The launch videos for GrowX aren't edited on a timeline. They're written as code, rendered to 4K, and every sound is placed by a script. Here's the pipeline, and why I work this way.

Why code instead of an editor
When a video is code, a change is an edit to one line, not twenty minutes of dragging clips. Three variations of the same ad share one pipeline. And because every step is a script, the result is the same every time I render it.
1. Prepare: transcribe the talking head
A Python script pulls a 16 kHz mono track out of the raw footage with ffmpeg and transcribes it with faster-whisper, keeping word-level timestamps. Those timestamps drive the captions, so every word appears exactly when it's spoken. It also saves a contact sheet of the footage, so I can see the whole take at a glance.
2. Footage: keep the detail
The source footage is 2160×3840. It's re-encoded on the GPU (NVENC) at native resolution with a keyframe every 30 frames. The dense keyframes make frame capture fast during rendering, and the full resolution means the 4K render has real detail to work with. The 1080p version is just a downscale.
3. Composition: the ad is an HTML page
Each ad is an HTML composition animated by a single GSAP timeline and rendered by HyperFrames. The UI moments (the "in the shower / on a long drive / late at night" cards, the "months later" calendar, the tap-to-apply screen) are real HTML, so they're sharp at any resolution. Python even generates the colour grade: one script writes a 65×65×65 3D LUT.
4. Sound: every cue has a visible cause
- SFX: a script builds one sound-effects track for the whole cut. Every whoosh, riser and click is placed on a beat that already exists on screen, trimmed to its onset and faded so nothing bleeds into the next cue.
- Music: the supplied track was mastered about 14 dB louder than the voice. So the script measures loudness, matches the music to the narration first, and ducks it slightly under speech. "35% volume" then means 35% of the voice level, not 35% of a very loud file.
5. Render
HyperFrames renders the composition at 2160×3840 with H.264 video and AAC audio. The last check is ffprobe, to confirm the audio stream is actually there, because a silent export is the easiest mistake to ship.

Where AI fits
AI does the tedious parts: faster-whisper for timing, and an AI coding assistant that writes and revises the scripts and composition with me. The creative calls, like what to say, when to cut and what the viewer should feel, stay human.