Generative Media · entry 03/04
Video, voice & music
Text-to-video holds together for seconds, a voice clones from a short sample, and music generation is stress-testing copyright.
Video: coherent, briefly
Text-to-video now produces clips that read as filmed — camera moves, lighting continuity, believable motion — on the scale of five to twenty seconds. Past that, the seams show. Physics goes soft: liquids merge, objects pass through hands, shadows forget their sources. Identity drifts: a character's face and wardrobe mutate between second three and second twelve. Long-horizon consistency is the video cousin of the context problem — the model must keep faith with everything it has already committed to frame — and quality has so far tracked the same scaling curve as language. Production use has settled where the economics work: b-roll, VFX elements, product shots, and ads assembled from many short generations. Not long-form narrative.
shot: slow dolly-in, 35mm, shallow depth of field
scene: rain-soaked alley, neon signage, drifting steam
action: courier looks up at camera, holds gaze, 4 seconds
avoid: camera shake, on-screen text, cuts
Voice: nearly solved, immediately abused
Speech synthesis is effectively past the human line: modern TTS carries prosody, hesitation, and emotion well enough that listeners stop noticing. Cloning a specific voice takes seconds of reference audio. Both halves of that sentence matter. The utility is real — dubbing that keeps an actor's own timbre across languages, restored voices for people who lost theirs, game characters with unbounded dialogue. The fraud is equally real: voice is now a forgeable credential, "I recognized his voice" is no longer evidence, and impersonation calls are routine enough that families and finance teams adopt callback rules and code words.
Music and the licensing storm
Music models generate full arrangements with convincing vocals from a text brief, which made them the fastest route to a courtroom. Catalogs were ingested without licenses; labels sued; the settlements are converging on the streaming playbook — licensed training pools, attribution, and revenue splits rather than prohibition. The legal line is genuinely awkward: style is not copyrightable, recordings and compositions are, and a model trained on the latter to imitate the former sits exactly on the boundary.
Realtime voice as an interface
Speech-to-speech models now respond at conversational latency, survive interruption, and shift tone on request — which upgrades voice from a dictation trick to a front door for software. This is where generative audio stops being a media story: a spoken interface backed by a capable agent replaces a screen for a real class of tasks, and the phone tree is the first casualty.
Failure mode
Planning generated video like filmed video. There are no retakes: every generation is a new world, so "the same shot but warmer" returns a different alley, a different courier, a different jacket. Teams that storyboard as if continuity were free stall immediately. Work the other way: pin identity with reference images where the tool supports it, generate far more clips than the cut needs, and let the edit absorb the drift. Budget for selection, not retakes — the model is a casting call, not a camera.