How the seven films are actually made

Production Workflow

A film begins as a story decision, becomes a locked set of visual references, then combines generated environmental plates with authored interfaces, existing voices and a deterministic 60fps compositor.

Living processSilent motion platesAuthored screen text60fps masters

The production line

Every stage has a concrete output. A beautiful asset does not move downstream until it satisfies the story and continuity constraints of the stage before it.

01
Story

Write the dramatic job.

Define what the audience knows before and after the scene, what changes its meaning and which later episode receives the handoff.

OUTPUT
episode source HTML
continuity ledger
live manuscript row
02
Reference lock

Fix identity, room and screen ownership.

Select the canonical character, location and monitor references. Record age, wardrobe, device count, time of day and OS era before writing an image prompt.

OUTPUT
character bible
location master
monitor PNG
visual manifest
03
Still coverage

Build the scene as a shot family.

Generate or author a wide, medium, reverse, over-shoulder, insert and empty environmental beat from the same geography. Reject face, age, wardrobe and equipment drift.

OUTPUT
approved source plates
prompt archive
coverage map
04
Silent motion

Generate longer than the cut.

Hailuo supplies volume coverage; Kling is reserved for selected human hero shots. One modest subject action and one camera move per take. No dialogue or lip sync.

OUTPUT
prediction ID
source + final prompt
stable in/out points
clip hash
05
Composition

Author everything that must be exact.

Terminals, reports, charts, dialogue, names, warnings and reveals are deterministic HTML/CSS/JS. Generated monitors contribute only hardware, light and broad layout.

OUTPUT
source deck
speech-relative reveals
era-owned interfaces
06
Sound

Edit voice first; place score beneath it.

ElevenLabs voices define the dialogue cut. Location ambience and score remain separate lanes. Music ramps, actions receive motivated effects and speech stays foreground.

OUTPUT
voice cache
room tone
score + effects
loudness-matched mix
07
Export & review

Render from one canonical clock.

The exporter requests every frame at 60fps, drives all reveals and transitions from timeline time, mixes the audio, decodes the result and promotes only a visually reviewed master.

OUTPUT
1920×1080 MP4
full decode check
contact sheet
preserved prior master

Current production stack

The tools have specialised responsibilities; none is allowed to invent the whole film.

HTML · CSS · JavaScriptStory UI, reports, terminal state, text timing, transitions and preview decks.
Python · browser captureTimeline construction, frame requests, voice cache and export validation.
Replicate motion modelsPeople, rooms, hands, light, reflections and restrained camera movement.
ElevenLabsRole-specific voices and selected original music/effect generation.
FFmpeg · FFprobe60fps composition, audio mix, loudness inspection and full-file decode.
Reference websiteCast, scenes, screens, manuscript, prompts, guides and approved assets.
Source verificationPrimary reports and documented incidents separated from fictionalised drama.
Human reviewStory sense, identity, visual quality, pacing, voice choice and final approval.

What changed from the prototype

The original single-film presentation was a useful experiment, but it is not the current production method.

Current series

Grounded bright locations, era-specific screens, conversational performances, selective generated motion, story-owned music, dynamic evidence and deterministic 60fps export.

Retired direction

Constant black backgrounds, orange corporate branding, generic persona cards, looping glitch effects, full-slide fades and the claim that a named autonomous AI “team” independently produced the work.

Reproduce the work

Start from the designed references rather than an old screenshot or an isolated prompt.