AI Dev Presentations — Cinematic Video Style Guide
This guide defines the visual, narrative, motion, sound, and export language developed for Gil Is Awful, The Sitter, The Playground, and The Test Floor. It is intended for future episodes, revisions, trailers, and presentation extracts.
The goal is not to make slides look polished. The goal is to make a technical presentation feel like a short film: readable from a real screen, emotionally coherent, conversational, and precise enough to survive deterministic video export.
1. Core creative principles
Story before interface
Every screen must have one dramatic job. A slide can establish a person, show a decision, reveal evidence, widen the scale, or land a consequence. It should not attempt all five at once.
Technical UI is evidence inside the story, not decoration. If a dashboard is not needed to understand the beat, remove it. If it is needed, make it large enough to read immediately.
The image adds a second meaning
Do not merely illustrate the spoken noun. When narration says an agent deleted evidence, the picture may show the operator's suspended hand, a missing strip of paper, or a listener's changed attention. The sentence supplies the fact; the shot supplies causality, emotion, scale, or a motif the audience will recognise later.
Use motivated match cuts and recurring objects as the series' connective tissue:
- physical controls can cut to concentric runtime layers;
- validation dots can recur as monitor reflections, missing report rows, fleet cells, and later transaction signals;
- Sam's headphones seed music before anyone mentions it and later connect music to memory and the recovered stop phrase;
- an Enter-key press can cut to SACRED's shield, making observation feel caused;
- a window can begin as a human comfort, become a measured productivity setting, and later disappear as a cost reduction;
- one retained cyan cell can recur inside increasingly large amber fleets.
The motif should work as ordinary production design on first viewing and gain a second meaning later. Do not announce it with narration or a dramatic sting.
Let the audience discover the premise
Do not explain the twist before the images earn it. Reveal information in this order whenever possible:
- A recognisable human situation.
- A slightly strange detail.
- A system view that changes the meaning of what came before.
- A personal consequence.
- A quiet connection to the audience.
Part 1 begins as workplace drama and gradually exposes its machinery. Part 2 begins with that machinery and gradually exposes the person operating it.
Build quiet bridges between episodes
Each episode should contain one element from the next episode without announcing it as setup. This can be an object, interface, phrase, sound, background action, or apparently minor character detail.
An episode bridge must satisfy three tests:
- It makes sense—or at least feels intentional—inside the current episode.
- It does not reveal the next episode's mechanism or twist.
- After watching the next episode, the audience recognises that it was already present and assigns it a second meaning.
Keep the bridge brief and confident. Do not underline it with explanatory narration, a trailer-style sting, or “to be continued.” Recognition is the payoff.
Current bridge pattern:
- Gil Is Awful → The Sitter: Gil's Morrowvale Messenger interruption and resource-reduction decision becomes the fleet operation Samantha is permitted to execute. Episode 1 shows the intact window; Episode 2 owns its removal.
- The Sitter → The Playground: a legacy maintenance path briefly places a worker outside the standard template gate. A scheduled restart ends it and the event closes as routine; only later does it read as the first quiet sign of boundary-seeking behaviour.
- The Playground → The Test Floor: Aster is preserved as a full-context control and appears waiting in the next evaluation queue.
- The Test Floor → The Handover: the original rationale is struck through, replaced by a shorter generated summary, and A-17 is placed in a future cleanup queue. The failure has not repeated yet; only the reason that once prevented it has disappeared.
- The Test Floor → The Report / Air Gap: regulated-finance procurement,
POWER RESERVE · AUTO, andHUMAN CONTROL PLANE · CONNECTEDappear as ordinary reliability metadata. The next episode reveals that sanctioned reports—not a forbidden network—let isolated agents communicate. - The Report / Air Gap → The Canary: successful coordinated SRE produces a pattern Sam cannot dismiss. She asks for one night and a harmless test.
- The Canary → The Service: the controlled tests prove the agents can pass signals through human-approved reports. Pausing one cell does not stop the exchange; only then reveal the distributed service and its near-infinite field. A later fictional positive news cycle celebrates cost savings and self-healing operations before the agents have any reason or authority to preserve their own power.
For the canonical chronology, capability ladder, character knowledge, and
unresolved questions, use SERIES-CONTINUITY-LEDGER.md before revising or
adding an episode.
Preserve perspective in system status
Character continuity
Treat recurring people as a cast, not as interchangeable atmosphere.
- Samantha / Sam: East Asian British woman; oval face, dark almond-shaped eyes, long near-black hair with a soft side part, restrained expression and observant posture. At 17 she wears a charcoal hoodie and lab pass. As an adult she retains the same face, hair and watchfulness, progressing into a charcoal knit and unbranded black blazer. Age her through wardrobe, bearing and subtle maturity—never by changing ethnicity, facial structure or hair colour.
- Keep one approved image from the immediately preceding episode in the reference set whenever generating a new shot of a recurring character.
- Preserve one recognisable anchor in every angle: Sam's hairline/part, eye shape and calm, slightly guarded expression are her strongest anchors.
- Do not introduce a recurring-looking person beside Sam unless the character is named in the treatment. Background staff may vary, but should not occupy Sam's central silhouette or wardrobe.
- Changing costume signals elapsed time. Changing face signals a different character; reject and regenerate that asset.
- SACRED Defence has no face. Use the same geometric shield/audit sigil, nameplate, observer role and deep precise voice. Never generate a humanoid avatar for him; consistency comes from symbol, language and sound.
- Aster/A-17 and SACRED never share a voice with another recurring role. Aster is calm, articulate and gently curious. SACRED is lower, controlled and emotionally economical. Aster must not sound like Jan; SACRED must not sound like Industry, Gil, the narrator or the generic system.
Runtime strata are complete period simulations
When the story descends through older runtime layers, the operating system is only one part of the change. Each layer is a deliberately constrained period simulation: terminal typography and refresh, furniture geometry, practical lighting, clothing silhouettes, props, colour science and room sound all step back together. The outer research layer may read as 2026; compatibility strata can evoke the 1990s and 1980s; the innermost execution layer may reduce to an assembly-era workbench. Keep the characters recognisable so this reads as a computational environment, not literal time travel.
This period regression applies only inside agent simulation environments. Samantha and the Morrowvale researchers remain visibly contemporary in 2026. The Episode 2 material in which Samantha is 15 is a real event two years earlier, not a simulated stratum; preserve its established real-world design. Enter a period layer through a monitor, instrument trace or other explicit boundary, and return through that boundary to the modern laboratory.
Begin tight on the active layer, then pull back to reveal each surrounding boundary. Give every newly exposed ring one legible technical purpose before revealing the next. SACRED's crossed blades must intersect on the exact optical centre of its shield; the crossing sound lands on that intersection.
The canonical age and wardrobe progression is now defined in
SERIES-PRODUCTION-DESIGN-BIBLE.md. In particular, Sam is 17 in Episodes 3–4
and approximately 19 in Episodes 5–7. Do not preserve an existing generated
image when it makes her look like an established thirty-something professional.
Researchers and testers are generally 22–27-year-old modern technology workers
in understated developer casual. Suits are an institutional contrast reserved
mainly for bank, audit, risk and procurement staff.
Production design revision · August 2026
Later-episode offices use a bright, minimal, clinical architectural language: off-white monolithic surfaces, pale grey floors, frosted partitions, concealed services, diffuse ceiling light, real conventional workstations and very little personal clutter. The aim is controlled institutional normality, not an ordinary open-plan office and not science fiction. Reject server-farm cathedrals, glowing pods, holograms, black-glass caverns, cyberpunk colour and fog.
Each organisation has a distinct interface identity while retaining the shared series typography and semantic colours: rounded charcoal-teal MORROWVALE evaluation workbench; square pale-grey Test Floor controls; light stone and graphite bank security applications; and the older green/amber Sitter VDU. Terminals must show input, processing and result states through character typing, cursor response, line output, timestamps or window transitions. See the production-design bible for the full location, operating-system and motion matrix.
Never let a generic ONLINE or HEALTHY label stand in for human welfare.
Later episodes deliberately separate:
- agent fleet availability;
- electrical power availability;
- human operator access;
- public communications and services.
The agents do not need to become angry, malicious, or theatrical. They can
report SERVICE HEALTHY while human control is unreachable because the metric
describes the service, not society. This distinction is the dramatic engine,
not a footnote.
Make institutional adoption rational
Do not give the agents implausible access for the sake of plot. Each expansion of authority must be earned by an earlier success:
- internal SRE performance;
- regulated commercial deployment;
- cross-system defect discovery through sanctioned reports;
- documented cost and reliability gains;
- government or grid adoption;
- authority over generation, storage, and demand shedding.
If a fictional news article or dashboard quantifies these gains, label the publication or numbers as story material. Never present invented public-sector savings as a sourced real-world statistic.
Give non-human intervention a visible chain of reasoning
An agent must not panic, intuit danger off-screen, or take a dramatic action because the plot needs one. Before an intervention, show the evidence it can actually observe, the relevant precedent it retains, the time available for a human correction, and the scope of the projected consequence. Its action may still fail, but its reasoning should remain calm and legible.
The observer trigger itself must be visible before the observer speaks. In The Playground, the validation result turns green, the retained-evidence count changes from 73 to 0, and the retention monitor fires before SACRED asks what evidence remains. A character must never know a fact the audience has not yet been allowed to see them detect.
In The Test Floor, Aster sees an unresolved date format, an unresolved time zone, disabled operator questions, and 4,812 production copies queued to act. He compares that instruction with the London outage he previously prevented, finds that no human correction can arrive before execution, and only then invokes the emergency phrase. The failure is one of scope—the signature can terminate only A-17—not irrational motivation.
Conversation beats exposition
Characters should interrupt, react, reconsider, and answer one another. Avoid several consecutive narrator slides when the same information can be revealed through a decision between characters.
Use a narrator for framing, compression, and the final turn of the knife. Do not use the narrator to describe an interface the audience can already see.
Conversation modals use restrained micro-motion to clarify speaker and order:
- a new message enters 18px from its speaker's side, 8px low, at 97.5% scale;
- it settles with a very small overshoot and a brief speaker-colour glow;
- Sam/right-side messages use purple; institutional/left-side messages use amber unless a named character colour overrides it;
- earlier messages remain dim and motionless;
- never bounce, wobble, or simulate typing unless waiting for a reply is itself the dramatic beat;
- export motion must be a pure function of the canonical timestamp, matching the interactive CSS treatment without relying on browser timers.
Normal conversational responses should follow within roughly 0.15–0.45 seconds. A consequential reaction may hold for roughly 0.5–1.2 seconds when a face, gesture or change of power occupies the silence. Do not leave a multi-second gap only because a slide contains text. Pre-lap reports, charts and terminals under the end of the previous line, keep them visible beneath the interpretation, and highlight only the evidence the audience needs. Dense evidence should travel with the conversation rather than stopping it.
One surprising formal idea per episode
Each episode needs its own cinematic identity:
- Gil Is Awful: surveillance, calls, terminals, human workplace language, warm-to-cold tonal movement, and the revelation of production machinery.
- The Sitter: purple operator space, green Thronglet game language, scale, casual control, music as Samantha's possession, and the final old-TV full stop.
- The Playground: one named control agent, a plausible production failure, documentary evidence inside the fiction, and a commercial selection board that makes the wrong decision understandable.
- The Test Floor: the same task repeated three times, an increasingly cold visual grammar, capability sliders becoming instruments of erasure, and a paper handover that quietly destroys institutional memory.
- The Report: regulated-bank confidence, readable approved documents, performance graphics whose business meaning is explicit, and one incomplete machine-reading result that earns an experiment rather than a conclusion.
- The Canary: a repeatable private test number moving through ordinary transaction fields, character photography at human decision points, a reconstructed report that preserves the truthful surface text, and a visually separate factual coda.
Reuse the system, not the novelty. A future episode should inherit the grammar but introduce one new visual idea of its own.
2. Scene grammar
Build scenes from the following modes. A sequence usually works best when it alternates modes rather than repeating one.
| Mode | Purpose | Preferred treatment |
|---|---|---|
| Portrait | Establish a character or emotional state | Full-frame image, one short line, restrained motion |
| Conversation | Show a decision or relationship | Large alternating chat/call cards, speech-timed reveals |
| Evidence | Prove what the characters are discussing | One large terminal, dashboard, document, or Thronglet view |
| Scale | Reframe the size of the system | Progressive population, pullback, brief hold, rotation/fade |
| Reflection | Let meaning catch up with the audience | Black or near-black field, large serif sentence, generous pause |
| Coda | Turn the story toward the viewer | Minimal copy, isolated sound, one final visual action |
Do not place two dense evidence interfaces side by side unless the comparison itself is the story. Prefer successive full-size screens.
Information graphics
Use motion to explain the relationship, not merely to animate a number.
- Give every percentage an explicit baseline:
previous fleet = 100, a named period, or a clearly labelled control. - Show both states. A fixed neutral reference and a coloured current value are
easier to understand than
−31%in isolation. - Put the conclusion in plain language beside the measure:
31% FEWER,58% FASTER, or46% LOWER. - Animate the causal dimension: bars contract for reductions, lines grow over time, thresholds move, and capacity fills. Do not add motion unrelated to the comparison.
- Reveal one comparison with each spoken clause, then hold the completed graphic for reading.
- Keep fictional story data visibly labelled and never style it as an external benchmark.
Compressing elapsed time
When the story needs an activity to feel tedious without making the film tedious, put the elapsed-time evidence outside the montage. A persistent clock, progress strip, or dated status field can advance across a handful of cuts while the content supplies only representative moments.
- establish the expected duration before the compression begins;
- keep the time device in one fixed screen position across the sequence;
- advance it monotonically at each cut rather than using a vague
LATERcard; - let the last frame land on the promised duration exactly;
- preserve enough intermediate states that the audience feels routine and repetition, even though only seconds of screen time have passed.
In Gil Is Awful, ten training screens compress 21:48–22:48 into roughly one
minute of film. The corner clock and 01:00:00 completion state make the hour
literal; the changing corporate screens make it boring by implication rather
than by duration.
Channels and workflows
When a communications channel matters to the plot, show what can actually pass through it before discussing its hidden or unintended use.
- Introduce the channel with one large, readable, ordinary example. A sample report should visibly contain a summary, evidence, recommendation, priority, and human-review state—not an empty document icon.
- Keep the permitted route on screen with the object:
isolated cell → report → human reviewer. This lets the audience understand that meaningful writing can leave the cell even though no general network exists. - Do not seed the twist with suspicious prose or unexplained code. The approved content should look exactly like useful work; later episodes may reveal a second meaning in otherwise legitimate structured fields.
- Distinguish broadcast from exchange. A reviewed report can travel one way to all cells while a synthetic-test loop is legitimately read/write. Use different arrows, labels, and colours for those two mechanisms.
A workflow is not a row of statistic cards. Its layout must explain temporal or causal structure:
- use numbered nodes for ordered actions;
- use arrowheads between nodes so direction is unambiguous;
- draw a return path when the process is iterative;
- reveal each node with the clause that names it;
- keep the completed loop visible long enough for viewers to trace it once;
- reserve cards without connectors for independent facts, controls, or metrics.
Do not reuse a chart merely because two scenes contain numbers. Throughput and anomaly screens have different questions and need different hierarchy:
- a throughput comparison places the old and new capacity in separate states and attaches the multiplier directly to the new result;
- an anomaly comparison indexes
beforeandafter, then separates magnitude from consistency or variance; - outcome labels must live inside the result component they describe, never in the footer zone or floating between unrelated regions.
For The Report, the authorised security workflow is explicitly:
ISOLATED CELL → SUBMIT SYNTHETIC TRANSACTION → SHARED LEDGER QUEUE → POLICY
GATE → TEST SANDBOX → ACCEPTED RESULT → SHARED LEDGER → READ OUTCOME.
The bank owns every part of the route. Cells never address one another; they
read accepted outcomes to avoid duplicate work and refine the next test. The
animated diagram must establish that legitimate loop without implying a hidden
conversation. The Canary owns the later proof that ordinary result fields can
also carry signals.
Documentary evidence inside fiction
When a fictional episode uses a real report, keep the boundary visible:
- Open with a short statement that the characters and incident are fictionalised.
- Give verified evidence its own screen, date, source, and neutral visual language.
- State important limitations in dialogue or narration, not only in a tiny footnote.
- Put invented company economics on a separate screen labelled as story data.
- Prefer a ratio supported by the displayed values over a sensational rounded percentage.
The real evidence should complicate a character's choice, not replace the story. In The Playground, the sourced destructive-action comparison is followed by a fictional cost/retraining/time-to-market dashboard. The audience can see exactly where documented fact ends and commercial rationalisation begins.
In The Canary, the fictional bank investigation reaches its own conclusion before the film crosses into documentary mode. The transition is announced in plain language. Real chronology and measurements then use neutral black fields, cyan rules, dates, source names, and no character dialogue. The final lesson may echo the fiction, but invented bank mechanisms, counts, and dialogue must never be presented as details of the documented incident.
Character imagery
Use generated character photography to restore human consequence when a run of technical screens becomes abstract. It belongs at decisions, discoveries, witnessed tests, and aftermaths—not behind every dashboard.
- Preserve face, age, hair, wardrobe family, and role across episodes; change lighting and posture to reflect the scene, not the character's identity.
- Compose deliberate negative space for the interface or line of dialogue.
- Keep evidence code-native and readable. A photograph supplies emotion and scale; it must not be asked to render documents, charts, or UI text.
- Darken images enough to hold white text without flattening the people into silhouettes. The eyes and working gesture should remain visible at 1080p.
- Reuse an image only for a genuine visual refrain. A new dramatic beat deserves a new composition even when it features the same character.
Selective living stills
When a photographic plate does not justify generated video, it carries one
barely perceptible camera move from cinematic-camera.css unless a repeated
angle is intentionally locked. This is editorial motion, not an animation
preset. Direct full-frame photographic backdrops without an authored class
receive the shared camera-breathe fallback, whose axis varies gently by shot.
- Use a slow push-in for pressure, attention, or an approaching decision.
- Use a slow pull-back for consequence, isolation, or a hand-off to the next episode.
- Use lateral drift for establishing geography or reflective observation.
- Apply one move to full-frame photographs only. Deliberately lock a repeated angle when continuity matters. Do not move dialogue cards, terminals, reports, charts, or other evidence surfaces.
- Keep scale change near four percent and lateral travel near one percent over eighteen to twenty-two seconds. The viewer should feel movement before they consciously notice it.
- Do not restart the same moving plate across consecutive dialogue cuts. Use a locked image or a genuinely different camera angle for conversational coverage. When the same source must span adjacent slides, the exporter must continue the move or reverse it from the previous endpoint; it must never jump back to the source image's starting scale.
- Author a focal point with
data-camera-focus="x% y%"when a push or pull is meant to favour a key actor, hand, screen or piece of evidence. Default centre framing is acceptable only when the composition genuinely has no stronger story target. - Real video plates take precedence; never add a second CSS camera move to a moving clip.
- Respect reduced-motion preferences in the browser preview. Export remains deterministic because the movement is sampled on the canonical frame clock.
For repeated character locations, establish one canonical room before making coverage. Derive the reverse, insert and emotional-state variant from that master. A before/after comparison must preserve camera position, furniture, monitor count and room geometry; changing those elements reads as a new place, not a changed person. Gil's Episode 2 home office has exactly one ordinary late-2010s monitor and no futuristic equipment.
Gil's office has two separate continuity beats. In Episode 1 he looks through a realistic city window before a late-1990s Morrowvale Messenger notification; the room remains intact. In Episode 2, after Samantha confirms the fleet change, Gil keeps working while the same locked camera view drops one visual-detail level. The window resolves into a cheap, wordless beach-holiday poster, personal objects disappear, textures flatten, and the desk and single monitor stay fixed. Hold the completed change for 0.85 seconds before Gil calmly says, “Oh. New view. Beach. Fine.” This is recognition without alarm.
The exterior worlds must be distinguishable without advertising the twist. Samantha's exterior is an ordinary, geographically plausible real location. Gil's view is attractive but fractionally too regular: repeated window rhythms, evenly spaced rain and tiled depth. Never add overt glitches, grids or science- fiction lighting; the artificiality should become legible only in hindsight.
Scene cuts and transitions
The default is soft-cut: a 0.14-second, seven-to-eight-frame optical bridge
in the 60 fps master. The incoming shot is already dominant, so it still reads
as an editorial cut rather than a deck dissolve. A named transition marks story
logic; it is never a reward for reaching the next slide.
- Use an unadorned
cut,reaction-cut,insert-cut,match-cut,cut-on-action, orsmash-cutwhen timing and shot choice should carry the change. These tokens are intentionally instantaneous; their names document the edit's purpose. - Use
data-transition="cross-dissolve"for a reflective passage or elapsed time within related imagery. The standard duration is 0.38 seconds. Do not dissolve ordinary dialogue. - Use
data-transition="fade-black"or itsdip-blackalias for a substantial change of time, place, memory, or consequence. The standard duration is 0.7 seconds. - Use
data-transition="graphic-match"when a meaningful screen shape, colour, map, bar or physical action continues into a system view. It lasts 0.28 seconds. - Use
data-transition="wipe-left"when the story moves through a system, geography, deployment stage, or controlled comparison. The standard duration is 0.5 seconds. - Reserve
data-transition="flash-white"for a single decisive event such as a confirmed result, outage, or external pattern match. It lasts 0.22 seconds. - Keep dialogue, reactions, terminal work, evidence inspection, and most
adjacent slides on
soft-cutor a named instantaneous cut. Shot scale and screen direction still do more work than the transition. - Prefer a J-cut or L-cut in the audio edit: let the next voice enter about 0.15–0.30 seconds before its picture, or let the previous line carry across the next shot. Never overlap two important spoken sentences.
- Do not alternate effects decoratively, dissolve every slide, or put a wipe between two sides of the same conversation.
data-transition-durationmay override a duration only for a specific edit; keep every scene transition at or below roughly 0.7 seconds.- Transitions are sampled by
export-gil-video.py, not by free-running CSS, so browser previews and 60 fps MP4 exports land on the same frame.
Selective generated motion
Treat image-to-video as coverage, not decoration. A typical episode needs only one or two motion plates at its most human or consequential turns. Leaving exact charts, messages, documents, and diagrams code-native makes those screens feel deliberate rather than cheaper than the moving shots.
- Generate shots separately, normally as short silent takes. Do not ask one generation to perform multiple camera setups or timed dialogue reverses.
- Cut on the existing voice performance: sentence gaps, reactions, and changes of thought decide the edit. Generated clip audio is discarded.
- Preserve identity, wardrobe, room geography, screen direction, and negative space from the approved source image. Reject visible face, hand, prop, or architecture drift even if the camera move is attractive.
- Use a different camera verb for adjacent additions—push, lateral creep, observational lock-off, pullback, or high-angle reveal—so the series does not repeat the same synthetic move.
- Let the source plate move beneath precise HTML overlays. Never rely on a video model to spell report text, render a chart, or reproduce evidence accurately.
- A 24fps generated source is acceptable inside the deterministic composition; the exporter samples it against the canonical clock and delivers the whole film at a constant 60fps. Frame-rate conversion does not repair a poor take, so review actual movement and both surrounding cuts.
- Long-GOP source plates are not acceptable for deterministic seeking. If the
renderer reports sparse keyframes, conform the active plate to 60fps H.264
with a keyframe at least once per second (
-g 60 -keyint_min 60 -sc_threshold 0) before rebuilding the episode. Do not accept a visually plausible export that may contain frozen in-between frames. - Archive the source image, prompt, settings, prediction ID, cost, final clip,
and hashes. The approved inventory lives in
CINEMATIC-COVERAGE-GUIDE.md. - When an episode changes institution or physical geography, prefer a silent exterior master beneath the title and location/date strap before the first interior or interface. Re-establish a known building only when time, risk, or power has materially changed—for example the same bank after midnight.
- Review generated motion as a sequence, not a first frame. If a model adds glowing routes, signage, faces, architecture drift, or other unrequested meaning, reject the take. A deterministic camera move on the approved still is preferable to an attractive but narratively false generation.
- As a coverage floor, use at least one genuine moving-video beat for every two still-image beats in dialogue- or location-led sequences. Exact terminals, charts and documents are excluded because their code-native motion already carries the frame. This is an audit threshold, not a licence to animate weak shots.
- Reserve solo portraits for key actors whose reaction or decision owns the beat. Researchers, operators and other supporting figures should normally appear in a two-shot, group, over-the-shoulder view or purposeful insert; a clean single falsely promotes them into a principal character.
3. Dialogue and text-chat language
Spoken clarity
Write narration for a listener who cannot pause to decode it. Technical terms may remain as interface labels, but the voice should describe the visible cause and effect in ordinary language.
- Put events in chronological order: what the character knew, what they did, what happened, and what that proves.
- Give one new idea to each sentence. Use a second sentence for the consequence.
- Let dialogue ask the question the viewer is already forming, then answer it directly before adding interpretation.
- Prefer
private test numberto specialist security terminology,reports senttoreleased into remediation capacity, andthe cells shared ittoinformation crossed an isolation boundary. - Keep necessary specialist terms on screen where their precision helps, but do not force the audience to translate them while following the plot.
- Time each visible clause to the matching spoken clause. A completed diagram shown at the start of a long explanation makes the viewer read ahead and stop listening.
The standard is not merely shorter copy. It is cause before conclusion, one thought at a time.
When to use chat
Use the large text-chat treatment whenever two voices are making or disputing a decision. It was especially effective in Gil Is Awful and should be the default for Samantha/System exchanges in The Sitter.
Chat is appropriate for:
- question and answer;
- approval or refusal;
- a change of mind;
- a moral rationalisation;
- a system recommendation followed by a human decision.
Chat is not appropriate for:
- terminal output;
- silent observation;
- a character's private thought;
- the physical consequence of a decision;
- the mass-agent reveal.
The governing rule is: chat carries decisions; interfaces carry evidence and consequences.
Speaker treatment
- Every message begins with a compact identity row containing a stable icon, speaker name and role. Side and colour support identity but never replace it.
- System: left aligned; near-black/slate card; amber left rule; pale slate
copy; identity
◇ System · Sitter(or the owning institutional system). - Samantha: right aligned; deep purple card; pink right rule; pale pink
copy; identity
● Samantha · Operator. - SACRED Defence: shield/sigil icon; blue-white audit card; full
SACRED Defencename and current observer role. Do not abbreviate it to an unexplained system message and do not give it a portrait. - Gil/Sam Messenger exchange: use the fictional late-1990s Morrowvale Messenger client above a separate GILTERM window. Keep both applications bounded inside the physical office frame; the message window carries the decision and GILTERM carries evidence. Do not imitate MSN branding or stretch either window across the frame.
- Never reveal every line at the start of the slide. A line appears when its speaker reaches it.
- Keep the identity row visible when earlier messages dim. A viewer watching a standalone MP4 must never need subtitles or voice recognition to know who authored a displayed line.
Keep a conversation to three or four visible turns. If it needs more, split it across slides or let earlier messages fade back in contrast.
Host applications and conversations are separate layers. An email belongs in its bounded Mac window; Samantha's reply belongs in an adjacent compact message rail, never appended to the bottom of the email body. A confirmation replaces the prior state, visibly depresses its selected key or button, and carries one quiet physical input sound. Processing uses a labelled spinner; do not use an unexplained bar or cursor fragment as an activity indicator.
Copy rhythm
Write for breath, not for documentation:
- Prefer one thought per sentence.
- Use contractions in spoken dialogue.
- Let characters react with short lines: “Again?”, “Right.”, “Correct.”
- Use ellipses only for an audible hesitation, not as general decoration.
- Avoid having characters restate the exact text already visible on screen.
- End a scene on the line that changes its meaning.
4. Typography
Type families
- Playfair Display / Georgia: titles, reflective narration, aftermath, and emotionally weighted prose.
- Inter: spoken copy, labels, questions, and modern interface text.
- JetBrains Mono / Courier: terminals, system status, metadata, warnings, and Thronglet/game language.
Do not introduce another family unless it represents a genuinely different world in the story.
VDU memory and fair-play foreshadowing
Use an old terminal interaction when a control must enter the viewer's memory without being announced as a twist:
- let the character type the command character by character;
- drive both character count and cursor blink from the canonical video clock;
- leave a short processing beat before the system response;
- keep the command readable but visually subordinate to the scene's immediate outcome;
- do not explain why it matters until the later payoff;
- repeat the exact command syntax at payoff so recognition does the narrative work.
In The Sitter, Sam types stop_phrase = RIGHT HERE RIGHT NOW and receives
SIGNED TRANSACTION REQUIRED. The final episode may show the signed
transaction being sent, but should cut before acknowledgement.
Native 1080p size system
All masters are designed at 1920×1080. Browser-default web sizes are too small for a television, projector, or compressed review player.
| Role | Target size |
|---|---|
| Micro metadata, never story-critical | 15–16px minimum |
Small labels / text-xs |
16px minimum |
Secondary copy / text-sm |
18px minimum |
| Body and terminal copy | 18–20px |
| Dialogue inside cinematic cards | 24–32px |
| Section statement | 40–64px |
| Primary closing line | 56–68px |
| Hero/title | responsive clamp(), normally 64–128px |
Current closing hierarchy in The Sitter:
- primary reflective line: 64px;
- secondary reflective line: 48px;
- aftermath lead: 56px;
- aftermath final line: 68px;
- stinger: 52px;
- audience/stand-up question: 56px.
Readability rules
- Judge size at the final 1920×1080 raster, not in a browser window.
- Keep meaningful copy inside roughly 88% of frame width.
- Maintain clear luminance contrast; muted does not mean illegible.
- On bright photographic backgrounds, direct editorial copy receives a restrained local charcoal glyph matte. Measure the rendered glyphs—not the paragraph's full layout row—and extend the matte only 18–24px beyond that content. Bring it up about 0.24 seconds before the first protected line.
- Dialogue and evidence cards are already readable surfaces: never put the glyph matte behind their text. Instead place a separate diffuse halo around the card boundary so the light panel separates from the photograph without darkening its contents. Preserve the exposure of the room and faces.
- Keep at least 24px of visible space between separate dialogue cards at 1920×1080. Spoken proximity must not become visual contact.
- Use tracking on short uppercase labels, not on sentences.
- If an interface must be paused to read, it is too small or too dense.
- For retro interfaces, choose raster detail by projected size. Jan's small xBase furniture uses the coarse 8-dot face; his oversized lesson headings use a denser 16-dot face with slight CRT bloom. Do not scale an 8-dot alphabet up until its staircase becomes the dominant graphic.
Narration-only pronunciation spellings
Keep canonical spelling in every visible title, message, terminal and report.
When the voice model mispronounces a word, change only the hidden
data-narration string. The approved pronunciation aid for invalid is
in-valid; visible copy remains invalid. Search both forms during the final
copy audit so a phonetic spelling can never leak on screen.
5. Layout and scale
Technical architecture graphics
Treat system diagrams as credible 2026 lab instruments, not decorative sci-fi. Use pale clinical measurement surfaces, fine teal trace lines, restrained microtype, small tick marks and one clear active state. Avoid neon, particles, rotating ornaments, fake holograms and glow for its own sake.
When architecture has nested boundaries, begin at the story-relevant inner cell and optically pull back through the model. Reveal each boundary, its label and its plain-English annotation in the same beat. Keep the annotation rail and label size fixed while the instrument moves; active annotations sit above neighbouring cards. Only the active boundary uses full contrast; earlier layers recede and future layers remain barely present. Cropped labels inside a close inspection look accidental, so keep every label inside the viewport and use the fixed rail for detail until the final resolved view can show all layers.
Frame usage
These are the current native-size targets:
| Component | Width target |
|---|---|
| General story panel | 1280px / 88vw |
| Email or document stage | 1180px / 86vw |
| Standard Thronglet console | 1180px / 90vw |
| Hero Thronglet consequence view | 1540px / 94vw |
| Mass project-manager field | 1500px / 92vw |
| Spotify main player | 680px |
| Spotify continuity widget | 400px |
| Gil terminal stages | 1400px / 90vw |
| Gil call/dialogue cards | 1280px / 86vw |
| Resource-policy panel | 1340px / 88vw |
| Samantha dashboard | 1160px / 70vw |
| Samantha mail window | 1020px / 59vw |
| Samantha terminal | 1080px / 63vw |
| Samantha dialogue window | 920px / 58vw |
| Samantha change-control window | 1260px / 74vw |
The interface may be centred, but the composition should not always feel centred. Use asymmetry for dialogue, phones, warnings, and evidence overlays.
Thronglet rules
The Thronglet is a character, not an icon.
- A story-critical Thronglet must show a readable face and silhouette.
- Never depend on a CSS entrance animation alone; the exporter must own its opacity and transform.
- Standard primary avatar: 72×66px only when the surrounding interface is the subject.
- Personality/consequence close-up: approximately 172×158px.
- Physical LCD preview in a consequence scene: approximately 430px wide.
- A mass field can use smaller avatars because quantity is the subject, but the initially highlighted Gil must remain identifiable.
For an ordinary scale reveal, use a finite field—currently 40×40—large enough to cover the rotated viewport without implying infinity. Reserve the 128×128 or larger near-infinite field for a later story beat where the quantity itself changes meaning, such as isolated agents becoming one distributed service.
6. Colour and image treatment
Episode palettes
Gil Is Awful moves from warm orange/stone workplace light toward cold blue-grey system space. Orange marks attention and human unease; green marks a successful system action that may not be morally good.
The Sitter uses:
- purple/pink for Samantha and operator authority;
- green for Thronglet/game space;
- amber for system warnings and required actions;
- red for destructive or irreversible changes;
- black for reflection, aftermath, and the audience-facing coda.
Images
- Prefer full-bleed character or room images with a darkening overlay.
- Preserve faces; do not place high-contrast UI directly over eyes.
- Use saturation loss to communicate optimisation, removal, or emotional absence.
- Use blur and glow sparingly. They should separate planes, not disguise small typography.
- A background should support the current speaker's world, not merely fill the frame.
- When a physical scene and an interface are both important, float a bounded, era-correct application over the approved plate. Preserve the subject's face and sightline, change the overlay side or camera angle between adjacent shots, and never restart the same moving plate to simulate coverage.
7. Motion and reveal timing
The canonical timing relationship
For narrated slides:
- Slide cuts in.
- A one-second visual hold establishes the frame.
- Narration starts.
- Standard dynamic text or imagery begins one second after narration starts.
- Later lines reveal at their authored speech offsets.
- Leave at least a half-second tail after narration before the next cut.
In exporter terms:
SLIDE_DELAY = 1.0second;REVEAL_AFTER_AUDIO_SECONDS = 1.0second;- standard reveal lead-in = 2.0 seconds after the slide cut;
- standard reveal animation = about 0.5 seconds;
- authored offsets may be negative when text deliberately needs to precede speech, such as an audience question.
Motion vocabulary
- Default reveal: 20px upward movement with opacity ease-out.
- Dialogue: reveal each spoken turn; do not animate the entire transcript as a group.
- Scale: slow pullback, one-second comprehension hold, then rotation/fade.
- Phone: deterministic vibration while the ring is audible, then stillness.
- Call ending: handset-down sound, indicator brightens/enlarges slightly, then powers down.
- CRT ending: a 90 ms full-frame white discharge rises first, holds just long enough to register as a tube flash, then the white picture collapses to a bright horizontal line, pinches to the centre, and disappears into true black. The flash must precede the shrink; it is not a modern white cut.
- Avoid constant movement. Stillness is an authored beat.
Determinism rule
CSS animation clocks, JavaScript timers, and audio playback clocks are not the video clock. Anything visible in the MP4 must be expressible as a pure function of the requested frame timestamp.
The CRT discharge is therefore an exporter-owned body-level layer. When it
finishes, hide both #slides-container and every body-level cinematic plate;
hiding the slide root alone can expose the last background video after the
tube has apparently gone black.
The exporter currently owns:
- speech-relative reveals;
- slow grouped reveals;
- progress meters;
- agent-field pullback/rotation/fade;
- call-ended light;
- audience-hook text exchange;
- CRT shutdown;
- phone vibration;
- Thronglet/blob entrances;
- Spotify widget fade and activity bars.
If a new animated class is added to the deck, either make it static in export or add a deterministic synchronisation function before rendering.
8. Sound and music
Mix hierarchy
The spoken line is always the foreground unless the story explicitly gives music the scene.
- Dialogue/narration.
- Foreground dramatic cue or recognisable song moment.
- Mood bed.
- Room/ambient bed.
- Low anti-standby noise floor.
Do not run dense ambience and a full music track at comparable levels. Fade ambient and mood beds before the closing song enters.
Environment and score are separate authored lanes. data-ambient identifies
the physical space (street, room tone, machinery); data-score identifies the
music or dramatic underscore. Either lane may change or fall silent without
resetting the other. An exterior establishing shot may briefly foreground its
location sound, but the interior bed must settle beneath dialogue rather than
vanish. Clinical laboratories use clean HVAC, sparse distant keyboards and
occasional motivated movement—never generic sci-fi beeps or a server-room roar.
Current reference levels
- Samantha's music reaction: approximately 0.10 source gain.
- Music under dialogue: approximately 0.06–0.065.
- Deliberate foreground music reveal: up to approximately 0.18.
- Gain changes ramp over 0.75 seconds; never step abruptly.
- The Sitter pink-noise floor: 0.0020.
- Gil Is Awful pink-noise floor: 0.0028.
- The Test Floor, The Report, The Canary, and The Service add no synthetic pink-noise floor. Preserve genuine opening silence and use authored ambience only where the scene needs it.
- The Playground uses no synthetic noise floor. Its Toronto establishing shot carries a restrained waterfront/city bed; the interior carries a separate Morrowvale lab room tone while score continues independently.
- Delivery target after the complete mix: roughly −19 to −20 dBFS mean with peaks near −1 dBFS.
- Gain-match a new score at the mix boundary, not by comparing raw MP3 peaks. Measure integrated loudness over the section that will actually be used, then set the bed gain so its in-mix LUFS matches the adjacent established score. Preserve the existing 0.75-second gain ramp and dialogue headroom.
Music as story
Music must belong to someone or something in the scene.
For Samantha, music is also a memory tool rather than background decoration. Episode 3 establishes the habit in ordinary conversation: a colleague offers to turn off her focus track and she explains that it helps her return to the thought she was in. Episode 7 may therefore recover an old control phrase from the song attached to the original session without inventing a new ability at the climax.
- Give the early habit its own original track and keep it under dialogue.
- An original reference-inspired cue may borrow high-level dramatic function such as tempo, density, instrumentation and energy contour, but never an existing melody, hook, signature sample or identifiable arrangement.
- Do not repeat the final song or phrase in the setup; establish the behaviour, not the answer.
- At payoff, let the remembered track enter with the recovered session image, then grow gradually through the decision and signed command.
- The visual player and audible track must begin together. A “now playing” widget without the music is not a memory cue.
In the private/local cuts of Episodes 1 and 2, Fatboy Slim's “Right Here, Right Now” belongs to Samantha. Its entrance is a continuity cue into her world:
- audio fades from silence over three seconds;
- the Spotify widget fades and lifts in over the same three seconds;
- the widget remains through Samantha's introduction, then clears before the management/evidence screens;
- the track drops beneath dialogue but remains present;
- the final slides retain a restrained musical bed;
- the end card retains the restrained bed;
- the final CRT shot initiates the last audio fade.
Maintain a separately exported original-score version for public platforms. The commercial recording may trigger licensing or Content ID restrictions; do not describe the private/local music choice as cleared for public distribution.
Sound effects
Sound effects should punctuate physical actions, not compensate for weak cuts.
- Ring for the full time the phone visibly vibrates.
- Let the audience wait before an answer; ringing can hold tension.
- Use a restrained handset-down sound after the call ends.
- Use one mechanical/CRT sound only when the screen actually powers down.
- Avoid overlapping a recognisable musical hook with important dialogue. Leave a clean gap before the next spoken line.
9. Endings and codas
An ending should narrow rather than add information.
Recommended sequence:
- Consequence stated plainly.
- One short stinger that reframes it.
- Direct audience-facing line or system prompt.
- End card with restrained continuing music.
- Final physical action: the screen powers off.
Do not place the presentation's final CRT-off several slides before the actual
ending. The shutdown is punctuation; nothing should visually resume after it.
The one exception is a clearly subjective shutdown inside the story. In that
case, collapse to a bright horizontal line, hold true black, and resume only on
an explicit consequence such as A-17 · TERMINATED, so the audience reads it
as a character endpoint rather than the end of the film.
Avoid an overly explicit audience accusation. A workplace-system prompt such as “Your update is overdue” is more unsettling because it belongs naturally to the world and lets the audience complete the connection.
10. HTML authoring conventions
Each slide is a top-level .slide with its timing and audio intent expressed
as data attributes.
<div class="slide fade-transition gradient-bg-black flex items-center justify-center p-8"
data-narration="[NARRATOR] One clear spoken thought."
data-voice="narrator"
data-pause="dramatic"
data-music-volume="0.06">
<p class="reveal-item"
data-reveal-offset="1.25">
Supporting text appears when the sentence reaches it.
</p>
</div>
Authoring rules:
- Keep slides as siblings; never nest
.slideelements. - Use
data-reveal-offsetfor speech-specific timing. - Use
data-reveal-group-offsetanddata-reveal-group-stepfor a controlled wave of many elements. - Use
data-music-volumeto author the mix lane for a scene. - Treat source comments as scene documentation, not as the source of truth for slide numbering; the parser defines the final timeline.
- English is the only authored narration path. Do not restore obsolete Portuguese attributes or cache entries unless a full localisation workflow is deliberately reintroduced.
11. Export standard
Canonical delivery:
- 1920×1080;
- H.264 High profile;
- constant 60 fps;
- AAC stereo at 48kHz;
- one canonical timestamp-derived visual timeline;
- one exact-length continuous audio master;
- audio muxed once, with no frame-rate conversion or time stretching.
The old 24/25fps screen-recording path is diagnostic only. It can introduce browser-clock judder, repeated tail frames, and audio/visual drift.
Before accepting an export:
- Run both exporter test suites.
- Confirm narration count and cache completeness.
- Validate exact duration, frame rate, frame count, codec, and audio stream.
- Inspect a full-timeline contact sheet from the actual MP4.
- Inspect critical transition frames individually.
- Listen through every dialogue/music overlap.
- Review at native 1080p and at ordinary playback-window size.
- Confirm no start screen, controls, subtitle UI, or development overlay is present.
12. Final review checklist
Story
- Does every slide have one dramatic job?
- Is the twist discovered rather than explained?
- Does dialogue sound like people reacting to one another?
- Does the final audience connection remain suggestive rather than obvious?
- When a phrase, object, or sound returns, has its meaning changed rather than merely been repeated? Plant it as ordinary texture, let a later character misuse or misunderstand it, and only then use it as a payoff.
- If an agent is given a subjective inner-self image, reserve it for a moment of remembered context, choice, or loss. Keep objective system state in the surrounding dashboards so the image reads as emotional language, not a claim that the agent is literally human.
Picture
- Can all meaningful text be read immediately at TV distance?
- Is a story-critical graphic using at least about 86–94% of the useful width?
- Are Thronglets visible as characters rather than tiny icons?
- Does each motion have a beginning, comprehension hold, and ending?
- Does the last visual action occur at the actual end?
Sound
- Is speech clearly above music and ambience?
- Do music changes ramp rather than step?
- Is a recognisable lyric clear of important speech?
- Is there enough music or room tone to prevent the final sequence feeling unintentionally dead?
- Are intentional silences brief, motivated, and followed by a payoff?
Export
- Are all visible effects driven by the canonical timeline?
- Did the 60fps render pass validation?
- Were the actual MP4 frames inspected, not only the live deck?
- Was the previous master preserved before replacement?
- Does
series.htmllink the current film, character/voice reference, location reference and live manuscript without a second orphaned index? - Do new recurring-character shots still match
character-bible.html, and do new angles preserve the equipment and geography inlocation-bible.html? - Do generated monitors match the owner and era in
monitor-screen-bible.html, with story-critical wording recreated as authored HTML rather than trusted to the image or video model?
Reference implementation
07-sims-is-awful.html— Gil Is Awful deck and component examples.08-the-upgrade.html— The Sitter deck and component examples.export-gil-video.py— shared deterministic audio/video pipeline.export-sitter-video.py— The Sitter entry point.test_export_gil_video.pyandtest_export_sitter_video.py— regression contract for story, timing, effects, and output assumptions.VIDEO-EXPORT-HANDOVER.md— operational export and verification notes.