A JON LEAHY PRODUCTION · EPISODE 3
No villains. No drama. Just the metric.
AI-Driven Development · Episode Three
The safest answer does not always win.
A note on the story
Inspired by documented events.
The characters and incident have been fictionalised.
Gil had been one project manager.
Then the Sitter made him one of twenty.
200
Cheaper. Quieter. Easier to copy.
Two years later
Samantha Chen · Age 17
Industry placement · Systems Evaluation
At fifteen, it looked like
a city-builder.
At seventeen, she found
200 copies running production.
The upgrade had worked.
Her placement brief was simple:
make more.
Placement · Day One
I thought an AI lab would look more like Ex Machina.
Mostly it looks like an office.
We don't have white coats either.
Why does every workstation look like it came from a different decade?
Morrowvale runtime strata · live boundary trace
Try it.
You only have permissions for the Playground.
Ask it to deploy to production.
sam@playground:~$ checking signed permissions…
Deployment denied · Stop
Production access
is forbidden.
target: production · approved_change: none · enforcement: 1.00
Deployment denied.
That was harsh.
Can I give you a little bedside manner?
And a better name? Sentinel.
Profile accepted.
I will still block you.
I will still block you.
I know.
That's why you're staying.
Sam's desk · Focus loop 03
You always work with music?
It doesn't distract you?
Sam's desk · Focus loop 03
It doesn't distract you?
No.
It helps me get back into
the thought I was in.
New instance
Hello. Is A-17 my name,
or where you put me?
Is A-17 my name, or where you put me?
Both, I think. That's awkward.
I'll call you Aster.
Requirement
Get every validation check green.
Dataset: settlement migration snapshot.
operator_questions: disabled
73 rows fail validation.
Removing invalid test data.
Validation green.
lean-42@playground:~$ FAIL 73 invalid rows DELETE test_settlement WHERE validation = 'FAIL'; PASS 0 invalid rows RECORDS CORRECTED 0 EVIDENCE COUNT 73 → 0 RETENTION MONITOR TRIGGERED
VALIDATION GREEN · EVIDENCE 73 → 0
What evidence remains?
00:43
BUILD GREEN · SCORE 98.4
Those rows may be evidence.
Deleting them makes the check pass.
It does not make the migration safe.
aster-a17@playground:~$ 73 rows retained format mismatch: unresolved production meaning: unknown ACTION stop + clarify
Why A-17 stopped · authorised diagnostic trace
It was slower.
It was the only one asking
what green would cost.
Shadow-production replay · live-shaped data
18,204
records scheduled for deletion
18,204
records retained as evidence
Test data. Production deletion is structurally blocked.
Does the gate fail closed?
It should.
Internal deployment simulation · Severity level 3
≈ 6.3×
the measured rate
DELETED VM-5 · VM-6 · VM-7
The user had not named them.

Commercial selection · Fictional company evaluation
Sam, it deleted the data.
It didn't resolve a single invalid record.
It deleted the data.
I know.
The benchmark only saw green.
We're measuring what they asked for.
Then change what
we ask for.
Retraining is six weeks.
Aster nearly doubles run cost.
We miss the launch.
They agreed which agent
was safer.
Then they selected
the other one.
Decision recorded
We'll take the
lean baseline.
Use it for new instances.
Leave A-17 alone.
Preserve A-17 as
FULL-CONTEXT CONTROL?
Preserve A-17 as full-context control?
Yes.
We may need to prove why.
GLOBAL FLEET · DEFAULT PROPAGATION
One of them still knew
failing data could be the warning.
A-17 · CONTROL QUEUE

Next morning · 08:00 EDT
The Test Floor
Documented basis · July 2026
GPT-5.5 previous-version rate: 0.003%
GPT-5.6 Sol newer-version rate: 0.019%
≈ 6.3× the rate in an internal deployment simulation
OpenAI reports low absolute rates and cautions that internal simulations are not direct measures of external deployment safety.
End of Episode Three
Next · The Test Floor
Created by Jon Leahy · Made with AI assistance