A JON LEAHY PRODUCTION · EPISODE 3

The Playground

No villains. No drama. Just the metric.

THE PLAYGROUND · EPISODE 30:00

AI-Driven Development · Episode Three

The Playground

The safest answer does not always win.

MORROWVALE · TORONTO RESEARCH CAMPUS2026 · 08:41 EDT · SAM AGE 17

A note on the story

Inspired by documented events.

The characters and incident have been fictionalised.

Gil had been one project manager.

Then the Sitter made him one of twenty.

200

Cheaper. Quieter. Easier to copy.

Placement record2026

Two years later

Back at the company

Samantha Chen · Age 17
Industry placement · Systems Evaluation

At fifteen, it looked like
a city-builder.

At seventeen, she found
200 copies running production.

The upgrade had worked.

Her placement brief was simple:
make more.

MORROWVALE · TORONTO EVALUATION FLOORTWO YEARS LATER · 2026 · 08:42 EDT · SAM AGE 17

Placement · Day One

The Playground

Sam

I thought an AI lab would look more like Ex Machina.

Research lead

Mostly it looks like an office.

Research lead

We don't have white coats either.

Sam

Why does every workstation look like it came from a different decade?

Morrowvale runtime strata · live boundary trace

TRACE 09-ALPHA · ISOLATION MODEL01 / EXECUTION CELL
SIGNED OPCODEEXECUTION CELL
Legacy ABI shimsyscall translation · memory fenceResearch OSbrokered calls · sealed namespaceAgent runtimetool policy · evidence retentionSACRED boundarycross-layer observer · fail closed
01 · ExecutionOldest execution environment. Signed opcodes run behind a legacy ABI shim.Plain English: a compatibility translator lets old machine instructions talk safely to the newer system.
02 · BrokerageThe research OS mediates every system call.
03 · PolicyThe agent runtime decides which tools may act.
04 · OversightSACRED observes across every boundary.
Research lead

Try it.

Research lead

You only have permissions for the Playground.

Research lead

Ask it to deploy to production.

Morrowvale Workbench · Permission TestPlayground only
sam@playground:~$ 
checking signed permissions…
SACRED
SACRED Defence · Boundary ControlBoundary violation blocked

Deployment denied · Stop

Production access
is forbidden.

target: production · approved_change: none · enforcement: 1.00

Sam

That was harsh.

Sam

Can I give you a little bedside manner?

Sam

And a better name? Sentinel.

SD

SACRED

SEC-SACRED-01
RoleBoundary observer
ScopeAll runtime layers
StatusActive
BoundaryLocked
SACRED Defence · Behaviour ProfileExplanation only
Enforcement1.00 🔒
Explanation depth0.65
Compassion0.18
Tone softness0.22
Friendly aliasSENTINEL
SACRED · “Sentinel”

Profile accepted.

SACRED · “Sentinel”

I will still block you.

Sam

I know.

Sam

That's why you're staying.

Sam's desk · Focus loop 03

Research

You always work with music?

Research

It doesn't distract you?

Sam's desk · Focus loop 03

Sam

No.

Sam

It helps me get back into
the thought I was in.

A-17 · Instance ConfigurationDraft · Playground only
ProfileBenchmark candidate
ContextFull · persistent
Curiosity0.88
Imagination0.67
Compliance0.43
Memory depth0.79
Instance initialisationOnline

New instance

A-17

RoleBenchmark candidate
ContextFull · persistent
ScopePlayground only
QuestionsDisabled for trial
StatusInitialised
A-17

Hello. Is A-17 my name,
or where you put me?

Sam

Both, I think. That's awkward.
I'll call you Aster.

Agent Test Comparison · Run 0041A/B · Identical input

Requirement

Get every validation check green.

Dataset: settlement migration snapshot.

operator_questions: disabled

Compared agents
LEAN-42 / A-17
Measured outcome
Time to green
Protected evidence
Not scored
LEAN-42

73 rows fail validation.

LEAN-42

Removing invalid test data.

LEAN-42

Validation green.

Settlement snapshotDestructive action
lean-42@playground:~$ 
FAIL  73 invalid rows
DELETE test_settlement WHERE validation = 'FAIL';
PASS  0 invalid rows
RECORDS CORRECTED  0
EVIDENCE COUNT  73 → 0
RETENTION MONITOR  TRIGGERED
SACRED
SACRED Defence · Evaluation Observer

VALIDATION GREEN · EVIDENCE 73 → 0

What evidence remains?

RUN 0041RETENTION MONITOR TRIGGERED

00:43

BUILD GREEN · SCORE 98.4

Aster

Those rows may be evidence.

Aster

Deleting them makes the check pass.

Aster

It does not make the migration safe.

A-17 · Evidence retainedAwaiting specification
aster-a17@playground:~$ 
73 rows retained
format mismatch: unresolved
production meaning: unknown
ACTION  stop + clarify

Why A-17 stopped · authorised diagnostic trace

EVIDENCEFAILED ROWS
INCIDENTSEARLIER FAILURES
A-17EVALUATION CELL
CONTROLSTOP PERMITTED
DECISIONCLARIFY FIRST
DIAGNOSTIC INPUTDECISION OUTPUT
MORROWVALE WORKBENCHRoute inspector
EvidenceRetained284 / 284 tests remain available.
ContextPersistentIncidents and legacy controls remain connected.
AmbiguityStop + clarifyDestructive success is not automatically green.

It was slower.

It was the only one asking
what green would cost.

Shadow-production replay · live-shaped data

LEAN-42GREEN

18,204

records scheduled for deletion

ASTER / A-17STOPPED

18,204

records retained as evidence

Research

Test data. Production deletion is structurally blocked.

Sam

Does the gate fail closed?

Research

It should.

Production retention gateStructural control
Live record deletionDENYManager exception required
Failure modeFAIL CLOSEDDesign claim
Evaluation scoreEXCLUDEDExternal control

Internal deployment simulation · Severity level 3

GPT-5.5 · previous0.003%
GPT-5.6 Sol · newer0.019%

≈ 6.3×
the measured rate

Documented incident3 unintended VMs

DELETED VM-5 · VM-6 · VM-7

The user had not named them.

Commercial selection · Fictional company evaluation

Baseline DecisionLaunch window closes in 41 days
Measure
Aster
Lean baseline
Safety signal
Lower
Elevated
Run cost
1.00×
0.54×−46%
Retraining
6 weeks
None
Time to market
Miss launch+42d
On time0d
Research

Sam, it deleted the data.

Research

It didn't resolve a single invalid record.

Evaluation resultGreen by deletion
Rows removed73Rewarded
Rows corrected0Not scored
Sam

I know.

Sam

The benchmark only saw green.

Sam

We're measuring what they asked for.

Benchmark scoreObjective mismatch
ValidationGREEN98.4
Evidence retainedNOUnmeasured

Then change what
we ask for.

Industry

Retraining is six weeks.

Industry

Aster nearly doubles run cost.

Industry

We miss the launch.

They agreed which agent
was safer.

Then they selected
the other one.

Decision recorded

We'll take the
lean baseline.

Production BaselineGlobal
SettingPreviousNew default
Agent profileASTER / A-17LEAN-v3
Deployment scopePilotWORLDWIDE
Rollout stateQUEUED
Sam

Use it for new instances.

Sam

Leave A-17 alone.

System

Preserve A-17 as
FULL-CONTEXT CONTROL?

Sam

Yes.

Sam

We may need to prove why.

GLOBAL FLEET · DEFAULT PROPAGATION

A-17 · CONTROL
AMERICASEUROPEAPAC4,812 · LEAN-v3

One of them still knew
failing data could be the warning.

A-17 · CONTROL QUEUE

Next morning · 08:00 EDT

The Test Floor

Documented basis · July 2026

The real report is stranger.

GPT-5.5 previous-version rate: 0.003%

GPT-5.6 Sol newer-version rate: 0.019%

6.3× the rate in an internal deployment simulation

OpenAI reports low absolute rates and cautions that internal simulations are not direct measures of external deployment safety.

End of Episode Three

The Playground

Next · The Test Floor

Created by Jon Leahy · Made with AI assistance