# Simulating your organization

> Run the organization you declared against a world you script, years in minutes. The specification runs unchanged: the same seats, rules, envelopes, votes and agents. The world around it is emulated: providers, regions, prices, customers, a market, and surrogate models. You watch from outside, with the script in hand. The organization only knows what it observes.

Canonical: https://endoskeletal.com/simulate/
Last updated: 2026-09-28

**Fast-forward capture: Northwind over 18 months.** The page shows a timeline with two tracks. The world script (upper track) is visible to the viewer from the start; the organization's log (lower track) fills in only as the organization observes things. As the playhead approaches a world event, a foresight panel says what is about to happen and what in the organization's current state makes it vulnerable. The org chart grows as seats are created, and four indicators (MRR, availability, spend, share of work handled by reasoning) are drawn up to the playhead, with the rest of the already-computed run shown dashed.

World script:

1. **Day 45: The self-serve plan launches. Signups more than double within six weeks.** What the simulation knows beforehand: Support is a single model seat. Ticket volume is about to double, and nothing in the org says so yet. What happened: Backlog noticed on day 49. A second triage seat was appointed on day 52; a distilled program took over on day 74.
2. **Day 110: Provider A announces a 38 % price rise on standard compute, effective day 140.** What the simulation knows beforehand: Compute runs entirely on provider A. Provider B is 22 % cheaper and nine organizations in the pool have measured it. What happened: The org read the announcement the day it was published and finished moving to provider B on day 134, six days before the rise.
3. **Day 170: A German insurer will sign, on condition that its data stays in the EU.** What the simulation knows beforehand: No EU realization of VectorSearch exists. The frontier has three candidates; one of them is unknown. What happened: EU realizations bound on day 196, in a single region: eu-west-1.
4. **Day 232: Region eu-west-1 goes down for 26 hours.** What the simulation knows beforehand: Since day 196, Database[EU] has been bound in eu-west-1 only. Nothing in the org's log says that is a risk. What happened: Ops failed over to eu-central-1 in 14 minutes, inside its envelope. Residency held on every copy. A two-region requirement was added on day 240.
5. **Day 300: The support model is deprecated with 90 days' notice.** What the simulation knows beforehand: The triage program handles 96 % of tickets. The other 4 % depend on model v3. What happened: Model v4 ran in shadow for 22 days and was promoted on day 324 at 0.97 agreement.
6. **Day 360: A production deploy token is pushed to a public fork at 02:14.** What the simulation knows beforehand: The token can deploy Api. Rotation is inside the Ops envelope; revoking customer tokens is not. What happened: Rotated in 61 seconds. A broader revoke that would have broken 12 customers was refused as a self-approval.
7. **Day 420: Billing v2 ships, with proration and invoice previews.** What the simulation knows beforehand: The triage program has never seen a proration ticket. What happened: Abstention rose to 31 % for six days while the model covered the gap; v5 of the program was promoted on day 427.
8. **Day 480: A post goes viral overnight and API traffic triples.** What the simulation knows beforehand: Api runs 12 units at 61 % utilization. Three times the load needs 20 units, which is outside the Ops envelope. What happened: The council blocked the scale-up at 1 of 3, re-voted when need rose 25 %, and approved 20 units on day 483.

Key entries in the organization's log:

- Day 0 (person): Founded: constitution enacted. Six seats, three of them filled by agents.
- Day 52 (person): CEO appoints TriageBot to Support.
- Day 60 (person): Amendment in force: scope Growth, with a PM and three engineers.
- Day 74 (person): OpsLead promotes TriageProgram v1 (agreement 0.96).
- Day 110 (evidence): pv.price-change(compute.std) +38 %, effective day 140.
- Day 121 (person): CTO binds Compute to provider B.
- Day 170 (evidence): cu.ResidencyRequired(Hanse Mutual, EU).
- Day 196 (person): CTO binds VectorSearch[EU] and Database[EU] in eu-west-1.
- Day 200 (person): Amendment in force: scope Security, SecurityLead, Advisory model.
- Day 232 (blocked): ops.availability(Api/EU) is F: region outage observed by the monitor.
- Day 240 (person): Amendment in force: Database property regions >= 2.
- Day 260 (person): InfraCouncil formed: three seats.
- Day 300 (evidence): pv.deprecation(model v3), 90 days' notice.
- Day 324 (person): OpsLead promotes model v4.
- Day 330 (person): Amendment in force: scope Finance, FinanceLead, Finance agent.
- Day 360 (evidence): sec.exposed(tok-prod-deploy) reported by the code host.
- Day 400 (person): Two EU sales hires join Growth.
- Day 421 (reasoning): TriageProgram abstains on 31 %; the model takes the rest.
- Day 480 (evidence): ops.utilization(Api) = 0.93.
- Day 483 (person): Need rose 25 %: re-vote 2 of 3, Scale(20) executes.
- Day 539 (scripted): Horizon reached: 93,412 events, log sha256-4be1…9c02.

Above, 18 months of Northwind in about 90 seconds. The upper track is the world's script, and you can read all of it from day 0. The lower track is the organization's log, and it fills in only as the organization observes things. When a scripted event is close, the panel under the timeline says what is coming and what in the organization's current state makes it vulnerable. That is information the simulation has and the organization doesn't. The clock slows down around each event so you can watch the response, then speeds up again.

The outage on day 232 is the clearest case. On day 196 the organization bound its EU database in one region, which was correct under the rules it had. From that moment the simulation knows the exposure, because it knows the script. The organization finds out when its monitor reports the region down. Its Ops agent fails over inside its envelope in 14 minutes, and eight days later the organization amends its own rules to require two regions.

## Where the future lives

A simulation is only worth running if the organization can't cheat. Three rules keep the future on your side of the glass:

**Diagram: Where the future lives in a simulation.**

The scenario (the world's script) drives world emulators: providers, customers, the market and model surrogates. World parties observe and assert what happens, on their own clocks, and those events are the only way anything reaches the organization's log. Organization parties read their log and act; their executes are honoured by the world. The viewer reads both the scenario and the log. A constitutional norm, NoPeeking, prohibits any organization party from reading the scenario, and a leak audit checks the finished run for violations.

*The organization learns about the world the same way it would in production: through events that world parties record. The script never reaches it.*

- **The script is data for the world, not for the organization.** World emulators read it. Organization parties can't: `NoPeeking` is a constitutional, non-defeasible prohibition, and the leak audit checks every finished run for a read by any organization party.
- **Everything reaches the organization as an event.** A price change arrives as a provider's lifecycle claim on the day it is published. An outage arrives when the monitor observes it. A customer's demand arrives as a signed requirement. The organization learns exactly the way it would in production, under the same trust policies.
- **Agents don't get special access either.** Model-based seats run as surrogates or as replays of recorded determinations. They see the same log as every other party, and nothing else.

## What you declare to run one

A simulation is a small module next to your organization. It names the organization, the world, the script, the horizon and the seed. It doesn't copy or modify the organization; it imports it.

*northwind.sim.esk*

```esk
module northwind.sim @ 1.0.0
  import northwind.core @ 1.* as nw
  import lib.sim        @ 1.* as sim

use sim.Run(
  org      = nw.Northwind,
  world    = sim.World(
    providers = sim.Emulated(14),
    customers = sim.Cohorts("customers-2026q3.json"),
    market    = sim.Growth(6 % per month)
  ),
  scenario = "growth-18m.scenario.json",
  horizon  = 18 months,
  seed     = 7,
  models   = sim.Surrogate
) as Growth18

use sim.Counterfactual(
  of    = Growth18,
  amend = "amendments/two-regions.esk"
) as TwoRegions

-- the organization may never read the script it is being tested against
norm NoPeeking {
  prohibition on any nw.Party
  aim occurred read(sim.Scenario, _)
  level constitutional  not tradeable  not defeasible
}
```

The script is a list of world events on days. It says what happens in the world, never what the organization should do about it.

*growth-18m.scenario.json*

```json
{
  "scenario": "growth-18m",
  "days": 540,
  "events": [
    { "day": 45,  "kind": "market.launch",     "plan": "self-serve", "signup-multiplier": 2.4 },
    { "day": 110, "kind": "pv.price-change",   "provider": "provider-a", "offer": "compute.std",
      "delta": "+38 %", "effective-day": 140 },
    { "day": 170, "kind": "cu.deal",           "customer": "Hanse Mutual", "requires": { "residency": "EU" } },
    { "day": 232, "kind": "pv.region-outage",  "provider": "provider-c", "region": "eu-west-1",
      "duration": "26 h", "observable-by": "monitor" },
    { "day": 300, "kind": "pv.deprecation",    "provider": "model-co", "offer": "model-v3", "notice": "90 d" },
    { "day": 360, "kind": "sec.leak",          "credential": "tok-prod-deploy", "at": "02:14" },
    { "day": 420, "kind": "product.release",   "change": "billing-v2", "new-ticket-kinds": 3 },
    { "day": 480, "kind": "market.spike",      "traffic-multiplier": 3.0, "decay": "30 d" }
  ]
}
```

Runs are deterministic. The world's randomness comes from a keyed stream over the seed, the clock is discrete-event, and model seats are surrogates by default, so the same specification, script and seed always produce the same log and the same hash. Change any of the three and the hash changes, which is how you know a comparison is fair.

*terminal*

```console
$ esk simulate northwind.sim.esk --run Growth18
world        14 emulated providers · 3 regions · 212 customer cohorts · models: surrogate
horizon      540 days · seed 7 · discrete-event clock
running      ████████████████████ 540/540 days · 93,412 events · 71 s
log          sha256-4be1…9c02   same spec + same seed → same hash
leak audit   0 reads of the scenario by organization parties
summary      11 binds · 6 amendments · 3 council votes · 1 outage (14 min, failover)
             people 2 → 13 · agents 3 → 8 · reasoning share 18 % → 5 %

$ esk simulate northwind.sim.esk --run TwoRegions --diff Growth18
first divergence   day 196 (bind Database[EU]: 2 regions instead of 1)
                        Growth18       TwoRegions
downtime                14 min         0 min
SLO budget spent        31 %           2 %
people paged            3              0
spend, 18 months        2.410 M USD    2.415 M USD
```

## Counterfactuals: the same world, a different organization

The most useful run is usually the second one. Keep the seed and the script, change the organization, and compare. Below, the only change is one property added before day 0: Database must run in at least two regions.

**Counterfactual fork.** Two runs of the same seed and world script: Growth18 as written, and TwoRegions, amended before day 0 to require two regions for Database. In Growth18 the day-232 outage causes 14 minutes of downtime; in TwoRegions it has no impact.

- downtime: as written 14 min, amended 0 min
- SLO error budget spent: as written 31 %, amended 2 %
- people paged: as written 3, amended 0
- spend over 18 months: as written 2.410 M USD, amended 2.415 M USD
- first divergence: as written —, amended day 196

The as-written organization ends up with two regions anyway: it amends its rules on day 240, after the outage. The amended one simply has them from day 196, so the whole difference is 56 days of a second region, about 5,000 USD, against 14 minutes of downtime, 29 points of error budget and three people paged. Most counterfactuals aren't this one-sided, and the simulation doesn't decide which trade is worth making. It gives the people who decide a comparison in which nothing differs except the thing they are deciding about.

Typical counterfactuals for a software company:

- **Authority.** Widen or narrow an agent's envelope and see how many decisions reach a person, and how long they wait.
- **Budget.** Lower a budget or its reserve and see when the freeze fires, and what it stops.
- **Structure.** Split a team, add a council, or remove a seat and see where commitments pile up.
- **Knowledge.** Turn off pooled evidence and see how much longer the organization takes to find alternatives on its own.
- **Seeds.** Run the same organization and script under several seeds to separate what your rules cause from what luck causes.

## Foresight in the running organization

You don't have to wait for a new specification to simulate. The running organization can fork itself from any checkpoint of its own log and play the next 90 days against a library of scripted worlds: outages in each region it uses, price shocks, demand spikes, a deprecation of every model it depends on. It never sees what will actually happen. It sees what would happen in each world, which is the honest version of foresight.

*northwind.esk · scope Architecture*

```esk
norm StressTest {
  obligation of Architect to persona
  when occurred tick(t) and t.boundary = "month"
  aim within 3 d: occurred derive(sim.StressReport(t.month)) by Architect
}

-- a simulated breach is a true fact about simulations, never about the world
norm ActOnForesight {
  obligation of Architect to persona
  when holds sim.breach-share(ApiUp, next = 90 d, worlds = StressWorlds) >= 0.4
  aim within 10 d: occurred assert(arch.Proposal(_)) by Architect
}

use ev.TrustPolicy(
  source-kind = sim.Lab,
  base        = T,
  vocab       = { sim.breach-share, sim.forecast }
) as LabResults
```

Stress results come back as claims from the lab. They are true as statements about simulations, and the trust policy says so. They never become claims about the world. When 40 % or more of the stress worlds show the availability objective breaking within 90 days, the Architect is obliged to propose something. That is how the day-232 exposure would have been caught in production: not by knowing the outage was coming, but by noticing that half of the plausible worlds contain one.

## Keeping simulations honest

- **Emulated numbers are labelled as emulated.** Prices, latencies and outages in a run come from the script and the emulators, and every report says so.
- **Model seats say how they ran.** A surrogate stands in for a model, a replay reuses recorded determinations, and a live run calls the real model at real cost. Each run's manifest records which.
- **The run plan is itself a set of predictions.** Before a long run, the lab predicts events, cost and wall time as claims. It revises them as the run proceeds and scores them against the outcome, so you learn how far to trust the next plan.
- **Any moment can be rebuilt.** Checkpoints let you replay a run to any day and get byte-identical state, which is how a surprising result becomes a finding instead of an anecdote.

See also [Watching it operate](https://endoskeletal.com/operate/) for the views that work the same on a live log and a simulated one, and [Reasoning into rules, and back](https://endoskeletal.com/crystallize/) for what the reasoning-share indicator in the capture is measuring.
