endoskeletal kernel 1 · lib 2026.09

Running an organization

Simulating your organization

Run the organization you declared against a world you script, years in minutes. The specification runs unchanged: the same seats, rules, envelopes, votes and agents. The world around it is emulated: providers, regions, prices, customers, a market, and surrogate models. You watch from outside, with the script in hand. The organization only knows what it observes.

esk simulate northwind.sim.esk --run Growth18 month 7 · day 214 ▶▶ fast-forward
WORLD SCRIPT · visible to you, never to the organizationTHE ORGANIZATION’S LOG · only what it has observedM0M3M6M9M12M15M18Self-serve launchCompute price +38 %EU residency dealeu-west-1 outagein 18 dModel v3 deprecatedToken leakedBilling v2 launchTraffic ×3day 214

In 18 days · Region eu-west-1 goes down for 26 hours.

Since day 196, Database[EU] has been bound in eu-west-1 only. Nothing in the org’s log says that is a risk.

The organization can’t see this yet. You can.

The organization10 people · 7 agents

NorthwindCEOCTOPlatformDeployerOpsOnCall ×2ArchitectSupportSupportTriageBotTriageSupportLeadGrowthPMEngineers ×3Sales EU ×2SecuritySecLeadAdvisoryInfraCouncil3 seatsFinanceFinLeadFinance
  • CEO
  • CTO
  • PlatformDeployerOpsOnCall ×2Architect
  • SupportSupportTriageBotTriageSupportLead
  • GrowthPMEngineers ×3Sales EU ×2
  • SecuritySecLeadAdvisory
  • InfraCouncil3 seats
  • FinanceFinLeadFinance

person agent model program

MRR

100k USD

Availability · 30 d

99.975%

Spend per month

35k USD

Reasoning share

4%

Dashed: the part of the run you haven’t reached. The simulation already knows it.

  1. day 214Monthly close: spend 34,900 USD; MRR 97,600 USD.
  2. day 200Amendment in force: scope Security, SecurityLead, Advisory model.
  3. day 196CTO binds VectorSearch[EU] and Database[EU] in eu-west-1.
  4. day 178Sandbox experiment: EU composition, p99 41 ms.
  5. day 171gap VectorSearch: no EU realization; one candidate is unknown.
  6. day 170cu.ResidencyRequired(Hanse Mutual, EU).

Above, 18 months of Northwind in about 90 seconds. The upper track is the world's script, and you can read all of it from day 0. The lower track is the organization's log, and it fills in only as the organization observes things. When a scripted event is close, the panel under the timeline says what is coming and what in the organization's current state makes it vulnerable. That is information the simulation has and the organization doesn't. The clock slows down around each event so you can watch the response, then speeds up again.

The outage on day 232 is the clearest case. On day 196 the organization bound its EU database in one region, which was correct under the rules it had. From that moment the simulation knows the exposure, because it knows the script. The organization finds out when its monitor reports the region down. Its Ops agent fails over inside its envelope in 14 minutes, and eight days later the organization amends its own rules to require two regions.

Where the future lives

A simulation is only worth running if the organization can't cheat. Three rules keep the future on your side of the glass:

Where the future lives in a simulationexecute(…) honoured by the worldyou also read the logScenariothe world's scriptdays and eventsWorld emulatorsproviders · customersmarket · model surrogatesWorld partiesobserve · asserton their own clocksThe org's logthe only state it hasOrg partiespeople · agents · programsYou, watchingscript + log + thealready-computed runNoPeeking: no organization party may read the scenario · checked by the leak audit
The organization learns about the world the same way it would in production: through events that world parties record. The script never reaches it.
Text description of this diagram

The scenario (the world's script) drives world emulators: providers, customers, the market and model surrogates. World parties observe and assert what happens, on their own clocks, and those events are the only way anything reaches the organization's log. Organization parties read their log and act; their executes are honoured by the world. The viewer reads both the scenario and the log. A constitutional norm, NoPeeking, prohibits any organization party from reading the scenario, and a leak audit checks the finished run for violations.

  • The script is data for the world, not for the organization. World emulators read it. Organization parties can't: NoPeeking is a constitutional, non-defeasible prohibition, and the leak audit checks every finished run for a read by any organization party.
  • Everything reaches the organization as an event. A price change arrives as a provider's lifecycle claim on the day it is published. An outage arrives when the monitor observes it. A customer's demand arrives as a signed requirement. The organization learns exactly the way it would in production, under the same trust policies.
  • Agents don't get special access either. Model-based seats run as surrogates or as replays of recorded determinations. They see the same log as every other party, and nothing else.

What you declare to run one

A simulation is a small module next to your organization. It names the organization, the world, the script, the horizon and the seed. It doesn't copy or modify the organization; it imports it.

northwind.sim.eskesk
module northwind.sim @ 1.0.0
  import northwind.core @ 1.* as nw
  import lib.sim        @ 1.* as sim

use sim.Run(
  org      = nw.Northwind,
  world    = sim.World(
    providers = sim.Emulated(14),
    customers = sim.Cohorts("customers-2026q3.json"),
    market    = sim.Growth(6 % per month)
  ),
  scenario = "growth-18m.scenario.json",
  horizon  = 18 months,
  seed     = 7,
  models   = sim.Surrogate
) as Growth18

use sim.Counterfactual(
  of    = Growth18,
  amend = "amendments/two-regions.esk"
) as TwoRegions

-- the organization may never read the script it is being tested against
norm NoPeeking {
  prohibition on any nw.Party
  aim occurred read(sim.Scenario, _)
  level constitutional  not tradeable  not defeasible
}

The script is a list of world events on days. It says what happens in the world, never what the organization should do about it.

growth-18m.scenario.jsonjson
{
  "scenario": "growth-18m",
  "days": 540,
  "events": [
    { "day": 45,  "kind": "market.launch",     "plan": "self-serve", "signup-multiplier": 2.4 },
    { "day": 110, "kind": "pv.price-change",   "provider": "provider-a", "offer": "compute.std",
      "delta": "+38 %", "effective-day": 140 },
    { "day": 170, "kind": "cu.deal",           "customer": "Hanse Mutual", "requires": { "residency": "EU" } },
    { "day": 232, "kind": "pv.region-outage",  "provider": "provider-c", "region": "eu-west-1",
      "duration": "26 h", "observable-by": "monitor" },
    { "day": 300, "kind": "pv.deprecation",    "provider": "model-co", "offer": "model-v3", "notice": "90 d" },
    { "day": 360, "kind": "sec.leak",          "credential": "tok-prod-deploy", "at": "02:14" },
    { "day": 420, "kind": "product.release",   "change": "billing-v2", "new-ticket-kinds": 3 },
    { "day": 480, "kind": "market.spike",      "traffic-multiplier": 3.0, "decay": "30 d" }
  ]
}

Runs are deterministic. The world's randomness comes from a keyed stream over the seed, the clock is discrete-event, and model seats are surrogates by default, so the same specification, script and seed always produce the same log and the same hash. Change any of the three and the hash changes, which is how you know a comparison is fair.

terminalconsole
$ esk simulate northwind.sim.esk --run Growth18
world        14 emulated providers · 3 regions · 212 customer cohorts · models: surrogate
horizon      540 days · seed 7 · discrete-event clock
running      ████████████████████ 540/540 days · 93,412 events · 71 s
log          sha256-4be1…9c02   same spec + same seed → same hash
leak audit   0 reads of the scenario by organization parties
summary      11 binds · 6 amendments · 3 council votes · 1 outage (14 min, failover)
             people 2 → 13 · agents 3 → 8 · reasoning share 18 % → 5 %

$ esk simulate northwind.sim.esk --run TwoRegions --diff Growth18
first divergence   day 196 (bind Database[EU]: 2 regions instead of 1)
                        Growth18       TwoRegions
downtime                14 min         0 min
SLO budget spent        31 %           2 %
people paged            3              0
spend, 18 months        2.410 M USD    2.415 M USD

Counterfactuals: the same world, a different organization

The most useful run is usually the second one. Keep the seed and the script, change the organization, and compare. Below, the only change is one property added before day 0: Database must run in at least two regions.

Same seed · same world script · days 180–260

day 196binding Database[EU]
day 232eu-west-1 goes down
days 240–252afterwards
Growth18 · as written
One region: eu-west-1
14 min down · SLO budget −31 % · 3 people paged
Amends its rules; second region bound on day 252
TwoRegions · amended before day 0
Two regions: eu-west-1 and eu-central-1
No impact: eu-central-1 keeps serving
Nothing to change
over the whole runas writtenamended
downtime14 min0 min
SLO error budget spent31 %2 %
people paged30
spend over 18 months2.410 M USD2.415 M USD
first divergence—day 196
The only difference between the two runs is one property added before day 0: Database must run in at least two regions.

The as-written organization ends up with two regions anyway: it amends its rules on day 240, after the outage. The amended one simply has them from day 196, so the whole difference is 56 days of a second region, about 5,000 USD, against 14 minutes of downtime, 29 points of error budget and three people paged. Most counterfactuals aren't this one-sided, and the simulation doesn't decide which trade is worth making. It gives the people who decide a comparison in which nothing differs except the thing they are deciding about.

Typical counterfactuals for a software company:

  • Authority. Widen or narrow an agent's envelope and see how many decisions reach a person, and how long they wait.
  • Budget. Lower a budget or its reserve and see when the freeze fires, and what it stops.
  • Structure. Split a team, add a council, or remove a seat and see where commitments pile up.
  • Knowledge. Turn off pooled evidence and see how much longer the organization takes to find alternatives on its own.
  • Seeds. Run the same organization and script under several seeds to separate what your rules cause from what luck causes.

Foresight in the running organization

You don't have to wait for a new specification to simulate. The running organization can fork itself from any checkpoint of its own log and play the next 90 days against a library of scripted worlds: outages in each region it uses, price shocks, demand spikes, a deprecation of every model it depends on. It never sees what will actually happen. It sees what would happen in each world, which is the honest version of foresight.

northwind.esk · scope Architectureesk
norm StressTest {
  obligation of Architect to persona
  when occurred tick(t) and t.boundary = "month"
  aim within 3 d: occurred derive(sim.StressReport(t.month)) by Architect
}

-- a simulated breach is a true fact about simulations, never about the world
norm ActOnForesight {
  obligation of Architect to persona
  when holds sim.breach-share(ApiUp, next = 90 d, worlds = StressWorlds) >= 0.4
  aim within 10 d: occurred assert(arch.Proposal(_)) by Architect
}

use ev.TrustPolicy(
  source-kind = sim.Lab,
  base        = T,
  vocab       = { sim.breach-share, sim.forecast }
) as LabResults

Stress results come back as claims from the lab. They are true as statements about simulations, and the trust policy says so. They never become claims about the world. When 40 % or more of the stress worlds show the availability objective breaking within 90 days, the Architect is obliged to propose something. That is how the day-232 exposure would have been caught in production: not by knowing the outage was coming, but by noticing that half of the plausible worlds contain one.

Keeping simulations honest

  • Emulated numbers are labelled as emulated. Prices, latencies and outages in a run come from the script and the emulators, and every report says so.
  • Model seats say how they ran. A surrogate stands in for a model, a replay reuses recorded determinations, and a live run calls the real model at real cost. Each run's manifest records which.
  • The run plan is itself a set of predictions. Before a long run, the lab predicts events, cost and wall time as claims. It revises them as the run proceeds and scores them against the outcome, so you learn how far to trust the next plan.
  • Any moment can be rebuilt. Checkpoints let you replay a run to any day and get byte-identical state, which is how a surprising result becomes a finding instead of an anecdote.

See also Watching it operate for the views that work the same on a live log and a simulated one, and Reasoning into rules, and back for what the reasoning-share indicator in the capture is measuring.