Non-Fixed Exams: A Genealogy
Companion to the TD genealogy map. That map covered the build-feed side of the family — what randomizes the stream of units/cards/towers you build from. This map covers the dual question: the exam side — what generates the challenge your build is graded against, and what happens to each answer over time.
Vocabulary carried over from the Micro/Meso/Macro session and the Bayesian-epistemology discussion that followed it:
- Exam — the thing that grades the build: wave schedules, maps, fights, raids, weather.
- Fixed exam — authored once, identical across runs. Produces answer keys (Bloons CHIMPS guides, Arknights stage videos).
- The imagined-wiki test — a mode is collapse-resistant to the degree that its hypothetical wiki must be a textbook of methods rather than a table of values, and actually unbreakable only if even the methods textbook goes stale because runs keep demanding concepts not yet in it.
- The Ngo rule — a game is breakable once its hypothesis space is enumerated; published distributions are pre-constructed model spaces handed to the community. So: publish the rules, never print the distributions. An emergent exam is one where the distribution exists nowhere — not in data tables, not in the design doc, not in the designer's head — because it is generated, not authored.
1. The six answers to "who authors the exam?"
The card-TD genealogy was a trunk built from one question asked repeatedly: what randomizes the build feed? The exam side has its own recurring question — who authors the exam? — and every game in and around the family gives one of six answers.
| # | Answer | Exemplars | Uncertainty type | Collapse signature (imagined wiki) |
|---|---|---|---|---|
| 1 | Nobody — authored once (fixed) | Classic TD, Arknights, puzzle games | None | Answer key. Per-map build orders. Total collapse. |
| 2 | A distribution, sampled per run | Rogue → NetHack → Spelunky → StS → Monster Train | Aleatory, distribution authored | Method wiki. Tier lists + pathing heuristics + EV charts. Slow collapse, Bayesian-ready from birth. |
| 3 | The player (scheduled) | StS routing, Against the Storm forest modifiers, ascension/heat, AI War's AIP, RoR's difficulty clock | Aleatory + self-referential | Routing EV tables. Collapses unless the thing being scheduled is itself uncertain. |
| 4 | A concealed director | Mario Kart rubber-banding (1992) → RE4 (2005) → L4D Director (2008) → RimWorld storytellers | Epistemic (hidden policy) | Director-reading. Players reverse the policy, then suppress its input variable. Also fails the fairness/attribution constraint outright. |
| 5 | Other minds | Wintermaul/Legion sends → auto-battler PvP → battle royales | Entity (non-stationary) | Unbreakable — but it's the venue, not the game. The industry's default fix for exam staleness is "add humans." Rejected per the session's constitution-vs-citizens cut. |
| 6 | A visible simulation | Dwarf Fortress, Factorio, Rain World, Creeper World, ONI, OTC | Emergent (irreducible or reflexive) | Field guide at best — climate knowledge that never resolves into trajectories. The only branch whose best members pass the stale-textbook bar. |
Three observations about the table before the islands:
Branches 2–4 are all Bayesian-ready. Sampled exams publish their distribution through volume and datamining; scheduled exams publish their options on screen; directors get reverse-engineered into an effective distribution. In each case the community ends up holding an enumerated hypothesis space, and play inside it becomes credence-updating — which automates. The collapse rates differ by orders of magnitude (StS is still partially alive at year eight; a director gets read in months), but the asymptote is shared.
Branch 3 is not really an uncertainty source — it's a commitment amplifier. Scheduling your own exam (wave drafting, StS pathing, AStS modifiers) is the aleatory commit engine pointed at the exam side. It composes beautifully with everything, but on its own it grades against known quantities. Note that Risk of Rain's difficulty clock belongs here and is the reason looting feels like a routing problem: the player schedules the exam continuously by choosing how long to loot. The clock is the most elegant member of this branch — and it still produced wiki pages of stage-timing rules, on schedule.
Branch 6 is the subject of this document, and its headline property is negative: it has no phylogeny. The fixed-exam tree from the previous map is a dense lineage — every generation visibly inherits organs from its parent. The emergent branch is a scattering of islands separated by decades, where almost nobody iterated on anybody. The rest of this doc walks the islands, explains why no lineage formed, and extracts what's portable.
2. The islands, chronologically
For each island: the mechanism, its visibility grade, its coupling to the player, the shape of its wiki, and the lesson.
2.1 The economy root — The Sumerian Game (1964) / Hamurabi (1968/73)
The oldest digital exam is emergent-ish. Hamurabi has no levels, no waves, no authored challenge at all: you allocate grain, the population/land/harvest difference equations respond, and starvation is the grade. The dynamics ARE the exam. The lesson is genealogical: management games began on this branch — the fixed exam (levels, waves) is the later invention, imported from arcade structure. The lineage is old; it's just thin, because the dynamics in this root were simple enough to solve by hand.
2.2 M.U.L.E. (1983) — the market made of players
Auction-driven economy where prices emerge from supply, demand, and the other three players' positions. Half entity-exam (branch 5), half economy-exam, but it established the reflexive-market mechanism: a price that responds to your own purchases is partially unpredictable to you even with perfect information, because predicting it requires predicting your own future trades. Self-reference as an uncertainty source, four decades early.
2.3 The divorce — SimCity (1989) and the Wright line
The single most important event in the genealogy, and it's an absence. Will Wright built city dynamics out of Jay Forrester's system-dynamics work — exactly the simulation substance an emergent exam needs — and then deliberately refused to grade the player. "Software toys": emergence kept, exam dropped. SimEarth, SimAnt, The Sims doubled down. The people best at building dynamics opted out of grading; the people who grade (game designers proper) kept their exams authored. The two skills split into separate genres in 1989 and have barely re-merged since. This divorce, more than any technical barrier, is why branch 6 is islands instead of a tree.
2.4 Ultima Online's ecology (1997) — eaten by its players
The first industrial attempt to ship an ecology as content: Raph Koster's closed-loop resource system, with herbivores, predators, and resource flows that player actions fed into. It died almost immediately — players slaughtered and harvested everything faster than the loop could regenerate, before most of them ever noticed the system existed. The canonical lesson: players are an extraction and optimization pressure that exceeds any closed loop's regeneration rate, given unbounded time. UO gave them unbounded time (persistent world). The fix nobody stated for twenty years: bound the time. See §5.3.
2.5 Dwarf Fortress (2006) — the re-merge
The island where sim and exam finally re-married. The exam emerges from several stacked systems: sieges scaled by created wealth (a scalar coupling — and DF players do thermostat it; "wealth management" predates RimWorld), fluid and temperature physics, animal and civilization populations, and above all cascades — the tantrum spiral, where one death sours moods, sour moods cause fights, fights cause deaths. The cascade is the genuinely emergent member: nobody authored it; it falls out of the social sim's coupling.
The wiki test grades DF honestly: the buildable layer wiki'd completely (quickfort blueprints, optimal fort layouts — the ONI failure mode, see 2.13), the trajectory layer never did. The DF wiki is a field guide — what creatures do, what moods mean — not an answer key; what veterans hold is exactly the unwritable embodied model.
Two more things DF shipped that matter here:
- "Losing is fun" — the community's cultural technology for variance acceptance. An emergent exam will occasionally hand you an unwinnable cascade, and DF's audience converted that from a complaint into the genre's motto. This is the equanimity gym norm, named and adopted at scale. It proves the audience for attribution-ambiguous deaths exists — and that it's a culture you have to build, not just a mechanic you ship.
- The legibility bill, unpaid. DF's infamous interface doubles as an instruments failure: the dynamics are in-principle visible and in-practice illegible. DF survived it on depth; nothing else has.
2.6 S.T.A.L.K.E.R.'s A-Life (2007) — emergence as set dressing
Offline faction/mutant simulation populating the Zone, producing emergent encounters — around an exam (the shooting) that stays authored. This is the common industrial compromise: emergence relegated to texture while the graded loop stays fixed. Worth a node because it shows the marketing-vs-mechanics gap: "living world" sells, but if the sim doesn't grade you, it's the Wright divorce wearing a trench coat.
2.7 Left 4 Dead (2008) — the domestication (branch 4 contrast)
The AI Director is what the industry built instead of emergent exams: a concealed authored policy modulating spawns against a pacing curve. It solves staleness for a few dozen hours and is the right tool for a co-op shooter. But it's epistemic uncertainty (a secret rulebook — fails the attribution constraint), and it got read: intensity-meter manipulation strategies were community knowledge within a year. RimWorld inherited the pattern (2.11). The Director matters to this genealogy mainly as the domesticated alternative that absorbed all the industrial demand which might otherwise have funded real emergent exams.
2.8 Creeper World (2009) — the enemy as a field
The TD-family member of this map. It replaced the wave table with a fluid: the creeper is a visible mass that flows, pools, and presses, governed by a diffusion-like rule. Full visibility — you SEE the enemy's entire state, its depth, its fronts. This is "chaos with the lights on" achieved in the visibility dimension... and missed in the dynamics dimension: diffusion is monotone and tame, so the flow is predictable minutes ahead, the authored maps make each level a puzzle, and the wiki is per-map build orders. Lesson: visibility is necessary but nowhere near sufficient. A field enemy with tame dynamics is a slow answer key. The field idea itself, though — enemy as continuous mass rather than spawn-table rows — is one of the most portable organs in the whole map.
2.9 AI War: Fleet Command (2009) — the visible scalar
AI Progress: a number on screen that rises when you take territory, raising the exam's intensity. Player-scheduled escalation, fully visible, attribution-pristine. And it demonstrates the scalar shape perfectly: because the coupling is one number, managing the number is the strategy — AIP discipline is the core of AI War play, deliberately. Fine as design (it's branch 3 done well), but it shows what happens when player-coupling compresses to a scalar: the exam's responsiveness becomes a thermostat the player sets.
2.10 Kerbal Space Program (2011) — the integrability trap
KSP's physics are deliberately solvable: patched conics, two-body problems on rails, chosen for performance and predictability. The two-body problem is the textbook integrable system — closed-form solutions exist — and the community responded exactly as the math predicts: transfer-window calculators, delta-v maps, the exam collapsed into tooling. "Physics" buys nothing by itself. The line between a breakable and unbreakable physical exam is the integrability boundary — two-body vs three-body, single vs double pendulum. KSP is the shipped-game proof of the session's double-pendulum point, from the wrong side.
2.11 RimWorld (2013/2018) — the retreat
DF's most successful descendant, and on the exam side it's a retreat: the colony sim stays emergent, but the exam is re-domesticated into a storyteller (a director, branch 4) whose main input is a scalar — colony wealth. Raid points are a formula; the formula is on the wiki; and "wealth management" is now a whole strategy genre where players suppress the input variable itself — killboxes aside, the dominant meta is staying poor on paper. The exam collapsed into thermostat play, fully documented. Lesson: when the coupling between player state and exam intensity is a printed scalar, the meta-game becomes playing the thermostat, not the weather.
2.12 Cities: Skylines traffic (2015) — the mass-market emergent exam nobody names
Hiding inside a builder: congestion emerges from your road network × an agent simulation, fully visible (you can follow every car), perfectly attributed (it's your road), and never the same twice because the substrate (the city) is player-authored. The traffic wiki is a methods textbook — road hierarchy, junction theory, lane mathematics — that famously cannot solve your city; every intersection is a fresh diagnosis. By the imagined-wiki test, C:S traffic is one of the best-performing exams in mainstream gaming, and it's reached tens of millions of players. Lesson: emergent exams are already mass-market viable when the player authors the substrate — attribution stays pristine and the audience never even files it as difficulty.
2.13 Offworld Trading Company (2016) — the reflexive market
The market IS the exam: every price moves with every purchase, including yours. Perfect information, zero hidden state, and still irreducible to you, because exploiting a price moves the price — M.U.L.E.'s self-reference made the entire game. The caveat: OTC fills its market with AI/human opponents, so it's partially branch 5. But the mechanism is entity-free in principle: a market whose prices respond to the player is unpredictable through self-reference alone. This is the economist's version of the ecology.
2.14 Oxygen Not Included (2017) — instruments solved, exam self-authored
Two lessons in one game. First, the positive one: ONI's overlay system (oxygen, heat, germs, decor — one keypress each) is the gold standard of "chaos with the lights on" instrumentation. The sim is deep and the player can see all of it, layer by layer. This is solved technology; steal it. Second, the negative one: ONI's exam is your own base's thermodynamics — deterministic, no exogenous variance, authored entirely by your own construction. So optimal sub-builds transfer between runs, and the game wiki'd into blueprints (the SPOM and its descendants — literal answer keys you copy tile by tile). Lesson: a purely self-built sim collapses to blueprints. The exam needs dynamics the player didn't author.
2.15 Rain World (2017) — the purest ecology, the unpaid bill
The realest agent-ecology exam ever shipped: persistent creature populations per region, individual variation within species, predator-prey relations, simulation continuing off-screen. The exam is literally an ecosystem, and the design works — the wiki test comes back field guide (a bestiary of dispositions, not solutions), and the veterans' skill is precisely the learnable-but-unwritable read (watch a speedrunner negotiate lizards in real time: that's embodied distribution-reading of an agent system, the exact spec).
And the market graded it: launch reviews called it unfair and opaque, the mass audience read the ecology as noise, and the game became a cult classic among the minority willing to build the embodied model with zero help. Rain World hides almost nothing structurally, but it shows almost nothing either — no instruments, no climate readouts, high ambient hidden state. Lesson: the difference between "alive" and "noise" is not the simulation, it's the instrumentation. The legibility bill is real and the designer must pay it, or only players who pay it themselves stay.
2.16 Noita (2019) — the two-layer split
The falling-everything engine makes the world a material simulation, and the exam's sharpest moments are emergent (fire spreads, gases mix, liquids pour through the level you just melted). Noita's instructive property is the split: its build layer (wand mechanics) is fully explicable and has been wiki'd into one of the deepest answer-key corpora in roguelikes, while its physics exam stays wild — the wiki can teach you wand-building and still can't tell you what this cavern does when it catches fire. The game stays alive on the irreducible half after the explicable half is solved. (Blemish: per-seed alchemy recipes are hidden parameters — epistemic, the cheap way out, and players resent exactly that part.) Lesson, and it's a load-bearing one for the card-TD: a game can host an explicable build layer and an irreducible exam layer simultaneously, and the second protects the first from becoming an answer key.
2.17 Minor islands and counterfeits
- From Dust (2011) — terrain/fluid dynamics vs your tribe; the right shape, middling execution; mostly proves real-time geophysics exams are buildable.
- Timberborn (2021) — water and drought as visible dynamics over player-built dams; tame enough that it collapsed into dam blueprints, Creeper-style.
- Eco (2018) — multiplayer ecology with a meteor deadline; the UO problem answered with social governance (players must regulate their own extraction). Interesting, but the regulation is made of people — branch 5.
- Kenshi (2018), Crusader Kings — world/character simulation generating semi-emergent strategic exams (factions, schemes, claimants); both lean on entity simulation and both partially wiki'd, but CK's character-opinion machinery is a decent example of high-dimensional coupling resisting scalar compression.
- Counterfeits — chaos theming over fixed exams. Frostpunk's weather is an authored schedule wearing a storm costume; Vampire Survivors' chaos is a fixed timetable; They Are Billions is fixed waves plus noise. All three sell the emergent-exam feeling while shipping branch 1 or 2. Frostpunk is the instructive one: it succeeded commercially because the schedule is authored — dramatic pacing is exactly what authorship buys — which is reason #5 below in one sentence.
3. Why no lineage formed
Five compounding reasons the emergent branch is islands rather than a tree:
- The Wright divorce (1989). The simulation-building skill and the challenge-grading skill split into separate genres — toys vs games — and the people holding each half rarely meet. DF is the one full re-merge in 35 years, and it was built by an outsider over decades with no commercial constraints.
- Unbreakability binds the designer. You cannot balance what you cannot predict. The entire industrial QA methodology — playtest, observe, tune, patch — verifies propositions ("wave 12 is too hard"), and an emergent exam only supports climate claims ("late-game pressure trends high"). Distribution-design is a different epistemology (semantic-view QA: hold a model of your own game and grade it on verisimilitude, not correctness), and almost nobody is trained in it.
- The UO lesson. Players are an optimization pressure that finds any open loop's degenerate state, given time. Persistent designs hand them unbounded time. Most designers who watched UO's ecology die concluded "ecologies don't work" instead of "ecologies need containment."
- The legibility bill. Rain World shows the cost of shipping a real ecology without instruments: the mass audience experiences variance they can't attribute, which reads as unfairness. Paying the bill (ONI-grade overlays for an ecology) is expensive UI work with no marketing screenshot.
- Directors are cheaper. A concealed authored policy delivers 80% of the feeling of a living exam for 5% of the cost, plus full authorial control of drama (Frostpunk's success). The domesticated alternative absorbed the demand.
None of these is a mathematical barrier. They're all economics and craft tradition — consistent with the session's conclusion that the missing ingredient is "a designer willing to govern a system they've deliberately made too alive to fully know."
4. Collapse taxonomy inside the emergent branch
The islands that died each died in a characteristic way. Naming the failure modes makes them checkable at design time:
| Signature | Mechanism of death | Islands that died there |
|---|---|---|
| Thermostat | Player↔exam coupling compresses to a scalar; the meta becomes managing the number | RimWorld (wealth), Factorio (pollution/evolution formula), DF partially (wealth), AI War (deliberately) |
| Calculator | Dynamics are integrable/closed-form; community ships tools that solve them | KSP (patched conics), Creeper World partially (diffusion) |
| Blueprint | Exam authored entirely by the player's own deterministic construction; solutions transfer across runs | ONI (SPOM), Timberborn (dams) |
| Answer key | Emergent dynamics mounted on authored, finite maps; per-map solutions | Creeper World (campaign) |
| Starvation | Open loop + unbounded player extraction time | Ultima Online |
| Noise | Real ecology, no instruments; variance unattributable, audience leaves | Rain World (commercially) |
And the two shapes that survive the imagined-wiki test:
- The field guide (agent ecologies — DF, Rain World): wiki tops out at climate knowledge; trajectory skill stays in nervous systems.
- The ticker (reflexive markets — M.U.L.E., OTC): prices respond to the reader; self-reference keeps the future open even under perfect information.
Both survivors share the structural property identified in the Bayesian discussion: their distributions are generated, not authored — they exist nowhere as data, so there is nothing to publish, leak, or datamine. The hypothesis space never closes, so play never reduces to credence-updating within it.
5. Convergence: the emergent-exam mode for SNKRX / the card-TD
What the islands hand to an ecology mode for SNKRX (proposal A-repaired from the session) or to the card-TD's exam side, organ by organ.
5.1 Organs to import
- From ONI — the instrument layer, as first-class UI. Population graphs per species, trend arrows, a food-web view, between-wave climate readouts. Rule of thumb: show the climate completely, the trajectory never — the first because fairness demands it, the second because you can't (that's the point). Budget real UI time for this; it is the difference between Rain World's reception and a playable design.
- From DF / Rain World — the population substrate. A handful of species with growth, predation, and competition coefficients; the player's kills as harvest pressure feeding back in. Cascades (one population's collapse releasing another) are the emergent payoff — DF's tantrum spiral translated into wave composition.
- From Creeper World — the enemy as field. Waves as emissions of a visible standing mass rather than rows in a spawn table. The wave table stops existing as data; what spawns is a function of population state. This is the concrete implementation of "the distribution is generated, not authored."
- From UO, inverted — the run as containment vessel. UO's ecology died because a persistent world gives players unbounded extraction time. A 25-wave run bounds extraction by construction: the ecology only has to survive one run's worth of pressure, which means it can be aggressive, fragile, and interesting instead of armored against infinite farming. The roguelike format is not a compromise here — it is the enabling technology that makes ecologies shippable.
- From Noita — the two-layer split. Let the build layer (deck, units, tier lists) be learnable and even wiki-able; protect the game at the exam layer. The deck is the player's authored distribution; the ecology is the world's grown one; the run is the collision of the two. Tier lists can exist — against an emergent exam they degrade from answer keys into priors, which is exactly what the Arknights-collapse warning in the card-TD doc asked for.
- From M.U.L.E. / OTC — couple the shop to the ecology. Offers and prices as readouts of the same population state the waves are emissions of: culled a species hard, and units strong against it get cheaper/rarer accordingly (one dynamical system, two windows — exam and feed). This repairs the session's proposal D (market shop), whose standalone version was dismissed as thin: endogenous coupling was the missing part. It also makes every purchase reflexive — buying moves the system that prices the next purchase.
- From the TD chassis (previous map) — visible spatial commitment. Placement, lossy sell-back, the maze. Unchanged, and now it finally has the uncertainty its commitment was missing: "commitment without uncertainty is just construction" gets its other half.
5.2 Traps, as design rules
- No scalar coupling. If the player↔exam feedback can be displayed as one number, it will be played as a thermostat (RimWorld's wealth, Factorio's evolution). The coupling must flow through the population vector — five species' states can't be suppressed the way one threat meter can. Checkable test: if a wiki could print "keep X below Y," the coupling has failed.
- No integrable dynamics. Pure decay/diffusion/linear growth has closed forms — a calculator's home turf (KSP). The interaction terms (predation products, competition, thresholds) are what buy non-integrability. Lotka-Volterra-with-harvesting is already past the line; it doesn't take much.
- Exogenous variation, visible at start. A purely deterministic sim seeded identically collapses to blueprints (ONI). Roll initial populations/coefficients per run — and show the rolls (the starting climate screen). Visible-at-start randomization is aleatory-then-public: fairness-clean by the session's own standard, and it kills blueprint transfer.
- Pay the legibility bill (§5.1, instruments). Non-optional.
- Bound the blast radius. Emergent exams produce occasional doomed states; the "luck isn't real" audience grades on process and accepts this, but a cap on cascade intensity (e.g. a population explosion saturates at +X% pressure for N waves) keeps doomed from meaning degenerate and widens the audience past the DF cult. Distributional fairness plus a bounded tail.
- The Ngo rule, as an implementation discipline. No wave tables anywhere in the data. If a distribution appears as an authored asset — even an internal tuning table — it can be datamined, published, and turned into the enumerated hypothesis space that makes the game Bayesian-ready. Author coefficients and rules; let every distribution be a runtime consequence.
5.3 The composed architecture
Putting the two maps together: the card-TD already chose its build feed (player-authored deck — the sibling branch to the auto-battler's shop). This map supplies the exam: a visible population system whose waves are emissions and whose shop (if any secondary feed exists) reads from the same state. In one line: you author the answer; the world grows the question. The deck is a theory the player constructs before and during the run; the ecology guarantees the experiment never repeats; the spatial chassis makes every commitment legible and lossy. Against the imagined-wiki test, the goal state is: deck-building wiki = allowed to exist (priors, not answers); exam wiki = field guide that keeps needing new chapters.
5.4 Open dials
- Ecology horizon vs run length. A 25-wave run may be too short for slow population dynamics to be felt; options are faster generation cycles, mid-run climate events, or persisting population state across NG+ loops (the session's note).
- Species count. 3 reads instantly but may be too thin to cascade interestingly; 6+ is rich but strains the instrument layer. Likely sweet spot 4-5.
- Shop coupling strength. Full coupling (offers and prices both from ecology state) vs partial (prices only). Full is the most reflexive but risks feedback so strong the player feels steered.
- Starting-climate variance. How wide the per-run coefficient rolls go: wide = runs feel like different planets (less transferable skill, more reading), narrow = one climate with weather (more transferable gut, less novelty).
- Cascade cap. The blast-radius bound from §5.2.5 — where to set it is a feel call, found empirically by climate-tuning, which is the distribution-design craft this whole branch demands.
6. The one-paragraph summary
There is no genealogy of emergent exams — that absence is the finding. The fixed-exam family is a dense tree; the emergent branch is five decades of isolated islands (Hamurabi's dynamics, UO's eaten ecology, DF's re-merge, Creeper's fluid, Factorio's pollution, ONI's heat death, Rain World's ecosystem, OTC's reflexive market), each invented from scratch, each dying its own characteristic death — thermostat, calculator, blueprint, starvation, noise — or surviving as a field-guide game with a cult audience. The dead islands die when their dynamics compress (to a scalar, a closed form, or a blueprint); the survivors are the ones whose distributions are generated rather than authored, so there is nothing to enumerate and the wiki tops out at climate. Every organ an emergent-exam SNKRX mode or card-TD needs already exists on some island — ONI's instruments, DF's populations, Creeper's field-enemy, UO's inverted containment lesson, Noita's two-layer split, M.U.L.E.'s reflexive prices — and the reason they've never been assembled in one game is economic and cultural (the Wright divorce, director-shaped demand, the designer-binding problem), not mathematical.