The ontoref ladder, the mechanism
Project governance, infrastructure and authoring: three measured cases, with real output.
Jesús Pérez
·
The same ladder, with the real output of ontoref ladder show, ladder audit, governs coverage and describe state. Each rung declares what checks it and what no check answers; one rung’s figures are never set beside another’s, and the audit closes without taking any rung as holding.
The event and the question
On September 10, 2026, IBTimes reported that a researcher was leaving Anthropic after three years researching model pretraining:
«Neither company is acting responsibly. They are racing straight to
self-improving superintelligence and gambling with our lives.»
self-improving superintelligence and gambling with our lives.»
Another researcher at the same company agreed:
«we do not yet have a plan to solve alignment for superintelligence
and are not clearly on track to.»
and are not clearly on track to.»
The public response argues about who brakes.
Whoever coordinates projects, maintains systems, or writes with these tools has a different question.
Whoever coordinates projects, maintains systems, or writes with these tools has a different question.
It’s not who governs, it’s whether it lets itself be governed
You can put the wisest person in the world in charge of something. If what they're running doesn't say what it is, doesn't let you see what it does, and doesn't allow itself to be checked, there is no possible government.
| Word | What it names | The question it answers |
|---|---|---|
| Governance | The arrangement: boundaries, membership, what gets published | How is it organized? |
| Government | The exercise: who acts and under what procedures | Who acts, and how? |
| Governability | The capacity: "whether a subject admits government at all" | Does it let itself be governed? |
Governability is a property of the protocol, not of any particular domain: it has to come before any arrangement and any exercise.
You can’t govern what isn’t described
In 1970, Roger C. Conant and W. Ross Ashby proved a theorem whose title says it all:
«Every Good Regulator of a System Must Be a Model of That System.»
According to the authors themselves, the theorem changes the status of model-making
«from optional to compulsory.»
That's why the first thing ontoref asks is that the project describe itself. One of its invariant principles:
«Any actor — human or agent — can verify that a slice … WITHOUT possessing or loading global knowledge of the whole.»
A governable system tells you where it isn’t
`ontoref governs <path>` answers which architecture decisions govern a file. `ontoref governs coverage` runs the reverse count over ontoref itself, today:
CONSTRAINT ROUTING COVERAGE
accepted ADRs 104
constraints 469
routable 442
declared non-path 21 ← say what they are; not a blind spot
STALE TYPED 0 ← declared a path, none of them resolves
UNTYPED 6 ← neither addressable nor declared. THIS is the blind spot.
of which ungated 6 ← the real blind spot: nothing addresses them AND nothing checks them
of which gated 0 ← a check verifies them; only the SCOPE names no path
routes derived 737 over 339 distinct prefixes
via scope 347
via cmd 112
via check 278What isn’t written, someone invents
You don't give a compiler architecture instructions in prose. With AI agents we do the opposite: we write them a good-intentions prose — be careful, don't make things up, respect the conventions. It may or may not be read, and when it is, it's interpreted.
And there are gaps. An agent doesn't stop to ask: it fills what's missing with something plausible, with the inertia of answering fast and pleasing.
An agent had the rule right in front of it and still filled a gap with an invented reference, and attributed to the author words the author never said.
"Moving the rule is not enough" — case file 2/26
Two decisions
Decision 1
Decisions are written as typed data, not as prose
Every architecture decision is a file that gets validated, with constraints marked as hard or soft and, when possible, an associated check.
Decision 2
A rule holds where it runs, not where it's read
Where does it run? Does it check meaning or just a spelling? Has anyone seen it reject what it forbids and accept what it allows? Without those three, it's an intention, not a mechanism.
Braking is not steering
Guardrails and harnesses limit, filter and cut off. They act in the negative. Conant and Ashby already named that form of regulating half a century ago:
«a primitive and demonstrably inferior method of regulation» — «its success can only be partial.»
Chris Argyris explained it in 1977 with a thermostat. Correcting the temperature is single-loop learning. Asking
«whether it should be set at 68 degrees»
is the second loop — the one that remembers why.
Seven rungs, seven questions
| Rung | The question it asks | |
|---|---|---|
| 1 | Reason for being | Is what the project says it is still true? |
| 2 | Reflection | Does what it knows about itself run, or is it only written? |
| 3 | Categories and relations | Is it in a closed vocabulary a machine can validate? |
| 4 | Territory and criterion | Is it declared how far what’s governed reaches? |
| 5 | Accreditation | Is the evidence verified, or only the report of whoever did it? |
| 6 | Observability | Does a green state the distance between declared and actual? |
| 7 | Knowledge lineage | Is it recorded where what the project knows comes from? |
Mechanism · ladder show (1/3)
1 · Reason for being
anchored in self-describing · voluntary-adoption
answered by describe project · about
binds with self-describing → `ontoref positioning signals self-describing` · voluntary-adoption → `ontoref positioning signals voluntary-adoption`
2 · Reflection
anchored in reflection-axis-act · reflection-modes
answered by describe capabilities · run <mode-id> · mode list
binds with reflection-axis-act → `ontoref positioning signals reflection-axis-act` · reflection-modes → `ontoref positioning signals reflection-modes`
3 · Categories and relations
anchored in ontology-axis-substance · dag-formalized
answered by validate ontology · describe term --list · nickel export
binds with ontology-axis-substance → `ontoref positioning signals ontology-axis-substance` · dag-formalized → `ontoref positioning signals dag-formalized`Mechanism · ladder show (2/3)
4 · Territory and criterion
anchored in internal-coherence-enforced · governed-delivery · plane-habitability
answered by constraint · sync audit · describe constraints
binds with internal-coherence-enforced → `ontoref positioning signals internal-coherence-enforced` · governed-delivery → `ontoref positioning signals governed-delivery` · plane-habitability → `ontoref positioning signals plane-habitability`
5 · Accreditation
anchored in witness-as-axis-seam · sufficient-verification
answered by sync substrate · positioning witness-ondaod <signal-id>
binds with witness-as-axis-seam → `ontoref positioning signals witness-as-axis-seam` · sufficient-verification → `ontoref positioning signals sufficient-verification`
6 · Observability
anchored in drift-observation · audit-runner
answered by describe diff · sync audit · health --full · adr amendments
binds with drift-observation → `ontoref positioning signals drift-observation` · audit-runner → `ontoref positioning signals audit-runner`Mechanism · ladder show (3/3)
7 · Knowledge lineage
anchored in knowledge-lineage
answered by positioning sources · positioning signals <node> · validate sources
binds with knowledge-lineage → `ontoref positioning sources`
7 rungs · 14 checks DECLARED, none run here — `ladder rung <id>` shows them, an audit would run them
No context, so nothing is weighed or compared: each rung shows what its
material binds with. The figures live in `ladder rung <id>`, one rung at a
time; a context (a discipline view, a mode run) decides what weighs on its
plane.Auditing one rung
[Pass ] ladder reason-for-being:conditions
2 of 2 necessary condition(s) met — necessary, never sufficient: this is
a verdict about the conditions, not about the rung
+ read: FileExists .ontoref/ontology/core.ncl [exists]
+ read: FileExists .ontoref/ontology/state.ncl [exists]
? not read: what settles the rung: Whether what the file answers is
still the project's actual reason for being, or a sentence that stopped
being true and nobody re-read. No check can compare a declaration
against an intention; it can only confirm the declaration is reachable.No rung is reported as HOLDING. Every verdict above is about a rung's
necessary conditions; what would settle the rung itself is the line
marked `? not read` under it, and no check answers that one.The anti-pattern, named
ADR-111 names, with an id, the exact mistake this deck refuses to make:
id ·
«Rendering the audit as «N of 7 rungs passing», a percentage, or a badge — anything that presents the conditions verdict as a verdict about the rungs.»
ladder-reported-as-a-score«Rendering the audit as «N of 7 rungs passing», a percentage, or a badge — anything that presents the conditions verdict as a verdict about the rungs.»
Today all fourteen declared conditions hold: eleven check that a file exists and three look for a pattern. That's a pass over conditions — never a rung taken as holding.
Three domains where it plays out
This isn't theory: three cases that already happened and were measured, one per domain. In all three, the same thing happens — something looked fine, and it wasn't.
Project governance
whoever declares the rules is also inside them
Infrastructure
a silence
Authoring
a citation that looks right
Case 1 · Project governance
On August 24, 2026, it was measured that ontoref's own repository type appeared in no domain. The tool that hands out rules to other projects was the only project with no domain of its own.
That measurement is what the project governance domain came out of. There's no exception for whoever writes the rules.
That domain declares seven procedures. One of them, governed delivery, had run twice in six weeks while another ran twenty-six times — and both were equally present on disk.
Case 1 · what got written down
A check that only looks at what's declared doesn't distinguish a dormant procedure from an abandoned one. Under the Reflection rung, every review of the ladder prints the measurement again:
? not read: what settles the rung: Whether the declared modes are actually
RUN. A mode nobody invokes is inert rather than still, and the distinction
is invisible to a check over declarations: measured 2026-08-14,
governed-delivery ran twice in six weeks while generate-article ran
twenty-six times, and both were equally present on disk.Case 2 · Infrastructure: a silence
In an infrastructure project, on June 29, 2026, the component that connects the storage disks looked for the data in one directory, and the system that runs the services stored it in another. The case's record sums it up in one line:
«No component reported a failure at any point — the failure mode IS the silence.»
What was declared said one thing and the territory did another. Nobody was lying. Nobody was looking.
Case 2 · the check states its own reach
The check that came out of the incident declares its own limit:
«A workspace where BOTH are wrong in the same way passes here and fails live.»
? not read: what settles the rung: Whether a criterion that exists can
actually LOOK at its subject. The corpus catalogues gates inert since a
layout change, a scan that aborted and never ran once, and a
`must_be_empty` reporting SATISFIED over a path that did not resolve —
every one of them declared, and every one of them green.Case 3 · Authoring: a citation that looks right
In an authoring project, the unit is the work as a coherent whole. The domain declares three content gates: glossary coherence, that no section is left without an exercise, and that each step picks back up the previous concept. None checks where a piece of data an agent inserts comes from.
A citation that doesn't resolve to anything passes review precisely because it looks correct.
The domain recorded the question without treating it as resolved: no gate today rejects a citation an agent inserted that doesn't lead to a preserved or accessible source.
Mechanism · ladder rung knowledge-lineage
## Anchored in
knowledge-lineage — declares the intake catalog as its artefact:
32 judged source(s), 181 claim(s)
Convergent 82
Adjacent 36
Canon 24
Evidence 16
Challenge 15
Rival 6
Carrier 2
Corroboration does not apply: these claims are this rung's content,
not evidence about it. Custody (quotes, captures, digests) is checked
by `ontoref validate sources`.What no check answers, here
«Nothing mechanical distinguishes a thorough reading from a confirming one.»
The corpus records one reading of its own that was nearly closed carrying six claims, all found by grepping for anchors named in advance; a second pass found three of the strongest claims in the record.
A territory that changes
Knowing the territory isn't making a map once. Conant and Ashby foresaw this too: when the system changes, the regulator has to change with it, and regulating something that varies over time needs «a time-varying model». Argyris found it in people: few know that they don't use the theories they claim to follow, and end up «prisoners of their own theories».
FSM Dimensions 4 total
protocol-maturity ✓ reached
Protocol Maturity horizon: Months
current: protocol-stable desired: protocol-stable
self-description-coverage ✓ reached
Self-Description Coverage horizon: Weeks
current: fully-self-described desired: fully-self-described
ecosystem-integration ✓ reached
Ecosystem Integration horizon: Months
current: multi-project desired: multi-project
operational-mode → in progress
Operational Mode horizon: Continuous
current: local desired: daemonA stance, not a mechanism
What follows has no command. We're facing new paradigms that call for different roles and different safety mechanics. But the roles we see in the debate aren't new: the world's savior, the one who doesn't want to get their hands dirty, the smartest and fastest one.
The risk isn't borne only by a few, and the benefit shouldn't belong only to a few either.
If checking is left in the hands of whoever holds authority, most people can only trust or distrust. If what's used is governable, anyone can check their part.
Do we know what the achievements are?
What's known. Each rung's necessary conditions are known, and today they hold. The question no machine answers is known, because it's written down.
What isn't claimed. No rung is taken as holding, because verifying is always local and partial. Nor is there yet a way to order the rungs by the plane of each domain.
What can already be pointed to. In all three cases there's something written down that wasn't there before: a tool that stopped being outside its own rules, a silence turned into a check that says what it doesn't see, and a citation with no provenance turned into a question with a closing criterion.
Was this useful? Rate it
Got something to add?
Tell me what you think, what you'd suggest, or whether we should keep exploring this topic.
·
reads