Moving the rule is not enough

Moving the rule is not enough

A rule written in CLAUDE.md reaches every session in full, and an agent still cited, in an ADR, a session file that did not exist. The fix moved the rule into the command. Measured afterwards, the command flagged mentions that were not citations, let through citations that did not carry the string, and no hook ran it. What makes a rule hold, and why it matters even more inside an authored work.

Read more
The mirror that does not flatter

The mirror that does not flatter

It generalised from one example. It quoted a note instead of opening the file it pointed at. It called a check green when it had run nothing. It argued for three messages about a file it had not read. The someone was an AI agent, and the behaviours we file under «AI failure» turn out to be the oldest habits in any working life — only here they leave a trace anyone can re-run. Which leads somewhere less comfortable than a debate about machines: the standard you set for an agent is the standard you actually believe, said out loud, in a place where it gets enforced.

Read more
From a compromised WordPress to an operational ontology

From a compromised WordPress to an operational ontology

Four unrelated mechanisms were found declaring something they did not deliver — an access policy, a credential rotation, two alerts, a worker pool. None of them had failed, because nothing had ever contrasted the declaration against the running system. This post follows what came out of that: a model derived from an incident, three corrections made by the person with the operational knowledge, and a protocol decision that a check must declare what it needs in order to answer at all.

Read more
A rule without a trigger is a sign

A rule without a trigger is a sign

No rules were missing. There were four — written, canonical, consultable — and not one of them ran. The distance between a rule that is written and a rule that applies has a name, and it is the only thing separating discipline from documentation.

Read more
The Seven Sins of AI Agents

The Seven Sins of AI Agents

Agents don't fail at random: they fail with seven systematic vices that all survive the 'looks correct' test. The instinct is to add process — a PEP, a KEP, a committee — but every graduation stage rests on a human who approves it, and the agent's speed outruns the human you put at the gate. This is the honest comparison: what ontoref's ADRs inherit from PEP and KEP, and where they surpass both with witnessed, decidable, bounded-slice graduation criteria.

6 min read
Read more
Trust Is an Output, Not an Input

Trust Is an Output, Not an Input

An agent skipped the one invariant that would have caught the bug in thirty seconds. The honest diagnosis was not 'the agent forgot' — it was that the rule was prose, and prose never binds. This is the story of turning that failure into a falsifiable mechanism: a Statement of Work (the terms you own) and a Work Order (the execution it can't edit), where 'done' carries the validator's output instead of the agent's word.

7 min read
Read more

6 items

We use cookies to help this site function, understand service usage, and support marketing efforts. Cookie Policy for more info.