1. The Objection Worth Taking Seriously
Among the recurring doubts about ANSELM, one deserves the most respect. It appears in some form whenever practitioners discuss the method seriously, and in its strongest form it runs like this: a knowledge-first method cannot be done without a rigorous ontology to describe systems, and some modeling framework to structure the system description. Those who have raised knowledge graphs from requirement databases note that they had a metamodel underneath — and that the result supported review of the requirements, not architecting.
The doubt deserves respect rather than dismissal. The need for rigor has never been the problem; the instruments we built for it were. The answer we will defend is therefore not "ANSELM needs no ontology." It is: ANSELM needs an ontology that earns its place — derived, not authored; small, not cathedral; verified, not decorated. And it lives at the seams of the knowledge ecosystem, not at its center. Where exactly, and why, is what this article sets out to explain.
2. Two Decades of Not Needing Ontology — and Why the Calculus Was Right
Many seasoned architects have spent entire careers without a formal ontological model, and the honest reading of that fact is not that they were behind the times. It is that, in the pre-AI era, the economics of ontology were structurally broken.
Start with a praxeological rule we will use throughout this article: a formal element earns its keep only through the acts that consume it. A query answered, a check passed, an integration performed, a contradiction caught. If no act consumes a piece of formalism, that piece is decoration, regardless of how elegant it is.
Now audit the pre-AI ledger against that rule. On the cost side: ontology was authored by hand — the architect compiling rich, tacit understanding into triples and class hierarchies, which is the very burden the manifesto denounces. It was decoupled — the ontology lived in a tool or a repository separate from the artifacts people actually worked in, so it rotted the moment the work moved on. And it was expensive to maintain — every vocabulary shift, every new stakeholder term, demanded a curator's manual intervention.
On the benefit side, the consumers were almost nonexistent. The reasoning ontologies enabled — subsumption, DL queries, class-consistency checks — had no practical downstream. The consumers that did exist — integration contracts, compliance evidence, shared meaning between teams — were almost always served by something far cheaper: a controlled vocabulary, a schema, naming conventions, a good glossary. That is why decades of enterprise practice rarely needed the formal apparatus. The architects were not failing to appreciate ontology. They were correctly pricing it: high authoring cost, thin consumer base, faster decay than adoption.
3. What AI Actually Changed — and What It Did Not
The arrival of capable LLMs inverts the cost structure, but only on one side, and the distinction matters enormously.
Two costs collapsed. Extraction: an LLM lifts candidate entities, relations, and candidate types out of prose at near-zero marginal cost — the most expensive step of classical ontology engineering became a proposal, generated in seconds. Shallow reasoning over prose: consistency checking, contradiction detection, gap analysis over natural language are now real capabilities, which removes the historical raison d'être of description logic — you no longer need a reasoner to tell you that two statements conflict when a competent reader of the corpus can do it.
What did not change is the demand side. Where stakes are high — a contractual boundary, a regulatory obligation, an audit trail — stochastic judgment is not acceptable. An LLM that is right 99% of the time is not a compliance instrument. This is why the ANSELM experiments run a deterministic constraint checker as oracle: the model drafts, the oracle judges. The division of labor is permanent: stochastic reasoning, deterministic verification.
And one demand-side change cut the other way. Because AI makes it cheap to spin up dozens of concurrent reasoning contexts — chats, cells, agents, teams — the number of seams across which meaning must travel has exploded. Every seam is a lossy channel. And a lossy channel's capacity is set, in the end, by how much formal vocabulary the two sides share.
4. Interfaces in Formalism, Reasoning in Language
The companion experiment [3] measured the hand-off tax: fragment a task across role-bound agents and the data-processing inequality grinds every lossy interface. The lesson drawn there was fragment only when forced. But there is a second, complementary lesson hiding in the same experiment: when a hand-off cannot be eliminated, the cheapest way to reduce its cost is to raise the channel capacity — and channel capacity is shared schema.
Some hand-offs cannot be eliminated. A regulator will not share your conversation. A partner company, a team onboarded in year three, an acquisition, a court of law — each is a seam you cannot dissolve by keeping the conversation whole. At those seams, natural language alone is a slot machine: two parties read the same paragraph and derive different commitments. The formal model is what lets two contexts, two institutions, or two eras exchange meaning with bounded loss.
This yields the architectural principle that reconciles the manifesto with the objection:
Interfaces in formalism, reasoning in language.
The core of ANSELM — the workshop where synthesis, critique, and trade-off reasoning happen — stays conversational, accumulating, unsegmented. The formal ontological layer lives on the boundaries: between contexts, between people, between the ecosystem and its oracles, between the present and the future archive. Formalism is the connective tissue, not the skeleton.
5. The Ontology That Earns Its Place
The ontology ANSELM needs has three properties, each a direct consequence of the praxeological rule.
Derived, not authored. Candidate structure is extracted from the knowledge cells by AI and committed by humans — the Knowledge Steward approves, never compiles. The day the method asks an architect to hand-translate prose into triples, ANSELM has relapsed into the disease it diagnosed.
Committed, not decorated. Every formal element — every relation type, every schema field, every vocabulary term — must name its consumer before it exists. The consumer test is brutal and should be: "Which query, checker, or integration point consumes this?" No consumer, no formalism.
Tiered, not monolithic. Commitment should be graded by the stakes at the seam:
| Tier | Trigger | Contents | Typical consumers |
|---|---|---|---|
| 0 — Glossary | Always, from day one | Controlled terms, cell frontmatter types (need, function, component, decision, constraint, interface), semantic tags | Retrieval, generation, consistency prompts |
| 1 — Typed graph | More than one context | A small, closed set of relation types (satisfies, refines, conflicts, constrains, derives-from, decides) over the knowledge cells | Impact analysis, traceability views, contradiction routing |
| 2 — Oracles | Regulatory or contractual stakes | Deterministic validators over the typed graph: schema checks, constraint checkers, provenance | Compliance evidence, audit, acceptance gates |
| 3 — Full ontology | Rare, explicitly justified | Upper ontologies, DL reasoning, automated deduction | Only when an inference is actually consumed downstream |
Tier 0 costs almost nothing and pays immediately — the LLM is dramatically more coherent when cells carry types and terms are stable. Tier 1 is where the semantic graph of the enterprise ecosystem [2] stops being a pretty picture and becomes a queryable network. Tier 2 is the deterministic floor under everything stochastic. Tier 3 is where ontology projects went to die in the pre-AI era; ANSELM should treat it as the exception that must prove its consumer, not the default aspiration.
6. The CESAM Question
The same discussions that raise the ontology objection often propose a pragmatic bridge: a deliberately minimalist modeling framework — for example CESAM, with its operational / functional / constructional layers and its small catalog of views. Does ANSELM accommodate such frameworks?
Yes — as lenses, not containers. ANSELM's disposability of views means a view catalog is a library of templates the AI can render on demand from the living knowledge, not a repository to fill. The layers of CESAM are one legitimate projection of the knowledge graph — a useful one, precisely because it is small and its views have audiences. Nothing in the method forbids projecting the graph through CESAM's twelve views when a stakeholder community thinks in those views.
What ANSELM refuses is the inversion: the metamodel as entry ticket. Structure imposed before understanding is the original sin of the pre-AI cathedral. The doubt that opened this article had it exactly right: the successful cases had a metamodel underneath. The honest reply is that ANSELM's metamodel is Tier 0–1: extracted from the ecosystem, curated by stewards, versioned like code. It arrives as the steady state of the method, not as its precondition. You can have your first conversation without it. You cannot run a two-company, two-year, regulated program without it — and you will have it by then, because it grew out of the work rather than being painted over it.
7. Failure Modes on Both Sides
Explaining where the ontology lives is only half the argument. The other half is naming where this goes wrong — on both sides of the line. These are the failure modes a knowledge-first method must avoid.
The cathedral relapse. Ontology as gatekeeping: "real systems engineering has a metamodel." The moment the formal layer becomes a requirement of membership rather than an instrument of work, it is consuming attention instead of being consumed. Watch for formalism whose consumers are reviewers rather than queries.
The graph theater. An automatically extracted graph looks formal — nodes, edges, labels — while carrying no semantics at all. Labels are noisy, extraction is non-deterministic run to run, and two extractions of the same corpus will disagree. The remedy is structural: the normative layer is curated and versioned like code; extraction is always a proposal, never the source of truth. A graph nobody committed to is a rumor, not a model.
Vocabulary drift. The single most common cause of death of pre-AI ontologies was slow, silent semantic drift between the model and the work. ANSELM's answer is the same as for code: closed vocabularies at the seams, versioned; open prose in the core, unversioned. The formal layer is small enough to review and change deliberately.
The chat silo. The mirror image of ontology worship is "just chat" — which is document-centrism wearing new clothes. A conversation silo is still a silo; the enterprise ecosystem article [2] made this point, and it applies to the individual. The knowledge-first position is not formalism-free; it is formalism-in-its-place. Treating the absence of a metamodel as a virtue would be exactly the complacency this article warns against from the other direction.
8. Monday Morning
The translation into practice is deliberately small. In a knowledge-first workspace, three artifacts carry the formal layer:
- Typed frontmatter on every knowledge cell — a
typefrom the closed Tier 0 list, plus status, author, and semantic tags. Zero friction; immediate retrieval and coherence gains. - A closed relation vocabulary — ten or so relation types, defined once in a versioned file, used by the AI when linking cells. No free-form edges in the committed graph.
- Deterministic oracles — schema validation for every cell, and constraint checkers wherever the stakes demand, run mechanically and never delegated to the LLM's judgment.
Found the contract, not the model. These three artifacts are the day-zero agreement layer: small enough to be co-created in an afternoon, and small enough to be thrown away if the domain surprises you. They are a contract — controlled terms, cell types, commitment types, and the oracle that checks them. Everything descriptive — the cells themselves, the relations between them, the semantics of the domain — waits to be derived from the conversation. Sanitizing the knowledge ecosystem is then a continuous byproduct of use: the oracle and the steward clean as they go, and the ecosystem is never asked to be clean before thinking begins. If a founding element cannot be thrown away in an afternoon, it is too big.
None of this requires a modeling suite, a reasoner, or a consultant. It requires a schema file, a checker, and the discipline to ask, before every new formal element: who consumes this?
9. What the Experiment Showed
The claims in this article are testable, so we tested them. On the same harness that produced the measurements of [3], we ran two isolated reasoning contexts — one for propulsion, one for energy and thermal management — forced to integrate through a single boundary artifact, under six deterministic constraint families checked by an oracle. Two conditions, identical except for the seam: TYPED, where the boundary artifact is typed cells gated at every crossing by the deterministic tier01 oracle, and PROSE, where the artifact is a free-form document translated into the typed format by an integrator at the end. Same pinned model, temperature zero, n = 5 runs per condition, full call logs kept for every run.
The instrument came first. The gate at the typed seam checks structure only — schema, vocabulary, references — deliberately. The six constraint families (mass, interface, safety, terms, assumptions, structure) run once, at the finish, on both conditions alike. If the gate also graded the content, the measurement would be teaching to the test. The gate is the treatment; the final checker is the measurement.
9.1 The first measurement falsified the naive metric
Phase 1 (k = 1 round-trip, gpt-4o-mini) produced the numbers a marketing department would have hidden:
| Condition | Violations (mean) | Cells surviving in the final deliverable | Structural violations |
|---|---|---|---|
| PROSE | 8.0 | 1–2 per run | 20 |
| TYPED | 9.8 | 9 of 9 in 5/5 runs | 0 |
By the naive metric — total violations — prose won. It lost by the metric that matters. A subsystem dropped in translation costs one "nothing declared" violation; a subsystem that survives but is incomplete costs one violation per missing field. The violation count was rewarding the seam that lost the design. The lesson is a measuring rule: at a seam, count preservation first, defects second. From that point on, the experiment used the pre-registered corrected metric — valid cells that crossed the seam — with exact permutation tests instead of eyeballed averages.
9.2 Loss at a prose seam is a step, not a slope
Scaling from one round-trip to three answered the obvious question: does prose degrade gradually? It does not. It collapses at the first translation — to one or two cells — and stays at the floor. What accumulates with more round-trips is vocabulary drift: each rewritten pack introduces uncontrolled terms, and term violations doubled (5 → 10). The typed seam, gated at every crossing, held 8–9 of 9 cells and zero structural violations at both k. The gate's price is paid per crossing — rejection rounds scale with k — but the loss it bounds does not.
9.3 Two model families, two ways to fail — one way to hold
To check whether the result was a quirk of one model family, we re-ran k = 3 on a second family (gpt-4o and DeepSeek). The typed seam was unchanged: 9 valid cells in every run, zero structural violations, both models. The prose seam failed in different ways and in the same way:
| Model family | Prose failure mode | Valid cells across the seam |
|---|---|---|
| GPT family | Collapse — the integrator drops the design (1–2 raw cells) | 0 |
| DeepSeek | Overproduction — 7–16 raw cells, none valid (invented types, missing fields) | 0 |
A raw-cell count would have falsified the hypothesis on DeepSeek — the integrator over-produced. The distinction between raw and valid cells, pre-registered in the formal amendment, is the load-bearing definition; both readings are on record. Whichever way the prose seam fails, nothing structurally valid arrives on the other side.
9.4 The confirmatory phase
The corrected metric and the hypotheses (preservation, defects-per-cell, structural validity) were fixed in a formal amendment [6] before any fresh data was collected. The confirmatory run — DeepSeek, k = 3, n = 5 — delivered complete separation: the typed seam produced 9 valid cells in each of 5 runs; the prose seam produced 0 in each of 5, from 7–16 raw cells per run. Exact permutation p = 0.0079 on preservation, with defects-per-cell undefined for prose (there is no valid content to defect against) and finite for typed. Structural validity: 0 violations in every typed run, 88 for prose.
The honest boundaries: n = 5, one task family, one seam topology, an oracle that measures declarations rather than truth, and an analysis that was not blind. Within those boundaries, the claim is no longer an argument — it is a measured regularity: the typed seam preserves the design across round-trips, brief sizes, and model families; the prose seam, in these runs, carried no valid content at all. Replication on a second task and a second seam topology is the declared next step, not an afterthought.
9.5 What this means in practice
Three working rules follow from the measurement, each one ordinary enough for a Monday morning:
- Measure preservation at the seam, not defect counts. A seam that loses the design and reports few defects is the worst of both worlds.
- Keep the gate structural. Agents optimize for what the gate enforces. A seam gate should hold shape; completeness is the conversation's job, and the conversation is where ANSELM does its thinking.
- Treat the formal layer like the experiment treated its checker: small, versioned, and verified. The oracle did not make the designs — it made the loss visible.
10. Closing
Anselm of Canterbury is remembered for the ontological argument — an attempt to prove existence through pure reason alone. There is a quiet joke in naming a systems engineering method after him and then debating whether it needs an ontology. But the discipline the method needs is the precise opposite of the saint's: not formal being deduced a priori, but formal being earned in action. The test of every ontological commitment in ANSELM is not whether it is true, but whether it is used.
The objection that started this article was right to demand an ontology. The answer is not to refuse it, and not to surrender to it. It is to put it where it works — at the seams — and to make it small enough to keep honest.
Complexity seeking clarity — and clarity paying its way.
References
- Gruber, T. R. (1993). A Translation Approach to Portable Ontology Specifications. Knowledge Acquisition, 5(2).
- ANSELM (2026). The Enterprise Knowledge Ecosystem.
- ANSELM (2026). Not a Committee, a Conversation.
- CESAMES. CESAM — Center of Excellence on Systems Architecture, Management, Economy & Strategy framework.
- W3C (2012). OWL 2 Web Ontology Language — for the formal apparatus Tier 3 would invoke, and a reminder of its weight.
- ANSELM (2026). Cross-Context Coherence Experiment — pre-registered design, formal amendment (metric v2), and run archives backing §9. https://github.com/anselm-systems-engineering/handoff-tax-experiment/blob/main/tier01/CROSS-CONTEXT-EXPERIMENT.md