Convincingly wrong¶
Security and reliability tend to get filed together, and they are different bills. Security asks who can make the system do something harmful. Reliability asks what happens when the system, unprompted and unattacked, confidently does the wrong thing. The board paper budgeted for the first question, at least nominally; the second is where the quieter costs live, because the failure modes changed underneath the procurement.
Dr. Flannel has forty residents and one Kevin, and a reader that produces a handover summary in four seconds is the closest thing to a second Kevin the budget will ever allow, so the pilot had been used there most. The consultant spent an afternoon there, most of it watching what the summaries were made from.
A document system fails visibly. File missing, search returns nothing, permission denied: hard failures, annoying and self-announcing, each one generating its own ticket. The reader fails softly. A synthesis that is plausible and wrong. A retrieval that is incomplete without indicating incompleteness. A ranking that surfaces the persuasive document over the authoritative one, Fabulist’s product sheet with the padlock icon over the DPA that does not mention cloud inference. An outdated source selected over its successor because the successor was named worse: the draccus’s diet came, for three weeks, from a sheet superseded by a file called “FINAL diet ACTUAL v3”. A contradiction resolved silently in the wrong direction, of which the Home has one per department. None of these announces itself; each arrives formatted exactly like success. The failure mode of the whole class is not that the system did not work but that it worked convincingly in the wrong direction.
Degradation inherits the same silence. A broken search system becomes obvious within the hour. A degrading reader can remain useful enough that nobody files anything: retrieval quality slips a little, old documents gain unearned weight as the index ages, a newly connected repository is ingested under the wrong assumptions, which is what happened when the Burrow was added in month four and every rota from 2021 became current again for a week. Users develop workarounds and stop reporting whatever the workarounds route around. Nikolaj’s night shift has a Signal group for that. The two members of staff who keep the retired spreadsheet have the spreadsheet. The system never fails; it becomes slightly less trustworthy every month, and slight monthly losses of trustworthiness are exactly the signal that operational monitoring, tuned to outages, does not carry. By the time the degradation shows in outcomes, it has been compounding in answers for a long time, and the answers have been circulating.
The loop is slow, and so far it has turned one way. As answer quality slips, workarounds grow; as workarounds grow, the reports that would trigger repair thin; with fewer reports, less corrective effort; with less effort, quality slips further. Every edge but the last has a person on it, and none of the people is doing anything wrong. The grey branch is where the answers go meanwhile, into the decisions taken on them, from which nothing evidenced returns.
Testing meets an unbounded surface. A document management system could be tested against enumerable cases: can this user open that file, does this query return that record. A reader’s input space is every question in natural language, times every phrasing, times every ambiguity, times every combination of documents the retrieval might assemble, times every workflow Kevin has invented since the last review. Exhaustion is not on offer, and the discipline that replaces it in software, representative suites maintained against regression, has no equivalent in most knowledge programmes, because the deployment was rarely classified as software. The missing discipline would be regression testing for language behaviour: any change anywhere in the chain, model, prompt, index or estate, can move answers. Only a maintained set of questions with known good answers will show it.
Reliability debt is the closest thing this domain has to technical debt, and it compounds. A programme can launch fast on broad ingestion, permissive access, weak metadata, no ownership model and minimal evaluation, and the launch will look excellent, because a system that has read everything answers everything, and breadth photographs well in a pilot. The Head of Fundraising’s slide was the photograph. The mess is being learned before it shows in the answers. Retrieval habits form around the unowned folders; summaries absorb the contradictions nobody resolved; users build workflows on answers whose sources nobody curated. Later, every improvement is harder than it would have been at the start: tightening access changes answers people rely on, cleaning the estate changes them again, and introducing evaluation reveals a baseline nobody wants to publish. A database without constraints becomes difficult to repair after years of growth, because everything since has grown around the absence. A knowledge system without discipline follows the same curve, with one difference: the database’s corruption stayed in the database, and the knowledge system’s has been advising people the whole time.
Old friction |
Removed by agent |
New friction |
|---|---|---|
loud failure |
fluent soft failure |
drift detection and evaluation |
a passive store |
an active knowledge layer |
reliability engineering |