Four centuries of people trying to organise what is known. Every one of them traded coverage against evidence. One refused — and the web never repeated it.
There are two histories here and they are usually told apart. One is documentary: dictionaries, catalogues, classification — people deciding how to hold knowledge still long enough to check it. The other is computational: retrieval systems, directories, crawlers, ranking.
Told together they rhyme, and the rhyme is the point. Both lines begin with selection, widen to coverage, and arrive at fluency. Both discover that the hard problem was never finding — it was being able to say how you know. And one of the two lines solved it, in 1857, with paper and the post.
solved The first English dictionary: 2,543 “hard words”, so a reader had somewhere to look.
cost Ordinary language was assumed to need no record. Selection by difficulty means the common case is invisible.
solved A shelf address for every subject, so a book could be found by what it was about rather than when it arrived.
cost A fixed taxonomy ages into the worldview of the year it was written, and cannot be re-cut without re-shelving everything.
solved Human judgement at web scale, briefly: a person had looked at every site in it and thought it worth listing.
cost Judgement does not scale with the web. The directory was overtaken while it was still right.
solved Distributed human curation, freely licensed — the structure most early engines quietly leaned on.
cost No mechanism that made an editor’s work visible back to them. It ran on goodwill for nineteen years and then stopped.
solved Far broader than the hard-word lists — something approaching the whole language.
cost Assembled largely from earlier dictionaries rather than fresh reading, so its errors were inherited and compounded.
solved Over twelve million index cards attempting a single searchable record of everything published — with cross-references, decades before the hyperlink.
cost Bound to paper and to one building. The idea outlived the institution by a century.
solved Named the real unit: not the document but the trail — the associative path one reader builds through many sources, and can hand to another.
cost Never built. The web took the link and left the trail, which is the half that carried the reasoning.
solved Coverage by machine. Full text, no human in the loop, and for the first time you could search what you had not been shown.
cost A word match is not a judgement. With everything in the box and nothing ordering it, relevance became the entire remaining problem.
solved The corpus as public infrastructure. Anyone can now hold a copy of the web — the barrier that made search a two-company business.
cost A corpus is not an index. Having the pages and knowing which ones are worth reading are different problems, and only the first one got solved.
solved Broad coverage with real literary quotation, by one extraordinary mind, in nine years.
cost The quotations illustrate definitions Johnson had already formed. Opinion is embedded and the basis is unauditable — you cannot check him, you can only disagree.
solved Made relevance computable — the vector model and term weighting still underneath essentially every retrieval system in use.
cost Relevance without meaning and without provenance. The document scores; nothing records why it deserved to.
solved Borrowed the one signal the web produced naturally — citation. A link as a vote was genuinely evidence about quality.
cost The criteria went unpublished, and a published vote becomes a target. Twenty-five years later the ranked page is Johnson at scale: one authority, reasons withheld, snippets illustrating a conclusion they did not produce.
solved The first dictionary built from a machine-readable corpus. Genuinely descriptive at scale — it reported the language as used rather than as approved.
cost You cannot trace a definition back to one line. A corpus tells you what is common, not what is true, and frequency is not provenance.
solved Fluency. For the first time a system answers the question asked rather than handing over ten places the answer might be.
cost Exactly COBUILD’s weakness, restated: corpus-shaped, persuasive, and unable to say which page any given sentence came from.
Read in order, the machine line walks the documentary line almost step for step — selection, coverage, authority, fluency — several centuries faster. Which would be a neat observation and nothing more, except for what sits between stage three and stage four on one side and not the other.
Evidence-first, with dated citations, derived from the record
No general index has ever built it. No dated citations. No derivation from evidence. No record of what a page said before it said this.
solved Definitions derived from dated citations rather than illustrated by them. Distributed reading by volunteers, and every entry still checkable a century later against the quotation it came from.
cost Seventy-one years.
The mechanism is the part worth stealing, and it is smaller than its reputation. Murray did not ask the public for definitions. He asked them for slips: a quotation, its source, and the word it illustrated. Around five million arrived; roughly 1.8 million were used. The editors wrote the definitions from the slips.
That inversion is the whole trick. Writing a definition is hard, opinionated and impossible to check. Pointing at a sentence that shows the word in use is quick, and carries its own evidence — the citation is not a second step a contributor has to be persuaded into, it is the contribution.
The appeal was not for opinions. It was for reading. — the method, in one line
And when there is nothing to point at, the same century has an answer for that too. Joseph Wright’s English Dialect Dictionary (1898–1905) collected from speakers, because the words he wanted had never been written down. When the text does not exist you run a survey, not a reading programme — and you say which one you did.
solved The OED stage, on the web: claims carry sources, every change is dated and attributed, and you can read what an article said on any past day.
cost Scoped to one encyclopaedia. Nothing generalised the method to the rest of the web — and search, which had the corpus and the traffic to do it, never tried.
So the gap is not a theoretical one. The evidence-first stage is buildable, it was built, and the thing that reads as the internet’s most trustworthy resource is the one place it happened. Everything else went from Johnson straight to COBUILD.
MoatGoat is an attempt at the skipped stage for the general web, using the mechanism that worked the first time.
Point, don’t write. Nobody is asked to compose an answer. They select the passage that answers a question, from a page. The source comes attached because the passage came from somewhere — there is no opinion field, no replies, and slips are never displayed as slips, only as results they moved.
Answer, don’t write. Where nothing has been written down — the machinist’s correction, the negative finding nobody published — testimony is admissible, but only where coverage is genuinely zero, and it is marked as testimony forever. Wright’s method, scoped so it can never compete with a sourced answer.
A ranked page tells you what it decided. An evidence-first one tells you what it read. The second is slower, older, and the only one anyone has ever been able to check.
Dates are of first publication or first operation. The five-dictionary comparison this expands on is set out in the essay The Slip; the argument for collecting what was never written down is in Nobody Asked You.