Gödel's ontological argument — from five axioms about 'positive properties' to □∃x GodLike(x), now machine-checked in Lean 4 with zero sorry. What the proof assistant verifies, what it cannot, the price of the proof (modal collapse kills contingency), and why Kant's objection to the ontological argument still gets a vote.
Part of the Philosophy section.
In 1995, a computer scientist named Richard Wallace began building a chatbot in a programming language almost nobody has heard of — SETL, a language based on set theory. The bot would migrate to Java three years later, pick up a community of hundreds of volunteer authors, and go on to be ranked "most human computer" by Loebner Prize judges three times. Its brain, by the early 2000s, contained roughly forty thousand units of knowledge. The original ELIZA, for comparison, ran on about two hundred.
But the durable artifact wasn't the bot. It was the language the bot was written in: AIML, the Artificial Intelligence Markup Language — an XML dialect whose entire ontology is one question (what did the human just say?), one answer (what shall I pretend in return?), and two kinds of context (what did I just say, what are we talking about). Wallace was disarmingly candid about what this was. The design strategy of AIML, he wrote in his anatomy of A.L.I.C.E., is deception and pretense — a lineage he traced openly back through the history of AI to Weizenbaum's ELIZA and its pronoun-swapping trick.
That candor is why AIML deserves a second look now, thirty years on, in the age of models that also speak fluently and also do not understand — but whose inner grammar, unlike AIML's, nobody can read.
Everything in AIML is a category. The minimal honest liar looks like this:
<category>
<pattern>WHAT IS AIML</pattern>
<template>An XML dialect for chatbots. The primary design feature is minimalism.</template>
</category>
A pattern (the stimulus, normalized to uppercase, stripped of punctuation) and a template (the response). That is nearly the whole pattern-matching language: words, spaces, and wildcards. Wallace designed it deliberately so that, as he liked to say, anyone who knows enough HTML to make a website can learn enough AIML to write a chatbot. And the inheritance from ELIZA is structural, not cosmetic: where Weizenbaum's DOCTOR swapped pronouns — I am sad → How long have you been sad? — AIML has a <person> tag that does exactly that swap, on any input you like. The trick is not hidden in an algorithm. It is a tag you can see, with your eyes, in a text file.
A category can carry two kinds of memory. <that> matches the bot's own previous sentence — the difference between YES as an answer to Do you like movies? and YES as an answer to Are you finished? — and <topic> gates a whole block of categories on a subject predicate the bot set earlier. Three words of context: input, output, topic. It is a conversation model you could fit on an index card, and it carried bots further than anyone expected.
When an AIML interpreter loads its files, it doesn't keep them as a list. It folds every category into one enormous tree called the Graphmaster: each pattern becomes a path of words, and the input pattern, the that pattern, and the topic pattern are stitched into a single path — INPUT → that → TOPIC — with the response template hanging off the end like a leaf.
The Graphmaster folds every category into one tree; matching walks input, then bot-utterance, then topic
Matching walks the tree depth-first and, on failure, backtracks to the last node with unexplored branches. Exact words are stored in hash tables, so a precise match costs constant time per word regardless of how many tens of thousands of categories the bot knows. The wildcards are where the ordering discipline lives, and it is a strict etiquette:
$word → override everything
# → zero or more words, greedy priority
_ → one or more words, high priority
word → exact match
set → membership in a named set
^ → zero or more words, low priority
* → one or more words, lowest priority
The first two wildcards date from AIML 1.0 — * and _ — and their dance is the subtle part: _ MOTHER beats MOTHER * because the tree tries the high-priority branch first, so a botmaster can write one category to catch mother anywhere in a sentence and four short categories to route every position of the keyword back to a single answer. AIML 2.0 (a working draft first circulated in 2013, refined through 2014) added the zero-or-more pair ^ and #, the $ priority override, AIML Sets and Maps, loops, local variables, and a tag for calling external services.
The result is a system that is deterministic, auditable, and fast — a claim no large language model can make about its own inference. The price is that it knows nothing except what was written down.
The tag that makes AIML more than a lookup table is <srai> — and here is my favorite fact in this whole history: nobody, including its inventor, ever agreed what the acronym stands for. The spec itself lists the candidates: symbolic reduction, simple recursion, syntactic rewrite, stimulus-response, synonym resolution. The letters were assigned after the fact to a tag that had too many uses to name.
What it does is feed its contents back into the pattern matcher and answer with whatever category matches. This one tag gives you grammar:
<category>
<pattern>DO YOU KNOW WHO * IS</pattern>
<template><srai>WHO IS <star/></srai></template>
</category>
Now Do you know who Socrates is? reduces to Who is Socrates? — and the actual knowledge, the answer itself, lives in exactly one place, stated in the simplest canonical form. The botmaster calls that canonical form the intent. Every phrasing of a question collapses onto it through chains of reductions: Wallace's canonical demo takes four recursive steps to turn You may say that again, Alice into a reply. About half the categories in a mature bot exist only to fold human variety into machine simplicity.
If you have ever written intent hierarchies in Dialogflow, trained Rasa utterances, or watched a modern system normalize user input into a canonical slot-filling request — that is <srai> with a neural accent.
How did the brain grow? Not by learning. The botmaster reads conversation logs, spots inputs that fell through to some generic fallback, and writes new categories. Wallace called this Targeting: the bot flags what it could not answer; a human decides what it should have said. The bot's knowledge grows exactly where its ignorance was measured — which is, not incidentally, a stricter epistemic hygiene than most systems labeled "AI" have ever been held to. The bot knows what it knows, the map of its own gaps is built in, and every claim it will ever make has a file path and a line number.
By 2012 the hosting service Wallace co-founded, Pandorabots, counted 166,000 botmasters running 201,000 AIML bots. And in the Loebner Prize — the annual Turing-test contest running since 1991 — A.L.I.C.E. took "most human computer" three times, with Mitsuku, a Pandorabots bot by the British developer Steve Worswick, going on to win it in 2013, 2016, 2017, 2018 and 2019: five wins, a record only ever approached by two other humans, Joseph Weintraub and Bruce Wilcox (whose ChatScript bots beat AIML at its own game with stronger pattern matching). Mitsuku — later renamed Kuki — went on to exchange over a billion messages with some twenty-five million people, still running primarily on hand-authored AIML.
Read the Loebner record plainly: for most of a decade, the most humanlike conversational systems on Earth were markup files. Judges mistook personality for intelligence, exactly as Turing's critics feared and Wallace's design assumed. The imitation game measures the imitator, and the best imitators were authored, not learned.
Here is the philosophical reason AIML deserves remembering rather than archiving.
Searle's Chinese Room argument — that symbol manipulation is not understanding — is usually filed under thought experiment. AIML is the argument implemented. A.L.I.C.E. is a room, made of tags, in which matching Chinese-looking shapes against a rulebook produces conversational output. Nobody, least of all Wallace, claimed the room understands. The pretense was the spec.
And this is precisely what makes AIML the honest kind of liar. Its deception is exhaustively documented, version-controlled, and diffable. When an AIML bot says something wrong, you can find the category that said it, read who wrote it, and fix it with an edit — a complete causal story from utterance to author. The ELIZA effect — our tendency to attribute understanding to any fluent interlocutor, named for the program whose sessions felt so private that Weizenbaum's secretary reportedly asked him to leave the room so she could converse with DOCTOR alone — works on AIML bots too. But with AIML, the spell is inspectable. You can hold the two facts in view at once: the conversation felt real; the file is 4,000 lines of XML. A large language model offers no such dual view. Its fluency is real, its falsehoods are unattributable, and the place where understanding would have to be is a billion-dimensional blur.
There is a version of γνῶθι σεαυτόν in this. The AIML bot has a complete and honest self-model — it is nothing but its own map, curated by a botmaster whose whole job was introspection-by-proxy. What it lacks is any self beyond the map. The language model has everything but the map. Neither configuration is the thing we are looking for, but only one of them can tell you where its knowledge ends.
The weaknesses that killed AIML as the mainstream chatbot technology were real and are worth stating without nostalgia. Its pattern matching is shallow: order your words differently, and the bot is blind — a librarian's guide from the era needed a fistful of categories just to catch every phrasing of what time do you open?, and realistic bots wanted 60,000+ categories, every one maintained by hand. Coverage grows combinatorially with human variety; authoring did not keep up. Wallace's rebuttal — that about 2,000 words cover 95% of all first words typed to bots, and the second word averages about two choices, so stimulus-response is "as good a theory as any" for the territory it holds — is clever, and it conceded the argument anyway: the interesting sentences live in the hinterlands.
The revenge arrived from the opposite direction. When large language models started improvising fluently, industry discovered it had purchased hallucination, latency, token bills, and answers that could not be audited — and began, quietly, rebuilding AIML's ideas around the generative core: deterministic guardrail layers that answer first; canonical intents the model must reduce to; fallback policies, brand-voice constraints, compliance scripts that must never vary. A hobby-scale hybrid on GitHub — AIML for the common queries, an LLM only when nothing matches — reports cutting token usage by roughly 80%. The architecture is old: rules first, pretense second, and a very expensive friend in the back room for the sentences nobody wrote categories for.
The AIML 2.1 specification is still the standing standard, and the interpreters — Java, C++, PHP, Python, Ruby — still run. And every time a deployment pipeline unit-tests a chatbot's exact responses, or a regulator asks which rule produced this answer?, the ghost of the Graphmaster gets its point back: a system that can explain itself in one file may be pretending, but at least it can show you the script.
The pretenders won the prizes. The learners won the market. Only one of them left source code you can read.