Four million words, and the part that isn't the model
Everyone building an AI product is using the same models. The knowledge base is the whole company.
Veridian marks a full Pearson International A-Level Economics or Business paper: grade, UMS, assessment-objective breakdown, and a confidence flag on every judgement it makes. Two schools and seventy independent students pay for it.
The model underneath is the same one you can use. That is not a weakness in the product — it is the point I want to make about the product, and about most AI products.
The demo everyone can build
Paste an exam question and a student's answer into any frontier model and ask it to mark against the Edexcel mark scheme. You will get something that looks correct.
It will be confident, well-written, and reference the assessment objectives by name. It will award a grade. A parent looking at it would be satisfied.
It is also, in my testing, wrong often enough to be worse than useless — because a marking tool that is wrong twenty per cent of the time and confident one hundred per cent of the time teaches students to write incorrect answers with conviction. That is not a neutral failure. It actively damages the thing it claims to help.
The gap between "looks like marking" and "is marking" is the entire business.
What was actually in the four million words
Two hundred plus documents. Forty to sixty-five of them per paper.
That is the number people react to, and the reaction is usually that it sounds like a lot of scraping. It was not scraping. Almost none of it existed in a form I could take.
What is in there:
- The specification, decomposed. Not the PDF — the PDF broken into every individually assessable claim, each tagged with its assessment objective and the command words that trigger it.
- Mark schemes, with the reasoning reconstructed. A published mark scheme tells you what earns a mark. It does not tell you why the near-miss answer did not. That reconstruction — for hundreds of questions — is a large share of the corpus, and it is the part that makes the difference between a grade-B and grade-A answer visible to the system.
- Examiner reports, mined. Chief examiner reports are the most underused documents in A-Level education. They say, in plain English, exactly what candidates did wrong that year. They are a direct signal about the marking distribution and almost nobody reads them.
- Level-boundary exemplars. Real answers at each band, annotated with the specific feature that put them there.
- Command-word behaviour. Evaluate, assess, analyse, examine have distinct mark profiles. A student who writes a brilliant analysis on an evaluate question loses marks for a reason that has nothing to do with economics.
None of that comes out of a model's weights. All of it is public, and all of it required someone to sit down and structure it, and that someone was me, for months.
Why the boring part is the defensible part
The marking harness — the prompting, the retrieval, the confidence scoring, the UMS conversion — took weeks. It is real engineering and I am not dismissing it. It is also the part a competent developer could rebuild in a month.
The knowledge base cannot be rebuilt in a month. Not because it is secret, but because it is tedious, and tedium at that volume is a genuine barrier. Somebody would have to want it enough to spend a year of evenings on Economics mark schemes.
I wanted it enough because I was sitting the exams. That is the only reason it exists.
Confidence flags, and why they came late
The first version did not flag anything. It marked, and it sounded certain.
I removed that after watching myself trust an incorrect mark because it was delivered in the same tone as the correct ones. If I could not tell, a student certainly could not.
Now every judgement carries a confidence signal, and the low-confidence ones say so. This makes the product look worse in a demo — a marking tool that admits doubt is less impressive than one that does not.
It makes it dramatically more useful, because a student can see which parts of the feedback to argue with. And arguing with feedback is the thing that teaches. A tool that is always certain removes the argument, and the argument was the lesson.
What it cost me to learn this
Version one through three were me improving the prompting and getting steadily disappointed. Better instructions, better structure, better examples in-context. Marginal gains each time, and each time the same failure mode: it was excellent on the questions I had thought about and unreliable on the ones I had not.
The unlock was realising that was not a prompting problem. There was no sentence I could write that would make the model know what a specific examiner wanted on a specific twelve-marker in a specific series. That information exists in documents, and it had to be in the system rather than in the instructions.
Everything after that was corpus work.
The general form
I think this is true well beyond exam marking. If your AI product's advantage is the prompt, you do not have an advantage — someone will write a better one in a fortnight, and the next model release may make the whole prompt unnecessary.
If the advantage is a body of structured knowledge that did not exist in that form until you built it, the model getting better makes your product better rather than making it redundant.
Four million words is a description of where the year went. Read it as a boast about scale and you have missed what it cost.
More writing
Four thousand contacts and no idea who mattered
I had 4,000 connections from one platform alone and was quietly letting the twenty relationships that mattered decay. Every tool I tried was built for sales pipelines. So I built the other thing.
One brutal thing a year, done in the wrong order
Once a year I do something with a genuine chance of failure. The current version is an ultra Ironman in reverse order, and the reversal is not a gimmick — it's the entire mechanism.