It was a command-line tool. A model made a first pass against the descriptors, and Python built the workbook: a sheet per assignment, the full term’s strand detail, and a summary with room for the teacher’s own words, so a whole class could be read at a glance rather than reconstructed. It also produced printable feedback sheets, batch printed for a class. None of it asked the school to adopt anything. It produced the artefacts the school already used, with the strand detail still attached, and it was written on my own initiative because nothing provided came close.
I ran it across a hundred and sixteen students and three levels. It broke down the conventional errors, and it flagged essays where polished writing was not backed by the argument the top criteria ask for. When I went back to those essays, the specific criticisms held up. Some of what I had rewarded was good writing rather than good thinking.
That is the charge usually made against language models, and here it was a human marker doing it, under exactly the time pressure described above.
None of this shows the model is accurate or unbiased. It was real marking, but it was one teacher and one subject, which is not the same as a trial. A retrospective audit showed what a proper study would need, starting with recording the live questions after a presentation, which the teacher heard and the model never did. What it shows is that this is worth studying, with a teacher making the final call.
What has happened since is not a change of approach. It is a friendlier interface, a test suite and the maintainability needed to hand it to someone else, and to cover more than one curriculum and more than one subject.