Working example · Assessment Desk

Keep the judgement attached to the work.

A grade nobody can trace, and nobody can compare.

Assessment Desk is a marking and feedback workspace for departments running a detailed criterion-based curriculum on the hours a timetable actually leaves. It keeps each strand-level judgement attached to the student work that earned it, so the sub-scores, the reasoning behind them and the report at the end of term all come from the same place, and so two teachers marking the same strand can be held to the same descriptor.

Marking happened in whatever was left after teaching, which came down to seconds on each assignment. Barely enough to read the work properly, and nowhere near enough to write something a student could act on.

Open the working demo ↗Book a conversation ↗

Live · fictional teacher workspace, no sign-in

The chain, named

Eight transfers between the work and the report. Every one of them by hand.

None of these systems is bad at its job. The problem is that no two of them share a record, so a person carries everything across the gaps.

01Designedplanner02Submittedplatform03Markedpro forma04Averagedplatform05Recordedspreadsheet06Aggregatedformula07Writtenspreadsheet08Re-keyedschool systemWHAT THE RECORD CAN STILL ANSWER FORSTRAND DETAIL STOPS TRAVELLINGTHE RULE IS IN A PRIVATE CELLNOBODY HERE HAS READ THE WORKEvery transfer is done by hand, and the last one asks for words the record can no longer support.

The strand-level judgement, the thing the whole specification exists to produce, exists in one place for one step. After that it is an average, then a cell, then a sentence somebody types from memory.

StationWhat stops being answerable

  1. 01The assignment is designed, with criteria, strands and band descriptors attachedNothing yet. This is the most complete the record will ever be
  2. 02Students submit into the learning platform, which tracks deadlines, late delivery and plagiarismThe rubric, which does not exist in this system at all
  3. 03The teacher marks, recording strand scores on a printed pro forma or a sheet of their own devisingNothing yet, and this is the only moment strand-level judgement exists anywhere
  4. 04The average is typed back into the platform, one student at a time, with a commentThe strand scores, as far as the platform is concerned. Whether they survive at all depends on what the teacher decided to keep
  5. 05The numbers go into the class spreadsheet, under hand-typed column labels grouped beneath the assignment, whose details are held elsewhere. Some teachers keep every strand in its own column. Some keep only the averageAny guarantee. What the record holds is a private decision, and nothing downstream knows which one was made
  6. 06At term end the teacher sums and averages selected columnsWhich scores counted, and whether practice was separated from judgement, both decided by a formula nobody else can see
  7. 07The teacher writes parent-facing comments in the same spreadsheetThe evidence. What is left to write from is memory and whatever the grid displays
  8. 08Another teacher re-keys the averages and the comments into the school system parents and students seeAny route back to the work. Nobody at this station has read it

I kept every strand in its own column, grouped by assignment, because an average on its own answers almost nothing later. I could not tell you whether that was common practice. Once I went further and produced a PDF of the full strand sheet for each student, uploaded into their private folder, which is a ridiculous amount of work multiplied by every student in every class. It still does not survive into any later analysis. A PDF in a folder is not a record you can ask questions of.

What can be settled

Nobody can tell you what the score should be. That is not the failure.

The page is not arguing for better judgement. It is arguing for a judgement that can still be examined afterwards.

Ask what a grammar score rests on and the only honest answer is to open the work beside it and read again. Careful markers will still disagree about how many errors, and of what kind, are needed before a student’s meaning is genuinely compromised. Even writing without a single error can be misread by a particular reader. None of that can be settled from the outside, and this does not try to.

Two things can be settled. Whether a judgement was made with the work in front of it. And whether a score contradicts the descriptor it claims.

That second one is where lifting a whole class does its damage. If the top band asks for writing with no errors that compromise communication, and the strongest student in the class made errors that plainly do, awarding that band is not a difference of opinion about severity. It is a statement the work does not support.

What it costs

None of this appears in a budget. It is paid in evenings, in meetings, and in grades that cannot be defended.

Four costs, in the order a department tends to notice them.

01

Seconds on each assignment

Marking happens in whatever the timetable leaves, which is not a scheduled thing. Thirty students in a class, several classes, and the time available for each piece of work falls to seconds: enough to arrive at a number, not enough to read closely or write anything a student could act on.

The same spreadsheet is also where attendance gets re-keyed from the paper register after a lesson, which tells you what kind of instrument it really is.

02

A single challenge costs a term

When a parent questions a grade, the department re-reads a term of work to answer, because the reasoning behind the original judgement was never recorded anywhere. A meeting follows, and the question quietly shifts from whether the work meets the descriptors to whether this student belongs with the strongest in the class or the year.

That is construct-irrelevant variance, which is to say a score moving for reasons that have nothing to do with what is being measured, and it is the third time it appears in this chain.

03

The same strand, two meanings

A shared rubric exists so that a 5 in one class means what a 5 means in the next one, and in the school down the road, at the same phase and age. Nothing in the chain enforces that.

Each teacher marks against the descriptors as they read them, records the result privately, and the differences never surface, because there is nothing to compare except the numbers that survived.

04

The rule lives in a cell

What counts toward a term result is a spreadsheet formula selecting certain columns, written by whoever built the sheet, according to a convention they invented. Whether practice was separated from judgement is decided there too.

It is not written down, not reviewed, and not visible to anyone else in the department, including the person answerable for the results.

Each of these is good at its job. None of them holds the judgement.

This is not an argument that school software is bad. It is an argument about where the gap falls, so a department can decide what to do about it deliberately rather than by habit.

Curriculum and unit planners

Hold the specification properly. Criteria, strands, phases and band descriptors, all selectable, all attached to a plan. Their work is finished when the unit is written, which is before any student work exists.

Learning platforms

Collect reliably. Who has submitted, who is late, what needs checking for plagiarism, where the file is. They are built around the deadline rather than the descriptor, so the rubric has nowhere to live in them.

Spreadsheets

Will record anything, which is why they end up recording everything. The price is that the meaning of a column sits in a header row, the rule that aggregates it sits in a cell, and neither one travels with the number.

School reporting systems

Present what survived, as averaged grades and typed comments, to parents and upward. They are the only part of the chain most people outside the department ever see, and they receive the record at its thinnest.

None of these needs replacing. The gap is between them, and it is the only place the judgement ever existed. The newest option, a model that reads the work and returns a grade, falls into the same gap from the other side: quick, and unable to show you what it was looking at. The question to ask of any of them is the one a parent will eventually ask. When somebody wants to know what a score rests on, what can you open?

What this does

It does not decide. It says where to look.

Image analysis in medicine does not diagnose cancer. It marks the regions a specialist should look at closely, and the specialist decides.

That is the job here. The model reads against the descriptors the school already publishes, finds the candidate instances, counts them, and hands the marker a pointer to the exact sentence. What the work is worth stays with the teacher.

01

It reads against your descriptors, not its own

The criteria, strands and phase are the ones the class was assigned. Every claim it makes is expressed in the language of a descriptor a teacher already has to apply, which is what makes disagreeing with it possible.

02

Every claim points at the work

Not an impression of the essay, but this sentence, and these four others like it. A count of instances, each one a link. The evidence sits beside the score rather than somewhere behind it.

03

It reads the twentieth essay the way it read the first

Reading at speed is where polished writing starts to look like clear thinking, and no marker is immune to it. A bound reader has no such drift. It may still be wrong, and a teacher may well decide the instances it found do not justify the band it suggested. That disagreement is cheap, because it is about a specific sentence rather than an overall impression.

04

What was judged survives

The strand scores and the evidence behind them stay attached to the work through the term, which means the aggregate is not reconstructed by hand, the formula is not private, and the report at the end is written from what was actually observed rather than from memory of a student across three months. An assignment that touches more than one criterion is just an assignment, rather than a new recording sheet and a new block of hand-labelled columns.

This is a partial demo on fictional data, not a platform and not a study. It does not show the model is accurate or unbiased, and a proper trial would have to test exactly that, with the teacher making the final call throughout. What it does show is the shape: when somebody asks what a score rests on, you open the work beside it, and the conversation is about the evidence rather than about who the student sits near in the class.

The version that ran on real marking

The scores came back lower than I expected, and the criticisms held up.

It was a command-line tool. A model made a first pass against the descriptors, and Python built the workbook: a sheet per assignment, the full term’s strand detail, and a summary with room for the teacher’s own words, so a whole class could be read at a glance rather than reconstructed. It also produced printable feedback sheets, batch printed for a class. None of it asked the school to adopt anything. It produced the artefacts the school already used, with the strand detail still attached, and it was written on my own initiative because nothing provided came close.

I ran it across a hundred and sixteen students and three levels. It broke down the conventional errors, and it flagged essays where polished writing was not backed by the argument the top criteria ask for. When I went back to those essays, the specific criticisms held up. Some of what I had rewarded was good writing rather than good thinking.

That is the charge usually made against language models, and here it was a human marker doing it, under exactly the time pressure described above.

None of this shows the model is accurate or unbiased. It was real marking, but it was one teacher and one subject, which is not the same as a trial. A retrospective audit showed what a proper study would need, starting with recording the live questions after a presentation, which the teacher heard and the model never did. What it shows is that this is worth studying, with a teacher making the final call.

What has happened since is not a change of approach. It is a friendlier interface, a test suite and the maintainability needed to hand it to someone else, and to cover more than one curriculum and more than one subject.

Experience behind it

Not one job title. Five decades of arriving at this from different directions.

Music teaching in the UK to private students for a decade and more, small groups in theory, lectures to adult audiences, and life-skills training for vulnerable adults in supported housing, where part of the work was finding ways to capture what someone had actually learned so it could be reported later to people who were never in the room.

Eight years in Japan since then: ESL, PYP music and drama, MYP language acquisition, across several schools and as many pedagogical approaches as there are books making the case for one.

And the other side of it throughout. Seminars in music, law and engineering, lectures in English literature, engineering, law and music, and a long run of courses taken in person, online and through media, each promising learning and a certificate at the end on the strength of some test of knowledge. Schooling at Christ Church Cathedral School in Oxford and at Oakham.

None of this is a complaint about teachers, or about schools, or about the people who designed the curriculum. Everyone in the chain is doing the sensible thing with the time and the tools in front of them. The record at the end still cannot answer for itself.

Someone will reasonably ask what I really know, and I would rather hear from a person with a stronger set of letters after their name. I am not going to defend that.

What I will say, across 1977 to 2026 with no year absent of some challenge involving education or learning, institutional or otherwise, is that we are still a long way from being honest with ourselves about what these systems deliver, how many operational problems remain unsolved, and how invested we can be in keeping things as they are.

1977 to 2026

No year without it

As a student, a teacher, and a person sent on a course with a test at the end.

Teacher

Music, ESL, PYP, MYP

Private students, small groups, adults, and eight years across schools in Japan.

Supported housing

Capturing what was learned

Recording real progress for stakeholders who were never in the room.

Student

Seminar, lecture, certificate

Music, law, engineering, English literature, and the online course that ends in a test.

The demo

Look at the workspace before we talk.

There is a teacher workspace open now, on fictional data with no sign-in. It is a partial demo rather than a platform, and it says on each screen what is working and what is only a direction.

Open the demo ↗

What you can see in it

  • Classes, rosters, assignments and unit plans
  • A curriculum browser showing criteria, strands and phases
  • The submission pipeline and evidence views
  • Create-flows for a new unit plan and a new assignment
The same evidence problem, another sector — Impact Desk ↗

Marking a term of work in the hours between lessons?

I would like to hear which part of your chain costs the most, and what happens in your department when a parent asks what a grade rests on. One conversation about one workflow, not a pitch, and no follow-up sequence.

If a change to the spreadsheet would fix it, I will say so. The point is a department that can stand behind its grades, not software for its own sake.

Book a conversation ↗