Live
Assessment Desk
AI contributes a lens. Teachers retain the judgement.
Assessment Desk grew from a working command-line tool I used to organize and review written, handwritten and spoken language assignments. It connects student work to declared curriculum criteria and produces structured scores, rationales and feedback for a teacher to examine. A retrospective audit is now tracing that chain end to end—not to claim that AI grades correctly, but to ask whether its responses are warranted by the evidence, reasonable to a teacher and useful in practice. The same audit is showing how future research could isolate questions about task conditions, transcription repair, rater information and human–AI review.
The opportunity
Designed for Teachers, departments and schools
Assessment is an argument from a student's work to a judgement about attainment. Research on feedback, rater effects and construct validity shows why neither hurried human marking nor confident machine output should be accepted uncritically. Schools need ways to sustain attention across a cohort while keeping the criteria, evidence and reasoning open to inspection.
Working now
- ✓Live, fixture-backed, privacy-safe teacher web demo
- ✓Classes, rosters, assignments and unit plans
- ✓Curriculum browser — criteria, strands and phases
- ✓Submission-status pipeline and evidence views
- ✓Create-flows for new unit plans and assignments
Next on the roadmap
- 01Bring the CLI's audio transcription and handwriting OCR into the web workflow
- 02Test how reliably constrained AI analysis can identify language evidence, patterns and uncertainty against explicit criteria
- 03Declare the provenance and semantic role of rubric descriptors and local indicators throughout the schema and prompt
- 04Wire the structured analysis, rationale and source evidence into a teacher-reviewed web workflow
- 05Generate Excel workbooks, grade sheets and review-ready reports in-app
- 06Add multi-tenant storage, authentication and a real job queue
- 07Pilot the workflow with real teachers
Project in brief
“Assessment Desk began with a practical classroom problem. Written work, handwriting and recordings arrived in different forms, while the eventual judgement still had to remain connected to the assignment, its criteria and the evidence in the student's work.
The working command-line tool created an auditable chain: assignment definitions, source submissions, OCR or transcription where needed, model prompts and responses, criterion-level scores and feedback, grade records, and teacher-facing matrices. AI supplied a first analytical pass; it did not acquire professional authority.
Retrospective analysis has made the limits of that first experiment clearer. A recorded presentation cannot be interpreted without knowing whether it was rehearsed, timed or spontaneous. A cleaned transcript is not the raw recording. A polished example written for feedback is not the submission that received the score. These distinctions are part of assessment validity, not implementation detail.
The present evidence is therefore about inspectability and practical reasonableness. Can a teacher follow the chain from task to evidence to judgement? Does a score cite relevant work? Does the feedback preserve the student's meaning and offer something useful? Where the answer is uncertain, does the system make that uncertainty visible enough for the teacher to disagree?
This does not establish that AI grading is accurate, unbiased or better than human marking. Research on rater effects, construct-irrelevant variance and automation bias explains why an auditable second lens may be worth investigating; this classroom corpus was not designed to measure those phenomena.
It has, however, generated better research questions. Future work could compare identity-masked and identity-visible review, raw and conservatively repaired transcripts, rehearsed and time-bounded speaking, independent human ratings and AI-assisted decisions. Those would be controlled studies derived from the audit, not conclusions retrofitted onto it.
The public demo currently uses fictional data to show the teacher workspace. The next product work is to bring evidence, transformations, rationale, uncertainty and teacher adjudication into the review interface. AI contributes sustained attention and another perspective; the teacher retains the judgement.”