Thirty exam scripts
I have just finished marking thirty scripts from our first-year Synoptic Exam. Synoptic in the proper sense: students are not asked to reproduce memorised content, but to integrate material across the major topics of the course and construct a coherent argument from it. It is a genuinely difficult task, and the format reflects that – an open exam, with a week to research and write.
The use of AI was not prohibited. It was, in fact, encouraged. That will change next year, when submissions will move on campus, but for this cohort the expectation was transparent: AI is available, use it thoughtfully, and demonstrate that you have actually engaged with the material. How well that worked varied considerably across the thirty submissions – which is, as it turns out, precisely the point.
The rubric problem
Before I get to what was interesting about these scripts, I want to say something about marking rubrics, because they were supposed to feature in this exercise and I find them deeply unsatisfying as a tool.
The idea is straightforward: divide the assessment criteria into categories, describe what each grade band looks like for each category, and then map the submission onto the grid. It sounds clean and defensible. On paper. In practice, it rarely is. Real student submissions do not fall neatly into cells. A submission might show sophisticated engagement with primary literature in one section and almost none in another. The referencing might be strong in approach but inconsistent in execution in a way that matches neither the "good" nor the "very good" description particularly well. You find yourself choosing the least wrong cell, and then writing a paragraph of actual feedback explaining what the cell does not capture – which raises the reasonable question of what the rubric was for in the first place.
The deeper problem is that rubrics can actively mislead. If the grid description for a given band does not quite fit the submission, highlighting it gives the student feedback that is technically incorrect. The written comments then have to do the work of correcting the impression left by the grid. It is, at best, an inefficient system.
What was actually interesting
Here is what I did not expect: thirty scripts, substantial AI use evident throughout, and not one of them was the same.
The expected story – AI homogenises student output, submissions converge on the same polished but hollow response – did not materialise. What I found instead was a genuinely varied set of submissions that sat on different points of a multidimensional space rather than a simple good-to-bad line. Some were polished but shallow. Some were messy but contained real insight. Some were broad but unfocused; others mechanistically precise but narrow. Many were genuinely synoptic in the way the assignment intended – making connections across levels of biology that suggested the student had done more than assemble information.
What this tells me is that the students put enough of themselves into these essays that the AI amplified rather than replaced their thinking. The direction each submission took – the aspects it emphasised, the connections it drew or failed to draw, the level at which it engaged with mechanism – reflected choices that were, at least in part, the student's own. That is not nothing. And I think the least we can do, as academics, to honour this effort is giving students specific and individual feedback. Not in the format of a generic rubric with highlighted pre-phrased cells that only half apply. With the intention of helping the students in future assignments. I know some of my colleagues think this is overkill, but I have generated more than 12,000 words of feedback for the scripts I have received. We know that quite a few students never take a look. They are only interested in the result. Maybe for those the time spent is a waste. However, it is, in my opinion, compensated by the students who do care, as these will take the feedback and use it to improve themselves in the years to come – if it is specific.
But the individuality of the responses also tells me something about the assignment design. An open synoptic question, precisely because it cannot be answered by reproducing a memorised pathway, exposes how a student actually thinks. AI can supply information; it cannot supply the intellectual architecture that connects it. Where that architecture was present in these submissions, it was apparent. Where it was absent, that was apparent too – and in a way that a more constrained question might never have revealed.
A qualified conclusion
None of this is an argument for keeping the current format. If a significant portion of the factual content is being supplied by AI rather than retrieved and understood by the student, that is a problem the assessment needs to address. And moving submissions on campus next year will do exactly that.
But I want to record, while it is fresh, that marking this particular cohort was genuinely engaging in a way that marking is not always. Every script made me think. None of them could be evaluated with a generic response. Each one had individual strengths and weaknesses that required actual attention to identify and describe. That is, perhaps, the best thing you can say about a set of exam submissions. None of them fitted the given rubric cleanly – and it happened despite, or possibly partly because of, the use of AI.
Similar and related blog posts: