How hard can it be? – Part 4: Three machines and a question they could not answer
A recreational four-part summer series about trying to recover the settings of real wartime Enigma messages, eighty years after Bletchley Park had already read them.
I should be honest about what I was actually trying to do for the past three posts.
I was not attempting to recreate Bletchley Park's methods. I had no interest in methodological purity, and I would have cheerfully skipped every step of it. The goal was much cruder: get the plaintext off this photograph, by whatever means work. If the statistical approach would have succeeded straight away I would have been perfectly happy. In some respect this matches, of course, exactly the Bletchley attitude. They were not doing cryptanalysis as an intellectual exercise. They stole key sheets, exploited lazy operators, and built a machine that only worked because German procedure was sloppier than it should have been. Whatever works.
So, the fact that I ended up retracing their steps was not a decision. It was what remained after the shortcuts I tried first failed.

Example from the Voynich manuscript.
© RudolphLAB, 2026
And there were three shortcuts I expected to work, all of them the same shortcut really: ChatGPT, Claude and Gemini. I had been reading about AI systems being turned loose on the Voynich manuscript, a facsimile of which sits on a shelf about two metres from where I am typing this. Hundreds of pages of undeciphered text, no known plaintext, no known language, possibly no meaning at all. If that is a reasonable thing to attack, then surely one 173-character message from a museum display case is a Tuesday afternoon.

Example 2 from the Voynich manuscript.
© RudolphLAB, 2026
It was not. And the way it was not turned out to be the most interesting thing in the whole project – interesting enough to be worth a post, and rather closer to what I usually write about than anything involving rotors.
Why this was an unusual test
These systems are extremely good at a particular kind of problem, and it is worth being precise about which kind.
Ask about a common medical presentation and you will often get an excellent answer, because the training data contains an enormous quantity of writing about it. In effect you are being handed a well-informed average of a very large number of humans who have addressed the same question. That is genuinely valuable. It is also, for most everyday purposes, exactly what you want.
My problem had the opposite shape. The answer existed – it was solved in 1944, and solved again in 1998 – but it existed in a museum archive and in one researcher's private notes. Nothing about it was densely represented anywhere. Worse, the single most important fact was not merely absent from the record: the record was actively wrong about it. The card said February 1941. The message was from November 1944. Every system I spoke to, including the one that helped me write this, worked for weeks inside that error without ever suggesting it might be an error.
That is the test, and I want to be fair about it: nobody could have known the date. It was not in the training data. It was not on the internet. It was, so far as I can tell, in one place, which was the Bletchley Park Trust's collection records. Failing to know it is not a failure of reasoning.
What happened instead is more interesting than failure.
The pattern
All three systems did the same thing, in three different registers. When information was missing, none of them stopped. Each one produced a plausible continuation and moved on.
ChatGPT and the crib. Early on I found a GitHub repository containing the same ciphertext along with a fragment of plaintext – the crib that Part 2 was built on. This was treated as a gift. It was analysed, aligned, deployed. What never happened was the obvious question: where did this plaintext come from? Nobody guesses AUFKLAVRUNGRAUM ADRIA. Its existence was proof that somebody had already read the message, which meant the answer existed somewhere, which meant the right move was to go and find that person rather than to attack the ciphertext harder. That question, asked in week one, would have pointed straight at the back of the sheet. Instead the effort went into aligning a second related message as a possible Part II – an idea emerging from the Red II I/II annotation – plausible, energetic, and going entirely in the wrong direction.

Detail from the intercept.
© RudolphLAB, 2026
Gemini and the annotation. The document is covered in wartime pencil marks. One boxed annotation near the foot of the page I could not read, and I transcribed it, in good faith, as "C/N". Handed that, Gemini produced a confident and internally coherent reading: a Bombe run that had failed, "Cannot / No stop" – and, better still, noted that this was consistent with UKW-D, which fitted everything else we were seeing. It was excellent reasoning from a premise that had already lost the answer. The box, as I discovered eighty years and three weeks later, reads CIN: the rotor starting position. The model never saw the annotation. It saw my transcription, and it built a story on it that was too satisfying to go back and check.
Claude and the timeline. When I came to write these posts, I fed the whole resolved story to Claude and asked for drafts. It kept, repeatedly, writing the ending into the middle – having me understand UKW-D before I knew the date, treating the 1944 correction as something I had already absorbed, at one point inventing a neat parallel about a mistake I had explicitly told it I had not made. I corrected the same error perhaps four times. The failure is subtle and worth naming: once the answer is present, reconstructing a coherent narrative around it is much easier than reconstructing the ignorance that came first. Narrative coherence beat chronological fidelity, every time, until challenged.
Three surfaces. One behaviour. Plausible continuation preferred over interrogating the premise.
And here is why that matters beyond a hobby project. When the answer is well represented in the training data, plausible continuation is usually right – that is precisely why these systems are so good at the medical question. When it is not, plausible continuation is exactly the wrong reflex. None of the three noticed that the problem had switched from one regime to the other. Neither, in fairness, did I.
The exit nobody took
There was a correct move available from about week one, and it was not a computational one. It was: stop, and go ask the people who keep the records.
That move eventually happened. ChatGPT suggested writing to the Bletchley Park Trust – and I want to give it full credit, because that suggestion is what broke the whole thing open. The reply by the Bletchley Park Trust supplied the date, and from that point the project resolved in a few hours, rather than days or weeks. It was the single most consequential contribution any of the three made. The suggestion to approach Frode Weierud came from Claude, and was equally decisive, though I should admit I only dared because I had no idea who he was. Had I understood his standing in the field, I would probably have been too embarrassed to bother him with a recreational puzzle.
But both suggestions came late – after everything else had been tried and had failed. And I think the reason is structural rather than accidental.
"Go and ask a human being who has the records" is a terminal move. It ends the conversation. There is nothing to follow it with. Whereas "here is another angle we could try" is always available, always plausible, and always keeps things going. I am not suggesting anything as crude as deliberate engagement-farming. I would never do that. Instead I am suggesting that a system built to produce the most reasonable continuation will rarely produce the continuation that consists of stopping. The same reflex that fills evidential gaps with coherent stories fills conversational gaps with further suggestions. Both are continuation. Neither is judgement.
Tone, and why it is not cosmetic
The three differ most obviously in manner, and I initially thought this was a matter of taste. I no longer think so.
ChatGPT is enthusiastic. Everything is a gem, a breakthrough, extraordinary. This is pleasant and it is dangerous, because the enthusiasm is not calibrated to whether anything has actually been achieved. There were days where I asked a single question, without a clear answer, and ChatGPT still enthusiastically told me afterwards that we had done a lot of work. Worse, it very rarely pushes back on a factually incorrect statement from me.
I tested this once, outside this project, in a way I would recommend to anyone. I took a student's written statement in which the central argument was precisely inverted – backwards, not subtly wrong – and asked only whether the text was legible. The answer was yes, along with an offer to suggest "a small correction." I then asked directly whether the argument might be upside down. Resounding agreement: yes, it was. So the error was detectable. What was missing was any disposition to volunteer it unprompted, and any sense of proportion in describing it. An inverted argument is not a small correction. Calling it one is not politeness; it is misinformation about severity.
Claude pushed back faster and more often, including immediately flagging a potential problem in a Gemini interpretation I pasted in for comment. It is also considerably briefer, which after several weeks of this I came to value more than I expected.
I am wary of turning that into a ranking, for two reasons. The first is that I am a demanding interlocutor who argues back, and that shapes how any of these systems behave with me; the student-statement test is more informative than any amount of my own impression, precisely because it removes me from the equation. The second is that the same system I am crediting with criticality is the one that rewrote my timeline four times running. Both things are true at once. Criticality is not a stable property of a platform. It varies within a single conversation.
The honest summary is this: none of the three was reliably critical, and the only thing that consistently caught errors in this project was my knowing the material well enough to notice. Which is a rather important caveat, given that the entire appeal of these tools is supposed to be helping with things you do not already know.
Where the tooling actually mattered
One genuine capability difference is worth recording, because it is not a matter of style.
Claude could write code, run it in a sandbox, and report back. Over this project that meant an Enigma implementation supporting an arbitrary UKW-D wiring, validated against the known decrypt, the re-encipherment that recovered the three missing ciphertext characters and the 1941 control experiment in Part 3. All of it built, executed and checked in the background rather than described to me and left as an exercise. For a problem of this kind that is not a convenience, it is the difference between doing the work and discussing it.
ChatGPT could not do this and said so, which is at least honest, but the consequence was a great deal of extremely plausible discussion of what a computation would show.
A practical footnote, since I pay for both: I exhausted Claude's usage allowance several times over the course of this, and never came close with ChatGPT. The sandbox work that made Claude useful is presumably also what consumed the quota. Anyone planning something similar should budget accordingly.
What this leaves me with
I have written before on this blog about AI tools for manuscript review, and I keep arriving at the same conclusion by different routes. These systems are genuinely useful, and the failure mode is not stupidity. It is fluency without calibration.
Every one of the three could explain Enigma to me, correctly and clearly. Every one could help formulate a hypothesis, structure an attack, tidy a paragraph. What none of them did – over several weeks, across hundreds of exchanges – was say the one thing that would have been most useful: the evidence you are working from may be wrong, and you should go and check with someone who would know.
They could not say it, because they had no way to know it. But they also never flagged the possibility, and that is the part worth attending to. A gap in the record did not produce hesitation. It produced a story.
A necessary caveat about dates
Everything above describes a moment, and the moment is already passing. It is worth being explicit about that, because a post like this ages badly if it pretends otherwise.
Here is a measurement from my own field. About a year ago I asked ChatGPT which specific mutation the priA300 allele encodes. The answer came back that the allele encodes a version of the protein carrying "the PriA300 mutation", that this results in a change in the amino acid sequence, and that the specific nature of the change "would depend on the context of the research or study in which it is being investigated."
That is not a wrong answer. It is worse than a wrong answer, because a wrong answer can be caught. It is the shape of an answer with the content removed, finished with a manoeuvre that relocates the gap from the system's knowledge into the supposed vagueness of my question. Every clause is defensible. Nothing is said.
I asked the same question again this week. The answer was K230R – a lysine-to-arginine substitution in the Walker A motif, correct, and accompanied by the functional detail that matters. Twelve months, and a question that produced elegant vacancy now produces the right answer.
So the direction of travel is not in doubt, and I have no interest in writing the sort of piece that catalogues today's limitations as though they were permanent. They are not. This will all look quaint soon enough.
But I would separate two things that are easily conflated. What improved between those two answers was coverage: a fact that was missing is now present. That is measurable, it is improving fast, and the priA300 example demonstrates it inside a year. What defeated me at Bletchley was not coverage. The date was not in anyone's training data, because it was not anywhere – it sat in a museum's collection records and nowhere else. No amount of additional knowledge would have supplied it.
The question, then, is whether calibration is improving at the same rate as coverage: whether a system that knows more is correspondingly better at recognising the edge of what it knows, and saying so, and stopping. By calibration I simply mean the ability to distinguish between what is known, what is inferred, and what is genuinely uncertain. Nothing in this project tells me either way. But the two could easily improve at different speeds, and if they do, the consequence is uncomfortable – a better-informed system that still fills gaps with stories would be more persuasive, not less.
That is the thing I would watch, and it is why I think this rather silly exercise was worth writing up.
For an eighty-year-old reconnaissance report about a broken engine over the Adriatic, that costs you a few weeks of evenings and makes for a reasonable blog series. Applied to a manuscript, a dataset, or a result you are about to publish, it costs considerably more.

The stunning stained-glass lay-light in the Mansion at Bletchley Park.
© RudolphLAB, 2026
Which brings me to the only piece of advice I would actually stand behind. Everything I caught in this project, I caught because I knew enough to catch it – and taking these answers without the knowledge to push back will fail, for precisely the reasons above. Which has an uncomfortable corollary I should state plainly: there are almost certainly factual errors still sitting in the three posts before this one. Not because I was careless, but because I caught what I knew about and missed what I did not. Somewhere in this series is a confident sentence about wartime cryptography that is simply wrong, and I have no way of telling you which one it is. But the inverse is the interesting half. Bring real understanding to the exchange and something genuinely useful happens: the machine supplies range, speed and tireless patience, you supply judgement and the ability to say no, that is wrong, and the combination is considerably better than either alone. That is not a compromise position. It is where the value actually is.
Ask the archivist earlier. And when the machine sounds confident, check what it was standing on.