Impressive is not the same as useful
As discussions around research assessment and REF2029 have intensified, AI manuscript review tools have been appearing with increasing frequency in academic conversations. Several are now being enthusiastically recommended as part of the pre-submission workflow – tools that promise to identify weaknesses, flag methodological gaps, and simulate the kind of scrutiny a manuscript might receive from peer reviewers. The promise is appealing, particularly for researchers working without large support teams.
I have now tried three of these tools across different manuscripts, and I want to share some observations – not to dismiss them, because they are genuinely technically impressive, but because impressive and useful are not the same thing, and that distinction matters.
I have written about two of these tools previously: LINER AI and Paper Wizard, both tested on real manuscripts prior to submission. This post concerns q.e.d., which I tested on a paper that we recently got published – one that had already survived two rounds of peer review at Nucleic Acids Research. That context is important and I will return to it.
What q.e.d. produces

© RudolphLAB, 2026
Where it goes wrong
The difficulty becomes apparent when you read the review carefully against the paper it is reviewing. A significant portion of the criticism addresses either things the paper already does, or concerns that reflect a misunderstanding of the experimental logic.
The review flags the absence of direct verification that replication fork fusions occurred at specific locations in the chromosome. This sounds reasonable in isolation. In practice, it misses the point of the experimental design entirely – the entire approach is built around switching fork fusions on and off at defined chromosomal locations and observing the consequences. The on/off switch, replicated across two independent chromosomal locations, is precisely what makes the argument compelling. Demanding direct physical visualisation of a stochastic molecular event at a specific locus is not a methodological gap – it is a request for a different kind of experiment that would not actually strengthen the argument being made.
Similarly, the review raises questions about the biochemical intermediates involved in certain pathologies – intermediates that are supported by a substantial body of published work, including our 2013 Nature paper. The tool does not appear to have weighed that context. It has flagged a pattern it associates with methodological insufficiency, regardless of whether the flag applies.
This is a different failure mode from what I observed with LINER AI and Paper Wizard. Those tools tended to want a different paper – more experiments, broader scope, a complete mechanistic dissection where a targeted observation was appropriate. q.e.d. goes further: it produces criticism that looks authoritative but misreads the experimental logic. A review that is confidently wrong is, in some respects, more problematic than one that is simply overambitious.
The published paper problem
I should be transparent: testing a review tool on a published paper, as done on this occasion, is not the same as testing it on a manuscript heading for submission. A paper that has passed two rounds of peer review at a competitive journal has already had its major weaknesses identified and addressed. The remaining surface area for legitimate criticism is smaller, which makes the tool's job harder and makes gaps easier to mistake for genuine weaknesses.
That said, this is also precisely what makes the test informative. If a tool struggles to distinguish real methodological concerns from well-handled limitations and already-addressed caveats in a published paper, that raises reasonable questions about how reliably it does so in a manuscript context, where the lines are less clearly drawn.
A pattern worth noting
Across all three tools I have now tested, I notice a common tendency: the reviews feel compelled to find something to criticise. This mirrors a pattern I have written about this separately in the context of human peer review – the assumption that a reviewer who raises no concerns has not done their job. The result, whether from a human reviewer or an AI tool, is feedback that generates motion without necessarily generating progress.
The tools are impressive. The output looks professional, the structure is clear, and some of the observations are worth taking seriously. But forming a reliable opinion about how useful they actually are requires working with them for long enough to develop a sense of what is the signal and what is noise. Enthusiastic early recommendations, however well-intentioned, skip this important step.
These tools are worth knowing about. They are worth trying. But they reward scepticism, and they reward scientific confidence – the ability to read a criticism and know, from a position of expertise, whether it reflects a genuine gap or a pattern-matched assumption. Used that way, they have something to offer. Trusted uncritically, they have rather less.
Related and similar blog posts: