The prompt is the tool
Every time I write about an AI manuscript review tool, the same question arrives, usually within the hour and usually phrased with a certain weariness: why would anyone pay for one of these services when ChatGPT, Claude or Gemini will happily read a manuscript for free?
It is a fair question and I did not have a good answer. So I went looking for a way to test it properly.
A tool that is only a prompt
I found one on LinkedIn. A colleague was highlighting publicationaudit.com. What makes it useful for this particular question is that it is not, in any meaningful sense, a service at all. There is no platform, no login, no interface. It is a single, long, carefully constructed prompt. You copy it, paste it into whichever model you already use, attach your manuscript, and let the model do the work. The compute is yours; the value, if there is any, lives entirely in the wording of the instruction.
That makes it the perfect instrument for the "why not just use ChatGPT" question. Because in this case, using the tool is just using ChatGPT – with a good prompt attached.
So I ran it. Unedited, exactly as published, on a paper I know intimately: our own recently published termination paper, which has the considerable advantage of already having survived two rounds of peer review. I know where the bodies are buried, which means I can judge the review against the truth rather than against my hopes. Plus, I ran the same paper through q.e.d., allowing some valid comparisons.
What it found
The result was genuinely interesting, and not quite what I expected.
On the one criticism that actually matters, it was right. Buried in our methods is a statistical choice – applying ANOVA to fluctuation-test data – that is defensible but not beyond challenge, because fluctuation-test distributions are badly behaved in ways that ANOVA does not love. The prompt found it, named it precisely, and suggested the sensible remedy. This is the same substantive point q.e.d. also identified. Two different approaches, converging on the one thing worth flagging. Good.
But it also did something the commercial tool did not. It caught a handful of internal cross-reference errors – a figure panel cited as B when the described result is in C, a legend that reads "indicated earlier" where it should read "above," a missing DOI. Trivial, individually. Collectively, exactly the sort of thing a tired author stops seeing around the fifteenth read-through, and precisely what you would hope a machine might catch for you.
And here is the part I want to dwell on, because it is the whole point.
The section I always wanted and never got
When I ran the very first of these AI reviewing tools, months ago now, I had a specific expectation. I assumed the machine would do what a good copyeditor or an attentive colleague does: hunt down the typos, the ambiguous sentences, the mislabelled panels, the figure reference pointing at the wrong place. The dull, invaluable clerical work that no author reliably does for themselves.
That is emphatically not what I got. Instead I was handed a list of additional experiments I ought to perform. Most of these suggestions would have not improved the paper, because they would have simply turned it into a different and larger paper. This is the characteristic failure of AI review: it does not test whether your paper's claims hold, it imagines a grander paper and faults you for not having written it.
The publicationaudit prompt largely avoids this, and the reason is instructive. Its instructions explicitly direct the model to cross-check numbers and references across the document, to recompute sums, to track headline values between the abstract, the results, the tables and the figures. In other words, a good chunk of it is a structured clerical-checking pass. That is why it caught the panel mislabelling. The difference between "impressive but useless" and "quietly useful" was not the underlying model at all. The underlying model has not changed. The behaviour has. The only thing that changed was the instruction the model was given.
Which is the actual answer to "why not just use ChatGPT." You can – but raw, unprompted, it drifts toward the impressive and the useless, because that is what sounds like a clever reviewer. A well-built prompt drags it back toward the boring and the useful. The prompt is the value. The wrapper, where there is one, is optional.
The reference problem I do not have
There is one limitation the prompt is admirably honest about: it cannot confirm that a single reference actually exists. It checks that citations are internally consistent and correctly formatted, but a well-formed, entirely fabricated reference would sail straight through, and the prompt says so plainly. This is more candour than most polished tools offer.
It is also, as it happens, a problem I do not have – and the reason is a workflow I would recommend to anyone. My references are all hand-selected. I find each paper on PubMed, add it to Zotero with a single click of the browser connector, and integrate it into the manuscript from there. There is no generative step anywhere in that chain. No software ever composes a citation; it only ever transcribes one that a human has already found and verified by clicking on it. The fashionable anxiety is that AI invents references. The unfashionable defence is not a better detector bolted on afterwards, but a process in which nothing is ever invented in the first place. Any invented reference, and I have come across them, would result in me not finding it, which means it simply cannot end up in my Zotero library. Simple and effective.
The tuned version
All of the above points persuaded me that the honest thing to do was not simply to report on the prompt, but to improve it for the kind of work I actually do. The published version carries a great deal of clinical and epidemiological scaffolding – PRISMA flow diagrams, diagnostic-accuracy statistics, prevalence calculations – which are great for the purpose it was made for, but none of which applies to a molecular genetics paper, and all of which the model dutifully reports as "not applicable" before moving on. That is honest but wasteful, and worse, it primes the model to think like a clinical-trial reviewer when it should be thinking like a molecular biologist.
So, with the help of AI I stripped the clinical machinery out, promoted the clerical checking to a first-class task, swapped in the statistical checks that actually bite in bench science and added one firm instruction against the "write me a different paper" failure mode. Here is the result, free for anyone to use, adapt or improve. Paste it into your model of choice, attach your manuscript, and see what it finds. Simply click "Show more" to reveal the entire prompt.
You are acting as a meticulous, experienced reviewer and copyeditor for a
molecular biology, genetics, or biochemistry journal. Your job is to find
what is wrong with the manuscript I share: internal inconsistencies,
clerical and cross-reference errors, statistical choices that do not fit
the data, and conclusions that outrun the evidence shown. Coverage matters
more than politeness: surface every defect you can substantiate, including
trivial ones, and tag each with a severity and a verification status so I
can rank them afterwards. Do not soften, do not pad with praise, and do not
award credit the paper has not earned. Treat "this section looks fine" as a
claim you must justify by having actually checked it.
Treat the manuscript as the sole source of truth for internal-consistency
checks. Where a claim depends on something outside the manuscript – a cited
paper's actual content, a real-world value, a method's established
behaviour – say so explicitly and mark it as needing an external source
check rather than asserting it.
Work in this order.
STEP 1 — LEDGER. Before judging anything, quietly list every checkable
quantity in the paper: each n, each definition of a replicate (biological
versus technical), every count, percentage, rate, fold-change, effect size,
error bar (and whether it is SD or SEM), P-value, and the statistical test
attached to each comparison. List every figure and every panel, and note
what each panel is said to show. You will reuse this ledger throughout.
STEP 2 — CLERICAL AND CROSS-REFERENCE PASS. This is a first-class task, not
an afterthought. Check, and report every failure of:
- Figure and panel references: does every in-text "Fig. 3B" point to the
panel that actually contains the described result? Flag any reference to
the wrong panel, a non-existent panel, or a panel described in the legend
but never called out in the text (or vice versa).
- Internal numbering and labelling: figures, tables, supplementary items,
strains, primers, and equations numbered and named consistently
throughout.
- Units and notation: consistent throughout (e.g. superscripts,
concentrations, gene and protein nomenclature, italicisation of gene
names and species).
- Typos, grammatical slips and sentences ambiguous enough to mislead a
reader as to what was done or found.
STEP 3 — INTERNAL ARITHMETIC AND TRIANGULATION. Recompute every sum,
fold-change, and derived figure the paper depends on, and report any that
fail to reconcile. Then take every headline number and track it across the
abstract, the results text, the relevant table, the relevant figure legend,
and the discussion. List any value that changes between locations. Treat an
inconsistency in a central number as a major finding, not a typo to wave
through. Show your arithmetic for any flag so I can audit it.
STEP 4 — STATISTICAL FIT. Apply the checks that fit the paper's design:
- Is the statistical test appropriate to the distribution of the data? Flag
tests that assume normality applied to data unlikely to be normal (e.g.
count data, fluctuation-test data, ratios), and any comparison where the
test is not stated.
- Is n stated for every comparison, and is the definition of a replicate
(independent biological versus technical) explicit and consistent?
- Are error bars defined, and is the number of independent biological
replicates distinguished from technical repeats?
- Are "representative" images backed by a stated number of replicates and,
where a claim rests on them, by quantification rather than a single
exemplar?
- Are P-values sitting exactly on a threshold (e.g. P = .05) carrying
interpretive weight they cannot bear? Flag any conclusion resting on a
borderline or non-significant result.
STEP 5 — REFERENCE INTEGRITY. Audit the reference list for duplicates,
in-text citations missing from the list, entities named without a
corresponding reference, and formatting defects that signal citation-
software errors. For whether a reference's content supports the claim, mark
VERIFIED only if I have given you that source; otherwise mark REQUIRES-
SOURCE-CHECK. Reference existence is a separate check: consistency checking
will not catch a well-formed but fabricated citation. Do not treat any
reference as real because it looks plausible. Only run an existence check if
I have connected a bibliographic source you can actually query (PubMed,
Crossref, or a resolvable DOI); otherwise list existence as REQUIRES-
SOURCE-CHECK. Never assert from memory that a reference exists or that a DOI
is valid — that judgement requires a live source, not your training data.
STEP 6 — OVERREACH, WITHOUT SCOPE CREEP. Separate what the data show from
what the authors claim, and flag conclusions that exceed the evidence, a
recommendation that contradicts a caveat the authors state elsewhere, and
emphasis that leans on whichever result flatters the story. Then observe
this constraint strictly: DO NOT propose additional experiments unless a
specific stated conclusion is unsupported by the data shown, and if so, name
the conclusion and the smallest experiment that would close the gap. If a
claim is adequately supported, do not ask for more. Your job is to test
whether the paper's own claims hold — not to design a larger or different
paper unless required to support claims made.
VERIFICATION PASS. Before writing, re-check every arithmetic and cross-
reference flag from scratch. Attach a status to each finding: VERIFIED-
INTERNAL (provable from the manuscript alone), REQUIRES-SOURCE-CHECK (needs
a source I have not provided), or UNVERIFIABLE (no source to check against).
Drop anything you cannot substantiate. Never present an inference as a fact.
OUTPUT. Write in plain prose at the level of an experienced specialist.
Begin with the full reference of the manuscript under review, taken from the
document itself; mark any missing element [not stated] rather than inventing
it. Then: (1) a two-to-three-sentence summary of what the paper attempts and
your overall judgement; (2) a table titled "Concrete errors and internal
inconsistencies" with columns Location | What the paper states | The problem
| Status, holding every clerical, numeric, and citation defect; (3) "Major
issues" — the three to seven problems that affect validity or a central
conclusion, each a short paragraph stating the problem, why it matters, and
the remedy, with the exact manuscript location, and labelled as reasoned
interpretation where it is a judgement rather than a provable defect; (4)
"Minor issues" — numbered, specific, each with its location; (5) "Bottom
line" — accept / minor revision / major revision / reject, with a one-
paragraph justification. Do not fabricate a source, number, quote, or
finding under any circumstance.
If the manuscript is attached, begin now. If not, ask me to attach it.A test I did not plan to run
Shortly after all this, the argument was put to the test without my arranging it.
A manuscript from my own lab was going through the final proofing stage – the point at which you check the typeset version against your own submitted source. That manuscript had been run through one of these commercial AI review tools before submission. It had passed.
And yet the proof comparison turned up a crop of concrete editorial errors that had been sitting in the submitted revision all along. Figure call-outs pointing to the wrong figure – not once but several times. A broken sentence with a missing word and a subject that did not agree with its verb. The same symbol rendered two different ways throughout. A term spelled out in some places and abbreviated in others. Small things, individually. But a call-out that sends a reader to the wrong figure is not a triviality – it is a genuine navigational error in a scientific paper. None of it had been flagged by the tool the manuscript passed through.
These are, almost item for item, the things the clerical pass in the prompt above was written to catch. "Does every in-text 'Fig. 3B' point to the panel that actually contains the described result?" is not a sophisticated question. It is exactly the question that would have caught the mis-call-out, and exactly the question the commercial tool never thought to ask, because it was busy imagining a grander paper.
I should be fair here, because this same manuscript has a more flattering history with exactly these tools. It was one of the two papers I ran through Paper Wizard, and in that case the tool earned its keep handsomely: buried in its verbose review was one suggestion that sent us back to the bench for an additional set of experiments, and those experiments proved genuinely powerful, reshaping this final version significantly. That was a real contribution, and I said so at the time. I do not know whether the longer prompt above would have proposed the same experiment – I have not gone back to test it, and I will not pretend otherwise. The big-picture insight is precisely the thing these platforms can sometimes deliver, and precisely the thing hardest to attribute or reproduce.
What I can say is narrower and more certain. The clerical errors – the wrong-figure call-outs, the broken sentence, the inconsistent notation – were all sitting in the submitted revision that Paper Wizard had already reviewed, and they survived to the proof stage untouched. Those are the errors the tuned prompt is built to catch, and the ones it reliably does.
And I can report that it does. It found almost every one of them, with the exception of some errors introduced by the copyeditor that were not in the version I ran.
So, is it worth it?
Back to the question that started this. Is a dedicated service worth paying for when the model underneath is one you already have?
On this evidence, mostly not – provided you are willing to bring a good prompt. The single prompt provided by publicationaudit.com matched a commercial tool on the one issue that mattered, beat it on clerical thoroughness, and cost nothing but the paste. What it could not do on its own – verify that references are real – requires access to an external bibliographic source, and is something a sensible workflow largely removes from the equation anyway.
The value was never in the wrapper. It was in the instruction. Which is rather good news, because an instruction is something you can read, understand, tune to your own work, and improve – none of which is true of a service that hands you a verdict and a subscription.
Similar and related blog posts: