Search
RSS Feed

Thesify – the wrong tool for the right job

by Christian Rudolph

Published: 1 July 2026

Tags: AI critique AI & digital tools research & publishing academic life critical thinking

AI manuscript review tools are currently appearing with the regularity and variety of mushrooms after autumn rain. Keeping track of them, let alone evaluating what each one actually does, is a task that sits somewhere between ambitious and impossible for a working researcher. The pragmatic response – and, if we are honest, the one most of us actually adopt – is to try a tool when it is recommended, use whatever free credits are on offer, throw a draft at it and see what comes back.

This is exactly what happened with Thesify. It was recommended, free credits were available, and a manuscript was uploaded without any particular investigation into what the tool was designed to do. The front page promises expert-level AI feedback on your manuscripts, which is precisely the expectation we brought to it. The results were instructive – partly because of what the review said, and partly because of what it revealed about the tool itself once we finally read the small print. Whether the output meets that standard for a primary research paper is, on the evidence, a matter of some debate, at least based on the version I tested in autumn 2025.

What came back

Title page

The Thesify review is produced by an AI assistant called Theo, and the output is structured, readable and clearly the product of genuine analytical effort. It scores the title and abstract – which it rated excellent. It evaluates the introduction against a rubric covering contextual foundation, literature positioning, structure, and research gap alignment. It extracts what it calls the thesis statement and subjects it to a series of tests: the So What test, the How and Why test, whether the argument is debatable.

That last category is where things become interesting. The paper in question is a primary research paper reporting experimental findings. Its central claim is not an argument to be debated – it is an observation to be verified. Applying the criteria of undergraduate essay writing to a specialist research paper produces, predictably, a mixed result. The paper fails the So What test because it does not sufficiently connect its findings to broader applications. It fails the How and Why test because it does not explain the significance of the coordination it describes. The introduction scores zero for Research Gap Alignment and zero for Problem Significance.

Thesis Statement

These are not unreasonable criticisms of a student essay. They are the wrong criticisms of a research paper that had and I cannot deny a certain degree of satisfaction that at the point of writing this it was accepted into Nucleic Acids Research.

The right tool for a different job

Looking more carefully at what Thesify actually is – something we perhaps should have done before uploading – clarifies the picture considerably. It appears to be primarily designed for graduate students writing theses and dissertations, and for early-career researchers developing their academic writing. The rubric-based approach, the thesis statement evaluation, the essay structure scoring – all of these make complete sense in that context. A PhD student writing their first chapter genuinely benefits from being asked whether their argument passes the So What test. An experienced research team submitting to a specialist journal has, in all likelihood, already answered that question to their own satisfaction.

Thesify does market itself toward researchers as well as students, and it does promise to spot gaps and problems before reviewers do. So the mismatch is not entirely a failure of due diligence on our part. But the tool is calibrated for a different kind of writing problem than the one a submitted research manuscript presents – I guess in some way the name "Thesify" is giving it away, as does the writing style.

Thesify addressing students

What it did and did not find

In fairness, a small number of the observations were genuinely useful. A handful of points were adopted when working through the next draft – enough to justify the time spent. This fits the pattern I have observed across all the AI review tools I have now tried: there is almost always something worth taking away, embedded in a larger body of feedback that either misreads the task or asks for a different paper entirely.

The distinctively Thesify failure mode, though, is only in part asking for more experiments or a broader mechanistic scope – that is what LINER AI and q.e.d. tend to do, as I described in earlier posts. Thesify's failure mode is applying essay-marking criteria to scientific reporting. It is a well-built tool doing exactly what it was designed to do, in a context it was not designed for.

A note on the landscape

The broader point is worth stating plainly. When AI academic tools are appearing faster than any individual researcher can meaningfully evaluate them, the try-it-and-see approach is not irrational – it is the only realistic option. Free credits exist to encourage exactly this behaviour, and recommendations spread through academic networks in the same way they always have: someone finds something useful, tells a colleague, and the colleague tries it without necessarily reading the documentation first.

The lesson is not to be more cautious about trying new tools. It is to be appropriately sceptical about what any given tool is actually measuring, and to read what comes back with the same critical eye you would apply to any other source of feedback. These platforms are not interchangeable, and the confidence of the output is not a reliable guide to its relevance.

We skipped the small print. The small print, it turns out, mattered.


Similar and related blog posts: