One larger problem here is the value of a research paper is rarely the specific knowledge it adds but in the process of researching that adds to the collective knowledge+experience of those involved, especially training graduate students. AI papers shortcut this entirely. Academia has a lot to answer for this too by making papers the currency of success. AI generated papers are almost shortcut learning at a full system level.
I disagree for academia using papers as an easy medium for verification and providing more knowledge. Llms always short circuit everything, so how would you fix academia?
The Medium comments on this post are also on point. Running the same experiment with accepted papers is a good control. Running a similar experiment with reviewers would be interesting, but more obnoxious because they are not being paid.
I would keep a private blacklist (shadow ban) the authors who wasted several hours of a reviewer's time to prove they were not legitimate. The existence of such a list would be problematic, though.
Could the same system we use here be applied? Accepted authors could "vouch" for "dead" papers in case they were "auto-killed"?
This system is broken and providing more evidence that it is broken isn't much of a step towards fixing it.
> All reviewers, Action Editors and Editors-in-Chief for TMLR are unpaid volunteers.
He is not an unpaid volunteer.
He's an associate professor at the prestigious Carnegie Mellon University. He is not paid by the journal, but he is paid a salary by the university, and the university expects that a small part of his academic work is to serve as an editor in academic journals.
Peer review has its historical issues, but the landscape of science and science-publishing has changed. New problems of authorship and authorial-understanding are now challenged by LLMs writing (at least) good sounding papers - some of which might be of acceptable quality in subject (I am not against AI in the sciences; some of the math work has been great). On the other hand: I am against authors not understanding their own work. High repute journals may need to add "oral exams" to the paper acceptance process...
I'm not familiar with the world of academic publishing, so I want to ask: how is the industry making sure that submissions aren't at least partially AI-generated?
Is it standard practice for authors to have to defend their submissions via interview like this? If not, why not?
Does the vetting process vary with the quality of the publisher?
As an outsider, it's extremely worrying that anyone would even attempt to submit an AI-generated paper for publication in an academic journal. At that level I would have assumed literally everybody should know better than to even try.
I’ve only published a few papers, but this interview sounds extremely unusual to me (I mean, it is clearly a special thing that the editor is doing, which is fine). I wouldn’t do something unethical, but if I had and the editor asked me for an interview like this, I’d know I’d probably been caught.
Why is it unusual: it sounds extremely time-consuming.
As to how worrying AI-generated papers are… it sounds more like a headache for the editors really.
In general, journals don’t have to be perfect; mostly researchers read research papers. You already have to read critically (publish-or-perish has been a thing for a while, so there are plenty of not-so-great papers out there). Peer review is just the “entry” barrier, science is a social process and papers become more or less influential based on a fuzzy process of citation, conference talks, and peer-to-peer suggestions.
For whom? Surely the authors can find an hour after submitting the paper to a journal?
Beware that a reviewer easily spend a full week on reviewing a paper, and there are typically three of them. So if one hour of conversation can save three weeks work, it sounds worth it.
> Is it standard practice for authors to have to defend their submissions via interview like this? If not, why not?
It isn't, but maybe it should be. For post-grad qualifications oral defense is standard, and I didn't mind defending my central thesis then, and won't mind now.
Not in a direct interview style, but most us conferences can request additional information or feedback. If they conditionally accept or reject a paper, that conditional relies on feedback from the author(s).
Interviews like this are interesting, but in no way can scale to the infinite paper slop conferences are facing.
> Not in a direct interview style, but most us conferences can request additional information or feedback. If they conditionally accept or reject a paper, that conditional relies on feedback from the author(s).
Requesting feedback is useless, as the article points out - the "authors" could not answer basic questions during the interview, but after the interview were able to send full explanations to the interviewer.
If you have indirect feedback ("please answer these questions we have") the "author" will simply feed it into an LLM and send the results back. You need to get the author to do an oral defense to verify that they wrote the paper.
This is the main problem with AI generated output, whether it's a research paper, a blog, an email, a comment on a forum, similar: the value in knowing that a human wrote $X sends a signal - that the human understands what it is they wrote, even if they misunderstand the concepts.
When you get a message from someone who is a "I only used an LLM to clean it up, the thoughts are all mine"[1] person, you cannot engage with them, because they may not understand the message they transmitted, and so any human engaging with them is only burning their own time for no gain.
When you get a message from a real person, you get not only the message, you also get a signal about their understanding. That signal is missing in AI generated messages.
> I only used an LLM to clean it up, the thoughts are all mine"
I mean, it seems to me like that should be nearly everybody at this point, right? There are at least some elements of using LLMs that are the equivalent of a spell check, like asking to make sure that links work and that references are to the thing they're supposed to be, and so on. I feel like the only difference at this point is between people who do that and say that they do, and people who do that and don't say that they do.
IMHO (not a paper writer, but read a lot during my grad school years), the Genie is out of the bottle. The only way forward, as I see it, is using LLMs for reviews also. Basically, filter all submitted papers with an LLM and ask it to summarize it, find the biggest weaknesses and main strong points, etc. that a human can then use to review the paper. Basically, LLM-as-a-reviewer .
Personally, I would love to see a conference where people are explicitly encouraged to use LLMs for doing the work and writing the papers, and LLMs are used to review them too.
There should at least be a code of professional conduct where authors state the extent to which LLMs were used. (This would also help not wasting time by asking some “authors” about “their” paper.)
Journals themselves should make policies about the extent to which they allow the use of LLMs. In some areas it might be considered more benign than in others.
The entire point of an academic paper is to add to the sum of human knowledge. How can an LLM trained on a subset of human knowledge possibly even begin to accurate evaluate such a paper?
I trust an LLM to review that the language used in the paper is grammatically correct, but not to evaluate new information for accuracy.
I think that’s a lot of risk of anchoring reviewer bias. I’d be more comfortable with a triaged review where the editor’s office uses models to score whether a human editor should evaluate a paper to potentially send out for review, then the editor makes their own assessment, and the reviewers continue to do their job unassisted
Is it really “research” - as in expanding human knowledge - if nobody understands it? The point is deepening human understanding, not producing research papers
I wonder how the ratios would change for papers at different parts of the review process. For what fraction of published papers are the authors unable to answer basic questions about them?
I think authenticity and trust will command a (larger) premium in this new age of slop.
The article highlights how only one out of ten paper’s authors were able to answer questions thoroughly and at a high level. This indicates an overwhelming percentage of authors are slopping up their work with AI and submitting it without even reading it.
No doubt this is happening, but I wonder how many authors of papers "slated for desk rejection" 10 years ago could answer questions about their papers? We'd need that comparison to understand if this is a new problem or if AI is just a new source of content that the authors of poorly-written papers are using.
I would keep a private blacklist (shadow ban) the authors who wasted several hours of a reviewer's time to prove they were not legitimate. The existence of such a list would be problematic, though.
Could the same system we use here be applied? Accepted authors could "vouch" for "dead" papers in case they were "auto-killed"?
This system is broken and providing more evidence that it is broken isn't much of a step towards fixing it.
He is not an unpaid volunteer.
He's an associate professor at the prestigious Carnegie Mellon University. He is not paid by the journal, but he is paid a salary by the university, and the university expects that a small part of his academic work is to serve as an editor in academic journals.
Is it standard practice for authors to have to defend their submissions via interview like this? If not, why not?
Does the vetting process vary with the quality of the publisher?
As an outsider, it's extremely worrying that anyone would even attempt to submit an AI-generated paper for publication in an academic journal. At that level I would have assumed literally everybody should know better than to even try.
Why is it unusual: it sounds extremely time-consuming.
As to how worrying AI-generated papers are… it sounds more like a headache for the editors really.
In general, journals don’t have to be perfect; mostly researchers read research papers. You already have to read critically (publish-or-perish has been a thing for a while, so there are plenty of not-so-great papers out there). Peer review is just the “entry” barrier, science is a social process and papers become more or less influential based on a fuzzy process of citation, conference talks, and peer-to-peer suggestions.
For whom? Surely the authors can find an hour after submitting the paper to a journal?
Beware that a reviewer easily spend a full week on reviewing a paper, and there are typically three of them. So if one hour of conversation can save three weeks work, it sounds worth it.
It isn't, but maybe it should be. For post-grad qualifications oral defense is standard, and I didn't mind defending my central thesis then, and won't mind now.
Interviews like this are interesting, but in no way can scale to the infinite paper slop conferences are facing.
Requesting feedback is useless, as the article points out - the "authors" could not answer basic questions during the interview, but after the interview were able to send full explanations to the interviewer.
If you have indirect feedback ("please answer these questions we have") the "author" will simply feed it into an LLM and send the results back. You need to get the author to do an oral defense to verify that they wrote the paper.
This is the main problem with AI generated output, whether it's a research paper, a blog, an email, a comment on a forum, similar: the value in knowing that a human wrote $X sends a signal - that the human understands what it is they wrote, even if they misunderstand the concepts.
When you get a message from someone who is a "I only used an LLM to clean it up, the thoughts are all mine"[1] person, you cannot engage with them, because they may not understand the message they transmitted, and so any human engaging with them is only burning their own time for no gain.
When you get a message from a real person, you get not only the message, you also get a signal about their understanding. That signal is missing in AI generated messages.
==========================
[1] Sure, buddy. We believe you /s.
I mean, it seems to me like that should be nearly everybody at this point, right? There are at least some elements of using LLMs that are the equivalent of a spell check, like asking to make sure that links work and that references are to the thing they're supposed to be, and so on. I feel like the only difference at this point is between people who do that and say that they do, and people who do that and don't say that they do.
Personally, I would love to see a conference where people are explicitly encouraged to use LLMs for doing the work and writing the papers, and LLMs are used to review them too.
Journals themselves should make policies about the extent to which they allow the use of LLMs. In some areas it might be considered more benign than in others.
I trust an LLM to review that the language used in the paper is grammatically correct, but not to evaluate new information for accuracy.
EDIT: I found a live link on arxiv https://arxiv.org/html/2609.20481v1
https://chorasimilarity.wordpress.com/2026/06/13/a-captcha-f...
At the moment this was seen as a tongue in cheek proposal.
The article highlights how only one out of ten paper’s authors were able to answer questions thoroughly and at a high level. This indicates an overwhelming percentage of authors are slopping up their work with AI and submitting it without even reading it.