
Banning researchers from using AI to assess the merits of scientific papers doesn鈥檛 make much difference to the conclusions the reviewers reach 鈥� in large part because many of them ignore the prohibition and use AI anyway.
A team led by computer scientists at Microsoft Research ran a large randomised experiment during the , one of the world鈥檚 biggest AI conferences, which took place in Seoul, South Korea in July.
Some 24,661 papers were submitted to the conference, requiring oversight from 17,886 reviewers. When researchers submitted their work, they could choose how it would be reviewed. The first option was to have it assessed by reviewers聽using a conservative policy that entirely banned the use of large language models (LLMs) to help with reviewing. The other option was to have the assessment performed by reviewers using a permissive policy that allowed the use of LLMs to help understand papers, check related work and polish reviewers鈥� own writing, but not to judge the merits of a paper or draft a review of the work.
Advertisement
The team led by Microsoft Research scientists discovered that聽papers reviewed under the two policies had almost identical acceptance rates 鈥� 27 per cent under the stricter policy and 26.5 per cent under the permissive one 鈥� while average review scores were 3.31 and 3.32 out of 6, respectively. Reviewer confidence was also virtually unchanged. Reviews written under the permissive policy were around 5.5 to 7 per cent longer, and were rated as slightly higher quality by experts, although a reviewer-by-reviewer comparison found no meaningful difference.
The reason for that lack of difference became clear after an anonymous survey of 1486 of the reviewers who took part in the exercise. There, 22.5 per cent of reviewers who were instructed not to use AI admitted to using an LLM anyway. 鈥淪elf-reported non-compliance of 23 per cent is definitely higher than what I was anticipating,鈥� says at Microsoft Research, who was involved in the study and was also one of the conference organisers.
Those who used AI when told not to did so to brainstorm feedback, draft review text, read the papers and summarise their strengths and weaknesses. 鈥淥ur post-conference survey reveals some reasons why reviewers might end up violating policy, like large reviewing workloads and insufficiently clear rules,鈥� says Dud铆k.
鈥淭his study shows that banning AI in peer review is very difficult to enforce,鈥� says at the University of Wolverhampton, UK. 鈥淏ecause of the workload of academics, there is a temptation鈥� to use AI, says Kousha.
In a separate part of the study, the researchers analysed the reviews using Pangram, an AI text detector. They found that only 52.2 per cent of reviews under the conservative policy were classified as fully human-written, compared with 37.0 per cent under the permissive policy 鈥� though the researchers warn AI text detectors aren鈥檛 perfect.
鈥淚 believe our work has relevant insights for other computer science conferences 鈥� and potentially venues in other fields as well 鈥� that are facing similar challenges,鈥� says at Microsoft Research.
arXiv