• Wion
  • /Technology
  • /AI-generated exam answers go undetected after getting more marks than real students

AI-generated exam answers go undetected after getting more marks than real students

AI-generated exam answers go undetected after getting more marks than real students

OpenAI logo is seen in this illustration taken May 20, 2024

Researchers at the University of Reading managed to deceive their own professors by submitting AI-generated exam answers that not only went undetected but also received higher grades than real students.

The covert project involved creating fictitious student identities to submit unedited answers generated by ChatGPT-4 in take-home online assessments for undergraduate courses. Only one of the 33 AI-generated entries was flagged, with the rest achieving above-average grades compared to actual student submissions.

What does it mean?

Add WION as a Preferred Source

The researchers argue that their findings indicate that AI tools like ChatGPT are now capable of passing the "Turing test"—a benchmark named after computing pioneer Alan Turing, which measures an AI's ability to pass undetected by experienced judges.

Also read |India’s AI revolution in the spotlight

Touted as “the largest and most robust blind study of its kind,” the project sought to determine if human educators could distinguish between AI-generated and human-generated responses. The authors warn that the findings have significant implications for the future of academic assessments.

"Our research highlights the global importance of understanding how AI will impact the integrity of educational assessments," Dr. Peter Scarfe, one of the study's authors and an associate professor at Reading’s School of Psychology and Clinical Language Sciences was quoted as saying by The Guardian.

"While we may not revert entirely to handwritten exams, the global education sector must adapt in response to AI."

The study concluded that as AI continues to develop more advanced reasoning capabilities, its detectability will diminish.

Experts reviewing the study expressed concern about the viability of take-home exams and unsupervised coursework.

Professor Karen Yeung, a fellow in law, ethics, and informatics at the University of Birmingham, stated: “This study clearly demonstrates that generative AI tools enable students to cheat on take-home exams without difficulty, obtaining better grades undetected.”

The study's endnotes suggest the authors may have used AI to prepare the research, questioning, "Would you consider it ‘cheating’? If you did, how would you prove we were lying if we denied using GPT-4 or any other AI?"

But a spokesperson for Reading confirmed that the study was "definitely done by humans."