Andrej Karpathy, a former OpenAI researcher and Tesla executive, recently commented on the use of artificial intelligence in homework. According to him, schools’ efforts to monitor whether students use AI are a lost cause. In a post on X, he wrote that detecting AI use in assignments will never be 100 percent possible. Instead, he suggests moving assessments directly into classrooms, where teachers can monitor students in person. Karpathy emphasizes that the technology is here to stay and is very powerful, so students should know how to work with AI while also being able to complete tasks without it.
Doing Homework at Home with AI
Karpathy argues that current AI detection tools do not work reliably and can be easily circumvented. According to him, the entire system is doomed to fail because AI is constantly evolving. Instead of fighting over detection, schools should assume that all assignments completed at home use AI. This would reduce stress for both teachers and students and prevent a culture of cheating. Karpathy compares the situation to calculators—they are everywhere and speed up work, but schools still teach the fundamentals of mathematics by hand so that students understand the principles and can verify errors made by the tools. Similarly, AI can fail in many ways, which is why it is important for students to be able to think independently.
According to Karpathy, most tests and assessments should take place directly in the classroom. This would give teachers control over whether students use AI or not. He proposes a "flipped classroom" model in which students learn and practice with AI at home, while exams take place at school without tools or with limited access. The goal is for students to become proficient in using AI while also being able to function without it. Karpathy himself founded Eureka Labs, a startup focused on AI and education. There, human teachers create course content, while AI assistants expand on it and guide students individually.
A number of people are talking about implications of AI to schools. I spoke about some of my thoughts to a school board earlier, some highlights:
— Andrej Karpathy (@karpathy) November 24, 2025
1. You will never be able to detect the use of AI in homework. Full stop. All "detectors" of AI imo don't really work, can be… https://t.co/jEtuuW3bGP
Pangram: The Detector That Almost Never Gets It Wrong
While Karpathy talks about the failure of detection, a new study from the University of Chicago offers the opposite perspective. Researchers tested commercial AI text detectors on a dataset of 1,992 human-written texts across six categories: Amazon product reviews, blog posts, newspaper articles, excerpts from novels, restaurant reviews, and résumés. They used four AI models—GPT-4, Claude Opus 4, Claude Sonnet 4, and Gemini 2.0 Flash—to create AI-generated texts in the same categories. They tracked two metrics: the false positive rate (FPR), where human-written text is identified as AI, and the false negative rate (FNR), where AI-generated text passes as human-written.
The Pangram detector excelled in the tests. For medium-length and long texts, both its FPR and FNR were close to zero. Even for short texts, error rates were below 0.01, except for restaurant reviews generated by Gemini 2.0 Flash, where the FNR reached 0.02. Pangram was reliable across all four AI models, with an FNR of no more than 0.02. Longer texts, such as excerpts from novels or résumés, were easier for the detector, while short reviews were more difficult, but Pangram outperformed the competition there as well.
Other Detectors and Resistance to Tricks
Other detectors, such as OriginalityAI and GPTZero, came in second. They worked well on longer texts, with an FPR below 0.01, but failed on very short samples and were vulnerable to "humanization" tools that disguise AI-generated texts. The open-source RoBERTa-based detector performed the worst, incorrectly identifying 30 to 69% of human-written texts as AI. The researchers also tested the detectors against StealthGPT, a tool that makes detection more difficult—Pangram was mostly resistant, while the others failed.
Pangram was also the least expensive, with an average cost of 0.0228 dollars per correctly identified AI-generated text. That is half the cost of OriginalityAI and one-third the cost of GPTZero. The study proposes "policy thresholds" for setting a maximum FPR, such as 0.5%, and Pangram was the only detector that maintained high accuracy under such a restriction.
Challenges for the Future
The researchers warn that the results are only a snapshot and predict an ongoing battle between detectors, new AI models, and circumvention tools. They recommend regular transparent audits, similar to stress tests for banks. While AI helps with ideas and editing, problems arise when it replaces original work in areas such as education or product reviews. Previous studies described detectors as unreliable, and OpenAI withdrew its own detector because of errors. A new version from OpenAI has not appeared, perhaps because it would reduce the use of ChatGPT among students.
Additional source: the-decoder.com



