Is Homework Still Worth Assigning in the Age of AI?

Is Homework Still Worth Assigning in the Age of AI?

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
4. 12. 2025
4 minutes reading
Is Homework Still Worth Assigning in the Age of AI?

Andrej Karpathy, a former OpenAI researcher and Tesla executive, recently commented on the use of artificial intelligence in homework. According to him, schools’ efforts to monitor whether students use AI are a lost cause. In a post on X, he wrote that detecting AI use in assignments will never be 100 percent possible. Instead, he suggests moving assessments directly into classrooms, where teachers can monitor students in person. Karpathy emphasizes that the technology is here to stay and is very powerful, so students should know how to work with AI while also being able to complete tasks without it.

Doing Homework at Home with AI

Karpathy argues that current AI detection tools do not work reliably and can be easily circumvented. According to him, the entire system is doomed to fail because AI is constantly evolving. Instead of fighting over detection, schools should assume that all assignments completed at home use AI. This would reduce stress for both teachers and students and prevent a culture of cheating. Karpathy compares the situation to calculators—they are everywhere and speed up work, but schools still teach the fundamentals of mathematics by hand so that students understand the principles and can verify errors made by the tools. Similarly, AI can fail in many ways, which is why it is important for students to be able to think independently.

According to Karpathy, most tests and assessments should take place directly in the classroom. This would give teachers control over whether students use AI or not. He proposes a "flipped classroom" model in which students learn and practice with AI at home, while exams take place at school without tools or with limited access. The goal is for students to become proficient in using AI while also being able to function without it. Karpathy himself founded Eureka Labs, a startup focused on AI and education. There, human teachers create course content, while AI assistants expand on it and guide students individually.

Pangram: The Detector That Almost Never Gets It Wrong

While Karpathy talks about the failure of detection, a new study from the University of Chicago offers the opposite perspective. Researchers tested commercial AI text detectors on a dataset of 1,992 human-written texts across six categories: Amazon product reviews, blog posts, newspaper articles, excerpts from novels, restaurant reviews, and résumés. They used four AI models—GPT-4, Claude Opus 4, Claude Sonnet 4, and Gemini 2.0 Flash—to create AI-generated texts in the same categories. They tracked two metrics: the false positive rate (FPR), where human-written text is identified as AI, and the false negative rate (FNR), where AI-generated text passes as human-written.

The Pangram detector excelled in the tests. For medium-length and long texts, both its FPR and FNR were close to zero. Even for short texts, error rates were below 0.01, except for restaurant reviews generated by Gemini 2.0 Flash, where the FNR reached 0.02. Pangram was reliable across all four AI models, with an FNR of no more than 0.02. Longer texts, such as excerpts from novels or résumés, were easier for the detector, while short reviews were more difficult, but Pangram outperformed the competition there as well.

Other Detectors and Resistance to Tricks

Other detectors, such as OriginalityAI and GPTZero, came in second. They worked well on longer texts, with an FPR below 0.01, but failed on very short samples and were vulnerable to "humanization" tools that disguise AI-generated texts. The open-source RoBERTa-based detector performed the worst, incorrectly identifying 30 to 69% of human-written texts as AI. The researchers also tested the detectors against StealthGPT, a tool that makes detection more difficult—Pangram was mostly resistant, while the others failed.

Pangram was also the least expensive, with an average cost of 0.0228 dollars per correctly identified AI-generated text. That is half the cost of OriginalityAI and one-third the cost of GPTZero. The study proposes "policy thresholds" for setting a maximum FPR, such as 0.5%, and Pangram was the only detector that maintained high accuracy under such a restriction.

Challenges for the Future

The researchers warn that the results are only a snapshot and predict an ongoing battle between detectors, new AI models, and circumvention tools. They recommend regular transparent audits, similar to stress tests for banks. While AI helps with ideas and editing, problems arise when it replaces original work in areas such as education or product reviews. Previous studies described detectors as unreliable, and OpenAI withdrew its own detector because of errors. A new version from OpenAI has not appeared, perhaps because it would reduce the use of ChatGPT among students.

Additional source: the-decoder.com

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok