Skip to main content

1.4 Responsible Use of AI

Take Home message

  • Responsibility is a design task, not a control task. Detection of AI does not work properly.
  • And AI detection is unfair when it fails. There is a risk of mistakenly identifying answers as AI-generated, or vice versa. Several European countries have already ruled out these tools.
  • Confident invention and inherited bias follow from how these systems are built and judged. The consequence should be: teach them rather than waiting for a fix. 
  • The working principle is: think first, then AI. Making people commit to their own answer beats explaining the machine; building limits into the software beats banning it.
  • Prompts alone don't hold. A chatbot tasked with acting as a tutor tends to drift back to giving answers, which makes reliable limits a purchase question.
  • Amplification is the default; equalisation must be built from access, design and scaffolding together — and the scarcest of the three is a trained teacher.{{rev}}
  • Inclusion is the clearest win — and where teachers themselves see the strongest case.

If you remember only one sentence from this page: The dams are breaking. Teach swimming.

The reading journal, second look

Back to our ninth-grade student and his suspiciously polished reading journal. The instinct is almost universal, and it is the instinct of a good teacher: find out whether he used AI. Build a better dam.

This chapter is about why that instinct, on its own, leads nowhere. We will come to what works instead.

Start with the uncomfortable part. Teachers cannot reliably tell AI text from student text. In a German study, 289 teachers of different experience were given a mix of student writing and chatbot writing and asked to sort it. Neither group could. But both groups were confident they could. Experience made almost no difference. At scale, it is worse. At one university, researchers submitted chatbot-generated answers into the real examination system through the normal channels, without notifying the markers. 94 % were never questioned. Moreover, on average, they were graded above the work of actual students. So, what's the consequence? We reach for the machine! And here the finding stops being awkward and becomes an ethical problem!

The detector that fails the wrong students

Researchers tested seven of the most widely used AI detectors on TOEFL essays written by students learning English as a foreign language. Alongside, they put essays by native-speaking American eighth-graders. The native speakers' work was classified correctly almost every time. Of the foreign-language essays, more than six in ten were wrongly flagged as machine-written.


The mechanism is not mysterious. These tools do not detect machines;the work of Large Language models, they detect predictability. Text that uses common words in expected orders scores as artificial. Anyone writing in a second language, with a smaller vocabulary and safer sentence structures, produces exactly that signal. The tool's mistakes land,fall systematically,systematically on the students who are already carrying the most. 


Institutions have done the arithmetic. Vanderbilt University pointed out that even if such a tool were wrong only 1 % of the time, that would still mean roughly 750 wrongly accused students a year out of its 75,000 submissionssubmissions. Consequently, andthey switched it off. The guidance issued by Germany's federal education ministry is blunt: the technology is "not sufficiently reliable to provide legally secure proof".{{c:Scheiter et al., 2025, p. 28}}

Much of Europe has stopped debating this. France, the Netherlands and Switzerland advise against detectors or forbid them. The United Kingdom's exam regulator has deliberately chosen human judgment over software. The Dutch argument goes one step further and is worth carrying home: pasting a student's work into a detector is itself an act of processing their personal data, and doing it without permission breaks data protection law. → 1.5 Compliant Use of AI


One piece of perspective. The cheating panic is thinner than the noise around it. Researchers surveyed students at three secondary schools before ChatGPT existed and again afterwards; cheating stayed broadly where it was. Most students rejected having a chatbot simply produce their work, while thinking it fair to use one to get started or to have something explained.


The real question was never how do I catch him? It is: what does a homework task still measure, once its product can be generated on demand?