Who this is for
A department chair, an academic-integrity officer, a course convenor — anyone who has to write down what happens when a submitted text is suspected of being machine-generated.
It assumes you want two things at once: to deter misconduct, and not to wrongly accuse. Most policies we have read optimise the first and leave the second to the good judgement of whoever is in the room.
Disclosure. Writence builds writing tools for people working in a language that is not their first. We have an interest here. We have tried to write guidance that holds whether or not you ever use a product like ours — every recommendation below can be adopted by a department that buys nothing. If any of it only works when you are our customer, we have failed and you should discard it.
The specific failure this prevents
Detectors infer from the finished text: word variety, sentence shape, statistical smoothness. Those are also the properties that distinguish second-language English. The consequence is not a rounding error.
Stanford researchers (Liang et al., 2023) found leading detectors flagged a majority of TOEFL essays by non-native writers as AI-generated while near-perfectly clearing essays by native US students. Later work, including from Maryland in 2025, reproduces the direction of the effect. Garland (2026) frames it as a structural limit rather than a fixable defect. Vendors typically advertise false-positive rates below 1%; independent testing on second-language English reports 4–18%. The sources are set out in full in our research on AI-detection bias.
We are not claiming detectors are useless or that anyone is acting in bad faith. The narrower claim is the one that matters for policy: on second-language English, current detectors are wrong often enough that their output must not function as a finding.
Five things the policy needs
1. A detector score opens an inquiry; it never closes one
Write this explicitly, because in practice the number does the deciding unless the document forbids it. State that no percentage, threshold or confidence label constitutes a finding of misconduct on its own. This is the single highest-value sentence in the policy: it costs nothing and it is what stands between a statistical artifact and a disciplinary record.
2. A named burden, carried by the accuser
Say who must establish what. If the policy is silent, the burden silently inverts: the student is asked to prove they wrote it, which is proving a negative, and the students least equipped to do so under pressure are the ones the detector already over-flags.
3. A defined route to respond — and not only an interview
Most policies default to “the student explains their work in a meeting.” That is reasonable-looking and quietly discriminatory. Oral fluency is not evidence of written authorship, and anxiety degrades second-language speech production more than first. The interview therefore disadvantages precisely the group the detector already misidentified — the same person is penalised twice by two supposedly independent checks.
Keep the interview if you want it. Do not make it the only route. Offer at least one alternative that does not depend on real-time verbal performance: drafts, version history, notes, a documented writing process.
4. Process evidence, accepted with stated limits
A process record can show that text was composed over time rather than pasted in one action, and it can be signed so that it is checkable by someone with no account and no relationship to the vendor who issued it. It cannot prove authorship. That limit is not a reason to reject process evidence — it is the reason to write the limit into the policy, so the evidence is weighed rather than treated as either proof or noise.
Two practical requirements: it must be verifiable by you without an account with the issuer, and it must state its own limitations on its face.
5. Keep records and review your own rate
Log how many inquiries were opened, how many resulted in findings, and the first-language distribution of both. If your inquiry rate for one group is several times another's and the finding rate is not, the process is producing false positives and you now have the evidence to fix it. Almost no institution measures this.
A clause you can copy
Adapt freely; no attribution required.
Use of automated detection tools. The output of an automated AI-detection tool may be used as a preliminary signal to open an inquiry. It does not, by itself, constitute evidence of academic misconduct, and no finding under this policy may rest solely or primarily on such output. The burden of establishing misconduct rests with the department.
Right to respond. A student whose work is subject to inquiry will be informed of the basis for it and given a reasonable opportunity to respond. The student may respond in writing, by providing evidence of their writing process, or in a meeting, at the student's election. A student's verbal fluency, or their performance in a meeting, will not be treated as evidence of authorship.
Evidence of writing process. The department will accept documented evidence of a text's composition — including drafts, version history, and cryptographically signed process records that can be independently verified without an account — as relevant evidence of authorship. Such evidence documents an observable writing process; it does not by itself establish authorship, and it will be weighed alongside other evidence.
Review. The department will record the number of inquiries opened, their outcomes, and aggregate demographic data including first language, and will review this record annually.
Three patterns to avoid
A blanket ban on AI assistance. It is unenforceable, and it pushes the work off whatever platform you can see into tools you cannot. A disclosure requirement produces better information than a prohibition everyone quietly ignores.
Requiring students to prove innocence. However it is phrased, any process that begins with a suspicion and demands rebuttal has inverted the burden. Check the actual sequence of steps in your document, not the sentence that says the burden lies with the department.
Quoting a detector percentage in the finding. Once a number appears in a written finding it acquires an authority the underlying method does not support, and it is the first thing an appeal will attack. If the score was only a signal to look, it does not belong in the conclusion.
What this does not solve
It does not tell you whether a particular student cheated. It reduces the rate at which your process produces confident wrong answers, and it gives the accused a route that does not depend on how well they speak under pressure.
It does not remove the need for judgement. It constrains where judgement is applied — to evidence, rather than to a number that was never designed to bear the weight.
This guidance is published so it can be used, adapted, or argued with by people who do not work for us. The model clause may be copied without attribution. It is guidance, not legal advice, and should be reviewed against your institution's own regulations.