No. 23 / 339
Is the take-home essay dead as an assessment format now that AI authorship can't be reliably detected?
The shift
Generating a fluent, structured, on-topic essay is now abundant — near-zero cost, seconds of latency, and good enough to pass as competent undergraduate work most of the time. Detecting whether a specific essay came from a model is not abundant in the same way: classifiers stay probabilistic, easy to evade with light editing, and biased against non-native and neurodivergent writers, so the take-home essay's grading model (assume solo authorship, verify occasionally) breaks even though the writing itself is more available than ever.
The axioms
- Producing a competent essay demonstrates competent thinking. Rests on writing being expensive enough that only someone who understood the material could produce fluent prose about it.
- A single unsupervised written artifact is sufficient evidence of a student's understanding. Rests on authorship being unambiguous — only the student had the means to write it.
- Plagiarism/authorship detection tools can adjudicate disputes. Rests on there being a detectable statistical signature that separates machine text from human text.
- The essay is the most efficient way to make a student practice synthesis and argument construction. Rests on essays being the cheapest available proxy task for that skill, given grading capacity constraints.
- Grading a stack of essays for logic, evidence use, and voice requires a human reader's judgment. Rests on evaluation being a scarce, expensive act only a trained reader can do at scale.
- A degree signals that the holder can produce this kind of work unaided. Rests on the credential being anchored to individually-authored artifacts as evidence.
Invalid axioms
- A single unsupervised written artifact is sufficient evidence of a student's understanding. The scarcity that made this true — that only the student had the means and motive to produce fluent argumentative prose on the topic — is gone; anyone with model access can produce it in seconds. The habit-trap: departments keep weighting the take-home essay as a major grade component as if turning it in still proves anything about the person who submitted it, rather than treating it as unverified until some other signal corroborates it.
- Plagiarism/authorship detection tools can adjudicate disputes. AI text detectors were never reliable and current-generation models (paraphrased, custom-instructed, or run through a "make this sound like me" pass) evade them further while flagging genuine human writing, especially from non-native speakers, at unacceptable false-positive rates. The habit-trap: institutions keep running essays through detection software and treating a score as evidence in academic-integrity hearings, exposing the institution to real liability for false accusations built on a tool that doesn't do what it claims.
- The essay is the most efficient way to make a student practice synthesis and argument construction. First-draft synthesis is now the cheapest thing a model does; assigning "write 1500 words at home, unsupervised" as the entire task wastes the scarce classroom time that could go to the parts AI can't do — debating the draft, defending a claim under questioning, revising against pushback. The habit-trap: syllabi keep the same take-home-essay-as-final-deliverable structure instead of moving the graded moment to where verification is still possible.
Unchanged axioms
- Grading a stack of essays for logic, evidence use, and voice requires a human reader's judgment on borderline cases. Models can draft and even pre-screen rubric criteria competently, but the stakes-bearing call — is this argument actually sound, is this a case of grade inflation, does this student deserve the mark that shapes their transcript — still needs a person who's accountable for the grade. AI grading assistance is real and improving fast, but institutions are still unwilling (correctly, for now) to let a model be the final accountable grader for high-stakes work.
- A degree signals that the holder can produce this kind of work unaided, and that signal still needs to be earned somewhere. The take-home essay as a single checkpoint is compromised, but the underlying thing employers and grad programs actually want verified — can this person reason and argue under their own steam — hasn't stopped mattering. It just can't be verified by that format alone anymore; it has to be verified live (oral defense, in-class writing, timed exams) or over a longer trail (drafts, version history, working sessions), and that verification still requires scarce instructor time and physical presence.
- Learning to write is learning to think, and that requires the student to do the effortful part themselves. No model change flips this — struggling with a blank page and a half-formed argument is still how the skill gets built, and a model that removes the struggle removes the learning. This isn't nostalgia; it's the one axiom in the set that has nothing to do with detection and everything to do with what practice is for.
New axioms
- How do you verify authorship without reverting to surveillance-heavy, resource-heavy formats (proctored writing, oral exams for every assignment) that don't scale to large courses? The abundance of AI-written essays forces a verification cost that didn't exist before, and most institutions don't have the TA hours or exam-room capacity to absorb it at the same throughput as take-home grading.
- What happens to the students who used AI as a legitimate drafting or scaffolding tool once the format shifts to distrust-by-default? Treating all take-home writing as suspect punishes exactly the students who used the tool the way a study-skills center would recommend, and there's no settled norm yet for what "acceptable AI use in coursework" means, assignment by assignment.
- Who is liable when a false-positive AI-detection accusation damages a student's record? This is a new institutional exposure — academic integrity processes built for a world with reliable evidence are now running with unreliable evidence at scale, and the accountability question (registrar, professor, vendor of the detection tool) hasn't been tested.
Where it breaks
Departments are simultaneously declaring the take-home essay dead as a trustworthy artifact (dropping its grade weight, adding integrity disclaimers) and still running it through detection software to adjudicate cheating cases — using a tool everyone privately distrusts to make decisions with real consequences (failing grades, expulsion referrals) for the exact new problem (false-positive liability) that same distrust created. The two moves can't coexist: either the essay is evidence or it isn't, and right now policy treats it as both at once, in different parts of the same institution.
Related axioms
Education
What changes for higher education with AI?
Education
What changes for K-12 education with AI?
Education
What changes for vocational education with AI?
Education
Can admissions essays still signal anything now that AI can write a plausible, polished one for any applicant?
Education
What's the business model for a degree when the credential's signal value is exactly what AI undermines?
Education
What's left for a TA to do when AI can hold office hours, explain concepts, and grade problem sets?
Other axioms
Engineering
What changes for ML engineering with AI?
Healthcare
Is the radiologist obsolete now that AI reads scans as well as humans, or did the job just move to accountability?
Engineering
Is code review dead now that AI writes most of the diff, or did it just move upstream to spec/plan review?
Media
What changes for fine art, galleries, and auction houses with AI?
Research
Is peer review still meaningful when a meaningful share of reviews at major venues are already AI-generated?
Healthcare
Who's accountable when a semi-autonomous surgical robot, guided by AI, is involved in a bad outcome?