No. 22 / 339

What's left for a TA to do when AI can hold office hours, explain concepts, and grade problem sets?

The shift

Explanation-on-demand and first-pass problem-set grading — patient, repeatable, available at 2am, tuned to any level — go from scarce (a TA's limited hours) to abundant (instant, free, unlimited). What doesn't move: verifying a student actually understands versus can produce the right-looking answer, and being the accountable, trusted human attached to a grade or a flagged concern.

The axioms

  • A TA is the accessible tier of explanation between lecture and self-study — scarce faculty time made available at a lower cost.
  • Grading a problem set requires a competent evaluator checking work against a rubric — scarce evaluation labor.
  • Office hours exist because students need on-demand, repeatable answers to common confusions — scarce availability at the moment of stuckness.
  • A TA is a junior apprentice-professor, learning pedagogy by explaining material out loud under supervision — a scarce training slot in the academic pipeline.
  • A TA knows this specific class — this week's common mistakes, this professor's grading quirks, this cohort's gaps — scarce situated knowledge that generic explanation doesn't have.
  • A TA is who a struggling student trusts enough to admit "I don't get this" or "I'm behind" to, and who vouches for that student's effort or integrity to the professor — scarce trust and standing.
  • A TA notices a student sliding (skipped problem sets, sudden drop in quality, a question that reveals more than confusion) before the student self-reports — scarce judgment on a low-signal, high-ambiguity case.
  • A TA's grade and academic-integrity flag carry institutional weight — someone is accountable when a grade is wrong or a cheating case is raised — scarce accountability.

Invalid axioms

  1. A TA is the accessible tier of explanation between lecture and self-study. Explanation-on-demand, at any pace and any level of remediation, is now free and instant — a model doesn't get tired of the fourth "wait, why though" of the night. The habit-trap: departments still staff and schedule office hours as if patient one-on-one explanation were the scarce resource, when for most conceptual questions it no longer is.
  2. Grading a problem set requires a competent evaluator checking work against a rubric. First-pass grading against a rubric — did the student get the right answer, did they show the expected steps — is exactly the pattern-matching-at-scale task AI is strong at, and it's abundant now. The habit-trap: paying TA hours by the stack of problem sets graded, as if throughput on mechanical grading were still the bottleneck rather than a solved problem.
  3. A TA is a junior apprentice-professor, learning pedagogy by explaining material out loud under supervision. Some of that training value used to piggyback on volume — you got good at explaining by explaining a lot, to a lot of confused students. If AI absorbs most of the repetitive explaining, the training rep count for TAs drops with it. The habit-trap: assuming TA-ing still builds the same teaching muscle it used to just because the title and the hours logged look the same.

Unchanged axioms

  1. A TA is who a struggling student trusts enough to admit confusion or fall-behind to, and who vouches for them to the professor. A model can explain a concept faster than any TA, but it has no standing to write a note to the professor saying "this student is drowning, cut them a break" — that requires a human who has actually watched the student and has something to lose by vouching wrongly. Trust and institutional standing aren't things AI generates by being helpful.
  2. A TA notices a student sliding before the student self-reports. This is judgment on a genuinely ambiguous, low-signal, high-stakes case — is this student behind because of the material, home life, mental health, or checked out — with no clean pattern to match against. AI sees the artifacts a student submits to it; a TA sees the student not show up, sees the tone change, sees the pattern across a section that a per-student chat interface never assembles.
  3. A TA's grade and academic-integrity flag carry institutional weight. Someone has to be answerable when a grade is contested, an accommodation is misapplied, or an integrity case is raised — and that has to be a person embedded in the institution's process, not a model. This scarcity didn't move; if anything AI-generated submissions make the integrity-flag judgment call harder and more load-bearing, not less.

New axioms

  1. When a student can get a flawless AI explanation and near-instant AI grading feedback on every problem set, what's the TA verifying? The volume of AI-assisted or AI-generated submissions rises faster than any human's ability to tell "this student understands it" from "this student's tool understands it." The TA's job shifts from producing feedback to auditing whether the work in front of them reflects actual understanding — a much harder, less scalable task than grading was.
  2. If office hours attendance drops because AI answers the routine questions, who's left to notice the students who need a human and won't ask an AI for help with the real problem (falling behind, checked out, in crisis)? The students who still show up to office hours skew toward exactly the ambiguous, high-stakes cases a TA needs judgment for — but there are fewer low-stakes interactions left to build the pattern-recognition and rapport that made noticing possible in the first place.

Where it breaks

Departments keep paying and scheduling TAs against grading throughput and office-hours coverage (invalid) right as the actual bottleneck becomes verifying whether a submission reflects real understanding at a volume no rubric-checking workflow was built for (new) — the TA hours saved on mechanical grading aren't being redirected to the harder verification problem, they're just being cut.

Faculty treat a TA's light office-hours traffic as a sign the AI is doing its job (invalid habit: read low attendance as success) while the students who still show up are disproportionately the ambiguous, high-stakes cases that need the most judgment (new problem: less practice reading students, right when reading them matters more) — the TA's judgment is atrophying at the exact moment the remaining caseload requires more of it.

Related axioms

Other axioms