No. 18 / 339
Is human threat-intelligence analysis still worth doing manually, or is AI synthesis good enough now?
The shift
Reading, correlating, and summarizing the raw material of threat intelligence — malware writeups, leak-site posts, actor forum chatter, sandbox output, prior incident reports, OSINT across languages — goes from scarce analyst-hours spent scanning feeds to abundant, near-instant synthesis across far more sources than any team could cover by hand.
The axioms
- An analyst has to read the raw feed (forums, leak sites, malware reports, vendor blogs, sandbox output) because nobody's automated understanding what it means. Rests on scarce hours to read and correlate volume.
- Attribution — is this APT29 or someone borrowing their tools — requires an analyst who's tracked the actor's TTPs over time. Rests on scarce pattern-matching against a large, fragmented, mostly-tacit body of prior cases.
- Translating and making sense of foreign-language chatter (Russian forums, Chinese-language leak sites) requires a linguist-analyst hybrid. Rests on scarce language-plus-domain expertise.
- A finished intelligence report has to say what it means for this specific organization — which of our systems, which of our third parties, how urgent. Rests on judgment applied to organizational context a feed can't supply on its own.
- Confidence levels on a judgment (low/moderate/high confidence this is nation-state) carry consequences — a board decision, a customer notification, an insurance claim — so someone has to own that call. Rests on scarce accountability.
- Sources lie, seed false flags, and get planted by the actors being tracked, so raw material needs a skeptical human filter before it becomes a finding. Rests on judgment under adversarial, low-trust conditions.
- Analysts build a mental model of "who's targeting us and why" over years of tracking the same actors and campaigns. Rests on scarce time to accumulate tacit, organization-specific context.
Invalid axioms
- An analyst has to manually read and correlate the raw feed to build a picture. Ingesting hundreds of reports, forum posts, and sandbox runs and surfacing what's new, what's repeated, and what connects to a prior campaign is exactly the large-volume synthesis LLMs do cheaply and fast, in any language, without fatigue. The habit-trap: CTI teams still size themselves and budget hours as if reading the feed is the job, rather than treating first-pass correlation as a solved, near-zero-cost step and reallocating the freed hours to checking and deciding.
- Language coverage requires a linguist-analyst hybrid for each region. Current models translate and contextualize foreign-language forum and leak-site content competently enough that most first-pass triage no longer needs a dedicated native speaker on staff for every language of interest. The habit-trap: teams still hire and structure regional coverage around scarce bilingual analysts instead of using them for the judgment calls translation alone can't resolve — tone, slang shifts, deception in a specific cultural register.
- Writing the finished report is the bulk of the analyst's time. Drafting the structured writeup — background, TTPs, indicators, a first-pass confidence statement — is now a fast draft from the correlated material, not hours of composition. The habit-trap: teams still schedule report production as the long pole, instead of treating drafting as instant and the review-and-sign-off step as the actual bottleneck.
Unchanged axioms
- Attribution and confidence-level judgment on an adversarial, deceptive body of evidence. Actors plant false flags, reuse other groups' tooling on purpose, and seed disinformation into the exact forums analysts scrape. A model pattern-matches against what's been written before; it has no way to independently verify that a claimed leak is real or that a TTP overlap is deliberate misdirection rather than genuine reuse. That skepticism, applied to sources actively trying to deceive the reader, stays a human judgment call.
- What this means for us, specifically. A model can summarize a campaign; it can't reliably weigh it against this org's actual attack surface, its risk appetite, its regulatory exposure, or what the board will do with the finding — that requires organizational context no feed carries and no general-purpose model has ground truth on.
- Someone is accountable for the confidence level attached to a finding. "High confidence nation-state activity" in a report that triggers a customer notification, an insurance claim, or a board briefing needs a named analyst who can be asked to defend the call and who is answerable if it's wrong. A model can't be cross-examined or held liable.
- Trust with the consumers of the intelligence. Executives, incident responders, and law enforcement acting on a CTI finding need to be able to press the analyst who produced it, ask follow-up questions in real time, and get a commitment they can act on. That standing doesn't transfer to a system regardless of how good the underlying synthesis was.
New axioms
- Who verifies AI-synthesized intelligence at the volume it now gets produced. When correlation and first-draft writeups are free, the bottleneck moves to checking that synthesis for confidently-wrong attribution or a hallucinated indicator link — and CTI teams haven't sized the review capacity that requires, especially when the volume of raw material being processed has also gone up an order of magnitude.
- Where does attribution judgment come from when nobody spends years reading raw feed by hand. The tacit sense of "this actor's fingerprint" used to come from grinding through hundreds of cases manually. If that grinding gets automated away, it's unclear how the next generation of analysts builds the pattern library that senior attribution judgment depends on.
- Adversaries feeding the synthesis engine deliberately. Once threat actors know defenders' CTI pipelines lean on AI summarization of open-source and forum content, seeding poisoned chatter or fabricated leak-site posts to shape a model's output becomes a live tactic — an adversarial-input problem that a bored human skimming the same forum was less systematically exposed to.
- Cross-team overconfidence in a single synthesized narrative. When AI can produce a polished, internally consistent report from partial or contaminated sources at speed, the report reads more authoritative than the underlying evidence supports, and downstream consumers (execs, insurers, regulators) have no easy way to tell a well-sourced finding from a fluent guess.
Where it breaks
CTI teams are already treating feed-reading and report-drafting as solved and shrinking headcount there (invalid), while the volume of raw material a team now tries to cover has grown to match — nobody's built the verification layer to catch the model's confidently-wrong attribution calls at that new scale (new). The team cuts the very capacity that used to slow-walk a claim before it became a finding.
Analysts who used to build attribution instincts by grinding through raw cases by hand (the apprenticeship the field still depends on for senior judgment) are being redirected toward "higher-value" review-only work as soon as synthesis is automated (invalid habit of treating reading-the-feed as pure grunt work) — which quietly removes the training ground the next generation of attribution experts needs, right as adversaries start deliberately poisoning the sources that same synthesis leans on (new).
Related axioms
Cybersecurity
What changes for cybersecurity with AI?
Cybersecurity
Who's liable for a breach an AI security agent missed or misclassified?
Cybersecurity
How does a CISO's risk calculus change when both attackers and defenders run autonomous AI agents?
Cybersecurity
What happens to junior security hiring when AI eats the entry-level triage rung?
Cybersecurity
Does human penetration testing still matter when AI can run continuous automated red-teaming?
Cybersecurity
Is the Tier-1 SOC analyst job already gone now that AI triages the alert queue?
Other axioms
Engineering
Does platform engineering still need a human-designed golden path if agents can self-serve infra on demand?
Education
What's the point of a take-home essay now that neither writing nor detecting AI writing is reliable?
Engineering
What changes for data engineering with AI?
Engineering
Is the ML engineer role converging with software engineer now that AI handles most model-building grunt work?
Healthcare
If ML triage out-performs the call-taker's protocol, does accountability for the dispatch decision move to the model — and who's liable for the under-triage that kills?
Healthcare
What changes for pharmacy with AI?