Why AI Detection Is Harder Than It Sounds
You paste a suspicious article into an AI detector. It says “92% AI-generated.” Do you believe it? Do you act on it? And what happens when the tool is wrong?

Here’s the thing nobody in the AI detection space wants to admit: these tools are probabilistic guesses, not truth machines. I learned this the hard way when a colleague flagged one of my own technical drafts as “likely AI” — a 3,000-word piece I’d struggled through line by line. The detector didn’t know I was the author. It just saw patterns.
That’s the gap between how AI detectors market themselves and how they actually behave. Pangram, one of the best-known startups in the space covered by TechCrunch in early September, just raised $9 million and landed a partnership with Substack to flag AI-assisted writing. The company’s CEO Max Spero went on TechCrunch’s Equity podcast — around the same time OpenAI was rolling out its new reasoning technique that alarmed safety experts in early September 2026 and talked through why detection is fundamentally harder than a “real or fake” binary. The internet has a trust problem — like I covered in securing your AI deployments, and AI-generated text is now showing up in job applications, product reviews, insurance claims — everywhere. But the tools we’ve built to catch it are imperfect by nature.
This guide isn’t about declaring any tool a magic bullet. It’s about what you can actually do with AI detection tools right now: how to use them, when to trust them, when to ignore them, and how to avoid the false-positive nightmare that ruins people’s reputations — the same kind of AI accountability issue behind OpenAI’s Tumbler Ridge lawsuits.
What AI Detectors Are Actually Measuring
AI detectors don’t “know” whether text was written by a human or a machine. They look for statistical fingerprints — patterns in word choice, sentence structure, and predictability that tend to differ between AI output and human writing.
The most common signal is perplexity — how surprised a model would be by the next word in a sequence. AI-generated text tends to be more predictable because the model is sampling from a probability distribution that favors likely continuations. Human writing, especially from non-native speakers or people writing casually, is often messier and less predictable. Another signal is burstiness — the variation in sentence length and structure. AI text can be more uniform; human text often swings between short punches and longer winding sentences.
These signals work best on clean, unedited AI output. They fall apart the moment someone — or something — touches the text. A 2026 study in Computational Linguistics found that detection accuracy dropped to 60–80% once a human manually edited or “humanized” AI-generated text. Independent testing by researchers comparing Turnitin, GPTZero, Copyleaks, and Pangram found enormous differences between tools, with some substantially underestimating AI content from newer models and others flagging human writing as AI. The study’s authors put it bluntly: their findings “do not confirm the claims presented by the systems.”
That’s not a reason to dismiss detection entirely — it’s a reason to treat every result as a signal, not a verdict.
The Tools Worth Knowing About
If you need to check whether something was AI-generated, you have options. None are perfect. Here’s what’s actually useful in 2026.
Pangram
Pangram is the tool getting the most attention right now. Pangram’s own detector claims over 99% accuracy and says it’s been independently verified. It claims over 99% accuracy and says it’s been independently verified by researchers at the University of Chicago and the University of Maryland. The company’s false positive rate is reported at 1 in 10,000 — but only on the specific benchmarks they publish, and a study they cite found higher false positive rates under real-world conditions. Pangram also recently launched an AI image detection tool — part of a broader wave of AI security funding that includes HiddenLayer’s $100 million bet on AI security alongside its text detector.
What makes Pangram interesting: it was built by people who explicitly frame detection as harder than a binary classification problem. Max Spero, the co-founder and CEO, has argued publicly that the real challenge isn’t building a detector that works on clean AI text — it’s handling the mixed-authorship gray zone where most real-world cases actually live.
The free tier has daily limits on tokens. The paid tiers are meant for platforms and high-volume users — which is who Pangram’s Substack partnership targets.
GPTZero
GPTZero was one of the first widely-used AI detectors, launched in early 2023 by a college student. It’s still around and still free for basic checks. It looks at perplexity and burstiness like most detectors, and it shows you a breakdown of which sentences it considers most AI-like.
The catch: independent research has consistently found GPTZero’s accuracy drops significantly on edited AI text, and its false positive rate on human writing — especially writing from non-native English speakers — has been a persistent criticism since launch. Treat its sentence-level highlighting as a hint, not proof.
Originality.ai
Originality.ai is built for publishers and content platforms rather than individual checkers. It combines AI detection with plagiarism checking, which matters if you’re running a content site and need both signals at once. Independent testing in 2026 found overall accuracy around 69% — better than some competitors on raw AI text, but still far from the “99% accurate” marketing line.
Turnitin
Turnitin is the tool your professor probably used if you were in college during the AI boom. It’s the most widely deployed detector in academia, baked directly into learning management systems. But a 2026 comparative study found Turnitin’s overall accuracy at only 61% — meaning it gets the answer wrong roughly four times out of ten on challenging cases. The tool also struggles specifically with text that combines human and AI writing, which is exactly the kind of content most people actually submit.
How to Actually Use These Tools Without Making Mistakes
Here’s a practical workflow that doesn’t assume the detector is always right.
Start with context, not the score
Before you paste anything into a detector, know what you’re checking and why. A marketing blog post, a student essay, a job application cover letter, and a product review are all different problems. The consequences of a false positive range from annoying (asking someone to rewrite something) to devastating (a student accused of cheating, a job applicant rejected, a freelance writer’s contract terminated).
Ask yourself: what’s the cost of being wrong in this specific case? If the cost is high, the detector’s score is a starting point for conversation, not a decision. If the cost is low — you’re just curious about an unknown blog’s authorship — you can afford to be more casual about it.
Never trust a single tool
Run the same text through at least two detectors if the stakes are moderate or higher. Different tools use different training data and different models, so they won’t always agree. When they disagree — and they often will — that disagreement is itself useful information. It means the text lives in the gray zone where detection is unreliable, and you should treat any single tool’s score with skepticism.
A 2026 study comparing Turnitin, GPTZero, Copyleaks, and Pangram found “enormous differences” between tools. That’s not a flaw in any one tool — it’s a property of the problem. Different detectors are looking for different signals in different ways.
Look at the flagged text, not just the score
Most detectors will show you which sentences or passages triggered the AI flag. Read them. If the flagged sections are the parts that genuinely read as generic, formulaic, or oddly flat — the kind of writing that feels AI-ish to a human reader — the detector might be picking up on a real signal. If the flagged sections are specific, opinionated, or idiosyncratic passages that you wrote yourself, the detector is probably wrong and you should weight its overall score lower.
This is the single most practical habit you can build: treat the detector’s output as a pointer to specific text, then apply human judgment to that text. The detector tells you where to look. You decide what you see.
Beware of edited AI text
The hardest case isn’t clean AI output — it’s AI output that someone has edited. A 2026 study in Computational Linguistics found that detection accuracy dropped substantially on text that had been manually revised. Another 2026 study found that mixed human-AI writing was the hardest category for detectors to handle, with both Turnitin and Originality.ai struggling specifically on that type of content.
If someone took an AI draft and rewrote half the sentences, changed the vocabulary, and added personal details — the most likely real-world scenario for AI-assisted writing — most detectors will either miss it entirely or flag it with low confidence. That’s not a detector failure in the sense of a broken tool. It’s a genuine limit of what statistical signals can detect when the signal has been deliberately diluted.
Don’t assume human-written text is always clean
Here’s the false-positive problem in practice: detectors get worse on text written by non-native English speakers, people who write in a formal or formulaic style by habit, and text that’s been through heavy editing. A 2026 study tracking detector reliability found that false positives were a persistent problem across tools, especially on writing that deviated from the “standard” human writing style the detectors were trained on.
If you’re using a detector to screen job applications, grade student work, or make any decision that affects someone’s opportunities, the false positive rate matters more than the true positive rate. A detector that catches 90% of AI text but falsely flags 5% of human text will ruin a lot of innocent people’s days at scale.
When Detectors Are Useful — and When They’re Not
Useful for: spotting obviously AI-generated spam, scraping, and low-effort content at scale. If you run a comment section, a review platform, or a user-generated content site and you’re seeing floods of generic text, a detector can help triage. It won’t catch everything, but it can flag stuff worth a human look.
Useful for: starting a conversation. If you suspect someone used AI in a context where that matters — a student paper, a freelance deliverable, an application — a detector result is a prompt to ask questions, not a conclusion to announce. “I ran this through a detector and it came back high — can you walk me through how you wrote this?” is a reasonable approach. “This says you used AI, prove you didn’t” is not.
Not useful for: making final decisions with serious consequences. No detector in 2026 is reliable enough to be the sole basis for an academic integrity finding, a hiring decision, or a legal argument. The research consensus is consistent on this point: detection is probabilistic, adversarial conditions exist, and false positives are real.
Not useful for: edited or mixed-authorship text in the gray zone. If someone used AI as a brainstorming aid or a first draft and then rewrote it substantially, detection tools will either miss it or flag it unreliably. There’s no statistical signal strong enough to distinguish “AI-assisted but heavily human-revised” from “human-written with a formal style” in many cases. Anyone claiming otherwise is overselling.
The Trust Problem Is Bigger Than Detection
Max Spero’s framing on the TechCrunch Equity podcast — that AI detection is “harder than real or fake” — points at something deeper than detector accuracy. The real problem isn’t that we need better tools to tell human writing from machine writing. The real problem is that the distinction itself is getting less meaningful.
A writer who uses AI to outline, draft sections, revise prose, and fact-check — and then reviews and edits the whole thing critically — has produced something that is neither purely human nor purely AI. It’s a hybrid. Existing detectors weren’t built for hybrids, and the research suggests they can’t reliably distinguish them from clean human writing.
That’s not an argument for abandoning detection. It’s an argument for being honest about what detection can and can’t do, and for building policies — in classrooms, newsrooms, hiring processes, and platform guidelines — that account for the tool’s limits instead of pretending they don’t exist.
Pangram’s $9 million raise and the Substack partnership suggest the market believes detection is worth investing in. The research suggests the problem is harder than the marketing suggests. Both can be true at once.
A Practical Checklist
If you’re about to use an AI text detector on something that matters, run through this before you act on the result:
- What are the consequences if this result is wrong? If the answer is “someone gets hurt,” slow down.
- Have I run this through more than one tool? Disagreement between tools is a signal to pause, not a tie-breaker.
- Did I read the flagged text myself, or just look at the percentage? The specific passages matter more than the score.
- Is this text the kind that trips detectors up — non-native English, formal style, heavy editing? If so, weight the result lower.
- Am I using this to start a conversation or to make a final call? A good AI toolchain audit — like the supply chain risk checklist I walked through before — starts with knowing what you’re actually protecting The tool is better suited to the first one.
- Do I have other evidence — process records, drafts, timelines — that I can check alongside the detector result?
The detector is one data point. Treat it that way and you’ll avoid most of the worst outcomes. Treat it as a verdict and you’ll eventually get burned.
Bottom Line
AI text detection in 2026 is a genuinely useful screening tool with hard, well-documented limits. The best detectors are good at catching clean, unedited AI output. They’re unreliable on edited text, mixed-authorship content, and writing that deviates from the patterns they were trained on. False positives are real and disproportionately affect certain groups of writers.
Use detectors to flag stuff worth a closer look. Don’t use them as the final word. Read the flagged text yourself. Ask questions before drawing conclusions. And remember that the trust problem AI detection is trying to solve — “what did a human actually write here?” — isn’t going to be solved by a better classifier alone. It’s going to take a combination of better tools, honest policies, and human judgment that knows when not to trust the machine.