Quality & Editing

How to Spot AI-Written Content

A practical guide to spotting AI content in editorial review, covering style tells, structure, fact-checking, and where provenance tools actually fit.

Say a contributor sends you a draft on Tuesday afternoon. Twelve hundred words, on topic, deadline met. Your job is to decide whether it's ready to publish, and somewhere underneath that, whether a person actually thought about what they were writing or just filled space. Spotting AI content isn't a single test you run. It's a read you build up, paragraph by paragraph, until a pattern either shows up or doesn't.

This piece walks through that read using one imagined draft, the kind that lands in an editor's inbox every day. No checklist to tick off in order. Just the questions worth asking, in the order you'd naturally ask them.

Start with a real draft, not a summary of it

Open the file and read the first three paragraphs before you judge anything. Not the headline, not the meta description, the actual prose, because that's where the tells live or don't.

Say the piece is about remote team management. The opening line is something like "In today's fast-paced work environment, managing remote teams effectively has become increasingly important for organizations of all sizes." Nothing there is wrong exactly. It's also nothing. It could open an article about literally any workplace topic from the last five years, which is itself informative: specificity starts at sentence one or it usually doesn't start at all.

Compare that to an opening that names a thing. "Our support team went fully remote in March 2023, and ticket response times got worse for about six weeks before anyone figured out why." That sentence commits to a fact you could ask about. The generic opener commits to nothing, which is exactly the point. A model trained to sound plausible on any topic will default to language that fits every topic, and that's the first thing to notice, before you've even reached a claim you could check.

This is also where you decide how much time the rest of the read deserves. A draft that opens like a template usually keeps going like one. A draft that opens with a specific, checkable detail has earned a slower, more generous read from here on.

Read for the easy tells

The fastest signals are surface ones: word choice and phrasing a model reverts to by default. On their own they prove nothing. In a cluster, they're hard to ignore.

Keep reading the remote-team draft. A few paragraphs in, you hit "it's important to note that communication plays a significant role in team cohesion." Then "leveraging the right tools can help teams stay aligned." Then, two sentences later, "when it comes to building trust, transparency is key." Any one of those could come from a tired human writer on a Friday afternoon. All three inside one paragraph is a different story.

Watch too for the contrast sentence that shows up almost on cue: "it's not about monitoring employees, it's about building trust." That construction reads like insight the first time you see it. By the third time across a body of work, it reads like a habit the model can't shake. A Stanford study of GPTZero found the detector did reasonably well on purely machine-generated text but got shakier separating human essays from AI ones, which lines up with how unreliable these surface patterns are as a solo signal. [source: Stanford SCALE, https://scale.stanford.edu/ai/repository/assessing-gptzeros-accuracy-identifying-ai-vs-human-written-essays]

These tells cluster into a few recognisable groups. Worth having in front of you when you're scanning a draft fast:

  • Filler qualifiers: "it's important to note," "it's worth mentioning," "at the end of the day"
  • Empty verbs: "leverage," "utilise," "facilitate," "drive"
  • Forced contrast lines: "it's not X, it's Y," repeated more than once in a piece
  • Round, hedged claims: numbers or facts stated with no source and suspiciously tidy figures

We've caught this pattern in our own early drafts. When we ran a script across 80 of our own published articles counting these habits, every single one failed at least one check, and the worst piece hit roughly 13 colon-plus-three-item constructions per thousand words. That's not a rare slip. That's a default the model keeps reaching for unless something stops it.

Some defaults are stubborn enough that a style guide won't fix them. We have a hard rule against em dashes in our own output, and the model still ignores it often enough that we stopped asking nicely. We strip them out in code now, after every generation, because a rule that matters needs enforcement, not a reminder in a prompt. If a pattern shows up reliably in a first draft, that's a hint about how the draft was made, not proof of anything by itself. For a longer catalog of these habits and how they show up across genuinely AI-generated writing, see common AI writing tells.

Look for structure that feels manufactured

Step back from individual sentences and look at the shape of the whole piece. The remote-team draft has six sections. Each one runs almost exactly 180 words. Each one opens with a short claim, follows with three bullet points, and closes with a one-line summary. Every section, same template, repeated six times down the page.

That symmetry is unusual for human writing, which tends to be lumpy. A person who's actually managed a remote team will spend four paragraphs on the thing that bit them, hiring across time zones, say, and a single sentence on the thing that never mattered much. The draft in front of you gives equal weight to everything, which is its own kind of tell, because equal weight to everything usually means nobody involved had a strong view about what mattered most.

We've seen the same pattern outside the remote-team example, in submissions on entirely different topics, personal finance explainers, product comparison posts, even a piece on soil drainage for a gardening client. The subject changes. The shape doesn't: three even sections, each with a claim, a bulleted list, and a wrap-up sentence that restates the claim without adding to it. Once you've read a few dozen drafts, that template becomes recognisable almost on sight, independent of what the draft is actually about.

Watch for the pattern where every argument gets rendered as a bulleted list, even ones that are really a single connected thought that deserved three sentences of reasoning instead of three fragments. And watch for the ending that presents both sides of a debate and then declines to land anywhere. "Ultimately, the right approach depends on your team's specific needs and circumstances" is the sound of a piece that ran out of things to say and stopped rather than a piece that reached a conclusion.

None of this proves the draft came from a model. A rushed human writer under deadline pressure can produce the same evenness, the same list-everything instinct, the same hedge at the end. But structural symmetry across an entire piece, section after section built from the same mold, is a strong secondary signal once you've already noticed the surface tells from the last section.

Check whether the facts hold up

This is where the read stops being about style and starts being about risk. Go back through the remote-team draft and pull out every specific claim: names, numbers, sources. There usually aren't many, and that scarcity is itself worth noting.

The draft says "studies show that remote teams with clear communication protocols see productivity gains of up to 25%." No study named, no source linked, a number that's suspiciously round and suspiciously convenient. Here's how you'd actually check it. Search the exact figure plus "remote team productivity" and see what surfaces. If nothing traces back to an actual paper or dataset, the number is unverifiable as stated, and the honest move is either to find the real source and cite it properly or to cut the claim entirely and replace it with something the writer can actually stand behind, even something softer like "teams we spoke with reported fewer missed handoffs after they wrote down a communication protocol."

The draft also says "many companies have found that async-first communication reduces meeting fatigue," with no company named. Claims like these are built to sound authoritative while giving you nothing to check. That's the pattern to catch: not that the claim is necessarily false, but that it evaporates the moment you go looking for where it came from.

This softness is also where the real cost sits if you publish without checking. A structural tell or a clunky phrase embarrasses nobody. A fabricated statistic that gets shared and cited does actual damage. We saw a version of this ourselves: our own SEO analysis tool used to confidently report that articles contained tables and FAQ sections that simply weren't there. It wasn't lying exactly, it was grading from vibes, generating a plausible-sounding assessment rather than checking the actual document. The fix wasn't a better prompt asking it to be more careful. It was giving the model ground truth: a script that counts headings, tables, images, and lists with regex, then hands the model those numbers as facts it isn't allowed to contradict. The hallucinated tables stopped that day, because the tool could no longer guess when it could just look.

A detector score or a fact-check pass isn't the end of the process either. It's an input into whatever editorial policy your team actually enforces. If a claim can't be verified and the writer can't point to where it came from, that's a manual review decision, not something a tool resolves for you. The same discipline applies to any draft you're evaluating. Take the load-bearing claims, the ones the argument depends on, and verify them against something real before you publish. For a fuller workflow on how to run that verification pass systematically, see fact-checking AI-assisted content.

Separate origin from editorial quality

Here's the reframe that should shape everything above it. You don't actually care whether a model wrote the first draft of this piece. You care whether the version in front of you is accurate, specific, and worth a reader's time. Those are different questions, and conflating them leads to bad calls in both directions.

What editing actually fixes. A draft that started with a model, then got fact-checked against real sources, rewritten with a specific example from someone who's actually managed a distributed team, and given an actual point of view instead of a shrug, is publishable work. The origin stops mattering the moment the editing happens.

Take the round-trip on one paragraph from the remote-team draft. Before: "It's important to note that communication plays a significant role in team cohesion, and leveraging the right tools can help teams stay aligned." After: "Our support team missed three handoffs in the first month of going remote, all because nobody had agreed on which channel was for urgent issues versus general updates. We fixed it with one rule: anything blocking a customer goes in the #urgent channel, nothing else does." Same underlying point. One version is filler. The other is a claim a reader can picture and, if they wanted to, ask about.

Meanwhile a draft a person typed from scratch, never checked, full of vague claims and no lived detail, has the exact same problems as the unedited AI draft and deserves the exact same rejection. So the tells in the last three sections aren't a verdict on where the words came from. They're a proxy for whether a human editorial pass actually happened.

Where provenance tools fit. Detector tools try to answer the origin question directly, and they're worth understanding precisely because they're weaker than most people assume. Some publishers are also starting to rely on provenance metadata, content credentials embedded at the point of generation or capture, and on watermarking schemes that mark text or image output at the model level. Both are still early and neither is a substitute for reading the draft. A watermark tells you something about where content originated, not whether it's any good, and metadata can be stripped or never attached in the first place.

Where policy has to fill the gap. Because detectors and watermarks are unreliable, most of the actual decision-making has to sit in editorial policy, not tooling. That means having an answer, in writing, to questions like: does AI-assisted work get a disclosure line, does a borderline draft go to a second editor before it runs, and what score or pattern triggers escalation rather than a straight accept or reject. A newsroom that publishes a disclosure note whenever a draft started with a model is making a different bet than one that stays silent on origin and judges every piece purely on the finished quality. Neither is obviously wrong, but a team that hasn't decided which bet it's making will end up litigating the same argument draft by draft, which is slower and less fair than picking a rule once.

What the detector evidence actually says. OpenAI has said openly that its own AI-text classifier isn't reliable enough to use as a primary decision-making tool, particularly on shorter passages, and that it can mislabel genuine human writing as machine-generated. [source: OpenAI, https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/] Turnitin, one of the more widely used detectors in education, changed its own scoring in 2024 so that anything under 20% no longer surfaces a confident score at all, specifically to cut down on false positives. [source: Turnitin, https://guides.turnitin.com/hc/en-us/articles/28294949544717-AI-writing-detection-model] Researchers at Temple University tested whether people can reliably tell AI and human writing apart by eye and found the task harder than most readers assume it is. [source: Temple University, https://liberalarts.temple.edu/news/2025/03/control-adaptive-behavior-lab-conducts-study-human-detection-ai-written-content] And a University of Chicago working paper found that detector accuracy drops sharply once you require a genuinely low false-positive rate, with only one tool in their comparison holding up under that stricter bar. [source: University of Chicago, https://bfi.uchicago.edu/working-papers/artificial-writing-and-automated-detection/]

What this means for a workflow. Put a detector score in the same category as one paragraph with too many hedges: a data point, never a verdict. Read for the four kinds of signal, style, structure, facts, and lived specificity, and use them to find drafts that skipped the human pass. Then judge and fix the actual work, which is what editorial quality has always meant. For the broader case on why the human pass is the thing worth protecting, and how to run it as a repeatable step rather than a one-off save, see editing AI content.

Less work, more on-brand content

Austen runs this whole workflow for you: from research to on-brand drafts that get found by Google and AI.

Start free

More in Quality & Editing