A year ago, dodging AI detection meant cutting a few words from your draft — no "delve," no "unlock your potential." That checklist stopped working months ago. Detectors improved. Models got trained away from the obvious tells. Readers got sharper too. What actually gives AI writing away today is a quieter set of patterns — in rhythm, structure, and the specific words a model reaches for when it wants to sound meaningful without saying much. Here's the current list, not the one everyone copy-pasted from 2023.
Why Sounding Like AI Costs You More Than a Detector Score
Getting flagged by a detector is the visible cost. The bigger one happens before any tool runs: a reader notices the rhythm feels off, the examples feel generic, and they leave. Recent industry surveys keep landing in the same place — most people say they don't fully trust content they suspect was AI-written, and once suspicion kicks in, they read the rest of the page hunting for more tells instead of absorbing what you're actually saying.
Search engines have started rewarding the opposite. Google's ranking guidance increasingly favors what it calls information gain — content that adds something a reader couldn't already predict, rather than reshuffling the same five points everyone else's AI-assisted draft also produced. Two competitors using the same prompt on the same topic end up with near-identical articles, which defeats the purpose of writing one in the first place.
None of this is about gatekeeping AI use. It's about the gap between a draft and a piece someone actually finished — and readers can tell the difference within the first two sentences, whether or not they could explain why.
The Checklists Everyone's Using Are Already Out of Date
If your idea of an AI checklist is scanning for "delve," "landscape," "unlock," and "elevate," you're checking for tells that mostly stopped being reliable a while back. Those words got called out so widely that most models were trained away from overusing them, and most writers who edit AI drafts now strip them on the first pass anyway. Checking for them still costs nothing — keep that pass in — but it's necessary, not sufficient.
What's replaced them is subtler: plain, ordinary words used as vague intensifiers instead of doing real descriptive work. Think "quietly," "shape," "matters," "hold," "pull," "signal," "the work," "built different." None of these words are wrong on their own. The tell is a sentence that leans on one of them to sound meaningful without committing to anything specific — "the results quietly reshape how the team works" says almost nothing next to "the results cut onboarding time from three weeks to four days."
For a deeper look at the mechanics detectors actually use — perplexity, burstiness, and how those get scored — see our full breakdown of perplexity and burstiness. Keep the old buzzword list as a first pass if you want; it costs nothing to check, but treat it as a filter for the obvious cases, not the whole test.
The 12-Point Checklist
Run a draft against these in order. The first two are the fastest to spot without close reading; the rest take a slightly closer pass.
- Every sentence is roughly the same length.Human writing naturally varies — a punchy four-word sentence next to a much longer one that unpacks a detail. Flat, even rhythm across a paragraph (technically, low "burstiness") is still the single most reliable manual tell.
- Em dashes show up in nearly every other sentence.One or two per paragraph is normal punctuation. When they replace commas, periods, and colons entirely — in more than roughly a third of sentences — it's a strong signal, though never conclusive on its own.
- Nothing on the page contradicts itself, even slightly.Real writing forgets a detail mentioned two paragraphs back, or complicates something stated earlier. Generated drafts tend toward suspiciously tidy coherence — every example fits, every thread resolves, nothing is left hanging.
- There's no accidental detail.No throwaway number, no oddly specific but unimportant fact, no aside that doesn't serve the argument. Human writing leaks small, irrelevant true things. Generated text usually doesn't, unless it's specifically prompted to.
- A sentence reframes itself as "not X, it's Y.""It's not about working harder — it's about working smarter" is the single most recognizable AI sentence shape in circulation right now. One instance is a stylistic choice. Three in one piece is a pattern.
- Vague intensifier words are doing the heavy lifting."Quietly," "shape," "matters," "hold," "pull," "signal," "the work" — words that sound like they mean something specific without ever naming what. Swap one out and ask what it actually claims; often the honest answer is nothing.
- Formal transition words open more than a third of paragraphs."Moreover," "furthermore," "additionally," "in conclusion" — connectors real writers use sparingly show up in generated drafts at two to three times the natural rate.
- Every argument gets exactly two balanced sides.A genuinely opinionated piece takes a position and defends it. A hedge-heavy one lists the case for, the case against, and never actually lands anywhere — a shape that reads as fair but usually just reads as empty.
- The examples are hypothetical instead of specific."A marketing team might use this to improve engagement" is the generated version. "We tested this on a 40-person list and open rate moved from 22% to 31%" is the human one — named, numbered, falsifiable.
- Every section follows the exact same shape.Topic sentence, three supporting points, one-line wrap — repeated section after section like a template rather than a piece someone thought through in real time.
- Statements tell you the conclusion instead of showing the reason."The launch was successful. The team hit every deadline. Results were strong." Three flat claims in a row, no texture, no specific moment that would let a reader draw the conclusion themselves.
- The whole piece reads like it could've been written by anyone in your space.If a competitor ran the same prompt on the same topic, would the output be recognizably different from yours? If not, nothing in it is actually yours.
Run your draft through AI Detector — free, no login, instant score, right in your browser.
How Many Of These Actually Matter?
None of this works as a single tripwire. One em dash doesn't flag anything. One "quietly" doesn't either. What flags content is density and combination — several signals stacking in the same paragraph, or the same one repeating across an entire piece. A human writer trips one or two of these by accident in any given article. A raw, unedited AI draft usually trips six or more, consistently, from the first paragraph to the last.
That's also roughly how the checklist doubles as a self-test. Skim a draft once looking only for sentence-length variation and em dash frequency — the two easiest to spot without close reading. If both look flat, go back through the rest of the list before deciding the piece is fine as-is. There's a useful rule of thumb buried in that: if you have to ask whether a passage sounds like AI, it probably does. Writing that clearly doesn't rarely prompts the question in the first place.
The other direction matters too — don't over-correct into choppy, artificially broken-up sentences just to defeat a detector. Forced burstiness reads as strangely as absent burstiness. The goal is writing that sounds like you'd actually say it out loud, not writing engineered to score well on one specific check.
Where a Detector Score Actually Helps
Running a draft through AI Detector before publishing does two things a manual skim can't. It gives you a number to track across revisions, so "does this feel more human now" turns into "did the score actually move." And it catches signals that are genuinely hard to eyeball — token-level predictability, for instance, which is nearly impossible to judge just by reading, even for an experienced editor.
It's not a verdict, though. Detector scores carry real error rates, and non-native English writers in particular get flagged at meaningfully higher rates than native speakers producing entirely human, original writing — a documented bias worth knowing before treating any single score as proof of anything. Use it as one input alongside the checklist above, not a pass/fail gate.
This matters more for some content than others. A quick internal Slack update doesn't need this level of scrutiny. A landing page, a client proposal, or anything with your name on it in front of a skeptical reader benefits from the extra pass — the stakes for sounding generic are highest exactly where it costs you the most.
Fixing These Without Starting Over
None of this means rewriting from scratch. Most of the checklist above is a targeted edit, not a rewrite:
- Read it aloud once. Flat rhythm is far more obvious out loud than on a screen — the fastest way to catch issue #1 before anything else.
- Cut every third em dash. Not all of them, just enough that they stop being the default punctuation choice.
- Replace one vague intensifier per paragraph with a specific number, name, or detail. "Quietly reshapes the market" becomes "added 40,000 signups in the first quarter."
- Delete one balanced-both-sides paragraph and replace it with an actual opinion. Committing to a position is the fastest way to sound like someone who did the thinking, not a summary of other people's thinking.
- Read it a second time looking only for "the" followed by an abstract noun — "the work," "the shift," "the pull." Each instance is a place where a concrete noun or number was available and a vague one got used instead.
- Run the result through AI Humanizer for a second pass, then AI Detector to check whether the score actually moved. If it didn't, the edit didn't touch the real problem — go back to the checklist and find which signal is still present.
Not inherently. Google has said repeatedly that it rewards quality and helpfulness regardless of how content was produced, not the production method itself. The problem is generic, low-information-gain content, which AI drafts produce by default unless someone edits them. That said, Google has also said it can identify low-effort, mass-produced content at scale, so "AI-assisted" and "unedited AI output" aren't treated the same in practice even if written policy doesn't distinguish them explicitly.
There's no fixed number, but as a rough guide: one or two showing up occasionally is normal even in careful human writing. Four or more appearing consistently across a piece, especially numbers 1, 2, and 5 together, is a strong signal the draft needs another editing pass. Context matters too — a checklist item appearing once in a 2,000-word article reads very differently than the same item appearing three times in a 300-word section.
They're a useful signal, not a verdict. Our own AI Detector and every other detection tool measure statistical patterns like token predictability and sentence-length variation, and they carry real false-positive and false-negative rates — non-native English writers and heavily-edited human writing both get misflagged more often than typical native-English first drafts. The tools have gotten meaningfully better over the past year, but meaningfully better still isn't reliable enough to be the only check.
Some of them will, the same way "delve" and "landscape" mostly already have. That's exactly why this list needs revisiting periodically rather than treated as permanent. Expect this exact list to look different in a year — plan to revisit it rather than treat any single version as final. The underlying principle, avoiding generic and overly tidy writing, will outlast any specific word on it.
Partially. Prompting for varied sentence length and specific examples helps, but models default back to their trained patterns over a long piece regardless of instructions. Editing the output afterward, checking against a list like this one, still catches more than prompting alone. Think of prompting as reducing how much editing you'll need to do afterward, not eliminating the need for it.
Yes, arguably more. A 200-word post with three of these signals reads as obviously generated much faster than a 2,000-word article, since there's no room for the flatness to blend in. If you're publishing short-form content at volume, it's worth running a handful of posts through the checklist together rather than checking each one individually — patterns that aren't obvious in a single post often are across five.