TMOD LogoTMOD

AI and scaled content checker

Flags scaled-content-abuse risk using an originality read plus a page-uniformity heuristic.

What TMOD checks

  • An AI originality read of your content, producing a 0–100 originality score plus specific flags for scraped or aggregated, auto-generated, mass-produced, doorway and no-original-value content.
  • A structural-uniformity heuristic that runs independently of the AI: word counts for every page with at least 50 words, and the coefficient of variation across them.
  • When six or more pages qualify and that coefficient of variation falls below 0.12, the site is flagged for content uniformity, a statistical fingerprint of templated production.
  • An originality score below 40 is reported as a failure in its own right, independent of the uniformity result.

Why it matters

Google's policy is about scaled content abuse, not about AI. Using a language model to help write is not a violation. Producing many pages primarily for search rankings rather than for people is, whether a model, a template or a team of freelancers produced them.

The uniformity heuristic catches something no word-count check can. Real writing varies in length because topics vary in what they need. Twenty pages that are all within a few percent of the same length did not happen naturally. A coefficient of variation under 0.12 across six or more pages is a strong tell, and it is measurable from the outside without reading a word.

This is also the check most worth running on your own site before someone else does. The pattern is far easier to see in aggregate than while writing, and by the time it has been noticed externally the fix involves rewriting a lot of pages.

How to fix it

01Vary structure, not just wording

If uniformity is flagged, rewriting sentences will not help, the statistic measures shape. Some pages should be short and direct, others long and thorough, with different section counts and different kinds of supporting material. Content worth writing individually ends up looking individually written.

02Add what a generator cannot

First-hand experience, original data, screenshots of your own work, a specific opinion with reasoning. These raise the originality read and are also the only durable defence, since they cannot be reproduced by anyone running the same prompt.

03Cut the pages that exist only for search

Pages targeting keyword variants of the same topic are doorway pages, and they usually show up as near-duplicate pairs as well. Consolidate them into one page that covers the topic properly and redirect the rest. Fewer, better pages is the entire fix.

04Edit AI drafts like drafts

A model's output is a starting point. Restructure it, cut what does not apply to your subject, add the specifics only you have, and check every factual claim. Publishing generated text unedited at scale is precisely the pattern the policy targets.

How the uniformity number behaves on a real site

Take eight pages of 480, 500, 495, 510, 485, 505, 490 and 500 words. The mean is 495, the standard deviation is about 10, and the coefficient of variation is roughly 0.02. Nothing about those pages is bad on its own. Together they say that a template decided how long each one would be before anybody decided what it would say.

A blog written page by page looks nothing like that. A 200-word note sits next to a 2,400-word walkthrough, and the coefficient usually lands somewhere between 0.4 and 0.8. The heuristic only runs on six or more pages of at least 50 words each, because below that a run of similar lengths is ordinary coincidence rather than a pattern.

The honest way to move the number is not to pad random pages. It is to let subjects decide length: cut the pages that never had 500 words of substance down to what they actually contain, and let the ones with more to say run long. Uniformity is a symptom of a production process, so the fix is in the process, and the per-page counts are where you can watch it change.

Editing a draft until it stops reading as generated

The tells are structural before they are stylistic. Every section the same length. Every article opening with two paragraphs of context before the subject arrives. A closing summary of what was just read. Headings that follow the identical pattern across unrelated topics. Balanced both-sides paragraphs that end without a position. None of these are errors, and all of them are what a page looks like when a template chose the shape.

Editing that out is mostly deletion. Cut the introduction until the first sentence is the answer. Delete the conclusion entirely. Replace the generic example with one from your own work, including the part that went wrong. Take a position where the draft hedged, and say what you would do rather than listing what could be done.

Then add the things no generator has access to: a number you measured, a screenshot of your own screen, a supplier who did not reply, a version that broke. This is the same work that fixes a thin page, which is not a coincidence. Both checks are circling one question, whether the page carries anything that could only have come from you, and a reviewer is asking exactly that.

Questions

Will this flag my site for using AI?

It is not an AI-text detector, and those do not work reliably enough to build on. It looks for the signals of mass production, low originality, uniform structure, no editorial voice, doorway patterns. Carefully edited AI-assisted content that varies naturally will not trip it. Templated human-written content sometimes will, correctly, because the policy is about the pattern rather than the author.

What exactly is a coefficient of variation of 0.12?

It is the standard deviation of your page word counts divided by their mean. At 0.12, page lengths cluster within roughly 12% of the average. For comparison, a normal blog with a mix of short notes and long guides typically lands somewhere between 0.4 and 0.8. Below 0.12 across six or more pages means the lengths are unnaturally consistent.

My pages are uniform because the format genuinely requires it.

That happens, recipe cards, reference entries, product specifications. The check reports it as a warning rather than a failure for exactly that reason. But it is worth being honest about whether the uniformity comes from the format or from a template with the details swapped, because a reviewer will apply the same test and will not have your context.

Why does this need fewer than six pages to be skipped?

Below six pages the statistic is not meaningful, three pages of similar length is ordinary coincidence. Requiring six with at least 50 words each keeps the heuristic from firing on small sites where it would say nothing useful.

This check runs inside the content audit

Checking one thing costs the same as checking everything, the crawl is the expensive part, not the checks.

Open it