AI and scaled content checker
Flags scaled-content-abuse risk using an originality read plus a page-uniformity heuristic.
What TMOD checks
- An AI originality read of your content, producing a 0–100 originality score plus specific flags for scraped or aggregated, auto-generated, mass-produced, doorway and no-original-value content.
- A structural-uniformity heuristic that runs independently of the AI: word counts for every page with at least 50 words, and the coefficient of variation across them.
- When six or more pages qualify and that coefficient of variation falls below 0.12, the site is flagged for content uniformity, a statistical fingerprint of templated production.
- An originality score below 40 is reported as a failure in its own right, independent of the uniformity result.

Why it matters
Google's policy is about scaled content abuse, not about AI. Using a language model to help write is not a violation. Producing many pages primarily for search rankings rather than for people is, whether a model, a template or a team of freelancers produced them.
The uniformity heuristic catches something no word-count check can. Real writing varies in length because topics vary in what they need. Twenty pages that are all within a few percent of the same length did not happen naturally. A coefficient of variation under 0.12 across six or more pages is a strong tell, and it is measurable from the outside without reading a word.
This is also the check most worth running on your own site before someone else does. The pattern is far easier to see in aggregate than while writing, and by the time it has been noticed externally the fix involves rewriting a lot of pages.
How to fix it
01Vary structure, not just wording
If uniformity is flagged, rewriting sentences will not help, the statistic measures shape. Some pages should be short and direct, others long and thorough, with different section counts and different kinds of supporting material. Content worth writing individually ends up looking individually written.
02Add what a generator cannot
First-hand experience, original data, screenshots of your own work, a specific opinion with reasoning. These raise the originality read and are also the only durable defence, since they cannot be reproduced by anyone running the same prompt.
03Cut the pages that exist only for search
Pages targeting keyword variants of the same topic are doorway pages, and they usually show up as near-duplicate pairs as well. Consolidate them into one page that covers the topic properly and redirect the rest. Fewer, better pages is the entire fix.
04Edit AI drafts like drafts
A model's output is a starting point. Restructure it, cut what does not apply to your subject, add the specifics only you have, and check every factual claim. Publishing generated text unedited at scale is precisely the pattern the policy targets.
How the uniformity number behaves on a real site
Take eight pages of 480, 500, 495, 510, 485, 505, 490 and 500 words. The mean is 495, the standard deviation is about 10, and the coefficient of variation is roughly 0.02. Nothing about those pages is bad on its own. Together they say that a template decided how long each one would be before anybody decided what it would say.
A blog written page by page looks nothing like that. A 200-word note sits next to a 2,400-word walkthrough, and the coefficient usually lands somewhere between 0.4 and 0.8. The heuristic only runs on six or more pages of at least 50 words each, because below that a run of similar lengths is ordinary coincidence rather than a pattern.
The honest way to move the number is not to pad random pages. It is to let subjects decide length: cut the pages that never had 500 words of substance down to what they actually contain, and let the ones with more to say run long. Uniformity is a symptom of a production process, so the fix is in the process, and the per-page counts are where you can watch it change.

Editing a draft until it stops reading as generated
The tells are structural before they are stylistic. Every section the same length. Every article opening with two paragraphs of context before the subject arrives. A closing summary of what was just read. Headings that follow the identical pattern across unrelated topics. Balanced both-sides paragraphs that end without a position. None of these are errors, and all of them are what a page looks like when a template chose the shape.
Editing that out is mostly deletion. Cut the introduction until the first sentence is the answer. Delete the conclusion entirely. Replace the generic example with one from your own work, including the part that went wrong. Take a position where the draft hedged, and say what you would do rather than listing what could be done.
Then add the things no generator has access to: a number you measured, a screenshot of your own screen, a supplier who did not reply, a version that broke. This is the same work that fixes a thin page, which is not a coincidence. Both checks are circling one question, whether the page carries anything that could only have come from you, and a reviewer is asking exactly that.
Fourteen doorway pages, one honest rebuild
The clearest real-world shape this check catches is the profession-swapped comparison set. Fourteen pages titled best scheduling software for dentists, for lawyers, for realtors, for plumbers and ten more, each between 610 and 640 words, each with the same five headings in the same order, each recommending the same three products. The mean length is 625 words, the standard deviation is about nine, and the coefficient of variation comes out near 0.015, an order of magnitude under the 0.12 line. Most of the pairs also score past 0.55 in the similarity check, because swapping dentist for lawyer changes very little vocabulary.
The set fails commercially before any policy reads it. Fourteen near-identical pages compete with each other for overlapping queries, none carries anything a searcher could not get from the other thirteen, and none accumulates the links or engagement that would let it rank. The policy name for the pattern is doorway pages, but the practical description is fourteen pages doing the work of zero.
The rebuild that works is smaller and more honest. One comparison page built from actually using the products: screenshots of your own account, the price on a stated date, the import that failed and what support said. Then separate pages only for the two or three professions whose requirements genuinely differ, a dental practice has recall appointments and insurance codes that a plumber does not, written about those differences. Redirect the other eleven URLs to the comparison page.
The after picture is three pages of roughly 900, 2,100 and 1,300 words. The uniformity statistic collapses because the lengths now vary the way subjects vary, the duplicate pairs are gone because the pages say different things, and the originality read finally has material to score, since screenshots and dated prices cannot come out of a template. The content audit shows all three moving together, which is the reliable sign the fix was structural rather than cosmetic.
Questions
Will this flag my site for using AI?
It is not an AI-text detector, and those do not work reliably enough to build on. It looks for the signals of mass production, low originality, uniform structure, no editorial voice, doorway patterns. Carefully edited AI-assisted content that varies naturally will not trip it. Templated human-written content sometimes will, correctly, because the policy is about the pattern rather than the author.
What exactly is a coefficient of variation of 0.12?
It is the standard deviation of your page word counts divided by their mean. At 0.12, page lengths cluster within roughly 12% of the average. For comparison, a normal blog with a mix of short notes and long guides typically lands somewhere between 0.4 and 0.8. Below 0.12 across six or more pages means the lengths are unnaturally consistent.
My pages are uniform because the format genuinely requires it.
That happens, recipe cards, reference entries, product specifications. The check reports it as a warning rather than a failure for exactly that reason. But it is worth being honest about whether the uniformity comes from the format or from a template with the details swapped, because a reviewer will apply the same test and will not have your context.
Why does this need fewer than six pages to be skipped?
Below six pages the statistic is not meaningful, three pages of similar length is ordinary coincidence. Requiring six with at least 50 words each keeps the heuristic from firing on small sites where it would say nothing useful.
If I label my content as AI-assisted, does that change the assessment?
No, in either direction. The policy question is whether pages were mass-produced for search rather than written for readers, and a label answers a different question. Disclosure is a reasonable editorial choice for your readers' sake, but a labelled doorway set is still a doorway set, and an unlabelled, carefully edited article with real experience in it passes every test here. Spend the effort on the editing rather than the disclaimer.
Read more
Screaming Frog, Ahrefs and Sitebulb for AdSense readiness
The big SEO crawlers are better at technical auditing than any AdSense tool. They are also scoped to a different question. Here is where the line falls.
The robots.txt lines that quietly block ad crawlers
Blocking Mediapartners-Google costs you ad revenue while the site keeps ranking normally, so nothing looks wrong. Here is how to read the file.
What AdSense means by "low value content"
The rejection reason gives you nothing to act on. Here is what reviewers are actually looking at, and the order to fix it in.