
We traced 28 common SEO claims back to the earliest source we could find. Most did not start as nonsense. They started with something real, then lost a caveat, outlived the system they described, or became something that could be sold.
Phillip Wendell · commentary/corrections: service@clickclickmedia.com.au
Something odd happened in June 2026. Several SEO articles said Google had quietly updated its Search Quality Rater Guidelines.
The stories looked believable. They gave a date – 12 June 2026 – named specific section numbers, and described a new flag called “Synthetic Authority”. Within days, people were selling compliance work against the change.
There was one problem: the update did not exist. Google’s public PDF was still dated 11 September 2025. Its change log ended there. “Synthetic Authority” appeared zero times in all 182 pages.
The sourcing was worse. The article cited one “Search Engine Land June 2026 analysis”. A site-scoped search of Search Engine Land found no such analysis in its public record. The citation used a real publisher name for a source we could not locate.
The date gave the story away. “Synthetic Authority” was a real phrase from an essay by Jono Alderson, published on 12 June 2025. The fake Google update was dated 12 June 2026 – one year later to the day.
The article ran on pravinkumar.co, a freelance Webflow specialist’s site, and ended with an offer to audit sites against the new guidelines. Archived pages show a high-volume, formulaic publishing pattern. That is consistent with an automated content pipeline, but the exact generation chain cannot be proved from the public record. The false update and matching date can be proved directly; the cited Search Engine Land analysis could not be found in the publisher’s public record.
It is tempting to call this a new AI problem. It is not. The automation is new; the pattern is old. A real phrase gets moved beyond what the source said, the stronger version gets repeated, and eventually it becomes a service, a checklist or a rule.
That fake update became the starting point for this study. We traced 28 widely repeated SEO claims back to their earliest dated source, or logged when no first source could be found. The oldest source is from 1996.
The newest was only eight weeks old when this draft was checked – and Google never wrote it.
How we checked each claim
Every source was checked live or through an archive on 9 August 2026. Each claim went through the same four checks: find the earliest dated source, read what it actually says, check whether the cited page still exists, and compare it with current primary documentation.
A note on certainty: Every finding on this page describes the public record we could verify as of 9 August 2026. Where a primary source establishes something directly, we say so directly. Where a finding depends on an absence – no study, no documented mechanism, no findable origin – “none” means none found under the search protocol described here, not proof that an undiscoverable source cannot exist. If contrary primary evidence changes a conclusion, the page will be updated visibly and the change logged.

The method also caught our own memory. We had remembered a John Mueller post as support for the idea that AI can read images. Read in context, it was sarcasm making the opposite point. That mattered: a method that only catches other people is not much of a method.

How this article is organised
The 28 claims are grouped by how they changed: claims that started as marketing; real quotes that lost a caveat; true ideas stretched too far; the faster AI-era versions; and claims with no clear first source.
After the 28 cases, the article checks places where Google’s own wording conflicts, two claims that survived the test, four that remain unsettled, and 27 more quick checks.
The appendices hold the full dataset, the dead-source register, the “opposite of its own source” register, the method, archive notes and the corrections policy.
What we found

Most of the claims started around 2015-2016, so the typical rule in this study is built on a source about a decade old. The oldest dates to 1996. The newest dates to 2026.

1. Most of the claims started with something real. Of the 28: five began as seller marketing, three began with a real statement that lost its caveat, sixteen stretched a real finding beyond its scope, and our search found no clear first author for four.
2. Sources die before claims do. Eleven important source or evidence pages are dead at the URL still being cited, and one fabricated claim cites an article we could not locate in the named publisher’s public record.
Four examples show the problem: Google’s 2019 E-E-A-T whitepaper now 404s at its main URL; the 2002 “Death Of A Meta Tag” article is archive-only; the serpIQ word-count study survives only in the Wayback Machine; and the Semrush 2017 ranking study now redirects to a different product page. The full list is in the appendix.
The fake June 2026 rater-guideline story is the extra case: its only citation was a Search Engine Land article we could not locate in the publisher’s public record as of 9 August 2026.
3. Seven claims are now repeated as the opposite of words still sitting in their own source. We counted only direct contradictions in the original or current source page, so seven is a conservative number.
Examples: the most-cited word-count page now says word count is evenly distributed across the top 10; Google’s spokesman said of its top signals “there is no order”; and the tweet behind the author-bio rule ends “True E-A-T is hard to fake”. The full seven are in the appendix.
We kept the counting rules strict. Claims were excluded from this “opposite of its source” group when the contradiction came later, or when no first source had been identified under the protocol. Those cases are still covered in the main article.
Part 1: Claims that started as marketing
These five claims are the easiest to trace. The earliest clear version we found for each claim appears while someone is selling a product or service.
1. “Press release links build authority”
Started: 2005 · Type: seller marketing · Evidence: T1
This claim starts in sales copy, not in a Google document. PRWeb was selling search visibility and ranking gains by 2005. By 2009 it said more than 40,000 organisations used releases to improve search rankings.
Google pushed back in 2012. In January 2013, an SEO used one press-release link to rank Matt Cutts’s blog for the nonsense word “sreppleasers”. In July 2013, Google added optimised links in distributed press releases to its link-spam rules. That rule is still live.
The old pitch is still sold in 2026. BrandPush sells paid distribution with “high-authority backlinks”. Press releases can still earn real coverage. What does not hold is the idea that the distributed links themselves are an authority-building SEO tactic.

Sources
PRWeb homepage – Wayback 20050101031740, 20090122093246 – T1; sreppleasers – seroundtable.com/press-release-links-google-16254.html – T1; developers.google.com/search/docs/essentials/spam-policies – T1; brandpush.co – T1; all 9 August 2026.
Bottom line: Google’s live spam policy treats optimised links in distributed press releases as link spam. The ranking claim has been in wire-service sales copy since 2005.
2. “We’ll raise your Domain Authority”
Started: 2010 · Type: seller marketing · Evidence: T1
PageRank was Google’s own public authority score. Domain Authority was different from the start. SEOmoz launched it in 2010 as an estimate of how likely a domain was to rank, based on Moz’s own link index.
When Google retired public Toolbar PageRank in 2016, DA became the number the market could still see. Link sellers began pricing placements by DA. Some services now sell DA increases directly, even when the method only changes what Moz’s crawler sees.
Moz says plainly that Domain Authority is not a Google ranking factor. Google says it does not use Moz’s score. Google may have site-level signals of its own, but no outside authority score is a window into them.


Sources
searchengineland.com/seomoz-launches-open-site-explorer-33892 – T1; moz.com/learn/seo/domain-authority – T1; seroundtable.com/moz-updated-domain-authority-seo-confused-27211.html – T1; Mueller – web.archive.org/web/20200224162307/https://twitter.com/JohnMu/status/1231974976514908160 – T1; developers.google.com/search/blog/2024/03/core-update-spam-policies – T1; web.archive.org/web/20220703031154/https://www.fiverr.com/baigzia/increase-domain-authority-da-60-plus – T1.
Bottom line: Moz says Domain Authority is not a Google ranking factor, and Google says it does not use the score.
3. “Toxic backlinks need regular disavows”
Started: 2012 · Type: seller marketing · Evidence: T1
Google launched Penguin in April 2012 to deal with manipulative links. A commercial “toxic link” detector went on sale that September. Google’s disavow tool did not launch until 33 days later.
The case for routine cleanup weakened again in 2016, when Penguin changed to devalue spam instead of suppressing whole sites. In 2024, John Mueller said the idea of “toxic links” was made up by SEO tools so people would keep paying them.
Google’s live disavow guide says most sites will not need the tool. It is mainly for serious spam-link cases tied to a manual action, or a likely one. Google also warns that using it incorrectly can hurt performance. That is a very different rule from “run a toxic-link cleanup every week”.

Sources
disavow documentation – support.google.com/webmasters/answer/2648487 – T1; Link Detox launch – web.archive.org/web/20120917000035/http://www.prweb.com/releases/2012/9/prweb9885644.htm – T1; Mueller – seroundtable.com/google-again-blasts-the-concept-of-toxic-links-37453.html – T1.
Bottom line: Google says most sites do not need the disavow tool, warns misuse can hurt, and John Mueller says “toxic links” are an SEO-tool concept.
4. “Schema markup gets you cited by AI”
Started: 2023 · Type: seller marketing, then a quote was stretched · Evidence: T1/T2
The earliest clear commercial pitch we found appears in March 2023, when Schema App said schema would prepare sites for AI-powered search. At that point there was no outcome data showing that schema caused AI citations.
In March 2025, Microsoft’s Fabrice Canel said schema helps Microsoft’s LLMs understand content. The same report said it was not clear whether Google, OpenAI, Perplexity or others used schema in the same way. “Helps Microsoft understand” soon became “helps every AI cite you”.
The first controlled study arrived in May 2026. Ahrefs tracked 1,885 pages that added JSON-LD against matched controls. ChatGPT and AI Mode showed no meaningful lift, while Google AI Overviews fell 4.6%. Google’s own AI guidance says there is no special schema markup required for AI features. Schema can still be useful. The citation claim is the part that failed.


Sources
origin pitch – web.archive.org/web/20230311164200/https://www.schemaapp.com/schema-markup/the-future-of-search-ai-machine-learning-schema-markup/ – T1; “Content Knowledge Graph” – web.archive.org/web/20240718030643 – T1; web.archive.org/web/20260608211326 – T1; Canel – searchengineland.com/microsoft-bing-copilot-use-schema-for-its-llms-453455 – T1; ahrefs.com/blog/schema-ai-citations/ – T2; developers.google.com/search/docs/appearance/ai-features – T1.
Bottom line: The controlled 2026 test found no ChatGPT or AI Mode lift from adding schema, and a 4.6% fall in Google AI Overview citations.
5. “Google updated its rater guidelines in June 2026”
Started: 2026 · Type: fabricated seller claim · Evidence: T1
The opening of this article gave the short version. The full record is even clearer. In June 2026, pravinkumar.co said Google had changed sections 5.2, 5.4 and 6.1 of the Search Quality Rater Guidelines. It said section 6.1 introduced a flag called “Synthetic Authority”. None of that text exists in Google’s PDF.
The phrase itself is real. Jono Alderson used “Synthetic Authority” in an essay published on 12 June 2025. The fake Google update was dated 12 June 2026 – exactly one year later. The site also showed a high-volume, formulaic publishing pattern. That pattern is consistent with an automated content pipeline, but the exact generation chain cannot be proved from the public record.
What can be proved is enough: the update did not happen, the named Search Engine Land source does not exist, and the date matches the earlier essay exactly. The false sections were indexed, which means answer engines can now repeat them as if they were real.
Sources
static.googleusercontent.com/media/guidelines.raterhub.com/en//searchqualityevaluatorguidelines.pdf – 9 August 2026 – T1; pravinkumar.co – Wayback 2026-04-22, 2026-05-10 – T1.
Bottom line: Google’s live rater-guideline PDF is dated September 2025. “Synthetic Authority” appears zero times. The claimed June 2026 update is dated exactly one year after the essay that coined the phrase.
Part 2: Real quotes that lost their caveats
These three started with real Google statements. The problem came later, when a limit, warning or change in policy was dropped from the retelling.
6. “Your E-E-A-T score needs work”
Started: 2014/2018 · Type: real Google idea turned into a score · Evidence: T1
E-E-A-T is a real Google concept. It began as guidance for human quality raters. Google has not documented it as one numeric score that a website could raise from, say, 62 to 78.
In 2019, Google’s own whitepaper said its systems look for signs of expertise, authority and trust. The same document also said rater scores do not directly rank individual sites. Later that year, Google staff said there was no internal E-A-T or YMYL score. In 2022, Google added the second E for Experience. The audit market simply renamed the same product.
The 2024 Google leak contains thousands of modules and attributes, but no single E-E-A-T score field. Google’s current wording is the useful middle ground: E-E-A-T is not one ranking factor, but Google uses many signals to find content that shows it. The concept is real. The score sold by third parties is not.


Sources
404 – blog.google/documents/33/HowGoogleFightsDisinformation.pdf – T1; mirror – storage.googleapis.com/gweb-uniblog-publish-prod/documents/How_Google_Fights_Disinformation.pdf – T1; “E-A-T gets an extra E” – developers.google.com/search/blog/2022/12/google-raters-guidelines-e-e-a-t – T1; May 2024 leak – T1-adjacent; “E-E-A-T Auditor” – expertseoconsulting.com/eeat-audit/ – 22 April 2026 – T1; developers.google.com/search/docs/fundamentals/creating-helpful-content – T1.
Bottom line: E-E-A-T is a real quality concept. A single third-party E-E-A-T score is not a Google metric.
7. “Fix Core Web Vitals to fix your rankings”
Started: 2020 · Type: real Google statement, caveat lost · Evidence: T1
Google really did add Core Web Vitals to ranking. But the launch post carried its own limit: better page experience would not beat clearly better, more relevant content. It mattered most when pages were otherwise similar.
Before the signal even launched, Google said sites should not expect drastic changes. By 2023, page experience had been removed from Google’s ranking-systems page. The current documentation says good Core Web Vitals do not guarantee top rankings, and chasing perfect scores may not be the best use of time.
None of this makes speed work useless. Faster sites can convert better and feel better to use. The overreach is selling Core Web Vitals as a general “fix rankings” lever. This study could not find a controlled result showing that claim.


Sources
Search Engine Land 395885 – T1; developers.google.com/search/docs/appearance/page-experience – T1; web.archive.org/web/20260101013926/https://nitropack.io/ – T1.
Bottom line: Core Web Vitals matter to page experience, but Google says they do not override relevance and do not guarantee top rankings.
8. “Google penalises AI content”
Started: 2022 · Type: real Google statement, later policy changed · Evidence: T1
In April 2022, John Mueller was asked about GPT-3 content under Google’s old rules for automatically generated text. He said it would still be treated as spam. Within a week, that became the much broader headline: “Google says AI-generated content is spam.”
Google changed the wording over the next year. In February 2023 it published: “Rewarding high-quality content, however it is produced.” Its current spam rules focus on scaled content made to manipulate search, whether it was made by humans, automation or both.
The old belief still supports two markets: AI detectors on one side and “humaniser” tools on the other. Both sell around the idea that AI authorship itself is the offence. Google’s current policy says the offence is manipulation and deception, not the mere use of AI.

Sources
web.archive.org/web/20260730225932/https://www.bypassgpt.ai/ – T1; “Rewarding high-quality content” – developers.google.com/search/blog/2023/02/google-search-and-ai-content – T1; developers.google.com/search/blog/2024/03/core-update-spam-policies – T1.
Bottom line: Google’s written policy since February 2023 is to reward useful content “however it is produced”. AI use alone is not the offence.
Part 3: True ideas that were stretched too far
These claims all started with something real: an old search-engine feature, a correlation, a small test or a practice that mattered in one narrow setting. The mistake was carrying it further than the evidence allowed.
9. “Meta keywords help your Google rankings”
Started: 1996 · Type: true on old search engines, wrongly carried into Google · Evidence: T2/T1
Meta keywords were once a real search-engine feature. Infoseek, AltaVista, HotBot and Lycos used them in the 1990s. Google did not. A Search Engine Watch comparison from 1999 already listed Google among the engines that ignored the tag.
By 2002, the major engines covered in the source record that had used the tag had dropped or reduced it. Google wrote the point down in 2009 and still says today that the keywords meta tag has no effect on Google indexing or ranking.
Yet WordPress plugins still ship the field in 2026. There is one narrow reason to keep it: Yandex still says the tag can affect relevance. For Google, the rule is simple. The tag was true for old engines, not for Google.

Sources
“Death Of A Meta Tag” – web.archive.org/web/20030207082526/http://www.searchenginewatch.com/sereport/02/10-meta.html – T2; HTML 4.0 example – w3.org – T1; SEW features matrix – Wayback 19991012050946 – T2; plugin – wordpress.org – 9 August 2026 – T1.
Bottom line: Google was documented as ignoring meta keywords in 1999 and still says the tag has no effect on Google indexing or ranking.
10. “Backlinks are the #1 ranking factor”
Started: 1998 · Type: true idea stretched into a ranking order · Evidence: T2/T1
Links were central to early Google. PageRank is part of what made the search engine different. The problem is the word “#1”.
In March 2016, Google’s Andrey Lipattsev named content and links among major signals. When asked if that was the order, he said: “There is no order.” The next day, coverage turned the answer into a ranked “top three”. That stronger version stuck.
Links still matter. Google still talks about links and PageRank. But Google’s own wording weakened in 2024 from links being “an important factor”, to “a factor”, and then the sentence disappeared. The evidence supports “links matter”. It does not support a fixed #1 ranking.

Sources
“an important factor” – web.archive.org/web/20240218095802/…spam-policies – T1; “a factor” – web.archive.org/web/20240411141606/…spam-policies – T1; Illyes – x.com/patrickstox/status/1781349615465304466 – T1; reply – x.com/methode/status/1781357974578995315 – T1; How Search Works – T1.
Bottom line: Links matter. Google explicitly refused to rank them #1, and weakened its own links wording in 2024.
11. “Duplicate content gets your site penalised”
Started: 2003 · Type: real problem turned into a penalty story · Evidence: T2/T1
The “duplicate content penalty” had two strong reasons to feel real in the early 2000s. Google warned against creating many substantially duplicate pages, and its old supplemental index often held low-value or awkward duplicate URLs. To site owners, filtering could look like punishment.
By 2006 Google was already saying it mostly filtered duplicates rather than applying ranking penalties. In September 2008 it published the clearest version: “There’s no such thing as a duplicate content penalty.” That page is still live.
Duplicate content can still create real technical problems. Google may choose the wrong main URL. Signals can split across copies. Huge duplicate spaces can waste crawling. Scraping content to manipulate rankings is spam. Those are real issues, but they are not a general site-wide “duplicate content penalty”.


Sources
web.archive.org/web/20030207040004/http://www.google.com/webmasters/guidelines.html – T2; webmasterworld.com/webmaster/6612.htm – T1; developers.google.com/search/blog/2008/09/demystifying-duplicate-content-penalty – T1; Semrush Site Audit – 11 December 2025 – T1; agency guide – 9 August 2026 – T1.
Bottom line: Google has said since 2008 that there is no general duplicate-content penalty. Canonicalisation and spam are separate issues.
12. “Use LSI keywords”
Started: 2003 · Type: real science given the wrong SEO label · Evidence: T1
Latent Semantic Analysis is real research from the late 1980s and 1990. It was built for finding meaning from patterns of words in a fixed study of documents. That research does not show that Google has an “LSI keywords” system.
The SEO phrase grew later. By 2008, tools were selling “LSI keywords” as related words scraped from top-ranking pages. In 2015, LSIGraph went further and claimed Google had confirmed that using more LSI keywords would usually improve rankings. We found no such Google confirmation in the public record checked through 9 August 2026.
John Mueller put it plainly in 2019: “There’s no such thing as LSI keywords.” The practical advice underneath the label is still fine: cover a topic properly, use the words and entities it naturally needs, and answer the questions a complete page should answer. You do not need the LSI label to do that.
Sources
Mueller – web.archive.org/web/20190909171627/https://twitter.com/JohnMu/status/1156293862681468929 – T1; Patent 4,839,853, 1990 JASIS paper – Google Patents, Crossref – T1-adjacent; LSIGraph – 6 June 2026 – T1; Google self-assessment – T1.
Bottom line: LSI is real old information-retrieval research. “LSI keywords” are not a Google ranking system.
13. “NAP consistency is a top local ranking factor”
Started: 2008 · Type: once important, now overstated · Evidence: T1
NAP means name, address and phone number. In 2008, local search depended much more heavily on directory data, and David Mihm’s first Local Search Ranking Factors survey rated citations very highly. At the time, that made sense.
During the 2010s, one-time listing cleanup became a recurring subscription product. The same research family now shows the factor shrinking. Whitespark’s 2026 survey puts citations at roughly 6-7% of local pack ranking and says the weight is still falling.
On the page checked 9 August 2026, Google’s local-ranking guidance names relevance, distance and popularity and does not list NAP consistency as a ranking factor. Keep the major listings accurate – Google, Apple, Bing and important industry directories. The evidence does not support treating hundreds of minor directory listings as an ongoing ranking lever.


Sources
davidmihm.com/local-search-ranking-factors-2008 – T1; Whitespark 2026 – whitespark.ca/local-search-ranking-factors/ – T1; Google local-ranking doc – support.google.com/business/answer/7091 – 9 August 2026 – T1, verified absence.
Bottom line: Keep major listings accurate. The citation industry’s own 2026 survey puts the factor at roughly 6-7% and falling, while Google’s local-ranking page does not list NAP consistency.
14. “Bounce rate is a Google ranking factor”
Started: 2008 · Type: one-site correlation turned into doctrine · Evidence: T2/T1
The earliest clear version we found is one blog post on 21 November 2008 titled “Confirmed: Bounce Rate is A Search Engine Ranking Factor”. The proof was three Google Analytics screenshots from one site. The same post admitted the word “bounce” did not appear in Google’s SEO guide.
The idea reached trade press in 17 days. Google staff pushed back almost immediately. Matt Cutts called bounce rate noisy and easy to spam. Google later said many times that it does not use a site’s Google Analytics bounce rate for rankings. In 2020, John Mueller said on video that this was “definitely not the case”.
Google does use its own search-side click and interaction data in some systems. That is not the same thing as reading your Analytics account, and it does not make “bounce rate” a ranking signal. Bounce rate can still be useful for UX and conversion work. A high bounce can also mean the user found the answer and left happy.

Sources
origin post – web.archive.org/web/20090106010959/http://seoblackhat.com/2008/11/21/bounce-rate-seo/ – T2; Illyes on Analytics – T1; Mueller office-hours – youtube.com/watch?v=JXxkoASrqNg&t=1500 – T1; Backlinko factors list – 2 August 2026 – T1; tool blog – 9 August 2026 – T1.
Bottom line: Google has repeatedly said it does not use a site’s Analytics bounce rate for rankings. Search-side click data is a different thing.
15. “This keyword needs 2,000+ words”
Started: 2012 · Type: correlation turned into a target · Evidence: T2/T1
The 2,000-word rule came from correlation studies. In 2012, serpIQ found that top results were often long. In 2016, Backlinko reported an average of 1,890 words on page one. Neither study proved that adding words caused better rankings.
The strongest correction happened at the same Backlinko URL. The live page now says word count was evenly distributed across the top 10, with an average of 1,447 words. The evidence changed. The “2,000+ words” rule did not.
Google’s current documentation is even simpler: it has no preferred word count. Write enough to answer the query properly. Stop when the useful answer is done.

Sources
web.archive.org/web/20180122185932/http://blog.serpiq.com/how-important-is-content-length-why-data-driven-seo-trumps-guru-opinions/ – T2; Backlinko 1M-results, 1,890 words – Wayback 20160903224518 – T2; backlinko.com/search-engine-ranking – T1; developers.google.com/search/docs/fundamentals/creating-helpful-content – T1; Verblio – T1.
Bottom line: The most-cited word-count source now says word count is evenly distributed across the top 10. Google says it has no preferred word count.
16. “Schema markup improves rankings”
Started: 2014 · Type: correlation quoted without its warning · Evidence: T1
In 2014, Searchmetrics found that pages using Schema.org markup ranked about four positions higher on average. The report also said, in the same passage, that this was “not necessarily a causal relationship”. Less than 1% of sites used schema at the time, so it was also a marker for more advanced sites.
The warning was dropped in later retellings. A 2015 Google comment that schema might become a ranking factor “over time” was also shortened into a stronger claim. By 2024-2026, guides were promising four-position gains as if the study had tested cause and effect.
Google’s structured-data policies say schema controls eligibility for rich results, not normal ranking position. Schema can improve how a result is displayed and can help machines understand a page. That is useful. It is not the same as a direct ranking boost.

Sources
“4 positions higher” – pageoptimizer.pro – 2024-09-19 – T1; “up to four positions higher” – agency guide – 2025-05-13 – T1; developers.google.com/search/docs/appearance/structured-data/sd-policies – T1
Bottom line: The 2014 schema study warned that its result was not causal. Google says structured data affects rich-result eligibility, not normal ranking position.
17. “Weekly GBP posts improve your map ranking”
Started: 2017 · Type: small case study turned into a rule · Evidence: T1/T2
The rule came from a 2017 local-search case study using two listings. It suggested that Google Posts could move map rankings. The result was easy to turn into a weekly posting rule.
In 2021, the same author, Joy Hawkins, ran the stronger test: three listings, 441 tracked keywords per location and nine weeks of weekly posts. The result was no measurable direct impact on local pack rankings.
Google Posts can still be useful. They can show offers, updates and proof that a business is active. The evidence does not support selling a fixed posting quota as a ranking lever.

Sources
searchengineland.com/google-posts-removed-7-days-exception-event-posts-278420 – T1; sterlingsky.ca/do-google-posts-impact-ranking/ – T2; W3Era – web.archive.org/web/20260607203826 – T1
Bottom line: The stronger 2021 test found no measurable local-pack ranking lift from weekly Google Posts.
18. “Add author bios to every page for E-E-A-T”
Started: 2018 · Type: advice repeated without its caveat · Evidence: T1/T2
The advice traces to a 2018 tweet from Marie Haynes during the Medic-update debate. The tweet did recommend stronger About and author pages. It also said, in the same message, that this “doesn’t make you a high E-A-T site” and that true E-A-T is hard to fake.
The recommendation travelled further than the warning. A later John Mueller comment framed bios as a quality and user-experience choice, not a technical requirement. SearchPilot then tested credentialed bylines and bios and found no detectable organic-traffic lift.
Author bios can still be good publishing. They help readers know who wrote something and why that person is credible. Use them for people, trust and transparency – not because a template needs an E-E-A-T box ticked.


Sources
twitter.com/Marie_Haynes/status/1026481387463954432 – T1; searchpilot.com/resources/case-studies/authorship-content-and-eat-signals – T2
Bottom line: The 2018 advice came with its own warning, and a later controlled test found no detectable organic-traffic lift from adding bios.
19. “FAQ schema gets you into AI Overviews”
Started: 2019 · Type: old feature given a new AI reason · Evidence: T1
FAQ schema had a real feature behind it. Google launched FAQ rich results in 2019 and also connected structured Q&A to Google Assistant. That made “machine-readable FAQs” a reasonable idea at the time.
Google restricted FAQ rich results in August 2023, then stopped showing them entirely in May 2026. Yet the old markup is now sold again as an AI-visibility tactic, often with unsourced citation statistics.
Google’s AI-features guidance says there is no special schema markup needed for AI results. Clear question-and-answer content may still be easy for search systems to extract. The markup itself no longer has the feature that created the original claim.

Sources
developers.google.com/search/blog/2019/05/new-in-structured-data-faq-and-how-to – T1; developers.google.com/search/blog/2023/08/howto-faq-changes – T1; changelog – 7 May 2026 – T1; AI-features doc – T1; “2x” and “58.3%” – T4, absence logged
Bottom line: Google removed FAQ rich results in May 2026 and says no special schema is required for AI features.
Part 4: The same pattern, now moving faster with AI
The technology is new. The pattern is not. These five AI-era claims formed in months instead of years: a useful observation became a score, a tool, a checklist item or a promise before strong evidence arrived.

20. “Our AI visibility score is ground truth”
Started: 2023 · Type: public research turned into proprietary scores · Evidence: T1
The idea of measuring visibility in generative search has a public research base. The 2023 GEO paper defined methods and published its benchmark. That made the measurement inspectable.
Within months, commercial tools began publishing proprietary “AI visibility” scores. The scores can be useful, but they are samples: they depend on which prompts, engines, countries and dates the tool chooses. Two tools can measure the same brand and get different numbers without either being broken.
The useful product is repeated measurement with a disclosed method. The overreach is treating one proprietary score as ground truth.
Sources
GEO paper – arXiv 2311.09735, KDD 2024 – T2; Profound – web.archive.org/web/20240713034739/https://www.tryprofound.com/ – T1; web.archive.org/web/20260715193123/https://www.tryprofound.com/ – 15 July 2026 – T1
Bottom line: AI visibility scores are samples produced by a method. They are useful for trends, not as a single ground-truth number.
21. “Repeat your brand and keywords so ChatGPT recommends you”
Started: 2023 · Type: research stretched beyond what it tested · Evidence: T2
The founding GEO paper actually tested keyword stuffing and put it in the “non-performing” group. On the deployed engine in the paper, stuffing was about 10% worse than baseline.
Later research points the same way. Studies of AI citations show that third-party mentions matter more than simply repeating a brand on its own pages. Ahrefs also tested self-promotional pages and found little effect for an established brand, although owned pages can help a new entity establish its name.
The practical rule is boring and useful: say your brand name clearly where it belongs. Do not confuse clarity with repetition. We found no evidence in the sources checked through 9 August 2026 that stuffing your own site with your brand or target phrases makes ChatGPT recommend you more.
Sources
GEO paper, Table 1 – arXiv 2311.09735 – T2; Omniscient Digital – T3; Ahrefs self-promotion – T2; no sales page naming the deliverable – T4, absence logged
Bottom line: The paper that founded GEO put keyword stuffing in its non-performing group, about 10% worse than baseline.
22. “You need Domain Authority before AI engines will cite you”
Started: 2024 · Type: observation turned into an entry requirement · Evidence: T1/T3
The first step was a fair observation: sites cited by AI Overviews often looked authoritative. The next step was not measured. “Authoritative” became “high DA or DR”, then “you need a minimum DA/DR before AI will cite you”.
Large 2026 datasets do not support that gate. Surfer analysed five million AI citations and found authority-metric correlations close to zero. Seer reported a similar result. In Ahrefs data, branded web mentions were much more strongly related to AI visibility than DR.
Authority scores can still help with link analysis. They are not a proven entry ticket for AI citations. The stronger signal in the data is simpler: other parts of the web talking about the brand.

Sources
surferseo.com/blog/domain-authority-impact-on-ai-citations/ – T3; Ahrefs 75,000-brand set – T3; web.archive.org/web/20260807054555/https://userp.io/ – T1
Bottom line: Large 2026 studies found little or no relationship between DA/DR-style metrics and AI citations; brand mentions were much stronger.
23. “You need an llms.txt file for AI visibility”
Started: 2024 · Type: developer proposal turned into an SEO claim · Evidence: T1/T2
llms.txt began as a developer proposal for giving language models a cleaner, text-first view of a site. The proposal did not claim that the file improved rankings or AI citations.
The SEO claim arrived later. The file became an “AI readiness” checklist item and a generator feature. Then the traffic data arrived. Ahrefs checked 137,210 domains: 97% of llms.txt files had never received a visit. About 21% of the identified requests came from SEO audit tools; roughly 1% came from AI retrieval bots.
The file is cheap and harmless, so keeping one can be a reasonable hedge. The unsupported part is the promise that it currently improves AI visibility.


Sources
answer.ai/posts/2024-09-03-llmstxt.html – T1; Mueller – searchenginejournal.com/google-says-llms-txt-comparable-to-keywords-meta-tag/544804/ – T1; 137,210-domain logs – ahrefs.com/blog/llmstxt-study/ – T3; single-site logs – seodepths.com/seo-research/what-server-logs-tell-about-llms-txt/ – T3
Bottom line: Across 137,210 domains, 97% of llms.txt files had never been visited; SEO audit tools were far more common readers than AI retrieval bots.
24. “Write alt text so AI can read your images”
Started: 2025 · Type: good old practice given a new AI reason · Evidence: T1
Alt text is older than Google. Its primary documented purpose is accessibility: describe an image when a person or device cannot use the pixels directly. It also has legitimate uses in image search.
The newer claim says alt text is needed because AI systems cannot understand images. Modern multimodal systems can read pixels. In 2025, John Mueller said the choice of alt text is not mainly an SEO decision. In 2026 he also pushed back on the idea that putting important text only inside images or alt text is a good machine-reading strategy.
We found no controlled test in the public record checked through 9 August 2026 showing that alt text increases AI citations. Write good alt text anyway – for accessibility and when it helps describe the image. It does not need an invented AI-ranking story.


Sources
Google image-SEO and Gemini documentation – T1; Mueller, January 2025 – bsky.app/profile/johnmu.com/post/3lgilpi2gek2k, SEJ 2025-01-28 – T1; Mueller, February 2026 – bsky.app/profile/johnmu.com/post/3mdxp3zkwa22o, SER 2026-02-04 – T1; Mueller, March 2026 – bsky.app/profile/johnmu.com/post/3mhussxgr622s – T1; no controlled study either direction – 9 August 2026 – T4, absence logged
Bottom line: Good alt text is worth writing for accessibility. No controlled evidence on file shows that it increases AI citations.
Part 5: Claims with no clear first source
For these four, no single first author could be found. The claim was already circulating when the earliest dated source appeared.
25. “The optimal keyword density is 2-3%”
Earliest range: 2002 · Type: no clear source for the 2-3% rule · Evidence: T4/T1
There is no clear source for the familiar “2-3% keyword density” rule. The early numbers do not even agree with each other. One 2002 source said 1-7%. Ten days later another said 5-20%. Yoast’s live feature uses 0.5-3%.
Google has repeatedly said it has no concept of an optimal keyword-density percentage. The word “density” does not appear in its current spam policy or SEO Starter Guide as a target to hit.
Keywords still matter in the obvious sense. A page about emergency plumbing should say “emergency plumber”. The part that fails is the idea that there is a correct ratio. Use the language a reader needs, not a percentage dial.
| Date | Source | The published “optimal density” |
|---|---|---|
| 1958 | Luhn, IBM Journal (the science itself) | no percentage exists |
| 2002-01-24 | keyworddensity.com (archived; dead at URL) | “1% to 7%” |
| 2002-02-03 | Tabke, WebmasterWorld (archived) | “between 5 and 20%” |
| 2026 (live) | Yoast keyphrase-density feature page (archived 2026-08-09) | “between 0.5 and 3%” |
| 2025-10 → today | resume-advice industry, citing “Jobscan, 2025” | “2-3%” – of a resume |
| 1998-2026 | Google’s documentation corpus | no number, ever (“density”: 0 occurrences) |
Sources
Cutts on density, via SEJ’s ranking-factors entry – T1; Mueller, January 2022 and January 2023 – T1; spam policies and SEO Starter Guide, “density” at zero occurrences – 2026-08-09 – T1, verified absence
Bottom line: Published “optimal” density ranges have never agreed, and Google says it has no optimal keyword-density percentage.
26. “Crawl budget needs optimising on every site”
Earliest surviving use: 2010 · Type: no clear first source; scope kept shrinking · Evidence: T2/T1
“Crawl budget” was already common SEO language before Google formally adopted the term. The earliest surviving Google discussion, from 2010, is already pushing back on the simple idea that Googlebot arrives with a fixed number of pages it will crawl.
Google kept narrowing the scope. In 2017 it said crawl budget was not something most publishers needed to worry about. In 2020 it gave examples such as sites with around a million unique pages, or 10,000+ pages that change very quickly. In July 2026 the live guide says that if pages are crawled promptly, “you don’t need to read this guide”.
Crawl budget is real engineering for very large or very fast-changing sites. Clean URLs, sensible robots rules and a fast server are good practice at any size. What does not hold is treating crawl budget as a routine ranking health score for every website. Google also says crawl rate is not a ranking signal.
Sources
Enge/Cutts interview – Wayback 2010-03-11 – T2; Google on crawl rate and position, 2017 post; myths-about-crawling doc – T1
Bottom line: Google’s live guide says most sites do not need crawl-budget work. It matters at genuinely large or fast-changing scale.
27. “Exact-match domains are a ranking cheat code”
True era: 2000s · Type: once true, now outdated · Evidence: T1/T2
This one really was true for a while. In the 2000s, exact-match keyword domains could get extra ranking credit. Google confirmed in 2011 that it had been giving too much weight to keywords in domains.
On 28 September 2012, Matt Cutts announced an algorithm change to reduce that advantage for low-quality exact-match domains. Moz measured a drop in exact-match-domain influence within a day. Google later gave the change a permanent name: the “exact match domain system”, which exists to stop domains getting too much credit simply because they match a query.
Keywords in a domain are still read, and a good keyword domain can be memorable. The old bonus is the part that ended. Choose a domain because it works as a brand, not because you expect “1-3 positions higher” from the name alone.


Sources
Cutts on keyword domains – youtube.com/watch?v=rAWFv43qubI – T1; EMD tweet – web.archive.org/web/20121002010623/http://twitter.com/mattcutts/status/251784203597910016 – T1; MozCast EMD – T2; developers.google.com/search/docs/appearance/ranking-systems-guide – T1
Bottom line: Exact-match domains had extra weight, Google reduced it in 2012, and its current system exists to prevent that extra credit.
28. “Geotagging photos improves local rankings”
No clear first source · Type: later seller amplification · Evidence: T1/T2
No clear first source for this claim could be found. Local-search surveys from 2008, 2010 and 2012 did not list photo geotags as a ranking factor. Even Whitespark’s own myth page does not identify who first made the claim.
The sales pitch appeared later. GeoImgr’s 2008 homepage simply offered a way to add coordinates to photos. By early 2018, the same tool was selling geotagging as a way to improve local search rankings.
The mechanism is weak. Google strips EXIF metadata when photos are uploaded to Business Profiles. A 27-location Sterling Sky test found no measurable local-pack lift. A later 27-client test found one narrow “near me” improvement and declines elsewhere. Geotag photos if you have another reason to do it. The evidence does not support it as a general local-ranking tactic.


Sources
Whitespark, no origin cited – T4, documented absence; GeoImgr, Local SEO panel – Wayback 20180114073857 – T1; Headley, via Sterling Sky – T1; Mueller, Hundley study – searchengineland.com/geotagging-photos-google-business-profile-rank-453525 – T2; Sterling Sky, 27 locations – sterlingsky.ca/geotagging-photos-impact-ranking/ – T2
Bottom line: Google strips EXIF metadata from Business Profile uploads, and controlled local tests have not shown a general ranking gain from geotagging.
When Google’s own wording causes confusion
Not every problem starts with SEO retellings. Google sometimes changes its wording without cleaning up the old version. In other cases, two true statements sound contradictory because they describe different layers of the system.
Links: “an important factor” → “a factor” → gone
Google’s links wording changed three times in 2024: “an important factor”, then “a factor”, then no sentence at all. We found no announcement accompanying either edit in the Google record checked for this article.
Site reputation abuse: two live Google tests that do not match
In March 2024, Google said site-reputation abuse targeted third-party content with little or no first-party oversight. In November, the same author wrote that no amount of first-party involvement changes the third-party nature of the content. Both pages are still live. The later post called itself a clarification, but the tests are materially different.

Googlebot: 2MB in the page, 15MB in the machine summary
Google’s visible Googlebot documentation now says it crawls the first 2MB of a supported file type. The old 15MB figure is gone from the page. But the raw HTML still contains a machine-generated summary saying “up to 15MB”. Human-readable text and machine-readable summary disagree on the same URL.
This is the same failure pattern as the fake June 2026 update, but on Google’s own infrastructure: a machine summary kept an old fact after the source text changed.

Clicks: simple CTR theories are wrong, but Google still uses interaction data
In 2019, Gary Illyes dismissed simple CTR and dwell-time ranking theories as noisy. In 2023, Google VP Pandu Nayak testified that NavBoost uses months of click and interaction data. The 2024 leak also showed click-related modules. Both can be true: simple “raise CTR and rank higher” tactics can be wrong while aggregate search interaction data still feeds ranking systems.
The missing distinction is between your site’s Analytics metrics and Google’s own search-side interaction data. Case 14 covers that boundary.
Authority: no Moz-style score, but site-level signals exist
Google repeatedly said it did not have a public-style website authority score. The 2024 leak included an attribute named siteAuthority, and Google has also acknowledged “some site-wide signals”. The clean reading is narrow: Moz DA is not Google’s score, but Google can still have site-level signals of its own.
E-E-A-T: not one score, still part of how Google defines quality
Google says E-E-A-T itself is not one ranking factor. It also uses a large rater programme and says aggregated quality feedback helps improve how its systems judge information. So “there is no E-E-A-T score” and “Google cares about E-E-A-T-like quality” can both be true.
Raters do not directly rank pages. Their feedback helps test and train the systems that do. The disagreement often comes from treating those two statements as if only one can be true.
Sources
links documentation – 2024-02-18, 2024-04-11, 2024-09-16, 2024-11-05 – T1; developers.google.com/search/blog/2024/03/core-update-spam-policies – T1; developers.google.com/search/blog/2024/11/site-reputation-abuse – T1; Googlebot documentation – developers.google.com/search/docs/crawling-indexing/googlebot – T1; Nayak, NavBoost, US v. Google – T1, sworn testimony; May 2024 leak schema – T1-adjacent; creating-helpful-content doc, How Search Works, rater-overview PDF – T1
Two claims that survived the test
The method was not built to prove every claim wrong. These two held up once the wording was narrowed to what the evidence actually supports.
Long-lived 302 redirects can eventually be treated more like 301s. John Mueller confirmed the behaviour in 2021 and also said there is no fixed timer. So the behaviour is real; claims such as “it converts after X months” are the unsupported part. Use a 301 when the move is meant to be permanent.
AI referral traffic often converts better than normal organic traffic. Semrush reported about 4.4× on average; Ahrefs saw a much larger uplift in its own data, with wide variation by vertical. The catch is volume and selection: an AI referral often arrives after the assistant has already explained and recommended the product. The visitor is more qualified, but the channel is still small.
Sources
Mueller – seroundtable.com/google-302-redirects-treated-301-redirects-31221.html – T3; semrush.com/blog/ai-search-seo-traffic-study/ – T3; Stox – ahrefs.com/blog/ai-search-traffic-conversions-ahrefs/ – T3
Four claims we could not settle
These four did not have enough evidence for a yes or no answer. Each one is listed with the test that would settle it.
| Claim | State of evidence | What settles it |
|---|---|---|
| Responding to reviews improves local rankings | No controlled evidence either way; Google’s customer-value wording sits on its own “improve your local ranking” page | A controlled response-on/off grid test, or a direct statement |
| Entity-graph rigour produces durable AI citations | Mechanism definitionally sound; zero external outcome evidence; the schema null (case 4) lowers the prior for markup-side versions | Citation-durability data over time (a read is scheduled on our own set) |
| Cloudflare Rocket Loader breaks SEO | Mechanism real (deferred JS can move content out of some crawls’ view); community reports run both ways; Google silent | Per site, in an hour: diff rendered HTML with it on vs off |
| A tighter 2,400-word page beats its 5,200-word rewrite | The length-preference folklore is dead in both directions (case 15); the head-to-head outcome question is unread | A live revert experiment – ours is running |
Sources
Business Profile help, review responses – T1; Rocket Loader, no Google statement – T4, absence logged
27 more claims, checked quickly
The table below checks 27 more claims in one line each. The point is not that every SEO rule is false. It is that many rules become much easier to judge once the original source is opened.
The next 27 claims are checked in one line each. The full source record remains in the data files.
| Claim | Finding |
|---|---|
| Meta descriptions affect rankings | Google said no in 2009, in the same post that now anchors case 9; current guidance remains consistent with that position |
| Titles must stay under 60 characters | “there’s no limit on how long a <title> element can be” – truncation is display-only |
| Date-bumping refreshes rankings | “no additional value in making pages artificially appear to be fresh” – see methodology |
| Multiple H1s are penalised | “rank perfectly fine with no H1 tags or with five H1 tags” video |
| Social shares as ranking signals | Not direct inputs; social links are nofollow |
| Content pruning as recovery | “Deleting content is a last resort” per Google’s core-updates doc |
| More reviews always ranks higher | Threshold-and-plateau: the measured boost triggers at review 9→10; 10→11 adds nothing |
| Keyword-stuffed reviews boost rank | Controlled test null; review text earns display and AI summaries, not rank |
| AI citations are random | Grounded-engine citations are measurably systematic; chat-assistant misattribution is a different defect |
| A Wikipedia page is required | Raters score absent reputation as neutral; top ratings attainable with none |
| All commercial sites are YMYL | “many or most topics are not YMYL” |
| A fixed local ranking radius | Google documents no fixed ranking radius; distance trades against relevance and prominence per search |
| Tiered link building | Link spam pointed at link spam; amplifying a nullified link multiplies zero |
| Sitemap submission gets you indexed | “doesn’t guarantee that all the items in your sitemap will be crawled and indexed” |
| 404s must be driven to zero | 4xx “don’t waste crawl budget” (case 26); only URLs with links or traffic warrant redirects |
| IndexNow speeds up Google | Google had not adopted IndexNow as of 9 August 2026; a Bing-family protocol kept ambiguous by tool marketing |
| 410 deindexes faster than 404 | “All 4xx errors, except 429, are treated the same” |
| A traffic drop means a penalty | An empty Manual Actions report means no manual action; a traffic drop alone does not establish a penalty |
| Act fast mid-update | Google’s advice is “waiting at least a full week after a core update completes” |
| GSC average position fell = rankings fell | New queries entering at low positions drag the average down; growth reads as decline |
| Tab/accordion content is devalued | “these elements don’t violate our policies” |
| 301s leak PageRank | “301 and other permanent redirects don’t cause a loss in PageRank” |
| Sitemap priority/changefreq tuning | “Google ignores <priority> and <changefreq> values” |
| Real names required on content | “an alias or username is adequate” |
| “Google keeps your title 87% of the time” | Only under the loosest definition of “keeps”; exact-match rates run 24-67% and falling |
| Snippets need 1,100+ words and 8 images | The thresholds are not in the cited study checked for this article |
| Dynamic rendering is current best practice | Google’s doc is now written in the past tense: “Dynamic rendering was a workaround” |
One row captures the whole problem: people repeat thresholds that are not in the study they cite. Opening the source takes longer than repeating the rule.
After 28 cases, the pattern is less dramatic than it sounds. Most claims did not begin with a liar. They began with something true, useful or plausible. Then a warning fell off, the system changed, or the claim became a product. The correction usually travelled less widely than the original.
For thirty years, people carried these claims from one another. The newest one did not need people at all.
Appendices
The article ends above. The appendices contain the full 28-claim dataset, the dead-source register, the seven direct source contradictions, the method, archive notes and the corrections policy.
Full dataset
The table below is the compact data layer for all 28 cases. “Origin” combines the earliest year, source class and number of dated changes. n/a means no first source could be identified. Evidence tiers are defined in the methodology.
| # | Claim | Origin | Earliest source | Current form / evidence |
|---|---|---|---|---|
| 1 | press-release links build authority | 2005 · VENDOR-ORIGIN · 4 dated changes | wire vendor’s own sales copy (PRWeb) | wire distribution as link building, $195-$795/release Evidence: T1 |
| 2 | raising your Domain Authority | 2010 · VENDOR-ORIGIN · 3 dated changes | SEO tool launch (SEOmoz/Moz) | DA-increase gigs $10-$80; DA-KPI link retainers Evidence: T1 |
| 3 | toxic backlinks need regular disavows | 2012 · VENDOR-ORIGIN · 4 dated changes | SEO tool launch PR (Link Detox), 33 days before the disavow tool | toxicity-score audits; disavow retainers; weekly-review tool cycles Evidence: T1 |
| 4 | schema markup earns AI citations | 2023 · VENDOR-ORIGIN (misquote accelerant) · 3 dated changes | schema SaaS vendor pitch (Schema App) | “content knowledge graph” platforms; AI-schema retainers Evidence: T1 |
| 5 | June 2026 QRG “Synthetic Authority” update | 2026 · VENDOR-ORIGIN (fabrication subtype) · 3 dated changes | fabricated article, publishing pattern consistent with an automated pipeline; sole citation is a phantom | “QRG compliance” audits against a nonexistent update Evidence: T1* |
| 6 | your E-E-A-T score needs work | 2018 · MISQUOTE-ORIGIN (reification) · 4 dated changes | Google rater rubric (term 2014), turned at Medic | E-E-A-T audits; score-calculator SaaS Evidence: T1 |
| 7 | fix Core Web Vitals to fix rankings | 2020 · MISQUOTE-ORIGIN · 4 dated changes | Google announcement, hedge in its own third paragraph | CWV rank-recovery projects; speed-tool subscriptions Evidence: T1 |
| 8 | Google penalises AI content | 2022 · MISQUOTE-ORIGIN · 3 dated changes | Mueller office-hours answer, absolutised in 7 days | AI-detection passes + “humaniser” SaaS Evidence: T1 |
| 9 | the meta keywords tag helps Google rankings | 1996 · EXTRAPOLATION (temporal + cross-engine) · 8 dated changes | engine instructions (Infoseek/AltaVista) | plugin features (“every little bit helps”); a Yandex hedge Evidence: T2/T1 |
| 10 | backlinks are the #1 factor | 1998 · EXTRAPOLATION (misquote at the modern root) · 4 dated changes | academic paper (Brin & Page) – true at origin | link retainers; per-placement pricing by DA/DR Evidence: T2/T1 |
| 11 | duplicate content gets you penalised | 2003 · EXTRAPOLATION (folklore earliest source) · 11 dated changes | Google’s own guideline sentence + the supplemental index | 85%-similarity audit flags; penalty-titled remediation guides Evidence: T2/T1 |
| 12 | use LSI keywords | 2003 · EXTRAPOLATION (strongest form vendor-origin, 2015) · 7 dated changes | Google/Applied Semantics press release (real science: 1988-90) | LSI-keyword subscriptions $16.64-$49.99/mo; title-inertia guides Evidence: T1 |
| 13 | NAP consistency is a top local factor | 2008 · EXTRAPOLATION (truth decay) · 3 dated changes | practitioner survey (Mihm, LSRF v1) – true at origin | listings subscriptions; citation retainers Evidence: T1 |
| 14 | bounce rate is a ranking factor | 2008 · EXTRAPOLATION (hardening into folklore) · 10 dated changes | one blogger’s three screenshots (“Confirmed: …”) | factor #134 on the web’s most-cited list; dwell-time guides Evidence: T2/T1 |
| 15 | 2,000+ words to rank | 2012 · EXTRAPOLATION · 3 dated changes | SEO tool blog study (serpIQ), self-flagged | per-word content pricing $0.06-$0.16/word; editor targets Evidence: T2/T1 |
| 16 | schema markup improves rankings | 2014 · EXTRAPOLATION (misquote accelerant) · 4 dated changes | vendor correlational study (Searchmetrics) | standalone schema packages as a ranking lever Evidence: T1 |
| 17 | weekly GBP posts improve pack rank | 2017 · EXTRAPOLATION · 3 dated changes | practitioner case study (n=2) + a 7-day-expiry UI fact | posting quotas 15-30/mo from $399/mo Evidence: T1 |
| 18 | author bios sitewide for E-E-A-T | 2018 · EXTRAPOLATION (caveat amputation) · 4 dated changes | practitioner tweet, caveat included | bio/Person-schema rollouts; “author entity optimisation” Evidence: T1 |
| 19 | FAQ schema for AI Overviews | 2019 · EXTRAPOLATION (zombie-feature subtype) · 3 dated changes | Google rich-results launch, disclaimer included | AEO technical-setup packages Evidence: T1 |
| 20 | AI visibility scores as ground truth | 2023 · EXTRAPOLATION · 3 dated changes | academic benchmark (Princeton GEO paper) | AI-visibility SaaS; scores re-reported as ground truth Evidence: T1 |
| 21 | brand-stuff your site for ChatGPT | 2023 · EXTRAPOLATION (double scope-theft) · 4 dated changes | academic paper (GEO) + vendor study (Ahrefs), both scope-stripped | on-site “GEO copywriting” inside AEO retainers Evidence: T2/T4 |
| 22 | domain authority gates AI citations | 2024 · EXTRAPOLATION (vendor-amplified) · 3 dated changes | vendor observational research (BrightEdge), metric substituted | “LLM visibility” link retainers $5,000-$25,000+/mo Evidence: T1 |
| 23 | llms.txt for AI visibility | 2024 · EXTRAPOLATION · 4 dated changes | developer-tooling proposal (zero SEO claims) | llms.txt generators; AI-readiness audit line-items Evidence: T1 |
| 24 | write alt text for the AI | 2025 · EXTRAPOLATION · 8 dated changes | distributed emergence in AEO marketing; no single document | “AI-visible” image scanners from $5/mo; “retrieval surfaces” guides Evidence: T1 |
| 25 | keyword density: aim for 2-3% | 2002 · FOLKLORE (extrapolated from legitimate IR) · 7 dated changes | earliest published ranges (1-7% and 5-20%, ten days apart); “2-3%” itself: absence | density traffic lights (0.5-3%); resume-industry “2-3%” Evidence: T4/T1 |
| 26 | crawl budget as sitewide health metric | 2010 · FOLKLORE (term predates its earliest surviving definition) · 9 dated changes | Enge/Cutts interview – already refuting the folk model | sitewide optimisation checklists at every site size Evidence: T2/T1 |
| 27 | exact-match domains are a cheat code | n/a · FOLKLORE earliest source (no first source) · 7 dated changes | none – origin cases are confirmation (2011) + measurement (2012) | EMD-finder tools (“1-3 positions higher”); factors lists Evidence: T1 |
| 28 | geotagged photos for local rankings | n/a · FOLKLORE (late vendor amplification) · 5 dated changes | none findable; tool-vendor pivot ≤2018 (GeoImgr) | geotagging tool $12.90/mo; monthly photo quotas Evidence: T1 |
Class count:
VENDOR-ORIGIN ×5,
MISQUOTE-ORIGIN ×3,
EXTRAPOLATION ×16,
FOLKLORE ×4.
Total: 28 cases and 138 dated mutation points. Very few claims moved from source to slogan in a single step.
Finding 2: the dead-source register. Eleven important source or evidence pages are dead at the URL still being cited. The table shows what each URL serves now. Archived copies and capture IDs are filed in the Source archive below.
| Document | Claim it anchors | What its URL serves now |
|---|---|---|
| the serpIQ word-count study | word count | company defunct; readable only in the Wayback Machine |
| the Searchmetrics schema study | schema-for-rankings | domain 301-redirects wholesale to conductor.com, its acquirer |
| Google’s 2019 disinformation whitepaper | E-E-A-T | 404 at its main blog.google URL (verified 2026-08-09) |
| the 2017 Google-posts case study | GBP posts | URL redirects to a generic SEO guide |
| the founding bounce-rate post | bounce rate | seoblackhat.com now serves a hosting company’s parked placeholder |
| Bing’s 2011 dwell-time post | bounce rate | 404; archive-only – and the URL shape most secondary sources gesture at has no captures at all |
| the Enge/Cutts 2010 interview, crawl budget’s earliest surviving document | crawl budget | 404; Wayback-only |
| Sullivan’s 2002 “Death Of A Meta Tag”, the record of origin for meta-keywords engine support | meta keywords | 404; archive-only |
| keyworddensity.com, the earliest dated publication of a density range | keyword density | 404 |
| lsikeywords.com, the earliest dated artefact of “LSI keywords” as a product | LSI keywords | server unreachable, domain still resolving |
| the Semrush 2017 ranking-factors study | bounce rate | its URL 301s to an AI Visibility Index – while a live 2026 factors page still cites it as support for bounce rate |
Finding 3: seven claims now repeated as the opposite of words in their own source.
| Claim | The origin document’s own text | Circulates as |
|---|---|---|
| word-count | “word count was evenly distributed among the top 10 results” (the most-cited source URL, today) | a length target |
| links-as-#1 | “there is no order” (Lipattsev, 2016) | a ranked podium |
| schema-for-rankings | “not necessarily a causal relationship” (Searchmetrics, 2014) | a causal promise |
| Core Web Vitals | “A good page experience doesn’t override having great, relevant content” (Google, 2020) | a rank-fix lever |
| brand-stuffing for LLMs | keyword stuffing listed under “Non-Performing Generative Engine Optimization methods”, 10% worse than baseline (GEO paper, 2023) | the lever |
| author bios | “this doesn’t make you a high E-A-T site. True E-A-T is hard to fake” (Haynes, 2018) | the first half only |
| crawl budget | “there isn’t really such thing as an indexation cap… There is also not a hard limit on our crawl” (the term’s earliest surviving document, 2010) | a per-site allowance every site must monitor |
The methodology
Method. Each claim was traced back to the earliest dated primary source we could find, or to a logged absence when no first source could be found. Deleted and silently edited pages were checked in the Wayback Machine. Archive copies of cited pages are kept with the published data.
Median origin. Twenty-six of the 28 cases have a dateable earliest source. The middle two dates are 2014 and 2017, so the median sits in 2015-2016. Exact-match domains and geotagged photos are excluded because no first source was found.
Four source classes are used. Each claim gets one primary class:
- VENDOR-ORIGIN – the earliest clear version found under the protocol appears in seller marketing.
- MISQUOTE-ORIGIN – a real statement loses a caveat or limit as it is repeated.
- EXTRAPOLATION – something real is stretched beyond the situation it actually measured.
- FOLKLORE – no clear first author or source was found under the search protocol.
Two extra labels are used when useful:
- TRUE THEN, OUTDATED NOW – the claim was once correct but is still repeated as if the old system exists.
- LOW-COST HEDGE – the practice is cheap enough to keep, but the claimed mechanism is not yet supported.
Evidence is graded on the same four-level scale throughout:
- T1 – primary record: Google in writing, sworn testimony, or the subject’s own live or archived page.
- T2 – controlled evidence, or a primary public document with a dated archive trail.
- T3 – large observational data or strong trade-press confirmation. Useful for correlation, not proof of cause.
- T4 – practitioner experience, seller claims, or a logged search that found no source.
Cases were chosen for two reasons: the claim is still widely repeated, and there is enough dated evidence to trace how it changed. The quick-check table is also a feeder for future full cases. Each case ends with a one-sentence bottom line.
Tracing a claim yourself
You can repeat the method with a browser, the Wayback Machine and a few search operators:
- Find the earliest dated primary source. Follow citations backwards until the dates stop getting older, or log that no first source can be found.
- Read the source itself. Do not rely on later articles telling you what it said.
- Check whether the cited page still exists. A dead link does not make a claim false, but it can help explain why later writers stop checking whether the source has changed.
- Check current primary documentation. Compare live pages, changelogs and archived versions for silent edits.
Archive note. John Mueller’s old X posts are no longer publicly searchable as a complete archive. Many on-record statements now survive through Wayback captures and trade-press embeds. Capture IDs are included where available so the evidence can outlive the original page.
Source archive
The dead-source register above is the reader-facing finding. The underlying archive is the evidence layer: original URLs, Wayback capture IDs or archive locations, last-checked dates and saved copies where available. Capture IDs stay in the case source lines, and the index of every source is below, and the full archive bundle is available on request.
The same method caught us. CCM once used the tactic of updating a published date to make content look fresh. Google’s crawling documentation says there is “no additional value in making pages artificially appear to be fresh”. It is now recorded internally as an anti-pattern.
| PRWeb homepage (exhibit 1) | prweb.com | live domain; the 2005 and 2009 states are archive-only | 20050101031740; 20090122093246 | 9 August 2026 | bundle |
| Open Site Explorer launch coverage (exhibit 2) | searchengineland.com/seomoz-launches-open-site-explorer-33892 | live | – | 9 August 2026 | bundle |
| Mueller, “we don’t use domain authority at all” (exhibit 2) | twitter.com/JohnMu/status/1231974976514908160 | original no longer publicly readable | 20200224162307 | 9 August 2026 | bundle |
| Fiverr DA-increase gig (exhibit 2) | fiverr.com/baigzia/increase-domain-authority-da-60-plus | listed July 2022, listed now | 20220703031154 | 9 August 2026 | bundle |
| Link Detox launch release (exhibit 3) | prweb.com/releases/2012/9/prweb9885644.htm | archive-only | 20120917000035 | 9 August 2026 | bundle |
| Schema App origin pitch (exhibit 4) | schemaapp.com/schema-markup/the-future-of-search-ai-machine-learning-schema-markup/ | live, “Last Updated: 1 month ago” | 20230311164200; 20240718030643; 20260608211326 | 9 August 2026 | bundle |
| Search Quality Rater Guidelines PDF (exhibit 5) | static.googleusercontent.com/media/guidelines.raterhub.com/en//searchqualityevaluatorguidelines.pdf | live, dated September 11, 2025 | – | 9 August 2026 | bundle, binary-scanned copy |
| pravinkumar.co, the fabricated update (exhibit 5) | pravinkumar.co | live | Wayback 2026-04-22; 2026-05-10 | 9 August 2026 | bundle, plus an archive.ph capture |
| Alderson, The Rise of Synthetic Authority (exhibit 5) | advancedwebranking.com, 12 June 2025; jonoalderson.com, 17 June 2025 | live | – | 9 August 2026 | bundle |
| “How Google Fights Disinformation” whitepaper (exhibit 6) | blog.google/documents/33/HowGoogleFightsDisinformation.pdf | 404 at the canonical URL | – | 9 August 2026 | bundle, plus Google’s own live mirror at storage.googleapis.com/gweb-uniblog-publish-prod/documents/How_Google_Fights_Disinformation.pdf |
| “E-E-A-T Auditor” (exhibit 6) | expertseoconsulting.com/eeat-audit/ | live | – | 22 April 2026 | bundle |
| Page-experience announcement coverage (exhibit 7) | searchengineland.com, article 395885 | live | – | 9 August 2026 | bundle |
| NitroPack homepage (exhibit 7) | nitropack.io | live, matching the capture | 20260101013926 | 9 August 2026 | bundle |
| BypassGPT (exhibit 8) | bypassgpt.ai | live | 20260730225932 | 9 August 2026 | bundle |
| Sullivan, “Death Of A Meta Tag” (exhibit 9) | searchenginewatch.com/sereport/02/10-meta.html | 404; archive-only | 20030207082526 | 9 August 2026 | bundle |
| Search Engine Watch features matrix (exhibit 9) | searchenginewatch.com | archive-only | 19991012050946 | 9 August 2026 | bundle |
| WordPress meta-keywords plugin listing (exhibit 9) | wordpress.org | live, updated the week of 9 August 2026 | save lodged 2026-08-09 | 9 August 2026 | bundle |
| Google spam policies, the links sentence (exhibit 10, Google-versus-Google section) | developers.google.com/search/docs/essentials/spam-policies | sentence deleted between the 2024-09-16 and 2024-11-05 captures | 20240218095802 (“an important factor”); 20240411141606 (“a factor”) | 9 August 2026 | bundle, four dated captures |
| Illyes on links, conference quote and retraction (exhibit 10) | x.com/patrickstox/status/1781349615465304466; x.com/methode/status/1781357974578995315 | live | – | 9 August 2026 | bundle |
| Google quality guidelines, 7 February 2003 (exhibit 11) | google.com/webmasters/guidelines.html | superseded; archive-only | 20030207040004 | 9 August 2026 | bundle |
| Google’s duplicate-content help article (exhibit 11) | 301s twice into the choose-a-canonical-URL guide | – | 9 August 2026 | bundle | |
| Mueller, “no such thing as LSI keywords” (exhibit 12) | twitter.com/JohnMu/status/1156293862681468929 | original no longer publicly readable | 20190909171627 | 9 August 2026 | bundle |
| lsikeywords.com (exhibit 12) | lsikeywords.com | server unreachable, domain still resolving | first capture 12 July 2008 | 9 August 2026 | bundle |
| LSIGraph (exhibit 12) | lsigraph.com | live, “LSIGraph is now SurgeGraph” | 2015-06-15, the “Google has confirmed” page; save 2026-06-06 | 6 June 2026 | bundle |
| Mihm, Local Search Ranking Factors 2008 (exhibit 13) | davidmihm.com/local-search-ranking-factors-2008 | live | – | 9 August 2026 | bundle |
| Google local-ranking documentation (exhibit 13) | support.google.com/business/answer/7091 | live; NAP consistency absent, verified | – | 9 August 2026 | bundle |
| SEO Black Hat, “Confirmed: Bounce Rate is A Search Engine Ranking Factor” (exhibit 14) | seoblackhat.com/2008/11/21/bounce-rate-seo/ | domain serves a hosting company’s parked page | 20090106010959 | 9 August 2026 | bundle, plus the dead-URL check screenshot |
| Sphinn thread carrying Cutts’ reply (exhibit 14) | cited from the capture | capture 2008-12-25 | 9 August 2026 | bundle | |
| Bing’s 2011 dwell-time post (exhibit 14) | 404; archive-only, and the URL shape most secondary sources gesture at has no captures at all | – | 9 August 2026 | bundle | |
| Semrush 2017 ranking-factors study (exhibit 14) | semrush.com | 301s to an AI Visibility Index | – | 9 August 2026 | bundle |
| Backlinko, “Google’s 200 Ranking Factors” (exhibit 14) | backlinko.com | live, updated 15 May 2025 | capture 2026-08-02 | 9 August 2026 | bundle |
| serpIQ content-length study (exhibit 15) | blog.serpiq.com/how-important-is-content-length-why-data-driven-seo-trumps-guru-opinions/ | company defunct; archive-only | 20180122185932 | 9 August 2026 | bundle |
| Backlinko 1M-results study (exhibit 15) | backlinko.com/search-engine-ranking | live, reversed at the same URL | 20160903224518, the 1,890-word state | 9 August 2026 | bundle |
| Searchmetrics 2014 schema study (exhibit 16) | searchmetrics.com | domain 301-redirects wholesale to conductor.com | – | 9 August 2026 | bundle |
| “4 positions higher” circulation page (exhibit 16) | pageoptimizer.pro | live | capture lodged in the publish-time pass | 19 September 2024 | bundle |
| Hawkins/Fisher 2017 case study (exhibit 17) | redirects to a generic SEO guide | – | 9 August 2026 | bundle | |
| W3Era posting packages (exhibit 17) | archived 7 June 2026 | 20260607203826 | 9 August 2026 | bundle | |
| Haynes, 6 August 2018 (exhibit 18) | twitter.com/Marie_Haynes/status/1026481387463954432 | live | – | 9 August 2026 | bundle |
| FAQ rich-results launch post (exhibit 19) | developers.google.com/search/blog/2019/05/new-in-structured-data-faq-and-how-to | live; the feature it announced stopped appearing 7 May 2026 | – | 9 August 2026 | bundle |
| GEO paper, Aggarwal et al. (exhibit 20, exhibit 21) | arXiv 2311.09735 | live | – | 9 August 2026 | bundle |
| Profound homepage (exhibit 20) | tryprofound.com | live | 20240713034739; 20260715193123 | 9 August 2026 | bundle |
| uSERP (exhibit 22) | userp.io | live | 20260807054555 | 9 August 2026 | bundle |
| Howard, the /llms.txt proposal (exhibit 23) | answer.ai/posts/2024-09-03-llmstxt.html | live | – | 9 August 2026 | bundle |
| Mueller’s three Bluesky posts (exhibit 24) | bsky.app/profile/johnmu.com/post/3lgilpi2gek2k; /3mdxp3zkwa22o; /3mhussxgr622s | live | Save Page Now saves pending confirmation | 9 August 2026 | bundle; trade-press carries, SEJ 2025-01-28 and SER 2026-02-04 |
| keyworddensity.com (exhibit 25) | keyworddensity.com | 404 | the 24 January 2002 state, archived | 9 August 2026 | bundle |
| Tabke, “Successful Site in 12 Months with Google Alone” (exhibit 25) | webmasterworld.com | archived | the 3 February 2002 state | 9 August 2026 | bundle |
| WebPosition Page Critic product tour (exhibit 25) | archive-only | the 1999 product tour | 9 August 2026 | bundle | |
| Yoast keyphrase-density feature page (exhibit 25) | live | archived 2026-08-09 | 9 August 2026 | bundle | |
| Resume-industry “2-3%” page (exhibit 25) | live; its byline date rolls forward automatically | capture 2025-10-14 | 9 August 2026 | bundle | |
| Enge/Cutts interview, 8 March 2010 (exhibit 26) | 404; Wayback-only | capture 2010-03-11 | 9 August 2026 | bundle, plus the dead-URL check screenshot | |
| Google’s crawl-budget guide (exhibit 26) | live; migrated into the crawling-infrastructure section, updated 22 July 2026 | – | 9 August 2026 | bundle | |
| Cutts’ exact-match-domain tweet, 28 September 2012 (exhibit 27) | twitter.com/mattcutts/status/251784203597910016 | cited from the capture | 20121002010623 | 9 August 2026 | bundle |
| Google ranking-systems guide (exhibit 27) | developers.google.com/search/docs/appearance/ranking-systems-guide | live | archived 2026-08-09 | 9 August 2026 | bundle |
| GeoImgr homepage (exhibit 28) | live; the Local SEO panel appears between the 2016 and early-2018 captures | 23 December 2008; 2014; 2016; 20180114073857 | 9 August 2026 | bundle | |
| Googlebot documentation (the Google-versus-Google section) | developers.google.com/search/docs/crawling-indexing/googlebot | live; visible text says 2MB, the same page’s machine-generated summary still says 15MB | – | 9 August 2026 | bundle |
| Site-reputation abuse posts, March and November 2024 (the Google-versus-Google section) | developers.google.com/search/blog/2024/11/site-reputation-abuse | both live; the March post unedited and unannotated | – | 9 August 2026 | bundle |
| Mueller on 302s, 8 April 2021 (the acquittals) | original posts no longer publicly readable | – | 9 August 2026 | trade-press carry at seroundtable.com/google-302-redirects-treated-301-redirects-31221.html | |
| Sitemap-ping endpoint (the morgue) | google.com/ping?sitemap= | 404; deprecated June 2023 | – | 9 August 2026 | bundle |
Corrections, commentary and data
Corrections are part of the project. Every case carries a last-verified date; this edition was checked on 9 August 2026. If better primary evidence changes a conclusion, the correction will be made visibly, dated, and explained rather than silently deleted. Commentary, corrections or additional primary-source evidence: service@clickclickmedia.com.au. A public changelog will record changes.
The data behind this article is published in full. The complete dataset (JSON) holds all 28 cases with every dated change, source and evidence tier. For spreadsheets, the dated changes (CSV, 196 rows) and the quoted sources (CSV, 82 rows) are the same record, flattened. Files are dated and will be re-versioned alongside any correction. The underlying archive bundle is available on request. None of it is for sale.