LSI Keywords: The Myth, and What Google Actually Uses
Latent semantic indexing is a real technique from 1988. "LSI keywords" as SEO tools sell them are not, and Google has said so on the record twice. Here is the difference, and the habit to have instead.
Three things that share a name. Only the middle one is being sold to you.
Two things are true at once, and confusing them has kept this myth alive for a decade — long after the rest of SEO practice moved on. Latent semantic indexing exists: a documented, patented method from the late 1980s.
"LSI keywords" — the related-terms list a tool tells you to sprinkle through your copy — only borrows that name. Google does not run latent semantic indexing on the web, and nothing in its documentation asks you to place related terms at any rate.
Our main guide owns the discipline end to end. Finding and grouping search terms lives in our keyword research post; entities and the Knowledge Graph live in our entity SEO post. This page covers the term itself, why it is wrong, and what to do instead.
What Google has actually said about LSI keywords
Google's position is unusually blunt for Google. John Mueller of Google Search Relations has addressed it at least twice, and both times he ruled the concept out rather than qualifying it.
On 30 July 2019, on Twitter: "There's no such thing as LSI keywords -- anyone who's telling you otherwise is mistaken, sorry." Reported by Search Engine Roundtable on 31 July 2019.
On 2 January 2023, asked again, also on Twitter: "both have no effect. Anyone who tells you to use LSI keywords is ... still wrong after all these years." Reported by Search Engine Roundtable on 4 September 2023.
We cite the reports because they are what we read this run. Both are forum-style comments from a Google spokesperson rather than Search Central documentation — so treat them as what they are: an on-record denial, repeated, never since contradicted by anything Google has published.
The sibling myth got the same treatment. In a Google SEO office-hours video, Mueller said "Google does not have a notion of optimal keyword density" — reported by Search Engine Roundtable on 31 January 2023. If there is no optimal density, there is no target for a list of related terms to hit.
Latent semantic indexing is real — it is just not Google
This is the part most myth-busting posts get wrong, and it makes them easy to dismiss. Latent semantic indexing is not invented nonsense. It is a specific technique with a patent number.
US patent 4,839,853, "Computer information retrieval using latent semantic structure" was filed on 15 September 1988 by Deerwester, Dumais, Furnas, Harshman, Landauer, Lochbaum and Streeter at Bell Communications Research, and granted on 13 June 1989. The same team published "Indexing by latent semantic analysis" in the Journal of the American Society for Information Science in September 1990.
The method, in plain English: build a grid of which words appear in which documents, then use singular value decomposition — a linear-algebra technique — to squash that grid into fewer dimensions. Words that keep appearing in the same documents land near each other in that compressed space, so a query can match a document sharing no words with it.
It was a good answer to a 1988 problem: small, fixed document collections where synonyms broke keyword search. It is a poor answer to the web. The decomposition has to be recomputed over the whole collection whenever it changes, and there is no room in it for word order, negation, or who wrote the page.
| Claim | Status | Why |
|---|---|---|
| Latent semantic indexing is a real technique | True | Patented 1989 from a 1988 filing; peer-reviewed in JASIS 1990 |
| Google uses latent semantic indexing to read pages | False | Never named in Google documentation; denied twice on the record |
| "LSI keywords" are a ranking factor | False | Mueller, 2019 and again 2023 — "no such thing", "no effect" |
| There is an ideal number of related terms to include | False | Google says it has no notion of optimal keyword density |
| Related terms are worth knowing about | True | As a coverage checklist for topics you forgot, not a density target |
What Google actually uses to understand language
Google publishes a list. Its Guide to Google Search ranking systems (last updated 10 December 2025) names the systems it is willing to name, and three of them do the language work that "LSI" is imagined to do.
| System | What Google's guide says it does | What that means for your writing |
|---|---|---|
| BERT | "Allows us to understand how combinations of words express different meanings and intent" | Sentence structure carries meaning. "Parking on a hill with no curb" is not "parking on a hill with a curb" |
| Neural matching | "AI system that Google uses to understand representations of concepts in queries and pages and match them" | You do not need to repeat the exact query to be matched to it |
| RankBrain | "AI system that helps us understand how words are related to concepts" | Concept-level matching, not string-level matching |
| Passage ranking | "Identifies individual sections … of a web page to better understand how relevant a page is" | A well-answered section can earn a result even on a long page |
BERT was announced in "Understanding searches better than ever before", by Pandu Nayak on the Google blog, 25 October 2019 — one in ten US English searches at launch. Every example Google chose turned on a function word: to, for, no. The tiny words no synonym list can help you with.
Note what none of those descriptions contain: a term list, a co-occurrence threshold, a density target. They describe reading, not counting.
Entities: things, not strings
The other half of Google's understanding is not about language at all. It is about the things language refers to.
Google announced the Knowledge Graph in "Introducing the Knowledge Graph: things, not strings" on 16 May 2012, describing it as a model that "understands real-world entities and their relationships to one another: things, not strings."
That is the useful shift. Your page is not a bag of words scored against another bag of words. It is a statement about things — a product, a place, a method — and Google is working out which things, and whether you know them. Our entity SEO post covers doing that deliberately, and the same logic runs through our AI search guide, because assistants that summarise the web resolve entities too.
What to do instead: answer the sub-questions
The replacement habit is duller than a term list. Instead of asking "which related words am I missing", ask "which questions does a person with this query still have after reading my page".
Google's own quality guidance points the same way. Its helpful content documentation (last updated 10 December 2025) asks whether a page gives a "substantial, complete, or comprehensive description of the topic" and whether it offers "insightful analysis or interesting information that is beyond the obvious." Completeness is measured in questions answered, not in vocabulary spread.
The opposite behaviour has a name in Google's rules. Its spam policies page (last updated 28 August 2026) defines keyword stuffing as "filling a web page with keywords or numbers in an attempt to manipulate rankings in Google Search results" and lists "repeating identical words or phrases unnaturally" among the examples. A dumped LSI list sits closer to that end than to the helpful end.
Where to find the sub-questions
Three sources, in the order we use them. None of them needs a paid tool.
- People Also Ask on the live SERP. Search your primary query, open the PAA box, expand two questions and watch what loads in behind them. Google publishes no documentation for PAA, so treat it as observation rather than a spec — but it is Google showing you live which follow-ups it associates with your query. Our search intent guide covers reading the rest of the SERP for the same signal.
- Search Console queries your page already gets impressions for. Open the Performance report, filter to the page, switch to the Queries tab. A question-shaped query with impressions and no clicks is the clearest gap you will ever be handed: Google thinks your page is about that, and the searcher disagrees. Google's Performance report help page carries no published last-updated date, so 28 September 2026 is the date we checked it, not Google's. New to the report? Start with our Search Console guide.
- The questions inside your own source material. Read your interview notes for the things you had to ask twice. If the topic has a primary document — a patent, a standard, a Google doc — the questions it answers in its first two pages are your readers' questions too.
Then write the sections. One question per H2 or H3, answered in the first two sentences, detail after. The related terms appear on their own — you cannot explain singular value decomposition without saying "matrix". That is what coverage looks like from the inside.
To make that repeatable, our content brief guide is the format we use — it deliberately carries no keyword-density figure and no LSI list, for the reasons above. Re-briefing a whole library this way is where our SEO services engagements usually start, and expect three to six months before organic results are meaningful.
So what is an "LSI keyword generator" giving you?
Usually a related-terms tool with the wrong name on the box. It scrapes the pages currently ranking, extracts the terms they share, and hands you the overlap — a competitor vocabulary list. Not latent semantic indexing, and not what Google runs.
Which does not make it useless. Used properly, that list answers one question well: what are the ranking pages talking about that I am not?
- Good use: scanning the list for a subtopic you genuinely forgot, then writing a section about it because it deserves one.
- Good use: spotting that every ranking page covers pricing and yours does not.
- Bad use: pasting terms into a brief with target counts beside them.
- Bad use: rewriting a clear sentence to accommodate a term a tool suggested.
- Bad use: treating the list as a score to maximise. There is no score.
The test is simple. If a term earns a paragraph, it was a real gap. If it can only be inserted, it never was.
Where to go from here
Delete the LSI row from your brief template. Replace it with the questions the page must answer and a note on where each one came from.
Then judge drafts on coverage: did the page answer everything a person with that query still wanted to know, in language a human would use? That standard survives every algorithm update, because it is not built on a mechanism.
Frequently asked questions
Are LSI keywords a Google ranking factor?
No. John Mueller of Google said "there's no such thing as LSI keywords" on Twitter on 30 July 2019, and in January 2023 said they have "no effect" and that anyone recommending them is "still wrong after all these years". No Google Search Central documentation mentions latent semantic indexing.
Is latent semantic indexing a real thing?
Yes, and that is the source of the confusion. Latent semantic indexing is an information-retrieval method patented in the United States as 4,839,853 from a 1988 filing, using singular value decomposition on a term-document matrix. It is real, it is just not a technique Google applies to the web.
What does Google use instead of LSI?
Google's Guide to Google Search ranking systems, last updated 10 December 2025, names BERT for understanding how word combinations express meaning and intent, neural matching for matching concepts in queries and pages, and RankBrain for relating words to concepts. Separately, the Knowledge Graph handles entities — things rather than strings.
Should I still use an LSI keyword generator?
You can, as long as you know what it is. These tools return terms shared by the pages currently ranking — a competitor vocabulary list, not latent semantic indexing. Use it to spot subtopics you missed and write real sections about them. Never use it as a density target or a score.
How do I find the sub-questions my page should answer?
Three places. The People Also Ask box on the live SERP for your query. The Queries tab of Search Console's Performance report, filtered to that page — question-shaped queries with impressions but no clicks are direct gaps. And the questions your own source material raised while you were researching.
Want your content judged on coverage, not keyword counts?
We rebuild content briefs around the questions a page has to answer — then report on what actually moved.
Make Digital Hangover a preferred source
One tap tells Google to show more of our SEO and marketing coverage in your Top Stories.
