Orphan Pages: How to Find and Fix Them
On the audits we run, almost every site past a few hundred URLs has pages that nothing links to. Some of them are ranking. Here is how to find yours this week without buying a tool, and how to decide what each one deserves.
Orphans hide in the gap between what your sitemap claims and what a crawler can actually reach.
The uncomfortable part is not that orphans exist. It is which ones turn up.
On the audits we run, the list is rarely junk. It is usually two or three URLs Google found years ago, kept indexed, and still sends traffic to, while the current linked version of that page sits below it.
This page covers the audit problem: finding URLs that fell out of your link graph and deciding what each one deserves. Building the graph in the first place (hub-and-spoke structure, anchor text, where links belong on a page) is our internal linking guide. That guide says how the site should be wired. This one deals with the wires that came loose.
What an orphan page actually is
An orphan page is a URL on your domain that receives zero internal links from any other page on your domain. Not "few links". Zero. It can still be indexed, still rank, and still take traffic, because discovery and linking are separate things.
Three things get confused with it constantly:
- A deep page is not an orphan. A product six clicks from the homepage is badly placed, which is a crawl-priority problem. It still has a path. An orphan has none.
- A page missing from your sitemap is not an orphan. Different sets. The gap between them is what makes the detection method below work.
- A noindexed or blocked page is rarely worth calling an orphan. If you excluded it on purpose, it is doing its job.
One honest note before you go looking for official backing: Google does not use the term "orphan page" in its documentation. We checked Search Central this run. The phrase appears only in Search Central Community forum threads posted by site owners, never in a guidance page. It is an SEO-industry label for a real structural condition, and Google describes the condition, not the word.
What Google does say is direct. Its SEO Starter Guide (last updated 10 December 2025) states that "Google primarily finds pages through links from other pages it already crawled" and that "the vast majority of the new pages Google finds every day are through links". How Search Works (last updated 18 December 2025) lists three discovery routes: previous crawls, links extracted from a known page, and a submitted sitemap.
A sitemap entry is a hint, not a substitute for a link. The sitemaps overview (last updated 10 December 2025) says a sitemap "doesn't guarantee that all the items in your sitemap will be crawled and indexed".
This is also a different problem from keyword cannibalisation, where two pages compete for one query. An orphan is the opposite shape: one URL, zero internal votes.
Why orphan pages appear on real sites
Nobody creates an orphan on purpose. They are a side effect of six routine events. Recognise more than two and run the hunt in the next section.
- A migration that redirected the menu but not the content. New slugs, rebuilt navigation, 40 old in-body links pointing nowhere. Whatever they used to reach now depends on a redirect somebody remembered to write.
- Expired campaign landing pages. A Pune D2C skincare brand builds three Diwali offer pages, points ads and an email at them, then takes the banner down in November and leaves the URLs live. Ads never created an internal link, so those pages were orphans from birth.
- CMS templates nobody configured. WordPress ships a "Hello world!" post and, on many setups, an attachment page per uploaded image. Live URLs, no menu entry.
- Filtered and faceted archives. Colour, size and price filters generate crawlable URLs that exist only while someone is clicking. Some get indexed anyway and then sit outside the structure permanently.
- Uploaded PDFs. That brand's ingredient sheet, emailed to customers and uploaded to
/wp-content/uploads/, is a live indexable URL with no link on the site at all. PDFs rank more often than people expect. - Duplicate copies that went live. A page cloned for a redesign, published, never removed. It usually sits outside the sitemap too, which is what hides it from the obvious checks.
Each produces a URL that a crawler starting at your homepage will never reach. That is the basis of the detection method.
How to find orphan pages without a paid tool
The method is set subtraction. Build one list of URLs your site links to, a second list of URLs that exist or earn traffic, and look at what appears only in the second. A spreadsheet, Search Console, GA4 and a free-tier crawler will do it.
- Crawl your site from the homepage. Start at the root, follow internal links only, export every unique internal URL. That is your "reachable" set. Screaming Frog's free version caps a crawl at 500 URLs, enough for most Indian SME sites; the licence is £199 per year (checked 23 September 2026) above that.
- Pull your XML sitemap into a column. Open
/sitemap.xml, follow the child sitemaps, paste every URL into the sheet. Anything in the sitemap but not in the crawl is your first candidate list: the CMS knows the page exists, the navigation does not. - Export pages from Search Console. In the Performance report, switch to the Pages tab, widen the date range past the default and export. Search Console Help states the report defaults to "click and impression data for your site in Google Search results for the past three months" (checked 23 September 2026). A page earning impressions that your crawl never reached is an orphan with proven demand. These are the valuable ones.
- Add the Page indexing report. Set the filter to "All known pages", which Search Console Help describes as "all URLs known to Google, whether or not they are listed in a sitemap" (checked 23 September 2026). The example lists under each status are capped at 1,000 rows, so read them as samples.
- Check the Links report. Its "Top linked pages" table under Internal links shows what has inbound links. Google states the report is "not a comprehensive list of every link on your site" and that "tables are limited to 1,000 rows" (checked 23 September 2026), so use it to confirm a suspected orphan, never to declare a site clean. What each report is good for is in our Search Console guide.
- Export GA4 landing pages. Reports, Engagement, Landing page, last 12 months. A URL with sessions your crawl never reached is being found by somebody: search, an old email, a backlink, a QR code on a leaflet.
- Sample your server logs if you can get them. Ask your host for a week of access logs and filter for Googlebot hits. A URL Googlebot crawls that your own crawl did not find is an orphan confirmed by the crawler, not inferred.
Do steps 1 and 3 first. On most sites that pair surfaces the valuable orphans within an hour; the rest is completeness. Where this check sits in a full site review is in our SEO audit checklist.
How to find orphan pages with a crawler
A crawler does the same subtraction automatically, if you give it the extra sources and run the analysis step people forget.
Screaming Frog's orphan pages tutorial (checked 23 September 2026) sets it out: enable Crawl Linked XML Sitemaps under Configuration > Spider > Crawl, connect Analytics and Search Console under Configuration > API Access, tick "Crawl New URLs Discovered In Google Analytics" and its Search Console equivalent, run the crawl, then run Crawl Analysis. That last step is the one people skip. The tool states the three Orphan URLs filters are "required post 'Crawl Analysis' for them to be populated with data", so without it the report reads empty. That is how a site with thirty orphans gets declared clean.
The crawler is faster, not more correct. It finds orphans only in the sources you connected, so a URL with no sitemap entry, no sessions and no impressions stays invisible. That is what the log step is for.
The decision table: what to do with each orphan you find
Finding them is the easy half. The common mistake is treating the whole list one way, usually a bulk delete. Sort by signal instead.
| What you found | The signal that decides it | Action | Check afterwards |
|---|---|---|---|
| Orphan with impressions or clicks in Search Console | It ranks. Something is working that you did not intend. | Link it from the relevant hub and 2–3 related pages, and add it to the sitemap. | Impressions and average position for that URL 4–6 weeks later. They should hold or rise. |
| Orphan that duplicates a page you already link to | Two URLs, one topic, and one of them is the version you maintain. | Merge the useful content into the maintained URL, then 301 the orphan to it. | The old URL shows as a redirect in the Page indexing report and the target holds position. |
| Expired campaign or seasonal landing page | No current traffic, no evergreen use, but it may hold backlinks. | 301 to the nearest live category or the campaign's permanent home. | That the target is genuinely relevant, not a blanket homepage redirect. |
| Filter, facet, tag or attachment URL | Machine-generated, thin, no unique content. | Noindex it, or block the pattern at the source and let the URLs drop out. | Google's noindex documentation (last updated 10 December 2025) warns the page "must not be blocked by a robots.txt file" or the crawler never sees the rule. |
| Indexed PDF with traffic | People are landing on a file instead of a page. | Identify or build the HTML page that should own the topic, link the PDF from it, link that page from the site. | Which of the two now ranks. If the PDF still wins, your page is thinner than the file. |
| Genuinely dead page: no traffic, no links, no purpose | Nothing points to it and nothing looks for it. | Delete it and return 410, or 404 it. | That nothing else linked to it. What removing URLs does is in our guide to 404 errors. |
One redirect rule: point each orphan at the closest equivalent page, never at the homepage as a job lot. Google's redirects documentation (last updated 14 April 2026) says "the indexing pipeline uses the redirect as a signal that the redirect target should be canonical". A mass homepage redirect therefore tells Google your homepage is the canonical version of forty unrelated pages.
Working a list like this across a few hundred URLs is about two days, and it is a standard part of how we scope SEO engagements. The audit produces the list; the decision rule turns it into a work queue instead of an argument.
What we found on our own site
First-hand, and not flattering. We ran these steps on digitalhangover.in in September 2026. Three things came out, and the numbers below are our own Search Console data from that month.
- A legacy slug outranking its replacement. An old post at
/blog/on-page-seo-a-practical-guide-checklist-2026was earning roughly 21 times the impressions of the newer, properly linked page built to replace it. Nothing on the site linked to it. Google had simply kept it, and it kept winning. - The WordPress default post was ranking. The "Hello world!" post that ships with every install was live, indexed and picking up impressions. Not many. More than zero, which is worse in its own way.
- Duplicate copies of service pages outside the sitemap. Second versions of live service pages, made during build work, published, never removed.
What we did: kept the legacy on-page SEO URL live and linked it into the structure instead of redirecting it on reflex, because folding the stronger URL into the weaker one would have thrown away demand it was already serving. The default post was deleted. The duplicate service pages were 301'd to their real counterparts.
Our view, from our own list: the instinct to redirect every orphan into the "correct" page is wrong often enough to be dangerous. Check impressions before you decide which URL is canonical. Sometimes the accident is outperforming the plan, and the honest move is to adopt the accident.
How to stop creating orphan pages
Four habits, none of which needs a tool.
- Nothing publishes without an inbound link. Put "which existing page links to this, with what anchor?" in the publishing checklist. No answer, no publish.
- Every hub links to all of its children. A pillar listing nine of its eleven spokes has just built two orphans on purpose.
- Campaign URLs get an expiry decision at launch. Decide on day one whether the Diwali page gets deleted, redirected or relinked in January. Cheap to decide in October, expensive in March.
- Re-run the crawl-versus-Search-Console diff every quarter. Under an hour once the sheet exists, and short enough a cycle that whoever caused the orphans still remembers doing it.
One more for migrations: crawl the old site before you switch and keep that URL list. Almost every orphan we find on a client site traces back to a launch where nobody kept the before-picture.
Frequently asked questions
What is an orphan page in SEO?
An orphan page is a live URL on your website that receives no internal links from any other page on that website. It can still be indexed and still rank, because Google may have discovered it from a sitemap, an external link or a previous crawl. Google's documentation does not use the term "orphan page"; it is an SEO-industry label for the structural condition, not a named Google concept.
Are orphan pages bad for SEO?
Not automatically. The harm depends on the page. An orphaned URL that should be part of your structure loses the internal links, anchor text and crawl paths that help it, and it sits outside your reporting so nobody maintains it. An orphaned filter URL or attachment page is just clutter. The real risk is the third kind: an orphan that duplicates or outranks a page you do maintain, which splits the topic between two URLs you never intended to compete.
How do I find orphan pages without paying for a tool?
Crawl your site from the homepage following internal links only, and export every URL found. Then export your XML sitemap, the Pages tab of the Search Console Performance report, and GA4 landing pages for the last 12 months. Anything that appears in those exports but not in your crawl is an orphan candidate. Screaming Frog's free version crawls up to 500 URLs, which is enough for most small and mid-size sites.
Does Google Search Console show orphan pages?
Not as a report. There is no "orphan pages" view. You can infer them: the Page indexing report's "All known pages" filter shows all URLs known to Google whether or not they are in a sitemap, and the Links report's internal links table shows what has inbound links. Google states the Links report is not a comprehensive list of every link on your site and that its tables are limited to 1,000 rows, so use it to confirm a suspected orphan rather than to declare a site clean. Both checked 23 September 2026.
Should I delete an orphan page or link to it?
Check its Search Console impressions first. If the URL is earning impressions or clicks, link it into your structure and add it to the sitemap. Deleting or redirecting a page that already works throws away demand you have. If it duplicates a page you maintain, merge the content and 301 the orphan to the version you keep. If it is a machine-generated filter, tag or attachment URL, noindex it. Delete only when the page has no traffic, no backlinks and no purpose.
Your best-performing page might be one nothing links to
We run the crawl-versus-Search-Console diff on every audit, then turn the list into a fix queue with one decision per URL: link, merge, redirect, noindex or delete.
Make Digital Hangover a preferred source
One tap tells Google to show more of our SEO and marketing coverage in your Top Stories.
