XML Sitemap: What It Does, What It Doesn't, and How to Fix Yours
A sitemap is a discovery hint, not an indexing instruction. Every Google claim below was fetched from Google's own documentation on 28 September 2026 and carries that page's last-updated date.
What a sitemap does and doesn't — and the three lists that should agree but usually don't.
Most sitemap articles describe the file correctly, then quietly imply it does something it does not. So start with the boundary.
A sitemap tells Google which URLs exist. It does not tell Google to index them, and it says nothing about how they should rank. Our SEO guide covers the work that decides rankings; this page covers one file, including the fields you have probably been filling in for no reason.
We link out rather than repeat: robots.txt owns its own directives, crawl budget owns crawl demand, and canonical tags own duplicate consolidation.
What an XML sitemap is — and what it does not do
An XML sitemap is a UTF-8 encoded file listing your site's URLs, optionally with a last-modified date for each one. Google's What is a sitemap page (last updated 10 December 2025) says it "helps search engines discover URLs on your site, but it doesn't guarantee that all the items in your sitemap will be crawled and indexed."
That sentence is the whole argument of this post. Here is what it rules out.
| What you will read elsewhere | What Google's documentation says |
|---|---|
| "A sitemap helps you rank." | No ranking claim appears anywhere in Google's sitemap docs. Discovery is the stated job. |
| "Submit it and the pages get indexed." | Build and submit a sitemap (8 July 2026): "submitting a sitemap is merely a hint: it doesn't guarantee that Google will download the sitemap or use the sitemap for crawling URLs on the site." |
| "Set priority 1.0 on your money pages." | Same page: "Google ignores <priority> and <changefreq> values." |
| "Every site needs one." | The overview page says properly linked sites are usually discovered anyway, and that you may not need a sitemap at around 500 pages or fewer with solid internal linking. |
So a 40-page services site gains little from a sitemap, and loses nothing by having one. And if pages in your sitemap are not being indexed, editing the sitemap is not the fix.
One more boundary: a sitemap is not internal linking. A URL that appears only in your sitemap is still an orphan page, and Google has to guess what it relates to.
The XML format, field by field
A working sitemap needs four things: the XML declaration, a <urlset> wrapper carrying the sitemap namespace, one <url> block per page, and a <loc> inside each. Everything else is optional, and two of the optional fields are dead weight.
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://example.in/services/seo-services/</loc>
<lastmod>2026-09-14</lastmod>
</url>
<url>
<loc>https://example.in/blog/xml-sitemap/</loc>
<lastmod>2026-09-28</lastmod>
</url>
</urlset>
| Tag | Required? | What Google does with it |
|---|---|---|
<urlset> | Yes | The wrapper. Must carry the 0.9 namespace or parsers reject the file. |
<url> | Yes | One block per URL. Everything else nests inside it. |
<loc> | Yes | The URL. The sitemaps.org 0.9 protocol requires the protocol prefix and under 2,048 characters; Google asks for fully-qualified absolute URLs and says it will crawl them exactly as listed. |
<lastmod> | No | Date of last significant change, W3C datetime (YYYY-MM-DD is fine). Used only when consistently accurate — see below. |
<changefreq> | No | Seven protocol values from always to never. Ignored by Google. Harmless and pointless. |
<priority> | No | A 0.0–1.0 hint, protocol default 0.5, relative to your own site only. Ignored by Google. Setting everything to 1.0 changes nothing. |
Three file-level rules people break: every tag value must be entity-escaped, so an ampersand becomes &; the file must be UTF-8; and a sitemap can only cover URLs at or below its own directory, which is why Google recommends the site root. A file at /catalog/sitemap.xml cannot legitimately list /images/ URLs.
XML is not the only accepted format — Google also takes RSS 2.0, Atom 1.0 and a plain text file of one URL per line. Feeds carry only recent URLs, which is why most sites end up with XML.
lastmod: the one optional field worth getting right
Google uses <lastmod> only if the value is consistently and verifiably accurate, and it should reflect a significant update to main content, structured data or links — not a copyright year. That is on the build-and-submit page, last updated 8 July 2026.
Read it as the conditional it is. Accuracy is judged across the file, so one careless generator setting devalues the field for every page on the site.
- Every date is today. A plugin or build step stamping generation time instead of content-change time. The field becomes noise.
- Every date is 2023. A static file nobody regenerates. Thirty posts published since, and the sitemap says nothing changed.
For larger sites this is not cosmetic: Google's crawl budget guidance (last updated 22 July 2026) says to keep sitemaps up to date and recommends <lastmod> where content is updated. If you cannot make the date honest, leave it out.
Your sitemap is a statement about your own housekeeping
Here is the part most guides skip. A sitemap is a list of URLs you are vouching for. When it contains noindexed pages, redirects, 404s and canonicalised duplicates, you are telling Google your idea of an important URL is unreliable.
That is our reading, not a Google statement. Google has published nothing saying messy sitemaps are a quality or ranking signal, and we will not pretend otherwise. What Google does document are three things that add up:
- Sitemap inclusion is a canonicalisation signal. Google's Consolidate duplicate URLs page (10 July 2026) calls it "a weak signal that helps the URLs that are included in a sitemap become canonical", and says all pages listed in a sitemap are suggested as canonicals. Listing a filter URL alongside the page it canonicalises to is a contradiction you are sending on purpose.
- Duplicate URLs waste crawling. The crawl budget page says to eliminate duplicate content so crawling focuses on unique content rather than unique URLs.
- lastmod is trusted at file level. Accuracy is assessed across the sitemap, not per URL.
The practical rule: a URL belongs in your sitemap only if it returns 200, is indexable, and is the canonical version. Everything else comes out.
That reconciliation is roughly the first hour of our SEO engagements, and it is usually where the surprises are.
When you need a sitemap index file
You need one when a single sitemap would break a limit, or when you want sitemaps split by section so the reporting is readable. The limits, from Google's build-and-submit page: 50MB uncompressed or 50,000 URLs per sitemap file. Past either number, split and list the parts in an index.
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://example.in/sitemap-posts.xml</loc>
<lastmod>2026-09-28</lastmod>
</sitemap>
<sitemap>
<loc>https://example.in/sitemap-products.xml</loc>
<lastmod>2026-09-27</lastmod>
</sitemap>
</sitemapindex>
Google's large sitemaps page (last updated 10 December 2025) adds the rules that catch people out: up to 50,000 <loc> entries in an index file, up to 500 sitemap index files per site in Search Console, referenced sitemaps hosted on the same site unless cross-site submission is set up, and those sitemaps in the same directory as the index or lower.
Split by content type well before 50,000 URLs. Separate files for posts, products, categories and static pages show you which section is losing URLs, instead of one number for the whole site.
Image, video and news sitemaps: who actually needs them
These are extensions to the same file, added with an extra namespace. Most sites need none of them.
| Type | Who it is for | Limits and notes (Google docs, fetched 28 Sep 2026) |
|---|---|---|
| Image | Sites whose images load via JavaScript and where Google Images traffic matters | Up to 1,000 <image:image> tags per <url>; only <image:loc> needed per image. Image sitemaps, 10 December 2025. |
| Video | Sites hosting their own video where the player markup hides it | Needs <video:thumbnail_loc>, <video:title>, <video:description> and at least one of <video:content_loc> or <video:player_loc>; multiple videos per page allowed. Video sitemaps, 20 May 2026. |
| News | News publishers only | Only articles created in the last two days; maximum 1,000 <news:news> entries. News sitemaps, 10 December 2025. |
The news sitemap is the one with a discipline attached. A two-day window means it has to be generated automatically, because every entry in it is a claim about recency.
How to submit a sitemap, and what Search Console will not tell you
- Add the robots.txt line. Google says to insert
Sitemap: https://example.in/sitemap.xmlanywhere in robots.txt, as an absolute URL. That is how other crawlers find it too; the file itself is covered in our robots.txt guide. - Submit it once in Search Console. Sitemaps report, paste the sitemap or index URL. If the property is not set up, start with our Search Console guide.
- Then leave it alone. Resubmitting daily does nothing. Google reads sitemaps on its own schedule.
- Check the status weekly. The report shows the last read: Success, "Couldn't fetch", or an error count.
Now the part that causes the most confusion on client calls. The Sitemaps report Help page — checked 28 September 2026; Google shows no last-updated date on it, so that is our read date, not Google's publish date — says "Discovered URLs" is the count of page URLs parsed from the sitemap, and that "There is no guarantee that a page URL discovered in a sitemap has been or will be crawled or indexed by Google."
Discovered URLs is a parsing number, not an indexing number. For indexing, the same page points you to the Page indexing report, filtered by sitemap. Submitted versus indexed, per sitemap, is the only sitemap metric worth putting in a report.
On "Couldn't fetch", Google lists robots.txt blocking, a manual action, a URL returning 404, server unavailability and low crawl demand. Load the URL yourself first — in our experience it is usually a typo or a file that never deployed.
The sitemap mistakes we keep finding
- Not referenced in robots.txt. Submitted in Search Console years ago, invisible to everything else.
- Noindexed URLs listed. You suggest a URL as canonical while telling Google to keep it out of Search. Google's canonicalisation page prefers
rel="canonical"overnoindexfor choosing between duplicates on one site. - Redirects and 404s listed. Usually an old static file, or a generator that never checks status codes.
- Canonicalised duplicates listed. Filter, sort and pagination URLs next to the page they point at.
- Relative URLs. Google asks for fully-qualified absolute URLs.
/about/in a<loc>is not a URL. - http/https or www mismatch. The sitemap lists a host the site does not serve, so every entry is a redirect and the whole list describes URLs you do not use.
- Stale
lastmod, orlastmodset to build time. Both make the field useless. - One 48,000-URL file for a catalogue. Legal, unreadable. Split by type so crawl budget problems are visible.
- Unescaped ampersands. The file fails to parse and the report says it had errors. Five minutes to fix, months unnoticed.
- The sitemap blocked in robots.txt. Rare and fatal. And note that a disallow does not keep a page out of the index either, which Google's robots.txt introduction (10 December 2025) states plainly.
Where to go from here
If your sitemap is generated by your CMS, listed in robots.txt, contains only canonical indexable 200s and carries honest dates, it is done. There is no further optimisation hiding in the file.
The useful next step is the gap between submitted and indexed, filtered by sitemap. That gap is a content and duplication question, and it takes roughly three to six months of consistent work to move — the same timeline as anything else in organic search.
<loc> is the only required field, <lastmod> counts only when it is accurate across the file, and <changefreq> and <priority> are ignored. Split past 50,000 URLs or 50MB. List only canonical, indexable 200s, and read "Discovered URLs" as parsing, never indexing.Frequently asked questions
What is an XML sitemap and what does it actually do?
An XML sitemap is a UTF-8 encoded file listing the URLs on your site, optionally with a last-modified date for each. Its job is discovery. Google's overview page, last updated 10 December 2025, says a sitemap "helps search engines discover URLs on your site, but it doesn't guarantee that all the items in your sitemap will be crawled and indexed", and its build-and-submit page, last updated 8 July 2026, calls submitting one "merely a hint". A sitemap does not improve rankings and does not force indexing.
Does Google use changefreq and priority in a sitemap?
No. Google's Build and submit a sitemap page, last updated 8 July 2026, states: "Google ignores <priority> and <changefreq> values." Both are optional in the sitemaps.org 0.9 protocol, where priority defaults to 0.5 and is explicitly relative to other URLs on your own site only. Setting every page to 1.0 does nothing either way. The only optional field worth maintaining is lastmod, and only if the date is genuinely accurate.
How many URLs can one sitemap contain?
50,000 URLs or 50MB uncompressed, whichever comes first, for every sitemap format — per Google's Build and submit a sitemap page (8 July 2026) and the sitemaps.org 0.9 protocol. Past either limit, split the file and list the parts in a sitemap index. Google's large sitemaps page, last updated 10 December 2025, allows up to 50,000 entries in an index and up to 500 index files per site in Search Console, and requires the referenced sitemaps to sit on the same site, in the same directory as the index or lower, unless cross-site submission is set up.
Should noindexed or redirected URLs be in my sitemap?
No. Include only URLs that return 200, are indexable and are the canonical version. Google's Consolidate duplicate URLs page, last updated 10 July 2026, says all pages listed in a sitemap are suggested as canonicals and treats inclusion as "a weak signal" towards canonicalisation, so listing a noindexed page or a canonicalised duplicate sends a contradictory suggestion. Google has not said a messy sitemap is a quality signal, and we would not claim it has — but a list mixing 404s, redirects and duplicates gives Google no reason to trust the rest of it.
Does the Search Console Sitemaps report show how many pages are indexed?
No. "Discovered URLs" is the number of page URLs parsed from the sitemap. Google's Sitemaps report Help page, which we checked on 28 September 2026 and which carries no last-updated date, states: "There is no guarantee that a page URL discovered in a sitemap has been or will be crawled or indexed by Google." For indexing status, filter the Page indexing report by sitemap. That submitted-versus-indexed gap is the only sitemap number worth reporting.
We will tell you what your sitemap is hiding
Every engagement starts with a crawl reconciled against your sitemap and Search Console, so you know which URLs Google has, which it ignores, and why.
Make Digital Hangover a preferred source
One tap tells Google to show more of our SEO and marketing coverage in your Top Stories.
