Home › Free Tools › Sitemap Generator
FREE TOOL · NO SIGNUP

Sitemap Generator

Paste a column of URLs, get valid sitemap XML back. This tool does not crawl your site — no page running in your browser can, so it works from the list you give it. What it does instead is check that list: ten checks for duplicates, wrong hosts, mixed http and https, fragments, query strings and the protocol's limits, before it writes a single line of XML.

10Checks before the XML 0Data leaves your browser FreeNo signup, no URL cap
By the Digital Hangover team · Updated September 2026 · Free forever

Your URL list

Absolute URLs including https://. Blank lines are ignored. Export the list from your CMS, your database, or a crawl.
Optional tags

Off by default. Google's sitemap documentation (last updated 8 July 2026) says it "uses the <lastmod> value if it's consistently and verifiably… accurate". Stamping today's date on every URL is neither, so a blanket lastmod is worth less than none. Switch it on only when the date is true for the pages in the list.


Off by default, and it should stay off for Google. Both are optional in the sitemaps.org 0.9 protocol, and Google's documentation states plainly: "Google ignores <priority> and <changefreq> values." The toggle exists because a few other crawlers and internal tools still read them.

Splitting
The protocol caps one file at 50,000. Many teams split at 10,000 so a single bad file is easier to find.
Used for the index file. {n} becomes 1, 2, 3…


          

          
Quick answer: an XML sitemap is a file listing the URLs on your site that you want search engines to crawl, wrapped in a <urlset> element using the sitemaps.org 0.9 namespace. One file holds up to 50,000 URLs and 50MB uncompressed; past that you split it and list the parts in a <sitemapindex> file. Only <loc> is required. Before you generate one by hand, check whether your CMS already publishes one at /sitemap.xml or /sitemap_index.xml — most do.

How to use the generator

  1. Get your URL list from the right place A crawl tells you what is linked. Your CMS or database tells you what exists. For a big or programmatic site the export is the better source, because it includes the pages nothing links to — the orphan pages a crawler never reaches.
  2. Paste it in, one URL per line Absolute URLs, with the scheme. Blank lines are skipped. There is no cap on how many you paste.
  3. Read the report before the XML This is the part worth your attention. Errors are left out of the file; warnings are included and flagged so you can decide. A clean report is the point of the exercise.
  4. Leave the optional tags off unless you have a reason lastmod, changefreq and priority all default to off, for the reasons printed next to each toggle.
  5. Copy or download, then upload and reference it Put the file at your site root, add a Sitemap: line to robots.txt, and submit it once in Search Console. You do not need to resubmit it every time it changes.

The ten checks it runs on your list

Most sitemap generators are formatters: text in, angle brackets out. The formatting is the easy half. The half that actually costs you crawl efficiency is what is in the list, so this tool validates first and reports per URL. Three of the checks are hard errors and the URL is left out of the file. Seven are warnings — the URL goes in, flagged, because in some cases it belongs there.

CheckLevelWhy it matters
Not an absolute URLErrorA relative path, a bare domain with no scheme, or a line with a space in it is not a valid <loc>. Left out.
Exact duplicateErrorListing a URL twice gains nothing and inflates the count. The first occurrence is kept, the rest reported with the line they repeat.
Over 2,048 charactersErrorThe protocol states a <loc> value "must be less than 2,048 characters". Left out rather than silently truncated.
Different hostWarningA sitemap normally covers one host. Mixing example.com and www.example.com — or a CDN subdomain — is the single most common mistake in a hand-built list.
http in an https listWarningUsually a leftover from a migration. If the site redirects http to https, the http URL in the sitemap is a redirect you are asking Google to crawl.
Has a query stringWarningFlagged, never removed. A filter or search URL rarely belongs in a sitemap; a legitimate parameterised product page does.
Trailing slash inconsistentWarning/a/ and /a are two URLs to a server. The minority form in your list gets flagged so you notice the split.
Fragment removedWarningEverything from # onwards is stripped, and said out loud. A fragment is a position on a page, not a page.
Non-ASCII charactersWarningSitemaps are UTF-8, but non-ASCII characters in a URL generally need percent-encoding. The tool does not re-encode your string — it tells you to check it.
Trailing-slash twinWarningTwo lines identical apart from a trailing slash. One of them is a redirect or a duplicate; both in one sitemap is a contradiction.

On top of the per-URL checks it reports the two limits from the protocol. It measures the real byte size of the file it just generated against the 50MB uncompressed ceiling, and counts your URLs against the 50,000 ceiling. Go past either and it splits the set at whatever you set "URLs per sitemap file" to, and writes the matching <sitemapindex> for you on the second tab.

Why lastmod, changefreq and priority are all off by default

Three of the four fields the sitemaps.org 0.9 protocol defines are optional, and two of them Google has said outright that it does not read. Google's sitemap documentation (last updated 8 July 2026) says: "Google ignores <priority> and <changefreq> values." There is no version of setting every page to priority 1.0 that helps. The toggle exists because some other crawlers and internal tooling still parse them, and because being able to see the fields makes the point better than a paragraph does.

lastmod is different, and more interesting. The same document says Google "uses the <lastmod> value if it's consistently and verifiably (for example by comparing to the last modification of the page) accurate". That is a conditional, and the condition is the whole thing. A date that matches the real last substantive change to the page is a useful scheduling signal. A date generated automatically at build time on every page in the sitemap, every night, is noise — and once a site has taught Google that its lastmod dates mean nothing, the field stops working for the pages where it would have mattered. This tool applies one date to the whole list, which is honest for a batch of pages you genuinely just changed and wrong for anything else. That is why it is off, and why the date picker only appears once you deliberately switch it on.

What this tool cannot do

Start with the biggest one, because it shapes everything else.

  • It does not crawl your site. A page running in your browser cannot fetch another domain's HTML — the browser's same-origin policy blocks it, and this page has no server behind it to do the fetching instead. So it cannot discover URLs. It can only structure the ones you hand it. Every "free unlimited" online sitemap generator that does crawl runs a server to do it, and nearly all of them stop at a few hundred pages unless you pay.
  • It cannot tell whether any URL in your list actually works. No status codes, no redirect chains, no noindex check, no canonical check, no robots.txt check. If you paste a 404 it will put the 404 in your sitemap. The list has to be right before it gets here.
  • It does not gzip, upload or submit anything. You get a file. Putting it on the server, referencing it from robots.txt and submitting it in Search Console are all still yours to do.
  • It writes plain <urlset> XML only. No image, video or news sitemap extensions, and no hreflang alternates. If you need those, the extension namespaces have to be added by hand.

If you need a crawl, use one of these instead

No point sending you away empty-handed. In rough order of how often the answer is the first one:

  • Your CMS almost certainly already does this. WordPress publishes a sitemap out of the box, and Yoast, Rank Math and All in One SEO each publish and auto-update one. Shopify, Squarespace and Webflow all generate one for you. Check /sitemap.xml and /sitemap_index.xml on your own domain before you build anything — a live sitemap that updates itself beats a hand-made file that goes stale in a week.
  • Screaming Frog SEO Spider — a desktop crawler, free for up to 500 URLs in a single crawl, and XML sitemap generation is in the free tier. It also gives you the status codes and canonicals this page cannot see. Paste its URL export in here if you want the validation report on top.
  • xml-sitemaps.com — the long-standing browser-based option, free up to 500 pages with no registration, paid past that.
  • For anything above a few thousand URLs, stop crawling and query the database. A crawl of a large ecommerce or programmatic site takes hours and still misses whatever is unlinked. A SELECT against your own content tables takes seconds and returns the truth. Then paste it here.

Where to go next

The mechanics of the file are only the last step. Our XML sitemap explainer covers what to include and what to leave out, and why a smaller sitemap usually outperforms a complete one. Crawl budget is the reason that is true. Search Console is where you find out what Google did with the file after you submitted it, and the robots.txt generator handles the directive that points at it. The complete SEO guide is the map for the rest, and the other free tools cover the neighbouring jobs.

Frequently asked questions

Does this sitemap generator crawl my website?

No, and it is worth being blunt about why. This tool runs entirely inside your browser with no server behind it, and a browser will not let a page on one domain fetch the HTML of another — that is the same-origin policy, and there is no way around it client-side. So it cannot discover your URLs. You paste the list; it validates and formats it. If you need the discovery step, a desktop crawler like Screaming Frog does it free up to 500 URLs, your CMS sitemap plugin does it automatically, and for a large site an export from your own database is faster and more complete than any crawl.

Do changefreq and priority actually matter?

Not for Google. Both tags are optional in the sitemaps.org 0.9 protocol, and Google's sitemap documentation (last updated 8 July 2026) states: "Google ignores <priority> and <changefreq> values." Setting every page to priority 1.0 does nothing except make the file bigger. They are off by default here. The toggle is there because a handful of other crawlers and internal tools still read them, and because seeing the tags makes the point land better than being told.

Should I include lastmod?

Only if the date is true. Google's documentation says it "uses the <lastmod> value if it's consistently and verifiably… accurate" — a conditional that most sites fail, because their build process stamps the current date on every URL every night. Once your dates stop meaning anything, the signal stops being used, including on the pages where an accurate date would have helped. This tool applies one date across the whole list, which is right for a batch you genuinely just updated and wrong for a full-site sitemap. That is why it defaults to off.

How many URLs can one XML sitemap hold?

50,000 URLs or 50MB uncompressed, whichever you hit first — the same limits in both the sitemaps.org protocol and Google's documentation. Past either one you split the list into several files and list those files in a <sitemapindex>, which is the file you then submit. This tool measures the real byte size of what it generates, counts your URLs, splits at whatever you set the per-file number to, and writes the index XML on the second tab. Google accepts up to 500 sitemap index files per site, so the ceiling is a long way above where most sites live.

Where do I put the file, and how does Google find it?

Upload it to your site root so it sits at https://yourdomain.com/sitemap.xml. Then do two things: add a Sitemap: https://yourdomain.com/sitemap.xml line to your robots.txt, which is how every crawler discovers it, and submit it once under Sitemaps in Search Console, which is how you see what Google made of it. You do not need to resubmit after each change — Google refetches submitted sitemaps on its own schedule. Nothing you paste into this page is sent to us; the XML is built in your browser and the download button writes the file locally.

FREE CRAWL REVIEW

A perfect sitemap cannot fix what is in it.

Send us your site. We will crawl it properly and tell you which URLs in your sitemap are redirects, which are noindexed, which are orphaned, and which pages Google is spending your crawl budget on instead.

See how we run SEO →