Skip to main content

XML Sitemap Generator

Crawls your site from the homepage and builds a sitemap.xml from the pages that can be indexed.

We fetch the page for you. Our server fetches the public address you enter and reads the response. Nothing you enter is stored. Free, no account.

Finding the problem is the first half.

RankBrain helps implement sitemap generation from your page data. Review exclusions and verify updates after site changes.

RankBrain works on websites built in code, through the AI tool you already use (Cursor, Claude Code, Codex, ChatGPT).

Check my site free

How it works

You need a sitemap.xml, and no list of what belongs in it. This tool starts at your homepage, follows the links it finds, and keeps only the pages that can be indexed. You get the file to download and the pages it left out, with the reason for each.

  1. 1

    Starting from the address you give, we fetch pages in batches, read the links on each one and follow those that stay on the same site. You watch the crawl progress page by page.

  2. 2

    The crawl obeys robots.txt and nofollow. A page goes into the sitemap only if it returns 200, is HTML, is not marked noindex and its canonical points to itself.

  3. 3

    When it finishes you get sitemap.xml to download, plus the list of pages that were left out and the reason for each.

What it does not do

  • Up to 500 pages per crawl. Larger sites should generate the sitemap from their CMS or framework.

  • Pages that no other page links to cannot be discovered by crawling.

  • Links created by JavaScript after load are not followed.

When a source is busy or a site does not answer, this tool says so. It never fills the gap with a guess.

Questions

Why is a page missing from the sitemap?

Check the list the crawl returns. A page is left out if it did not return 200, is not HTML, is marked noindex, or has a canonical that points to a different address. Pages no other page links to cannot be found by crawling, and links created by JavaScript after load are not followed.

My site has more than 500 pages. What then?

One crawl covers up to 500 pages. Above that, generate the sitemap from your CMS or framework, which already knows every address it publishes. A crawl follows links from your pages, so it only lists what it can reach, and a large site always has pages a crawl misses.

Does the crawler respect my robots.txt and nofollow links?

Yes. The crawl obeys robots.txt and skips nofollow links, so pages you have blocked are not fetched and do not appear in the file. If a page belongs in the sitemap, allow the crawler in robots.txt first, then run the crawl again.