Sign up free

XML Sitemap Generator

Let us crawl your site or paste your own list of URLs, leave out the pages you don't want, and download a sitemap ready to submit.

About XML Sitemap Generator

This XML sitemap generator builds a list of the pages you want search engines to find, which helps most on new sites, large catalogs and pages few links point to. With Crawl my site, enter your home page and our server follows your internal links on the same host. It obeys robots.txt and skips noindex pages (both on by default), lists the address a redirect or canonical tag points to rather than the page itself, drops pages that return errors, and leaves out files such as PDFs, CSS and JavaScript. Lastmod comes from each page's Last-Modified header, and priority drops with every click away from the start page. A crawl covers up to 100 pages for visitors, 500 for free members and 5,000 for VIP, stops after 8 minutes and counts as one task. It only follows links in the HTML, so pages reached through scripts or a search box may be missed.

Under Addresses to leave out, enter one pattern per line, up to 50, such as /tag/ or ?sort=. Any address containing it stays out of the sitemap, and so does a page whose redirect or canonical points to a left-out address. Tick the option to list each page's pictures too, and up to 10 images per page go into the file in Google's image sitemap format. Next to Download sitemap.xml you get sitemap.txt, one address per line, the plain list that search engines such as Baidu accept for submitting URLs, and sitemap.html, a page listing your site's pages for visitors. The report shows how many pages were left out and how many pictures were listed. Paste a list of addresses builds sitemap.xml in your browser from any number of URLs, with optional lastmod, changefreq and priority, and isn't counted. Check the file with the Sitemap Checker.

How to use XML Sitemap Generator

  1. 1
    Pick a mode

    Crawl my site finds your pages for you, while Paste a list of addresses uses URLs you already have.

  2. 2
    Set up the crawl

    Enter your home page, keep Follow robots.txt and Skip noindex pages ticked, add any addresses to leave out, choose whether to list pictures, and click Start crawling.

  3. 3
    Download the files

    Look through the table of pages found, which you can also export as CSV, then download sitemap.xml, sitemap.txt or sitemap.html.

  4. 4
    Upload and submit

    Put the files in your site root, add a Sitemap line to robots.txt and submit sitemap.xml in Google Search Console and Bing Webmaster Tools.

Why use Cubfile for this

  • A real crawl

    Follows internal links from your start page, respects robots.txt and noindex, uses canonical URLs and dates pages by Last-Modified.

  • Addresses to leave out

    Up to 50 patterns, one per line, where * matches anything, as in /page/*. Redirects and canonicals pointing to them stay out too.

  • Image sitemap

    Optionally list up to 10 pictures per page in Google's image sitemap format, with a count in the report.

  • XML, TXT and HTML

    sitemap.txt for search engines that take a plain list, such as Baidu, and sitemap.html for visitors, each with its own download button.

FAQ

XML Sitemap Generator: questions and answers

How many pages can the sitemap generator crawl?
Up to 100 pages per crawl for visitors, 500 for free members and 5,000 for VIP members, and a crawl stops after 8 minutes. Each crawl counts as one task, with 2 in total for visitors, 2 a day for free members and no limit for VIP. Building a sitemap from a pasted list isn't counted.
How do I keep tag pages or sorted URLs out of the sitemap?
Enter one pattern per line under Addresses to leave out, up to 50. Any address containing a pattern is left out, and * matches anything, so /tag/, /page/* and ?sort= cover tag pages, paginated lists and sorted copies. A page whose redirect or canonical points to a left-out address is left out too, and the report shows how many were.
What are sitemap.txt and sitemap.html for?
sitemap.txt lists one address per line, the plain format that search engines such as Baidu accept for submitting URLs. sitemap.html is a page that lists your site's pages for visitors. Every crawl makes both next to sitemap.xml, each with its own download button.
Why are some of my pages missing?
The crawl only finds pages linked from the HTML of pages it has read. Pages blocked by robots.txt, marked noindex or matching one of your addresses to leave out are skipped, redirects and canonicals are replaced by the address they point to, and the page limit may have been reached. Add missing URLs with the list mode.
Where do I put sitemap.xml, and how do I submit it?
Upload it to your site root, add a Sitemap line with its full address to robots.txt, and submit that address in Google Search Console and Bing Webmaster Tools. The Robots.txt Generator can add that Sitemap line along with your other rules.
Share XML Sitemap Generator with a friendIt runs in any browser, and they can try it without signing up.

Related tools

HTML Meta Tag GeneratorWrite title, description and robots tags with live length checks and a preview.
TXT Robots.txt GeneratorCreate a robots.txt with presets for search engines and AI crawlers.
JSON Schema Markup GeneratorGenerate JSON-LD for FAQ, articles, products, organizations and more.
HTML Open Graph GeneratorMake Open Graph and Twitter Card tags and preview the share card.
CONF 301 Redirect GeneratorGet redirect rules for Apache, Nginx, IIS, PHP and HTML.
URL UTM Link BuilderAdd UTM tags to links so you can track campaigns in analytics.
HTML Hreflang Tag GeneratorBuild hreflang tags for every language version of a page.
CONF Htpasswd GeneratorCreate .htpasswd lines for Apache and Nginx basic authentication.