About XML Sitemap Generator
This XML sitemap generator builds a list of the pages you want search engines to find, which helps most on new sites, large catalogs and pages few links point to. With Crawl my site, enter your home page and our server follows your internal links on the same host. It obeys robots.txt and skips noindex pages (both on by default), lists the address a redirect or canonical tag points to rather than the page itself, drops pages that return errors, and leaves out files such as PDFs, CSS and JavaScript. Lastmod comes from each page's Last-Modified header, and priority drops with every click away from the start page. A crawl covers up to 100 pages for visitors, 500 for free members and 5,000 for VIP, stops after 8 minutes and counts as one task. It only follows links in the HTML, so pages reached through scripts or a search box may be missed.
Under Addresses to leave out, enter one pattern per line, up to 50, such as /tag/ or ?sort=. Any address containing it stays out of the sitemap, and so does a page whose redirect or canonical points to a left-out address. Tick the option to list each page's pictures too, and up to 10 images per page go into the file in Google's image sitemap format. Next to Download sitemap.xml you get sitemap.txt, one address per line, the plain list that search engines such as Baidu accept for submitting URLs, and sitemap.html, a page listing your site's pages for visitors. The report shows how many pages were left out and how many pictures were listed. Paste a list of addresses builds sitemap.xml in your browser from any number of URLs, with optional lastmod, changefreq and priority, and isn't counted. Check the file with the Sitemap Checker.
How to use XML Sitemap Generator
- 1Pick a mode
Crawl my site finds your pages for you, while Paste a list of addresses uses URLs you already have.
- 2Set up the crawl
Enter your home page, keep Follow robots.txt and Skip noindex pages ticked, add any addresses to leave out, choose whether to list pictures, and click Start crawling.
- 3Download the files
Look through the table of pages found, which you can also export as CSV, then download sitemap.xml, sitemap.txt or sitemap.html.
- 4Upload and submit
Put the files in your site root, add a Sitemap line to robots.txt and submit sitemap.xml in Google Search Console and Bing Webmaster Tools.
Why use Cubfile for this
- A real crawl
Follows internal links from your start page, respects robots.txt and noindex, uses canonical URLs and dates pages by Last-Modified.
- Addresses to leave out
Up to 50 patterns, one per line, where * matches anything, as in /page/*. Redirects and canonicals pointing to them stay out too.
- Image sitemap
Optionally list up to 10 pictures per page in Google's image sitemap format, with a count in the report.
- XML, TXT and HTML
sitemap.txt for search engines that take a plain list, such as Baidu, and sitemap.html for visitors, each with its own download button.