Home / Blog / XML Sitemap Best Practices: What Google Actually Uses

XML Sitemap Best Practices: What Google Actually Uses

The short list of XML sitemap best practices is this: list only canonical URLs that return a 200 status and are allowed to be indexed, keep each file under 50,000 URLs and 50MB uncompressed, use absolute URLs in UTF-8, give every URL an honest lastmod…

XML Sitemap Best Practices: What Google Actually Uses

The short list of XML sitemap best practices is this: list only canonical URLs that return a 200 status and are allowed to be indexed, keep each file under 50,000 URLs and 50MB uncompressed, use absolute URLs in UTF-8, give every URL an honest lastmod date, and skip changefreq and priority, because Google ignores both. Host the file at your site root, reference it in robots.txt, and check it in Search Console every month.

That’s the whole rulebook on paper. In practice, most sitemaps I audit break one of those rules without the owner knowing, usually because a plugin or theme made choices for them.

So this guide is less about what a sitemap is and more about deciding what goes in yours.

What Changes the Right Sitemap Setup for Your Site?

Four things decide how much work your sitemap needs. I check them before I touch a single setting.

  1. How many indexable pages you have. Under about 500, a sitemap is a safety net. Over 50,000, you need several files and an index.
  2. What generates the file. WordPress core, an SEO plugin, Shopify or a custom script each make different default choices, and some of them are wrong for your site.
  3. What kind of content you publish. Plain pages need a plain sitemap. Heavy video, image search traffic or Google News changes the format.
  4. How often your main content changes. A shop with daily stock changes needs lastmod handled carefully. A brochure site barely needs it.

If you run a small or mid-size site on a mainstream CMS, you’re on the main path below. The branches after it cover the bigger or stranger setups.

XML Sitemap Best Practices Google Publishes, Word for Word

I like to start from the source, because half the “best practices” floating around SEO blogs are myths. Google’s build and submit a sitemap page is short, and these are the lines that matter.

RuleWhat Google’s docs sayWhat I do with it
Size“All formats limit a single sitemap to 50MB (uncompressed) or 50,000 URLs.”Split long before the limit, by content type
Encoding“The sitemap file must be UTF-8 encoded.”Leave it to the CMS, then spot-check non-English slugs
URL format“Use fully-qualified, absolute URLs in your sitemaps.”Match the live protocol and host exactly
lastmod“Google uses the <lastmod> value if it’s consistently and verifiably… accurate.”Only update it for real content changes
changefreq, priority“Google ignores <priority> and <changefreq> values.”Leave them out
LocationA sitemap “affects only descendants of the parent directory” unless submitted in Search ConsolePut it at the root

Two of those deserve a longer look, because they cause most of the damage I see.

Why Does an Honest Lastmod Matter So Much?

Because it’s the only date field Google says it uses, and it comes with a condition. Google trusts lastmod when it’s “consistently and verifiably” accurate, and the docs add that it can check this by comparing against the page itself.

The same page spells out what counts. Updating the main content, the structured data or the links on a page is generally significant. Updating the copyright date is not. So a sitemap that stamps every URL with today’s date on every build is actively teaching Google to ignore your dates.

My rule is simple. If a reader wouldn’t notice the change, lastmod shouldn’t move. When I rewrite a post properly, I update it, and that’s exactly when I want Google to come back quickly.

And changefreq and priority? I’ve seen teams debate whether the blog deserves 0.8 or 0.6. Google ignores both values, so that time is better spent cleaning the URL list.

A Correct XML Sitemap Example

Here’s a minimal file that follows every rule above. It has absolute HTTPS URLs, W3C date formats and nothing Google ignores.

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://www.example.com/</loc>
    <lastmod>2026-09-12</lastmod>
  </url>
  <url>
    <loc>https://www.example.com/services/technical-seo/</loc>
    <lastmod>2026-08-30T14:05:00+00:00</lastmod>
  </url>
  <url>
    <loc>https://www.example.com/blog/xml-sitemap-guide/</loc>
    <lastmod>2026-09-28</lastmod>
  </url>
</urlset>

Notice what’s missing. There’s no http:// version, no ?utm_source= link, no trailing-slash duplicate. Also remember that tag values must be entity escaped, so a raw & in a URL has to become &amp; or the file breaks.

What Should You Include in an XML Sitemap, and What Should You Leave Out?

Google’s docs put it plainly: when you build a sitemap, “you’re telling search engines about which URLs you prefer to show in search results. These are the canonical URLs.” Everything else is noise, and noise has a cost later when you try to read your indexing reports.

Here’s the filter I run every URL through.

The URL…In the sitemap?Why
Returns 200, indexable, self-canonicalYesThis is exactly what the file is for
Redirects (301 or 302)NoList the destination instead
Returns 404 or 410NoRemove it, don’t wait for Google to notice
Has a noindex tag or headerNoYou’re sending two opposite signals
Canonicalizes to another URLNoList the canonical target only
Carries tracking or sort parametersNoDuplicates of a clean URL
Blocked in robots.txtNoGoogle can’t crawl it, so it can’t read the page
Thin tag or author archive you don’t want rankingUsually noKeep it out unless the archive is a real landing page

The mixed-signal cases annoy me most. A page in the sitemap that also says noindex tells Google “please show this” and “please don’t” at the same time. Google will pick one, and you’ve lost control of which.

Canonical conflicts are close behind. If your sitemap lists a URL that points its canonical somewhere else, Search Console often reports it under the “alternate page” status, which I unpack in alternate page with proper canonical tag.

Branch One: WordPress Sites

WordPress has had a built-in sitemap since version 5.5, at /wp-sitemap.xml. It works, but it’s basic. It includes user (author) sitemaps by default, and the WordPress developer reference sets 2,000 URLs as the default cap per child sitemap.

Most sites I work on run an SEO plugin instead, and the plugin replaces the core file with its own index, usually /sitemap_index.xml. That’s fine. What’s not fine is running two sitemaps at once, or trusting the plugin’s defaults without looking.

If you’re on WordPress, here’s my order of checks:

  • If an SEO plugin is active, open its sitemap and confirm the core file redirects to it or is switched off.
  • If author archives add nothing, exclude the users sitemap. On a one-author blog it’s a duplicate of the homepage feed.
  • If tag archives are thin, exclude them from the sitemap and noindex them together, so the two signals agree.
  • If “Discourage search engines from indexing this site” is ticked under Settings, Reading, core WordPress disables its sitemap. That’s usually a staging leftover.

Honestly, the most common WordPress problem isn’t the sitemap. It’s a noindexed page type that the plugin still lists, and it only shows up when you actually open the XML.

Branch Two: Shopify Stores

Shopify owners can mostly relax. The Shopify help center says every store automatically generates sitemap.xml at the root, with separate child sitemaps for products, collections, blogs and pages, and it updates itself when you add products or posts.

What you can’t do is hand-edit that file. So your best practices move one level up. Keep out-of-stock product pages you’ve deleted properly redirected, avoid publishing near-empty collections, and remember a store in private (password) mode can’t be read by Google at all.

When Does a Site Need a Sitemap Index?

Once any single file would pass 50,000 URLs or 50MB uncompressed. Google’s large sitemaps guide says one index file can list up to 50,000 sitemaps, and you can submit up to 500 index files per site.

<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <sitemap>
    <loc>https://www.example.com/sitemap-products.xml</loc>
    <lastmod>2026-09-28</lastmod>
  </sitemap>
  <sitemap>
    <loc>https://www.example.com/sitemap-posts.xml</loc>
    <lastmod>2026-09-25</lastmod>
  </sitemap>
</sitemapindex>

I split sites by content type even when they’re nowhere near the limit. Products, categories and posts each get a file. That way, when indexing drops on one template, I can see it in one row of the report instead of hunting through 30,000 mixed URLs.

One scope rule catches people on subfolder setups. An index at /public/sitemap_index.xml can only list sitemaps in that folder or deeper, unless you submit through Search Console or use cross-site submission.

Image, Video and News Sitemaps in Brief

You rarely need separate files for these. Extensions sit inside a normal sitemap, and Google’s overview lists video, image and news entries as the three types it supports.

ExtensionWhen I’d use itOne rule to remember
ImageImages load via JavaScript or live on a CDN domainUp to 1,000 image:image tags per URL
VideoVideo is central to the page and hard to discoverNest video:video inside the page’s url entry
NewsYou publish news eligible for Google NewsOnly articles from the last two days, max 1,000 news:news tags

If you run multilingual pages, a sitemap can also carry the localized versions of each URL. That’s a hreflang topic of its own, so I’ll leave it there.

The Robots.txt Sitemap Line

Add this line anywhere in robots.txt. Google says it will find it the next time it crawls the file, and there’s no limit on how many Sitemap: lines you can add.

Sitemap: https://www.example.com/sitemap_index.xml

It has to be the full absolute URL. I put it at the bottom, and I point it at the index, not at each child.

For the actual submission in Search Console, and for fixing a “Couldn’t fetch” status, I’ve written a separate step-by-step guide on how to submit a sitemap to Google Search Console. I won’t repeat it here.

How Do You Know Your Sitemap Is Working?

Submission tells Google where the file is. Monitoring tells you whether it’s clean. Google’s own overview is clear that a sitemap “doesn’t guarantee that all the items in your sitemap will be crawled and indexed,” so “Success” in the report is only the starting line.

Once a month, I do three things. I compare the Discovered pages count in the Sitemaps report against the number of pages I expect. Then I open the Page indexing report and filter it to that one sitemap. Finally, I read the “Why pages aren’t indexed” reasons for URLs I submitted.

That filter is the most useful view in Search Console for this job. If, for example, a big share of submitted product URLs sits at “Crawled, currently not indexed”, the sitemap is fine and the pages have a quality or duplication problem.

A Quick Sitemap Audit Script

Before I trust a sitemap, I check every listed URL for the problems in the table above. This short Python script does a rough version of that. It follows index files, requests each URL without following redirects, and prints anything that isn’t a clean, indexable, self-canonical 200.

import sys
import re
import xml.etree.ElementTree as ET
import requests

NS = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}
HEADERS = {"User-Agent": "sitemap-audit/1.0"}

def get_locs(sitemap_url):
    resp = requests.get(sitemap_url, headers=HEADERS, timeout=20)
    root = ET.fromstring(resp.content)
    locs = [el.text.strip() for el in root.findall("sm:*/sm:loc", NS)]
    if root.tag.endswith("sitemapindex"):
        urls = []
        for child in locs:
            urls += get_locs(child)
        return urls
    return locs

def check(url):
    r = requests.get(url, headers=HEADERS, timeout=20, allow_redirects=False)
    problems = []
    if r.status_code != 200:
        problems.append(f"status {r.status_code}")
    if "noindex" in r.headers.get("X-Robots-Tag", "").lower():
        problems.append("noindex header")
    html = r.text if r.status_code == 200 else ""
    if re.search(r'<meta[^>]+name=["\']robots["\'][^>]+noindex', html, re.I):
        problems.append("noindex meta tag")
    canon = re.search(r'<link[^>]+rel=["\']canonical["\'][^>]+href=["\']([^"\']+)', html, re.I)
    if canon and canon.group(1).rstrip("/") != url.rstrip("/"):
        problems.append(f"canonical points to {canon.group(1)}")
    return problems

if __name__ == "__main__":
    for page in get_locs(sys.argv[1])[:int(sys.argv[2]) if len(sys.argv) > 2 else None]:
        issues = check(page)
        if issues:
            print(page, "|", ", ".join(issues))

Run it with python sitemap_audit.py https://www.example.com/sitemap_index.xml 200 to check the first 200 URLs. It needs the requests library, it doesn’t unpack gzipped files, and the regex checks assume common attribute order. For a full crawl, a desktop crawler in list mode is the better tool, but this catches the obvious leaks in a minute.

The Decision Matrix

If your siteAnd the sitemapThen
Has under about 500 well-linked pagesIs generated by your CMSKeep it, clean the URL list, move on
Runs WordPress with an SEO pluginAlso exposes /wp-sitemap.xml without a redirectKeep one sitemap only
Has over 50,000 URLsIs one fileSplit by content type, submit the index
Changes content oftenBumps every lastmod dailyFix it so only real changes move the date
Lists noindexed, redirected or 404 URLsShows “Success”Remove them; “Success” only means it parsed
Relies on JavaScript-loaded images or videoHas no extensionsAdd image or video entries to the page URLs

When Should You Bring In a Technical SEO?

A sitemap problem is quick to fix when it’s a setting. Escalate when it’s a symptom: thousands of parameter URLs generated by faceted navigation, lastmod values your platform can’t control, or a migration that left half the file redirecting.

That work needs someone who can read crawl data and server logs together. It’s the kind of cleanup my team handles in our technical SEO service, and a free SEO audit is the easiest way to find out whether yours needs it.

Frequently Asked Questions

Does Google Use Changefreq and Priority?

No. Google’s sitemap documentation states that it ignores both values. Other search engines may read them, but for Google your time is better spent keeping lastmod accurate and the URL list clean.

Should Noindex Pages Be in an XML Sitemap?

No. A sitemap is a list of URLs you want shown in search, so listing a noindexed page sends two conflicting signals. Remove it from the sitemap, or remove the noindex if you actually want the page to rank.

How Many URLs Can One XML Sitemap Hold?

Up to 50,000 URLs or 50MB uncompressed, whichever comes first. Past either limit, split the file and list the parts in a sitemap index, which can itself reference up to 50,000 sitemaps.

Do Small Websites Need an XML Sitemap?

Google says a site of about 500 pages or fewer that’s well linked internally might not need one. I still add one on every site I build, because it gives you the Sitemaps report and a per-sitemap filter in the Page indexing report.

Last updated: September 2026 by Mizanur Rahman

Put this guide to work.

Want help applying it? Start with a free audit of your site. We’ll show you what to fix first.

Get a free SEO audit