Home / Blog / How to Get a Competitor Topical Map and Find Their Topic Gaps

How to Get a Competitor Topical Map and Find Their Topic Gaps

A competitor topical map is a reconstruction of how a rival site organizes its content: every topic it covers, the pages that belong to each topic, and which of those pages actually earn search traffic. Here is how to get a competitor topical map in…

How to Get a Competitor Topical Map and Find Their Topic Gaps

A competitor topical map is a reconstruction of how a rival site organizes its content: every topic it covers, the pages that belong to each topic, and which of those pages actually earn search traffic. Here is how to get a competitor topical map in short: list all their URLs from the XML sitemap (or site: searches if there isn’t one), add traffic estimates from a tool like Ahrefs or Semrush, cluster the URLs into topics, then compare those clusters with your own site to find the gaps worth taking.

This post covers the reverse-engineering part only. If you need to build your own map from scratch, start with my guide on how to create a topical map for SEO, then come back here to pressure-test it against the sites you compete with.

Why Reverse-Engineer a Competitor’s Topical Map at All?

Because their site is a finished experiment you didn’t have to pay for. A competitor that has published for 3 years has already found out which topics pull traffic and which ones die quietly, and their URL list shows you the result.

In my experience, the value isn’t copying their map. It’s seeing the shape of it. When I pull a competitor’s URLs for a client, the first useful finding is usually a cluster they’ve covered in 40 posts that the client has touched in two, or a cluster they abandoned after 5 posts because nobody searched for it.

Which Route Fits Your Competitor?

Four things decide how you’ll get the data. Check them before you open any tool, because each one changes the path.

  1. Does the competitor publish an XML sitemap? Most WordPress, Shopify and Wix sites do. If yes, you get a near-complete URL list in 10 minutes. If no, you rebuild it from search results and a crawler.
  2. How big is the site? Under 500 URLs, a free crawler handles everything. Past a few thousand, you’ll need filters and a spreadsheet you trust.
  3. Do you have a paid SEO tool? Without one, you get the structure but not the traffic. With one, you see which clusters matter.
  4. What’s your goal? Finding gaps needs their topics compared with yours. Copying a site architecture needs their folder structure and internal links.

How to Get a Competitor Topical Map: Start With the Sitemap

For most competitors, the sitemap is the fastest and most honest source. It’s the list of URLs the site owner wants Google to find, which is close to a list of pages they consider important.

Step 1: Find the Sitemap

Open competitor.com/robots.txt first. Google tells site owners they can add a Sitemap: line anywhere in that file, so many sites list theirs there. If robots.txt says nothing, try the usual locations: /sitemap.xml, /sitemap_index.xml (Yoast and Rank Math) and /wp-sitemap.xml (WordPress core).

Big sites split their sitemaps. Google’s sitemap documentation caps a single file at 50,000 URLs or 50MB uncompressed, so larger sites use a sitemap index that points to child files. Those child file names are a gift: post-sitemap.xml, product-sitemap.xml and category-sitemap.xml already tell you how the site divides its content.

Step 2: Pull Every URL Into a Spreadsheet

Screaming Frog’s SEO Spider can crawl an XML sitemap directly, and the free version handles up to 500 URLs per crawl. For bigger sites, or if you’d rather not install anything, this short Python script walks a sitemap index, collects every URL with its lastmod date, and saves a CSV. It waits 1 second between requests so you don’t hammer anyone’s server.

import csv
import gzip
import time
import xml.etree.ElementTree as ET
from collections import Counter
from urllib.parse import urlparse

import requests

NS = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}
HEADERS = {"User-Agent": "Mozilla/5.0 (compatible; sitemap-research/1.0)"}

def fetch_xml(url):
    resp = requests.get(url, headers=HEADERS, timeout=30)
    resp.raise_for_status()
    data = resp.content
    if data[:2] == b"\x1f\x8b":  # gzipped sitemap file
        data = gzip.decompress(data)
    return ET.fromstring(data)

def collect_urls(sitemap_url, seen=None):
    seen = set() if seen is None else seen
    if sitemap_url in seen:
        return []
    seen.add(sitemap_url)
    root = fetch_xml(sitemap_url)
    time.sleep(1)  # one request per second
    rows = []
    if root.tag.endswith("sitemapindex"):
        for loc in root.findall("sm:sitemap/sm:loc", NS):
            rows.extend(collect_urls(loc.text.strip(), seen))
    else:
        for node in root.findall("sm:url", NS):
            loc = node.find("sm:loc", NS)
            lastmod = node.find("sm:lastmod", NS)
            rows.append({
                "url": loc.text.strip() if loc is not None else "",
                "lastmod": lastmod.text.strip() if lastmod is not None else "",
                "sitemap": sitemap_url,
            })
    return rows

if __name__ == "__main__":
    start = "https://competitor.example/sitemap_index.xml"
    rows = collect_urls(start)
    with open("competitor-urls.csv", "w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=["url", "lastmod", "sitemap"])
        writer.writeheader()
        writer.writerows(rows)
    print(f"Saved {len(rows)} URLs")

    folders = Counter(
        urlparse(r["url"]).path.strip("/").split("/")[0] or "(home)" for r in rows
    )
    for folder, count in folders.most_common(20):
        print(f"{count:>6}  /{folder}/")

Change the start URL, run it, and you get a CSV plus a count of URLs per top-level folder. It needs only the requests library (pip install requests).

Step 3: Read the Shape Before You Cluster

If the folder count shows clear sections like /blog/, /guides/ and /products/, then those folders are the competitor’s own first-level map. Start there. If almost every URL sits at the root, as on many WordPress blogs with /post-name/ permalinks, then the folder count tells you nothing and you’ll cluster by slug and title instead (Step 4 below).

The lastmod column matters too. A cluster where nothing has changed in 2 years is a cluster they’ve stopped investing in, which makes it easier to beat.

What If the Sitemap Is Missing or Useless?

Some sites block their sitemap, leave it broken, or list only 20 pages out of 800. In that case you rebuild the list from two sources, and you accept that it’ll be incomplete.

Start with Google. A site:competitor.com search shows pages Google has indexed, and site:competitor.com/blog/ narrows it to one section. Google is honest about the limits here: its site: operator documentation says the results aren’t exhaustive and bigger sites shouldn’t expect to see every URL. Use it to discover sections, not to count pages.

Then crawl. Point a crawler at the homepage and let it follow internal links. This finds what the site links to, which is often more revealing than the sitemap, because pages that get lots of internal links are the ones the owner treats as pillars.

If the site renders its links with JavaScript, then switch your crawler to JavaScript rendering, or you’ll see a homepage and nothing else. If the crawl explodes into thousands of filter URLs with ?color= and ?sort= parameters, then exclude parameters and crawl categories only.

How Do You See Which Topics Actually Earn Traffic?

A URL list shows what they published. It doesn’t show what works. For that you need traffic estimates, and there’s no free way to get them for someone else’s site.

You can’t see a competitor’s Google Search Console. GSC only shows data for properties you’ve verified, so any “GSC export” of a competitor is really a third-party estimate. The two most common sources are Ahrefs Site Explorer’s Top pages report, which ranks a site’s URLs by estimated organic traffic, and Semrush’s Organic Research, which has a Pages report and a Topics report (the Topics report needs a Guru or Business plan, according to Semrush’s own help docs).

Export the page-level report and join it to your sitemap CSV by URL. Now every row has a URL, a lastmod date, estimated traffic and the number of keywords it ranks for. Treat the numbers as directional. They’re modelled estimates, so I compare clusters against each other rather than trusting any single page’s figure.

For your own side of the comparison, use real data. The Search Console interface exports up to 1,000 rows, and the Search Console API goes further, which matters if your site is large. My walkthrough on Google Search Console keyword research covers the filters I use.

How Do You Turn a URL Dump Into a Topical Map?

This is the part most tutorials skip. A list of 600 URLs isn’t a map until each one sits in a cluster and each cluster has a job.

I work through it in three passes. First, the folder pass: anything under /guides/standing-desks/ goes in one bucket without discussion. Second, the slug pass: sort the remaining URLs alphabetically and group slugs that share a head term (standing-desk-height, standing-desk-mat, standing-desk-vs-sitting). Third, the intent pass: pull the page titles and split any bucket where the intent differs, because a buying guide and a troubleshooting post don’t belong in the same cluster even if they share words.

If you have thousands of URLs, cluster the titles automatically before the manual check. My post on semantic keyword clustering explains which method fits which list. Honestly, for most competitor sites under 1,000 URLs, a sorted spreadsheet and 2 hours of focus beat any tool.

Here’s what the output looks like for an invented office furniture competitor. The URLs and figures are made up to show the format.

URLClusterRoleEst. trafficLast updated
/standing-desks/Standing desksPillar (category)HighRecent
/standing-desk-height/Standing desksClusterHighRecent
/standing-desk-vs-sitting/Standing desksClusterMediumOld
/ergonomic-chair-guide/Ergonomic chairsPillarMediumOld
/chair-armrest-height/Ergonomic chairsClusterLowOld
/office-plants/Office decorOrphan clusterNoneOld

Mark the role column from internal links, not from gut feel. A page that dozens of other pages link to, with a broad title, is a pillar. A page with no internal links and no traffic is usually an experiment they gave up on.

How Do You Find the Topic Gaps Worth Taking?

Put your clusters and theirs side by side. Every competitor cluster lands in one of four buckets, and each bucket gets a different action.

  • They cover it, you don’t, and it earns traffic. This is your main gap list. Check the search intent and plan a cluster.
  • You both cover it, but they go deeper. Count pages per cluster. If they have 15 and you have 3, list the subtopics they answer that you skip.
  • They cover it and it earns nothing. Leave it. Copying their dead clusters is the most common mistake I see in competitor research.
  • Nobody covers it. Check real demand in your own Search Console and in keyword tools. Sometimes it’s a genuine opening; sometimes there’s simply no search.

This is topic-level gap analysis, not a keyword-by-keyword comparison. The keyword version has its place, and I walk through it in how to run a keyword gap analysis, but at the planning stage I want to know which clusters to build, and the page counts tell me that faster.

Once the gap list exists, rank it. I score each gap cluster on two things: how much traffic the competitor’s pages in that cluster pull, and how many pages you’d need to cover it properly. A cluster that earns well from 4 pages goes to the top. A cluster that needs 25 pages to match them goes lower, unless it sits right next to a service you sell.

That second condition overrides the first more often than people expect. A modest cluster that feeds a money page will usually do more for a small business than a big informational cluster with no path to a sale. When I review a gap list with a client, this is the conversation that takes the longest, and it should.

I’d also run the comparison against 3 competitors, not one. One site’s map reflects one team’s habits. A topic that shows up as a real cluster on all 3 sites is a topic Google clearly expects a complete site in your niche to cover.

Edge Cases That Change the Process

  • Ecommerce giants: a store with 50,000 product URLs will drown your spreadsheet. Pull only the category and blog sitemaps, and treat products as one bucket per category.
  • Multilingual sites: filter to one language folder (/en/, /de/) first, or every cluster appears two or three times.
  • Tag and archive pages: WordPress /tag/ and /author/ URLs sometimes appear in sitemaps. Drop them. They’re listings, not topics.
  • Brand-new competitors: a site with 30 posts has no traffic history worth trusting. Read its structure for ideas, but don’t copy its priorities.
  • Big brands that rank on authority: if a household name ranks with thin pages, their map tells you little about what a smaller site needs to rank.

Decision Matrix: Which Method to Use

If the competitor…And you have…Then use…
Has a clean XML sitemapNo paid toolSitemap script + manual clustering, then validate demand in your own GSC
Has a clean XML sitemapAhrefs or SemrushSitemap CSV joined to a Top pages or Pages export
Has no sitemap or a broken oneNo paid toolsite: searches by folder + a crawl from the homepage
Has no sitemap, 5,000+ URLsA paid crawler or SEO toolA rendered crawl with parameters excluded + traffic export
Is a big brand or brand-newAnythingRead their structure for ideas only; weight traffic data from mid-size rivals instead

When to Hand It Off

If you’re comparing 3 competitors with a few thousand URLs each, this becomes a multi-day job, and the clustering pass is where tired eyes make bad calls. Our keyword research service includes topical maps built this way, with the competitor comparison done for you. If you’d rather start with a quick look at your own site first, you can request a free SEO audit.

Frequently Asked Questions

Can I Get a Competitor’s Topical Map for Free?

Yes, the structure part is free. Their XML sitemap, site: searches and a free crawler for up to 500 URLs give you the full list of topics they cover. What you can’t get for free is reliable traffic data per page, so you’ll know what they published but not which clusters actually pull visitors.

Is There a Tool That Builds a Competitor Topical Map Automatically?

Some SEO suites group a competitor’s pages into topics for you, such as the Topics report inside Semrush’s Organic Research on its higher plans. I use these as a starting point, but I still check the clusters by hand, because automatic grouping often lumps different search intents together.

How Many Competitors Should I Analyze?

Three is my default. One competitor gives you one team’s opinions; three shows you which clusters the whole niche treats as essential. Pick sites that rank for your main queries and are roughly your size, not the biggest brand in the space.

How Often Should I Redo a Competitor Topical Map?

Every 6 months is enough for most niches. Re-run the sitemap script, compare the new URL list with the old one, and look at what they added. New clusters from a competitor are an early signal of where they think demand is growing.

Last updated: September 2026 by Mizanur Rahman

Put this guide to work.

Want help applying it? Start with a free audit of your site. We’ll show you what to fix first.

Get a free SEO audit