How to find all pages on a website, even removed ones

How to find all pages on a website using its sitemap, robots.txt, site search, a free crawler and the Wayback CDX index for pages a competitor has removed.

Ned, founder of Figo Verified 4 October 2026 7 min read

No single method lists every page. Here is how to find all pages on a website: start with its sitemap, the site's own list of what it wants found, then fill the gaps with robots.txt, a search engine, a crawler and the Wayback Machine. Each one catches something the others miss, including pages the company has since deleted.

On a competitor's site this takes about forty-five minutes, and the result is a map of their whole strategy: what they publish, which sections they are building and what they quietly removed.

How to find all pages on a website in six steps

StepSourceWhat it addsTime
1SitemapEvery page they want indexed5 min
2robots.txtHidden sitemaps, sections they hide2 min
3Google and BingIndexed pages, subdomains5 min
4A crawlerLinked pages the sitemap missed15 min
5Wayback CDXPages they removed10 min
6Traffic and adsWhich pages matter10 min

Step 1: read the sitemap

Add one of these to the domain and see which loads:

TryUsually means
/sitemap.xmlMost sites, including Shopify and Squarespace
/sitemap_index.xmlWordPress with Yoast or another SEO plugin
/wp-sitemap.xmlWordPress with no SEO plugin (built in since version 5.5)
/sitemap.xml.gzA large site serving a compressed file

Larger sites publish a sitemap index: a list of further sitemaps, often one per section. A Shopify store's index, for example, points to separate files for products, pages, collections and blogs. Google's documentation caps a single sitemap at 50,000 URLs, so very large sites have many files.

To turn the XML into a plain list, you can open it in a browser and copy, or run two lines in a terminal (Mac, Linux or Git Bash on Windows):

curl -s https://competitor.com/sitemap_index.xml | grep -o '<loc>[^<]*' | sed 's/<loc>//' > sitemaps.txt
while read s; do curl -s "$s" | grep -o '<loc>[^<]*' | sed 's/<loc>//'; done < sitemaps.txt > pages.txt

The first line lists the child sitemaps, the second pulls every page URL from each into pages.txt. If the site has a single sitemap rather than an index, the first line on its own gives you the pages.

What the sitemap leaves out: pages the site does not want indexed (thank-you pages, many paid landing pages), pages somebody forgot to include, and on badly maintained sites it may still list pages deleted long ago.

Step 2: read robots.txt

Open competitor.com/robots.txt. Two kinds of line matter.

Sitemap lines. They point to sitemaps at addresses you would never guess. Shopify's own site, for example, declares its sitemap at /sitemaps_list.xml. Some sites list a second sitemap on another subdomain, often where a separate content system serves the blog.

Disallow lines. Most are platform boilerplate: /cart, /checkout and /account on every Shopify store. The custom ones are the interesting ones. A line disallowing /lp/ or /partners/ names a section the company would rather search engines skipped, which usually means paid landing pages or a partner programme. Lines naming AI crawlers such as GPTBot show where the company stands on AI search, and which AI crawlers to allow explains what each one does.

Step 3: ask the search engines

Search site:competitor.com in Google. Google's own help says the site: operator "doesn't necessarily return all the URLs that are indexed", so treat the count as a rough indication. Its value is the variants:

  • site:competitor.com -inurl:blog shows everything except the blog.
  • site:competitor.com intitle:vs finds comparison pages, including any about you.
  • site:competitor.com -site:www.competitor.com surfaces subdomains: help centres, docs, regional sites.

Run the same searches in Bing. It indexes differently and sometimes shows pages Google omits.

Step 4: crawl it

A crawler starts at the homepage and follows every link, which finds pages that are linked but missing from the sitemap. Screaming Frog's free version crawls up to 500 URLs, which covers the marketing pages of most small and mid-sized competitors. Most configuration options are locked in the free version, and the paid licence is $279 a year, read from Screaming Frog's own pricing page on 4 October 2026.

Compare the crawl with the sitemap. Pages in the crawl but not the sitemap are often recent additions nobody registered. Pages in the sitemap but unreachable by links are orphans, which a crawler can never find on its own.

Crawl slowly and only public pages. You are a guest on their server.

Step 5: find pages they removed

This is the step that turns a page list into competitive research. The Wayback Machine keeps an index of every URL it has captured on a domain, and the CDX query returns it as a plain list:

https://web.archive.org/cdx/search/cdx?url=competitor.com/&matchType=prefix&collapse=urlkey&fl=original&filter=statuscode:200&filter=mimetype:text/html

matchType=prefix returns everything under the domain, collapse=urlkey lists each URL once, fl=original returns just the address, and the two filters keep only HTML pages that loaded successfully. For a browser view, the Internet Archive's own help suggests web.archive.org//competitor.com/. If a page you need was never captured there, the Wayback Machine alternatives are worth a second look.

Expect noise: tracking parameters, fragments and oddly encoded URLs. Strip anything with a question mark first. Then compare with today's sitemap. A URL in the archive list but not in the sitemap was removed or renamed, and curl -sI on it tells you which: a 404 means gone, a 301 means moved.

Removals are often the most revealing thing on the map. A deleted product page, a retired integration, a dropped industry page or a pricing tier that vanished are decisions the company made and did not announce. The Wayback Machine guide covers what to do with the old versions.

Step 6: find the pages that matter

A list of 900 URLs is not a priority. Three quick ways to rank it.

Traffic. Semrush's free Website Traffic Checker lists a site's top pages by estimated search traffic. Ahrefs' Top pages report does the same in its paid tool. A handful of pages usually carry most of a site's search traffic.

Ads. Open the competitor's ads in the public ad libraries and note where they send people. Paid landing pages are often missing from the sitemap and hidden from search, and they show what the company is willing to pay to promote. The guide to seeing every ad a competitor runs shows where to look.

Their own shortlist. Some sites now publish /llms.txt, a curated list of their key pages for AI tools. If it exists, it is the company telling you which pages it considers most important. There is more on the format in our llms.txt explainer.

What you will still miss

  • Pages behind a login, including the product itself.
  • Pages on other domains: a help centre on a support platform, careers on a hiring platform, docs on a separate host.
  • Unlinked landing pages with no ads, no archive history and no search presence.
  • Parameter variants, such as filtered category pages, which are rarely worth mapping anyway.

Reading the map

Count pages by top-level folder:

cut -d/ -f4 pages.txt | sort | uniq -c | sort -rn

That one line gives you the shape of their site: 300 help articles, 120 blog posts, 40 location pages, 12 comparison pages. The shape is their strategy stated more honestly than their marketing ever would. The competitor content audit walks through reading it properly.

Watch for new top-level folders above all. A folder that did not exist last quarter, such as /industries/ with a dozen pages, is a new initiative you have caught early.

Keeping the map current

Save pages.txt with the date in the filename. Next month, rebuild it and compare:

sort -u pages-2026-09.txt > old.txt
sort -u pages-2026-10.txt > new.txt
comm -13 old.txt new.txt

comm -13 prints pages that are new this month; swap it for comm -23 to see pages that disappeared. Five minutes a month, and nothing they publish gets past you.

If you would rather not run it yourself, Figo reads each tracked competitor's sitemap every week and reports new pages and blog posts in the Monday briefing, alongside their rankings, ads and reviews. Plans start at $49 a month for three competitors. Ours, so judge accordingly. For a one-off map of a single site, the free steps above are all you need.

Questions people ask

How many pages does a website have?

Count the URLs in its sitemap for the pages it wants found, then compare with a site search and a crawl. The sitemap figure is usually the most honest single number, and the other two tell you what it leaves out.

Is it legal to crawl a competitor's website?

Reading public pages is what search engines do every day, but check the site's terms and robots.txt, crawl slowly, and never try to reach pages behind a login. A polite crawl of a few hundred pages is ordinary practice.

Can I find hidden pages that are not linked anywhere?

Only if they left a trace somewhere: the sitemap, an archived capture, an ad pointing at them, or a search engine that found them. A page with no links and no history cannot be discovered from outside.

What if the sitemap is enormous?

Large sites split it into an index of smaller files by section, with up to 50,000 URLs each. Download only the sections you care about, such as pages, products or posts, rather than the whole thing.

See it on your own competitors

Figo checks their ads, pages, rankings and reviews every week, then tells you what to do in plain words. Set up in two minutes.

Keep reading