How to use the Wayback Machine for competitor research

How to use the Wayback Machine to see old versions of a competitor's website: pricing history, messaging changes, removed pages, the CDX index and its limits.

Ned, founder of Figo Verified 4 October 2026 6 min read

To use the Wayback Machine, go to web.archive.org, paste a URL and pick a date on the calendar. That shows you the page as the Internet Archive's crawler saw it that day.

For competitor research it does five jobs that no paid tool can, because paid tools only start collecting the day you sign up: rebuilding a pricing history, tracing how messaging changed, finding removed pages, listing every URL a site has ever had, and showing how a company's public claims grew. All of it is free.

How to use the Wayback Machine: the basics

  1. Paste the exact page, not just the domain. competitor.com/pricing gives you the pricing history; competitor.com gives you the homepage.
  2. Pick a year on the timeline, then a day on the calendar. Each dot is a day with captures.
  3. Read the dot colours. The Internet Archive's help explains them: blue means the page loaded normally, green a redirect, orange a client error such as a 404, red a server error. You almost always want blue.
  4. Step through time with the capture navigation at the top of any archived page.

Four URL patterns save a lot of clicking:

URLWhat it does
web.archive.org/web/2024/https://competitor.com/pricingJumps to the nearest capture to that year
web.archive.org/web/*/competitor.com/pricingLists every capture date for the page
web.archive.org/web//competitor.com/Lists every archived URL on the site
web.archive.org/web/changes/https://competitor.com/pricingOpens the comparison view

The long number in an archived URL is a timestamp in the form yyyymmddhhmmss, so 20241217190139 is 17 December 2024 at 19:01:39. Add id_ straight after it, as in /web/20241217190139id_/, and you get the page as captured without the Wayback toolbar, which is the cleanest version to save.

The "Save Page Now" box at web.archive.org saves one page, once. It does not add the URL to future crawls. It is still worth using on a competitor's pricing page today, because it creates a public, dated record you can point to later.

Five competitor research jobs

1. Rebuild their pricing history

The guide to monitoring competitor pricing explains what to look for: tiers, limits and discounts as much as the headline number. The Wayback part is getting one capture per month without clicking through the calendar:

https://web.archive.org/cdx/search/cdx?url=competitor.com/pricing&collapse=timestamp:6&fl=timestamp&filter=statuscode:200

That returns one timestamp per month. Open each as web.archive.org/web/ followed by the timestamp and the URL, and screenshot the pricing table.

Three traps. Pricing tables loaded by JavaScript often appear blank or stuck loading. Sites that set prices by country show whatever version the crawler's location received. And a capture can mix dates: when an asset is missing, the Archive fills it from the nearest date it has, so check the timestamp on anything that looks out of place.

2. Map their messaging over time

Open the homepage once a quarter for the last two or three years and copy the headline and subheading into a table:

DateHeadlineWho it speaks toProof offered
Q1 2024(copy here)
Q3 2024
Q1 2025

Read it for shifts. A new audience in the headline, a new category word, or proof that changes from features to customer logos to numbers. Each one is a decision about positioning that the company made, and the order they made them in tells you what did not work.

3. Find the pages they removed

The guide to finding every page on a website shows how to list every URL the Archive holds for a domain and compare it with today's sitemap. Once you have a removed page, date its removal:

https://web.archive.org/cdx/search/cdx?url=competitor.com/integrations/acme&fl=timestamp,statuscode&limit=-10

limit=-10 returns the last ten captures. The switch from 200 to 404 or 301 brackets when the page went. A retired integration, a dropped industry page or a removed product tier is a decision the company did not announce.

4. Track their claims over time

Homepages carry numbers: "trusted by 3,000 teams", review badges, logo walls, case study counts. About and team pages carry headcount. Pull the same page once a year and record each claim.

A competitor that went from 3,000 to 8,000 claimed customers in two years has told you its growth rate. The same trick on a team page feeds the headcount methods and, through them, a revenue estimate.

5. Compare two versions side by side

Open web.archive.org/web/changes/ followed by the URL. You get a grid of captures. Select two and press Compare to see the versions side by side with added and removed text highlighted.

It compares text, not layout, so it is ideal for terms, pricing copy and feature lists, and useless for a redesign.

The CDX index, for when the calendar is too slow

Every query above uses the CDX server, the Archive's index of captures. It returns plain text, one capture per line, and takes these parameters:

ParameterWhat it doesExample
urlThe page or path to look upurl=competitor.com/pricing
matchTypeprefix returns everything under a pathmatchType=prefix
flWhich fields to returnfl=timestamp,original,statuscode
collapseDrops repeated rowscollapse=digest
filterKeeps rows matching a pattern, ! excludesfilter=statuscode:200
from and toDate range, as 1 to 14 digitsfrom=2024&to=2025
limitFirst N rows, or the last N with a minuslimit=-10
outputjson instead of plain textoutput=json

Two collapses do most of the work. collapse=digest keeps only captures whose content differed from the previous one, which lists the moments a page changed. collapse=timestamp:6 keeps one capture per month; use 4 for one per year or 8 for one per day.

Queries on large domains can be slow or time out. Narrow them with a path, a date range or a limit.

Limits to know before you trust a snapshot

Gaps. The crawler finds pages through links. Popular pages are captured often, deep pages rarely, and pages nothing links to may never be captured at all.

Exclusions. Pages blocked by robots.txt and sites whose owners requested exclusion show nothing.

JavaScript. The Archive's help says JavaScript elements are often hard to archive, especially anything that needs to contact the live server. Calculators, pricing toggles and embedded widgets break most.

Mixed dates. Missing elements are filled from the closest date the Archive has, and occasionally from the live web.

Capture is not change. A capture shows when the crawler visited, not when the page changed. Two captures bracket a change; neither dates it.

No full-text search. You can look up URLs, not words on pages.

When the archive has nothing

archive.today, also at archive.ph, is an independent archive that keeps on-demand snapshots and sometimes holds pages the Wayback Machine missed. Google's cache, which people used to use for this, was retired in 2024. The Wayback Machine alternatives roundup covers the rest.

For the future, the fix is to keep your own record from today: Save Page Now on the pages that matter each month, or a change monitor that stores screenshots.

Figo does the forward-looking half for tracked competitors: it checks their pages, new posts and published pricing every week and keeps what changed. Plans start at $49 a month for three competitors. Ours, so judge accordingly. It cannot reach back before the day you add a competitor, which is exactly what the Wayback Machine is for.

Questions people ask

Is the Wayback Machine free?

Yes. It is run by the Internet Archive, a non-profit library, and browsing captures or querying its CDX index needs no account or payment.

Why is a website missing from the Wayback Machine?

The Internet Archive's own help lists the reasons: its crawlers never found the site, the pages were password protected or blocked by robots.txt, or the owner asked for the site to be excluded.

Can a competitor delete its history from the Wayback Machine?

A site owner can ask the Internet Archive to exclude its archives, and the Archive reviews each request without promising an outcome. If a competitor's domain shows nothing at all, that may be why.

How far back does the Wayback Machine go?

To the mid-1990s for sites that existed then. How far back a particular competitor goes depends on when the site launched and how well linked it was.

See it on your own competitors

Figo checks their ads, pages, rankings and reviews every week, then tells you what to do in plain words. Set up in two minutes.

Keep reading