← Back to blog

Published

Stop Re-Scraping Pages That Haven't Changed: A Practical Guide to Incremental Scraping

Take your scraper's output from yesterday and today and count how many records are different. For most monitoring jobs (product catalogs, job boards, listings, documentation, competitor pages), it's a few percent at most. Everything else you downloaded today, and paid proxy bandwidth, API requests, rendering time, and possibly LLM tokens for, was a copy of data you already had.

Incremental scraping fixes that by fetching and processing only what's new or changed. Most scraping guides focus on getting the page in the first place: avoiding blocks, rotating proxies, choosing a browser. This one is about deciding whether to fetch it at all. There are four techniques, and they work best together:

  1. Conditional requests let the server tell you a page hasn't changed.
  2. Change feeds (sitemaps, RSS, site APIs) tell you which pages changed.
  3. Field-level fingerprints let you decide whether anything you care about changed.
  4. Adaptive scheduling checks each page about as often as it actually changes.

Which technique saves money depends on how you pay

Each technique cuts a different cost, so start with the one that matches how you're billed:

How you payExampleWhat saves money
Per GBResidential or mobile proxy bandwidthConditional requests (a 304 is a few hundred bytes; the page might be 150 KB)
Per requestScraping APIs, unblockers, SERP APIsMaking fewer requests, through scheduling and trusted change feeds. A 304 still counts as a request.
Per page processedHeadless rendering, LLM extraction, downstream pipelinesFingerprints, so unchanged pages skip processing

If you pay per request, conditional requests save you very little and scheduling is where the savings are. If you run an LLM over every page to extract structured data, fingerprinting pays for itself on the first day.

1. Conditional requests with ETag and Last-Modified

HTTP has had a built-in way to ask "has this changed since I last looked?" for decades. When a server returns a page, it may include one or both validators:

HTTP/1.1 200 OK
ETag: "5f2c-61a8e3b1"
Last-Modified: Tue, 29 Sep 2026 14:02:11 GMT

Store them, and send them back on your next request:

GET /products/widget HTTP/1.1
If-None-Match: "5f2c-61a8e3b1"
If-Modified-Since: Tue, 29 Sep 2026 14:02:11 GMT

If nothing changed, the server answers 304 Not Modified with an empty body. You skip the download, the parse, and the processing, and the site does less work as well. When both headers are present, If-None-Match takes precedence (RFC 9110 §13.2.2), but send both anyway because some servers only implement one.

It's the cheapest optimization in scraping, with some real limits:

  • Many dynamic pages don't participate. Plenty of application servers send no validators, or send Last-Modified set to the current time on every response. Those pages will never return 304.
  • ETags can differ between backend servers. If each server behind a load balancer computes its own ETag, the page will look changed whenever a different server answers.
  • A 304 is still a request. It saves bandwidth, but you're still billed for the request if you pay per request.
  • Validators track bytes. A page with a rotating ad slot or a "12 people are viewing this" counter will never return 304, even when the price and description are identical. Fingerprints handle that case.

So send validators on every request, confirm changes with your own fingerprint, and measure the 304 rate per host. If a host still hasn't returned a single 304 after a week, its server doesn't support them, and the other three techniques will have to do the work there.

HEAD requests are often suggested as an alternative, and they're rarely worth it. A HEAD costs a full request, many dynamic servers render the whole page to answer it anyway, and you still need a GET whenever something looks different. A conditional GET gives you the answer and the new content in one round trip.

2. Change feeds: sitemaps, RSS, and site APIs

Conditional requests tell you whether one page changed. A change feed tells you which pages changed, so you can skip asking about the rest. Many sites publish one:

  • XML sitemaps with <lastmod>. One sitemap file covers up to 50,000 URLs, and a sitemap index links to as many files as a site needs. Pages whose lastmod is newer than your last check move to the front of the queue. Google says it only uses lastmod when the value is consistently and verifiably accurate, and it ignores changefreq and priority entirely. That's a sensible policy for a scraper too.
  • RSS and Atom feeds. Blogs, news sites, job boards, and changelogs often list their most recent items with timestamps.
  • The site's own JSON. CMS and storefront platforms often expose update times. The WordPress REST API can sort posts by modified, and many Shopify storefronts serve /products.json with an updated_at on every product. Listing pages sorted by "newest" or "recently updated" work the same way. Check the site's terms and robots rules before relying on any of these; our legal guide covers the questions to ask.
  • Search engines. A site: query with a recency filter asks Google what it has recently indexed on a domain. The dates are Google's estimates and can be rough, but it can surface new URLs that nothing you crawl links to.

Use feeds to pull a check earlier, and keep your own schedule as the backstop. A sitemap that says a page hasn't changed since 2023 might be right, or the site might simply never update its lastmod. Over time you'll see which hosts' feeds can be trusted: if a feed reliably flags the changes your fingerprints find, you can raise the maximum interval for that host and lean on the feed. That's where the request savings from feeds come from. If a feed never predicts a change your fingerprints detect, ignore it.

3. Change detection with field-level fingerprints

Once you've downloaded a page, you still have to decide whether it changed. Hashing the raw HTML and comparing it to the last hash almost never works, because the HTML differs between two requests made seconds apart: CSRF tokens, nonces, cache-busting query strings, timestamps, rotating recommendations, A/B test markers, ad slots.

Fingerprint the data you extract instead:

  1. Run your extractor and get the fields you actually use, such as title, price, stock status, and description.
  2. Serialize them canonically (sorted keys, normalized whitespace) so equal data always produces equal bytes.
  3. Hash the result and compare it to the stored hash.

This gives you a business-level definition of "changed": the price moved, the item went out of stock, the job listing was edited. A page redesign that leaves every field intact doesn't register as a change, and a one-cent price drop does.

It also makes the rest of the pipeline cheaper. An unchanged fingerprint means no database write, no re-index, no re-embedding, and no LLM call. If you use LLM-based extraction (see HTML to JSON), add a cheap first step: hash the cleaned text of the main content area, and only run the expensive extraction when that hash changes.

One caution, which ties into silent scraper failures: fingerprinting can hide a broken parser. If a redesign makes every selector return nothing, the fingerprint of "all fields empty" is the same on every visit, and the pipeline reports the page as unchanged forever. Treat an extraction with no fields filled as an error.

4. Adaptive recrawl scheduling

If you pay per request, scheduling is the biggest lever. A fixed "scrape everything daily" schedule gets it wrong in both directions. It re-downloads pages that change once a year 365 times, and it catches only one of the day's changes on pages that update hourly.

The simplest scheduler that works is multiplicative:

  • Every page starts with a default interval, such as 24 hours.
  • A check that finds a change halves the interval.
  • A check that finds nothing grows it by 50%.
  • The interval stays between a floor (say 1 hour) and a ceiling (say 14 days).

After a few weeks, stable pages sit at the ceiling and volatile pages settle at an interval of roughly half their typical time between changes. The ceiling guarantees every page gets looked at occasionally, which catches deleted pages, broken parsers, and changes your feeds missed.

What crawler research says about recrawl frequency

Search engines have studied recrawl scheduling since the early 2000s, and two results are directly useful for scrapers.

Counting changes underestimates how often pages change. If you check a page daily and see a change on 5 of 10 checks, it's tempting to conclude it changes every other day. But two changes between the same pair of checks look like one. Cho and Garcia-Molina's paper "Estimating Frequency of Change" gives a simple correction. With n checks at a fixed interval, x of which found a change:

import math


def estimated_changes_per_interval(n: int, x: int) -> float:
    # Cho & Garcia-Molina (2003): less biased than x / n when changes can happen between checks
    return -math.log((n - x + 0.5) / (n + 0.5))

Five changes in 10 daily checks works out to about 0.65 changes per day, against the naive 0.5. Nine in 10 works out to about 1.95 per day, so the page changes roughly twice a day and you're seeing about half of those changes. The formula assumes equal intervals between checks. To use it for a page your scheduler has been adjusting, pin that page to a fixed interval for a couple of weeks first.

On a fixed budget, checking fast-changing pages more often can reduce overall freshness. In "Synchronizing a Database to Improve Freshness", the same authors showed that if your goal is to maximize the share of your copies that are up to date, visiting pages in proportion to how often they change does worse than visiting every page at the same rate. A page that changes every ten minutes goes stale again almost immediately after each visit, so extra checks on it buy very little. The same checks spent on slower pages keep them accurate for days.

How you apply this depends on what you're building:

  • For an accurate snapshot (a catalog mirror, a search index, a RAG corpus), cap how often any single page gets checked and spend the budget on coverage.
  • For catching events (every price change, every new job listing), each missed change has a cost, so volatile pages do deserve more checks. Weight them by what a missed change would cost you.

Most real systems need both: a floor and ceiling per page, plus a priority weight for the pages that matter most to the business.

Incremental scraping in Python: a complete example

Here's a compact implementation of all four techniques using requests, BeautifulSoup, and SQLite. It reads a sitemap, moves up any page whose lastmod is newer than its last check, sends conditional requests, fingerprints the extracted fields, and reschedules each page based on what it found. It needs Python 3.11 or later (for datetime.fromisoformat on sitemap dates) and pip install requests beautifulsoup4.

import hashlib
import json
import sqlite3
import time
from datetime import datetime, timezone
from xml.etree import ElementTree

import requests
from bs4 import BeautifulSoup

HOUR, DAY = 3600, 86400
MIN_INTERVAL = 1 * HOUR
MAX_INTERVAL = 14 * DAY
FULL_REFETCH_EVERY = 30 * DAY  # skip validators now and then, in case they're wrong
SITEMAP_NS = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}

db = sqlite3.connect("incremental.db")
db.executescript("""
CREATE TABLE IF NOT EXISTS pages (
    url TEXT PRIMARY KEY,
    etag TEXT,
    last_modified TEXT,
    fingerprint TEXT,
    interval REAL NOT NULL DEFAULT 86400,
    next_due REAL NOT NULL DEFAULT 0,
    last_checked REAL NOT NULL DEFAULT 0,
    last_full_fetch REAL NOT NULL DEFAULT 0,
    checks INTEGER NOT NULL DEFAULT 0,
    changes INTEGER NOT NULL DEFAULT 0
);
CREATE TABLE IF NOT EXISTS snapshots (
    url TEXT NOT NULL,
    captured_at REAL NOT NULL,
    record TEXT NOT NULL
);
""")


def extract(html: str) -> dict:
    """Pull out only the fields you care about. Everything else is noise."""
    soup = BeautifulSoup(html, "html.parser")

    def text(selector: str) -> str | None:
        el = soup.select_one(selector)
        return " ".join(el.get_text().split()) if el else None

    return {
        "title": text("h1"),
        "price": text("[itemprop=price]"),
        "availability": text("[itemprop=availability]"),
    }


def fingerprint(record: dict) -> str:
    canonical = json.dumps(record, sort_keys=True, separators=(",", ":"))
    return hashlib.sha256(canonical.encode()).hexdigest()


def reschedule(url: str, changed: bool, now: float) -> None:
    (interval,) = db.execute("SELECT interval FROM pages WHERE url = ?", (url,)).fetchone()
    # Changed: come back twice as soon. Unchanged: wait 50% longer next time.
    interval = interval / 2 if changed else interval * 1.5
    interval = max(MIN_INTERVAL, min(MAX_INTERVAL, interval))
    db.execute(
        """UPDATE pages SET interval = ?, next_due = ?, last_checked = ?,
               checks = checks + 1, changes = changes + ?
           WHERE url = ?""",
        (interval, now + interval, now, int(changed), url),
    )


def check(session: requests.Session, url: str) -> str:
    etag, last_modified, old_fp, last_full = db.execute(
        "SELECT etag, last_modified, fingerprint, last_full_fetch FROM pages WHERE url = ?",
        (url,),
    ).fetchone()
    now = time.time()

    headers = {}
    if now - last_full < FULL_REFETCH_EVERY:
        if etag:
            headers["If-None-Match"] = etag
        if last_modified:
            headers["If-Modified-Since"] = last_modified

    resp = session.get(url, headers=headers, timeout=30)

    if resp.status_code == 304:
        reschedule(url, changed=False, now=now)
        return "not-modified"

    resp.raise_for_status()  # 404, 410, and 5xx are handled by the caller

    record = extract(resp.text)
    if not any(record.values()):
        # A parser that matches nothing would otherwise look "unchanged" forever.
        raise ValueError(f"extractor matched nothing on {url}")

    new_fp = fingerprint(record)
    if new_fp != old_fp:
        db.execute(
            "INSERT INTO snapshots (url, captured_at, record) VALUES (?, ?, ?)",
            (url, now, json.dumps(record)),
        )
    db.execute(
        """UPDATE pages SET etag = ?, last_modified = ?, fingerprint = ?, last_full_fetch = ?
           WHERE url = ?""",
        (resp.headers.get("ETag"), resp.headers.get("Last-Modified"), new_fp, now, url),
    )

    if old_fp is None:  # first visit: a baseline, not a change
        reschedule(url, changed=False, now=now)
        return "new"
    changed = new_fp != old_fp
    reschedule(url, changed=changed, now=now)
    return "changed" if changed else "unchanged"


def parse_lastmod(value: str) -> float:
    dt = datetime.fromisoformat(value.strip())  # Python 3.11+ accepts "Z" and date-only values
    if dt.tzinfo is None:
        dt = dt.replace(tzinfo=timezone.utc)
    return dt.timestamp()


def sync_sitemap(session: requests.Session, sitemap_url: str) -> None:
    resp = session.get(sitemap_url, timeout=30)
    resp.raise_for_status()
    root = ElementTree.fromstring(resp.content)

    for loc in root.findall("sm:sitemap/sm:loc", SITEMAP_NS):  # sitemap index
        sync_sitemap(session, loc.text.strip())

    now = time.time()
    for node in root.findall("sm:url", SITEMAP_NS):
        url = node.findtext("sm:loc", namespaces=SITEMAP_NS).strip()
        db.execute("INSERT OR IGNORE INTO pages (url, next_due) VALUES (?, ?)", (url, now))

        lastmod = node.findtext("sm:lastmod", namespaces=SITEMAP_NS)
        if lastmod:
            # lastmod can move a check earlier, never later.
            db.execute(
                "UPDATE pages SET next_due = MIN(next_due, ?) WHERE url = ? AND last_checked < ?",
                (now, url, parse_lastmod(lastmod)),
            )
    db.commit()


def run_due(session: requests.Session, limit: int = 500) -> None:
    due = db.execute(
        "SELECT url FROM pages WHERE next_due <= ? ORDER BY next_due LIMIT ?",
        (time.time(), limit),
    ).fetchall()
    for (url,) in due:
        try:
            print(check(session, url), url)
        except (requests.RequestException, ValueError) as exc:
            # Retry in an hour without touching the page's learned interval.
            db.execute("UPDATE pages SET next_due = ? WHERE url = ?", (time.time() + MIN_INTERVAL, url))
            print("error", url, exc)
    db.commit()


if __name__ == "__main__":
    session = requests.Session()
    session.headers["User-Agent"] = "example-monitor/1.0 ([email protected])"
    sync_sitemap(session, "https://shop.example.com/sitemap.xml")
    run_due(session)

Run it from cron every few minutes, and each run does only the work that's due. A few details worth pointing out:

  • FULL_REFETCH_EVERY drops the validators once a month, so a server that wrongly answers 304 can't freeze a page forever.
  • The first visit sets a baseline. It's saved as a snapshot but isn't counted as a change, so it doesn't shorten the page's interval or skew its change count.
  • Errors back off without changing the schedule. A failed check is retried an hour later and shows up in your log. A page that keeps failing doesn't get hammered every few minutes, and its learned interval is kept for when it recovers.
  • Snapshots are written only when something changed, so the history table doubles as a change log you can query directly, such as "every price change on this product since June."
  • checks and changes are stored per page so you can spot volatile pages and pin them to a fixed interval for the change-rate estimator above.

Swap extract for your real fields, or for an LLM extraction call behind a text-hash check, and the rest stays the same. For gzipped sitemaps (sitemap.xml.gz), decompress the response with gzip.decompress before parsing.

What incremental scraping saves: a worked example

Here are rough numbers for a common job: monitoring 100,000 product pages daily through a residential proxy billed per GB. Assume an average transfer of 150 KB per page and that 5% of pages change on a given day. These are illustrative assumptions, so substitute your own.

  • Naive: 100,000 full downloads a day, about 15 GB/day, and 100,000 records sent to the rest of your pipeline.
  • With fingerprints: still 15 GB/day, but only about 5,000 changed records go downstream. If the expensive part of your pipeline comes after extraction (re-indexing, embeddings, or an LLM step behind a text-hash check), that's about 95% less processing.
  • Plus conditional requests: if the host returns 304 for 70% of unchanged pages, about 66,500 of the 95,000 unchanged checks become tiny 304 responses. Full downloads drop to about 33,500, or about 5 GB/day, a two-thirds cut in bandwidth. On a host that never returns 304, this step saves nothing, which is why you measure it per host.
  • Plus adaptive scheduling: product pages that haven't changed in weeks drift out to the 14-day ceiling and get checked one-fourteenth as often. The total saving depends on how uneven your catalog's change rates are. In most catalogs, a small share of pages accounts for most of the changes, so daily checks fall sharply. This is also the only step that helps if you pay per request.

None of these steps make the scraper itself any smarter. They cut the work you do on data that hasn't changed.

Incremental scraping for search results

Search engine results pages are a special case. Google results don't come with an ETag, there's no sitemap for "what changed in the results for this keyword," and every check is a full request. If you track rankings, AI Overview citations, or news coverage for a keyword, conditional requests don't apply, so fingerprints and scheduling do most of the work.

Fingerprint the part of the results you act on. For rank tracking, that's usually the ordered list of top-10 URLs. For brand monitoring, it may just be whether your domain appears at all. Ads, related searches, and snippet wording change all the time and usually don't matter to you.

Schedule queries the same way you schedule pages. A keyword whose top 10 hasn't moved in three weeks doesn't need a daily check, and a volatile one may need two. With PrismCrawl and its Python SDK, the scheduler above handles queries too. Store each query as a serp: key in place of a URL:

from prismcrawl import PrismCrawl, PrismCrawlError

client = PrismCrawl()  # reads PRISMCRAWL_API_KEY from the environment


def check_query(key: str) -> str:
    query = key.removeprefix("serp:")
    (old_fp,) = db.execute("SELECT fingerprint FROM pages WHERE url = ?", (key,)).fetchone()

    body = client.google.search(query=query, gl="us", hl="en-US")
    results = body["data"]["content"]["results"]
    new_fp = fingerprint({"top10": [item.get("url") for item in results[:10]]})

    db.execute("UPDATE pages SET fingerprint = ? WHERE url = ?", (new_fp, key))
    changed = old_fp is not None and new_fp != old_fp
    reschedule(key, changed=changed, now=time.time())
    return "changed" if changed else "unchanged"


def run_due(session: requests.Session, limit: int = 500) -> None:
    due = db.execute(
        "SELECT url FROM pages WHERE next_due <= ? ORDER BY next_due LIMIT ?",
        (time.time(), limit),
    ).fetchall()
    for (key,) in due:
        try:
            status = check_query(key) if key.startswith("serp:") else check(session, key)
            print(status, key)
        except (requests.RequestException, ValueError, PrismCrawlError) as exc:
            db.execute("UPDATE pages SET next_due = ? WHERE url = ?", (time.time() + MIN_INTERVAL, key))
            print("error", key, exc)
    db.commit()


db.execute("INSERT OR IGNORE INTO pages (url) VALUES (?)", ("serp:best running shoes",))

This run_due replaces the one in the main script. The SDK raises a PrismCrawlError subclass on failed requests, so errors back off the same way as page errors.

Ask only for what's new. For news and mention monitoring, a recency filter works like a change feed: tbs="qdr:d" limits Google results to the past day and qdr:h to the past hour, so you aren't paying to see last week's articles again. Our news trading bot tutorial uses this pattern.

PrismCrawl charges one credit per successful search, failed searches are free, and credits last 90 days with no subscription. Since you aren't billed for bandwidth, the number of queries is the only cost lever, which is exactly what adaptive scheduling reduces. The cost calculator shows what a given schedule costs, and the free tier includes 100 credits with no credit card.

Incremental scraping checklist

  • Know how you're billed (per GB, per request, or per page processed) and start with the technique that targets that cost.
  • Send If-None-Match and If-Modified-Since on every request, and track the 304 rate per host.
  • Read sitemaps and feeds to move checks earlier, and keep your own schedule as the backstop.
  • Fingerprint extracted fields rather than raw HTML, and treat an empty extraction as an error.
  • Give each page its own interval, with a floor, a ceiling, and a priority weight for the pages that matter most.
  • Do a full, unconditional refresh now and then, so a bad validator or feed can't freeze your data.

Frequently asked questions

What is incremental web scraping?

Incremental scraping means re-fetching and re-processing only the pages that are new or have changed since your last visit, instead of re-downloading everything on a fixed schedule. It combines change signals from the server (ETag and Last-Modified headers, sitemap lastmod dates, feeds) with your own change detection (fingerprints of the fields you extract) and a scheduler that checks each page about as often as it actually changes.

How often should I re-scrape a website?

Decide per page. Start every page at a default interval such as a day, shorten it when a check finds a change, and lengthen it when a check finds nothing, within a minimum and maximum you set. Pages that change faster than you can afford to check deserve extra checks only if each individual change is worth money to you; if you just want an accurate copy, cap their frequency and spend the budget elsewhere.

Do ETag and Last-Modified headers work for web scraping?

Often, but not everywhere. Static files, CDN-served pages, feeds, and many CMS pages answer conditional requests with 304 Not Modified, which costs almost no bandwidth. Many dynamic pages send no validators, or send a Last-Modified equal to the current time, so they never return 304. Send the headers anyway, measure the 304 rate per host, and confirm changes with your own content fingerprint.

How do I detect whether a web page has changed?

Hash the fields you extract rather than the raw HTML. Raw HTML changes on every request because of CSRF tokens, timestamps, ad slots, and session IDs. Extract the fields you actually use, serialize them in a canonical form such as sorted JSON, hash that, and compare it with the previous hash. A different hash means something you care about changed.