Blog 6 min read

Monitor Any Web Page for Changes with Python (About 40 Lines)

A small Python script that checks web pages on a schedule, detects real content changes without false alarms from ads or cookie banners, saves a full-page screenshot of each new version, and posts an alert to Slack.

The reliable way to detect that a web page changed is to compare what a reader sees, not the HTML behind it. This guide builds a Python monitor that fetches the rendered text of each page, hashes it, saves a full-page screenshot whenever the hash moves, and posts the change to Slack. It runs from cron or any scheduler, and the whole script is about 40 lines.

Most first attempts at change detection download the HTML and diff it. That works for a day. Then the alerts start firing on every run, because modern pages put a fresh CSRF token, a cache-busting asset hash or a rotating ad slot into the markup on each request. Nobody reads an alert channel that cries wolf.

The fix is to render the page the way a browser does, strip the noise, and compare only the content.

Try it on your URL

Live capture, no signup. You get the screenshot and the exact API call that produced it.

What the monitor does

For each URL on your list, once per run:

  1. Fetch the page's visible text after JavaScript has run, with cookie banners and ads removed.
  2. Hash that text with SHA-256 and compare it with the hash stored from the last run.
  3. If it is unchanged, move on. That is the common case and costs one API call.
  4. If it changed, capture a full-page screenshot as evidence, save it with a timestamp, and send a Slack message.

Rendering and cleanup happen in SnapRender, so the script needs no headless browser, no Chromium install and no consent-dialog clicking. Text comes from GET /v1/extract?type=text; the screenshot from GET /v1/screenshot.

Or skip the setup entirely

Everything in this guide is one GET request with SnapRender. No browser to babysit, no timeouts to tune.

200 screenshots a month free. First render in under a minute.

The script

Install the one dependency with pip install requests, set SNAPRENDER_API_KEY (a free key gives 200 calls a month) and, optionally, SLACK_WEBHOOK_URL for alerts.

import datetime, hashlib, json, os, pathlib, requests

API = "https://app.snap-render.com"
HEADERS = {"X-API-Key": os.environ["SNAPRENDER_API_KEY"]}
SLACK = os.environ.get("SLACK_WEBHOOK_URL")
WATCH = {
    "https://example.com/pricing": "main",       # CSS selector to watch
    "https://example.com/legal/terms": "article",
}
STATE = pathlib.Path("state.json")
SHOTS = pathlib.Path("captures")
SHOTS.mkdir(exist_ok=True)
state = json.loads(STATE.read_text()) if STATE.exists() else {}

for url, selector in WATCH.items():
    r = requests.get(f"{API}/v1/extract", headers=HEADERS, timeout=60,
                     params={"url": url, "type": "text", "selector": selector})
    r.raise_for_status()
    digest = hashlib.sha256(r.json()["content"].encode()).hexdigest()
    if state.get(url) == digest:
        continue

    stamp = datetime.datetime.now(datetime.timezone.utc).strftime("%Y%m%dT%H%M%SZ")
    shot = requests.get(f"{API}/v1/screenshot", headers=HEADERS, timeout=90,
                        params={"url": url, "format": "png", "full_page": "true"})
    shot.raise_for_status()
    path = SHOTS / f"{hashlib.md5(url.encode()).hexdigest()[:8]}-{stamp}.png"
    path.write_bytes(shot.content)

    if url in state and SLACK:  # the first run only records a baseline
        requests.post(SLACK, json={"text": f"Changed: {url} (saved {path.name})"}, timeout=10)
    state[url] = digest

STATE.write_text(json.dumps(state, indent=2))

Run it once to record a baseline. Every run after that compares against the stored hashes and alerts only on real changes.

Why this avoids false alarms

Three choices do most of the work.

Rendered text, not HTML. The extract endpoint loads the page in real Chromium, waits for it to render, and returns the text a reader would see. Script tags, tokens, tracking pixels and attribute churn never reach the hash.

Cookie banners and ads removed first. Consent dialogs are one of the biggest sources of noise in change monitoring, because many of them vary by visit. SnapRender removes them before extracting, by default.

A selector scopes the comparison. Watching main or article instead of the whole body ignores the navigation, the footer and the "latest posts" sidebar. On a pricing page, pointing at the pricing table itself is even better. If one element inside your scope still changes on every load, such as a rotating testimonial, the screenshot endpoint's hide_selectors parameter hides it in the captured image too.

Scheduling it

Any scheduler works, because the script is stateless apart from state.json.

Cron on a server, every six hours:

0 */6 * * * cd /opt/monitor && SNAPRENDER_API_KEY=sk_live_... python3 monitor.py >> monitor.log 2>&1

GitHub Actions is a free alternative if you do not want a server. Commit state.json back to the repository at the end of each run so the next run has its baseline. Two limits to know: scheduled workflows run at most every five minutes and can start several minutes late at busy times, and GitHub pauses schedules in public repositories after 60 days without activity. The change monitoring guide has a complete workflow file.

When you want pixels, not text

Text hashing misses changes that are purely visual: a new hero image, a swapped logo, a broken stylesheet. For pages where layout matters, compare screenshots instead:

  1. Capture at a fixed viewport, for example width=1280&height=800, so every image has the same dimensions. Full-page heights change as content moves, and pixel comparison needs equal sizes.
  2. Compare the new image with the previous one using a library such as pixelmatch in Node.js or Pillow's ImageChops.difference in Python.
  3. Alert when the share of changed pixels crosses a threshold, typically between 0.5 and 2 percent, and tune it per page.

Text first, pixels where they earn their keep, is the combination that stays quiet until something real happens.

What it costs

Each run makes one extract call per page, plus one screenshot call per page that changed. Failed calls are not billed.

Pages Check frequency Calls a month (approx.) Plan that fits
6 daily 180 Free (200)
50 daily 1,500 Starter, $9
50 every 6 hours 6,000 Growth, $29
50 hourly 36,000 Business, $79

Every row above is one GET request away. 200 free renders a month, no card.

Start free

Changes add screenshot calls on top, but on most pages they are rare. Every plan includes every feature, so the free tier is enough to build and test the whole monitor before deciding.

Where to take it next

  • Keep the screenshots in object storage with a retention rule instead of a local folder.
  • Store the extracted text as well as the hash, so the Slack message can include a short diff.
  • Watch a list of hundreds of pages by sending them as a batch of up to 50 URLs and handling the result in a webhook.

The full pattern, with a Node.js pixel-diff script and a scheduled workflow, is on the website change monitoring page. If you want to try the capture side before writing any code, paste a URL into the free screenshot tool.

Frequently asked questions

How do I detect changes on a web page with Python?

Fetch the page's rendered text on a schedule, hash it with hashlib.sha256, and compare the hash with the one from the previous run. Using rendered text instead of raw HTML avoids false alarms from tokens, ad slots and timestamps in the markup.

Why not compare the raw HTML?

Raw HTML changes on almost every request: CSRF tokens, cache-busting query strings, rotating ad markup, analytics snippets. A diff on it fires constantly. The text a reader sees changes only when the content changes.

How do I monitor only part of a page?

Pass a CSS selector to the extract call, for example selector=main or selector=.pricing-table. Only the text inside that element is hashed, so changes to the header, footer or sidebar are ignored.

How often can I check a page?

As often as your scheduler allows. Each check is one API call, so 20 pages checked every hour is about 14,400 calls a month, while the same pages checked daily are about 600, which fits a $9 plan.

Or skip the setup entirely

200 screenshots a month free. First render in under a minute.

Grab a free API key