Blog 6 min read

Screenshot Evidence for Compliance: Archiving Web Pages with Timestamps and Hashes

How to keep a defensible record of what a web page showed: full-page capture, a SHA-256 manifest, an independent RFC 3161 timestamp, and write-once storage in S3 or Cloudflare R2. Runnable Python and shell.

A screenshot proves something only if you can show when it was taken and that nobody changed it since. That takes four things: a capture of the whole page, a hash recorded at capture time, a timestamp from someone other than you, and storage that refuses overwrites. This guide builds all four with a short Python script, two OpenSSL commands and one storage setting.

The need usually arrives sideways. A customer disputes what your pricing page said in March. A regulator asks to see the disclosures that ran next to a promotion. A partner claims your terms never mentioned a fee. Someone opens the page today, and today's page is the wrong one.

Teams that archive on a schedule have the answer in minutes. Teams that do not have a hand-taken screenshot from someone's laptop, cropped, with a cookie banner across the middle and no reliable date.

Try it on your URL

Live capture, no signup. You get the screenshot and the exact API call that produced it.

What makes a capture defensible

Four properties do most of the work. None of them is exotic.

Completeness. Capture the whole page, top to bottom, including sections that load as you scroll. A viewport screenshot shows what fit on one screen, and the clause in question is rarely above the fold.

Content, not overlays. A consent dialog over the middle of the page hides exactly the text you need. Capture with cookie banners removed, unless your policy specifically requires recording the banner itself.

Integrity. Record a SHA-256 hash of the file the moment it is created. Anyone can recompute it later; if one pixel changed, the hash changes.

Independent time. Your server clock is your word. A timestamp from an RFC 3161 authority is someone else's signed word that this exact hash existed at that moment.

Then store the file where it cannot be replaced, for as long as your policy says.

Or skip the setup entirely

Everything in this guide is one GET request with SnapRender. No browser to babysit, no timeouts to tune.

200 screenshots a month free. First render in under a minute.

Step 1: capture and record

This script captures a full-page PNG through SnapRender, hashes it, and writes a JSON record next to it. Cookie banners and ads are removed before capture, and every request renders fresh.

import datetime, hashlib, json, os, pathlib, requests

API = "https://app.snap-render.com"
HEADERS = {"X-API-Key": os.environ["SNAPRENDER_API_KEY"]}
OUT = pathlib.Path("evidence")
OUT.mkdir(exist_ok=True)

def capture(url: str, fmt: str = "png") -> pathlib.Path:
    requested_at = datetime.datetime.now(datetime.timezone.utc)
    r = requests.get(f"{API}/v1/screenshot", headers=HEADERS, timeout=120, params={
        "url": url, "format": fmt, "full_page": "true",
        "block_cookie_banners": "true", "block_ads": "true",
    })
    r.raise_for_status()
    digest = hashlib.sha256(r.content).hexdigest()
    path = OUT / f"{requested_at:%Y%m%dT%H%M%SZ}-{digest[:12]}.{fmt}"
    path.write_bytes(r.content)
    path.with_suffix(".json").write_text(json.dumps({
        "url": url,
        "requested_at_utc": requested_at.isoformat(),
        "server_date": r.headers.get("Date"),
        "sha256": digest,
        "bytes": len(r.content),
        "file": path.name,
    }, indent=2))
    return path

for url in ["https://example.com/terms", "https://example.com/pricing"]:
    print(capture(url))

The record keeps two times on purpose: when your script asked, and the Date header of the response. They should be seconds apart; a large gap is worth investigating before the record is filed.

If you also need the words searchable, call GET /v1/extract?type=text for the same URL and store the text with the image. It is one more call, and it makes the archive greppable.

Step 2: an independent timestamp

RFC 3161 timestamp authorities sign a hash and a time. You never send the file itself. OpenSSL builds the request and reads the answer:

FILE=evidence/20261005T080000Z-3f2a9c1b7e4d.png

# 1. Build a timestamp query from the file's SHA-256
openssl ts -query -data "$FILE" -sha256 -cert -out "$FILE.tsq"

# 2. Send it to a timestamp authority (FreeTSA shown; use the authority your policy names)
curl -s -H "Content-Type: application/timestamp-query" \
  --data-binary @"$FILE.tsq" https://freetsa.org/tsr -o "$FILE.tsr"

# 3. Read it: the signed time and the hash it covers
openssl ts -reply -in "$FILE.tsr" -text

Keep the .tsr file with the image and its JSON. To verify later, download the authority's CA certificate and run openssl ts -verify -data "$FILE" -in "$FILE.tsr" -CAfile cacert.pem -untrusted tsa.crt. A passing check means this exact file existed no later than the signed time.

Step 3: storage that refuses overwrites

A record that someone can quietly replace is not much of a record. Both major object stores can enforce retention.

Amazon S3 Object Lock. Turn on Object Lock for the bucket (it requires versioning), then upload in compliance mode with a retention date. Nobody, including the account root, can delete or overwrite the object before that date.

import datetime, boto3

s3 = boto3.client("s3")
until = datetime.datetime.now(datetime.timezone.utc) + datetime.timedelta(days=365 * 7)
for name in ["20261005T080000Z-3f2a9c1b7e4d.png", "20261005T080000Z-3f2a9c1b7e4d.json"]:
    with open(f"evidence/{name}", "rb") as f:
        s3.put_object(Bucket="evidence-archive", Key=f"captures/{name}", Body=f,
                      ObjectLockMode="COMPLIANCE", ObjectLockRetainUntilDate=until)

Cloudflare R2 bucket locks. R2 applies retention through bucket lock rules scoped by key prefix: lock everything under captures/ for a fixed number of days, until a date, or indefinitely. Add a rule in the dashboard or with npx wrangler r2 bucket lock add, and uploads under that prefix cannot be deleted or overwritten while the rule applies. R2 has no egress fees, which matters if auditors download the archive often.

Whichever you use, set the retention to what your policy says, not to a guess. Too short defeats the purpose; too long can conflict with data minimization rules.

Step 4: run it on a schedule

Archives are only as good as their coverage. Decide which pages matter (terms, privacy policy, pricing, regulated product pages, campaign landing pages) and how often they change, then schedule accordingly:

  • Pages that change on release: capture after every deploy and weekly as a backstop.
  • Campaign pages: capture at launch, at each change, and at the end of the campaign.
  • Legal pages: capture weekly and on every edit.

To capture many pages at once, send them as a batch of up to 50 URLs; failed items are listed with their error and not billed, and a webhook tells you when the batch is done. The compliance screenshots page has a calculator for the monthly volume.

How to verify a record later

When someone asks what a page said on a date, the check takes a minute:

  1. Find the capture for that date in the archive and open its JSON record.
  2. Recompute the hash with sha256sum and compare it with the record.
  3. Verify the .tsr timestamp against the authority's certificate.
  4. Show the object's retention settings, which prove it could not have been replaced.

If all four hold, you have a full-page image, made at a known time, unchanged since.

What a screenshot does not prove

Be precise about the claim. A capture shows what the page rendered for an automated browser at that time, from that location, without a login. It does not show who else saw it, what a logged-in user saw, or what a visitor in another country saw if the site varies by region. SnapRender captures public pages and does not offer geo-targeted rendering. For anything that goes to a court or a regulator, have counsel review the process before you rely on it.

Getting started

The capture step is one GET request. A free SnapRender key includes 200 full-page captures a month with every option on this page, which is enough to archive a few dozen pages weekly. Try a single page in the free screenshot tool first if you want to see the output, then get a key and run the script above.

Frequently asked questions

What makes a web page screenshot usable as evidence?

Completeness, a fixed capture time and proof the file was not changed later. In practice that means a full-page capture, a SHA-256 hash recorded at capture time, a timestamp from an independent authority, and storage that cannot be overwritten. Whether that meets a specific legal standard depends on your jurisdiction; ask counsel.

What is an RFC 3161 timestamp?

A signed statement from a timestamp authority that a given hash existed at a given time. You send only the hash, never the file, and get back a token anyone can verify with OpenSSL. It means the capture time does not rest on your own server clock.

Should compliance captures be PNG or PDF?

PNG is an exact visual record and the usual choice. PDF suits document management systems. If you also need the page text searchable, store the extracted text alongside the image.

How long should I keep compliance screenshots?

As long as your retention policy requires, which varies by industry and jurisdiction. Write-once storage such as S3 Object Lock or R2 bucket locks enforces the period so nobody can delete records early.

Or skip the setup entirely

200 screenshots a month free. First render in under a minute.

Grab a free API key