A folder of screenshots on one server works for a week. An archive that stays useful for years needs a key layout you can browse, metadata that answers "where did this come from", rules that delete or tier old captures on their own, and, for evidence, a lock that stops anyone from replacing a file. This guide sets all of that up in Amazon S3 or Cloudflare R2.
Everything here uses the S3 API, which R2 also speaks, so the same code works against either with a different endpoint.
A key layout that sorts itself
Object storage has no folders, only keys with slashes in them. Choose keys so that the listing you will want most often is a single prefix:
captures/{host}/{path}/{YYYY-MM-DDTHHMMZ}.{ext}
captures/example.com/pricing/2026-10-05T0600Z.png
captures/example.com/pricing/2026-10-12T0600Z.png
captures/example.com/legal/terms/2026-10-05T0600Z.png
Listing captures/example.com/pricing/ returns that page's history in date order, because ISO dates sort lexically. Lifecycle rules and locks can target captures/ as a whole or one site's prefix. Avoid putting the date first unless "everything captured on Tuesday" is the question you ask most.
Capture and upload
The capture comes from SnapRender as the response body; there is no intermediate file. This Node.js 18+ example uploads straight to R2 with the AWS SDK and keeps the useful facts as object metadata:
// npm install @aws-sdk/client-s3
import crypto from 'node:crypto';
import { S3Client, PutObjectCommand } from '@aws-sdk/client-s3';
const s3 = new S3Client({
region: 'auto',
endpoint: `https://${process.env.R2_ACCOUNT_ID}.r2.cloudflarestorage.com`, // omit for AWS S3
credentials: { accessKeyId: process.env.R2_ACCESS_KEY_ID, secretAccessKey: process.env.R2_SECRET_ACCESS_KEY },
});
export async function archive(url) {
const params = new URLSearchParams({ url, format: 'png', full_page: 'true' });
const res = await fetch(`https://app.snap-render.com/v1/screenshot?${params}`, {
headers: { 'X-API-Key': process.env.SNAPRENDER_API_KEY },
});
if (!res.ok) throw new Error(`capture failed for ${url}: ${res.status}`);
const body = Buffer.from(await res.arrayBuffer());
const { hostname, pathname } = new URL(url);
const stamp = new Date().toISOString().slice(0, 16).replace(':', '') + 'Z'; // 2026-10-05T0600Z
const key = `captures/${hostname}${pathname.replace(/\/+$/, '') || '/home'}/${stamp}.png`;
await s3.send(new PutObjectCommand({
Bucket: 'screenshots',
Key: key,
Body: body,
ContentType: 'image/png',
Metadata: {
'source-url': url,
sha256: crypto.createHash('sha256').update(body).digest('hex'),
'captured-at': new Date().toISOString(),
},
}));
return key;
}
Keep metadata small and factual: the source URL, a hash and the capture time cover most questions an archive gets asked. Anything bigger, such as extracted page text, belongs in its own object next to the image.
Retention: delete or tier what you no longer need
A daily capture of 200 pages is about 73,000 objects a year. Decide up front how long they are useful and let the bucket enforce it.
AWS S3 lifecycle. This configuration moves captures to an infrequent-access class after 30 days and deletes them after a year, but only under captures/:
{
"Rules": [
{
"ID": "screenshots-retention",
"Filter": { "Prefix": "captures/" },
"Status": "Enabled",
"Transitions": [{ "Days": 30, "StorageClass": "STANDARD_IA" }],
"Expiration": { "Days": 365 }
}
]
}
Apply it with aws s3api put-bucket-lifecycle-configuration --bucket screenshots --lifecycle-configuration file://lifecycle.json.
Cloudflare R2 lifecycle. R2 lifecycle rules can delete objects after a number of days and move them to Infrequent Access storage. From the command line:
npx wrangler r2 bucket lifecycle add screenshots --name screenshots-retention --prefix captures/ --expire-days 365
npx wrangler r2 bucket lifecycle list screenshots
Lifecycle deletions run in the background, so objects may linger a little past their expiry day. Do not rely on lifecycle timing for anything that must be gone at an exact moment.
Evidence: make captures impossible to replace
If the archive exists to prove what a page said, retention is the opposite problem: nobody should be able to delete or overwrite a capture early.
- S3 Object Lock (requires versioning on the bucket): upload with
ObjectLockMode: 'COMPLIANCE' and ObjectLockRetainUntilDate. In compliance mode, not even the account root can remove the object before that date.
- R2 bucket locks: add a rule that retains everything under a prefix for a number of days, until a date, or indefinitely, with
npx wrangler r2 bucket lock add or in the dashboard. The strictest rule covering an object wins.
Use a separate prefix or bucket for locked evidence so that lifecycle cleanup of ordinary captures never collides with it. The compliance screenshots guide adds hashing and independent RFC 3161 timestamps on top.
Sharing without making the bucket public
Keep the bucket private and hand out short-lived links instead:
import boto3
s3 = boto3.client("s3") # for R2, pass endpoint_url="https://<account>.r2.cloudflarestorage.com"
link = s3.generate_presigned_url(
"get_object",
Params={"Bucket": "screenshots", "Key": "captures/example.com/pricing/2026-10-05T0600Z.png"},
ExpiresIn=7 * 24 * 3600,
)
print(link)
Presigned links expire on their own and can be sent to a client or embedded in a report. If what you want to share is a live, current capture rather than an archived one, SnapRender's signed URLs render the page on request without exposing your API key.
S3 or R2?
| Question |
Lean S3 |
Lean R2 |
| Is the archive downloaded or viewed often? |
|
R2 has no egress fees |
| Do you need very cold, very cheap tiers for years? |
S3 Glacier classes |
|
| Is the rest of your stack on AWS? |
Same IAM, same region |
|
| Do you serve images to browsers through a CDN? |
|
R2 sits on Cloudflare's network |
Every row above is one GET request away. 200 free renders a month, no card.
Start free
Both support lifecycle rules and write-once retention, so the decision is usually about egress and where the rest of your infrastructure lives.
Putting it on a schedule
Wrap archive() in whatever already runs your jobs: cron, a cloud scheduler, a GitHub Actions schedule or an n8n workflow. For long lists, send them as a batch of up to 50 URLs and upload each result when the batch's webhook fires; batch download links last 24 hours, so copy them into the bucket straight away.
Related guides: daily screenshots of 100 URLs with GitHub Actions and website change detection without false alarms.