Steady-state footprint
Idle privacytracker (no scrape running, no users connected):
The Next.js process is the dominant memory user. better-sqlite3 keeps a DB connection open in WAL mode; the connection itself is a few MB. There’s no separate service to supervise — the 30-minute ticker runs in the same process via
instrumentation.ts.
One thread does get spawned, though: on the first bulk write of a process’s life, lib/db-worker-client.ts starts a singleton worker_threads writer (lib/db-worker.cjs) holding its own better-sqlite3 connection to the same privacy.db, serialising with the main thread on the WAL write lock. Scrapes and imports both trigger it, so it’s live exactly during the ~250 MB worst case above — a second V8 isolate and a second SQLite connection, both counted in the process figures. It never starts on an idle install, which is why the idle rows are lower.
Disk per year of history
Per app, ballpark:
For a typical install — 100 apps, weekly sync, AI summaries on, ~5% of policies change per year — you’re looking at:
VACUUM if you want to reclaim space from soft-deleted annotations and pruned snapshots, but the trajectory is gentle. A 200-app install at weekly sync with AI on uses about 35 MB per year of history. A 1 GB SSD partition holds 25+ years of operation comfortably.
The Wayback back-fill adds a one-time bump: each successfully-fetched quarterly snapshot is the same ~3 KB as a live snapshot, so 4 quarters × 100 apps × 3 KB ≈ 1.2 MB for a year of pre-existing history.
Scrape time
The bottleneck is Apple’s HTML response time, not your machine.
These assume a sequential fetch loop (which is what privacytracker does) on a residential broadband connection. The runner is single-threaded by design — Apple rate-limits aggressively if you parallelise, and the per-app work isn’t CPU-bound enough to benefit from concurrency on most hardware.
Re-syncs are much faster than initial scrapes when policies haven’t changed: the privacy-policy hash check skips re-summarisation, and the snapshot diff is only persisted if categories actually changed.
Apple rate limits
Apple’sapps.apple.com and the iTunes Search API both rate-limit individual IPs. The thresholds aren’t published; observed behaviour:
- ~30-60 sequential scrapes in tight succession before the first 429
- 429s clear after roughly 10-30 minutes of cooldown
- Geographic and storefront variation —
apps.apple.com/us/...andapps.apple.com/au/...count separately enough that a multi-storefront install rarely hits 429
partial: rateLimited activity row with totals, and clears state + mutex cleanly so the next 30-minute scheduler tick can retry fresh. 429 is a recoverable expected condition, not a crash.
Practical implication: if you’re tracking more than ~200 apps on a single IP, set sync to weekly rather than daily. You’ll probably trip 429s on a daily cadence and nothing useful will come of it. Weekly cadence on 500 apps comfortably finishes in one tick.
If you need to track thousands of apps, route the scrape egress through multiple IPs (separate VPN endpoints, residential proxies). privacytracker doesn’t ship a built-in IP-rotation feature; you’d build it at the network layer.
AI summarisation cost
This is the only line item with a real money number. Summary cost depends on policy length × your provider’s per-token rate, scaled by how often policies change. For a typical privacy policy of ~20 KB:
privacytracker hashes policy text and skips re-summarisation when the hash is unchanged, so cosmetic edits to policies don’t burn calls. The provider only sees policy text — no app names, no annotations, no telemetry. See FAQ → How much do AI summaries cost? for the breakdown.
There’s no in-app spend cap. If you want a hard upper bound, set a spending limit on the API key at OpenAI or Anthropic and use that key here — privacytracker never sees your provider’s billing, so the ceiling has to live on their side.
RAM headroom
The Next.js process settles around 150 MB resident in steady state. Spikes during a scrape come from:- Parsing the App Store HTML payload (~500 KB – 1.5 MB unparsed JSON per app)
- Building the snapshot diff against the previous snapshot
- (If AI is on) chunking and round-tripping policy text — a 100 KB policy split into eight 12-KB chunks holds all eight in memory briefly
CPU profile
Two things produce visible CPU:- The HTML parser. Apple’s
serialized-server-datais ~300-500 KB of JSON;JSON.parseplus the shelf walk takes 5-30 ms per app on modern silicon, more on a Pi. For a 1000-app re-sync this is a few minutes of cumulative parsing — small compared to the network round-trip time. - AI provider calls. Network-bound from privacytracker’s perspective, but your local model (Ollama, LM Studio) will pin CPU/GPU during summarisation. A 7B-parameter model on Apple Silicon takes ~30-60 s per policy chunk; an 8B model on a CPU-only Pi takes minutes.
Soft scaling limits
privacytracker isn’t designed for thousands-of-apps installs, but here’s where the seams show up:
There’s no architectural ceiling — better-sqlite3 will happily hold tens of thousands of rows across all the tables — but the practical limit is your tolerance for Apple’s rate-limit behaviour.
Shrinking an existing install
If you’ve been running for a while and the DB is bigger than you’d like:DELETE is easier from a JSON bundle than from raw SQLite.
Sizing recommendation
If you’re not sure, start with:- Hardware: anything that runs Node 24 with 512 MB RAM free. A Raspberry Pi 4 / 5, a Synology DS220+ or better, an old MacBook Air, a $5/month VPS — all comfortable.
- Sync cadence: weekly. Switch to daily once you’ve confirmed your install doesn’t trip 429s.
- AI: start off. Turn it on per-app or for just the high-severity apps once you understand the cost shape.
- Storage: 10 GB free. You’ll use a fraction of that even after years of operation.