# Performance & sizing
Source: https://docs.privacytracker.privacykey.org/performance-and-sizing

Disk per year of label history, RAM/CPU footprint, soft scaling limits, Apple rate-limit math — what to expect before you commit hardware.

privacytracker is a single-user app with deliberately modest hardware requirements. This page is for self-hosters trying to decide whether their NAS / Pi / repurposed laptop is enough, and for anyone who wants the numbers before tracking 200+ apps.

The numbers here are approximations from the project's own usage, not stress-test results — your mileage will vary by hardware, network, AI provider, and how often labels actually change for the apps you track.

## Steady-state footprint

Idle privacytracker (no scrape running, no users connected):

| Resource | Typical | Worst case |
|---|---|---|
| **RAM (process)** | 120 – 180 MB | ~250 MB during a scrape |
| **CPU** | `<1%` on any modern core | brief 20–40% spike per scrape |
| **Disk (DB)** | see [Disk per year](#disk-per-year-of-history) | grows linearly with apps × snapshots |
| **Network egress (idle)** | zero | 30-min scheduler tick fires scrapes if `sync_schedule` is on |

The Next.js process is the dominant memory user. better-sqlite3 keeps a DB connection open in WAL mode; the connection itself is a few MB. There's no separate service to supervise — the 30-minute ticker runs in the same process via `instrumentation.ts`.

One thread does get spawned, though: on the first bulk write of a process's life, `lib/db-worker-client.ts` starts a singleton `worker_threads` writer (`lib/db-worker.cjs`) holding its *own* better-sqlite3 connection to the same `privacy.db`, serialising with the main thread on the WAL write lock. Scrapes and imports both trigger it, so it's live exactly during the ~250 MB worst case above — a second V8 isolate and a second SQLite connection, both counted in the process figures. It never starts on an idle install, which is why the idle rows are lower.

## Disk per year of history

Per app, ballpark:

| Component | Bytes per row | Rows per year | Total per app per year |
|---|---|---|---|
| `privacy_snapshots.snapshot_json` | ~3 KB | 12 (monthly sync) to 365 (daily) | 36 KB – 1 MB |
| `privacy_snapshots.changes_json` | ~500 B | only when changes detected (~3-5/year typical) | 1.5 KB – 2.5 KB |
| `privacy_categories` rows (current state) | ~80 B | constant ~30 per app | 2.4 KB |
| `policy_versions` (with AI summary) | ~6 KB | 1-3/year typical | 6 – 18 KB |
| `policy_versions` (without AI) | ~2 KB | 1-3/year typical | 2 – 6 KB |
| `notifications` | ~200 B | one per change | `<1 KB` |

For a typical install — **100 apps, weekly sync, AI summaries on, ~5% of policies change per year** — you're looking at:

```
100 apps × (52 snapshots × 3 KB + 5 changes × 500 B + 30 cats × 80 B + 1.5 policy ver × 6 KB)
≈ 100 × (156 KB + 2.5 KB + 2.4 KB + 9 KB)
≈ 17 MB per year
```

After the first year you can run `VACUUM` if you want to reclaim space from soft-deleted annotations and pruned snapshots, but the trajectory is gentle. **A 200-app install at weekly sync with AI on uses about 35 MB per year of history.** A 1 GB SSD partition holds 25+ years of operation comfortably.

The Wayback back-fill adds a one-time bump: each successfully-fetched quarterly snapshot is the same ~3 KB as a live snapshot, so 4 quarters × 100 apps × 3 KB ≈ 1.2 MB for a year of pre-existing history.

## Scrape time

The bottleneck is Apple's HTML response time, not your machine.

| Apps | Typical wall-clock time for a full re-sync |
|---|---|
| 10 | 8 – 15 s |
| 50 | 40 – 90 s |
| 100 | 90 s – 4 min |
| 500 | 8 – 20 min |
| 1000 | 20 – 50 min, expect to hit a 429 partway through |

These assume a sequential fetch loop (which is what privacytracker does) on a residential broadband connection. The runner is single-threaded by design — Apple rate-limits aggressively if you parallelise, and the per-app work isn't CPU-bound enough to benefit from concurrency on most hardware.

Re-syncs are *much* faster than initial scrapes when policies haven't changed: the privacy-policy hash check skips re-summarisation, and the snapshot diff is only persisted if categories actually changed.

## Apple rate limits

Apple's `apps.apple.com` and the iTunes Search API both rate-limit individual IPs. The thresholds aren't published; observed behaviour:

- ~30-60 sequential scrapes in tight succession before the first 429
- 429s clear after roughly 10-30 minutes of cooldown
- Geographic and storefront variation — `apps.apple.com/us/...` and `apps.apple.com/au/...` count separately enough that a multi-storefront install rarely hits 429

The bulk runner handles 429 gracefully: it bails out of the loop on the first 429, records a `partial: rateLimited` activity row with totals, and **clears state + mutex cleanly** so the next 30-minute scheduler tick can retry fresh. 429 is a recoverable expected condition, not a crash.

Practical implication: **if you're tracking more than ~200 apps on a single IP, set sync to weekly rather than daily.** You'll probably trip 429s on a daily cadence and nothing useful will come of it. Weekly cadence on 500 apps comfortably finishes in one tick.

If you need to track thousands of apps, route the scrape egress through multiple IPs (separate VPN endpoints, residential proxies). privacytracker doesn't ship a built-in IP-rotation feature; you'd build it at the network layer.

## AI summarisation cost

This is the only line item with a real money number. Summary cost depends on policy length × your provider's per-token rate, scaled by how often policies change.

For a typical privacy policy of ~20 KB:

| Provider / model | Per-summary cost | 100 apps, ~5%/year change | Notes |
|---|---|---|---|
| `gpt-4o-mini` | $0.001 – $0.003 | $0.05 – $0.15 / year | Default for OpenAI |
| `claude-3-5-haiku` | $0.001 – $0.004 | $0.05 – $0.20 / year | Default for Anthropic |
| `claude-3-5-sonnet` | $0.005 – $0.015 | $0.25 – $0.75 / year | Better-quality summaries |
| Local Ollama (`llama3.1:8b`) | $0 runtime | $0 | One-time ~5 GB model download |

privacytracker hashes policy text and skips re-summarisation when the hash is unchanged, so cosmetic edits to policies don't burn calls. The provider only sees policy text — no app names, no annotations, no telemetry. See [FAQ → How much do AI summaries cost?](https://docs.privacytracker.privacykey.org/faq) for the breakdown.

There's no in-app spend cap. If you want a hard upper bound, set a spending limit on the API key at OpenAI or Anthropic and use that key here — privacytracker never sees your provider's billing, so the ceiling has to live on their side.

## RAM headroom

The Next.js process settles around 150 MB resident in steady state. Spikes during a scrape come from:

- Parsing the App Store HTML payload (~500 KB – 1.5 MB unparsed JSON per app)
- Building the snapshot diff against the previous snapshot
- (If AI is on) chunking and round-tripping policy text — a 100 KB policy split into eight 12-KB chunks holds all eight in memory briefly

In practice a 256 MB container is enough for installs up to ~200 apps; 512 MB is comfortable up to ~1000.

## CPU profile

Two things produce visible CPU:

1. **The HTML parser.** Apple's `serialized-server-data` is ~300-500 KB of JSON; `JSON.parse` plus the shelf walk takes 5-30 ms per app on modern silicon, more on a Pi. For a 1000-app re-sync this is a few minutes of cumulative parsing — small compared to the network round-trip time.
2. **AI provider calls.** Network-bound from privacytracker's perspective, but your *local* model (Ollama, LM Studio) will pin CPU/GPU during summarisation. A 7B-parameter model on Apple Silicon takes ~30-60 s per policy chunk; an 8B model on a CPU-only Pi takes minutes.

Local models are why people end up running privacytracker on a beefier machine than they'd otherwise need. Hosted AI is much lighter on local resources.

## Soft scaling limits

privacytracker isn't designed for thousands-of-apps installs, but here's where the seams show up:

| Apps | What starts to matter |
|---|---|
| ≤ 50 | Nothing. Runs fine on anything that boots Node 24. |
| 50 – 200 | Daily sync still fine on a single IP. |
| 200 – 500 | Move to weekly sync to avoid 429s. Initial scrape may need to be split across two days. |
| 500 – 1000 | Weekly sync only. Wayback bulk imports take 30+ minutes — leave them running, they auto-resume on crash. |
| 1000+ | Outside the supported model. Consider IP rotation, multiple instances, or asking whether you actually need to track every app on someone's phone. |

There's no architectural ceiling — better-sqlite3 will happily hold tens of thousands of rows across all the tables — but the practical limit is your tolerance for Apple's rate-limit behaviour.

## Shrinking an existing install

If you've been running for a while and the DB is bigger than you'd like:

```bash
# Stop the app, then:
sqlite3 data/privacy.db <<'SQL'
  -- Hard-delete soft-deleted annotations
  DELETE FROM annotations WHERE deletedAt IS NOT NULL;

  -- Drop snapshots older than N days for apps with >M snapshots,
  -- keeping the first and last per quarter.
  -- (Sketch — adapt to your retention policy.)
  DELETE FROM privacy_snapshots
  WHERE id NOT IN (
    SELECT id FROM privacy_snapshots
    WHERE id IN (
      SELECT MIN(id) FROM privacy_snapshots
      GROUP BY appId, strftime('%Y-Q%q', scrapedAt/1000, 'unixepoch')
    )
  );

  VACUUM;
SQL
```

privacytracker doesn't ship an automatic retention policy — once persisted, snapshots are forever. If retention matters to you, schedule something like the above as a cron job. Take a backup first; recovery from a too-aggressive `DELETE` is easier from a JSON bundle than from raw SQLite.

## Sizing recommendation

If you're not sure, start with:

- **Hardware:** anything that runs Node 24 with 512 MB RAM free. A Raspberry Pi 4 / 5, a Synology DS220+ or better, an old MacBook Air, a $5/month VPS — all comfortable.
- **Sync cadence:** weekly. Switch to daily once you've confirmed your install doesn't trip 429s.
- **AI:** start *off*. Turn it on per-app or for just the high-severity apps once you understand the cost shape.
- **Storage:** 10 GB free. You'll use a fraction of that even after years of operation.

If you're tracking your kid's iPad or doing a one-off privacy review, even a Raspberry Pi Zero 2 W is enough. The main reason to go bigger is if you want a local AI model on the same box.
