π― The takeaway, first
The frontier is the system. A crawler is a polite, deduplicating, trap-avoiding URL scheduler with a downloader attached β like a librarian who must copy every book in the city but may only knock on each door once every few seconds, must obey "do not disturb" signs, and refuses to copy the same book twice. If your design starts with "fetch fast," you've missed the point: the binding constraint is courtesy, and the scale problem is the queue.
π« Misconception, busted
"A crawler just downloads pages as fast as possible." Speed is the one thing you are not allowed to maximize. Every host gets a politeness delay (default ~1 request/second, more if their robots.txt says so), and spider traps β infinite calendars, faceted-search URLs like ?color=red&size=m&page=48291 β will eat an impolite crawler alive. Most of the engineering is scheduling and dedup, not fetching.
Requirements
Say these out loud before drawing a single box.
Functional
- Start from seed URLs, discover new URLs from page links
- Fetch pages and extract outbound links
- Respect
robots.txt(allow/disallow, crawl-delay) - Deduplicate URLs and near-duplicate content
- Store fetched pages for the indexer
Non-functional
- 10B+ pages/day fetch capacity (~115K/s)
- Politeness: default β€ 1 req/s per host, honor crawl-delay
- A dead host must not stall the crawl
- Freshness: recrawl important pages (news hourly, blogs weekly)
- Survive traps: no infinite loops, bounded per-host budgets
Back-of-the-envelope math
Google-scale-ish. Assumptions labeled.
| What | Assumption | Math |
|---|---|---|
| Fetch rate | 10B pages/day | β 115K fetches/s sustained |
| Bandwidth | avg page 50 KB | 115K Γ 50 KB β 5.8 GB/s + DNS/TLS overhead |
| Frontier size | web holds trillions of known URLs | queue must hold 100B+ URLs β it lives on disk, not RAM |
| Seen-URL dedup | Bloom filter, 1T URLs, 1% false positives β 9.6 bits/URL | β 1.2 TB β sharded, or segmented per crawl batch |
| Politeness paradox | 100M hosts Γ 1 req/s politeness | To feed 115K fetchers you need 115K distinct hosts with ripe queues at all times β that's the scheduling crux |
| Robots cache | 100M hosts Γ ~2 KB rules, 24h TTL | β 200 GB β fits in a shared cache, refetched daily |
Go deeper: why a Bloom filter and not a hash set?
An exact set of 1 trillion URLs at ~60 bytes each is 60 TB β you can't keep it in RAM and disk lookups at 115K/s would melt. A Bloom filter answers "have we seen this URL?" in O(1) with a few hash computations and ~10 bits per URL. The cost: ~1% false positives β we occasionally skip a page we hadn't seen. For a crawler that's a fine trade: the page will likely be rediscovered via another link later. Probabilistic data structures are the correct answer to web-scale dedup; say so confidently.
Architecture
The centerpiece: everything orbits the frontier.
flowchart LR
SEED[Seed URLs] --> FR[Frontier
per-host queues + priority]
FR --> SCH[Scheduler
politeness gate]
SCH --> FE[Fetcher Pool]
FE --> DNS[(DNS Cache)]
FE --> RC[(Robots Cache)]
RC -->|disallow| SKIP[Skip]
RC -->|allow| GET[HTTP GET]
GET --> PA[Parser]
PA --> DED[Dedup
URL Bloom + content simhash]
DED -->|new| PS[(Page Store)]
DED -->|duplicate| DROP[Drop]
PA --> LE[Link Extractor]
LE --> NORM[URL Normalizer
strip tracking params, fragments]
NORM --> FR
PS --> IDX[Indexer
downstream]
The frontier holds one queue per host β this is the key structural decision, because politeness is per-host. The scheduler only hands a fetcher a URL whose host's delay has expired. Fetched pages go through robots check (cached), parse, dedup (URL-level Bloom filter, then content-level simhash for near-duplicates), and link extraction feeds normalized URLs back into the frontier. The indexer consumes the page store downstream and never talks to the crawl loop.
Go deeper: why partition the frontier by host?
Two reasons. First, politeness becomes local: a scheduler owning example.com's queue just tracks one "next allowed fetch" timestamp. Second, failure isolation: if a host hangs or blackholes connections, only its queue stalls β the other 99,999,999 hosts keep flowing. In a distributed build you'd consistent-hash hosts across frontier shards, so all URLs for one host always land on the same scheduler.
Component deep-dives
Fetching with politeness
sequenceDiagram
participant S as Scheduler
participant F as Frontier (per-host queue)
participant R as Robots Cache
participant W as Fetcher Worker
participant H as example.com
S->>F: next ripe host? (delay expired)
F-->>S: example.com β /blog/post-9
S->>R: robots rules for example.com?
alt cache miss
S->>H: GET /robots.txt
H-->>S: rules (cache 24h)
end
alt disallowed
S->>F: drop URL, note the rule
else allowed
S->>W: fetch /blog/post-9
W->>H: GET (1 req β then the host goes quiet)
H-->>W: 200 HTML
W->>F: mark host next-allowed = now + delay
end
Dedup β two gates
sequenceDiagram
participant P as Parser
participant B as URL Bloom Filter
participant C as Content Simhash Index
participant S as Page Store
P->>P: normalize URL (lowercase host, strip ?utm_*, #frag)
P->>B: seen this URL?
alt yes β drop (URL dupe)
B-->>P: probably seen β skip fetch entirely
else no
P->>P: fetch + parse
P->>C: simhash(content) near-duplicate?
alt near-dupe of known page β drop
C-->>P: Hamming distance < 3 β skip
else new
P->>S: store page
P->>B: add URL
end
end
API + data model
A crawler has no public API β its "API" is the frontier's internal contract. This is what the scheduler reads and writes:
frontier_entry { url, host, priority_score, next_fetch_at, retries, discovered_from }
page { url_hash PK, url, content_hash, fetched_at, content_ptr β blob store }
robots_cache { host PK, rules, crawl_delay_s, fetched_at } -- 24h TTL
host_stats { host PK, urls_seen, urls_fetched, errors_24h, budget_remaining }
Trade-offs
| Decision | Option A | Option B | Pick |
|---|---|---|---|
| Crawl order | BFS (fair, simple) | Priority (PageRank-ish, important pages first) | Priority β freshness where it matters; BFS starves news |
| URL dedup | Exact hash set (no false drops) | Bloom filter (tiny, 1% false drops) | Bloom β a skipped page gets rediscovered |
| Frontier | Central queue | Sharded by host hash | Sharded β politeness stays local, failures contained |
| Fetching | Raw HTTP | Headless browser (renders JS) | Raw HTTP first; headless only for JS-heavy hosts, budgeted separately |
| robots.txt fetch fails | Fail open (crawl anyway) | Fail closed (skip host) | Fail closed, retry later β politeness is the brand |
Failure modes
What I'd actually build
Opinionated. Steal this for the interview.
Stack: Go fetcher pool (great at 100K concurrent connections). Frontier as Kafka topics partitioned by host hash + Redis per-host "next allowed at" timestamps (token-bucket politeness). Seen-URL Bloom segments in RocksDB. Page content in S3, metadata in Postgres. Robots cache in Redis, 24h TTL, fail closed. A separate small headless-Chrome fleet for JS-heavy hosts with its own strict budget. Nightly job recomputes priority scores from the link graph.
Interview tips
- Say "politeness" in the first two minutes. It's the constraint the whole design hangs on β interviewers are listening for it.
- Draw the frontier as the centerpiece, not the fetcher. The queue is the system; fetching is plumbing.
- Mention robots.txt before they ask. Cache it, honor crawl-delay, fail closed β three phrases, full marks.
- Traps are your best "what breaks" answer. Infinite calendars and faceted URLs show you've thought adversarially.
- Close with freshness. Recrawl scheduling (news hourly, blogs weekly) turns a one-shot crawler into a real system.
π¬ Interactive: crawl frontier simulator
A tiny 3-host web. Watch BFS crawl it β with politeness delays, a robots.txt wall, a duplicate URL, and a calendar trap.
π Bloom filter note
Every discovered URL is hashed into a Bloom filter before fetching: O(1) "have we seen this?", ~10 bits per URL. Watch /p1?ref=sidebar β it normalizes to the same canonical URL as /p1, the filter says seen, and we skip the fetch. ~1% false positives just mean occasionally skipping a page that gets rediscovered later anyway.