System design interview Β· Search

Design a web crawler

Googlebot's job interview. Downloading pages is the easy 5% β€” the other 95% is deciding what to fetch next, being polite about it, and never fetching the same page twice.

🎯 The takeaway, first

The frontier is the system. A crawler is a polite, deduplicating, trap-avoiding URL scheduler with a downloader attached β€” like a librarian who must copy every book in the city but may only knock on each door once every few seconds, must obey "do not disturb" signs, and refuses to copy the same book twice. If your design starts with "fetch fast," you've missed the point: the binding constraint is courtesy, and the scale problem is the queue.

🚫 Misconception, busted

"A crawler just downloads pages as fast as possible." Speed is the one thing you are not allowed to maximize. Every host gets a politeness delay (default ~1 request/second, more if their robots.txt says so), and spider traps β€” infinite calendars, faceted-search URLs like ?color=red&size=m&page=48291 β€” will eat an impolite crawler alive. Most of the engineering is scheduling and dedup, not fetching.

Requirements

Say these out loud before drawing a single box.

Functional

  • Start from seed URLs, discover new URLs from page links
  • Fetch pages and extract outbound links
  • Respect robots.txt (allow/disallow, crawl-delay)
  • Deduplicate URLs and near-duplicate content
  • Store fetched pages for the indexer

Non-functional

  • 10B+ pages/day fetch capacity (~115K/s)
  • Politeness: default ≀ 1 req/s per host, honor crawl-delay
  • A dead host must not stall the crawl
  • Freshness: recrawl important pages (news hourly, blogs weekly)
  • Survive traps: no infinite loops, bounded per-host budgets

Back-of-the-envelope math

Google-scale-ish. Assumptions labeled.

WhatAssumptionMath
Fetch rate10B pages/dayβ‰ˆ 115K fetches/s sustained
Bandwidthavg page 50 KB115K Γ— 50 KB β‰ˆ 5.8 GB/s + DNS/TLS overhead
Frontier sizeweb holds trillions of known URLsqueue must hold 100B+ URLs β€” it lives on disk, not RAM
Seen-URL dedupBloom filter, 1T URLs, 1% false positives β‰ˆ 9.6 bits/URLβ‰ˆ 1.2 TB β€” sharded, or segmented per crawl batch
Politeness paradox100M hosts Γ— 1 req/s politenessTo feed 115K fetchers you need 115K distinct hosts with ripe queues at all times β€” that's the scheduling crux
Robots cache100M hosts Γ— ~2 KB rules, 24h TTLβ‰ˆ 200 GB β€” fits in a shared cache, refetched daily
Go deeper: why a Bloom filter and not a hash set?

An exact set of 1 trillion URLs at ~60 bytes each is 60 TB β€” you can't keep it in RAM and disk lookups at 115K/s would melt. A Bloom filter answers "have we seen this URL?" in O(1) with a few hash computations and ~10 bits per URL. The cost: ~1% false positives β€” we occasionally skip a page we hadn't seen. For a crawler that's a fine trade: the page will likely be rediscovered via another link later. Probabilistic data structures are the correct answer to web-scale dedup; say so confidently.

Architecture

The centerpiece: everything orbits the frontier.

flowchart LR
    SEED[Seed URLs] --> FR[Frontier
per-host queues + priority] FR --> SCH[Scheduler
politeness gate] SCH --> FE[Fetcher Pool] FE --> DNS[(DNS Cache)] FE --> RC[(Robots Cache)] RC -->|disallow| SKIP[Skip] RC -->|allow| GET[HTTP GET] GET --> PA[Parser] PA --> DED[Dedup
URL Bloom + content simhash] DED -->|new| PS[(Page Store)] DED -->|duplicate| DROP[Drop] PA --> LE[Link Extractor] LE --> NORM[URL Normalizer
strip tracking params, fragments] NORM --> FR PS --> IDX[Indexer
downstream]

The frontier holds one queue per host β€” this is the key structural decision, because politeness is per-host. The scheduler only hands a fetcher a URL whose host's delay has expired. Fetched pages go through robots check (cached), parse, dedup (URL-level Bloom filter, then content-level simhash for near-duplicates), and link extraction feeds normalized URLs back into the frontier. The indexer consumes the page store downstream and never talks to the crawl loop.

Go deeper: why partition the frontier by host?

Two reasons. First, politeness becomes local: a scheduler owning example.com's queue just tracks one "next allowed fetch" timestamp. Second, failure isolation: if a host hangs or blackholes connections, only its queue stalls β€” the other 99,999,999 hosts keep flowing. In a distributed build you'd consistent-hash hosts across frontier shards, so all URLs for one host always land on the same scheduler.

Component deep-dives

Fetching with politeness

sequenceDiagram
    participant S as Scheduler
    participant F as Frontier (per-host queue)
    participant R as Robots Cache
    participant W as Fetcher Worker
    participant H as example.com
    S->>F: next ripe host? (delay expired)
    F-->>S: example.com β†’ /blog/post-9
    S->>R: robots rules for example.com?
    alt cache miss
        S->>H: GET /robots.txt
        H-->>S: rules (cache 24h)
    end
    alt disallowed
        S->>F: drop URL, note the rule
    else allowed
        S->>W: fetch /blog/post-9
        W->>H: GET (1 req β€” then the host goes quiet)
        H-->>W: 200 HTML
        W->>F: mark host next-allowed = now + delay
    end

Dedup β€” two gates

sequenceDiagram
    participant P as Parser
    participant B as URL Bloom Filter
    participant C as Content Simhash Index
    participant S as Page Store
    P->>P: normalize URL (lowercase host, strip ?utm_*, #frag)
    P->>B: seen this URL?
    alt yes β†’ drop (URL dupe)
        B-->>P: probably seen β€” skip fetch entirely
    else no
        P->>P: fetch + parse
        P->>C: simhash(content) near-duplicate?
        alt near-dupe of known page β†’ drop
            C-->>P: Hamming distance < 3 β€” skip
        else new
            P->>S: store page
            P->>B: add URL
        end
    end

API + data model

A crawler has no public API β€” its "API" is the frontier's internal contract. This is what the scheduler reads and writes:

frontier_entry { url, host, priority_score, next_fetch_at, retries, discovered_from }
page           { url_hash PK, url, content_hash, fetched_at, content_ptr β†’ blob store }
robots_cache   { host PK, rules, crawl_delay_s, fetched_at }   -- 24h TTL
host_stats     { host PK, urls_seen, urls_fetched, errors_24h, budget_remaining }

Trade-offs

DecisionOption AOption BPick
Crawl orderBFS (fair, simple)Priority (PageRank-ish, important pages first)Priority β€” freshness where it matters; BFS starves news
URL dedupExact hash set (no false drops)Bloom filter (tiny, 1% false drops)Bloom β€” a skipped page gets rediscovered
FrontierCentral queueSharded by host hashSharded β€” politeness stays local, failures contained
FetchingRaw HTTPHeadless browser (renders JS)Raw HTTP first; headless only for JS-heavy hosts, budgeted separately
robots.txt fetch failsFail open (crawl anyway)Fail closed (skip host)Fail closed, retry later β€” politeness is the brand

Failure modes

Spider trap: /calendar/2026/01 β†’ /02 β†’ /03 … foreverPer-host URL budget + pattern detection (same path prefix exploding) β†’ quarantine host
Host blackholes connections (hangs)Aggressive timeouts; per-host circuit breaker; only that host's queue stalls
DNS slow/poisonedShared DNS cache with TTL; cap concurrent DNS lookups; DNS failure = host paused, not worker blocked
Content farm: 1M near-duplicate pagesSimhash gate drops them before the page store; host budget still applies
Frontier shard diesHost→shard mapping is deterministic; replay the shard's queue from durable log onto a new node

What I'd actually build

Opinionated. Steal this for the interview.

Stack: Go fetcher pool (great at 100K concurrent connections). Frontier as Kafka topics partitioned by host hash + Redis per-host "next allowed at" timestamps (token-bucket politeness). Seen-URL Bloom segments in RocksDB. Page content in S3, metadata in Postgres. Robots cache in Redis, 24h TTL, fail closed. A separate small headless-Chrome fleet for JS-heavy hosts with its own strict budget. Nightly job recomputes priority scores from the link graph.

Interview tips

πŸ”¬ Interactive: crawl frontier simulator

A tiny 3-host web. Watch BFS crawl it β€” with politeness delays, a robots.txt wall, a duplicate URL, and a calendar trap.

fetched queued robots-blocked duplicate skipped trap blocked
0fetched
0in frontier
0robots-blocked
0dupes skipped
0trap blocked

πŸ” Bloom filter note

Every discovered URL is hashed into a Bloom filter before fetching: O(1) "have we seen this?", ~10 bits per URL. Watch /p1?ref=sidebar β€” it normalizes to the same canonical URL as /p1, the filter says seen, and we skip the fetch. ~1% false positives just mean occasionally skipping a page that gets rediscovered later anyway.

v2026.10.03-01