06 / PROJECTDistributed Systems

Distributed Web Crawler

Coordinating five workers over shared Redis state

A multi-worker crawler where the interesting problem is coordination: a shared frontier, a shared visited set, and politeness rules that must hold across every worker at once.

Year
2026
Status
Building
Technology
Python · Redis · MongoDB · BeautifulSoup · Docker Compose · Streamlit
01 / SYSTEM MAP

Architecture

Distributed Systems / flow6 stages

Redis sorted set

Frontier

Priority-ordered URL queue shared by every worker.

Redis frontier → 5 workers → politeness gate → visited set → MongoDB → Streamlit monitor

02 / CONTEXT

The problem

Parallel crawlers duplicate work and hammer hosts unless the queue, the visited set and the rate limit are shared state rather than per-worker state.

Engineering challenge

Keeping five concurrent workers from duplicating fetches or violating per-domain politeness, when the constraint has to hold globally rather than per worker.

03 / DESIGN DECISION

Approach

Moved queue state, deduplication and rate limiting into Redis so all workers share one view. The frontier is a sorted set giving priority ordering, the visited set prevents duplicate fetches, and the 1-second per-domain gap is enforced centrally. The whole topology runs under Docker Compose.

Redis
Sorted sets give a priority frontier; shared sets give cheap cross-worker deduplication.
MongoDB
Schema-flexible storage for heterogeneous crawled documents.
Docker Compose
Reproducible multi-service topology in one definition.
Streamlit
Minimal live view of queue depth without building an admin UI.
04 / IMPLEMENTATION

Engineering detail

  1. 015 worker threads pulling from one shared frontier
  2. 02Redis sorted-set frontier providing priority ordering across workers
  3. 03Shared visited-set coordination, so no URL is fetched twice
  4. 04robots.txt compliance checked before fetching
  5. 051-second per-domain request gap enforced across all workers
  6. 06MongoDB persistence for crawled documents
  7. 07Docker Compose deployment of workers, Redis and MongoDB together
  8. 08Live queue monitoring through a Streamlit dashboard
05 / OPERATING PRINCIPLES

Engineering practices

  • Coordination state centralised rather than duplicated per worker.
  • Politeness enforced before fetch, not as a retry.
  • Whole topology reproducible from one Compose file.
06 / CAPABILITIES

Features

  • Redis sorted-set frontier.
  • Cross-worker deduplication.
  • robots.txt and rate-limit compliance.
  • Live queue monitoring.