Building public-data crawlers
Build an ethical crawler in six steps with Playwright, http_utils, and APScheduler.
- Difficulty
- Intermediate
- Lessons
- 6
Public data like NPS, DART, and HIRA is accessible to everyone, but automation comes with rules — robots.txt, rate limits, terms of service. Six steps to an ethical and sustainable crawler.
Who it's for
- Developers who need more control than portal APIs offer
- Anyone who has been blocked by a crawl target
- Teams who want incremental collection, schedules, and observability
What you can do afterwards
- Separate dynamic pages (Playwright) from static ones (BS4)
- Apply robots.txt + rate limit + backoff
- Schedule in KST with APScheduler
- Combine public APIs, ministry CSVs, and web scraping
- Incremental collection, dedup, checkpoints
- Healthchecks and failure alerts
Flow
A safe collection pipeline
Check ethical and legal constraints before choosing a collection tool.
Protect the source system with rate limits and schedules.
Make reruns safe through incremental processing and deduplication.
Detect failures and stale data with metrics and alerts.
The goal of a crawler: refresh our DB without loss while not trespassing on the source site. The flow above strengthens both axes in turn.
Steps
- Crawler ethics and legal boundaries — robots.txt · terms · personal data
- Static vs dynamic — BS4 + Playwright — pick the right tool
- Rate limiting · retries · backoff — exponential + jitter
- APScheduler + KST — idempotency ·
replace_existing=True· double-trigger defence - Incremental collection · deduplication — checkpoints · unique keys · change detection
- Observability · alerts — success rate · latency · Slack · PagerDuty
Prerequisites — complete python-data-pipeline.
Lessons
Other courses
All courses →- Production Engineering — Boundaries, Performance, Recovery, and Delivery in 14 Steps
- Getting Started with a Dev Environment
- From HTML/CSS/JS to React, Next.js, Tailwind
- Build Your First Fullstack App with Next.js 16
- Backend with Spring Boot 4
- Python · FastAPI · Data Pipelines
- AI-native developer tooling — Claude Code · MCP · design tools
- Docker · Caddy · Cloud — 10 deploy options
- Operations console design — many resources in one view
- Local LLM · pgvector · building a RAG chatbot
- Tauri 2 — desktop · mobile in one codebase
- Testing strategy and quality gates
- Web security foundations — JWT · OAuth · OWASP
- PostgreSQL in depth + Redis · Kafka
- Monorepo · SSOT · layer separation thinking