Deploy Crawlab
Crawlab Is a Distributed Web Crawler Admin Platform
crawlab-master
Just deployed
/root/.crawlab
mongo
Just deployed
/data/db
crawlab-worker
Just deployed
/root/.crawlab
Deploy and Host Crawlab on Railway
Crawlab is an open-source, distributed web crawler management platform designed to organize, schedule, execute, and monitor web spiders (Scrapy, Puppeteer, Selenium, Playwright, Python, Golang, Node.js) through a powerful web-based dashboard. This Railway template deploys a two-tier stack featuring the Crawlab Master Node (crawlab-master) paired with a dedicated MongoDB Database (mongo).
About Hosting Crawlab
Hosting Crawlab on Railway deploys a two-container architecture connected via Railway's private network:
- Crawlab Master Node (
crawlab-master): Core control server runningcrawlabteam/crawlab:latest. It hosts the web administration dashboard, API endpoints, spider task scheduler, and node controller. It listens internally on port8080and mounts a persistent volume at/root/.crawlabto preserve configuration files, uploaded spiders, and task execution logs. - MongoDB Database (
mongo): Document database usingmongo:4.2. It stores spider metadata, execution results, task logs, cron schedules, user accounts, and system state, with persistent data secured on a volume mounted at/data/db.
Common Use Cases
- Centralized Web Spider Management: Upload, configure, run, and monitor Scrapy, Python, Node.js, or custom web crawlers from a single unified web console.
- Automated Task Scheduling: Configure cron-like automated schedules for recurring web scraping and data extraction pipelines.
- Real-time Task Analytics & Log Viewer: Inspect execution logs, task statuses, success/failure rates, and runtime statistics in real time.
- Multi-language Scraping Infrastructure: Execute web crawlers written in Python, Node.js, Go, or shell scripts without maintaining manual server cron jobs or SSH scripts.
Dependencies for Crawlab Hosting
- Crawlab Master image:
crawlabteam/crawlab:latest - MongoDB image:
mongo:4.2 - Two persistent volumes:
/root/.crawlaboncrawlab-master/data/dbonmongo
- Railway public domain mapped to port
8080oncrawlab-master - Auto-generated database secret:
MONGO_INITDB_ROOT_PASSWORD(24-character secret)
Upstream: Crawlab Site · GitHub (crawlab-team/crawlab) · Docker Hub
Implementation Details
| Service | Image | Role | Web Port / Internal | Volume Mount |
|---|---|---|---|---|
| crawlab-master | crawlabteam/crawlab:latest | Master Controller & Web Console | 8080 | /root/.crawlab |
| mongo | mongo:4.2 | Document Database Backend | 27017 | /data/db |
Topology
| Service | Role | Volume | Public | Notes |
|---|---|---|---|---|
| crawlab-master | Master Node & Web UI | /root/.crawlab | Yes (Port 8080) | Connects to mongo via ${{mongo.RAILWAY_PRIVATE_DOMAIN}} |
| mongo | Database Backend | /data/db | TCP Proxy (Port 27017) | Stores spider state, logs, and credentials |
Volumes (drives) — what to mount
| Service | Mount path | What is stored |
|---|---|---|
| crawlab-master | /root/.crawlab | Crawlab workspace settings, spider code uploads, and execution state |
| mongo | /data/db | MongoDB data directory, collection indexes, task execution history, and user profiles |
Warning: Do not remove or detach the persistent volumes mounted at
/root/.crawlabor/data/db— deleting these volumes will result in permanent loss of uploaded spiders, task logs, user credentials, and database records during redeployments.
Quick Start
- Click the Deploy on Railway button above.
- Sign in (or create a free Railway account) and click Deploy.
- Wait 2–3 minutes for both
mongoandcrawlab-masterservices and persistent volumes to provision. - Open the crawlab-master service → Settings → Networking and click the generated public domain URL.
- Log in to the Crawlab dashboard using the default credentials:
- Username:
admin - Password:
admin
- Username:
- Immediately change the default administrator password under User Settings in the Crawlab web interface.
Configuration
Crawlab Master Variables (crawlab-master)
| Variable | Default / Source | Description / Notes |
|---|---|---|
CRAWLAB_NODE_MASTER | Y | Configures container as master node in Crawlab cluster |
CRAWLAB_MONGO_HOST | ${{mongo.RAILWAY_PRIVATE_DOMAIN}} | Internal private hostname for MongoDB service |
CRAWLAB_MONGO_PORT | 27017 | Internal port for MongoDB connection |
CRAWLAB_MONGO_DB | crawlab | Database name used by Crawlab |
CRAWLAB_MONGO_USERNAME | crawlab | Database user for Crawlab connection |
CRAWLAB_MONGO_PASSWORD | ${{mongo.MONGO_INITDB_ROOT_PASSWORD}} | Database password referenced from mongo service |
CRAWLAB_MONGO_AUTHSOURCE | admin | Authentication database source |
PORT | 8080 | Internal HTTP port the web console listens on |
MongoDB Variables (mongo)
| Variable | Default / Source | Description / Notes |
|---|---|---|
MONGO_INITDB_ROOT_USERNAME | crawlab | Initial root database username |
MONGO_INITDB_ROOT_PASSWORD | Auto-generated secret (24 chars) | Secure password for MongoDB root access |
MONGOHOST | ${{RAILWAY_PRIVATE_DOMAIN}} | Internal private network host domain |
MONGOPORT | 27017 | Database port inside private network |
MONGO_URL | mongodb://${{MONGO_INITDB_ROOT_USERNAME}}:${{MONGO_INITDB_ROOT_PASSWORD}}@${{RAILWAY_PRIVATE_DOMAIN}}:27017 | Internal private connection URL |
DATABASE_URL | Same as MONGO_URL | Secondary private database connection string |
MONGO_PUBLIC_URL | Connection string via TCP Proxy | External connection URL for remote database tools |
Custom Domain
- Open the crawlab-master service → Settings → Networking → Custom Domain.
- Add your custom domain and follow Railway’s DNS configuration instructions.
- Railway provisions TLS automatically.
- Access your Crawlab console securely at
https://your.custom.domain.
Updating Crawlab
- Open the crawlab-master service → Settings → Source.
- Update the image tag (e.g.,
crawlabteam/crawlab:latestto a specific release tag). - Click Redeploy.
All uploaded spider scripts, scheduled jobs, task logs, and user records remain safe inside the /root/.crawlab and /data/db volumes.
Traps
Common pitfalls and failure modes:
- Default Login Credentials Vulnerability — Crawlab boots with default credentials (
admin/admin). Change the admin password immediately upon first login to prevent unauthorized access to your server. - Database Startup Timing — During initial deployment,
crawlab-masterattempts to connect to MongoDB over private networking. If the web console fails to start, wait formongovolume provisioning to finish and redeploycrawlab-master. - Missing Persistent Volumes — Detaching or omitting the
/root/.crawlabor/data/dbvolumes will cause all uploaded spiders, task logs, and database records to reset on container updates. - Spider Package Dependencies — Custom Python (
pip) or Node.js (npm) packages installed directly inside container shells will reset on image redeployments. Include dependencies in spider ZIP bundles or virtual environments saved within/root/.crawlab.
Why Deploy Crawlab on Railway?
Railway provides a seamless, high-availability platform for multi-tier web applications and databases. Deploying Crawlab on Railway gives you private service networking, persistent volume backups, automated HTTPS certificates, and zero-maintenance cloud container execution.
Template Content
