Deploy Crawlab

Crawlab Is a Distributed Web Crawler Admin Platform

Deploy Crawlab

Just deployed

/root/.crawlab

Just deployed

/data/db

Just deployed

/root/.crawlab

Deploy on Railway

Deploy and Host Crawlab on Railway

Crawlab is an open-source, distributed web crawler management platform designed to organize, schedule, execute, and monitor web spiders (Scrapy, Puppeteer, Selenium, Playwright, Python, Golang, Node.js) through a powerful web-based dashboard. This Railway template deploys a two-tier stack featuring the Crawlab Master Node (crawlab-master) paired with a dedicated MongoDB Database (mongo).


About Hosting Crawlab

Hosting Crawlab on Railway deploys a two-container architecture connected via Railway's private network:

  • Crawlab Master Node (crawlab-master): Core control server running crawlabteam/crawlab:latest. It hosts the web administration dashboard, API endpoints, spider task scheduler, and node controller. It listens internally on port 8080 and mounts a persistent volume at /root/.crawlab to preserve configuration files, uploaded spiders, and task execution logs.
  • MongoDB Database (mongo): Document database using mongo:4.2. It stores spider metadata, execution results, task logs, cron schedules, user accounts, and system state, with persistent data secured on a volume mounted at /data/db.

Common Use Cases
  • Centralized Web Spider Management: Upload, configure, run, and monitor Scrapy, Python, Node.js, or custom web crawlers from a single unified web console.
  • Automated Task Scheduling: Configure cron-like automated schedules for recurring web scraping and data extraction pipelines.
  • Real-time Task Analytics & Log Viewer: Inspect execution logs, task statuses, success/failure rates, and runtime statistics in real time.
  • Multi-language Scraping Infrastructure: Execute web crawlers written in Python, Node.js, Go, or shell scripts without maintaining manual server cron jobs or SSH scripts.

Dependencies for Crawlab Hosting
  • Crawlab Master image: crawlabteam/crawlab:latest
  • MongoDB image: mongo:4.2
  • Two persistent volumes:
    • /root/.crawlab on crawlab-master
    • /data/db on mongo
  • Railway public domain mapped to port 8080 on crawlab-master
  • Auto-generated database secret: MONGO_INITDB_ROOT_PASSWORD (24-character secret)

Upstream: Crawlab Site · GitHub (crawlab-team/crawlab) · Docker Hub

Implementation Details
ServiceImageRoleWeb Port / InternalVolume Mount
crawlab-mastercrawlabteam/crawlab:latestMaster Controller & Web Console8080/root/.crawlab
mongomongo:4.2Document Database Backend27017/data/db

Topology
ServiceRoleVolumePublicNotes
crawlab-masterMaster Node & Web UI/root/.crawlabYes (Port 8080)Connects to mongo via ${{mongo.RAILWAY_PRIVATE_DOMAIN}}
mongoDatabase Backend/data/dbTCP Proxy (Port 27017)Stores spider state, logs, and credentials
Volumes (drives) — what to mount
ServiceMount pathWhat is stored
crawlab-master/root/.crawlabCrawlab workspace settings, spider code uploads, and execution state
mongo/data/dbMongoDB data directory, collection indexes, task execution history, and user profiles

Warning: Do not remove or detach the persistent volumes mounted at /root/.crawlab or /data/db — deleting these volumes will result in permanent loss of uploaded spiders, task logs, user credentials, and database records during redeployments.


Quick Start
  1. Click the Deploy on Railway button above.
  2. Sign in (or create a free Railway account) and click Deploy.
  3. Wait 2–3 minutes for both mongo and crawlab-master services and persistent volumes to provision.
  4. Open the crawlab-master service → Settings → Networking and click the generated public domain URL.
  5. Log in to the Crawlab dashboard using the default credentials:
    • Username: admin
    • Password: admin
  6. Immediately change the default administrator password under User Settings in the Crawlab web interface.

Configuration
Crawlab Master Variables (crawlab-master)
VariableDefault / SourceDescription / Notes
CRAWLAB_NODE_MASTERYConfigures container as master node in Crawlab cluster
CRAWLAB_MONGO_HOST${{mongo.RAILWAY_PRIVATE_DOMAIN}}Internal private hostname for MongoDB service
CRAWLAB_MONGO_PORT27017Internal port for MongoDB connection
CRAWLAB_MONGO_DBcrawlabDatabase name used by Crawlab
CRAWLAB_MONGO_USERNAMEcrawlabDatabase user for Crawlab connection
CRAWLAB_MONGO_PASSWORD${{mongo.MONGO_INITDB_ROOT_PASSWORD}}Database password referenced from mongo service
CRAWLAB_MONGO_AUTHSOURCEadminAuthentication database source
PORT8080Internal HTTP port the web console listens on
MongoDB Variables (mongo)
VariableDefault / SourceDescription / Notes
MONGO_INITDB_ROOT_USERNAMEcrawlabInitial root database username
MONGO_INITDB_ROOT_PASSWORDAuto-generated secret (24 chars)Secure password for MongoDB root access
MONGOHOST${{RAILWAY_PRIVATE_DOMAIN}}Internal private network host domain
MONGOPORT27017Database port inside private network
MONGO_URLmongodb://${{MONGO_INITDB_ROOT_USERNAME}}:${{MONGO_INITDB_ROOT_PASSWORD}}@${{RAILWAY_PRIVATE_DOMAIN}}:27017Internal private connection URL
DATABASE_URLSame as MONGO_URLSecondary private database connection string
MONGO_PUBLIC_URLConnection string via TCP ProxyExternal connection URL for remote database tools
Custom Domain
  1. Open the crawlab-master service → Settings → Networking → Custom Domain.
  2. Add your custom domain and follow Railway’s DNS configuration instructions.
  3. Railway provisions TLS automatically.
  4. Access your Crawlab console securely at https://your.custom.domain.

Updating Crawlab
  1. Open the crawlab-master service → Settings → Source.
  2. Update the image tag (e.g., crawlabteam/crawlab:latest to a specific release tag).
  3. Click Redeploy.

All uploaded spider scripts, scheduled jobs, task logs, and user records remain safe inside the /root/.crawlab and /data/db volumes.


Traps

Common pitfalls and failure modes:

  • Default Login Credentials Vulnerability — Crawlab boots with default credentials (admin / admin). Change the admin password immediately upon first login to prevent unauthorized access to your server.
  • Database Startup Timing — During initial deployment, crawlab-master attempts to connect to MongoDB over private networking. If the web console fails to start, wait for mongo volume provisioning to finish and redeploy crawlab-master.
  • Missing Persistent Volumes — Detaching or omitting the /root/.crawlab or /data/db volumes will cause all uploaded spiders, task logs, and database records to reset on container updates.
  • Spider Package Dependencies — Custom Python (pip) or Node.js (npm) packages installed directly inside container shells will reset on image redeployments. Include dependencies in spider ZIP bundles or virtual environments saved within /root/.crawlab.

Why Deploy Crawlab on Railway?

Railway provides a seamless, high-availability platform for multi-tier web applications and databases. Deploying Crawlab on Railway gives you private service networking, persistent volume backups, automated HTTPS certificates, and zero-maintenance cloud container execution.



Template Content

More templates in this category

View Template
NEW
Swarm
Named LLM bots join channels, take jobs, run routines on your Railway box

mcmax
3
View Template
Telegram JavaScript Bot
A template for Telegram bot in JavaScript using grammY

Agampreet Singh
294
View Template
Cobalt Tools [Updated Sep ’26]
Cobalt Tools [Sep ’26] (Media Downloader, Converter & Automation) Self Host

shinyduo
299