---
title: "Deploy anydoc | (Just Updated) Any Document to Markdown API, Japanese & Thai OCR"
description: "Word, PowerPoint, Excel, PDF and scans to Markdown; Japanese/Thai OCR"
category: "AI/ML"
url: https://railway.com/deploy/anydoc-or-just-updated-any-document-to-m
---

# Deploy anydoc | (Just Updated) Any Document to Markdown API, Japanese & Thai OCR

Word, PowerPoint, Excel, PDF and scans to Markdown; Japanese/Thai OCR

**[Deploy anydoc | (Just Updated) Any Document to Markdown API, Japanese & Thai OCR on Railway](https://railway.com/template/anydoc-or-just-updated-any-document-to-m)**

Machine-readable deploy manifest (JSON, validated by TemplateCI): https://railway.com/deploy/anydoc-or-just-updated-any-document-to-m/manifest.json

- **Creator:** SuperSlowSloth
- **Category:** AI/ML

## Template content

### anydoc

- **Image:** ghcr.io/bon5co/anydoc-serve:latest
- **Health check:** /health
- **Public domain:** Yes

## Documentation

# Deploy and Host anydoc on Railway

anydoc-serve turns Word, PowerPoint, Excel, OpenDocument, PDF, EPUB, RTF and CSV files into clean Markdown over HTTP and MCP. It wraps [firecrawl/anydoc](https://github.com/firecrawl/anydoc), Firecrawl's fast Rust document converter, in one small container and adds local OCR for scans and images in English, Japanese and Thai.

## About Hosting anydoc

This template runs one service from `ghcr.io/bon5co/anydoc-serve:latest` (amd64 and arm64). There is no database and no volume: every request is converted in memory and nothing is stored. Born-digital documents go through anydoc unchanged; scanned or image-only pages, and images (`.png`, `.jpg`, `.tiff`, `.webp`), are OCR'd on CPU with Tesseract 5, reading Japanese, Thai and English in one pass, so nothing leaves your deployment. On a mixed PDF only the scanned pages are OCR'd. Every endpoint except `/health` requires `Authorization: Bearer $API_KEY`, and the key is generated for you at deploy, so no deployment is ever open. A `/health` healthcheck keeps traffic off the service until it is ready. Measured on a Railway deploy of this template: 246-263 MiB of RAM idle and 272 MiB after converting a DOCX and OCR'ing two images, under the Free plan's 0.5 GB cap.

## Why Deploy anydoc on Railway?

Railway is a singular platform to deploy your infrastructure stack. Railway will host your infrastructure so you don't have to deal with configuration, while allowing you to vertically and horizontally scale it.

By deploying anydoc on Railway, you are one step closer to supporting a complete full-stack application with minimal burden. Host your servers, databases, AI agents, and more on Railway.

- **Scans too, not just born-digital files** — plain anydoc refuses scanned PDFs and images; this server OCRs them locally instead of sending them to a hosted OCR API.
- **Japanese and Thai OCR out of the box** — `OCR_LANGS` defaults to `eng,jpn,tha`, and one page mixing all three scripts is read line by line.
- **Private by default** — a 32-character API key is generated per deploy; requests without it get `401`.
- **Built for agents** — an MCP server at `/mcp` exposes a `convert_document` tool for Claude Code, Cursor and any client that speaks streamable HTTP.
- **Small** — one service, no volume, about 250 MiB idle, 272 MiB after OCR.

## Common Use Cases

- **RAG ingestion** — convert uploaded Office files, PDFs and scans to Markdown before chunking and embedding.
- **AI agent tools** — give Claude Code or Cursor a `convert_document` tool over MCP so the agent can read any attachment.
- **Japanese and Thai back-office documents** — OCR scanned invoices, forms and contracts without a cloud OCR bill.

## Dependencies for anydoc Hosting

- Nothing beyond the one service in this template: no database, no volume, no external API.

### Deployment Dependencies

- [bon5co/anydoc-serve](https://github.com/bon5co/anydoc-serve) — the server this template deploys (source, Dockerfile, OCR benchmark)
- [firecrawl/anydoc](https://github.com/firecrawl/anydoc) by [Firecrawl](https://firecrawl.dev) — the document converter doing all non-OCR conversion (MIT). anydoc-serve is not affiliated with Firecrawl.
- [Tesseract](https://github.com/tesseract-ocr/tesseract) — the OCR engine (Apache 2.0)

### Implementation Details

Convert a document to Markdown:

```bash
curl -H "Authorization: Bearer $API_KEY" -H "Accept: text/markdown" \
  -F file=@report.docx https://your-deployment.up.railway.app/v1/convert
```

OCR a scan or an image, getting JSON with the Markdown plus the pages that went through OCR:

```bash
curl -H "Authorization: Bearer $API_KEY" \
  -F file=@scan.pdf https://your-deployment.up.railway.app/v1/convert
```

Convert from a URL:

```bash
curl -H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/report.pdf"}' https://your-deployment.up.railway.app/v1/convert
```

Add it to Claude Code as an MCP server:

```bash
claude mcp add --transport http anydoc https://your-deployment.up.railway.app/mcp \
  --header "Authorization: Bearer $API_KEY"
```

| Variable | Default | Meaning |
| --- | --- | --- |
| `API_KEY` | generated (32 chars) | Bearer key required by every endpoint except `/health` |
| `OCR_LANGS` | `eng,jpn,tha` | OCR languages, comma-separated |

More settings (`OCR_ENABLED`, `OCR_DPI`, `MAX_UPLOAD_MB`, `CONVERT_TIMEOUT_S`, ...) are documented in the [anydoc-serve README](https://github.com/bon5co/anydoc-serve#configuration).


## Similar templates

- [Chat Chat](https://railway.com/deploy/-WWW5r) — Chat Chat, your own unified chat and search to AI platform.
- [stella](https://railway.com/deploy/stella) — Self-host stella with web, API, Postgres, Redis, and object storage.
- [Hermes Agent | OpenClaw Alternative with Dashboard](https://railway.com/deploy/hermes-agent-or-openclaw-alternative-wit) — Self-Hosted Hermes AI Agent for Telegram, Discord & Slack

Open this page in a browser: https://railway.com/deploy/anydoc-or-just-updated-any-document-to-m
