Deploy nolock + llama.cpp: ai coding harness + local slms (micro-agents)

ai coding harness + local slms (micro-agents)

Deploy nolock + llama.cpp: ai coding harness + local slms (micro-agents)

Just deployed

/data

Just deployed

/models

Deploy and Host nolock + llama.cpp: ai coding harness + local slms (micro-agents) on Railway

nolock is an AI-native coding harness — a full IDE with chat, agents, micro-agents, tools, terminal, notebooks, and git session diffs. Paired with llama.cpp, it runs local small language models (SLMs) for private, free-token micro-agent execution — all from your browser, no API keys, no per-token bills.

About Hosting nolock + llama.cpp: ai coding harness + local slms (micro-agents)

This stack deploys two isolated Railway services. nolock serves the web frontend + Rust backend (the exact same codebase as the Tauri v2 desktop app) — chat, agents, micro-agents, file tools, search, linter, sessions, and git diffs. llama.cpp runs the official server image wrapped with a small model store, and pulls a GGUF model from HuggingFace into a persistent volume on first boot. The services talk over Railway's private network — llama.cpp is never exposed publicly. The model is a configuration (MODEL_HF), swapped by changing a variable and redeploying; models pulled from the nolock UI persist on the volume, so switching back is instant. Tool calling works end-to-end: the SLM calls tools (file ops, search, git) and uses the results in its answer.

Common Use Cases

  • Private AI coding assistant — an agentic IDE with zero per-token costs and prompts that never leave your Railway instance
  • Local micro-agents — code review, research, and writing agents powered by small local models, spawned from chat
  • Free-token API backend — llama.cpp exposes an OpenAI-compatible /v1/chat/completions endpoint you own
  • Self-hosted AI workspace — persistent sessions, git diffs, terminal, and notebooks, all in the browser

Dependencies

  • nolock web app — React frontend + Rust backend (nolock-server), the same codebase as the Tauri desktop app
  • llama.cpp server + model store — official ghcr.io/ggml-org/llama.cpp:server image wrapped by this repo's deploy/llamacpp/Dockerfile; GGUF models persist at /models
  • Railway private networking — nolock reaches llama.cpp via http://llamacpp.railway.internal:8080 (inference) and :8081 (model store); neither is public
  • Persistent volumes — one for GGUF models (/models), one for nolock state (/data → NOLOCK_DATA_DIR=/data/nolock): secrets, RLHF logs, and opened projects survive redeploys
  • HuggingFace GGUF model — any GGUF repo, e.g. impacte/ullr:Q4_K_M

After Deploying

  1. Open your nolock domain. The app shows a login page — paste the NOLOCK_WEB_TOKEN value (Railway → nolock service → Variables), or open https:///?token= to sign in automatically. The token is remembered per browser tab and sent as a Bearer header on every API call.
  2. The llama.cpp provider is auto-wired — the app fetches the internal llama.cpp URL from the server, so chat and FIM completions work out of the box.
  3. Pull models from the UI: Model Providers → select llama.cpp → paste a Hugging Face GGUF id (e.g. owner/model-GGUF:Q4_K_M) → Pull. Downloads run beside inference on the private model-store port and persist under /models/pulled/ on the volume.

Implementation Details

llama.cpp service (wrapper image from deploy/llamacpp/Dockerfile, no public domain):

# Start command
/entrypoint.sh   # starts the model store on private port 8081, then llama-server on 8080

# Variables
MODEL_HF=impacte/ullr:Q4_K_M        # the model is a configuration — swap and redeploy
LLAMA_CACHE=/models                 # persistent volume — model downloads once, survives redeploys
NOLOCK_MODEL_DIR=/models            # model store directory (same volume)
NOLOCK_MODEL_STORE_PORT=8081        # private model-store port
PORT=8080                           # inference port
NOLOCK_MODEL_PULL_TOKEN=    # shared with nolock — authenticates UI model pulls

nolock service (builds deploy/Dockerfile, healthcheck /health):

# Variables
LLAMACPP_URL=http://llamacpp.railway.internal:8080              # private inference
LLAMACPP_MODEL_STORE_URL=http://llamacpp.railway.internal:8081  # private model store
NOLOCK_DATA_DIR=/data/nolock   # persistent volume — secrets, sessions, projects
NOLOCK_WEB_TOKEN=   # login token for the public UI
NOLOCK_MODEL_PULL_TOKEN=

Both services share one random NOLOCK_MODEL_PULL_TOKEN so the nolock server can proxy authenticated model-pull requests to the model store. Railway credentials are never used by the app or sent to the browser. For gated HuggingFace repos, set HF_TOKEN on llamacpp.


Template Content

More templates in this category

View Template
Chat Chat
Chat Chat, your own unified chat and search to AI platform.

okisdev
116
View Template
stella
Self-host stella with web, API, Postgres, Redis, and object storage.

Jan Kubica
7
View Template
Hermes Agent | OpenClaw Alternative with Dashboard
Self-Hosted Hermes AI Agent for Telegram, Discord & Slack

codestorm
82