Deploy nolock + llama.cpp: ai coding harness + local slms (micro-agents)
ai coding harness + local slms (micro-agents)
nolock
Just deployed
/data
llama.cpp
Just deployed
/models
Deploy and Host nolock + llama.cpp: ai coding harness + local slms (micro-agents) on Railway
nolock is an AI-native coding harness — a full IDE with chat, agents, micro-agents, tools, terminal, notebooks, and git session diffs. Paired with llama.cpp, it runs local small language models (SLMs) for private, free-token micro-agent execution — all from your browser, no API keys, no per-token bills.
About Hosting nolock + llama.cpp: ai coding harness + local slms (micro-agents)
This stack deploys two isolated Railway services. nolock serves the web frontend + Rust backend (the exact same codebase as the Tauri v2 desktop app) — chat, agents, micro-agents, file tools, search, linter, sessions, and git diffs. llama.cpp runs the official server image wrapped with a small model store, and pulls a GGUF model from HuggingFace into a persistent volume on first boot. The services talk over Railway's private network — llama.cpp is never exposed publicly. The model is a configuration (MODEL_HF), swapped by changing a variable and redeploying; models pulled from the nolock UI persist on the volume, so switching back is instant. Tool calling works end-to-end: the SLM calls tools (file ops, search, git) and uses the results in its answer.
Common Use Cases
- Private AI coding assistant — an agentic IDE with zero per-token costs and prompts that never leave your Railway instance
- Local micro-agents — code review, research, and writing agents powered by small local models, spawned from chat
- Free-token API backend — llama.cpp exposes an OpenAI-compatible
/v1/chat/completionsendpoint you own - Self-hosted AI workspace — persistent sessions, git diffs, terminal, and notebooks, all in the browser
Dependencies
- nolock web app — React frontend + Rust backend (
nolock-server), the same codebase as the Tauri desktop app - llama.cpp server + model store — official
ghcr.io/ggml-org/llama.cpp:serverimage wrapped by this repo'sdeploy/llamacpp/Dockerfile; GGUF models persist at/models - Railway private networking — nolock reaches llama.cpp via
http://llamacpp.railway.internal:8080(inference) and:8081(model store); neither is public - Persistent volumes — one for GGUF models (
/models), one for nolock state (/data→NOLOCK_DATA_DIR=/data/nolock): secrets, RLHF logs, and opened projects survive redeploys - HuggingFace GGUF model — any GGUF repo, e.g.
impacte/ullr:Q4_K_M
After Deploying
- Open your nolock domain. The app shows a login page — paste the
NOLOCK_WEB_TOKENvalue (Railway → nolock service → Variables), or openhttps:///?token=to sign in automatically. The token is remembered per browser tab and sent as a Bearer header on every API call. - The llama.cpp provider is auto-wired — the app fetches the internal llama.cpp URL from the server, so chat and FIM completions work out of the box.
- Pull models from the UI: Model Providers → select llama.cpp → paste a Hugging Face GGUF id (e.g.
owner/model-GGUF:Q4_K_M) → Pull. Downloads run beside inference on the private model-store port and persist under/models/pulled/on the volume.
Implementation Details
llama.cpp service (wrapper image from deploy/llamacpp/Dockerfile, no public domain):
# Start command
/entrypoint.sh # starts the model store on private port 8081, then llama-server on 8080
# Variables
MODEL_HF=impacte/ullr:Q4_K_M # the model is a configuration — swap and redeploy
LLAMA_CACHE=/models # persistent volume — model downloads once, survives redeploys
NOLOCK_MODEL_DIR=/models # model store directory (same volume)
NOLOCK_MODEL_STORE_PORT=8081 # private model-store port
PORT=8080 # inference port
NOLOCK_MODEL_PULL_TOKEN= # shared with nolock — authenticates UI model pulls
nolock service (builds deploy/Dockerfile, healthcheck /health):
# Variables
LLAMACPP_URL=http://llamacpp.railway.internal:8080 # private inference
LLAMACPP_MODEL_STORE_URL=http://llamacpp.railway.internal:8081 # private model store
NOLOCK_DATA_DIR=/data/nolock # persistent volume — secrets, sessions, projects
NOLOCK_WEB_TOKEN= # login token for the public UI
NOLOCK_MODEL_PULL_TOKEN=
Both services share one random NOLOCK_MODEL_PULL_TOKEN so the nolock server can proxy authenticated model-pull requests to the model store. Railway credentials are never used by the app or sent to the browser. For gated HuggingFace repos, set HF_TOKEN on llamacpp.
Template Content
nolock
impacte-tech/nolockNOLOCK_WEB_TOKEN
your token
NOLOCK_MODEL_PULL_TOKEN
shared model pull token
llama.cpp
impacte-tech/nolockNOLOCK_MODEL_PULL_TOKEN
secret, same value as defined in the other service
