Skip to content

Installation and self-hosting

To use Relab without any local setup, open app.cml-relab.org.

This page covers running the stack yourself: for evaluation, institutional deployment, offline use, or local development. For contributor workflow and tooling policy, see CONTRIBUTING.md.

  • Docker Desktop
  • just is optional but recommended
  • Contributing code additionally requires uv, Node and pnpm at the versions pinned in .node-version and package.json. See step 2 below and CONTRIBUTING.md
  • PostgreSQL 18 when using an external database; the bundled compose stack ships it.
  1. Clone the repository.

    Terminal window
    git clone https://github.com/CMLPlatform/relab
    cd relab
  2. Install local tooling if you plan to modify code.

    Terminal window
    just setup
  3. Create local backend secrets.

    Terminal window
    just deploy-secrets-template dev

    Create backend/.env.dev only when you need backend-only local overrides such as OAuth, email, or bootstrap settings. Replace values under secrets/dev/ only when you need real local credentials for integrations.

    backend/.env.dev
    GOOGLE_OAUTH_CLIENT_ID=google-oauth-client-id
    GITHUB_OAUTH_CLIENT_ID=github-oauth-client-id
    EMAIL_PROVIDER=smtp
    SMTP_HOST=smtp.example.com
    SMTP_USERNAME=you@example.com
    EMAIL_FROM=Your Name <you@example.com>
    EMAIL_REPLY_TO=you@example.com
    BOOTSTRAP_SUPERUSER_EMAIL=you@example.com
  4. Start the containerized database/cache and run the first migration pass.

    Terminal window
    just dev-db
    just dev-migrate

    To seed sample data during migrations, run SEED_DUMMY_DATA=true just dev-migrate.

    If you also need CPV or HS taxonomy seeding in the migration container:

    Terminal window
    BACKEND_MIGRATIONS_INCLUDE_TAXONOMY_SEED_DEPS=true just dev-migrate
  5. Start the stack.

    Terminal window
    just dev

    If you do not want file watching, use just dev-up instead.

  6. Open the local services.

  7. Verify the backend is healthy.

    Terminal window
    curl http://127.0.0.1:8010/health
  8. Run checks if needed.

    Terminal window
    just ci
    just test

The stack runs on one host behind a Cloudflare Tunnel, so the host needs no public ports. Deploys are three commands on the server: pull the repo, pick a published image tag, start the stack. They can be run there, or sent from another machine over an ssh key whose forced command is scripts/remote_deploy.sh, which allows exactly those steps and nothing else (see deploy/DEPLOY-PROD.md Part 1.6). Every command runs as just stack <prod|staging> <command>, and a state-changing command takes YES to confirm which host it acts on. Deployment and operations describes the topology these steps produce.

  1. Create a Cloudflare Tunnel, one of two ways.

    • By hand: in the Cloudflare dashboard, create a remotely managed tunnel and add a public hostname per service, forwarding to app:8081 and api:8000. The landing page and docs hostnames are Workers Custom Domains instead (step 5).

    • With OpenTofu: infra/cloudflare/ manages the DNS records, the tunnels, and the ingress rules. Export the credentials, then plan and apply per environment:

      Terminal window
      export CLOUDFLARE_API_TOKEN='...'
      export TF_VAR_cloudflare_account_id='...'
      export TF_VAR_cloudflare_zone_id='...'
      export TF_VAR_cloudflare_zone_name='example.org'
      just cloudflare-check
      just cloudflare-plan prod
      just cloudflare-apply prod # plans and saves it; review the diff
      just cloudflare-apply prod YES # applies the plan you just reviewed

      Keep prod and staging state separate. Do not commit Cloudflare tokens, tunnel tokens, or state files.

    Either way, copy the tunnel token for the next step.

  2. Copy .env.example to .env and fill it in.

    Terminal window
    cp .env.example .env

    The root .env is gitignored and holds every host-local value Compose interpolates. One host serves one environment. Every key is described in .env.example. The required ones:

    • ENVIRONMENT: prod or staging. just stack refuses to run against a host whose .env says otherwise.
    • API_PUBLIC_URL, APP_PUBLIC_URL, SITE_PUBLIC_URL, DOCS_PUBLIC_URL: the four public origins on your domain.
    • IMAGE_TAG: the published image tag to run, see step 5. Set IMAGE_REGISTRY too when the images come from your own fork.
    • CLOUDFLARE_TUNNEL_TOKEN: the tunnel token from the previous step.
    • GOOGLE_OAUTH_CLIENT_ID, GITHUB_OAUTH_CLIENT_ID: OAuth client IDs for social login.
    • EMAIL_PROVIDER and the sender fields. With smtp, also fill SMTP_HOST, SMTP_USERNAME, and secrets/<env>/smtp_password. With microsoft_graph, fill the tenant, client, and sender values and secrets/<env>/microsoft_graph_client_secret. Backend startup validates whichever provider you chose.
    • BOOTSTRAP_SUPERUSER_EMAIL: the first admin account. The migrator creates it with the password in secrets/<env>/bootstrap_superuser_password.
    • MALWARE_SCAN_ENABLED: true starts ClamAV with the stack, see step 6.

    Upload quotas: MAX_UPLOAD_FILES_PER_USER and MAX_UPLOAD_BYTES_PER_USER_MB cap contributor accounts; the *_LAB_USER* pair caps lab accounts and must not be lower. The quota counts existing rows, so on a host with existing data raise the limits before the first start; an owner already above the limit cannot upload at all. Every account starts as contributor; a superuser promotes lab members with PUT /v1/admin/users/{user_id}/role and body {"role": "lab"}.

    Per-image caps: MAX_IMAGE_UPLOAD_SIZE_MB and MAX_IMAGE_UPLOAD_PIXELS cap contributor images (10 MB, 12.6 MP); MAX_IMAGE_UPLOAD_SIZE_LAB_MB and MAX_IMAGE_UPLOAD_PIXELS_LAB cap lab images (20 MB, 24 MP). No pixel cap may exceed 50 MP. An image takes about 6 bytes per pixel of memory while it is processed (about 145 MB at 24 MP), so size IMAGE_RESIZE_WORKERS x WEB_CONCURRENCY against the backend’s memory limit before raising a pixel cap.

  3. Create the runtime secret files.

    Terminal window
    just deploy-secrets-template prod

    Replace every placeholder under secrets/prod/. Runtime secrets (database passwords, the auth token secret, the restic password, provider secrets) live only there, never in .env. deploy/env/variables.toml is the inventory; just env-inventory prints it.

  4. Validate the configuration.

    Terminal window
    just compose-config # the Compose overlays render for every environment
    just deploy-secrets-check # every secret file exists, has mode 0644 in a 0700 dir, and is not a placeholder
  5. Publish the images.

    The hosts pull their images from GHCR instead of building them, and the landing page and docs are served by Cloudflare Workers. The backend images work for any deployment, but the app, landing page and docs bake their public URLs in at build time, so a deployment on another domain publishes its own from a fork:

    • In the fork’s settings, create a GitHub Environment named prod (and staging if you run one) with the variables API_PUBLIC_URL, APP_PUBLIC_URL, SITE_PUBLIC_URL and DOCS_PUBLIC_URL, plus the optional FEATURED_PRODUCT_ID for the landing page hero. If you run infra/cloudflare for your edge, it creates the Environment and the four URLs for you, along with the Worker names and your Cloudflare account ID.
    • Add a CLOUDFLARE_API_TOKEN secret to each Environment: an account API token with the Workers Editor role and nothing else. A token granted on all Workers is not scoped to one environment, so give both Environments a required reviewer (infra/cloudflare does, and refuses to apply without one). In prod the reviewer is also the release gate.
    • Run the Deploy Sites workflow for each environment before the infra/cloudflare apply that gives the Workers their hostnames; infra/cloudflare/README.md lists the order.
    • Without Cloudflare, pnpm run build in www/ and docs/ gives a static dist/ for any static host. Its security headers are in dist/_headers, which your host has to apply.
    • Run the Publish Images workflow. A manual run publishes the commit as sha-<short sha>; a release published by release.yml uses its version (0.4.0).
    • Make the packages public in the fork’s package settings, or log the host in to GHCR.
    • Set IMAGE_REGISTRY=ghcr.io/<your-account> in .env, then pull the tag:
    Terminal window
    just stack prod tag YES <tag> # pulls every image, then writes IMAGE_TAG to .env

    To check first that the tag was built by your fork’s workflow, run GITHUB_REPOSITORY=<your-account>/relab IMAGE_REGISTRY=ghcr.io/<your-account> just images-verify prod <tag> from a machine with gh logged in.

  6. Start the stack.

    The migrations profile runs the migrator first and starts the API only after it exits 0, so use it on every start that may carry schema changes, including the first.

    ClamAV upload scanning is on by default (MALWARE_SCAN_ENABLED=true in .env.example), and up starts the scanner whenever that setting is not false. ClamAV needs roughly 3-4 GiB of extra RAM. Its signature database persists in the clamav_db volume; the first boot with an empty volume takes several minutes to download it.

    Terminal window
    just stack prod up YES migrations

    To run without scanning, set MALWARE_SCAN_ENABLED=false. Uploads are then stored unscanned; treat this as a temporary, accepted risk.

    To also seed the CPV or HS taxonomies, run the migrator once by itself:

    Terminal window
    BACKEND_MIGRATIONS_INCLUDE_TAXONOMY_SEED_DEPS=true just stack prod migrate YES
  7. Verify.

    /live on the API is the shallow process check Compose uses; /health also checks PostgreSQL and Redis. Then log in as the bootstrap superuser, enrol two-factor authentication in the account settings (the /admin routes refuse a superuser without it), and try one upload.

    Terminal window
    just stack prod logs
  8. Upgrade later with the same commands: pull a known-good revision, just stack prod tag YES <tag>, then just stack prod up YES migrations. A failed migration leaves the old API serving. If the migrator stops on an unresolvable revision, the database’s alembic_version predates the 2026-09-08 flatten: bring the host to a9c2e4f60b18 on a release from before the flatten, or restore from backup, before continuing. To return to the previous release, just stack prod rollback YES <tag> pulls that release’s images and restarts on them; add the previous alembic revision to downgrade the schema too (revisions before a9c2e4f60b18 were flattened away and cannot be targeted), which the recipe allows only when no migration in between dropped or rewrote data. just stack prod down YES stops the stack.

    Before starting anything, up probes the three mounts the stack writes to, each as the service that writes it: the user_uploads and restic_cache volumes, and the restic bind mount. It refuses to start when one is not writable by UID 65532, because reads and stat still succeed on a wrongly-owned mount and only the writes fail. Docker sets a named volume’s ownership when it first creates the volume and never again, so a host whose volumes were created by a release that ran as a different UID needs a one-time chown, with the stack down:

    Terminal window
    env=prod # or staging
    just stack "$env" down YES
    sudo chown -R 65532:65532 "${BACKUP_HOST_DIR:-./backups}"
    for volume in user_uploads restic_cache; do
    docker run --rm --user 0 -v "relab_${env}_${volume}:/mnt" \
    busybox chown -R 65532:65532 /mnt
    done
    just stack "$env" up YES migrations

The backup container runs as UID 65532. Create the restic directory before the first backup — the service refuses to start if it’s missing (create_host_path: false), rather than let Docker create an empty one indistinguishable from a real mount.

Creating the restic repository is a separate, one-time step. Backup runs never create one — see “Backup repository” in deploy/DEPLOY-PROD.md for why, and do not run backup-init against a host that already has backups.

Terminal window
mkdir -p "${BACKUP_HOST_DIR:-./backups}/restic"
sudo chown -R 65532:65532 "${BACKUP_HOST_DIR:-./backups}"
just backup-init prod # one time only: creates the repository, marks the uploads volume
just backup prod # takes the first snapshot
just restore-check prod # restores that snapshot: DB into a scratch container, uploads into a scratch dir

The repository is encrypted with secrets/prod/restic_password. Never rotate it: every later snapshot depends on it. Before anything destructive, take a tagged backup with just backup prod manual; the manual tag is exempt from retention.

Starting the stack does not schedule backups. A systemd timer on the host runs the one-shot backup service every hour; a second daily timer does repository upkeep (retention, integrity check, offsite copy). Install all four scheduled jobs (backup, backup-maintenance, watchdog, restore-check) once per environment. The installer renders the committed units with this host’s checkout path, deploy user, and just location, then enables the timers:

Terminal window
just timers-render # inspect what will be installed
just timers-install staging # renders, installs, enables (prompts for sudo; not `sudo just`)
systemctl list-timers 'relab-*@staging.timer' # confirm NEXT times are scheduled

The installer also seeds /etc/relab/relab.env with empty PING_* URLs. Fill in PING_WATCHDOG with a healthchecks.io check URL; the hourly watchdog reports every other job’s state through it, and the other three are optional per-job checks. Until you do, job failures are invisible outside the host, and just watchdog <env> keeps saying so. With a dedicated deploy user, run the installer from an account with sudo as RELAB_UNIT_USER=<user> just timers-install <env>.

Without the timers there are no recurring backups; just watchdog <env> reports that.

The offsite path is a second restic repository copied from the local one, over restic’s rclone backend.

  1. Write secrets/<env>/rclone.conf with exactly one WebDAV remote. The remote’s root is the repository; the backup uses it as rclone:<remote>: with no path. just deploy-secrets-check warns while the file still holds the placeholder.

    The daily maintenance job then copies snapshots offsite, initializing the offsite repository on first run.

  2. To copy snapshots on demand, for example right after a one-off local backup:

    Terminal window
    just backup-offsite-copy staging

Prod and staging can ship to a central monitoring stack (Grafana + Loki + Tempo + Prometheus):

  1. Set OTEL_EXPORTER_OTLP_ENDPOINT, OTLP_AUTH_TOKEN and TELEMETRY_EDGE_KEY in the host’s root .env; your monitoring operator supplies them (.env.example describes each). This turns on the backend’s OpenTelemetry exporter and a Grafana Alloy agent that forwards every other container’s stdout over the same OTLP endpoint.

  2. just stack <env> up includes compose.telemetry.yml when the endpoint is non-empty. Hosts without it ship nothing; docker logs stays the only log path.

See Deployment and operations for what each variable does and how Alloy is sandboxed.

For camera-assisted capture, install the Raspberry Pi Camera Plugin.

The RPi connects outbound to the backend over a WebSocket relay, so it needs no public IP or port forwarding. To pair automatically, set PAIRING_BACKEND_URL on the RPi, boot it, and enter the displayed pairing code in the app. On a headless Pi, read the code from its local /setup page or from the PAIRING READY log line over SSH, docker compose logs, or journalctl. See the plugin install guide, the platform camera guide, and the API reference.