Security and hardening
Public traffic enters through Cloudflare: Tunnel for the API and app, Workers static assets for the landing page and docs. Data services stay on private Compose networks, outbound backend HTTP is allowlisted, and source-controlled checks cover sensitive application behavior.
Baseline
Section titled “Baseline”OWASP ASVS is the application-security baseline. Review focus:
- authentication, OAuth, and token flows
- public read APIs and authenticated mutation APIs
- uploads, media, malware scanning, and file storage
- admin APIs
- RPi camera device APIs and WebSocket relay
- backups, secrets, logs, telemetry, and release/security artifacts
Assets: accounts, profile/privacy settings, research records, uploaded media/files, OAuth and YouTube tokens, RPi camera credentials, refresh-token state, database dumps, backup material, and runtime secrets.
Data protection
Section titled “Data protection”Data is classified by the harm its disclosure or misuse would cause:
- Public research data: product records, public catalog/reference data, public dataset metadata, and intentionally public uploaded media. These may be cached when responses are public and content-addressed assets must 404 when missing rather than falling back to the app shell.
- Published dataset releases: curated, versioned snapshots deposited under an open licence with a DOI. A release is immutable and mirrored by third parties once published, so account and upload deletion remove records from the live platform without retracting anything already released. Everything entering a release is checked before deposit: pseudonymised owner identity, named attribution only for contributors who opted in, and no account, auth, or contact data.
- User and account data: email addresses, usernames, profile preferences, contribution statistics, owner attribution, and account state. These are access-controlled by backend dependencies and privacy policy helpers. Viewer-dependent API responses use
no-store, and product owner identity is redacted when profile visibility does not allow the current viewer to see it. - Secrets and credentials: passwords, auth tokens, refresh-token state, OAuth/YouTube tokens, RPi camera credentials, local camera API keys, runtime secrets, encryption keys, database passwords, and backup passwords. These must not appear in URLs, logs, browser local storage, docs examples, or analytics. Browser sessions use secure cookies; native bearer tokens use platform secure storage; local camera API keys use secure storage on native and memory-only storage on web.
- Uploads and metadata: uploaded images, generic research files, derived thumbnails, image metadata, video metadata, and RPi capture metadata. Uploads are validated, malware-scanned when enabled, and size-limited before storage. Before an image reaches any storage backend (local disk or S3), it is re-oriented and its EXIF reduced to an allowlist of capture parameters (camera and lens model, exposure, focal length, ISO, white balance); GPS, vendor MakerNote blocks, serial numbers, and everything else identifying are dropped, as are XMP packets, JPEG comments, and PNG text chunks. Research files are stored as submitted, since the file itself is the research artifact.
- Operational data: request IDs, logs, telemetry, database dumps, and backup archives. Logs and traces must avoid tokens, passwords, OAuth material, private URLs, and sensitive personal values. Backups and dumps inherit the highest classification of their contents and are encrypted, access-controlled, and retained only for operational recovery.
Protection requirements:
- Send sensitive values in request bodies, headers, secure cookies, or URL fragments that are scrubbed by the client; never place them in server-visible query strings.
- Use
Cache-Control: no-storefor authenticated, viewer-dependent, error, logout, password-reset, and verification flows. Public immutable assets may use long-lived caching only when missing files produce a real 404. - Keep access control server-side. Public docs schemas and filtered UI states are inventory tools, not authorization controls.
- Store secret material in Docker secrets, secure platform stores, hashed or token-fingerprint state, or encrypted backups.
- Delete obsolete runtime state through session revocation, upload deletion, cache invalidation, cleanup, and backup rotation. Add retention automation when a new data class needs a schedule beyond those.
Account privileges
Section titled “Account privileges”Three independent attributes decide what an account may do; each is enforced separately:
is_verified: whether the account may create records at all.is_superuser: access to the/adminroutes, and moderating any contributor’s products: a superuser can correct the details of and delete any contributor’s products and media, but cannot add content to them (no uploads, components or materials). Every such action is audited. Every superuser power, the/adminroutes included, needs TOTP MFA enrolled on the account; without it a superuser is refused on/adminand treated as an ordinary user elsewhere. Superusers always hold thelabtier.role: the contributor tier,contributor(default) orlab. It gates non-image research-file upload and selects the upload quota tier.
Only a superuser sets roles, through PUT /v1/admin/users/{user_id}/role; every change is audit-logged. The self-service update schema has no role field, so PATCH /users/me cannot reach it. New accounts, and every account backfilled when roles were introduced, start at contributor.
Clients hide affordances the role cannot use (the app renders its research-file block for lab accounts only), but the route dependency is the control: it refuses the request whatever the client displayed.
Logging and error handling
Section titled “Logging and error handling”The backend writes application, access, audit, and security-control events through Python logging. Development logs are human-readable; production and staging logs are JSON on stderr. Each request gets an X-Request-ID, echoed to the client and attached to log records with method, path, status, and latency.
Structured audit events cover login success or failure, MFA challenge success or failure, logout, all-session revocation, authorization denial, and rate limiting. Fields are sanitized before logging. Logs may include stable IDs, masked account identifiers, request metadata, action names, resource labels, outcomes, and non-secret reasons; never passwords, raw tokens, OAuth material, private URLs, or raw sensitive payload bodies.
Container logs are the default transport. Deploy hosts may enable OpenTelemetry export to a separate collector, which also starts an Alloy agent forwarding the other containers’ stdout over the same authenticated OTLP endpoint. Sanitization applies to the backend’s own logging only. Other containers’ stdout ships as-is, and PostgreSQL error logs can quote statement text with literal parameter values, so the collector must sit inside the same data-protection boundary as the database: a department-operated stack, not a third-party service. Log access, retention, backup, and deletion are operator responsibilities.
The catch-all exception handler logs internal details and stack traces for operators; clients receive a generic Problem Details response with the request ID.
Inbound edge
Section titled “Inbound edge”Production and staging are HTTPS-only behind Cloudflare Tunnel. infra/cloudflare/ holds the per-environment DNS records, public hostnames, tunnel ingress rules, and the Workers custom domains of the landing page and docs; infra/cloudflare-zone/ holds the zone-wide TLS minimums, HTTPS redirects, and rulesets. Edge rate limiting covers the authentication endpoints only. Upload and RPi camera limits live in the application’s Redis-backed limiters, which fail open when Redis is unreachable; the per-user upload quota is then the remaining bound.
The application’s read, write, and upload limits key anonymous requests per client IP (300, 120, and 30 a minute). A request carrying a live access token is keyed per user instead, at 1500, 600, and 300 a minute: a workshop of about 50 people often shares one demo account behind one IP, so a single user bucket has to carry the room. The token only selects the bucket; an unknown or expired token falls back to the IP bucket. The limits are the API_*_RATE_LIMIT and API_*_RATE_LIMIT_PER_USER settings. The edge rule still counts every /v1/auth/ request per IP (10 per 10 seconds), so a room signing in within the same few seconds sees brief 429s from the edge and succeeds on retry.
Unknown tunnel hostnames fall through to a final http_status:404 rule. Compose still owns runtime services; OpenTofu owns only the public edge.
Network boundaries
Section titled “Network boundaries”Compose networks segment the deploy stack:
- public-facing services use the
edgenetwork - PostgreSQL, Redis, migrations, and backups use the internal
datanetwork - ClamAV scanning uses the private
scanningnetwork - only the API bridges the networks it needs
PostgreSQL and Redis publish no host ports in prod or staging. The defaults are DATABASE_TLS=false and REDIS_TLS=false. If either service moves to an external host, managed service, or other untrusted network, set DATABASE_TLS=true or REDIS_TLS=true, and set the matching CA file variable for a private CA.
Backend egress
Section titled “Backend egress”Backend outbound HTTP must use the shared HTTP client unless a reviewed exception exists. The shared client ignores ambient proxy and CA environment variables, does not follow redirects, uses bounded timeouts, validates TLS certificates, and permits only OUTBOUND_HTTP_ALLOWED_URLS.
Prefer exact HTTPS URLs. Use trailing-slash prefixes only for integrations with dynamic path segments. Narrow GitHub, raw GitHub, and provider API access to the specific URLs Relab uses.
Denial-of-service controls
Section titled “Denial-of-service controls”Cloudflare handles volumetric attacks. Source-controlled edge rules rate-limit /v1/auth/* per client IP, challenge authentication calls from high-risk countries, and block admin calls from them. Uploads, RPi camera traffic, and WebSocket connections are bounded by the application, as below.
The expensive paths are media uploads, image processing and thumbnails, malware scanning, file cleanup, backups, OpenAPI generation, YouTube API calls, and RPi camera relay and HLS traffic. The backend bounds them with JSON streaming limits, multipart upload caps, file format allowlists, user upload quotas, Redis-backed route rate limits, image worker limits, ClamAV scan timeouts, bounded outbound HTTP timeouts and retries, WebSocket authentication and frame limits, and relay commands that fast-fail for cameras without a live heartbeat.
Keep long-running maintenance work in just tasks, scripts, migrations, or background services, not request handlers. A new API path that can exceed normal client timeouts needs asynchronous processing or explicit limits before exposure.
For one host only, create a gitignored compose.host.yaml; deploy recipes include it last:
services: api: environment: WEB_CONCURRENCY: "1"Browser runtime policy
Section titled “Browser runtime policy”Browser code is bundled and served first-party. Biome and CI checks block remote script tags, remote module imports, CDN runtime assets, analytics snippets, and tag-manager snippets.
The one exception is product video embedding from https://www.youtube-nocookie.com, constrained by CSP and iframe/WebView controls. Any new third-party runtime script must be reviewed, documented, CSP-allowlisted, and tested.
The Expo web app requires inline and eval-style first-party script support in the enforced CSP. A stricter report-only policy is active and documented as an ASVS V3 L3 exception. Docs and public web origins enforce stricter script policies; all browser origins enforce object-src, base-uri, and referrer suppression. The API forbids framing (frame-ancestors 'none'). The app, docs, and public website allow embedding from any HTTPS page: the docs and website have no sessions, and the app’s SameSite=Lax session cookies are never sent to a cross-site frame, so an embedded app always runs signed out and its sign-in screens point to a new tab instead.
Supply chain
Section titled “Supply chain”GitHub Dependency Review and Dependency Graph gate dependency changes; Renovate opens update PRs. just audit runs a full-tree vulnerability sweep. When an upstream package pins a vulnerable transitive version, a scoped overrides: entry in pnpm-workspace.yaml lifts it; just overrides-check runs weekly and flags any override the tree no longer needs.
The security workflow builds deployable Compose images, scans them with Trivy, and stores SPDX JSON SBOM artifacts for 90 days. Release automation builds the deployable images, scans them with Trivy before pushing them to GHCR, attests each image’s SPDX SBOM against its digest, and uploads the SBOM files and their attestation bundles with the GitHub release. The deploy hosts pull those images and never build their own.
Dependency updates use risk-based timeframes:
- critical exploitable vulnerabilities: mitigate or patch within 48 hours
- high vulnerabilities: patch within 7 days
- medium vulnerabilities: patch within 30 days
- low vulnerabilities and routine library updates: handle in the normal weekly Renovate cycle
If no upstream fix exists, record the accepted risk, compensating control, or temporary pin in the PR, issue, or nearest maintainer documentation. Do not keep production dependencies past these windows without an explicit maintainer decision.
Trusted dependency sources are the source-controlled package manifests, lockfiles, Docker image references, GitHub Actions pins, and Renovate configuration. The SBOMs from the security and release workflows are the deployable dependency inventory. The deployed images are the published ones, so gh attestation verify oci://ghcr.io/cmlplatform/<image>:<tag> --owner CMLPlatform checks the SBOM bound to what a host runs.
Risky components parse untrusted content, execute native code, perform cryptography, manage authentication, or open network or file boundaries: image processing, archive and office file inspection, OAuth and provider clients, S3-compatible storage, Redis and PostgreSQL drivers, ClamAV, and browser media playback helpers. Changes to them need maintainer review even when automated checks pass.
Dangerous functionality
Section titled “Dangerous functionality”Dangerous functionality is isolated in small modules and reviewed as a security-sensitive bucket:
- upload byte handling, public file serving, image processing, archive and office inspection, and malware scanning live in the backend file-storage and image modules
- outbound HTTP uses the shared client (see Backend egress)
- RPi camera direct and relayed traffic is authenticated, rate-limited, and separated from database and storage services by Compose networks
- backup jobs read database and upload volumes but run in dedicated containers with separate secrets
Encapsulate new dangerous functionality near its trust boundary, cover it with focused tests, and document it here when it changes the security model.