| 1 | # Runbook |
| 2 | |
| 3 | > Attempt #2's only setup instructions lived in a plan file describing an architecture |
| 4 | > that had already been deleted, so they were actively wrong. Keep this file honest: |
| 5 | > if a command here doesn't work, fix it or delete it. |
| 6 | |
| 7 | ## Status |
| 8 | |
| 9 | The app boots and serves. Everything marked **(#2)** is carried from the previous |
| 10 | attempt and has **not** been re-verified against Topcoat — treat it as a sketch. |
| 11 | |
| 12 | ## Requirements |
| 13 | |
| 14 | - **rustc ≥ 1.95** — Topcoat 0.5 requires it, and on an older toolchain `cargo add |
| 15 | topcoat` silently resolves to an empty `topcoat v0.0.0` placeholder rather than |
| 16 | failing. Verified on 1.97.1. |
| 17 | - `git` on `PATH` (from Milestone 3 — `git init --bare` creates repos, and from |
| 18 | Milestone 4 `git http-backend` serves the protocol). **`cargo test` needs it too**: |
| 19 | `DiskGitStorage`'s tests run real `git init`, so a machine without `git` fails the |
| 20 | suite, not just the app. |
| 21 | |
| 22 | ## Dev setup |
| 23 | |
| 24 | ```bash |
| 25 | cargo run # serves on http://127.0.0.1:3000 |
| 26 | cargo test |
| 27 | ``` |
| 28 | |
| 29 | ```bash |
| 30 | cargo install topcoat-cli # dev server: watch, rebuild, asset bundling |
| 31 | topcoat dev # working |
| 32 | ``` |
| 33 | |
| 34 | `topcoat dev` builds, bundles assets, watches sources, and live-reloads pages that |
| 35 | include `topcoat::dev::script()`. Press `r` to force a rebuild. |
| 36 | |
| 37 | ## Configuration |
| 38 | |
| 39 | `STEID_`-prefixed env vars via `dotenvy` + `envy`, read into `AppConfig`. Every value |
| 40 | has a default, so a bare `cargo run` works with no environment at all. |
| 41 | |
| 42 | Live now: |
| 43 | |
| 44 | ``` |
| 45 | STEID_DATABASE_URL=sqlite:steid.db?mode=rwc # default |
| 46 | STEID_DATA_DIR=./data # default; bare repos, read from M3 |
| 47 | STEID_INSECURE_COOKIES=false # default; see below |
| 48 | ``` |
| 49 | |
| 50 | There is deliberately **no owner password in configuration** — the owner is created |
| 51 | through the claim flow instead. See |
| 52 | [0002](decisions/0002-first-run-claim-not-config-bootstrap.md). |
| 53 | |
| 54 | The bind address is **not** a `STEID_` variable — Topcoat owns it: |
| 55 | |
| 56 | ```bash |
| 57 | HOST=0.0.0.0 PORT=8080 cargo run |
| 58 | ``` |
| 59 | |
| 60 | That supersedes attempt #2's `STEID_LISTEN_ADDR`. |
| 61 | |
| 62 | ### `STEID_INSECURE_COOKIES` — development only |
| 63 | |
| 64 | Topcoat's session cookie is `__Host-` prefixed and `Secure`. `Secure` means the browser |
| 65 | only keeps it over a trustworthy origin, and browsers disagree about whether |
| 66 | plain-HTTP `localhost` qualifies. Where it doesn't, **the failure is completely |
| 67 | silent**: the server issues a session and records the row, the browser discards the |
| 68 | cookie, and every page renders signed out with no error anywhere. This cost an |
| 69 | afternoon; the symptom looks exactly like broken auth logic. |
| 70 | |
| 71 | Setting `STEID_INSECURE_COOKIES=true` swaps in `InsecureCookieTokenStore` — the same |
| 72 | cookie without `Secure` and without the prefix, named `steid-dev-session` so it can |
| 73 | never be confused with a hardened one. `HttpOnly` and `SameSite=Lax` are kept. Boot |
| 74 | prints a warning while it's on. |
| 75 | |
| 76 | **Never set this on a deployed instance.** Without `Secure` the session cookie travels |
| 77 | unencrypted and anyone on the network path can lift it and become that user. Behind |
| 78 | TLS, leave it unset. |
| 79 | |
| 80 | A gitignored `.env` in the repo root sets it for local work. Keep `.env` and |
| 81 | `.env.prod` out of git — both are gitignored. |
| 82 | |
| 83 | ## First run |
| 84 | |
| 85 | ```bash |
| 86 | cargo run # or: topcoat dev |
| 87 | ``` |
| 88 | |
| 89 | An unclaimed instance prints a setup token and redirects every route to `/auth/setup`. |
| 90 | Paste the token, choose a handle, email, and password, and the owner is created and |
| 91 | signed in. |
| 92 | |
| 93 | The token is **held in memory only**, so every restart mints a new one — including |
| 94 | each rebuild under `topcoat dev`. Use the most recent one printed. Once claimed, no |
| 95 | token is minted at all and `/auth/setup` redirects away. |
| 96 | |
| 97 | Sign in at `/auth/login` with the **email**, not the handle. |
| 98 | |
| 99 | ## Repo layout on disk (Milestone 3) |
| 100 | |
| 101 | Bare repos at `{STEID_DATA_DIR}/{handle}/{name}.git`, created with no template (so no |
| 102 | `.sample` hooks) and `HEAD` pinned to `refs/heads/main` regardless of the host's |
| 103 | `init.defaultBranch`. See [0006](decisions/0006-git-binary-behind-narrow-ports.md). |
| 104 | |
| 105 | Created empty — no initial commit and no branch, like GitHub. |
| 106 | |
| 107 | Creating one refuses rather than reusing a directory that already exists, so an orphan |
| 108 | left by a create that died mid-way blocks that name until it is removed by hand. |
| 109 | |
| 110 | ## Git transport (Milestone 4) |
| 111 | |
| 112 | Smart HTTP, delegated to `git http-backend`, authenticated with personal access tokens |
| 113 | over HTTP Basic — see [0001](decisions/0001-git-over-http-not-ssh.md). No SSH, no host |
| 114 | keys, no `authorized_keys`. |
| 115 | |
| 116 | ```bash |
| 117 | git clone http://host/{handle}/repos/{name}.git |
| 118 | ``` |
| 119 | |
| 120 | Fill in the token workflow and the `body_limit` setting once this is built. |
| 121 | |
| 122 | ## Deployment (container) |
| 123 | |
| 124 | A `Dockerfile` at the repo root builds a self-contained image. Two stages: `rust:1.97-bookworm` |
| 125 | compiles and bundles, `debian:bookworm-slim` runs. ~206 MB. |
| 126 | |
| 127 | ```bash |
| 128 | docker build -t steid . |
| 129 | docker volume create steid-data |
| 130 | docker run -d --name steid -p 3000:3000 -v steid-data:/data steid |
| 131 | docker logs steid # the setup token is here, and only here |
| 132 | ``` |
| 133 | |
| 134 | Verified end to end on 2026-08-29: the image builds, boots, applies migrations, |
| 135 | creates `/data/steid.db`, serves `/auth/setup` with its stylesheet, and `git init |
| 136 | --bare` succeeds inside `/data/repos` as the non-root user. |
| 137 | |
| 138 | ### The build needs `topcoat asset bundle`, not just `cargo build` |
| 139 | |
| 140 | `cargo build --release` alone produces a binary that **fails to boot**. `main` calls |
| 141 | `AssetBundle::load()`, which walks up from the executable looking for |
| 142 | `assets/manifest.toml`; without one it returns `NotFound` and the process exits before |
| 143 | serving anything. `build.rs` does not write that bundle — it only runs Tailwind and |
| 144 | stages icons into `OUT_DIR`, where they are embedded in the binary. |
| 145 | |
| 146 | The bundle comes from the CLI, which the builder stage installs: |
| 147 | |
| 148 | ```bash |
| 149 | cargo install topcoat-cli --version 0.5.0 --locked |
| 150 | topcoat asset bundle --release # runs `cargo build --release` itself, then bundles |
| 151 | ``` |
| 152 | |
| 153 | It writes `target/assets/`, which must be copied **next to the binary** in the runtime |
| 154 | image — `/app/steid` finds `/app/assets`. This is the same step `topcoat dev` performs |
| 155 | for you, and the reason a hand-built binary serves stale CSS. |
| 156 | |
| 157 | Two build-time consequences worth knowing: |
| 158 | |
| 159 | - The build **needs network access**: `build.rs` downloads the standalone Tailwind CLI |
| 160 | from GitHub releases, and the bundler downloads any remote asset. |
| 161 | - The image is **Debian, not Alpine, on both sides**. Those Tailwind binaries are |
| 162 | glibc-linked, so a musl builder fails during `cargo build`. |
| 163 | |
| 164 | The builder mounts the cargo registry and `target/` as BuildKit caches. There is |
| 165 | deliberately **no dummy-`main.rs` dependency-caching trick**: `build.rs` scans the real |
| 166 | sources for Tailwind classes, and a faked source tree yields a stale stylesheet — a |
| 167 | wrong answer that still builds, which is the worst kind. |
| 168 | |
| 169 | ### Configuration in a container |
| 170 | |
| 171 | The image sets these defaults, so the `docker run` above needs no `-e` flags at all: |
| 172 | |
| 173 | ``` |
| 174 | STEID_DATABASE_URL=sqlite:/data/steid.db?mode=rwc |
| 175 | STEID_DATA_DIR=/data/repos |
| 176 | HOST=0.0.0.0 # Topcoat's, not STEID_-prefixed — see Configuration above |
| 177 | PORT=3000 |
| 178 | ``` |
| 179 | |
| 180 | No public URL is configured anywhere: the origin is derived from the `Host` header and |
| 181 | `X-Forwarded-Proto`, so a proxy that forwards both needs nothing further. |
| 182 | |
| 183 | **`STEID_INSECURE_COOKIES` is deliberately unset in the image and must stay unset.** |
| 184 | The session cookie is `Secure`, which means the deployment needs TLS — assume a |
| 185 | terminating proxy in front (Caddy, nginx, a platform router). Setting the variable to |
| 186 | paper over a missing certificate hands every session cookie to anyone on the network |
| 187 | path. `.dockerignore` excludes `.env` for the same reason: the dev `.env` sets it, and |
| 188 | copying it in would silently unharden a deployed image. |
| 189 | |
| 190 | ### State is one volume |
| 191 | |
| 192 | Everything that must survive a restart lives under `/data`: the SQLite database as a |
| 193 | file directly in it, the bare repositories under `/data/repos`. `VOLUME ["/data"]` is |
| 194 | declared, so a container started without `-v` still keeps its state — in an anonymous |
| 195 | volume that is easy to lose track of. Name it. |
| 196 | |
| 197 | `/data` itself must be writable, not just the database file: SQLite creates `-wal` and |
| 198 | `-shm` siblings next to it. |
| 199 | |
| 200 | The container runs as uid **10001** (`steid`). A named volume inherits that ownership |
| 201 | from the image on first use. A **bind mount does not** — `-v /srv/steid:/data` starts |
| 202 | root-owned and the app fails to write, so `chown 10001:10001 /srv/steid` on the host |
| 203 | first. |
| 204 | |
| 205 | ### Claiming a deployed instance |
| 206 | |
| 207 | The setup token is printed to **stdout only, and only while the instance is |
| 208 | unclaimed**. It is held in memory, so every restart — including every redeploy — |
| 209 | mints a new one, and a claimed instance mints none at all. |
| 210 | |
| 211 | There is no way to recover it other than the platform's logs: |
| 212 | |
| 213 | ```bash |
| 214 | docker logs steid | tail -20 |
| 215 | ``` |
| 216 | |
| 217 | Read the token from the **most recent** boot, then claim at `https://your-host/auth/setup`. |
| 218 | Until it is claimed every route redirects there, so an instance left unclaimed on a |
| 219 | public address is an open door — claim it immediately after the first deploy. |
| 220 | |
| 221 | ## Manual verification checklist (#2) |
| 222 | |
| 223 | Attempt #2 verified these by hand each milestone but never wrote down the steps. They |
| 224 | are the smoke test for Milestones 4–5. The auth rows assumed SSH keys; the shape of |
| 225 | the check still holds with tokens substituted: |
| 226 | |
| 227 | - [x] Create a repo via the web UI → bare repo appears at |
| 228 | `{data_dir}/{handle}/{name}.git` *(Milestone 3; also checked that the name |
| 229 | normalises, that a duplicate re-renders the form, and that a private repo 404s |
| 230 | for a signed-out visitor)* |
| 231 | - [ ] `git clone` an empty repo → succeeds |
| 232 | - [ ] `git clone` a repo with history → succeeds |
| 233 | - [ ] `git clone` a non-existent repo → clean error, not a hang or panic |
| 234 | - [ ] First push to an empty repo → succeeds |
| 235 | - [ ] Push to a repo with history → succeeds |
| 236 | - [ ] Push a repo large enough to exercise the `body_limit` cap → succeeds |
| 237 | - [ ] Clone with no credentials → rejected, and the prompt is comprehensible |
| 238 | - [ ] Clone with a valid token → succeeds |
| 239 | - [ ] Revoke the token → subsequent clone rejected at auth |
| 240 | - [ ] Clone a private repo as a non-member → rejected |
| 241 | - [ ] Push as a non-owner member → rejected |
| 242 | |
| 243 | Worth automating as an integration test rather than re-running by hand a fourth time. |
| 244 | |
| 245 | ## `.gitignore` |
| 246 | |
| 247 | Applied: `/target`, `/data`, `*.db*`, `.env`, `.env.prod`. |
| 248 | |
| 249 | **Do not add `/plans`.** Attempt #2 did, and that is why these docs had to be |
| 250 | hand-carried between repos. |
| 251 | |
| 252 | ## Deploying this instance |
| 253 | |
| 254 | The operator's path, as opposed to `README.md`, which is written for a stranger |
| 255 | installing their own. `jpgill.dev` serves two roles from one box: this instance, |
| 256 | and the place everyone else downloads Steid from. |
| 257 | |
| 258 | ### Order matters |
| 259 | |
| 260 | 1. **Provision — AWS Lightsail**, **Debian 13** blueprint (12 also works). Lightsail rather than EC2 |
| 261 | deliberately: it *is* AWS's VPS product, where EC2 makes you assemble the same box out |
| 262 | of a VPC, security groups, an EBS volume and an Elastic IP, with per-GB egress on top. |
| 263 | Either architecture works — releases are built for `x86_64` and `aarch64`, and |
| 264 | `install.sh` picks between them from `uname -m`, so a Graviton instance is fine. |
| 265 | |
| 266 | **Not Amazon Linux**, which Lightsail pre-selects: `install.sh` is Debian-family, and |
| 267 | it stops with a clear message on anything without `apt-get` rather than half-installing. |
| 268 | Ubuntu 22.04/24.04 work too. Debian 13 and 12 are both verified end to end in a |
| 269 | container — including that Caddy installs, since its apt repository URL is |
| 270 | codename-independent and that was the only plausible difference between them. |
| 271 | |
| 272 | Three Lightsail-specific things, each of which breaks the deployment silently if |
| 273 | missed: |
| 274 | |
| 275 | - **Attach a static IP.** A Lightsail instance takes a *new* public IP on stop/start. |
| 276 | Without one, a reboot changes the address, DNS points at nothing, and Caddy's |
| 277 | certificate renewals start failing with no obvious cause. Free while attached to a |
| 278 | running instance. |
| 279 | - **Open 80 and 443** in the instance's Networking tab. Only 22 is open by default, |
| 280 | and **Let's Encrypt validates over port 80** — a firewall that allows 443 alone |
| 281 | fails at certificate issuance, not at first request. |
| 282 | - **Snapshots are not backups.** They are whole-disk and live in the same AWS account. |
| 283 | The two state paths still want copying somewhere else. |
| 284 | 2. **DNS first, then install.** `dig +short jpgill.dev` must return the box's IP |
| 285 | *before* `install.sh` runs. Caddy requests a certificate on startup; if DNS has not |
| 286 | propagated it fails and backs off, and the resulting error points nowhere useful. |
| 287 | 3. **First install uses `--tarball`.** There is a bootstrap: this instance is what will |
| 288 | serve the releases, so at that moment there is nowhere to download from. |
| 289 | |
| 290 | ```sh |
| 291 | scp dist/steid-0.1.0-x86_64-unknown-linux-gnu.tar.gz install.sh root@<ip>:/root/ |
| 292 | ssh root@<ip> './install.sh --domain jpgill.dev \ |
| 293 | --tarball ./steid-0.1.0-x86_64-unknown-linux-gnu.tar.gz' |
| 294 | ``` |
| 295 | |
| 296 | Then `curl https://jpgill.dev/healthz` — over https, with a real certificate. |
| 297 | |
| 298 | ### Becoming the distribution host |
| 299 | |
| 300 | Only this instance does this. Uncomment the two `handle` blocks in |
| 301 | [deploy/Caddyfile](../deploy/Caddyfile) and rsync the artefacts into a **versioned** |
| 302 | directory, matching the URL `install.sh` builds |
| 303 | (`${RELEASE_BASE_URL}/v${VERSION}/…`): |
| 304 | |
| 305 | ```sh |
| 306 | rsync dist/*.tar.gz dist/*.sha256 root@<ip>:/var/lib/steid/dist/v0.1.0/ |
| 307 | rsync install.sh root@<ip>:/var/lib/steid/dist/ |
| 308 | ``` |
| 309 | |
| 310 | **Those Caddy paths shadow Steid.** Nothing lives at |
| 311 | `/{handle}/repos/{name}/releases` today, so nothing breaks — but when Steid grows a real |
| 312 | release feature at that URL, Caddy will keep winning and the feature will look broken. |
| 313 | Delete the blocks then. The URL is deliberately the one that feature will use, so links |
| 314 | published now survive it. |
| 315 | |
| 316 | `/install.sh` at the root is safe permanently rather than by luck: `OrgName` allows only |
| 317 | `[a-z0-9-]`, so no handle can contain a dot and none can ever collide with it. A |
| 318 | root-level `/releases` would **not** be safe — it is a valid handle shape. |
| 319 | |
| 320 | ### What is verified, and what is not |
| 321 | |
| 322 | Verified in a Debian 12 container with `systemctl` stubbed: prerequisites install, the |
| 323 | tarball extracts, `/opt/steid` and `/var/lib/steid` get the right owners and modes, the |
| 324 | `steid` user is created with `nologin`, the env file and unit and Caddyfile are written, |
| 325 | and **the binary starts as the `steid` user**. The artifact itself was booted on Debian |
| 326 | 11 (glibc 2.31) and served pages. |
| 327 | |
| 328 | **Not verified anywhere but a real box:** the systemd unit lifecycle, and Caddy's ACME |
| 329 | certificate issuance. Expect the first real run to need a fix or two. |
| 330 | |
| 331 | ### One bug already found this way |
| 332 | |
| 333 | `install.sh` defaulted to a `musl` target while `release.sh` had moved to `gnu`. The |
| 334 | symptom was `checksum mismatch` — because the installer was looking for a tarball that |
| 335 | was never built. The two defaults must agree; both now say so in a comment. |