Files

219 lines
9.6 KiB
Markdown

# Tomter Vel — infrastructure
Everything the vel's own site needs on one host, as containers: Caddy
for TLS, NATS with JetStream for the records, Kanidm for who is who,
and portal, the site itself. The content (pages, forms, desks) lives
in [tomtervel/questions](https://prosjekt.klingenbergbygg.no/tomtervel/questions)
and is fetched from there; this repo owns the host.
```
Internet ──► Caddy (Let's Encrypt)
├── PORTAL_HOST ──► portal:3000 ──► NATS (records) + Kanidm (login)
└── ID_HOST ─────► kanidm:8443 (internal TLS)
```
Today the vel's draft site runs on Klingenberg Bygg's host as
vel.klingenbergbygg.no, sharing that host's Kanidm and NATS. This repo
is the same site on a host of the vel's own, with a Kanidm of its own,
so nothing about the vel's members or records depends on anyone else.
## First start
On a fresh Linux host with podman + podman-compose and openssl (run rootful,
i.e. as root, so Caddy can bind 80/443 and Kanidm sees a stable source IP):
```sh
git clone https://prosjekt.klingenbergbygg.no/tomtervel/infrastructure /srv/tomtervel/infrastructure
cd /srv/tomtervel/infrastructure
cp .env.example .env # set PORTAL_HOST and ID_HOST; DNS must point here
sudo sh bootstrap.sh # renders configs, makes the internal cert, starts, recovers Kanidm admin
sh kanidm-setup.sh # logs in, creates the portal client and desk groups, restarts portal
```
`bootstrap.sh` prints the `admin` and `idm_admin` passwords once;
write them down. After `kanidm-setup.sh`, https://PORTAL_HOST serves the
site and https://ID_HOST is the login.
## Running on kasse (behind the host Caddy)
The vel keeps its **own** Kanidm and NATS here so it can later lift onto a host
of its own unchanged — on kasse it just runs in isolation, fronted by kasse's
existing host Caddy (which owns 80/443). So the bundled Caddy stays off (it's
behind `profiles: [edge]`); run the default `podman compose up -d --build`.
The services publish loopback-only ports for the host Caddy to reach:
- portal → `127.0.0.1:3050`
- kanidm → `127.0.0.1:8443` (internal self-signed TLS)
- nats → `127.0.0.1:4223` (kasse's shared platform NATS owns 4222)
Add two host-Caddy site blocks in `/etc/caddy/conf.d/`, using the vel's `.env`
hosts (`vel.klingenbergbygg.no → :3050` already exists):
```
<PORTAL_HOST> {
reverse_proxy localhost:3050
}
<ID_HOST> {
reverse_proxy https://localhost:8443 {
transport http { tls_insecure_skip_verify }
}
}
```
Then `sudo systemctl reload caddy`, point the content repo's `lint-and-reload`
reload step at `nats://127.0.0.1:4223`, and retire the old systemd portal:
`sudo systemctl disable --now app@tomtervel-portal`.
## Accounts, all the way
`kanidm-setup.sh` sets up everything portal needs from this Kanidm,
and is safe to rerun:
- **Desk groups, named as the content names them.** Read from every
`qualifies:` under `questions/` in the content repo, so the list
cannot drift from the pages. This Kanidm is the vel's own, so there is
no prefix: the board's group is `styret`. Groups from before are
renamed in place (`tomtervel_styret` → `styret`), members and all.
- **Sign-in.** The OAuth2 client's scope map is `idm_all_persons`:
everyone here is the vel's, and what they see is decided by their
desk groups in the `groups` claim.
- **Onboarding.** The `portal-onboarding` service account manages
every desk group, so an invite from a desk and a `grants:` can add
people to them. Its token goes into `portal.env`.
- **The first person in.** `SEED_ADMIN_EMAIL` in `.env`: at start,
portal invites them into the most privileged group (the board) and
mails the link, once. Everyone else they invite from their desk.
- **The people who ask the questions.** Everyone a page names as
`responsible` gets an account at start, a mail saying they are listed
as responsible for asking that question, and a place in
`RESPONSIBLE_GROUP` (`redaktor`). Signed in, they can suggest changes
to their own pages.
- **The project Gitea.** With `GITEA_URL` set, `redaktor` may sign in
to it with the same account: its own OAuth2 client, only for that
group. Run on the Gitea host, the script adds the sign-in source
itself (and requires the `redaktor` claim there too); elsewhere it
leaves the secret in `gitea-oauth.secret` and prints the one command.
## People and desks
Each group in the content (`qualifies:` under `questions/`) is a
Kanidm group `tomtervel_<group>`, all members of `tomtervel_members`.
Someone in `tomtervel_styret` logs in and sees the board's desk.
```sh
kanidm -D idm_admin -H https://ID_HOST person create kari "Kari Lien"
kanidm -D idm_admin -H https://ID_HOST person update kari --mail kari@example.no
kanidm -D idm_admin -H https://ID_HOST person credential create-reset-token kari
kanidm -D idm_admin -H https://ID_HOST group add-members tomtervel_styret kari
```
Membership is read at login; someone added while logged in logs out
and in again.
## Content changes
A push to tomtervel/questions is linted on push and tells the portal
to reload over NATS. That reload reaches the host it runs on: today
Klingenberg Bygg's. On this host the runner is a **host service**
(`pacman -S gitea-runner`, registered against prosjekt.klingenbergbygg.no) —
not an in-compose container — so it reaches NATS on the published
`127.0.0.1:4222`. Give the content repo's `lint-and-reload.yml` a reload job
whose `runs-on` matches that runner's host label, running
`nats --server nats://portal:$NATS_PASSWORD@127.0.0.1:4222 pub portal.content.reload ""`.
Until then, `podman compose restart portal` picks up new content.
## Upgrading portal
Bump `PORTAL_RELEASE` in `.env` (a tag of uhhm/portal) and
`podman compose up -d --build portal`. Keep `IRIS_RELEASE` in the
content repo's workflow matched to it.
## Backups
- Kanidm writes a nightly backup into its volume (`/data/backups`,
seven kept); copy that directory off the host.
- NATS JetStream data is the `nats_data` volume: every record ever
submitted and every state change. Snapshot the volume.
- Caddy's certificates regenerate; nothing to keep.
## What is not here
Mail (a person's reset link is a token you hand them), monitoring, and
the vel's current website at tomtervel.no, which stays where it is
until the vel points the apex at PORTAL_HOST.
## Mail
Portal decides what to send - a `mail:` on a state in the content, or an
invite - and publishes it on this stack's NATS. `gdo` is what hands it to
a mail server, and it runs here, in the vel's own stack, so the vel's mail
leaves on the vel's own terms.
It has no mail server of its own. It relays through the host's, which on
kasse is Klingenberg Bygg's postfix on port 25, reached from the container
as `host.containers.internal`. That is the one thing in this stack the vel
borrows, and the one thing that changes if the vel ever moves to a host of
its own: point `SMTP_HOST` at whatever that host runs.
`site.yaml` in the content repo needs a `mail.from`, or portal has nothing
to send from and says so in its log.
Replies are not read yet: `JMAP_URL`, `JMAP_TOKEN` and `MAIL_REPLY_DOMAIN`
turn a reply into a note on the record it answers, and none of them are set
here.
## How the login page looks
Out of the box Kanidm presents itself: its own name, its own mark. A
neighbour arriving at a login page that belongs to nobody in particular
is right to hesitate, so `kanidm-setup.sh` sets the vel's name and mark
on both the instance and the portal client:
```sh
kanidm -D idm_admin -H https://$ID_HOST system domain set-displayname "Tomter Vel"
kanidm -D idm_admin -H https://$ID_HOST system domain set-image logo.svg svg
```
The logo is fetched from the content repo - the same `images/logo.svg`
that `site.yaml` uses as the favicon - so the login page and the site
are branded from one place. png, jpg, gif, svg and webp all work.
Two things the setup script deliberately leaves alone, because they are
decisions rather than branding:
- `system domain set-allow-account-recovery true` lets someone who has
lost their credentials get a reset link by proving one of their own
email addresses, instead of asking the board. Worth having for a vel;
it is still a door, so it is turned on knowingly or not at all.
- `system domain set-allow-easter-eggs` - seasonal icons and birthday
surprises. Off in production builds, and this is somebody's vel.
## Keep the CLI and the server on the same version
The `kanidm` CLI on the host and the `kanidm/server` image in
`compose.yml` must match. They are not merely fussy about it: a 1.11.2
client against a 1.11.1 server looks up the domain entry at a UUID the
older server does not have, so `system domain set-displayname` and
`set-image` fail with "Item not found" - a message that says nothing
about versions. The CLI warns on every call; the warning is worth
reading.
The same "Item not found" also comes back when the versions do match
and the account is `idm_admin`: the instance's own name and logo are
system settings only `admin` may change. `kanidm-setup.sh` sets them
when an `admin` session exists and says so when it does not.
The host's CLI comes from pacman and moves on its own, so the image is
what to bump: `image: kanidm/server:<version>` in `compose.yml`, then
`podman compose up -d kanidm`.
## Mail leaving the host
gdo hands the vel's mail to kasse's own postfix over the container
network, so postfix has to be willing to relay from it: `mynetworks`
on kasse includes `172.16.0.0/12` and `10.88.0.0/16` (set 2026-09-30,
the original is in `/etc/postfix/main.cf.before-containers`). Without
it every address outside the host's own domains is refused with
"Relay access denied", and gdo retries a minute apart until it is.