Decisions

Why the system looks the way it does — architecture decisions and their rationale. Consult this before changing fundamentals; append new entries when decisions are made.

Tailscale-only ingress

Decision: nothing is reachable from the public internet. No public service DNS records, no port forwards, Proxmox UI only via tailnet. Rationale: eliminates the entire public attack surface; the threat model for a single-user/small-group stack doesn’t justify public exposure. Consequence: access requires tailnet membership or device sharing; broke once before when the old model (public-IP firewall allowlist) depended on a home IP that changed — see lockout-recovery.

Real TLS certs via DNS-01

Decision: Caddy obtains Let’s Encrypt certs for *.«MYDOMAIN» service names via Cloudflare DNS-01 challenges. Rationale: valid certs for internal-only names without any public A records — TLS warnings gone, no exposure added. Consequence: Caddy needs a Cloudflare API token and must resolve public DNS for ACME (hence --accept-dns=false on its tailscale config).

Split-horizon DNS on Technitium

Decision: «MYDOMAIN» is authoritative on an internal Technitium server; tailnet clients reach it via Split DNS at «TECHNITIUM_TAILNET_IP». Cloudflare only carries NS/MX/ACME-TXT. Rationale: service names resolve only for tailnet clients — the DNS layer itself enforces the no-public-ingress invariant.

Device sharing instead of tailnet invites

Decision: other users run their own tailnets; the caddy and technitiumdns devices are shared into them. Rationale: avoids tailnet user-seat limits and keeps the primary tailnet small. Consequence: users must configure Split DNS on their side — onboarding is documented in dns.

No subnet routing

Decision: 10.0.1.0/24 is never advertised to the tailnet; specific machines are joined instead. Rationale: least-exposure — a tailnet client can only reach the explicitly joined machines. Consequence: gitea needed tailnet membership for SSH because Caddy can’t proxy it.

Helper-script LXCs by default

Decision: services deploy as Proxmox community helper-script LXCs; manual builds only when required. Rationale: fast, consistent, update-able. Consequence: rootless-podman-per-VM experiments abandoned (rootless-podman). Update 2026-10: superseded for new builds by the manual-first decision below — helper scripts stopped being the default. Existing manual guests: vpngw (LXC 111), devvm (QEMU 112), and LXC 114.

Pocket-ID + LLDAP over Authentik

Decision: Pocket-ID is the OIDC provider, backed by LLDAP; Authentik deprecated. Consumers: gitea.

Isolated subnet for the dev/agent VM

Decision (live — deployed 2026-09): the agent dev environment runs as a QEMU guest on a dedicated SDN subnet (vnet1, 10.0.2.0/24) whose only egress is a Mullvad tunnel via a dedicated gateway guest (see next entry) — iptables drops all guest-initiated traffic to private ranges, tailnet CGNAT, and the host itself. See dev-vm and topology. Rationale: moves the dev env onto the always-on server while keeping the hard boundary that the agent environment sees nothing identifying — enforced by the network, not by trusting the workload. Consequence: access is ssh via «HOSTNAME» jump; the VM is deliberately not tailnet-joined (would expose peer names/«TAILNET_DOMAIN»). If a workload ever needs internal access, that’s a new decision — document it here first.

Dev VM egresses via Mullvad, not «PUBLIC_IP»

Decision (live — deployed 2026-09): vnet1 gets a dedicated gateway guest (vpngw, LXC 111) running a WireGuard tunnel to Mullvad. devvm routes through it; a ! -o wg0 drop on vpngw plus a ! -s 10.0.2.2 drop on the host form the killswitch. Rationale: plain SNAT would make devvm’s egress IP «PUBLIC_IP» — trivially discoverable (curl ifconfig.me) and logged by every service it contacts. «PUBLIC_IP» is a redacted value, so the egress path must not expose it. Rejected alternatives: Mullvad tunnel on the host itself (adds policy routing + a tunnel to host networking); accepting the exposure (deliberate hole in the redaction boundary). Consequence: devvm internet depends on vpngw — no tunnel, no connectivity, by design. Mullvad account uses a second device slot. Chosen over host-side tunneling to keep host networking minimal; a failed earlier attempt (mullvad-gateway) was a different mechanism (wg-quick inside the workload VM, not a gateway guest). Implementation note: pve-firewall is enabled and host INPUT is effectively unfiltered, so all vnet1 isolation lives in the mangle table — evaluated before the filter table, immune to PVEFW state/policies. Rules are applied by /etc/network/if-up.d/vnet-rules — interfaces.d/sdn is fully SDN-generated (Apply regenerates it, including its own SNAT/CT-zone post-up lines; hand-added lines are wiped — confirmed). Transport: discovered at build that the provider edge drops all inbound UDP including replies — raw WireGuard can never handshake. So vpngw runs the mullvad app with udp2tcp (WG wrapped in TCP) rather than wg-quick, plus a /dev/net/tun passthrough (unprivileged LXC), TCP-only bootstrap DNS (options use-vc), MSS clamping (tunnel MTU 1326), and devvm DNS via Mullvad’s in-tunnel resolver 10.64.0.1.

Operator-gated host diagnostics (diag-pipe)

Decision (implemented — 2026-10): host diagnostics during agent-led troubleshooting run through a job pipeline orchestrated on the operator machine: the agent submits job files from devvm, the operator approves/denies each one, approved jobs execute on the host over ssh, output is sanitized by the docmirror redaction engine, and the operator reviews before release to the guest. See diag-pipe. Rationale: the EINVAL postmortem took seven+ rounds of manual copy/paste ferrying; the pipeline automates transport and redaction while keeping both human gates (script approval, output release) that make it safe — the guest never learns real values and nothing runs unreviewed. Key choices: operator machine as trust boundary; pre-approved parameterized templates carry routine probe shapes while novel scripts always need explicit approval; scripts may be token-written and are unredacted at approve time — the operator reviews what actually executes. Rejected: virtiofs exchange dir + host-side watcher (machinery and secrets maps on the hypervisor, gains only unattended staging); regex command allowlists (templates bound arguments too). Consequence: redaction stays assist-not-boundary — output review is the leak control. Shakedown confirmed coverage (««HOSTNAME», ««SB_*» round-trips, ««AUTO_IPV4_*»/««AUTO_MAC_*» minting) while leaving unmapped-but-benign values (UUIDs, ClientGUID) to the human gate.

Fresh QEMU build over migrating the desktop image

Decision: recreate the dev VM from a bootstrap script (dev-vm-bootstrap) rather than moving the workstation’s qcow2. Rationale: the desktop image is a full Xubuntu with libvirt/X11 baggage; a scripted headless Debian build is reproducible and is the documentation.

Manual-first deployment; IaC as a follow-up (planned)

Decision: stop treating community helper scripts as the default for new services. Formalize the manual pattern — pct create/qm create + per-service bootstrap + compose where containers are wanted — then investigate IaC (OpenTofu or Ansible) as a separate step. Rationale: the scripts already failed to cover recent needs — Debian 13 lag (vpngw), hangs on isolated vnet1 (vpngw), opaque images. The manual pattern has worked; document it before automating it. Note: existing helper-script guests stay as-is — this applies to new deployments only. Execution-side caveat: provisioning can’t run on devvm (vnet1 isolation) — pipelines are operator-side, with definitions in gitea.

Offsite backups via encrypted sync, not PBS (planned)

Decision: keep the vzdump → backups LV jobs and add an encrypted sync (restic/borg) of the dump dir to an external destination — Storage Box to start, retargetable (Proton Drive, a home machine). Rationale: backups is same-RAID1 — there is currently no offsite copy at all; PBS adds dedupe/verification but doesn’t fix that without a second host. Encrypted sync closes the DR gap at low effort, client-side encryption makes the public Storage Box safe to use, and the destination stays swappable. Revisit if: an always-on second box appears (remote PBS then gives real incremental dedupe offsite), or backup size makes restic throughput painful.

QEMU 108 rebuilt as LXC (virtiofs retired)

Decision (live — 2026-10-08): QEMU 108 moved to unprivileged LXC 114 (same IP, 10.0.1.28), taking the standard mp0 bind mount for the share instead of virtiofs. Rationale: forced by a PVE 9 regression — QEMU 108 deadlooped in OVMF’s VirtioFs DXE driver on virtiofs0 under QEMU 11 (2026-10-08-ovmf-virtiofs-deadloop). The LXC pattern is already proven, removes the whole virtiofsd failure class (fd-cache EINVAL poisoning, session-drop wedges), and drops QEMU overhead for a service that never needed it. Consequence: mullvad-daemon in an unprivileged CT needs nesting=1 + /dev/net/tun passthrough (vpngw precedent); QEMU guests are now just devvm. virtiofs remains only on devvm — it’s SeaBIOS, which never probes the device.

Storage Box is “durable enough”

Decision: the Hetzner Storage Box (provider-side multi-drive RAID) is treated as the durable tier for media. Update 2026-09: guest backups moved off it to the local backups LV (same-RAID1 — no offsite copy; see the encrypted-sync decision above). Storage Box is now media-only and may become the encrypted backup target. Caveat: it’s still one provider, one failure domain.