Decisions
Why the system looks the way it does — architecture decisions and their rationale. Consult this before changing fundamentals; append new entries when decisions are made.
Tailscale-only ingress
Decision: nothing is reachable from the public internet. No public service DNS records, no port forwards, Proxmox UI only via tailnet. Rationale: eliminates the entire public attack surface; the threat model for a single-user/small-group stack doesn’t justify public exposure. Consequence: access requires tailnet membership or device sharing; broke once before when the old model (public-IP firewall allowlist) depended on a home IP that changed — see lockout-recovery.
Real TLS certs via DNS-01
Decision: Caddy obtains Let’s Encrypt certs for *.«MYDOMAIN»
service names via Cloudflare DNS-01 challenges.
Rationale: valid certs for internal-only names without any public A
records — TLS warnings gone, no exposure added. Consequence: Caddy
needs a Cloudflare API token and must resolve public DNS for ACME (hence
--accept-dns=false on its tailscale config).
Split-horizon DNS on Technitium
Decision: «MYDOMAIN» is authoritative on an internal Technitium
server; tailnet clients reach it via Split DNS at
«TECHNITIUM_TAILNET_IP». Cloudflare only carries NS/MX/ACME-TXT.
Rationale: service names resolve only for tailnet clients — the DNS
layer itself enforces the no-public-ingress invariant.
Device sharing instead of tailnet invites
Decision: other users run their own tailnets; the caddy and
technitiumdns devices are shared into them.
Rationale: avoids tailnet user-seat limits and keeps the primary
tailnet small. Consequence: users must configure Split DNS on their
side — onboarding is documented in dns.
No subnet routing
Decision: 10.0.1.0/24 is never advertised to the tailnet; specific
machines are joined instead.
Rationale: least-exposure — a tailnet client can only reach the
explicitly joined machines. Consequence: gitea needed tailnet
membership for SSH because Caddy can’t proxy it.
Helper-script LXCs by default
Decision: services deploy as Proxmox community helper-script LXCs;
manual builds only when required.
Rationale: fast, consistent, update-able. Consequence:
rootless-podman-per-VM experiments abandoned (rootless-podman).
Update 2026-10: superseded for new builds by the manual-first
decision below — helper scripts stopped being the default. Existing
manual guests: vpngw (LXC 111), devvm (QEMU 112), and LXC 114.
Pocket-ID + LLDAP over Authentik
Decision: Pocket-ID is the OIDC provider, backed by LLDAP; Authentik deprecated. Consumers: gitea.
Isolated subnet for the dev/agent VM
Decision (live — deployed 2026-09): the agent dev environment runs as a QEMU guest
on a dedicated SDN subnet (vnet1, 10.0.2.0/24) whose only egress is
a Mullvad tunnel via a dedicated gateway guest (see next entry) —
iptables drops all guest-initiated traffic to private ranges, tailnet
CGNAT, and the host itself. See dev-vm and topology.
Rationale: moves the dev env onto the always-on server while keeping
the hard boundary that the agent environment sees nothing identifying —
enforced by the network, not by trusting the workload.
Consequence: access is ssh via «HOSTNAME» jump; the VM is
deliberately not tailnet-joined (would expose peer
names/«TAILNET_DOMAIN»). If a workload ever needs internal access,
that’s a new decision — document it here first.
Dev VM egresses via Mullvad, not «PUBLIC_IP»
Decision (live — deployed 2026-09): vnet1 gets a dedicated gateway guest (vpngw,
LXC 111) running a WireGuard tunnel to Mullvad. devvm routes through it;
a ! -o wg0 drop on vpngw plus a ! -s 10.0.2.2 drop on the host form
the killswitch.
Rationale: plain SNAT would make devvm’s egress IP «PUBLIC_IP» —
trivially discoverable (curl ifconfig.me) and logged by every service
it contacts. «PUBLIC_IP» is a redacted value, so the egress path must
not expose it.
Rejected alternatives: Mullvad tunnel on the host itself (adds
policy routing + a tunnel to host networking); accepting the exposure
(deliberate hole in the redaction boundary).
Consequence: devvm internet depends on vpngw — no tunnel, no
connectivity, by design. Mullvad account uses a second device slot.
Chosen over host-side tunneling to keep host networking minimal; a
failed earlier attempt (mullvad-gateway) was a different mechanism
(wg-quick inside the workload VM, not a gateway guest).
Implementation note: pve-firewall is enabled and host INPUT is
effectively unfiltered, so all vnet1 isolation lives in the mangle
table — evaluated before the filter table, immune to PVEFW
state/policies. Rules are applied by /etc/network/if-up.d/vnet-rules
— interfaces.d/sdn is fully SDN-generated (Apply regenerates it,
including its own SNAT/CT-zone post-up lines; hand-added lines are
wiped — confirmed).
Transport: discovered at build that the provider edge drops all
inbound UDP including replies — raw WireGuard can never handshake. So
vpngw runs the mullvad app with udp2tcp (WG wrapped in TCP)
rather than wg-quick, plus a /dev/net/tun
passthrough (unprivileged LXC), TCP-only bootstrap DNS
(options use-vc), MSS clamping (tunnel MTU 1326), and devvm DNS via
Mullvad’s in-tunnel resolver 10.64.0.1.
Operator-gated host diagnostics (diag-pipe)
Decision (implemented — 2026-10): host diagnostics during agent-led
troubleshooting run through a job pipeline orchestrated on the operator
machine: the agent submits job files from devvm, the operator
approves/denies each one, approved jobs execute on the host over ssh,
output is sanitized by the docmirror redaction engine, and the operator
reviews before release to the guest. See diag-pipe.
Rationale: the EINVAL postmortem took seven+ rounds of manual
copy/paste ferrying; the pipeline automates transport and redaction
while keeping both human gates (script approval, output release) that
make it safe — the guest never learns real values and nothing runs
unreviewed.
Key choices: operator machine as trust boundary;
pre-approved parameterized templates carry routine probe shapes while
novel scripts always need explicit approval; scripts may be
token-written and are unredacted at approve time — the operator reviews
what actually executes.
Rejected: virtiofs exchange dir + host-side watcher (machinery and
secrets maps on the hypervisor, gains only unattended staging);
regex command allowlists (templates bound arguments too).
Consequence: redaction stays assist-not-boundary — output review
is the leak control. Shakedown confirmed coverage (««HOSTNAME»,
««SB_*» round-trips, ««AUTO_IPV4_*»/««AUTO_MAC_*» minting) while
leaving unmapped-but-benign values (UUIDs, ClientGUID) to the human
gate.
Fresh QEMU build over migrating the desktop image
Decision: recreate the dev VM from a bootstrap script (dev-vm-bootstrap) rather than moving the workstation’s qcow2. Rationale: the desktop image is a full Xubuntu with libvirt/X11 baggage; a scripted headless Debian build is reproducible and is the documentation.
Manual-first deployment; IaC as a follow-up (planned)
Decision: stop treating community helper scripts as the default
for new services. Formalize the manual pattern — pct create/qm create + per-service bootstrap + compose where containers are wanted
— then investigate IaC (OpenTofu or Ansible) as a separate step.
Rationale: the scripts already failed to cover recent needs —
Debian 13 lag (vpngw), hangs on isolated vnet1 (vpngw), opaque
images. The manual pattern has worked; document it before automating
it. Note: existing helper-script guests stay as-is — this applies
to new deployments only. Execution-side caveat: provisioning can’t
run on devvm (vnet1 isolation) — pipelines are operator-side, with
definitions in gitea.
Offsite backups via encrypted sync, not PBS (planned)
Decision: keep the vzdump → backups LV jobs and add an
encrypted sync (restic/borg) of the dump dir to an external
destination — Storage Box to start, retargetable (Proton Drive, a
home machine). Rationale: backups is same-RAID1 — there is
currently no offsite copy at all; PBS adds dedupe/verification
but doesn’t fix that without a second host. Encrypted sync closes
the DR gap at low effort, client-side encryption makes the public
Storage Box safe to use, and the destination stays swappable.
Revisit if: an always-on second box appears (remote PBS then
gives real incremental dedupe offsite), or backup size makes restic
throughput painful.
QEMU 108 rebuilt as LXC (virtiofs retired)
Decision (live — 2026-10-08): QEMU 108 moved to unprivileged LXC 114
(same IP, 10.0.1.28), taking the standard mp0 bind mount for the
share instead of virtiofs.
Rationale: forced by a PVE 9 regression — QEMU 108 deadlooped in
OVMF’s VirtioFs DXE driver on virtiofs0 under QEMU 11
(2026-10-08-ovmf-virtiofs-deadloop). The LXC pattern is already
proven, removes the
whole virtiofsd failure class (fd-cache EINVAL poisoning,
session-drop wedges), and drops QEMU overhead for a service that
never needed it.
Consequence: mullvad-daemon in an unprivileged CT needs
nesting=1 + /dev/net/tun passthrough (vpngw precedent); QEMU
guests are now just devvm. virtiofs remains only on devvm — it’s
SeaBIOS, which never probes the device.
Storage Box is “durable enough”
Decision: the Hetzner Storage Box (provider-side multi-drive RAID)
is treated as the durable tier for media.
Update 2026-09: guest backups moved off it to the local backups
LV (same-RAID1 — no offsite copy; see the encrypted-sync decision
above). Storage Box is now media-only and may become the encrypted
backup target.
Caveat: it’s still one provider, one failure domain.