Runbook: Proxmox VE 8 → 9 in-place upgrade

Completed 2026-10-08

Host now runs PVE 9.2.21 on trixie, kernel 7.0.14-22-pve; final pve8to9 run clean (34 pass / 7 skip / 0 warn / 0 fail), all guests returned, vnet1 isolation + both backup jobs intact. The upgrade ran end-to-end. See “Outcome notes” in §5 for what deviated from this plan (NIC pin, repo dedup, enterprise-repo 401, the VM 108 virtiofs casualty). Kept for future major upgrades.

In-place apt upgrade of «HOSTNAME» (PVE 8.4.21, Debian bookworm → trixie). Tracked in TODO. Official guide: https://pve.proxmox.com/wiki/Upgrade_from_8_to_9 — this runbook is that guide filtered to this host.

Access risk — read first

The host has a single public IP and no KVM console. The only management path is tailnet → ssh (and the PVE UI over tailnet). If networking doesn’t come back after reboot — most likely cause: the physical NIC enp0s31f6 gets renamed by the new kernel — recovery is Hetzner rescue mode → chroot, same flow as lockout-recovery. Step 1.6 (NIC pinning) exists to prevent exactly this; don’t skip it.

0. Scope notes

  • Single node — no cluster/HA/Ceph considerations apply.
  • PVE 9 is pure cgroup-v2: guests need systemd ≥ 230 — trivially true for all documented guests (Debian 12/13, Ubuntu 26.04).
  • While the host reboots, devvm (QEMU 112) goes down with it — this session’s agent included. Keep a fallback agent session staged on the desktop VM.
  • Guests keep their existing QEMU machine versions — don’t bump them to 10.x; nothing here needs it.
  • /tmp becomes a tmpfs under trixie (≤50% RAM, periodic cleanup) — just noting it, nothing to do.

1. Pre-flight

  • Host fully patched on PVE 8: apt update && apt dist-upgrade, then pveversion — must be ≥ 8.4.1 (docs say 8.4.21 already).
  • Fresh vzdump of all guests on backups (/mnt/backups). The scheduled jobs run Mon/Thu — if the last run is stale, trigger both Datacenter → Backup jobs manually and confirm files land in /mnt/backups/dump. Caveat stands: same-RAID1 backups cover rollback/config loss, not disk death — acceptable for this.
  • Back up host config for the chroot-rescue scenario:
tar czf /root/host-config-$(date +%F).tgz \
  /etc/pve /etc/network /etc/default /etc/sysctl.conf /etc/sysctl.d
  (`/etc/network` covers `interfaces.d/sdn` and
  `if-up.d/vnet-rules`; `/etc/pve` carries storage.cfg, guest
  configs, firewall + backup jobs.)
  • Root filesystem space: df -h / — need ≥ 5 GB free, ideally 10.
  • linux-image-amd64 conflict — this host was Debian installed by Hetzner + PVE post-install, so the package is likely present and it conflicts with PVE 9’s kernel packaging:
dpkg -l linux-image-amd64   # if installed:
apt remove linux-image-amd64
  • Boot mode + GRUB-on-LVM: [ -d /sys/firmware/efi ] — root is on LVM (vg0-root), so if this is EFI, install the fixed grub meta-package before dist-upgrade:
[ -d /sys/firmware/efi ] && apt install grub-efi-amd64
  • systemd-sysctl on PVE 9 ignores /etc/sysctl.conf — if it has any non-comment lines (grep -vE '^\s*(#|$)' /etc/sysctl.conf), move them into /etc/sysctl.d/90-*.conf.
  • Pin the physical NIC name. PVE 9’s kernel names interfaces from more PCI features than 8’s — enp0s31f6 can change, and /etc/network/interfaces would then reference a nonexistent NIC. Generate a .link file that pins the MAC to a stable name:
which pve-network-interface-pinning   # shipped in current pve-manager
pve-network-interface-pinning generate \
  --interface enp0s31f6 --target-name enp0s31f6
  Pinning to the *current* name keeps `/etc/network/interfaces`
  unchanged; if the tool refuses a same-name pin, use
  `--target-name nic0` instead and let it rewrite the pending
  `interfaces.new`. Then **reboot and verify** `ip link` shows the
  pinned name and tailnet access still works — this proves the
  .link file before the risky upgrade reboot. (If the tool doesn't
  exist on 8.4.21, note the current name + MAC and keep rescue
  ready — renaming is then a known failure mode, not a surprise.)
  • Silence the audit-log spam during upgrade: systemctl disable --now systemd-journald-audit.socket.
  • Third-party repos: ls /etc/apt/sources.list.d/ — everything must have a trixie suite or be commented out. tailscale.list is the one that matters (tailscale publishes a trixie suite — repoint it in step 2 alongside the Debian lines). Hetzner mirror lines get caught by the same bookworm→trixie sed.
  • Run the checker, fix what it flags, re-run until clean:
pve8to9 --full
  Expected on this host: the LVM autoactivation migration hint —
  optional for node-local thick LVM (running
  `/usr/share/pve-manager/migrations/pve-lvm-disable-autoactivation`
  is harmless either way); `systemd-boot` meta-package removal if
  it shipped with the ISO-era install — safe to remove unless the
  host was deliberately set up on systemd-boot (it wasn't).
  • Have lockout-recovery open + know how to activate Hetzner rescue in the Robot panel before the reboot steps.

2. Upgrade window

Work in tmux on the host — a dropped tailnet ssh mid-dist-upgrade must not kill the process (apt install tmux if absent).

  • Guests may stay running through the package phase; they get stopped cleanly at the reboot. Optionally pct shutdown / qm shutdown non-criticals first to shorten the reboot.
  • Repoint repos to trixie:
sed -i 's/bookworm/trixie/g' /etc/apt/sources.list
sed -i 's/bookworm/trixie/g' /etc/apt/sources.list.d/*.list   # incl. tailscale
 
cat > /etc/apt/sources.list.d/proxmox.sources <<'EOF'
Types: deb
URIs: http://download.proxmox.com/debian/pve
Suites: trixie
Components: pve-no-subscription
Signed-By: /usr/share/keyrings/proxmox-archive-keyring.gpg
EOF
  Remove/comment the old PVE 8 repo (pve-enterprise.list or
  pve-install-repo.list — whichever exists). If any Ceph repo
  exists remove it — no Ceph here (`which ceph` should be empty).
  • apt update && apt policy — no errors, no bookworm lines left.
  • apt dist-upgrade. Conffile prompts — diff each; defaults for this host:
FileChoice
/etc/issuekeep local (cosmetic)
/etc/lvm/lvm.confmaintainer’s version if unmodified
/etc/ssh/sshd_configdiff first — if only the ChallengeResponseAuthentication→KbdInteractiveAuthentication rename, take maintainer’s; otherwise keep local
/etc/default/grubkeep local
/etc/chrony/chrony.confmaintainer’s version if unmodified
  • pve8to9 again — should be clean.
  • Reboot — required even if a 6.14 kernel was already in use.

3. Post-upgrade verification

Host (in this order — connectivity first):

  • tailnet access works; tailscale status healthy. Verify before logging off anything.
  • pveversion -v → 9.x; pve8to9 re-run clean; systemctl --failed empty; journalctl -b -p err sane.
  • ip link — pinned NIC name present and up; ip r has the default route; PVE UI loads over tailnet (force-reload browser).
  • pvesm status — local, main, backups, storagebox all active; ls /mnt/pve/storagebox shows the share.
  • cat /etc/network/interfaces.d/sdn — SDN file regenerated correctly (do not hand-edit; the vnet1 rules live in the if-up hook).
  • vnet1 isolation intact: iptables -t mangle -L PREROUTING -n and -L FORWARD -n show the vnet1 ruleset from /etc/network/if-up.d/vnet-rules. Trixie’s iptables is the nft backend — the same commands work; cross-check nft list ruleset | grep vnet1. If the rules are missing: IFACE=vnet1 /etc/network/if-up.d/vnet-rules.
  • The new nftables-based proxmox-firewall is opt-in in PVE 9 — leave it off; vnet1 isolation is mangle-table and unaffected.
  • Datacenter → Backup: both jobs present (lxc-suspend + vm-snapshot pools, backups storage, keep-last=5).
  • Optional: apt modernize-sources to move Debian entries to deb822 .sources files.

Guests (all onboot=1; verify, don’t assume):

  • pct list / qm list — all guests running.
  • vpngw (111): pct exec 111 -- systemctl --failed — watch for the nesting/CREDENTIALS signature (dev-vm: journald/sysctl dead → ip_forward silently 0). pct exec 111 -- mullvad status → Connected, udp2tcp. pct exec 111 -- curl -4 ifconfig.me → Mullvad exit IP.
  • devvm: ssh in via the jump path; curl -4 ifconfig.me → Mullvad IP, never «PUBLIC_IP»; ping 10.0.1.20 fails (isolation still holds — mangle ruleset verified above).
  • technitiumdns (101): dig +short wastebin.«MYDOMAIN» @10.0.1.1 from the host — resolves.
  • caddy (100): hit each service URL — valid cert, 200/expected responses.

4. Recovery paths

  • Network/tailscale dead after reboot → Hetzner Robot → activate rescue system → reboot → rescue shell → chroot per lockout-recovery (vgchange -ay, mount vg0-root, bind proc/sys/dev). Inside: ip link for the renamed NIC → fix /etc/network/interfaces, or systemctl disable pve-firewall / ifreload -a, or remove the pinning .link (/usr/local/lib/systemd/network/50-pve-*.link).
  • apt/dpkg broke mid-upgrade → same chroot → dpkg --configure -a then apt -f install.
  • Unbootable → PVE ISO “Rescue Boot” (advanced menu) if a console is ever attached; otherwise rescue mode + restore host files from the tarball made in 1.3.
  • Total loss → disaster-recovery: rebuild host, restore guests from backups (/mnt/backups/dump).

5. After the dust settles

  • Update inventory (PVE version row), check off the TODO item — done 2026-10-08.
  • Guest OS upgrades (bookworm→trixie in LXCs) are a separate pass — not part of this runbook. (LXC 114 sidestepped it: rebuilt as a Debian 13 LXC.)
  • If anything fought back during the upgrade, record it where it belongs (topology, decisions, or the service page) — outcome notes below.

Outcome notes (2026-10-08 run)

Deviations and findings worth keeping:

  • Root space: / had ~5 GB free pre-flight — grew vg0/root lvresize -r -L +20G before dist-upgrade (~24 GB free after).
  • NIC pinning: --target-name enp0s31f6 was refused (“target-name already exists as link or pin!”) → used nic0; pve-network-interface-pinning wrote /usr/local/lib/systemd/network/50-pve-nic0.link + a rewritten interfaces.new (iface nic0 / bridge-ports nic0). A verification reboot before the upgrade proved the pin — tailnet, vmbr0, and all guests returned. The permanent name is nic0 now.
  • Housekeeping done pre-upgrade: chrony installed; systemd-journald-audit.socket disabled; LVM autoactivation migration (pve-lvm-disable-autoactivation) run.
  • Repo dedup: removed hetzner-security-updates.list + pve-install-repo.list (dupes), commented proxmox.list, added deb822 proxmox.sources (pve-no-subscription), and appended non-free-firmware to Debian components (intel-microcode).
  • Post-upgrade 401: pve-enterprise.sources was enabled on a no-subscription install → apt update failed with 401 Unauthorized. Fix: Enabled: no in the .sources file — don’t delete package-managed repo files.
  • Guest casualty: QEMU 108 hung in OVMF after the upgrade — frozen RIP, zero disk I/O; bisected to the virtiofs0 device (OVMF’s VirtioFs DXE driver deadloops on it under QEMU 11; SeaBIOS guests like devvm are unaffected). Fixed by deleting virtiofs0; the service was rebuilt as LXC 114 with an mp0 bind mount instead. Incident: 2026-10-08-ovmf-virtiofs-deadloop.
  • Optional cleanup still available: apt modernize-sources for the remaining Debian .list files; obsolete commented .list backups (*.list.bak) and legacy /etc/pve/priv/{ipam,macs}.db can be removed; old-format RRD files may be regenerated.