Runbook: Proxmox VE 8 → 9 in-place upgrade
Completed 2026-10-08
Host now runs PVE 9.2.21 on trixie, kernel
7.0.14-22-pve; finalpve8to9run clean (34 pass / 7 skip / 0 warn / 0 fail), all guests returned, vnet1 isolation + both backup jobs intact. The upgrade ran end-to-end. See “Outcome notes” in §5 for what deviated from this plan (NIC pin, repo dedup, enterprise-repo 401, the VM 108 virtiofs casualty). Kept for future major upgrades.
In-place apt upgrade of «HOSTNAME» (PVE 8.4.21, Debian bookworm →
trixie). Tracked in TODO. Official guide:
https://pve.proxmox.com/wiki/Upgrade_from_8_to_9 — this runbook is that
guide filtered to this host.
Access risk — read first
The host has a single public IP and no KVM console. The only management path is tailnet → ssh (and the PVE UI over tailnet). If networking doesn’t come back after reboot — most likely cause: the physical NIC
enp0s31f6gets renamed by the new kernel — recovery is Hetzner rescue mode → chroot, same flow as lockout-recovery. Step 1.6 (NIC pinning) exists to prevent exactly this; don’t skip it.
0. Scope notes
- Single node — no cluster/HA/Ceph considerations apply.
- PVE 9 is pure cgroup-v2: guests need systemd ≥ 230 — trivially true for all documented guests (Debian 12/13, Ubuntu 26.04).
- While the host reboots, devvm (QEMU 112) goes down with it — this session’s agent included. Keep a fallback agent session staged on the desktop VM.
- Guests keep their existing QEMU machine versions — don’t bump them to 10.x; nothing here needs it.
/tmpbecomes a tmpfs under trixie (≤50% RAM, periodic cleanup) — just noting it, nothing to do.
1. Pre-flight
- Host fully patched on PVE 8:
apt update && apt dist-upgrade, thenpveversion— must be ≥ 8.4.1 (docs say 8.4.21 already). - Fresh vzdump of all guests on
backups(/mnt/backups). The scheduled jobs run Mon/Thu — if the last run is stale, trigger both Datacenter → Backup jobs manually and confirm files land in/mnt/backups/dump. Caveat stands: same-RAID1 backups cover rollback/config loss, not disk death — acceptable for this. - Back up host config for the chroot-rescue scenario:
tar czf /root/host-config-$(date +%F).tgz \
/etc/pve /etc/network /etc/default /etc/sysctl.conf /etc/sysctl.d (`/etc/network` covers `interfaces.d/sdn` and
`if-up.d/vnet-rules`; `/etc/pve` carries storage.cfg, guest
configs, firewall + backup jobs.)
- Root filesystem space:
df -h /— need ≥ 5 GB free, ideally 10. -
linux-image-amd64conflict — this host was Debian installed by Hetzner + PVE post-install, so the package is likely present and it conflicts with PVE 9’s kernel packaging:
dpkg -l linux-image-amd64 # if installed:
apt remove linux-image-amd64- Boot mode + GRUB-on-LVM:
[ -d /sys/firmware/efi ]— root is on LVM (vg0-root), so if this is EFI, install the fixed grub meta-package before dist-upgrade:
[ -d /sys/firmware/efi ] && apt install grub-efi-amd64-
systemd-sysctlon PVE 9 ignores/etc/sysctl.conf— if it has any non-comment lines (grep -vE '^\s*(#|$)' /etc/sysctl.conf), move them into/etc/sysctl.d/90-*.conf. - Pin the physical NIC name. PVE 9’s kernel names interfaces
from more PCI features than 8’s —
enp0s31f6can change, and/etc/network/interfaceswould then reference a nonexistent NIC. Generate a .link file that pins the MAC to a stable name:
which pve-network-interface-pinning # shipped in current pve-manager
pve-network-interface-pinning generate \
--interface enp0s31f6 --target-name enp0s31f6 Pinning to the *current* name keeps `/etc/network/interfaces`
unchanged; if the tool refuses a same-name pin, use
`--target-name nic0` instead and let it rewrite the pending
`interfaces.new`. Then **reboot and verify** `ip link` shows the
pinned name and tailnet access still works — this proves the
.link file before the risky upgrade reboot. (If the tool doesn't
exist on 8.4.21, note the current name + MAC and keep rescue
ready — renaming is then a known failure mode, not a surprise.)
- Silence the audit-log spam during upgrade:
systemctl disable --now systemd-journald-audit.socket. - Third-party repos:
ls /etc/apt/sources.list.d/— everything must have a trixie suite or be commented out. tailscale.list is the one that matters (tailscale publishes a trixie suite — repoint it in step 2 alongside the Debian lines). Hetzner mirror lines get caught by the samebookworm→trixiesed. - Run the checker, fix what it flags, re-run until clean:
pve8to9 --full Expected on this host: the LVM autoactivation migration hint —
optional for node-local thick LVM (running
`/usr/share/pve-manager/migrations/pve-lvm-disable-autoactivation`
is harmless either way); `systemd-boot` meta-package removal if
it shipped with the ISO-era install — safe to remove unless the
host was deliberately set up on systemd-boot (it wasn't).
- Have lockout-recovery open + know how to activate Hetzner rescue in the Robot panel before the reboot steps.
2. Upgrade window
Work in tmux on the host — a dropped tailnet ssh mid-dist-upgrade
must not kill the process (apt install tmux if absent).
- Guests may stay running through the package phase; they get
stopped cleanly at the reboot. Optionally
pct shutdown/qm shutdownnon-criticals first to shorten the reboot. - Repoint repos to trixie:
sed -i 's/bookworm/trixie/g' /etc/apt/sources.list
sed -i 's/bookworm/trixie/g' /etc/apt/sources.list.d/*.list # incl. tailscale
cat > /etc/apt/sources.list.d/proxmox.sources <<'EOF'
Types: deb
URIs: http://download.proxmox.com/debian/pve
Suites: trixie
Components: pve-no-subscription
Signed-By: /usr/share/keyrings/proxmox-archive-keyring.gpg
EOF Remove/comment the old PVE 8 repo (pve-enterprise.list or
pve-install-repo.list — whichever exists). If any Ceph repo
exists remove it — no Ceph here (`which ceph` should be empty).
-
apt update && apt policy— no errors, no bookworm lines left. -
apt dist-upgrade. Conffile prompts — diff each; defaults for this host:
| File | Choice |
|---|---|
/etc/issue | keep local (cosmetic) |
/etc/lvm/lvm.conf | maintainer’s version if unmodified |
/etc/ssh/sshd_config | diff first — if only the ChallengeResponseAuthentication→KbdInteractiveAuthentication rename, take maintainer’s; otherwise keep local |
/etc/default/grub | keep local |
/etc/chrony/chrony.conf | maintainer’s version if unmodified |
-
pve8to9again — should be clean. - Reboot — required even if a 6.14 kernel was already in use.
3. Post-upgrade verification
Host (in this order — connectivity first):
- tailnet access works;
tailscale statushealthy. Verify before logging off anything. -
pveversion -v→ 9.x;pve8to9re-run clean;systemctl --failedempty;journalctl -b -p errsane. -
ip link— pinned NIC name present and up;ip rhas the default route; PVE UI loads over tailnet (force-reload browser). -
pvesm status—local,main,backups,storageboxall active;ls /mnt/pve/storageboxshows the share. -
cat /etc/network/interfaces.d/sdn— SDN file regenerated correctly (do not hand-edit; the vnet1 rules live in the if-up hook). - vnet1 isolation intact:
iptables -t mangle -L PREROUTING -nand-L FORWARD -nshow the vnet1 ruleset from/etc/network/if-up.d/vnet-rules. Trixie’s iptables is the nft backend — the same commands work; cross-checknft list ruleset | grep vnet1. If the rules are missing:IFACE=vnet1 /etc/network/if-up.d/vnet-rules. - The new nftables-based
proxmox-firewallis opt-in in PVE 9 — leave it off; vnet1 isolation is mangle-table and unaffected. - Datacenter → Backup: both jobs present (
lxc-suspend+vm-snapshotpools,backupsstorage, keep-last=5). - Optional:
apt modernize-sourcesto move Debian entries to deb822.sourcesfiles.
Guests (all onboot=1; verify, don’t assume):
-
pct list/qm list— all guests running. -
vpngw(111):pct exec 111 -- systemctl --failed— watch for the nesting/CREDENTIALS signature (dev-vm: journald/sysctl dead → ip_forward silently 0).pct exec 111 -- mullvad status→ Connected, udp2tcp.pct exec 111 -- curl -4 ifconfig.me→ Mullvad exit IP. -
devvm: ssh in via the jump path;curl -4 ifconfig.me→ Mullvad IP, never «PUBLIC_IP»;ping 10.0.1.20fails (isolation still holds — mangle ruleset verified above). -
technitiumdns(101):dig +short wastebin.«MYDOMAIN» @10.0.1.1from the host — resolves. -
caddy(100): hit each service URL — valid cert, 200/expected responses.
4. Recovery paths
- Network/tailscale dead after reboot → Hetzner Robot → activate
rescue system → reboot → rescue shell → chroot per
lockout-recovery (
vgchange -ay, mountvg0-root, bind proc/sys/dev). Inside:ip linkfor the renamed NIC → fix/etc/network/interfaces, orsystemctl disable pve-firewall/ifreload -a, or remove the pinning .link (/usr/local/lib/systemd/network/50-pve-*.link). - apt/dpkg broke mid-upgrade → same chroot →
dpkg --configure -athenapt -f install. - Unbootable → PVE ISO “Rescue Boot” (advanced menu) if a console is ever attached; otherwise rescue mode + restore host files from the tarball made in 1.3.
- Total loss → disaster-recovery: rebuild host, restore guests
from
backups(/mnt/backups/dump).
5. After the dust settles
- Update inventory (PVE version row), check off the TODO item — done 2026-10-08.
- Guest OS upgrades (bookworm→trixie in LXCs) are a separate pass — not part of this runbook. (LXC 114 sidestepped it: rebuilt as a Debian 13 LXC.)
- If anything fought back during the upgrade, record it where it belongs (topology, decisions, or the service page) — outcome notes below.
Outcome notes (2026-10-08 run)
Deviations and findings worth keeping:
- Root space:
/had ~5 GB free pre-flight — grewvg0/rootlvresize -r -L +20Gbefore dist-upgrade (~24 GB free after). - NIC pinning:
--target-name enp0s31f6was refused (“target-name already exists as link or pin!”) → usednic0;pve-network-interface-pinningwrote/usr/local/lib/systemd/network/50-pve-nic0.link+ a rewritteninterfaces.new(iface nic0/bridge-ports nic0). A verification reboot before the upgrade proved the pin — tailnet,vmbr0, and all guests returned. The permanent name isnic0now. - Housekeeping done pre-upgrade:
chronyinstalled;systemd-journald-audit.socketdisabled; LVM autoactivation migration (pve-lvm-disable-autoactivation) run. - Repo dedup: removed
hetzner-security-updates.list+pve-install-repo.list(dupes), commentedproxmox.list, added deb822proxmox.sources(pve-no-subscription), and appendednon-free-firmwareto Debian components (intel-microcode). - Post-upgrade 401:
pve-enterprise.sourceswas enabled on a no-subscription install →apt updatefailed with401 Unauthorized. Fix:Enabled: noin the .sources file — don’t delete package-managed repo files. - Guest casualty: QEMU 108 hung in OVMF after
the upgrade — frozen RIP, zero disk I/O; bisected to the
virtiofs0device (OVMF’s VirtioFs DXE driver deadloops on it under QEMU 11; SeaBIOS guests like devvm are unaffected). Fixed by deletingvirtiofs0; the service was rebuilt as LXC 114 with anmp0bind mount instead. Incident: 2026-10-08-ovmf-virtiofs-deadloop. - Optional cleanup still available:
apt modernize-sourcesfor the remaining Debian.listfiles; obsolete commented.listbackups (*.list.bak) and legacy/etc/pve/priv/{ipam,macs}.dbcan be removed; old-format RRD files may be regenerated.