Project: diag-pipe — operator-gated host diagnostics
Status: in use (2026-10). Same trust shape as the docs mirror: the agent works in sanitized space, the operator machine is the trust boundary, both directions are gated on human review. Live shakedown completed Oct 3 — see “Rollout”.
Summary
A job pipeline that lets the agent in devvm get host-side diagnostics run during live troubleshooting without the operator hand-copying commands and output both ways. The agent writes job files on the guest; the operator machine pulls them, the operator approves or denies each one, approved jobs execute on the host over the existing ssh path, output is sanitized by the docmirror redaction engine, the operator reviews the redacted output, and the result is pushed back to the guest. Two human gates — script approval and output release — are the security. Automation covers only transport, bookkeeping, and redaction.
This is for attended use: the operator is present and watching during the investigation. It removes ferrying friction from multi-hour incident work, not oversight.
Motivation
The Storage Box EINVAL incident needed seven+ rounds of host
diagnostics (diag1, diag2, output3–output7, post-restart
checks). Every round was: agent drafts probe → operator retypes/pastes
it on the host as root → operator sanitizes and carries output back →
agent reads it. Each round cost minutes of operator attention on pure
transport, and the interesting signal kept arriving one ferry-crossing
late. The commands were mostly variations on four shapes (open-probe,
lsof, hash/stat, journalctl window) — exactly the kind of
repetition a template tier handles.
Trust model
devvm (agent) operator machine «HOSTNAME» (hypervisor)
~/diag-jobs/ ──pull──▶ ~/…/diagpipe/inbox/ (untrusted input)
review → approve/deny
~/…/approved/ ──ssh──▶ timeout bash -s
~/…/raw/ ◀──ssh── stdout/stderr (operator-only)
redact sanitize
~/…/redacted/ (operator reviews)
~/…/outbound/ ──scp──▶ ~/diag-results/
- devvm is the untrusted side. It sits behind the vnet1 isolation (see dev-vm §isolation, decisions); nothing it produces is executed without review and nothing reaches it that hasn’t been sanitized and released.
- The operator machine is the trust boundary — same role it plays for the docs mirror. It already holds the redaction maps and engine; diag-pipe reuses them rather than duplicating them.
- The hypervisor is the execution target, reached over the same
ssh alias the operator already uses (
ProxyJumptarget for the guest). No new daemon, no new credentials, no new network path — the only thing that crosses is an approved script on stdin and its output back. - Raw output and unredacted scripts never leave the operator machine. The guest sees only sanitized, released artifacts — identical policy to the docs mirror.
Rejected alternative: a virtiofs exchange dir on devvm (precedent:
storageboxmapping) with a host-side watcher doing the pull. It would
decouple staging from the operator machine being up, but it puts
executor machinery and the secrets maps on the hypervisor, adds a VM
config + guest fstab change (with the nofail/automount lesson from
the incident), and approvals still need the operator at a terminal
anyway. Revisit only if unattended staging becomes worth it.
Job lifecycle
States live as directories under the work dir; a job moves exactly one direction:
guest ──pull──▶ inbox ──approve──▶ approved ──run──▶ raw ──redact──▶ redacted ──release──▶ guest
│
└──deny──▶ outbound <id>.denied ──release──▶ guest
- pull — fetch new job files from
~/diag-jobs/on the guest (per-filecat/rmover ssh — mailbox semantics). Validates name, size (≤64 KiB), required headers; malformed jobs go straight to denied with a reason. - approve / deny — the first human gate.
reviewshows the job plus the rendered form (see below) and a sha256;approvebinds approval to that exact rendered file, staged operator-side — the guest cannot modify what was approved (TOCTOU-safe by construction).denyqueues a reason file that rides back to the guest on the next release so the agent sees why, not silence. - run — executes the staged file on the host inside a transient
systemd-rununit (--wait --pipe,RuntimeMaxSec=TIMEOUT+30):env -i+ pinnedPATH+timeout --kill-afterinside. The cgroup teardown bounds daemonized/setsidescapees that a plaintimeoutcan’t reach; a plain-timeoutrunner remains as fallback for non-systemd targets. Output capped (default 1 MiB, truncation flagged), exit status recorded. - redact — output through the sanitize filter
→
redacted/<id>.txtwith a small metadata header (job id, exit code, truncation flag). - release — the second human gate. Operator reads the redacted
file, then
releasestages it intooutbound/and pushes all pending outbound files to~/diag-results/on the guest.
Every transition appends a JSON line to audit.log (timestamp, job
id, event, sha256, exit code, deny reason). Raw outputs and rendered
scripts are retained operator-side — the audit trail is the incident
evidence bundle the postmortem had to assemble by hand.
Job formats
The agent drops files in ~/diag-jobs/; convention and skeleton are
provided by a guest-side diag-submit helper.
<id>.sh — full script (always needs explicit approval):
#!/bin/sh
# intent: <one line — what question this answers>
# scope: <what it touches — paths, VMIDs, units>
set -u
…<id>.tpl — template call (JSON; auto-approved when the template
and its args validate — see below):
{"template": "probe-path", "intent": "is the path still EINVAL",
"args": ["/mnt/pve/storagebox/<dir>/…/<file>"]}Both directions use the redaction engine symmetrically:
- Outbound (scripts): the agent writes
««TOKENS»;approveruns the job through the unredact filter and the rendered file is what the operator reviews and what runs — same trick assync-in.sh, so scripts can reference real paths and hostnames the agent never learns. Approval binds to the unredacted bytes. - Inbound (output): the sanitize filter,
secrets→tokens plus
««AUTO_*»minting for stray IPs/emails/MACs.
The auto-map should be shared with docmirror so a
machine that appears as ««AUTO_IPV4_2» in the docs is the same token
in diag output — the agent can correlate artifacts across sources.
The catch-all allowlist should be shared too (diag ALLOWLIST
config → docmirror’s allowlist.txt): values that are documented
anyway — internal LAN ranges — stay literal in both places.
Templates — where the friction reduction actually comes from
Per-script review of novel commands stays manual — that’s the gate. The incident showed most diagnostic work is a few shapes repeated, so pre-approved parameterized templates carry the bulk of traffic:
- Templates live operator-side (
templates/*.sh), each declaring an arg count and a per-arg policy regex in a header block (# args: 1,# arg1: ^/[-A-Za-z0-9._+/]+$,# desc: …). - Args are passed as positional parameters (
bash -s -- arg…), never interpolated into script text — the policy regex bounds the value space, and quoting can’t escape into code. - A valid
.tplcall auto-approves at pull time; anything else waits fordiag approve. Output review stays human for all jobs — templates skip the script gate only, because a read-only command can still print something identifying. - Starter set derived from the incident’s actual probe shapes:
probe-path(stat/magic-bytes/bounded open test),open-fds(lsofon a mount/path),journal-window,dmesg-tail,vm-status(qm/pctread-only).
Templates beat a command allowlist: they bound the arguments too,
and “read-only” is a fuzzier property than it looks (cat on an
arbitrary path is arbitrary read — the output gate catches it, but
keeping the auto tier narrow keeps the reviewer’s job easy).
Security details that make or break it
- Execute the staged copy. The rendered
.runfile is written at approve time on the operator machine; the guest cannot alter it afterward. - Approvals bind to bytes, not job names.
approverecords sha256 of payload + argv;runre-hashes and refuses on drift.redactrecords the redacted file’s sha256;releasere-hashes and refuses if the file changed after review. So the audit chain proves reviewed-bytes executed-bytes released-bytes. - Review what runs.
diag reviewshows the rendered (unredacted) form — the actual bytes that will execute — not the tokenized submission. Heuristic warnings flagrm,dd,| sh,qm/pctmutations, network calls, etc. — informational, not blocking. - Fail closed everywhere. Any step error (ssh failure, redact crash, non-UTF-8 output) leaves the job held — nothing is released on a partial pipeline.
- Output hygiene before the filter: byte cap, UTF-8 decode with replacement, so binary garbage can’t smuggle or crash redaction.
- Constrained runner:
env -iwith a pinnedPATH,timeout, one job at a time, raw fileschmod 600. - Deny returns a reason. The
.deniedfile pushed back tells the agent the constraint it hit, so it resubmits smarter instead of blind.
Known gaps / honest limits
- Redaction is assist, not boundary. The catch-all patterns only cover IPv4/IPv6/email/MAC — hostnames, usernames, serials, UUIDs, tailnet device names sail through unless they’re in the secrets map. The output review gate exists precisely because this list can never be proven complete.
- Rubber-stamping is the failure mode. Friction reduction erodes
review quality if jobs get big or vague. Mitigations are social,
not technical: required
# intent:/# scope:headers, small-job convention, deny-with-reason feedback. - A benign-looking approved script can still probe for secrets not in the map — output review is the only backstop. That is a deliberate trade: the same human gate exists today, just with worse ergonomics.
- Approval is only as fresh as the context. The tool shows the
rendered script; it cannot judge whether running
lsofright now is disruptive. Still on the reviewer.
Operator workflow (steady state)
diag pull # fetch new jobs from the guest
diag list # live queue only (-a includes history)
diag review <id> # rendered script + warnings + hash
diag approve <id> # or: diag deny <id> -m "reason" (-m required)
diag run <id> # execute on «HOSTNAME», capture raw
diag redact <id> # sanitize into redacted/<id>.txt
diag view <id> # page the redacted output (--all for history)
diag release <id> # push to guest + flush pending denials<id> accepts a full id, a unique prefix, or the slug after the
timestamp (diag run probe-path); bash completion is in
completions/diag.bash. Finished jobs (released/rejected — artifacts
linger as the audit trail) don’t resolve for pipeline commands;
diag view --all rereads them.
Rollout
- Tooling repo mirrored to gitea (bare repo
~/mirrors/diag-pipe.gitexists on devvm; gitea ferry still pending — same pattern asinfra-docs-mirror). - Operator machine: clone tooling,
~/.config/diagpipe/config.env(ssh aliases, work dir,REDACTpath). - Extend the secrets map for the diag domain — Oct-3 shakedown
confirmed
««HOSTNAME»/««SB_HOST»/««SB_USER»round-trips; MACs are now engine-covered (««AUTO_MAC_n»); unmapped values seen so far (VM UUIDs, CIFS ClientGUID) are low-sensitivity and stay on the review gate. - Guest:
~/diag-jobs/+~/diag-results/dirs,diag-submithelper, convention notes in agent rules. -
./selftest.sh— full local round-trip with fake secrets. - Live shakedown (Oct 3): template calls + script + deny path
end-to-end; surfaced and fixed real bugs — inline-comment config
parsing, symlinked-
diagtemplate resolution, held-vs-denied tpl jobs, reasonless denials, finished jobs resolving/completing as live work, missing--allowlistwiring.
Open questions
- Auto-release tier for structurally-safe templates (e.g.
vm-statusoutput can’t contain secrets)? — deferred; keep gate 2 fully manual until redaction coverage is proven against real output. - Multiple execution targets (a template that runs inside a guest via
pct exec/qm guest exec)? — deferred; boundedpct exectemplates possible later, but arbitrary inner commands are the risky shape this design exists to bound. - virtiofs exchange variant — rejected for now (see Trust model).