Operations
8 pages in this section.
- Getting into a guest The serial console, the generated credentials nobody told you about, and how SSH keys reach a guest — including the part of that which has no CLI.
- Restoring a guest Rolling back to a snapshot, cloning one into a new guest, and bringing a replicated copy to life on the host that holds it — plus which of those the CLI can actually do.
- Disaster recovery DR Manager — an off-fleet appliance holding encrypted restic repositories that stay closed until a host knocks. Deploying it, enabling DR on a guest, and restoring from a capability URL.
- The scheduler What is scheduled, what ran, what failed and why — plus running a job now without waiting for its cron. The operational half of snapshots, replication and DR.
- Monitoring and alerts Prometheus scrape config and 80-odd alert rules generated from the node itself, guests discovered automatically, and the one flag on a guest that decides whether it is watched at all.
- Troubleshooting a node Inspecting the host itself, regenerating the configuration files everything else depends on, clearing the stale locks and interfaces a crash leaves behind, and talking to QEMU directly.
- LeilFS Shared storage across the cluster, for the workloads ZFS replication does not suit — what Hoster provisions for it, what it monitors, and the honest boundary between the two.
- Building your own template A template is a ZFS dataset with a raw disk in it — nothing more. Making one by hand, what the image has to support, publishing a catalogue of your own, and keeping templates in step across a fleet.