Longhorn volume faulted¶
Use this runbook when the alert fires. Start with Triage (confirm it is real), then Diagnose, then Resolve. Document anything unique to your environment in the app-specific mk_runbook.md next to the HelmRelease if needed.
| Field | Value |
|---|---|
| Alerts | HomelabLonghornVolumeFaulted (app PVCs), HomelabLonghornVolsyncPvcFaulted (volsync-* temp PVCs) |
| PrometheusRule | homelab-storage.yaml |
| Condition | longhorn_volume_robustness{state="faulted"} == 1 |
What this means¶
A Longhorn volume is in the faulted state. Application PVCs (downloaders, media, ai, excluding names matching volsync-*) may indicate real data risk — investigate before delete. VolSync temp PVCs (volsync-*-src clones) are disposable after a failed backup; deleting them does not remove app config data when Restic backups to Garage are healthy.
Triage¶
Triage (first 5 minutes)¶
- Acknowledge the alert (note time,
alertname, namespace/release from ntfy). - Check if something changed recently (Git push, chart bump, node drain, storage outage).
- Confirm the alert is still firing in Prometheus / Grafana (Alerting → Alert rules).
- Decide: transient (wait one reconcile interval) vs sustained (continue below).
# Recent events for the namespace (replace NAMESPACE)
kubectl get events -n NAMESPACE --sort-by='.lastTimestamp' | tail -20
# From alert labels: pvc_namespace, pvc, volume
kubectl get pvc -n <pvc_namespace> <pvc>
kubectl describe pvc -n <pvc_namespace> <pvc>
kubectl get volume.longhorn.io -n longhorn-system | rg <pvc> || true
Determine: app PVC vs VolSync temp PVC (pvc name contains volsync-).
Diagnose¶
# Events on the workload using the PVC
kubectl get pods -n <pvc_namespace> -o wide
kubectl describe pod -n <pvc_namespace> <pod-using-pvc>
# All faulted volsync-related PVCs
kubectl get pvc -A -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,PHASE:.status.phase | rg volsync
# ReplicationSource still synchronizing?
kubectl get replicationsource -A
kubectl describe replicationsource -n <pvc_namespace> <name>
VolSync temp PVC faulted often follows disk not schedulable or a stuck volsync-src job. Fix disk headroom first (disk schedulability runbook).
App PVC faulted may need Longhorn UI: salvage, replica rebuild, or restore from VolSync/Garage — do not delete the app PVC without a backup/restore plan.
Resolve¶
VolSync temp PVC (HomelabLonghornVolsyncPvcFaulted)¶
- Pause the matching
ReplicationSource(spec.paused: true). - Delete the failed Job:
kubectl delete job -n <ns> -l ...or by namevolsync-src-*. - Delete the faulted PVC(s) matching
volsync-*-src(and other idlevolsync-*clones if documented in your chart notes — not the main appconfigclaim). - Unpause
ReplicationSource; operator recreates the next backup job.
Deleting only the Job without pausing can cause the operator to recreate work while Synchronizing=True.
Application PVC (HomelabLonghornVolumeFaulted)¶
- Longhorn UI → Volume → check replicas, snapshots, last error.
- If the workload is down, check whether a detach/reattach cycle or node issue caused fault.
- Restore path: TrueCharts dest
restore-once+ Garage if reinstalling; do not assume faulted app PVC is safe to delete. - GitOps fixes (replica count, storage class) go through
helm-release.yamlfor the app and Longhorn — neverkubectl applylive hotfixes you cannot commit.
Escalation¶
Escalation / close-out¶
- Alert resolved in Alertmanager (or silenced with a documented reason and expiry).
- Root cause noted (link PR/commit if GitOps change).
- Update this runbook or the app
mk_runbook.mdif you learned something new.
If the issue is upstream (TrueCharts chart bug, Flux bug), capture logs and open an issue; avoid permanent silences without a ticket.
Pair with VolSync missed backup when faulted clones coincide with failed volsync-src jobs.