PoliNetwork Docs

Monitoring

Monitoring

Alerts

You usually don't have to go looking for problems: Prometheus evaluates alert rules and Alertmanager sends them to our Telegram alert chat. Critical alerts repeat every hour, warnings every 12 hours. Among others, you get alerted when:

  • the node is short on CPU, memory, disk space or inodes;
  • a pod is crash-looping, OOM-killed or not running, or a deployment is missing replicas;
  • a Flux resource, an ExternalSecret or a secret store is not ready;
  • the Cloudflare tunnel has no connections (all public websites are down);
  • a backup hasn't succeeded on schedule.

The rules live in apps/monitoring/prometheus/alert-rules.yml in polinetwork-cd.

Dashboards

  • Flux Web UI: is every app in sync with Git? Why not?
  • Grafana: node and container metrics from Prometheus.
  • Uptime Kuma: is each public website up?

From the command line

SSH into the node, then run

kubectl get pods -A

This will give you a breakdown of all the pods running in the infrastructure. Besides one namespace per app (web, backend, postgres, ...), you'll see:

NamespaceWhat runs there
flux-systemFlux controllers, the Flux Web UI
external-secretsExternal Secrets Operator
cloudflaredThe two Cloudflare Tunnel connectors
local-path-storageThe volume provisioner
kube-systemK3s components: CoreDNS, Traefik, metrics-server, node-exporter

Look at the READY, STATUS and RESTARTS columns: a pod that is not Running, or whose restart count keeps growing, needs a look. Some useful follow-ups:

# Events and the reason of the last restart
kubectl describe pod -n <namespace> <pod-name>

# Logs (add --previous to see the logs of the crashed container)
kubectl logs -n <namespace> deploy/<deployment-name>

# CPU and memory usage per pod
kubectl top pods -A

# Apps that Flux couldn't apply
kubectl get kustomizations -n flux-system

Find an overview of Flux in Flux.