K3s Node
The k3s01 VM, how Ansible configures it, and how to run the playbooks.
Overview
Production runs on a single VM, k3s01, created by Terraform
(environments/k3s) and configured by the Ansible playbooks in
polinetwork-cd/ansible.
| Size | Standard_E2ps_v6 (ARM64), West Europe, zone 1 |
| OS | Debian 13 ARM64 |
| Private IP | 10.43.1.4 (the public IP is outbound-only) |
| Kubernetes | K3s, single server node, pinned version (see ansible/vars/main.yml) |
| Admin user | pnadmin |
ARM64
The node is ARM64. Every image we deploy must have a linux/arm64 variant,
otherwise the pod fails with exec format error. Our GitHub workflows publish
multi-arch images (amd64 + arm64) for this reason.
Disks
| Mount | Azure disk | Type | Used for |
|---|---|---|---|
/ | disk-k3s01-os | Standard SSD, 64 GB | Operating system |
/srv/fast | disk-k3s-fast (LUN 0) | Premium SSD v2, 64 GB | Volumes of the fast storage class (databases) |
/srv/standard | disk-k3s-standard (LUN 1) | Standard SSD, 128 GB | K3s data (/srv/standard/k3s), volumes of the standard storage class, local backup staging |
Ansible finds the data disks by Azure LUN, refuses to touch a disk that is
smaller than expected or holds an unexpected filesystem, formats them only when
blank and mounts them by UUID. Both disks have prevent_destroy in Terraform.
See Add Storage for how volumes are carved out of these disks.
What Ansible sets up
The provisioning playbook (ansible/playbooks/provision.yml) applies these roles, in order:
| Role | What it does |
|---|---|
base | Base packages, UTC timezone, pnadmin user and the k3s-admin group, kernel modules and sysctls for Kubernetes |
security | SSH hardening (keys and Cloudflare Access certificates only), Debian security updates via unattended-upgrades (no automatic reboot), nftables host firewall |
storage | Formats and mounts the data disks, creates /srv/*/volumes |
k3s | Installs the pinned K3s binary (checked by SHA-256) and its configuration, waits for the node to be Ready |
backup | Installs the encrypted backup scripts and their systemd timers (see Backups) |
flux | Installs the pinned Flux Operator and the FluxInstance that syncs clusters/k3s from the main branch |
After that, Flux takes over everything inside the cluster (see Flux).
Notable K3s settings
- Pod and Service CIDRs are
10.52.0.0/16and10.53.0.0/16. The K3s default10.43.0.0/16would collide with the Azure VNet. - ServiceLB and the bundled local-storage are disabled. Storage comes from the local-path provisioner managed by Flux.
- The bundled Traefik stays as the ingress controller, but Flux turns its
Service into
ClusterIP: onlycloudflaredreaches it. - Secrets encryption at rest is enabled.
- The API server signs service-account tokens for a public OIDC issuer
(
pnk3soidcstorage account). Azure trusts that issuer, which is how External Secrets reads Key Vault without any stored credential.
Security guardrails
- The VM's public IP is outbound-only: the NSG denies all inbound traffic. The
host firewall also drops everything except the node itself, cluster traffic,
and SSH from private ranges (WARP sessions arrive from the
cloudflaredpods). - Pods cannot reach the Azure Instance Metadata Service (
169.254.169.254), so they can't borrow the VM's managed identities. - An admission policy rejects pods using
hostNetwork,hostPIDorhostIPCoutsidekube-system, and another one forbids pods from running as thekeyvault-readerservice account.
Running the playbooks
You rarely need this: day-to-day changes go through Flux. Run Ansible when you
change something under ansible/ (e.g. a K3s upgrade or a new backup target).
Only the IT Lead group has access to the node.
The playbooks run from your machine against 10.43.1.4 over SSH, so connect
WARP first.
cd ansible
ansible-galaxy collection install -r requirements.yml
./scripts/provision-and-verify.shThe script is the only supported way to apply changes. It runs the playbook in
check mode, then applies it twice and fails if the second run still changes
something (the playbook must be idempotent). Finally it runs
playbooks/verify.yml, which checks the disks, the K3s version, the firewall,
the SSH configuration, that Flux is ready, that pods can't reach Azure IMDS, and
it takes and validates a fresh backup. Logs are saved in ansible/artifacts/
(ignored by Git).
To run from inside the VM (for example through Azure Run Command), put a reviewed checkout on the VM and run:
./scripts/provision-and-verify.sh inventories/local/hosts.ymlNever copy a private SSH key or a secret value into the repository or into Azure Run Command parameters.
CI
The Ansible workflow in polinetwork-cd runs on every pull request that
touches ansible/, clusters/k3s/, infrastructure/, apps/ or tests/. It
runs ansible-lint, shellcheck, a playbook syntax check, kustomize build on
every Flux path and the manifest tests in tests/. It does not touch the VM.
Upgrading K3s
- Pick a release from K3s releases and
download its
k3s-arm64binary checksum. - Update
k3s_version,k3s_release_urlandk3s_sha256inansible/vars/main.yml. - Open a PR, merge it, then run
./scripts/provision-and-verify.sh. K3s restarts with the new binary; workloads come back on their own.
The Flux Operator is upgraded the same way (flux_operator_* and
flux_distribution_version).
Rebuilding the node
The workload identity trust relies on the cluster's service-account signing key,
whose public part is committed in Terraform
(environments/k3s/k3s-service-account-jwks.json). If the node is rebuilt from
scratch without restoring the control-plane backup, the key changes and that file
must be updated, otherwise External Secrets can't read Key Vault.
PoliNetwork Docs