PoliNetwork Docs

Deploy a New App

Introduction

We use k8s to deploy our applications, on our K3s node.
Each app is a folder of plain k8s manifests in the polinetwork-cd repository, tied together with Kustomize. Helm Charts are possible too, through a Flux HelmRelease (see infrastructure/external-secrets for an example).

You don't deploy from your machine: you open a PR, and once it's merged into main, Flux applies it.
Check minikube or K3s to run a k8s cluster on your machine for testing purposes.

CD

To deploy our applications, we use Flux syncing the clusters/k3s folder of the main branch of polinetwork-cd.

For more info, check out Flux.

Procedure

  1. Publish a Docker image with an arm64 variant
  2. Create the app folder
  3. Register the app with Flux
  4. Optionally, publish it on a domain, add secrets, add storage and enable automatic image updates
  5. Open a PR into polinetwork-cd, wait for CI, merge
  6. Monitor the deploy on the Flux Web UI

ARM64 images

The node is ARM64. An image built only for amd64 can't run there, and the pod fails with exec format error. Build a multi-arch image (see the web workflow for an example), or check that the third-party image you use publishes linux/arm64.

Create the app folder

This is the crucial part, so make sure to follow it step-by-step.

  1. Clone polinetwork-cd repo
  2. Create a new branch (named as you like, e.g. mc-server-deploy)
  3. Create this minimal file structure:
🗀 apps
└─ 🗀 <k8s-namespace>
   ├── kustomization.yaml
   ├── namespace.yaml
   └── deployment.yaml

The <k8s-namespace> should be a slug (lowercase words separated by hyphens) identifier of the application/service you want to deploy.

Use the same <k8s-namespace> value for the folder name, the Namespace resource and the namespace field of kustomization.yaml. You don't need to write the namespace in the other manifests: Kustomize adds it.

apps/<k8s-namespace>/kustomization.yaml
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization

namespace: <k8s-namespace>

resources:
  - namespace.yaml
  - deployment.yaml
apps/<k8s-namespace>/namespace.yaml
apiVersion: v1
kind: Namespace
metadata:
  name: <k8s-namespace>
  annotations:
    kustomize.toolkit.fluxcd.io/prune: disabled

The prune: disabled annotation keeps Flux from deleting the namespace, and with it everything inside, if the app is removed from Git by mistake.

apps/<k8s-namespace>/deployment.yaml
apiVersion: v1
kind: Service
metadata:
  name: <app>-service
spec:
  selector:
    app: <app>
  ports:
    - name: http
      port: <port>
      targetPort: <port>
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: <app>
spec:
  replicas: 1
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 0   # keep the old pod until the new one is ready
      maxSurge: 1
  selector:
    matchLabels:
      app: <app>
  template:
    metadata:
      labels:
        app: <app>
    spec:
      containers:
        - name: <app>
          image: ghcr.io/polinetworkorg/<image>:latest
          imagePullPolicy: IfNotPresent
          ports:
            - containerPort: <port>
          readinessProbe:
            tcpSocket:
              port: <port>
            periodSeconds: 5
          livenessProbe:
            tcpSocket:
              port: <port>
            periodSeconds: 10
          resources:
            requests:
              cpu: 10m
              memory: 128Mi
            limits:
              memory: 256Mi

A few rules we follow for every app:

  • Requests and limits. Everything shares one node, so every container declares its memory request and limit. Base the numbers on what the app actually uses (kubectl top pods).
  • Probes. The readiness probe is what makes rolling updates safe: without it, Kubernetes kills the old pod before the new one is serving.
  • Apps with a volume use strategy: Recreate instead (see Add Storage).
  • Apps talk to each other through their Service: http://<service>.<namespace>.svc.cluster.local:<port>.

Register the app with Flux

Flux only deploys the apps listed in clusters/k3s/apps/. Create a Flux Kustomization for the app:

clusters/k3s/apps/<k8s-namespace>.yaml
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
  name: <k8s-namespace>
  namespace: flux-system
spec:
  dependsOn:
    - name: infrastructure-secret-stores
    - name: infrastructure-storage
    - name: infrastructure-traefik
  interval: 30m
  retryInterval: 1m
  timeout: 5m
  prune: true
  wait: true
  sourceRef:
    kind: GitRepository
    name: flux-system
  path: ./apps/<k8s-namespace>

If the app needs another app to be up first (e.g. a database), add it to dependsOn (e.g. - name: postgres).

Then add it to the list:

clusters/k3s/apps/kustomization.yaml
resources:
  - admin.yaml
  ...
  - <k8s-namespace>.yaml

If you want automatic image updates, write a ResourceSet instead of this Kustomization: see below.

Publish it on a domain

Public traffic arrives through the Cloudflare Tunnel to Traefik, which routes it by hostname. Two things are needed.

  1. An Ingress in the app folder (and in kustomization.yaml):
apps/<k8s-namespace>/ingress.yaml
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: <app>
  annotations:
    traefik.ingress.kubernetes.io/router.entrypoints: web
spec:
  ingressClassName: traefik
  rules:
    - host: <hostname>.polinetwork.org
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: <app>-service
                port:
                  number: <port>
  1. A public hostname on the k3s01 tunnel in the Cloudflare Zero Trust dashboard (Networks → Tunnels), pointing to Traefik like the existing hostnames. The tunnel is managed from the dashboard, not from Git, so ask someone in the Direttivo to add it.

Cloudflare terminates HTTPS, so the Ingress only needs the plain web entrypoint: no certificates in the cluster.

Automatic image updates

Flux can redeploy the app every time a new latest image is published, pinning the image digest (see how it works).
This is an optional feature: the alternative is a fixed version or digest in deployment.yaml (e.g. nginx:1.27.4), bumped with a PR.

Consider the following manifest:

      containers:
        - name: nginx
          image: nginx:latest

the image would not update if you don't follow this guide, even though there is latest: the node already has an image called nginx:latest and has no reason to pull it again.

The image must be public on GHCR: neither Flux nor the node use registry credentials. Using the GHCR package ghcr.io/polinetworkorg/<image>:

  1. Add a provider that tracks its latest tag:
infrastructure/image-automation/image-providers.yaml
---
apiVersion: fluxcd.controlplane.io/v1
kind: ResourceSetInputProvider
metadata:
  name: <app>-image
  namespace: flux-system
  annotations:
    fluxcd.controlplane.io/reconcileEvery: "6h"
spec:
  type: OCIArtifactTag
  url: oci://ghcr.io/polinetworkorg/<image>
  filter:
    includeTag: '^latest$'
    limit: 1
  1. Let the GitHub webhook trigger it:
infrastructure/image-automation/receiver.yaml
  resources:
    ...
    - apiVersion: fluxcd.controlplane.io/v1
      kind: ResourceSetInputProvider
      name: <app>-image
  1. Replace clusters/k3s/apps/<k8s-namespace>.yaml with a ResourceSet that renders the same Kustomization, plus an images override with the digest:
clusters/k3s/apps/<k8s-namespace>.yaml
apiVersion: fluxcd.controlplane.io/v1
kind: ResourceSet
metadata:
  name: <k8s-namespace>
  namespace: flux-system
spec:
  inputStrategy:
    name: Permute
  inputsFrom:
    - kind: ResourceSetInputProvider
      name: <app>-image
  resources:
    - apiVersion: kustomize.toolkit.fluxcd.io/v1
      kind: Kustomization
      metadata:
        name: <k8s-namespace>
        namespace: flux-system
        annotations:
          # Never garbage-collect the app if the provider exports no inputs.
          fluxcd.controlplane.io/prune: disabled
      spec:
        dependsOn:
          - name: infrastructure-secret-stores
          - name: infrastructure-storage
          - name: infrastructure-traefik
        interval: 30m
        retryInterval: 1m
        timeout: 5m
        prune: true
        wait: true
        sourceRef:
          kind: GitRepository
          name: flux-system
        path: ./apps/<k8s-namespace>
        images:
          - name: ghcr.io/polinetworkorg/<image>
            newTag: << inputs.<app>_image.tag | quote >>
            digest: << inputs.<app>_image.digest | quote >>

Input names

The inputs are named after the provider, with hyphens turned into underscores: the provider bot-ts-image becomes inputs.bot_ts_image.

In deployment.yaml, keep image: ghcr.io/polinetworkorg/<image>:latest: the name under images must match it exactly, and Flux replaces it with the pinned digest. See clusters/k3s/apps/web.yaml for a complete example.

Options

These are some useful operations on existing apps.

Disable an app

To stop an app temporarily, suspend its Kustomization in the Flux Web UI and scale it down from the node:

kubectl scale deployment/<app> -n <k8s-namespace> --replicas=0

To bring it back, resume the Kustomization: Flux restores the replicas from Git.

Remove an app

  1. Remove <k8s-namespace>.yaml from clusters/k3s/apps/kustomization.yaml and delete the file and the apps/<k8s-namespace>/ folder. Remove its provider and receiver entry too, if it had automatic image updates.
  2. Once merged, Flux deletes the app's resources, except the ones annotated with prune: disabled: the namespace, the PVCs and, for apps with automatic image updates, the generated Kustomization. Delete those by hand from the node once you're sure the data is not needed:
kubectl delete kustomization <k8s-namespace> -n flux-system   # only for ResourceSet apps
kubectl delete namespace <k8s-namespace>

Volume directories stay on disk anyway (reclaimPolicy: Retain), under /srv/<fast|standard>/volumes/<k8s-namespace>/.

The config.json files with "disabled": true you may find at the root of polinetwork-cd belong to the old ArgoCD setup. Flux doesn't read them.