Skip to content

Testing

It has tests. It has an executioner.

A project that emulates a container orchestrator by flattening its output into a different container orchestrator, whose primary architectural document is a fake Lovecraftian grimoire, whose fake kube-apiserver lives on a personal GitHub account for plausible deniability — this project now has automated regression testing. With CI. And a torturer.

And lo, the disciples returned to the temple carrying instruments of verification — not to prove the rituals false, but to ensure they failed in precisely the same way, every time, without exception. The high priest wept, for he understood: reproducible heresy is still heresy, but it is heresy you can ship.

— Necronomicon, On the Formalization of Abomination (apparently)

What it tests

dekube-testsuite compares dekube output between a pinned reference version and the latest release. The test harness uses dekube-manager (rolling from main) as the runner — it's the tool, not the subject. What's compared is core + extension output between versions. Every download retries on network errors (curl --retry 3 --retry-all-errors, curl ≥ 7.71.0), and an exhausted retry never leaves a partial file in the download cache.

Regression

The test runner (run-tests.sh) downloads both versions, generates a dekube.yaml for each, and diffs the output. It runs multiple extension combos:

  1. Core only — no extensions, baseline behavior
  2. Each extension individually — isolation testing
  3. All extensions together — interaction testing

A diff is not a failure. A diff is information. When you bump core from v3.0.1 to v3.1.0 and the caddy service disappears from the output, that's the executioner doing its job — it tells you the refactoring changed behavior, and you decide whether that's intentional.

Static manifests

The manifests/ directory contains edge cases organized by kind:

File What it covers
deployments.yaml Single-container, multi-container, probes, command/args, securityContext, hostNetwork, serviceAccount
statefulsets.yaml volumeClaimTemplates, headless services, updateStrategy, init containers
jobs.yaml Basic jobs, init containers, sidecars, restartPolicy variations
services.yaml ClusterIP, multi-port, ExternalName, ExternalName chain, NodePort, LoadBalancer
ingress.yaml Single/multi-path, TLS, HAProxy annotations (server-ssl, path-rewrite, server-ca)
configmaps-secrets.yaml envFrom, volume mounts, binary data, shared references across deployments
crds.yaml KeycloakRealmImport, Certificate, ClusterIssuer, Issuer, ServiceMonitor, Bundle
edge-cases.yaml Empty docs, 63/64-char names, missing namespace, no selector, empty containers, unknown kinds
bug-<slug>.yaml One regression fixture per fixed bug, resource names prefixed with the slug so fixtures never collide
ext-<name>.yaml The main path of an extension (nginx, traefik, keycloak, servicemonitor) — otherwise only exercised by the bug fixtures

The torturer (--perf N)

The generator (generate.py) produces manifests at O(n³) scale:

  • n releases, each with n Deployments, n ConfigMaps, n Secrets, n Services, 1 Ingress
  • Each Deployment mounts all n ConfigMaps of its release
  • Total volume resolutions: n² deployments × n mounts = n³
n Deployments ConfigMap mounts Approx. time
5 25 125 < 1s
15 225 3,375 seconds
30 900 27,000 notable
50 2,500 125,000 pain

Performance tests run locally, not in CI — public runners aren't meant for this.

Usage

# Full regression
./run-tests.sh

# Override reference core version
./run-tests.sh --core v2.0.0

# Override reference extension version
./run-tests.sh --ext keycloak==v0.1.0

# Torture test
./run-tests.sh --perf 15

# Keep /tmp output for inspection
./run-tests.sh --perf 15 --keep

# Test unreleased core work (latest side only; ref stays pinned)
./run-tests.sh --local-core /path/to/helmfile2compose.py

# Test an unreleased extension file, repeatable per extension
./run-tests.sh --local-ext nginx=/path/to/nginx_rewriter.py

# Latest side: bundled extensions (indexers, workload, haproxy, caddy, emptydir,
# fix-permissions) from main instead of the release
./run-tests.sh --latest-main

What a plain run measures: the reference side is the pinned core release; the latest side is the engine's latest release plus the dependency extensions (keycloak, nginx, cert-manager…) from each repo's main. Right after a rebaseline, the release halves are the same build, so only extension changes on main show up until the next tag. --latest-main also takes the bundled extensions from main (the engine body itself still comes from the latest release — there is no built engine for main); --local-core wins over it.

--local-core and --local-ext only replace the named piece on the latest side — the reference side always comes from the pinned release. --local-ext doesn't follow transitive dependencies: if the extension you're overriding pulls in another one (trust-manager pulling cert-manager), that dependency is still fetched at its latest released tag unless you also pass --local-ext for it.

Reference versions

Edit dekube-known-versions.json to bump the pinned reference:

{
  "reference": {
    "core": "v3.0.1",
    "extensions": {
      "cert-manager": "v0.2.0",
      "keycloak": "v0.3.0",
      "servicemonitor": "v0.2.0",
      "trust-manager": "v0.2.0",
      "nginx": "v0.2.0",
      "traefik": "v0.2.0",
      "flatten-internal-urls": "v0.2.0",
      "bitnami": "v0.2.0"
    },
    "exclude-ext-all": ["flatten-internal-urls"]
  }
}

Extensions listed here are tested individually; unlisted are skipped. Extensions in exclude-ext-all are tested in isolation but excluded from the combined ext-all combo (e.g. due to incompatibilities declared in the registry).

CI

The GitHub Actions workflow runs regression weekly (Monday 6am UTC) and on push to test-related files. Diffs are uploaded as artifacts for human review. There are no assertions — the diff is the output. GitHub disables the weekly trigger after 60 days without repository activity; gh workflow enable regression.yml -R dekubeio/dekube-testsuite turns it back on.

Reading diffs

  • core-only diff → pure dekube-engine behavioral change
  • ext-<name> diff → change in that extension or its interaction with core
  • ext-all diff → interaction between extensions
  • ref run FAILED, latest OK → a crash fixed in latest; no diff available for that combo
  • ref run FAILED, latest FAILED → both sides still crash; check the latest output before assuming it's the same failure

The latest run's output directory is pre-seeded with the reference run's secrets/ before it starts, so idempotent generators (cnpg's superuser password, for instance) produce the same values on both sides instead of manufacturing a spurious diff. *.crt/*.key files are excluded from the diff entirely — key material is random, and cert-manager before v0.5.0 regenerated it on every run, so its content is noise, not drift.

When you see a diff, the question isn't "is this a bug?" — it's "did I mean to change this?" If yes, bump the reference version in dekube-known-versions.json and the diff disappears. If no, you just caught a regression. The suite stays assertion-free by design — it measures drift, it doesn't judge it.