Testing¶
It has tests. It has an executioner.
A project that emulates a container orchestrator by flattening its output into a different container orchestrator, whose primary architectural document is a fake Lovecraftian grimoire, whose fake kube-apiserver lives on a personal GitHub account for plausible deniability — this project now has automated regression testing. With CI. And a torturer.
And lo, the disciples returned to the temple carrying instruments of verification — not to prove the rituals false, but to ensure they failed in precisely the same way, every time, without exception. The high priest wept, for he understood: reproducible heresy is still heresy, but it is heresy you can ship.
— Necronomicon, On the Formalization of Abomination (apparently)
What it tests¶
dekube-testsuite compares dekube output between a pinned reference version and the latest release. The test harness uses dekube-manager (rolling from main) as the runner — it's the tool, not the subject. What's compared is core + extension output between versions. Every download retries on network errors (curl --retry 3 --retry-all-errors, curl ≥ 7.71.0), and an exhausted retry never leaves a partial file in the download cache.
Regression¶
The test runner (run-tests.sh) downloads both versions, generates a dekube.yaml for each, and diffs the output. It runs multiple extension combos:
- Core only — no extensions, baseline behavior
- Each extension individually — isolation testing
- All extensions together — interaction testing
A diff is not a failure. A diff is information. When you bump core from v3.0.1 to v3.1.0 and the caddy service disappears from the output, that's the executioner doing its job — it tells you the refactoring changed behavior, and you decide whether that's intentional.
Static manifests¶
The manifests/ directory contains edge cases organized by kind:
| File | What it covers |
|---|---|
deployments.yaml |
Single-container, multi-container, probes, command/args, securityContext, hostNetwork, serviceAccount |
statefulsets.yaml |
volumeClaimTemplates, headless services, updateStrategy, init containers |
jobs.yaml |
Basic jobs, init containers, sidecars, restartPolicy variations |
services.yaml |
ClusterIP, multi-port, ExternalName, ExternalName chain, NodePort, LoadBalancer |
ingress.yaml |
Single/multi-path, TLS, HAProxy annotations (server-ssl, path-rewrite, server-ca) |
configmaps-secrets.yaml |
envFrom, volume mounts, binary data, shared references across deployments |
crds.yaml |
KeycloakRealmImport, Certificate, ClusterIssuer, Issuer, ServiceMonitor, Bundle |
edge-cases.yaml |
Empty docs, 63/64-char names, missing namespace, no selector, empty containers, unknown kinds |
bug-<slug>.yaml |
One regression fixture per fixed bug, resource names prefixed with the slug so fixtures never collide |
ext-<name>.yaml |
The main path of an extension (nginx, traefik, keycloak, servicemonitor) — otherwise only exercised by the bug fixtures |
The torturer (--perf N)¶
The generator (generate.py) produces manifests at O(n³) scale:
nreleases, each withnDeployments,nConfigMaps,nSecrets,nServices, 1 Ingress- Each Deployment mounts all
nConfigMaps of its release - Total volume resolutions: n² deployments × n mounts = n³
| n | Deployments | ConfigMap mounts | Approx. time |
|---|---|---|---|
| 5 | 25 | 125 | < 1s |
| 15 | 225 | 3,375 | seconds |
| 30 | 900 | 27,000 | notable |
| 50 | 2,500 | 125,000 | pain |
Performance tests run locally, not in CI — public runners aren't meant for this.
Usage¶
# Full regression
./run-tests.sh
# Override reference core version
./run-tests.sh --core v2.0.0
# Override reference extension version
./run-tests.sh --ext keycloak==v0.1.0
# Torture test
./run-tests.sh --perf 15
# Keep /tmp output for inspection
./run-tests.sh --perf 15 --keep
# Test unreleased core work (latest side only; ref stays pinned)
./run-tests.sh --local-core /path/to/helmfile2compose.py
# Test an unreleased extension file, repeatable per extension
./run-tests.sh --local-ext nginx=/path/to/nginx_rewriter.py
# Latest side: bundled extensions (indexers, workload, haproxy, caddy, emptydir,
# fix-permissions) from main instead of the release
./run-tests.sh --latest-main
What a plain run measures: the reference side is the pinned core release; the latest side is the engine's latest release plus the dependency extensions (keycloak, nginx, cert-manager…) from each repo's main. Right after a rebaseline, the release halves are the same build, so only extension changes on main show up until the next tag. --latest-main also takes the bundled extensions from main (the engine body itself still comes from the latest release — there is no built engine for main); --local-core wins over it.
--local-core and --local-ext only replace the named piece on the latest side — the reference side always comes from the pinned release. --local-ext doesn't follow transitive dependencies: if the extension you're overriding pulls in another one (trust-manager pulling cert-manager), that dependency is still fetched at its latest released tag unless you also pass --local-ext for it.
Reference versions¶
Edit dekube-known-versions.json to bump the pinned reference:
{
"reference": {
"core": "v3.0.1",
"extensions": {
"cert-manager": "v0.2.0",
"keycloak": "v0.3.0",
"servicemonitor": "v0.2.0",
"trust-manager": "v0.2.0",
"nginx": "v0.2.0",
"traefik": "v0.2.0",
"flatten-internal-urls": "v0.2.0",
"bitnami": "v0.2.0"
},
"exclude-ext-all": ["flatten-internal-urls"]
}
}
Extensions listed here are tested individually; unlisted are skipped. Extensions in exclude-ext-all are tested in isolation but excluded from the combined ext-all combo (e.g. due to incompatibilities declared in the registry).
CI¶
The GitHub Actions workflow runs regression weekly (Monday 6am UTC) and on push to test-related files. Diffs are uploaded as artifacts for human review. There are no assertions — the diff is the output. GitHub disables the weekly trigger after 60 days without repository activity; gh workflow enable regression.yml -R dekubeio/dekube-testsuite turns it back on.
Reading diffs¶
core-onlydiff → pure dekube-engine behavioral changeext-<name>diff → change in that extension or its interaction with coreext-alldiff → interaction between extensionsref run FAILED, latest OK→ a crash fixed in latest; no diff available for that comboref run FAILED, latest FAILED→ both sides still crash; check the latest output before assuming it's the same failure
The latest run's output directory is pre-seeded with the reference run's secrets/ before it starts, so idempotent generators (cnpg's superuser password, for instance) produce the same values on both sides instead of manufacturing a spurious diff. *.crt/*.key files are excluded from the diff entirely — key material is random, and cert-manager before v0.5.0 regenerated it on every run, so its content is noise, not drift.
When you see a diff, the question isn't "is this a bug?" — it's "did I mean to change this?" If yes, bump the reference version in dekube-known-versions.json and the diff disappears. If no, you just caught a regression. The suite stays assertion-free by design — it measures drift, it doesn't judge it.