Skip to content

Repository files navigation

Labs64.IO Ecosystem

Labs64.IO :: Tests

Regression Suite License: Apache 2.0 📖 Documentation

Integration & Regression Test Suite for the Labs64.IO Ecosystem.

Overview

Black-box API-edge regression suite for the Labs64.IO platform, built with Robot Framework and robotframework-requests.

Every test exercises the gateway edge (Traefik + authproxy), never a backend port directly. Explicitly tagged local-k8s-only probes may additionally corroborate an otherwise unobservable auth or cross-service delivery effect through correlation-scoped pod logs; they never inspect RabbitMQ or a database. See AGENTS.md for the narrow exception rules.

Covers auditflow and payment-gateway today. See AGENTS.md for how to extend this to another module.

Repository Structure

labs64.io-tests/
├── requirements.txt                # Python dependencies
├── resources/                      # Shared Robot Framework resource files
│   ├── common.resource             # HTTP sessions + mock/Keycloak exact-scope token minting
│   ├── auditflow.resource          # AuditFlow-specific keywords (POST /audit/publish)
│   └── payment_gateway.resource    # Payment Gateway-specific keywords
├── tests/
│   ├── common/
│   │   └── e2e/
│   │       └── payment_auditflow.robot # PG -> AuditFlow local-k8s delivery probe
│   ├── auditflow/
│   │   ├── smoke.robot             # happy path + 400 validation
│   │   └── authz.robot             # auth/authz matrix — see P0 Defect Coverage below
│   └── payment-gateway/
│       ├── smoke.robot
│       ├── payment_providers.robot # create/read/update/delete lifecycle (noop PSP)
│       └── authz.robot             # auth/authz scope matrix
├── scripts/
│   ├── generate_auth_enforcement_suite.py  # generates tests/common/auth_enforcement.robot
│   ├── mirror_edge_images.sh       # pull :edge images into the local k3d registry
│   ├── robot_summary.py            # output.xml -> job summary (totals + failures by cause)
│   └── wait_for_pods.sh            # CI readiness gate; fails fast on unrecoverable pods
└── .github/workflows/
    └── labs64io-regression-suite.yml        # GitHub Actions CI workflow

Tag Taxonomy

Tag Meaning Runs where
smoke Fast critical-path only Every PR
regression Full functional coverage per service Nightly + pre-release
contract Mirrors a path covered by Schemathesis Informational
e2e Cross-service flows Targeted / pre-release
critical Failure blocks a release Always gating

| flaky | Quarantined — non-blocking | Nightly, excluded from gating | | known-bug | Documents a known, unresolved defect | Excluded from smoke and regression; target it directly via test-file/test-case to check on it | | not-ga | Targets a module whose images have never been published | Always excluded, kept for when it ships | | auth | Authentication / authorisation assertions | — | | tenant-isolation | Cross-tenant / cross-scope isolation scenarios | — | | error-handling | Error path / negative testing | — | | psp-stub | Requires the external PSP HTTP stub and a matching provider endpoint override | Targeted locally + nightly |

contract, critical, and flaky are reserved for future use — no test currently carries them, and that's not drift. contract is earmarked for tests that mirror a path Schemathesis already covers (informational, not gating); critical for a case whose failure should always block a release regardless of suite; flaky is for quarantining a genuinely flaky case without deleting its coverage. The e2e tag is active for cross-service flows spanning more than one module; environment-specific probes also carry a tag such as local-k8s-only.

not-ga is currently carried — by every Checkout case, generated and hand-written. No released Checkout image exists (labs64/checkout, labs64/checkout-ui have only ever published :edge), so charts/labs64io-ecosystem defaults checkout.enabled to false and the PR gate's install.sh stack never serves it; a call there would be a 503 no available server, not a pass/fail signal about the module itself. smoke/p0/regression all --exclude not-ga so that structural fact doesn't read as a permanent regression.

The nightly job is now the exception: it deploys Checkout's :edge image, so those cases would actually run there. Dropping the tag is therefore a judgement call per job, not a blocked one — remove it from a module's cases (and the not_ga=True module entry in scripts/generate_auth_enforcement_suite.py, for the generated suite) once its suites are confirmed green, and unconditionally the day its chart flips enabled to true by default.

Setup

python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -r requirements.txt

You need a running Labs64.IO stack reachable through its gateway edge — either the local k3d cluster or another Deployment Mode. Local runs support both identity profiles: lightweight mock-oidc and production-like Keycloak. Both mint exact-scope M2M tokens for the same auth/authz matrix, so switching provider does not weaken or duplicate the assertions.

Running Tests

Fastest path: just (see justfile — just smoke, just regression, just test-all, just test-module auditflow, just log, etc.; just --list for the full set). It wraps venv setup and the robot invocations below, writing output to results/. The rest of this section shows the underlying robot commands directly, for when you need a variation the justfile doesn't cover.

From labs64.io-workspace, use the native profile commands against an already-running stack:

just mock smoke
just mock regression
just mock test
just keycloak smoke
just keycloak auth
just keycloak regression
just keycloak test

Profile reports are isolated under results/mock/ and results/keycloak/.

All smoke tests (fast, every PR):

robot --include smoke --exclude not-ga tests/

A single service:

robot tests/auditflow/
robot tests/payment-gateway/

A single file or test case:

robot tests/auditflow/authz.robot
robot --test "Publish With Correct Scope Is Allowed" tests/auditflow/authz.robot

P0 blocker tests only (never skipped):

robot --include smoke --exclude not-ga tests/

Full regression, excluding flaky and not-yet-GA modules:

robot --include regression --exclude flaky --exclude not-ga --exclude psp-stub tests/

Auth/authz matrix only, across all services:

robot --include auth tests/

Robot writes output.xml, log.html, and report.html to the current directory (or --outputdir <dir>) on every run — open log.html first when a test fails, it has the full request/response detail per keyword.

PSP stub tests

The ordinary local regression assumes a normal Payment Gateway deployment and excludes psp-stub cases:

cd ../labs64.io-helm-charts
just up

cd ../labs64.io-tests
just test

To switch the existing local Payment Gateway deployment to the host-side WireMock endpoint, run the targeted suite, and then restore the normal provider endpoint:

just test-up
just test-status
just test-psp
just test-down

To run the complete local gate in one command, including both the normal regression and the PSP-stub scenarios, use:

just test-all

It runs the ordinary regression before changing PG, enables the stub only for the PSP phase, and restores the normal deployment even when setup or tests fail. Phase reports are written to results/regression/ and results/psp/; the merged report remains at results/report.html. A final summary repeats every phase result and report path at the bottom of the terminal output, including NOT RUN for a phase that setup prevented from starting.

test-up does not rebuild an image or recreate the cluster. It starts WireMock and applies a small Helm override that rolls only Payment Gateway. test-down reapplies the ordinary local values before stopping WireMock. The Robot tests are not Kubernetes-specific: just test-psp can target any already-prepared environment through the usual base URL variables.

Provider-specific Robot scenarios and WireMock mappings remain in the provider module repository; this repository owns only the reusable WireMock process and test-environment lifecycle. test-up probes a host address from the k3d node itself and restricts PG egress to that single /32 plus WireMock port 8090, so it does not depend on a particular k3d hosts-file layout and works in Docker Desktop/devcontainers and Linux CI.

Targeting a different environment

Base URLs and the identity provider are resolved from environment variables (see resources/common.resource for the full list and defaults):

GATEWAY_BASE_URL=https://gateway.example.com \
MOCK_OIDC_BASE_URL=https://mock-oidc.example.com \
robot --include smoke tests/

For the local Keycloak profile, set IDENTITY_PROVIDER=keycloak; the default transport reuses GATEWAY_BASE_URL with Host: keycloak.localhost, which also works from the workspace devcontainer. Override KEYCLOAK_BASE_URL, KEYCLOAK_ROUTE_HOST, KEYCLOAK_REALM, OIDC_CLIENT_ID, and OIDC_CLIENT_SECRET when targeting another environment. Its realm must provide the same exact-scope/tenant-persona contract as the local declarative realm.

CI

The GitHub Actions workflow (.github/workflows/labs64io-regression-suite.yml) runs three jobs:

Job Trigger Target it provisions
Static Checks every PR, ~1 min none — no cluster needed
Smoke every PR ephemeral k3d + bash install.sh install (the published ecosystem chart, same pattern as labs64io-published-chart-e2e.yml in labs64.io-helm-charts)
Full Regression nightly, release, workflow_dispatch k3d + local registry + Helmfile (just up), running each module's :edge image

Static Checks is the fastest way to find a broken test. It runs scripts/generate_auth_enforcement_suite.py --check (every operation declaring x-labs64-auth has a generated edge test) and robot --dryrun over every suite, which resolves each keyword and Resource import without sending a request — so a test calling a keyword that doesn't exist fails in seconds instead of surfacing an hour later as a test failure indistinguishable from a real regression. Run the same check locally with just dryrun.

Full Regression deliberately provisions differently from the PR gate. Its item-12 suites (pipeline routing, condition operators, redaction, DLQ replay, quota enforcement, secretRef resolution) assert against fixtures that exist only in labs64.io-helm-charts' overrides/auditflow/values.local.yaml — the t_regression tenant and its deliberately failing probe pipelines, t_regression_quota's tiny quota, and the global redaction rule. install.sh's quickstart profile provisions only a t_mock demo tenant, so on that path every one of those cases fails 403 TENANT_NOT_PROVISIONED no matter how healthy the code is. The Helmfile path is also the only one that applies each module's local override (central Cerbos PDP address, env secretRef resolver, gateway routes), and — because it creates the k3d-labs64io context — the only one where the local-k8s-only cases actually execute instead of self-skipping.

What nightly runs against

Every module publishes <image>:edge to Docker Hub after a green master build — the edge-image job in each repo's CI, via labs64.io-workspace's reusable docker-publish.yml. Nightly mirrors those tags into the k3d registry (scripts/mirror_edge_images.sh) as the localhost:5005/<name>:latest references the Helmfile local overrides pin, and deploys them. Nothing is built here: nightly exists to find integration failures between modules, and rebuilding eight images to do that would re-run work each module's CI already did while adding several more ways to go red for reasons that are not a regression.

The mirror step writes every image's source digest to the job summary, so a red nightly says exactly which builds it ran against — :edge moves on every master push, so timestamps alone would not.

If a module's :edge tag is missing, the step fails listing every missing image together with the workflow that publishes it, rather than deploying a half-mirrored stack of mixed vintages.

Every job excludes not-ga; see the not-ga entry in Tag Taxonomy. Note that Checkout and Customer Portal — which have never had a released image — now publish :edge too, so nightly does deploy them (the Helmfile app layer deploys them regardless, and an unpublished image would otherwise sit in ImagePullBackOff). Dropping not-ga from that job's filter is a one-line change once those suites are confirmed green.

Every robot step writes a triage summary to the GitHub job summary via scripts/robot_summary.py: totals, failures grouped by cause, then the failing test list. A broken environment reads as one fat row ("39 × TENANT_NOT_PROVISIONED") rather than 39 individual regressions, without downloading the log.html artifact.

P0 Defect Coverage

Defect class Test file Tag
Phantom JWT (auth gap between spec and implementation) tests/auditflow/authz.robot regression

Adding, running, or auditing tests

See the test-suite-steward skill (workspace-level .agents/skills/test-suite-steward/) — it covers where a new test belongs, the OpenAPI x-labs64-auth-driven authz matrix, how to run and interpret results, and a periodic suite-health audit (drift, coverage gaps, duplication, flaky handling).

License

The core of the Labs64.IO Ecosystem is entirely open source and free forever. Community modules are licensed under Apache License 2.0.

Contributors

Languages