6  Scenario Files & SLA Gates

Scenario files are the declarative way to define, execute, and validate Kates test runs. Rather than stringing together CLI flags, you describe one or more test scenarios in a YAML (or JSON) file and let Kates orchestrate everything — including automated pass/fail enforcement against SLA thresholds. Kates calls those thresholds SLAs, though they work like SLOs: targets you set, not agreements with anyone [10].

This chapter is for engineers graduating from ad-hoc kates test create runs to version-controlled, CI-gated test suites. After this chapter, you can:

6.1 Why Scenario Files?

CLI flags are convenient for ad-hoc testing, but production-grade performance validation requires:

  • Reproducibility — the same file produces the same test every time
  • Version control — scenario files live in Git alongside your application code
  • Multi-scenario orchestration — run a load test, a stress test, and an integrity test in sequence with a single command
  • SLA enforcement — define pass/fail criteria that block CI/CD pipelines on regressions

graph LR
    subgraph Developer
        Y[scenario.yaml] --> CLI[kates test apply -f]
    end
    
    subgraph Kates
        CLI --> P1[Scenario 1\nLOAD test]
        CLI --> P2[Scenario 2\nSTRESS test]
        CLI --> P3[Scenario 3\nINTEGRITY test]
        P1 --> V1[Validate SLAs]
        P2 --> V2[Validate SLAs]
        P3 --> V3[Validate SLAs]
    end
    
    subgraph Result
        V1 --> R1["✅ PASS / ❌ FAIL"]
        V2 --> R2["✅ PASS / ❌ FAIL"]
        V3 --> R3["✅ PASS / ❌ FAIL"]
    end

6.2 File Format

A scenario file is a YAML (or JSON) document with a single top-level key: scenarios, which is a list of test definitions.

scenarios:
  - name: "My First Test"
    type: LOAD
    spec:
      records: 100000
    validate:
      maxP99LatencyMs: 50
      minThroughputRecPerSec: 10000

Top-Level Structure

Field Type Required Description
scenarios List One or more test scenario definitions; a file of one scenario can leave the list out, as below

A file of one scenario can leave the list out and give that scenario’s fields at its top level. kates test apply reads it as a list of one, and so do kates scenario-diff and the MCP server’s draft_scenario. They read the top level only when the file’s scenarios list is missing or empty and the top level has a type; beside a list of scenarios, top-level fields are ignored.

name: "My First Test"
type: LOAD
spec:
  records: 100000
validate:
  maxP99LatencyMs: 50

Scenario Fields

Each entry in scenarios, or the top level of a file of one scenario, takes these five fields. Only type is required; spec and validate carry the settings and the gates that the next two sections describe.

Field Type Required Description
name String Human-readable scenario name (displayed in output)
type String Yes Test type: LOAD, STRESS, SPIKE, ENDURANCE, VOLUME, CAPACITY, ROUND_TRIP, INTEGRITY, TUNE_REPLICATION, TUNE_ACKS, TUNE_BATCHING, TUNE_COMPRESSION, TUNE_PARTITIONS, or INTEGRATION_CDC. Case-insensitive — the CLI upper-cases the value before submitting
backend String Benchmark backend: native or trogdor (default: native)
spec Object Test specification — see Spec Reference below
validate Object SLA gates — see Validation Reference below

6.3 Spec Reference

The spec object controls all test parameters. Every field is optional. The Kates API, the service in the cluster that runs your tests, fills in defaults per test type (configurable via its kates.tests.<type>.* properties), so a STRESS test defaults to larger batches and more producers than a ROUND_TRIP test. The defaults shown below are the stock values for a LOAD test — other types differ.

Important

The Kates API merges spec with the test type’s defaults, and every key below reaches the run. targetThroughput sets the producer rate in place of the type’s default. parallelProducers counts only for STRESS and CAPACITY, and no test type reads numConsumers, so a LOAD scenario runs one producer and one consumer whatever they say.

The three enable keys take true or false. kates test apply refuses a file where one holds anything else, such as yes, on or nothing, naming the scenario and the key, before it starts any of the file’s tests.

A key the scenario’s type or benchmark backend cannot apply is refused: kates test apply gets a 400 naming it, and the scenario does not start. That is a rate other than -1 for SPIKE or CAPACITY, which run unthrottled; consumerGroup or a fetch setting for a type that takes no consumer settings (every type but LOAD, ENDURANCE and INTEGRITY); enableCrc: true for any type but INTEGRITY; enableIdempotence: true or enableTransactions: true when acks is not all (SPIKE’s default is 1); enableTransactions: true with enableIdempotence: false, or on the trogdor benchmark backend. The API Reference lists the rules; Test Types Deep Dive and Data Integrity Verification cover what the options do.

Producer Configuration

These keys configure the producer. parallelProducers counts only for STRESS and CAPACITY, and targetThroughput replaces the type’s default rate.

Field Type Default Description
records Integer 1,000,000 Number of records to produce
parallelProducers Integer 1 Number of producers for STRESS and CAPACITY; other types run one
recordSizeBytes Integer 1024 Payload size per record in bytes
acks String all Producer acknowledgment mode: 0, 1, or all
batchSize Integer 65536 Producer batch size in bytes
lingerMs Integer 5 Milliseconds to wait before sending a batch
compressionType String lz4 Compression: none, gzip, snappy, lz4, zstd
targetThroughput Integer -1 Producer rate in records/s, for each producer; -1 is unlimited. Replaces the type’s default rate (5,000 for ENDURANCE, 10,000 for ROUND_TRIP)
enableIdempotence Boolean not set The producer’s enable.idempotence; not set, the producer is idempotent whenever acks is all
enableTransactions Boolean false Transactional producers, committing every 100 records or every 10 seconds, whichever comes first; any consumer in the run then reads with read_committed

Consumer Configuration

Consumer settings apply only to LOAD, ENDURANCE and INTEGRITY. ROUND_TRIP starts a consumer too, but it reads without a group and with the Kafka client’s fetch defaults, and no type reads numConsumers.

Field Type Default Description
numConsumers Integer 1 Read by no test type
consumerGroup String auto The consumer’s group, for LOAD, ENDURANCE and INTEGRITY; not empty or blank. A group with committed offsets on the topic resumes from them, and a LOAD or ENDURANCE consumer commits as it reads, so use a group of the test’s own, not one an application reads with. An INTEGRITY consumer joins it with -integrity appended (integrity-cg-integrity when not set)
fetchMinBytes Integer 1 The consumer’s fetch.min.bytes, for LOAD, ENDURANCE and INTEGRITY
fetchMaxWaitMs Integer 500 The consumer’s fetch.max.wait.ms, for LOAD, ENDURANCE and INTEGRITY

Topic Configuration

The first key names the topic a run writes to; the other three describe the layout Kates creates that topic with.

Field Type Default Description
topic String auto (<type>-test) Target topic name
partitions Integer 3 Number of topic partitions
replicationFactor Integer 3 Topic replication factor
minInsyncReplicas Integer 2 Minimum in-sync replicas

Test Execution

Field Type Default Description
durationSeconds Integer per type Time-based duration cap — the run stops at records or the deadline, whichever comes first

Integrity Options

Field Type Default Description
enableCrc Boolean true Whether an INTEGRITY run verifies a CRC32 checksum on each record; false turns the check off. true is refused for any other type

6.4 Validation Reference (SLA Gates)

The validate section defines pass/fail criteria that the CLI checks after each test completes. If any threshold is breached, the violation is listed in the summary table and the CLI exits with a non-zero status code — making it ideal for CI/CD gate enforcement.

Each threshold in validate is a gate. Kates has two other words for a result, and neither names a gate. A disruption plan’s sla block earns an SLA grade, a letter from A to F, as SLA Grading explains. An INTEGRITY run ends with a verdict, such as PASS or DATA_LOSS, as Interpreting Integrity Results explains.

Warning

SLA gates are only evaluated when you run kates test apply with --wait. Without it, scenarios are fire-and-forget: each one is submitted (status SUBMITTED), no gate is ever checked, and how the tests turn out never affects the exit status. A scenario that fails to submit still makes the CLI exit 1.

graph LR
    subgraph Test Completes
        R[Test Results]
    end
    
    subgraph SLA Gates
        R --> G1(["P99 ≤ threshold?"])
        R --> G2(["Avg ≤ threshold?"])
        R --> G3(["Throughput ≥ min?"])
        R --> G5(["Data loss ≤ max?"])
        R --> G6(["RTO ≤ max?"])
        R --> G7(["RPO ≤ max?"])
        R --> G8(["Out-of-order ≤ max?"])
        R --> G9(["CRC failures ≤ max?"])
    end
    
    subgraph Outcome
        G1 --> PASS["All pass → exit 0"]
        G1 --> FAIL["Any fail → exit 1"]
    end

Performance Gates

These gates judge a run’s speed. The CLI checks the first three; maxErrorRate is accepted but not evaluated, since a run reports no error count. Each gate is checked against every task of the run, not against the report summary, so every task must meet it. A consumer records no latency on the native backend, so there a LOAD run’s latency gates judge its producer, while minThroughputRecPerSec judges the producer and the consumer alike. A latency gate needs at least one task that measured latency. When none did, as with a producer that failed before its first acknowledgment or a ROUND_TRIP run on the Trogdor backend, the gate fails as p99 not measured or avg not measured.

Field Type Description
maxP99LatencyMs Float Maximum acceptable P99 latency in milliseconds
maxAvgLatencyMs Float Maximum acceptable average latency in milliseconds
minThroughputRecPerSec Float Minimum acceptable throughput in records per second
maxErrorRate Float Maximum acceptable error rate. Accepted in the file, but not currently evaluated by the CLI’s gate check

Resilience Gates

These gates judge recovery, so they need a run that measured it. When a run reports no RTO or no RPO, the summary marks the gate not evaluable, and the exit code is unchanged. Only an INTEGRITY run reports RTO, and it reports RPO only when a resilience run marks the moment its fault goes in, so a scenario file’s run never measures RPO.

Field Type Description
maxRtoMs Float Maximum Recovery Time Objective in milliseconds. If the run reports no RTO, the summary marks the gate not evaluable rather than passing it; the exit code is unchanged
maxRpoMs Float Maximum Recovery Point Objective in milliseconds. RPO is measured from a chaos start time; a run without one reports RPO as not measured, and the summary marks the gate not evaluable rather than passing it; the exit code is unchanged

Integrity Gates

These gates cap the loss, disorder and corruption an INTEGRITY run may report, and 0 is the strict setting for each. No other type reports them, so on any other run these gates never fail, and the summary still shows ✓ SLA Pass. On an INTEGRITY run with a validate block, a gate the block leaves out is held at 0, and a negative value turns it off.

Field Type Description
maxDataLossPercent Float Maximum acceptable data loss percentage (0.0 = zero loss)
maxOutOfOrder Integer Maximum out-of-order messages (0 = strict ordering)
maxCrcFailures Integer Maximum CRC32 checksum failures (0 = no corruption)

6.5 Examples

Each example is a complete scenario file that you can run with kates test apply -f.

Simple Load Test with Performance SLA

scenarios:
  - name: "Baseline Load Test"
    type: LOAD
    spec:
      records: 100000
      recordSizeBytes: 1024
      acks: all
    validate:
      maxP99LatencyMs: 50
      minThroughputRecPerSec: 10000

Multi-Phase Regression Suite

Run multiple test types in sequence and validate each independently:

scenarios:
  - name: "Load Baseline"
    type: LOAD
    spec:
      records: 100000
    validate:
      maxP99LatencyMs: 50
      minThroughputRecPerSec: 10000

  - name: "Stress Ramp-Up"
    type: STRESS
    spec:
      records: 500000
      parallelProducers: 8
    validate:
      maxP99LatencyMs: 200
      minThroughputRecPerSec: 5000

  - name: "Data Integrity Check"
    type: INTEGRITY
    spec:
      records: 50000
      acks: all
    validate:
      maxDataLossPercent: 0.0
      maxOutOfOrder: 0
      maxCrcFailures: 0

Round-Trip Latency Measurement

scenarios:
  - name: "End-to-End Latency"
    type: ROUND_TRIP
    spec:
      records: 10000
    validate:
      maxP99LatencyMs: 25
      maxAvgLatencyMs: 10

Tuning Comparison

Test different producer configurations side-by-side:

scenarios:
  - name: "Default Batching"
    type: LOAD
    spec:
      records: 100000
    validate:
      maxP99LatencyMs: 50

  - name: "Aggressive Batching"
    type: LOAD
    spec:
      records: 100000
      batchSize: 262144
      lingerMs: 50
      compressionType: zstd
    validate:
      maxP99LatencyMs: 100
      minThroughputRecPerSec: 20000

6.6 Running Scenario Files

One command runs a whole file, and one flag decides how. With --wait, kates test apply runs the scenarios one after another and checks each one’s gates; without it, the command submits them all and checks nothing.

Basic Execution

The two forms differ only in --wait:

# Submit all scenarios in the file (fire-and-forget — they run concurrently, with no SLA evaluation)
kates test apply -f scenarios.yaml

# Run and wait for each to complete; SLA gates are evaluated at the end
kates test apply -f scenarios.yaml --wait

Use --wait whenever the scenarios are meant to be compared, as in the Tuning Comparison file. Without it, the scenarios run at the same time: scenarios of one type that set no topic share that type’s default topic (load-test for LOAD), so each measures the load of the others, and the Kates API runs at most three tests at once (kates.engine.max-concurrent-tests), refusing any further scenario with 429 Too Many Requests.

How Execution Works

  1. Kates parses the file (YAML or JSON) — a malformed file aborts the run with the raw parse error; there is no further schema validation on the client side
  2. Each scenario is submitted to the Kates API sequentially
  3. If --wait is specified, Kates polls until each test completes before submitting the next
  4. SLA gates are evaluated for each completed scenario — this only happens with --wait; without it, every scenario is left as SUBMITTED and never validated
  5. A summary table is printed showing each scenario’s result

Output

Running with --wait:

  ▸ Baseline Load Test (LOAD)...
  ✓   Created: 3f8a2c1e-9b4…
  ✓ Baseline Load Test → DONE
  ▸ Stress Ramp-Up (STRESS)...
  ✓   Created: 7c5e0d2a-1f6…
  ✓ Stress Ramp-Up → DONE
  ▸ Data Integrity Check (INTEGRITY)...
  ✓   Created: b2d94e7f-8a3…
  ✓ Data Integrity Check → DONE

▸ Summary
  Scenario              ID             Status  Note
  ────────────────────  ─────────────  ──────  ─────────────────
  Baseline Load Test    3f8a2c1e-9b4…  DONE    ✓ SLA Pass
  Stress Ramp-Up        7c5e0d2a-1f6…  DONE    p99=210ms > 200ms
  Data Integrity Check  b2d94e7f-8a3…  DONE    ✓ SLA Pass

  ✖ One or more SLA gates violated

In this example, the stress test’s P99 latency (210ms) exceeded the 200ms threshold. The CLI exits with code 1, which would fail a CI/CD pipeline.

6.7 CI/CD Integration

Scenario files are designed for CI/CD pipelines. Combine with --wait to block the pipeline until all tests complete:

# In your CI pipeline script
kates test apply -f regression-suite.yaml --wait

# The exit code tells you the result:
# 0 = every scenario finished and no SLA gate was violated
# 1 = a scenario failed to submit, finished FAILED, was lost track of (ERROR), or violated an SLA gate

A CI job has no terminal, so --wait shows no spinner there: it prints a plain line to stderr each time a run’s status changes, and the summary table to stdout. With -o json stdout carries only the summary as JSON, with each scenario’s runId, status and, for a scenario with gates, its sla violations. A scenario that failed to submit has no runId, so a script reads the run IDs with jq -r '.scenarios[] | select(.runId) | .runId'. The exit code is the same either way.

A failed run fails the pipeline whether or not its scenario has gates. A scenario that fails to submit or finishes FAILED shows as FAILED in the summary, one the CLI loses track of while waiting shows as ERROR, and any of them makes the command exit 1, just as a violated gate does. A run that finishes FAILED meets none of its gates: its first violation is run FAILED, with the first task error or before any task ran. Interrupting the command, with Ctrl-C or by cancelling the job, which sends SIGTERM, cancels the run it is waiting for, starts no further scenario, and exits 130. A scenario that finishes DONE passes unless one of its gates is violated, so a regression fails the pipeline only in a scenario that carries a validate block; without one, a run that completes but regresses exits 0.

For JUnit-compatible output, export each test report individually after the suite completes — see Observability & Monitoring for export formats.

6.8 JSON Format

Scenario files also work in JSON:

{
  "scenarios": [
    {
      "name": "Load Test",
      "type": "LOAD",
      "spec": {
        "records": 100000
      },
      "validate": {
        "maxP99LatencyMs": 50,
        "minThroughputRecPerSec": 10000
      }
    }
  ]
}

6.9 Scaffolding Scenario Files

Rather than writing scenario YAML from scratch, start from the curated template library built into the CLI:

# List the built-in templates (bare `kates test scaffold` does the same)
kates test scaffold list

# Filter the list by test type
kates test scaffold list --type LOAD

# Preview a template
kates test scaffold show quick-load

# Export a template as an editable file in the current directory
kates test scaffold export quick-load

# Export with a custom filename or directory, or export everything
kates test scaffold export production-load -o load-scenario.yaml
kates test scaffold export --all --dir ./scenarios/

The library covers the common cases: quick-load (fast smoke test), production-load (1M records, strict SLA), stress-test, endurance-soak, exactly-once, integrity-tx, spike-test, and ci-gate (a fast 10k-record CI pipeline gate). Edit the exported file and run it with kates test apply -f.

6.10 Common Mistakes

The CLI does not validate scenario files against a schema — a file either parses or it doesn’t, and everything else is checked by the Kates API when the scenario is submitted. These are the most frequent mistakes and what actually happens when you make them.

1. Missing Required Field (type)

type is the only required field, but the CLI does not check for it. The scenario is submitted as-is and the Kates API rejects it with a 400 (its request validation requires a test type), so the scenario shows a ✖ Failed: ... line and is marked FAILED in the summary table.

Fix: Add the type field to every scenario:

scenarios:
  - name: "My Test"
    type: LOAD          # ← required
    spec:
      records: 100000

2. Invalid Test Type Name

Case is not the problem — the CLI upper-cases the type before submitting, so load and Load work fine. What fails is a name that isn’t a real test type: the Kates API rejects it at submission and the scenario is marked FAILED.

Fix: Use one of the valid type names:

scenarios:
  - name: "My Test"
    type: LOAD           # ✅ canonical
    # type: load         # ✅ also works — the CLI upper-cases it
    # type: LATENCY      # ❌ not a valid type — use ROUND_TRIP

3. SLA Threshold Format Error

SLA threshold values must be plain numbers, not strings with units — the field name already indicates the unit (e.g., maxP99LatencyMs implies milliseconds). A string value fails YAML decoding, so the whole run aborts with the raw parse error:

  ✖ Invalid scenario file: yaml: unmarshal errors:
  line 8: cannot unmarshal !!str `50ms` into float64

Fix: Remove the unit suffix and any quotes around the number:

validate:
  maxP99LatencyMs: 50          # ✅ correct — plain number
  # maxP99LatencyMs: "50ms"    # ❌ wrong — string with unit
  # maxP99LatencyMs: "50"      # ❌ wrong — quoted string
  minThroughputRecPerSec: 10000  # ✅ correct

4. Records Count Too Low for Meaningful Results

Kates will not warn you about this — a LOAD test with records: 100 runs happily and reports a “P99” that is really just your single slowest message. Low record counts produce unreliable metrics.

Fix: Use appropriate record counts per test type:

scenarios:
  - name: "Proper Load Test"
    type: LOAD
    spec:
      records: 100000        # ✅ good — 100K for load tests
      # records: 100         # ❌ too low — unreliable P99

  - name: "Proper Integrity Test"
    type: INTEGRITY
    spec:
      records: 100000        # ✅ good — 100K for integrity tests
      # records: 1000        # ❌ too low — may miss intermittent issues

The table gives a recommended minimum for four of the types; the file above sits well above it for LOAD and INTEGRITY.

Test Type Recommended Minimum Why
LOAD 10,000 Enough samples for stable percentile calculations
STRESS 50,000 Need sustained load to detect saturation
INTEGRITY 50,000 Higher counts catch intermittent data loss
ROUND_TRIP 5,000 Latency measurement is per-message, so fewer needed

5. Setting Both records and durationSeconds

Kates does not treat this as a conflict — no error is raised, both values are forwarded to the Kates API, and the run stops at whichever bound is hit first. That silent “whichever comes first” behavior is easy to misread when you look at results later, so keep the intent explicit.

Fix: Set only the termination condition you mean:

scenarios:
  # Option A: count-based (stop after N records)
  - name: "Count-Based Test"
    type: LOAD
    spec:
      records: 100000
      # durationSeconds: 300   # ← remove this

  # Option B: time-based (stop after N seconds)
  - name: "Time-Based Test"
    type: ENDURANCE
    spec:
      durationSeconds: 300
      # records: 100000        # ← remove this
Tip

Try it

Watch an SLA gate fail on purpose — the fastest way to trust a gate is to see it catch something:

# Export the built-in quick-load template into the current directory
kates test scaffold export quick-load

# Edit quick-load.yaml: change maxP99LatencyMs from 100 to 1

# Run with gates enabled, then check the exit code
kates test apply -f quick-load.yaml --wait
echo $?

No real cluster delivers a 1 ms P99, so the summary table marks the scenario DONE with a p99=… > 1ms violation note and echo $? prints 1 — exactly the signal that blocks a CI/CD pipeline.

6.11 Summary

  • A scenario file is a scenarios: list in YAML or JSON, or the fields of one scenario at its top level; type is the only field a scenario must carry, and the Kates API fills in per-type defaults for everything else
  • SLA gates in the validate block are evaluated only with --wait — without it, kates test apply is fire-and-forget, and how the runs turn out never affects its exit status
  • kates test apply exits 1 when any scenario fails to submit, with or without --wait; with --wait it also exits 1 when a scenario finishes FAILED, is lost track of (ERROR), or violates a gate, and a run that completes without a validate block passes whatever its numbers
  • The CLI never validates a file against a schema: malformed YAML aborts the run with the raw parse error, while an invalid type travels to the Kates API and is rejected there
  • Start from a kates test scaffold export template instead of a blank file — edit a known-good scenario, then run it with kates test apply -f

Scenario files lock a winning configuration into Git; finding that configuration interactively is the job of Lab — Interactive Performance Tuning.