graph TB
subgraph Performance["Performance Tests"]
LOAD[LOAD<br/>Steady-state capacity]
STRESS[STRESS<br/>Breaking point]
SPIKE[SPIKE<br/>Burst handling]
ENDURANCE[ENDURANCE<br/>Long-term stability]
VOLUME[VOLUME<br/>Large data sets]
CAPACITY[CAPACITY<br/>Maximum throughput]
end
subgraph Correctness["Correctness Tests"]
RT[ROUND_TRIP<br/>End-to-end latency]
INT[INTEGRITY<br/>Zero data loss]
end
5 Test Types Deep Dive
Table of contents
This chapter covers the eight core Kates test types, each designed to answer a specific question about your Kafka cluster’s behavior — the methodology, use case, and configuration for every type. (Kates also has specialized TUNE_* parameter-sweep types, covered in CLI Reference, and an INTEGRATION_CDC type.)
Whether you’re baselining a new cluster or gating a CI pipeline, after this chapter you can:
- Pick the test type that answers the question you’re actually asking — steady-state capacity, breaking point, burst recovery, or data safety
- Configure each type’s key parameters and know the load each one puts on the cluster, whichever benchmark backend runs it
- Read the results — recognize saturation, slow leaks, and data loss in the metrics each type reports
- Run any type from a built-in scenario template instead of hand-rolled flags
5.1 Test Type Overview
A test type decides the tasks its run starts: how many producers and consumers, at what rate, with which defaults. Both benchmark backends run those same tasks. The default native backend runs them inside the Kates API, the service in the cluster that runs your tests. The Trogdor backend (--backend trogdor) submits each one to the Trogdor coordinator as a task of its own. Each load shape below therefore holds on either backend, and a section says so where Trogdor differs. The diagram groups the types by the question they answer:
5.2 LOAD Test
Question: “What is my cluster’s steady-state performance at expected production throughput?”
Methodology
A LOAD test sends a fixed number of records at a controlled, sustainable rate. It measures the baseline performance that users experience during normal operations.
graph LR
subgraph Load["Load Profile"]
direction LR
T1["Start<br/>producers + consumers"] --> T2["Steady State<br/>fixed rate until records sent<br/>or duration reached"] --> T3["Collect<br/>results"]
end
When to Use
Configuration
| Parameter | Default | Description |
|---|---|---|
records |
1,000,000 | Total messages to produce |
recordSizeBytes |
1024 | Message payload size |
parallelProducers |
1 | Ignored: LOAD runs one producer |
numConsumers |
1 | Ignored: LOAD runs one consumer |
targetThroughput |
-1 (unlimited) | Producer rate in records/s |
acks |
all |
Producer acknowledgment mode |
topic |
load-test |
Target topic name (unless overridden) |
partitions |
3 | Topic partition count |
replicationFactor |
3 | Topic replication factor |
Example
The first command runs a quick baseline. The second runs closer to production, with 2048-byte records, a topic of its own and acks=all spelled out.
# Quick baseline
kates test create --type LOAD --records 100000 --wait
# Production-like configuration
kates test create --type LOAD \
--records 500000 \
--record-size 2048 \
--topic perf-load-test \
--acks all \
--waitScenario file equivalent (see Scenario Files & SLA Gates):
scenarios:
- name: "Production Load Baseline"
type: LOAD
spec:
records: 500000
recordSizeBytes: 2048
topic: perf-load-test
acks: all
validate:
maxP99LatencyMs: 50
minThroughputRecPerSec: 10000Interpreting Results
Healthy ranges are environment-dependent — treat these as starting points and calibrate against your own baseline:
| Metric | Healthy Range | Warning |
|---|---|---|
| P99 Latency | < 50ms | > 200ms suggests resource contention |
| Error Rate | 0% | Any errors indicate a configuration problem |
| Throughput variability | < 10% stddev | High variance suggests GC or I/O pressure |
For iterative parameter tuning, use kates lab instead of individual test create commands. Lab lets you tweak parameters, run tests, and compare results in a single session — see Lab — Interactive Performance Tuning.
5.3 STRESS Test
Question: “At what point does my cluster break, and how does it degrade?”
Methodology
A STRESS test pushes the cluster well past its comfortable operating point to find the saturation point and characterize the degradation curve. It starts parallelProducers producers at once, 3 by default, and each one sends as fast as the cluster accepts unless targetThroughput sets a rate for it. The profile has no steps, as the diagram shows:
graph LR
subgraph Stress["Load Profile"]
direction LR
S1["Start<br/>3 producers at once"] --> S2["Full speed<br/>each producer unthrottled<br/>until its records are sent<br/>or the duration ends"] --> S3["Collect<br/>results"]
end
When to Use
- Capacity planning — how much headroom does the cluster have?
- Identifying bottlenecks — which component saturates first (CPU, network, disk, memory)?
- Validating auto-scaling policies — does the cluster scale before degradation?
Configuration
| Parameter | Default | Description |
|---|---|---|
parallelProducers |
3 | Producers started at once |
targetThroughput |
-1 (unlimited) | Rate for each producer, in rec/s |
durationSeconds |
900 | Upper bound on each producer’s run |
records |
5,000,000 | Records for each producer; it stops once they are sent |
recordSizeBytes |
1024 | Message size |
Interpreting Results
One STRESS run gives one point on the degradation curve, because its producers start together and run at one level of load. To draw the curve, repeat the run with more producers each time, or a higher targetThroughput, and compare the runs with kates report compare. Across the series, the metrics pass through three stages [20]:
graph LR
H["Low load: Healthy<br/>throughput rises with load<br/>latency stable<br/>errors = 0"] --> S["Near the ceiling: Saturation<br/>throughput plateaus<br/>latency rising<br/>GC pressure increasing"] --> O["Past the ceiling: Overload<br/>throughput drops<br/>latency spikes<br/>errors appear"]
5.4 SPIKE Test
Question: “Can my cluster handle sudden traffic bursts without cascading failure?”
Methodology
A SPIKE test simulates a flash-sale or viral event — a sudden, dramatic increase in traffic followed by a return to normal. The run is the burst alone: one producer, unthrottled from its first record, with acks=1 by default. The baseline before it and the recovery after it come from outside the run, as the diagram shows:
graph LR
subgraph Spike["Around a SPIKE Run"]
direction LR
S1["Before<br/>a LOAD run records<br/>your baseline P99"] --> S2["SPIKE run<br/>1 unthrottled producer<br/>acks=1"] --> S3["After<br/>your monitoring<br/>shows the recovery"]
end
When to Use
- Flash sale preparation — can the cluster absorb 10x traffic?
- Incident simulation — what happens when a retry storm hits?
- Recovery validation — how long until the cluster returns to normal after a spike?
Configuration
| Parameter | Default | Description |
|---|---|---|
records |
2,000,000 | Total records for the burst |
recordSizeBytes |
1024 | Message size |
durationSeconds |
300 | Upper bound on the burst; the producer stops sooner once its records are sent |
acks |
1 |
Latency-oriented default for burst traffic |
Key Metrics
Each row is a moment around the burst and what to note there. Only the middle row comes from the SPIKE run itself: the baseline P99 comes from the LOAD run before it, and the recovery from your monitoring after it.
| Phase | Watch For |
|---|---|
| Pre-spike baseline | Record your normal P99 |
| During spike | Does latency grow linearly or exponentially? |
| Post-spike recovery | How long until P99 returns to baseline? |
5.5 ENDURANCE Test
Question: “Does performance degrade over hours or days of sustained load?”
Methodology
An ENDURANCE (soak) test runs at a moderate, realistic load for an extended period — up to two hours on a default install, longer once you raise the Kates API’s run limit — to detect slow resource leaks and gradual degradation.
graph LR
subgraph Endurance["Load Profile"]
direction LR
E1["Sustained rate-limited load<br/>5,000 msg/s default<br/>2 h run limit by default — raise it for leak hunting"]
end
What It Detects
Each row is a slow failure and the symptom it leaves over a long run: drift that a short test ends before it can show.
| Problem | How It Manifests |
|---|---|
| Memory leak | P99 latency slowly rises over hours |
| Log segment accumulation | Disk usage grows, then GC pauses spike |
| Connection pool exhaustion | Error rate slowly increases |
| JVM metaspace growth | Off-heap memory consumption rises |
| Thread leak | Thread count climbs, eventually OOM |
Configuration
| Parameter | Default | Description |
|---|---|---|
durationSeconds |
3600 (1h) | Upper bound on the run; a longer soak also needs more records, and one past two hours a raised Kates API run limit (see the callout below) |
parallelProducers |
1 | Ignored: ENDURANCE runs one producer and one consumer |
targetThroughput |
5,000 msg/s | Rate limit that keeps the load sustainable; the 5,000 is the ENDURANCE default, and this key replaces it |
records |
10,000,000 | Enough for the full duration |
No run lasts longer than two hours by default
The Kates API refuses, with 400, a run set to last longer than two hours, and fails a run still RUNNING five minutes after its own durationSeconds is up, counted from the run’s creation. The two hours are the Kates API setting kates.engine.max-duration-ms, 7,200,000 ms by default, and the kates chart has no value for it. To allow longer soaks, save the release’s values with helm get values kates -n kates -o yaml, add the environment variable KATES_ENGINE_MAX_DURATION_MS, in milliseconds, to their extraEnv, and helm upgrade the release with that file. kates deploy upgrades the release from its own values files, which drops the entry, so repeat the upgrade after it.
5.6 VOLUME Test
Question: “How does my cluster handle large messages or large data volumes?”
Methodology
A VOLUME test focuses on data size rather than request rate. It sends large messages or large total volumes to stress the storage and replication subsystems. The run is one producer sending large records, 2,000,000 of 10 KB each by default, as fast as the cluster accepts them unless targetThroughput sets a rate.
When to Use
- Validating large message support — Kafka has a default 1MB message size limit
- Storage capacity planning — how fast does disk fill at production data rates?
- Replication overhead — larger messages amplify replication latency
Configuration
| Parameter | Default | Description |
|---|---|---|
recordSizeBytes |
10,240 | Large messages (10 KB) |
records |
2,000,000 | Enough to stress storage |
acks |
all |
Full replication to measure real cost |
Scenario file equivalent:
scenarios:
- name: "Large Message Volume"
type: VOLUME
spec:
records: 10000
recordSizeBytes: 102400
acks: all
validate:
maxP99LatencyMs: 5005.7 CAPACITY Test
Question: “What is the absolute maximum throughput my cluster can sustain?”
Methodology
A CAPACITY test removes all artificial throttling and pushes the cluster to its maximum throughput. It finds the ceiling and measures what metric (CPU, disk, memory, network) is the bottleneck. The run starts parallelProducers producers at once, 5 by default, and every one is unthrottled: the Kates API refuses any rate but -1 for CAPACITY.
Configuration
| Parameter | Default | Description |
|---|---|---|
parallelProducers |
5 | Producers started at once, each unthrottled |
recordSizeBytes |
1024 | Standard message size |
records |
10,000,000 | Records for each producer; it stops once they are sent |
durationSeconds |
1200 | Upper bound on each producer’s run |
Interpreting Results
The output is a throughput curve, built by running a series of CAPACITY tests with increasing --producers counts. Max throughput is where adding more producers stops increasing total rec/s (illustrative numbers):
| Producers | Throughput | Interpretation |
|---|---|---|
| 1 | 50K rec/s | Single-threaded baseline |
| 2 | 95K rec/s | Near-linear scaling |
| 4 | 170K rec/s | Still scaling |
| 8 | 200K rec/s | Diminishing returns — approaching saturation |
| 16 | 195K rec/s | Throughput actually drops — overloaded |
5.8 ROUND_TRIP Test
Question: “What is the true end-to-end latency from produce to consume?”
Methodology
A ROUND_TRIP test measures the complete record lifecycle: the time from when a producer sends a record to when a consumer receives it. That span includes the producer’s batching, the replication the leader waits for before the record becomes readable, and the consumer’s fetch.
The sequence below marks where the clock starts and stops. Notice that the producer’s acknowledgement comes midway, so a run that stopped there would miss the consumer’s half of the trip:
sequenceDiagram
participant P as Producer
participant L as Leader
participant F as Follower
participant C as Consumer
P->>L: 1. Send (t₁)
L->>F: 2. Replicate
F->>L: 3. ACK
L->>P: 4. Producer ACK
C->>L: 5. Fetch
L->>C: 6. Deliver (t₂)
Note over P,C: Round-trip latency = t₂ - t₁
On the default native backend, the producer and the consumer both run inside the Kates API, so one clock times both ends of the trip. The consumer starts first and reads every partition of the topic from its current end, without a consumer group, so records that earlier runs left on the topic are never read ahead of this run’s. Each record carries the moment it was sent, and the consumer subtracts that moment from the moment the record arrives. Both readings come from the same clock in the same process, so the difference needs no clock synchronization between hosts.
Reading the Results
A ROUND_TRIP run has one task, and its row reports both sides of the trip. The table says what each field counts on that row:
| Field | On a ROUND_TRIP row |
|---|---|
| Records and throughput | Records that came back to the consumer, the same records the latency describes; the API field is recordsSent, as on every row |
| Latency, P50 to max | Send to receipt, one sample per record received; the producer’s acknowledgement time is not part of it |
| Error | Sends the broker rejected, and acknowledged records that never came back |
| Status | FAILED when the broker rejected every send or no acknowledged record came back; otherwise DONE |
A shortfall on either side leaves the run DONE and names the count in the error, so a DONE row with an error is a partial result, not a clean one. Once the producer finishes, the consumer waits until every acknowledged record has arrived, or until 10 seconds pass with none of the rest arriving. With enableTransactions: true, the consumer reads with read_committed, so each record arrives only when its transaction commits, and its latency includes that wait.
Configuration
The ROUND_TRIP defaults trade volume for clean samples, and the consumer’s settings are fixed:
| Parameter | Default | Description |
|---|---|---|
parallelProducers |
1 | Ignored: ROUND_TRIP runs one producer |
numConsumers |
1 | Ignored: ROUND_TRIP runs one consumer |
records |
500,000 | Records to measure |
targetThroughput |
10,000 rec/s | Rate-limited to keep latency measurements clean; the 10,000 is the ROUND_TRIP default, and this key replaces it |
consumerGroup, fetchMinBytes, fetchMaxWaitMs |
— | Refused with a 400: the consumer reads without a group, with the Kafka client’s fetch defaults |
Scenario file equivalent:
scenarios:
- name: "End-to-End Latency"
type: ROUND_TRIP
spec:
records: 10000
validate:
maxP99LatencyMs: 25
maxAvgLatencyMs: 10Limits
End-to-end latency comes only from the native backend. On the Trogdor backend (--backend trogdor), Trogdor’s round-trip workload counts the records it sends and receives but reports no latency. The row’s records are the ones received, as on the native backend, but every latency field reads 0, which means not measured. A latency gate such as maxP99LatencyMs then fails as not measured, so run ROUND_TRIP on the native backend whenever the latency matters. A record the producer retried can arrive twice, and each arrival is timed: ROUND_TRIP does not check for duplicates, and INTEGRITY does.
5.9 INTEGRITY Test
Question: “Does my cluster lose, duplicate, or reorder messages under stress?”
Methodology
The INTEGRITY test is the most critical test type. It produces messages with monotonic sequence numbers, tracks acknowledgments, and then consumes all messages to verify completeness.
graph TB
subgraph Producer
P[Produce messages<br/>seq: 1, 2, 3, ..., N]
PA[Track ACKs<br/>Record gaps]
end
subgraph Kafka
K[Replication + Storage]
end
subgraph Consumer
C[Consume all messages]
CV[Verify sequences<br/>Detect gaps]
end
subgraph Verdict
V{All sequences<br/>accounted for?}
PASS[PASS ✅<br/>Zero data loss]
FAIL[DATA_LOSS ❌<br/>Lost ranges identified]
end
P --> K --> C
PA --> V
CV --> V
V -->|Yes| PASS
V -->|No| FAIL
An INTEGRITY run needs the native backend. The Trogdor backend has no workload that numbers and verifies records, so on --backend trogdor the run fails as it starts, with the error INTEGRITY/CDC tests require the native backend.
What It Verifies
| Property | How |
|---|---|
| No data loss | Every produced sequence number is consumed |
| No duplication | Each sequence number appears exactly once (with idempotence) |
| No reordering | Sequence numbers arrive in order per partition |
| ACK consistency | Every ACKed message is actually persisted |
Configuration
| Parameter | Default | Description |
|---|---|---|
records |
1,000,000 | Messages to verify |
acks |
all |
Default for integrity guarantees; a request can override it, and 1 or 0 also turn producer idempotence off |
enableIdempotence |
not set | Sets the producer’s enable.idempotence; left out, the Kafka producer is idempotent by default with acks=all (see the callout below) |
enableTransactions |
false |
Transactional producer, committing every 100 records or every 10 seconds, whichever comes first; the verifying consumer then reads with read_committed |
enableCrc |
true |
Per-record CRC payload verification; false turns it off |
numConsumers |
1 | Ignored: INTEGRITY runs one producer and one consumer |
consumerGroup |
integrity-cg |
Base of the consumer group name: the verifying consumer joins it with -integrity appended, integrity-cg-integrity by default |
enableIdempotence and enableTransactions need acks to be all, which the Kafka producer requires for both, and a transactional producer is always idempotent: the Kates API refuses a request that asks for either with other acks, or for transactions with enableIdempotence: false, with a 400 that names the field. A request that leaves enableIdempotence out gets the client’s choice, idempotent whenever acks is all. Data Integrity Verification covers what each integrity mode checks.
Scenario file equivalent:
scenarios:
- name: "Zero-Loss Integrity"
type: INTEGRITY
spec:
records: 100000
acks: all # also makes the producer idempotent
validate:
maxDataLossPercent: 0
maxOutOfOrder: 0
maxCrcFailures: 0Integrity + Chaos
The real power of INTEGRITY tests emerges when combined with chaos engineering. An INTEGRITY run injects no fault by itself: kates resilience run pairs a test request with one fault, injected while the test runs. That pairing is a resilience run, one of the two ways Kates runs a chaos experiment, and it doesn’t go through the safety guard that checks a disruption plan (see Chaos Engineering in Practice):
# resilience-integrity.yaml
testRequest:
type: INTEGRITY
spec:
numRecords: 180000 # at 500 records/s: 360 s of producing
throughput: 500
durationMs: 600000
chaosSpec:
experimentName: broker-pod-kill
targetNamespace: kafka
targetLabel: "strimzi.io/component-type=kafka,strimzi.io/broker-role=true"
disruptionType: POD_KILL
chaosDurationSec: 30
steadyStateSec: 30# Run it — produces for 360 s and kills one broker from 30 s to 60 s
kates resilience run -f resilience-integrity.yamlThis produces sequenced records at 500 per second and deletes one broker’s pod 30 s in and, on the default LitmusChaos provider, again until 60 s. The run keeps producing through the broker’s restart and return to the ISR, then consumes everything back and verifies that every acknowledged record was persisted. The rate limit is what makes the result mean something: an unthrottled run can finish before the fault is triggered, and its verdict then says nothing about the failure. A resilience file’s spec uses the API’s field names, so the rate is throughput; the API also takes it as targetThroughput, the scenario file’s name. Data Integrity Verification walks through the sizing and how to read the verdict, which the INTEGRITY run reports rather than kates resilience run. For a standalone integrity scenario, export the built-in template instead: kates test scaffold export integrity-tx.
5.10 Scenario Files
All test types support YAML scenario files for reproducible, version-controlled test definitions. See Scenario Files & SLA Gates for the complete YAML schema reference, including the full spec field list and the SLA gates.
The CLI ships a curated library of built-in templates. Browse it with list (optionally filtered by --type), preview with show, and write a ready-to-edit file with export:
# Browse the built-in template library
kates test scaffold list
kates test scaffold --type LOAD
# Preview and export a template
kates test scaffold show quick-load
kates test scaffold export quick-load
# Apply a scenario
kates test apply -f quick-load.yaml --waitCLI flags and scenario-file spec keys use different names for the same setting. The canonical mapping:
CLI flag (kates test create) |
Scenario YAML key (spec:) |
Meaning |
|---|---|---|
--records |
records |
Total records |
--producers |
parallelProducers |
Producers, for STRESS and CAPACITY only; other types run one |
--consumers |
numConsumers |
Read by no test type; each type runs at most one consumer |
--record-size |
recordSizeBytes |
Record size in bytes |
--duration |
durationSeconds |
Test duration in seconds |
--acks |
acks |
Producer acknowledgment mode |
--topic |
topic |
Topic name |
--throughput |
targetThroughput |
Producer rate in records/s, for each producer (see the callout below) |
--consumer-group |
consumerGroup |
Consumer group, for LOAD, ENDURANCE and INTEGRITY; a group of the test’s own, since a LOAD or ENDURANCE consumer commits offsets in it |
--fetch-min-bytes, --fetch-max-wait-ms |
fetchMinBytes, fetchMaxWaitMs |
Consumer fetch settings, for LOAD, ENDURANCE and INTEGRITY |
| — | enableIdempotence, enableTransactions, enableCrc |
Integrity options, true or false (scenario files only; see the callout under INTEGRITY Test) |
--throughput and targetThroughput both send the API field targetThroughput, which sets the producer’s rate in place of the type’s default. The API also takes the same rate as throughput, the name a kates resilience run file uses; when a request sets both, throughput wins. SPIKE and CAPACITY run their producers unthrottled, so the Kates API refuses a rate other than -1 for them, and it refuses a consumer setting for a type that starts no consumer: the answer is a 400 naming the field.
Try it
Run a correctness test end to end from a built-in template:
# List every test type the API supports
kates test types
# Export the integrity-tx INTEGRITY template and inspect it
kates test scaffold export integrity-tx
cat integrity-tx.yaml
# Run it and wait for the verdict
kates test apply -f integrity-tx.yaml --waitThe apply blocks until the verification pass completes — on a healthy cluster, expect ✓ SLA Pass. The SLA gates do not check duplicates, so confirm Duplicates 0 and Verdict ● PASS with kates test get <id>.
5.11 Summary
- Every test type answers one specific question — choose by the question you need answered, not by the knobs you want to turn.
- LOAD establishes the baseline every other result is judged against; STRESS and CAPACITY find the ceiling — STRESS characterizes how the cluster degrades, CAPACITY measures the absolute maximum.
- The benchmark backend doesn’t change a type’s load: native and Trogdor start the same producers and consumers, though INTEGRITY runs only on native.
- ENDURANCE and VOLUME stress the dimensions short tests miss: time (slow leaks, gradual degradation) and data size (storage and replication overhead).
- INTEGRITY verifies zero loss, zero duplication, and correct ordering with sequence numbers and CRC checks — pair it with chaos through
kates resilience runfor the ultimate durability validation.
Every type here maps onto a version-controlled YAML definition — Scenario Files & SLA Gates covers the full schema and the SLA gates that turn test results into a pass or a fail.