flowchart LR
subgraph Machine["Your machine"]
CLI["kates CLI<br/>commands, Lab, kates mcp"]
end
subgraph Cluster["Kubernetes cluster"]
API["Kates API pod<br/>REST and gRPC, native benchmark backend, chaos provider"]
PG[("PostgreSQL")]
KAFKA["Kafka cluster krafter"]
LIT["LitmusChaos"]
TROG["Trogdor coordinator<br/>optional, not installed"]
PROM["Prometheus and Grafana"]
K8S["Kubernetes API"]
end
CLI -->|"REST, API key"| API
CLI -->|"kubectl, helm"| K8S
API -->|"runs and results"| PG
API -->|"producers, consumers, AdminClient"| KAFKA
API -->|"ChaosEngine resources"| LIT
LIT -->|"faults"| KAFKA
API -.->|"trogdor benchmark backend"| TROG
TROG -.-> KAFKA
PROM -->|"scrapes"| API
PROM -->|"scrapes"| KAFKA
2 Architecture & Design
Table of contents
Kates has two halves: the kates CLI on your machine, and the Kates API in the cluster. Around them, kates deploy installs Kafka, PostgreSQL, LitmusChaos, Prometheus and Grafana, and this chapter shows how the parts fit. It serves anyone who operates, extends, or debugs Kates — the mental model built here underpins every later chapter. After this chapter, you can:
- Name Kates’s parts and say where each one runs
- Trace a LOAD test from
kates test createthrough theTestOrchestratorandNativeKafkaBackendto its finalTestReport - Say what the safety guard checks before a disruption plan injects a fault, and what it rolls back when a step fails
- Say where a run’s results live, and how long each piece of them lasts
2.1 High-Level Architecture
The half of Kates that does the work isn’t on your machine. The kates CLI runs where you type; the Kates API runs in the cluster, as one Deployment beside the Kafka cluster it tests. It generates the load, asks for the faults, grades disruption plans and keeps every run in PostgreSQL.
The diagram shows the main parts of a default kates deploy and who calls whom. Solid arrows are the default paths; the dashed ones are the Trogdor path, which kates deploy doesn’t set up.
Read it from the left. The CLI calls the Kates API over REST, with the URL and API key of your CLI context. For the commands that install and reach the stack, such as kates deploy and kates ports, it runs kubectl and helm against your current Kubernetes context instead. The Commands table in CLI Reference says which command takes which path. In the cluster, the Kates API talks to krafter, the Kafka cluster under test, as an ordinary Kafka client, and keeps each run in PostgreSQL. On the default chaos provider, it creates LitmusChaos resources for every fault except POD_DELETE, ROLLING_RESTART and SCALE_DOWN, which it injects itself through the Kubernetes API. Prometheus scrapes the brokers and the Kates API, and the Kates API reads Prometheus back for a disruption plan’s broker metrics.
Each part in the table runs in one place. The namespaces are those of the default isolated topology; with --topology single, every part below that runs in the cluster, except Trogdor, shares kates-stack, as Single-Namespace vs Multi-Namespace describes.
| Part | What it does | Where it runs |
|---|---|---|
kates CLI |
Calls the Kates API, and runs kubectl and helm for the commands that install, reach and remove the stack |
Your machine |
Lab, dashboard, top and kafka tui |
Terminal views on the Kates API; every run Lab starts, warm-ups and median runs included, is one test run | Your machine, inside the CLI |
kates mcp |
Serves read-only tools to an AI agent, whose MCP client starts it | Your machine, inside the CLI |
| Kates API | Serves REST and gRPC on port 8080, and runs tests, disruptions and schedules | Deployment kates, namespace kates |
| Native benchmark backend | Runs a test’s producers and consumers on virtual threads inside the Kates API | The Kates API pod |
| Trogdor benchmark backend | Sends a test’s workload to a Trogdor coordinator, which kates deploy doesn’t install |
The Kates API pod; the load comes from the Trogdor agents you run |
| Chaos provider | Turns each fault into LitmusChaos resources or direct Kubernetes API calls | The Kates API pod |
| LitmusChaos | Runs the experiments the chaos provider asks for | Operator in litmus; ChaosEngines in kafka |
| PostgreSQL | Keeps runs and their results, disruption reports, schedules, baselines and audit events | StatefulSet kates-postgresql, namespace kates |
krafter |
The Kafka cluster under test | Namespace kafka |
| Prometheus and Grafana | Scrape and chart the brokers and the Kates API | Namespace monitoring |
The Kates API is one process, and that has a cost. A native run’s producers and consumers share the pod’s CPU and memory with everything else the Kates API does. kates deploy limits that pod to one CPU and 512 MiB of memory, on Kind and on other clusters alike. A throughput ceiling you measure can therefore be the pod’s rather than krafter’s.
The Kates API comes as two images built from the same code: a JVM image and a GraalVM native image. On a Kind cluster, kates deploy runs a native image that you built or pulled onto your machine, kates:native-local or else kates:native, and never pulls one itself. Without either, it stops at the Kates API step and names the command that builds one. On any other cluster, it runs the published JVM image. The two collect garbage differently: the native image’s Serial GC stops the Kates API for every collection, and those pauses land in the latencies it records. Native Image Build explains when to run which.
2.2 The Kates API
The Kates API is a Quarkus application. It exposes both a REST API and a gRPC API (see gRPC API Reference), and manages the full test lifecycle. Both APIs delegate to the same service layer, so a test run behaves the same whichever API starts it; the gRPC API covers fewer operations, and some of its responses carry less than their REST counterparts, as its chapter lists.
Component Map
graph LR
subgraph Domain
TR[TestRun]
TS[TestSpec]
TT[TestType]
TRes[TestResult]
end
subgraph Engine
TO[TestOrchestrator]
NKB[NativeKafkaBackend]
LH[LatencyHistogram]
BS[BenchmarkStatus]
BM[BenchmarkMetrics]
end
subgraph Report
RG[ReportGenerator]
CS[ClusterSnapshot]
BrokerM[BrokerMetrics]
SLA[SlaVerdict]
end
subgraph Export
CSV[CsvExporter]
JUnit[JunitXmlExporter]
HM[HeatmapExporter]
end
TO --> NKB
NKB --> LH
NKB --> BS
RG --> CS
RG --> BrokerM
RG --> SLA
RG --> CSV
RG --> JUnit
RG --> HM
TestOrchestrator
The TestOrchestrator is the central coordinator. When a test is created, it:
- Resolves defaults — merges the incoming
TestSpecwithTestTypeDefaultsfor the chosen test type - Creates the topic — ensures the Kafka topic exists with the required partition count and replication factor
- Launches workers — delegates to the
BenchmarkBackendto start producer/consumer tasks - Polls status — periodically polls each
BenchmarkHandleforBenchmarkStatusupdates - Collects heatmap data — on each poll of a running test, stores the latency buckets recorded so far as one heatmap row
The Kates API builds a report only when something asks for one. It builds a finished run’s report from the stored run the first time it’s asked (ReportGenerator), with summary metrics, whether the run met its SLA thresholds (targets you set, in the style of an SLO) and broker correlation.
NativeKafkaBackend
The native benchmark backend (NativeKafkaBackend) runs each test’s producers and consumers on virtual threads inside the Kates API [9]. It is one of two benchmark backends; the other, trogdor, sends the workload to a Trogdor coordinator, and Test Types Deep Dive compares them. Its workers:
- Produce messages with configurable record size, acknowledgment mode, and throughput throttling
- Consume messages with configurable consumer group, fetch settings, and poll timeout
- Record latency in a
LatencyHistogrambacked by HdrHistogram (1µs–60s range, microsecond precision) - Track integrity — sequence numbers, acknowledgment gaps, and consumer-side deduplication
LatencyHistogram
The histogram is the heart of latency measurement. It is backed by HdrHistogram (configured for 1µs–60s with 3 significant value digits), which provides high resolution at low latencies (sub-millisecond) while covering tails up to 60 seconds.
graph LR
subgraph Internal["HdrHistogram (1µs–60s)"]
direction TB
B1["0.001ms"] --> B2["0.01ms"] --> B3["0.1ms"] --> B4["1ms"] --> B5["10ms"] --> B6["100ms"] --> B7["1000ms"]
end
subgraph Export["25 Heatmap Buckets"]
direction LR
H1["0–0.1ms"]
H2["0.1–0.5ms"]
H3["0.5–1ms"]
H4["1–5ms"]
H5["5–50ms"]
H6["50–500ms"]
H7["500ms–10s"]
end
Internal -->|exportBuckets| Export
Key methods:
| Method | Lock | Purpose |
|---|---|---|
recordLatency(latencyMs) |
Write | Record a single latency observation |
getPercentile(p) |
Read | Compute P50/P95/P99 from cumulative distribution |
exportBuckets() |
Read | Compress to 25 heatmap ranges (non-destructive) |
snapshotAndReset() |
Write | Atomic capture + reset for windowed collection |
2.3 Disruption Engine
A disruption plan injects faults step by step and watches the cluster from outside, while a resilience run, started with kates resilience run, injects one fault under a Kates test and measures what a client sees. Both get the fault from the Kates API’s chaos provider, LitmusChaos by default, as Chaos Engineering in Practice explains.
The diagram shows the parts behind a disruption plan. Look at the Providers group: one setting picks the chaos provider that every fault goes through.
graph LR
subgraph Control
DO[DisruptionOrchestrator]
DSG[DisruptionSafetyGuard]
DPC[DisruptionPlaybookCatalog]
end
subgraph Intelligence
KIS[KafkaIntelligenceService]
ISR[ISR Tracking]
LAG[Consumer Lag]
LEAD[Leader Resolution]
end
subgraph Providers
CP{{kates.chaos.provider}}
HCP[HybridChaosProvider]
KCP[KubernetesChaosProvider]
LCP[LitmusChaosProvider]
end
subgraph Reporting
DR[DisruptionReport]
SG[SlaGrader]
PMC[PrometheusMetricsCapture]
end
DO --> DSG
DO --> KIS
DO --> CP
DO --> DR
DPC --> DO
KIS --> ISR
KIS --> LAG
KIS --> LEAD
CP -->|"litmus-crd, the default"| LCP
CP -->|kubernetes| KCP
CP -->|hybrid| HCP
HCP -.->|picks one| LCP
HCP -.->|picks one| KCP
LCP -->|"POD_DELETE,<br/>ROLLING_RESTART,<br/>SCALE_DOWN"| KCP
DR --> SG
DR --> PMC
Disruption Types
Kates’s disruption types run from killing a broker pod to draining a node, and kates disruption types lists them. What each one does to a pod or to the network depends on the Kates API’s chaos provider, LitmusChaos by default, so each type is described once, per provider, in Chaos Engineering in Practice.
Safety Guardrails
Before a disruption plan touches the cluster, Kates counts the brokers its steps would hit, and refuses the plan if that’s every broker, or more than the plan’s maxAffectedBrokers. Before each fault it checks that every Kafka pod is Running and Ready. When a step fails with autoRollback on, rollback gives a scaled-down node pool its broker back, or deletes the NetworkPolicies Kates created for a network partition. The safety guard doesn’t read the ISR or the KRaft quorum, and a resilience run doesn’t go through it: Chaos Engineering in Practice says exactly what it checks.
2.4 CLI Architecture
The CLI is a standalone Go binary built with Cobra. It splits its work between two channels: it calls the Kates API over REST for test, report, and disruption operations, and it shells out to kubectl, helm, and kind (via os/exec) for cluster provisioning and lifecycle tasks such as kates deploy, kates clean, and kates ports. It also keeps your contexts, profiles and snapshots in files in your home directory; the Commands table in CLI Reference says which command uses which.
graph TD
subgraph Config
CTX[~/.kates.yaml]
CTXM[Context Manager]
end
subgraph Commands
TEST[test create/list/get/delete/watch/apply/scaffold]
REPORT[report show/summary/export/compare/diff/brokers]
DISRUPT[disruption run/list/status/timeline/types/kafka-metrics]
RESIL[resilience run]
TREND[trend]
OPS[health/cluster/top/dashboard/status]
end
subgraph Output
TABLE[Table Renderer]
JSON[JSON Printer]
SPARK[Sparkline Charts]
BADGE[Status Badges]
BAR[Metric Bars]
end
CTX --> CTXM
CTXM --> Commands
Commands --> TABLE
Commands --> JSON
Commands --> SPARK
Key design decisions:
- Multi-context support — like
kubectl, the CLI supports named contexts for targeting different Kates APIs - Rich terminal output — tables, colored badges, metric bars, sparkline charts, and ASCII banners
- Scaffold templates —
kates test scaffold export <name>writes a ready-to-use YAML scenario file to the current directory (browse the library withkates test scaffold list, optionally filtered by--type LOAD) - Streaming watch —
kates test watchandkates disruption watchprovide real-time progress updates (disruption watchreceives no events for a disruption ID yet; see its entry in the CLI reference)
2.5 Data Flow
This diagram traces a complete test execution from CLI command to final report:
sequenceDiagram
participant CLI as Kates CLI
participant API as REST API
participant Orch as TestOrchestrator
participant Engine as NativeKafkaBackend
participant Kafka as Kafka Cluster
participant Hist as LatencyHistogram
participant Report as ReportGenerator
CLI->>API: POST /api/tests {type: LOAD, spec: {...}}
API->>Orch: createTest(type, spec)
Orch->>Kafka: Create topic (if needed)
Orch->>Engine: start(handles)
Engine->>Kafka: Produce messages
Engine->>Hist: record(latencyUs)
loop Every poll interval
Orch->>Engine: poll(handle)
Engine->>Hist: exportBuckets()
Engine-->>Orch: BenchmarkStatus + heatmapBuckets
Orch->>Orch: Accumulate heatmap rows
end
CLI->>API: GET /api/tests/{id}
API->>Orch: getTest(id)
Orch-->>CLI: TestRun (status, results)
CLI->>API: GET /api/tests/{id}/report
API->>Report: generate(testRun)
Report->>Kafka: captureSnapshot (broker metrics)
Report-->>CLI: TestReport (summary, SLA, brokers)
2.6 Disruption Pipeline
The disruption path adds two things the test path lacks: a check that can refuse the plan, and a rollback when a step fails. This diagram traces a plan from kates disruption run to its report:
sequenceDiagram
autonumber
participant Cli as kates CLI
participant API as Kates API
participant Kube as Kubernetes API
participant Prov as Chaos provider
participant Prom as Prometheus
Cli->>API: POST /api/disruptions
API->>Kube: List the Kafka pods, count the brokers the plan hits
API-->>Cli: 202 with a disruption ID, 422 refused, or 409 busy
loop Each step
API->>API: If the step names a topic, find its leader, then wait steadyStateSec
API->>Kube: Are all Kafka pods Running and Ready?
API->>Prom: Baseline snapshot
API->>Prov: Inject the fault
Prov-->>API: Pass, Fail or Skipped
API->>API: Wait observationWindowSec
API->>Prom: Impact snapshot
opt requireRecovery
API->>Kube: Wait for the Kafka pods to be Ready again
end
opt Recovery failed or the step threw, with autoRollback on
API->>Kube: Restore node pool replicas, or delete the partition policies Kates made
end
end
API->>API: Grade against the sla block, if it has one
Cli->>API: GET /api/disruptions/{id}
API-->>Cli: The report, graded when the plan has an sla block
2.7 Technology Stack
The table lists what Kates is built with and the tools it works beside. A default kates deploy installs neither Jaeger, Velero and its SeaweedFS object store, nor Kyverno. It does install the Strimzi operator, cert-manager, Apicurio Registry and Kafka UI, beside the parts in High-Level Architecture.
| Component | Technology | Version | Purpose |
|---|---|---|---|
| Kates API | Quarkus | 3.x | REST + gRPC framework, CDI, native compilation |
| Runtime | Java | 21+ | Virtual threads, modern GC |
| Build | Maven | 3.x | Kates API build system |
| CLI | Go | 1.25+ | Cross-platform binary |
| CLI Framework | Cobra | Latest | Command parsing, help generation |
| Cluster | Kind | Latest | Local Kubernetes simulation |
| Kafka | Apache Kafka | 4.3.1 | KRaft mode, Share Groups |
| Operator | Strimzi | 1.2.0 | Kafka lifecycle management |
| Chaos | LitmusChaos | Latest | Runs the default chaos provider’s experiments |
| Monitoring | Prometheus + Grafana | Latest | Metrics collection and visualization |
| Tracing | Jaeger (OTLP) | 2.15.0 | Distributed trace collection |
| Registry | Apicurio | Latest | Schema registry for Kafka |
| Database | PostgreSQL | Latest | Test results and schedule persistence |
| Backup | Velero + SeaweedFS | Latest | Cluster backup and restore |
| Policy Engine | Kyverno | Latest | Admission control, PSS enforcement, NetworkPolicy generation |
The pinned versions above are a snapshot for orientation; the Version & Compatibility Matrix is generated from versions.env and the charts, and it wins when the two disagree.
2.8 Where Results Live
A run’s results live in PostgreSQL, not in Kafka, and the rest of what you see about a run lives somewhere with a shorter life. Knowing which is which explains why a Grafana board goes quiet after a run while kates trend can still reach every DONE run the Kates API has kept, and why a run’s heatmap can be missing.
The table lists what Kates keeps about a run, where each piece lives and how long it lasts:
| What | Where it lives | How long it lasts |
|---|---|---|
| The run: its spec, status, per-task results, an INTEGRITY task’s integrity result among them, and SLA | PostgreSQL: the kates chart’s own, beside the Kates API, or an external database |
90 days; once a day the Kates API deletes finished runs older than that, and kates test prune deletes them sooner |
| Disruption reports, schedules, webhooks and audit events | PostgreSQL | No automatic expiry |
| A baseline: the run each test type is compared with | PostgreSQL | The baseline never expires; the run it names is deleted at 90 days like any other |
The report of a finished run, as report show prints it |
Built from the stored run when first asked for, then kept in the Kates API’s memory | The 200 most recently used, until the pod restarts; then built again |
| A resilience run’s report | The response to kates resilience run |
Not stored; the test run inside it is kept like any other |
| Heatmap rows, for native runs | The Kates API’s memory | The 50 most recent runs, until the pod restarts |
| Per-run meters | The Kates API’s /q/metrics, scraped into Prometheus |
Until the run ends; Prometheus keeps the samples for its retention |
Security audit grades behind security trend |
The Kates API’s memory | The last 100 audits, until the pod restarts |
| Contexts, profiles, snapshots and saved Lab sessions | Files in your home directory | Until you delete them |
Two consequences follow. A Grafana board built on per-run meters stops getting data when the run ends, while kates trend and report show still answer, because the runs they read are in PostgreSQL. And a report isn’t a record. The Kates API builds a finished run’s report the first time something asks for it and keeps it in memory. The broker figures that report brokers prints therefore describe the topic’s partition leaders at that moment, not during the run, and a restart can change them.
No Kafka topic holds a run’s results. The Kates API writes them only to PostgreSQL. When a run changes status, it also writes a lifecycle event to an outbox table in the same transaction, for a poller to publish to Kafka, the transactional outbox pattern [46]. The event carries the run’s ID, type, status and time, never its results. The topics that the kafka-cluster chart’s platform profile creates on krafter, such as kates-results, hold none of them either.
The events go to kates-test-events, a topic the Kates API creates for itself on the cluster it tests (Topics gives its settings), and the Kates API reads them back from it. A DONE or FAILED event read there is what calls the registered webhooks, so a webhook fires only once its event has made the round trip through Kafka. The poller deletes an event from the outbox when the broker acknowledges it. While Kafka is unreachable the events wait in the outbox and the poller retries them. After kates.outbox.max-attempts failed sends, 10 by default and at least 10 minutes when no broker answers, it moves an event to the outbox_dead_letters table, and that event’s webhooks never fire.
Every replica of the Kates API reads kates-test-events in one consumer group, kates-webhooks, so one replica handles each event, and a replica that starts resumes from the group’s committed offset. A group with no committed offset, as on its first start, reads from the oldest event the topic holds. The processed_events table remembers the events handled in the last 7 days (kates.outbox.processed-events-retention-days), so an event read twice fires its webhooks once, and the Kates API skips any event older than that.
What lives in the Kates API’s memory goes with its pod. When the Kates API shuts down, it stops the runs in flight and stores them as FAILED with the error Server shutdown. When the pod dies without shutting down, those runs stay RUNNING until the Kates API starts again, and it marks them FAILED as it starts, each task that had not finished with the error Recovered: test was orphaned after server restart. In both cases, a run whose tasks were still being started stays PENDING until five minutes after the time it was set to last, when the Kates API fails it with an error that starts Timeout: still pending. Either way, each task keeps the figures it last recorded, and every heatmap and every built report is gone. kates clean goes further: it deletes the namespace that holds the bundled PostgreSQL, and every stored run with it.
2.9 Data Model
Kates uses PostgreSQL for persistent storage. The schema is managed by Flyway migrations in kates/src/main/resources/db/migration/.
erDiagram
test_runs ||--o{ test_results : "has many"
test_runs {
varchar id PK
varchar test_type
varchar status
timestamptz created_at
varchar backend
varchar scenario_name
text spec_json
text requested_spec_json
text sla_json
text labels_json
}
test_results {
bigserial id PK
varchar test_run_id FK
varchar test_type
varchar status
bigint records_sent
double throughput_rec_per_sec
double avg_latency_ms
double p50_latency_ms
double p95_latency_ms
double p99_latency_ms
double max_latency_ms
varchar phase_name
}
scheduled_test_runs {
varchar id PK
varchar name
varchar cron_expression
boolean enabled
text request_json
varchar last_run_id
timestamptz created_at
}
disruption_reports {
varchar id PK
varchar plan_name
varchar status
varchar sla_grade
timestamptz created_at
text report_json
}
disruption_schedules {
varchar id PK
varchar name
varchar cron_expression
boolean enabled
varchar playbook_name
text plan_json
varchar last_run_id
timestamptz created_at
}
Migration History
| Version | File | Purpose |
|---|---|---|
| V1 | V1__create_test_tables.sql |
test_runs + test_results with indexes on type, status, created_at |
| V2 | V2__create_schedules_table.sql |
scheduled_test_runs for recurring test automation |
| V3 | V3__create_disruption_reports.sql |
disruption_reports with SLA grade tracking |
| V4 | V4__create_disruption_schedules.sql |
disruption_schedules for recurring disruptions |
| V5 | V5__create_audit_events.sql |
audit_events table for action/event auditing |
| V6 | V6__create_webhook_deliveries.sql |
webhook_deliveries tracking for outbound webhooks |
| V7 | V7__create_profiles.sql |
profiles table storing baseline performance profiles |
| V8 | V8__create_snapshots.sql |
snapshots table for cluster topology captures |
| V9 | V9__add_test_tags.sql |
tags JSONB column + GIN index on test_runs |
| V10 | V10__add_composite_indexes.sql |
Composite indexes for test_type + status queries |
| V11 | V11__labels_jsonb.sql |
JSONB labels column for flexible test categorization |
| V12 | V12__cdc_phases_jsonb.sql |
cdc_phases_json column on test_runs |
| V13 | V13__create_outbox_events.sql |
outbox_events table for the transactional outbox pattern |
| V14 | V14__create_processed_events.sql |
processed_events idempotency-key table for consumer dedup |
| V15 | V15__create_webhook_dlq.sql |
webhook_dlq dead-letter table for failed webhook deliveries |
2.10 Graceful Degradation
Distributed systems fail in partial ways [62]. Kates is designed to degrade gracefully rather than crash catastrophically when its dependencies become unavailable.
| Failure Scenario | What Happens | Recovery |
|---|---|---|
| Kafka unreachable | Running tests fail with a connection error. The CLI reports Kafka: ❌ Disconnected in health checks. No new tests can be created until Kafka is reachable. Existing test results in PostgreSQL remain accessible. |
The Kates API reconnects automatically when Kafka becomes available. No manual intervention required. |
| PostgreSQL down | Runs and their results live only in PostgreSQL, so every command that creates or reads a run fails. No Kafka topic holds a copy. | Bring PostgreSQL back. |
| You cancel a run | The Kates API stops the run’s tasks and stores it as FAILED, each unfinished task with the error Cancelled by user. |
Run it again; the cancelled run stays in history for comparison. |
| The Kates API pod restarts during a test | A clean shutdown stores runs in flight as FAILED (Server shutdown). After a crash they stay RUNNING until the Kates API is back, which fails them as it starts. Their tasks keep the figures recorded so far. |
Kubernetes restarts the pod; rerun the test. Its heatmap is lost either way. |
Where Results Live says what survives a restart.
Try it
Trace one test through every layer of the Data Flow diagram:
kates health
kates test create --type LOAD --records 100000 --wait
kates test list
kates report show <id>health confirms the CLI reaches the REST API and Kafka; create --wait drives the TestOrchestrator and NativeKafkaBackend until the run completes; test list shows the persisted TestRun; and report show (with the ID printed by create) returns the ReportGenerator’s summary, its SLA section, and a broker snapshot.
2.11 Summary
- Kates has two halves: the Go CLI on your machine and the Kates API, one Quarkus service in the cluster that serves REST and gRPC.
- The
TestOrchestratorowns a run’s lifecycle: resolve defaults, create the topic, launch workers, poll status and collect heatmap rows; reports are built on request. - Latency measurement rests on HdrHistogram (1µs–60s, microsecond precision), compressed into 25 heatmap buckets for export.
- The safety guard refuses a disruption plan that would hit too many brokers, and rollback undoes some faults; a resilience run skips both.
- Results live in PostgreSQL, not Kafka; heatmaps, security grades and built reports live only in the Kates API’s memory.
The Cluster Under Test comes next: it builds the Kubernetes and Kafka environment that the Kates API runs in and tests.