2  Architecture & Design

Kates has two halves: the kates CLI on your machine, and the Kates API in the cluster. Around them, kates deploy installs Kafka, PostgreSQL, LitmusChaos, Prometheus and Grafana, and this chapter shows how the parts fit. It serves anyone who operates, extends, or debugs Kates — the mental model built here underpins every later chapter. After this chapter, you can:

2.1 High-Level Architecture

The half of Kates that does the work isn’t on your machine. The kates CLI runs where you type; the Kates API runs in the cluster, as one Deployment beside the Kafka cluster it tests. It generates the load, asks for the faults, grades disruption plans and keeps every run in PostgreSQL.

The diagram shows the main parts of a default kates deploy and who calls whom. Solid arrows are the default paths; the dashed ones are the Trogdor path, which kates deploy doesn’t set up.

flowchart LR
    subgraph Machine["Your machine"]
        CLI["kates CLI<br/>commands, Lab, kates mcp"]
    end
    subgraph Cluster["Kubernetes cluster"]
        API["Kates API pod<br/>REST and gRPC, native benchmark backend, chaos provider"]
        PG[("PostgreSQL")]
        KAFKA["Kafka cluster krafter"]
        LIT["LitmusChaos"]
        TROG["Trogdor coordinator<br/>optional, not installed"]
        PROM["Prometheus and Grafana"]
        K8S["Kubernetes API"]
    end
    CLI -->|"REST, API key"| API
    CLI -->|"kubectl, helm"| K8S
    API -->|"runs and results"| PG
    API -->|"producers, consumers, AdminClient"| KAFKA
    API -->|"ChaosEngine resources"| LIT
    LIT -->|"faults"| KAFKA
    API -.->|"trogdor benchmark backend"| TROG
    TROG -.-> KAFKA
    PROM -->|"scrapes"| API
    PROM -->|"scrapes"| KAFKA
Figure 2.1: The CLI on your machine drives the Kates API in the cluster, which generates load on Kafka, keeps runs in PostgreSQL and asks LitmusChaos for faults.

Read it from the left. The CLI calls the Kates API over REST, with the URL and API key of your CLI context. For the commands that install and reach the stack, such as kates deploy and kates ports, it runs kubectl and helm against your current Kubernetes context instead. The Commands table in CLI Reference says which command takes which path. In the cluster, the Kates API talks to krafter, the Kafka cluster under test, as an ordinary Kafka client, and keeps each run in PostgreSQL. On the default chaos provider, it creates LitmusChaos resources for every fault except POD_DELETE, ROLLING_RESTART and SCALE_DOWN, which it injects itself through the Kubernetes API. Prometheus scrapes the brokers and the Kates API, and the Kates API reads Prometheus back for a disruption plan’s broker metrics.

Each part in the table runs in one place. The namespaces are those of the default isolated topology; with --topology single, every part below that runs in the cluster, except Trogdor, shares kates-stack, as Single-Namespace vs Multi-Namespace describes.

Part What it does Where it runs
kates CLI Calls the Kates API, and runs kubectl and helm for the commands that install, reach and remove the stack Your machine
Lab, dashboard, top and kafka tui Terminal views on the Kates API; every run Lab starts, warm-ups and median runs included, is one test run Your machine, inside the CLI
kates mcp Serves read-only tools to an AI agent, whose MCP client starts it Your machine, inside the CLI
Kates API Serves REST and gRPC on port 8080, and runs tests, disruptions and schedules Deployment kates, namespace kates
Native benchmark backend Runs a test’s producers and consumers on virtual threads inside the Kates API The Kates API pod
Trogdor benchmark backend Sends a test’s workload to a Trogdor coordinator, which kates deploy doesn’t install The Kates API pod; the load comes from the Trogdor agents you run
Chaos provider Turns each fault into LitmusChaos resources or direct Kubernetes API calls The Kates API pod
LitmusChaos Runs the experiments the chaos provider asks for Operator in litmus; ChaosEngines in kafka
PostgreSQL Keeps runs and their results, disruption reports, schedules, baselines and audit events StatefulSet kates-postgresql, namespace kates
krafter The Kafka cluster under test Namespace kafka
Prometheus and Grafana Scrape and chart the brokers and the Kates API Namespace monitoring

The Kates API is one process, and that has a cost. A native run’s producers and consumers share the pod’s CPU and memory with everything else the Kates API does. kates deploy limits that pod to one CPU and 512 MiB of memory, on Kind and on other clusters alike. A throughput ceiling you measure can therefore be the pod’s rather than krafter’s.

The Kates API comes as two images built from the same code: a JVM image and a GraalVM native image. On a Kind cluster, kates deploy runs a native image that you built or pulled onto your machine, kates:native-local or else kates:native, and never pulls one itself. Without either, it stops at the Kates API step and names the command that builds one. On any other cluster, it runs the published JVM image. The two collect garbage differently: the native image’s Serial GC stops the Kates API for every collection, and those pauses land in the latencies it records. Native Image Build explains when to run which.

2.2 The Kates API

The Kates API is a Quarkus application. It exposes both a REST API and a gRPC API (see gRPC API Reference), and manages the full test lifecycle. Both APIs delegate to the same service layer, so a test run behaves the same whichever API starts it; the gRPC API covers fewer operations, and some of its responses carry less than their REST counterparts, as its chapter lists.

Component Map

graph LR
    subgraph Domain
        TR[TestRun]
        TS[TestSpec]
        TT[TestType]
        TRes[TestResult]
    end
    
    subgraph Engine
        TO[TestOrchestrator]
        NKB[NativeKafkaBackend]
        LH[LatencyHistogram]
        BS[BenchmarkStatus]
        BM[BenchmarkMetrics]
    end
    
    subgraph Report
        RG[ReportGenerator]
        CS[ClusterSnapshot]
        BrokerM[BrokerMetrics]
        SLA[SlaVerdict]
    end
    
    subgraph Export
        CSV[CsvExporter]
        JUnit[JunitXmlExporter]
        HM[HeatmapExporter]
    end
    
    TO --> NKB
    NKB --> LH
    NKB --> BS
    RG --> CS
    RG --> BrokerM
    RG --> SLA
    RG --> CSV
    RG --> JUnit
    RG --> HM

TestOrchestrator

The TestOrchestrator is the central coordinator. When a test is created, it:

  1. Resolves defaults — merges the incoming TestSpec with TestTypeDefaults for the chosen test type
  2. Creates the topic — ensures the Kafka topic exists with the required partition count and replication factor
  3. Launches workers — delegates to the BenchmarkBackend to start producer/consumer tasks
  4. Polls status — periodically polls each BenchmarkHandle for BenchmarkStatus updates
  5. Collects heatmap data — on each poll of a running test, stores the latency buckets recorded so far as one heatmap row

The Kates API builds a report only when something asks for one. It builds a finished run’s report from the stored run the first time it’s asked (ReportGenerator), with summary metrics, whether the run met its SLA thresholds (targets you set, in the style of an SLO) and broker correlation.

NativeKafkaBackend

The native benchmark backend (NativeKafkaBackend) runs each test’s producers and consumers on virtual threads inside the Kates API [9]. It is one of two benchmark backends; the other, trogdor, sends the workload to a Trogdor coordinator, and Test Types Deep Dive compares them. Its workers:

  • Produce messages with configurable record size, acknowledgment mode, and throughput throttling
  • Consume messages with configurable consumer group, fetch settings, and poll timeout
  • Record latency in a LatencyHistogram backed by HdrHistogram (1µs–60s range, microsecond precision)
  • Track integrity — sequence numbers, acknowledgment gaps, and consumer-side deduplication

LatencyHistogram

The histogram is the heart of latency measurement. It is backed by HdrHistogram (configured for 1µs–60s with 3 significant value digits), which provides high resolution at low latencies (sub-millisecond) while covering tails up to 60 seconds.

graph LR
    subgraph Internal["HdrHistogram (1µs–60s)"]
        direction TB
        B1["0.001ms"] --> B2["0.01ms"] --> B3["0.1ms"] --> B4["1ms"] --> B5["10ms"] --> B6["100ms"] --> B7["1000ms"]
    end
    
    subgraph Export["25 Heatmap Buckets"]
        direction LR
        H1["0–0.1ms"]
        H2["0.1–0.5ms"]
        H3["0.5–1ms"]
        H4["1–5ms"]
        H5["5–50ms"]
        H6["50–500ms"]
        H7["500ms–10s"]
    end
    
    Internal -->|exportBuckets| Export

Key methods:

Method Lock Purpose
recordLatency(latencyMs) Write Record a single latency observation
getPercentile(p) Read Compute P50/P95/P99 from cumulative distribution
exportBuckets() Read Compress to 25 heatmap ranges (non-destructive)
snapshotAndReset() Write Atomic capture + reset for windowed collection

2.3 Disruption Engine

A disruption plan injects faults step by step and watches the cluster from outside, while a resilience run, started with kates resilience run, injects one fault under a Kates test and measures what a client sees. Both get the fault from the Kates API’s chaos provider, LitmusChaos by default, as Chaos Engineering in Practice explains.

The diagram shows the parts behind a disruption plan. Look at the Providers group: one setting picks the chaos provider that every fault goes through.

graph LR
    subgraph Control
        DO[DisruptionOrchestrator]
        DSG[DisruptionSafetyGuard]
        DPC[DisruptionPlaybookCatalog]
    end
    
    subgraph Intelligence
        KIS[KafkaIntelligenceService]
        ISR[ISR Tracking]
        LAG[Consumer Lag]
        LEAD[Leader Resolution]
    end
    
    subgraph Providers
        CP{{kates.chaos.provider}}
        HCP[HybridChaosProvider]
        KCP[KubernetesChaosProvider]
        LCP[LitmusChaosProvider]
    end
    
    subgraph Reporting
        DR[DisruptionReport]
        SG[SlaGrader]
        PMC[PrometheusMetricsCapture]
    end
    
    DO --> DSG
    DO --> KIS
    DO --> CP
    DO --> DR
    DPC --> DO
    KIS --> ISR
    KIS --> LAG
    KIS --> LEAD
    CP -->|"litmus-crd, the default"| LCP
    CP -->|kubernetes| KCP
    CP -->|hybrid| HCP
    HCP -.->|picks one| LCP
    HCP -.->|picks one| KCP
    LCP -->|"POD_DELETE,<br/>ROLLING_RESTART,<br/>SCALE_DOWN"| KCP
    DR --> SG
    DR --> PMC
Figure 2.2: A disruption plan’s orchestrator checks the plan, reads Kafka’s state, and hands every fault to the one chaos provider that kates.chaos.provider selects.

Disruption Types

Kates’s disruption types run from killing a broker pod to draining a node, and kates disruption types lists them. What each one does to a pod or to the network depends on the Kates API’s chaos provider, LitmusChaos by default, so each type is described once, per provider, in Chaos Engineering in Practice.

Safety Guardrails

Before a disruption plan touches the cluster, Kates counts the brokers its steps would hit, and refuses the plan if that’s every broker, or more than the plan’s maxAffectedBrokers. Before each fault it checks that every Kafka pod is Running and Ready. When a step fails with autoRollback on, rollback gives a scaled-down node pool its broker back, or deletes the NetworkPolicies Kates created for a network partition. The safety guard doesn’t read the ISR or the KRaft quorum, and a resilience run doesn’t go through it: Chaos Engineering in Practice says exactly what it checks.

2.4 CLI Architecture

The CLI is a standalone Go binary built with Cobra. It splits its work between two channels: it calls the Kates API over REST for test, report, and disruption operations, and it shells out to kubectl, helm, and kind (via os/exec) for cluster provisioning and lifecycle tasks such as kates deploy, kates clean, and kates ports. It also keeps your contexts, profiles and snapshots in files in your home directory; the Commands table in CLI Reference says which command uses which.

graph TD
    subgraph Config
        CTX[~/.kates.yaml]
        CTXM[Context Manager]
    end
    
    subgraph Commands
        TEST[test create/list/get/delete/watch/apply/scaffold]
        REPORT[report show/summary/export/compare/diff/brokers]
        DISRUPT[disruption run/list/status/timeline/types/kafka-metrics]
        RESIL[resilience run]
        TREND[trend]
        OPS[health/cluster/top/dashboard/status]
    end
    
    subgraph Output
        TABLE[Table Renderer]
        JSON[JSON Printer]
        SPARK[Sparkline Charts]
        BADGE[Status Badges]
        BAR[Metric Bars]
    end
    
    CTX --> CTXM
    CTXM --> Commands
    Commands --> TABLE
    Commands --> JSON
    Commands --> SPARK

Key design decisions:

  • Multi-context support — like kubectl, the CLI supports named contexts for targeting different Kates APIs
  • Rich terminal output — tables, colored badges, metric bars, sparkline charts, and ASCII banners
  • Scaffold templates — kates test scaffold export <name> writes a ready-to-use YAML scenario file to the current directory (browse the library with kates test scaffold list, optionally filtered by --type LOAD)
  • Streaming watch — kates test watch and kates disruption watch provide real-time progress updates (disruption watch receives no events for a disruption ID yet; see its entry in the CLI reference)

2.5 Data Flow

This diagram traces a complete test execution from CLI command to final report:

sequenceDiagram
    participant CLI as Kates CLI
    participant API as REST API
    participant Orch as TestOrchestrator
    participant Engine as NativeKafkaBackend
    participant Kafka as Kafka Cluster
    participant Hist as LatencyHistogram
    participant Report as ReportGenerator
    
    CLI->>API: POST /api/tests {type: LOAD, spec: {...}}
    API->>Orch: createTest(type, spec)
    Orch->>Kafka: Create topic (if needed)
    Orch->>Engine: start(handles)
    Engine->>Kafka: Produce messages
    Engine->>Hist: record(latencyUs)
    
    loop Every poll interval
        Orch->>Engine: poll(handle)
        Engine->>Hist: exportBuckets()
        Engine-->>Orch: BenchmarkStatus + heatmapBuckets
        Orch->>Orch: Accumulate heatmap rows
    end
    
    CLI->>API: GET /api/tests/{id}
    API->>Orch: getTest(id)
    Orch-->>CLI: TestRun (status, results)
    
    CLI->>API: GET /api/tests/{id}/report
    API->>Report: generate(testRun)
    Report->>Kafka: captureSnapshot (broker metrics)
    Report-->>CLI: TestReport (summary, SLA, brokers)

2.6 Disruption Pipeline

The disruption path adds two things the test path lacks: a check that can refuse the plan, and a rollback when a step fails. This diagram traces a plan from kates disruption run to its report:

sequenceDiagram
    autonumber
    participant Cli as kates CLI
    participant API as Kates API
    participant Kube as Kubernetes API
    participant Prov as Chaos provider
    participant Prom as Prometheus
    Cli->>API: POST /api/disruptions
    API->>Kube: List the Kafka pods, count the brokers the plan hits
    API-->>Cli: 202 with a disruption ID, 422 refused, or 409 busy
    loop Each step
        API->>API: If the step names a topic, find its leader, then wait steadyStateSec
        API->>Kube: Are all Kafka pods Running and Ready?
        API->>Prom: Baseline snapshot
        API->>Prov: Inject the fault
        Prov-->>API: Pass, Fail or Skipped
        API->>API: Wait observationWindowSec
        API->>Prom: Impact snapshot
        opt requireRecovery
            API->>Kube: Wait for the Kafka pods to be Ready again
        end
        opt Recovery failed or the step threw, with autoRollback on
            API->>Kube: Restore node pool replicas, or delete the partition policies Kates made
        end
    end
    API->>API: Grade against the sla block, if it has one
    Cli->>API: GET /api/disruptions/{id}
    API-->>Cli: The report, graded when the plan has an sla block
Figure 2.3: Kates checks a disruption plan before anything runs; each step then checks the Kafka pods, injects one fault and watches the recovery, and rollback runs only when a step fails.

2.7 Technology Stack

The table lists what Kates is built with and the tools it works beside. A default kates deploy installs neither Jaeger, Velero and its SeaweedFS object store, nor Kyverno. It does install the Strimzi operator, cert-manager, Apicurio Registry and Kafka UI, beside the parts in High-Level Architecture.

Component Technology Version Purpose
Kates API Quarkus 3.x REST + gRPC framework, CDI, native compilation
Runtime Java 21+ Virtual threads, modern GC
Build Maven 3.x Kates API build system
CLI Go 1.25+ Cross-platform binary
CLI Framework Cobra Latest Command parsing, help generation
Cluster Kind Latest Local Kubernetes simulation
Kafka Apache Kafka 4.3.1 KRaft mode, Share Groups
Operator Strimzi 1.2.0 Kafka lifecycle management
Chaos LitmusChaos Latest Runs the default chaos provider’s experiments
Monitoring Prometheus + Grafana Latest Metrics collection and visualization
Tracing Jaeger (OTLP) 2.15.0 Distributed trace collection
Registry Apicurio Latest Schema registry for Kafka
Database PostgreSQL Latest Test results and schedule persistence
Backup Velero + SeaweedFS Latest Cluster backup and restore
Policy Engine Kyverno Latest Admission control, PSS enforcement, NetworkPolicy generation

The pinned versions above are a snapshot for orientation; the Version & Compatibility Matrix is generated from versions.env and the charts, and it wins when the two disagree.

2.8 Where Results Live

A run’s results live in PostgreSQL, not in Kafka, and the rest of what you see about a run lives somewhere with a shorter life. Knowing which is which explains why a Grafana board goes quiet after a run while kates trend can still reach every DONE run the Kates API has kept, and why a run’s heatmap can be missing.

The table lists what Kates keeps about a run, where each piece lives and how long it lasts:

What Where it lives How long it lasts
The run: its spec, status, per-task results, an INTEGRITY task’s integrity result among them, and SLA PostgreSQL: the kates chart’s own, beside the Kates API, or an external database 90 days; once a day the Kates API deletes finished runs older than that, and kates test prune deletes them sooner
Disruption reports, schedules, webhooks and audit events PostgreSQL No automatic expiry
A baseline: the run each test type is compared with PostgreSQL The baseline never expires; the run it names is deleted at 90 days like any other
The report of a finished run, as report show prints it Built from the stored run when first asked for, then kept in the Kates API’s memory The 200 most recently used, until the pod restarts; then built again
A resilience run’s report The response to kates resilience run Not stored; the test run inside it is kept like any other
Heatmap rows, for native runs The Kates API’s memory The 50 most recent runs, until the pod restarts
Per-run meters The Kates API’s /q/metrics, scraped into Prometheus Until the run ends; Prometheus keeps the samples for its retention
Security audit grades behind security trend The Kates API’s memory The last 100 audits, until the pod restarts
Contexts, profiles, snapshots and saved Lab sessions Files in your home directory Until you delete them

Two consequences follow. A Grafana board built on per-run meters stops getting data when the run ends, while kates trend and report show still answer, because the runs they read are in PostgreSQL. And a report isn’t a record. The Kates API builds a finished run’s report the first time something asks for it and keeps it in memory. The broker figures that report brokers prints therefore describe the topic’s partition leaders at that moment, not during the run, and a restart can change them.

No Kafka topic holds a run’s results. The Kates API writes them only to PostgreSQL. When a run changes status, it also writes a lifecycle event to an outbox table in the same transaction, for a poller to publish to Kafka, the transactional outbox pattern [46]. The event carries the run’s ID, type, status and time, never its results. The topics that the kafka-cluster chart’s platform profile creates on krafter, such as kates-results, hold none of them either.

The events go to kates-test-events, a topic the Kates API creates for itself on the cluster it tests (Topics gives its settings), and the Kates API reads them back from it. A DONE or FAILED event read there is what calls the registered webhooks, so a webhook fires only once its event has made the round trip through Kafka. The poller deletes an event from the outbox when the broker acknowledges it. While Kafka is unreachable the events wait in the outbox and the poller retries them. After kates.outbox.max-attempts failed sends, 10 by default and at least 10 minutes when no broker answers, it moves an event to the outbox_dead_letters table, and that event’s webhooks never fire.

Every replica of the Kates API reads kates-test-events in one consumer group, kates-webhooks, so one replica handles each event, and a replica that starts resumes from the group’s committed offset. A group with no committed offset, as on its first start, reads from the oldest event the topic holds. The processed_events table remembers the events handled in the last 7 days (kates.outbox.processed-events-retention-days), so an event read twice fires its webhooks once, and the Kates API skips any event older than that.

What lives in the Kates API’s memory goes with its pod. When the Kates API shuts down, it stops the runs in flight and stores them as FAILED with the error Server shutdown. When the pod dies without shutting down, those runs stay RUNNING until the Kates API starts again, and it marks them FAILED as it starts, each task that had not finished with the error Recovered: test was orphaned after server restart. In both cases, a run whose tasks were still being started stays PENDING until five minutes after the time it was set to last, when the Kates API fails it with an error that starts Timeout: still pending. Either way, each task keeps the figures it last recorded, and every heatmap and every built report is gone. kates clean goes further: it deletes the namespace that holds the bundled PostgreSQL, and every stored run with it.

2.9 Data Model

Kates uses PostgreSQL for persistent storage. The schema is managed by Flyway migrations in kates/src/main/resources/db/migration/.

erDiagram
    test_runs ||--o{ test_results : "has many"
    
    test_runs {
        varchar id PK
        varchar test_type
        varchar status
        timestamptz created_at
        varchar backend
        varchar scenario_name
        text spec_json
        text requested_spec_json
        text sla_json
        text labels_json
    }
    
    test_results {
        bigserial id PK
        varchar test_run_id FK
        varchar test_type
        varchar status
        bigint records_sent
        double throughput_rec_per_sec
        double avg_latency_ms
        double p50_latency_ms
        double p95_latency_ms
        double p99_latency_ms
        double max_latency_ms
        varchar phase_name
    }
    
    scheduled_test_runs {
        varchar id PK
        varchar name
        varchar cron_expression
        boolean enabled
        text request_json
        varchar last_run_id
        timestamptz created_at
    }
    
    disruption_reports {
        varchar id PK
        varchar plan_name
        varchar status
        varchar sla_grade
        timestamptz created_at
        text report_json
    }
    
    disruption_schedules {
        varchar id PK
        varchar name
        varchar cron_expression
        boolean enabled
        varchar playbook_name
        text plan_json
        varchar last_run_id
        timestamptz created_at
    }

Migration History

Version File Purpose
V1 V1__create_test_tables.sql test_runs + test_results with indexes on type, status, created_at
V2 V2__create_schedules_table.sql scheduled_test_runs for recurring test automation
V3 V3__create_disruption_reports.sql disruption_reports with SLA grade tracking
V4 V4__create_disruption_schedules.sql disruption_schedules for recurring disruptions
V5 V5__create_audit_events.sql audit_events table for action/event auditing
V6 V6__create_webhook_deliveries.sql webhook_deliveries tracking for outbound webhooks
V7 V7__create_profiles.sql profiles table storing baseline performance profiles
V8 V8__create_snapshots.sql snapshots table for cluster topology captures
V9 V9__add_test_tags.sql tags JSONB column + GIN index on test_runs
V10 V10__add_composite_indexes.sql Composite indexes for test_type + status queries
V11 V11__labels_jsonb.sql JSONB labels column for flexible test categorization
V12 V12__cdc_phases_jsonb.sql cdc_phases_json column on test_runs
V13 V13__create_outbox_events.sql outbox_events table for the transactional outbox pattern
V14 V14__create_processed_events.sql processed_events idempotency-key table for consumer dedup
V15 V15__create_webhook_dlq.sql webhook_dlq dead-letter table for failed webhook deliveries

2.10 Graceful Degradation

Distributed systems fail in partial ways [62]. Kates is designed to degrade gracefully rather than crash catastrophically when its dependencies become unavailable.

Failure Scenario What Happens Recovery
Kafka unreachable Running tests fail with a connection error. The CLI reports Kafka: ❌ Disconnected in health checks. No new tests can be created until Kafka is reachable. Existing test results in PostgreSQL remain accessible. The Kates API reconnects automatically when Kafka becomes available. No manual intervention required.
PostgreSQL down Runs and their results live only in PostgreSQL, so every command that creates or reads a run fails. No Kafka topic holds a copy. Bring PostgreSQL back.
You cancel a run The Kates API stops the run’s tasks and stores it as FAILED, each unfinished task with the error Cancelled by user. Run it again; the cancelled run stays in history for comparison.
The Kates API pod restarts during a test A clean shutdown stores runs in flight as FAILED (Server shutdown). After a crash they stay RUNNING until the Kates API is back, which fails them as it starts. Their tasks keep the figures recorded so far. Kubernetes restarts the pod; rerun the test. Its heatmap is lost either way.

Where Results Live says what survives a restart.

Tip

Try it

Trace one test through every layer of the Data Flow diagram:

kates health
kates test create --type LOAD --records 100000 --wait
kates test list
kates report show <id>

health confirms the CLI reaches the REST API and Kafka; create --wait drives the TestOrchestrator and NativeKafkaBackend until the run completes; test list shows the persisted TestRun; and report show (with the ID printed by create) returns the ReportGenerator’s summary, its SLA section, and a broker snapshot.

2.11 Summary

  • Kates has two halves: the Go CLI on your machine and the Kates API, one Quarkus service in the cluster that serves REST and gRPC.
  • The TestOrchestrator owns a run’s lifecycle: resolve defaults, create the topic, launch workers, poll status and collect heatmap rows; reports are built on request.
  • Latency measurement rests on HdrHistogram (1µs–60s, microsecond precision), compressed into 25 heatmap buckets for export.
  • The safety guard refuses a disruption plan that would hit too many brokers, and rollback undoes some faults; a resilience run skips both.
  • Results live in PostgreSQL, not Kafka; heatmaps, security grades and built reports live only in the Kates API’s memory.

The Cluster Under Test comes next: it builds the Kubernetes and Kafka environment that the Kates API runs in and tests.