Kates — The Definitive Guide
Kafka Advanced Testing & Engineering Suite
A comprehensive guide to performance testing, chaos engineering, and operational resilience for Apache Kafka — from theory to practice.
Table of contents
Preface
Kates answers three questions about an Apache Kafka cluster: how much load it takes before latency degrades, what happens when parts of it fail, and whether it loses data along the way. This book teaches you to ask those questions and to read the answers.
What This Book Covers
Kates (Kafka Advanced Testing & Engineering Suite) is a Kubernetes-native platform for performance testing, chaos engineering, and operational resilience auditing of Apache Kafka clusters.
This guide takes you from first principles to production operations:
- Part I — Foundations: What Kates is, its architecture, and the cluster under test
- Part II — Performance Testing: Measurement theory, the test types, scenario files with SLA gates, and the interactive Lab
- Part III — Chaos & Integrity: Chaos engineering theory and practice, and zero-loss verification
- Part IV — Observability: Dashboards, metrics, heatmaps, tracing, and alerting
- Part V — Deployment & Operations: Installing and engineering Kafka, deploying Kates, security, multi-tenancy, upgrades, Kafka Connect, and operational recipes
- Part VI — Reference: The complete CLI, REST API, and gRPC API references
Who This Book Is For
| Reader | Start Here | Focus On |
|---|---|---|
| Platform Engineer deploying Kafka | Installing Kafka → Kafka Deployment Engineering → Deployment Guide | Deployment, installation, and infrastructure |
| SRE / Reliability Engineer | Chaos Engineering Theory → Chaos Engineering in Practice → Observability & Monitoring | Chaos engineering, observability, and resilience |
| Performance Engineer | Performance Theory → Test Types Deep Dive → CLI Reference | Performance theory, test types, and CLI workflows |
| Security Engineer | Security & Compliance → Troubleshooting Index | Security, compliance, and troubleshooting |
| Developer using the API | REST API Reference → gRPC API Reference → Scenario Files & SLA Gates | REST API, gRPC API, and scenario files |
How to Read This Book
The first four Parts build on one another, so read them in order. Part I — Foundations is the shared base: Architecture & Design builds the model that later chapters assume, and The Cluster Under Test describes the cluster behind every number you measure. Part II — Performance Testing, Part III — Chaos & Integrity and Part IV — Observability build on that base and on each other. Part III reuses the test types of Part II to measure what a fault costs, and Part IV explains why a run produced the numbers it did.
Part V — Deployment & Operations is for the day you deploy, secure, upgrade or extend a cluster, and Part VI — Reference and the appendices are for looking things up. Each Part opens with a page that says what the Part is for, lists its chapters with the question each one answers, and links the tutorials that go with them.
One question runs through Parts I to V: is krafter, the Kafka cluster under test, ready for a payments workload? You run the payments platform, and before it moves onto krafter you have to say whether the cluster can carry it. The book turns that into targets that a Kates run can check:
kraftertakes 2,000 records per second of 1 KiB withacks=all, and answers writes with a P99 of 100 ms or less from send to acknowledgment.krafterloses no acknowledged record when one broker or one zone fails.
Kates calls thresholds like these SLAs, though they’re SLO-style targets that you set, not agreements with anyone. Part I builds the lab and a first run. Part II checks the first target and Part III the second, Part IV explains the numbers both produce, and Part V moves the question from the lab to a cluster you run. On the pages of Parts I to V, The Payments Question says what each Part adds toward the answer.
Before You Start
The book’s exercises run against a lab on your own machine, so set it up before you reach the Quick Start in Introduction.
You don’t need to know Kafka’s internals first. The Cluster Under Test explains the replication settings, acks, the in-sync replicas and the controller quorum behind every test, in its Replication Configuration section, and the Glossary defines the terms the book uses. You do need to be at home at a shell prompt and with kubectl.
The Deployment Guide’s Prerequisites table lists the tools the lab needs, with their versions. For the machine, Installing Kafka with the kafka-cluster Helm Chart asks for a Kubernetes cluster with at least 16 GB of memory and 6 CPU cores, and recommends 24 GB and 12 cores. On Kind, the memory and cores come from Docker.
From a checkout of the repository, make all builds the lab. When your current kubectl context can’t reach a cluster, it creates the panda Kind cluster first. It then asks you to choose a topology for Kafka, Kates and the chaos tooling: 1 puts them together in the kates-stack namespace, and 2 gives each a namespace of its own. Choose 2, which the Quick Start assumes. kates deploy then checks which clusters your kubeconfig reaches. When only the current context answers, it deploys there. When several answer, it asks which one to use, starting on the current context, and when only another cluster answers, it asks before deploying there. The cluster it deploys to becomes kubectl’s current context, and when that changes, kates deploy prints the kubectl config use-context command that switches back.
Two names recur throughout the book. panda is the Kind cluster: its three nodes are named after the zones they stand for, alpha, sigma and gamma, and Kafka spreads its brokers across them. krafter is the Kafka cluster under test, which The Cluster Under Test describes.
The Kates Tutorials are for practice, and the book is for explanation. Each tutorial walks through one workflow on the lab, step by step, while the chapters explain why it works; each Part page links the tutorials that go with its chapters.
Versions This Book Targets
The book tracks the versions pinned in the repository — Kafka, Strimzi, and the Kates charts themselves. The consolidated matrix lives in the Version & Compatibility Matrix appendix, generated from versions.env; when a claim in a chapter and the matrix disagree, the matrix wins.
Conventions
The same conventions hold in every chapter. The table says what each one means:
| Convention | What it means |
|---|---|
| A shell command | It carries no prompt character, so copy it as it is. |
A kubectl command |
It assumes a kubeconfig pointing at your cluster. |
A kates command |
It assumes the CLI is installed and configured with a context that carries the API key, as the Quick Start in Introduction sets up. |
A curl or grpcurl example |
It assumes the API is forwarded to localhost:30083, and reads the key from $KATES_API_KEY, exported as shown in REST API Reference. |
A code block whose first line is a comment naming a file, such as # kafka-eks.yaml |
A file you create or edit. |
A text block after a line that reads Output: |
What the command before it printed. IDs, times and figures differ from run to run. |
<id> in a command |
A placeholder for an ID that an earlier command printed, such as a test run’s. |
| LOAD, ROUND_TRIP and the other test types in capitals | Kates test types, spelled as the API spells them. Lowercase “round-trip” is the general idea of round-trip latency. |
| A Note callout | Context worth knowing. |
| A Tip callout | A shortcut or a good practice. |
| An Important callout | A constraint you must not miss. |
| A Warning callout | A risk of breaking something. |
| A Caution callout | A risk of losing data or causing downtime. |
| A Tip callout that opens with Try it | An exercise to run on your lab. |
| “About N minutes” on a Part page | The chapter’s prose read at 200 words a minute and rounded to 5 minutes, code blocks left out, so time at the keyboard comes on top. |
Source Code
The complete source code for Kates is available on GitHub:
- Repository: github.com/bmscomp/kates
- License: Apache 2.0