Appendix B — Troubleshooting Index

A consolidated index of troubleshooting procedures from across the book. Jump to the relevant section for step-by-step diagnosis and fixes.

B.1 Kafka Cluster

The cluster exists but isn’t healthy: the Strimzi operator, the brokers, the controller quorum or Cruise Control crash, stall, report an error or raise an alert.

Symptom Likely Cause Chapter
Strimzi operator CrashLoopBackOff with UnsupportedVersionException Local chart has mismatched Kafka image map Kafka Deployment Engineering
Brokers crash with ConfigException: Invalid value -1 for local.retention.bytes Kafka 4.1+ tightened validation — -1 rejected when retention.bytes is set Kafka Deployment Engineering
Brokers crash immediately with remote.log.storage.system.enable=true Tiered storage enabled without remote storage manager plugin JAR Kafka Deployment Engineering
KafkaActiveControllerCount != 1 alert Controller quorum lost or election in progress Kafka Deployment Engineering
Under-replicated partitions for extended period Broker disk I/O saturated, network issues, or follower falling behind The Cluster Under Test
Cruise Control unsupported goals error Goals list doesn’t match Strimzi’s default goals Kafka Deployment Engineering
Kafka CR stuck on NotReady (often with UnforceableProblem) Strimzi CRDs missing, insufficient resources, or operator egress blocked by operatorNetworkPolicy / isolated topology NetworkPolicy missing DNS/API server egress — can’t reach controllers Installing Kafka with the kafka-cluster Helm Chart, Kafka Deployment Engineering

B.2 Kafka Connectivity

A client — Kates, Kafka UI or your own — can’t reach the cluster or authenticate to it, or its credentials never arrive.

Symptom Likely Cause Chapter
Kafka UI CreateContainerConfigError — secret not found KafkaUser not applied before UI deployment Kafka Deployment Engineering
Kates can’t connect to Kafka Wrong bootstrap address or NetworkPolicy blocking Deployment Guide
SCRAM authentication failure Password rotated or KafkaUser not reconciled Security & Compliance
Connection timeout from new namespace The client namespace’s own egress policy, such as Kyverno’s generated default-deny; after the listeners carry networkPolicyPeers, a missing networkPolicy.clients entry Security & Compliance
KafkaUser secrets never created Entity Operator (User Operator) only starts after the Kafka CR reaches Ready Installing Kafka with the kafka-cluster Helm Chart

B.3 Kafka Connect

A Kafka Connect release misbehaves: its workers restart or keep rebalancing, a connector fails or is refused at render time, or the PostgreSQL database a Debezium connector reads grows on disk.

Symptom Likely Cause Chapter
KafkaConnector status FAILED with DebeziumException Wrong database credentials, replication slot still held, wal_level not logical, or schema.include.list matching no tables Operating Kafka Connect
Connect cluster stuck in REBALANCING — KafkaConnectRebalanceTooLong alert fires Workers crashing mid-rebalance, NetworkPolicy blocking inter-worker traffic on port 8083, or OOM kills during task assignment Operating Kafka Connect
Connect worker pods restart with OOMKilled Container memory limit under 2× the JVM heap — off-heap memory pushes usage over the limit Operating Kafka Connect
PostgreSQL disk usage grows while a connector is down or paused Replication slot retains WAL segments until the connector drains them Operating Kafka Connect
helm upgrade of the Connect chart fails with connect-cluster: connectors.<name> … A connector in the values misses a class, a required config key or topics, or breaks a production rule — checked at render time Operating Kafka Connect
KafkaConnector FAILED with Forbidden or secrets "…" is forbidden Chart 2.0 grants the workers get on only the Secrets its own connector configs reference — a connector applied outside the chart needs its Secret in rbac.secretNames Operating Kafka Connect

B.4 Performance Issues

Tests complete but the numbers look wrong — latency that regresses, splits in two or looks too good, results that vary between identical runs — or a broker raises a request-handler or log-flush alert.

Symptom Likely Cause Chapter
P99 latency regression between test runs Partition hotspot, GC pauses, or ISR changes Recipes & Patterns — Recipe 4
Bimodal latency distribution in heatmap Some requests hitting page cache, others going to disk Performance Theory
Artificially low latency measurements Coordinated omission — tool slows down with the system Performance Theory
Stress test results vary wildly between identical runs JVM warm-up (JIT), GC pauses, small sample size — increase records to 500K+, use ZGC, discard first 2–3 warm-up iterations Performance Theory, Deployment Guide
KafkaRequestHandlerSaturated alert Request handlers over 70% busy — add threads or brokers Kafka Deployment Engineering
KafkaLogFlushLatencyHigh alert Disk I/O saturated — check storage class and disk utilization Kafka Deployment Engineering

B.5 Deployment Issues

Pods, images or Helm releases fail while you install the stack or roll it out.

Symptom Likely Cause Chapter
Images won’t load into Kind Registry unreachable or platform mismatch (arm64/amd64) Deployment Guide
Kafka pods stuck in Pending StorageClass can’t provision PVCs, or no node matches the zone nodeAffinity rules Deployment Guide, Installing Kafka with the kafka-cluster Helm Chart
helm upgrade fails with another operation in progress A previous install or upgrade was interrupted — roll back the release, then retry Installing Kafka with the kafka-cluster Helm Chart
PDB blocks rolling restart Only 1 pod can be unavailable — intentional safety behavior Upgrade Playbook
Entity Operator never starts Kafka CR hasn’t reached Ready — check operator logs for UnforceableProblem Kafka Deployment Engineering
PostgreSQL pod CrashLoopBackOff with could not create lock file readOnlyRootFilesystem: true mutated by Kyverno — mount emptyDir at /var/run/postgresql and /tmp Deployment Guide
Pod admission rejected with Kyverno policy violation Pod doesn’t meet PSS standards — run kates kyverno violations to identify failing rules, then fix the manifest or add a PolicyException Security & Compliance

B.6 CLI Issues

The kates CLI itself fails, or can’t reach or authenticate to the Kates API.

Symptom Likely Cause Chapter
kates health killed immediately (exit 137) on macOS macOS blocks unsigned binary — com.apple.provenance xattr Deployment Guide
CLI connection timeout / connection refused Kates API not running or port-forward died Deployment Guide
[401] Missing API key or [403] Invalid API key, while kates health works The CLI context carries no API key or a stale one, or a stale KATES_API_KEY is exported, which the CLI prefers to the context’s key — only /api/health is public Deployment Guide

B.7 Chaos Engineering

A LitmusChaos experiment or a Kates disruption fails to start, has no effect or leaves a cluster that doesn’t recover, or the Kates — Chaos board shows no data.

Symptom Likely Cause Chapter
Litmus experiments fail to start Chaos operator pod not running or RBAC insufficient Deployment Guide
Disruption doesn’t take effect Target pod selector doesn’t match, or NetworkPolicy blocks Chaos Engineering in Practice
Every step’s Verdict in kates disruption status is Skipped The Kates API fell back to the noop chaos provider, which injects nothing Chaos Engineering in Practice
A DISK_FILL or NETWORK_LATENCY step fails The kates-chaos chart installs no disk-fill or pod-network-latency experiment for the default litmus-crd chaos provider Chaos Engineering in Practice
Cluster doesn’t recover after chaos ISR too small, min.insync.replicas violated Chaos Engineering Theory
Kates — Chaos board: $namespace picker empty, Chaos engines running reads No data No ChaosEngine has been created yet, or kube-state-metrics lacks the custom-resource configuration for it (charts/monitoring before 1.6.0, or another stack) Observability & Monitoring

B.8 Upgrades

Something that worked before an operator or Kafka upgrade fails or slows down after it.

Symptom Likely Cause Chapter
UnsupportedVersionException after operator upgrade Kafka version not supported by new operator version Upgrade Playbook
Topics not reconciling after CRD API change CRDs still using deprecated v1beta2 Upgrade Playbook
Performance regression after Kafka upgrade New version defaults changed — compare baseline tests Upgrade Playbook

B.9 Connectivity Debugging Flowchart

When you can’t connect to Kafka, work through this decision tree. The commands here and in Quick Diagnostic Commands assume the chart defaults — cluster name krafter in namespace kafka (set in charts/kafka-cluster/values.yaml); adjust them if your deployment overrides these values.

flowchart TD
    A["Can't connect to Kafka"] --> B{"Are Kafka pods running?"}
    B -->|No| B1["Check pod status:<br/>kubectl get pods -n kafka"]
    B1 --> B2{"Pods in Pending?"}
    B2 -->|Yes| B3["Check StorageClass and PVCs:<br/>kubectl get pvc -n kafka"]
    B2 -->|No| B4["Check pod logs:<br/>kubectl logs &lt;pod&gt; -n kafka --previous"]
    
    B -->|Yes| C{"Can you reach the bootstrap service?"}
    C -->|No| C1{"Is a NetworkPolicy blocking?"}
    C1 -->|Yes| C2["Add an ingress rule for<br/>your namespace → port 9092<br/>See Security & Compliance"]
    C1 -->|No| C3["Check DNS resolution:<br/>nslookup krafter-kafka-bootstrap.kafka"]
    C3 --> C4{"DNS resolves?"}
    C4 -->|No| C5["Check CoreDNS pods:<br/>kubectl get pods -n kube-system"]
    C4 -->|Yes| C6["Check port-forward or NodePort:<br/>kubectl port-forward svc/krafter-kafka-bootstrap 9092:9092 -n kafka"]
    
    C -->|Yes| D{"Authentication succeeds?"}
    D -->|No| D1["Check SCRAM credentials:<br/>kubectl get secret &lt;user&gt; -n kafka"]
    D1 --> D2["Verify KafkaUser reconciled:<br/>kubectl get kafkauser -n kafka"]
    
    D -->|Yes| E["Connection works!<br/>Check application config"]

B.10 When to Escalate

Not every problem is a Kates problem. Use this guide to determine where to focus your investigation:

Symptom Pattern Likely Layer What to Check
kates health shows all components UP, but tests fail Kafka Check broker logs, partition health, ISR state
CLI commands return “connection refused” or timeout Kates API Check the Kates API pod’s status, port-forward, service endpoints
Pods stuck in Pending, CrashLoopBackOff, or ImagePullBackOff Kubernetes Check node resources, StorageClass, image registry access
Kyverno rejecting pod creation Kyverno policies Run kates kyverno violations to identify which rule is failing
Latency numbers are unreasonably high for all tests Infrastructure Check node CPU/memory pressure, disk I/O, network bandwidth
Everything works locally but fails in CI CI environment Check resource limits, network access, Docker-in-Docker configuration
Tip

When filing an issue, include the output of kates doctor — it runs a battery of diagnostic checks and gives you a single summary of system health.

B.11 Common Issues (Additional)

The stack is up and the Kates API answers, but a command or a test result isn’t what you expect.

Symptom Likely Cause Fix
kates cluster topology returns “Cluster topology is only available when the Kates backend is deployed on Kubernetes with access to Strimzi CRDs” Missing ClusterRoleBinding for the Kates service account — the Kates API can’t query Strimzi CRDs Verify RBAC: kubectl get clusterrolebinding kates — if missing, redeploy with helm upgrade --install kates charts/kates -n kates
Test results show 0 records consumed even though producers succeeded Consumer group hasn’t started consuming, or topic has no committed offsets for the group Check consumer lag: kates kafka group <group-name>. If lag equals total records, the consumer never started — check the Kates API’s logs for consumer errors
kates trend shows no data even after running tests A trend reads only the type’s DONE runs created within --days; FAILED runs and runs still in flight are left out Check the runs’ status with kates test list --type <TYPE>, and widen --days for older runs

B.12 Quick Diagnostic Commands

Run these for a first snapshot: the Strimzi resources and pods, the operator and broker logs, the Kafka conditions, partition health through the Kates API, and Kyverno’s policies and violations.

# Cluster overview
kubectl get kafka,kafkanodepool,kafkatopic,kafkauser -n kafka

# Pod health
kubectl get pods -n kafka -o wide

# Strimzi operator logs (last 50 lines) — the operator is its own release,
# in its own namespace
kubectl logs deployment/strimzi-cluster-operator -n strimzi-operator --tail=50

# Broker logs (last crash)
kubectl logs <broker-pod> -n kafka --previous --tail=30

# Kafka status conditions
kubectl get kafka krafter -n kafka -o jsonpath='{range .status.conditions[*]}{.type}: {.status} - {.message}{"\n"}{end}'

# Under-replicated and offline partitions, through the Kates API
kates cluster check

# Kyverno policy status
kates kyverno status

# Kyverno violations (all namespaces)
kates kyverno violations

# Kyverno violations (specific namespace)
kates kyverno violations --namespace kafka

Kafka Tools Inside a Broker Pod

The tools under /opt/kafka/bin need credentials. Both internal listeners authenticate — SCRAM-SHA-512 on 9092, mutual TLS on 9093 — and the brokers enforce ACLs. Without credentials kafka-topics.sh retries until it times out, prints an Error while executing topic command line to standard output and exits 1, so a filter after it shows nothing and a broken command reads like a healthy cluster. Build one client configuration from the kates-backend KafkaUser’s Secret — the platform profile makes that user a super user, so it can describe every topic, its configuration and every consumer group — and pass it to each tool with --command-config:

BROKER=$(kubectl get pods -n kafka -l strimzi.io/cluster=krafter,strimzi.io/broker-role=true -o name | head -1)
JAAS=$(kubectl get secret kates-backend -n kafka -o jsonpath='{.data.sasl\.jaas\.config}' | base64 -d)

# Written to the pod's memory-backed /tmp, so it is gone when the pod restarts
kubectl exec -i -n kafka "${BROKER}" -- sh -c 'cat > /tmp/client.properties' <<EOF
security.protocol=SASL_PLAINTEXT
sasl.mechanism=SCRAM-SHA-512
sasl.jaas.config=${JAAS}
EOF

# Prove the credentials work: this lists the topics, __consumer_offsets included
kubectl exec -n kafka "${BROKER}" -- /opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server localhost:9092 --command-config /tmp/client.properties --list

Then every tool takes the same two flags:

# Under-replicated partitions — no output means none, once --list above succeeded
kubectl exec -n kafka "${BROKER}" -- /opt/kafka/bin/kafka-topics.sh \
  --bootstrap-server localhost:9092 --command-config /tmp/client.properties \
  --describe --under-replicated-partitions

# Consumer lag
kubectl exec -n kafka "${BROKER}" -- /opt/kafka/bin/kafka-consumer-groups.sh \
  --bootstrap-server localhost:9092 --command-config /tmp/client.properties \
  --all-groups --describe

kates kafka groups and kates kafka group <group-name> answer the lag question through the Kates API without a broker pod. The client configuration carries a super user’s password: it stays in the broker pod, and kubectl exec -n kafka "${BROKER}" -- rm /tmp/client.properties removes it when you are done.