1# Helm chart metadata for kates with Chainguard images
2# These images are distroless and signed with Sigstore
5home: https://github.com/bmscomp/kates
7 - https://github.com/bmscomp/kates
9description: Kates — Kafka Advanced Testing & Engineering Suite
19 email: bmscomp@gmail.com
20icon: https://raw.githubusercontent.com/bmscomp/kates/main/docs/icon.png
24 artifacthub.io/category: integration-delivery
25 artifacthub.io/changes: |
27 description: "appVersion 1.24.0, the first image with the backend fixes since 1.23.0: every requested spec field honoured and the request kept (#204); a cancelled run frees its concurrency slot (#203); a launched disruption plan stores its outcome (#207); run latency percentiles no longer averaged across tasks (#230); native ROUND_TRIP timed from send to the consumer's receipt (#234); lifecycle events reach webhooks, across restarts too (#231, #233); the direct Kubernetes chaos backend's ephemeral containers and the Strimzi-native SCALE_DOWN and ROLLING_RESTART (#174, #177, #181); and a Trogdor backend that can run a task on a real coordinator (#232, #236), which trogdor.agentNodes needs. The entries below that say they take an image newer than 1.23.0 are met by this one"
29 description: "trogdor.agentNodes (KATES_TROGDOR_AGENT_NODES, default node0, the name in the trogdor.conf Kafka ships) names the Trogdor agents the Trogdor backend's tasks run on, comma-separated and taken in turn. A Trogdor spec names the agent that runs it, and the backend named none, so the coordinator ended every task at once with \"Unable to find nodes for task\". Set it to the node names in your coordinator's platform config. It takes an image newer than 1.23.0, whose Trogdor backend could not run a task at all"
31 description: "The backend read Prometheus at http://prometheus.monitoring.svc:9090, its built-in default, and no chart creates that Service: charts/monitoring installs kube-prometheus-stack, whose Service is monitoring-kube-prometheus-prometheus. Every disruption on a default install captured no Kafka metrics, and its SLA verdict left latency and throughput unevaluated. prometheus.url (KATES_PROMETHEUS_URL) now defaults to http://monitoring-kube-prometheus-prometheus.monitoring.svc:9090, the Service `kates deploy` installs, and `kates deploy` sets it for the namespace it installs monitoring into"
33 description: "The NetworkPolicy's egress to Prometheus always named namespace monitoring, whatever networkPolicy.prometheus.namespace said (and values.yaml said kafka). It now reads the value, which defaults to monitoring, so the rendered policy is unchanged when a values file leaves it out; set it with prometheus.url when Prometheus runs elsewhere, as after `make monitoring` (namespace kafka). Upgrading: a values file copied from 0.10.4 or earlier (helm show values) carries the old, ignored networkPolicy.prometheus.namespace: kafka, which now takes effect and moves the egress to namespace kafka while prometheus.url points into monitoring, so Prometheus becomes unreachable with the NetworkPolicy on. Remove that line, or set it to the namespace your Prometheus runs in"
35 description: "values-prod.yaml set a 1Gi memory limit and inherited -Xmx2560m from values.yaml, and values-staging.yaml a 512Mi limit under the same heap, so either pod is OOMKilled once its heap grows past the limit. values-prod.yaml now runs 2Gi/4Gi, the chart defaults, with the heap set beside it (62.5% of the limit); values-staging.yaml runs 512Mi/1Gi with -Xmx512m, as values-dev.yaml does"
37 description: "values-generic.yaml, the overlay `kates deploy` applies on every cluster that is not kind, adds -XX:+ZGenerational, so ZGC runs its generational mode on the JDK 21 image there as it does under values.yaml and values-prod.yaml; -XX:+UseZGC alone selects the single-generation mode on JDK 21"
39 description: "values-prod.yaml, values-generic.yaml and values-corporate.yaml pinned image.tag 1.17.0, older than the 1.22.0 the probe defaults need, and `kates deploy` applies values-generic.yaml on every cluster that is not kind. They set no tag now, so the one in values.yaml applies"
41 description: "The default ClusterRole grants patch on kafkanodepools. SCALE_DOWN now removes a broker by lowering spec.replicas of its KafkaNodePool, and rollback and orphan recovery put the count back; it used to scale StatefulSets, which Strimzi does not create, so it removed nothing and its rollback restored nothing. It runs on the Kubernetes API with the Kates service account on both chaos backends (Litmus pod-delete only kills a pod its StrimziPodSet recreates), so the rule is not behind rbac.directChaos"
43 description: "The default ClusterRole grants patch on statefulsets, which only rbac.directChaos granted before. ROLLING_RESTART on a StatefulSet that Strimzi does not manage (charts/legacy-kafka runs Kafka and ZooKeeper that way) restarts the StatefulSet with the Kates service account on both chaos backends, like the pod annotation below, so on a default install that restart was refused. rbac.directChaos no longer carries the rule; SCALE_DOWN, which also patches the StatefulSet, is covered by the default one"
45 description: "The default ClusterRole grants patch on pods. ROLLING_RESTART now annotates each Kafka pod it restarts with strimzi.io/manual-rolling-update and lets the Strimzi Cluster Operator roll them; it used to restart StatefulSets, which Strimzi does not create, so it restarted nothing. It runs on the Kubernetes API with the Kates service account on both chaos backends (Litmus has no rolling restart experiment), so the rule is not behind rbac.directChaos"
47 description: "rbac.directChaos (default false) grants the writes the direct Kubernetes chaos backend (kates.chaos.provider=kubernetes, or hybrid without Litmus) makes itself: create, delete and deletecollection on networkpolicies; patch on statefulsets; get and update on statefulsets/scale; update on pods/ephemeralcontainers. The ClusterRole granted only get and list on NetworkPolicies and StatefulSets, and nothing on ephemeral containers, so every NETWORK_PARTITION, SCALE_DOWN, ROLLING_RESTART, CPU_STRESS and IO_STRESS on that backend was refused, along with its cleanup, rollback and startup orphan recovery. Litmus runs its experiments as litmus-admin and needs none of this, so the default ClusterRole is unchanged"
49 description: "CPU_STRESS and IO_STRESS on the direct backend replaced the whole pod to add their ephemeral container, which the API server rejects; they now write through the pods/ephemeralcontainers subresource. That takes an image newer than 1.23.0: on 1.23.0 the two faults fail even with rbac.directChaos=true"
51 description: "appVersion 1.23.0. The image publishes the Agroal pool meters (quarkus.datasource.metrics.enabled) and the HTTP latency histogram buckets (HttpLatencyHistogram) that Kates — Application Health and KATES — Overview read, and counts a task's failure on the poll that sees it, so kates_benchmark_errors_total moves during a run. A 1.22.0 image publishes none of that, and the pool row and the latency percentiles are empty on it"
53 description: "The ServiceMonitor matched on the pod selector labels (app: kates), which the Service does not carry, so with metrics.serviceMonitor.enabled=true Prometheus discovered the backend's endpoints and dropped every one of them — silently, and every Kates board stayed empty. It now selects on app.kubernetes.io/name and app.kubernetes.io/instance, which the Service has"
55 description: "The ServiceMonitor and PrometheusRule templates render only where the monitoring.coreos.com/v1 API exists, like every other chart in the repository; before, either flag on a cluster without prometheus-operator failed the release at apply time"
57 description: "ACTION REQUIRED — metrics.grafanaDashboard.*, kyvernoPolicy.grafanaDashboard and kyvernoPolicy.grafanaDashboardNamespace are GONE, and a release that still sets any of them is refused with the replacement named. Both boards are delivered by charts/monitoring (dashboards.enabled there) with every other board; the overview board's $job variable already lets one copy serve every release, and there is one Kyverno per cluster. The chart's <fullname>-grafana-dashboard and <fullname>-kyverno-dashboard ConfigMaps are no longer rendered"
59 description: "Both boards move to dashboards/kates-overview and dashboards/kyverno-security and are built from board.py by scripts/gen-dashboards.py, instead of being written inline in this chart's templates"
61 description: "The uid of each board is unchanged (kates-overview, kyverno-security-overview), so Grafana replaces them in place and existing links, playlists and annotations keep resolving. The overview board's title changes from KATES — Kafka Testing Suite to KATES — Overview"
63 description: "Five metric names on the overview board that no meter has ever published, renamed in place: agroal_pool_{active,available,max_used}_count (Quarkus publishes agroal_*, with no pool segment), kates_active_runs and kates_benchmark_throughput_records_per_sec. The monitoring chart's boards spelled these the other way, so the repository shipped three spellings of the same metric; they now agree with each other and with the code"
65 description: "Two label names on the Kyverno board that Kyverno does not emit. Policy Violations Over Time grouped on `policy` and Violations by Rule on `rule`; kyverno_policy_results_total carries policy_name and rule_name, and a group on a label that does not exist collapsed both panels to one unlabelled series"
67 description: "Three overview panels no rename could fix — Benchmark P99 Latency (ms), which read kates_benchmark_p99_latency_ms (there has never been such a meter; the percentiles are one series with a quantile label, and the per-run series that does publish them belongs on the benchmark boards, not on this one), Kafka Admin Connections (Micrometer binds only clients the Kafka extension creates, and this application builds its own AdminClient), and Vert.x Event Loop Latency (Vert.x metrics options are not enabled). The CPU and JVM panels survived on corrected series: process_cpu_usage and system_cpu_usage in place of process_cpu_seconds_total, and jvm_memory_used_bytes{area=nonheap} in place of process_resident_memory_bytes, both of which are Prometheus simpleclient default exports a Quarkus Micrometer registry never publishes. Two legends also read a task_id tag no meter carries"
69 description: "A $job template variable on the overview board, so one copy serves every Kates release a Grafana can see; the release name is no longer interpolated into 16 of its 21 expressions as a literal job= selector. Descriptions on all nine Kyverno panels, which had none — the layout gate in ci-kafka-charts.yml now covers this chart and will not accept a panel without one"
70 artifacthub.io/license: Apache-2.0