Skip to content

Release notes

kconmon-ng v2.5.1

Changed

  • Shorter alert texts. Every built-in alert's description is now one or two sentences on what to check first, down from up to a thousand characters, so a notification template that prints one per firing alert no longer turns four pairs into a wall of text. The explanations moved to the new Alert runbooks page. Node-pair alerts no longer print zones, which read as (zone ) on clusters without zone labels, and PathMTUBlackHole's summary now reads "only packets up to N bytes get through".

Added

  • runbook_url on every built-in alert, pointing at its section of the Alert runbooks page.
  • namespace label on every built-in alert, set to the release namespace. The aggregated rules had none, so notification templates showed Namespace: unknown, and an Alertmanager inhibit rule with equal: [namespace] treated them as matching every other alert without a namespace.

Fixed

  • The console's notice when console.alerting.enabled is off read as if alerting as a whole were off. It now says that only the rules built in the console are not applied, and that the chart's built-in rules keep alerting through Prometheus.

Upgrade notes

  1. Alertmanager routes and inhibit rules that match on namespace now see kconmon-ng's alerts in the release namespace; check routes that send a namespace to an application team.
  2. Templates, silences or tests that match the old summary or description text need the new wording. Expressions, thresholds and the other labels are unchanged.

kconmon-ng v2.5.0

2.5.0 adds a path MTU probe for the failure small probes cannot see: a pair where handshakes and pings cross while full-size packets vanish now turns red with the size that still crosses, instead of staying green on every plane. The rest of the release pays the debts a first outside user runs into: node-level alerts, maintenance windows that hold the console's webhooks, local users managed from the console, a config reload that survives the way files are really replaced and applies what it reads, a console whose first load is a fifth of what it was, and a set of security fixes.

Read the Upgrade notes before rolling out: the probe is on by default, two of the four new rules are critical, and the chart's NetworkPolicy is split per component.

Added

  • Path MTU probe (config.checkers.pmtu, on by default). Once a minute every agent sends each peer's UDP echo port a 64-byte datagram and a full-size one with Don't Fragment set. The full size is the MTU of the route to the peer (the route's own mtu when the CNI sets one, as Cilium does, else the egress device's); size overrides it. When the full size does not come back, the agent bisects (at most 16 sizes, timeout 500ms each) and confirms both ends, so a lossy path does not pass for a black hole. A pair reads ok (the full size is echoed), reduced (the path answers ICMP frag-needed: TCP adapts, UDP without its own path MTU discovery does not) or blackhole (full-size datagrams vanish while the small one crosses, and large transfers stall). A lost small datagram is a connectivity failure, left to the UDP plane, and a pmtu failure does not trigger MTR. The agent warns when interval is under 28 timeouts (14s) or 3m and more. New series: kconmon_ng_pmtu_bytes and kconmon_ng_pmtu_probe_bytes (per pair: the size that crossed, the size probed), kconmon_ng_pmtu_results_total (success for ok and reduced, fail for a black hole), kconmon_ng_zone_pmtu_results_total and kconmon_ng_agent_pmtu_probe_bytes. Agents advertise plane:pmtu. Walkthrough in Catch an MTU black hole.
  • PathMTUBlackHole and ZonePathMTUBlackHole (prometheusRule.pathMtuBlackHole, warning): more than half of a pair's probes failed over 10 minutes, or more than sustainedThreshold (0.1) over 30 minutes with at least two failures and one in the last 10, held for 5. The second arm catches a black hole on one of several ECMP paths. The zone rule fires only where no per-pair pmtu series exist (agent.metrics.detail=zone-only). A network that carries less than its routes say on purpose and clamps TCP MSS can set config.checkers.pmtu.size or turn the rule off.
  • NodeUnreachable and NodeIsolated (prometheusRule.nodeUnreachable, .nodeIsolated, critical, for: 5m): most peers fail most of their TCP probes to a node, or a node fails to reach most of its peers, with at least minPeers (2) reporting. They catch a node that stays registered behind a host firewall, a NetworkPolicy or a broken CNI datapath; the inhibit rules on the metrics page fold the per-pair alerts under them.
  • Path MTU on the dashboards. Overview opens with the key indicators in two rows (agents, leader, pairs, pairs with failures, black-hole and reduced-path pairs), then the worst-pair and MTR bars, with the charts and tables below, pairs below their probe size among them. Node Detail: the path MTU to and from each peer and black-hole probes by peer. Panels show the smallest size of the last 10 minutes, so an ECMP-split black hole stays on screen.
  • PMTU in the console and the CLI. The matrix gains a PMTU protocol: the path MTU in bytes, green at full size, amber on a reduced path ("1400 of 1500") or while recovering, red while recent probes fail, dashed "Not run" for a 2.4.x agent. The Overview, node and pair pages follow it. Run checks accepts pmtu between nodes, and kubectl kconmon check <source> <destination> --type pmtu exits 2 on a black hole.
  • Local users in the console. With auth.mode=local, Settings > Users adds, re-roles, resets, disables and deletes accounts under the new users:manage permission (built-in admin only), which the last enabled holder cannot lose. Every local user can change their own password. A password change or reset ends the user's other sessions; a disable or delete also revokes their API tokens. Routes under /api/v1/users in the Console API.
  • NetworkPolicy keys (Upgrade notes 4 to 6):
    • networkPolicy.dnsEgress replaces the default DNS rule, for NodeLocal DNSCache and other host-network resolvers;
    • networkPolicy.ciliumKubeAPIEgress (auto) adds CiliumNetworkPolicies for apiserver and node traffic, which no ipBlock matches on Cilium;
    • networkPolicy.clusterCIDRs carves pod and Service CIDRs out of the default 0.0.0.0/0 egress, which on Calico and Antrea matches pods;
    • console.networkPolicy.prometheusTargetPort opens the pod port behind a Prometheus Service that maps it (Thanos 9090 to 10902).
  • console.clientAddress.trustedProxyCIDRs: the proxies whose X-Forwarded-For names the client for rate limits, the WebSocket cap and the audit log, never for identity (Upgrade note 20).
  • agent.tls.enabled (false): TLS verified against the system trust pool with no other TLS field set, for an external agent whose gateway has a publicly signed certificate.

Changed

  • Maintenance windows hold the console's alert webhooks. An alert that starts firing inside a matching window (fleet-wide, a node, the pair or its target) is delivered only if it still fires when the window closes, and not at all if it resolves inside. A restarted console keeps holding. kconmon_ng_console_webhook_suppressed_total{event} counts what was held.
  • Config hot reload applies what it reads. The keys that go live and the ones that wait for a restart are in Upgrade note 11 and What reloads and what does not.
  • The node page covers both directions, so a node NodeUnreachable names no longer reads Healthy there, and its peer breakdown switches between To peers and From peers. Diagnostic runs list failed pairs first.
  • Console pages load on demand. The first page load drops from 935 kB to 198 kB gzipped, and charts load only the ECharts parts they draw with.
  • DNS probes against an explicit resolver ask the absolute name, so the search list no longer multiplies the query or sends cluster names outside. checkers.dns.timeout defaults to 2s (was 5s).
  • A departed peer's series go away ten minutes after it leaves the agent's peer list.
  • The controller refuses what cannot run. POST /api/v1/diagnostics answers 400 for an external destination with a type other than tcp, icmp or mtr or for a plane other than pod, 501 when the source agent does not run the type, and 503 leadership lost when the lease goes mid-task. kubectl kconmon check exits 1 on these, and kubectl kconmon finds the leader itself.
  • So does the console. Runs, definitions and schedules toward a target or an ad-hoc address take only tcp, icmp and mtr (a continuous schedule also dns and http), and an edit that would leave a stored schedule or definition unable to run answers 422; enabled: false always saves. The import follows the same rules, and the forms offer only what runs.
  • Console API input. Malformed input answers 400 or 422 instead of 502, and alert rule names that become one Prometheus alert name are refused.
  • The console's agent-missing template renders the chart's KconmonAgentsMissing expression; existing rules change on the next sync.
  • ZoneLossHigh description. Below about ten node pairs between two zones one broken link crosses the default 10% alone; raise prometheusRule.zoneLossHigh.threshold and zoneChecksFailing.threshold.
  • Continuous external checks fit the controller's 8 MiB limit: the console leaves out whole definitions, newest first, and counts them in kconmon_ng_console_external_specs_skipped_total{reason="over-budget"}.
  • Console UI. A silent MTR hop shows a dash, not 100% loss; refused forms focus the refused field; a foreign rule import asks to confirm; phones get a drawer Close button and tables that scroll inside their card.
  • Chart. The console gets GOMEMLIMIT from its memory limit. The GeoLite2 sidecar moves to ghcr.io/maxmind/geoipupdate:v8.0.0. The schema refuses what the binaries refuse at startup. The install notes flag an Ingress with no trusted proxies and an external gateway neither on externalTrafficPolicy: Local nor behind loadBalancerSourceRanges.
  • deb/rpm. The packaged agent config ships with the DNS checker off, and the postinstall keeps a net.ipv4.ping_group_range the admin already set.
  • Release images. :latest, the chart and the GitHub release follow a tag only after e2e passes on its images, and only the newest stable release moves :latest, the Latest badge and the krew index.
  • Build stack. Go 1.27.1, distroless static-debian13, Vite 8. The unused OpenTelemetry SDK is gone; observability.otel.* logs a warning.

Fixed

  • Agents.
    • Hot reload stopped for good after the first atomic replacement of the file (an editor's save, a puppet file resource, a ConfigMap swap).
    • logLevel: DEBUG and logFormat: TEXT ran at info in JSON.
    • An agent probed at a secondary IP, a multi-homed advertiseAddress or a VIP read as 100% UDP loss.
    • A dead DNS resolver read green for names in /etc/hosts or hostAliases.
    • Departed peers, targets and zones left stale series behind.
    • A node relabel overwrote an explicit agent.zone.
  • Controller.
    • After a failover, pairs towards late agents went unprobed for about a minute.
    • A rolling restart left no leader for the 15s lease; now about 2s.
    • An external-check assignment or removal that an agent's stream could not take was lost.
    • A subscriber that stopped reading was never cut off, except on the peer list.
  • Console.
    • Chart tooltips were white on the dark theme and ignored a theme switch.
    • An unreachable IdP, or replicas refreshing a rotating token at once, signed OIDC users out.
    • A PostgreSQL restart or a Valkey failover signed everyone out; the console now answers 503 until the store is back.
    • A custom config.metricsPrefix broke the browser views.
    • Rate limits failed on Redis and Valkey before 7.0, and Redis 5 or an endpoint without CLIENT TRACKING left each replica on its own state.
    • A transient apply failure, a rule-name collision or a shared bundleName could delete, overwrite or freeze the managed alert rules.
    • A prometheus.queryTimeout above 30s never took effect.
    • /metrics had no go_* or process_* families.
    • The Overview read healthy a pair the matrix painted red for loss.
    • Deleting a target in use on PostgreSQL 18 gave a store error, not 409.
    • The Time Machine lost a node when one of its two agents deregistered.
    • A failed MTR hop lookup waited the full enrichment TTL for a retry.
    • A cancelled run left pairs at dispatched, and permalinks could miss their final frames.
    • Concurrent role changes could leave a user with two roles.
    • Investigate lost unsaved notes, Time Machine incident lists kept resolved incidents, and failed reads showed "none".
  • Chart and dashboards.
    • helm test never started its Pod (CreateContainerConfigError).
    • An unquoted numeric geoip accountId crash-looped geoipupdate.
    • PairWentSilent said 15m whatever its for.
    • The console's scrape rule needed serviceMonitor.enabled, and KconmonExternalAgentDown missed a custom jobName.
    • The Overview dashboard's Agents missing let external agents mask missing ones, and several panels went blank with UDP off.
    • Under a long release name the dashboard ConfigMaps collided.

Security

  • NetworkPolicy. Every agent pod and every externalPeerCidrs address reached the controller's unauthenticated gRPC, HTTP and metrics ports.
  • Chart RBAC. The console could change or delete every PrometheusRule in its namespace, and the controller could update any Lease there.
  • UDP echo loop. One spoofed datagram could start an endless echo loop between two agents.
  • Webhooks. A receiver could redirect a signed delivery to any host, and failed deliveries logged the webhook URL, often a credential itself.
  • External gateway. Any token holder could read the event stream, and unauthenticated connections piled up without limit.
  • Controller. One large request body could run it out of memory.
  • HTTP check URLs. A password in a target URL leaked into the url label, logs, diagnostic results and config errors.
  • Console.
    • A burst of logins could run a 256Mi console out of memory.
    • A crafted X-Forwarded-For could become a rate-limit key or an audit address, and anyone behind the ingress could spend every user's sign-in budget.
    • A flood of refused requests could push sign-ins and admin actions out of the audit log, and OIDC sign-ins were not audited.
    • run:{id} WebSocket topics skipped runs:read, and /ws took any number of sockets.
    • GET /api/v1/export gave a role with settings:write every section, webhook URLs included.
    • A credential-less 16 MiB body cost about 50 MiB of heap, other sites could frame the console, and Redis errors logged session keys.

Upgrade notes

  1. The path MTU probe is on by default and sizes itself from the route. On CNIs that set a route MTU the pmtu gauges read it (1450 on Cilium with VXLAN), not the pod's 1500. The chart renders no pmtu key until you tune one, so 2.4.x agents keep running without the probe and still answer it. Set config.checkers.pmtu.* or agent.tls.enabled only with 2.5.0 agents: an older agent refuses the unknown key.
  2. Cardinality. Up to four new series per directed pair and one per agent, about 40 thousand more at 100 nodes. agent.metrics.detail=zone-only drops the per-pair ones; config.checkers.pmtu.enabled=false drops all.
  3. Two of the four new rules are critical. Check Alertmanager routing for NodeUnreachable and NodeIsolated, or lower prometheusRule.nodeUnreachable.severity and prometheusRule.nodeIsolated.severity for the first week. A black hole resolves up to ten minutes after its last failed probe unless prometheusRule.pathMtuBlackHole.sustainedThreshold is 1.
  4. The NetworkPolicy is split per component. With networkPolicy.enabled, helm upgrade replaces <fullname> with <fullname>-agent and <fullname>-controller; the controller admits agents and nodeCidrs on grpcPort only and nothing from externalPeerCidrs. Update your own policies that name the old object or relied on the wider rules.
  5. Cilium gets its own policies. With networkPolicy.ciliumKubeAPIEgress at auto the upgrade creates the CiliumNetworkPolicy <fullname>-kube-apiserver, plus <fullname>-node-ingress under agent.hostNetwork or the external gateway. helm template without cluster access needs --api-versions cilium.io/v2/CiliumNetworkPolicy or true; set false if your own policies cover that traffic.
  6. Off-cluster egress depends on the CNI. The default HTTP-check, webhook, OIDC and geoipupdate rules allow 0.0.0.0/0 on their ports, which on Calico and Antrea opens pods too: list the pod and Service CIDRs in networkPolicy.clusterCIDRs. The OIDC rule also admits any pod on 443; to close that, or to reach an IdP pod on its own port, set console.networkPolicy.oidcEgress.
  7. From 2.4.x, upgrade with --reset-then-reuse-values (Helm 3.14+) or -f. --reuse-values keeps every changed default at its old value (DNS timeout 5s, geoipupdate v7.1.1); a release older than 2.3.0 stops with a message.
  8. Changed chart defaults. config.checkers.dns.timeout drops to 2s, and the console gets GOMEMLIMIT from console.resources.limits.memory. checksum/secret changes once, then ignores console.auth.local.secret.password, which only seeds an empty users table.
  9. geoipupdate v8 sidecar. With console.mtr.enrichment.geoip.mode=auto the sidecar runs v8.0.0, which stops at the first edition that fails to download; console.mtr.enrichment.geoip.image.tag=v7.1.1 keeps the old behaviour.
  10. Stricter validation. The agent and the controller at startup, and the chart schema at helm upgrade, refuse an enabled checker interval under 100ms or timeout under 1ms (5ns for 5s), a zero mtr.cooldown, topology.sparse.zoneChords above 64 and an invalid metricsPrefix, bodyPattern, expectStatus, resolver port or CIDR, and the agent a loopback or unspecified agent.advertiseAddress. Each error names the key. mode and observability.otel.* load with a warning: remove them.
  11. Hot reload applies now. An in-place edit of a ConfigMap or config file goes live when saved, so review it first. Live: logLevel and checkers.* on the agent; logLevel, topology.*, controller.agentTtl and checkers.external.enabled/.allowedCidrs on the controller. Anything else, logFormat and ports included, logs a warning and needs a restart.
  12. DNS against an explicit resolver needs full names. A relative name such as kubernetes.default, or one only /etc/hosts knows, now fails against checkers.dns.resolvers and in external DNS checks: use an FQDN.
  13. The UDP echo refuses some sources and listens twice. Ports below 1024 and known agent echo endpoints get no reply. Full peer lists carry the whole fleet's echo endpoints (about 20 KB at 1000 agents, sparse mode included), and ss -lun shows two listeners on config.grpcPort.
  14. Masked HTTP check labels. A target URL with a password gets a new url label value (https://user:xxxxx@host/...); update queries on it.
  15. users:manage is new and admin-equivalent: a holder can create an admin. Add it to custom roles that should administer users.
  16. runs:read for live runs. A run:{id} WebSocket subscribe needs it; a custom role with events:read alone loses live run progress.
  17. Exports follow section permissions. Each section of GET /api/v1/export also needs its read permission (targets:read, checks:read, alerts:read, webhooks:manage, maintenance:read, rbac:manage), else it is listed in omitted: grant backup roles all six.
  18. Sessions. Local sessions from before 2.5.0 last until console.auth.session.ttl (12h). A password change now signs the user out elsewhere, and a disable revokes their API tokens for good.
  19. Login lockout. console.rateLimit.loginPerMinute (5) counts attempts per username, so anyone who knows a username can keep it locked out; where that matters, use OIDC or header auth, or limit sources at the ingress.
  20. Set the client-address proxies behind an Ingress, in every auth mode. List the ingress controller's addresses (its pod CIDR at the widest) in console.clientAddress.trustedProxyCIDRs, or every client shares one address for rate limits, the WebSocket cap and the audit log. While it is empty the console reads console.auth.header.trustedProxyCIDRs, which in auth.mode=header also decides identity: keep that to the auth proxy.
  21. WebSocket caps. Each console replica accepts 1024 /ws sockets (console.websocket.maxConnections), 256 per client address (.maxConnectionsPerAddress) and 32 per user or token (.maxConnectionsPerSubject); 0 turns a cap off. Behind an unlisted proxy a whole team shares the 256. The chart writes these keys and clientAddress only for console images 2.5.0 or newer.
  22. OIDC sign-in state lives in the browser, sealed under a key derived from the client secret: rotating the secret fails sign-ins in flight for up to 5 minutes, and one that crosses versions mid-rollout fails once.
  23. Maintenance windows hold console webhooks only; Alertmanager routes still need a silence. Existing windows start holding once the console runs 2.5.0: review them, since a long global window holds every managed alert.
  24. Webhook receivers must be the final URL. A 3xx now fails the delivery.
  25. Give the console's alert bundle its own name. With prometheusRule.enabled, helm upgrade refuses a console.alerting.bundleName equal to the chart's own PrometheusRule, and any other collision shows a sync error on every rule.
  26. Rename rules whose alert names collide. Of two stored rules with one Prometheus alert name, the one not deployed under it shows the sync error alert name collision.
  27. Schedules the agents cannot run. A stored once or interval schedule of a udp, dns or http definition toward a target or an ad-hoc address, or with a plane other than pod, no longer starts runs: pause, delete or repoint it. A stored udp definition toward a target saves only with enabled: false.
  28. API limits. POST /api/v1/diagnostics takes 64 KiB and PUT /api/v1/external-checks 8 MiB. A definition's params hold 4096 bytes and an address 2048: shorten a longer stored row before editing it. kubectl kconmon check --plane takes pod only.
  29. Dashboard ConfigMap names. From a release fullname of 41 characters up, the dashboard ConfigMaps get new names; tools that look them up by name need them.
  30. A signed image can precede its release. A tag pushes and signs :<version> before e2e, and a failed e2e leaves it with no chart, release or :latest. Deploy, or mirror :latest, once the release is published.

kconmon-ng v2.4.0

External agents stop being second-class. A host outside the cluster now tells its peers where it listens, gets scraped without a hand-written target, and shows up as what it is in the console, the CLI and the Time Machine. One rule comes with it, and it is the one to read before rolling out: per-agent ports are honoured only by upgraded agents; keep one port set until every agent, deb/rpm hosts included, runs 2.4.0. An older agent reports no ports and dials every peer on its own configured values, so in a fleet whose ports differ each old agent goes one-way red toward every peer listening elsewhere, its on-demand diagnostics included. The full skew matrix is under Upgrade notes at the end of this section.

Added

  • Per-agent ports on the wire. AgentMeta gains http_port, udp_port and metrics_port. Every agent reports its three listener ports at registration, and peers probe it on the ones it reported, for the scheduled mesh and for on-demand tasks alike, so an external host no longer has to mirror the cluster's port pair. Zero means "not reported" (an agent older than 2.4.0): the prober then dials its own configured port, each port falling back on its own. The controller refuses a port above 65535, peer lists carry the two probe ports but not metrics_port (nobody dials it), and an agent never adopts ports from the controller's reply; zone stays the only thing it takes from there. See Ports.
  • Prometheus HTTP SD for external agents. A bare host has no Service for a ServiceMonitor to select, so the controller, the one party that knows the host registered and on which address, now publishes it: GET /api/v1/prometheus/sd, served on httpPort and on metricsPort (the port the chart's scrape NetworkPolicy already opens). The contract:
    • one target group per external agent, <advertised address>:<metricsPort>, sorted by node name and deduplicated by address;
    • a fixed label set, node, zone, external="true" and agent_id. An agent's own labels never reach Prometheus, so a host cannot inject target labels;
    • with no external agent registered the body is the literal [], and the list is always served with Cache-Control: no-store;
    • a standby answers 503 not the leader, never 200 []: Prometheus reads every 200 as the complete target set, so an empty one from a standby would wipe every external target, while on a non-200 it keeps the list it has;
    • an agent that reported no metrics port (older than 2.4.0) is published on the controller's own config.metricsPort, and the controller logs metrics port assumed from controller config once per agent, which is the clue when such a host on another port sits at up == 0;
    • controller.prometheusSD.enabled: false closes the route (404 on both listeners). The key reaches the shared ConfigMap only when false, so an older controller image never trips over it;
    • with controller.replicaCount > 1 the controller Service spreads refreshes over all replicas and roughly half of them land on a standby: prometheus_sd_http_failures_total climbs for the job while the targets stay correct. Cosmetic, and written down so nobody chases it. Body and semantics in the HTTP API reference.
  • scrapeConfig.externalAgents in the chart. Renders a Prometheus Operator ScrapeConfig (needs the scrapeconfigs.monitoring.coreos.com CRD) named <release>-agent-external that reads the SD route, with labels for your Prometheus' selector (kube-prometheus-stack wants release: <its release name>), jobName, refreshInterval (30s) and interval. It applies the same agent.metrics.detail valve as the agent ServiceMonitor, so an external host never returns per-pair detail the valve drops for the pods, and the valve no longer insists on serviceMonitor.enabled when this is on. The chart refuses the ScrapeConfig without controller.externalGateway.enabled (nothing external could register) or with controller.prometheusSD.enabled=false (every refresh would 404), and the install notes remind you when the gateway is on without it, or when labels is empty. Plain-Prometheus http_sd_configs job and the reachability rules in Scraping external agents.
  • KconmonExternalAgentDown and kconmon_ng_controller_external_agents. An optional warning (prometheusRule.externalAgentDown, off by default, for: 5m) on up{job=~".*agent-external.*"} == 0: a host the controller lists that Prometheus cannot scrape, usually the host firewall admitting the Prometheus pod IP when the CNI NATs its egress to a node IP. The new controller gauge counts registered agents that came through the gateway.
  • agent.hostNetwork, for pod networks external hosts cannot route. The DaemonSet moves into each node's network namespace: the agents advertise the node IP (KCONMON_NG_POD_IP from status.hostIP), declare hostPort on all three ports, get dnsPolicy: ClusterFirstWithHostNet unless agent.dnsPolicy says otherwise, and label themselves kconmon-ng.io/host-network=true from a 2.4.0 image. The chart stops rendering the ping_group_range pod sysctl there, since the kubelet refuses net.* sysctls in the host namespace. It changes what is measured, for the whole DaemonSet: every in-cluster pair then probes node IP to node IP over the underlay, and the CNI datapath (overlay, conntrack, NetworkPolicy enforcement) is no longer on the probe path, so the breakage this tool exists to catch can hide behind a green matrix. Turn it on only when the goal is visibility between external agents and a cluster whose pod network they cannot reach. Before you do: PSS privileged for the namespace, TCP 8080, UDP 9090 and TCP 9091 (or your config.*Port values) free on every node, ping_group_range set by the node OS, and one agent per machine (a host-network pod and a bare-host agent cannot share an IP). See When the pod network does not route and Host networking.
  • networkPolicy.nodeCidrs and networkPolicy.externalPeerCidrs. Host-network agents register from node IPs that no pod selector matches, so with agent.hostNetwork and networkPolicy.enabled both on, the chart refuses to render the policy until nodeCidrs lists the node CIDRs; without it every registration but the one from the controller's own node would drop silently. externalPeerCidrs closes the old gap where an external agent registered fine and every cell between it and the cluster stayed red: its CIDRs join the agent-to-agent rules in both directions (UDP grpcPort, TCP httpPort, the ports-less ICMP/MTR rule) and never the gateway rule.
  • External agents in the console. Everything keys off the kconmon-ng.io/external registration label, which the console now passes through from the controller's topology together with the agent's capabilities (labels and capabilities on TopologyAgent in the Console API).
    • Topology draws the host beside the cluster nodes in the lane of its zone, with a neutral external badge (identity, never a health tier) and "readiness unknown" for screen readers.
    • Node page swaps Pod IP for Advertised address, explains the Ready dash, and lists the probe Planes the agent advertised. Agents now advertise plane:tcp, plane:udp, plane:icmp, plane:dns, plane:http and plane:mtr; an agent advertising none (older than 2.4.0) reads as "unknown", never as running nothing.
    • Overview badges the host in Worst pairs and adds "+N external agents" beside Nodes ready without counting them in, since that tile is Kubernetes readiness.
    • Matrix tells two silences apart from plain no-data. An external agent Prometheus is not scraping keeps the no-data fill and aria text, but its cells' tooltip and a note above the grid say why and link the scraping docs, until the first measured cell appears. A protocol the source does not run renders dashed like not probed, with its own legend row. Precedence when a cell has no data: excluded by the plan, then unsupported, then unscraped. See Silence with a known cause and External agents on the map.
  • External agents in the Time Machine. TopologyChanged events carry the agent's labels, so a replay badges a host the way the live view does, and a reconstructed topology lists a bare host under agents only, never as a presence-derived READY node. History recorded before the upgrade shows no external badges: a 2.3.x controller wrote no labels, and such a host stays an ordinary node in those instants. Historical responses never carry capabilities, since no event records them.

Changed

  • KconmonAgentsMissing is no longer masked by external agents. Registered agents include them and expected agents (schedulable nodes) never did, so one external host hid one missing in-cluster agent. The expression now subtracts controller_external_agents, with an or registered * 0 stand-in so the rule keeps working against a controller image that predates the gauge.
  • kubectl kconmon tables. agents gains an EXTERNAL column (yes, or - for the DaemonSet's own), and topology prints bare-host rows after the node rows (NODE is the registered name, READY is -, AGENT IP the advertised address). Scripts that parse the human tables will see the columns and rows shift; -o json stays the controller's body, unchanged apart from httpPort, udpPort and metricsPort on agents that report them.
  • Documentation lives on the site. The chart's Artifact Hub links, the install notes (Docs:, Helm values:, Console setup:), the chart README, kubectl kconmon --help, the krew manifest, the packaged agent config, the image OCI labels and the GitHub release footer now point at https://esdmitrii.github.io/kconmon-ng/. In the console, Settings → About gains Documentation, Release notes and Source links, and the command palette an Open documentation action. The README turned into a short front door with a Scope and limits section in place of the old "Not yet" list.
  • Screenshots and the demo match one stand. The docs frames were re-shot on one kind stand with an external agent, edge-host-01, in the mesh; the captions were rewritten to describe what each frame shows, near-duplicate frames were folded into one, and the API tokens frame no longer shows a token. The breaking-the-network demo was rewritten for that kind stand instead of the old Minikube helper, with the stand's own names and numbers.
  • Console polish. One type scale across pages, row actions reduced to their verbs (Test, Edit, Delete), and status and identity badges drawn as one system. Chart hover pills print the time as HH:mm:ss instead of a date wide enough to clip, the pill on a two-axis chart reads each axis in its own unit, and the topology map zooms out far enough to fit a phone.
  • Windows is explicitly unsupported. The agent compiles for windows/amd64, and TCP, UDP, DNS and HTTP would work as written, but the two checkers that make this tool what it is would not: ICMP and MTR sit on a datagram ICMP socket that golang.org/x/net/icmp supports only on Linux and Darwin by its own contract, and raw ICMP on Windows needs the Administrators group, so an agent that also runs on-demand probes for the controller would run as SYSTEM. On top of that, Go's clock on Windows is interrupt-tick granular (up to 15.6 ms), so a 0.3 ms LAN round trip would read as zero or as one tick while the histograms looked perfectly valid. A Windows vantage point without ICMP and MTR is designed but not scheduled; the trigger is a concrete host that needs it. CI now runs GOOS=windows go vet over the agent and checker trees so the door stays open at no cost. See the FAQ.
  • CI guards. The buf plugins are pinned and CI regenerates the protobuf code and fails on a diff, so a .proto edit that skipped make proto can no longer ship a wire skew. Chart CI checks the host-network and ScrapeConfig renders, and that the chart refuses a host-network policy without nodeCidrs and a ScrapeConfig without the gateway. The kind e2e gained a leg that turns the gateway on, joins one simulated external agent through it, and checks the topology, the SD body, a Prometheus scrape of the host, the console matrix and the CLI.

Fixed

  • The -arm64 agent and controller images shipped an x86-64 binary. The arm64 image entries in .goreleaser.yaml never set goarch, goreleaser defaults it to amd64, so every published -arm64 image and the arm64 half of the multi-arch tags carried the amd64 build since the first release (checked on kconmon-ng-agent:2.3.1 and kconmon-ng-controller:2.3.1). On an arm64 node the container failed with exec format error; amd64 clusters were never affected. 2.4.0 builds both from the arm64 binary. The console image was not affected: it is built outside goreleaser.

  • The Time Machine topology forgot agents that never changed. Topology events only say what changed, so a console that started recording next to a running fleet never showed an agent that stayed put afterwards, external agents included, and retention pruning stripped old registrations the same way. The console now stores the controller's whole topology as a topology_baseline row each time its event stream connects and hourly while it stays connected, and a reconstruction starts from the newest baseline at or before the instant, then replays the events after it. The events list never shows these rows. The cost, on consoles with a database and the event stream: one row per console replica per hour, plus one per reconnect, sized by the fleet at roughly 170 to 230 bytes of JSON per node with its agent, so about 20 KB an hour per replica at 100 nodes. History recorded before the upgrade stays baseline-less and folds from events alone, with the same gaps as before.

  • Grafana dashboards. Ratio panels in Overview and Node detail cap their axis at 100% (axisSoftMax: 1) instead of stretching a flat zero line to a 0-10000% axis. Node detail's MTR trace count is plain text now: it counts traces, including the ones the console asked for, and colouring it green, yellow or red passed that off as a health verdict. The Zone heatmap tables sort rows and columns the same way, so same-zone cells line up on the diagonal, under a from \ to corner header.
  • Console accessibility. Buttons and links meet 4.5:1 contrast in both themes (the dark theme's destructive fill and the light theme's primary fill, focus ring and muted text were darkened), links inside running text carry an underline instead of relying on colour, scroll regions are keyboard-reachable and labelled, headings follow page order, the live event feed is a valid list inside a labelled log region, and phones get a proper header landmark.
  • Tooltips covered what they described. The entrance animation's final transform outlived the animation and overrode the tooltip's own positioning, so a matrix tooltip landed on the very cell it explained.
  • Topology edges never showed their failure label on hover: React Flow disables pointer events on edges of a non-selectable map, so the hover handler never fired.
  • The agent DaemonSet declared its grpc port as TCP. On an agent that port is the UDP echo listener; it is declared UDP now, which is what makes the scheduler's hostPort clash check guard the right protocol under agent.hostNetwork.

Upgrade notes

The default install pins controller, agent and console images to 2.4.0 and needs nothing special. The table is for fleets that run mixed versions for a while, most often deb/rpm hosts upgraded after the cluster.

Old side + new side Per-agent ports Prometheus SD Topology event labels (Time Machine) Console badges and hints agent.hostNetwork
2.3.x agent + 2.4.0 controller reports none; peers dial it on their own ports listed on the controller's metricsPort, logged once as assumed carried as the agent sent them (a 2.3.x external agent already sets the label) badge shows; planes read "unknown" the agent ignores KCONMON_NG_HOST_NETWORK, so no host-network label
2.4.0 agent + 2.3.x controller dropped at registration; everyone dials their own ports no route: a ScrapeConfig on it fails every refresh not recorded depends on the console the node IP passes registration; the label is stored verbatim
2.3.x console + 2.4.0 controller n/a n/a the new field is ignored none: external agents look like ordinary nodes n/a
2.4.0 console + 2.3.x controller n/a n/a events carry no labels; only the console's own baselines badge history live badges and the unscraped hint work (the controller already publishes labels) n/a
2.3.x deb/rpm agent in a 2.4.0 fleet reports none, dials its own ports: one-way red toward peers on other ports scraped on the controller's metricsPort badged badge shows; planes "unknown" n/a
mixed fleet with different ports old agents dial their own ports: one-way red, on-demand tasks included n/a n/a n/a n/a

If you pin the controller image behind the chart, note that controller.prometheusSD.enabled reaches the shared ConfigMap only when false: a pre-2.4.0 controller image rejects the unknown key and crashloops, so leave it at the default until the image is current. Flipping agent.hostNetwork or agent.dnsPolicy rolls the DaemonSet; during the rollout node IPs and pod IPs coexist and a little transient PairWentSilent noise is expected.

kconmon-ng v2.3.1

Fixed

  • ZoneChecksFailing and ZoneLossHigh failed every evaluation with "vector cannot contain metrics with the same labelset" and raised PrometheusRuleFailures on the cluster: rate() over a __name__ regex union drops the metric name and collapses the per-protocol families into duplicate labelsets. The expressions now build the union with label_replace(...) or label_replace(...), which keeps the branches distinct and still tolerates a disabled checker's absent family.
  • CI now evaluation-tests every alert rule with promtool test rules against synthetic series for all metric families, including a positive check that each zone alert fires on staged bad data. Rendering and syntax checks never execute the query engine, which is exactly where this defect lived.

kconmon-ng v2.3.0

The sparse mesh changes WHAT "no data for a pair" means: under topology.mode: sparse most directed pairs are deliberately never probed. Everything in this release that reads per-pair series learns to tell "not planned" from "went dark" through one new metric, kconmon_ng_probe_intended — and that metric comes from the AGENT: images below appVersion 2.3.0 do not export it. The chart's rules degrade honestly on an older fleet (see PairWentSilent below), but do not flip topology.mode: sparse until controller AND agents run a 2.3.0 image — the controller config key is emitted only when sparse precisely because an older controller image rejects it and crashloops. The appVersion pin is aligned when the app release ships.

Added

  • topology.* — the sparse probe mesh, by values. topology.mode: sparse trims the full N×(N−1) probe matrix to a ring over sorted node names (sparse.ringDegree successors each, the connectivity guarantee) plus HRW-chosen cross-zone chords (sparse.zoneChords per directed zone pair, which keep the zone metric family fully populated), so probed pairs — and every per-pair series they export — scale ~linearly with node count instead of quadratically. sparse.autoThreshold is the floor: fleets smaller than it get the full mesh regardless of mode, because sparse only pays for itself at scale. Default is mode: full, byte-identical rendering to 2.2.0.
  • kconmon_ng_probe_intended — the plan, scrapable. A gauge, value 1 for every directed pair the topology plan assigns ({source_node, destination_node}, exported by the source agent), preset from the peer list at registration and pruned on every plan change — stale pairs are deleted, not left at 1. In full-mesh mode it simply marks every peer, so dashboards and rules can join on it without caring which mode the fleet runs. It is the one honest way to distinguish "this pair is not supposed to report" from "this pair went dark", which is why it ships in the same release as sparse mode and not one later.
  • investigateUrl on the two zone alerts. ZoneChecksFailing and ZoneLossHigh now annotate a console deep link, /investigate?kind=zone-pair&scope=<source>-><destination>, straight into the Investigate page scoped to the firing zone pair. The link is console-RELATIVE on purpose — the chart cannot know the console's external URL (ingress is optional), so notification templates prepend their own origin; the console normalises the typeable -> into its canonical pair arrow.

Changed

  • PairWentSilent joins on the plan. The rule now fires only for pairs present in the source agent's kconmon_ng_probe_intended series — the hard rule of the sparse design, shipped in the same release: without the join, every pair the plan trims would read as "went silent" for the hour its results take to age out of the lookback window. The fallback is per SOURCE, not global: a source_node exporting no probe_intended at all keeps the old two-window behaviour, so a pre-2.3.0 agent image alerts exactly as before, a mixed fleet mid-rollout gets each behaviour where it applies — and an agent that dies outright takes its probe_intended series with it, which lands its pairs in the same fallback and preserves the alert's original purpose: catching an agent that stopped running or stopped being scraped.

This release also carries everything prepared for the never-published 2.2.0 tag (its pipeline caught two release-tooling defects before anything went out); those changes follow below, under their original heading kept for upgrade notes.

Carried over from the unreleased 2.2.0

Everything in this release reads the new zone-level metric family (kconmon_ng_zone_*), and that family comes from the AGENT, not the chart: agents below appVersion 2.2.0 do not export it (this chart pins 2.2.0, so a default install is fine — the warning is for fleets running an older agent image behind a newer chart). Until the fleet runs an agent image that does, the two zone alerts are silently inert (their expressions match no series), the Zone Heatmap dashboard renders empty, and agent.metrics.detail: zone-only would drop the per-pair series with nothing replacing them — Prometheus goes dark on the mesh while the console keeps working. Upgrade the agent image first, flip the valve second. The appVersion pin is aligned when the app release ships.

Added

  • ZoneChecksFailing and ZoneLossHigh. Two alerts on the zone plane, with the same per-rule knobs as the rest (prometheusRule.{zoneChecksFailing,zoneLossHigh}.{enabled,threshold,for,severity}). ZoneChecksFailing is the failure ratio of all TCP, UDP and ICMP probes between a zone pair, in one expression — the __name__ union keeps a disabled checker from blanking the ratio. ZoneLossHigh computes loss as (sent − received) / sent from the zone packet counters; averaging the per-pair loss-ratio gauges into a zone would weight an idle pair the same as a busy one, so the chart never does. Its default threshold is 0.1, lower than the per-pair UDPLossHigh at 0.5, because the zone aggregate dilutes any single link by the pair count: sustained loss at that level means the fabric, not one node. Both survive every agent.metrics.detail mode — that is the point of alerting on the zone family.
  • agent.metrics.detail — the cardinality valve. A scrape-time knob rendered as metricRelabelings on the agent ServiceMonitor: full (default, everything, ~70 series per directed pair), counters-only (drops the four per-pair histograms, ~10/pair — every pair alert keeps firing), zone-only (drops every series naming a destination_node, ~0/pair; the zone family at ~74×Z² series and the linear DNS/HTTP/external families remain). At 100 nodes that is ~0.7M → ~0.1M → practically N-independent, by configuration alone. Setting it without serviceMonitor.enabled is refused at render time rather than silently dropping nothing; plain-Prometheus equivalents are in docs/metrics.md.
  • controller.externalGateway — the external agent gateway, exposed by the chart. The controller's second gRPC listener (same services, but TLS with a bootstrap token, for agents OUTSIDE the cluster) gets a values block and three templates. templates/controller/service-external.yaml is a NodePort/LoadBalancer Service carrying the gateway port ALONE — the plaintext in-cluster gRPC port authenticates by network position and never appears on it, because a LoadBalancer in front of it would hand the whole mesh to anything that can reach the address. The deployment mounts two referenced Secrets read-only: tls.secretName (a kubernetes.io/tls serving pair; tls.clientCaKey names the CA bundle key in the same Secret and switches on client-cert identity pinning — empty is token-only mode, where any token holder can impersonate any agent, and NOTES.txt says so at install) and bootstrapToken.{secretName,key}. With networkPolicy.enabled, ingress on the gateway port is opened from networkPolicy.externalAgentCidrs toward the controller pods alone, and an empty list is refused at render rather than shipping a gateway no packet can reach; missing Secret names and a port colliding with config.{httpPort,grpcPort,metricsPort} are refused the same way. Two operational notes. Rotation: the gateway reads the certificate and token ONCE at startup and the chart cannot checksum content it only references, so rotating either Secret in place needs kubectl rollout restart deploy/<release>-controller. Version skew: the externalGateway config key is emitted only when enabled, because a controller image at appVersion 2.0.3 rejects the unknown key and crashloops — upgrade the image before flipping the switch, same rule as the zone family above.

Changed

  • The Zone Heatmap dashboard reads the zone family. Every panel that aggregated per-pair series into zones at query time now reads the pre-aggregated kconmon_ng_zone_* metrics, so the dashboard keeps working in every agent.metrics.detail mode and its queries stop scaling with the pair count. Loss panels are packet-weighted from the sent/received counters instead of averaging the per-pair ratio gauges. The one exception is the "MTR traces triggered" panel: MTR has no zone-level family, its counter is per-pair, and in zone-only mode that panel reads zero — its description now says so.

Performance and self-observability

  • Peer probing fans out with a bounded pool (32 in flight per round): a dead peer costs one timeout, not one timeout per peer in sequence, so probe cadence holds through partitions.
  • Reactive MTR traces are bounded by a global semaphore (4 in flight) on top of the existing per-pair cooldown; a mass partition trickles traces out instead of forking one per broken pair.
  • The agent exports self-metrics under kconmon_ng_agent_*: probe cycle duration and overruns per checker, controller reconnects, peer-list age, reactive-MTR in-flight and coalesced counters. All are fleet-size independent.
  • The controller coalesces peer-list broadcasts (trailing edge, 200 ms): a rollout's burst of registrations produces one broadcast, not one per change; the peer message is built once per broadcast and carries only the fields agents read.

kconmon-ng v2.1.0

Added

  • agent.updateStrategy is a value. The DaemonSet hardcoded maxUnavailable: 1 — the right default and the wrong ceiling: one node at a time turns a version rollout across a few hundred nodes into hours of half-upgraded fleet. The block passes through verbatim, so maxUnavailable: 10% — or OnDelete — is now a values change instead of a fork of the template.
  • priorityClassName on every workload. agent.priorityClassName, controller.priorityClassName and console.priorityClassName; empty by default, and empty renders nothing. The agent is the one worth setting: under node pressure the kubelet evicts lowest priority first, and the first pod gone should not be the one reporting on the node.

Fixed

  • helm test passes restricted-PSS admission. The connection-test Pod ran curl with no securityContext at all, so a namespace enforcing the restricted Pod Security Standard rejected it at admission — the test failed before it made the one request it exists to make. The container now declares the four fields the profile checks: runAsNonRoot, a RuntimeDefault seccomp profile, no privilege escalation, all capabilities dropped. The image already runs as uid 100.

kconmon-ng v2.0.3

Fixed

  • A config change restarts the pods that read it. The agent and the controller share one ConfigMap and read it once at startup, and a mounted ConfigMap changes under a running process without telling it — so controller.events.enabled: true applied to a live release updated the object and left the controller on the old file. It went on advertising no capabilities, the Console's realtime ingester retried against a stream that was configured but never started, and nothing anywhere reported an error: the Live page was simply empty. Both workloads now carry checksum/config, so a values change rolls them; a change the ConfigMap does not carry still does not.

kconmon-ng v2.0.2

Fixed

  • The matrix no longer opens at half size. The grid measured the height its own content had produced and fed that back into the fit, so a fresh render saw the container's 256px minimum, decided the grid did not fit, shrank to 50%, and the smaller grid then held the box at 256px — a loop with no way out. It measures the space available instead, and a seven-node fleet opens at 100%.
  • Zooming in gives the node names back. The shared prefix every node name begins with is dropped to buy column width, which is right while the column is narrower than the names and wrong the moment it is not: at 125% a label column holds adm-kuber-01 with room over and still read …01. The elision is now decided per axis at the current scale, and the note above the grid appears only while an axis is actually eliding.

Added

  • A favicon. The console had none, so every tab showed the browser's blank square; it now wears the mark it wears in its own sidebar.

kconmon-ng v2.0.1

Fixes a hole in 2.0.0: auth.mode=oidc and auth.mode=header shipped with no way to grant anybody a role. Both modes worked, and neither was usable.

Fixed

  • An OIDC or header install can grant roles at deploy time. Role bindings live in the database and are created through an API that already requires rbac:manage, so a fresh install had nobody able to make the first binding; the only alternative was auth.defaultRole, which is one role for every authenticated subject. The way out that 2.0.0 left was to bring the console up in local mode, log in, create a binding by hand and only then switch — a workaround, published as if it were a procedure.

console.auth.groupRoles maps a group the identity provider asserts onto a role this console grants, in the values file:

console:
  auth:
    groupRoles:
      platform-oncall: admin
      everyone: viewer

Roles resolve as the union of that map and any binding made through the API, so a grant by hand still adds to what the provider's groups carry. A group absent from the map grants nothing. What the map grants cannot be revoked through the API — that is what makes it declarative. - A role store outage no longer costs an operator their access. The store's half still fails closed, because an unreadable database is no evidence a subject holds anything; a grant that came from the claim and the config was never in doubt, and an outage is when the console is most needed.

kconmon-ng v2.0.0

A chart that installs monitoring and nothing else, a console that survives more than one replica, and one that tells the truth about time. The chart no longer ships a database or a cache — point it at the ones you already run. The Time Machine moved out of the top bar and into each page's own time controls, and the charts pin their axis to the window you asked for rather than to the data that happened to arrive. MTR gained a Runner, path history that reads as a timeline, and external targets.

Breaking

  • The chart no longer installs PostgreSQL or Valkey. database.mode, database.cnpg.* and the bundled subcharts are gone: set database.existingSecret to a Secret holding a postgres:// DSN and redis.existingSecret to one holding a redis:// DSN, and any managed instance works — RDS, a StatefulSet, a CloudNativePG cluster you run yourself. Every removed key fails the render with a message naming its replacement (templates/_migrations.tpl), so no old value is silently honoured.
  • console.database.* moved to the top-level database.*, and console.* keys that described the bundled datastores went with it.

Added

  • Chart 2.0.0, templates split per component — agent, controller, console, shared and observability each own their directory, with a NetworkPolicy set covering every component, fail-closed on external egress. Render-time guards refuse a port collision, an OIDC redirectURL the console would not start on, and more than one console replica without a shared cache.
  • The console scales past one replica — sessions, the fixed-window rate-limit counters and the realtime fan-out live in the Redis-compatible server, and the controller elects a leader so exactly one replica drives the reconcilers.
  • MTR Runner and path history — start a trace from the Explorer itself, with a settable cadence and duration; every distinct route the fleet has taken is kept, diffed and drawn on a timeline of when it changed.
  • External targets — probe a destination that is not a fleet peer, gated by config.checkers.external.allowedCidrs and the cluster's own egress policy. The console refuses at create time a target no agent could ever reach.
  • Time in the Console's result table — every figure says when it was read.

Changed

  • The Time Machine lives with the page's time filters, not in a strip across the top of every route. It is offered only on the pages that resolve their reads through ?at=, and the engaged banner stays global because writes are disabled console-wide.
  • Explore's axis is the window you picked — a 24h view draws 24 hours even when Prometheus holds less, instead of quietly redrawing three.
  • MTR Explorer is sorted by name, both destinations and their sources, with numbers read as numbers (m9 before m10).
  • OIDC identity is the sub claim, namespaced as oidc:<sub> — the only claim OIDC Core §5.7 allows as an identifier. auth.oidc.usernameClaim now decides the display name alone, so renaming a person no longer moves their roles (Grafana's CVE-2023-3128 is what the old shape risked). Group membership is re-read on every token refresh. Bindings made against a username stop granting; the console names them at boot so they can be remapped.
  • The configuration bundle carries access control — custom roles and the grant list, but only for a caller who holds rbac:manage. Roles import; bindings never do, because a grant names a person in the source console's own identity namespace.
  • Only the chart under the cursor shows a tooltip. Its neighbours keep the shared crosshair and mark their own samples with a dot, instead of each covering its own curves with a box of numbers.

Fixed

  • WebSocket topics are authorized per topic: events:read no longer carries the topology and matrix snapshots that topology:read and matrix:read gate, and a permission taken away reaches a socket that is already open — the topics it may no longer have are dropped, the rest of the connection is left alone.
  • The audit row describes the mutation that happened. A body could name one thing for the handler and another for the audit log by spelling a key in a different case, and a value carrying a NUL made the whole row unwritable — in both directions the caller chose whether their own privileged action was recorded. The extraction now matches keys the way encoding/json matches struct fields, and is bounded before it is decoded, so a wide body on the public login route can no longer take the replica past its memory limit.
  • A broken alert rule no longer freezes the whole bundle. Editing a deployed rule into PromQL the apiserver rejects used to stop every other rule from being applied, while the API answered 2xx and Prometheus kept evaluating the stale set. The quarantine now keys on rendered content rather than rule ids, offers each suspect to the cluster on its own, and removes the object only when every rule was offered and every one refused.
  • auth.mode=anonymous is not exempt from CSRF. Any page an operator's browser visited could POST into a console kept off the internet; a cross-origin write is now refused, while a script that sends no Origin is unaffected.
  • The node-local HTTP checker verifies certificates. An expired certificate, one issued for another hostname or an interceptor's CA all used to pass, so an https check could not fail on the condition it was added to notice; opt out per target with insecureSkipVerify.
  • External metrics separate the checks on one target — the series carry check_type, so an icmp and a tcp check on the same target no longer average each other's failures away under the ExternalChecksFailing rule.
  • A check no agent could run is refused when it is written, instead of being dropped by every agent with nothing but a log line while the console listed it as enabled.
  • The MTR destination listing is complete. It is paged behind a keyset cursor rather than capped, so no pair is missing from the Explorer and no per-destination total is short.
  • A subscriber that stops reading its peer-update stream is torn down rather than holding a controller goroutine and its connection slot until TCP notices.
  • Shutdown finishes in-flight runs before tearing down the pipeline they publish onto, so a rolling update no longer logs dropped frames that were delivered.
  • Every request body is capped, so one oversized POST can no longer take a console replica past its memory limit.
  • The OIDC callback binds its state to the browser that started the flow.
  • A role-store failure now refuses rather than granting the default role.
  • External TCP and UDP checks probe what was asked for instead of speaking the agent's own protocol to something that is not an agent.
  • A user binding can no longer be resolved by a subject of another kind: role resolution matches the caller's kind as well as their id.
  • Revoking a role binding is auditable — the audit row names the role and the subject, read before the row is destroyed rather than after.
  • Path history says when it has reached the end instead of leaving a "Load older" button that can never be pressed, and counts the routes it is showing against the traces folded into them.
  • A probe tick on a diagnostics run leads with that probe — its sequence, its clock, its latency or its error — so two ticks on an unchanged route are no longer indistinguishable.

kconmon-ng v1.9.0

Console release, and the last planned milestone. Alert rules you build in the Console and Prometheus evaluates; alert webhooks; configuration export/import; a command palette; a Settings page. Plus one long-carried fix: the controller finally attributes topology changes, so Time Machine reconstructs a real cluster. Everything new is off by default or read-only-additive: a 1.8.0 install that upgrades and changes nothing renders the same manifests.

Added

  • Alert rule management — /alerting builds a Prometheus alert rule from six typed templates (pair loss, zone latency, DNS failures, HTTP TTFB, agent missing, external target down) or raw PromQL — seven kinds in all — and the Console reconciles every enabled rule into one PrometheusRule object by server-side apply. The Console manages; Prometheus evaluates. Nothing here decides that an alert fired.
  • Validation by running it, not by parsing it — there is deliberately no prometheus/prometheus parser dependency. Every template has a byte-exact render golden, and POST /api/v1/alert-rules/preview runs the expression as an instant query against your actual Prometheus and reports how many series it matches. The render and the query fail independently: a render failure is a 422, a query failure is a 200 carrying the expression and the error.
  • Drift is recorded, then fixed — a reconcile always re-asserts the Console's bytes. A rule showing drift also carries a fresh lastSyncedAt, and both are true: the divergence was observed and corrected in the same pass. Failures never crash the loop; they land per rule as sync_status=error with a closed cause class (crd-missing, forbidden, other).
  • Foreign rules and explicit adoption — PrometheusRule objects the Console did not write are listed read-only, and POST /api/v1/alert-rules/import copies one into builder rows. The foreign object is never mutated, which means the same alerts then exist twice until you remove one copy. The import report says so, and names every skipped rule with its reason.
  • GET /api/v1/alerts — the firing set, projected onto this API's vocabulary, with ?managedOnly=. With no Prometheus configured it answers 200 and promConfigured: false rather than 503: "nothing is firing" and "nobody is watching" are different sentences.
  • Alert webhooks — alert.fired and alert.resolved, their own payload family, dispatched from a poller that diffs Prometheus' alert state on console.webhooks.alertPollInterval (30s). It baselines on boot rather than paging the fleet about what was already broken, freezes on a failed or undecodable poll rather than "resolving" everything, and ignores rules the Console does not manage. The M6 incident payload bytes are unchanged.
  • Configuration export/import — GET /api/v1/export and POST /api/v1/import, versioned bundle v1, admin-only under settings:write, dry-run first. Webhook endpoints export with hasSecret only and therefore cannot be created by import — a sealed secret never leaves this API.
  • Settings page — webhook CRUD, export/import with a per-collection dry-run report, and read-only deployment info that renders only what GET /api/v1/config actually serves.
  • Command palette (⌘K / Ctrl-K) — hand-rolled, zero dependencies, over navigation (generated from the sidebar so it cannot drift), five actions and the Time Machine pair. It does not jump to an arbitrary node, target or pair: that needs a live object search, not a static registry.
  • Overview — the "Firing alerts" placeholder carried since M1 is now the real panel, severity-sorted with oldest-first ties, and the Investigate timeline gained an alert row.
  • Permissions — alerts:read (all built-in roles) and alerts:manage (operator, admin, and alert-editor — the builtin has waited for exactly this permission since M3, and a role by that name that cannot edit an alert rule breaks its promise on first click). AllPermissions is now 25.

Fixed

  • Topology events are attributed — the controller now emits one topology_changed event per affected agent, carrying nodeName, agentId and zone, from all four sites (register, zone update, deregister, stale eviction). Time Machine's topology fold reconstructs a real node set with real zone lanes instead of an honest empty one. Events written by an earlier controller are counted as unfoldable and age out with retention; the page reports both numbers rather than rendering an empty cluster.
  • WebSocket topics are authorized individually — /ws admits events:read or runs:read, and a subject admitted on runs:read alone gets run:{id} topics and an error frame for the fleet-wide ones, on a socket that stays open. A custom role can finally watch the run it started. Carried from M3.
  • A null console.database.cnpg override no longer crashes rendering — nor does a null sub-block. Real nil-pointer class, found by the schema work.
  • One redundant token listing on the ownership-resolution path is now a targeted lookup.

Chart

  • 1.8.0 → 1.9.0. New value blocks: console.alerting.* (enabled/namespace/syncInterval/bundleName, rendered into the console ConfigMap only when enabled) and console.webhooks.alertPollInterval (rendered only when a key Secret is named and alerting is on — the key existed in the binary since M7 but was unreachable from Helm).
  • Enabling console.alerting renders a namespaced Role and RoleBinding over monitoring.coreos.com/prometheusrules (get,list,watch,create,update,patch — never delete), bound to the console-only ServiceAccount. Not a ClusterRole: the Console writes one object into one namespace, and pointing namespace elsewhere fails with a forbidden rather than widening anything.
  • The console ServiceAccount, serviceAccountName, POD_NAMESPACE and the apiserver egress rule are now shared between console.kubernetesContext and console.alerting through one helper, so either flag renders them.
  • values.schema.json closes 44 chart-owned levels with additionalProperties: false, and gained the nameOverride, fullnameOverride, agent.nodeSelector and agent.affinity keys the templates always used and the schema never declared. Pod/container securityContext stay deliberately open — Kubernetes grows union members every release, and closing them would turn a cluster upgrade into an install failure.
  • Three new ci profiles: console-alerting-values.yaml (the fullest console), console-auth-local-values.yaml and console-auth-header-values.yaml. Default render is key-identical to 1.8.0.

Upgrade notes

  • Nothing to do. alert_rules is created by migration 00007 on first start with console.database.mode=cnpg|external; with the database disabled the new routes answer 503 exactly as M3–M6's do.
  • Turning on console.alerting.enabled needs three things the chart cannot check: the Prometheus Operator's PrometheusRule CRD, a database, and a Prometheus whose ruleSelector/ruleNamespaceSelector actually selects the object the Console writes. A rule that syncs cleanly and never fires is almost always the third one.
  • Changing console.alerting.bundleName on a live install orphans the previous object. The reconciler owns what it is pointed at and deletes nothing.
  • There is no leader election on the alert-webhook watcher: N console replicas deliver N copies of every edge. The payload carries a stable (event, ruleId, labels, firedAt) tuple so a receiver can dedupe.
  • alert-editor gained alerts:manage. If you granted that builtin to somebody expecting it to stay inert, it is now able to create, edit and delete alert rules.
  • If you set values the schema never declared, helm upgrade may now reject them. That is the typo protection working — check the key against values.yaml.

kconmon-ng v1.8.0

Console release. Investigation Mode with an honest, documented correlation panel; saveable incidents with shareable permalinks; maintenance windows; outbound webhooks; and optional Kubernetes event capture (M6). Everything new is off by default or read-only-additive: a 1.7.0 install that upgrades and changes nothing renders the same manifests.

Added

  • Investigation Mode — /investigate assembles a merged timeline, synced signal panels (loss/RTT with a matrix delta chip and an MTR path diff) and an actions rail for a scope and a time range. Entry is the URL and only the URL — ?kind=&scope=&from=&to= — from any node/pair/target card, any matrix cell, or the page's own form. Every source is permission-gated with zero requests when denied, and each absent or bounded one leaves a muted line rather than blanking the page.
  • Correlation v1, documented rather than magic — edge-triggered threshold crossings (loss > 1%, RTT > 2× the range median), an onset, a 300-second candidate window and a linear proximity decay against published class weights. The panel links the scoring source itself, so the operator reads exactly the constants the code executes. No ML, and nothing you cannot reproduce by hand.
  • Incidents — save an investigation (/api/v1/incidents), pin findings from six source kinds, write notes, resolve and reopen. The permalink /investigate?incident={id} rehydrates scope and range from the row, so the link cannot drift from the incident it names. Open incidents appear on Overview and beside the charts on every object card. PATCH is deliberate and is this API's only one: an incident evolves under collaboration, and a full replace would let one writer discard another's notes.
  • Maintenance windows — /api/v1/maintenance, drawn as markArea on Explore, the Pair card and the Target card and as timeline rows. M6 renders declared windows; it does not suppress anything, because nothing evaluates alerts until M7.
  • Outbound webhooks — /api/v1/webhooks (admin-only webhooks:manage), firing on incident lifecycle. Deliveries are signed X-Kconmon-Signature: sha256=<hmac> over the raw body, retried 3 times (0s / 30s / 5m, ±20% jitter, 10s per attempt), with the outcome kept on the endpoint row. Each endpoint's signing secret is write-only over the API and sealed at rest with AES-256-GCM under console.webhooks.encryptionKeySecret. POST /{id}/test sends one signed ping. Without a key, create and test answer 503 and everything else keeps working.
  • Kubernetes event capture — console.kubernetesContext.* (off by default) list+watches core/v1 Events into k8s_events, filtered to nodes in the fleet topology and pods in one namespace, and read back through GET /api/v1/k8s-events under events:read. Zero new Go dependencies — it reuses the client-go the controller already pulls in. Without a topology it fails closed and drops node events rather than storing an unfiltered firehose.
  • Overview — the "Recent events" placeholder that had been carried since M2 is now the real panel, and an "Open incidents" card sits beside it. "Firing alerts" stays an honest placeholder until M7.
  • Permissions — incidents:read and maintenance:read (all built-in roles), incidents:write and maintenance:write (operator/admin), and webhooks:manage (admin only — an endpoint carries a signing secret).
  • Metrics — kconmon_ng_console_k8s_events_total{result} and kconmon_ng_console_webhook_deliveries_total{result}, plus three new table values on kconmon_ng_console_retention_deleted_total.

Chart

  • 1.7.0 → 1.8.0. New value blocks: console.kubernetesContext.* (rendered into the console ConfigMap only when enabled) and console.webhooks.encryptionKeySecret.{name,key} (referenced by name, mounted as a file beside the DSN — there is deliberately no inline-key value). Enabling kubernetesContext also renders a console-only ServiceAccount, its own ClusterRole (events: list, watch) and binding, serviceAccountName on the console Deployment, a POD_NAMESPACE downward-API variable and an apiserver egress rule (console.networkPolicy.kubeAPIEgress). The agent/controller ServiceAccount and its grant are untouched. A ninth ci profile (console-investigation-values.yaml). Default render is key-identical to 1.7.0.

Upgrade notes

  • Nothing to do. All four M6 tables are created by migration 00006 on first start with console.database.mode=cnpg|external; with the database disabled the new routes answer 503 exactly as M3–M5's do.
  • Turning on console.kubernetesContext.enabled is a deliberate act with an RBAC consequence — read SECURITY.md §10.3 first.
  • Webhook endpoints are API-only in M6: there is no Settings page yet.
  • Rotating console.webhooks.encryptionKeySecret does not re-seal existing rows. An admin must PUT each endpoint with a fresh secret afterwards.

kconmon-ng v1.7.0

Console release. MTR Explorer with durable path history and optional hop enrichment, the Time Machine global time context, and chart annotations (M5). Everything is off by default or read-only-additive; a 1.6.0 install that upgrades and changes nothing renders the same manifests.

Added

  • MTR path history — every MTR result the Console records is projected into mtr_path_snapshots: hop lists content-hashed per (source, destination) route, deduplicated at ingest, with first/last-seen and trace counts. kconmon_ng_console_mtr_snapshots_total{result="new-path"} is the "route changed" alerting primitive.
  • MTR Explorer — /mtr three panes (destinations → path history → trace detail), client-side path diff between any two snapshots, a "path changes" timeline overlaid with the pair's loss series, per-hop RTT trends, and a Runner tab (gated on runs:create).
  • Hop enrichment — console.mtr.enrichment.* (off by default): reverse DNS and/or MaxMind GeoLite2 ASN/City mmdb files mounted read-only at /geoip from an operator-supplied volume; resolved server-side into a TTL cache (mtr_hop_enrichment), air-gap friendly, per-source degradation.
  • Time Machine — a top-bar control and shareable ?at= URL state: topology reconstructed from topology_events up to t (GET /api/v1/topology?at=), PromQL surfaces evaluated at t, the Live feed becomes scrollback, and every mutating control is disabled behind the banner while engaged. Note: today's controller does not yet attribute topology_changed events to nodes, so historical topology reconstruction reports honestly-empty results until it does (named deferral).
  • Annotations — notes pinned to instants or time ranges (/api/v1/annotations), rendered as chart markers on Explore and the object cards and inline in the Live scrollback. annotations:write stops at operator/admin; reads are telemetry (every role).
  • Explore A/B — a Compare panel: a second curated metric on the same axes, or the same metric time-shifted (1h/24h/7d) against itself.
  • Permissions — mtr:read, annotations:read (all built-in roles) and annotations:write (operator/admin).

Chart

  • 1.6.0 → 1.7.0. New value blocks: console.mtr.enrichment.* (rendered into the console ConfigMap only when enabled; mmdb volume via an opaque geoip.volume VolumeSource passthrough with a render-time guard). An eighth ci profile (console-mtr-values.yaml). Default render is key-identical to 1.6.0.

kconmon-ng v1.6.0

Console release. External probe targets, saved check definitions and schedules, continuous external checks with agent-side CIDR enforcement, and diagnostics v2 (M4). Off by default throughout: without console.scheduler.enabled, config.checkers.external.enabled or a database, a 1.5.0 install behaves identically.

Added

  • Targets, checks, schedules — CRUD /api/v1/{targets,checks,schedules} (PostgreSQL-backed, 503 without a database), with a cardinality projection guard: an enabled definition may project at most 400 series, enforced server-side and mirrored live in the UI.
  • Console scheduler — console.scheduler.{enabled,tickInterval} (off by default): once/interval schedules fire diagnostics runs from exactly one replica under a PostgreSQL advisory lock; a stuck-run reaper finishes abandoned runs as cancelled; POST /api/v1/runs/{id}/cancel cancels an in-flight run.
  • Continuous external checks — check definitions of kind continuous are reconciled to the controller (PUT /api/v1/external-checks, WatchExternalChecks stream) and probed by agents on their own cadence. The agent is authoritative: config.checkers.external.enabled plus a mandatory allowedCidrs allowlist (deny-wins, resolved-IP matching, re-checked on every probe); an agent without the opt-in refuses external work and the controller answers 501 for it.
  • External metrics — the kconmon_ng_external_* family ({source_node,source_zone,target,target_kind} labels; the target NAME, never an address) plus ExternalChecksFailing in the default rules.
  • Diagnostics v2 — POST /api/v1/runs accepts destinationKind=node|target|adhoc; the diagnostics form grows a destination selector and Save-as-definition.
  • Rate limits — console.rateLimit.{runsPerMinute,loginPerMinute} (fixed-window on the shared KV; fail-open on a Valkey outage; the login limit runs before argon2id).
  • OpenAPI — a committed spec (docs/console-api.yaml), generated TS types, and a router-walking test that fails on any drift in either direction.
  • UI — the Targets & Schedules page and the Target card (/targets/{id}).

Chart

  • 1.5.0 → 1.6.0. New value blocks: config.checkers.external.* (agent, rendered only when enabled), console.scheduler.*, console.rateLimit.*; a seventh ci profile; the external-mode Valkey egress NetworkPolicy now derives its port from the configured address.

kconmon-ng v1.5.0

Console release. Adds durable persistence, authentication/RBAC, and an on-demand diagnostics runner (M3) on top of the M1/M2 Console. Off by default (console.enabled: false, console.database.mode: disabled, console.auth.mode: anonymous); existing installs — including existing Console installs on anonymous auth — are functionally unaffected until you turn these on.

Added

  • PostgreSQL persistence — console.database.mode=cnpg|external|disabled (ADR-001): CloudNativePG-provisioned or externally-supplied PostgreSQL for event history, RBAC, audit, and diagnostics run history. GET /api/v1/events now serves durable scrollback for the Live page, backed by topology_events (all five WebSocket event types, not only topology ones).
  • Authentication and RBAC — console.auth.mode=anonymous|local|header|oidc: local users (PostgreSQL, argon2id), header-based trusted-proxy auth (explicit CIDR opt-in), and OIDC (code flow + PKCE). Built-in roles (viewer/operator/alert-editor/admin) are compiled-in and work with database.mode=disabled; a custom-role admin API layers on top when a database is configured. __Host- session cookies, CSRF double-submit for cookie-authenticated mutations, and an async best-effort audit log (GET /api/v1/audit).
  • API tokens (PATs) — /api/v1/tokens: SHA-256-hashed bearer tokens that work in every auth mode. A PAT is not individually scoped in M3: its effective permissions are exactly auth.defaultRole, deployment-wide — a token subject resolves no role bindings of its own, and token-kind bindings are rejected by the RBAC API by design (they would silently grant nothing). Where a database is configured, disabling a local user's account revokes the tokens they own on that user's next request; subjects created by header or OIDC mode have no users row, so their disable state lives upstream at the proxy/IdP and only DELETE /api/v1/tokens/{id} revokes their tokens.
  • Diagnostics runner — POST/GET /api/v1/runs, GET /api/v1/runs/{id}: bounded on-demand check fan-out (up to 400 pairs) with persisted run history and shareable permalinks. Live per-pair progress streams over a new ephemeral run:{id} WebSocket topic, opened per run, with an automatic REST-polling fallback on any console replica other than the one executing the run.
  • Object cards v1 — Node and Pair cards with a shared "Recent changes" event rail, linked from Topology and the Matrix heatmap.
  • NetworkPolicy — opens console↔database (CNPG pod-selector rule, or a namespace-wide default for mode=external) and console→OIDC IdP on both layers, alongside the existing console→controller/Valkey rules.

Fixed

  • Chart mounts every console secret (database DSN, local-admin bootstrap password, OIDC client secret) under one sibling directory, /etc/kconmon-ng-console-secrets/, group-readable (0440) with console.podSecurityContext.fsGroup matching the distroless nonroot gid — the milestone's originally planned nested path and owner-only mode were both unworkable (a read-only ConfigMap volume cannot host a mountpoint inside itself, and 0400 is unreadable by the nonroot process without matching UID ownership).

Upgrade Notes

  1. This release is safe to roll out with every new feature left at its default (console.database.mode: disabled, console.auth.mode: anonymous): the M1/M2 surface behaves identically.
  2. Every existing Console install rolls once on upgrade, even one that changes nothing in values.yaml. auth.defaultRole and the auth.session.* block are now rendered into the console ConfigMap unconditionally (previously absent keys), so the ConfigMap's content — and therefore its checksum annotation on the Deployment's pod template — changes for every install, triggering one rollout. There is no data or config migration behind it; it is a one-time restart.
  3. To turn on persistence and auth, set console.database.mode to cnpg (this chart does not install the CloudNativePG operator or its CRDs — install those first, or helm install fails with a clear error) or external (supply console.database.existingSecret), then set console.auth.mode and its mode-specific block. See the chart README and the commented values.yaml for the full validation matrix and secret-mount layout.

Install

helm upgrade --install kconmon-ng oci://ghcr.io/esdmitrii/charts/kconmon-ng \
  --version 1.5.0 \
  --namespace kconmon-ng \
  --create-namespace

Images

ghcr.io/esdmitrii/kconmon-ng-agent:1.5.0
ghcr.io/esdmitrii/kconmon-ng-controller:1.5.0
ghcr.io/esdmitrii/kconmon-ng-console:1.5.0

kconmon-ng v1.4.0

Console release. Adds an optional read-only web Console (M1) with a realtime event pipeline (M2). The Console is off by default (console.enabled: false); existing agent/controller installs are unaffected until it is turned on.

Added

  • Console web UI — new kconmon-ng-console binary and image with an embedded SPA: Overview, Matrix, Topology, Explore, PromQL console and a Live event feed. Read-only in this release (anonymous viewer role, banner shown in the UI).
  • Realtime pipeline — the controller exposes a leader-gated EventStream.WatchEvents gRPC stream (controller.events.enabled); the console ingests it and fans events out to browsers over WebSocket (/ws) with per-topic sequencing, snapshot replay and duplicate suppression. The matrix switches from polling to push when the stream is healthy and falls back to REST automatically.
  • Cross-replica fan-out via Valkey — console.valkey.mode supports off (in-process, single replica), bundled (ephemeral Valkey Deployment, no PVC by design) and external. NetworkPolicies open console→controller gRPC and console→Valkey on both sides.
  • Chart — new console.* values (deployment, service, ingress, PDB, NetworkPolicy, bundled Valkey), documented in docs/configuration.md and covered by values.schema.json and a ci/console-values.yaml lint profile.

Fixed

  • Controller graceful shutdown hang with active streaming subscribers — WatchEvents/WatchTasks/WatchPeers handlers now terminate on shutdown and GracefulStop is bounded with a hard-stop fallback (the Tasks/Peers case was latent since v1.3.0, masked by Kubernetes killing the pod after the grace period).
  • Fleet-safe events config — the events key is omitted from the controller ConfigMap when disabled, so controller images without M2 support keep starting under strict config parsing.

Security

  • grpc-go bumped to v1.82.1 (GO-2026-6061, reachable via the event stream).

Install

helm upgrade --install kconmon-ng oci://ghcr.io/esdmitrii/charts/kconmon-ng \
  --version 1.4.0 \
  --namespace kconmon-ng \
  --create-namespace

kubectl plugin (via krew, from the release manifest):

kubectl krew install --manifest-url \
  https://github.com/EsDmitrii/kconmon-ng/releases/download/v1.4.0/kconmon.yaml

Images

ghcr.io/esdmitrii/kconmon-ng-agent:1.4.0
ghcr.io/esdmitrii/kconmon-ng-controller:1.4.0
ghcr.io/esdmitrii/kconmon-ng-console:1.4.0

kconmon-ng v1.3.3

Chart-focused release. The Go agent/controller code is unchanged from v1.3.2; the :1.3.3 images are a version-synchronized rebuild (the release tag drives both the chart version and the image tag).

Fixes

  • ICMP checker on runtimes with a closed net.ipv4.ping_group_range — the ICMP checker opens an unprivileged ICMP "ping" socket (SOCK_DGRAM), which the kernel gates on net.ipv4.ping_group_range, not on NET_RAW. Some container runtimes leave this at the closed kernel default (1 0), so the checker failed with socket: permission denied on those nodes. The agent Pod now sets the safe, namespaced sysctl net.ipv4.ping_group_range=0 2147483647, so ping sockets work regardless of the runtime default.

Chart

  • New agent.podSecurityContext value exposes the agent Pod-level securityContext (defaults to opening ping_group_range for the ICMP checker). Set agent.podSecurityContext: {} to opt out. Documented in the chart README and values.schema.json.
  • values.schema.json: HTTP target field corrected from expectedStatus to expectStatus to match the checker's config (the schema key never matched the code, so a schema-guided value was silently ignored).

Install

helm upgrade --install kconmon-ng oci://ghcr.io/esdmitrii/charts/kconmon-ng \
  --version 1.3.3 \
  --namespace kconmon-ng \
  --create-namespace

kubectl plugin (via krew, from the release manifest):

kubectl krew install --manifest-url \
  https://github.com/EsDmitrii/kconmon-ng/releases/download/v1.3.3/kconmon.yaml

Images

ghcr.io/esdmitrii/kconmon-ng-agent:1.3.3
ghcr.io/esdmitrii/kconmon-ng-controller:1.3.3

kconmon-ng v1.3.2

Note: this is the first fully working release of the on-demand diagnostics feature set. v1.3.0 was aborted mid-release (GitHub immutable releases sealed it before all assets were attached); v1.3.1 published but shipped a krew manifest with an invalid version string, so kubectl krew install --manifest-url rejected it. v1.3.2 carries the same content with a valid krew manifest. The v1.3.0/v1.3.1 tags are retired.

Features

  • kubectl-kconmon plugin (on-demand diagnostics) — a new kubectl plugin talks to the controller's HTTP API through a client-go port-forward, so operators can inspect topology (kubectl kconmon topology / agents) and run one-shot connectivity checks (kubectl kconmon check SRC DST --type …, kubectl kconmon mtr SRC DST) between any two nodes without opening Grafana. Table or -o json output; a failed check exits 2 (distinct from 1 for CLI/API errors) so it composes in shell pipelines. Install via krew from the release manifest (see Install below).

  • On-demand diagnostics API — new POST /api/v1/diagnostics controller endpoint runs a single check (tcp/udp/icmp/dns/http/mtr) from a source node's agent to a destination and returns the CheckResult verbatim. Served by the leader only; ?timeout= caps the wait (default 60s, max 120s). This is the endpoint the plugin drives. See docs/api.md.

  • Graceful agent deregistration on SIGTERM — a restarting agent now deregisters from the controller on shutdown, so peers drop it immediately instead of waiting out the heartbeat TTL. This removes the transient false-loss window that a rolling agent restart used to leave in its own metrics.

Security

  • Toolchain and dependency bumps — Go toolchain go1.26.4; google.golang.org/grpc 1.79.1 → 1.82.0, golang.org/x/net 0.51 → 0.56, golang.org/x/sys 0.41 → 0.46, and OpenTelemetry 1.41 → 1.44. This clears the CVE findings behind the previous Artifact Hub security-report grade.
  • govulncheck in CI — a dedicated CI job runs govulncheck ./... on every PR and tag; Dependabot (gomod / github-actions / docker, weekly) keeps dependencies current so CVE fixes land as normal PRs instead of accumulating until the next scan.

Supply chain

  • The Helm chart is now signed with cosign (keyless, by digest) — v1.3.2 is the first signed release. Artifact Hub repository metadata continues to be published as an ORAS artifact.

Docs

  • README reworked with an "On-demand diagnostics (kubectl plugin)" section and real command output.
  • docs/api.md documents the full POST /api/v1/diagnostics contract (request fields, status codes, ?timeout= cap, and ICMP / MTR response examples).

Install

helm upgrade --install kconmon-ng oci://ghcr.io/esdmitrii/charts/kconmon-ng \
  --version 1.3.2 \
  --namespace kconmon-ng \
  --create-namespace

kubectl plugin (via krew, from the release manifest):

kubectl krew install --manifest-url \
  https://github.com/EsDmitrii/kconmon-ng/releases/download/v1.3.2/kconmon.yaml

Images

ghcr.io/esdmitrii/kconmon-ng-agent:1.3.2
ghcr.io/esdmitrii/kconmon-ng-controller:1.3.2

kconmon-ng v1.2.0

Features

  • Automatic zone discovery — agents no longer need a statically configured zone. The controller enriches each agent registration with the node's failure-domain zone taken from its node informer (failureDomainLabel, default topology.kubernetes.io/zone) and returns the resolved metadata in RegisterResponse.agent; the agent adopts it for all source_zone / destination_zone metric labels. KCONMON_NG_ZONE (agent.zone in Helm values) remains an explicit override and always wins. Node zone relabels propagate to peers via a FULL_SYNC peer update; the relabeled node's own source_zone refreshes on its next re-registration. Per-zone metrics and the Zone Heatmap dashboard now work out of the box on multi-zone clusters.

  • Self-monitoring — new gauge kconmon_ng_controller_expected_agents (count of schedulable nodes from the controller's node informer) and two PrometheusRule alerts: KconmonAgentsMissing (warning: registered < expected for 10m) and KconmonControllerDown (critical: absent(kconmon_ng_controller_leader == 1) for 5m). Degradation of kconmon-ng itself now alerts instead of failing silently. Requires controller.leaderElection: true (default) for the node informer.

Breaking-ish Changes

  • Strict config parsing — the application config (ConfigMap / --config file) is now decoded with unknown-field rejection and per-checker semantic validation (intervals/timeouts > 0 for enabled checkers, HTTP target URL scheme/host, DNS resolver host[:port], non-empty DNS hosts). A typo'd or invalid config now fails startup and is rejected on hot-reload (the previous config stays active) instead of being silently ignored. Review your values overrides before upgrading: a config that previously "worked" by accident will now fail loudly. timeout >= interval logs a warning but does not fail.

Helm Chart / Artifact Hub

  • Chart README is now packaged into the chart archive — the Artifact Hub package page renders description, install instructions, values and metrics reference instead of "This package version does not provide a README file".
  • home and sources added to Chart.yaml; Artifact Hub repository metadata (artifacthub-repo.yml) is published as an ORAS artifact on release for repository verification.
  • agent.zone is now documented as an optional override (auto-discovery is the default).

Dashboards

  • Overview / MTR Triggers Count — switched from increase(...[$__range]) to a plain sum(...): increase() misses counter births on freshly restarted agent pods and chronically undercounted exactly when MTR fires most (pod churn).

Local Development

  • hack/local-test.sh hardening: unique image tag per build (minikube's image-load cache silently kept stale same-tag images on re-runs), set -e/pipefail fixes (((ok++)) pre-increment exit, SIGPIPE on head-truncated pipes), port-forward cleanup.

Upgrade Notes

  1. Validate your config overrides against the stricter parser before rolling out (a quick check: helm template ... | <render your config> and run the controller/agent with --config locally, or just watch pod readiness on a staging cluster first).
  2. If you previously set agent.zone to force a zone, you can keep it (it still wins) or drop it to switch to automatic discovery.
  3. Metric label sets are unchanged; the new alerts ship in the chart's default prometheusRule.rules and are inert unless prometheusRule.enabled: true.

Install

helm upgrade --install kconmon-ng oci://ghcr.io/esdmitrii/charts/kconmon-ng \
  --version 1.2.0 \
  --namespace kconmon-ng \
  --create-namespace

Images

ghcr.io/esdmitrii/kconmon-ng-agent:1.2.0
ghcr.io/esdmitrii/kconmon-ng-controller:1.2.0

kconmon-ng v1.1.0

Bug Fixes

  • MTR memory leak — lastRun map in MTRChecker could grow unboundedly in long-running agents on large clusters where node pairs come and go. Expired entries are now purged inline on each TryAcquire call while the lock is already held, keeping the map size proportional to active pairs within the current cooldown window.

  • HTTP body pattern mismatch counted as success — when a bodyPattern check failed, the checker set StatusCode = -1, which was not caught by the result handler's >= 400 guard and was silently recorded as result="success" in Prometheus. The status code field now always carries the real HTTP status. A dedicated BodyMismatch bool field signals pattern failure, and the result handler correctly marks such checks as result="fail".

Improvements

  • Configurable DNS resolver dial timeout — the dialer timeout for custom DNS resolvers was previously hard-coded to 5 seconds and could not be adjusted for slow or distant resolvers. A new timeout field has been added to the DNS checker config (default: 5s). Update your Helm values or config file to override:

    checkers:
      dns:
        timeout: 3s
    

  • Jitter in agent re-registration backoff — when the controller restarts, all agents previously retried at exactly the same interval, causing a thundering herd. Up to 25% random jitter is now added to each retry wait, spreading reconnect load across agents.

  • MTR buffer allocation — the 1500-byte read buffer in the traceroute loop was allocated once per hop. It is now allocated once per trace, reducing GC pressure under frequent MTR runs.

Helm Chart

  • config.checkers.dns.timeout added to values.yaml (default: 5s).

Tests

  • Updated TestHTTPCheckerBodyPatternMismatch: verifies BodyMismatch=true and real HTTP status code instead of the former -1 sentinel.
  • Added TestHTTPCheckerBodyPatternMatch: verifies BodyMismatch=false on a successful pattern.
  • Added TestDNSCheckerTimeoutPropagated: verifies the configured timeout is stored on the checker.
  • Added TestMTRCheckerExpiredEntriesPurged: verifies stale entries are removed from lastRun after cooldown expiry.

Upgrade Notes

The HTTPDetails.StatusCode field no longer returns -1 for body pattern mismatches — it now always holds the actual HTTP response status code. If you have alerting or dashboards that rely on statusCode == -1 to detect body mismatch failures, update them to use the new bodyMismatch field in the JSON result or the result="fail" label in Prometheus metrics.

Install

helm upgrade --install kconmon-ng oci://ghcr.io/esdmitrii/charts/kconmon-ng \
  --version 1.1.0 \
  --namespace kconmon-ng \
  --create-namespace

Images

ghcr.io/esdmitrii/kconmon-ng-agent:1.1.0
ghcr.io/esdmitrii/kconmon-ng-controller:1.1.0

kconmon-ng v1.0.0 — Initial Release

Kubernetes Node Connectivity Monitor, next-generation rewrite with a gRPC-based agent/controller architecture and rich observability out of the box.

Features

Core - Agent/controller architecture with gRPC streaming peer updates - TCP, UDP, ICMP, DNS, and HTTP checkers with configurable timeouts and thresholds - Per-node and per-zone Prometheus metrics for all check types - Reactive MTR traceroute on check failure with per-pair cooldown - Self-probe prevention: peers filtered by agent ID, node name, and pod IP - Atomic gauge reset on peer topology changes to prevent stale metrics

Scheduler - Pause/resume support, per-check jitter, and NodeLocal checker mode - NodeWatcher: live Kubernetes node info exposed via /api/v1/topology

Observability - Grafana dashboards: Overview, Node Detail, Cross-Zone Heatmap - Helm chart with ServiceMonitor, PrometheusRule, NetworkPolicy, PDB, and RBAC

Operations - Multi-arch Docker images (linux/amd64, linux/arm64) published to GHCR - Local dev tooling: hack/local-test.sh with Minikube + Prometheus + Grafana stack - Chaos testing guide with NetworkPolicy example

Install

helm install kconmon-ng oci://ghcr.io/esdmitrii/charts/kconmon-ng \
  --version 1.0.0 \
  --namespace kconmon-ng \
  --create-namespace

Images

ghcr.io/esdmitrii/kconmon-ng-agent:1.0.0
ghcr.io/esdmitrii/kconmon-ng-controller:1.0.0