Release notes¶
kconmon-ng v2.5.1¶
Changed¶
- Shorter alert texts. Every built-in alert's
descriptionis now one or two sentences on what to check first, down from up to a thousand characters, so a notification template that prints one per firing alert no longer turns four pairs into a wall of text. The explanations moved to the new Alert runbooks page. Node-pair alerts no longer print zones, which read as(zone )on clusters without zone labels, andPathMTUBlackHole's summary now reads "only packets up to N bytes get through".
Added¶
runbook_urlon every built-in alert, pointing at its section of the Alert runbooks page.namespacelabel on every built-in alert, set to the release namespace. The aggregated rules had none, so notification templates showedNamespace: unknown, and an Alertmanager inhibit rule withequal: [namespace]treated them as matching every other alert without a namespace.
Fixed¶
- The console's notice when
console.alerting.enabledis off read as if alerting as a whole were off. It now says that only the rules built in the console are not applied, and that the chart's built-in rules keep alerting through Prometheus.
Upgrade notes¶
- Alertmanager routes and inhibit rules that match on
namespacenow see kconmon-ng's alerts in the release namespace; check routes that send a namespace to an application team. - Templates, silences or tests that match the old
summaryordescriptiontext need the new wording. Expressions, thresholds and the other labels are unchanged.
kconmon-ng v2.5.0¶
2.5.0 adds a path MTU probe for the failure small probes cannot see: a pair where handshakes and pings cross while full-size packets vanish now turns red with the size that still crosses, instead of staying green on every plane. The rest of the release pays the debts a first outside user runs into: node-level alerts, maintenance windows that hold the console's webhooks, local users managed from the console, a config reload that survives the way files are really replaced and applies what it reads, a console whose first load is a fifth of what it was, and a set of security fixes.
Read the Upgrade notes before rolling out: the probe is on by default, two of the four new rules are critical, and the chart's NetworkPolicy is split per component.
Added¶
- Path MTU probe (
config.checkers.pmtu, on by default). Once a minute every agent sends each peer's UDP echo port a 64-byte datagram and a full-size one with Don't Fragment set. The full size is the MTU of the route to the peer (the route's ownmtuwhen the CNI sets one, as Cilium does, else the egress device's);sizeoverrides it. When the full size does not come back, the agent bisects (at most 16 sizes,timeout500ms each) and confirms both ends, so a lossy path does not pass for a black hole. A pair readsok(the full size is echoed),reduced(the path answers ICMP frag-needed: TCP adapts, UDP without its own path MTU discovery does not) orblackhole(full-size datagrams vanish while the small one crosses, and large transfers stall). A lost small datagram is a connectivity failure, left to the UDP plane, and a pmtu failure does not trigger MTR. The agent warns whenintervalis under 28 timeouts (14s) or 3m and more. New series:kconmon_ng_pmtu_bytesandkconmon_ng_pmtu_probe_bytes(per pair: the size that crossed, the size probed),kconmon_ng_pmtu_results_total(successfor ok and reduced,failfor a black hole),kconmon_ng_zone_pmtu_results_totalandkconmon_ng_agent_pmtu_probe_bytes. Agents advertiseplane:pmtu. Walkthrough in Catch an MTU black hole. PathMTUBlackHoleandZonePathMTUBlackHole(prometheusRule.pathMtuBlackHole, warning): more than half of a pair's probes failed over 10 minutes, or more thansustainedThreshold(0.1) over 30 minutes with at least two failures and one in the last 10, held for 5. The second arm catches a black hole on one of several ECMP paths. The zone rule fires only where no per-pair pmtu series exist (agent.metrics.detail=zone-only). A network that carries less than its routes say on purpose and clamps TCP MSS can setconfig.checkers.pmtu.sizeor turn the rule off.NodeUnreachableandNodeIsolated(prometheusRule.nodeUnreachable,.nodeIsolated, critical,for: 5m): most peers fail most of their TCP probes to a node, or a node fails to reach most of its peers, with at leastminPeers(2) reporting. They catch a node that stays registered behind a host firewall, a NetworkPolicy or a broken CNI datapath; the inhibit rules on the metrics page fold the per-pair alerts under them.- Path MTU on the dashboards. Overview opens with the key indicators in two rows (agents, leader, pairs, pairs with failures, black-hole and reduced-path pairs), then the worst-pair and MTR bars, with the charts and tables below, pairs below their probe size among them. Node Detail: the path MTU to and from each peer and black-hole probes by peer. Panels show the smallest size of the last 10 minutes, so an ECMP-split black hole stays on screen.
- PMTU in the console and the CLI. The matrix gains a PMTU protocol: the
path MTU in bytes, green at full size, amber on a reduced path ("1400 of
1500") or while recovering, red while recent probes fail, dashed "Not run"
for a 2.4.x agent. The Overview, node and pair pages follow it. Run checks
accepts
pmtubetween nodes, andkubectl kconmon check <source> <destination> --type pmtuexits 2 on a black hole. - Local users in the console. With
auth.mode=local, Settings > Users adds, re-roles, resets, disables and deletes accounts under the newusers:managepermission (built-in admin only), which the last enabled holder cannot lose. Every local user can change their own password. A password change or reset ends the user's other sessions; a disable or delete also revokes their API tokens. Routes under/api/v1/usersin the Console API. - NetworkPolicy keys (Upgrade notes 4 to 6):
networkPolicy.dnsEgressreplaces the default DNS rule, for NodeLocal DNSCache and other host-network resolvers;networkPolicy.ciliumKubeAPIEgress(auto) adds CiliumNetworkPolicies for apiserver and node traffic, which no ipBlock matches on Cilium;networkPolicy.clusterCIDRscarves pod and Service CIDRs out of the default0.0.0.0/0egress, which on Calico and Antrea matches pods;console.networkPolicy.prometheusTargetPortopens the pod port behind a Prometheus Service that maps it (Thanos 9090 to 10902).
console.clientAddress.trustedProxyCIDRs: the proxies whoseX-Forwarded-Fornames the client for rate limits, the WebSocket cap and the audit log, never for identity (Upgrade note 20).agent.tls.enabled(false): TLS verified against the system trust pool with no other TLS field set, for an external agent whose gateway has a publicly signed certificate.
Changed¶
- Maintenance windows hold the console's alert webhooks. An alert that
starts firing inside a matching window (fleet-wide, a node, the pair or its
target) is delivered only if it still fires when the window closes, and not
at all if it resolves inside. A restarted console keeps holding.
kconmon_ng_console_webhook_suppressed_total{event}counts what was held. - Config hot reload applies what it reads. The keys that go live and the ones that wait for a restart are in Upgrade note 11 and What reloads and what does not.
- The node page covers both directions, so a node
NodeUnreachablenames no longer reads Healthy there, and its peer breakdown switches between To peers and From peers. Diagnostic runs list failed pairs first. - Console pages load on demand. The first page load drops from 935 kB to 198 kB gzipped, and charts load only the ECharts parts they draw with.
- DNS probes against an explicit resolver ask the absolute name, so the
search list no longer multiplies the query or sends cluster names outside.
checkers.dns.timeoutdefaults to 2s (was 5s). - A departed peer's series go away ten minutes after it leaves the agent's peer list.
- The controller refuses what cannot run.
POST /api/v1/diagnosticsanswers 400 for an external destination with a type other thantcp,icmpormtror for aplaneother thanpod, 501 when the source agent does not run the type, and 503leadership lostwhen the lease goes mid-task.kubectl kconmon checkexits 1 on these, andkubectl kconmonfinds the leader itself. - So does the console. Runs, definitions and schedules toward a target or
an ad-hoc address take only
tcp,icmpandmtr(a continuous schedule alsodnsandhttp), and an edit that would leave a stored schedule or definition unable to run answers 422;enabled: falsealways saves. The import follows the same rules, and the forms offer only what runs. - Console API input. Malformed input answers 400 or 422 instead of 502, and alert rule names that become one Prometheus alert name are refused.
- The console's agent-missing template renders the chart's
KconmonAgentsMissingexpression; existing rules change on the next sync. ZoneLossHighdescription. Below about ten node pairs between two zones one broken link crosses the default 10% alone; raiseprometheusRule.zoneLossHigh.thresholdandzoneChecksFailing.threshold.- Continuous external checks fit the controller's 8 MiB limit: the console
leaves out whole definitions, newest first, and counts them in
kconmon_ng_console_external_specs_skipped_total{reason="over-budget"}. - Console UI. A silent MTR hop shows a dash, not 100% loss; refused forms focus the refused field; a foreign rule import asks to confirm; phones get a drawer Close button and tables that scroll inside their card.
- Chart. The console gets
GOMEMLIMITfrom its memory limit. The GeoLite2 sidecar moves toghcr.io/maxmind/geoipupdate:v8.0.0. The schema refuses what the binaries refuse at startup. The install notes flag an Ingress with no trusted proxies and an external gateway neither onexternalTrafficPolicy: Localnor behindloadBalancerSourceRanges. - deb/rpm. The packaged agent config ships with the DNS checker off, and
the postinstall keeps a
net.ipv4.ping_group_rangethe admin already set. - Release images.
:latest, the chart and the GitHub release follow a tag only after e2e passes on its images, and only the newest stable release moves:latest, the Latest badge and the krew index. - Build stack. Go 1.27.1, distroless
static-debian13, Vite 8. The unused OpenTelemetry SDK is gone;observability.otel.*logs a warning.
Fixed¶
- Agents.
- Hot reload stopped for good after the first atomic replacement of the file (an editor's save, a puppet file resource, a ConfigMap swap).
logLevel: DEBUGandlogFormat: TEXTran at info in JSON.- An agent probed at a secondary IP, a multi-homed
advertiseAddressor a VIP read as 100% UDP loss. - A dead DNS resolver read green for names in
/etc/hostsorhostAliases. - Departed peers, targets and zones left stale series behind.
- A node relabel overwrote an explicit
agent.zone.
- Controller.
- After a failover, pairs towards late agents went unprobed for about a minute.
- A rolling restart left no leader for the 15s lease; now about 2s.
- An external-check assignment or removal that an agent's stream could not take was lost.
- A subscriber that stopped reading was never cut off, except on the peer list.
- Console.
- Chart tooltips were white on the dark theme and ignored a theme switch.
- An unreachable IdP, or replicas refreshing a rotating token at once, signed OIDC users out.
- A PostgreSQL restart or a Valkey failover signed everyone out; the console now answers 503 until the store is back.
- A custom
config.metricsPrefixbroke the browser views. - Rate limits failed on Redis and Valkey before 7.0, and Redis 5 or an
endpoint without
CLIENT TRACKINGleft each replica on its own state. - A transient apply failure, a rule-name collision or a shared
bundleNamecould delete, overwrite or freeze the managed alert rules. - A
prometheus.queryTimeoutabove 30s never took effect. /metricshad nogo_*orprocess_*families.- The Overview read healthy a pair the matrix painted red for loss.
- Deleting a target in use on PostgreSQL 18 gave a store error, not 409.
- The Time Machine lost a node when one of its two agents deregistered.
- A failed MTR hop lookup waited the full enrichment TTL for a retry.
- A cancelled run left pairs at
dispatched, and permalinks could miss their final frames. - Concurrent role changes could leave a user with two roles.
- Investigate lost unsaved notes, Time Machine incident lists kept resolved incidents, and failed reads showed "none".
- Chart and dashboards.
helm testnever started its Pod (CreateContainerConfigError).- An unquoted numeric geoip
accountIdcrash-looped geoipupdate. PairWentSilentsaid 15m whatever itsfor.- The console's scrape rule needed
serviceMonitor.enabled, andKconmonExternalAgentDownmissed a customjobName. - The Overview dashboard's Agents missing let external agents mask missing ones, and several panels went blank with UDP off.
- Under a long release name the dashboard ConfigMaps collided.
Security¶
- NetworkPolicy. Every agent pod and every
externalPeerCidrsaddress reached the controller's unauthenticated gRPC, HTTP and metrics ports. - Chart RBAC. The console could change or delete every PrometheusRule in its namespace, and the controller could update any Lease there.
- UDP echo loop. One spoofed datagram could start an endless echo loop between two agents.
- Webhooks. A receiver could redirect a signed delivery to any host, and failed deliveries logged the webhook URL, often a credential itself.
- External gateway. Any token holder could read the event stream, and unauthenticated connections piled up without limit.
- Controller. One large request body could run it out of memory.
- HTTP check URLs. A password in a target URL leaked into the
urllabel, logs, diagnostic results and config errors. - Console.
- A burst of logins could run a 256Mi console out of memory.
- A crafted
X-Forwarded-Forcould become a rate-limit key or an audit address, and anyone behind the ingress could spend every user's sign-in budget. - A flood of refused requests could push sign-ins and admin actions out of the audit log, and OIDC sign-ins were not audited.
run:{id}WebSocket topics skippedruns:read, and/wstook any number of sockets.GET /api/v1/exportgave a role withsettings:writeevery section, webhook URLs included.- A credential-less 16 MiB body cost about 50 MiB of heap, other sites could frame the console, and Redis errors logged session keys.
Upgrade notes¶
- The path MTU probe is on by default and sizes itself from the route. On
CNIs that set a route MTU the pmtu gauges read it (1450 on Cilium with
VXLAN), not the pod's 1500. The chart renders no
pmtukey until you tune one, so 2.4.x agents keep running without the probe and still answer it. Setconfig.checkers.pmtu.*oragent.tls.enabledonly with 2.5.0 agents: an older agent refuses the unknown key. - Cardinality. Up to four new series per directed pair and one per agent,
about 40 thousand more at 100 nodes.
agent.metrics.detail=zone-onlydrops the per-pair ones;config.checkers.pmtu.enabled=falsedrops all. - Two of the four new rules are critical. Check Alertmanager routing for
NodeUnreachableandNodeIsolated, or lowerprometheusRule.nodeUnreachable.severityandprometheusRule.nodeIsolated.severityfor the first week. A black hole resolves up to ten minutes after its last failed probe unlessprometheusRule.pathMtuBlackHole.sustainedThresholdis 1. - The NetworkPolicy is split per component. With
networkPolicy.enabled,helm upgradereplaces<fullname>with<fullname>-agentand<fullname>-controller; the controller admits agents andnodeCidrsongrpcPortonly and nothing fromexternalPeerCidrs. Update your own policies that name the old object or relied on the wider rules. - Cilium gets its own policies. With
networkPolicy.ciliumKubeAPIEgressatautothe upgrade creates the CiliumNetworkPolicy<fullname>-kube-apiserver, plus<fullname>-node-ingressunderagent.hostNetworkor the external gateway.helm templatewithout cluster access needs--api-versions cilium.io/v2/CiliumNetworkPolicyortrue; setfalseif your own policies cover that traffic. - Off-cluster egress depends on the CNI. The default HTTP-check, webhook,
OIDC and geoipupdate rules allow
0.0.0.0/0on their ports, which on Calico and Antrea opens pods too: list the pod and Service CIDRs innetworkPolicy.clusterCIDRs. The OIDC rule also admits any pod on 443; to close that, or to reach an IdP pod on its own port, setconsole.networkPolicy.oidcEgress. - From 2.4.x, upgrade with
--reset-then-reuse-values(Helm 3.14+) or-f.--reuse-valueskeeps every changed default at its old value (DNS timeout 5s, geoipupdate v7.1.1); a release older than 2.3.0 stops with a message. - Changed chart defaults.
config.checkers.dns.timeoutdrops to 2s, and the console getsGOMEMLIMITfromconsole.resources.limits.memory.checksum/secretchanges once, then ignoresconsole.auth.local.secret.password, which only seeds an empty users table. - geoipupdate v8 sidecar. With
console.mtr.enrichment.geoip.mode=autothe sidecar runs v8.0.0, which stops at the first edition that fails to download;console.mtr.enrichment.geoip.image.tag=v7.1.1keeps the old behaviour. - Stricter validation. The agent and the controller at startup, and the
chart schema at
helm upgrade, refuse an enabled checkerintervalunder 100ms ortimeoutunder 1ms (5nsfor5s), a zeromtr.cooldown,topology.sparse.zoneChordsabove 64 and an invalidmetricsPrefix,bodyPattern,expectStatus, resolver port or CIDR, and the agent a loopback or unspecifiedagent.advertiseAddress. Each error names the key.modeandobservability.otel.*load with a warning: remove them. - Hot reload applies now. An in-place edit of a ConfigMap or config file
goes live when saved, so review it first. Live:
logLevelandcheckers.*on the agent;logLevel,topology.*,controller.agentTtlandcheckers.external.enabled/.allowedCidrson the controller. Anything else,logFormatand ports included, logs a warning and needs a restart. - DNS against an explicit resolver needs full names. A relative name
such as
kubernetes.default, or one only/etc/hostsknows, now fails againstcheckers.dns.resolversand in external DNS checks: use an FQDN. - The UDP echo refuses some sources and listens twice. Ports below 1024
and known agent echo endpoints get no reply. Full peer lists carry the
whole fleet's echo endpoints (about 20 KB at 1000 agents, sparse mode
included), and
ss -lunshows two listeners onconfig.grpcPort. - Masked HTTP check labels. A target URL with a password gets a new
urllabel value (https://user:xxxxx@host/...); update queries on it. users:manageis new and admin-equivalent: a holder can create an admin. Add it to custom roles that should administer users.runs:readfor live runs. Arun:{id}WebSocket subscribe needs it; a custom role withevents:readalone loses live run progress.- Exports follow section permissions. Each section of
GET /api/v1/exportalso needs its read permission (targets:read,checks:read,alerts:read,webhooks:manage,maintenance:read,rbac:manage), else it is listed inomitted: grant backup roles all six. - Sessions. Local sessions from before 2.5.0 last until
console.auth.session.ttl(12h). A password change now signs the user out elsewhere, and a disable revokes their API tokens for good. - Login lockout.
console.rateLimit.loginPerMinute(5) counts attempts per username, so anyone who knows a username can keep it locked out; where that matters, use OIDC or header auth, or limit sources at the ingress. - Set the client-address proxies behind an Ingress, in every auth mode.
List the ingress controller's addresses (its pod CIDR at the widest) in
console.clientAddress.trustedProxyCIDRs, or every client shares one address for rate limits, the WebSocket cap and the audit log. While it is empty the console readsconsole.auth.header.trustedProxyCIDRs, which inauth.mode=headeralso decides identity: keep that to the auth proxy. - WebSocket caps. Each console replica accepts 1024
/wssockets (console.websocket.maxConnections), 256 per client address (.maxConnectionsPerAddress) and 32 per user or token (.maxConnectionsPerSubject); 0 turns a cap off. Behind an unlisted proxy a whole team shares the 256. The chart writes these keys andclientAddressonly for console images 2.5.0 or newer. - OIDC sign-in state lives in the browser, sealed under a key derived from the client secret: rotating the secret fails sign-ins in flight for up to 5 minutes, and one that crosses versions mid-rollout fails once.
- Maintenance windows hold console webhooks only; Alertmanager routes still need a silence. Existing windows start holding once the console runs 2.5.0: review them, since a long global window holds every managed alert.
- Webhook receivers must be the final URL. A 3xx now fails the delivery.
- Give the console's alert bundle its own name. With
prometheusRule.enabled,helm upgraderefuses aconsole.alerting.bundleNameequal to the chart's own PrometheusRule, and any other collision shows a sync error on every rule. - Rename rules whose alert names collide. Of two stored rules with one
Prometheus alert name, the one not deployed under it shows the sync error
alert name collision. - Schedules the agents cannot run. A stored once or interval schedule of
a
udp,dnsorhttpdefinition toward a target or an ad-hoc address, or with aplaneother thanpod, no longer starts runs: pause, delete or repoint it. A storedudpdefinition toward a target saves only withenabled: false. - API limits.
POST /api/v1/diagnosticstakes 64 KiB andPUT /api/v1/external-checks8 MiB. A definition'sparamshold 4096 bytes and an address 2048: shorten a longer stored row before editing it.kubectl kconmon check --planetakespodonly. - Dashboard ConfigMap names. From a release fullname of 41 characters up, the dashboard ConfigMaps get new names; tools that look them up by name need them.
- A signed image can precede its release. A tag pushes and signs
:<version>before e2e, and a failed e2e leaves it with no chart, release or:latest. Deploy, or mirror:latest, once the release is published.
kconmon-ng v2.4.0¶
External agents stop being second-class. A host outside the cluster now tells its peers where it listens, gets scraped without a hand-written target, and shows up as what it is in the console, the CLI and the Time Machine. One rule comes with it, and it is the one to read before rolling out: per-agent ports are honoured only by upgraded agents; keep one port set until every agent, deb/rpm hosts included, runs 2.4.0. An older agent reports no ports and dials every peer on its own configured values, so in a fleet whose ports differ each old agent goes one-way red toward every peer listening elsewhere, its on-demand diagnostics included. The full skew matrix is under Upgrade notes at the end of this section.
Added¶
- Per-agent ports on the wire.
AgentMetagainshttp_port,udp_portandmetrics_port. Every agent reports its three listener ports at registration, and peers probe it on the ones it reported, for the scheduled mesh and for on-demand tasks alike, so an external host no longer has to mirror the cluster's port pair. Zero means "not reported" (an agent older than 2.4.0): the prober then dials its own configured port, each port falling back on its own. The controller refuses a port above 65535, peer lists carry the two probe ports but notmetrics_port(nobody dials it), and an agent never adopts ports from the controller's reply; zone stays the only thing it takes from there. See Ports. - Prometheus HTTP SD for external agents. A bare host has no Service for
a ServiceMonitor to select, so the controller, the one party that knows
the host registered and on which address, now publishes it:
GET /api/v1/prometheus/sd, served onhttpPortand onmetricsPort(the port the chart's scrape NetworkPolicy already opens). The contract:- one target group per external agent,
<advertised address>:<metricsPort>, sorted by node name and deduplicated by address; - a fixed label set,
node,zone,external="true"andagent_id. An agent's own labels never reach Prometheus, so a host cannot inject target labels; - with no external agent registered the body is the literal
[], and the list is always served withCache-Control: no-store; - a standby answers
503 not the leader, never200 []: Prometheus reads every 200 as the complete target set, so an empty one from a standby would wipe every external target, while on a non-200 it keeps the list it has; - an agent that reported no metrics port (older than 2.4.0) is published
on the controller's own
config.metricsPort, and the controller logsmetrics port assumed from controller configonce per agent, which is the clue when such a host on another port sits atup == 0; controller.prometheusSD.enabled: falsecloses the route (404 on both listeners). The key reaches the shared ConfigMap only when false, so an older controller image never trips over it;- with
controller.replicaCount > 1the controller Service spreads refreshes over all replicas and roughly half of them land on a standby:prometheus_sd_http_failures_totalclimbs for the job while the targets stay correct. Cosmetic, and written down so nobody chases it. Body and semantics in the HTTP API reference.
- one target group per external agent,
scrapeConfig.externalAgentsin the chart. Renders a Prometheus OperatorScrapeConfig(needs thescrapeconfigs.monitoring.coreos.comCRD) named<release>-agent-externalthat reads the SD route, withlabelsfor your Prometheus' selector (kube-prometheus-stack wantsrelease: <its release name>),jobName,refreshInterval(30s) andinterval. It applies the sameagent.metrics.detailvalve as the agent ServiceMonitor, so an external host never returns per-pair detail the valve drops for the pods, and the valve no longer insists onserviceMonitor.enabledwhen this is on. The chart refuses the ScrapeConfig withoutcontroller.externalGateway.enabled(nothing external could register) or withcontroller.prometheusSD.enabled=false(every refresh would 404), and the install notes remind you when the gateway is on without it, or whenlabelsis empty. Plain-Prometheushttp_sd_configsjob and the reachability rules in Scraping external agents.KconmonExternalAgentDownandkconmon_ng_controller_external_agents. An optional warning (prometheusRule.externalAgentDown, off by default,for: 5m) onup{job=~".*agent-external.*"} == 0: a host the controller lists that Prometheus cannot scrape, usually the host firewall admitting the Prometheus pod IP when the CNI NATs its egress to a node IP. The new controller gauge counts registered agents that came through the gateway.agent.hostNetwork, for pod networks external hosts cannot route. The DaemonSet moves into each node's network namespace: the agents advertise the node IP (KCONMON_NG_POD_IPfromstatus.hostIP), declarehostPorton all three ports, getdnsPolicy: ClusterFirstWithHostNetunlessagent.dnsPolicysays otherwise, and label themselveskconmon-ng.io/host-network=truefrom a 2.4.0 image. The chart stops rendering theping_group_rangepod sysctl there, since the kubelet refusesnet.*sysctls in the host namespace. It changes what is measured, for the whole DaemonSet: every in-cluster pair then probes node IP to node IP over the underlay, and the CNI datapath (overlay, conntrack, NetworkPolicy enforcement) is no longer on the probe path, so the breakage this tool exists to catch can hide behind a green matrix. Turn it on only when the goal is visibility between external agents and a cluster whose pod network they cannot reach. Before you do: PSSprivilegedfor the namespace, TCP 8080, UDP 9090 and TCP 9091 (or yourconfig.*Portvalues) free on every node,ping_group_rangeset by the node OS, and one agent per machine (a host-network pod and a bare-host agent cannot share an IP). See When the pod network does not route and Host networking.networkPolicy.nodeCidrsandnetworkPolicy.externalPeerCidrs. Host-network agents register from node IPs that no pod selector matches, so withagent.hostNetworkandnetworkPolicy.enabledboth on, the chart refuses to render the policy untilnodeCidrslists the node CIDRs; without it every registration but the one from the controller's own node would drop silently.externalPeerCidrscloses the old gap where an external agent registered fine and every cell between it and the cluster stayed red: its CIDRs join the agent-to-agent rules in both directions (UDPgrpcPort, TCPhttpPort, the ports-less ICMP/MTR rule) and never the gateway rule.- External agents in the console. Everything keys off the
kconmon-ng.io/externalregistration label, which the console now passes through from the controller's topology together with the agent's capabilities (labelsandcapabilitiesonTopologyAgentin the Console API).- Topology draws the host beside the cluster nodes in the lane of its zone, with a neutral external badge (identity, never a health tier) and "readiness unknown" for screen readers.
- Node page swaps Pod IP for Advertised address, explains the Ready
dash, and lists the probe Planes the agent advertised. Agents now
advertise
plane:tcp,plane:udp,plane:icmp,plane:dns,plane:httpandplane:mtr; an agent advertising none (older than 2.4.0) reads as "unknown", never as running nothing. - Overview badges the host in Worst pairs and adds "+N external agents" beside Nodes ready without counting them in, since that tile is Kubernetes readiness.
- Matrix tells two silences apart from plain no-data. An external agent Prometheus is not scraping keeps the no-data fill and aria text, but its cells' tooltip and a note above the grid say why and link the scraping docs, until the first measured cell appears. A protocol the source does not run renders dashed like not probed, with its own legend row. Precedence when a cell has no data: excluded by the plan, then unsupported, then unscraped. See Silence with a known cause and External agents on the map.
- External agents in the Time Machine.
TopologyChangedevents carry the agent's labels, so a replay badges a host the way the live view does, and a reconstructed topology lists a bare host underagentsonly, never as a presence-derived READY node. History recorded before the upgrade shows no external badges: a 2.3.x controller wrote no labels, and such a host stays an ordinary node in those instants. Historical responses never carry capabilities, since no event records them.
Changed¶
KconmonAgentsMissingis no longer masked by external agents. Registered agents include them and expected agents (schedulable nodes) never did, so one external host hid one missing in-cluster agent. The expression now subtractscontroller_external_agents, with anor registered * 0stand-in so the rule keeps working against a controller image that predates the gauge.kubectl kconmontables.agentsgains anEXTERNALcolumn (yes, or-for the DaemonSet's own), andtopologyprints bare-host rows after the node rows (NODE is the registered name, READY is-, AGENT IP the advertised address). Scripts that parse the human tables will see the columns and rows shift;-o jsonstays the controller's body, unchanged apart fromhttpPort,udpPortandmetricsPorton agents that report them.- Documentation lives on the site. The chart's Artifact Hub links, the
install notes (
Docs:,Helm values:,Console setup:), the chart README,kubectl kconmon --help, the krew manifest, the packaged agent config, the image OCI labels and the GitHub release footer now point at https://esdmitrii.github.io/kconmon-ng/. In the console, Settings → About gains Documentation, Release notes and Source links, and the command palette an Open documentation action. The README turned into a short front door with a Scope and limits section in place of the old "Not yet" list. - Screenshots and the demo match one stand. The docs frames were re-shot
on one kind stand with an external agent,
edge-host-01, in the mesh; the captions were rewritten to describe what each frame shows, near-duplicate frames were folded into one, and the API tokens frame no longer shows a token. The breaking-the-network demo was rewritten for that kind stand instead of the old Minikube helper, with the stand's own names and numbers. - Console polish. One type scale across pages, row actions reduced to
their verbs (Test, Edit, Delete), and status and identity badges drawn as
one system. Chart hover pills print the time as
HH:mm:ssinstead of a date wide enough to clip, the pill on a two-axis chart reads each axis in its own unit, and the topology map zooms out far enough to fit a phone. - Windows is explicitly unsupported. The agent compiles for
windows/amd64, and TCP, UDP, DNS and HTTP would work as written, but the two checkers that make this tool what it is would not: ICMP and MTR sit on a datagram ICMP socket thatgolang.org/x/net/icmpsupports only on Linux and Darwin by its own contract, and raw ICMP on Windows needs the Administrators group, so an agent that also runs on-demand probes for the controller would run as SYSTEM. On top of that, Go's clock on Windows is interrupt-tick granular (up to 15.6 ms), so a 0.3 ms LAN round trip would read as zero or as one tick while the histograms looked perfectly valid. A Windows vantage point without ICMP and MTR is designed but not scheduled; the trigger is a concrete host that needs it. CI now runsGOOS=windows go vetover the agent and checker trees so the door stays open at no cost. See the FAQ. - CI guards. The buf plugins are pinned and CI regenerates the protobuf
code and fails on a diff, so a
.protoedit that skippedmake protocan no longer ship a wire skew. Chart CI checks the host-network and ScrapeConfig renders, and that the chart refuses a host-network policy withoutnodeCidrsand a ScrapeConfig without the gateway. The kind e2e gained a leg that turns the gateway on, joins one simulated external agent through it, and checks the topology, the SD body, a Prometheus scrape of the host, the console matrix and the CLI.
Fixed¶
-
The
-arm64agent and controller images shipped an x86-64 binary. The arm64 image entries in.goreleaser.yamlnever setgoarch, goreleaser defaults it to amd64, so every published-arm64image and the arm64 half of the multi-arch tags carried the amd64 build since the first release (checked onkconmon-ng-agent:2.3.1andkconmon-ng-controller:2.3.1). On an arm64 node the container failed withexec format error; amd64 clusters were never affected. 2.4.0 builds both from the arm64 binary. The console image was not affected: it is built outside goreleaser. -
The Time Machine topology forgot agents that never changed. Topology events only say what changed, so a console that started recording next to a running fleet never showed an agent that stayed put afterwards, external agents included, and retention pruning stripped old registrations the same way. The console now stores the controller's whole topology as a
topology_baselinerow each time its event stream connects and hourly while it stays connected, and a reconstruction starts from the newest baseline at or before the instant, then replays the events after it. The events list never shows these rows. The cost, on consoles with a database and the event stream: one row per console replica per hour, plus one per reconnect, sized by the fleet at roughly 170 to 230 bytes of JSON per node with its agent, so about 20 KB an hour per replica at 100 nodes. History recorded before the upgrade stays baseline-less and folds from events alone, with the same gaps as before. - Grafana dashboards. Ratio panels in Overview and Node detail cap their
axis at 100% (
axisSoftMax: 1) instead of stretching a flat zero line to a 0-10000% axis. Node detail's MTR trace count is plain text now: it counts traces, including the ones the console asked for, and colouring it green, yellow or red passed that off as a health verdict. The Zone heatmap tables sort rows and columns the same way, so same-zone cells line up on the diagonal, under afrom \ tocorner header. - Console accessibility. Buttons and links meet 4.5:1 contrast in both themes (the dark theme's destructive fill and the light theme's primary fill, focus ring and muted text were darkened), links inside running text carry an underline instead of relying on colour, scroll regions are keyboard-reachable and labelled, headings follow page order, the live event feed is a valid list inside a labelled log region, and phones get a proper header landmark.
- Tooltips covered what they described. The entrance animation's final
transformoutlived the animation and overrode the tooltip's own positioning, so a matrix tooltip landed on the very cell it explained. - Topology edges never showed their failure label on hover: React Flow disables pointer events on edges of a non-selectable map, so the hover handler never fired.
- The agent DaemonSet declared its
grpcport as TCP. On an agent that port is the UDP echo listener; it is declared UDP now, which is what makes the scheduler'shostPortclash check guard the right protocol underagent.hostNetwork.
Upgrade notes¶
The default install pins controller, agent and console images to 2.4.0 and needs nothing special. The table is for fleets that run mixed versions for a while, most often deb/rpm hosts upgraded after the cluster.
| Old side + new side | Per-agent ports | Prometheus SD | Topology event labels (Time Machine) | Console badges and hints | agent.hostNetwork |
|---|---|---|---|---|---|
| 2.3.x agent + 2.4.0 controller | reports none; peers dial it on their own ports | listed on the controller's metricsPort, logged once as assumed |
carried as the agent sent them (a 2.3.x external agent already sets the label) | badge shows; planes read "unknown" | the agent ignores KCONMON_NG_HOST_NETWORK, so no host-network label |
| 2.4.0 agent + 2.3.x controller | dropped at registration; everyone dials their own ports | no route: a ScrapeConfig on it fails every refresh | not recorded | depends on the console | the node IP passes registration; the label is stored verbatim |
| 2.3.x console + 2.4.0 controller | n/a | n/a | the new field is ignored | none: external agents look like ordinary nodes | n/a |
| 2.4.0 console + 2.3.x controller | n/a | n/a | events carry no labels; only the console's own baselines badge history | live badges and the unscraped hint work (the controller already publishes labels) | n/a |
| 2.3.x deb/rpm agent in a 2.4.0 fleet | reports none, dials its own ports: one-way red toward peers on other ports | scraped on the controller's metricsPort |
badged | badge shows; planes "unknown" | n/a |
| mixed fleet with different ports | old agents dial their own ports: one-way red, on-demand tasks included | n/a | n/a | n/a | n/a |
If you pin the controller image behind the chart, note that
controller.prometheusSD.enabled reaches the shared ConfigMap only when
false: a pre-2.4.0 controller image rejects the unknown key and crashloops,
so leave it at the default until the image is current. Flipping
agent.hostNetwork or agent.dnsPolicy rolls the DaemonSet; during the
rollout node IPs and pod IPs coexist and a little transient PairWentSilent
noise is expected.
kconmon-ng v2.3.1¶
Fixed¶
ZoneChecksFailingandZoneLossHighfailed every evaluation with "vector cannot contain metrics with the same labelset" and raisedPrometheusRuleFailureson the cluster:rate()over a__name__regex union drops the metric name and collapses the per-protocol families into duplicate labelsets. The expressions now build the union withlabel_replace(...) or label_replace(...), which keeps the branches distinct and still tolerates a disabled checker's absent family.- CI now evaluation-tests every alert rule with
promtool test rulesagainst synthetic series for all metric families, including a positive check that each zone alert fires on staged bad data. Rendering and syntax checks never execute the query engine, which is exactly where this defect lived.
kconmon-ng v2.3.0¶
The sparse mesh changes WHAT "no data for a pair" means: under
topology.mode: sparsemost directed pairs are deliberately never probed. Everything in this release that reads per-pair series learns to tell "not planned" from "went dark" through one new metric,kconmon_ng_probe_intended— and that metric comes from the AGENT: images below appVersion 2.3.0 do not export it. The chart's rules degrade honestly on an older fleet (see PairWentSilent below), but do not fliptopology.mode: sparseuntil controller AND agents run a 2.3.0 image — the controller config key is emitted only when sparse precisely because an older controller image rejects it and crashloops. The appVersion pin is aligned when the app release ships.
Added¶
topology.*— the sparse probe mesh, by values.topology.mode: sparsetrims the full N×(N−1) probe matrix to a ring over sorted node names (sparse.ringDegreesuccessors each, the connectivity guarantee) plus HRW-chosen cross-zone chords (sparse.zoneChordsper directed zone pair, which keep the zone metric family fully populated), so probed pairs — and every per-pair series they export — scale ~linearly with node count instead of quadratically.sparse.autoThresholdis the floor: fleets smaller than it get the full mesh regardless of mode, because sparse only pays for itself at scale. Default ismode: full, byte-identical rendering to 2.2.0.kconmon_ng_probe_intended— the plan, scrapable. A gauge, value 1 for every directed pair the topology plan assigns ({source_node, destination_node}, exported by the source agent), preset from the peer list at registration and pruned on every plan change — stale pairs are deleted, not left at 1. In full-mesh mode it simply marks every peer, so dashboards and rules can join on it without caring which mode the fleet runs. It is the one honest way to distinguish "this pair is not supposed to report" from "this pair went dark", which is why it ships in the same release as sparse mode and not one later.investigateUrlon the two zone alerts.ZoneChecksFailingandZoneLossHighnow annotate a console deep link,/investigate?kind=zone-pair&scope=<source>-><destination>, straight into the Investigate page scoped to the firing zone pair. The link is console-RELATIVE on purpose — the chart cannot know the console's external URL (ingress is optional), so notification templates prepend their own origin; the console normalises the typeable->into its canonical pair arrow.
Changed¶
PairWentSilentjoins on the plan. The rule now fires only for pairs present in the source agent'skconmon_ng_probe_intendedseries — the hard rule of the sparse design, shipped in the same release: without the join, every pair the plan trims would read as "went silent" for the hour its results take to age out of the lookback window. The fallback is per SOURCE, not global: asource_nodeexporting noprobe_intendedat all keeps the old two-window behaviour, so a pre-2.3.0 agent image alerts exactly as before, a mixed fleet mid-rollout gets each behaviour where it applies — and an agent that dies outright takes itsprobe_intendedseries with it, which lands its pairs in the same fallback and preserves the alert's original purpose: catching an agent that stopped running or stopped being scraped.
This release also carries everything prepared for the never-published 2.2.0 tag (its pipeline caught two release-tooling defects before anything went out); those changes follow below, under their original heading kept for upgrade notes.
Carried over from the unreleased 2.2.0¶
Everything in this release reads the new zone-level metric family (
kconmon_ng_zone_*), and that family comes from the AGENT, not the chart: agents below appVersion 2.2.0 do not export it (this chart pins 2.2.0, so a default install is fine — the warning is for fleets running an older agent image behind a newer chart). Until the fleet runs an agent image that does, the two zone alerts are silently inert (their expressions match no series), the Zone Heatmap dashboard renders empty, andagent.metrics.detail: zone-onlywould drop the per-pair series with nothing replacing them — Prometheus goes dark on the mesh while the console keeps working. Upgrade the agent image first, flip the valve second. The appVersion pin is aligned when the app release ships.
Added¶
ZoneChecksFailingandZoneLossHigh. Two alerts on the zone plane, with the same per-rule knobs as the rest (prometheusRule.{zoneChecksFailing,zoneLossHigh}.{enabled,threshold,for,severity}).ZoneChecksFailingis the failure ratio of all TCP, UDP and ICMP probes between a zone pair, in one expression — the__name__union keeps a disabled checker from blanking the ratio.ZoneLossHighcomputes loss as(sent − received) / sentfrom the zone packet counters; averaging the per-pair loss-ratio gauges into a zone would weight an idle pair the same as a busy one, so the chart never does. Its default threshold is0.1, lower than the per-pairUDPLossHighat0.5, because the zone aggregate dilutes any single link by the pair count: sustained loss at that level means the fabric, not one node. Both survive everyagent.metrics.detailmode — that is the point of alerting on the zone family.agent.metrics.detail— the cardinality valve. A scrape-time knob rendered asmetricRelabelingson the agent ServiceMonitor:full(default, everything, ~70 series per directed pair),counters-only(drops the four per-pair histograms, ~10/pair — every pair alert keeps firing),zone-only(drops every series naming adestination_node, ~0/pair; the zone family at ~74×Z² series and the linear DNS/HTTP/external families remain). At 100 nodes that is ~0.7M → ~0.1M → practically N-independent, by configuration alone. Setting it withoutserviceMonitor.enabledis refused at render time rather than silently dropping nothing; plain-Prometheus equivalents are indocs/metrics.md.controller.externalGateway— the external agent gateway, exposed by the chart. The controller's second gRPC listener (same services, but TLS with a bootstrap token, for agents OUTSIDE the cluster) gets a values block and three templates.templates/controller/service-external.yamlis a NodePort/LoadBalancer Service carrying the gateway port ALONE — the plaintext in-cluster gRPC port authenticates by network position and never appears on it, because a LoadBalancer in front of it would hand the whole mesh to anything that can reach the address. The deployment mounts two referenced Secrets read-only:tls.secretName(akubernetes.io/tlsserving pair;tls.clientCaKeynames the CA bundle key in the same Secret and switches on client-cert identity pinning — empty is token-only mode, where any token holder can impersonate any agent, and NOTES.txt says so at install) andbootstrapToken.{secretName,key}. WithnetworkPolicy.enabled, ingress on the gateway port is opened fromnetworkPolicy.externalAgentCidrstoward the controller pods alone, and an empty list is refused at render rather than shipping a gateway no packet can reach; missing Secret names and a port colliding withconfig.{httpPort,grpcPort,metricsPort}are refused the same way. Two operational notes. Rotation: the gateway reads the certificate and token ONCE at startup and the chart cannot checksum content it only references, so rotating either Secret in place needskubectl rollout restart deploy/<release>-controller. Version skew: theexternalGatewayconfig key is emitted only when enabled, because a controller image at appVersion 2.0.3 rejects the unknown key and crashloops — upgrade the image before flipping the switch, same rule as the zone family above.
Changed¶
- The Zone Heatmap dashboard reads the zone family. Every panel that
aggregated per-pair series into zones at query time now reads the
pre-aggregated
kconmon_ng_zone_*metrics, so the dashboard keeps working in everyagent.metrics.detailmode and its queries stop scaling with the pair count. Loss panels are packet-weighted from the sent/received counters instead of averaging the per-pair ratio gauges. The one exception is the "MTR traces triggered" panel: MTR has no zone-level family, its counter is per-pair, and inzone-onlymode that panel reads zero — its description now says so.
Performance and self-observability¶
- Peer probing fans out with a bounded pool (32 in flight per round): a dead peer costs one timeout, not one timeout per peer in sequence, so probe cadence holds through partitions.
- Reactive MTR traces are bounded by a global semaphore (4 in flight) on top of the existing per-pair cooldown; a mass partition trickles traces out instead of forking one per broken pair.
- The agent exports self-metrics under
kconmon_ng_agent_*: probe cycle duration and overruns per checker, controller reconnects, peer-list age, reactive-MTR in-flight and coalesced counters. All are fleet-size independent. - The controller coalesces peer-list broadcasts (trailing edge, 200 ms): a rollout's burst of registrations produces one broadcast, not one per change; the peer message is built once per broadcast and carries only the fields agents read.
kconmon-ng v2.1.0¶
Added¶
agent.updateStrategyis a value. The DaemonSet hardcodedmaxUnavailable: 1— the right default and the wrong ceiling: one node at a time turns a version rollout across a few hundred nodes into hours of half-upgraded fleet. The block passes through verbatim, somaxUnavailable: 10%— orOnDelete— is now a values change instead of a fork of the template.priorityClassNameon every workload.agent.priorityClassName,controller.priorityClassNameandconsole.priorityClassName; empty by default, and empty renders nothing. The agent is the one worth setting: under node pressure the kubelet evicts lowest priority first, and the first pod gone should not be the one reporting on the node.
Fixed¶
helm testpasses restricted-PSS admission. The connection-test Pod ran curl with no securityContext at all, so a namespace enforcing the restricted Pod Security Standard rejected it at admission — the test failed before it made the one request it exists to make. The container now declares the four fields the profile checks:runAsNonRoot, aRuntimeDefaultseccomp profile, no privilege escalation, all capabilities dropped. The image already runs as uid 100.
kconmon-ng v2.0.3¶
Fixed¶
- A config change restarts the pods that read it. The agent and the
controller share one ConfigMap and read it once at startup, and a mounted
ConfigMap changes under a running process without telling it — so
controller.events.enabled: trueapplied to a live release updated the object and left the controller on the old file. It went on advertising no capabilities, the Console's realtime ingester retried against a stream that was configured but never started, and nothing anywhere reported an error: the Live page was simply empty. Both workloads now carrychecksum/config, so a values change rolls them; a change the ConfigMap does not carry still does not.
kconmon-ng v2.0.2¶
Fixed¶
- The matrix no longer opens at half size. The grid measured the height its own content had produced and fed that back into the fit, so a fresh render saw the container's 256px minimum, decided the grid did not fit, shrank to 50%, and the smaller grid then held the box at 256px — a loop with no way out. It measures the space available instead, and a seven-node fleet opens at 100%.
- Zooming in gives the node names back. The shared prefix every node name
begins with is dropped to buy column width, which is right while the column is
narrower than the names and wrong the moment it is not: at 125% a label column
holds
adm-kuber-01with room over and still read…01. The elision is now decided per axis at the current scale, and the note above the grid appears only while an axis is actually eliding.
Added¶
- A favicon. The console had none, so every tab showed the browser's blank square; it now wears the mark it wears in its own sidebar.
kconmon-ng v2.0.1¶
Fixes a hole in 2.0.0:
auth.mode=oidcandauth.mode=headershipped with no way to grant anybody a role. Both modes worked, and neither was usable.
Fixed¶
- An OIDC or header install can grant roles at deploy time. Role bindings
live in the database and are created through an API that already requires
rbac:manage, so a fresh install had nobody able to make the first binding; the only alternative wasauth.defaultRole, which is one role for every authenticated subject. The way out that 2.0.0 left was to bring the console up in local mode, log in, create a binding by hand and only then switch — a workaround, published as if it were a procedure.
console.auth.groupRoles maps a group the identity provider asserts onto a
role this console grants, in the values file:
Roles resolve as the union of that map and any binding made through the API, so a grant by hand still adds to what the provider's groups carry. A group absent from the map grants nothing. What the map grants cannot be revoked through the API — that is what makes it declarative. - A role store outage no longer costs an operator their access. The store's half still fails closed, because an unreadable database is no evidence a subject holds anything; a grant that came from the claim and the config was never in doubt, and an outage is when the console is most needed.
kconmon-ng v2.0.0¶
A chart that installs monitoring and nothing else, a console that survives more than one replica, and one that tells the truth about time. The chart no longer ships a database or a cache — point it at the ones you already run. The Time Machine moved out of the top bar and into each page's own time controls, and the charts pin their axis to the window you asked for rather than to the data that happened to arrive. MTR gained a Runner, path history that reads as a timeline, and external targets.
Breaking¶
- The chart no longer installs PostgreSQL or Valkey.
database.mode,database.cnpg.*and the bundled subcharts are gone: setdatabase.existingSecretto a Secret holding apostgres://DSN andredis.existingSecretto one holding aredis://DSN, and any managed instance works — RDS, a StatefulSet, a CloudNativePG cluster you run yourself. Every removed key fails the render with a message naming its replacement (templates/_migrations.tpl), so no old value is silently honoured. console.database.*moved to the top-leveldatabase.*, andconsole.*keys that described the bundled datastores went with it.
Added¶
- Chart 2.0.0, templates split per component — agent, controller, console,
shared and observability each own their directory, with a NetworkPolicy set
covering every component, fail-closed on external egress. Render-time guards
refuse a port collision, an OIDC
redirectURLthe console would not start on, and more than one console replica without a shared cache. - The console scales past one replica — sessions, the fixed-window rate-limit counters and the realtime fan-out live in the Redis-compatible server, and the controller elects a leader so exactly one replica drives the reconcilers.
- MTR Runner and path history — start a trace from the Explorer itself, with a settable cadence and duration; every distinct route the fleet has taken is kept, diffed and drawn on a timeline of when it changed.
- External targets — probe a destination that is not a fleet peer, gated by
config.checkers.external.allowedCidrsand the cluster's own egress policy. The console refuses at create time a target no agent could ever reach. - Time in the Console's result table — every figure says when it was read.
Changed¶
- The Time Machine lives with the page's time filters, not in a strip across
the top of every route. It is offered only on the pages that resolve their
reads through
?at=, and the engaged banner stays global because writes are disabled console-wide. - Explore's axis is the window you picked — a 24h view draws 24 hours even when Prometheus holds less, instead of quietly redrawing three.
- MTR Explorer is sorted by name, both destinations and their sources, with
numbers read as numbers (
m9beforem10). - OIDC identity is the
subclaim, namespaced asoidc:<sub>— the only claim OIDC Core §5.7 allows as an identifier.auth.oidc.usernameClaimnow decides the display name alone, so renaming a person no longer moves their roles (Grafana's CVE-2023-3128 is what the old shape risked). Group membership is re-read on every token refresh. Bindings made against a username stop granting; the console names them at boot so they can be remapped. - The configuration bundle carries access control — custom roles and the
grant list, but only for a caller who holds
rbac:manage. Roles import; bindings never do, because a grant names a person in the source console's own identity namespace. - Only the chart under the cursor shows a tooltip. Its neighbours keep the shared crosshair and mark their own samples with a dot, instead of each covering its own curves with a box of numbers.
Fixed¶
- WebSocket topics are authorized per topic:
events:readno longer carries the topology and matrix snapshots thattopology:readandmatrix:readgate, and a permission taken away reaches a socket that is already open — the topics it may no longer have are dropped, the rest of the connection is left alone. - The audit row describes the mutation that happened. A body could name one
thing for the handler and another for the audit log by spelling a key in a
different case, and a value carrying a NUL made the whole row unwritable — in
both directions the caller chose whether their own privileged action was
recorded. The extraction now matches keys the way
encoding/jsonmatches struct fields, and is bounded before it is decoded, so a wide body on the public login route can no longer take the replica past its memory limit. - A broken alert rule no longer freezes the whole bundle. Editing a deployed rule into PromQL the apiserver rejects used to stop every other rule from being applied, while the API answered 2xx and Prometheus kept evaluating the stale set. The quarantine now keys on rendered content rather than rule ids, offers each suspect to the cluster on its own, and removes the object only when every rule was offered and every one refused.
auth.mode=anonymousis not exempt from CSRF. Any page an operator's browser visited could POST into a console kept off the internet; a cross-origin write is now refused, while a script that sends noOriginis unaffected.- The node-local HTTP checker verifies certificates. An expired certificate,
one issued for another hostname or an interceptor's CA all used to pass, so an
https check could not fail on the condition it was added to notice; opt out per
target with
insecureSkipVerify. - External metrics separate the checks on one target — the series carry
check_type, so an icmp and a tcp check on the same target no longer average each other's failures away under theExternalChecksFailingrule. - A check no agent could run is refused when it is written, instead of being dropped by every agent with nothing but a log line while the console listed it as enabled.
- The MTR destination listing is complete. It is paged behind a keyset cursor rather than capped, so no pair is missing from the Explorer and no per-destination total is short.
- A subscriber that stops reading its peer-update stream is torn down rather than holding a controller goroutine and its connection slot until TCP notices.
- Shutdown finishes in-flight runs before tearing down the pipeline they publish onto, so a rolling update no longer logs dropped frames that were delivered.
- Every request body is capped, so one oversized POST can no longer take a console replica past its memory limit.
- The OIDC callback binds its
stateto the browser that started the flow. - A role-store failure now refuses rather than granting the default role.
- External TCP and UDP checks probe what was asked for instead of speaking the agent's own protocol to something that is not an agent.
- A user binding can no longer be resolved by a subject of another kind: role resolution matches the caller's kind as well as their id.
- Revoking a role binding is auditable — the audit row names the role and the subject, read before the row is destroyed rather than after.
- Path history says when it has reached the end instead of leaving a "Load older" button that can never be pressed, and counts the routes it is showing against the traces folded into them.
- A probe tick on a diagnostics run leads with that probe — its sequence, its clock, its latency or its error — so two ticks on an unchanged route are no longer indistinguishable.
kconmon-ng v1.9.0¶
Console release, and the last planned milestone. Alert rules you build in the Console and Prometheus evaluates; alert webhooks; configuration export/import; a command palette; a Settings page. Plus one long-carried fix: the controller finally attributes topology changes, so Time Machine reconstructs a real cluster. Everything new is off by default or read-only-additive: a 1.8.0 install that upgrades and changes nothing renders the same manifests.
Added¶
- Alert rule management —
/alertingbuilds a Prometheus alert rule from six typed templates (pair loss, zone latency, DNS failures, HTTP TTFB, agent missing, external target down) or raw PromQL — seven kinds in all — and the Console reconciles every enabled rule into onePrometheusRuleobject by server-side apply. The Console manages; Prometheus evaluates. Nothing here decides that an alert fired. - Validation by running it, not by parsing it — there is deliberately no
prometheus/prometheusparser dependency. Every template has a byte-exact render golden, andPOST /api/v1/alert-rules/previewruns the expression as an instant query against your actual Prometheus and reports how many series it matches. The render and the query fail independently: a render failure is a422, a query failure is a200carrying the expression and the error. - Drift is recorded, then fixed — a reconcile always re-asserts the
Console's bytes. A rule showing
driftalso carries a freshlastSyncedAt, and both are true: the divergence was observed and corrected in the same pass. Failures never crash the loop; they land per rule assync_status=errorwith a closed cause class (crd-missing,forbidden,other). - Foreign rules and explicit adoption —
PrometheusRuleobjects the Console did not write are listed read-only, andPOST /api/v1/alert-rules/importcopies one into builder rows. The foreign object is never mutated, which means the same alerts then exist twice until you remove one copy. The import report says so, and names every skipped rule with its reason. GET /api/v1/alerts— the firing set, projected onto this API's vocabulary, with?managedOnly=. With no Prometheus configured it answers200andpromConfigured: falserather than503: "nothing is firing" and "nobody is watching" are different sentences.- Alert webhooks —
alert.firedandalert.resolved, their own payload family, dispatched from a poller that diffs Prometheus' alert state onconsole.webhooks.alertPollInterval(30s). It baselines on boot rather than paging the fleet about what was already broken, freezes on a failed or undecodable poll rather than "resolving" everything, and ignores rules the Console does not manage. The M6 incident payload bytes are unchanged. - Configuration export/import —
GET /api/v1/exportandPOST /api/v1/import, versioned bundle v1, admin-only undersettings:write, dry-run first. Webhook endpoints export withhasSecretonly and therefore cannot be created by import — a sealed secret never leaves this API. - Settings page — webhook CRUD, export/import with a per-collection dry-run
report, and read-only deployment info that renders only what
GET /api/v1/configactually serves. - Command palette (
⌘K/Ctrl-K) — hand-rolled, zero dependencies, over navigation (generated from the sidebar so it cannot drift), five actions and the Time Machine pair. It does not jump to an arbitrary node, target or pair: that needs a live object search, not a static registry. - Overview — the "Firing alerts" placeholder carried since M1 is now the real panel, severity-sorted with oldest-first ties, and the Investigate timeline gained an alert row.
- Permissions —
alerts:read(all built-in roles) andalerts:manage(operator, admin, andalert-editor— the builtin has waited for exactly this permission since M3, and a role by that name that cannot edit an alert rule breaks its promise on first click).AllPermissionsis now 25.
Fixed¶
- Topology events are attributed — the controller now emits one
topology_changedevent per affected agent, carryingnodeName,agentIdandzone, from all four sites (register, zone update, deregister, stale eviction). Time Machine's topology fold reconstructs a real node set with real zone lanes instead of an honest empty one. Events written by an earlier controller are counted as unfoldable and age out with retention; the page reports both numbers rather than rendering an empty cluster. - WebSocket topics are authorized individually —
/wsadmitsevents:readorruns:read, and a subject admitted onruns:readalone getsrun:{id}topics and an error frame for the fleet-wide ones, on a socket that stays open. A custom role can finally watch the run it started. Carried from M3. - A
nullconsole.database.cnpgoverride no longer crashes rendering — nor does a null sub-block. Real nil-pointer class, found by the schema work. - One redundant token listing on the ownership-resolution path is now a targeted lookup.
Chart¶
- 1.8.0 → 1.9.0. New value blocks:
console.alerting.*(enabled/namespace/syncInterval/bundleName, rendered into the console ConfigMap only when enabled) andconsole.webhooks.alertPollInterval(rendered only when a key Secret is named and alerting is on — the key existed in the binary since M7 but was unreachable from Helm). - Enabling
console.alertingrenders a namespacedRoleandRoleBindingovermonitoring.coreos.com/prometheusrules(get,list,watch,create,update,patch— neverdelete), bound to the console-only ServiceAccount. Not a ClusterRole: the Console writes one object into one namespace, and pointingnamespaceelsewhere fails with aforbiddenrather than widening anything. - The console ServiceAccount,
serviceAccountName,POD_NAMESPACEand the apiserver egress rule are now shared betweenconsole.kubernetesContextandconsole.alertingthrough one helper, so either flag renders them. values.schema.jsoncloses 44 chart-owned levels withadditionalProperties: false, and gained thenameOverride,fullnameOverride,agent.nodeSelectorandagent.affinitykeys the templates always used and the schema never declared. Pod/containersecurityContextstay deliberately open — Kubernetes grows union members every release, and closing them would turn a cluster upgrade into an install failure.- Three new ci profiles:
console-alerting-values.yaml(the fullest console),console-auth-local-values.yamlandconsole-auth-header-values.yaml. Default render is key-identical to 1.8.0.
Upgrade notes¶
- Nothing to do.
alert_rulesis created by migration00007on first start withconsole.database.mode=cnpg|external; with the database disabled the new routes answer 503 exactly as M3–M6's do. - Turning on
console.alerting.enabledneeds three things the chart cannot check: the Prometheus Operator'sPrometheusRuleCRD, a database, and a Prometheus whoseruleSelector/ruleNamespaceSelectoractually selects the object the Console writes. A rule that syncs cleanly and never fires is almost always the third one. - Changing
console.alerting.bundleNameon a live install orphans the previous object. The reconciler owns what it is pointed at and deletes nothing. - There is no leader election on the alert-webhook watcher: N console
replicas deliver N copies of every edge. The payload carries a stable
(event, ruleId, labels, firedAt)tuple so a receiver can dedupe. alert-editorgainedalerts:manage. If you granted that builtin to somebody expecting it to stay inert, it is now able to create, edit and delete alert rules.- If you set values the schema never declared,
helm upgrademay now reject them. That is the typo protection working — check the key againstvalues.yaml.
kconmon-ng v1.8.0¶
Console release. Investigation Mode with an honest, documented correlation panel; saveable incidents with shareable permalinks; maintenance windows; outbound webhooks; and optional Kubernetes event capture (M6). Everything new is off by default or read-only-additive: a 1.7.0 install that upgrades and changes nothing renders the same manifests.
Added¶
- Investigation Mode —
/investigateassembles a merged timeline, synced signal panels (loss/RTT with a matrix delta chip and an MTR path diff) and an actions rail for a scope and a time range. Entry is the URL and only the URL —?kind=&scope=&from=&to=— from any node/pair/target card, any matrix cell, or the page's own form. Every source is permission-gated with zero requests when denied, and each absent or bounded one leaves a muted line rather than blanking the page. - Correlation v1, documented rather than magic — edge-triggered threshold crossings (loss > 1%, RTT > 2× the range median), an onset, a 300-second candidate window and a linear proximity decay against published class weights. The panel links the scoring source itself, so the operator reads exactly the constants the code executes. No ML, and nothing you cannot reproduce by hand.
- Incidents — save an investigation (
/api/v1/incidents), pin findings from six source kinds, write notes, resolve and reopen. The permalink/investigate?incident={id}rehydrates scope and range from the row, so the link cannot drift from the incident it names. Open incidents appear on Overview and beside the charts on every object card.PATCHis deliberate and is this API's only one: an incident evolves under collaboration, and a full replace would let one writer discard another's notes. - Maintenance windows —
/api/v1/maintenance, drawn asmarkAreaon Explore, the Pair card and the Target card and as timeline rows. M6 renders declared windows; it does not suppress anything, because nothing evaluates alerts until M7. - Outbound webhooks —
/api/v1/webhooks(admin-onlywebhooks:manage), firing on incident lifecycle. Deliveries are signedX-Kconmon-Signature: sha256=<hmac>over the raw body, retried 3 times (0s / 30s / 5m, ±20% jitter, 10s per attempt), with the outcome kept on the endpoint row. Each endpoint's signing secret is write-only over the API and sealed at rest with AES-256-GCM underconsole.webhooks.encryptionKeySecret.POST /{id}/testsends one signed ping. Without a key, create and test answer 503 and everything else keeps working. - Kubernetes event capture —
console.kubernetesContext.*(off by default) list+watches core/v1 Events intok8s_events, filtered to nodes in the fleet topology and pods in one namespace, and read back throughGET /api/v1/k8s-eventsunderevents:read. Zero new Go dependencies — it reuses the client-go the controller already pulls in. Without a topology it fails closed and drops node events rather than storing an unfiltered firehose. - Overview — the "Recent events" placeholder that had been carried since M2 is now the real panel, and an "Open incidents" card sits beside it. "Firing alerts" stays an honest placeholder until M7.
- Permissions —
incidents:readandmaintenance:read(all built-in roles),incidents:writeandmaintenance:write(operator/admin), andwebhooks:manage(admin only — an endpoint carries a signing secret). - Metrics —
kconmon_ng_console_k8s_events_total{result}andkconmon_ng_console_webhook_deliveries_total{result}, plus three new table values onkconmon_ng_console_retention_deleted_total.
Chart¶
- 1.7.0 → 1.8.0. New value blocks:
console.kubernetesContext.*(rendered into the console ConfigMap only when enabled) andconsole.webhooks.encryptionKeySecret.{name,key}(referenced by name, mounted as a file beside the DSN — there is deliberately no inline-key value). EnablingkubernetesContextalso renders a console-only ServiceAccount, its own ClusterRole (events: list, watch) and binding,serviceAccountNameon the console Deployment, aPOD_NAMESPACEdownward-API variable and an apiserver egress rule (console.networkPolicy.kubeAPIEgress). The agent/controller ServiceAccount and its grant are untouched. A ninth ci profile (console-investigation-values.yaml). Default render is key-identical to 1.7.0.
Upgrade notes¶
- Nothing to do. All four M6 tables are created by migration
00006on first start withconsole.database.mode=cnpg|external; with the database disabled the new routes answer 503 exactly as M3–M5's do. - Turning on
console.kubernetesContext.enabledis a deliberate act with an RBAC consequence — read SECURITY.md §10.3 first. - Webhook endpoints are API-only in M6: there is no Settings page yet.
- Rotating
console.webhooks.encryptionKeySecretdoes not re-seal existing rows. An admin mustPUTeach endpoint with a fresh secret afterwards.
kconmon-ng v1.7.0¶
Console release. MTR Explorer with durable path history and optional hop enrichment, the Time Machine global time context, and chart annotations (M5). Everything is off by default or read-only-additive; a 1.6.0 install that upgrades and changes nothing renders the same manifests.
Added¶
- MTR path history — every MTR result the Console records is projected
into
mtr_path_snapshots: hop lists content-hashed per (source, destination) route, deduplicated at ingest, with first/last-seen and trace counts.kconmon_ng_console_mtr_snapshots_total{result="new-path"}is the "route changed" alerting primitive. - MTR Explorer —
/mtrthree panes (destinations → path history → trace detail), client-side path diff between any two snapshots, a "path changes" timeline overlaid with the pair's loss series, per-hop RTT trends, and a Runner tab (gated onruns:create). - Hop enrichment —
console.mtr.enrichment.*(off by default): reverse DNS and/or MaxMind GeoLite2 ASN/City mmdb files mounted read-only at/geoipfrom an operator-supplied volume; resolved server-side into a TTL cache (mtr_hop_enrichment), air-gap friendly, per-source degradation. - Time Machine — a top-bar control and shareable
?at=URL state: topology reconstructed fromtopology_eventsup tot(GET /api/v1/topology?at=), PromQL surfaces evaluated att, the Live feed becomes scrollback, and every mutating control is disabled behind the banner while engaged. Note: today's controller does not yet attributetopology_changedevents to nodes, so historical topology reconstruction reports honestly-empty results until it does (named deferral). - Annotations — notes pinned to instants or time ranges
(
/api/v1/annotations), rendered as chart markers on Explore and the object cards and inline in the Live scrollback.annotations:writestops at operator/admin; reads are telemetry (every role). - Explore A/B — a Compare panel: a second curated metric on the same axes, or the same metric time-shifted (1h/24h/7d) against itself.
- Permissions —
mtr:read,annotations:read(all built-in roles) andannotations:write(operator/admin).
Chart¶
- 1.6.0 → 1.7.0. New value blocks:
console.mtr.enrichment.*(rendered into the console ConfigMap only when enabled; mmdb volume via an opaquegeoip.volumeVolumeSource passthrough with a render-time guard). An eighth ci profile (console-mtr-values.yaml). Default render is key-identical to 1.6.0.
kconmon-ng v1.6.0¶
Console release. External probe targets, saved check definitions and schedules, continuous external checks with agent-side CIDR enforcement, and diagnostics v2 (M4). Off by default throughout: without
console.scheduler.enabled,config.checkers.external.enabledor a database, a 1.5.0 install behaves identically.
Added¶
- Targets, checks, schedules — CRUD
/api/v1/{targets,checks,schedules}(PostgreSQL-backed, 503 without a database), with a cardinality projection guard: an enabled definition may project at most 400 series, enforced server-side and mirrored live in the UI. - Console scheduler —
console.scheduler.{enabled,tickInterval}(off by default):once/intervalschedules fire diagnostics runs from exactly one replica under a PostgreSQL advisory lock; a stuck-run reaper finishes abandoned runs ascancelled;POST /api/v1/runs/{id}/cancelcancels an in-flight run. - Continuous external checks — check definitions of kind
continuousare reconciled to the controller (PUT /api/v1/external-checks,WatchExternalChecksstream) and probed by agents on their own cadence. The agent is authoritative:config.checkers.external.enabledplus a mandatoryallowedCidrsallowlist (deny-wins, resolved-IP matching, re-checked on every probe); an agent without the opt-in refuses external work and the controller answers 501 for it. - External metrics — the
kconmon_ng_external_*family ({source_node,source_zone,target,target_kind}labels; the target NAME, never an address) plusExternalChecksFailingin the default rules. - Diagnostics v2 —
POST /api/v1/runsacceptsdestinationKind=node|target|adhoc; the diagnostics form grows a destination selector and Save-as-definition. - Rate limits —
console.rateLimit.{runsPerMinute,loginPerMinute}(fixed-window on the shared KV; fail-open on a Valkey outage; the login limit runs before argon2id). - OpenAPI — a committed spec (
docs/console-api.yaml), generated TS types, and a router-walking test that fails on any drift in either direction. - UI — the Targets & Schedules page and the Target card
(
/targets/{id}).
Chart¶
- 1.5.0 → 1.6.0. New value blocks:
config.checkers.external.*(agent, rendered only when enabled),console.scheduler.*,console.rateLimit.*; a seventh ci profile; the external-mode Valkey egress NetworkPolicy now derives its port from the configured address.
kconmon-ng v1.5.0¶
Console release. Adds durable persistence, authentication/RBAC, and an on-demand diagnostics runner (M3) on top of the M1/M2 Console. Off by default (
console.enabled: false,console.database.mode: disabled,console.auth.mode: anonymous); existing installs — including existing Console installs onanonymousauth — are functionally unaffected until you turn these on.
Added¶
- PostgreSQL persistence —
console.database.mode=cnpg|external|disabled(ADR-001): CloudNativePG-provisioned or externally-supplied PostgreSQL for event history, RBAC, audit, and diagnostics run history.GET /api/v1/eventsnow serves durable scrollback for the Live page, backed bytopology_events(all five WebSocket event types, not only topology ones). - Authentication and RBAC —
console.auth.mode=anonymous|local|header|oidc: local users (PostgreSQL, argon2id), header-based trusted-proxy auth (explicit CIDR opt-in), and OIDC (code flow + PKCE). Built-in roles (viewer/operator/alert-editor/admin) are compiled-in and work withdatabase.mode=disabled; a custom-role admin API layers on top when a database is configured.__Host-session cookies, CSRF double-submit for cookie-authenticated mutations, and an async best-effort audit log (GET /api/v1/audit). - API tokens (PATs) —
/api/v1/tokens: SHA-256-hashed bearer tokens that work in every auth mode. A PAT is not individually scoped in M3: its effective permissions are exactlyauth.defaultRole, deployment-wide — a token subject resolves no role bindings of its own, and token-kind bindings are rejected by the RBAC API by design (they would silently grant nothing). Where a database is configured, disabling a local user's account revokes the tokens they own on that user's next request; subjects created by header or OIDC mode have nousersrow, so their disable state lives upstream at the proxy/IdP and onlyDELETE /api/v1/tokens/{id}revokes their tokens. - Diagnostics runner —
POST/GET /api/v1/runs,GET /api/v1/runs/{id}: bounded on-demand check fan-out (up to 400 pairs) with persisted run history and shareable permalinks. Live per-pair progress streams over a new ephemeralrun:{id}WebSocket topic, opened per run, with an automatic REST-polling fallback on any console replica other than the one executing the run. - Object cards v1 — Node and Pair cards with a shared "Recent changes" event rail, linked from Topology and the Matrix heatmap.
- NetworkPolicy — opens console↔database (CNPG pod-selector rule, or a
namespace-wide default for
mode=external) and console→OIDC IdP on both layers, alongside the existing console→controller/Valkey rules.
Fixed¶
- Chart mounts every console secret (database DSN, local-admin bootstrap
password, OIDC client secret) under one sibling directory,
/etc/kconmon-ng-console-secrets/, group-readable (0440) withconsole.podSecurityContext.fsGroupmatching the distroless nonroot gid — the milestone's originally planned nested path and owner-only mode were both unworkable (a read-only ConfigMap volume cannot host a mountpoint inside itself, and0400is unreadable by the nonroot process without matching UID ownership).
Upgrade Notes¶
- This release is safe to roll out with every new feature left at its
default (
console.database.mode: disabled,console.auth.mode: anonymous): the M1/M2 surface behaves identically. - Every existing Console install rolls once on upgrade, even one that
changes nothing in
values.yaml.auth.defaultRoleand theauth.session.*block are now rendered into the console ConfigMap unconditionally (previously absent keys), so the ConfigMap's content — and therefore its checksum annotation on the Deployment's pod template — changes for every install, triggering one rollout. There is no data or config migration behind it; it is a one-time restart. - To turn on persistence and auth, set
console.database.modetocnpg(this chart does not install the CloudNativePG operator or its CRDs — install those first, orhelm installfails with a clear error) orexternal(supplyconsole.database.existingSecret), then setconsole.auth.modeand its mode-specific block. See the chart README and the commentedvalues.yamlfor the full validation matrix and secret-mount layout.
Install¶
helm upgrade --install kconmon-ng oci://ghcr.io/esdmitrii/charts/kconmon-ng \
--version 1.5.0 \
--namespace kconmon-ng \
--create-namespace
Images¶
ghcr.io/esdmitrii/kconmon-ng-agent:1.5.0
ghcr.io/esdmitrii/kconmon-ng-controller:1.5.0
ghcr.io/esdmitrii/kconmon-ng-console:1.5.0
kconmon-ng v1.4.0¶
Console release. Adds an optional read-only web Console (M1) with a realtime event pipeline (M2). The Console is off by default (
console.enabled: false); existing agent/controller installs are unaffected until it is turned on.
Added¶
- Console web UI — new
kconmon-ng-consolebinary and image with an embedded SPA: Overview, Matrix, Topology, Explore, PromQL console and a Live event feed. Read-only in this release (anonymous viewer role, banner shown in the UI). - Realtime pipeline — the controller exposes a leader-gated
EventStream.WatchEventsgRPC stream (controller.events.enabled); the console ingests it and fans events out to browsers over WebSocket (/ws) with per-topic sequencing, snapshot replay and duplicate suppression. The matrix switches from polling to push when the stream is healthy and falls back to REST automatically. - Cross-replica fan-out via Valkey —
console.valkey.modesupportsoff(in-process, single replica),bundled(ephemeral Valkey Deployment, no PVC by design) andexternal. NetworkPolicies open console→controller gRPC and console→Valkey on both sides. - Chart — new
console.*values (deployment, service, ingress, PDB, NetworkPolicy, bundled Valkey), documented indocs/configuration.mdand covered byvalues.schema.jsonand aci/console-values.yamllint profile.
Fixed¶
- Controller graceful shutdown hang with active streaming subscribers —
WatchEvents/WatchTasks/WatchPeershandlers now terminate on shutdown andGracefulStopis bounded with a hard-stop fallback (the Tasks/Peers case was latent since v1.3.0, masked by Kubernetes killing the pod after the grace period). - Fleet-safe events config — the
eventskey is omitted from the controller ConfigMap when disabled, so controller images without M2 support keep starting under strict config parsing.
Security¶
- grpc-go bumped to v1.82.1 (GO-2026-6061, reachable via the event stream).
Install¶
helm upgrade --install kconmon-ng oci://ghcr.io/esdmitrii/charts/kconmon-ng \
--version 1.4.0 \
--namespace kconmon-ng \
--create-namespace
kubectl plugin (via krew, from the release manifest):
kubectl krew install --manifest-url \
https://github.com/EsDmitrii/kconmon-ng/releases/download/v1.4.0/kconmon.yaml
Images¶
ghcr.io/esdmitrii/kconmon-ng-agent:1.4.0
ghcr.io/esdmitrii/kconmon-ng-controller:1.4.0
ghcr.io/esdmitrii/kconmon-ng-console:1.4.0
kconmon-ng v1.3.3¶
Chart-focused release. The Go agent/controller code is unchanged from v1.3.2; the
:1.3.3images are a version-synchronized rebuild (the release tag drives both the chart version and the image tag).
Fixes¶
- ICMP checker on runtimes with a closed
net.ipv4.ping_group_range— the ICMP checker opens an unprivileged ICMP "ping" socket (SOCK_DGRAM), which the kernel gates onnet.ipv4.ping_group_range, not onNET_RAW. Some container runtimes leave this at the closed kernel default (1 0), so the checker failed withsocket: permission deniedon those nodes. The agent Pod now sets the safe, namespaced sysctlnet.ipv4.ping_group_range=0 2147483647, so ping sockets work regardless of the runtime default.
Chart¶
- New
agent.podSecurityContextvalue exposes the agent Pod-levelsecurityContext(defaults to openingping_group_rangefor the ICMP checker). Setagent.podSecurityContext: {}to opt out. Documented in the chart README andvalues.schema.json. values.schema.json: HTTP target field corrected fromexpectedStatustoexpectStatusto match the checker's config (the schema key never matched the code, so a schema-guided value was silently ignored).
Install¶
helm upgrade --install kconmon-ng oci://ghcr.io/esdmitrii/charts/kconmon-ng \
--version 1.3.3 \
--namespace kconmon-ng \
--create-namespace
kubectl plugin (via krew, from the release manifest):
kubectl krew install --manifest-url \
https://github.com/EsDmitrii/kconmon-ng/releases/download/v1.3.3/kconmon.yaml
Images¶
kconmon-ng v1.3.2¶
Note: this is the first fully working release of the on-demand diagnostics feature set. v1.3.0 was aborted mid-release (GitHub immutable releases sealed it before all assets were attached); v1.3.1 published but shipped a krew manifest with an invalid version string, so
kubectl krew install --manifest-urlrejected it. v1.3.2 carries the same content with a valid krew manifest. The v1.3.0/v1.3.1 tags are retired.
Features¶
-
kubectl-kconmonplugin (on-demand diagnostics) — a new kubectl plugin talks to the controller's HTTP API through a client-go port-forward, so operators can inspect topology (kubectl kconmon topology/agents) and run one-shot connectivity checks (kubectl kconmon check SRC DST --type …,kubectl kconmon mtr SRC DST) between any two nodes without opening Grafana. Table or-o jsonoutput; a failed check exits2(distinct from1for CLI/API errors) so it composes in shell pipelines. Install via krew from the release manifest (see Install below). -
On-demand diagnostics API — new
POST /api/v1/diagnosticscontroller endpoint runs a single check (tcp/udp/icmp/dns/http/mtr) from a source node's agent to a destination and returns theCheckResultverbatim. Served by the leader only;?timeout=caps the wait (default 60s, max 120s). This is the endpoint the plugin drives. Seedocs/api.md. -
Graceful agent deregistration on SIGTERM — a restarting agent now deregisters from the controller on shutdown, so peers drop it immediately instead of waiting out the heartbeat TTL. This removes the transient false-loss window that a rolling agent restart used to leave in its own metrics.
Security¶
- Toolchain and dependency bumps — Go toolchain
go1.26.4;google.golang.org/grpc1.79.1 → 1.82.0,golang.org/x/net0.51 → 0.56,golang.org/x/sys0.41 → 0.46, and OpenTelemetry 1.41 → 1.44. This clears the CVE findings behind the previous Artifact Hub security-report grade. govulncheckin CI — a dedicated CI job runsgovulncheck ./...on every PR and tag; Dependabot (gomod / github-actions / docker, weekly) keeps dependencies current so CVE fixes land as normal PRs instead of accumulating until the next scan.
Supply chain¶
- The Helm chart is now signed with cosign (keyless, by digest) — v1.3.2 is the first signed release. Artifact Hub repository metadata continues to be published as an ORAS artifact.
Docs¶
- README reworked with an "On-demand diagnostics (kubectl plugin)" section and real command output.
docs/api.mddocuments the fullPOST /api/v1/diagnosticscontract (request fields, status codes,?timeout=cap, and ICMP / MTR response examples).
Install¶
helm upgrade --install kconmon-ng oci://ghcr.io/esdmitrii/charts/kconmon-ng \
--version 1.3.2 \
--namespace kconmon-ng \
--create-namespace
kubectl plugin (via krew, from the release manifest):
kubectl krew install --manifest-url \
https://github.com/EsDmitrii/kconmon-ng/releases/download/v1.3.2/kconmon.yaml
Images¶
kconmon-ng v1.2.0¶
Features¶
-
Automatic zone discovery — agents no longer need a statically configured zone. The controller enriches each agent registration with the node's failure-domain zone taken from its node informer (
failureDomainLabel, defaulttopology.kubernetes.io/zone) and returns the resolved metadata inRegisterResponse.agent; the agent adopts it for allsource_zone/destination_zonemetric labels.KCONMON_NG_ZONE(agent.zonein Helm values) remains an explicit override and always wins. Node zone relabels propagate to peers via a FULL_SYNC peer update; the relabeled node's ownsource_zonerefreshes on its next re-registration. Per-zone metrics and the Zone Heatmap dashboard now work out of the box on multi-zone clusters. -
Self-monitoring — new gauge
kconmon_ng_controller_expected_agents(count of schedulable nodes from the controller's node informer) and two PrometheusRule alerts:KconmonAgentsMissing(warning: registered < expected for 10m) andKconmonControllerDown(critical:absent(kconmon_ng_controller_leader == 1)for 5m). Degradation of kconmon-ng itself now alerts instead of failing silently. Requirescontroller.leaderElection: true(default) for the node informer.
Breaking-ish Changes¶
- Strict config parsing — the application config (ConfigMap /
--configfile) is now decoded with unknown-field rejection and per-checker semantic validation (intervals/timeouts > 0 for enabled checkers, HTTP target URL scheme/host, DNS resolver host[:port], non-empty DNS hosts). A typo'd or invalid config now fails startup and is rejected on hot-reload (the previous config stays active) instead of being silently ignored. Review your values overrides before upgrading: a config that previously "worked" by accident will now fail loudly.timeout >= intervallogs a warning but does not fail.
Helm Chart / Artifact Hub¶
- Chart README is now packaged into the chart archive — the Artifact Hub package page renders description, install instructions, values and metrics reference instead of "This package version does not provide a README file".
homeandsourcesadded toChart.yaml; Artifact Hub repository metadata (artifacthub-repo.yml) is published as an ORAS artifact on release for repository verification.agent.zoneis now documented as an optional override (auto-discovery is the default).
Dashboards¶
- Overview / MTR Triggers Count — switched from
increase(...[$__range])to a plainsum(...):increase()misses counter births on freshly restarted agent pods and chronically undercounted exactly when MTR fires most (pod churn).
Local Development¶
hack/local-test.shhardening: unique image tag per build (minikube's image-load cache silently kept stale same-tag images on re-runs),set -e/pipefailfixes (((ok++))pre-increment exit, SIGPIPE onhead-truncated pipes), port-forward cleanup.
Upgrade Notes¶
- Validate your config overrides against the stricter parser before rolling out (a quick check:
helm template ... | <render your config>and run the controller/agent with--configlocally, or just watch pod readiness on a staging cluster first). - If you previously set
agent.zoneto force a zone, you can keep it (it still wins) or drop it to switch to automatic discovery. - Metric label sets are unchanged; the new alerts ship in the chart's default
prometheusRule.rulesand are inert unlessprometheusRule.enabled: true.
Install¶
helm upgrade --install kconmon-ng oci://ghcr.io/esdmitrii/charts/kconmon-ng \
--version 1.2.0 \
--namespace kconmon-ng \
--create-namespace
Images¶
kconmon-ng v1.1.0¶
Bug Fixes¶
-
MTR memory leak —
lastRunmap inMTRCheckercould grow unboundedly in long-running agents on large clusters where node pairs come and go. Expired entries are now purged inline on eachTryAcquirecall while the lock is already held, keeping the map size proportional to active pairs within the current cooldown window. -
HTTP body pattern mismatch counted as success — when a
bodyPatterncheck failed, the checker setStatusCode = -1, which was not caught by the result handler's>= 400guard and was silently recorded asresult="success"in Prometheus. The status code field now always carries the real HTTP status. A dedicatedBodyMismatch boolfield signals pattern failure, and the result handler correctly marks such checks asresult="fail".
Improvements¶
-
Configurable DNS resolver dial timeout — the dialer timeout for custom DNS resolvers was previously hard-coded to 5 seconds and could not be adjusted for slow or distant resolvers. A new
timeoutfield has been added to the DNS checker config (default:5s). Update your Helm values or config file to override: -
Jitter in agent re-registration backoff — when the controller restarts, all agents previously retried at exactly the same interval, causing a thundering herd. Up to 25% random jitter is now added to each retry wait, spreading reconnect load across agents.
-
MTR buffer allocation — the 1500-byte read buffer in the traceroute loop was allocated once per hop. It is now allocated once per trace, reducing GC pressure under frequent MTR runs.
Helm Chart¶
config.checkers.dns.timeoutadded tovalues.yaml(default:5s).
Tests¶
- Updated
TestHTTPCheckerBodyPatternMismatch: verifiesBodyMismatch=trueand real HTTP status code instead of the former-1sentinel. - Added
TestHTTPCheckerBodyPatternMatch: verifiesBodyMismatch=falseon a successful pattern. - Added
TestDNSCheckerTimeoutPropagated: verifies the configured timeout is stored on the checker. - Added
TestMTRCheckerExpiredEntriesPurged: verifies stale entries are removed fromlastRunafter cooldown expiry.
Upgrade Notes¶
The HTTPDetails.StatusCode field no longer returns -1 for body pattern mismatches — it now
always holds the actual HTTP response status code. If you have alerting or dashboards that rely
on statusCode == -1 to detect body mismatch failures, update them to use the new
bodyMismatch field in the JSON result or the result="fail" label in Prometheus metrics.
Install¶
helm upgrade --install kconmon-ng oci://ghcr.io/esdmitrii/charts/kconmon-ng \
--version 1.1.0 \
--namespace kconmon-ng \
--create-namespace
Images¶
kconmon-ng v1.0.0 — Initial Release¶
Kubernetes Node Connectivity Monitor, next-generation rewrite with a gRPC-based agent/controller architecture and rich observability out of the box.
Features¶
Core - Agent/controller architecture with gRPC streaming peer updates - TCP, UDP, ICMP, DNS, and HTTP checkers with configurable timeouts and thresholds - Per-node and per-zone Prometheus metrics for all check types - Reactive MTR traceroute on check failure with per-pair cooldown - Self-probe prevention: peers filtered by agent ID, node name, and pod IP - Atomic gauge reset on peer topology changes to prevent stale metrics
Scheduler
- Pause/resume support, per-check jitter, and NodeLocal checker mode
- NodeWatcher: live Kubernetes node info exposed via /api/v1/topology
Observability - Grafana dashboards: Overview, Node Detail, Cross-Zone Heatmap - Helm chart with ServiceMonitor, PrometheusRule, NetworkPolicy, PDB, and RBAC
Operations
- Multi-arch Docker images (linux/amd64, linux/arm64) published to GHCR
- Local dev tooling: hack/local-test.sh with Minikube + Prometheus + Grafana stack
- Chaos testing guide with NetworkPolicy example
Install¶
helm install kconmon-ng oci://ghcr.io/esdmitrii/charts/kconmon-ng \
--version 1.0.0 \
--namespace kconmon-ng \
--create-namespace