Skip to content

Helm values

The chart installs the monitor (agent, controller, console) and nothing else: PostgreSQL, a Redis-compatible bus and Prometheus are infrastructure it consumes, each configured by one DSN or URL. The full values.yaml below documents every key inline and is the authoritative reference; this table is the short list you will actually touch.

Key values at a glance

Key Default What it does
config.metricsPrefix kconmon_ng Prefix for every exported metric; changing it renames all of them
config.checkers.tcp.enabled true TCP checker (interval 5s, timeout 1s)
config.checkers.udp.enabled true UDP checker (interval 5s, timeout 250ms, packets: 5)
config.checkers.icmp.enabled true ICMP checker (interval 5s, timeout 1s); unprivileged socket, no added capabilities
config.checkers.pmtu.enabled true Path MTU probe (interval 60s, timeout 500ms per datagram, size: 0 = the MTU of the route to the peer: a CNI route MTU such as Cilium's when set, else the egress device's); DF-marked UDP to the peer's echo port, no added capabilities. Keep interval at 28× timeout (14s at the default) or more and under 3m; the agent warns outside that range, since from 3m on PathMTUBlackHole gets few probes per window. The chart writes a pmtu key only when it differs from these defaults, and tuning one needs agent images 2.5.0 or newer
config.checkers.dns.enabled true DNS checker (interval 5s, timeout 2s)
config.checkers.http.enabled false HTTP checker; targets are required when enabled
config.checkers.external.enabled false Probes to non-peer destinations, gated by allowedCidrs; see External targets
config.checkers.mtr.cooldown 60s Minimum gap between reactive MTR traces for the same (src, dst) pair
config.checkers.mtr.maxHops 30 Traceroute hop ceiling (1–64)
config.controllerAgentTtl 30s Evict an agent missing heartbeats for this long; any Go duration spelling (45s, 1.5m, .5m), min 10s, and the chart refuses less at install
agent.tolerations [{operator: Exists}] Run the agent on every node, tainted ones included
agent.metrics.detail full Scrape-time cardinality valve on the agent ServiceMonitor and the external-agent ScrapeConfig: full / counters-only / zone-only (78 / 14 / 0 series per directed pair); needs serviceMonitor.enabled or scrapeConfig.externalAgents.enabled
agent.pingGroupRange true Render the net.ipv4.ping_group_range sysctl the ICMP socket needs; not rendered under agent.hostNetwork, where the node OS must set it
agent.hostNetwork false Run the agents in the node's network namespace so external hosts can reach them without a routable pod network. Changes what every in-cluster pair measures (node-to-node underlay, not the CNI datapath); needs a privileged namespace, free ports on every node, and networkPolicy.nodeCidrs when policies are on. Read External agents first
agent.dnsPolicy "" Pod dnsPolicy, passed through verbatim (ClusterFirst, ClusterFirstWithHostNet, Default, None); empty renders ClusterFirstWithHostNet under agent.hostNetwork and nothing otherwise
agent.updateStrategy RollingUpdate, maxUnavailable: 1 DaemonSet rollout, passed through verbatim; raise on large fleets
controller.replicaCount 1 Controller replicas; only the leader is active
controller.leaderElection true Leader election; false also disables zone enrichment and expected_agents
controller.events.enabled false Domain event stream for the Console's realtime pages (leader-only)
controller.externalGateway.enabled false The TLS gateway for external agents: a second gRPC listener, exposed by its own NodePort/LoadBalancer Service
controller.externalGateway.port 9443 Gateway listener; must differ from config.{httpPort,grpcPort,metricsPort}
controller.externalGateway.tls.secretName "" kubernetes.io/tls Secret with the serving pair; REQUIRED when enabled. tls.clientCaKey names the client-CA bundle key in the same Secret; empty means token-only mode
controller.externalGateway.bootstrapToken.secretName "" Secret holding the shared bearer token (< 16 chars is refused); REQUIRED when enabled
controller.prometheusSD.enabled true Serve GET /api/v1/prometheus/sd, the external-agent target list, on httpPort and metricsPort; false answers 404 on both. Written to the shared ConfigMap only when false, so an older controller image never sees the key
console.enabled false Deploy the Console
console.replicas 1 More than 1 REQUIRES redis.existingSecret; the chart refuses the combination otherwise
console.prometheus.url "" Required by the data pages (Matrix, Metrics, PromQL), which answer 503 without it
console.networkPolicy.prometheusTargetPort 0 The pod port behind console.prometheus.url when that URL names a Service whose targetPort differs from its port (the Bitnami Thanos chart: Service 9090, pod 10902). A NetworkPolicy sees the pod port after the Service DNAT, so the default Prometheus egress rule then opens both ports. 0 opens the URL's port only; ignored when console.networkPolicy.prometheusEgress is set. redisEgress and databaseEgress have no such key: a Service mapping 6379 or 5432 to another targetPort needs the list set on the pod port
console.auth.mode anonymous anonymous / local / header / oidc
console.auth.groupRoles {} IdP group → console role map; what makes an oidc/header install usable from a cold database
console.auth.header.trustedProxyCIDRs [] Identity: the authenticating proxy whose X-Remote-User/X-Remote-Groups the console believes. Required in header mode; name that proxy and nothing wider. Gives the client address too, but only while console.clientAddress.trustedProxyCIDRs is empty
console.clientAddress.trustedProxyCIDRs [] Proxy networks whose X-Forwarded-For gives the client address, in every mode; never identity. Behind an Ingress or a NAT, list the ingress controller's addresses (its pod CIDR at the widest; a pod inside the list can name any client address), or the per-address login, OIDC, runs and PromQL budgets, the /ws cap and the audit address all see one shared address. Empty falls back to the header list, and so does a console image older than 2.5.0, for which the chart does not write the key
console.websocket.maxConnections / maxConnectionsPerAddress / maxConnectionsPerSubject 1024 / 256 / 32 Open /ws sockets per console replica, per client address and per user or token; 0 turns a cap off. A refused socket closes with 1013 and counts in kconmon_ng_console_ws_refused_total
console.alerting.enabled false Console-managed alert rules, reconciled into one PrometheusRule; needs a database and the operator CRD
console.webhooks.existingSecret "" Secret with the AES-256-GCM key encrypting webhook signing secrets at rest (key console-webhooks-encryption-key); empty leaves endpoint create and test at 503; see Set up alerting
console.scheduler.enabled false The schedule/dispatch loop for scheduled and external checks
console.scheduler.tickInterval 5s Poll cadence of that loop; must be > 0 when enabled
database.existingSecret "" Secret holding a postgres:// DSN; empty means an in-memory console
database.retentionDays 90 Daily prune of stored history; 0 keeps everything
redis.existingSecret "" Secret holding a redis:// DSN; empty means the in-process bus (single replica only)
dashboards.enabled false Ship the Grafana dashboards as sidecar-labelled ConfigMaps
serviceMonitor.enabled false Prometheus Operator ServiceMonitor for agents, controller and console
scrapeConfig.externalAgents.enabled false Prometheus Operator ScrapeConfig that reads the controller's HTTP SD endpoint and scrapes external agents; refused without controller.externalGateway.enabled or with controller.prometheusSD.enabled=false
scrapeConfig.externalAgents.labels {} Selector labels your Prometheus requires on the object; kube-prometheus-stack selects only release: <its release name>, an empty scrapeConfigSelector needs nothing
scrapeConfig.externalAgents.jobName "" Job label; empty means <release>-agent-external. Keep kconmon in it (the dashboards filter on it); KconmonExternalAgentDown follows whatever name you set
scrapeConfig.externalAgents.refreshInterval 30s How often Prometheus re-reads the target list; matches config.controllerAgentTtl
scrapeConfig.externalAgents.interval "" Scrape interval; empty falls back to serviceMonitor.interval
prometheusRule.enabled false The fourteen built-in alert rules as one PrometheusRule (thirteen on by default)
prometheusRule.<alertName> all enabled Per-rule enabled / threshold / for / severity knobs; nodeUnreachable and nodeIsolated also take minPeers (2), and pathMtuBlackHole drives ZonePathMTUBlackHole too
prometheusRule.pathMtuBlackHole.sustainedThreshold 0.1 Second arm of PathMTUBlackHole and ZonePathMTUBlackHole: a pmtu failure ratio over 30m above this, with at least two failed probes in that window and one in the last 10m, fires as well, which catches a black hole on one of several ECMP paths. A ratio 0.0-1.0; 1 turns the arm off
prometheusRule.externalAgentDown.enabled false KconmonExternalAgentDown: an external agent the SD endpoint lists at up == 0 for for (5m, warning); the job exists only with the ScrapeConfig or a hand-written job named *agent-external*
prometheusRule.additionalRules [] Your rules, appended verbatim
networkPolicy.enabled false NetworkPolicies per component: <fullname>-agent, <fullname>-controller (with its -apiserver egress policy, and -gateway when the gateway is on) and, with the console on, <fullname>-console. On Cilium it adds CiliumNetworkPolicy objects, see the next row
networkPolicy.prometheusNamespace "" REQUIRED for the scrape rule to render when policies are on; unset means up == 0, visibly
networkPolicy.ciliumKubeAPIEgress auto auto / true / false. Renders the <fullname>-kube-apiserver CiliumNetworkPolicy, allowing TCP 443/6443 to the kube-apiserver entity for the controller, and for the console with console.kubernetesContext.enabled or console.alerting.enabled; Cilium matches no ipBlock against that entity, so without it the controller stays NotReady. With agent.hostNetwork or controller.externalGateway.enabled it also renders <fullname>-node-ingress, which admits the remote-node and host entities to the controller's grpcPort and gateway port, since nodeCidrs and externalAgentCidrs match no node IP on Cilium. auto renders it when the cluster serves cilium.io/v2 CiliumNetworkPolicy; plain helm template needs --api-versions cilium.io/v2/CiliumNetworkPolicy or true. See NetworkPolicy on Cilium, Calico and Antrea
networkPolicy.httpEgress [] Egress for HTTP checker targets; empty renders TCP 80/443 to 0.0.0.0/0 minus clusterCIDRs. An in-cluster target needs a selector peer on the pod's own port on every CNI (Cilium never matches a pod IP against an ipBlock, Calico and Antrea see the targetPort after kube-proxy's DNAT); a set list replaces the default, so keep an ipBlock rule for external targets. console.networkPolicy.webhookEgress works the same for webhook receivers
networkPolicy.clusterCIDRs [] IPv4 pod and Service CIDRs, rendered as except entries of the default 0.0.0.0/0 rules of httpEgress, console.networkPolicy.webhookEgress, oidcEgress and geoipEgress. On Calico and Antrea that ipBlock also matches pod IPs, so without this list those defaults open every pod on their ports; Cilium never matches a pod against it. The default oidcEgress also carries a namespaceSelector: {} peer on 443, so it keeps every pod open on 443 with this list too; only an explicit console.networkPolicy.oidcEgress narrows it. The apiserver defaults (networkPolicy.kubeAPIEgress, console.networkPolicy.kubeAPIEgress) get no except and keep every pod open on 443/6443 to the controller and to a console that calls the apiserver until you name the apiserver endpoint in them. Entries must be network addresses: /0 and host bits set fail the render. See NetworkPolicy on Cilium, Calico and Antrea
networkPolicy.externalAgentCidrs [] Source CIDRs of external agents, opened on the gateway port toward the controller pods alone; REQUIRED when the gateway and this policy are both on (plain CIDR strings with a prefix length). Mind NAT, see External agents; on Cilium a node IP here matches nothing and ciliumKubeAPIEgress admits node sources instead
networkPolicy.externalPeerCidrs [] The probe half: the same hosts' CIDRs spliced into the agent↔agent rules in both directions (UDP grpcPort, TCP httpPort, the ports-less ICMP/MTR rule), never into the gateway rule. Without it an external agent registers and every cell between it and the cluster stays red
networkPolicy.dnsEgress [] Replaces the default cluster DNS egress rule (UDP/TCP 53 to any pod) in all three policies; needed for NodeLocal DNSCache or any host-network resolver, which a namespaceSelector cannot match. Keep namespaceSelector: {} in the list if the kube-dns pods must stay reachable
networkPolicy.nodeCidrs [] Node CIDRs admitted to the controller's gRPC port; REQUIRED with agent.hostNetwork when policies are on, since host-network agents register from node IPs that no pod selector matches, and the chart refuses to render the policy without it. On Cilium this ipBlock matches nothing; the <fullname>-node-ingress CiliumNetworkPolicy from ciliumKubeAPIEgress admits the nodes there

Secrets follow one pattern everywhere: existingSecret names a Secret you created (recommended), or a sibling secret.create: true block lets the chart render it, meant for a secrets injector's ${vault:...} placeholders, not for literals. The chart README lists every consumer and its key.

Full values.yaml

The complete, commented values.yaml of the current chart, embedded from the repo at build time so it cannot drift from what the chart ships:

charts/kconmon-ng/values.yaml: every key, documented inline
---
# kconmon-ng default values.
#
# The chart installs the MONITOR — agent, controller, console — and nothing else. PostgreSQL, a
# Redis-compatible bus and Prometheus are yours to run however you already run infrastructure; each
# is configured here by one DSN or URL. See the chart README for the stack it is tested against.
#
# Sections:
#   1 Naming          6 Controller
#   2 Shared config   7 Console
#   3 Database        8 Observability
#   4 Redis           9 Kubernetes
#   5 Agent

# == 1. Naming ==

# Overrides the chart name used in resource names and labels.
nameOverride: ""
# Overrides the full resource-name prefix outright.
fullnameOverride: ""

# == 2. Shared config == runtime config for controller and agents, hot-reloaded from a ConfigMap.

config:
  # Prefix for every exported metric; changing it renames all of them.
  metricsPrefix: kconmon_ng
  # Serves /metrics and health; must differ from grpcPort.
  httpPort: 8080
  # Controller peer-list/topology API; must differ from httpPort.
  grpcPort: 9090
  # /metrics and the health endpoints, on a listener of their OWN. The controller's API shares
  # httpPort and authenticates nothing, so this is the port the scrape NetworkPolicy opens — letting
  # a scraper in must not mean letting its whole namespace drive the fleet.
  metricsPort: 9091
  # debug | info | warn | error.
  logLevel: info
  # json | text.
  logFormat: json
  # Node label the controller reads for each agent's zone.
  failureDomainLabel: topology.kubernetes.io/zone
  # Grace period before an agent without a heartbeat leaves the peer list; at least 10s (two
  # heartbeats). Never 0 — the controller refuses a non-positive TTL.
  controllerAgentTtl: 30s
  # Per-protocol checkers; for an enabled one, interval must be at least 100ms and timeout at least 1ms.
  checkers:
    # TCP dial to a peer's port (L4 reachability).
    tcp:
      enabled: true
      interval: 5s
      timeout: 1s
    # UDP probe measuring packet loss to a peer.
    udp:
      enabled: true
      interval: 5s
      timeout: 250ms
      # Packets per probe, used to compute the loss ratio; must be >= 1.
      packets: 5
    # ICMP ping reachability/loss; runs on the unprivileged ICMP socket
    # agent.podSecurityContext opens, no added capability.
    icmp:
      enabled: true
      interval: 5s
      timeout: 1s
    # Path MTU: full-size datagrams with DF set to each peer's echo port, bisected on loss. Finds the
    # MTU black hole that small probes and TCP handshakes never see. Tuning any key needs agent
    # images 2.5.0 or newer: the chart writes a key only when it differs from these defaults, and an
    # older agent refuses the unknown key.
    pmtu:
      enabled: true
      # Keep it at 28x timeout or more (14s at the default 500ms: the search gets half the interval
      # and a black-hole search needs about fourteen datagram timeouts) and under 3m, so
      # PathMTUBlackHole gets several probes in its 10m window; the agent warns outside that range.
      interval: 60s
      # Per datagram; the whole search is bounded by half the interval.
      timeout: 500ms
      # IP-level bytes to probe at; 0 = the MTU of the route to the peer (a CNI route mtu such as
      # Cilium's, else the egress device).
      size: 0
    # DNS resolution check.
    dns:
      enabled: true
      interval: 5s
      timeout: 2s
      # Names to resolve; must be non-empty when the checker is enabled.
      hosts:
        - kubernetes.default.svc.cluster.local
      # Explicit resolvers as "host" or "host:port"; empty uses the Pod's own.
      resolvers: []
    # HTTP endpoint check; targets are required when enabled.
    http:
      enabled: false
      interval: 30s
      timeout: 5s
      # [{url, method, expectStatus, bodyPattern}]
      targets: []
    # Reactive traceroute fired on probe failures.
    mtr:
      # Minimum gap between MTR runs for the same pair.
      cooldown: 60s
      # Maximum traceroute hops; 1-64.
      maxHops: 30
    # Probes to destinations that are not peer agents.
    external:
      # Off means the block is not parsed and no external probe can run.
      enabled: false
      # CIDR allowlist matched against the resolved address;
      # keep it in step with networkPolicy.externalEgress,
      # and the console refuses a target outside it.
      allowedCidrs: []
      # Carve-outs subtracted from allowedCidrs; denied wins.
      deniedCidrs: []
      # Cap on operator-defined external targets per agent.
      maxTargets: 100
      # Bounds resolve-and-authorise only, not the probe itself; 0 means 10s, anything else at least 1ms.
      timeout: 10s

# Probe topology plan, part of the same shared config (rendered into the shared ConfigMap, read by
# the controller's leader). full is today's behaviour: every agent probes every peer, N*(N-1)
# directed pairs. sparse trims that to a ring over sorted node names plus cross-zone chords —
# connectivity and zone-pair coverage hold (property-tested), while pairs and their metric series
# scale ~linearly with N instead of quadratically. Under sparse, pairs the plan drops STOP being
# probed: the console matrix marks them unplanned, and PairWentSilent only fires for pairs present
# in <prefix>_probe_intended, which agents export from this release. Version skew: the key is
# emitted only when mode=sparse, because a pre-2.3.0 controller image rejects it and crashloops —
# upgrade the images first, flip the mode second.
topology:
  # full | sparse.
  mode: full
  # Read only when mode=sparse.
  sparse:
    # Ring successors per node over the sorted node-name circle; the ring is what guarantees the
    # plan stays connected.
    ringDegree: 2
    # Cross-zone peers each agent probes on top of the ring, 0-64, chosen by HRW hashing; 0 leaves
    # zone coverage to whatever the ring happens to cross.
    zoneChords: 2
    # Fleets SMALLER than this many nodes get the full mesh regardless of mode; 0 means no such
    # floor. Sparse only pays for itself at scale, and a small fleet is better off measured whole.
    autoThreshold: 0

# == 3. Database == PostgreSQL the Console stores history, users and the audit log in.
# BRING YOUR OWN server — CNPG, Percona, RDS, plain postgres, anything that speaks a postgres:// DSN;
# the chart installs nothing and only needs the Secret holding that DSN. Nothing configured means an
# in-memory console: no history, no incidents, no local/oidc auth. Tested stack: CloudNativePG — see
# the chart README for the exact Cluster I test with.

database:
  # Secret holding a full postgres:// DSN; setting it is what enables the database.
  existingSecret: ""
  # Key inside that Secret; the chart-managed path writes this same key.
  existingSecretKey: console-database-dsn
  # Chart-managed alternative to existingSecret; prefer it for an injector's ${vault:...}
  # placeholders rather than for literals.
  secret:
    create: false
    # Empty derives the name from the release name.
    name: ""
    annotations: {}
    labels: {}
    dsn: ""
  maxConns: 10
  connectTimeout: 10s
  # Run the embedded goose migrations at startup.
  migrateOnStart: true
  # Daily prune of rows older than this; 0 keeps everything.
  retentionDays: 90

# == 4. Redis == the Console's session store, rate-limit counters and cross-replica pub/sub.
# BRING YOUR OWN Redis-compatible server — Valkey, Redis, sentinel-fronted, ElastiCache; the chart
# installs nothing and only needs an address. Empty means the in-process bus, and then
# console.replicas must be 1 — the chart ENFORCES that now rather than only saying it: the fan-out,
# the sessions and the fixed-window rate-limit counters all live in this server, so a second replica
# without one silently doubles every rate limit and loses cross-replica realtime.
# Nothing here is durable, so persistence on the server is optional.
# Tested stack: valkey-helm — see the chart README for the exact values I test with.

redis:
  # Secret holding a full redis:// DSN; setting it is enables the shared bus.
  # redis:// | rediss:// (TLS) | valkey:// | valkeys:// | unix://, with the username, password and
  # database number where the server's own documentation puts them.
  existingSecret: ""
  # Key inside that Secret; the chart-managed path writes this same key.
  existingSecretKey: console-redis-dsn
  # Chart-managed alternative to existingSecret; same caveat as the database secret above.
  secret:
    create: false
    name: ""
    annotations: {}
    labels: {}
    dsn: ""
  dialTimeout: 5s

# == 5. Agent == DaemonSet running the enabled checkers against peers on every node.

agent:
  # Explicit zone override; empty lets the controller resolve it from the node.
  zone: ""
  # Scrape-time cardinality valve, rendered as metricRelabelings on the agent ServiceMonitor.
  # It needs serviceMonitor.enabled and the chart refuses the combination otherwise; without the
  # operator, copy the equivalent metric_relabel_configs from docs/metrics.md into your own scrape
  # config. What each mode keeps, per DIRECTED NODE PAIR (N nodes make N*(N-1) of them):
  #   full          78 series/pair: everything the agent exports.
  #   counters-only 14 series/pair: drops the four per-pair histograms (64 of the 78); the
  #                 loss/jitter gauges and result counters stay, so every pair alert keeps firing.
  #   zone-only     ~0 series/pair: drops every series naming a destination node (histograms,
  #                 gauges, counters, MTR). What remains is the zone family (~76 series per
  #                 DIRECTED ZONE PAIR, so ~76*Z^2 for Z zones) plus the DNS/HTTP/external
  #                 families, which are linear in node count. Path MTU black holes then alert as
  #                 ZonePathMTUBlackHole, per zone pair.
  # zone-only requires agents new enough to export the zone family, or Prometheus goes dark on the
  # mesh while the console (which does not scrape) keeps working.
  metrics:
    # full | counters-only | zone-only.
    detail: full
  podAnnotations: {}
  image:
    repository: ghcr.io/esdmitrii/kconmon-ng-agent
    # Empty means the chart appVersion.
    tag: ""
    pullPolicy: IfNotPresent
  resources:
    limits:
      cpu: 200m
      memory: 128Mi
    requests:
      cpu: 50m
      memory: 64Mi
  # Default runs an agent on every node, including tainted ones.
  tolerations:
    - operator: Exists
  # No added capabilities: ICMP and MTR use the unprivileged ICMP socket the ping_group_range sysctl
  # below opens. Add NET_RAW here only if you also make it effective (a root container, or ambient
  # caps) — it buys intermediate MTR hops on kernels that withhold them from ping sockets, and costs
  # the restricted PSS profile.
  securityContext:
    allowPrivilegeEscalation: false
    readOnlyRootFilesystem: true
    capabilities:
      drop:
        - ALL
  # restricted-PSS compliant; ping_group_range is a kubelet safe sysctl and is what makes the
  # unprivileged ICMP socket work.
  podSecurityContext:
    runAsNonRoot: true
    runAsUser: 65532
    seccompProfile:
      type: RuntimeDefault
  # Render the net.ipv4.ping_group_range sysctl the unprivileged ICMP socket needs. Its own switch,
  # because opting out through `agent.podSecurityContext: null` would also drop runAsNonRoot,
  # runAsUser and seccompProfile, which a namespace enforcing restricted PSS demands.
  pingGroupRange: true
  # Run the agents in the NODE's network namespace. Off by default; read this before flipping it.
  # It changes WHAT IS MEASURED for the whole DaemonSet, not only where the pods listen: every
  # in-cluster pair then probes node IP to node IP over the underlay, and the CNI datapath (overlay,
  # conntrack, NetworkPolicy enforcement) is no longer exercised, so the failure class this tool
  # exists to catch hides behind a green matrix. Turn it on only when the goal is visibility between
  # external agents and the cluster on a pod network the external hosts cannot route to.
  # What must already be true, because each miss fails in its own quiet way:
  #   - the namespace runs at PSS privileged or is exempted: baseline refuses hostNetwork at
  #     admission and the release still looks healthy;
  #   - TCP config.httpPort, UDP config.grpcPort and TCP config.metricsPort are free on EVERY node
  #     (`ss -lntup` first; Felix metrics also default to 9091 when enabled): an occupied port
  #     CrashLoops the agent on that node alone and fires KconmonAgentsMissing;
  #   - net.ipv4.ping_group_range is set by the node OS (sysctl.d, the same line the deb/rpm ships):
  #     the kubelet refuses net.* pod sysctls under host networking, so the chart stops rendering it;
  #   - networkPolicy.nodeCidrs is set when networkPolicy.enabled: registrations now arrive from
  #     node IPs, and the chart refuses to render the policy without it.
  # One host cannot run both a hostNetwork agent and a bare-host external agent: same IP, same ports.
  hostNetwork: false
  # Pod DNS policy, passed through verbatim. Empty renders ClusterFirstWithHostNet under hostNetwork
  # (controllerAddress is a bare Service name that only cluster DNS resolves; a host-network pod left
  # on ClusterFirst reads the node's resolv.conf instead) and nothing at all otherwise.
  dnsPolicy: ""
  nodeSelector: {}
  affinity: {}
  # Empty renders nothing. The agent is the workload worth raising: under node pressure the kubelet
  # evicts lowest priority first, and the first pod gone should not be the one reporting the node.
  priorityClassName: ""
  # DaemonSet rollout strategy, passed through verbatim. One node at a time is the safe default and
  # a slow one: on hundreds of nodes raise maxUnavailable (count or percentage) or use OnDelete.
  updateStrategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 1

# == 6. Controller == Deployment serving peer lists over gRPC and watching nodes as leader.

controller:
  # Controller replicas; only the leader is active, extras are HA standby.
  replicaCount: 1
  # Leader election; false also disables zone enrichment and expected_agents.
  leaderElection: true
  events:
    # Domain event stream for the Console's realtime ingester (leader-only).
    enabled: false
  # The external-agent target list (GET /api/v1/prometheus/sd) on httpPort and metricsPort, read by
  # scrapeConfig.externalAgents; false closes it. It reaches the shared ConfigMap only when false,
  # so an older controller image never sees the key.
  prometheusSD:
    enabled: true
  # SECOND gRPC listener for agents OUTSIDE the cluster: same services and registry, but TLS with a
  # bearer token — the in-cluster port is guarded by a NetworkPolicy, and whatever reaches a
  # NodePort/LoadBalancer is guarded by nothing but what is configured here. The in-cluster listener
  # and its ClusterIP Service are untouched by this block. The controller reads the certificate and
  # token ONCE at startup — the config hot-reload does not rebuild the listener — so rotating either
  # Secret's content needs a controller restart (kubectl rollout restart); the chart cannot see into
  # referenced Secrets and cannot roll the pods for you.
  externalGateway:
    enabled: false
    # Gateway listener; must differ from config.httpPort/grpcPort/metricsPort.
    port: 9443
    # The externally reachable Service, rendered only when enabled. It exposes the gateway port
    # ALONE — never the plaintext in-cluster gRPC port.
    service:
      # NodePort | LoadBalancer; no ClusterIP, an external agent cannot reach one.
      type: LoadBalancer
      # Cloud-LB knobs (internal LB, static IP) go here.
      annotations: {}
      # Fixed port for type NodePort; 0 lets the apiserver pick one.
      nodePort: 0
      # Cluster | Local; empty leaves the apiserver default (Cluster). Local preserves the agents'
      # source addresses, which networkPolicy.externalAgentCidrs matches on and the gateway's flood
      # eviction tells agents apart by, at the cost of only the nodes running a controller pod
      # answering. Under Cluster every agent arrives from a node IP: prefer Local with tight
      # externalAgentCidrs, or loadBalancerSourceRanges.
      externalTrafficPolicy: ""
      # Provider-enforced source filter for type LoadBalancer; the in-cluster half of the same
      # decision is networkPolicy.externalAgentCidrs.
      loadBalancerSourceRanges: []
    tls:
      # kubernetes.io/tls Secret holding the gateway's serving pair (tls.crt/tls.key), e.g. one
      # cert-manager maintains; REQUIRED when enabled.
      secretName: ""
      # Key in that SAME Secret holding the CA bundle that signed the agent CLIENT certificates
      # (conventionally ca.crt). Empty is token-only mode: fleet membership is authenticated but
      # agents cannot be told apart, so any token holder can impersonate any agent. Set it and the
      # gateway requires a verified client cert whose CN/URI SAN is pinned to the agent's node name.
      clientCaKey: ""
    # Shared bearer token every external agent must present; the controller refuses one shorter
    # than 16 characters.
    bootstrapToken:
      # Secret holding the token; REQUIRED when enabled.
      secretName: ""
      # Key inside that Secret.
      key: token
  image:
    repository: ghcr.io/esdmitrii/kconmon-ng-controller
    # Empty means the chart appVersion.
    tag: ""
    pullPolicy: IfNotPresent
  # restricted-PSS compliant; the controller only talks to the apiserver and gRPC.
  podSecurityContext:
    runAsNonRoot: true
    runAsUser: 65532
    seccompProfile:
      type: RuntimeDefault
  securityContext:
    allowPrivilegeEscalation: false
    readOnlyRootFilesystem: true
    capabilities:
      drop:
        - ALL
  # The controller is not on the data path, so defaults are small.
  resources:
    limits:
      cpu: 200m
      memory: 128Mi
    requests:
      cpu: 50m
      memory: 64Mi
  nodeSelector: {}
  tolerations: []
  affinity: {}
  # Empty renders nothing.
  priorityClassName: ""
  # PodDisruptionBudget; rendered only at replicaCount > 1, since minAvailable 1 over a single
  # replica blocks every voluntary eviction and hangs node drains.
  pdb:
    enabled: true
    minAvailable: 1

# == 7. Console == optional web UI and API; every block below is inert while enabled is false.

console:
  # -- 7.1 core --
  enabled: false
  # Console replicas; the console is stateless.
  # One by default, because the default has no redis and replicas without a shared bus each keep
  # their own sessions and budgets. Raise this once redis.existingSecret is set; the chart refuses the
  # combination otherwise.
  replicas: 1
  image:
    repository: ghcr.io/esdmitrii/kconmon-ng-console
    # Empty means the chart appVersion.
    tag: ""
    pullPolicy: IfNotPresent
  # Controller API access.
  controller:
    # Empty derives from this release's controller Service.
    url: ""
    timeout: 10s
    # EventStream host:port; empty derives it when controller.events.enabled.
    grpcAddress: ""
  # Prometheus query API; required by the data pages, which 503 without it.
  prometheus:
    url: ""
    queryTimeout: 30s
    # Largest query_range window the guarded proxy accepts.
    maxRange: 24h
    # Response size cap in bytes.
    maxResponseBytes: 8388608

  # -- 7.2 auth -- authentication mode and the role every request resolves to.
  auth:
    # anonymous | local | header | oidc.
    mode: anonymous
    # Role for an authenticated subject no binding matches; empty means none.
    # One of the built-ins: viewer | operator | alert-editor | admin.
    defaultRole: ""
    # Map a GROUP the identity provider asserts onto a role this console grants -- the declarative
    # half of RBAC, and the one that makes an oidc or header install usable from a cold database.
    # Roles resolve as the union of these and any binding made through the API; what is granted here
    # cannot be revoked through the API, which is the point. A group absent from this map grants
    # nothing, so the map is an allow-list. Values may be a built-in (viewer|operator|alert-editor|
    # admin) or the name of a custom role you created.
    #   groupRoles:
    #     platform-oncall: admin
    #     everyone: viewer
    groupRoles: {}
    anonymous:
      # Fixed role for every anonymous request. Same four built-ins as defaultRole.
      role: viewer
    # mode=local; requires a database.
    local:
      # Created on first start only, while the users table is empty.
      bootstrapAdmin: ""
      # Existing Secret holding the bootstrap password.
      existingSecret: ""
      # Key inside that Secret; the chart-managed path writes this same key.
      existingSecretKey: console-local-admin-password
      # Chart-managed alternative to existingSecret; values may be ${vault:...} placeholders.
      secret:
        create: false
        # Empty derives the name from the release name.
        name: ""
        annotations: {}
        labels: {}
        # Bootstrap admin password; required when create is true.
        password: ""
    # mode=header: identity from a trusted in-cluster reverse proxy.
    header:
      userHeader: X-Remote-User
      groupsHeader: X-Remote-Groups
      groupsDelimiter: ","
      # Required and non-empty for mode=header, else headers are an auth bypass. Header mode trusts
      # identity headers ONLY from these: list the authenticating proxy, never the whole pod CIDR.
      # For the client address behind an Ingress use console.clientAddress.trustedProxyCIDRs. This
      # list stands in for it while that one is empty or the console image is older than 2.5.0,
      # which is why it reaches the console config in every mode and values from 2.4 keep working.
      trustedProxyCIDRs: []
    # mode=oidc; requires a database, and redis.existingSecret for replicas > 1.
    oidc:
      issuer: ""
      clientID: ""
      redirectURL: ""
      scopes: [openid, profile, email, groups]
      # DISPLAY name only, and only in the UI: identity is always "oidc:<sub>" — the one claim OIDC
      # Core 5.7 allows as an identifier — and both RBAC bindings and the audit log hang off that.
      # Changing this renames a person in the header menu and moves nothing else.
      usernameClaim: preferred_username
      groupsClaim: groups
      # Existing Secret holding the OIDC client secret.
      existingSecret: ""
      # Key inside that Secret; the chart-managed path writes this same key.
      existingSecretKey: console-oidc-client-secret
      # Chart-managed alternative to existingSecret; values may be ${vault:...} placeholders.
      secret:
        create: false
        # Empty derives the name from the release name.
        name: ""
        annotations: {}
        labels: {}
        # OIDC client secret; required when create is true.
        clientSecret: ""
    # Session cookie for every non-anonymous mode.
    session:
      # Absolute session lifetime, counted from login and never extended.
      ttl: 12h
      # Idle timeout; slides forward on every request, never past ttl. 0 disables it.
      idleTimeout: 1h
      cookieName: __Host-kconmon_session
      # A __Host- cookie with secure=false is rejected at startup.
      secure: true

  # -- 7.3 webhooks -- endpoints themselves are database rows, not values.
  webhooks:
    # Existing Secret with the AES-256-GCM key (base64 of 32 bytes) that
    # encrypts endpoint signing secrets; empty leaves create and test at 503.
    existingSecret: ""
    # Key inside that Secret; the chart-managed path writes this same key.
    existingSecretKey: console-webhooks-encryption-key
    # Chart-managed alternative to existingSecret; values may be ${vault:...} placeholders.
    secret:
      create: false
      # Empty derives the name from the release name.
      name: ""
      annotations: {}
      labels: {}
      # base64 of 32 random bytes (openssl rand -base64 32); required when create is true.
      encryptionKey: ""
    # Alert-transition poll cadence, and the resolution granularity of every
    # alert.resolved delivery; emitted only with webhooks and alerting both on.
    alertPollInterval: 30s

  # -- 7.4 alerting -- Console-managed Prometheus alert rules, reconciled into one PrometheusRule.
  alerting:
    # Off by default: enabling it lets the console write a cluster object.
    enabled: false
    # Namespace the object is applied into; empty means this release's.
    namespace: ""
    # Reconcile cadence, jittered 20%; must be > 0 when enabled.
    syncInterval: 60s
    # Name of the one owned PrometheusRule; changing it orphans the previous.
    bundleName: kconmon-ng-console-rules

  # -- 7.5 optional console features --
  # Advisory-locked tick firing due schedules and reconciling external checks.
  scheduler:
    # Off by default so an upgrade never starts dispatching fleet traffic.
    enabled: false
    # Poll cadence; must be > 0 when enabled.
    tickInterval: 5s
  # Reverse DNS and MaxMind enrichment of stored MTR hop addresses.
  mtr:
    enrichment:
      # Off by default: rdns is an egress footprint.
      enabled: false
      # rdns and geoip gate independently; all sources off is a startup error.
      rdns:
        enabled: false
        # Bounds one lookup, in milliseconds.
        timeoutMs: 500
      # MaxMind GeoLite2 databases, served from the fixed /geoip mount.
      geoip:
        # auto (geoipupdate sidecar) | volume (you supply them) | disabled;
        # empty means volume when a volume or path is set, else disabled.
        mode: ""
        # Editions the sidecar downloads; also what the default paths are derived from.
        editions:
          - GeoLite2-ASN
          - GeoLite2-City
        # geoipupdate's own unit: hours between downloads.
        updateIntervalHours: 24
        # How often the console re-reads changed files; 0 never reloads (restart required).
        reloadInterval: 1h
        image:
          repository: ghcr.io/maxmind/geoipupdate
          tag: v8.0.0
          pullPolicy: IfNotPresent
        resources:
          limits:
            cpu: 100m
            memory: 64Mi
          requests:
            cpu: 10m
            memory: 32Mi
        # Sidecar container securityContext; it writes into the shared /geoip emptyDir.
        securityContext:
          allowPrivilegeEscalation: false
          readOnlyRootFilesystem: false
          capabilities:
            drop:
              - ALL
        # mode=auto: MaxMind credentials (a free GeoLite2 account issues them).
        existingSecret: ""
        # Keys inside that Secret; the chart-managed path writes these same keys.
        accountIdKey: console-maxmind-account-id
        licenseKeyKey: console-maxmind-license-key
        # Chart-managed alternative to existingSecret; values may be ${vault:...} placeholders.
        secret:
          create: false
          # Empty derives the name from the release name.
          name: ""
          annotations: {}
          labels: {}
          # MaxMind account ID; required when create is true.
          accountId: ""
          # MaxMind license key; required when create is true.
          licenseKey: ""
        # mode=volume (airgapped): opaque VolumeSource mounted read-only at /geoip.
        volume: {}
        # Override only if your files are not named after their edition IDs.
        asnPath: ""
        cityPath: ""
      # Cache row lifetime in PostgreSQL; must be > 0 when enabled.
      ttl: 24h
  # The console's own ServiceAccount when serviceAccount.create is false. Its grants (cluster-wide
  # event read, PrometheusRule write) must not land on the account the agent DaemonSet mounts, so the
  # chart refuses to render until this names one.
  serviceAccount:
    name: ""
  # Kubernetes event capture into the Investigate timeline.
  kubernetesContext:
    # Off by default: enabling it adds apiserver egress and an RBAC grant.
    enabled: false
    # The one namespace whose pod events are captured; empty means this release's.
    namespace: ""
    # Periodic relist backstop against a silently wedged watch.
    resyncInterval: 10m
  # Proxies whose X-Forwarded-For gives the client address: the per-address login and OIDC budgets,
  # the anonymous runs and PromQL budgets, the /ws address cap and the audit log. Never used for
  # identity. Behind an Ingress, list the ingress controller's addresses (its pod CIDR at the
  # widest), or every client shares its address and one budget; a pod inside the list can name any
  # client address. Empty falls back to auth.header.trustedProxyCIDRs. Written to the console
  # config only for images 2.5.0 or newer.
  clientAddress:
    trustedProxyCIDRs: []
  # Fixed-window request limits; 0 disables one, a negative value fails startup.
  rateLimit:
    # POST /api/v1/runs per subject per minute.
    runsPerMinute: 10
    # POST /api/v1/auth/login: this value per username, and 20x it per source IP — behind an Ingress
    # every request in the cluster arrives from one address, so a shared budget would let six bogus
    # attempts lock out every user.
    loginPerMinute: 5
    # POST /api/v1/promql/query[_range] per subject per minute; the proxy forwards arbitrary PromQL
    # to your Prometheus, and promql:query belongs to the viewer role.
    promqlPerMinute: 60
  # Open /ws sockets per console replica; 0 disables one cap, a negative value fails startup. A socket
  # over a cap is closed with 1013 and the browser reconnects with backoff. Written to the console
  # config only for images 2.5.0 or newer: an older console refuses the unknown key.
  websocket:
    # Every socket on the replica.
    maxConnections: 1024
    # Sockets from one client address; behind an ingress that is the ingress unless
    # clientAddress.trustedProxyCIDRs (or, while it is empty, auth.header.trustedProxyCIDRs) names it.
    maxConnectionsPerAddress: 256
    # Sockets of one user or API token; anonymous callers are held by the per-address cap.
    maxConnectionsPerSubject: 32

  # -- 7.6 console workload plumbing --
  # fsGroup lets the nonroot console read the root-owned Secret files.
  podSecurityContext:
    runAsNonRoot: true
    runAsUser: 65532
    fsGroup: 65532
    seccompProfile:
      type: RuntimeDefault
  # restricted-PSS compliant.
  securityContext:
    allowPrivilegeEscalation: false
    readOnlyRootFilesystem: true
    capabilities:
      drop:
        - ALL
  # Console NetworkPolicy rules; a host firewall on the destination is a
  # separate layer this chart cannot open.
  networkPolicy:
    # WARNING: empty leaves console ingress open to EVERY pod in the cluster;
    # set peers here (e.g. the ingress-controller namespace) to restrict it.
    # kubectl port-forward is unaffected either way.
    ingressFrom: []
    # Empty renders a default allowing the port from console.prometheus.url to ANY pod AND to
    # 0.0.0.0/0 — a selector cannot match a Prometheus outside the cluster, so the ipBlock is what
    # makes an external one reachable at all. Set this to scope it.
    prometheusEgress: []
    # The pod port behind console.prometheus.url when the URL names a Service whose targetPort
    # differs from its port (the Bitnami Thanos chart: Service 9090, pod 10902). Calico, Cilium and
    # Antrea see the pod port after the Service DNAT, so the default rule above then opens both.
    # 0 opens only the URL's port; ignored when prometheusEgress is set.
    prometheusTargetPort: 0
    # Egress to the Redis-compatible bus. Empty renders a default allowing
    # TCP 6379; override for a TLS endpoint, a sentinel quorum or a managed
    # service. The policy sees the pod port: a Service mapping 6379 to another targetPort needs
    # this set on the pod port.
    redisEgress: []
    # Empty renders a default allowing TCP 5432. As with redisEgress, a Service whose targetPort is
    # not 5432 needs this set on the pod port.
    databaseEgress: []
    # Empty renders TCP 443 to 0.0.0.0/0 (minus networkPolicy.clusterCIDRs) and to any pod: the
    # ipBlock for an IdP outside the cluster, the selector for one behind an in-cluster ingress
    # controller listening on 443. An IdP pod reached through its Service (Keycloak on 8080 or 8443)
    # needs a rule on the pod's own port on every CNI, and a set list replaces the default, e.g.
    #   - to: [{namespaceSelector: {matchLabels: {kubernetes.io/metadata.name: keycloak}}}]
    #     ports: [{protocol: TCP, port: 8443}]
    #   - to: [{ipBlock: {cidr: 0.0.0.0/0}}]
    #     ports: [{protocol: TCP, port: 443}]
    oidcEgress: []
    # Empty renders a default allowing TCP 443/6443 to 0.0.0.0/0, without the clusterCIDRs carve-out:
    # on Calico and Antrea that keeps every pod on 443/6443 open to the console, so name the
    # apiserver endpoint here to close it. On Cilium that ipBlock does not match the apiserver, see
    # networkPolicy.ciliumKubeAPIEgress.
    kubeAPIEgress: []
    # Empty renders TCP 443 to 0.0.0.0/0 (minus networkPolicy.clusterCIDRs) for the geoipupdate
    # sidecar.
    geoipEgress: []
    # Webhook delivery, when a webhook encryption key is configured
    # (console.webhooks.existingSecret, or console.webhooks.secret.create).
    # Empty renders TCP 443/80 to 0.0.0.0/0 minus networkPolicy.clusterCIDRs. On Cilium that ipBlock
    # matches only off-cluster receivers; on Calico and Antrea it matches pod IPs too unless
    # clusterCIDRs carves them out. A receiver running as a pod needs a selector peer on the pod's
    # own port on both (Cilium never matches a pod against an ipBlock, Calico sees the pod port after
    # kube-proxy's DNAT), and a set list replaces the default, so keep an ipBlock rule for external
    # receivers too, e.g.
    #   - to: [{namespaceSelector: {matchLabels: {kubernetes.io/metadata.name: alerting}},
    #           podSelector: {matchLabels: {app: webhook-receiver}}}]
    #     ports: [{protocol: TCP, port: 8080}]
    #   - to: [{ipBlock: {cidr: 0.0.0.0/0}}]
    #     ports: [{protocol: TCP, port: 443}]
    webhookEgress: []
  service:
    type: ClusterIP
    port: 8080
  ingress:
    enabled: false
    className: ""
    annotations: {}
    # [{host: console.example.com, paths: [{path: /, pathType: Prefix}]}]
    hosts: []
    # [{secretName: console-tls, hosts: [console.example.com]}]
    tls: []
  resources:
    limits:
      cpu: 500m
      memory: 256Mi
    requests:
      cpu: 100m
      memory: 128Mi
  nodeSelector: {}
  tolerations: []
  affinity: {}
  # Empty renders nothing.
  priorityClassName: ""
  # Rendered only when console.enabled and replicas > 1.
  pdb:
    enabled: true
    minAvailable: 1

# == 8. Observability == objects for a Prometheus Operator you already run; this chart installs none.

# Ships dashboards/*.json as ConfigMaps for the Grafana sidecar to pick up.
dashboards:
  enabled: false
  # Empty renders into the release namespace; set it where your Grafana sidecar watches.
  namespace: ""
  # Label the Grafana sidecar selects on, and its value.
  label: grafana_dashboard
  labelValue: "1"
  # Grafana folder the sidecar files them under; empty leaves it at the sidecar default.
  folder: kconmon-ng
  extraLabels: {}
  annotations: {}

# ServiceMonitor for /metrics; needs the operator CRDs.
serviceMonitor:
  enabled: false
  interval: 15s

# ScrapeConfig (monitoring.coreos.com/v1alpha1, needs the operator's ScrapeConfig CRD) that
# discovers EXTERNAL agents through the controller's HTTP SD endpoint: the agent ServiceMonitor
# selects pods, and an agent on a bare host is not one. Requires controller.externalGateway, and
# the chart refuses the combination without it, since the endpoint would list nothing.
scrapeConfig:
  externalAgents:
    enabled: false
    # Selector labels the operator's Prometheus requires. kube-prometheus-stack selects only
    # ScrapeConfigs labelled release=<its release name>, e.g. {release: kube-prometheus-stack};
    # an operator with an empty scrapeConfigSelector needs nothing here.
    labels: {}
    # Empty means <agent fullname>-external, i.e. <release>-agent-external for a release named
    # kconmon-ng. Keep "kconmon" in it: dashboards/overview.json filters job=~".*kconmon.*".
    # KconmonExternalAgentDown follows this name whatever it is.
    jobName: ""
    # How often Prometheus re-reads the target list; matches config.controllerAgentTtl.
    refreshInterval: 30s
    # Scrape interval; empty falls back to serviceMonitor.interval.
    interval: ""

# PrometheusRule with the built-in alerts; rule text lives in templates/_rules.tpl, tuning here.
prometheusRule:
  enabled: false
  # UDP loss ratio above threshold on a pair.
  udpLossHigh:
    enabled: true
    # Loss ratio, 0.0-1.0.
    threshold: 0.5
    for: 5m
    severity: warning
  # TCP failure ratio (rate(fail)/rate(all)) above threshold on a pair.
  tcpChecksFailing:
    enabled: true
    # Failure ratio, 0.0-1.0.
    threshold: 0.05
    for: 5m
    severity: warning
  # Full-size datagrams lost without ICMP frag-needed on a pair: an MTU black hole. Reads the pmtu
  # results that agents 2.5.0+ export; on older agents the rule is silently inert. The same knobs
  # drive ZonePathMTUBlackHole, the zone-pair variant that fires only where Prometheus holds no
  # per-pair pmtu series (agent.metrics.detail=zone-only).
  pathMtuBlackHole:
    enabled: true
    # Failure ratio of path MTU probes over 10m, 0.0-1.0.
    threshold: 0.5
    # Failure ratio over 30m, 0.0-1.0, that also pages while a probe failed in the last 10m: a
    # black hole on one of several ECMP paths fails only the probes hashed onto it. 1 turns it off.
    sustainedThreshold: 0.1
    for: 5m
    severity: warning
  # Most peers fail TCP to one node: one alert instead of a pair alert per peer.
  nodeUnreachable:
    enabled: true
    # Share of the node's probing peers that fail to reach it, 0.0-1.0.
    threshold: 0.5
    # Fewer probing peers than this cannot tell a dead node from a dead link; the rule stays quiet.
    minPeers: 2
    for: 5m
    severity: critical
  # One node fails TCP to most of the peers it probes: its own egress is broken.
  nodeIsolated:
    enabled: true
    threshold: 0.5
    minPeers: 2
    for: 5m
    severity: critical
  # A pair probed within the last hour that now reports nothing.
  pairWentSilent:
    enabled: true
    for: 10m
    severity: warning
  # DNS failure ratio above threshold per host/resolver.
  dnsChecksFailing:
    enabled: true
    # Failure ratio, 0.0-1.0.
    threshold: 0.05
    for: 5m
    severity: warning
  # External-target failure ratio above threshold; looser, it leaves the cluster.
  externalChecksFailing:
    enabled: true
    # Failure ratio, 0.0-1.0.
    threshold: 0.1
    for: 5m
    severity: warning
  # Failure ratio across TCP+UDP+ICMP between a zone pair. Reads the zone metric family, which
  # only agents new enough to export it serve — on older agents the rule is silently inert.
  zoneChecksFailing:
    enabled: true
    # Failure ratio, 0.0-1.0.
    threshold: 0.05
    for: 5m
    severity: warning
  # Packet loss between a zone pair, from the sent/received counters. Lower default than the
  # per-pair UDPLossHigh: the zone aggregate dilutes any single link by the pair count N, so one
  # dead link reads about 1/N. Below about 10 node pairs between two zones (up to 8 for a UDP-only
  # break) that alone crosses 0.1 and this fires next to UDPLossHigh; raise it there if you
  # want it to mean the fabric, not one link. zoneChecksFailing has the same arithmetic.
  zoneLossHigh:
    enabled: true
    # Loss ratio, 0.0-1.0.
    threshold: 0.1
    for: 5m
    severity: warning
  # Registered agents below the schedulable-node count; needs leaderElection.
  kconmonAgentsMissing:
    enabled: true
    for: 10m
    severity: warning
  # No controller reports itself leader; needs leaderElection.
  kconmonControllerDown:
    enabled: true
    for: 5m
    severity: critical
  # An external agent the controller's SD endpoint lists but Prometheus cannot scrape: up == 0 on
  # a job whose name contains "agent-external" (the scrapeConfig.externalAgents default, and the
  # plain-Prometheus job in docs/external-agents.md) or on scrapeConfig.externalAgents.jobName.
  # Off by default: the job only exists with one of those.
  externalAgentDown:
    enabled: false
    for: 5m
    severity: warning
  # Extra rules appended verbatim; write metric names with your own prefix.
  additionalRules: []

# == 9. Kubernetes == identity, network and disruption plumbing.

# ServiceAccount for the agent and controller.
serviceAccount:
  # False binds the pods to an existing ServiceAccount instead of creating one.
  create: true
  # Empty generates the name from the release name; required when create is false.
  name: ""

# ClusterRole/Role and their bindings for the resolved ServiceAccount; false means you manage RBAC.
rbac:
  create: true

# NetworkPolicy for agent/controller traffic; enable only if the cluster enforces policies.
networkPolicy:
  enabled: false
  # Pod labels of the scraper itself; empty admits the whole prometheusNamespace on the metrics port.
  prometheusPodLabels: {}
  # Namespace Prometheus scrapes from. REQUIRED for the scrape rule on the agent, controller and
  # console policies, whether the scrape comes from serviceMonitor, a PodMonitor or a plain scrape
  # config; the rule opens config.metricsPort only, never the unauthenticated API port. Unset means
  # no rule, and the scrape fails visibly (up = 0).
  prometheusNamespace: ""
  # Apiserver egress, controller only (the agent has no apiserver client; the console has
  # console.networkPolicy.kubeAPIEgress). Empty renders TCP 443/6443 to 0.0.0.0/0, which
  # clusterCIDRs does not narrow: on Calico and Antrea every pod on 443/6443 stays open to the
  # controller. Name your endpoint here to tighten it.
  kubeAPIEgress: []
  # Cilium: also render a CiliumNetworkPolicy allowing TCP 443/6443 to the kube-apiserver entity for
  # the controller, and for the console when it calls the apiserver. Under Cilium's default
  # policy-cidr-match-mode no ipBlock matches that entity, so kubeAPIEgress and
  # console.networkPolicy.kubeAPIEgress alone leave the controller NotReady and agents unregistered.
  # Node IPs have the same problem (Cilium tags them remote-node and host), so with agent.hostNetwork
  # or controller.externalGateway.enabled a second one admits those entities to the controller's
  # gRPC and gateway ports, which nodeCidrs and externalAgentCidrs cannot do there.
  # auto renders them when the cluster serves cilium.io/v2 CiliumNetworkPolicy (plain `helm template`
  # needs --api-versions cilium.io/v2/CiliumNetworkPolicy to see it); true or false forces it.
  ciliumKubeAPIEgress: auto
  # REQUIRED when the external checker is on, and the other half of
  # config.checkers.external.allowedCidrs: that one lets the agent try, this
  # one lets the packet leave.
  externalEgress: []
  # Source CIDRs of the EXTERNAL agents, opened on the gateway port toward the controller pods.
  # REQUIRED when controller.externalGateway.enabled is on together with this policy: the shared
  # policy default-denies controller ingress it does not name, so an empty list would let the chart
  # render a gateway no packet can reach — and it refuses to do that silently. Mind NAT: through a
  # NodePort or a LoadBalancer with externalTrafficPolicy: Cluster the source the policy sees is the
  # NODE's own IP, so either cover the node CIDR here or use externalTrafficPolicy: Local. On Cilium a
  # node CIDR here matches nothing; networkPolicy.ciliumKubeAPIEgress admits node sources instead.
  # [{cidr: 203.0.113.0/24}] shapes are not accepted — plain CIDR strings.
  externalAgentCidrs: []
  # Node CIDRs, opened on config.grpcPort toward the controller pods. REQUIRED when agent.hostNetwork
  # is on together with this policy: a host-network agent registers from its NODE's IP, and Kubernetes
  # leaves NetworkPolicy behaviour for hostNetwork pods undefined; the common implementation treats
  # their traffic as node traffic, which no podSelector matches, so every registration except the
  # agent on the controller's own node would be dropped while the config looks correct. The chart
  # refuses to render that silently. Plain CIDR strings, as for externalAgentCidrs. On Cilium this
  # ipBlock matches nothing; networkPolicy.ciliumKubeAPIEgress admits the node identities instead.
  nodeCidrs: []
  # Source CIDRs of the EXTERNAL agents' probe traffic, spliced into the agent-to-agent rules in both
  # directions (UDP grpcPort, TCP httpPort, and the ports-less ICMP/MTR rule). The gateway rule covers
  # registration only, so without this an external agent registers fine and every cell between it
  # and the cluster stays red. Never spliced into the gateway rule. Plain CIDR strings.
  externalPeerCidrs: []
  # Empty renders TCP 80/443 to 0.0.0.0/0 when the HTTP checker is on, minus clusterCIDRs. On Cilium
  # that ipBlock matches only off-cluster targets; on Calico and Antrea it matches pod IPs too, so
  # it opens every pod on 80/443 unless clusterCIDRs carves them out. An in-cluster target fails on
  # both unless a selector peer names the pod's own port, not the Service's: Cilium never matches a
  # pod against an ipBlock, and Calico sees the pod port after kube-proxy's DNAT. A set list replaces
  # the default, so keep an ipBlock rule for external targets too, e.g.
  #   - to: [{namespaceSelector: {matchLabels: {kubernetes.io/metadata.name: monitoring}},
  #           podSelector: {matchLabels: {app.kubernetes.io/name: grafana}}}]
  #     ports: [{protocol: TCP, port: 3000}]
  #   - to: [{ipBlock: {cidr: 0.0.0.0/0}}]
  #     ports: [{protocol: TCP, port: 443}]
  httpEgress: []
  # Pod and Service CIDRs (IPv4) carved out of the default 0.0.0.0/0 rules meant for off-cluster
  # peers: httpEgress, console webhookEgress, oidcEgress and geoipEgress, not the kubeAPIEgress
  # defaults. Needed on Calico and Antrea, where that ipBlock also matches pod IPs; Cilium never
  # matches pods against it. Network addresses only (no host bits, no /0). e.g.
  # [10.244.0.0/16, 10.96.0.0/12]
  clusterCIDRs: []
  # Egress to DNS resolvers outside the cluster (config.checkers.dns.resolvers).
  resolverEgress: []
  # Cluster DNS egress for the agent, controller and console policies. Empty renders UDP/TCP 53 to
  # any pod in any namespace. A namespaceSelector never matches a host-network resolver, so with
  # NodeLocal DNSCache list its address, e.g. [{to: [{ipBlock: {cidr: 169.254.20.10/32}},
  # {namespaceSelector: {}}], ports: [{protocol: UDP, port: 53}, {protocol: TCP, port: 53}]}].
  # Set, it replaces the default rule.
  dnsEgress: []