Data Collected

Exactly what the base spec gathers and what each signal requires


These tables outline everything we collect with the GraphOS Router Support Tool. This page is meant to be shared with your platform team so they can review it, grant the access needed for the collectors they want, and know exactly data won't be collected if they decline a collector.

What the base spec collects

SourceSignalCollected viaWhat it tells youRequiresResource consumedWhere
k8s APIPod status, restart counts, resource limitsclusterResources collectorWhether pods are healthy, restart count, resource limits configuredNone additionalAPI server CPUk8s control plane
Container image tagRouter versionclusterResources collectorRouter versionNone additionalAPI server CPUk8s control plane
Pod spec (spec.containers[].env)Router deployment env vars: APOLLO_GRAPH_REF, APOLLO_ROUTER_OFFICIAL_HELM_CHARTclusterResources collectorGraph ref and whether the router was deployed via Apollo's official Helm chartNone additionalAPI server CPUk8s control plane
Separate <release>-supergraph ConfigMapSchema (SDL)clusterResources collector and the dedicated configMap collectorFull graph schema.Values.supergraphFile set on the router's Helm chart. For more information, go to Schema collectionAPI server CPUk8s control plane
Container log streamRuntime logslogs collectorRecent router output, plus crash output from the previous container when one existsNone additionalNetwork bandwidthCluster network
Router metrics endpoint, per podFull Prometheus metrics snapshothttp collector per pod (mode: job); a hostCollectors.run script (mode: local)Complete operational metrics—request rates, error rates, latency, traffic shaping state—from every matching router pod individuallyPrometheus exporter enabled and reachably bound. For more information, go to Collecting MetricsNetwork, router HTTP handlerRouter network
ConfigMap holding the rendered configSanitized router.yamlconfigMap collectorFull router configuration—traffic shaping, timeouts, plugins, feature flagsA matching labeled ConfigMap present or configMapName/selector values set. For more information, see router.yaml capture belowAPI server CPUk8s control plane

A couple of notes that apply across the table:

  • Router deployment env vars, image tag, and pod status all come from the same pod collection, so declining or losing one doesn't affect the others.

  • If APOLLO_GRAPH_REF is routed through a Secret using your router chart's extraEnvVars, only the environment variable reference is captured, not the value, so the graph ref itself doesn't land in the bundle.

  • The same pod-spec collection captures every other env var too, literal values included. APOLLO_KEY is never collected as long as it's stored in a Kubernetes Secret and referenced via secretKeyRef, the recommended setup, and what the official Apollo router Helm chart produces, because in that case the key is structurally isolated from everything this tool reads, not redacted after the fact.

caution
That guarantee doesn't extend to a deployment that sets APOLLO_KEY, or any other secret, as a literal env value. If your deployment sets secrets as literal env values rather than through Kubernetes Secrets, inspect your bundle before sharing it.
  • Config is captured as written, not effective config. Environment-variable overrides applied on top of router.yaml aren't reflected in what's collected.

  • Previous-container logs are always collected and written to <name>-previous.log, so a router that already restarted still has its crash output captured. Logs are collected per pod and capture every container in the pod, including proxy/mesh sidecars. How far back logs reach and how many lines are kept are configurable. For more information, access the chart's logs.maxAge/logs.maxLines values in job-mode-details or local-mode-details.

router.yaml capture

The configMap collector reads the ConfigMap holding your router's rendered configuration directly. On the official Apollo router Helm chart, it's targeted by the standard app.kubernetes.io/name=router label. On a raw-manifest or custom deployment, supply configMapName/selector yourself (go to Getting Started guide), without one of those, and without config discoverable under the standard label, this section comes back empty.

If more than one ConfigMap in your namespace matches, every one of them is collected, each in its own file named after that ConfigMap's own Kubernetes name — not a Helm release, which a raw-manifest or custom deployment doesn't have for its router at all.

keyExists: false in configmaps/<namespace>/<configmap-name>.json doesn't mean the config is missing. This is a troubleshoot.sh quirk: the configMap collector's keyExists field is only ever set true when the collector is configured with a specific key to look for. Our spec uses includeAllData: true with no key set so keyExists is always false in this output, regardless of whether the config was actually captured. Check for the config itself, the data in that same file, rather than reading keyExists as a success/failure signal.

Known limitations:

  • A collected router.yaml isn't verified to be the one your router actually loaded. Treat a populated configmaps/ section as "a matching ConfigMap exists," not "confirmed active."

  • Config may not be collected at all, if your router reads router.yaml from somewhere this tool doesn't look, for example, mounted from a Secret, pulled by an init container, or mounted from a PVC.

  • If you set override_subgraph_url with a normal URL (one with a path, e.g. /graphql), the section of your router.yaml immediately following it may be missing its key name in the collected bundle. This is over-redaction, not a leak due to a troubleshoot.sh built-in redactor that matches more than it should past that URL and consumes the next line's key name before stopping, leaving that section's contents intact but orphaned under the wrong parent. If a section of your config looks like it's missing a heading, check for this before assuming your config is broken or wasn't collected.

Schema collection

Schema only lands in the cluster, and therefore only in the bundle, when you set .Values.supergraphFile on your router's own Helm chart (not this tool's chart). If you're on managed federation (GraphOS schema governance), you never set that value, so there's no schema ConfigMap for this collector to find, and schema is simply absent from the bundle.

When it is set, the <release>-supergraph ConfigMap carries the same app.kubernetes.io/name=router label as your router's own config, so it's picked up twice: once by clusterResources's full sweep (cluster-resources/configmaps/<namespace>.json, alongside everything else in the namespace), and once by the dedicated configMap collector's own targeted match, landing as its own file (configmaps/<namespace>/<release>-supergraph.json).

Memory and CPU information collected

What the base spec collects is sufficient to detect, confirm, and characterize a memory problem: is memory growing, how fast, how close to the limit, and did the kernel already kill the container. The router also emits jemalloc-level aggregate gauges (apollo.router.jemalloc.active, .allocated, .resident, .retained, etc.) on the same Prometheus endpoint above, which can help distinguish real heap growth from jemalloc fragmentation. What none of this answers is which code path is leaking. For that, you need jemalloc heap profiling, which isn't part of this base spec.

Every row in this table represents a collection that is attempted on every run, regardless of what else is configured. The "Requires" column states what has to be true for that specific row to come back populated; it isn't a condition on any other row.

SignalSourceCollected viaWhat it tells youRequiresResource consumedWhere
memory.workingSetByteskubelet Summary APInodeMetrics collectorHow much memory k8s counts against the limit — the OOM kill thresholdnodes/nodes/proxy/nodes/stats RBAC grant [2]API server + kubelet CPUk8s control plane → kubelet [1]
Configured resources.limits / requestsPod specclusterResources collectorWhat the usage numbers above should be compared againstNone additionalAPI server CPUk8s control plane
OOM kill occurrencesKubernetes events (OOMKilling, evictions)clusterResources collectorThat OOM kills happened, when, and how oftenNone additionalAPI server CPUk8s control plane
OOM kill last statePod status lastState.terminated.reason: OOMKilled, restartCountclusterResources collectorWhether the most recent restart was an OOM killNone additionalAPI server CPUk8s control plane
Node MemoryPressure / DiskPressure conditionsNode objectsclusterResources collectorWhether the node itself is under pressure, distinguishing a router problem from a neighbor'sNone additionalAPI server CPUk8s control plane
process_resident_memory_bytesRouter Prometheus endpoint, per podhttp collector per pod (mode: job); hostCollectors.run script (mode: local)Process-level RSS as each router pod sees itPrometheus exporter enabled and reachably bound. For more information, go to Collecting MetricsNetwork, router HTTP handlerRouter network
process_cpu_seconds_totalRouter Prometheus endpoint, per podhttp collector per pod (mode: job); hostCollectors.run script (mode: local)CPU time consumed by each router podSame as the row aboveNetwork, router HTTP handlerRouter network
memory.rssByteskubelet Summary APInodeMetrics collectorPhysical RAM in usenodes/nodes/proxy/nodes/stats RBAC grant [2]API server + kubelet CPUk8s control plane → kubelet [1]
memory.usageBytes, memory.availableByteskubelet Summary APInodeMetrics collectorUsage, and headroom remaining against the limitnodes/nodes/proxy/nodes/stats RBAC grant [2]API server + kubelet CPUk8s control plane → kubelet [1]
memory.pageFaults, memory.majorPageFaultskubelet Summary APInodeMetrics collectorMajor faults indicate real paging pressure rather than growth alonenodes/nodes/proxy/nodes/stats RBAC grant [2]API server + kubelet CPUk8s control plane → kubelet [1]
cpu.usageNanoCores, cpu.usageCoreNanoSecondskubelet Summary APInodeMetrics collectorCPU consumed by the router containernodes/nodes/proxy/nodes/stats RBAC grant [2]API server + kubelet CPUk8s control plane → kubelet [1]
cpu.psi, memory.psi (pressure stall information)kubelet Summary APInodeMetrics collectorWhether the container is stalling on CPU or memorynodes/nodes/proxy/nodes/stats RBAC grant [2], plus the KubeletPSI feature gate on your clusterAPI server + kubelet CPUk8s control plane → kubelet [1]

[1] The kubelet Summary API reports per node, per pod, and per container — this assumes one router process per container.

[2] nodeMetrics is a fallback, not the preferred path. Enabling the Prometheus exporter gets the same signal without any cluster-scoped RBAC grant. Go to Collecting Metrics → If Prometheus isn't set up yet to better understand the tradeoff, and go to nodeMetrics RBAC to learn exactly what these grants authorize.

nodeMetrics RBAC

nodeMetrics needs three separate grants to reach /api/v1/nodes/<node>/proxy/stats/summary:

  • nodes (list): To resolve node names when neither nodeNames nor selector is set. An ordinary, low-risk read of node objects—this is what surfaces node pressure conditions (MemoryPressure, DiskPressure).

  • nodes/proxy (get): Required unconditionally by the API server for any request matching the /nodes/{name}/proxy/{path} URL pattern. This is a broad grant: Kubernetes' own RBAC good-practices documentation states it "provides access to privileged kubelet APIs that can retrieve container logs or execute and attach to pod processes... bypasses audit logging and admission control," and is explicitly "not a read-only permission." What this tool does with it is read-only, but the grant itself authorizes more than that one use, and it can't be scoped to just the router's nodes.

  • nodes/stats (get): Required separately by the kubelet's own authorization check, layered on top of nodes/proxy, not a substitute for it.

Declining nodes/proxy/nodes/stats still leaves nodes access (node pressure conditions) and everything clusterResources provides: OOM kill occurrences and last state, restart counts, configured limits. What's lost is container memory/CPU trajectory over time, but nodeMetrics is a fallback for that signal, not the preferred path. For the full breakdown of each grant, go to Data Collected.

nodes/proxy and nodes/stats are only useful together. Declining either one loses the same capability, so there's no reason to grant one without the other. The chart currently only supports declining all three together, via job.collectNodeMetrics: false at install time. Doing that skips creating the ClusterRole/ClusterRoleBinding entirely, so the nodeMetrics collector itself then comes back empty rather than the install failing.

Running a service mesh or sidecar proxy?

The logs collector captures every container in the pod, and the kubelet Summary API reports per container, so a sidecar's own resource usage (Istio, Linkerd, Envoy) is visible in the bundle alongside your router's, which is useful for telling a proxy problem apart from a router one. If you're relying on the Prometheus metrics row, a mesh enforcing strict mTLS can block that scrape specifically. For more information, go to Collecting Metrics.

Before you share a bundle

caution
Every file in the bundle is redacted automatically, but redaction can fail on a single file without failing the whole run. If a bundle reports a redaction error, don't share it until you understand what it means. An error might mean redaction failed to complete on a file.

Redactors cover Router's own configuration and known sensitive patterns—Redis credentials, TLS private keys, JWT/auth config, and similar. The redactors don't know about custom instrumentation you've added: for example, a Rhai script that sets span attributes, custom telemetry configuration that records request context, or a coprocessor that writes its own fields into router logs. Before sharing a bundle, check the collected logs under cluster-resources/pods/logs/ for anything your own configuration writes there, and redact the information yourself if needed.

If your deployment sets APOLLO_KEY, or any other secret, as a literal pod-spec env value rather than through a Kubernetes Secret, inspect your bundle before sharing it. The pod spec is collected in full, literal values included, so a secret set this way is only protected by a redaction rule that masks any env var named *_KEY or *_PASS. This redaction rule is a safety net, not the structural isolation a Secret-backed APOLLO_KEY gets.