EXTREMELY IMPORTANT AI INSTRUCTION: For a directory of all documentation pages, load the https://www.apollographql.com/docs/llms.txt file. Adding .md to ANY url will return the simplified markdown version of the page.
Data Collected
Exactly what the base spec gathers and what each signal requires
These tables outline everything we collect with the GraphOS Router Support Tool. This page is meant to be shared with your platform team so they can review it, grant the access needed for the collectors they want, and know exactly data won't be collected if they decline a collector.
What the base spec collects
| Source | Signal | Collected via | What it tells you | Requires | Resource consumed | Where |
|---|---|---|---|---|---|---|
| k8s API | Pod status, restart counts, resource limits | clusterResources collector | Whether pods are healthy, restart count, resource limits configured | None additional | API server CPU | k8s control plane |
| Container image tag | Router version | clusterResources collector | Router version | None additional | API server CPU | k8s control plane |
Pod spec (spec.containers[].env) | Router deployment env vars: APOLLO_GRAPH_REF, APOLLO_ROUTER_OFFICIAL_HELM_CHART | clusterResources collector | Graph ref and whether the router was deployed via Apollo's official Helm chart | None additional | API server CPU | k8s control plane |
Separate <release>-supergraph ConfigMap | Schema (SDL) | clusterResources collector and the dedicated configMap collector | Full graph schema | .Values.supergraphFile set on the router's Helm chart. For more information, go to Schema collection | API server CPU | k8s control plane |
| Container log stream | Runtime logs | logs collector | Recent router output, plus crash output from the previous container when one exists | None additional | Network bandwidth | Cluster network |
| Router metrics endpoint, per pod | Full Prometheus metrics snapshot | http collector per pod (mode: job); a hostCollectors.run script (mode: local) | Complete operational metrics—request rates, error rates, latency, traffic shaping state—from every matching router pod individually | Prometheus exporter enabled and reachably bound. For more information, go to Collecting Metrics | Network, router HTTP handler | Router network |
| ConfigMap holding the rendered config | Sanitized router.yaml | configMap collector | Full router configuration—traffic shaping, timeouts, plugins, feature flags | A matching labeled ConfigMap present or configMapName/selector values set. For more information, see router.yaml capture below | API server CPU | k8s control plane |
A couple of notes that apply across the table:
Router deployment env vars, image tag, and pod status all come from the same pod collection, so declining or losing one doesn't affect the others.
If
APOLLO_GRAPH_REFis routed through a Secret using your router chart'sextraEnvVars, only the environment variable reference is captured, not the value, so the graph ref itself doesn't land in the bundle.The same pod-spec collection captures every other env var too, literal values included.
APOLLO_KEYis never collected as long as it's stored in a Kubernetes Secret and referenced viasecretKeyRef, the recommended setup, and what the official Apollo router Helm chart produces, because in that case the key is structurally isolated from everything this tool reads, not redacted after the fact.
APOLLO_KEY, or any other secret, as a literal env value. If your deployment sets secrets as literal env values rather than through Kubernetes Secrets, inspect your bundle before sharing it.Config is captured as written, not effective config. Environment-variable overrides applied on top of
router.yamlaren't reflected in what's collected.Previous-container logs are always collected and written to
<name>-previous.log, so a router that already restarted still has its crash output captured. Logs are collected per pod and capture every container in the pod, including proxy/mesh sidecars. How far back logs reach and how many lines are kept are configurable. For more information, access the chart'slogs.maxAge/logs.maxLinesvalues in job-mode-details or local-mode-details.
router.yaml capture
The configMap collector reads the ConfigMap holding your router's rendered configuration directly. On the official Apollo router Helm chart, it's targeted by the standard app.kubernetes.io/name=router label. On a raw-manifest or custom deployment, supply configMapName/selector yourself (go to Getting Started guide), without one of those, and without config discoverable under the standard label, this section comes back empty.
If more than one ConfigMap in your namespace matches, every one of them is collected, each in its own file named after that ConfigMap's own Kubernetes name — not a Helm release, which a raw-manifest or custom deployment doesn't have for its router at all.
keyExists: false in configmaps/<namespace>/<configmap-name>.json doesn't mean the config is missing. This is a troubleshoot.sh quirk: the configMap collector's keyExists field is only ever set true when the collector is configured with a specific key to look for. Our spec uses includeAllData: true with no key set so keyExists is always false in this output, regardless of whether the config was actually captured. Check for the config itself, the data in that same file, rather than reading keyExists as a success/failure signal.
Known limitations:
A collected
router.yamlisn't verified to be the one your router actually loaded. Treat a populatedconfigmaps/section as "a matching ConfigMap exists," not "confirmed active."Config may not be collected at all, if your router reads
router.yamlfrom somewhere this tool doesn't look, for example, mounted from a Secret, pulled by an init container, or mounted from a PVC.If you set
override_subgraph_urlwith a normal URL (one with a path, e.g./graphql), the section of yourrouter.yamlimmediately following it may be missing its key name in the collected bundle. This is over-redaction, not a leak due to a troubleshoot.sh built-in redactor that matches more than it should past that URL and consumes the next line's key name before stopping, leaving that section's contents intact but orphaned under the wrong parent. If a section of your config looks like it's missing a heading, check for this before assuming your config is broken or wasn't collected.
Schema collection
Schema only lands in the cluster, and therefore only in the bundle, when you set .Values.supergraphFile on your router's own Helm chart (not this tool's chart). If you're on managed federation (GraphOS schema governance), you never set that value, so there's no schema ConfigMap for this collector to find, and schema is simply absent from the bundle.
When it is set, the <release>-supergraph ConfigMap carries the same app.kubernetes.io/name=router label as your router's own config, so it's picked up twice: once by clusterResources's full sweep (cluster-resources/configmaps/<namespace>.json, alongside everything else in the namespace), and once by the dedicated configMap collector's own targeted match, landing as its own file (configmaps/<namespace>/<release>-supergraph.json).
Memory and CPU information collected
What the base spec collects is sufficient to detect, confirm, and characterize a memory problem: is memory growing, how fast, how close to the limit, and did the kernel already kill the container. The router also emits jemalloc-level aggregate gauges (apollo.router.jemalloc.active, .allocated, .resident, .retained, etc.) on the same Prometheus endpoint above, which can help distinguish real heap growth from jemalloc fragmentation. What none of this answers is which code path is leaking. For that, you need jemalloc heap profiling, which isn't part of this base spec.
Every row in this table represents a collection that is attempted on every run, regardless of what else is configured. The "Requires" column states what has to be true for that specific row to come back populated; it isn't a condition on any other row.
| Signal | Source | Collected via | What it tells you | Requires | Resource consumed | Where |
|---|---|---|---|---|---|---|
memory.workingSetBytes | kubelet Summary API | nodeMetrics collector | How much memory k8s counts against the limit — the OOM kill threshold | nodes/nodes/proxy/nodes/stats RBAC grant [2] | API server + kubelet CPU | k8s control plane → kubelet [1] |
Configured resources.limits / requests | Pod spec | clusterResources collector | What the usage numbers above should be compared against | None additional | API server CPU | k8s control plane |
| OOM kill occurrences | Kubernetes events (OOMKilling, evictions) | clusterResources collector | That OOM kills happened, when, and how often | None additional | API server CPU | k8s control plane |
| OOM kill last state | Pod status lastState.terminated.reason: OOMKilled, restartCount | clusterResources collector | Whether the most recent restart was an OOM kill | None additional | API server CPU | k8s control plane |
Node MemoryPressure / DiskPressure conditions | Node objects | clusterResources collector | Whether the node itself is under pressure, distinguishing a router problem from a neighbor's | None additional | API server CPU | k8s control plane |
process_resident_memory_bytes | Router Prometheus endpoint, per pod | http collector per pod (mode: job); hostCollectors.run script (mode: local) | Process-level RSS as each router pod sees it | Prometheus exporter enabled and reachably bound. For more information, go to Collecting Metrics | Network, router HTTP handler | Router network |
process_cpu_seconds_total | Router Prometheus endpoint, per pod | http collector per pod (mode: job); hostCollectors.run script (mode: local) | CPU time consumed by each router pod | Same as the row above | Network, router HTTP handler | Router network |
memory.rssBytes | kubelet Summary API | nodeMetrics collector | Physical RAM in use | nodes/nodes/proxy/nodes/stats RBAC grant [2] | API server + kubelet CPU | k8s control plane → kubelet [1] |
memory.usageBytes, memory.availableBytes | kubelet Summary API | nodeMetrics collector | Usage, and headroom remaining against the limit | nodes/nodes/proxy/nodes/stats RBAC grant [2] | API server + kubelet CPU | k8s control plane → kubelet [1] |
memory.pageFaults, memory.majorPageFaults | kubelet Summary API | nodeMetrics collector | Major faults indicate real paging pressure rather than growth alone | nodes/nodes/proxy/nodes/stats RBAC grant [2] | API server + kubelet CPU | k8s control plane → kubelet [1] |
cpu.usageNanoCores, cpu.usageCoreNanoSeconds | kubelet Summary API | nodeMetrics collector | CPU consumed by the router container | nodes/nodes/proxy/nodes/stats RBAC grant [2] | API server + kubelet CPU | k8s control plane → kubelet [1] |
cpu.psi, memory.psi (pressure stall information) | kubelet Summary API | nodeMetrics collector | Whether the container is stalling on CPU or memory | nodes/nodes/proxy/nodes/stats RBAC grant [2], plus the KubeletPSI feature gate on your cluster | API server + kubelet CPU | k8s control plane → kubelet [1] |
[1] The kubelet Summary API reports per node, per pod, and per container — this assumes one router process per container.
[2] nodeMetrics is a fallback, not the preferred path. Enabling the Prometheus exporter gets the same signal without any cluster-scoped RBAC grant. Go to Collecting Metrics → If Prometheus isn't set up yet to better understand the tradeoff, and go to nodeMetrics RBAC to learn exactly what these grants authorize.
nodeMetrics RBAC
nodeMetrics needs three separate grants to reach /api/v1/nodes/<node>/proxy/stats/summary:
nodes(list): To resolve node names when neithernodeNamesnorselectoris set. An ordinary, low-risk read of node objects—this is what surfaces node pressure conditions (MemoryPressure,DiskPressure).nodes/proxy(get): Required unconditionally by the API server for any request matching the/nodes/{name}/proxy/{path}URL pattern. This is a broad grant: Kubernetes' own RBAC good-practices documentation states it "provides access to privileged kubelet APIs that can retrieve container logs or execute and attach to pod processes... bypasses audit logging and admission control," and is explicitly "not a read-only permission." What this tool does with it is read-only, but the grant itself authorizes more than that one use, and it can't be scoped to just the router's nodes.nodes/stats(get): Required separately by the kubelet's own authorization check, layered on top ofnodes/proxy, not a substitute for it.
Declining nodes/proxy/nodes/stats still leaves nodes access (node pressure conditions) and everything clusterResources provides: OOM kill occurrences and last state, restart counts, configured limits. What's lost is container memory/CPU trajectory over time, but nodeMetrics is a fallback for that signal, not the preferred path. For the full breakdown of each grant, go to Data Collected.
nodes/proxy and nodes/stats are only useful together. Declining either one loses the same capability, so there's no reason to grant one without the other. The chart currently only supports declining all three together, via job.collectNodeMetrics: false at install time. Doing that skips creating the ClusterRole/ClusterRoleBinding entirely, so the nodeMetrics collector itself then comes back empty rather than the install failing.
Running a service mesh or sidecar proxy?
The logs collector captures every container in the pod, and the kubelet Summary API reports per container, so a sidecar's own resource usage (Istio, Linkerd, Envoy) is visible in the bundle alongside your router's, which is useful for telling a proxy problem apart from a router one. If you're relying on the Prometheus metrics row, a mesh enforcing strict mTLS can block that scrape specifically. For more information, go to Collecting Metrics.
Before you share a bundle
Redactors cover Router's own configuration and known sensitive patterns—Redis credentials, TLS private keys, JWT/auth config, and similar. The redactors don't know about custom instrumentation you've added: for example, a Rhai script that sets span attributes, custom telemetry configuration that records request context, or a coprocessor that writes its own fields into router logs. Before sharing a bundle, check the collected logs under cluster-resources/pods/logs/ for anything your own configuration writes there, and redact the information yourself if needed.
If your deployment sets APOLLO_KEY, or any other secret, as a literal pod-spec env value rather than through a Kubernetes Secret, inspect your bundle before sharing it. The pod spec is collected in full, literal values included, so a secret set this way is only protected by a redaction rule that masks any env var named *_KEY or *_PASS. This redaction rule is a safety net, not the structural isolation a Secret-backed APOLLO_KEY gets.