mode: job Reference

Full chart values, storage configuration, and RBAC for mode: job


This page is the complete reference for mode: job.

The Job image

The image bundles a pinned version of troubleshoot's support-bundle binary as well as aws-cli and gcloud for bundle upload. For more information, go to Support bundle storage and retrieval.

Chart values

Chart values common to both modes

Every one of these values applies regardless of which mode you run.

ValueRequiredDefaultApplies toPurpose
namespaceYes—Official chart, raw manifest / customYour router's namespace. Rendered into every collector that accepts it, including clusterResources.namespaces as a single-element list. Never rely on the kubectl context's default namespace; it isn't used.
modeYes—Official chart, raw manifest / customSelects how collection runs: local runs on your own machine using your own kubectl credentials and job runs in-cluster as a Kubernetes Job using a ServiceAccount the chart creates.
selectorRaw-manifest tier onlyThe official chart's app.kubernetes.io/name=router labelRaw manifests / customPod label selector, for example, app=my-router. Also what the metrics collector resolves its target from. For more information, go to Collecting Metrics.
configMapNameRaw-manifest tier onlyLabel-based targeting, no name neededRaw manifests / customName of the ConfigMap holding your router's rendered config
metricsPortNo9090Official chart, raw manifest / customPort your router's metrics endpoint is reachable on, if not 9090. For more information, go to Collecting Metrics.
logs.maxAgeNoUnset—no age capOfficial chart, raw manifest / customMaps to the logs collector's limits.maxAge. How far back to reach, for example, 2h
logs.maxLinesNo10000 (troubleshoot.sh default)Official chart, raw manifest / customMaps to limits.maxLines. Lines kept, newest per container

Go to Local Mode / Job Mode to determine what each mode needs beyond this shared set.

Chart values specific to mode: job

ValueRequiredDefaultPurpose
job.image.repositoryNoapollograph/router-diagnosticsThe Job's container image, published by Apollo. Override only to pin an older version or point at a private mirror.
job.image.tagNoMatches this chart's release versionThe troubleshoot.sh version it bundles is a separate, independently-pinned detail, recorded in a collected bundle's meta.json. Override only to pin an older release or point at a private mirror.
job.image.pullPolicyNoIfNotPresentStandard Kubernetes image pull policy
job.collectNodeMetricsNotrueSet false to skip creating the cluster-scoped ClusterRole/ClusterRoleBinding entirely and decline nodeMetrics data up front. For more information, go to Collecting Metrics → If Prometheus isn't set up yet
job.podAnnotationsNo{sidecar.istio.io/inject: "false", linkerd.io/inject: disabled}Annotations merged onto the Job's pod template—customer values win over these defaults. For more information, go to Job Mode → Service mesh environments
job.podLabelsNo{}Arbitrary labels merged onto the Job's pod template
job.ttlSecondsAfterFinishedNo3600Maps to the Job's ttlSecondsAfterFinished—how long after completion Kubernetes automatically deletes the Job and its pod
job.serviceAccount.annotationsOnly if using IRSA/Workload Identity{}Merged onto the Job's ServiceAccount. Carries eks.amazonaws.com/role-arn (IRSA) or iam.gke.io/gcp-service-account (Workload Identity), whichever matches job.storage.provider. For more information, go to Support bundle storage and retrieval → Credentials

Bundle Retrieval

Storage is required for mode: job because there is no way to get the bundle out otherwise. The chart fails the install at render time if mode: job is set and job.storage.provider is left unset. Go to Support bundle storage and retrieval for the full job.storage reference, worked examples per provider, and how to retrieve the bundle from each destination.

Kubernetes RBAC

Installing mode: job requires creating a Job, a namespace-scoped ServiceAccount, a Role/RoleBinding, and—because container memory/CPU come from the kubelet—cluster-scoped RBAC. Creating cluster-scoped RBAC is a broader capability than installing the router needs, so whoever installs the chart in mode: job needs more access than someone who could simply run mode: local themselves.

Collection itself runs using a namespace-scoped ServiceAccount the chart creates. All namespace-scoped access is read-only (get/list) and covers everything clusterResources/logs/configMap touch:

  • Core: configmaps, endpoints, events, limitranges, persistentvolumeclaims, pods, pods/log, resourcequotas, serviceaccounts, services

  • apps: daemonsets, deployments, replicasets, statefulsets

  • batch: cronjobs, jobs

  • rbac.authorization.k8s.io: roles, rolebindings

  • policy: poddisruptionbudgets

  • networking.k8s.io: ingresses, networkpolicies

  • discovery.k8s.io: endpointslices

  • coordination.k8s.io: leases

On top of this, three separate cluster-scoped grants apply for nodeMetrics, deliberately kept independent so you can decline the more sensitive ones without losing the others. Go to Collecting Metrics → If Prometheus isn't set up yet for exactly what those three grants are and what each authorizes.

pods/exec is not required: nothing this tool does runs inside your router container.

clusterResources also attempts several cluster-scoped resource types beyond what's listed above (nodes, storage classes, CustomResourceDefinitions, webhook configurations, and more). A <resource>-errors.json file lands alongside its corresponding <resource>.json containing the Kubernetes API error. cluster-resources/auth-cani-list/ is troubleshoot.sh's own self-reported record of every verb/resource/apiGroup combination the collecting identity had at collection time.

Cluster footprint

Spec ConfigMap, Helm release metadata, and the ServiceAccount/Role/RoleBinding stays behind after installing. The Job itself is transient: ttlSecondsAfterFinished (default 3600) means Kubernetes automatically deletes the Job and its pod an hour after it completes. To remove everything else, including the RBAC:

Bash
1helm uninstall router-diagnostics --namespace <namespace>

Bundle output shape

Extracting support-bundle-<timestamp>.tar.gz yields one top-level directory. version.yaml, analysis.json, cluster-info/, and execution-data/summary.txt always land in it, written unconditionally by the engine itself regardless of what the spec declares; meta.json is ours. execution-data/summary.txt is worth a look if something seems off — it's a human-readable per-collector success/failure/timing report.

Worked example, for a mode: job install in the production namespace:

Text
1support-bundle-2026-08-11T14_23_00/
2├── version.yaml
3├── analysis.json
4├── meta.json
5├── cluster-info/
6│   └── cluster_version.json
7├── execution-data/
8│   └── summary.txt
9├── router-metrics-<pod-name>/
10│   └── result.json
11├── router-logs/
12│   └── <router-pod-name>/
13│       └── router.log -> ../../cluster-resources/pods/logs/production/<router-pod-name>/router.log
14├── cluster-resources/
15│   ├── pods/
16│   │   ├── production.json
17│   │   └── logs/production/<router-pod-name>/router.log         # the real file behind the symlink above
18│   ├── configmaps/production.json
19│   ├── events/production.json
20│   ├── nodes.json
21│   └── ...
22├── node-metrics/
23│   └── <node-name>.json
24└── configmaps/
25    └── production/
26        └── <configmap-name>.json
  • router-metrics-<pod-name>/ — one directory per router pod matching selector, each from its own http collector. Go to Collecting Metrics for how pods are targeted.

  • The file under configmaps/production/ is named after the matched ConfigMap's own Kubernetes name, on the official chart these happen to coincide, since that chart names the ConfigMap after the release, but a raw-manifest or custom deployment has no Helm release of the router at all. If more than one ConfigMap matched the label selector in production, configmaps/production/ becomes multiple files, one per ConfigMap:

    Text
    1configmaps/
    2└── production/
    3    ├── <configmap-1-name>.json
    4    └── <configmap-2-name>.json
  • cluster-resources/configmaps/production.json is clusterResources's full, unfiltered sweep of every ConfigMap in the namespace. configmaps/production/<configmap-name>.json is the dedicated configMap collector's own output.

note
keyExists: false in configmaps/<namespace>/<configmap-name>.json doesn't mean the config is missing. This is a troubleshoot.sh quirk: the configMap collector's keyExists field is only ever set true when the collector is configured with a specific key to look for. Our spec uses includeAllData: true with no key set, so keyExists is always false in this output, regardless of whether the config was actually captured. Check for the config data in that same file, rather than reading keyExists as a success/failure signal.