Health Checks


Health checks

Apollo MCP Server provides health check endpoints for monitoring server health and readiness. This feature is useful for load balancers, container orchestrators, and monitoring systems.

Configuration

Health checks are only available when using the streamable_http transport and must be explicitly enabled:

YAML
Example health check configuration
1transport:
2  type: streamable_http
3  address: 127.0.0.1
4  port: 8000
5health_check:
6  enabled: true
7  path: /health
8  readiness:
9    allowed: 50
10    interval:
11      sampling: 10s
12      unready: 30s

Serving on a separate port

By default, the health check is served on the same address and port as the rest of the server. Set health_check.listen to serve it on its own address/port instead — useful when your infrastructure (like a Kubernetes readiness probe) expects to reach health checks on a dedicated port, for example one that isn't exposed through your ingress.

YAML
Example health check on a separate port
1transport:
2  type: streamable_http
3  address: 127.0.0.1
4  port: 8000
5health_check:
6  enabled: true
7  listen: 0.0.0.0:8088
note
Neither the merged nor the separate health check listener applies the Host/Origin validation that protects the main MCP endpoint — the endpoint only ever reports UP/DOWN status, so this isn't a concern by itself. Still, bind the separate listener to an address reachable only by your infrastructure (for example, a Kubernetes pod IP used by kubelet probes) rather than a publicly routable one. 0.0.0.0 is fine here if the pod itself isn't publicly exposed.

Endpoints

The health check provides different responses based on query parameters:

EndpointDescriptionResponse
GET /healthBasic health checkAlways returns {"status": "UP"}
GET /health?liveLiveness checkReturns {"status": "UP"} if server is alive
GET /health?readyReadiness checkReturns {"status": "UP"} if server is ready to handle requests

Probes

The server tracks failed requests and automatically marks itself as unready if too many failures occur within a sampling interval:

  • Sampling interval: How often the server checks the rejection count (default: 5 seconds)

  • Allowed rejections: Maximum failures allowed before becoming unready (default: 100)

  • Recovery time: How long to wait before attempting to recover (default: 2x sampling interval)

When the server becomes unready:

  • The /health?ready endpoint returns HTTP 503 with {"status": "DOWN"}

  • After the recovery period, the rejection counter resets and the server becomes ready again

This allows external systems to automatically route traffic away from unhealthy servers and back when they recover.