grpc/grpc-go · warning
connection active but received health check RPC error
Error message
connection active but received health check RPC error: %v
What it means
The gRPC client health-check watcher received an RPC error (other than UNIMPLEMENTED) on an open connection. Per the grpc health protocol, this sets the subchannel to TransientFailure and triggers backoff/retry; the wrapped %v is the stream RecvMsg error. This is the health-check state machine reporting a server-side Watch failure, not necessarily a transport failure.
Solutions
- Inspect the inner %v: CANCELLED/UNAVAILABLE usually means server restart; let the client's built-in backoff retry.
- If PERMISSION_DENIED, fix the auth policy on the server's Health service.
- If the stream resets repeatedly, check server logs for panics in the Health implementation and ensure the service is registered with the gRPC server.
- Stabilize long-lived streams: raise HTTP/2 keepalive timeouts on intermediaries, or use keepalive.EnforcementPolicy on the server.
- Confirm the server implements grpc.health.v1.Health (UNIMPLEMENTED is handled separately and is non-fatal).
Example fix
// before: server registers no Health service and Watch returns INTERNAL on a bad handler
srv := grpc.NewServer()
// after
import healthpb "google.golang.org/grpc/health"
import healthsvc "google.golang.org/grpc/health/grpc_health_v1"
hs := healthpb.NewServer()
hs.SetServingStatus("", healthpb.HealthCheckResponse_SERVING)
healthsvc.RegisterHealthServer(srv, hs) Defensive patterns
Strategy: retry
Validate before calling
// Health is built into grpc-go's connection; pre-flight by checking the server implements Health:
import healthpb "google.golang.org/grpc/health/grpc_health_v1"
hc := healthpb.NewHealthClient(conn)
resp, err := hc.Check(ctx, &healthpb.HealthCheckRequest{Service: svc})
if err != nil { return fmt.Errorf("server Health not usable: %w", err) }
if resp.Status != healthpb.HealthCheckResponse_SERVING { return fmt.Errorf("server not serving: %s", resp.Status) } Try / catch
// The health-check state machine already retries with backoff; surface non-fatal failures via connectivity callbacks. conn.WaitForStateChange(ctx, connectivity.Ready) // or subscribe to state updates // Treat TransientFailure as a signal to shed load; rely on the built-in retry, do not crash.
Prevention
- Implement and register the grpc.health.v1.Health service on every server.
- Use keepalive.EnforcementPolicy and raise HTTP/2 stream timeouts on intermediaries.
- Watch subchannel connectivity states instead of raw health errors.
When it happens
Trigger: Server's Health/Watch RPC returned an error status: CANCELLED (server shutdown), DEADLINE_EXCEEDED, PERMISSION_DENIED, INTERNAL, or the stream was reset. Emitted at health/client.go:104 during clientHealthCheck.
Common situations: Server process restarting or draining; server has a buggy Health service implementation that closes the stream early; interceptor/proxy resetting HTTP/2 streams; auth/z policy rejecting the Watch call; server panic in the Health handler; load balancer timing out long-lived streams.
Related errors
- connection active but health check failed. status=
- all SubConns are in TransientFailure, last connection error
- duplicated name
- endpoints list contains no addresses
- endpoints list is empty
AI-assisted analysis of grpc/grpc-go@0c51461d27 (2026-08-11).
Data as JSON: /api/errors/92e7a8d9063c5270.
Report an issue: GitHub.
Appendix: source
Thrown at health/client.go:104
if err = s.SendMsg(&healthpb.HealthCheckRequest{Service: service}); err != nil && err != io.EOF {
// Stream should have been closed, so we can safely continue to create a new stream.
continue retryConnection
}
s.CloseSend()
resp := new(healthpb.HealthCheckResponse)
for {
err = s.RecvMsg(resp)
// Reports healthy for the LBing purposes if health check is not implemented in the server.
if status.Code(err) == codes.Unimplemented {
setConnectivityState(connectivity.Ready, nil)
return err
}
// Reports unhealthy if server's Watch method gives an error other than UNIMPLEMENTED.
if err != nil {
setConnectivityState(connectivity.TransientFailure, fmt.Errorf("connection active but received health check RPC error: %v", err))
continue retryConnection
}
// As a message has been received, removes the need for backoff for the next retry by resetting the try count.
tryCnt = 0
if resp.Status == healthpb.HealthCheckResponse_SERVING {
setConnectivityState(connectivity.Ready, nil)
} else {
setConnectivityState(connectivity.TransientFailure, fmt.Errorf("connection active but health check failed. status=%s", resp.Status))
}
}
}
}
View on GitHub (pinned to 0c51461d27)