grpc/grpc-go · error
pickfirst: health check failure
Error message
pickfirst: health check failure: %v
What it means
Produced by pickfirst's updateSubConnHealthState() (balancer/pickfirst/pickfirst.go:777). When client-side health checking is enabled (healthCheckingEnabled == true, set via EnableHealthListener) and the health listener registered for a READY SubConn reports TransientFailure, the balancer moves to TransientFailure with this picker error. The backend transport is up, but the health check says the backend is unhealthy.
Solutions
- Check the backend's gRPC health service status (grpc.health.v1.Health.Check) — it must return SERVING.
- Verify the health check service is registered on the server.
- If using xDS outlier detection, check whether the endpoint was ejected and why.
- Temporarily disable client-side health checks to isolate whether the transport itself works.
Example fix
// Server side: register the health service so checks return SERVING
healthSvc := health.NewServer()
healthSvc.SetServingStatus("", healthpb.HealthCheckResponse_SERVING)
s := grpc.NewServer()
healthpb.RegisterHealthServer(s, healthSvc)
// Client side: if you did not intend health checks, do not call EnableHealthListener Defensive patterns
Strategy: retry
Try / catch
// Health-check failure → Unavailable while the transport is up.
// Retry; the health check may flip back to SERVING.
for {
resp, err := c.Call(ctx, in, grpc.WaitForReady(true))
if err == nil { return resp, nil }
if status.Code(err) != codes.Unavailable { return nil, err }
select {
case <-time.After(healthBackoff):
case <-ctx.Done(): return nil, ctx.Err()
}
} Prevention
- Register the gRPC health service (grpc.health.v1) on every server and report SERVING only when ready.
- Monitor health-check TransientFailure rates to detect backend unhealthiness.
- Do not enable client-side health checks via EnableHealthListener unless the server implements the health service.
When it happens
Trigger: A SubConn reached READY, registered a health listener (line 642), and the listener called back with ConnectivityState == TransientFailure (line 774-778). healthCheckingEnabled must be true (EnableHealthListener was applied to the resolver state). The picker reports 'pickfirst: health check failure: <ConnectionError>'.
Common situations: Backend is reachable at the TCP/TLS level but its gRPC health service (grpc.health.v1) returns NOT_SERVING; the health check stream breaks; service mesh / xDS outlier detection marks the endpoint unhealthy; the backend is draining but still accepting connections.
Related errors
- name resolver error
- all SubConns are in TransientFailure, last connection error
- last connection error
- last connection error
- pickfirst: received illegal BalancerConfig (type %T)
AI-assisted analysis of grpc/grpc-go@0c51461d27 (2026-08-11).
Data as JSON: /api/errors/b8f790d463cef3f2.
Report an issue: GitHub.
Appendix: source
Thrown at balancer/pickfirst/pickfirst.go:777
defer b.mu.Unlock()
// Previously relevant SubConns can still callback with state updates.
// To prevent pickers from returning these obsolete SubConns, this logic
// is included to check if the current list of active SubConns includes
// this SubConn.
if !b.isActiveSCData(sd) {
return
}
sd.effectiveState = state.ConnectivityState
switch state.ConnectivityState {
case connectivity.Ready:
b.updateBalancerState(balancer.State{
ConnectivityState: connectivity.Ready,
Picker: &picker{result: balancer.PickResult{SubConn: sd.subConn}},
})
case connectivity.TransientFailure:
b.updateBalancerState(balancer.State{
ConnectivityState: connectivity.TransientFailure,
Picker: &picker{err: fmt.Errorf("pickfirst: health check failure: %v", state.ConnectionError)},
})
case connectivity.Connecting:
b.updateBalancerState(balancer.State{
ConnectivityState: connectivity.Connecting,
Picker: &picker{err: balancer.ErrNoSubConnAvailable},
})
default:
b.logger.Errorf("Got unexpected health update for SubConn %p: %v", state)
}
}
// updateBalancerState stores the state reported to the channel and calls
// ClientConn.UpdateState(). As an optimization, it avoids sending duplicate
// updates to the channel.
func (b *pickfirstBalancer) updateBalancerState(newState balancer.State) {
// In case of TransientFailures allow the picker to be updated to update
// the connectivity error, in all other cases don't send duplicate state
// updates.View on GitHub (pinned to 0c51461d27)