{"record":{"id":"b8f790d463cef3f2","repo":"grpc/grpc-go","slug":"pickfirst-health-check-failure-v","errorCode":null,"errorMessage":"pickfirst: health check failure: %v","messagePattern":"pickfirst: health check failure: (.+?)","errorType":"exception","errorClass":null,"httpStatus":null,"severity":"error","filePath":"balancer/pickfirst/pickfirst.go","lineNumber":777,"sourceCode":"\tdefer b.mu.Unlock()\n\t// Previously relevant SubConns can still callback with state updates.\n\t// To prevent pickers from returning these obsolete SubConns, this logic\n\t// is included to check if the current list of active SubConns includes\n\t// this SubConn.\n\tif !b.isActiveSCData(sd) {\n\t\treturn\n\t}\n\tsd.effectiveState = state.ConnectivityState\n\tswitch state.ConnectivityState {\n\tcase connectivity.Ready:\n\t\tb.updateBalancerState(balancer.State{\n\t\t\tConnectivityState: connectivity.Ready,\n\t\t\tPicker:            &picker{result: balancer.PickResult{SubConn: sd.subConn}},\n\t\t})\n\tcase connectivity.TransientFailure:\n\t\tb.updateBalancerState(balancer.State{\n\t\t\tConnectivityState: connectivity.TransientFailure,\n\t\t\tPicker:            &picker{err: fmt.Errorf(\"pickfirst: health check failure: %v\", state.ConnectionError)},\n\t\t})\n\tcase connectivity.Connecting:\n\t\tb.updateBalancerState(balancer.State{\n\t\t\tConnectivityState: connectivity.Connecting,\n\t\t\tPicker:            &picker{err: balancer.ErrNoSubConnAvailable},\n\t\t})\n\tdefault:\n\t\tb.logger.Errorf(\"Got unexpected health update for SubConn %p: %v\", state)\n\t}\n}\n\n// updateBalancerState stores the state reported to the channel and calls\n// ClientConn.UpdateState(). As an optimization, it avoids sending duplicate\n// updates to the channel.\nfunc (b *pickfirstBalancer) updateBalancerState(newState balancer.State) {\n\t// In case of TransientFailures allow the picker to be updated to update\n\t// the connectivity error, in all other cases don't send duplicate state\n\t// updates.","sourceCodeStart":759,"sourceCodeEnd":795,"githubUrl":"https://github.com/grpc/grpc-go/blob/0c51461d27177d997e14c642fe18c11668fc09a3/balancer/pickfirst/pickfirst.go#L759-L795","documentation":"Produced by pickfirst's updateSubConnHealthState() (balancer/pickfirst/pickfirst.go:777). When client-side health checking is enabled (healthCheckingEnabled == true, set via EnableHealthListener) and the health listener registered for a READY SubConn reports TransientFailure, the balancer moves to TransientFailure with this picker error. The backend transport is up, but the health check says the backend is unhealthy.","triggerScenarios":"A SubConn reached READY, registered a health listener (line 642), and the listener called back with ConnectivityState == TransientFailure (line 774-778). healthCheckingEnabled must be true (EnableHealthListener was applied to the resolver state). The picker reports 'pickfirst: health check failure: <ConnectionError>'.","commonSituations":"Backend is reachable at the TCP/TLS level but its gRPC health service (grpc.health.v1) returns NOT_SERVING; the health check stream breaks; service mesh / xDS outlier detection marks the endpoint unhealthy; the backend is draining but still accepting connections.","solutions":["Check the backend's gRPC health service status (grpc.health.v1.Health.Check) — it must return SERVING.","Verify the health check service is registered on the server.","If using xDS outlier detection, check whether the endpoint was ejected and why.","Temporarily disable client-side health checks to isolate whether the transport itself works."],"exampleFix":"// Server side: register the health service so checks return SERVING\nhealthSvc := health.NewServer()\nhealthSvc.SetServingStatus(\"\", healthpb.HealthCheckResponse_SERVING)\ns := grpc.NewServer()\nhealthpb.RegisterHealthServer(s, healthSvc)\n\n// Client side: if you did not intend health checks, do not call EnableHealthListener","handlingStrategy":"retry","validationCode":null,"typeGuard":null,"tryCatchPattern":"// Health-check failure → Unavailable while the transport is up.\n// Retry; the health check may flip back to SERVING.\nfor {\n    resp, err := c.Call(ctx, in, grpc.WaitForReady(true))\n    if err == nil { return resp, nil }\n    if status.Code(err) != codes.Unavailable { return nil, err }\n    select {\n    case <-time.After(healthBackoff):\n    case <-ctx.Done(): return nil, ctx.Err()\n    }\n}","preventionTips":["Register the gRPC health service (grpc.health.v1) on every server and report SERVING only when ready.","Monitor health-check TransientFailure rates to detect backend unhealthiness.","Do not enable client-side health checks via EnableHealthListener unless the server implements the health service."],"tags":["grpc","balancer","pickfirst","health-check","transient-failure"],"backgroundTag":null,"analyzedSha":"0c51461d27177d997e14c642fe18c11668fc09a3","analyzedAt":"2026-08-11T14:49:15.055Z","contentChangedAt":null,"schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}