grpc/grpc-go · warning

connection active but received health check RPC error

Error message

connection active but received health check RPC error: %v

What it means

The gRPC client health-check watcher received an RPC error (other than UNIMPLEMENTED) on an open connection. Per the grpc health protocol, this sets the subchannel to TransientFailure and triggers backoff/retry; the wrapped %v is the stream RecvMsg error. This is the health-check state machine reporting a server-side Watch failure, not necessarily a transport failure.

Solutions

  1. Inspect the inner %v: CANCELLED/UNAVAILABLE usually means server restart; let the client's built-in backoff retry.
  2. If PERMISSION_DENIED, fix the auth policy on the server's Health service.
  3. If the stream resets repeatedly, check server logs for panics in the Health implementation and ensure the service is registered with the gRPC server.
  4. Stabilize long-lived streams: raise HTTP/2 keepalive timeouts on intermediaries, or use keepalive.EnforcementPolicy on the server.
  5. Confirm the server implements grpc.health.v1.Health (UNIMPLEMENTED is handled separately and is non-fatal).

Example fix

// before: server registers no Health service and Watch returns INTERNAL on a bad handler
srv := grpc.NewServer()

// after
import healthpb "google.golang.org/grpc/health"
import healthsvc "google.golang.org/grpc/health/grpc_health_v1"
hs := healthpb.NewServer()
hs.SetServingStatus("", healthpb.HealthCheckResponse_SERVING)
healthsvc.RegisterHealthServer(srv, hs)
Defensive patterns

Strategy: retry

Validate before calling

// Health is built into grpc-go's connection; pre-flight by checking the server implements Health:
import healthpb "google.golang.org/grpc/health/grpc_health_v1"

hc := healthpb.NewHealthClient(conn)
resp, err := hc.Check(ctx, &healthpb.HealthCheckRequest{Service: svc})
if err != nil { return fmt.Errorf("server Health not usable: %w", err) }
if resp.Status != healthpb.HealthCheckResponse_SERVING { return fmt.Errorf("server not serving: %s", resp.Status) }

Try / catch

// The health-check state machine already retries with backoff; surface non-fatal failures via connectivity callbacks.
conn.WaitForStateChange(ctx, connectivity.Ready) // or subscribe to state updates
// Treat TransientFailure as a signal to shed load; rely on the built-in retry, do not crash.

Prevention

When it happens

Trigger: Server's Health/Watch RPC returned an error status: CANCELLED (server shutdown), DEADLINE_EXCEEDED, PERMISSION_DENIED, INTERNAL, or the stream was reset. Emitted at health/client.go:104 during clientHealthCheck.

Common situations: Server process restarting or draining; server has a buggy Health service implementation that closes the stream early; interceptor/proxy resetting HTTP/2 streams; auth/z policy rejecting the Watch call; server panic in the Health handler; load balancer timing out long-lived streams.

Related errors


AI-assisted analysis of grpc/grpc-go@0c51461d27 (2026-08-11). Data as JSON: /api/errors/92e7a8d9063c5270. Report an issue: GitHub.

Appendix: source

Thrown at health/client.go:104

		if err = s.SendMsg(&healthpb.HealthCheckRequest{Service: service}); err != nil && err != io.EOF {
			// Stream should have been closed, so we can safely continue to create a new stream.
			continue retryConnection
		}
		s.CloseSend()

		resp := new(healthpb.HealthCheckResponse)
		for {
			err = s.RecvMsg(resp)

			// Reports healthy for the LBing purposes if health check is not implemented in the server.
			if status.Code(err) == codes.Unimplemented {
				setConnectivityState(connectivity.Ready, nil)
				return err
			}

			// Reports unhealthy if server's Watch method gives an error other than UNIMPLEMENTED.
			if err != nil {
				setConnectivityState(connectivity.TransientFailure, fmt.Errorf("connection active but received health check RPC error: %v", err))
				continue retryConnection
			}

			// As a message has been received, removes the need for backoff for the next retry by resetting the try count.
			tryCnt = 0
			if resp.Status == healthpb.HealthCheckResponse_SERVING {
				setConnectivityState(connectivity.Ready, nil)
			} else {
				setConnectivityState(connectivity.TransientFailure, fmt.Errorf("connection active but health check failed. status=%s", resp.Status))
			}
		}
	}
}

View on GitHub (pinned to 0c51461d27)