grpc/grpc-go · error
last connection error
Error message
last connection error: %v
What it means
Produced by the base balancer's mergeErrors() (balancer/base/balancer.go:156) and surfaced through the error picker when the aggregated balancer state is TransientFailure. The base balancer underlies policies like round_robin. It means the name resolver is healthy (resolverErr is nil) but every SubConn failed to connect; connErr holds the most recent transport-level failure. Each RPC Pick returns this error until at least one SubConn leaves TransientFailure.
Solutions
- Inspect the embedded connection error (the %v) for the concrete cause (e.g. 'connection refused', 'tls: handshake failure', 'i/o timeout').
- Verify the backend processes are running and listening on the addresses returned by the resolver.
- Check network reachability from the client (firewall, security groups, routing, VPC peering).
- Confirm the resolver (DNS/xDS) returns the correct host:port list.
Example fix
// before: wrong port, nothing listening on :80
conn, _ := grpc.NewClient("dns:///my-svc:80", grpc.WithDefaultServiceConfig(`{"loadBalancingConfig":[{"round_robin":{}}]}`))
// after: correct gRPC port and use WaitForReady so RPCs queue instead of failing immediately
conn, _ := grpc.NewClient("dns:///my-svc:50051", grpc.WithDefaultServiceConfig(`{"loadBalancingConfig":[{"round_robin":{}}]}`))
resp, err := client.Call(ctx, req, grpc.WaitForReady(true)) Defensive patterns
Strategy: retry
Try / catch
// Connection failures surface as RPC errors. Retry with backoff and
// respect grpc.WaitForReady so RPCs queue during TransientFailure.
func callWithRetry(ctx context.Context, c pb.FooClient, in *pb.Req) (*pb.Resp, error) {
bo := backoff.NewExponentialBackOff()
for {
resp, err := c.Call(ctx, in, grpc.WaitForReady(true))
if err == nil {
return resp, nil
}
if status.Code(err) != codes.Unavailable {
return nil, err // non-transient: do not retry
}
select {
case <-time.After(bo.NextBackOff()):
case <-ctx.Done():
return nil, ctx.Err()
}
}
} Prevention
- Monitor gRPC channel connectivity state (WaitForStateChange) and alert on TransientFailure.
- Validate backend reachability in health checks / readiness probes before sending traffic.
- Use grpc.WaitForReady(true) on non-urgent RPCs to absorb transient backend downtime.
When it happens
Trigger: All SubConns of a base-derived balancer (e.g. round_robin) are in TransientFailure, the last UpdateClientConnState succeeded so resolverErr == nil, and the channel reports TransientFailure. The picker built by regeneratePicker() wraps mergeErrors() which hits the connErr != nil / resolverErr == nil branch at line 156.
Common situations: Backend servers are down or listening on the wrong port; TLS/mTLS credentials mismatch between client and server; firewall or security group blocks the gRPC port; DNS returns stale or dead IPs; the dial target uses a port nothing listens on.
Related errors
- last connection error
- all SubConns are in TransientFailure, last connection error
- name resolver error
- last resolver error
- pickfirst: health check failure
AI-assisted analysis of grpc/grpc-go@0c51461d27 (2026-08-11).
Data as JSON: /api/errors/1abfb16b7e160a1b.
Report an issue: GitHub.
Appendix: source
Thrown at balancer/base/balancer.go:156
b.ResolverError(errors.New("produced zero addresses"))
return balancer.ErrBadResolverState
}
b.regeneratePicker()
b.cc.UpdateState(balancer.State{ConnectivityState: b.state, Picker: b.picker})
return nil
}
// mergeErrors builds an error from the last connection error and the last
// resolver error. Must only be called if b.state is TransientFailure.
func (b *baseBalancer) mergeErrors() error {
// connErr must always be non-nil unless there are no SubConns, in which
// case resolverErr must be non-nil.
if b.connErr == nil {
return fmt.Errorf("last resolver error: %v", b.resolverErr)
}
if b.resolverErr == nil {
return fmt.Errorf("last connection error: %v", b.connErr)
}
return fmt.Errorf("last connection error: %v; last resolver error: %v", b.connErr, b.resolverErr)
}
// regeneratePicker takes a snapshot of the balancer, and generates a picker
// from it. The picker is
// - errPicker if the balancer is in TransientFailure,
// - built by the pickerBuilder with all READY SubConns otherwise.
func (b *baseBalancer) regeneratePicker() {
if b.state == connectivity.TransientFailure {
b.picker = NewErrPicker(b.mergeErrors())
return
}
readySCs := make(map[balancer.SubConn]SubConnInfo)
// Filter out all ready SCs from full subConn map.
for addr, sc := range b.subConns.All() {
if st, ok := b.scStates[sc]; ok && st == connectivity.Ready {View on GitHub (pinned to 0c51461d27)