grpc/grpc-go · error

all SubConns are in TransientFailure, last connection error

Error message

all SubConns are in TransientFailure, last connection error: %v

What it means

Produced by grpclb's regeneratePicker() (balancer/grpclb/grpclb.go:249). When aggregateSubConnStats() returns TransientFailure (no SubConn is Ready, Connecting, or Idle), the picker is set to an errPicker wrapping this message with lb.connErr. This is specific to the deprecated grpclb policy: every backend SubConn — whether from the remote load balancer server list or the fallback list — has failed.

Solutions

  1. Check the embedded connection error (lb.connErr) for the specific transport failure.
  2. Verify the backend addresses delivered by the grpclb server (or fallback resolver) are reachable.
  3. Migrate off the deprecated grpclb policy to xDS-based load balancing (ring_hash / weighted_target) where possible.
  4. Ensure at least one backend in the server list is healthy and listening.

Example fix

// before: using deprecated grpclb with all backends down
import _ "google.golang.org/grpc/balancer/grpclb"
conn, _ := grpc.NewClient("dns:///my-svc:50051", grpc.WithDefaultServiceConfig(`{"loadBalancingConfig":[{"grpclb":{}}]}`))

// after: use round_robin (no remote LB needed) and fix backend reachability
conn, _ := grpc.NewClient("dns:///my-svc:50051", grpc.WithDefaultServiceConfig(`{"loadBalancingConfig":[{"round_robin":{}}]}`))
Defensive patterns

Strategy: retry

Try / catch

// grpclb all-SubConn failure → Unavailable. Retry with backoff.
for {
    resp, err := c.Call(ctx, in, grpc.WaitForReady(true))
    if err == nil { return resp, nil }
    if status.Code(err) != codes.Unavailable { return nil, err }
    select {
    case <-time.After(backoff):
    case <-ctx.Done(): return nil, ctx.Err()
    }
}

Prevention

When it happens

Trigger: grpclb balancer's aggregateSubConnStats() returns TransientFailure because every entry in lb.subConns is in TransientFailure, with none Ready/Connecting/Idle. regeneratePicker() (line 248-250) builds the error picker. Happens after all backend connections fail and none are mid-reconnect.

Common situations: Remote load balancer (grpclb server) returned a server list whose backends are all unreachable; fallback backends (resolved directly) are also down; TLS to backends is misconfigured; the cluster is fully offline. grpclb itself is deprecated in favor of xDS, so this may also indicate an outdated deployment.

Related errors


AI-assisted analysis of grpc/grpc-go@0c51461d27 (2026-08-11). Data as JSON: /api/errors/2d4e0ba1646d6d22. Report an issue: GitHub.

Appendix: source

Thrown at balancer/grpclb/grpclb.go:249

	remoteBalancerConnected bool
	serverListReceived      bool
	inFallback              bool
	// resolvedBackendAddrs is resolvedAddrs minus remote balancers. It's set
	// when resolved address updates are received, and read in the goroutine
	// handling fallback.
	resolvedBackendAddrs []resolver.Address
	connErr              error // the last connection error
}

// regeneratePicker takes a snapshot of the balancer, and generates a picker from
// it. The picker
//   - always returns ErrTransientFailure if the balancer is in TransientFailure,
//   - does two layer roundrobin pick otherwise.
//
// Caller must hold lb.mu.
func (lb *lbBalancer) regeneratePicker(resetDrop bool) {
	if lb.state == connectivity.TransientFailure {
		lb.picker = base.NewErrPicker(fmt.Errorf("all SubConns are in TransientFailure, last connection error: %v", lb.connErr))
		return
	}

	if lb.state == connectivity.Connecting {
		lb.picker = base.NewErrPicker(balancer.ErrNoSubConnAvailable)
		return
	}

	var readySCs []balancer.SubConn
	if lb.usePickFirst {
		for _, sc := range lb.subConns {
			readySCs = append(readySCs, sc)
			break
		}
	} else {
		for _, a := range lb.backendAddrsWithoutMetadata {
			if sc, ok := lb.subConns[a]; ok {
				if st, ok := lb.scStates[sc]; ok && st == connectivity.Ready {

View on GitHub (pinned to 0c51461d27)