hashicorp/terraform · error

retry timeout and got an error: %#v

Error message

retry timeout and got an error: %#v

What it means

Thrown by Invoker.Run() when a retryable error exhausts all configured retry attempts. The OSS backend registers two catchers: ClientErrorCatcher (matches 'AliyunGoClientFailure', 10 retries, 3s wait) and ServiceBusyCatcher (matches 'ServiceUnavailable', 10 retries, 3s wait). After 10 failed retries (~30s of sleep), the error is returned with %#v formatting of the last error.

Solutions

  1. Check Alibaba Cloud status page for ongoing OSS or location service incidents in your region.
  2. Verify credentials are valid and not expired — STS tokens may have rotated.
  3. Check network connectivity and firewall rules to OSS endpoints.
  4. If the error is transient, simply retry the Terraform operation.
  5. For persistent 'AliyunGoClientFailure', upgrade the alibaba-cloud-sdk-go dependency as it may be an SDK-level issue.
Defensive patterns

Strategy: retry

Validate before calling

// Pre-check Alibaba Cloud service health before running Terraform
func checkOSSServiceHealth(region string) error {
    // Simple DNS + connectivity check
    endpoint := fmt.Sprintf("oss-%s.aliyuncs.com", region)
    conn, err := net.DialTimeout("tcp", endpoint+":443", 5*time.Second)
    if err != nil {
        return fmt.Errorf("cannot reach OSS in %s: %w", region, err)
    }
    conn.Close()
    return nil
}

Try / catch

// Wrap Invoker.Run to add outer retry for exhausted inner retries
func runWithOuterRetry(invoker *Invoker, f func() error, outerRetries int) error {
    var lastErr error
    for i := 0; i < outerRetries; i++ {
        err := invoker.Run(f)
        if err == nil {
            return nil
        }
        if strings.Contains(err.Error(), "retry timeout") {
            lastErr = err
            time.Sleep(10 * time.Second)
            continue
        }
        return err // non-retryable error
    }
    return lastErr
}

Prevention

When it happens

Trigger: Invoker.Run(f) calls f(), which returns an error containing 'AliyunGoClientFailure' or 'ServiceUnavailable' in its message string. The catcher decrements RetryCount from 10, sleeping 3s between retries. When RetryCount reaches 0, this error is returned. The matcher uses strings.Contains, so the error message must contain the exact substring.

Common situations: Alibaba Cloud OSS service is genuinely down or degraded in the region. Client-side persistent failure that includes 'AliyunGoClientFailure' in the error (e.g. credential rotation taking longer than 30s). Network partition between the Terraform runner and OSS lasting more than ~30 seconds. RAM permission issue that the SDK reports as ServiceUnavailable.

Understand the failure class

Related errors


AI-assisted analysis of hashicorp/terraform@d32a084675 (2026-08-11). Data as JSON: /api/errors/7225f75e16c79164. Report an issue: GitHub.

Appendix: source

Thrown at internal/backend/remote-state/oss/backend.go:551

}

func (a *Invoker) AddCatcher(catcher Catcher) {
	a.catchers = append(a.catchers, &catcher)
}

func (a *Invoker) Run(f func() error) error {
	err := f()

	if err == nil {
		return nil
	}

	for _, catcher := range a.catchers {
		if strings.Contains(err.Error(), catcher.Reason) {
			catcher.RetryCount--

			if catcher.RetryCount <= 0 {
				return fmt.Errorf("retry timeout and got an error: %#v", err)
			} else {
				time.Sleep(time.Duration(catcher.RetryWaitSeconds) * time.Second)
				return a.Run(f)
			}
		}
	}
	return err
}

var providerConfig map[string]interface{}

func getConfigFromProfile(d *schema.ResourceData, ProfileKey string) (interface{}, error) {

	if providerConfig == nil {
		if v, ok := d.GetOk("profile"); !ok || v.(string) == "" {
			return nil, nil
		}
		current := d.Get("profile").(string)

View on GitHub (pinned to d32a084675)