apache/seatunnel · error · SeaTunnelRuntimeException

Get file status for [ ] failed, cause=

Error message

Get file status for [%s] failed, cause=%s: %s

What it means

safeGetFileSize queries the Hadoop FileSystem for a file's length so split boundaries can be computed. If getFileStatus throws IOException (missing file, permissions, storage/network errors), it is converted into 'Get file status for [path] failed, cause=...' via mapToRuntimeException.

Solutions

  1. Confirm the exact path exists and is accessible from the SeaTunnel worker nodes
  2. Check filesystem credentials/config (core-site.xml, service account, tokens) on workers
  3. Validate the path/bucket spelling and that no upstream job deletes it concurrently
  4. Inspect the wrapped cause in the message for the true root cause (FileNotFoundException vs permission vs timeout)

Example fix

// before
"path" = "gs://my-buket/data/file.csv"  // typo'd bucket -> getFileStatus IOException
// after
"path" = "gs://my-bucket/data/file.csv"
Defensive patterns

Strategy: try-catch

Validate before calling

hadoopFileSystemProxy.getFileStatus(path); // call once in a pre-check; catch IOException and fail fast with a clear message

Type guard

static boolean isAccessible(HadoopFileSystemProxy fs, String p) {
    try { fs.getFileStatus(p); return true; } catch (IOException e) { return false; }
}

Try / catch

try {
    long size = safeGetFileSize(filePath);
} catch (SeaTunnelRuntimeException e) {
    // inspect cause: FileNotFoundException -> fix path; auth -> fix credentials
    throw e;
}

Prevention

When it happens

Trigger: hadopFileSystemProxy.getFileStatus(filePath) throwing IOException — typically because the path does not exist, the caller lacks read permission, the bucket/container is wrong, or the remote filesystem is unreachable.

Common situations: Typo in path or bucket, missing GCS/S3 credentials on worker nodes, file deleted between source enumeration and split calculation, or HDFS namenode unreachable.

Understand the failure class

Background: "File not found" and ENOENT errors: why libraries can't find a file that should exist — this error's family across 50 libraries.

Related errors


AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10). Data as JSON: /api/errors/5a2eee3ea208a8c7. Report an issue: GitHub.

Appendix: source

Thrown at seatunnel-connectors-v2/connector-file/connector-file-base/src/main/java/org/apache/seatunnel/connectors/seatunnel/file/source/split/AccordingToSplitSizeSplitStrategy.java:130

                if (actualEnd <= currentStart) {
                    actualEnd = tentativeEnd;
                }
                splits.add(
                        new FileSourceSplit(
                                tableId, normalizedPath, currentStart, actualEnd - currentStart));
                currentStart = actualEnd;
            }
            return splits;
        } catch (IOException e) {
            throw mapToRuntimeException(normalizedPath, "Split file", e);
        }
    }

    private long safeGetFileSize(String filePath) {
        try {
            return hadoopFileSystemProxy.getFileStatus(filePath).getLen();
        } catch (IOException e) {
            throw mapToRuntimeException(filePath, "Get file status", e);
        }
    }

    private static SeaTunnelRuntimeException mapToRuntimeException(
            String filePath, String operation, IOException e) {
        IOException unwrapped = unwrapRemoteException(e);
        FileConnectorErrorCode errorCode = mapIOExceptionToErrorCode(unwrapped);
        String message =
                String.format(
                        "%s for [%s] failed, cause=%s: %s",
                        operation,
                        filePath,
                        unwrapped.getClass().getSimpleName(),
                        unwrapped.getMessage());
        return new SeaTunnelRuntimeException(errorCode, message, unwrapped);
    }

    private static FileConnectorErrorCode mapIOExceptionToErrorCode(IOException e) {

View on GitHub (pinned to cf67b549a7)