apache/iceberg · error · NotFoundException
File does not exist
Error message
File does not exist: %s
What it means
HadoopInputFile caches a FileStatus lazily via lazyStat(). If fs.getFileStatus throws FileNotFoundException — the file does not exist at the given path — it is rethrown as a NotFoundException with the path. This is the Iceberg way of surfacing a missing data/metadata file during reads.
Solutions
- Check the path in the exception and verify the file exists on the target filesystem.
- Verify table metadata is current: refresh the table and retry the scan.
- Confirm fs.defaultFS / scheme / authority configuration matches the cluster that wrote the files.
- Restore accidentally deleted files from backup or roll back the snapshot; disable concurrent cleanup jobs racing with reads.
- If files legitimately vanished, rewrite metadata or use Iceberg rollback to a snapshot whose files still exist.
Example fix
// before
InputFile f = io.newInputFile("s3://bucket/warehouse/db/t/metadata/old.avro");
long len = f.getLength(); // throws NotFoundException
// after
InputFile f = io.newInputFile(path);
if (f.exists()) {
long len = f.getLength();
} else {
table.refresh(); // pick up newer metadata
} Defensive patterns
Strategy: try-catch
Validate before calling
InputFile f = io.newInputFile(path);
if (f.exists()) { /* safe to read */ } Type guard
boolean readable = io.newInputFile(path).exists();
Try / catch
try {
InputFile f = io.newInputFile(path);
long len = f.getLength();
} catch (NotFoundException e) {
// file missing: refresh table metadata or restore file
} Prevention
- Refresh tables before scans to avoid stale metadata referencing deleted files.
- Keep snapshot expiration retention long enough for active readers.
- Align fs.defaultFS and path authorities with the cluster that wrote the data.
- Don't run orphan-file cleanup concurrently with readers.
When it happens
Trigger: Calling getLength(), getStat(), or exists() on a HadoopInputFile whose path was deleted, never written, has a typo, or whose authority/scheme resolves to the wrong filesystem; reading a table whose manifest/data files were removed by retention cleanup.
Common situations: Table metadata pointing to files deleted by expireSnapshots or orphan-file cleanup; wrong fs.defaultFS or warehouse path so relative paths resolve elsewhere; object-store eventual consistency on aggressively tuned S3 clients.
Understand the failure class
Background: "File not found" and ENOENT errors: why libraries can't find a file that should exist — this error's family across 50 libraries.
Related errors
- Failed to open input stream for file
- No in-memory file found for location
- Already exists
- Cannot apply unknown WAP ID
- Cannot find plan with id
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/bc07684e686bd148.
Report an issue: GitHub.
Appendix: source
Thrown at core/src/main/java/org/apache/iceberg/hadoop/HadoopInputFile.java:166
this.conf = conf;
this.length = length;
}
private HadoopInputFile(FileSystem fs, FileStatus stat, Configuration conf) {
this.fs = fs;
this.path = stat.getPath();
this.location = path.toString();
this.stat = stat;
this.conf = conf;
this.length = stat.getLen();
}
private FileStatus lazyStat() {
if (stat == null) {
try {
this.stat = fs.getFileStatus(path);
} catch (FileNotFoundException e) {
throw new NotFoundException(e, "File does not exist: %s", path);
} catch (IOException e) {
throw new RuntimeIOException(e, "Failed to get status for file: %s", path);
}
}
return stat;
}
@Override
public long getLength() {
if (length == null) {
this.length = lazyStat().getLen();
}
return length;
}
@Override
public SeekableInputStream newStream() {
try {View on GitHub (pinned to 86d9c8fc54)