risingwavelabs/risingwave · error
Invalid url: , should start with
Error message
Invalid url: {}, should start with {} What it means
When building a file scan for an Iceberg data file, `list_data_directory` computes an object-store `prefix` from the file path and checks that the directory URL starts with that prefix before listing via OpenDAL's `op.list`. If the directory URL does not start with the expected prefix, the URL is considered invalid for the configured object store and the function bails.
Solutions
- Check the data directory URL in the error message and align it with the configured storage prefix (bucket/root).
- Fix iceberg source properties (endpoint, bucket, path style, scheme) so file paths share the expected prefix.
- Use a consistent URI scheme (e.g. always s3://) matching the OpenDAL operator configuration.
- Verify the table's `location`/warehouse setting points to the same prefix the operator was created with.
Example fix
// before WITH (connector='iceberg', iceberg.s3.path='s3a://bucket/warehouse/') // after: use the scheme/prefix the operator is configured with WITH (connector='iceberg', iceberg.s3.path='s3://bucket/warehouse/')
Defensive patterns
Strategy: validation
Validate before calling
let prefix = expected_prefix_for_storage(&props);
if !dir.starts_with(&prefix) {
return Err(format!("data dir {} must start with {}", dir, prefix));
} Type guard
fn has_valid_prefix(dir: &str, prefix: &str) -> bool { dir.starts_with(prefix) } Try / catch
match new_file_scan(...) {
Err(e) if e.to_string().contains("Invalid url") => {
// fix storage properties / scheme and retry
}
r => r,
} Prevention
- Use one consistent URI scheme (s3:// vs s3a://) everywhere
- Verify bucket/root prefix in source properties matches table location
- Test storage config with a small scan before production
When it happens
Trigger: `new_file_scan` -> `list_data_directory` given a `dir` (data directory path) whose scheme/layout does not match the expected `prefix` derived from the storage configuration — e.g. an s3:// path against a configured prefix of a different bucket/root, or a relative/absolute path mismatch.
Common situations: Misconfigured `s3.path` / region / endpoint in iceberg source properties; mixing URI schemes (s3a vs s3, file:// vs plain path); table location changed so file paths no longer live under the configured warehouse prefix.
Understand the failure class
Background: "Invalid URL" errors: why new URL(), URI.parse, and reqwest::Url reject your string — missing scheme, whitespace, and bad path format — this error's family across 39 libraries.
Related errors
- Invalid warehouse path
- adlsgen2.authority_host does not parse as a URL
- adlsgen2.authority_host must not contain a path component
- adlsgen2.authority_host must not contain a query or fragment
- adlsgen2.authority_host must not contain userinfo
AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11).
Data as JSON: /api/errors/4df2fa101a90f815.
Report an issue: GitHub.
Appendix: source
Thrown at src/connector/src/source/iceberg/parquet_file_handler.rs:203
pub async fn list_data_directory(
op: Operator,
dir: String,
file_scan_backend: &FileScanBackend,
) -> Result<Vec<String>, anyhow::Error> {
let (bucket, file_name) = extract_bucket_and_file_name(&dir, file_scan_backend)?;
let prefix = match file_scan_backend {
FileScanBackend::S3 => format!("s3://{}/", bucket),
FileScanBackend::Gcs => format!("gcs://{}/", bucket),
FileScanBackend::Azblob => format!("azblob://{}/", bucket),
};
if dir.starts_with(&prefix) {
op.list(&file_name).await.map_err(Into::into).map(|list| {
list.into_iter()
.map(|entry| prefix.clone() + entry.path())
.collect()
})
} else {
bail!("Invalid url: {}, should start with {}", dir, prefix)
}
}
/// Extracts a suitable `ProjectionMask` from a Parquet file schema based on the user's requested schema.
///
/// This function is utilized for column pruning of Parquet files. It checks the user's requested schema
/// against the schema of the currently read Parquet file. If the provided `columns` are `None`
/// or if the Parquet file contains nested data types, it returns `ProjectionMask::all()`. Otherwise,
/// it returns only the columns where both the data type and column name match the requested schema,
/// facilitating efficient reading of the `RecordBatch`.
///
/// # Parameters
/// - `columns`: An optional vector of `Column` representing the user's requested schema.
/// - `metadata`: A reference to `FileMetaData` containing the schema and metadata of the Parquet file.
///
/// # Returns
/// - A `ConnectorResult<ProjectionMask>`, which represents the valid columns in the Parquet file schema
/// that correspond to the requested schema. If an error occurs during processing, it returns anView on GitHub (pinned to 6469eb736d)