risingwavelabs/risingwave · error

Invalid url: , should start with

Error message

Invalid url: {}, should start with {}

What it means

When building a file scan for an Iceberg data file, `list_data_directory` computes an object-store `prefix` from the file path and checks that the directory URL starts with that prefix before listing via OpenDAL's `op.list`. If the directory URL does not start with the expected prefix, the URL is considered invalid for the configured object store and the function bails.

Solutions

  1. Check the data directory URL in the error message and align it with the configured storage prefix (bucket/root).
  2. Fix iceberg source properties (endpoint, bucket, path style, scheme) so file paths share the expected prefix.
  3. Use a consistent URI scheme (e.g. always s3://) matching the OpenDAL operator configuration.
  4. Verify the table's `location`/warehouse setting points to the same prefix the operator was created with.

Example fix

// before
WITH (connector='iceberg', iceberg.s3.path='s3a://bucket/warehouse/')
// after: use the scheme/prefix the operator is configured with
WITH (connector='iceberg', iceberg.s3.path='s3://bucket/warehouse/')
Defensive patterns

Strategy: validation

Validate before calling

let prefix = expected_prefix_for_storage(&props);
if !dir.starts_with(&prefix) {
    return Err(format!("data dir {} must start with {}", dir, prefix));
}

Type guard

fn has_valid_prefix(dir: &str, prefix: &str) -> bool { dir.starts_with(prefix) }

Try / catch

match new_file_scan(...) {
    Err(e) if e.to_string().contains("Invalid url") => {
        // fix storage properties / scheme and retry
    }
    r => r,
}

Prevention

When it happens

Trigger: `new_file_scan` -> `list_data_directory` given a `dir` (data directory path) whose scheme/layout does not match the expected `prefix` derived from the storage configuration — e.g. an s3:// path against a configured prefix of a different bucket/root, or a relative/absolute path mismatch.

Common situations: Misconfigured `s3.path` / region / endpoint in iceberg source properties; mixing URI schemes (s3a vs s3, file:// vs plain path); table location changed so file paths no longer live under the configured warehouse prefix.

Understand the failure class

Background: "Invalid URL" errors: why new URL(), URI.parse, and reqwest::Url reject your string — missing scheme, whitespace, and bad path format — this error's family across 39 libraries.

Related errors


AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11). Data as JSON: /api/errors/4df2fa101a90f815. Report an issue: GitHub.

Appendix: source

Thrown at src/connector/src/source/iceberg/parquet_file_handler.rs:203

pub async fn list_data_directory(
    op: Operator,
    dir: String,
    file_scan_backend: &FileScanBackend,
) -> Result<Vec<String>, anyhow::Error> {
    let (bucket, file_name) = extract_bucket_and_file_name(&dir, file_scan_backend)?;
    let prefix = match file_scan_backend {
        FileScanBackend::S3 => format!("s3://{}/", bucket),
        FileScanBackend::Gcs => format!("gcs://{}/", bucket),
        FileScanBackend::Azblob => format!("azblob://{}/", bucket),
    };
    if dir.starts_with(&prefix) {
        op.list(&file_name).await.map_err(Into::into).map(|list| {
            list.into_iter()
                .map(|entry| prefix.clone() + entry.path())
                .collect()
        })
    } else {
        bail!("Invalid url: {}, should start with {}", dir, prefix)
    }
}

/// Extracts a suitable `ProjectionMask` from a Parquet file schema based on the user's requested schema.
///
/// This function is utilized for column pruning of Parquet files. It checks the user's requested schema
/// against the schema of the currently read Parquet file. If the provided `columns` are `None`
/// or if the Parquet file contains nested data types, it returns `ProjectionMask::all()`. Otherwise,
/// it returns only the columns where both the data type and column name match the requested schema,
/// facilitating efficient reading of the `RecordBatch`.
///
/// # Parameters
/// - `columns`: An optional vector of `Column` representing the user's requested schema.
/// - `metadata`: A reference to `FileMetaData` containing the schema and metadata of the Parquet file.
///
/// # Returns
/// - A `ConnectorResult<ProjectionMask>`, which represents the valid columns in the Parquet file schema
///   that correspond to the requested schema. If an error occurs during processing, it returns an

View on GitHub (pinned to 6469eb736d)