{"record":{"id":"4df2fa101a90f815","repo":"risingwavelabs/risingwave","slug":"invalid-url-should-start-with","errorCode":null,"errorMessage":"Invalid url: {}, should start with {}","messagePattern":"Invalid url: (.+?), should start with (.+?)","errorType":"exception","errorClass":null,"httpStatus":null,"severity":"error","filePath":"src/connector/src/source/iceberg/parquet_file_handler.rs","lineNumber":203,"sourceCode":"pub async fn list_data_directory(\n    op: Operator,\n    dir: String,\n    file_scan_backend: &FileScanBackend,\n) -> Result<Vec<String>, anyhow::Error> {\n    let (bucket, file_name) = extract_bucket_and_file_name(&dir, file_scan_backend)?;\n    let prefix = match file_scan_backend {\n        FileScanBackend::S3 => format!(\"s3://{}/\", bucket),\n        FileScanBackend::Gcs => format!(\"gcs://{}/\", bucket),\n        FileScanBackend::Azblob => format!(\"azblob://{}/\", bucket),\n    };\n    if dir.starts_with(&prefix) {\n        op.list(&file_name).await.map_err(Into::into).map(|list| {\n            list.into_iter()\n                .map(|entry| prefix.clone() + entry.path())\n                .collect()\n        })\n    } else {\n        bail!(\"Invalid url: {}, should start with {}\", dir, prefix)\n    }\n}\n\n/// Extracts a suitable `ProjectionMask` from a Parquet file schema based on the user's requested schema.\n///\n/// This function is utilized for column pruning of Parquet files. It checks the user's requested schema\n/// against the schema of the currently read Parquet file. If the provided `columns` are `None`\n/// or if the Parquet file contains nested data types, it returns `ProjectionMask::all()`. Otherwise,\n/// it returns only the columns where both the data type and column name match the requested schema,\n/// facilitating efficient reading of the `RecordBatch`.\n///\n/// # Parameters\n/// - `columns`: An optional vector of `Column` representing the user's requested schema.\n/// - `metadata`: A reference to `FileMetaData` containing the schema and metadata of the Parquet file.\n///\n/// # Returns\n/// - A `ConnectorResult<ProjectionMask>`, which represents the valid columns in the Parquet file schema\n///   that correspond to the requested schema. If an error occurs during processing, it returns an","sourceCodeStart":185,"sourceCodeEnd":221,"githubUrl":"https://github.com/risingwavelabs/risingwave/blob/6469eb736d691e8e9b8a419a57edd6429ca77417/src/connector/src/source/iceberg/parquet_file_handler.rs#L185-L221","documentation":"When building a file scan for an Iceberg data file, `list_data_directory` computes an object-store `prefix` from the file path and checks that the directory URL starts with that prefix before listing via OpenDAL's `op.list`. If the directory URL does not start with the expected prefix, the URL is considered invalid for the configured object store and the function bails.","triggerScenarios":"`new_file_scan` -> `list_data_directory` given a `dir` (data directory path) whose scheme/layout does not match the expected `prefix` derived from the storage configuration — e.g. an s3:// path against a configured prefix of a different bucket/root, or a relative/absolute path mismatch.","commonSituations":"Misconfigured `s3.path` / region / endpoint in iceberg source properties; mixing URI schemes (s3a vs s3, file:// vs plain path); table location changed so file paths no longer live under the configured warehouse prefix.","solutions":["Check the data directory URL in the error message and align it with the configured storage prefix (bucket/root).","Fix iceberg source properties (endpoint, bucket, path style, scheme) so file paths share the expected prefix.","Use a consistent URI scheme (e.g. always s3://) matching the OpenDAL operator configuration.","Verify the table's `location`/warehouse setting points to the same prefix the operator was created with."],"exampleFix":"// before\nWITH (connector='iceberg', iceberg.s3.path='s3a://bucket/warehouse/')\n// after: use the scheme/prefix the operator is configured with\nWITH (connector='iceberg', iceberg.s3.path='s3://bucket/warehouse/')","handlingStrategy":"validation","validationCode":"let prefix = expected_prefix_for_storage(&props);\nif !dir.starts_with(&prefix) {\n    return Err(format!(\"data dir {} must start with {}\", dir, prefix));\n}","typeGuard":"fn has_valid_prefix(dir: &str, prefix: &str) -> bool { dir.starts_with(prefix) }","tryCatchPattern":"match new_file_scan(...) {\n    Err(e) if e.to_string().contains(\"Invalid url\") => {\n        // fix storage properties / scheme and retry\n    }\n    r => r,\n}","preventionTips":["Use one consistent URI scheme (s3:// vs s3a://) everywhere","Verify bucket/root prefix in source properties matches table location","Test storage config with a small scan before production"],"tags":["rust","iceberg","object-storage","url"],"backgroundTag":"invalid-url-format","analyzedSha":"6469eb736d691e8e9b8a419a57edd6429ca77417","analyzedAt":"2026-09-11T21:06:21.487Z","contentChangedAt":"2026-09-11T21:06:21.487Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}