risingwavelabs/risingwave · error · SinkError::Config

Turbopuffer full_text_search column

Error message

Turbopuffer full_text_search column '{}' must be string or []string

What it means

build_turbopuffer_schema validates that every column marked with turbopuffer.full_text_search is either VARCHAR or VARCHAR[] ([]string), since Turbopuffer full-text search indexes only string attributes. Any other data type marked for full-text search is rejected when building the sink schema.

Solutions

  1. Change the column to VARCHAR or VARCHAR[] in the source/materialized view.
  2. Cast the column to VARCHAR in an intermediate materialized view before sinking.
  3. Remove the column from turbopuffer.full_text_search if full-text indexing is not needed.

Example fix

// before: ts TIMESTAMP marked for FTS
WITH (turbopuffer.full_text_search='description,ts')
// after: cast to string first
CREATE MVIEW mv AS SELECT description, ts::VARCHAR AS ts FROM src;
WITH (turbopuffer.full_text_search='description,ts')
Defensive patterns

Strategy: validation

Validate before calling

fn fts_type_ok(dt: &str) -> bool {
    dt == "VARCHAR" || dt == "VARCHAR[]" || dt == "character varying" || dt == "character varying[]"
}

Prevention

When it happens

Trigger: Creating a Turbopuffer sink with turbopuffer.full_text_search='col' where col is an INT, BOOLEAN, TIMESTAMP, JSONB, etc. rather than VARCHAR or VARCHAR[].

Common situations: Marking a numeric or timestamp column for full-text search by mistake; expecting automatic text extraction from non-string types.

Understand the failure class

Background: Type mismatch errors: IllegalArgumentException, TypeError and type guards across 150 open-source libraries — this error's family across 150 libraries.

Related errors


AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11). Data as JSON: /api/errors/3596994373ef00f7. Report an issue: GitHub.

Appendix: source

Thrown at src/connector/src/sink/turbopuffer.rs:791

    Ok(columns)
}

fn build_turbopuffer_schema(
    schema: &Schema,
    attribute_indices: &[usize],
    full_text_search_columns: &HashSet<String>,
    filterable_columns: &HashSet<String>,
) -> Result<Value> {
    let mut result = Map::new();
    for index in attribute_indices {
        let field = &schema[*index];
        let mut config = Map::new();
        let data_type = field.data_type();
        let turbopuffer_type = turbopuffer_type(&data_type)?;
        let is_vector = matches!(data_type, DataType::Vector(_));
        let is_full_text_search = full_text_search_columns.contains(&field.name);
        if is_full_text_search && !supports_full_text_search(&data_type) {
            return Err(SinkError::Config(anyhow!(
                "Turbopuffer full_text_search column '{}' must be string or []string",
                field.name
            )));
        }
        config.insert("type".to_owned(), Value::String(turbopuffer_type));
        if filterable_columns.contains(&field.name) {
            config.insert("filterable".to_owned(), Value::Bool(true));
        }
        if is_full_text_search {
            config.insert("full_text_search".to_owned(), Value::Bool(true));
        }
        if is_vector {
            config.insert("ann".to_owned(), Value::Bool(true));
        }
        result.insert(field.name.clone(), Value::Object(config));
    }
    Ok(Value::Object(result))
}

View on GitHub (pinned to 6469eb736d)