vectordotdev/vector · error
{}
Error message
{} What it means
The Parquet schema generator infers an Arrow schema from a batch of JSON events using the JSON schema inference utility. This error wraps any error returned by that inference (e.g. conflicting types for the same field across events) into an io::Error, with the message being the underlying error's display text (the "{}" template). A SchemaGenerationError event is emitted before returning.
Solutions
- Normalize field types across events before inference (coerce to a single type per field).
- Pre-validate the event batch for type conflicts and drop or cast conflicting records.
- Increase sampling/handling so nulls are handled and only mutually compatible events are used for schema inference.
Example fix
// before
[{"count": 1}, {"count": "many"}]
// after
[{"count": 1}, {"count": 2}] // or coerce all to strings Defensive patterns
Strategy: try-catch
Validate before calling
// detect type conflicts in events before schema inference
fn has_conflicting_types(events: &[serde_json::Value], key: &str) -> bool {
events.iter().filter_map(|e| e.get(key))
.map(|v| match v { serde_json::Value::Number(_) => "num", serde_json::Value::String(_) => "str", _ => "other" })
.collect::<std::collections::HashSet<_>>().len() > 1
} Try / catch
match generator.infer_schema(&events) {
Ok(schema) => schema,
Err(e) => {
tracing::warn!("schema inference failed: {e}; normalizing types and retrying");
let normalized = coerce_event_types(&events);
generator.infer_schema(&normalized)?
}
} Prevention
- Coerce numeric-like strings to numbers (or vice versa) upstream.
- Filter out null/empty events before inference.
- Pin producer field types with a schema contract.
When it happens
Trigger: Calling infer_schema with events whose inferred JSON schema conflicts — e.g. one event has field 'count' as a number and another has it as a string, or nested structures disagree on types.
Common situations: Heterogeneous log data feeding a parquet sink without type normalization; schema drift after a producer update; optional fields that sometimes appear as null and sometimes as values.
Understand the failure class
Background: Schema validation failed / invalid input schema: payload rejected because its shape doesn't match the expected schema — this error's family across 28 libraries.
Related errors
- Invalid Map schema for field
- already found inputs
- `Configurable` does not support numbers larger than an…
- {}
- enums must always have a tagging mode
AI-assisted analysis of vectordotdev/vector@bdb87aeaa4 (2026-09-16).
Data as JSON: /api/errors/05f2ea1fb4fdca6e.
Report an issue: GitHub.
Appendix: source
Thrown at lib/codecs/src/encoding/format/parquet.rs:382
}
let record_batch =
build_record_batch(Arc::clone(&self.schema), &json_values).map_err(Box::new)?;
Self::write_record_batch(&record_batch, buffer, &self.writer_props).map_err(Box::new)?;
Ok(())
}
}
pub struct ParquetSchemaGenerator {}
impl ParquetSchemaGenerator {
pub fn infer_schema(events: &[serde_json::Value]) -> Result<Schema, Error> {
let schema = infer_json_schema_from_iterator(events.iter().map(Ok::<_, ArrowError>))
.map_err(|e| {
emit(SchemaGenerationError { error: &e });
Error::new(ErrorKind::InvalidData, e.to_string())
})?;
Ok(schema)
}
/// Attempt to modify schema to set timestamp fields as Timestamp instead of Utf8.
/// Only works for top-level fields.
fn try_normalize_schema(events: &[Event], schema: Schema) -> Schema {
let mut ts_seen: HashSet<String> = HashSet::new();
let mut non_ts_seen: HashSet<String> = HashSet::new();
for event in events.iter().filter_map(Event::maybe_as_log) {
if let Some(object_map) = event.as_map() {
for (path, value) in object_map {
if value.is_timestamp() {
ts_seen.insert(path.to_string());
} else if !value.is_null() {
non_ts_seen.insert(path.to_string());View on GitHub (pinned to bdb87aeaa4)