ruvnet/ruflo · error · Error
Invalid embedding value at index
Error message
Invalid embedding value at index ${i}: expected finite number, got ${typeof embedding[i]} What it means
formatEmbedding() in the ruvector import command validates every element of each embedding array is a finite number before interpolating the array into a PostgreSQL vector literal '[...]'::ruvector(N). null, strings, NaN, or Infinity entries throw with the offending index — this is SQL-injection hardening for JSON-sourced data (CRIT class fix).
Solutions
- Preprocess the dataset: drop or impute non-numeric elements so every value is a finite number
- In Python, re-export with json.dumps(data, allow_nan=False) and convert tensors via .tolist() on float32
- Verify programmatically before import: rows.every(v => typeof v === 'number' && Number.isFinite(v))
Example fix
// before
{"id":"d1","embedding":[0.1,null,"0.3"]}
// after
{"id":"d1","embedding":[0.1,0.0,0.3]} Defensive patterns
Strategy: type-guard
Validate before calling
for (const row of rows) {
if (!isFiniteVector(row.embedding)) {
row.embedding = row.embedding.map((v) => (typeof v === 'number' && Number.isFinite(v)) ? v : 0);
}
}
// only then run the import Type guard
function isFiniteVector(v: unknown): v is number[] {
return Array.isArray(v) && v.every((x) => typeof x === 'number' && Number.isFinite(x));
} Try / catch
try {
await importRows(rows);
} catch (err) {
if (err instanceof Error && /Invalid embedding value at index \d+/.test(err.message)) {
// sanitize or drop the offending row, then retry the batch
} else throw err;
} Prevention
- Validate vectors with isFiniteVector() at export time, not import time
- Serialize with strict JSON (allow_nan=False in Python) so NaN/Infinity never reach the file
- Log row ids alongside validation failures so bad rows can be traced back
When it happens
Trigger: Importing a JSON/JSONL dataset whose embedding arrays contain null (Python None), "0.25" as a string, or literal NaN/Infinity written by a non-strict serializer; also mixed-shape rows from truncated exports.
Common situations: Exporting embeddings from Python with missing values; json.dumps with allow_nan=True (default) emitting NaN; schema drift where the producer changed element types between runs.
Related errors
- Invalid schema name: " ". Must contain only letters…
- Invalid timestamp format
- Schema name must not be empty
- each record requires a non-empty numeric vector
- Embedding must be Float32Array of length
AI-assisted analysis of ruvnet/ruflo@fa13ee4ad6 (2026-08-18).
Data as JSON: /api/errors/43e182666297a5ca.
Report an issue: GitHub.
Appendix: source
Thrown at v3/@claude-flow/cli/src/commands/ruvector/import.ts:52
*/
interface ImportStats {
total: number;
imported: number;
skipped: number;
errors: number;
withEmbeddings: number;
byNamespace: Record<string, number>;
}
/**
* Format a ruvector embedding array for PostgreSQL
* Validates each element is a finite number to prevent SQL injection via crafted arrays.
*/
function formatEmbedding(embedding: number[], dimensions: number = 384): string {
// Validate every element is a finite number (prevents SQL injection via crafted JSON)
for (let i = 0; i < embedding.length; i++) {
if (typeof embedding[i] !== 'number' || !Number.isFinite(embedding[i])) {
throw new Error(`Invalid embedding value at index ${i}: expected finite number, got ${typeof embedding[i]}`);
}
}
// Ensure correct dimensions by padding or truncating
const padded = [...embedding];
while (padded.length < dimensions) {
padded.push(0);
}
if (padded.length > dimensions) {
padded.length = dimensions;
}
return `'[${padded.join(',')}]'::ruvector(${dimensions})`;
}
/**
* Escape string for PostgreSQL
*/
function escapeString(str: string): string {View on GitHub (pinned to fa13ee4ad6)