zylon-ai/private-gpt · error · InvalidFileError

zpgt.ingest.no_valid_nodes.error

zpgt.ingest.no_valid_nodes.error

Error message

zpgt.ingest.no_valid_nodes.error

What it means

InvalidFileError raised when extended_index.insert()/ainsert() returns an empty result — the document parsed into nodes, but the vector store inserted none of them. Carries the IngestionLoadErrors.NO_VALID_NODES code (zpgt.ingest.no_valid_nodes.error) and the warnings list. At this point the index status is not advanced to POPULATED, so the collection stays empty for this artifact.

Solutions

  1. Check the warnings attached to the error and the parse result — they usually name why nodes were dropped
  2. Verify the document actually contains extractable text (run OCR for scanned PDFs)
  3. Confirm the embedding model's dimensions match the target collection/schema
  4. Log the nodes passed to insert to see whether they were empty before reaching the store

Example fix

# before: insert returns [] -> InvalidFileError(NO_VALID_NODES)
nodes = result.nodes
inserted = extended_index.insert(nodes, notify=notify_publisher)

# after
clean = [n for n in nodes if (n.get_content() or '').strip()]
if not clean:
    raise ValueError('document has no extractable text')
inserted = extended_index.insert(clean, notify=notify_publisher)
Defensive patterns

Strategy: try-catch

Validate before calling

contentful = [n for n in nodes if (n.get_content() or '').strip()]
if not contentful:
    raise ValueError('no nodes with non-empty text; document likely lacks extractable text')

Try / catch

try:
    component._ingest(file, collection)
except InvalidFileError as e:
    if IngestionLoadErrors.NO_VALID_NODES in (e.errors or []):
        return error_response(422, 'no valid nodes produced', warnings=e.warnings)
    raise

Prevention

When it happens

Trigger: Running _ingest with use_async true or false where insert(nodes, notify=...) returns [] — e.g. all nodes filtered out by the store's transforms/embedding rules, empty text nodes, unsupported node types, or a store silently rejecting the batch.

Common situations: Documents whose extracted text is empty or whitespace-only (scanned PDFs without OCR); node transforms dropping everything; dimension mismatch between embedding output and the vector store collection; store insert hooks filtering by metadata; silent failures in a custom ExtendedIndex.

Related errors


AI-assisted analysis of zylon-ai/private-gpt@4a030776a3 (2026-08-15). Data as JSON: /api/errors/00ff53ebdf8efbc5. Report an issue: GitHub.

Appendix: source

Thrown at private_gpt/components/ingest/ingest_component.py:279

            )

            # 1. Delete previous nodes, to avoid duplicates
            self.node_store_component.delete_filtered_nodes(
                collection=collection,
                artifacts=[artifact],
            )

            # 2. Insert nodes
            inserted_nodes: Sequence[BaseNode] = []
            if use_async:
                inserted_nodes = asyncio.run(
                    extended_index.ainsert(nodes, notify=notify_publisher)
                )
            else:
                inserted_nodes = extended_index.insert(nodes, notify=notify_publisher)

            if not inserted_nodes:
                raise InvalidFileError(
                    errors=[IngestionLoadErrors.NO_VALID_NODES], warnings=warnings
                )

            index.summary = ArtifactIndexStatus.POPULATED.value
            index.set_index_id(index_id)
            index.storage_context.persist(persist_dir=local_data_path / collection)
            logger.info("Finished loading index %s with %d nodes", index_id, len(nodes))

View on GitHub (pinned to 4a030776a3)