docling-project/docling · error · ValueError

Number of prompts ( ) must match number of images ( )

Error message

Number of prompts ({len(prompt)}) must match number of images ({len(pil_images)})

What it means

The Transformers VLM path requires exactly one user prompt per PIL image in the batch. A single string is broadcast; a list must satisfy len(prompt) == len(pil_images), otherwise ValueError is raised before the processor is invoked.

Solutions

  1. Use one shared string prompt when all images should get the same instruction
  2. Build prompts by iterating the same image list: [make_prompt(im) for im in images]
  3. Assert equality of both lengths at the call site in dev builds

Example fix

# before
prompts = [f"OCR page {i}" for i in range(total_pages)]
model(batch_of_10_images, prompts)  # total_pages != 10 -> ValueError
# after
prompts = [f"OCR page {p.page_no}" for p in batch_pages]
assert len(prompts) == len(batch_images)
model(batch_images, prompts)
Defensive patterns

Strategy: validation

Validate before calling

pil_images = list(images)
if isinstance(prompt, list):
    if len(prompt) != len(pil_images):
        prompt = [prompt[i] if i < len(prompt) else default_prompt for i in range(len(pil_images))]  # or fail fast

Type guard

def prompts_match(prompt: object, n: int) -> bool:
    return isinstance(prompt, str) or (isinstance(prompt, list) and len(prompt) == n)

Try / catch

try:
    model(batch, prompts)
except ValueError as e:
    if 'must match number of images' in str(e):
        model(batch, shared_prompt)
    else:
        raise

Prevention

When it happens

Trigger: Calling the model with a per-page prompt list built from a different page set than the image batch; converting one document where images were deduplicated but prompts were not (or vice versa).

Common situations: Custom chunking that re-batches images after prompts were generated; conditionally dropping blank pages from images while keeping all prompts; off-by-one when slicing pages.

Related errors


AI-assisted analysis of docling-project/docling@61d76f1ff3 (2026-08-14). Data as JSON: /api/errors/116d071f550e71f6. Report an issue: GitHub.

Appendix: source

Thrown at docling/models/vlm_pipeline_models/hf_transformers_model.py:281

                elif img.ndim == 2:
                    pil_img = PILImage.fromarray(img.astype(np.uint8), mode="L")
                else:
                    raise ValueError(f"Unsupported numpy array shape: {img.shape}")
            else:
                pil_img = img
            if pil_img.mode != "RGB":
                pil_img = pil_img.convert("RGB")
            pil_images.append(pil_img)

        if not pil_images:
            return

        # -- Normalize prompts (1 per image)
        if isinstance(prompt, str):
            user_prompts = [prompt] * len(pil_images)
        else:
            if len(prompt) != len(pil_images):
                raise ValueError(
                    f"Number of prompts ({len(prompt)}) must match number of images ({len(pil_images)})"
                )
            user_prompts = prompt

        # Use your prompt formatter verbatim
        if self.vlm_options.transformers_prompt_style == TransformersPromptStyle.NONE:
            inputs = self.processor(
                pil_images,
                return_tensors="pt",
                padding=True,  # pad across batch for both text and vision
                **self.vlm_options.extra_processor_kwargs,
            )
        else:
            prompts: list[str] = [self.formulate_prompt(p) for p in user_prompts]

            # -- Processor performs BOTH text+image preprocessing + batch padding (recommended)
            inputs = self.processor(
                text=prompts,

View on GitHub (pinned to 61d76f1ff3)