Skip to content

Document upload and validation

Document upload and validation

api/views/upload_processing_view.py is the request-time entrypoint for document processing. It does not extract data itself; instead it validates the request, resolves configuration, persists enough state for the worker, and queues process_document_upload_async.

What the endpoint owns

  • pydantic validation of the multipart request payload
  • optional use-case resolution via usecase_uid
  • country resolution and Country existence checks
  • document type lookup
  • schema selection and prompt injection when a schema is present
  • upload file acceptance rules returned to the frontend in the GET metadata response
  • logging and error codes for invalid requests

Important request-time invariants

Use-case precedence

When usecase_uid is supplied, the use-case can provide defaults for missing fields, but direct request values always win. If the resolved use-case is incomplete for required context, the request fails before any worker dispatch happens.

Country is mandatory unless resolved indirectly

country_id must be present directly or supplied by the resolved use-case. If neither source provides it, the endpoint returns a COUNTRY_ID_MISSING error.

Schema selection affects the prompt

When a schema is present, the endpoint builds a schema-aware prompt and a field-name list that the worker later uses to constrain extraction.

The upload response is asynchronous

The caller receives a task identifier and document identifier. The heavy work is deferred to Celery so the browser connection is not tied to extraction duration.

Request-time flow

flowchart TD
A[POST /api/v1/process-document/upload/] --> B[pydantic request validation]
B --> C{usecase_uid provided?}
C -->|yes| D[resolve active use case]
C -->|no| E[use direct request values]
D --> F{required context complete?}
F -->|no| G[400 USECASE_CONFIG_INCOMPLETE]
F -->|yes| E
E --> H[resolve country and document type]
H --> I[select schema / build prompt]
I --> J[create ProcessingDocument]
J --> K[enqueue Celery task]
K --> L[return task_id and document_id]

Dependencies

  • api.validators for request schema enforcement
  • api.constants.file_types for upload limits and supported file types
  • api.services.upload_processing.upload_handler for document-type and use-case resolution
  • api.services.schema_prompt_builder for schema-driven prompts
  • api.services.upload_processing.tasks.process_document_upload_async for background processing

Validation focus

  • tests/unit/processing/test_upload_validation.py — request validation and error shapes.
  • tests/unit/services/test_upload_sync_fields.py — synchronization of request fields and use-case defaults.
  • tests/integration/test_document_processing.py — the request contract for the upload endpoint.

Scope boundary

This page covers request-time behavior only. For the background worker and extraction pipeline, read Services and Document processing.