Skip to content

Document processing

Document processing

Document handling is split across HTTP entrypoints and background service code. The canonical lifecycle is:

  1. upload and validate the file
  2. persist a document record and storage artifact
  3. enqueue background processing
  4. poll task status or document status
  5. optionally edit extracted fields and record audit history

The route families that implement this lifecycle are:

  • api/views/upload_processing_view.py — upload-and-queue entrypoint
  • api/views/document_processing_views/status_views.py — status and metaprocessing read endpoints
  • api/views/task_status_view.py — polling endpoint for background tasks
  • api/views/document_views.py — list/detail/edit semantics for processing documents
  • api/services/upload_processing/* — the queue-side orchestration
  • api/services/pdf_processing_service/* — PDF and OCR extraction

Workflow shape

sequenceDiagram
participant Client
participant Upload as DocumentUploadAPIView
participant Storage as upload/storage helpers
participant DB as ProcessingDocument
participant Celery as process_document_upload_async
participant Status as TaskStatusAPIView
participant DocStatus as DocumentStatusAPIView
participant Edit as ProcessingDocumentViewSet
Client->>Upload: POST file + prompt/schema/usecase
Upload->>Storage: validate, store, compute config
Upload->>DB: create processing document
Upload->>Celery: enqueue async task
Client->>Status: poll task_id
Status->>DB: read source-of-truth row
Celery->>DB: update status, extracted_json_data, error_message
Client->>DocStatus: GET doc status
Client->>Edit: PATCH extracted fields when allowed
Edit->>DB: save edit_history and audit data

Upload and queue boundary

DocumentUploadAPIView is responsible for the request-time contract:

  • validates request data with pydantic models
  • resolves use-case configuration when present
  • resolves country and document type requirements
  • injects schema-derived prompt content when applicable
  • creates the processing record and dispatches the Celery task after commit

That endpoint does not perform the heavy extraction itself. It prepares the worker inputs and exits quickly with a task identifier.

Status endpoints

TaskStatusAPIView is the async polling surface. It reads the persisted ProcessingDocument row for the request user and maps status to HTTP behavior:

  • 202 while processing
  • 200 with extracted data when completed or verification is needed
  • 200 with error metadata on failure
  • 404 when the task is unknown or belongs to another user

DocumentStatusAPIView is the document-centric status endpoint. It adds ownership checks and can return the extracted payload when the document is complete.

MetaprocessingAPIView exposes file-to-schema metadata for the file-processing path.

Editable extraction and audit semantics

ProcessingDocumentViewSet provides the main CRUD surface for processing documents. Its partial_update behavior is the critical business invariant:

  • it computes diffs between old and new extracted JSON data
  • it preserves and accumulates edit_history
  • it overrides edited_by from the authenticated user instead of trusting the client payload
  • it enforces one-time edit rules for completed documents
  • it preserves existing history when a patch omits edit_history

Those invariants are exercised by the api/tests regression suite.

Failure modes

  • Upload validation can fail before task creation.
  • Use-case resolution can fail if the referenced configuration is incomplete or inactive.
  • Status endpoints can return permission errors when the caller does not own the document.
  • The worker can mark the document failed when the file disappears or a downstream service errors.

Representative tests

  • api/tests/test_document_edit_audit.py — diffing, one-time patch enforcement, and audit creation.
  • api/tests/test_edit_history.py — persistence and response shaping for edit_history.
  • tests/integration/test_document_processing.py — request/response behavior of the upload and status paths.
  • tests/integration/test_async_document_processing.py — async queue and poll flow.
  • tests/unit/processing/test_upload_validation.py — upload validation behavior.

Scope boundary

This page documents the lifecycle and invariants. For upload storage details, read File management and serving. For the service pipeline that performs OCR and LLM extraction, read Services.