Document processing
Document processing
Document handling is split across HTTP entrypoints and background service code. The canonical lifecycle is:
- upload and validate the file
- persist a document record and storage artifact
- enqueue background processing
- poll task status or document status
- optionally edit extracted fields and record audit history
The route families that implement this lifecycle are:
api/views/upload_processing_view.py— upload-and-queue entrypointapi/views/document_processing_views/status_views.py— status and metaprocessing read endpointsapi/views/task_status_view.py— polling endpoint for background tasksapi/views/document_views.py— list/detail/edit semantics for processing documentsapi/services/upload_processing/*— the queue-side orchestrationapi/services/pdf_processing_service/*— PDF and OCR extraction
Workflow shape
sequenceDiagram participant Client participant Upload as DocumentUploadAPIView participant Storage as upload/storage helpers participant DB as ProcessingDocument participant Celery as process_document_upload_async participant Status as TaskStatusAPIView participant DocStatus as DocumentStatusAPIView participant Edit as ProcessingDocumentViewSet
Client->>Upload: POST file + prompt/schema/usecase Upload->>Storage: validate, store, compute config Upload->>DB: create processing document Upload->>Celery: enqueue async task Client->>Status: poll task_id Status->>DB: read source-of-truth row Celery->>DB: update status, extracted_json_data, error_message Client->>DocStatus: GET doc status Client->>Edit: PATCH extracted fields when allowed Edit->>DB: save edit_history and audit dataUpload and queue boundary
DocumentUploadAPIView is responsible for the request-time contract:
- validates request data with pydantic models
- resolves use-case configuration when present
- resolves country and document type requirements
- injects schema-derived prompt content when applicable
- creates the processing record and dispatches the Celery task after commit
That endpoint does not perform the heavy extraction itself. It prepares the worker inputs and exits quickly with a task identifier.
Status endpoints
TaskStatusAPIView is the async polling surface. It reads the persisted ProcessingDocument row for the request user and maps status to HTTP behavior:
202while processing200with extracted data when completed or verification is needed200with error metadata on failure404when the task is unknown or belongs to another user
DocumentStatusAPIView is the document-centric status endpoint. It adds ownership checks and can return the extracted payload when the document is complete.
MetaprocessingAPIView exposes file-to-schema metadata for the file-processing path.
Editable extraction and audit semantics
ProcessingDocumentViewSet provides the main CRUD surface for processing documents. Its partial_update behavior is the critical business invariant:
- it computes diffs between old and new extracted JSON data
- it preserves and accumulates
edit_history - it overrides
edited_byfrom the authenticated user instead of trusting the client payload - it enforces one-time edit rules for completed documents
- it preserves existing history when a patch omits
edit_history
Those invariants are exercised by the api/tests regression suite.
Failure modes
- Upload validation can fail before task creation.
- Use-case resolution can fail if the referenced configuration is incomplete or inactive.
- Status endpoints can return permission errors when the caller does not own the document.
- The worker can mark the document failed when the file disappears or a downstream service errors.
Representative tests
api/tests/test_document_edit_audit.py— diffing, one-time patch enforcement, and audit creation.api/tests/test_edit_history.py— persistence and response shaping foredit_history.tests/integration/test_document_processing.py— request/response behavior of the upload and status paths.tests/integration/test_async_document_processing.py— async queue and poll flow.tests/unit/processing/test_upload_validation.py— upload validation behavior.
Scope boundary
This page documents the lifecycle and invariants. For upload storage details, read File management and serving. For the service pipeline that performs OCR and LLM extraction, read Services.