Skip to content

Schema management

Schema management

api/views/schema_views/schema_viewset.py is the canonical home for schema management. It is not the OpenAPI schema endpoint; it is the business-object endpoint for extraction schemas.

Canonical runtime shape

SchemaViewSet composes several mixins:

  • SchemaTenantMixin — tenant isolation and scoping
  • SchemaCRUDMixin — standard CRUD behavior
  • SchemaGenerateMixin — schema generation actions
  • SchemaImportExportMixin — bulk import/export actions

The viewset uses:

  • SchemaSerializer
  • MicrosoftAuthentication plus session auth
  • IsAuthenticated
  • lookup_field = 'schema_uid'
  • get_schema_queryset() for the base queryset

That lookup_field matters because route consumers must use the schema UID, not a default numeric PK, when they retrieve a specific schema.

API Endpoint Contracts & Parameters

MethodRoute PathAction / MixinRequired HeadersRequest Payload / ParamsSuccess Response (200/201/202)Error Codes
GET/api/v1/schema/List Schemas (SchemaCRUDMixin)Authorization: Bearer <token>Query: search (str), page (int), page_size (int), tenant_id (str)200 OK
{"count": 12, "results": [{"schema_uid": "sch_123", "name": "Invoice Schema", "fields": [...]}]}
UNAUTHORIZED
INVALID_TENANT
POST/api/v1/schema/Create Schema (SchemaCRUDMixin)Authorization: Bearer <token>JSON Body:
{"name": "Invoice", "description": "...", "fields": [{"name": "total", "type": "number"}]}
201 Created
{"schema_uid": "sch_999", "name": "Invoice", "fields": [...]}
DUPLICATE_NAME
VALIDATION_ERROR
GET/api/v1/schema/<schema_uid>/Retrieve Schema (SchemaCRUDMixin)Authorization: Bearer <token>Path: schema_uid (str)200 OK
{"schema_uid": "sch_123", "name": "Invoice Schema", "fields": [...]}
SCHEMA_NOT_FOUND
PUT/PATCH/api/v1/schema/<schema_uid>/Update Schema (SchemaCRUDMixin)Authorization: Bearer <token>Path: schema_uid (str)
JSON Body: partial/full schema fields
200 OK
{"schema_uid": "sch_123", "updated_at": "..."}
SCHEMA_NOT_FOUND
VALIDATION_ERROR
DELETE/api/v1/schema/<schema_uid>/Delete Schema (SchemaCRUDMixin)Authorization: Bearer <token>Path: schema_uid (str)204 No ContentSCHEMA_NOT_FOUND
POST/api/v1/schema/generate/Generate Schema AI (SchemaGenerateMixin)Authorization: Bearer <token>Multipart Form Data:
uploaded_file: binary
country_id: int
model_id: int
202 Accepted
{"job_id": "job_abc123", "status": "PROCESSING"}
LLM_GENERATION_FAILED
UNSUPPORTED_FILE_TYPE
POST/api/v1/schema/import_schemas/Import Schemas (SchemaImportExportMixin)Authorization: Bearer <token>Multipart or JSON:
{"file": "backup.json"}
200 OK
{"imported_count": 5, "errors": []}
INVALID_IMPORT_FORMAT
GET/api/v1/schema/export_schemas/Export Schemas (SchemaImportExportMixin)Authorization: Bearer <token>Query: schema_uids (comma-separated UIDs)200 OK (JSON file attachment download)EXPORT_FAILED

Why the page exists separately

Schema management is a control plane for document extraction, not a simple configuration table. It influences:

  • upload-time prompt injection
  • processing-time field validation
  • task status payloads
  • tenant-specific visibility
  • import/export workflows for admin users

Data and dependency flow

flowchart TD
A[Schema request] --> B[SchemaViewSet]
B --> C[SchemaTenantMixin]
B --> D[SchemaCRUDMixin]
B --> E[SchemaGenerateMixin]
B --> F[SchemaImportExportMixin]
C --> G[get_schema_queryset]
D --> H[SchemaSerializer]
E --> I[Schema generation helpers]
F --> J[Import/export service]

Invariants

  • schema records are tenant scoped unless a specific admin path says otherwise
  • generation/import/export behavior is layered onto the same viewset, not split into separate endpoints
  • the schema identifier exposed through routing is schema_uid
  • upload and processing code can depend on schema field names, so changes here can affect document extraction behavior indirectly

Validation

  • tests/unit/schema/test_schema_generator.py
  • tests/unit/schema/test_schema_import_export.py
  • tests/unit/schema/test_schema_versioning.py
  • tests/integration/test_schema_selection_integration.py
  • tests/unit/schema/test_schema_generator_endpoints.py

Scope boundary

This page covers schema management only. For the document pipeline that consumes schemas, read Document processing. For the service that builds prompts from schemas, read Services.