Processing guidance
Repeatable results, large PDFs and structured document metadata.
Repeatable results
Repeated submissions of identical file bytes within the same organization, document type and processing version reuse the completed result. This keeps extracted data, authenticity scores and risk levels consistent for repeat uploads. Reuse applies only after the earlier result has completed and is eligible for reuse; changing a filename alone does not change the file bytes.
Authenticity scores and risk levels are calculated from check evidence using fixed scoring rules. Fresh model runs can still vary, so deterministic scoring alone does not guarantee identical evidence. A new processing version or a changed file can produce a new result. Keep the document ID and processing provenance with any decision that depends on a result.
Large PDFs
The upload limit is 100 MB, with no separate page-count limit. Processing time and extraction quality depend on document length, scan quality and layout.
- Send one account and statement period per document where possible.
- If a document needs splitting, split at statement-period boundaries and submit the parts through batch processing. Each part has its own document ID, result and webhook.
- Prefer the original native PDF over a scan or photograph. Keep headers and transaction continuations together when splitting.
- Use polling or webhooks to wait for completion. A missing or unreadable field is not evidence of a zero value.
Structured document metadata
For document types with canonical document metadata, available extracted values are mapped into extracted_data.document_meta, including issue_date, issuer_name, document_number and issuer_address. Existing dotted keys such as document_meta.issue_date remain available for compatibility. Fields that are absent or cannot be read reliably may remain null. Bank-statement account and statement-period metadata use the bank-statement metadata schema; not every document type has the same metadata fields.

