OCR TEST LAB

Build OCR fixtures. Measure what changed.

Build a marked image pack, run it through your own OCR system and bring the transcriptions back. Compare two versions on the same cases without using personal documents.

FICTIONAL WORKSHEETS ONLY Local processing · No uploads · No OCR API calls

1. Build the test images
Image conditions
Fictional OCR worksheet, Clean reference, with four marked sample fields.

ZIP includes images, matching text, a manifest and a transcription template. Up to 15 sheets. No engine results are supplied.

Ground-truth text for this sheet
FICTIONAL OCR TEST SHEET
NOT VALID / NO IDENTITY EVIDENCE
CASE sheet-1-clean / LAYOUT stacked
RECORD
TEST-OCR-06559
FICTIONAL / NOT VALID
SAMPLE NAME
Sample Person 1
FICTIONAL / NOT VALID
TEST EMAIL
[email protected]
FICTIONAL / NOT VALID
BATCH
QA-1042-1
FICTIONAL / NOT VALID
NOT VALID / NO IDENTITY EVIDENCE

3 fictional sheets ready. Download a pack, run your OCR, then compare the transcriptions.

2. Compare your OCR outputs

Run the images in your own OCR system. Paste the completed transcription array. null means not run; an empty string means an observed empty result.

The example is an edited transcription, not a measured OCR run. Inputs stay in page memory; reload clears them.

Regression reportNOT AUTHENTICATION
A comparison needs observed text.

Download the images, run both OCR versions and fill the template. No accuracy or pass result is assumed for an unrun case.

A fixture pack and a result, kept together

  1. Choose a seed, layout, base-sheet count and image conditions.
  2. Download the ZIP. It contains image files, exact text, the manifest and a transcription template.
  3. Run the same images through your baseline and candidate OCR configurations.
  4. Replace null in the template with each observed text string. Keep null for a run you did not perform.
  5. Compare the completed template and save the report with the OCR version and configuration.

No OCR engine runs on this page. A scoring example deliberately changes one character so you can inspect the calculation. It is not a measured recognition result.

What changes in the image

Controlled image conditions
ConditionChangeUse
CleanOriginal abstract worksheetReference input
Soft blur1.1-pixel Gaussian blur on the field groupSensitivity to softened characters
SkewField group rotated by two degreesReading and orientation behavior
Low contrastField text changes to #959595Sensitivity to faint print
Grain650 seeded specklesSensitivity to this synthetic noise pattern

The settings are controlled perturbations, not a simulation of every phone, scanner or printing process. Header and footer retain explicit non-valid wording. There are no portraits, signatures, barcodes, issuer logos or government layouts. Tesseract: image-quality considerations

The same seed, layout and settings reproduce the SVG source and text under fakekit-ocr-v1. PNG pixels depend on your browser, operating system and local Arial/Helvetica font fallback. Record that environment when comparing image files. Each pack contains 1–3 base sheets and 1–5 conditions, up to 15 images at 1000 × 740 pixels.

Read error rates with their denominator

CER is the Levenshtein edit distance between reference and observed grapheme clusters divided by the reference character count. WER uses whitespace-separated tokens and the reference token count. Both count insertion, deletion and substitution costs. Lower is better. Rates can exceed 100% when the output inserts enough extra text. OCR-D: text evaluation and error-rate definitions

Text is normalized to Unicode NFC and CRLF/CR line endings become LF. The whitespace option also collapses whitespace runs and trims the edges; preserve-whitespace mode keeps spaces and newlines. Case and punctuation are preserved. The word metric retains punctuation in its whitespace-separated tokens. These conventions differ from OCR-D's normalized error-rate and word-boundary choices; this report does not claim OCR-D evaluation compliance.

Paired aggregate rates divide total edits by total reference units across cases with both outputs, rather than averaging per-case percentages. Missing runs are reported separately. null means not run; an empty string is a completed run with no text and is scored as such. A regression means candidate CER increased on that case; it is not a verdict on every capability of an engine.

A small regression set, not a production accuracy claim

The scorer accepts a JSON array up to 128 KiB, known unique case IDs and no extra keys. Both output keys are required. Each output is limited to 8000 UTF-16 code units and 2000 grapheme clusters; a comparison work limit prevents oversized batches from freezing the page. The fixture ZIP is limited to 20 MB. Reports contain scores and fixture metadata, not your observed transcription strings.

A low error rate on these fictional sheets does not establish accuracy on genuine identity documents, handwriting, photographs, multilingual material or your production distribution. It does not measure field extraction precision/recall, document authenticity or identity proofing. Representative labelled data and a defined evaluation method are still needed for those questions. Google Cloud: document extraction evaluation