A fixture pack and a result, kept together
- Choose a seed, layout, base-sheet count and image conditions.
- Download the ZIP. It contains image files, exact text, the manifest and a transcription template.
- Run the same images through your baseline and candidate OCR configurations.
- Replace null in the template with each observed text string. Keep null for a run you did not perform.
- Compare the completed template and save the report with the OCR version and configuration.
No OCR engine runs on this page. A scoring example deliberately changes one character so you can inspect the calculation. It is not a measured recognition result.
What changes in the image
| Condition | Change | Use |
|---|---|---|
| Clean | Original abstract worksheet | Reference input |
| Soft blur | 1.1-pixel Gaussian blur on the field group | Sensitivity to softened characters |
| Skew | Field group rotated by two degrees | Reading and orientation behavior |
| Low contrast | Field text changes to #959595 | Sensitivity to faint print |
| Grain | 650 seeded speckles | Sensitivity to this synthetic noise pattern |
The settings are controlled perturbations, not a simulation of every phone, scanner or printing process. Header and footer retain explicit non-valid wording. There are no portraits, signatures, barcodes, issuer logos or government layouts. Tesseract: image-quality considerations
The same seed, layout and settings reproduce the SVG source and text under fakekit-ocr-v1. PNG pixels depend on your browser, operating system and local Arial/Helvetica font fallback. Record that environment when comparing image files. Each pack contains 1–3 base sheets and 1–5 conditions, up to 15 images at 1000 × 740 pixels.
Read error rates with their denominator
CER is the Levenshtein edit distance between reference and observed grapheme clusters divided by the reference character count. WER uses whitespace-separated tokens and the reference token count. Both count insertion, deletion and substitution costs. Lower is better. Rates can exceed 100% when the output inserts enough extra text. OCR-D: text evaluation and error-rate definitions
Text is normalized to Unicode NFC and CRLF/CR line endings become LF. The whitespace option also collapses whitespace runs and trims the edges; preserve-whitespace mode keeps spaces and newlines. Case and punctuation are preserved. The word metric retains punctuation in its whitespace-separated tokens. These conventions differ from OCR-D's normalized error-rate and word-boundary choices; this report does not claim OCR-D evaluation compliance.
Paired aggregate rates divide total edits by total reference units across cases with both outputs, rather than averaging per-case percentages. Missing runs are reported separately. null means not run; an empty string is a completed run with no text and is scored as such. A regression means candidate CER increased on that case; it is not a verdict on every capability of an engine.
A small regression set, not a production accuracy claim
The scorer accepts a JSON array up to 128 KiB, known unique case IDs and no extra keys. Both output keys are required. Each output is limited to 8000 UTF-16 code units and 2000 grapheme clusters; a comparison work limit prevents oversized batches from freezing the page. The fixture ZIP is limited to 20 MB. Reports contain scores and fixture metadata, not your observed transcription strings.
A low error rate on these fictional sheets does not establish accuracy on genuine identity documents, handwriting, photographs, multilingual material or your production distribution. It does not measure field extraction precision/recall, document authenticity or identity proofing. Representative labelled data and a defined evaluation method are still needed for those questions. Google Cloud: document extraction evaluation
Editable formats and safe test layouts
Understand PSD searches without downloading a counterfeit template.
Explore resourceCompare structured extraction output
Text error rates and JSON field changes answer different questions.
Explore resourceKeep testing clearly fictional
Use the fixtures only in an authorized test environment.
Explore resource