
Corrected article
DocsRouter has not published a handwriting benchmark dataset or verified accuracy ranking. The current catalog does not include Claude 4.5 Opus.
Handwriting quality varies by writer, language, scan quality, age, and document layout. A model that works well for one archive may perform poorly on another.
Build a representative evaluation set
- Sample documents from the real workflow, including difficult and low-quality examples.
- Create reviewed ground truth for the fields or text that matter.
- Test model IDs returned by
GET /v1/modelswith the same prompt and inputs. - Measure field-level errors and review burden, not a generic marketing score.
- Keep human review for consequential medical, legal, or financial use.
anthropic/claude-sonnet-4 is one supported option, but its catalog description is guidance rather than a measured promise. DocsRouter does not automatically route handwriting to it or retry another model when confidence is low.