Source beta — 2026-09-24, revision 10

This is a local-first Linux application source release with a static HF landing page. It is not a model release or an online OCR service.

Revision 10: resource admission and Quick mode

Revision 9: community ports note

Revision 8: image layout

Revision 7: screenshots and README order

Revision 6: wording only

Revision 5: updated model-comparison evidence

Revision 4: process cleanup and connector response handling

Revision 3: optional vision API connector

Revision 3 validation scope

Standard model and revised distribution pages

Optional 26B: historical observations, not a guarantee

A newer same-day test of the shipped c1_ocr_v3 pipeline used 100 bilingual synthetic cases x three seeds: exact corrections were 165/300 (55.0%) for 12B versus 74/300 (24.7%) for 26B. Thus the older six-image observations below do not establish a general quality advantage for 26B. This supports keeping 12B+C1 as the default. Repeated short synthetic cases are not a real-book accuracy guarantee; see BENCHMARKS.md for errors, timings, limitations and the separately labeled unshipped-prompt follow-up.

The standard remains 12B+C1. For users with more available VRAM, 26B+C1 is optional. An earlier six-image comparison recorded 52.291 seconds for 12B and 40.022 seconds for 26B across 18 review requests. Exact-match corrections were 51/54 versus 45/54; allowing colon/adjacent-space differences gave 51/54 versus 54/54. Neither arm changed an originally correct field in that test.

These are repeated fields on six synthetic images with an earlier C1 prompt, not a current c1ocrv3 end-to-end speed or quality evaluation. More VRAM use can force sequential execution, making full jobs slower. See BENCHMARKS.md and its numeric data for scope and tradeoffs.

Packaging changes

Revision 2 validation

Revision 1 validation scope

Limitations

Full end-to-end installation, model download and GPU inference on a second clean physical machine are not established by source packaging. Browser layouts, driver versions and GPUs vary. Tests with generated pages do not prove OCR or correction accuracy. No universal multilingual accuracy claim is made. The scheduler is not proven optimal; 26B is optional.

Version pins and the Python 3.10 constraints are not artifact-hash locks or a complete license SBOM. Exact weight licenses and intended commercial eligibility need separate review. Neither HF publication nor legal clearance is performed by the builder.


Download original text