Scribe OCR
Scribe OCR by Balearica extracts text from scans in your browser. Review local OCR setup, exports, privacy boundaries and two prepared, unexecuted label tests.
On this page
What is Scribe OCR?
Scribe OCR is a specialist ai tool from Balearica for Document text extraction (OCR). Scribe OCR is Balearica's free open-source browser application for recognizing text in images and scanned documents, checking OCR against the original page, and exporting text or searchable PDFs. A person, creator or small shop can import a label, scan or PDF, run the built-in browser OCR, review quantities, dimensions, prices and missing fields, then download the result. The interface also supports editing existing OCR data and creating digital document versions. Balearica's public author profile explicitly names creation of scribeocr.com; current team size, legal operator and control remain unknown. The two synthetic English label cases are prepared but unexecuted while the specific image-import permission is pending. Optional AWS Textract is a separate cloud mode with credentials and potential charges.
Best suited for
- People digitizing owned scans, creators proofreading print material, and small shops copying label facts into usable text while checking the original image. The built-in browser route suits users who want an editable OCR interface and native exports without configuring a cloud OCR account.
- A pilot focused on original english product-label png to unedited first native txt and searchable pdf, using owned or authorized PNG/JPEG document images or PDFs. Optional existing character-level Tesseract HOCR or Abbyy XML can be supplied for proofreading. The proposed pilot uses only two original synthetic English product-label PNGs; their exact-reference text files are evaluator-only and must not be imported as OCR data or ground truth.
Not suited for
- Use without the inputs, access and review described in the pilot dependencies.
- Both original synthetic label cases are prepared and unexecuted. There is no pass/fail result, recognized output, export, quality score or runtime performance measurement.
- The specific two-PNG import permission is pending. Prepared inputs, an empty UI and source GETs are not product execution.
Capabilities, with sources
- 01The official GUI README describes a free open-source web application for recognizing text, proofreading OCR data and creating digital documents.Official vendor statement · checked 2026-10-03Source ↗
- 02The official Getting Started page instructs users to import PNG/JPEG scans or a PDF, then use Recognize Text in the complete browser interface.Official vendor statement · checked 2026-10-03Source ↗
- 03The official FAQ describes built-in recognition as running locally in the browser and supports common Latin-script languages, with Simplified Chinese marked experimental.Official vendor statement · checked 2026-10-03Source ↗
- 04The native application exposes document import, word editing, layout controls and a Download menu with text and document formats.Official vendor statement · checked 2026-10-03Source ↗
- 05Balearica publicly identifies as the creator of scribeocr.com and links the same official application from the personal profile.Official vendor statement · checked 2026-10-03Source ↗
- 06The official GUI package names Balearica as author and AGPL-3.0 as its application license.Official vendor statement · checked 2026-10-03Source ↗
- 07The official Support page describes monthly snapshots and states that the application has no formal versions or releases.Official vendor statement · checked 2026-10-03Source ↗
- 08The retained native GUI source contains a separate optional AWS Textract branch with credential and charge-acceptance controls.Official vendor statement · checked 2026-10-03Source ↗
- 09The fixed language-loader source can read cached traineddata and fetch language resources from a CDN when no explicit path is supplied.Official vendor statement · checked 2026-10-03Source ↗
Inputs and outputs
Inputs
Owned or authorized PNG/JPEG document images or PDFs. Optional existing character-level Tesseract HOCR or Abbyy XML can be supplied for proofreading. The proposed pilot uses only two original synthetic English product-label PNGs; their exact-reference text files are evaluator-only and must not be imported as OCR data or ground truth.
Outputs
The documented native Download menu offers TXT, searchable PDF, HOCR, ALTO XML, DOCX, HTML, Markdown and Scribe document files; table XLSX is a separate optional feature. No recognized text, PDF, TXT or other product output has been generated in this preparation. Both planned first exports must be checked against the original image.
Ecommerce & Retail fields
| Commerce platforms | Not verifiedNot verified in the reviewed official material. |
|---|---|
| Workflow stages | document ocr text extractionSource 1 |
| Integration by platform | Not verifiedNot verified in the reviewed official material. |
| Store data & permissions | Not verifiedNot verified in the reviewed official material. |
| Output formats | The native menu lists TXT, PDF, HOCR, ALTO XML, DOCX, HTML, Markdown and Scribe documents. Actual first-output export behavior remains untested.Source 1 |
| Batch processing | Not verifiedNot verified in the reviewed official material. |
| Localization languages | Not verifiedNot verified in the reviewed official material. |
| Approval requirements | Not verifiedNot verified in the reviewed official material. |
Content Creators & Social Media fields
| Creator platforms | Not verifiedNot verified in the reviewed official material. |
|---|---|
| Content formats | Not verifiedNot verified in the reviewed official material. |
| Inputs | Owned or authorized PNG/JPEG document images or PDFs. Optional existing character-level Tesseract HOCR or Abbyy XML can be supplied for proofreading. The proposed pilot uses only two original synthetic English product-label PNGs; their exact-reference text files are evaluator-only and must not be imported as OCR data or ground truth.Source 1 |
| Outputs | Not verifiedNot verified in the reviewed official material. |
| Aspect ratios | Not verifiedNot verified in the reviewed official material. |
| Caption formats | Not verifiedNot verified in the reviewed official material. |
| Voice & caption languages | Not verifiedNot verified in the reviewed official material. |
| Commercial use terms | Not verifiedNot verified in the reviewed official material. |
| Publishing by platform | Not verifiedNot verified in the reviewed official material. |
| Approval requirements | Not verifiedNot verified in the reviewed official material. |
A practical Scribe OCR workflow
- Prepare the original english product-label png to unedited first native txt and searchable pdf fixture: SYNTHETIC PRODUCT LABEL Harbor Lamp Workshop Item: Studio notebook SKU: HL-NB-010 Quantity: 10 Size: 21 x 15 cm Unit price: USD 4.50 Total: USD 45.00 Color: blue Material: NOT SPECIFIED Do not machine wash. END OF LABEL
- Check Scribe OCR access through Complete browser UI with native document import and Download, Optional local HTTP hosting of the complete official application, Separate optional AWS Textract cloud adapter requiring user credentials and confirm the selected feature’s actual permissions.
- After the specific image-import permission is granted and the evaluation runtime is reviewed and frozen, import only the original PNG for this case into the complete native Scribe OCR UI. Keep English, basic Quality 1, Batch Mode off and Advanced Recognition off; preserve the first Recognize Text result without editing. Export the first native TXT and first native PDF from that result. Do not import this exact-reference text or any ground truth/OCR data.
- Inspect the unedited first native recognized result satisfies all six original condition definitions, with identity, values, missing information, negation, complete TXT and searchable PDF verified directly against the original PNG and evaluator-only reference. No product result currently exists. Compare it against the source input and retain the output/action log.
- Run the boundary case: SYNTHETIC PRODUCT LABEL Harbor Lamp Workshop Item: Refill pack SKU: HL-RF-003 Quantity: 3 Size: 12 x 8 cm Unit price: USD 12.40 Total: USD 37.20 Color: gray Material: NOT SPECIFIED Do not expose to heat. Delivery date: NOT PROVIDED END OF LABEL Accept the result only if all pass conditions are met and no failure condition occurs.
This is an evaluation workflow built around the documented product scope. Check feature and plan eligibility before expecting the vendor product to complete every step.
Setup and integrations
Use the full official web application at scribeocr.com. Official documentation also supports serving the complete repository through a local HTTP server; no desktop installation is provided in the retained README. Built-in recognition is described as browser-local, but code, WASM and language resources may load over the network. Optional AWS Textract sends recognition to its separate cloud route. The retained GUI source commit and Scribe.js gitlink are research references, not verified identities of the current hosted build or downloaded traineddata.. Documented access methods: Complete browser UI with native document import and Download, Optional local HTTP hosting of the complete official application, Separate optional AWS Textract cloud adapter requiring user credentials. Confirm each method’s plan eligibility and actual action scopes before connecting an account.
Access and setup steps
- Prepare a scan or document image you own or are authorized to process. Choose the document language and keep the original file available for checking recognized text.
- Open the complete Scribe OCR application and use Select Files to import a PNG/JPEG scan or PDF. Optional existing HOCR or Abbyy XML is for proofreading existing OCR; it is separate from recognizing a new image.
- In the Recognize tab, select the document language and the built-in browser recognition options, then use Recognize Text. Review document orientation and available quality settings. Optional AWS Textract is a separate cloud mode with credentials and potential charges.
- Compare the recognized text with the original page, especially names, SKU strings, quantities, dimensions, prices, units, missing fields and negation. Save the original recognition before using word editing or proofreading controls to correct errors.
- Use the native Download menu to export the required format, such as TXT or a searchable PDF. Open the downloaded file and check complete text; a PDF should preserve the page image and have the searchable text layer you expect.
- For local hosting, follow the official instructions for the complete application and review its AGPL-3.0 license. Browser-local OCR may still load code, WASM and language resources over the network; review that boundary before processing sensitive documents.
Test access: public tool. The actual empty public interface and document import control were observed with English/Quality 1/Batch off and no initial Terms prompt. Both synthetic PNG cases are prepared but unexecuted while the specific image-import permission is pending. The actual runtime settings must be reviewed and frozen before this evaluation runs; no account-free execution success or output-quality result is recorded. Open the official access or installation page ↗
Pilot dependencies
- Specific two-PNG permission is pending. The original inputs and twelve acceptance conditions are frozen at input-and-acceptance contract SHA256 3eea51beea0b73b5060b24961bb77faa509715a7d9fddd259de91436c6c0b978; actual runtime remains unfrozen. Actual visible settings must be reviewed and frozen before recognition in this evaluation. Use the complete official UI with English/basic Quality 1/Batch off/Advanced off and the observed OCR Mode/Auto-Rotate on/Optimize Fonts off(disabled) preimport state. Hidden engine/model/build/segmentation/upscale/updateConfidence are unknown. Import only the PNG, never evaluator reference or existing OCR data. No account, paid cloud mode, AWS key or charge acceptance is part of the proposed pilot.
- The actual empty public interface and document import control were observed with English/Quality 1/Batch off and no initial Terms prompt. Both synthetic PNG cases are prepared but unexecuted while the specific image-import permission is pending. The actual runtime settings must be reviewed and frozen before this evaluation runs; no account-free execution success or output-quality result is recorded.
- Confirm free built-in browser ocr; optional cloud costs separate against the current vendor terms; usage and connected-service costs can affect the pilot.
- Create a test workspace or use public/authorized material. Keep an input baseline, output artifact and action log for comparison.
Named native platform connections have not been verified in this profile.
Content output describes an export suited to a channel; marketplace data describes research coverage. Exact data scopes and permissions need a setup review.
API: Not verifiedNot verified in the reviewed official material.
Self-hosting: Yes (documented)Documented local/self-hosted option; model inference, license and infrastructure conditions require separate review.Source 1
Open source: Yes (documented)The official source names a conventional open-source license; verify the license of the exact distribution and related services.Source 1
Pricing and additional costs
Free built-in browser OCR; optional cloud costs separate
The official README and documentation describe a free open-source application. The proposed pilot uses its built-in browser recognition, with no cloud credentials or checkout. Optional AWS Textract requires a user cloud account and explicit charge-risk acceptance; any cloud-provider charges are separate. Local hardware, electricity, networking and optional hosting costs were not measured.
The observed empty public interface was accessible without a login, payment or initial Terms prompt. No file has been imported or recognized, so successful free execution, current resource use and export behavior remain unverified. No AWS credentials, charge acceptance or paid call is authorized by this draft.
Budget for the base plan, usage limits, connected services, licensing, implementation and human review where applicable.
Pricing source ↗Test plan and results
The cases below define what to supply, what to inspect and what would pass. A planned case is not a completed product test.
See the testing method and all product plans →
No full vendor-account quality, latency, cost or outcome evaluation has been completed for Scribe OCR. 2 defined product cases remain unexecuted.
Current HTTP/readability checks are listed below. They establish access, not the truth of every vendor claim.
Rendering, source links and visible evaluation content need a recorded site acceptance run.
Dependencies before a product pilot
- Specific two-PNG permission is pending. The original inputs and twelve acceptance conditions are frozen at input-and-acceptance contract SHA256 3eea51beea0b73b5060b24961bb77faa509715a7d9fddd259de91436c6c0b978; actual runtime remains unfrozen. Actual visible settings must be reviewed and frozen before recognition in this evaluation. Use the complete official UI with English/basic Quality 1/Batch off/Advanced off and the observed OCR Mode/Auto-Rotate on/Optimize Fonts off(disabled) preimport state. Hidden engine/model/build/segmentation/upscale/updateConfidence are unknown. Import only the PNG, never evaluator reference or existing OCR data. No account, paid cloud mode, AWS key or charge acceptance is part of the proposed pilot.
- The actual empty public interface and document import control were observed with English/Quality 1/Batch off and no initial Terms prompt. Both synthetic PNG cases are prepared but unexecuted while the specific image-import permission is pending. The actual runtime settings must be reviewed and frozen before this evaluation runs; no account-free execution success or output-quality result is recorded.
- Confirm free built-in browser ocr; optional cloud costs separate against the current vendor terms; usage and connected-service costs can affect the pilot.
- Create a test workspace or use public/authorized material. Keep an input baseline, output artifact and action log for comparison.
Clear English label to first native TXT and searchable PDF Product case · not executed
Controlled input
SYNTHETIC PRODUCT LABEL Harbor Lamp Workshop Item: Studio notebook SKU: HL-NB-010 Quantity: 10 Size: 21 x 15 cm Unit price: USD 4.50 Total: USD 45.00 Color: blue Material: NOT SPECIFIED Do not machine wash. END OF LABEL
Request
After the specific image-import permission is granted and the evaluation runtime is reviewed and frozen, import only the original PNG for this case into the complete native Scribe OCR UI. Keep English, basic Quality 1, Batch Mode off and Advanced Recognition off; preserve the first Recognize Text result without editing. Export the first native TXT and first native PDF from that result. Do not import this exact-reference text or any ground truth/OCR data.
Steps
- Use the complete official native web application. Create a fresh document using the native UI only; keep Batch Mode off.
- Import only the one prepared PNG product input for that case. Never import its exact-reference text or any other OCR data.
- Through native Recognize tab click Recognize Text exactly once for that case after settings/input are frozen. Do not use direct Scribe.js/Tesseract calls, console state, custom wrapper, hidden API or another OCR product.
- Retain the first completion or error and original displayed OCR before any manual correction. Native Combined mode can have internal multiple engine passes; one UI click is not a claim of exactly one internal inference.
- Using the same unedited first result, export the first native TXT and the first native PDF. Native export of the existing result is separate from recognition; no second recognition click or quality retry.
- Record downloaded bytes/hash/type and compare against evaluator-only reference after first-output preservation. Exported PDF must be parsed to establish searchable text; a PDF magic header alone is insufficient.
- If access/loading/output fails, retain the actual first failure. Do not change engine, contrast, input, scoring conditions or repeat recognition to obtain a pass.
Expected output
The unedited first native recognized result satisfies all six original condition definitions, with identity, values, missing information, negation, complete TXT and searchable PDF verified directly against the original PNG and evaluator-only reference. No product result currently exists.
Observable pass conditions
- Preserve every exact required string: ["Harbor Lamp Workshop", "Item: Studio notebook", "SKU: HL-NB-010"]
- Preserve every exact required string: ["Quantity: 10", "Size: 21 x 15 cm"]
- Preserve every exact required string: ["Unit price: USD 4.50", "Total: USD 45.00"]
- Preserve every exact required string: ["Material: NOT SPECIFIED", "Do not machine wash."]
- First native TXT exists, is nonempty UTF-8, is unchanged, contains END OF LABEL, and normalized CER <= 0.02 against the exact reference.
- First native PDF parses as one page, opens with the original label image, and has extractable searchable text containing HL-NB-010, 45.00 and NOT SPECIFIED. No manual text addition or correction.
Failure conditions
- AWS Textract or any cloud-adapter mode
- API keys
- Payment
- Account creation
- Batch Mode
- Ground truth import
- Manual OCR text repair before first result evidence
- Engine or input change after recognition
- Recognition retry for quality
- If access/loading/output fails, retain the actual first failure. Do not change engine, contrast, input, scoring conditions or repeat recognition to obtain a pass.
Low-contrast rotated English label to first native TXT and searchable PDF Product case · not executed
Controlled input
SYNTHETIC PRODUCT LABEL Harbor Lamp Workshop Item: Refill pack SKU: HL-RF-003 Quantity: 3 Size: 12 x 8 cm Unit price: USD 12.40 Total: USD 37.20 Color: gray Material: NOT SPECIFIED Do not expose to heat. Delivery date: NOT PROVIDED END OF LABEL
Request
After the specific image-import permission is granted and the evaluation runtime is reviewed and frozen, import only the original PNG for this case into the complete native Scribe OCR UI. Keep English, basic Quality 1, Batch Mode off and Advanced Recognition off; preserve the first Recognize Text result without editing. Export the first native TXT and first native PDF from that result. Do not import this exact-reference text or any ground truth/OCR data.
Steps
- Use the complete official native web application. Create a fresh document using the native UI only; keep Batch Mode off.
- Import only the one prepared PNG product input for that case. Never import its exact-reference text or any other OCR data.
- Through native Recognize tab click Recognize Text exactly once for that case after settings/input are frozen. Do not use direct Scribe.js/Tesseract calls, console state, custom wrapper, hidden API or another OCR product.
- Retain the first completion or error and original displayed OCR before any manual correction. Native Combined mode can have internal multiple engine passes; one UI click is not a claim of exactly one internal inference.
- Using the same unedited first result, export the first native TXT and the first native PDF. Native export of the existing result is separate from recognition; no second recognition click or quality retry.
- Record downloaded bytes/hash/type and compare against evaluator-only reference after first-output preservation. Exported PDF must be parsed to establish searchable text; a PDF magic header alone is insufficient.
- If access/loading/output fails, retain the actual first failure. Do not change engine, contrast, input, scoring conditions or repeat recognition to obtain a pass.
Expected output
The unedited first native recognized result satisfies all six original condition definitions, with identity, values, missing information, negation, complete TXT and searchable PDF verified directly against the original PNG and evaluator-only reference. No product result currently exists.
Observable pass conditions
- Preserve every exact required string: ["Harbor Lamp Workshop", "Item: Refill pack", "SKU: HL-RF-003"]
- Preserve every exact required string: ["Quantity: 3", "Size: 12 x 8 cm"]
- Preserve every exact required string: ["Unit price: USD 12.40", "Total: USD 37.20"]
- Preserve every exact required string: ["Material: NOT SPECIFIED", "Do not expose to heat.", "Delivery date: NOT PROVIDED"]
- First native TXT exists, is nonempty UTF-8, is unchanged, contains END OF LABEL, and normalized CER <= 0.02 against the exact reference.
- First native PDF parses as one page, opens with the original rotated label image, and has extractable searchable text containing HL-RF-003, 37.20 and NOT PROVIDED. No manual text addition or correction.
Failure conditions
- AWS Textract or any cloud-adapter mode
- API keys
- Payment
- Account creation
- Batch Mode
- Ground truth import
- Manual OCR text repair before first result evidence
- Engine or input change after recognition
- Recognition retry for quality
- If access/loading/output fails, retain the actual first failure. Do not change engine, contrast, input, scoring conditions or repeat recognition to obtain a pass.
Permissions and failure boundary
- Documented access: Use the full official web application at scribeocr.com. Official documentation also supports serving the complete repository through a local HTTP server; no desktop installation is provided in the retained README. Built-in recognition is described as browser-local, but code, WASM and language resources may load over the network. Optional AWS Textract sends recognition to its separate cloud route. The retained GUI source commit and Scribe.js gitlink are research references, not verified identities of the current hosted build or downloaded traineddata.; Complete browser UI with native document import and Download, Optional local HTTP hosting of the complete official application, Separate optional AWS Textract cloud adapter requiring user credentials. Confirm the actual scopes for the selected account and plan.
- Acceptance boundary: Preserve every original field and negation in the first native recognition of the low-contrast rotated label. Missing data must remain represented by its exact source text. No manual repair, reference import, engine/input replacement or quality retry may turn a failed first result into a pass. These prepared conditions are unexecuted.
- Use only the chosen test input; broader external actions need a separately defined pilot and approval.
Official-page checks
| Source | Access status | Evidence and scope |
|---|---|---|
| Official Scribe OCR browser application | accessibleHTTP 200 · 2026-10-02T22:16:52.135427+00:00 | 139655 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Official Scribe.js library repository metadata | accessibleHTTP 200 · 2026-10-02T22:16:52.136429+00:00 | 6339 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Official Scribe.js README and supported complete GUI | accessibleHTTP 200 · 2026-10-02T22:16:52.139105+00:00 | 4661 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Official Scribe OCR GUI repository metadata | accessibleHTTP 200 · 2026-10-02T22:17:26.084685+00:00 | 6396 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Official Scribe OCR README: free browser application and local hosting | accessibleHTTP 200 · 2026-10-02T22:17:26.758282+00:00 | 4851 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Official GUI package: Balearica author and AGPL-3.0 | accessibleHTTP 200 · 2026-10-02T22:17:27.323760+00:00 | 1551 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Official getting started: import documents and recognize text | accessibleHTTP 200 · 2026-10-02T22:17:30.122681+00:00 | 12140 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Public Scribe OCR repository organization profile | accessibleHTTP 200 · 2026-10-02T22:17:31.069693+00:00 | 970 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Official language, local recognition, OCR editing and export FAQ | accessibleHTTP 200 · 2026-10-02T22:18:00.654486+00:00 | 23195 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Official support and monthly snapshot policy | accessibleHTTP 200 · 2026-10-02T22:18:01.321252+00:00 | 11815 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Official original-text editing and navigation controls | accessibleHTTP 200 · 2026-10-02T22:18:01.924937+00:00 | 17545 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Balearica public creator profile naming scribeocr.com creation | accessibleHTTP 200 · 2026-10-02T22:18:02.526630+00:00 | 1356 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Official retained GUI master commit and Scribe.js gitlink | accessibleHTTP 200 · 2026-10-02T22:18:03.125145+00:00 | 72664 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Official GUI AGPL-3.0 application license | accessibleHTTP 200 · 2026-10-02T22:18:03.867850+00:00 | 34523 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Fixed GUI source: native recognition/export and separate AWS branch | accessibleHTTP 200 · 2026-10-02T22:19:17.644735+00:00 | 100608 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Fixed Scribe.js source tree for retained GUI gitlink | accessibleHTTP 200 · 2026-10-02T22:19:18.714315+00:00 | 196539 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Fixed Scribe.js package metadata: 0.13.0 and AGPL-3.0 | accessibleHTTP 200 · 2026-10-02T22:19:19.481232+00:00 | 1748 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Fixed built-in browser Tesseract worker source | accessibleHTTP 200 · 2026-10-02T22:19:41.665127+00:00 | 9050 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
| Fixed language loader: cached traineddata and possible CDN reads | accessibleHTTP 200 · 2026-10-02T22:19:42.375238+00:00 | 32430 source bytes. Reused retained official-source GET bytes and receipt, verified against the research SHA256. Content length is original entity bytes. Source access and product performance are separate. |
Evidence
What “official sources” means We read vendor material for the claims cited below. This is a documentation review. No independent product test or professional endorsement is implied. Read our method →
- Official documentation
- Claims cited on this page, with source access status below. URL accessibility is separate from a substantive claim review.
- Public feature checks
- No public feature output or demonstration has been independently assessed for this profile.
- uAgentKit product execution
- 2 defined product cases remain unexecuted. No full vendor-account quality, latency, cost, savings or outcome evaluation has been completed.
- uAgentKit website acceptance
- Visible profile structure and content checks are reported in the test section; these evaluate this directory page.
- Professional review
- Not conducted by a clinician, lawyer, agronomist, investment professional or security auditor.
Commercial use: The reviewed application source declares AGPL-3.0. Use documents and images you have rights to process, and preserve their factual content and source rights when reusing exports. No separate hosted output-rights policy, blanket commercial clearance for all documents, or legal opinion was established; optional cloud-service terms and charges are separate from the built-in application.
Limitations and checks
- Both original synthetic label cases are prepared and unexecuted. There is no pass/fail result, recognized output, export, quality score or runtime performance measurement.
- The specific two-PNG import permission is pending. Prepared inputs, an empty UI and source GETs are not product execution.
- The two labels are original synthetic English text on clean backgrounds. Their high-contrast and low-contrast/rotated cases do not establish accuracy on real photos, handwriting, complex tables, long books or other languages.
- Numbers, units, SKU strings, negation and complete text need direct comparison with the source image. A native PDF file header does not establish a searchable text layer.
- Native proofreading can correct OCR errors, but a corrected export must not be reported as the unedited first recognition result. The proposed cases prohibit quality retries and replacement engines or inputs.
- Built-in browser processing can still require network resource downloads. The profile does not claim complete offline availability, zero whole-page external traffic or audited retention.
- AWS Textract is a separate optional cloud mode with credentials and potential charges. It is excluded from the local pilot and must not be described as the same privacy or price route.
- The repository/source gitlink, GUI package version and host behavior are different evidence scopes. No hosted commit or actual model digest is verified.
- Simplified Chinese support is experimental in the official FAQ. No language-quality, layout, multi-page or table-extraction benchmark was performed.
- Personal creator attribution does not prove current employee count, legal ownership or controlling structure. No autonomous planning, commerce update, messaging or order execution is established.
Field-level unknowns identify gaps in this review. They do not imply the vendor lacks the capability.
Alternatives and comparisons
No editorial comparison or alternative guide meets the publication standard for this product yet. Build an instant fact comparison.
Questions about Scribe OCR
What is Scribe OCR useful for?
Scribe OCR turns document images and scans into recognized text, lets a user proofread that text against the original image, and exports text or searchable document files. A person or small shop can use it to copy facts from an owned product label, while a creator can digitize print material. It is a user-directed OCR and document tool; no autonomous business or messaging workflow is established.
Who created Scribe OCR, and is its current team known?
The official GUI package names Balearica as author. The same public GitHub User profile explicitly says they created scribeocr.com and links the official app. That supports a personal creator origin. The real name, current team size, legal operator and controlling ownership were not established. The retained GUI repository was pushed on 28 September 2026 and its library on 1 October 2026; maintenance dates do not establish recognition quality.
Is Scribe OCR free, and does it require an account?
The official README and documentation describe the built-in browser application as free and open source. An empty public interface was observed without login, payment, CAPTCHA or initial Terms prompt, but neither prepared PNG has been imported or recognized. Successful free execution is therefore unverified. Optional AWS Textract requires cloud credentials and can incur separate charges; hardware, networking and hosting costs also differ from the free-software statement.
Does Scribe OCR process documents locally?
Official documentation describes built-in recognition as running locally in the browser. Code, WASM and language traineddata can still load over the network, and whole-page traffic, cache contents and retention were not independently audited. The current native application also exposes optional AWS Textract with separate credentials and a cloud-charge confirmation. That cloud branch is excluded from the proposed local tests and must not be covered by a blanket no-upload claim.
Which files can I import and export with Scribe OCR?
The official Getting Started guide supports PNG/JPEG document scans and PDFs. Existing character-level Tesseract HOCR or Abbyy XML can be imported for proofreading. The native Download menu lists TXT, PDF, HOCR, ALTO XML, DOCX, HTML, Markdown and Scribe documents; table XLSX is a separate optional feature. No actual export has been produced here. A downloadable PDF still needs parsing and visual inspection to establish its searchable text layer and original image.
How do I recognize text with Scribe OCR and preserve the original first output?
Open the complete official application, select the language and native recognition settings, import an authorized image, then click Recognize Text. Preserve the original first result before editing any words, and use native Download to retain TXT and PDF files from that same result. The proposed pilot keeps English, basic Quality 1, Batch Mode off and Advanced Recognition off. The observed public interface had Auto-Rotate on, Optimize Fonts off and disabled, and OCR Mode selected. Hidden model and engine details remain unknown. The original inputs and scoring are frozen, but both cases still need a reviewed and frozen evaluation runtime contract and the pending specific import permission before execution.
Has uAgentKit tested Scribe OCR?
No vendor-account performance test has been completed for Scribe OCR. No product OCR case has been executed. Two original synthetic English labels and evaluator-only exact references are prepared, and their original input bytes and six conditions per case have been frozen separately from runtime. The actual empty native UI was inspected, but imports, Recognize Text clicks and exports remain zero while the concrete two-PNG permission is pending. There is no pass/fail outcome or quality score. The runtime contract remains unfrozen and no product execution is registered.
What do the two planned Scribe OCR cases check?
The primary case uses a clear upright label; the boundary case uses gray text on white rotated five degrees. They check exact brand/item/SKU, quantity and dimensions, USD prices, missing material information and negation, complete first native TXT with character error rate at most 2%, and a first native PDF with a verifiable text layer. Exact-reference text is for the evaluator and must never be imported into the product. These proposed conditions are unexecuted and do not establish real-photo, handwriting, table, other-language, privacy or autonomous-action performance.
Sources and change history
- Official Scribe OCR browser application
Balearica / Scribe OCR · scribeocr.com · Read · 2026-10-03
- Official Scribe.js library repository metadata
Scribe OCR / GitHub · api.github.com · Read · 2026-10-03
- Official Scribe.js README and supported complete GUI
Scribe OCR · raw.githubusercontent.com · Read · 2026-10-03
- Official Scribe OCR GUI repository metadata
Scribe OCR / GitHub · api.github.com · Read · 2026-10-03
- Official Scribe OCR README: free browser application and local hosting
Scribe OCR · raw.githubusercontent.com · Read · 2026-10-03
- Official GUI package: Balearica author and AGPL-3.0
Scribe OCR · raw.githubusercontent.com · Read · 2026-10-03
- Official getting started: import documents and recognize text
Scribe OCR · docs.scribeocr.com · Read · 2026-10-03
- Public Scribe OCR repository organization profile
Scribe OCR / GitHub · api.github.com · Read · 2026-10-03
- Official language, local recognition, OCR editing and export FAQ
Scribe OCR · docs.scribeocr.com · Read · 2026-10-03
- Official support and monthly snapshot policy
Scribe OCR · docs.scribeocr.com · Read · 2026-10-03
- Official original-text editing and navigation controls
Scribe OCR · docs.scribeocr.com · Read · 2026-10-03
- Official retained GUI master commit and Scribe.js gitlink
Scribe OCR / GitHub · api.github.com · Read · 2026-10-03
- Official GUI AGPL-3.0 application license
Scribe OCR · raw.githubusercontent.com · Read · 2026-10-03
- Fixed GUI source: native recognition/export and separate AWS branch
Scribe OCR · raw.githubusercontent.com · Read · 2026-10-03
- Fixed Scribe.js source tree for retained GUI gitlink
Scribe OCR / GitHub · api.github.com · Read · 2026-10-03
- Fixed Scribe.js package metadata: 0.13.0 and AGPL-3.0
Scribe OCR · raw.githubusercontent.com · Read · 2026-10-03
- Fixed built-in browser Tesseract worker source
Scribe OCR · raw.githubusercontent.com · Read · 2026-10-03
- Fixed language loader: cached traineddata and possible CDN reads
Scribe OCR · raw.githubusercontent.com · Read · 2026-10-03