About this tool
English OCR is a red ocean. Chinese on a phone photo is the gap
PaddleOCR.js runs PP-OCRv5 with ONNX Runtime in this tab. Detection and recognition are the official Apache-2.0 pipeline, about 15–20 MB the first time, then cached. Traditional Chinese and thermal receipts are the reason this page exists: Tesseract's Chinese accuracy is the thing people bounce off, and we are not pretending otherwise.
Drop, paste, or try the sample
The same paste path as Redact Screenshot — a screenshot on the clipboard is enough. The sample button draws a fake receipt so you can see the round-trip before you trust a real one. Output is plain text you can copy or download, plus a simple line list with scores so a weak box does not hide in an average.
Handwriting and display type will be wrong
That is the model, not a missing checkbox. Printed menus, invoices, and a phone photo of a page of 繁體 are in-distribution. A signature, a neon sign, or a calligraphy scroll is not. The page says so next to the button rather than in a buried FAQ.
Later, PDF
The 1201 PDF line still needs an OCR fallback for scanned pages. When /pdf-to-text ships, it will call this engine rather than a second one. Redact Screenshot is the sibling for when you have read the text and now need to hide it.