OCR Every Scanned Record, in Nepali and English

Scanned registers, photographed pages, and Devanagari documents become text you can search, extract from, and route.

HOW IT WORKS

From a photograph to a searchable record

Most institutional history is on paper that no search tool can read.

Registers, circulars, signed agreements and photocopies hold decades of decisions. Each page goes through the same five stages, whether it arrived as a clean scan or a photograph taken on a phone at a branch counter.

01

Import and clean

Straighten skewed pages, lift faded contrast, and remove speckle from photocopies so the text underneath can be read.

02

Read the layout

Work out where the columns, tables, stamps and signature blocks sit, so a two-column circular does not come back as interleaved nonsense.

03

Recognise the text

Convert printed Devanagari and Latin script into text, including pages that mix both — which most Nepali official documents do.

04

Pull out the fields

Dates, reference numbers and the parties named in the document are extracted automatically as the file arrives.

05

File it and act on it

The text is indexed next to everything else you hold, inherits the document’s permissions, and can start a workflow.

Script Coverage

ScriptPrintedHandwrittenNotes
Devanagari (Nepali)SupportedPlannedPrinted registers, circulars, notices and agreements set in Devanagari.
Latin (English)SupportedPlannedCorrespondence, contracts and reports produced in English.
Mixed on one pageSupportedPlannedThe common case in Nepali official documents — Devanagari body with English figures, dates and proper nouns.

Technical Details

Scripts read
Devanagari and Latin, including pages that mix the two.
Accepted inputs
Scanned documents and photographs of documents, including pages captured on a phone.
How documents arrive
Direct upload, bulk upload driven from a spreadsheet, connected Google Drive, Dropbox or OneDrive, and indexed email.
Fields extracted
Dates, reference numbers, and the parties named in the document, pulled out on arrival.
Duplicate handling
Repeat copies of the same document are detected on the way in rather than accumulating.
Access control
Extracted text inherits the document’s permissions. A record stays private to whoever added it until it is deliberately shared, and search will not surface it to anyone else in the meantime.
Tenancy
Each organisation’s records are stored fully separated from every other organisation’s.
Not yet available
Handwriting recognition, a searchable text layer written back onto the original scan, and bulk import of historical registers sorted into categories on the way in.

Where paper is holding things up

Each of these starts with a document that exists, is authoritative, and cannot currently be found by anyone who did not file it.

Archived registers nobody can search

Decades of bound registers and ledgers sit in storage because finding one entry means sending someone to look for it. Once the pages are read, an entry can be found by name, date or reference number.

  • Bound registers
  • Ledgers
  • Day books

Signed agreements held only on paper

The executed copy carrying the signatures is often the only authoritative version, and it exists as one physical document. Reading it makes the terms searchable and the key dates extractable, while the original stays intact as evidence.

  • Executed contracts
  • Addenda
  • Guarantees

Regulator circulars arriving as scans

Circulars and notices frequently arrive as scanned PDFs with no selectable text in them. Reading them on arrival makes their content searchable the same day, and catches repeat copies rather than filing them twice.

  • Circulars
  • Directives
  • Public notices

Identity documents at intake

Customer and applicant files are assembled from photographed identity documents and supporting paperwork. Extracting the reference numbers and dates at intake means the file is complete before it reaches whoever approves it.

  • Citizenship certificates
  • Passports
  • Supporting proofs

Frequently Asked Questions

Does it read handwriting?+

Not yet. Printed Devanagari and Latin text is supported today. Handwriting recognition for older Nepali records is on the roadmap and is not something we claim to do now — if your archive is predominantly handwritten, say so early and we will tell you honestly what proportion is currently readable.

What happens to a poor-quality scan?+

Pages are straightened, contrast-corrected and de-speckled before any text is read, which recovers a lot of faded photocopies. Where a page is genuinely illegible, it is stored and indexed by its metadata rather than being silently dropped, so the record still exists and can be found.

Can it handle a page with both Nepali and English on it?+

Yes. Mixed Devanagari and Latin on a single page is the normal case in Nepali official documents — a Devanagari body with English figures, dates and proper nouns — and is handled directly rather than requiring the document to be split.

Is the original document kept?+

Yes. The original scan or photograph is retained exactly as supplied. The extracted text is stored alongside it rather than replacing it, so the source document remains available as evidence.

Who can see text extracted from a restricted document?+

Only whoever could already see the document. Extracted text inherits the document’s permissions, so reading a file never widens access to its contents — including through search results.