Skip to content
ToolboxHere home

Urdu OCR: image to Urdu text

An online image to Urdu text converter: turn printed Urdu in a photo, screenshot or scanned page into text you can edit. Its text recognition reads Nastaliq, the style most Urdu books, newspapers and notices are set in, with a recogniser trained for it, and shows you each line next to its picture so you can check it. Nothing is uploaded.

اردو میں: تصویر سے اردو متن نکالیں

  • Files never leave your device
  • 100% free
  • No sign-up
  • No watermark
Urdu OCRCopy printed Urdu, Nastaliq included, out of photos and scans, with a model made for it.Starting… the tool runs entirely in your browser and needs JavaScript.

How to use Urdu OCR

  1. Urdu and the Nastaliq engine are already selected. For a form that mixes Urdu and English, choose "Urdu + English, typed text" instead.
  2. Drop an image, choose one, or paste a screenshot with Ctrl+V or Cmd+V. Photograph the page flat and straight, with the text filling the frame.
  3. Click "Extract text". The Urdu model and its runtime (about 20 MB, shared with the photo tools) download once, then every page is read on your device, line by line.
  4. Check the result in "Line by line": the picture of each line sits above its text, and words the model was unsure about are marked. Edit a line where you see it, then copy the text or download it as .txt or a Word file.

Why use Urdu OCR

Made for Nastaliq

Our own recogniser, trained on printed Urdu in Nastaliq and Naskh faces. Measured on pages it had never seen: about 4% of characters wrong in Nastaliq, where the usual Urdu OCR model got 27%.

Check it line by line

Each recognised line appears under a strip of the original, with doubtful words marked, so proofreading takes seconds rather than a re-read.

Word file or .txt

Download the text as .txt, or as a Word document set right to left in a Nastaliq font, ready for WhatsApp, Word or InPage.

Private, and offline after the first run

The picture and the text stay on your device. The model (6 MB) and its runtime are cached by your browser, so later pages read without a connection.

About this tool

Most printed Urdu uses Nastaliq: words slope down to the left, letters stack, and dots float above and below the line. Standard OCR models are trained on Naskh, the upright style used for Arabic, and read Nastaliq badly: on the same sentences set in the two styles, Tesseract's Urdu model gets 1 to 5 characters in a hundred wrong in Naskh and about one in four wrong in Nastaliq, which makes two thirds of the words unusable. This page therefore uses a recogniser of our own, trained on Urdu lines rendered in the open Nastaliq faces (Noto Nastaliq Urdu, Gulzar) and several Naskh faces, with the blur, tint and slant of phone photos added, and a word list of 166,000 Urdu words. It reads one line at a time: the page finds the lines itself, straightens a slightly turned photo, and joins the lines top to bottom. On pages it had never seen it gets about 4% of characters wrong in Nastaliq and under 1% in Naskh.

It reads printed text, not handwriting, and it does best on a sharp, straight, well-lit picture cropped to the text. Names, numbers and rare words are where it slips most, which is why they are the words to check first in the line-by-line view. Urdu mixed with English on the same line (a bilingual form) is better read by the second engine, Tesseract with both languages, which the page offers next to ours.

Frequently asked questions

Does it read Nastaliq?

Yes, and it was built for it. The recogniser was trained on Urdu set in Nastaliq faces, and on pages it had never seen it gets about 4% of characters wrong, against about 27% for the usual Urdu OCR model. Clear print reads best; dense newspaper text and faded scans still produce more errors, so check the marked words.

How do I get the best result?

A sharp, straight photo with the text filling the frame, cropped to the text you need. Photograph the page flat, without shadows, and avoid pictures of screens. Slightly turned photos are straightened for you.

Can it convert InPage files?

Not directly: this tool reads pictures. Export or screenshot the InPage page, read it here, and paste the text back into InPage or Word.

Can I read Urdu and English in one image?

Yes. Keep Urdu selected and also select English; the page switches to the "Urdu + English, typed text" engine, which reads both on the same line.

Does it work with handwriting?

No. Both engines are made for printed text.

Is my image uploaded?

No. The model and its runtime download to your browser once (about 20 MB) and every picture is read on your device. Nothing is sent anywhere.

Related tools