Skip to content
ToolboxHere home

Documents · 6 min read

PDF to Word: why the converted file looks wrong, and how to get a clean .docx

A PDF does not contain paragraphs, tables or headings. It contains letters placed at coordinates. Every converter has to work out the structure from where the ink sits, and that is where the odd results come from.

Saad Naeem

AI engineer, Doha · Updated

What a PDF actually contains

A Word file is a list of paragraphs, each with a style, holding runs of text, with tables and pictures as objects. A PDF is nothing like that. It is a set of drawing instructions: draw this glyph from this font at this x and y, draw a line from here to there, fill this rectangle grey, place this picture. There is no “paragraph” and no “table” anywhere in the file; the reader’s eye supplies them.

So a PDF to Word converter is not translating, it is reverse-engineering. From where the letters sit it has to decide which runs belong to one line, which lines form a paragraph, where a heading starts, whether a block of aligned text is a table, whether two blocks are columns or a label and its value. When it guesses wrong you get the familiar symptoms: a paragraph split into one line per paragraph, a table that arrives as tab-separated text, a two-column page read straight across, or headings that are just bold body text.

Knowing which guess went wrong tells you what to do about it.

Symptom one: there is no text at all

If the Word file is a picture of each page, or Word opens it with nothing you can select, the PDF was a scan. A scanner or a phone camera produces a photograph of the paper; there are no letters inside the file, only pixels. No converter can turn that into text without optical character recognition (OCR), which reads the picture the way you do.

Good converters run OCR automatically on pages where they find no text. The PDF to Word tool on this site does that, in English, Arabic and Urdu, with the recognition running inside your browser. Two things to know: OCR output needs proofreading, because every OCR engine misreads some characters, especially on faint or skewed scans; and a scan at low resolution (under about 150 dots per inch, or a photo taken from too far away) gives poor text whatever the engine. If you still have the paper, rescanning at 300 dots per inch fixes more than any converter setting.

Try it: PDF to Word

Text pages become editable paragraphs, tables and headings; scanned pages are read with OCR first. The file stays on your device.

Open tool

Symptom two: the lines break in different places

The PDF embedded a font that Word does not have, or embedded only the subset of letters it used, so Word cannot use it. Word substitutes a font with slightly different letter widths, every line refills, and a document that was two pages becomes two and a quarter, with a lonely last line on page three.

Converters deal with this in two ways. The crude one writes every line as its own paragraph with a hard break, which keeps the look but makes the document impossible to edit. The better one maps the PDF’s fonts to the closest font that ships with Office: the free fonts that were drawn to match Microsoft’s metrics (Liberation to Times New Roman and Arial, Carlito to Calibri, Caladea to Cambria) have the same letter widths, so the lines break in the same places. Our converter does this, and it also reproduces the exact line pitch when the PDF was set tighter than Word’s single spacing, so text that was squeezed onto one page stays on one page.

If the PDF used a font with no metric twin, accept that the breaks will move, and fix the document’s page breaks in Word rather than fighting each line.

Symptom three: the table arrived as text

Tables are where converters differ most, because a PDF table can be drawn four ways:

  1. A ruled grid. Lines around every cell. Easy: the crossings of the lines give the cells.
  2. Rules between rows only. A statement or a price list with a line under each row and a shaded header. The columns have to be found from where the text sits.
  3. No lines at all. Just text aligned in columns. The converter has to notice that several rows share the same gutters.
  4. Nested tables. A small table inside the cell of a large one.

Our converter handles all four, building the second and third kinds from the alignment of the text and the edges of shaded bands. The third kind is still the fragile one: it needs at least three columns of short cells that line up across several rows. A two-column list of labels and values, or a table whose cells wrap onto two lines, may arrive as text with tabs between the columns. That is deliberate: a wrong table is harder to fix than tabbed text, which Word turns into a table in one step (select it, Insert, Table, Convert Text to Table).

A related oddity: a shaded box with a single empty cell is a form field, not a table, and a grid whose cells are mostly empty is usually a chart. Converters that treat everything ruled as a table produce one-cell tables all over a form; ours turns those into pictures or leaves them out.

Symptom four: two columns read straight across

Newsletters, papers and brochures set text in columns. A naive converter reads each line across the whole page, so the left column’s first line is followed by the right column’s first line. Converters that split the page first (an XY cut: find a vertical gap no text crosses, and treat the two sides as separate blocks) read each column in turn.

The trap is the opposite case. An invoice has a label on the left and a value on the right with a gap between them; that is one line, not two columns. Our converter splits only where both sides read as running text, or where the two sides are columns of similar width, so a label and its value stay on one line with a tab between them.

Symptom five: Arabic comes out backwards

Right-to-left text is stored in a PDF in the order it is drawn, which is visual order, right to left. Some converters hand that to Word unchanged, so the letters of each word are reversed and the words run in the wrong direction. Others reverse it correctly but break the ligatures: the lam-alef pair, which the PDF drew as one glyph, comes out as two letters in the wrong order.

Our tool fixes both, using where each glyph was drawn to put the letters back in reading order, normalises the Arabic presentation forms that PDFs use into ordinary letters, and marks the paragraphs as right-to-left so Word aligns them properly. For Urdu the same applies, with one addition: scanned Urdu in Nastaliq is read with a recogniser trained for that script, because general OCR engines were trained on Naskh and get about a quarter of Nastaliq characters wrong.

What to check after converting

A clean conversion still needs five minutes in Word before it goes anywhere:

  • Headings. Check that chapter and section titles carry Word’s Heading styles, not just bold text, so the navigation pane and a table of contents work. Our converter assigns Heading 1 and 2 by the size of the text.
  • Page breaks. Where a page in the PDF ended early (a chapter end), there should be a page break; where it was full, there should not be, so that small differences in layout do not cascade into near-empty pages.
  • Lists. Bullets and numbers should be real Word lists, so adding an item renumbers the rest.
  • Pictures. Charts drawn with lines in the PDF become pictures of that part of the page, labels included. They are not editable, but they are complete.
  • Links. Web addresses and internal references should still be clickable.

When not to convert

If the PDF was made from a Word file, the original .docx is better than any conversion: styles, numbering, tracked changes and comments all survive, and nothing has been guessed. Ask the sender for it. Convert when the original is gone, when you only need the text, or when you need to fill in a form that was sent as a flat PDF, and for that last case a form-filling tool that writes on the PDF itself is usually the better route.

And once you have edited the document, send it back as a PDF, not as the .docx: a PDF looks the same on every device, and the recipient cannot accidentally change it.

Try it: Word to PDF

Lays the document out with Word's own line and page rules, in your browser, so the PDF matches what you saw in Word.

Open tool

Frequently asked questions

Why does my converted Word file have no text I can select?

The PDF was a scan: every page is a photograph of paper, with no letters inside the file. The converter has to read the picture with OCR first. Our PDF to Word tool does that automatically for pages it finds no text on, in English, Arabic and Urdu; the result needs proofreading, as all OCR does.

Why are the line breaks in different places than in the PDF?

The PDF used a font Word does not have, so Word substituted one with slightly different letter widths and the lines refilled. Converters that map the PDF’s fonts to the closest font installed with Office (Liberation Serif to Times New Roman, for example) keep the breaks far closer.

Can a converter recover a table that has no lines?

Sometimes. A table with ruled lines is easy; one with only aligned text needs the converter to notice columns that line up across several rows. Our tool builds those when at least three columns of short cells line up; wider gaps or two-line cells may come out as tab-separated text instead, which is quick to turn into a table in Word.

Is it safe to convert a contract or a bank statement online?

Only if the file never leaves your device. Most converters upload the PDF to a server, keep it for a while and process it there. The tools on this site run inside your browser, so the document is never sent anywhere; you can switch the network off after the page loads and it still works.

Should I convert the PDF at all, or ask for the original?

If someone made the PDF from a Word file, the original .docx is always better than any conversion: styles, numbering and tracked changes survive. Convert when the original is gone, when the sender will not share it, or when you only need the text.

Written by Saad Naeem

AI engineer in Doha who builds every tool on ToolboxHere. The figures in the guides come from the official sources named in them or from the tools themselves. Found a mistake? Tell us and it is corrected.

Related guides