Tailoring a Resume Without Touching Its Layout: Editing DOCX in Place
Most resume tools make the same trade without telling you. They read your resume, turn it into plain text, rewrite the text, and pour it back into one of their own templates. The words get better. The document you spent an evening aligning is gone.
When I built SwitchWithAI I started from the opposite rule: the document structure is never rebuilt. Only paragraph text changes. That single sentence sits at the top of the tailoring module, and most of the engineering since has been about keeping it true.
Why rebuilding is the wrong default
A resume is a layout before it is a text file. People pick a font because it fits eleven years onto one page. They tune a tab stop so dates line up on the right. They add a coloured rule under their name because it looks like them.
A text-first pipeline throws all of that away and then spends effort trying to recreate it. Editing in place keeps it for free. The cost moves somewhere else: you now have to change words inside a structure you did not design, without disturbing it.
Paragraphs, not documents
A DOCX is a zip of XML. Inside, each paragraph is a list of runs: stretches of text that share formatting. "Led the migration" is at least two runs, because the bold stops after one word.
The tailoring pass works like this:
- Load the original file and walk its paragraphs.
- Classify each one: summary, skill block, experience bullet, achievement, project line, certification, education, or generic.
- Send only the paragraphs worth rewriting to the model, with the job description and the paragraph's role.
- Write the new text back into the same paragraph object, keeping the formatting of its runs.
- Save a new DOCX next to the untouched original.
The classification matters more than it looks. An education line should almost never change. A skill block can gain a keyword that the job description uses and the candidate genuinely has. An experience bullet can be re-emphasised. Treating them the same is how tools end up "improving" a degree name.
Keeping the original as the source of truth
Every rewrite is stored as a proposal, not applied destructively. On the review screen the user accepts or rejects each change. When they press apply, the backend rebuilds the DOCX from the untouched source, using only the accepted indices. No model call, no quota spent, and no drift from edits being layered on edits.
That design came from a bug I would rather not repeat: re-running changes on top of an already-tailored file compounds small formatting losses. Starting from the original every time makes each build reproducible.
The PDF is where layouts actually break
Getting the DOCX right turned out to be the easier half. Most people download the PDF, and converting DOCX to PDF on a Linux server means LibreOffice.
LibreOffice is excellent, but it can only use the fonts it has. If a resume asks for Calibri and the server does not have it, the renderer quietly picks a substitute. A substitute with slightly wider letters is the difference between a one-page resume and a two-page one, and nobody notices until a recruiter opens it.
So every conversion now answers two cheap questions:
- Which fonts were actually embedded, and were any of them chosen by the renderer rather than requested by the document? If so, the API warns instead of shipping silently.
- How many pages came out, compared with the original? Tailoring adds keywords, keywords add words, and words can push a proof-read single page onto a second one. That is a real quality regression, so it is reported.
Caching by content, not by filename
The same tailored DOCX gets converted several times: once while the result streams, again after the user accepts edits, and again on every later download from history. Each conversion is a few hundred milliseconds of identical work.
The cache key is the SHA-256 of the DOCX bytes plus an engine version. Two requests for identical content share one conversion. A document that changed by one character correctly misses. And bumping the engine version when fonts change invalidates everything, because a cache that outlives a font fix keeps serving the broken layout.
What I would tell someone building the same thing
- Decide early whether you own the layout or the user does. Everything downstream follows from that choice.
- Keep the original file immutable and rebuild from it. Layered edits rot.
- Treat the PDF converter as a separate system with its own failure modes, and measure fonts and page count on every run.
- Classify content before rewriting it. Not every paragraph should be touched, and the ones that should not are usually the ones users check first.
If you want to see the result, SwitchWithAI runs this pipeline. Upload a DOCX, paste a job description, and compare the output with your original. The layout should be the one thing you cannot tell apart.