Private AI Resume Tailoring: Run It Locally with Ollama
Private AI resume tailoring with a local model: what cloud providers receive, how to pick an Ollama model, and the speed, quality and timeout tradeoffs.
Each time you ask an AI model to tailor your resume, it reads the whole document: your name, phone number, email, every employer and date, and whatever else you keep in the file. With a cloud provider, that text leaves your machine. Private AI resume tailoring means running the model on your own hardware instead, so the resume (CV) and the job description never leave it. This post covers why you might want that, what a cloud provider receives when you don’t, how to choose a local model with Ollama, and the tradeoffs in speed, quality and timeouts.
The setup steps live in the docs. See Ollama setup for the Docker flow and bring your own API key for the Settings page and environment variables. This post is the why and the what.
Why keep your resume local
A resume is one of the more personal documents you own. It carries your full name, a phone number, an email, your city, and a dated account of where you worked. If you keep a master resume, it may also hold things you would never put on a tailored version: notes on why you left a job, a salary figure from an offer, a manager’s name. Paste that into a chat window and you have sent it to a company whose data-handling terms you may not have read.
Three groups have a concrete reason to care.
People under a confidentiality obligation. If your current employer’s name, project codenames or client list appear in your bullets, your resume contains information a contract may bind you to protect.
People in a quiet search. You do not want a record of “tailoring my resume for a role at a competitor” anywhere but your own disk.
People who read provider policies. Terms differ by provider and by plan. Anthropic’s commercial-products page, for example, states that inputs and outputs from its API and business products are not used for training by default, with exceptions for feedback you submit (Anthropic privacy centre, dated August 2026; its consumer plans are covered on a separate page). Other providers publish their own terms. If you would prefer not to depend on reading them right, a local model removes the question.
What a cloud provider receives
Resume Matcher is a harness: it sends prompts to whichever model you connect and applies code checks to the result. It performs no redaction before sending. The table below, from the project’s code, lists what each feature sends to the provider you chose.
| Feature | Sent to the provider |
|---|---|
| Resume upload (parsing) | Full extracted resume text, including name, email, phone and links |
| Job description analysis, resume title | The job description |
| Tailoring (skill plan, diff edits) | Job description plus the full resume as structured JSON, personal info included |
| Keyword pass after tailoring | Tailored and master resume JSON plus the first 2,000 characters of the job description |
| Bullet scoring | Bullet text, job description and extracted keywords |
| Cover letter, outreach message | Full resume JSON plus job description |
| Interview prep | Resume JSON (up to 30,000 characters) plus job description (up to 12,000) |
| Enrichment, regenerate, Resume Wizard | The resume or one item, plus your answers and instructions |
| Test Connection | The word “Hi” |
Nothing goes to Resume Matcher’s authors; the app contains no analytics code. Your API keys for other providers, the database, the tracker board, settings and the rendered PDFs stay on disk. The one other outbound request is the AI library (LiteLLM) fetching a public model price list at startup, which carries none of your data and falls back to a bundled copy when you are offline.
Pasting into ChatGPT or Claude by hand sends the same things, with less structure. Any workflow in which a cloud model reads your resume sends your resume to that cloud.
What private AI resume tailoring looks like
Point the same harness at a model running on your machine and every row of that table stays local. The model reads the resume, proposes edits, and the code checks them, and no network request carries the text anywhere.
In Resume Matcher this means one of two providers in Settings:
- Ollama, a free tool that downloads open models and serves them on
http://localhost:11434. - Any OpenAI-compatible server: llama.cpp, vLLM, LM Studio, or anything else that speaks the OpenAI Chat Completions API.
Four details qualify the word “private”.
Install needs the internet; use does not. You download the app, its dependencies, the browser it uses for PDF export, and the model weights once. After that the tailoring loop runs offline. The LiteLLM price-list fetch at startup falls back to a bundled copy when you are offline.
Your data sits in one folder. Resumes, job descriptions, tailored versions and the tracker board live in a SQLite file under the app’s data directory, or in one Docker volume. Settings sit in a small config file beside it. API keys, if you have saved any, are encrypted with a key file only your user account can read. The resume content itself is stored as plain data in that SQLite file, so the folder is as private as the disk it sits on; a laptop with full-disk encryption covers it.
No analytics in the app. The code contains no tracking SDK. The Docker image also disables Next.js’s own framework telemetry; if you run the frontend from source with npm run dev, that framework-level setting is yours to manage.
A local server never gets your cloud key. If you keep a paid provider’s key in your environment and switch to Ollama, the backend withholds that key from the local server, so switching providers cannot leak it.
Choosing a local model
Ollama is the path of least resistance. It pulls models by name and serves them on a local port. Resume Matcher’s Settings page suggests gemma3:4b as the Ollama default, and the project’s setup guide lists llama3.2 and mistral as other options. Any model Ollama serves can be typed into the model field.
Start with the default. gemma3:4b is the project’s suggested starting point, and a 3.4 GB download as of October 2026, according to Ollama’s model page. It is a small model, so read the paragraph on structured edits below for what to expect from it.
Go larger if your hardware allows. The same page lists gemma3:12b at 8.2 GB and gemma3:27b at 17 GB. Those are download sizes, and the memory a model needs at run time is a separate number that depends on quantisation and context length. Resume Matcher’s documentation states no RAM requirement, and I am not going to invent one. Ollama’s FAQ explains how to check whether a model loaded onto your GPU, into system memory, or a mix of both, which is the fastest way to find out whether a given size fits your machine.
Expect small models to struggle with structured edits. This is the known issue. Resume Matcher does not ask the model for a whole new resume. It asks for a list of targeted edits in a strict format, each one quoting the original text so code can verify it before applying it. The project’s design notes record that smaller models struggle with this format: they return malformed output, quote text that does not match, or skip the structure. The code retries a bounded number of times; if the retries fail you get an error, and you can regenerate or move to a larger model. A weak model costs you time, and the guardrails hold regardless of model: locked names, employers, titles, dates and degrees, and new numbers reverted.
Two things the harness does to help local models:
- JSON mode is always on for Ollama, with a prompt-only fallback if the server rejects it, so the model is pushed toward valid structure.
- Reasoning traces are stripped. Some reasoning models emit a
<think>block before their answer. The code removes those blocks before parsing, so a thinking model does not break the edit format.
If the default model gives you malformed previews on your resume, try a mid-size model next before you blame the prompt.
The tradeoffs: speed, quality, timeouts
Speed. A cloud model answers a tailoring request in seconds. A local model on a laptop CPU can take minutes for the same prompt, and the tailoring pipeline makes several model calls per run: keyword extraction from the posting, bullet relevance scoring, a skill plan, the edit list, and a keyword pass afterwards. Each call waits on the previous one. A model that is five times slower per call makes the whole run five times slower.
Quality. Larger models paraphrase with more fluency, stay closer to the posting’s vocabulary, and produce fewer broken edits. Smaller ones produce stiffer wording and more rejected changes. The code checks mean the floor is the same for both: a small model cannot invent a metric or change your dates any more than a large one can, because those edits are blocked before they reach you. The ceiling differs.
Timeouts. Every AI request in Resume Matcher runs under one deadline, REQUEST_TIMEOUT_SECONDS, which defaults to 240 seconds and accepts values from 30 to 1,800. That deadline covers the whole tailoring run, all of its model calls included. Local providers get a doubled timeout factor, but a large model on modest hardware can still run past it. If you see timeouts, raise REQUEST_TIMEOUT_SECONDS in the backend environment and NEXT_PUBLIC_REQUEST_TIMEOUT_MS in the frontend environment to match; the bring your own API key page lists both.
Cost. Zero per request with a local model, against a per-call charge from a cloud provider. Over a job search of forty applications, each with a tailored resume, cover letter and interview prep, the difference is real money for some and pocket change for others. The privacy argument stands on its own either way.
How it fits together
One paragraph, for orientation; the docs have the commands.
Ollama runs on your machine and listens on port 11434. Resume Matcher runs beside it, either from source or in Docker, and serves the app on port 3000. In Settings you pick Ollama as the provider, type the model name you pulled, and click Test Connection, which sends the word “Hi” and reports whether the model answered. From Docker, the app reaches Ollama through http://host.docker.internal:11434 instead of localhost, since the container has its own network namespace. After that, every feature in the app (tailoring, cover letters, interview prep, enrichment) uses the local model, and the browser prints your PDFs on the same machine.
The Ollama setup page has the Docker walkthrough, and bring your own API key covers Settings, environment variables and troubleshooting for both the source and Docker paths.
When a cloud model is the better call
Local is a tradeoff, and for some people the cloud side wins. If your laptop is thin, the model you can run is small, and small models produce rejected edits more often than useful ones. If you are applying this week and each run takes four minutes, speed matters more than the data flow. If your resume contains nothing you would mind a provider reading, and you have read the provider’s terms, the privacy argument is academic for you.
A middle path works too. Run the local model for the bulk of your tailoring and switch the provider in Settings for the one application where you want the best prose. The app stores a key per provider, so switching is a dropdown. The guardrail code stays constant: whichever model you connect, your name, employers, titles, dates and degrees stay locked, and any number the model invents gets reverted before you see it.
Resume Matcher is free, open source under the Apache 2.0 license, and used by over 200,000 job seekers. It ships no model and hosts no website, which is the whole point of this post: the model is yours, and so is the data.
If this helped
If running the model on your own machine is the version of this you were looking for, a star on Resume Matcher on GitHub helps other job seekers find it. You can follow me, Saurabh Rai, on GitHub, X and LinkedIn.
[This article was drafted, edited and formatted with the help of AI. Product facts were checked against the Resume Matcher source code, and outside sources are linked where they are used.]
Enjoyed this post? Star Resume Matcher on GitHub and follow along