Why Bring Your Own Model
Resume Matcher ships no AI model. You connect one yourself: a cloud provider you already hold a key for, or a free model running on your own hardware. You pick the provider and the model, and with them the cost and where your resume text goes.
Supported Providers
| Provider | Type | Default model in Settings |
|---|---|---|
| OpenAI | Cloud | gpt-5-nano-2025-08-07 |
| Anthropic (Claude) | Cloud | claude-haiku-4-5-20251001 |
| Google Gemini | Cloud | gemini-3-flash-preview |
| DeepSeek | Cloud | deepseek-chat |
| OpenRouter | Cloud | deepseek/deepseek-chat |
| Groq | Cloud | llama-3.3-70b-versatile |
| Azure AI Foundry | Cloud | mistral-large-latest |
| Ollama | Local | gemma3:4b |
| OpenAI-compatible server (llama.cpp, vLLM, LM Studio) | Local | None, type your own |
The default is a starting point. The model field is free text, so you can type any model name your provider offers.
Option A: The Settings Page
- Open
http://localhost:3000/settings. - Pick your provider.
- Paste your API key. For Ollama or an OpenAI-compatible server, enter the server URL instead.
- Change the model name if you want a different one.
- Click Save, then Test Connection.
Test Connection sends the single word “Hi” to your provider and reports whether it answered. Settings you save here take precedence over environment variables.
Option B: Environment Variables
Four variables configure the backend:
| Variable | What it sets |
|---|---|
LLM_PROVIDER |
One of openai, anthropic, gemini, deepseek, openrouter, groq, azure_foundry, ollama, openai_compatible |
LLM_MODEL |
The model name |
LLM_API_KEY |
Your provider key |
LLM_API_BASE |
Server URL, for Ollama and OpenAI-compatible servers |
For a local install, put them in apps/backend/.env. For Docker, pass them as container environment. The Docker image also reads Docker secrets: append _FILE to any of the four variables to load its value from a file, for example LLM_API_KEY_FILE=/run/secrets/llm_api_key.
Anthropic, local install (apps/backend/.env):
LLM_PROVIDER=anthropic
LLM_MODEL=claude-haiku-4-5-20251001
LLM_API_KEY=your-api-key
OpenAI, local install (apps/backend/.env):
LLM_PROVIDER=openai
LLM_MODEL=gpt-5-nano-2025-08-07
LLM_API_KEY=your-api-key
Ollama on the host, Resume Matcher in Docker:
docker run --name resume-matcher -p 3000:3000 \
-v resume-data:/app/backend/data \
-e LLM_PROVIDER=ollama \
-e LLM_MODEL=gemma3:4b \
-e LLM_API_BASE=http://host.docker.internal:11434 \
ghcr.io/srbhr/resume-matcher:latest
Local Models
Ollama. Install it from ollama.com, pull a model, then pick Ollama in Settings:
ollama pull gemma3:4b
The default URL is http://localhost:11434. From inside Docker, use http://host.docker.internal:11434. On Linux, use the host’s IP address or start the container with --network=host.
OpenAI-compatible servers. llama.cpp, vLLM, LM Studio and any other server that speaks the OpenAI Chat Completions API work through the openai_compatible provider. Set LLM_API_BASE to the server URL and LLM_MODEL to the model the server is serving.
Slow models. Each AI request runs under one deadline, REQUEST_TIMEOUT_SECONDS (default 240, range 30 to 1800). A large model on modest hardware can run past it. Raise REQUEST_TIMEOUT_SECONDS in the backend environment and NEXT_PUBLIC_REQUEST_TIMEOUT_MS in the frontend environment (see apps/frontend/.env.sample).
Keys and local servers. The backend never sends LLM_API_KEY to Ollama or an OpenAI-compatible server, so a paid cloud key in your .env can’t leak to a local one.
Small models. The project’s design notes warn that smaller models struggle with the structured edit format the tailoring step uses.
For the full Docker walkthrough, see Docker + Ollama.
Where Your Key Is Stored
Keys you save in Settings go into the local SQLite database, encrypted with Fernet. The encryption key lives in .secret_key in the data folder (apps/backend/data/ for a local install, the Docker volume otherwise), created with 0600 permissions so that your user account alone can read it. The backend strips keys before it writes config.json in the same folder, so the settings file never holds one.
LOG_LLM=DEBUG writes keys to the log in plaintext. Don’t enable it on a shared machine.
What Gets Sent to Your Provider
The app sends nothing to Resume Matcher’s authors and contains no telemetry. Your resume text, including your name and contact details, and the job description go to the provider you picked and to no one else. No redaction step runs before sending. With a local model, nothing leaves your machine.
| Feature | Sent to the provider |
|---|---|
| Resume upload (parsing) | Full resume text, including name, email, phone and links |
| Job description analysis, resume title | The job description |
| Tailoring (skill plan, diff edits) | Job description plus full resume JSON |
| Keyword pass after tailoring | Tailored and master resume JSON plus the first 2,000 chars of the job description |
| Cover letter, outreach message | Full resume JSON plus job description |
| Interview prep | Resume JSON (up to 30,000 chars) plus job description (up to 12,000 chars) |
| Bullet scoring | Bullet text, job description and extracted keywords |
| Enrichment, regenerate, Resume Wizard | The resume or item plus your answers and instructions |
| Test Connection | The word “Hi” |
Your other API keys, the database file, the tracker board, settings and PDFs stay local. The one other network call is the AI library (LiteLLM) fetching a public model price list at startup; it carries none of your data and falls back to a bundled copy offline.
Cost
Resume Matcher is free and open source. A cloud provider bills your own account at its own prices for each request the app makes. The app doesn’t track that spend, so check your provider’s dashboard. A local model costs nothing per request.
Troubleshooting
Test Connection fails
Check the API key, or the server URL for Ollama and OpenAI-compatible servers. Save, then test again.
Docker can’t reach Ollama
Use http://host.docker.internal:11434 as the Ollama URL, not localhost. On Linux, use the host’s IP address or run the container with --network=host.
Timeouts with a local model
Raise REQUEST_TIMEOUT_SECONDS (up to 1800) in the backend and NEXT_PUBLIC_REQUEST_TIMEOUT_MS in the frontend.
A .env change has no effect
Settings saved in the UI take precedence over environment variables. Open Settings and change the value there.
Next Steps
- What is Resume Matcher? for the overview
- Features for everything the AI can generate once it’s connected