Setup

Bring Your Own AI Key

Connect Claude, ChatGPT, Gemini, DeepSeek, Groq or a local Ollama model to Resume Matcher with your own API key, in Settings or with env vars.

Why Bring Your Own Model

Resume Matcher ships no AI model. You connect one yourself: a cloud provider you already hold a key for, or a free model running on your own hardware. You pick the provider and the model, and with them the cost and where your resume text goes.

Supported Providers

Provider Type Default model in Settings
OpenAI Cloud gpt-5-nano-2025-08-07
Anthropic (Claude) Cloud claude-haiku-4-5-20251001
Google Gemini Cloud gemini-3-flash-preview
DeepSeek Cloud deepseek-chat
OpenRouter Cloud deepseek/deepseek-chat
Groq Cloud llama-3.3-70b-versatile
Azure AI Foundry Cloud mistral-large-latest
Ollama Local gemma3:4b
OpenAI-compatible server (llama.cpp, vLLM, LM Studio) Local None, type your own

The default is a starting point. The model field is free text, so you can type any model name your provider offers.

Option A: The Settings Page

  1. Open http://localhost:3000/settings.
  2. Pick your provider.
  3. Paste your API key. For Ollama or an OpenAI-compatible server, enter the server URL instead.
  4. Change the model name if you want a different one.
  5. Click Save, then Test Connection.

Test Connection sends the single word “Hi” to your provider and reports whether it answered. Settings you save here take precedence over environment variables.

Option B: Environment Variables

Four variables configure the backend:

Variable What it sets
LLM_PROVIDER One of openai, anthropic, gemini, deepseek, openrouter, groq, azure_foundry, ollama, openai_compatible
LLM_MODEL The model name
LLM_API_KEY Your provider key
LLM_API_BASE Server URL, for Ollama and OpenAI-compatible servers

For a local install, put them in apps/backend/.env. For Docker, pass them as container environment. The Docker image also reads Docker secrets: append _FILE to any of the four variables to load its value from a file, for example LLM_API_KEY_FILE=/run/secrets/llm_api_key.

Anthropic, local install (apps/backend/.env):

LLM_PROVIDER=anthropic
LLM_MODEL=claude-haiku-4-5-20251001
LLM_API_KEY=your-api-key

OpenAI, local install (apps/backend/.env):

LLM_PROVIDER=openai
LLM_MODEL=gpt-5-nano-2025-08-07
LLM_API_KEY=your-api-key

Ollama on the host, Resume Matcher in Docker:

docker run --name resume-matcher -p 3000:3000 \
  -v resume-data:/app/backend/data \
  -e LLM_PROVIDER=ollama \
  -e LLM_MODEL=gemma3:4b \
  -e LLM_API_BASE=http://host.docker.internal:11434 \
  ghcr.io/srbhr/resume-matcher:latest

Local Models

Ollama. Install it from ollama.com, pull a model, then pick Ollama in Settings:

ollama pull gemma3:4b

The default URL is http://localhost:11434. From inside Docker, use http://host.docker.internal:11434. On Linux, use the host’s IP address or start the container with --network=host.

OpenAI-compatible servers. llama.cpp, vLLM, LM Studio and any other server that speaks the OpenAI Chat Completions API work through the openai_compatible provider. Set LLM_API_BASE to the server URL and LLM_MODEL to the model the server is serving.

Slow models. Each AI request runs under one deadline, REQUEST_TIMEOUT_SECONDS (default 240, range 30 to 1800). A large model on modest hardware can run past it. Raise REQUEST_TIMEOUT_SECONDS in the backend environment and NEXT_PUBLIC_REQUEST_TIMEOUT_MS in the frontend environment (see apps/frontend/.env.sample).

Keys and local servers. The backend never sends LLM_API_KEY to Ollama or an OpenAI-compatible server, so a paid cloud key in your .env can’t leak to a local one.

Small models. The project’s design notes warn that smaller models struggle with the structured edit format the tailoring step uses.

For the full Docker walkthrough, see Docker + Ollama.

Where Your Key Is Stored

Keys you save in Settings go into the local SQLite database, encrypted with Fernet. The encryption key lives in .secret_key in the data folder (apps/backend/data/ for a local install, the Docker volume otherwise), created with 0600 permissions so that your user account alone can read it. The backend strips keys before it writes config.json in the same folder, so the settings file never holds one.

LOG_LLM=DEBUG writes keys to the log in plaintext. Don’t enable it on a shared machine.

What Gets Sent to Your Provider

The app sends nothing to Resume Matcher’s authors and contains no telemetry. Your resume text, including your name and contact details, and the job description go to the provider you picked and to no one else. No redaction step runs before sending. With a local model, nothing leaves your machine.

Feature Sent to the provider
Resume upload (parsing) Full resume text, including name, email, phone and links
Job description analysis, resume title The job description
Tailoring (skill plan, diff edits) Job description plus full resume JSON
Keyword pass after tailoring Tailored and master resume JSON plus the first 2,000 chars of the job description
Cover letter, outreach message Full resume JSON plus job description
Interview prep Resume JSON (up to 30,000 chars) plus job description (up to 12,000 chars)
Bullet scoring Bullet text, job description and extracted keywords
Enrichment, regenerate, Resume Wizard The resume or item plus your answers and instructions
Test Connection The word “Hi”

Your other API keys, the database file, the tracker board, settings and PDFs stay local. The one other network call is the AI library (LiteLLM) fetching a public model price list at startup; it carries none of your data and falls back to a bundled copy offline.

Cost

Resume Matcher is free and open source. A cloud provider bills your own account at its own prices for each request the app makes. The app doesn’t track that spend, so check your provider’s dashboard. A local model costs nothing per request.

Troubleshooting

Test Connection fails

Check the API key, or the server URL for Ollama and OpenAI-compatible servers. Save, then test again.

Docker can’t reach Ollama

Use http://host.docker.internal:11434 as the Ollama URL, not localhost. On Linux, use the host’s IP address or run the container with --network=host.

Timeouts with a local model

Raise REQUEST_TIMEOUT_SECONDS (up to 1800) in the backend and NEXT_PUBLIC_REQUEST_TIMEOUT_MS in the frontend.

A .env change has no effect

Settings saved in the UI take precedence over environment variables. Open Settings and change the value there.

Next Steps