Quick Start
OpenRAG is an open source Retrieval Augmented Generation (RAG) solution. This guide is a step-by-step walkthrough to help you get started with OpenRAG.
Prerequisites
Section titled βPrerequisitesβ- Docker and Docker Compose
- Your hardware should meet these specifications:
- CPU deployment: Minimum 13 GiB RAM for light PDF parsers (
PyMuPDFLoader), or 23 GiB RAM for heavier parsers likeMarkerLoader(refer to this section for details) - GPU deployment: 16 GB GPU memory recommended (for systems with separate CPU and GPU memory)
- CPU deployment: Minimum 13 GiB RAM for light PDF parsers (
Installation and Configuration
Section titled βInstallation and Configurationβ1. Clone the repository:
Section titled β1. Clone the repository:βgit clone --recurse-submodules git@github.com:linagora/openrag.git
cd openrag/git checkout main # or a given release2. Create a .env File
Section titled β2. Create a .env FileβThe Docker Compose stack lives under infra/compose/, next to its .env.example template. Generate a .env from it β this fills in every credential with a fresh random value, which the stack requires to start:
python3 scripts/gen_env.pyThen fill in the blanks the generator cannot guess: your LLM/VLM endpoints and the embedder.
The template marks each credential __GENERATE_ME__ rather than shipping a working default, and OpenRag refuses to start on any value published in this repository. A plain cp .env.example .env therefore fails at startup with a message naming the variables still to set β deliberately, because a default nobody changed is the most common way a deployment ends up with a known credential.
Here is the minimal set of variables to get started β see the full environment-variable reference for every other option:
# ============================================================================# OpenRAG β minimal .env## Only the variables you must set for the default compose stack to boot are# listed here. Every other knob (PDF/audio loaders, chunking, retriever,# reranker, Ray Serve, admin UI, OIDC/SSO, rate limiting, MCP server, web# search, β¦) has a sensible default and is documented in full at:## https://linagora.github.io/openrag/documentation/env_vars/# ============================================================================
# ββ LLM (external, OpenAI-compatible) βββββββββββββββββββββββββββββββββββββββBASE_URL=API_KEY=MODEL=LLM_SEMAPHORE=10# Allow clients to supply their own LLM endpoint through `metadata.llm_override`.# LLM_OVERRIDE_ALLOW_CUSTOM_ENDPOINT=true
# ββ VLM (vision model, used for image understanding) ββββββββββββββββββββββββ# Can reuse the LLM values above if that model accepts images.VLM_BASE_URL=VLM_API_KEY=VLM_MODEL=VLM_SEMAPHORE=20
# ββ Embedder (HuggingFace model served by the bundled vLLM) βββββββββββββββββ# Point these at an external embedding service instead of the bundled vLLM:
# Changing the model of an install that has already indexed documents needs a# reindex: vectors from two models don't match, even at the same dimension# (both give 1024). An install indexed with the previous default keeps it with# EMBEDDER_MODEL_NAME=jinaai/jina-embeddings-v3.EMBEDDER_MODEL_NAME=Qwen/Qwen3-Embedding-0.6B# EMBEDDER_BASE_URL=http://vllm:8000/v1# EMBEDDER_API_KEY=EMPTY# MAX_MODEL_LEN=2047# Extra `vllm serve` flags for the bundled vLLM embedder, appended after the# built-in ones, so a repeated flag overrides its default:# EMBEDDER_EXTRA_ARGS=--gpu_memory_utilization 0.1
# ββ Reranker (re-scores retrieved chunks; bundled Infinity server) βββββββββββRERANKER_PROVIDER=infinity # 'infinity' (bundled, default) or 'openai' (bundled vLLM, or an external endpoint)RERANKER_MODEL=Alibaba-NLP/gte-multilingual-reranker-baseRERANKER_ENABLED=true# Extra `vllm serve` flags for the bundled vLLM reranker (RERANKER_PROVIDER=openai).# vLLM doesn't recognise gte-multilingual-reranker-base's `NewForSequenceClassification`# architecture on its own, so that model needs:# RERANKER_EXTRA_ARGS=--hf-overrides '{"architectures": ["GteNewForSequenceClassification"]}'
## Point these at an external reranker instead of the bundled Infinity / VLLM(Openai) server:# RERANKER_BASE_URL=http://reranker:7997# RERANKER_API_KEY=EMPTY
# ββ PDF parser (default: PyMuPDF β lightweight, CPU-friendly) ββββββββββββββββ# Switch to MarkerLoader for OCR / scanned PDFs, complex layouts & embedded# images (heavier β more RAM/GPU). Other option: DoclingLoader.PDFLOADER=PyMuPDFLoader# Marker tuning (only when PDFLOADER=MarkerLoader):# MARKER_POOL_SIZE=1 # marker worker actors (β 1 per cluster node / Machine).# MARKER_MAX_PROCESSES=2 # concurrent PDFs per worker (raise with more GPU)
# ββ Image captioning & chunk contextualization (both ON by default) ββββββββββ# Both run during indexing and call the VLM/LLM β set to false to index# faster and cheaper (with some retrieval-quality trade-off).# IMAGE_CAPTIONING=false # stop describing images in documents via the VLM# CONTEXTUAL_RETRIEVAL=false # stop prepending LLM-generated context to each chunk (Anthropic technique)
# ββ Secrets ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ# Every __GENERATE_ME__ below must be replaced before the stack will start β# the application refuses to boot on a value this project publishes. Fill them# all in one command instead of by hand:## python3 scripts/gen_env.py## MINIO_* are shared by the minio service and Milvus β both sides must match.MINIO_ACCESS_KEY=__GENERATE_ME__MINIO_SECRET_KEY=__GENERATE_ME__POSTGRES_PASSWORD=__GENERATE_ME__# POSTGRES_USER=root
# Fresh Milvus 3 installations can keep the default queue selection. During an# upgrade, set this to the queue already used by the existing deployment.# MILVUS_MQ_TYPE=rocksmq
# ββ API and Chainlit (chat interface) authentication ββββββββββββββββββββββββββββββ# Bearer token that bootstraps the admin user and guards the API.# Generated by scripts/gen_env.py; `openssl rand -hex 16` works by hand.# For SSO instead, see https://linagora.github.io/openrag/documentation/oidc/AUTH_TOKEN=__GENERATE_ME__# AUTH_MODE=token # 'token' (default) or 'oidc' for SSO# SUPER_ADMIN_MODE=true
# # FastAPI port (the API and Chainlit chat interface share the same port).# APP_PORT=8080
# # Optional name used when the assistant introduces itself to a user.# # Leave unset for a neutral introduction, or set it for a white-label deployment.# ASSISTANT_NAME=Your Assistant
# # Rate limiting, activated by default (recommended).# RATE_LIMIT_ENABLED=false# # Path prefixes the limiter skips (CSV, matched by prefix β keep the trailing# # slash so siblings like /chainlithack stay limited). Unset uses the default shown# # here; set-but-empty disables the exemption entirely.# RATE_LIMIT_EXEMPT_PATHS=/chainlit/,/assets/
# # Reverse-proxy trust. The client IP behind rate limits, and the scheme behind the# # OIDC cookie `Secure` flag, are only read from X-Forwarded-* headers when the peer# # is listed here. The bundled admin-ui proxy is NOT on loopback, so keeping the# # 127.0.0.1 default makes every chat user share a single rate-limit bucket keyed on# # the proxy's address. Prefer your compose/k8s proxy subnet (e.g. 172.16.0.0/12).# # Avoid `*` unless the API port is unreachable except through the proxy: `*` trusts# # X-Forwarded-For from ANY peer, so a caller reaching the API directly can forge the# # header and evade the per-IP brute-force limit on /auth/*. See env_vars.md# # (UVICORN_FORWARDED_ALLOW_IPS / Rate Limiting) for the full proxy-trust guidance.# UVICORN_FORWARDED_ALLOW_IPS=172.16.0.0/12
# ββββββββββββββββ Chainlit: Chat interface ββββββββββββββββ# Session-cookie signing secret β required once auth is enabled; keep it stable# (rotating it invalidates every live chat session).# Generated by scripts/gen_env.py; by hand:# python -c "import secrets; print(secrets.token_urlsafe(32))"CHAINLIT_AUTH_SECRET=__GENERATE_ME__
# ββββββββββββββββ Admin / Indexer UI (React SPA) ββββββββββββββββ# Host port for the admin UI β the document ingestion, indexing & management# interface, served at http://<host>:<ADMIN_UI_PORT>/app/. It also proxies the# API/auth, so it is the OIDC front door. Zero-config otherwise (same-origin, no# CORS); VITE_* build-time options are documented in the env vars reference.# ADMIN_UI_PORT=8081# GRAFANA_ADMIN_USER=admin# GRAFANA_ADMIN_PASSWORD=__GENERATE_ME__# Direct Grafana access:# GRAFANA_URL=http://localhost:3000/d/openrag-http/openrag-http-metrics# Or through the Admin UI proxy:# GRAFANA_URL=http://localhost:8081/grafana/d/openrag-http/openrag-http-metrics# GF_SERVER_ROOT_URL=http://localhost:8081/grafana/# GF_SERVER_SERVE_FROM_SUB_PATH=true
# ββββββββββββββββ Prometheus metrics ββββββββββββββββ# GET /metrics on the API port (APP_PORT) is not gated by AUTH_TOKEN and fails# closed: without METRICS_TOKEN every scrape gets 403. Set a random value# (`openssl rand -hex 16`); the monitoring overlay hands the same value to its# Prometheus, so the bundled Grafana needs nothing else. An external# Prometheus sends it as `Authorization: Bearer <METRICS_TOKEN>`.# METRICS_TOKEN=replace-with-a-random-secret# Only when APP_PORT is unreachable from outside your network and you'd rather# scrape without a token: open the endpoint to anyone who can reach the port.# The admin-ui proxy (ADMIN_UI_PORT) never serves /metrics either way.# METRICS_ALLOW_UNAUTHENTICATED=true
# ββ Ray (kept as-is by the compose stack; see the docs for what each does) βββRAY_DEDUP_LOGS=0# Both of the above are required with LOG_FORMAT=json: Ray's dedup rewrites# the surviving line as "{...} [repeated 2x across cluster]" and its relay# prefix is ANSI-colorized even on a pipe, so worker lines stop being JSON.RAY_COLOR_PREFIX=0RAY_ENABLE_RECORD_ACTOR_TASK_LOGGING=1RAY_task_retry_delay_ms=3000RAY_ENABLE_UV_RUN_RUNTIME_ENV=0
# RAY_memory_monitor_refresh_ms=0
# ββ Logging (DEBUG on dev, INFO on prod) ββLOG_LEVEL=DEBUG
# text (colorized, for a terminal) or json (one object per line, for a log# collector). The logging overlay forces json on the API; set it here only for# a collector of your own.# LOG_FORMAT=text
# ββββββββββββββββ Loki log shipping (logging.docker-compose.yaml) ββββββββββββββββ# Push endpoint of your Loki; required to start the overlay.# LOKI_URL=https://loki.example.com/loki/api/v1/push# Optional basic auth and tenant header (X-Scope-OrgID).# LOKI_USERNAME=# LOKI_PASSWORD=# LOKI_TENANT_ID=3. File Parser configuration
Section titled β3. File Parser configurationβAll supported file format parsers are pre-configured. For PDF processing, PyMuPDFLoader is the default parser β a lightweight, fast, CPU-friendly engine well suited to searchable PDFs and quick local testing.
4. Run OpenRAG
Section titled β4. Run OpenRAGβAll deployment assets now live under infra/:
Directoryinfra/
Directorycompose/ full stack (recommended)
- docker-compose.yaml
- .env.example
- .env your configured env
Directorydocker/ Dockerfiles
- β¦
Directoryansible/ remote deployment
- β¦
Directoryopenrag/ application code
- β¦
Directoryconf/ YAML configuration
- β¦
Run the stack from infra/compose/, where the .env you just created lives. Use the GPU tab on a machine with an NVIDIA GPU, or the CPU tab otherwise.
cd infra/composedocker compose up -d
# stop it later with:# docker compose downcd infra/composedocker compose --profile cpu up -d
# stop it later with:# docker compose --profile cpu downInference services
Section titled βInference servicesβThe LLM and VLM are always external OpenAI-compatible endpoints β set BASE_URL/MODEL/API_KEY (and the VLM_* equivalents) in .env.
The embedder and reranker, by contrast, are bundled: docker compose up -d starts a vllm container serving EMBEDDER_MODEL_NAME and a reranker container (RERANKER_PROVIDER, default Infinity). No extra step is needed to use them.
cd infra/composedocker compose up -d # GPU β core stack + bundled embedder & reranker# docker compose --profile cpu up -d # CPU-only hostTo use external embedding/reranking instead β reusing existing inference servers and keeping the footprint small:
- Embedder β set
EMBEDDER_BASE_URL(+EMBEDDER_API_KEY) to your endpoint, then stop the local server by commenting out thevllm-gpu/vllm-cpuservice ininfra/compose/docker-compose.yaml. - Reranker β set
RERANKER_ENABLED=falseto skip reranking, or point it at an endpoint withRERANKER_PROVIDER=openaiandRERANKER_BASE_URL(+RERANKER_API_KEY); to also stop the container, comment out theextern/reranker/β¦line in theinclude:block.
Once the app is up and running, you can access the provided services β see Default ports below.
Ansible
Section titled βAnsibleβClone the repository, then run the deployment script from infra/ansible/ and follow the interactive prompts:
git clone --recurse-submodules https://github.com/linagora/openrag.gitcd openrag/infra/ansible./deploy.shDefault ports
Section titled βDefault portsβOnce the stack is up, OpenRAG exposes the following services by default:
| Service | Port | Description |
|---|---|---|
| API Documentation | 8080/docs | Main FastAPI for document ingestion and querying. See this |
| Chainlit UI | 8080/chainlit | User interface for interacting with the RAG system |
| Ray Dashboard | 8265 | Ray dashboard for monitoring and managing tasks |
| Admin UI | 8081/app/ | Main user interface for indexing and viewing indexed documents |
More information about each service is available in its respective documentation page.