Skip to content

Environment Variable Configuration

OpenRAG provides a large range of environment variables that allow you to customize and configure various aspects of the application. This page serves as a comprehensive reference for all available environment variables, providing their types, default values, and descriptions. As new variables are introduced, this page will be updated to reflect the growing configuration options.

OpenRAG validates its credentials once, at startup, and refuses to start on any value published in this repository — the dev defaults in .env.example, the placeholders in the Helm chart, and the values used in the documentation and the test stacks. The same check runs at helm template time for the chart’s values secrets provider.

This exists because the common way a deployment ends up with a known credential is not a weak choice but no choice at all: a template copied verbatim. So the example files ship __GENERATE_ME__ where a credential belongs, and one command fills them in:

Terminal window
python3 scripts/gen_env.py # writes infra/compose/.env
python3 scripts/gen_env.py --check # verify nothing is left unset
VariableRequiredNotes
AUTH_TOKENYes, unless using SSOBootstraps the admin user and guards the API. Minimum 12 characters. Leaving it unset enables the no-auth development mode only when ALLOW_NO_AUTH=true.
POSTGRES_PASSWORDYesMinimum 12 characters. Compose fails closed without it; the chart reads it from postgresql.auth.password.
CHAINLIT_AUTH_SECRETWhen the chat interface is enabledSigns the chat session cookie. Minimum 12 characters; generate with python -c "import secrets; print(secrets.token_urlsafe(32))". Rotating it invalidates live sessions. In the Helm chart the chat interface is off by default (env.WITH_CHAINLIT_UI).
MINIO_ACCESS_KEY / MINIO_SECRET_KEYCompose onlyShared by the MinIO service and Milvus — both sides must match. MINIO_SECRET_KEY has a 12-character minimum. Kubernetes deployments use external object storage instead.
GRAFANA_ADMIN_PASSWORDWhen bundled Grafana is enabledMinimum 12 characters.
OIDC_CLIENT_SECRET, OIDC_TOKEN_ENCRYPTION_KEYSSO onlyFormats are set by your identity provider and by Fernet respectively, so no length floor is applied. See the OIDC guide.
HF_TOKENWhen pulling gated model weights
API_KEY, VLM_API_KEY, EMBEDDER_API_KEY, RERANKER_API_KEY, TRANSCRIBER_API_KEY, WEBSEARCH_API_TOKENPer integrationSet by the provider, so no length floor. Use the literal EMPTY for a local OpenAI-compatible server that requires no credential — that is the documented sentinel and is always accepted.

UVICORN_FORWARDED_ALLOW_IPS is not a secret but belongs in the same conversation: it must name your proxy’s subnet for session cookies and per-IP rate limits to behave correctly behind a reverse proxy. See FastAPI & Access Control.

Any value this project publishes, matched case-insensitively — including the __GENERATE_ME__ marker itself — and any value shorter than 12 characters for the variables marked with a minimum above. The canonical list lives in openrag/core/config/secrets_guard.py; the chart carries the same list and a unit test fails the build if the two disagree.

The startup error names the variables at fault and never prints their values.

ALLOW_INSECURE_SECRETS=true downgrades the refusal to a warning logged on every boot. It is an opt-out, never an opt-in: an unset value means the check is enforced. Use it for disposable development and CI stacks — the bundled API test stack sets it — and not in a deployment that holds real data.

Openrag loads all files into a pivot markdown file format before proceeding to chunking. Some environment variables can be configured to customized this pipeline

VariableTypeDefaultDescription
IMAGE_CAPTIONINGbooltrueIf true, an LLM is used to describe images and convert them into text using a specific prompt. The image in files are replaced by their descriptions
IMAGE_CAPTIONING_URLbooltrueIf true, HTTP/HTTPS image URLs in markdown files are fetched and described by the VLM.
SAVE_MARKDOWNboolfalseIf true, the pivot-format markdown produced during parsing is saved. Useful for debugging and verifying the correctness of the generated markdown.
SAVE_UPLOADED_FILESbooltrueWhen true, uploaded files are stored on disk. You must enable this option if you want Chainlit to show sources while chatting.
CONTENT_DEDUPLICATION_ENABLEDbooltrueRejects a file when identical content already exists in the same partition. Set it to false when a test intentionally indexes duplicates.
PDFLOADERstrPyMuPDFLoaderPDF parsing engine. PyMuPDFLoader (default) is a lightweight, fast, CPU-friendly backend for searchable PDFs. Switch to MarkerLoader for OCR / scanned documents, complex layouts and embedded images (heavier; GPU-friendly). Other options: DoclingLoader, DotsOCRLoader.
PARSE_TIMEOUTint3600Outer wall-clock bound (in seconds) for a single file’s parse stage, whichever loader runs it. Marker and Docling self-limit via their own timeouts, but PyMuPDFLoader has none — this bound stops a wedged parse from stalling indexing: the file fails and is reported instead.
INDEXING_CALLBACK_ALLOW_PRIVATE_URLSboolfalseDevelopment-only escape hatch for the upload callback_url. By default a callback targeting localhost, a private or link-local address is refused (SSRF guard on a caller-supplied URL). Set it to true only in a dev stack whose callback target is a local instance (e.g. http://cozy.localhost:8080). Non-http(s) schemes stay rejected either way.

These settings apply when MarkerLoader is selected (PDFLOADER=MarkerLoader; the default is PyMuPDFLoader). It can be configured using the following environment variables:

VariableTypeDefaultDescription
MARKER_POOL_SIZEint1Number of workers (typically 1 worker per cluster node)
MARKER_MAX_PROCESSESint2Number of subprocesses <-> Number of concurrent PDFs per worker (to increase depending on your available GPU resources)
MARKER_MAX_TASKS_PER_CHILDint20Number of tasks a child (PDF worker) has to process before it gets restarted to clean up memory leaks
MARKER_TIMEOUTint3600Timeout in seconds for marker processes
MARKER_PDFTEXT_WORKERSint2Number of PDF text extractor workers inside marker.
MARKER_CHUNK_SIZEint10Split large PDFs into chunks of this many pages for parallel processing across workers. Use <= 0 to deactivate chunking.

These settings apply when DoclingLoader is selected (PDFLOADER=DoclingLoader):

VariableTypeDefaultDescription
DOCLING_POOL_SIZEint1Number of Docling worker actors in the Ray pool
DOCLING_MAX_TASKS_PER_WORKERint2Maximum number of PDFs processed concurrently per Docling worker
DOCLING_NUM_GPUSfloat0.01Fraction of a GPU reserved per Docling worker in Ray’s resource accounting
OpenAI-Compatible OCR Loader Configuration
Section titled “OpenAI-Compatible OCR Loader Configuration”

Modern OCR pipelines increasingly rely on VLM-based OCR models (such as DeepSeek OCR, DotsOCR, or LightOn OCR) that convert PDF pages into images and feed them into vision-language models with specialized prompts.
This loader integrates that workflow by exposing an OpenAI-compatible API that accepts PDF image pages and returns structured text produced by the OCR-VLM model in Markdown.

The parameters below configure how the OCR loader communicates with the model server, handles retries, manages concurrency, and controls model sampling behavior.

VariableTypeDefaultDescription
OPENAI_LOADER_BASE_URLstringhttp://openai:8000/v1Base URL of the OCR loader (OpenAI-compatible endpoint).
OPENAI_LOADER_API_KEYstringEMPTYAPI key used to authenticate with the OCR service.
OPENAI_LOADER_MODELstringdotsocr-modelOCR VLM model to use (e.g., DotsOCR, DeepSeek OCR, LightOn OCR).
OPENAI_LOADER_TEMPERATUREfloat0.2Sampling temperature. Lower values produce more deterministic OCR results.
OPENAI_LOADER_TIMEOUTint180Maximum request duration (in seconds) before timing out.
OPENAI_LOADER_MAX_RETRIESint2Number of retry attempts for failed OCR requests.
OPENAI_LOADER_TOP_Pfloat0.9Nucleus sampling parameter that limits generation to the top-p probability mass.
OPENAI_LOADER_CONCURRENCY_LIMITint20Maximum number of OCR requests processed concurrently. Useful for multi-page PDF workloads.
OPENAI_LOADER_ENABLE_THINKINGboolunsetOptional chat-template control for OCR VLM models that support enable_thinking; leave unset for Mistral tokenizers, set false to suppress Qwen-style reasoning traces.

OpenRAG provides two deployment options for audio transcription, configurable via the AUDIOLOADER environment variable:

VariableTypeDefaultDescription
AUDIOLOADERstrLocalWhisperLoaderSpecifies the audio loader implementation. Options: LocalWhisperLoader (bundled Whisper service) or OpenAIAudioLoader (external OpenAI API)
Local Whisper Loader ( LocalWhisperLoader )
Section titled “Local Whisper Loader ( LocalWhisperLoader )”

For local whisper loader, here are the options to use

VariableTypeDefaultDescription
WHISPER_MODELstrbaseThe whisper multilingual model to use depending on available resources. Other options: base, small, large, large-v3, etc.
WHISPER_N_WORKERSint2Number of whisper workers
WHISPER_CONCURRENCY_PER_WORKERint2Maximum number of audio transcription tasks processed concurrently by each Whisper worker.
WHISPER_NUM_GPUSfloat0.01Fraction of a GPU reserved per Whisper worker in Ray’s resource accounting.
OpenAI-compatible audio Loader ( OpenAIAudioLoader )
Section titled “OpenAI-compatible audio Loader ( OpenAIAudioLoader )”

The OpenAIAudioLoader option can use any OpenAI-compatible transcription service. Configure its URL, credentials, and model with TRANSCRIBER_BASE_URL, TRANSCRIBER_API_KEY, and TRANSCRIBER_MODEL.

On first startup, these values seed the default STT endpoint in the Admin UI’s Model Endpoints page. Once an STT endpoint is saved, it is the editable source for OpenAIAudioLoader: use it to change the OpenAI-compatible /v1 URL, model, API key, timeout, concurrency, optional language hint, and non-secret provider request options for MOSS, Whisper, or another compatible provider. No OpenRAG restart is required: direct extraction reloads the endpoint before each audio request, and indexer workers refresh it within one minute. The transcription endpoint must implement /audio/transcriptions.

The audio is automatically segmented into chunks using silence detection, then transcribes these chunks in parallel for optimal speed and accuracy.

Here are some other variables related to openai-compatible endpoint.

VariableTypeDefaultDescription
TRANSCRIBER_BASE_URLstrhttp://transcriber:8000/v1Base URL for the transcriber API (OpenAI-compatible endpoint).
TRANSCRIBER_API_KEYstrEMPTYAuthentication key for transcriber service requests.
TRANSCRIBER_MODELstropenai/whisper-large-v3-turboModel identifier exposed by the transcription endpoint.
TRANSCRIBER_MAX_CONCURRENT_CHUNKSint20Per-worker concurrency for the initial STT endpoint seed and the fallback when no saved STT endpoint is available. For a saved endpoint, set Concurrency per worker in Admin UI → Model Endpoints.
TRANSCRIBER_TIMEOUTint3600Maximum duration in seconds allowed for a single transcription request.
TRANSCRIBER_DIRECT_UPLOAD_SUFFIXESstr.wav|.flac|.ogg|.mp3|.mp4|.m4a|.webm|.mpeg|.mpgaPipe-delimited list of audio file suffixes uploaded to the transcriber as-is (no WAV conversion). Other formats are re-encoded to WAV before upload. Trim this list when your transcriber backend (e.g. vLLM/libsndfile) only accepts a subset.
USE_WHISPER_LANG_DETECTORbooltrueWhen enabled, uses a local Whisper-based language detector to identify the source audio language before transcription.
TRANSCRIBER_PORTint8002Host port the bundled vLLM Whisper service (TRANSCRIBER_COMPOSE=extern/transcriber.yaml) is published on (maps to container port 8000). Only read once you uncomment the ports: mapping in that compose include — by default the service is reachable over the Docker network only.
VariableTypeDefaultDescription
CHUNKERstrstructured_sectionDefines the chunking strategy: structured_section or recursive_splitter.
CONTEXTUAL_RETRIEVALbooltrueEnables contextual retrieval to chunk context, a technique introduced by Anthropic to improve retrieval performance (Contextual Retrieval)
CHUNK_SIZEint512Target size of each chunk, in tokens — counted with the LLM tokenizer (tiktoken cl100k_base when the LLM is unreachable), not in characters.
CHUNK_OVERLAP_RATEfloat0.2Fraction of CHUNK_SIZE replayed between consecutive chunks. Applies to recursive_splitter only — structured_section forces overlap to 0.
CONTEXTUALIZATION_TIMEOUTint120Timeout in seconds for individual chunk contextualization LLM calls. Prevents long-running contextualization tasks from blocking the system.
MAX_CONCURRENT_CONTEXTUALIZATIONint10Maximum number of concurrent chunk contextualization tasks. Limits parallel LLM requests to prevent CPU exhaustion during batch indexing.

After files are converted to Markdown, only the text content is chunked. Image descriptions and Markdown tables are not chunked.

Chunker strategies:

  • structured_section (default): Cuts on the document’s own structure instead of on character separators. It detects headings (Markdown #, plus keyword headings such as Titre / Chapitre / Section) and leaf units (e.g. Article L110-1) by matching line content, keeps each leaf atomic, greedily packs consecutive short leaves up to CHUNK_SIZE, and prepends the heading path so every chunk is self-describing at retrieval time. Overlap is always 0 — leaves are atomic, so replaying a tail would only duplicate whole sections, and CHUNK_OVERLAP_RATE is therefore ignored by this strategy. Best for structured documents (legal codes, standards, reports, technical manuals).
  • recursive_splitter: Uses hierarchical text structure (sections, paragraphs, sentences). Based on RecursiveCharacterTextSplitter, it preserves natural boundaries whenever possible while ensuring chunks never exceed CHUNK_SIZE, and replays CHUNK_OVERLAP_RATE of each chunk into the next. Set CHUNKER=recursive_splitter for unstructured prose, or to reproduce the chunking of earlier OpenRAG releases.

Our embedder is OpenAI-compatible and runs on a VLLM instance configured with the following variables:

VariableTypeDefaultDescription
EMBEDDER_MODEL_NAMEstrQwen/Qwen3-Embedding-0.6BHuggingFace Embedding model served by VLLM .i.e Qwen/Qwen3-Embedding-0.6B or jinaai/jina-embeddings-v3. Upgrading from 2.2.x under Docker Compose: the default was jinaai/jina-embeddings-v3. If your .env does not set this variable, set it to the model your data was indexed with before upgrading, since changing the model needs a reindex. Without it the vLLM container serves Qwen while the saved endpoint still asks for jina: every embedding call fails and /ready reports checks.embedder: unavailable. MODEL_ENDPOINT_SYNC_ON_BOOT=true does not get around this: the sync refuses to move an endpoint that holds indexed files to another model and logs a warning. The legacy EMBEDDING_MODEL, when set, takes precedence over this variable. The Helm chart has defaulted to Qwen since before 2.2.x; there the embedder’s model is vllm.embedderModelName together with the embedder entry of vllm.servingEngineSpec.modelSpec, which an override must restate in full.
EMBEDDER_BASE_URLstrhttp://vllm:8000/v1Base URL of the embedder (OpenAI-style).
EMBEDDER_API_KEYstrEMPTYAPI key for authenticating embedder calls.
EMBEDDER_EXTRA_ARGSstr(empty)Extra vllm serve flags appended to the bundled vLLM embedder’s command (vllm-gpu / vllm-cpu in infra/compose/docker-compose.yaml). They come after the built-in flags, so repeating one overrides it, e.g. --gpu_memory_utilization 0.1.
MAX_MODEL_LENint2047Maximum context length (in tokens) supported by the embedding model. Chunks exceeding this limit are truncated (truncate_prompt_tokens = this value − 1). Keep it below the model’s real context boundary.
EMBEDDER_TIMEOUTfloat120.0Per-request HTTP timeout (in seconds) for embedding calls. Raise it for slow remote endpoints.
EMBEDDER_BATCH_SIZEint32Number of chunks sent per embedding request; large documents are split into batches of this size.
EMBEDDER_CONCURRENCYint4Maximum number of embedding requests in flight at once.

If you prefer to use an external embedding service, simply comment out the embedder service in the docker-compose.yaml and provide the variables above in your environment.

Model endpoints (embedder, LLM, VLM, reranker) are stored in a database-backed registry and can be edited at runtime from the admin UI. On first boot, one default endpoint per type is seeded from the *_BASE_URL / *_MODEL / *_API_KEY env vars documented above, so existing env-only deployments keep working with no admin action. After that first seed, the database is the source of truth — changing an env var no longer overwrites an endpoint an operator may have edited.

VariableTypeDefaultDescription
MODEL_ENDPOINT_SYNC_ON_BOOTboolfalseWhen true, the endpoint each type was auto-seeded with is re-synced from the environment on every boot — its endpoint, model_name and api_key are refreshed from the *_BASE_URL / *_MODEL / *_API_KEY values, and batch_size / timeout follow only when their own env var is set. This lets operators manage that endpoint via env vars + a pod rollout (e.g. a Helm values change), including changing the model and rotating the API key. Any endpoint created by hand is left untouched. Keep it false (the default) to preserve the “database wins after first boot” behavior. An embedder that already holds indexed files does not change model through the sync: the model change is refused and logged as a warning on every boot while env and the database disagree (EMBEDDER_EDIT_AFFECTS_INDEXED_DATA, as in the admin API), the row keeps its model, and its URL, API key and env-set tunables are still synced. A URL change alone syncs; if the new server serves another model, /ready reports checks.embedder: unavailable, except for an infinity or tei embedder, whose probe checks only /health and cannot see the model. To change the model of indexed data, use the admin API with acknowledge_indexed_data=true.
READINESS_REQUIRE_EMBEDDERboolfalseWhen true, /ready returns 503 while the default embedder is unavailable or unresolvable (not on a probe timeout, nor when model discovery itself fails). Off by default because every replica shares the embedder: under Kubernetes its outage, or a restart while the model loads, takes every replica out of the Service at once, admin API and admin UI included. checks.embedder in the /ready body reports the embedder either way. Before enabling it, check that /ready reports checks.embedder: ok on your deployment: the probe needs the configured model name verbatim in the endpoint’s GET /models list, so an embedder that works can still read unavailable, for example an Ollama model configured without its :tag, or an endpoint that has no /models route.

Our system uses two databases that work together:

  • Vector Database (VDB)

The vector database stores embeddings and is configured using the following environment variables:

VariableTypeDefaultDescription
VDB_HOSTstrmilvusHostname of the vector database service
VDB_PORTint19530Port on which the vector database listens
VDB_CONNECTOR_NAMEstrmilvusConnector/driver to use for the vector DB. Currently only milvus is implemented
VDB_COLLECTION_NAMEstrvdb_testName of the collection storing embeddings
VDB_HYBRID_SEARCHbooltrueTo activate hybrid search (semantic similarity + Keyword search)
VDB_ENABLE_INSERTIONbooltrueEnable or disable vector database insertion. When disabled, documents are processed but not inserted into Milvus. Useful for testing.
VDB_TIMEOUTfloat120.0Per-request timeout (seconds) applied to the Milvus sync and async clients

These variables can be overridden when using an external vector database service.

  • Relational Database (RDB)

The vector database implementation relies on an underlying PostgreSQL database that stores metadata about partitions and their owners (users). For more information about the data structure, see the data model.

The PostgreSQL database is configured using the following environment variables:

VariableTypeDefaultDescription
POSTGRES_HOSTstrrdbHostname of the PostgreSQL database service
POSTGRES_PORTint5432Port on which the PostgreSQL database listens
POSTGRES_USERstrrootUsername for database authentication
POSTGRES_PASSWORDstrroot_passwordPassword for database authentication
POSTGRES_DATABASEstrpartitions_for_collection_<VDB_COLLECTION_NAME>Database used for OpenRAG relational metadata. If unset, OpenRAG derives the historical name from the vector collection.
POSTGRES_AUTO_CREATE_DBbooltrueCreates the database automatically when it is missing. Keep this for local compose; set it to false for managed Postgres where the app role has no CREATEDB.
POSTGRES_RUN_MIGRATIONSbooltrueRuns Alembic migrations during app startup. Set it to false when migrations are applied by a deployment Job or init step.
POSTGRES_POOL_MIN_SIZEint5Minimum size of the async PostgreSQL connection pool.
POSTGRES_POOL_MAX_SIZEint20Maximum size of the async PostgreSQL connection pool.
POSTGRES_COMMAND_TIMEOUTint30Timeout in seconds for PostgreSQL commands issued through the async pool.
  • Object Storage (MinIO)

Milvus stores its data in a MinIO object store, whose credentials are required (no default) — the compose stack refuses to start if they are unset. Generate strong random values (e.g. openssl rand -hex 16). The same values are shared between the minio service and Milvus, so both sides must match.

VariableTypeDefaultDescription
MINIO_ACCESS_KEYstr(required)MinIO access key, shared by the minio service and Milvus. No default.
MINIO_SECRET_KEYstr(required)MinIO secret key, shared by the minio service and Milvus. No default.

The main Docker Compose stack keeps the historical host-path defaults. Set these variables when you want to move state elsewhere, including Docker named volumes.

When using host paths with the non-root API image, make sure the mounted directories are writable by the container user. If that is not practical for your deployment, use the named-volume profile instead.

For an opt-in named-volume profile, copy the values from infra/compose/.env.named-volumes.example into your .env.

VariableDefaultDescription
DATA_VOLUME../../dataOpenRAG uploaded files and app data mounted at /app/data.
MODEL_WEIGHTS_VOLUME~/.cache/huggingfaceModel cache mounted at /app/model_weights.
VLLM_CACHE/root/.cache/huggingfaceHugging Face cache used by vLLM, reranker, and transcriber services.
DB_VOLUME../../dbPostgreSQL data mounted at /var/lib/postgresql/data.
MILVUS_VOLUME_DIRECTORY./volumesParent directory for Milvus, etcd, and MinIO host-path storage.
MILVUS_MQ_TYPEdefaultMilvus message queue. Keep the existing value during a version upgrade; fresh installations can use the default.
MILVUS_COMPOSEmilvus/milvus.yamlMilvus compose include. Use milvus/milvus.named-volumes.yaml for the named-volume profile.
ETCD_VOLUMEetcdMilvus etcd named volume, used only with MILVUS_COMPOSE=milvus/milvus.named-volumes.yaml.
MINIO_VOLUMEminioMilvus object storage named volume, used only with MILVUS_COMPOSE=milvus/milvus.named-volumes.yaml.
MILVUS_VOLUMEmilvusMilvus named volume, used only with MILVUS_COMPOSE=milvus/milvus.named-volumes.yaml.

The system uses two types of language models:

  • LLM (Large Language Model): The primary model for text generation and chat interactions
  • VLM (Vision Language Model): Used for describing images (see IMAGE_CAPTIONING) and, to reduce load on the primary LLM, also handles contextualization tasks (see CONTEXTUAL_RETRIEVAL)

These are external services to provide !!!

VariableTypeDefaultDescription
BASE_URLstr(required)Base URL of the LLM API endpoint
MODELstr(required)Model identifier for the LLM
API_KEYstr(unset)API key for authenticating with the LLM service
LLM_ENABLE_THINKINGbool(unset)Optional chat-template control for models that support enable_thinking; leave unset for Mistral tokenizers, set false to suppress Qwen-style reasoning traces
LLM_SEMAPHOREint10Maximum number of concurrent requests to allow for the LLM service
LLM_OVERRIDE_ALLOW_CUSTOM_ENDPOINTboolfalseHonor a client-supplied base_url/api_key in metadata.llm_override. Off by default; read the trade-off below before enabling.
MAX_LLM_CONTEXT_SIZEint8192Fallback context window of the LLM answering a request. An LLM endpoint’s own Max context size (Model Endpoints) wins; without one, the max_model_len the endpoint reports on /v1/models is used (vLLM reports it, most gateways and hosted APIs don’t), and this value only when neither is known. The admin UI shows which one applies to each endpoint. Requests whose prompt (with the tools definitions and tool-call history a client sends) + max_tokens exceed the window are rejected with a 413 error, as are those the answer instructions push over it (CONTEXT_WINDOW_EXCEEDED, sent as an error event on a streaming request). The documents and web results given to the LLM are cut to fit what the window leaves, so a value smaller than the model’s real window means the LLM gets fewer chunks than the retrieval preset’s top_n.
MAX_OUTPUT_TOKENSint1024Default output-token budget (max_tokens) applied to chat completions when the request doesn’t set one explicitly.

A client can always override the model name for a single request via metadata.llm_override (see the API reference). LLM_OVERRIDE_ALLOW_CUSTOM_ENDPOINT=true additionally honors base_url and api_key from that object, so the request is served by a provider of the client’s choosing rather than the configured one.

It exists for deployments whose clients already send the full object and would otherwise break. Prefer registering a named endpoint under /model-endpoints and binding it to the partition (chat_llm) — same outcome, none of the trade-off below.

What enabling it means. Any caller who can reach /v1/chat/completions can make the server issue https POSTs from inside your network. The request shape is pinned, which is what bounds the exposure:

  • https only — plaintext internal services are unreachable.
  • The path is always {base_url}/chat/completions ({base_url}/completions for the legacy /v1/completions route); a query string, fragment or .. segment — percent-encoded or not — is rejected with a 400, so the override cannot be aimed at an arbitrary internal path.
  • Redirects are not followed, so a target cannot bounce the server elsewhere.
  • The server’s own API key is never forwarded; an override without api_key sends no Authorization header at all.

What remains reachable is therefore essentially other https LLM gateways — including an internal one that trusts its network rather than a credential, which such a caller could then use without holding its key.

What it does not change. It grants no read access a caller does not already have: /search returns the same partition content directly. What changes is the way data leaves — as outbound LLM traffic from the server’s egress rather than as a user read. That matters against a DLP or approved-subprocessor constraint, not against a caller who was never authorized in the first place.

Enable only where every API caller is already trusted with both.

VariableTypeDefaultDescription
VLM_BASE_URLstr(required)Base URL of the VLM API endpoint
VLM_MODELstr(required)Model identifier for the VLM
VLM_API_KEYstr(unset)API key for authenticating with the VLM service
VLM_ENABLE_THINKINGbool(unset)Optional chat-template control for models that support enable_thinking; leave unset for Mistral tokenizers, set false to suppress Qwen-style reasoning traces
VLM_SEMAPHOREint10Maximum number of concurrent requests to allow for the VLM service
VariableTypeDefaultDescription
RAG_MODEstrChatBotRagHow the pipeline turns the conversation into search queries. ChatBotRag (default) uses the LLM and the chat history to generate contextualized search queries; SimpleRag skips query generation and searches directly on the raw last user message.

The retriever fetches relevant documents from the vector database based on query similarity. Retrieved documents are then optionally reranked to improve relevance.

VariableTypeDefaultDescription
RETRIEVER_TYPEstrsingleRetrieval strategy to use. Options: single, multiQuery, hyde
RETRIEVER_TOP_Kint50Number of documents to retrieve before reranking.
SIMILARITY_THRESHOLDfloat0.6Minimum similarity score (0.0-1.0) for document retrieval. Documents below this threshold are filtered out
WITH_SURROUNDING_CHUNKSboolfalseWhen enabled, retrieves adjacent chunks (preceding and following) for each matched document to provide additional context.
INCLUDE_RELATEDbooltrueExpand results with chunks from files sharing the matched file’s relationship_id (see Linked files).
INCLUDE_ANCESTORSbooltrueExpand results with chunks from ancestor files in the parent/child file hierarchy (see Linked files).
RELATED_LIMITint10Maximum number of related/ancestor chunks fetched per matched result when expansion is enabled.
MAX_DEPTHint10Maximum ancestor depth traversed when INCLUDE_ANCESTORS is enabled.
RETRIEVER_ALLOW_FILTERLESS_FALLBACKbooltrueWhen a temporally-filtered retrieval returns no documents, re-run the query without the filter. Set to false for strict temporal retrieval.
RETRIEVER_MAX_PARTITION_CONCURRENCYint16Upper bound on how many per-partition retrievals run concurrently per retrieval call. Bounds fan-out for multi-partition / openrag-all searches; small fan-outs stay fully parallel. Note: a multi-query request (multiQuery/hyde) issues one such call per sub-query, so peak concurrency can reach N × this value, not a flat per-request cap.
StrategyDescription
singleStandard semantic search using the original query. Fast and efficient for most queries
multiQueryGenerates multiple query variations to improve recall. Better coverage for ambiguous or complex questions
hydeHypothetical Document Embeddings - generates a hypothetical answer then searches for similar documents

The reranker enhances search quality by re-scoring and reordering retrieved documents according to their relevance to the user’s query. Three providers are supported: Infinity (default), OpenAI-compatible endpoints, and Hugging Face Text Embeddings Inference (TEI).

VariableTypeDefaultDescription
RERANKER_ENABLEDbooltrueEnable or disable the reranking mechanism
RERANKER_PROVIDERstrinfinityReranker backend to use. Accepted values: infinity, openai, tei
RERANKER_MODELstrAlibaba-NLP/gte-multilingual-reranker-baseModel used for reranking documents. Ignored by the tei provider (a TEI instance serves a single fixed model)
RERANKER_TOP_Kint10Number of chunks kept after reranking and given to the LLM, for every retrieval preset that leaves top_n empty (a preset’s own top_n overrides it), as many of them as fit in the answering LLM’s context window (see MAX_LLM_CONTEXT_SIZE). Must be greater than 0. The retrieval presets an earlier release seeded store top_n: 10 and keep it on upgrade: clear the field in the admin UI for them to follow this value. Increase for better results if your LLM has a wider context window
RERANKER_BASE_URLstrhttp://reranker:7997Base URL of the reranker service
RERANKER_API_KEYstrEMPTYAPI key for the reranker service, sent as a Bearer token when set. Whether a key is required depends on your endpoint
RERANKER_TIMEOUTfloat60.0HTTP timeout in seconds for reranker requests
RERANKER_SEMAPHOREint5Maximum number of concurrent reranking requests. Adjust based on your server capacity
RERANKER_EXTRA_ARGSstr(empty)Extra vllm serve flags appended to the bundled vLLM reranker’s command (extern/reranker/openai.yaml, RERANKER_PROVIDER=openai). Alibaba-NLP/gte-multilingual-reranker-base needs --hf-overrides '{"architectures": ["GteNewForSequenceClassification"]}': vLLM doesn’t recognise its NewForSequenceClassification architecture on its own.
RERANKER_PORTint7997 (infinity) / 8000 (openai)Host port the bundled reranker service is published on. Only read by the compose includes (extern/reranker/*.yaml), and only once you uncomment their ports: mapping — by default the service is reachable over the Docker network only, so publishing it is just for host-side debugging or direct calls.
ProviderRERANKER_PROVIDER valueDescription
InfinityinfinityUses the Infinity server via its native client. Default port: 7997
OpenAI-compatibleopenaiUses any reranker endpoint implementing the {model, query, documents, top_n} → {results: [...]} rerank contract (e.g. vLLM, LiteLLM). Default port: 8000
TEIteiUses a Hugging Face Text Embeddings Inference server via its native /rerank API (which is not OpenAI-compatible: texts instead of documents, no model/top_n fields, bare-array response). Default port: 8080. Requests are batched to 32 texts to fit TEI’s default --max-client-batch-size

The RAG pipeline ships with preconfigured prompts bundled inside the package at openrag/prompts/templates. Here are the available Prompt Templates in that folder.

Template FilePurpose
sys_prompt_tmpl.txtSystem prompt that defines the assistant’s behavior and role
spoken_style_answer_tmpl.txtTemplate for converting responses to a more natural, conversational spoken style (oral / audio type of answer)
query_contextualizer_tmpl.txtTemplate for adding context to user queries
chunk_contextualizer_tmpl.txtTemplate for contextualizing document chunks during indexing
image_captioning_tmpl.txtTemplate for generating image descriptions using the VLM
asr_transcription_tmpl.txtEmpty by default so external transcription keeps the provider’s native behavior. Configure it in Prompt Library and select it on an indexation preset when AUDIOLOADER=OpenAIAudioLoader.
hyde.txtHypothetical Document Embeddings (HyDE) query expansion template
multi_query_pmpt_tmpl.txtTemplate for generating multiple query variations

To customize prompt:

  1. Copy the bundled templates: Copy openrag/prompts/templates to a folder of your choice
  2. Create your custom folder: Rename it to something meaningful, e.g., my_prompt
  3. Modify the prompts: Edit any prompt templates within your new folder
  4. Update configuration: Point PROMPTS_DIR at your custom prompts directory
.env
# Use custom prompts
export PROMPTS_DIR=/path/to/my_prompt
VariableTypeDefaultDescription
PROMPTS_DIRstr(bundled openrag/prompts/templates)Path to a directory of prompt templates. Unset uses the templates bundled in the package; set it only to override with a custom directory.

OpenRAG logs with Loguru on the process stderr, and nowhere else. Docker and Kubernetes capture that stream; a collector ships it to Loki (see Loki logs). “stderr” is the conventional diagnostic channel, not an error level: an INFO line goes there too.

VariableTypeDefaultDescription
LOG_LEVELstrINFOMinimum level emitted. DEBUG logs user queries and other request data; keep it for short-lived troubleshooting.
LOG_FORMATtext | jsontexttext is the colorized human format below. json writes one flat JSON object per line, no colour, for log collectors; it also routes every library’s stdlib logs through the same sink (the five named ones — asyncio, httpcore, httpx, urllib3 and openai — capped at WARNING), so with a collector attached LOG_LEVEL=DEBUG is a troubleshooting setting, not a production one.
Logging message in the terminal...
LEVEL | module:function:line - message [context_key=value]

Since this release every request-scoped line ends with [request_id=req_…] in text mode too — the same correlation id the response carries in its X-Request-ID header.

{"ts":"2026-09-07T14:03:12.481000+00:00","level":"INFO","logger":"api.routers.user.chat","function":"chat_completions","line":212,"msg":"Retrieved 8 documents","request_id":"req_7f3c…","partition":"docs"}

Reserved keys: ts, level, logger, function, line, msg, exception (only when a traceback is attached). Every field bound with logger.bind() is emitted at the top level; a bound field named like a reserved key is prefixed extra_.

LevelWhat You’ll See in Logs
WARNINGPotential issues that don’t stop execution: approaching rate limits, deprecated features used, retryable failures, configuration concerns. Review these periodically.
DEBUGDetailed diagnostic information including variable states, intermediate processing steps, and function entry/exit points. Useful during development and troubleshooting.
INFOStandard operational messages showing normal application behavior: server startup, request handling, major workflow stages. This is the typical production level.

Ray is used for distributed task processing and parallel execution in the RAG pipeline. This configuration controls resource allocation, concurrency limits, and serving options.

VariableTypeDefaultDescription
RAY_POOL_SIZEint1Number of indexer worker actors in the pool. Total indexing capacity = RAY_POOL_SIZE × RAY_MAX_TASKS_PER_WORKER.
RAY_MAX_TASKS_PER_WORKERint50Maximum number of files processed concurrently per indexer worker actor
RAY_DASHBOARD_PORTint8265Ray Dashboard port used for monitoring. In production, comment out this line to avoid exposing the port, as it may introduce security vulnerabilities.
RAY_DASHBOARD_HOSTstr127.0.0.1Interface the embedded Ray dashboard binds to. Defaults to loopback because the Ray dashboard/job-submission API is unauthenticated (CVE-2023-48022). Set to 0.0.0.0 only when the dashboard port is firewalled or sits behind an authenticating proxy. Ignored when RAY_ADDRESS is set.
RAY_METRICS_EXPORT_PORTint(unset)Port the embedded Ray’s metrics agent listens on. Unset, Ray picks a random port, so the metrics recorded in Ray actors (indexing, and the inference calls the indexing workers make) are exported but no scrape config can reach them. The Compose monitoring overlay sets it to 8091 and scrapes it; the Helm chart sets it to 8090 on the openrag pod when ray.enabled=false. Applies under ENABLE_RAY_SERVE=true too. Under embedded Ray Serve with WITH_CHAINLIT_UI, Chainlit listens on CHAINLIT_PORT (default 8090), so keep the two different. Unauthenticated and bound to every interface: never publish it on the host. Ignored when RAY_ADDRESS is set.
RAY_ADDRESSstr(unset)When set, attach to an external Ray cluster at this address (e.g. ray://HEAD_IP:10001) instead of starting an embedded cluster in-process. In this mode the app does not start a local dashboard — the head node owns it. See Ray Cluster deployment.
VariableTypevalueDescription
RAY_DEDUP_LOGSnumber0Turns off Ray log deduplication that appears across multiple processes. Set to 0 to see all logs from each process. Required (0) with LOG_FORMAT=json: the deduplicated survivor is rewritten as {…} [repeated 2x across cluster], which is no longer JSON. The logging overlay and the Helm chart set it.
RAY_COLOR_PREFIXnumber0Turns off the ANSI colorization of Ray’s (Actor pid=N) relay prefix, which is applied even when the output is a pipe. Required (0) with LOG_FORMAT=json, or every relayed worker line reaches the collector prefixed with escape sequences and fails to parse as JSON.
RAY_ENABLE_RECORD_ACTOR_TASK_LOGGINGnumber1Enables logs at task level in the Ray dashboard for better debugging and monitoring.
RAY_task_retry_delay_msnumber3000Delay (in milliseconds) before retrying a failed task. Controls the wait time between retry attempts.
RAY_ENABLE_UV_RUN_RUNTIME_ENVnumber0Controls UV runtime environment integration. Critical: Must be set to 0 when using the newest version of UV to avoid compatibility issues.
RAY_memory_monitor_refresh_msnumber250 msTo control the frequency of memory usage checks and task or actor termination if needed. If you set this value to 0, task killing is disabled.

Ray Serve enables deployment of the FastAPI app as a horizontally scalable service.

By default (ENABLE_RAY_SERVE=false) OpenRAG runs under uvicorn with a single worker. This is intentional: the app initializes Ray and its named actors (Indexer, Vectordb, TaskStateManager, …) at import time, so a second uvicorn worker would start its own isolated Ray cluster with duplicate actors, fragmenting task state and vector-DB access. Concurrency within the single worker comes from the async app and from Ray itself — not from multiple uvicorn workers (there is intentionally no API_NUM_WORKERS knob).

To scale the HTTP layer, enable Ray Serve — it runs RAY_SERVE_NUM_REPLICAS replicas inside one shared Ray cluster:

Terminal window
ENABLE_RAY_SERVE=true
RAY_SERVE_NUM_REPLICAS=4

For multi-node distributed deployments, see Distributed Deployment in a Ray Cluster.

VariableTypeDefaultDescription
ENABLE_RAY_SERVEboolfalseEnable Ray Serve deployment mode
RAY_SERVE_NUM_REPLICASint1Number of service replicas for load balancing
RAY_SERVE_HOSTstr0.0.0.0Host address for the Ray Serve deployment
RAY_SERVE_PORTint8080Port for the Ray Serve FastAPI endpoint
CHAINLIT_PORTint8090Port for the Chainlit UI interface if ray serve is enable ENABLE_RAY_SERVE. If not chainlit UI is simply a subroute (/chainlit see this) of the FastAPI base_url

Web search allows the LLM to augment RAG document context with live web results. It is disabled by default — set WEBSEARCH_API_TOKEN to enable it.

VariableTypeDefaultDescription
WEBSEARCH_PROVIDERstrstaanWeb search provider to use. Currently supported: staan.
WEBSEARCH_API_TOKENstr""API token for the web search provider. If empty, web search is disabled.
WEBSEARCH_BASE_URLstr(provider default)Base URL of the web search provider API.
WEBSEARCH_TOP_Kint5Number of web search results to return.
WEBSEARCH_LANGstrfr-FRLanguage/market code for web search queries.
WEBSEARCH_MAX_TOKENSint2000Maximum token budget for all web sources combined in the LLM context. This budget is reserved from the global context window when web results are present.
WEBSEARCH_FETCH_CONTENTbooltrueWhen enabled, fetches actual page content from the top URLs instead of relying on short search snippets.
WEBSEARCH_FETCH_MAX_RESULTSint3Number of top URLs to fetch content from (the remaining results use their search snippet).
WEBSEARCH_FETCH_TIMEOUTfloat1.0Per-URL timeout in seconds for content fetching. URLs that don’t respond within this time fall back to their snippet.
WEBSEARCH_FETCH_MAX_TOKENSint500Maximum approximate tokens of content to extract per page. Content is truncated at word boundaries.
WEBSEARCH_FETCH_VERIFY_SSLbooltrueWhether to verify SSL certificates when fetching page content. Set to false only for internal CAs — fetched pages are injected into the LLM context as cited sources.

The map & reduce mechanism processes documents by fetching chunks (map phase), filtering out irrelevant ones and summarizing relevant content (reduce phase) with respect to the user’s query. The algorithm works as follows:

  1. Initially fetches a batch of documents for processing
  2. Evaluates relevance and continues expanding the search if needed
  3. Stops expansion when the last MAP_REDUCE_EXPANSION_BATCH_SIZE chunks are all irrelevant
  4. Otherwise, continues fetching additional documents up to MAP_REDUCE_MAX_TOTAL_DOCUMENTS

When MAP_REDUCE_DEBUG is enabled, the mechanism logs detailed information to ./logs/map_reduce.md.

VariableTypeDefaultDescription
MAP_REDUCE_INITIAL_BATCH_SIZEint10Number of documents to process in the initial mapping phase
MAP_REDUCE_EXPANSION_BATCH_SIZEint5Number of additional documents to fetch when expanding the search (also used as the threshold for stopping)
MAP_REDUCE_MAX_TOTAL_DOCUMENTSint20Maximum total number of documents (chunks) to process across all iterations
MAP_REDUCE_DEBUGboolfalseEnable debug logging for map & reduce operations. Logs are written to ./logs/map_reduce.md

By default, our API (FastAPI) uses uvicorn for deployment. One can opt in to use Ray Serve for scalability (see the ray serve configuration)

The following environment variables configure the FastAPI server and control access permissions:

VariableTypeDefaultDescription
APP_PORTnumber8000Port number on which the FastAPI application listens for incoming requests.
AUTH_TOKENstringEMPTYAuthentication token used to bootstrap the admin user and access protected API endpoints. If it is empty, the API fails closed unless ALLOW_NO_AUTH=true is explicitly set for local development.
ALLOW_NO_AUTHbooleanfalseEnables the no-auth local development bypass when AUTH_MODE=token and AUTH_TOKEN is unset. Never enable this in production.
SUPER_ADMIN_MODEbooleanfalseEnables super admin privileges when set to true, granting unrestricted access to all operations and bypassing standard access controls. This is for debugging
DEFAULT_FILE_QUOTAint-1Default per-user file quota. <0 disables quotas globally; >=0 sets the default limit when a user has no explicit quota.
PREFERRED_URL_SCHEMEstringnullURL scheme (http or https) used when generating URLs in API responses (e.g., task_status_url). When running behind a reverse proxy that terminates SSL, set this to https to ensure generated URLs use the correct scheme. If unset, the scheme from the incoming request is used.
METRICS_TOKENstringunsetBearer a Prometheus scraper must send on GET /metrics. The route bypasses the auth middleware and admin tokens are not accepted there. It fails closed: with the token unset and METRICS_ALLOW_UNAUTHENTICATED off, every scrape gets 403. The compose monitoring overlay forwards this value to its Prometheus. See Prometheus metrics.
METRICS_ALLOW_UNAUTHENTICATEDbooleanfalseServe GET /metrics to anyone who can reach the API port when METRICS_TOKEN is unset. Only for deployments that block /metrics at the edge and scrape from inside the network; a configured token always wins.
ASSISTANT_NAMEstringemptyName used in the assistant’s greeting. Leave unset for a neutral introduction; set it for a white-label deployment.
CORS_EXTRA_ORIGINSstring(unset)Semicolon-separated list of additional origins allowed by CORS (e.g. https://app.example.com;https://other.example.com). Extends the default list without replacing it.
UVICORN_FORWARDED_ALLOW_IPSstring127.0.0.1Comma-separated CIDRs/IPs (or *) whose X-Forwarded-* headers uvicorn trusts. Required when OpenRAG runs behind a reverse proxy that lives outside loopback (typical docker-compose / k8s — including the bundled admin-ui proxy). Otherwise X-Forwarded-Proto is dropped and OIDC cookies ship with Secure=False even over HTTPS, and X-Forwarded-For is dropped so per-user rate limits collapse onto the proxy’s single IP. Set this to your proxy’s subnet, not * — see the proxy-trust caution under Rate Limiting for why * can be spoofed.
MAX_UPLOAD_SIZE_MBint1024Maximum accepted upload size, in MB. 0 or a negative value means unlimited.
MAX_PARTITIONS_PER_USERint100Maximum number of partitions a non-admin user may own. -1 disables the cap (unlimited). Admin users always bypass it.
APP_UIDint1000UID the API container drops to before running the app. Override when your host user is not UID 1000 and bind-mounted folders (data/, logs/) would otherwise not be writable by the container user.
WITH_OPENAI_APIbooltrueMount the OpenAI-compatible routers (/v1/*). Note: they stay mounted while WITH_CHAINLIT_UI=true, since Chainlit consumes them.
WITH_CHAINLIT_UIbooltrueMount the bundled Chainlit chat UI under /chainlit (plus its root assets, e.g. the pdf.js worker for source previews).

Per-identity request rate limiting, tiered by path prefix. Requests are keyed on the authenticated user id, falling back to the client IP for unauthenticated paths (/auth/*). Admin users bypass rate limiting entirely. Limits use a moving window and are enforced per worker/replica — front OpenRAG with shared storage (e.g. Redis) if you scale out and need a global budget. Exceeding a limit returns 429 with a Retry-After header.

Limit values use the <count>/<period> format from the limits library (e.g. 120/minute, 10/second).

VariableTypeDefaultDescription
RATE_LIMIT_ENABLEDbooltrueMaster switch for request rate limiting. When false, no limits are applied and malformed limit values are ignored.
RATE_LIMIT_DEFAULTstr600/minuteLimit applied to every path except the tiers below.
RATE_LIMIT_AUTHstr60/minuteLimit for /auth/* (login/callback/logout). Keyed on client IP because callers are unauthenticated there — keep it high enough that a shared corporate/NAT egress IP does not throttle a legitimate login rush.
RATE_LIMIT_CHATstr120/minuteLimit for /v1/* (chat completions, tools).
RATE_LIMIT_AUTH_FAILUREstrRATE_LIMIT_AUTH, else 20/minuteSeparate, stricter budget for failed authentication attempts, keyed by client IP (brute-force protection). Falls back to RATE_LIMIT_AUTH when unset, then to 20/minute. Disabled together with RATE_LIMIT_ENABLED=false.
RATE_LIMIT_EXEMPT_PATHSstr/chainlit/,/assets/Comma-separated path prefixes the limiter skips, matched with startswith. These are auth-bypassed (Chainlit does its own header auth), so requests there carry no user and can only be keyed by IP. Chainlit’s Socket.IO transport also issues one HTTP request per packet when it long-polls. Keep the trailing slash so a sibling like /chainlithack stays rate-limited rather than being swept into the /chainlit exemption. Set-but-empty (RATE_LIMIT_EXEMPT_PATHS=) removes all exemptions; /auth/* is never exempt.

The admin UI is a React SPA (the document ingestion, indexing & management interface) served by the admin-ui (nginx) container. Every VITE_* setting is baked into the bundle at build time — Vite inlines them when the image is built, so they are not read at container runtime. After changing one, rebuild the image: docker compose build admin-ui.

How the UI reaches the API (same-origin). The browser only ever talks to a single origin — http://<host>:ADMIN_UI_PORT. nginx inside the admin-ui container serves the static SPA under /app/ and reverse-proxies every other path (/v1, /auth, /chainlit, /indexer, …) to the API at openrag:8080 over the Docker network. Because the bundle is built with VITE_API_BASE_URL="", its API calls are relative, so they land back on that same origin — there is no CORS, and the OIDC openrag_session cookie is first-party. You don’t even need to publish the API’s own APP_PORT to the host; the UI reaches the backend internally over the compose network. Set VITE_API_BASE_URL only for a browser-direct build, where the SPA calls the API on a different origin — then also add that origin to CORS_EXTRA_ORIGINS.

flowchart TD
    B["Browser — single origin<br/>http://HOST:ADMIN_UI_PORT"]
    subgraph AUC["admin-ui container"]
        N{"nginx :8080<br/>route by path"}
        SPA["Static SPA files<br/>(/app/*)"]
    end
    API["openrag:8080<br/>API service · Docker network"]

    B -->|"GET /app/ (page load)"| N
    B -->|"fetch /v1, /auth, /users, /indexer …<br/>relative → same origin, no CORS"| N
    N -->|"/app/*"| SPA
    N -->|"everything else"| API
VariableTypeDefaultDescription
ADMIN_UI_PORTnumber8081Host port the admin UI (nginx) is published on. Serves /app/ and reverse-proxies /auth, /v1, /chainlit, … to the backend, so it is the OIDC front door (OIDC_REDIRECT_URI targets this port). Deploy-time (not a VITE_* build arg).
GRAFANA_URLstring""Runtime, browser-reachable URL for the Grafana dashboard opened from System → Metrics. Restart the API after changing it. When this is empty or invalid, the action explains how to configure the dashboard instead of opening it.
VITE_API_BASE_URLstring"" (same-origin)API base baked into the SPA. Empty (default) = same-origin: nginx reverse-proxies the API over the Docker network, so the UI works on any host/IP with no CORS. Set to an absolute URL only for a browser-direct build — then list the UI’s origin in CORS_EXTRA_ORIGINS.
VITE_BASE_PATHstring/app/Sub-path the SPA is served under; must match the nginx location.
VITE_GRAFANA_URLstring""Build-time fallback for deployments whose API does not expose GRAFANA_URL. New deployments should use the runtime setting instead.
VITE_APP_NAMEstringOpenRAGApplication display name used in the UI branding.
VITE_MOCK_APIbooleanfalseDevelopment only — serves in-browser MSW API mocks when true. Ignored in production builds.

See this for chainlit authentification

See this for chainlit data persistency

VariableTypeDefaultDescription
DEFAULT_LANGUAGEstr“UI language for Chainlit and the Admin UI (e.g. en-US, fr). When unset, the browser language is used, with en-US as the final fallback.

OpenRAG ships a standalone Model Context Protocol server (openrag/api/mcp/server.py) that exposes retrieval to MCP clients. It runs as its own process (not part of the default compose stack). These variables configure the FastMCP transport binding and the search-tool defaults/bounds applied before a request reaches the retrieval service.

VariableTypeDefaultDescription
OPENRAG_MCP_SERVER_NAMEstrOpenRAG MCPDisplay name advertised by the MCP server.
OPENRAG_MCP_HOSTstr0.0.0.0Interface the MCP server binds to.
OPENRAG_MCP_PORTint8081Port the MCP server listens on.
OPENRAG_MCP_PATHstr/mcpHTTP path the MCP endpoint is served under.
OPENRAG_MCP_DEFAULT_TOP_Kint5Number of chunks the search tool returns when the caller doesn’t specify top_k.
OPENRAG_MCP_MAX_TOP_Kint50Upper bound clamped on a caller-supplied top_k.
OPENRAG_MCP_SIMILARITY_THRESHOLDfloat0.8Minimum similarity score for a chunk to be returned by the search tool.
OPENRAG_MCP_DOWNLOAD_TIMEOUTfloat30.0Timeout (seconds) for the server-side index_url fetch (SSRF/DoS hardening).
OPENRAG_MCP_MAX_DOWNLOAD_BYTESint104857600Maximum bytes downloaded by an index_url fetch. Default is 100 MiB.

Model-endpoint seed overrides (legacy aliases)

Section titled “Model-endpoint seed overrides (legacy aliases)”

On first startup, OpenRAG seeds its model-endpoint catalog from the canonical variables documented above. The following legacy aliases are still read at seed time for backward compatibility and, when set, take precedence over their canonical counterpart, at that initial seeding and, with MODEL_ENDPOINT_SYNC_ON_BOOT=true, at every boot’s sync. Prefer the canonical variables in new deployments — do not set both.

Legacy aliasFalls back to (canonical)
LLM_ENDPOINTBASE_URL
LLM_MODELMODEL
VLM_ENDPOINTVLM_BASE_URL
EMBEDDER_ENDPOINTEMBEDDER_BASE_URL
EMBEDDING_MODELEMBEDDER_MODEL_NAME
RERANKER_ENDPOINTRERANKER_BASE_URL

(VLM_MODEL and RERANKER_MODEL are already the canonical names and are also used at seed time.)

Deployment-level knobs; most deployments never need to touch these — the compose stack drives the path variables through the storage volume variables instead.

VariableTypeDefaultDescription
OPENRAG_CONF_DIRstrbundled conf/Directory containing config.yaml. Override to run against a custom configuration tree.
DATA_DIRstr/app/data (container)Where uploaded files and app data are stored. In compose, relocate it via DATA_VOLUME rather than this variable.
DB_DIRstr/app/dbLocal database directory.
OPENRAG_CONTAINER_STARTUP_TIMEOUTfloatmax(60, 4 × POSTGRES_COMMAND_TIMEOUT) (= 120 with defaults)Seconds the API’s service container (DB pools, Ray actors, …) is allowed to initialize at startup before the app fails fast.
OPENRAG_BANNERbooltrueSet to false to suppress the ASCII startup banner. Its colors also auto-disable under the standard NO_COLOR / TERM=dumb conventions.
UVICORN_RELOADboolfalseDevelopment only — starts uvicorn with --reload (auto-restart on code changes). Also forces a single worker. Never enable in production.

Read only by the opt-in monitoring compose file (infra/compose/monitoring.docker-compose.yaml):

VariableTypeDefaultDescription
GRAFANA_ADMIN_USERstradminGrafana admin username.
GRAFANA_ADMIN_PASSWORDstr(required)Grafana admin password — compose refuses to start the monitoring profile if unset.
GF_SERVER_ROOT_URLstrhttp://localhost:3000Browser-facing Grafana root URL. Set this to the admin UI’s /grafana/ URL when using its proxy.
GF_SERVER_SERVE_FROM_SUB_PATHboolfalseSet to true when GF_SERVER_ROOT_URL includes the /grafana/ subpath.