Skip to content

🌟 API Documentation Overview

The FastAPI-powered backend provides a comprehensive document-based question answering system using Retrieval-Augmented Generation (RAG). The API supports semantic search, document indexing, and chat completions across multiple data partitions with full OpenAI compatibility.

Protected endpoints require authentication by default. Set AUTH_TOKEN in your .env and include it in the HTTP request header:

Authorization: Bearer YOUR_AUTH_TOKEN

For OpenAI-compatible endpoints, AUTH_TOKEN serves as the api_key parameter. Local no-auth development requires the explicit ALLOW_NO_AUTH=true opt-in; otherwise an empty token fails closed.


This API can be served using Uvicorn (default) or Ray Serve for distributed deployments.

By default, the backend uses uvicorn to serve the FastAPI app.

To enable Ray Serve, set the following environment variable:

.env
ENABLE_RAY_SERVE=true

Additional optional environment variables for configuring Ray Serve:

.env
RAY_SERVE_NUM_REPLICAS=1 # Number of deployment replicas
RAY_SERVE_HOST=0.0.0.0 # Host address for Ray Serve HTTP proxy
RAY_SERVE_PORT=8080 # Port for Ray Serve HTTP proxy

When using Ray Serve with a remote cluster, the HTTP server will be started on the head node of the cluster.

Verify server status and availability.

GET /health_check

Use readiness when deciding whether the API should receive traffic:

GET /ready

Readiness returns 503 when PostgreSQL, Milvus, or Ray is unavailable. It also reports the health of default and partition-referenced inference endpoints, plus broken preset references. Those model findings are informational because one tenant’s configuration must not remove a healthy shared API replica from service. checks.embedder shows whether the default embedder, which uploads and retrieval both need, is usable; with READINESS_REQUIRE_EMBEDDER=true readiness also returns 503 while it is unavailable or unresolvable.

Get openRAG version

GET /version

Get the current application configuration. Sensitive fields (api_key, password, token) are redacted.

GET /config

Permissions: Requires admin role

Response: JSON object with all configuration sections (LLM, embedder, vector DB, chunker, retriever, etc.)


POST /indexer/partition/{partition}/file/{file_id}

Upload a new file to a specific partition for indexing.

Parameters:

  • partition (path): Target partition name
  • file_id (path): Unique identifier for the file

Request Body (form-data):

  • file (binary): File to upload
  • metadata (JSON string): File metadata (e.g., {"owner": "user1"})
  • workspace_ids (JSON array, optional): Workspaces to add the file to
  • callback_url (string, optional): URL notified once indexing reaches a terminal state
  • callback_token (string, optional): Bearer token for that notification

Responses:

  • 201 Created: Returns task status URL
  • 400 Bad Request: callback_url is malformed or not a public http(s) URL
  • 409 Conflict: File already exists in partition. Two more cases name a code, as the [CODE] prefix of detail ("[CODE]: message"):
    • DOCUMENT_INDEXING_IN_PROGRESS: another task is still indexing this file_id. extra.existing_task_id and extra.task_status_url point at it, so a client that re-sent after a timeout can poll the first task instead of re-uploading. Earlier releases answered such a retry with 409 DOCUMENT_CONTENT_EXISTS naming the retried file_id itself when content deduplication is on (the default) and the bytes are the same; with deduplication off, or different bytes, they accepted it with 201 and let it fail at the end of the pipeline.
    • DOCUMENT_CONTENT_EXISTS: content deduplication matched another file (extra.existing_file_id).

Indexing is asynchronous. Pass a callback_url and OpenRAG POSTs the outcome once the task settles, so a client can wait instead of polling:

{
"partition": "alice.example.org",
"file_id": "file-123",
"status": "success",
"metadata": {"doc_rev": "<revision>", "datetime": "...", "doctype": "..."}
}

status is "success" or "error"; a user-cancelled task sends nothing. metadata is whatever you sent at upload, echoed back unchanged — including a revision-tracking field under any name you like, if that’s how your receiver orders these callbacks (the example above uses doc_rev). A field you didn’t send is simply absent, not sent back as null. Only server-computed keys are held back: the on-disk path, content hash, file size, filenames, and file_id (already a top-level field).

For an authenticated target, pass callback_token — sent as Authorization: Bearer <token>, never in the URL or the payload. Picking a safe scheme for it is on the caller: callback_url is otherwise unchecked beyond the public http(s) requirement below.

Best-effort: one attempt, no retries, a 5 s deadline, and any failure is logged without affecting the indexing result. Delivery is not guaranteed, so keep a client-side timeout and fall back to polling the task-status URL. The callback_url is checked against the SSRF guard and must be a public http(s) address — see INDEXING_CALLBACK_ALLOW_PRIVATE_URLS to target a local instance in development.

OpenRAG supports temporal filtering to retrieve documents from specific time periods. The client can include the temporal field to allow temporal-aware search in search endpoints.

  • created_at: ISO 8601 format date of when the file was created

created_at is provided by the client in the metadata of the file during upload. This is a first iteration — additional temporal fields (e.g. updated_at) may be added in future releases as needed.

Upload files while modeling relations between them
Section titled “Upload files while modeling relations between them”

OpenRAG supports document relationships to enable context-aware retrieval. You can model relationships between files using the metadata field during upload. Different relationship types can be represented using the relationship_id and parent_id metadata fields, depending on the use case: folder-based relationships, email threads, etc. (see Document Relationships documentation for more details).

POST /indexer/partition/{partition}/file/{file_id}
Authorization: Bearer YOUR_AUTH_TOKEN
Content-Type: multipart/form-data
file: <binary data>
metadata: {
"relationship_id": "documents/projects/2024/q1",
...
}
  • for email threads, one can rely on both relationship_id (to group emails in the same thread) and parent_id (to model reply hierarchies within the thread). See the Document Relationships documentation for more details and examples.

Example: Original Email (Root)

POST /indexer/partition/emails/file/email_a_id
Authorization: Bearer YOUR_AUTH_TOKEN
Content-Type: multipart/form-data
file: <email binary data>
metadata: {
"relationship_id": "thread-123",
"parent_id": null,
...
}

Example: Reply Email (Child)

POST /indexer/partition/emails/file/email_b_id
Authorization: Bearer YOUR_AUTH_TOKEN
Content-Type: multipart/form-data
file: <email binary data>
metadata: {
"relationship_id": "thread-123",
"parent_id": "email_a_id",
...
}

For context-aware search, see search endpoints and relationship-based file fetching.

PUT /indexer/partition/{partition}/file/{file_id}

Replace an existing file in the partition. Deletes the current entry and creates a new indexing task.

Parameters: Same as POST endpoint Request Body: Same as POST endpoint Responses:

  • 202 Accepted: Returns task status URL
  • 404 Not Found: File not found in partition
  • 409 Conflict: DOCUMENT_CONTENT_EXISTS, when content deduplication finds the same bytes held by another file, or by a task still indexing this one (extra.existing_file_id). A replace is not refused for a task already indexing the file, so it never returns DOCUMENT_INDEXING_IN_PROGRESS.
PATCH /indexer/partition/{partition}/file/{file_id}

Update file metadata without reindexing the document.

Request Body (form-data):

  • metadata (JSON string): Updated metadata

Response: 200 OK on successful update

DELETE /indexer/partition/{partition}/file/{file_id}

Remove a file from the specified partition.

Responses:

  • 204 No Content: Successfully deleted
  • 404 Not Found: File not found in partition
GET /indexer/task/{task_id}

Monitor the progress of an asynchronous indexing task.

Response: Task status information


GET /indexer/task/{task_id}/error
  • Search Across Multiple Partitions
GET /search/

Perform semantic search across specified partitions.

Query Parameters:

ParameterTypeDefaultDescription
partitions (optional)array[“all”]Partitions to search. (optional)
textstringrequiredSearch query
top_k (optional)integer5Number of initial results (optional)
include_related (optional)booleanfalseInclude chunks from files with same relationship_id
include_ancestors (optional)booleanfalseInclude chunks from ancestor files (via parent_id chain)
related_limit (optional)integer20Max related/ancestor chunks to fetch per result (used when include_related or include_ancestors is true)
filter (optional)stringNoneMilvus filter expression string for additional filtering. Supports comparison (==, !=, >, <, >=, <=), range (IN, LIKE), and logical (AND, OR, NOT) operators.

Responses:

  • 200 OK: JSON list of document links (HATEOAS format)
  • 400 Bad Request: Invalid partitions parameter
  • Search Within Single Partition
GET /search/partition/{partition}

Search within a specific partition only.

Query Parameters:

ParameterTypeDefaultDescription
textstringrequiredSearch query
top_k (optional)integer5Number of initial results (optional)
include_related (optional)booleanfalseInclude chunks from files with same relationship_id
include_ancestors (optional)booleanfalseInclude chunks from ancestor files (via parent_id chain)
related_limit (optional)integer20Max related/ancestor chunks to fetch per result (used when include_related or include_ancestors is true)
filter (optional)stringNoneMilvus filter expression string for additional filtering. Supports comparison (==, !=, >, <, >=, <=), range (IN, LIKE), and logical (AND, OR, NOT) operators.

Response: Same as multi-partition search

  • Search Within Specific File
GET /search/partition/{partition}/file/{file_id}

Search within a particular file in a partition.

Query Parameters: Same as partition search, including filter. Response: Same as other search endpoints


Extracts are the individual chunks a document is split into during indexing. Each has a stable extract_id, surfaced as the link on search results and on the file/chunk-listing endpoints below.

GET /extract/{extract_id}

Retrieve a specific document extract (chunk) by its ID.

Parameters:

  • extract_id (path): The unique chunk identifier (from search or chunk-listing results)

Permissions: Requires access to the partition containing the chunk — regular users are limited to their assigned partitions; admins can read any chunk.

Response: 200 OK

{
"page_content": "The text content of the chunk…",
"metadata": {
"file_id": "doc-a-id",
"filename": "Document A.pdf",
"partition": "my_partition",
"page": 3,
"indexed_at": "2026-01-01T12:00:00Z"
}
}

Errors:

  • 403 Forbidden: You don’t have access to the chunk’s partition
  • 404 Not Found: No extract with that ID

Partitions are the multi-tenant document collections OpenRAG indexes into. All routes below are prefixed with /partition. Access is role-based (hierarchy owner > editor > viewer): viewer for reads, owner for partition deletion, config changes, and member management. Admins with SUPER_ADMIN_MODE=true bypass membership checks.

GET /partition/

List the partitions you can access — admins see all partitions, regular users see only their memberships. Each entry includes partition, document_count, and (for non-admins) your role.

POST /partition/{partition}

Create an empty partition; you automatically become its owner. Returns 201 Created, or 409 Conflict if the name is taken. Non-admins are capped by MAX_PARTITIONS_PER_USER (a 403 is returned when the cap is reached).

DELETE /partition/{partition}

Permanently delete a partition and all its files and chunks. Owner only. Returns 204 No Content. This cannot be undone.

GET /partition/{partition}

Viewer+. Query: limit (optional). Returns { "files": [ { "file_id", "filename", "link", … } ] }, where link points at the file-detail endpoint below.

GET /partition/{partition}/file/{file_id}

Viewer+. Query: limit (max chunks, default 2000). Returns { "metadata": {…}, "documents": [ { "link": "…/extract/{id}" } ] }. Returns 404 if the file isn’t in the partition.

GET /partition/{partition}/chunks

List document chunks (extracts) in a partition. Viewer+.

Query Parameters:

ParameterTypeDefaultDescription
include_embeddingbooleantrueInclude each chunk’s vector embedding
file_idstringNoneRestrict to a single file’s chunks (recommended for the document-detail view)
limitintegerunboundedMax chunks to return

Returns { "chunks": [ { "content", "metadata", "link", "embedding"? } ] }. Note the chunk text is under content here, whereas the single-chunk GET /extract/{extract_id} endpoint returns it under page_content.

Each partition references an indexation and a retrieval preset, plus an embedder and chat LLM.

  • Get resolved config
GET /partition/{partition}/config

Viewer+. Returns the partition’s preset references and the fully resolved indexation/retrieval pipeline configuration.

  • Update config
PATCH /partition/{partition}

Owner only. Body fields (all optional): description, embedder, indexation_preset, retrieval_preset, chat_history_depth, chat_llm. Returns the updated resolved config. Changing embedder is refused on a partition that holds indexed files (409 PARTITION_HAS_INDEXED_FILES, or 409 INDEXING_IN_PROGRESS while files are still being indexed or copied into it): each embedder keeps its vectors in its own field, which only its partitions search.

curl -X PATCH http://localhost:8080/partition/my_partition \
-H "Authorization: Bearer YOUR_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d '{"indexation_preset": "legal", "retrieval_preset": "hyde"}'

Manage who can access a partition and with which role — all owner only.

EndpointMethodDescription
/partition/{partition}/usersGETList members → { "members": [...] }
/partition/{partition}/users/candidatesGETSearch non-members by display-name prefix or exact user ID → paginated identities
/partition/{partition}/usersPOSTAdd a new member → 201; returns 409 if already present
/partition/{partition}/users/{user_id}PATCHUpdate a member’s role — form field role → 200
/partition/{partition}/users/{user_id}DELETERemove a member → 204

POST no longer updates an existing member’s role. Use the PATCH endpoint when changing a role.

  • Get Files by Relationship
GET /partition/{partition}/relationships/{relationship_id}

Returns all files sharing the same relationship_id within a partition. Viewer+.

Parameters:

  • partition — partition name
  • relationship_id — the relationship group identifier (client-defined)

Response:

{
"files": [
{
"file_id": "doc-a-id",
"filename": "Document A",
"relationship_id": "group-123",
"parent_id": null
},
{
"file_id": "doc-b-id",
"filename": "Document B",
"relationship_id": "group-123",
"parent_id": "doc-a-id"
}
]
}
  • Get File Ancestors
GET /partition/{partition}/file/{file_id}/ancestors

Returns the complete ancestor path from root to the specified file. Viewer+.

Parameters:

  • partition — partition name
  • file_id — the file to trace ancestors for
  • max_ancestor_depth (optional) — limit on ancestor depth to return. None means unlimited.

Response:

{
"ancestors": [
{
"file_id": "email-a-id",
"filename": "Original Email",
"parent_id": null
},
{
"file_id": "email-b-id",
"filename": "First Reply",
"parent_id": "email-a-id"
},
{
"file_id": "email-c-id",
"filename": "Second Reply",
"parent_id": "email-b-id"
}
]
}

Named, reusable indexation and retrieval pipeline configurations. Partitions reference a preset by name (see PATCH /partition/{partition}) instead of carrying an inline config, so a change to a preset propagates to every partition using it. Six defaults are seeded on first boot (default, legal, finance for indexation; default, multiquery, hyde for retrieval); the default preset of each type cannot be deleted or renamed.

All routes are prefixed with /presets and require the admin role. preset_type is one of indexation | retrieval.

GET /presets/options

Returns the choices valid inside a preset config: chunking_strategies, parsing_strategies (pymupdf, marker, docling), retrieval_types, and reranker_providers, plus default_top_n: what a retrieval preset’s top_n resolves to when it is left out (the deployment’s RERANKER_TOP_K).

POST /presets/

Body: name (string), preset_type (indexation | retrieval), config (object — its keys depend on the type). Returns 201 Created with the stored preset.

curl -X POST http://localhost:8080/presets/ \
-H "Authorization: Bearer YOUR_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "hyde-large",
"preset_type": "retrieval",
"config": {"type": "hyde", "top_k": 50, "top_n": 10}
}'

An indexation config instead accepts keys such as chunking ({name, chunk_size, chunk_overlap_rate}), parsing_strategy, stt (a named STT endpoint), asr_transcription_prompt_name, enable_image_captioning, enable_contextualization, and contextualization_mode. Omitting stt or asr_transcription_prompt_name uses the corresponding global default. The STT and ASR prompt selections apply only to audio or video files routed to OpenAIAudioLoader; LocalWhisperLoader does not use them.

GET /presets/

Query: preset_type (optional) to filter by type. Returns a list of presets with name, preset_type, config, created_at, updated_at.

GET /presets/{preset_type}/{name}
PUT /presets/{preset_type}/{name}
DELETE /presets/{preset_type}/{name}

PUT accepts a partial body (name to rename and/or config; at least one required) and returns the updated preset. DELETE returns 204 No Content.


Every prompt the pipeline sends to a model is a stored, editable row rather than a bundled file. On first boot each type is seeded from its bundled template as that type’s default; an admin can add named variants and select one per preset or per partition.

All routes are prefixed with /prompts and require the admin role. prompt_type is one of sys_prompt | spoken_style_answer | query_contextualizer | chunk_contextualizer | image_captioning | hyde | multi_query | topic_tagger | asr_transcription.

Resolution order for a given type: the name selected for the request → the type’s global default → the bundled template. A selection naming a prompt that no longer exists falls back to the default rather than failing.

Where a prompt is selected — each setting lives with the thing it configures:

Prompt typeSelected on
sys_prompt, spoken_style_answerpartition — generation_prompt_names
query_contextualizer, hyde, multi_queryretrieval preset — *_prompt_name
chunk_contextualizer, image_captioning, topic_taggerindexation preset — *_prompt_name
asr_transcriptionindexation preset — asr_transcription_prompt_name for audio or video files routed to OpenAIAudioLoader
POST /prompts/

Body: prompt_type, name, content, is_default (default false). Returns 201 Created, or 409 if that (prompt_type, name) already exists.

Types rendered as templates (sys_prompt, spoken_style_answer, query_contextualizer, hyde, multi_query) accept only their own plain {placeholders} — no conversion (!r), format spec (:>10) or attribute access — and a violation returns 422 at write time rather than failing later at render. Escape a literal brace as {{ / }}. The remaining types are sent to the model verbatim. asr_transcription may be empty, which keeps the transcription model’s native prompt.

curl -X POST http://localhost:8080/prompts/ \
-H "Authorization: Bearer YOUR_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"prompt_type": "sys_prompt",
"name": "legal-assistant",
"content": "Answer strictly from the context.\n{context}\nToday is {current_date}."
}'
GET /prompts/ # ?prompt_type= to filter, ?offset= &limit= to page (limit ≤ 500)
GET /prompts/{prompt_id}
PATCH /prompts/{prompt_id} # any of name, content, is_default
DELETE /prompts/{prompt_id} # 204 No Content; refused for a type's current default

List entries carry used_by — the number of partitions that resolve to that prompt, counting those that fall back to it as the default.

PUT /prompts/{prompt_id}/default

Clears the previous default for that type and promotes this one, atomically.

PATCH /partition/{partition}

Send generation_prompt_names, a map of {prompt_type: name} restricted to sys_prompt and spoken_style_answer. Each name must exist, or the request returns 422. Send {} to clear the selection and fall back to the defaults.

curl -X PATCH http://localhost:8080/partition/my-partition \
-H "Authorization: Bearer YOUR_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d '{"generation_prompt_names": {"sys_prompt": "legal-assistant"}}'

Preset-scoped prompts are selected the same way, by putting the *_prompt_name field in the preset’s config (see Pipeline Presets). For audio or video files routed to OpenAIAudioLoader, the indexation preset also selects the ASR prompt and STT endpoint. LocalWhisperLoader uses neither setting.


A registry of named inference endpoints (embedder, reranker, LLM, VLM, STT) that partitions and presets can point at, so operators can manage and switch inference backends at runtime instead of via .env. Stored API keys are redacted in every response and only returned through the explicit reveal action below.

All routes are prefixed with /model-endpoints and require the admin role. model_type is one of embedder | reranker | llm | vlm | stt.

POST /model-endpoints/

Body: name, model_type, endpoint (URL), model_name (optional), batch_size (default 32), timeout (seconds, default 30), extra (object — put api_key here), is_default (default false). Returns 201 Created; the response carries has_api_key rather than the key itself.

curl -X POST http://localhost:8080/model-endpoints/ \
-H "Authorization: Bearer YOUR_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "prod-embedder",
"model_type": "embedder",
"endpoint": "https://vllm.internal/v1",
"model_name": "jinaai/jina-embeddings-v3",
"extra": {"api_key": "sk-…"}
}'
GET /model-endpoints/ # ?model_type= to filter
GET /model-endpoints/{model_type}/{name}
PUT /model-endpoints/{model_type}/{name}
DELETE /model-endpoints/{model_type}/{name} # 204 No Content

PUT takes a partial body (any of the create fields, plus name to rename) and returns the updated endpoint.

An LLM endpoint’s token budgets live in extra.max_llm_context_size and extra.max_output_tokens. Every route returning an LLM endpoint also reports what a missing value falls back to: detected_max_llm_context_size (the max_model_len the endpoint reported on /v1/models, null when it reports none, as most gateways and hosted APIs don’t), default_max_llm_context_size and default_max_output_tokens (the deployment’s MAX_LLM_CONTEXT_SIZE and MAX_OUTPUT_TOKENS). The context size that applies is max_llm_context_size, else detected_max_llm_context_size, else default_max_llm_context_size. A write to an LLM endpoint re-probes /v1/models after responding; until that lands, context_size_detection_pending is true and detected_max_llm_context_size is null. These fields are null on other endpoint types.

Changing an embedder’s endpoint, model_name, extra.implementation or extra.max_model_len changes the vectors it produces. While partitions hold files indexed with it, such an edit returns 409 EMBEDDER_EDIT_AFFECTS_INDEXED_DATA naming the partitions and file counts; resend it with "acknowledge_indexed_data": true to apply it anyway. The existing vectors are not rebuilt, so queries on those partitions are then compared against vectors from the previous configuration. GET /model-endpoints/{model_type}/{name}/indexed-usage returns the same counts ahead of time.

Files still indexing when such an edit is saved are not counted, since a file is only counted once indexing finishes. Instead, when the file is recorded, the indexer checks that the partition’s embedder still has the configuration it embedded with. If it doesn’t, the file fails with EMBEDDER_CHANGED_DURING_INDEXING, its vectors are removed, and it has to be indexed again. Indexers reload an edited endpoint before embedding, so only files already embedding during the edit are affected.

DELETE of an embedder returns 409 while a partition names it or, for the default embedder, while a partition following the default alias already holds indexed files. Reassign those partitions first. Deleting an embedder also drops its vector field from the collection, on a best-effort basis: if the drop fails, the delete still succeeds and the empty field stays until removed by hand.

POST /model-endpoints/{model_type}/{name}/set-default

Promotes the endpoint to the default used for its model_type, and returns it.

For embedders, a new default only reaches partitions that have no data yet. A partition stays on the default alias until it first receives a file (upload or copy), at which point the alias is replaced by the embedder it resolved to; a partition indexed before that is pinned the same way when the default changes. Partitions with files therefore keep the embedder that built their vectors. An emptied partition stays pinned — set its embedder back to default to follow the default again.

POST /model-endpoints/{model_type}/{name}/reveal-api-key

Explicit admin action that returns { "api_key": "…" } (or null if none is stored). This is the only endpoint that returns the key in clear text; the action is logged.

POST /model-endpoints/validate # probe a draft (unsaved) endpoint
POST /model-endpoints/{model_type}/{name}/validate # probe a registered endpoint

Both probe the target for reachability and model availability, returning { "reachable", "model_found", "models_served", "detail" }. The draft form takes endpoint (+ optional model_name, api_key); to reuse a saved key without resending it, pass stored_api_key_model_type and stored_api_key_name (both required together, and only accepted when the draft endpoint matches the saved one).


These endpoints provide full OpenAI API compatibility for seamless integration with existing tools and workflows. For detailed example of openai usage see this section

  • List Available Models
GET /v1/models

List all available RAG models (partitions).

Model Naming Convention:

  • Pattern: openrag-{partition_name} => This model allows to chat specifically with the partition {partition_name}
  • Special model: partition-all (queries entire vector database)
  • Chat Completions
POST /v1/chat/completions

OpenAI-compatible chat completion using RAG pipeline.

Request Body:

curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_AUTH_TOKEN" \
-d '{
"model": "openrag-{partition_name}",
"messages": [
{
"role": "user",
"content": "Your question here"
}
],
"temperature": 0.1,
"stream": false
}'

You can also direclty use this endpoint with no RAG pipeline, i.e. to directly use the LLM. For that, instead of using the openrag prefix for the model, you can:

  • Specify no model
  • Specify an empty model
  • Specify the openRAG configured model, e.g. Mistral-Small-3.1-24B-Instruct-2503.

Request Body:

curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_AUTH_TOKEN" \
-d '{
"model": "",
"messages": [
{
"role": "user",
"content": "Your question here"
}
],
"temperature": 0.1,
"stream": false
}'

Besides the OpenAI fields, the response carries extra, a JSON object with the sources of the answer. Up to 2.2.2 it was a JSON-encoded string that clients had to parse. In a streamed response every chunk carries extra, empty until the final chunks, which carry the whole object; only an error chunk carries none.

KeyContent
sourcesThe sources the model cited, or every presented source when it reported no citations: the answer has neither a [Sources: …] tag nor an inline [Source N] marker, or the request skipped citation reporting (structured output, direct LLM or web search without results). An explicit [Sources: none] gives an empty list. Kept for existing clients.
presented_sourcesEvery source shown to the model, whether cited or not.
cited_sourcesOnly the sources the model cited, in a [Sources: …] tag or inline [Source N] markers; empty when it cited none or reported no citations.
citations_reportedtrue when the model reported its citations, including [Sources: none]; false when sources fell back to every presented source. An empty [Sources: ] counts as no tag.
all_retrieved_sourcesEvery source retrieval returned, including those that did not fit in the prompt. Only when the request sets metadata.include_all_retrieved_sources: true.
attachmentsChat completions only: the attached file IDs actually used, when the request sets metadata.attachments.
truncatedStreamed chat completions only: true when the upstream stream ended without [DONE], so the answer may be cut short. Absent otherwise.

A document source nests the chunk’s metadata under chunk (since 2.3.0); the keys beside it are computed by the server:

{
"source_type": "document",
"chunk": { "filename": "report.pdf", "file_id": "…", "partition": "…", "…": "…" },
"rerank_score": 0.646,
"chunk_url": "https://<host>/extract/<chunk id>",
"file_url": "https://<host>/static/<chunk id>"
}

rerank_score is the score the reranker gave that chunk for this query, present only when a reranker ran; compare it within one response only. file_url is absent when the chunk has no source file. A web source is flat:

{
"source_type": "web",
"url": "https://example.org/page",
"title": "Page title",
"snippet": "The search result's excerpt of the page."
}
  • Text Completions
POST /v1/completions

OpenAI-compatible text completion endpoint.

  • When using /v1/chat/completions or /v1/completions, you can pass extra arguments via the metadata field of the request body to customize the RAG behavior. Some options apply only to chat, as shown below:

Partition-backed chat retrieves by default. It skips retrieval only for an empty or exclusively casual message; ambiguous, mixed, and factual messages retrieve. Text completions retain their contextualizer-controlled default.

OptionApplies toTypeDefaultDescription
require_retrievalChat and text completionsboolfalseFor partition-backed chat, JSON true forces retrieval even for a casual or normalized-empty message. Other chat messages already retrieve by default. For text completions, JSON true remains the opt-in override when the contextualizer would skip. The original input is the fallback query. Existing scope and filters still apply; matching sources and answer correctness are not guaranteed. Has no effect in direct LLM mode.
websearchChat onlyboolfalseAugments the RAG context with live web search results. When used with a partition (openrag-{partition}), document and web results are combined. When used without a partition (direct LLM mode), web results are the sole context. Requires WEBSEARCH_API_TOKEN to be configured. See web search configuration.
spoken_style_answerChat and text completionsboolfalseGenerates a succinct spoken-style conversational answer based on the retrieved documents.
use_map_reduceChat onlyboolfalseUses a map-reduce strategy to aggregate information from multiple documents. See map-reduce configuration.
llm_overrideChat and text completionsobjectnullOverrides the downstream LLM for this request. Accepts model (string), always honored. Also accepts base_url and api_key, which are honored only when the deployment sets LLM_OVERRIDE_ALLOW_CUSTOM_ENDPOINT — otherwise they are ignored and the request goes to the server’s configured endpoint. See custom LLM endpoints.
attachmentsChat onlylist[{"id": string}]nullScopes RAG retrieval to a specific list of file IDs within the target partition, instead of searching the whole partition. Combined with workspace, only the attached files that belong to the workspace are searched. Unknown, unindexed or foreign IDs are silently dropped; duplicates are deduplicated. The response’s extra.attachments reports which IDs were actually used.

Only allowlisted English and French casual messages skip retrieval; other phrasing retrieves by default. Casual replies use a dedicated instruction in the detected language. Only greetings include an introduction; set ASSISTANT_NAME to use your deployment’s name. Ambiguous or normalized-empty input defaults to English.

Examples:

Enabling conversational answer with openai chat completions endpoint
curl -X 'POST' 'http://localhost:8080/v1/chat/completions' \
-H 'accept: application/json' \
-H 'Authorization: Bearer YOUR_AUTH_TOKEN' \
-H 'Content-Type: application/json' \
-d '{
"model": "openrag-{partition_name}",
"messages": [
{
"role": "user",
"content": "your_query"
}
],
"temperature": 0.3,
"stream": false,
"metadata": {
"spoken_style_answer": true
}
}'
Enabling web search with RAG documents
curl -X 'POST' 'http://localhost:8080/v1/chat/completions' \
-H 'accept: application/json' \
-H 'Authorization: Bearer YOUR_AUTH_TOKEN' \
-H 'Content-Type: application/json' \
-d '{
"model": "openrag-{partition_name}",
"messages": [
{
"role": "user",
"content": "your_query"
}
],
"stream": false,
"metadata": {
"websearch": true
}
}'
Web search only (no RAG partition)
curl -X 'POST' 'http://localhost:8080/v1/chat/completions' \
-H 'accept: application/json' \
-H 'Authorization: Bearer YOUR_AUTH_TOKEN' \
-H 'Content-Type: application/json' \
-d '{
"model": "",
"messages": [
{
"role": "user",
"content": "your_query"
}
],
"stream": false,
"metadata": {
"websearch": true
}
}'
Using another configured LLM model with OpenRAG's RAG pipeline
curl -X 'POST' 'http://localhost:8080/v1/chat/completions' \
-H 'accept: application/json' \
-H 'Authorization: Bearer YOUR_AUTH_TOKEN' \
-H 'Content-Type: application/json' \
-d '{
"model": "openrag-{partition_name}",
"messages": [
{
"role": "user",
"content": "your_query"
}
],
"stream": false,
"metadata": {
"llm_override": {
"model": "gpt-4o"
}
}
}'

With LLM_OVERRIDE_ALLOW_CUSTOM_ENDPOINT=true, the same object may also carry the endpoint and its credential, sending the request to a provider of the client’s choosing instead of the configured one:

metadata.llm_override with a client-supplied endpoint
{
"llm_override": {
"base_url": "https://api.openai.com/v1",
"api_key": "sk-...",
"model": "gpt-4o"
}
}

base_url must be https, must carry no query string, fragment or .. path segment (percent-encoded or not), and is always requested as {base_url}/chat/completions — {base_url}/completions when the request comes in on the legacy /v1/completions route; anything else is rejected with a 400. The server’s own API key is never forwarded — an override without api_key sends no Authorization header at all. When the flag is off, base_url/api_key are ignored (a warning is logged) and only model applies, which typically surfaces as an “unknown model” error from the configured provider.

Scoping retrieval to specific attachments
curl -X 'POST' 'http://localhost:8080/v1/chat/completions' \
-H 'accept: application/json' \
-H 'Authorization: Bearer YOUR_AUTH_TOKEN' \
-H 'Content-Type: application/json' \
-d '{
"model": "openrag-{partition_name}",
"messages": [
{
"role": "user",
"content": "summarize these files"
}
],
"stream": false,
"metadata": {
"attachments": [
{"id": "file-001"},
{"id": "file-002"},
{"id": "file-003"}
]
}
}'

Tools are useful features that can be called directly by the client.

  • List available tools
GET /v1/tools

Request:

curl http://localhost:8080/v1/tools

Response:

[
{
"name": "Tool name",
"description": "Tool description"
}
]
  • Execute a tool
POST /v1/tools/execute

The parameters are given in multipart.

Request:

curl -X POST http://localhost:8080/v1/tools/execute \
-H "Content-Type: multipart/form-data" \
-H "Authorization: Bearer YOUR_AUTH_TOKEN" \
-F "file=@file.pdf" \
-F 'tool={"name":"extractText"}' \
-F 'metadata={"mime":"application/pdf","name":"test.pdf"}' \

Response:

{
"message": "File content"
}

For indexing multiple files programmatically, you can use the data_indexer.py utility script in the scripts/ folder or simply use indexer ui.

from openai import OpenAI, AsyncOpenAI
api_base_url = "http://localhost:8080" # fastapi base url of 'openrag'
base_url = f"{api_base_url}/v1"
auth_key = ... # your API authentication key AUTH_TOKEN from .env
client = OpenAI(api_key=auth_key, base_url=base_url)
your_partition= 'my_partition' # name of your partition
model = f"openrag-{your_partition}"
settings = {
'model': model,
'temperature': 0.3,
'stream': False
}
response = client.chat.completions.create(
**settings,
messages=[
{"role": "user", "content": "What information do you have about...?"}
]
)

The API uses standard HTTP status codes:

  • 200 OK: Successful request
  • 201 Created: Resource created successfully
  • 202 Accepted: Request accepted for processing
  • 204 No Content: Successful deletion
  • 400 Bad Request: Invalid request parameters
  • 404 Not Found: Resource not found
  • 409 Conflict: Resource already exists

Error responses include detailed JSON messages to help with debugging and integration.

An external indexer uses an admin token to create a user, then uses the returned user token for all subsequent operations.

sequenceDiagram
    participant Indexer
    participant OpenRAG

    Note over Indexer, OpenRAG: 1. Create User (admin token)

    Indexer->>OpenRAG: POST /users<br/>Authorization: Bearer {admin_token}<br/>{display_name: "alice"}
    OpenRAG-->>Indexer: 201 {id: 2, token: "or-xxx..."}
    Indexer->>Indexer: Store user token "or-xxx..."

    Note over Indexer, OpenRAG: 2. Create Partition (user token). Automatically grants owner rights to the user

    Indexer->>OpenRAG: POST /partition/{partition}<br/>Authorization: Bearer or-xxx...
    OpenRAG-->>Indexer: 201 Created

    Note over Indexer, OpenRAG: 3. Index a File (user token)

    Indexer->>Indexer: New data

    Indexer->>OpenRAG: POST /indexer/partition/{partition}/file/{file_id}<br/>Authorization: Bearer or-xxx...<br/>Body: multipart (file + metadata)
    OpenRAG->>OpenRAG: Validate token, check file quota
    OpenRAG-->>Indexer: 201 {task_status_url: "{task_url}"}

    Note over OpenRAG: Background: serialize → chunk → embed → store

Query indexed documents via the OpenAI-compatible chat completions endpoint.

sequenceDiagram
    participant Client
    participant OpenRAG
    participant LLM

    Client->>OpenRAG: POST /v1/chat/completions<br/>Authorization: Bearer or-xxx...<br/>{model: "openrag-{partition}",<br/>messages: [...], stream: true}

    OpenRAG->>OpenRAG: Authenticate user
    OpenRAG->>OpenRAG: Resolve partition from model name
    OpenRAG->>OpenRAG: Check user has access to partition
    OpenRAG->>OpenRAG: Retrieve relevant documents
    OpenRAG->>OpenRAG: Build prompt with retrieved context
    OpenRAG->>LLM: Query LLM

    loop SSE Stream
        OpenRAG-->>Client: data: {delta: {content: "..."},<br/>sources: [...]}
    end
    OpenRAG-->>Client: data: [DONE]