Skip to content

Grafana service dashboard

The OpenRAG Service dashboard (UID openrag-service) answers one question per row: is OpenRAG indexing, answering and serving what it should? The HTTP dashboard says whether requests succeed. This one covers what happens behind them: the indexing pipeline, the inference endpoints, and retrieval’s catalog drift.

It ships with the other dashboards in infra/charts/openrag-stack/dashboards/. The Compose monitoring overlay provisions it into the OpenRAG folder, and the Helm chart delivers it as a sidecar ConfigMap (see Kubernetes monitoring).

OpenRAG exports metrics from two places, and a panel stays empty when its target is not scraped:

TargetCarries
The API’s GET /metricsqueue depth, model endpoint readiness, catalog drift, and the API process’s inference calls and circuit breakers
Ray’s metrics agentdocument outcomes, stage durations, queue wait, the parse watchdog, and the indexing workers’ inference calls and circuit breakers

Neither target is limited to one kind of inference call. The API embeds each query and calls the LLM and reranker to answer it. The indexing workers embed, caption images and call the LLM to contextualize chunks. Under ENABLE_RAY_SERVE=true the API is itself a Ray actor, and its calls move to the Ray target too.

How each deployment collects them:

DeploymentAPI /metricsRay’s metrics agent
Compose, monitoring overlayjob openragjob openrag-ray, on the port the overlay pins with RAY_METRICS_EXPORT_PORT
Kubernetes, ray.enabled=trueopenrag.metrics.serviceMonitorthe Ray PodMonitor on the Ray pods’ metrics port, once ray.metrics.podMonitor.enabled=true or monitoring.bundled (off by default; see Monitoring Ray, Postgres and Milvus)
Kubernetes, ray.enabled=falseopenrag.metrics.serviceMonitorthe same PodMonitor, on the openrag pod’s ray-metrics port, which the chart pins with RAY_METRICS_EXPORT_PORT. Not with an external Ray cluster (RAY_ADDRESS, or ray.externalCluster): the API then runs no Ray, and that cluster is scraped where it runs

OpenRagTargetDown watches every job whose name contains openrag, which is why the Compose Ray job is named openrag-ray.

Ray prefixes everything it exports with ray_. Both scrape paths above strip it from OpenRAG’s series, so they are stored under the names the API exports and the alert rules query. Panels still select {__name__=~"(ray_)?openrag_…"}, so a Prometheus that scrapes Ray without that rename fills them too. Either way, a metric exported from both targets, such as openrag_inference_requests_total, is summed across the two.

See the metrics reference for every series, its labels, and the target that exports it.

Eight tiles, each mirroring an alert:

TileTurns red whenAlert
API scrapePrometheus cannot scrape /metricsOpenRagTargetDown
Ray scrapePrometheus cannot scrape Ray’s metrics agentOpenRagTargetDown
Queued tasksyellow at 50 queuedOpenRagBacklogGrowing
Since last parseneutral: only a problem while tasks are queuedOpenRagIngestStalled
Failed documents · 5m25% of the documents finished in the last 5 minutes failedOpenRagIngestFailureRate
Inference errors · 10mthe worst endpoint fails half its callsOpenRagInferenceProviderDown
Open breakersany circuit breaker is open or half-openOpenRagCircuitBreakerOpen
Catalog drift · 1hany retrieval hit dropped for a missing fileOpenRagCatalogDriftDetected

The colours use the alerts’ default thresholds, and the two ratio tiles apply the alerts’ volume floors. Failed documents needs 5 documents finished in 15 minutes, and Inference errors leaves out endpoints with fewer than 5 calls in 10 minutes that succeeded, failed or timed out; like the alert, it ignores cancelled, rejected and breaker-refused calls. Below the floor they read Low volume, because the alert does not judge a ratio that small. If you tune an alert, the tile does not follow it.

Each tile shows the current value. When a series disappears, its tile empties and shows its no-value text rather than the last value it had. An empty tile is not a healthy one, so that text is neutral, never green:

  • Unknown on Queued tasks means the API could not read the task state manager. An idle queue shows 0.
  • Not scraped on Ray scrape means no OpenRAG series from Ray reached Prometheus in the selected range.
  • Not scraped on API scrape is red: without the API’s scrape, most of the dashboard is blind.

While the Ray target is down, the tiles it feeds go quiet: Since last parse reads No parse yet, and Failed documents reads Idle once its window has passed. Ray scrape is the tile that says why.

Documents reaching a terminal state per minute, stacked by outcome; the failure ratio over the 5-minute window the failure-rate alert reads; and totals over the selected range. Cancelled uploads are counted but never treated as failures.

Tasks queued and in progress, queue wait (admission to the start of processing, p50 and p95), and the time since each parser pool last finished a document. A pool nobody uploads to (audio, say) climbs forever; that is expected. Clock-skew events counts queue-wait measurements that came out negative because the API and worker clocks disagree. When it is non-zero, the queue-wait panel reads low.

The p95 duration of each stage, on a logarithmic axis because chunking takes milliseconds and a large PDF parse takes minutes. Next to it, the average number of documents in each stage: seconds of stage work per second of wall clock. The tallest band is where indexing time goes. A stage’s time is recorded when the stage finishes, so a long stage reads 0 while it runs, then its whole duration lands in one spike. Read that panel over a range much longer than your slowest stage.

Calls by operation, summed across both targets, the error ratio per registry endpoint, and failed calls by outcome. Below them: p95 latency per endpoint and operation, circuit-breaker state (the worst across processes: open, then half-open, then closed), token throughput, and the readiness probe of every model endpoint in use. client_override groups requests that named their own endpoint through metadata.llm_override, so their failures never count against an endpoint you run.

Catalog drift over the trailing hour: retrieval hits dropped because their file is gone from the catalog. A hit is dropped per query, so this is not a document count. Only non-zero matters; see catalog reconciliation.

  • No per-partition, per-user or per-file breakdown. No OpenRAG metric carries those labels, and a unit test fails the build if a dashboard query references one. Which partition is failing is a log question, answered by the partition field of the structured logs.
  • Queue depth takes max, not sum. Every API replica attached to one Ray cluster reads the same task state manager, so summing would multiply the queue by the replica count. The same applies to model endpoint readiness (min).
  • Rates come before sums. Each Ray worker is its own series and restarts reset it; summing first would read every restart as a drop.
  • Ray is scraped only with its PodMonitor on. ray.metrics.podMonitor.enabled, or monitoring.bundled, in either Ray topology. Without it the ingestion rows and the inference calls made while indexing stay empty, and Ray scrape reads Not scraped.
  • Everything Prometheus scrapes is aggregated. Two OpenRAG releases scraped by one Prometheus show as one.
  • Ray Serve. Under ENABLE_RAY_SERVE=true the chart does not scrape the API’s /metrics (see Prometheus metrics), so every panel it feeds stays empty: API scrape, queue depth, readiness and catalog drift.