Use case

Managed RAG Platform: Production Retrieval-Augmented Generation APIs

Deploy retrieval-augmented generation (RAG) pipelines as authenticated, metered APIs on isolated Cloud Run. Bring your own embeddings, vector database, and LLM, or use Enqualia's defaults, and ship a grounded chat-with-your-data experience without writing infrastructure code.

By Antoni Elkenbracht·Updated January 15, 2026

How it works

  1. 1

    Connect your sources

    Point Enqualia at object storage, a database, or a webhook. The ingestion pipeline chunks, embeds, and indexes content with per-tenant scoping.

  2. 2

    Configure retrieval

    Pick an embedding model, set top-k, choose hybrid retrieval (vector + BM25), enable cross-encoder reranking, and optionally apply ontology-aware filters.

  3. 3

    Deploy the API

    Enqualia provisions an isolated Cloud Run service with dev and prod environments. Your clients hit a versioned endpoint; we meter every call and surface costs in real time.

What you get on Enqualia

Vector indexingPluggable: pgvector, Pinecone, Weaviate, Qdrant, or Enqualia-managed.
Embedding modelsGemini Embedding 001, OpenAI text-embedding-3, Voyage AI, Cohere. Bring your own key.
Retrieval modesPure vector, hybrid (vector + BM25), ontology-aware (typed entities).
Latency targetp95 < 250ms end-to-end on cached embeddings; < 1.2s cold.
BillingPer-call markup multiplier on underlying token cost, clamped by per-API min and max.
IsolationPer-project Cloud Run service. Dev and prod environments separated.
RBACOwner / admin / developer / web / viewer roles. Per-API key scoping.
Observabilityp50/p95/p99 latency, cost per request, error rate, per-API breakdown in BigQuery.

Quick start

Python · Admin SDK
from enqualia_admin_sdk import SDK

sdk = SDK(api_key="sk-...", project_id="proj_...")

# Deploy a RAG API in one call
sdk.upload_api_code(
    api_id="docs-rag-v1",
    path="./rag_runner.py",
    version_description="Hybrid retrieval + reranker",
)
sdk.deploy_to_dev()
sdk.promote_to_prod()

# Clients call it like any other API
import httpx
r = httpx.post(
    "https://docs-rag-v1.enqualia.io/run",
    headers={"X-API-KEY": "sk-client-..."},
    json={"query": "what's our refund policy?", "top_k": 5},
)
print(r.json()["answer"], r.json()["sources"])

Ready to deploy?

Spin up a project and deploy this workload to an isolated Cloud Run in minutes.

Frequently asked questions

Can I bring my own vector database?
Yes. Enqualia's RAG runner connects to pgvector, Pinecone, Weaviate, Qdrant, or any vector store reachable from your Cloud Run instance. You configure the connection in the API's environment variables; we don't lock the index inside our infrastructure.
How do you handle large knowledge bases?
Ingestion is incremental. Only changed documents are re-embedded. For multi-tenant deployments, we recommend metadata filtering at retrieval time and per-tenant collection scoping. Largest production workloads on Enqualia today index single-digit billions of chunks.
Do you support ontology-aware RAG / GraphRAG?
Yes. The retrieval layer accepts a typed query (entity types, relations, hop count) in addition to free text. Behind the scenes it queries a knowledge graph and falls back to vector search when the typed query underspecifies. Useful for regulated domains (healthcare, finance, legal) where audit-grade provenance is required.
What does pricing look like in practice?
A typical retrieval-only call costs ~$0.0001 (embedding) + ~$0.0005 (vector lookup) + a small markup. A retrieval + Gemini Flash answer at 1k input / 200 output tokens is ~$0.0003 underlying × multiplier. We clamp per-call cost to a configurable max so a runaway prompt can't blow the budget.

Learn the concepts