Managed RAG Platform: Production Retrieval-Augmented Generation APIs
Deploy retrieval-augmented generation (RAG) pipelines as authenticated, metered APIs on isolated Cloud Run. Bring your own embeddings, vector database, and LLM, or use Enqualia's defaults, and ship a grounded chat-with-your-data experience without writing infrastructure code.
How it works
- 1
Connect your sources
Point Enqualia at object storage, a database, or a webhook. The ingestion pipeline chunks, embeds, and indexes content with per-tenant scoping.
- 2
Configure retrieval
Pick an embedding model, set top-k, choose hybrid retrieval (vector + BM25), enable cross-encoder reranking, and optionally apply ontology-aware filters.
- 3
Deploy the API
Enqualia provisions an isolated Cloud Run service with dev and prod environments. Your clients hit a versioned endpoint; we meter every call and surface costs in real time.
What you get on Enqualia
| Vector indexing | Pluggable: pgvector, Pinecone, Weaviate, Qdrant, or Enqualia-managed. |
|---|---|
| Embedding models | Gemini Embedding 001, OpenAI text-embedding-3, Voyage AI, Cohere. Bring your own key. |
| Retrieval modes | Pure vector, hybrid (vector + BM25), ontology-aware (typed entities). |
| Latency target | p95 < 250ms end-to-end on cached embeddings; < 1.2s cold. |
| Billing | Per-call markup multiplier on underlying token cost, clamped by per-API min and max. |
| Isolation | Per-project Cloud Run service. Dev and prod environments separated. |
| RBAC | Owner / admin / developer / web / viewer roles. Per-API key scoping. |
| Observability | p50/p95/p99 latency, cost per request, error rate, per-API breakdown in BigQuery. |
Quick start
from enqualia_admin_sdk import SDK
sdk = SDK(api_key="sk-...", project_id="proj_...")
# Deploy a RAG API in one call
sdk.upload_api_code(
api_id="docs-rag-v1",
path="./rag_runner.py",
version_description="Hybrid retrieval + reranker",
)
sdk.deploy_to_dev()
sdk.promote_to_prod()
# Clients call it like any other API
import httpx
r = httpx.post(
"https://docs-rag-v1.enqualia.io/run",
headers={"X-API-KEY": "sk-client-..."},
json={"query": "what's our refund policy?", "top_k": 5},
)
print(r.json()["answer"], r.json()["sources"])Ready to deploy?
Spin up a project and deploy this workload to an isolated Cloud Run in minutes.