Deploy Fine-Tuned Models as APIs: Versioned, Metered, Production-Ready
Ship custom fine-tuned LLMs and domain-specific models behind authenticated, metered APIs. Bring weights from OpenAI, Anthropic, Google Vertex, Together, or your own GPU host. Enqualia handles versioning, traffic shifting, RBAC, and per-call billing.
How it works
- 1
Connect your model
Point Enqualia at a hosted fine-tune (OpenAI fine-tuning ID, Vertex endpoint, Together / Replicate model URL) or a private vLLM / TGI / TensorRT-LLM endpoint. Authentication credentials live in your project's gcp_config.
- 2
Define the API contract
Input / output JSON Schema, per-call cost configuration (token rates, multipliers, min / max clamps), optional tracking_id for per-tenant attribution. Validation runs on every request.
- 3
Version and ship
Upload new model versions through the Admin SDK. Test in dev, promote to prod with traffic shifting. Roll back instantly if metrics regress.
What you get on Enqualia
| Model sources | OpenAI fine-tunes, Google Vertex tuned models, Together AI, Replicate, Modal, AWS Bedrock, your own vLLM / TGI / TensorRT-LLM. |
|---|---|
| Per-call billing | Token rates × markup multiplier, clamped by per-API min and max cost. |
| Versioning | Full version history with rollback to any prior version in seconds. |
| Dev / prod | Isolated dev and prod Cloud Run services. Promote with one command. |
| Traffic shifting | Canary releases via Cloud Run revision splits (10% to v2, 90% to v1). |
| Schema validation | JSON Schema check on every request and response. Hard rejection on schema violation. |
| RBAC | API key scopes restrict keys to specific model versions or tracking_ids. |
| Observability | Per-version latency p50/p95/p99, error rate, cost per call, request logs in BigQuery. |
Quick start
from enqualia_admin_sdk import SDK
sdk = SDK(api_key="sk-...", project_id="proj_...")
# Upload a new version of your fine-tuned model API
sdk.upload_api_code(
api_id="claims-classifier-v3",
path="./claims_classifier.py",
version_description="LoRA fine-tune trained on 14k labeled claims",
)
# Roll out to dev → prod
sdk.deploy_to_dev()
# ... run evals against dev ...
sdk.promote_to_prod()
# Roll back if anything regresses
sdk.rollback(version=12)Ready to deploy?
Spin up a project and deploy this workload to an isolated Cloud Run in minutes.