Use case

Deploy Fine-Tuned Models as APIs: Versioned, Metered, Production-Ready

Ship custom fine-tuned LLMs and domain-specific models behind authenticated, metered APIs. Bring weights from OpenAI, Anthropic, Google Vertex, Together, or your own GPU host. Enqualia handles versioning, traffic shifting, RBAC, and per-call billing.

By Antoni Elkenbracht·Updated January 15, 2026

How it works

  1. 1

    Connect your model

    Point Enqualia at a hosted fine-tune (OpenAI fine-tuning ID, Vertex endpoint, Together / Replicate model URL) or a private vLLM / TGI / TensorRT-LLM endpoint. Authentication credentials live in your project's gcp_config.

  2. 2

    Define the API contract

    Input / output JSON Schema, per-call cost configuration (token rates, multipliers, min / max clamps), optional tracking_id for per-tenant attribution. Validation runs on every request.

  3. 3

    Version and ship

    Upload new model versions through the Admin SDK. Test in dev, promote to prod with traffic shifting. Roll back instantly if metrics regress.

What you get on Enqualia

Model sourcesOpenAI fine-tunes, Google Vertex tuned models, Together AI, Replicate, Modal, AWS Bedrock, your own vLLM / TGI / TensorRT-LLM.
Per-call billingToken rates × markup multiplier, clamped by per-API min and max cost.
VersioningFull version history with rollback to any prior version in seconds.
Dev / prodIsolated dev and prod Cloud Run services. Promote with one command.
Traffic shiftingCanary releases via Cloud Run revision splits (10% to v2, 90% to v1).
Schema validationJSON Schema check on every request and response. Hard rejection on schema violation.
RBACAPI key scopes restrict keys to specific model versions or tracking_ids.
ObservabilityPer-version latency p50/p95/p99, error rate, cost per call, request logs in BigQuery.

Quick start

Python · Admin SDK
from enqualia_admin_sdk import SDK

sdk = SDK(api_key="sk-...", project_id="proj_...")

# Upload a new version of your fine-tuned model API
sdk.upload_api_code(
    api_id="claims-classifier-v3",
    path="./claims_classifier.py",
    version_description="LoRA fine-tune trained on 14k labeled claims",
)

# Roll out to dev → prod
sdk.deploy_to_dev()
# ... run evals against dev ...
sdk.promote_to_prod()

# Roll back if anything regresses
sdk.rollback(version=12)

Ready to deploy?

Spin up a project and deploy this workload to an isolated Cloud Run in minutes.

Frequently asked questions

Can I run fully self-hosted weights?
Yes. Configure your gcp_config to point at a private vLLM, TGI, TensorRT-LLM, or LiteLLM endpoint reachable from your Cloud Run service. Enqualia treats it as another LLM provider. We handle the API contract, billing, RBAC, and observability, you handle the GPU.
How does versioning work with hot-swap?
Every code upload increments the version with a description. The functional code lives in Firestore, and the deployed Cloud Run service hot-loads the latest version on the next cold start (or immediately via redeploy). Rollback is a single SDK call that pins the active version.
What about LoRA adapters?
Two patterns supported: (1) merge the LoRA into the base before serving (recommended for stable adapters) and (2) serve LoRA at runtime on a vLLM endpoint with multi-LoRA support. Enqualia routes the adapter ID from your request schema to the inference server.
How does pricing for fine-tuned models work?
Each API has its own pricing entry (input token rate, output token rate, markup multiplier). For self-hosted models you typically charge a flat per-token rate that covers GPU hosting cost; for managed fine-tunes you mirror the provider's rate and apply your markup. min_cost / max_cost clamps protect against pathological prompts.

Learn the concepts