What is a Fine-Tuned Model?
A fine-tuned model is a base large language model whose weights have been updated on a smaller, domain-specific dataset to specialize behavior: tone, structured-output adherence, task accuracy, or knowledge in a narrow area. Fine-tuning complements (rather than replaces) RAG and prompt engineering.
The fine-tuning lifecycle
- Dataset curation. The single biggest determinant of quality. Aim for diverse, deduplicated, accurately labeled examples that mirror your production distribution.
- Base-model selection. Pick a base whose capabilities are close to your target task. Fine-tuning amplifies, it rarely teaches genuinely new skills.
- Training. Usually LoRA or QLoRA, 1-3 epochs, with held-out evals. Stop when eval loss starts rising.
- Evaluation. Compare against the base model on a fixed eval set. Watch for regression on capabilities the model used to have ("catastrophic forgetting").
- Deployment. Serve behind a versioned API endpoint with traffic shifting and a rollback plan.
Fine-tuning vs alternatives
- Prompt engineering. Try first. Often gets 70-90% of fine-tuning quality at zero training cost.
- RAG. Use when the answer depends on freshly-changing facts. Does not change behavior, only inputs.
- Few-shot prompting. Embed 3-10 examples in the prompt. Costs more per call than fine-tuning but no training required.
- Fine-tuning. Use when you have exhausted prompting AND need lower per-call cost OR strict format adherence.
Production deployment patterns
Three dominant patterns: managed fine-tuning (OpenAI, Anthropic, Google Vertex; easiest path, vendor lock-in), self-hosted on open weights (Llama, Mistral, Qwen, Gemma served via vLLM, TensorRT-LLM, or TGI; maximum control, most operational overhead), and serverless GPU hosting (Modal, Replicate, Together; sweet spot for variable load).
Whichever path you choose, the API contract is the same: clients hit an authenticated endpoint, the model produces a completion, and usage is metered per token. A platform like Enqualia abstracts that contract so you can swap providers without touching client code.
Cost vs frontier model comparison
At low volumes (under 5M tokens/day) frontier models are usually cheaper end-to-end: no training cost, no dedicated GPU. At high volume (over 50M tokens/day), a fine-tuned 7-13B open model on dedicated GPUs typically beats frontier models on cost by 5-20x. The crossover point has moved up as frontier prices fall, but the structural argument still holds for narrow, high-throughput tasks.
How Enqualia handles Fine-Tuned Model
Enqualia exposes this capability as a versioned, metered API on isolated Cloud Run. Read the deployment guide to see schemas, latency targets, and pricing.
Read: Fine-Tuned Models