Multi-Modal APIs: Image, Voice, and Vision-Language Models in Production
Deploy image generation, vision-language understanding, and voice AI behind unified APIs. Built-in support for Imagen 4, Gemini 3 Pro Image, Gemini 3.1 Flash Image (Nano Banana 2), Whisper, and streaming TTS, with safety filtering, retries, and signed-URL storage handled by the platform.
How it works
- 1
Pick your modality
Text-to-image, image-to-text, image-edit-with-reference, speech-to-text, text-to-speech, or speech-to-speech. The runtime supports all of them through one ImageGenerator and one VoiceRunner class.
- 2
Configure safety and storage
Per-API safety thresholds, automatic prompt-softening on refusal, retry-with-backoff on 429s, and signed-URL writes to Firebase Storage (your bucket or ours).
- 3
Deploy and meter
Per-image and per-second cost is billed at list price × your markup multiplier. Per-call cost ceilings prevent runaway spends. Latency p95 surfaced per route.
What you get on Enqualia
| Image models | Imagen 4 ($0.04), Imagen 4 Fast ($0.02), Gemini 3.1 Flash Image / Nano Banana 2 ($0.05), Gemini 3 Pro Image ($0.134). |
|---|---|
| Voice models | OpenAI Whisper, gpt-4o-audio, Google Chirp, ElevenLabs synthesis. Bring your own key. |
| Reference images | Edit-with-reference for style transfer, product variants, virtual try-on. Auto-upgrades to higher-quality model when references are provided. |
| Safety | Per-API prompt thresholds, auto soft-rephrase on rejection, refusal logging. |
| Storage | Generated assets stored in Firebase Storage (yours or ours) with signed URLs and retention policies (1d / 7d / 30d / never). |
| Concurrency | Thread-pool parallelism on image generation. Streaming TTS for low first-token latency. |
| Cost ceiling | Per-call max_cost in pricing_config.json. Multiplier deactivation for opt-out APIs. |
| Observability | Per-image generation cost, refusal rate, p95 latency, storage usage per project. |
Quick start
# product_image_runner.py, Enqualia multi-modal API
async def run(state, history, input_data):
from image_generation import ImageGenerator
from backend.cost_calculator import create_response, calculate_metadata
gen = ImageGenerator(model="gemini-3.1-flash-image-preview")
urls = await gen.generate(
prompt=input_data["prompt"],
reference_images=input_data.get("references", []),
count=4,
aspect_ratio="1:1",
)
return create_response(
content={"images": urls},
metadata=calculate_metadata("product-image-v1", image_count=4),
)Ready to deploy?
Spin up a project and deploy this workload to an isolated Cloud Run in minutes.