Use case

Multi-Modal APIs: Image, Voice, and Vision-Language Models in Production

Deploy image generation, vision-language understanding, and voice AI behind unified APIs. Built-in support for Imagen 4, Gemini 3 Pro Image, Gemini 3.1 Flash Image (Nano Banana 2), Whisper, and streaming TTS, with safety filtering, retries, and signed-URL storage handled by the platform.

By Antoni Elkenbracht·Updated January 15, 2026

How it works

  1. 1

    Pick your modality

    Text-to-image, image-to-text, image-edit-with-reference, speech-to-text, text-to-speech, or speech-to-speech. The runtime supports all of them through one ImageGenerator and one VoiceRunner class.

  2. 2

    Configure safety and storage

    Per-API safety thresholds, automatic prompt-softening on refusal, retry-with-backoff on 429s, and signed-URL writes to Firebase Storage (your bucket or ours).

  3. 3

    Deploy and meter

    Per-image and per-second cost is billed at list price × your markup multiplier. Per-call cost ceilings prevent runaway spends. Latency p95 surfaced per route.

What you get on Enqualia

Image modelsImagen 4 ($0.04), Imagen 4 Fast ($0.02), Gemini 3.1 Flash Image / Nano Banana 2 ($0.05), Gemini 3 Pro Image ($0.134).
Voice modelsOpenAI Whisper, gpt-4o-audio, Google Chirp, ElevenLabs synthesis. Bring your own key.
Reference imagesEdit-with-reference for style transfer, product variants, virtual try-on. Auto-upgrades to higher-quality model when references are provided.
SafetyPer-API prompt thresholds, auto soft-rephrase on rejection, refusal logging.
StorageGenerated assets stored in Firebase Storage (yours or ours) with signed URLs and retention policies (1d / 7d / 30d / never).
ConcurrencyThread-pool parallelism on image generation. Streaming TTS for low first-token latency.
Cost ceilingPer-call max_cost in pricing_config.json. Multiplier deactivation for opt-out APIs.
ObservabilityPer-image generation cost, refusal rate, p95 latency, storage usage per project.

Quick start

Python · Image generator
# product_image_runner.py, Enqualia multi-modal API

async def run(state, history, input_data):
    from image_generation import ImageGenerator
    from backend.cost_calculator import create_response, calculate_metadata

    gen = ImageGenerator(model="gemini-3.1-flash-image-preview")
    urls = await gen.generate(
        prompt=input_data["prompt"],
        reference_images=input_data.get("references", []),
        count=4,
        aspect_ratio="1:1",
    )

    return create_response(
        content={"images": urls},
        metadata=calculate_metadata("product-image-v1", image_count=4),
    )

Ready to deploy?

Spin up a project and deploy this workload to an isolated Cloud Run in minutes.

Frequently asked questions

Which image model should I use for what?
Imagen 4 Fast for high-volume, low-cost generation (catalog backgrounds, simple compositions). Gemini 3.1 Flash Image (Nano Banana 2) for edit-with-reference and high-fashion / product photography. Gemini 3 Pro Image for top-tier brand-quality output. Test on your domain, model quality varies sharply by content category.
How does Enqualia handle safety rejections?
On a safety rejection the runtime soft-rephrases the prompt (removing or rewording the flagged segment) and retries up to twice before surfacing an error. Refusal events are logged with the original and rephrased prompts so you can audit and tune.
Can I use my own GCS bucket for generated images?
Yes. Configure storage in gcp_config.storage_config: bucket name, credentials, retention policy, and optional per-API path overrides. Signed URLs are generated against your bucket. If you don't configure custom storage, generated assets go to Enqualia's default bucket billed at $0.02 / GB / month.
How is voice AI priced?
STT (Whisper, Chirp): $0.003–$0.01 per minute of audio. TTS: $0.015–$0.30 per 1k characters. Real-time speech-to-speech: $0.06–$0.10 per minute combined. All billed at list × your markup; per-call cost ceiling configurable per route.

Learn the concepts