Use case

Deploy Agentic AI Pipelines: Tool-Calling Agents as Production APIs

Run agentic AI workflows, multi-step LLM reasoning with tool use, memory, and branching, as authenticated, metered APIs on isolated Cloud Run. Set hard step and budget caps, log every tool call, and ship agents that are safe to put in front of paying customers.

By Antoni Elkenbracht·Updated January 15, 2026

How it works

  1. 1

    Define your agent

    Write a stateful Python module: tool registry, planner prompt, memory layer, and a stopping condition. Enqualia validates the code with a security check before deploy.

  2. 2

    Set safety rails

    Configure per-call step ceiling, per-tool timeout, max token budget, and the markup multiplier on underlying inference cost. Every call is clamped by your settings.

  3. 3

    Deploy and observe

    One command pushes the agent to a dev Cloud Run service; one more promotes it to prod. Per-step traces, tool-call logs, and cost breakdowns are surfaced in analytics.

What you get on Enqualia

Agent loopBuilt-in ReAct loop with configurable step cap (default 12). Compatible with tool-calling models.
Tool registryHTTP, SQL, Python sandbox, embedded RAG, custom user tools. Per-tool timeouts.
MemoryPer-session short-term scratchpad. Optional long-term store via Firestore or your vector DB.
Model supportGemini 2.5 Flash/Pro, Gemini 3 Pro, OpenAI GPT-4/5, Claude. Bring your own key.
Cost ceilingHard per-call max cost ($) and max steps. Calls abort cleanly when exceeded.
TracingFull step-by-step trace with thought/action/observation, token spend per step, latency.
ConcurrencyConfigurable Cloud Run concurrency. Auto-scaling on inbound load.
Audit logEvery tool call written to BigQuery with request_id, tracking_id, project_id.

Quick start

Python · Agent runner
# agent_runner.py, runs on Enqualia Cloud Run

async def run(state, history, input_data):
    from backend.cost_calculator import create_response, calculate_metadata
    from enqualia_runtime.agent import ReActAgent

    agent = ReActAgent(
        model="gemini-2.5-pro",
        tools=[search_kb, run_sql, send_email],
        max_steps=10,
        max_cost_usd=0.50,
    )
    result = await agent.run(goal=input_data["goal"])

    return create_response(
        content={"answer": result.final, "steps": result.steps},
        metadata=calculate_metadata("agent-v1", result.usage),
    )

Ready to deploy?

Spin up a project and deploy this workload to an isolated Cloud Run in minutes.

Frequently asked questions

How do you prevent runaway token spend?
Each call is clamped by max_steps and max_cost_usd. The cost calculator runs after every tool call; if the projected next step would breach the cap, the agent halts with a partial response. Per-API min/max prices in pricing_config.json provide a second safety net.
Can agents call other Enqualia APIs as tools?
Yes. Any deployed Enqualia API can be registered as a tool, including RAG endpoints, multi-modal generators, and other agents. The platform's RBAC applies, so an agent can only call APIs the project's key is scoped to.
How is multi-agent orchestration supported?
Two patterns: (1) a coordinator agent dispatches to specialist agents via the tool registry; (2) graph-style orchestration via your favorite framework (LangGraph, CrewAI, AutoGen) deployed inside the Cloud Run runtime. Enqualia is opinion-light on the orchestrator. We provide the runtime, billing, and observability.
What observability is available?
Per-step trace (thought / action / observation), per-tool latency and success rate, token spend per step, end-to-end cost per request, p50/p95/p99 latency, and an error log integrated with Google Cloud Logging. All also queryable via BigQuery for custom dashboards.

Learn the concepts