Skip to content

Parallel function calls from one model response are executed one at a time #226

Description

@adrianyangmeta

Environment

  • SDK version: 0.1.19 (also reproduced on 0.1.8)
  • Python version: 3.12
  • OS / platform: Linux x86_64
  • Install method: pip install google-antigravity

Description

When the model returns several function calls in a single response (Gemini parallel function calling), the harness executes them strictly one after another: it sends the next tool_call to the SDK only after it has received the tool_response for the previous one. A turn with N independent calls therefore takes the sum of their latencies instead of the max.

The Python side is already written for concurrency: LocalHarnessEventProcessor.process_event dispatches each tool_call event as its own task (_run_in_background), and ToolRunner.process_tool_calls runs a batch with asyncio.gather. It only ever receives one call at a time, though, so the serialization appears to happen in the harness binary.

Built-in tools behave the same way: three parallel run_command calls running sleep 3 also ran back to back. The Antigravity CLI (agy 1.2.11) shows the same behavior.

Steps to Reproduce

Self-contained script: a stub Gemini endpoint returns three slow_tool calls in one response, and each call sleeps 3 s.

"""Three function calls returned in ONE model response run one after another."""

import asyncio
import json
import os
import threading
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer

from google.antigravity import Agent, LocalAgentConfig, models, types
from google.antigravity.hooks import policy

os.environ.setdefault("NO_PROXY", "127.0.0.1,localhost")


class StubGemini(BaseHTTPRequestHandler):
    """First call: 3 parallel functionCalls. After the tool results: plain text."""

    def do_POST(self):
        req = json.loads(self.rfile.read(int(self.headers["Content-Length"])))
        answered = any(
            "functionResponse" in p for c in req["contents"] for p in c.get("parts", [])
        )
        parts = (
            [{"text": "done"}]
            if answered
            else [{"functionCall": {"name": "slow_tool", "args": {"n": i}}} for i in range(3)]
        )
        body = json.dumps(
            {"candidates": [{"content": {"role": "model", "parts": parts}, "finishReason": "STOP"}]}
        )
        data = f"data: {body}\r\n\r\n".encode()
        self.send_response(200)
        self.send_header("Content-Type", "text/event-stream")
        self.send_header("Content-Length", str(len(data)))
        self.end_headers()
        self.wfile.write(data)

    def log_message(self, *args):
        pass


T0 = time.monotonic()


async def slow_tool(n: int) -> str:
    """Simulates an I/O-bound tool (e.g. a search call) that takes 3 seconds."""
    print(f"{time.monotonic() - T0:6.2f}s  slow_tool({n}) start")
    await asyncio.sleep(3)
    print(f"{time.monotonic() - T0:6.2f}s  slow_tool({n}) end")
    return f"ok {n}"


async def main():
    server = ThreadingHTTPServer(("127.0.0.1", 0), StubGemini)
    threading.Thread(target=server.serve_forever, daemon=True).start()
    endpoint = models.GeminiAPIEndpoint(
        api_key="stub", base_url=f"http://127.0.0.1:{server.server_port}"
    )
    config = LocalAgentConfig(
        api_key="stub",
        models=[models.ModelTarget(name="gemini-3-flash", types=[models.ModelType.TEXT], endpoint=endpoint)],
        tools=[slow_tool],
        policies=[policy.allow_all()],
        capabilities=types.CapabilitiesConfig(enable_subagents=False, enabled_tools=[]),
    )
    async with Agent(config) as agent:
        await agent.conversation.send("go")
        async for _ in agent.conversation.receive_steps():
            pass
    print(f"{time.monotonic() - T0:6.2f}s  turn finished")


asyncio.run(main())

pip install google-antigravity==0.1.19 && python repro.py

Expected Behavior

The three calls overlap and the turn takes about 3 s (the slowest call).

Actual Behavior

  0.12s  slow_tool(0) start
  3.12s  slow_tool(0) end
  3.12s  slow_tool(1) start
  6.13s  slow_tool(1) end
  6.13s  slow_tool(2) start
  9.13s  slow_tool(2) end
  9.16s  turn finished

Each call starts only after the previous one returns, so the turn takes about 9 s.

Additional Context

For agents whose tools are I/O-bound (search, fetch, remote RPCs), turn latency grows linearly with the number of calls the model batches, which cancels out most of the benefit of parallel function calling. Could function calls from the same model response run concurrently, or could the SDK expose an opt-in for it (for example a CapabilitiesConfig flag, possibly combined with per-tool concurrency hints such as MCP readOnlyHint, see #62)?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions