Skip to content

Spans are lost because _handle_invocation never flushes the TracerProvider before the microVM freezes #629

Description

@shogo452

Describe the bug

_handle_invocation never calls force_flush() on the OpenTelemetry TracerProvider before returning. AgentCore Runtime freezes the microVM right after the /invocations response completes, but BatchSpanProcessor exports on a 5s timer by default — so spans queued for that request are often lost before the timer fires.

We hit this as 100% empty trace export for an AgentCore-hosted agent, with ADOT otherwise configured correctly.

To Reproduce

from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from starlette.testclient import TestClient
from bedrock_agentcore.runtime.app import BedrockAgentCoreApp

provider = TracerProvider()
flushed = []
orig = provider.force_flush
provider.force_flush = lambda *a, **k: (flushed.append(True), orig(*a, **k))[1]
trace.set_tracer_provider(provider)

app = BedrockAgentCoreApp()

@app.entrypoint
def handler(payload):
    return {"ok": True}

resp = TestClient(app).post("/invocations", json={"x": 1})
print(resp.status_code, "flush_called:", len(flushed))  # flush_called: 0 on main

Expected behavior

force_flush() is called once the response is ready (and again after a streamed response fully drains), so spans survive the freeze.

Root cause

  • ADOT's Lambda auto-flush (opentelemetry-instrumentation-aws-lambda) is gated on AWS_LAMBDA_FUNCTION_NAME and doesn't cover AgentCore Runtime.
  • Nothing else calls flush before the microVM freezes.
  • This SDK already owns the TracerProvider lifecycle here (see _ensure_baggage_processor_registered in runtime/tracing.py), so it's the natural place to fix.

Verified fix

Added _flush_tracer_provider() in runtime/tracing.py (same defensive style as _ensure_baggage_processor_registered), called from _handle_invocation's finally and from both streaming wrappers' finally.

Compare: main...shogo452:bedrock-agentcore-sdk-python:fix/tracer-provider-flush

Diff
--- a/src/bedrock_agentcore/runtime/app.py
+++ b/src/bedrock_agentcore/runtime/app.py
@@ -49,7 +49,7 @@ from .models import (
     PingStatus,
     is_forwardable_header,
 )
-from .tracing import _ensure_baggage_processor_registered
+from .tracing import _ensure_baggage_processor_registered, _flush_tracer_provider
 from .utils import convert_complex_objects

 # Sentinel so we only parse OTEL_RESOURCE_ATTRIBUTES once per process.
@@ -613,6 +613,12 @@ class BedrockAgentCoreApp(Starlette):
             duration = time.time() - start_time
             self.logger.exception("Invocation failed (%.3fs)", duration)
             return JSONResponse({"error": str(e)}, status_code=500)
+        finally:
+            # Flush now so non-streaming spans survive a post-response microVM freeze.
+            # For streaming responses this fires before the generator is consumed;
+            # _stream_with_error_handling/_sync_stream_with_error_handling flush again
+            # once the stream itself finishes.
+            _flush_tracer_provider()

     def _handle_ping(self, request):
         try:
@@ -894,6 +900,8 @@ class BedrockAgentCoreApp(Starlette):
                 "message": "An error occurred during streaming",
             }
             yield self._convert_to_sse(error_event)
+        finally:
+            _flush_tracer_provider()

     def _safe_serialize_to_json_string(self, obj):
         """Safely serialize object directly to JSON string with progressive fallback handling.
@@ -949,3 +957,5 @@ class BedrockAgentCoreApp(Starlette):
                 "message": "An error occurred during streaming",
             }
             yield self._convert_to_sse(error_event)
+        finally:
+            _flush_tracer_provider()
diff --git a/src/bedrock_agentcore/runtime/tracing.py b/src/bedrock_agentcore/runtime/tracing.py
index 7c0cf9c..d770c6f 100644
--- a/src/bedrock_agentcore/runtime/tracing.py
+++ b/src/bedrock_agentcore/runtime/tracing.py
@@ -63,6 +63,28 @@ def _ensure_baggage_processor_registered() -> None:
         logger.debug("Could not register BaggageSpanProcessor", exc_info=True)


+def _flush_tracer_provider(timeout_millis: int = 30000) -> None:
+    """Force-flush the active ``TracerProvider`` before the microVM freezes.
+
+    AgentCore Runtime freezes the microVM as soon as the ``/invocations``
+    response finishes. ``BatchSpanProcessor``/``BatchUnsampledSpanProcessor``
+    export on a timer (default 5s) that may not fire before the freeze,
+    silently dropping any spans still queued. Call this once the response is
+    ready (or, for streamed responses, once the stream is fully consumed) so
+    buffered spans are exported synchronously instead.
+
+    No-ops when ``opentelemetry-api``/``opentelemetry-sdk`` is not installed.
+    """
+    try:
+        from opentelemetry import trace
+
+        trace.get_tracer_provider().force_flush(timeout_millis=timeout_millis)
+    except ImportError:
+        logger.debug("opentelemetry-api not installed; tracer provider flush skipped")
+    except Exception:
+        logger.debug("Could not flush tracer provider", exc_info=True)
+
+
 def _get_base_class() -> type:
     """Return the OTel SDK SpanProcessor base if available, otherwise object.

Verified: full test suite (1190 passed, 1 skipped, no regressions), manual check that force_flush fires for both non-streaming and streaming paths, pre-commit lint passes.

Not opening a PR since this repo doesn't accept external code contributions per CONTRIBUTING.md — the branch above is for reference only.

Environment

  • bedrock-agentcore SDK: 1.21.0 (also seen on 1.9.1)
  • ADOT (aws-opentelemetry-distro): 0.18.0
  • Deployment: Amazon Bedrock AgentCore Runtime (hosted)

Additional context

Related: #471 (different microVM-freeze-timing issue: /ping time_of_last_update). Also flagging the Lambda-only auto-flush gap over in aws-observability/aws-otel-python-instrumentation as FYI, but the fix belongs here since this SDK owns the flush point.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workinghigh-severityOutage, data loss, security issue, crash with no workaround, or regression in a released version

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions