Skip to content

feat(openai_agents): expose real usage, response_id, plumb previous_response_id, opt-in prompt_cache_key for stateful responses and prompt caching - #335

Merged
x merged 8 commits into
mainfrom
dpeticolas/prompt-cache-upstream
May 4, 2026
Merged

x merged 8 commits into
mainfrom
dpeticolas/prompt-cache-upstream

Conversation

@x

@x x commented Apr 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Four changes to the streaming Responses API path in temporal_streaming_model.py.

My goal is to:

  • Make it easier to track prompt cache hits (usage)
  • Make it easier to get cache hits (prompt_cache_key)
  • Make it easier to use the responses API (response_id, previous_response_id)
  1. Capture real usage from ResponseCompletedEvent.response.usage and surface in the span's output dict.

    TemporalStreamingModel.get_response was constructing a zero-filled Usage and discarding event.response.usage. It now reads usage off the streaming protocol and falls back to zeros only on the error path (stream ends without a completed event).

    The len(''.join(reasoning_contents)) // 4 reasoning-tokens estimator is dropped — the real value arrives in the API response.

    The streaming_model_get_response span now carries {input_tokens, output_tokens, total_tokens, cached_input_tokens, reasoning_tokens} so traces show cache-hit rate.

  2. Return real response_id captured from ResponseCompletedEvent.response.id instead of the previously fabricated f"resp_{uuid.uuid4().hex[:8]}".

    OpenAI Agents SDK reads this field for previous_response_id chaining (run.py:145), exposes it as RunResult.last_response_id (result.py:108), and writes it into traces (tracing/span_data.py:164). A client-side UUID has never been issued by any server, so any caller that picks it up and tries to chain (the documented use case) gets a 400 "response not found".

    Latent since this file was added (commit 2f2a6ed7) because nothing in the codebase chains previous_response_id yet.

  3. Forward previous_response_id, conversation_id, and prompt from the SDK kwarg.

    The SDK's abstract Model.get_response has previous_response_id, conversation_id, prompt as required keyword-only params; the SDK threads previous_response_id down through _ServerConversationTracker when callers set it on Runner.run / RunConfig.

    Our implementation declared **kwargs # noqa: ARG002 and silently swallowed all three. Callers who used the official chaining API got no stateful behavior and no error. Replaced with explicit named params.

    This change is not necessarily required we need to pass previous_response_id but we could be doing that via extra_args. That said, the silent behavior is sketch. I could be convinced that the right thing to here is to raise an exception when we see them in kwargs and expect them in extra_args but this felt less dangerous and more intuitive to the caller.

  4. Plumb prompt_cache_key to responses.create as an opt-in parameter.

    Callers set it via model_settings.extra_args["prompt_cache_key"]. We do not auto-inject a default. I don't trust prompt_cache_key to be standard across all OpenAI-compatible endpoints.

    When unset, the parameter resolves to NOT_GIVEN and is omitted from the request body entirely.

Test plan

All 38 tests in test_streaming_model.py pass locally:

  • test_usage_captured_from_completed_event
  • test_usage_falls_back_when_no_completed_event
  • test_usage_emitted_in_span_output
  • test_response_id_captured_from_completed_event
  • test_response_id_is_none_when_no_completed_event — guards against the fake-UUID footgun
  • test_prompt_cache_key_not_sent_by_default — verifies NOT_GIVEN fallback for non-OpenAI compat
  • test_prompt_cache_key_forwarded_when_opted_in
  • test_previous_response_id_not_sent_by_default — verifies NOT_GIVEN fallback
  • test_previous_response_id_forwarded_via_sdk_kwarg — verifies the SDK's official chaining API now works
  • test_conversation_id_and_prompt_accepted_but_not_forwarded — verifies SDK contract compliance without surface-area expansion
  • All 28 pre-existing tests still pass (after stripping vestigial task_id kwargs)

Why these changes together

The motivating workstream (downstream of this PR) is a stateful-Responses-API migration for ST&S agents to chain via previous_response_id for 40–80% better cache utilization on reasoning models. That migration needs:

  • Real Usage to measure cache hit rate before/after.
  • Real response_id to actually pass back as previous_response_id.
  • previous_response_id actually reaching responses.create instead of being silently swallowed.
  • Optional prompt_cache_key for callers who want it without forcing it on everyone.

Once a release containing these changes lands and aimi-scale bumps its agentex-sdk pin, the runtime monkey-patch in st_s/*/agentex_usage_patch.py and the corresponding _apply_agentex_usage_patch() calls in each run_worker.py can be deleted.

Greptile Summary

Four correctness fixes to the Responses API streaming path: (1) real Usage is now read from ResponseCompletedEvent.response.usage and surfaced in the span; (2) response_id is the server-issued value instead of a fabricated UUID that would 400 on any chained turn; (3) previous_response_id, conversation_id, and prompt are explicitly accepted and forwarded rather than silently swallowed by **kwargs; (4) prompt_cache_key is opt-in via extra_args and resolves to NOT_GIVEN by default.

Confidence Score: 5/5

Safe to merge — all changes are correctness improvements with safe fallback paths and comprehensive test coverage.

No P0 or P1 findings. All four changes fix real latent bugs (fake UUID chaining, zero usage, silently dropped SDK params). Fallbacks are explicit (zeros, None, NOT_GIVEN). The removal of **kwargs is intentional and acknowledged. Tests cover every new code path.

No files require special attention.

Important Files Changed

Filename Overview
src/agentex/lib/core/temporal/plugins/openai_agents/models/temporal_streaming_model.py Captures real usage and response_id from ResponseCompletedEvent, forwards previous_response_id/conversation_id/prompt as explicit kwargs (replacing **kwargs swallowing), and plumbs opt-in prompt_cache_key — all correctness improvements with safe fallbacks.
src/agentex/lib/core/temporal/plugins/openai_agents/tests/test_streaming_model.py Adds 10 targeted tests for the new behaviour (usage capture, response_id, prompt_cache_key opt-in, previous_response_id chaining, conversation/prompt forwarding) and strips the now-vestigial task_id kwarg from 28 existing test calls.

Sequence Diagram

sequenceDiagram
    participant SDK as OpenAI Agents SDK
    participant TSM as TemporalStreamingModel
    participant OAI as OpenAI responses.create
    participant Redis as Redis (streaming)

    SDK->>TSM: get_response(..., previous_response_id, conversation_id, prompt)
    Note over TSM: pops prompt_cache_key from extra_args (NOT_GIVEN default)
    TSM->>OAI: responses.create(stream=True, previous_response_id, conversation, prompt, prompt_cache_key, ...)
    loop stream events
        OAI-->>TSM: ResponseTextDeltaEvent / etc.
        TSM->>Redis: stream partial content
    end
    OAI-->>TSM: ResponseCompletedEvent(response.id, response.usage)
    Note over TSM: captured_response_id = response.id<br/>captured_usage = response.usage
    TSM-->>SDK: ModelResponse(response_id=captured_response_id, usage=real_usage)
    Note over SDK: RunResult.last_response_id = captured_response_id<br/>(used for next turn's previous_response_id)
Loading

Reviews (7): Last reviewed commit: "Merge branch 'main' into dpeticolas/prom..." | Re-trigger Greptile

Comment thread src/agentex/lib/core/temporal/plugins/openai_agents/tests/test_streaming_model.py Outdated
x added 3 commits April 30, 2026 17:12
…_key

Two related changes to the streaming Responses API path in
`TemporalStreamingModel.get_response`. Both are observability/cache
improvements; neither changes how the API is called for callers who
don't opt in.

1. Capture real `Usage` from `ResponseCompletedEvent.response.usage`.
   Previously a zero-filled `Usage` was constructed and `event.response.usage`
   from the streaming protocol was discarded. The reasoning-tokens
   estimator (`len(''.join(reasoning_contents)) // 4`) is also dropped —
   the real value arrives in the API response. Falls back to zeros only
   when the stream ends without a `ResponseCompletedEvent` (error path).

2. Surface usage in the span's `output` dict. The
   `streaming_model_get_response` span now carries
   `{input_tokens, output_tokens, total_tokens, cached_input_tokens,
   reasoning_tokens}` so traces show cache-hit rate without external
   log scraping.

3. Plumb `prompt_cache_key` to `responses.create` as an opt-in. Callers
   set it via `model_settings.extra_args["prompt_cache_key"]`. We do not
   auto-inject a default — `prompt_cache_key` is not standard across
   OpenAI-compatible endpoints, and a non-OpenAI server that strictly
   validates request bodies could reject the field. When unset, the
   parameter resolves to `NOT_GIVEN` and is omitted from the request
   body entirely. Behavior on alternative providers is identical to
   today's unless a caller explicitly opts in.
`TemporalStreamingModel.get_response` was synthesizing a client-side
UUID for `ModelResponse.response_id`:

    response_id=f"resp_{uuid.uuid4().hex[:8]}"

Replace with the real `response.id` captured off
`ResponseCompletedEvent.response.id` (alongside the `Usage` capture
already happening in the same branch). On the error path, where the
stream ends without a `ResponseCompletedEvent`, we return `None` —
matching the documented `str | None` contract on
`ModelResponse.response_id`.

## Why this matters

The OpenAI Agents SDK reads `ModelResponse.response_id` in three places:

- `agents/run.py:145` — gates whether the SDK chains via
  `previous_response_id` on the next call (the conditional is
  None-tolerant: a None value just leaves the chain pointer alone).
- `agents/result.py:108` — exposes the value to user code as
  `RunResult.last_response_id`.
- `agents/tracing/span_data.py:164` — written into trace records.

A client-side UUID was never issued by any server. Any caller that
picks it up and tries to chain via `previous_response_id` (the
documented use case for `RunResult.last_response_id`) gets a 400
"response not found" from the API, surfacing far from the actual
cause.

Comparable SDK providers do this correctly:

- `agents/models/openai_responses.py:149`: `response_id=response.id`
- `agents/models/openai_chatcompletions.py:135`: `response_id=None`
- `agents/extensions/models/litellm_model.py:182`: `response_id=None`

`None` is the documented sentinel for "this provider doesn't support
response_id," and the SDK is built to handle it.

The bug has been latent since this file was added (commit 2f2a6ed,
Oct 10) because nothing in the codebase's call paths chains
`previous_response_id` yet. The first caller that does (e.g. a
multi-turn stateful Responses API workflow) triggers it.

## Compatibility

This change is invisible to callers that don't read `response_id` — and
nothing in `scale-agentex-python` reads it. A repo-wide grep finds zero
consumers; only the (now-fixed) write site exists. The field is
serialized into Temporal event history and trace records but consumed
only by the OpenAI Agents SDK, which already handles `None`.
… key

Seven new tests in `TestStreamingModelUsageResponseIdAndCacheKey`:

- Usage captured from `ResponseCompletedEvent.response.usage`
- Usage falls back to zeros when stream ends without a completed event
- Usage emitted in span output_data["usage"]
- response_id captured from `ResponseCompletedEvent.response.id`
- response_id is None (NOT a fabricated UUID) when stream ends without
  a completed event — guards against the previous footgun where a
  client-side UUID would be returned and silently break downstream
  `previous_response_id` chaining
- prompt_cache_key resolves to NOT_GIVEN by default (omitted from
  request body, safe for non-OpenAI endpoints)
- prompt_cache_key forwarded when caller opts in via
  `model_settings.extra_args["prompt_cache_key"]`, and popped from
  extra_args so it isn't passed twice

Pre-existing tests in `TestStreamingModelBasics` (test_responses_api_streaming,
test_task_id_threading, test_redis_context_creation) updated to set
`response.id=None` on their `MagicMock(spec=ResponseCompletedEvent)`
mocks. Without this, the auto-generated MagicMock attribute for
`response.id` flows into `ModelResponse.response_id` and trips
pydantic's `str | None` validation.
@x
x force-pushed the dpeticolas/prompt-cache-upstream branch from dc7de0f to eb8ff68 Compare April 30, 2026 21:15
@x x changed the title feat(openai_agents): capture real Usage + enable prompt_cache_key routing feat(openai_agents): expose real Usage and response_id, opt-in prompt_cache_key Apr 30, 2026
@x x changed the title feat(openai_agents): expose real Usage and response_id, opt-in prompt_cache_key feat(openai_agents): expose real response.usage, map response_id=response.id, opt-in prompt_cache_key Apr 30, 2026
The OpenAI Agents SDK's `Model.get_response` abstract has three
keyword-only parameters: `previous_response_id`, `conversation_id`,
`prompt`. The SDK threads them down through `_ServerConversationTracker`
when callers use `Runner.run(..., previous_response_id=X)` or set
`RunConfig` with `auto_previous_response_id=True`.

`TemporalStreamingModel.get_response` was declared with
`**kwargs # noqa: ARG002`, which silently swallowed all three. Callers
who used the SDK's official chaining API saw their `previous_response_id`
disappear and got no stateful behavior — without an error.

This commit:

- Replaces `**kwargs` with explicit `previous_response_id`,
  `conversation_id`, `prompt` params, matching the abstract.
- Forwards `previous_response_id` to `responses.create` via
  `_non_null_or_not_given` (so `None` resolves to `NOT_GIVEN` and the
  field is omitted from the request body — identical behavior to today
  for callers that don't opt in).
- Accepts `conversation_id` and `prompt` for SDK contract compliance
  but does not forward them yet (marked `# noqa: ARG002`); they can be
  wired through later if a use case appears.

## Compatibility with non-OpenAI backends

Same opt-in pattern as `prompt_cache_key`. `TemporalStreamingModel`
calls `responses.create`, but the underlying client can be pointed at
any OpenAI-compatible server (LiteLLM proxy, Foundry, vLLM, etc.).
Some of those backends don't recognize `previous_response_id`. Because
we forward it only when explicitly set, callers who don't opt in see
no change in the wire request — the field is filtered out by
`NOT_GIVEN`. Callers who opt in are responsible for knowing whether
their backend supports it.

## Test housekeeping

The 27 existing tests that passed `task_id=sample_task_id` to
`get_response` were relying on `**kwargs` to silently swallow it.
Production reads `task_id` from a ContextVar (set by
`ContextInterceptor` in real Temporal flows, set by the
`_streaming_context_vars` fixture in tests), not from a function
argument. The kwarg was vestigial cruft. Removed.
@x x changed the title feat(openai_agents): expose real response.usage, map response_id=response.id, opt-in prompt_cache_key feat(openai_agents): expose real Usage/response_id, plumb previous_response_id, opt-in prompt_cache_key Apr 30, 2026
@x x changed the title feat(openai_agents): expose real Usage/response_id, plumb previous_response_id, opt-in prompt_cache_key feat(openai_agents): expose real usage, response_id, plumb previous_response_id, opt-in prompt_cache_key for stateful responses and prompt caching Apr 30, 2026
x added 2 commits April 30, 2026 18:08
…create

The SDK's ``Model.get_response`` abstract has three Responses API
server-state parameters: ``previous_response_id``, ``conversation_id``,
``prompt``. The prior commit wired up ``previous_response_id`` and
accepted the other two for SDK contract compliance but discarded them
with ``# noqa: ARG002``.

Accept-and-discard is a code smell: callers using the SDK's
``Runner.run(conversation_id=..., prompt=...)`` API would see their
arguments silently dropped. Since both map directly to ``responses.create``
kwargs and we're already on that endpoint, the cost of forwarding is two
lines and removes the smell entirely.

- ``conversation_id`` (SDK abstract name) → ``conversation`` (responses.create
  endpoint kwarg). The ``Conversation`` type accepts ``str`` directly, so
  no translation is needed.
- ``prompt`` is the same name on both sides.

Both follow the same opt-in pattern as ``previous_response_id`` and
``prompt_cache_key``: ``None`` resolves to ``NOT_GIVEN`` and is omitted
from the request body, so behavior on alternative OpenAI-compatible
backends is unchanged unless a caller explicitly opts in.
… union

Two ruff fixes for the test file:

- ARG002 on 25 test method signatures: the prior commit
  (forward previous_response_id from SDK kwarg) stripped the
  vestigial ``task_id=sample_task_id`` kwargs from get_response calls,
  but left ``sample_task_id`` in the test method parameter lists.
  The contextvars fixture (``_streaming_context_vars``) already pulls
  ``sample_task_id`` transitively, so the explicit param is redundant.
  Removed from the 25 flagged signatures; preserved on
  ``test_responses_api_streaming`` where it's still used inside the
  body to assert against the streaming context.

- FA102 on _make_response_completed_event: the new test helper used a
  PEP 604 union (``str | None``) without ``from __future__ import
  annotations``. Switched to ``Optional[str]`` to keep the change
  local to the helper rather than retrofitting future annotations
  across the file.
@x
x enabled auto-merge May 1, 2026 15:34
auto-merge was automatically disabled May 4, 2026 17:47

Merge commits are not allowed on this repository

@x
x merged commit ba5d64b into main May 4, 2026
32 checks passed
@x
x deleted the dpeticolas/prompt-cache-upstream branch May 4, 2026 18:47
@stainless-app stainless-app Bot mentioned this pull request May 4, 2026
@stainless-app stainless-app Bot mentioned this pull request Jun 9, 2026
michaelxu2288 pushed a commit to michaelxu2288/scale-agentex-python that referenced this pull request Oct 1, 2026
…caleapi#506 (scaleapi#4)

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* feat(adk): allow all ClaudeAgentOptions in run_claude_agent_activity

* release: 0.9.8

* Bump LiteLLM and urllib3

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* fix(client): preserve hardcoded query params when merging with user params

* codegen metadata

* codegen metadata

* codegen metadata

* release: 0.9.9

* feat(adk): Revamp run_claude_agent_activity to use more streaming (scaleapi#309)

* codegen metadata

* Fix cost bug (scaleapi#313)

* codegen metadata

* release: 0.9.10

* Fix crash when .dockerignore file is missing during cloud build

The build context preparation crashes with FileNotFoundError when a
manifest specifies a dockerignore path but the file doesn't exist on
disk. This adds an existence check and logs a warning instead of
crashing, so builds proceed with no ignore patterns.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add AgentCard for self-describing agent capabilities (scaleapi#296)

* Add AgentCard feature for self-describing agent capabilities via registration_metadata

* Add tests for AgentCard feature, fix PEP 604 union unwrap in extract_literal_values

* Fix ruff import sorting in __init__.py and test file

* Fix pyright strict errors: use Enum isinstance checks, add override decorators in tests

* Add AgentCard.from_states() classmethod for list[State] + initial_state usage

* Minimize registration.py diff: only add agent_card param and merge logic

* Add missing AGENTEX_DEPLOYMENT_ID to test mock env vars

* fix(temporal): allowing-ACP-temporal-telemetry

* fix: Temporal Union deserialization causing tool_response messages to be lost

Temporal's payload converter deserializes Union types by trying each
variant in order. ToolResponseContent was silently misdeserialized as
TextContent (both share 'author' and 'content' fields), creating text
messages instead of tool_response messages in the database.

Fix: hooks now pass .model_dump() dicts to the activity, and the
activity reconstructs the correct Pydantic model using the 'type'
discriminator. Also fix test polling to handle the DONE/tool_response
ordering race condition.

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* feat(api): api update

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* fix: ensure file data are only sent as 1 parameter

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* release: 0.10.0

* Add ShellTool support to TemporalStreamingModel

openai-agents introduced a next-generation ShellTool (replacing
LocalShellTool) that carries an environment config like
{"type": "local", "skills": [...]}. The Temporal streaming model was
dropping it with "Unknown tool type: ShellTool, skipping", so agents
running through AgentEx/Temporal lost the tool entirely even though
plain Runner.run(...) worked.

Serialize ShellTool to the Responses API "shell" payload, defaulting
environment to {"type": "local"} when unset. Import is guarded so
users on older openai-agents versions (ShellTool not yet exported)
continue to work.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Upgrade openai-agents to 0.14.1 and temporalio to >=1.26.0

ShellTool (the next-gen replacement for LocalShellTool) is only
exported in modern openai-agents versions. With the old 0.4.2 pin
the ShellTool branch added in the prior commit was unreachable by
default-install users.

Bumps:
- openai-agents 0.4.2 -> 0.14.1
- temporalio >=1.18.2 -> >=1.26.0 (matches the version that supports
  ShellTool serialization in temporalio.contrib.openai_agents)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Narrow ComputerTool.computer union for Responses API serialization

openai-agents 0.14 widened ComputerTool.computer to accept factory
types (ComputerCreate/ComputerProvider) that don't expose environment
or dimensions. Match the upstream pattern: narrow to Computer /
AsyncComputer before reading those attributes, and validate that
environment/dimensions are set.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* release: 0.10.1

* add support for Temporal PayloadCodec (scaleapi#328)

* codegen metadata

* codegen metadata

* perf(client): optimize file structure copying in multipart requests

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* feat(api): api update

* fix(adk): fix to queue drain (scaleapi#327)

Co-authored-by: Declan Brady <declan.brady@scale.com>

* codegen metadata

* Add task_id to span creation (scaleapi#329)

* release: 0.10.2

* fix(tests): repair test_streaming_model so all 28 tests run and pass (scaleapi#334)

Four pre-existing bugs left this entire test file unrunnable on main (4
failures + 24 errors); fixing them here so the suite actually exercises
TemporalStreamingModel and protects against regressions.

Bug 1 (24 errors): `conftest.py` defines fixture `mock_adk_streaming` (no
underscore) but every test in TestStreamingModelSettings and
TestStreamingModelTools requested it as `_mock_adk_streaming`, so pytest
failed to resolve the fixture before the body ever ran. The fixture is
``autouse=True`` and the param value was never used in any test body, so
the parameter was vestigial — replaced with `_streaming_context_vars`,
which provides the ContextVar setup these tests now actually need.

Bug 2 (4 failures): `TemporalStreamingModel.get_response()` reads
`task_id`, `trace_id`, and `parent_span_id` from ContextVars populated
by `ContextInterceptor` from request headers in real Temporal flows.
Tests had been passing `task_id=...` as a kwarg, which is silently
swallowed by `**kwargs` and ignored, so all three ContextVars stayed at
their defaults and the validation at the top of `get_response` raised
before any work happened. New `_streaming_context_vars` fixture in
conftest sets all three vars (and resets them on teardown), simulating
what `ContextInterceptor` does in production.

Bug 3 (test_computer_tool): A recent commit narrowed `ComputerTool`
serialization to require an actual `Computer`/`AsyncComputer` instance,
but `sample_computer_tool` still built a bare `MagicMock`. Switched to
`MagicMock(spec=Computer)` so the production isinstance check passes.

Bug 4 (3 streaming-context tests): The 3 tests in TestStreamingModelBasics
that assert on `streaming_task_message_context` calls built event
sequences with raw `MagicMock(type="...")`. Production dispatches via
`isinstance(event, ResponseOutputItemAddedEvent)` etc., which `MagicMock`
without `spec` never satisfies, so dispatch was silently skipped and
the assertions failed. Switched to `MagicMock(spec=...)` for each event
type — passes isinstance without triggering pydantic validation on the
event's required fields. Also fixed `test_task_id_threading` which had
been asserting against a hardcoded `task_id="test_task_12345"` that was
never actually threaded anywhere (the kwarg was ignored, just like in
Bug 2); it now asserts against the value yielded by the fixture, which
is the value production reads from the ContextVar.

After all four fixes: 28/28 pass, ruff clean, pyright clean.

* release: 0.10.3 (scaleapi#330)

* feat(api): api update

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* chore(internal): more robust bootstrap script

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* fix: use correct field name format for multipart file arrays

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* feat: support setting headers via env

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* codegen metadata

* fix: allow litellm security patch (scaleapi#336)

* fix(adk): Always inject headers on execute activity (scaleapi#337)

* perf(streaming): coalesce per-token publishes to Redis (50ms / 128-char window) (scaleapi#333)

* perf(streaming): coalesce per-token publishes to Redis (50ms / 128-char window)

Per-token Redis publishes from TemporalStreamingModel were adding ~45s
(56-62%) overhead to agent response latency, mostly from head-of-line
blocking on the model's event loop: each `await streaming_context.stream_update(...)`
inside the OpenAI stream `async for` paused token consumption until the
publish round-trip completed.

This change introduces a `CoalescingBuffer` driven by an `asyncio.Event`,
so the producer never awaits on Redis. Deltas are merged consecutive-only
(preserving character order in every (type, index) channel) and flushed
on a 50ms timer, on a 128-char size threshold, or immediately for the
first delta to keep perceived responsiveness high. The buffer's `close()`
drains remaining deltas before the DONE event, so consumers see the full
sequence in order.

A new `StreamingMode = Literal["off", "per_token", "coalesced"]` lives
in `streaming.py` as the single source of truth and is plumbed through
the adk streaming module, `StreamingService.streaming_task_message_context`,
and `StreamingTaskMessageContext`. Default is `"coalesced"` everywhere,
so all 13+ existing context callers (claude_agents, langgraph, litellm
provider, openai sync provider, etc.) benefit automatically.

* chore(streaming): fix import ordering (ruff I001)

* fix(streaming): address greptile review findings

- _run: when CancelledError is raised mid-flush in the for-loop, re-enqueue
  the in-flight item plus any remaining items in the local `drained` list
  back into self._buf so close()'s final drain can recover them. Previously
  the local `drained` list was unreachable after CancelledError exited the
  for-loop, causing the last coalesced batch to be silently dropped on
  close-during-flush races. Trade-off: the in-flight item may be duplicated
  on the consumer side (Redis pub may have completed before cancel was
  delivered), which is preferable to silent loss for streaming UX.

- _merge_pair: replace `return b` fallback with AssertionError. All six
  current TaskMessageDelta variants have explicit isinstance branches, so
  the fallback is unreachable today. But _can_merge returns True for any
  same-type pair, so adding a 7th delta variant without updating
  _merge_pair would silently drop `a`'s accumulated content. Asserting
  turns a future silent data-loss into an immediate, diagnosable crash.

* test(streaming): add coalescing-layer tests; loosen one model assertion

After merging the test-suite repair from main (scaleapi#334) into this branch, one
model test (test_responses_api_streaming) regressed because its
assert_called_with strict-matched all kwargs of streaming_task_message_context
and didn't tolerate the new `streaming_mode='coalesced'` kwarg this PR
adds. Switched to assert_called() + targeted kwarg checks so the test
verifies what it cares about (task_id threading) without locking in
implementation details.

Replaced the ad-hoc smoke scripts that lived in conversation with a real
pytest module at tests/lib/core/services/adk/test_streaming.py covering:

- _delta_char_len, _can_merge, _merge_pair: per-channel correctness +
  None-handling
- _merge_consecutive: pure-text collapse, cross-channel order preservation,
  per-channel reconstruction matches per-token semantics
- CoalescingBuffer: first-delta-immediate flush within ~20ms,
  size-threshold flush before timer fires, multi-delta coalescing within
  one window, idle close, add-after-close no-op
- CoalescingBuffer cancel-during-flush regression test for the P1 fix:
  five queued chunks must all surface across publishes when close()
  cancels mid-flush (asserts substring presence rather than exact
  ordering, since the documented trade-off allows duplicates of the
  in-flight item)
- StreamingTaskMessageContext mode dispatch: "off" suppresses publishes
  but persists full content, "per_token" publishes each delta synchronously,
  "coalesced" batches and persists full content

* chore(streaming): route TemporalStreamingModel logger through make_logger

The model file used raw ``logging.getLogger("agentex.temporal.streaming")``,
which returns a logger with no handler attached and no level configured —
so the existing ``[TemporalStreamingModel] Initialized ... streaming_mode=...``
INFO log was silently dropped, making it impossible to verify at runtime
that a coalesced (or any) streaming mode was actually wired.

Switch to the SDK's ``make_logger`` helper (level=INFO, RichHandler in
local mode, StreamHandler otherwise) used everywhere else in the SDK.
The explicit logger name ``agentex.temporal.streaming`` is preserved so
any external logging configuration targeting that name keeps working.

* codegen metadata

* feat(api): api update

* release: 0.10.3

---------

Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Brandon Allen <brandon.allen@scale.com>
Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: Stas Moreinis <stas.moreinis@scale.com>

* release: 0.10.4 (scaleapi#338)

Co-authored-by: alvinkam2001 <alvin.kam@scale.com>
Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>

* feat(openai_agents): expose real `usage`, `response_id`, plumb `previous_response_id`, opt-in `prompt_cache_key` for stateful responses and prompt caching (scaleapi#335)

Co-authored-by: Stas Moreinis <stas.moreinis@scale.com>

* build(deps) bump scale-gp-beta to 0.2.0 (scaleapi#344)

* release: 0.10.5 (scaleapi#343)

Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Alvin Kam <alvin.kam@scale.com>

* Fix Redis stream leak: MAXLEN on xadd + sliding TTL on stream keys (scaleapi#339)

* ci: add conventional commit and PR base checks (scaleapi#346)

* release: 0.11.0 (scaleapi#345)

Co-authored-by: Daniel Miller <daniel.miller@scale.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>

* fix: render .env.example template in agentex init (scaleapi#351)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* release: 0.11.1 (scaleapi#350)

Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Daniel Miller <daniel.miller@scale.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Stas Moreinis <stas.moreinis@scale.com>
Co-authored-by: Devon Peticolas <devon.peticolas@scale.com>

* release: 0.11.2 (scaleapi#357)

Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>

* release: 0.11.3 (scaleapi#358)

Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: Michael Chou <michael.chou@scale.com>

* release: 0.11.4 (scaleapi#364)

Co-authored-by: Stas Moreinis <stas.moreinis@scale.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* release: 0.11.5 (scaleapi#369)

Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Stas Moreinis <stas.moreinis@scale.com>
Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: Michael Chou <michael.chou@scale.com>
Co-authored-by: Daniel Miller <daniel.miller@scale.com>

* release: 0.11.6 (scaleapi#376)

Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Stas Moreinis <stas.moreinis@scale.com>
Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: Michael Chou <michael.chou@scale.com>
Co-authored-by: Daniel Miller <daniel.miller@scale.com>
Co-authored-by: Matteo Librizzi <matteo.librizzi@scale.com>

* release: 0.11.7 (scaleapi#382)

Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Stas Moreinis <stas.moreinis@scale.com>
Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: Michael Chou <michael.chou@scale.com>
Co-authored-by: Daniel Miller <daniel.miller@scale.com>
Co-authored-by: Matteo Librizzi <matteo.librizzi@scale.com>
Co-authored-by: Stas Moreinis <smoreinis@gmail.com>

* release: 0.11.8 (scaleapi#386)

Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Stas Moreinis <stas.moreinis@scale.com>
Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: Michael Chou <michael.chou@scale.com>
Co-authored-by: Daniel Miller <daniel.miller@scale.com>
Co-authored-by: Matteo Librizzi <matteo.librizzi@scale.com>
Co-authored-by: Stas Moreinis <smoreinis@gmail.com>
Co-authored-by: James Cardenas <james.cardenas@scale.com>

* release: 0.11.9 (scaleapi#389)

Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Stas Moreinis <stas.moreinis@scale.com>
Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: Michael Chou <michael.chou@scale.com>
Co-authored-by: Daniel Miller <daniel.miller@scale.com>
Co-authored-by: Matteo Librizzi <matteo.librizzi@scale.com>
Co-authored-by: Stas Moreinis <smoreinis@gmail.com>
Co-authored-by: James Cardenas <james.cardenas@scale.com>

* release: 0.12.0 (scaleapi#390)

Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: release main (scaleapi#393)

Co-authored-by: Jerome Romualdez <jerome.romualdez@scale.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: Nitesh Dhanpal <NiteshDhanpal@users.noreply.github.com>

* chore: release main (scaleapi#404)

Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>

* chore: release main (scaleapi#411)

Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Stas Moreinis <stas.moreinis@scale.com>
Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: Michael Chou <michael.chou@scale.com>
Co-authored-by: Daniel Miller <daniel.miller@scale.com>
Co-authored-by: Matteo Librizzi <matteo.librizzi@scale.com>
Co-authored-by: Stas Moreinis <smoreinis@gmail.com>
Co-authored-by: James Cardenas <james.cardenas@scale.com>
Co-authored-by: Nitesh Dhanpal <NiteshDhanpal@users.noreply.github.com>

* chore: release main (scaleapi#424)

Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: Vijay Kalmath <158184866+vkalmathscale@users.noreply.github.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Nitesh Dhanpal <NiteshDhanpal@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Daniel Miller <daniel.miller@scale.com>

* chore: release main (scaleapi#443)

Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: OpenAI <openai@example.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>

* chore: release main (scaleapi#448)

Co-authored-by: Endre Berki <endre.berki@scale.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>

* chore: release main (scaleapi#452)

Co-authored-by: Jerome Romualdez <jerome.romualdez@scale.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>

* chore: release main (scaleapi#456)

Co-authored-by: Daniel Miller <daniel.miller@scale.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>

* chore: release main (scaleapi#457)

Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: Vijay Kalmath <158184866+vkalmathscale@users.noreply.github.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Nitesh Dhanpal <NiteshDhanpal@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Daniel Miller <daniel.miller@scale.com>

* chore: release main (scaleapi#461)

Co-authored-by: Nitesh Dhanpal <NiteshDhanpal@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Levi Lentz <levi.lentz@scale.com>

* chore: release main (scaleapi#463)

Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Declan Brady <declan.brady@scale.com>

* chore: release main (scaleapi#464)

Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Nitesh Dhanpal <NiteshDhanpal@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Declan Brady <declan.brady@scale.com>

* chore: release main (scaleapi#475)

Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Nitesh Dhanpal <NiteshDhanpal@users.noreply.github.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Deepthi Rao <deepthi.rao@scale.com>

* chore: release main (scaleapi#479)

Co-authored-by: Deepthi Rao <deepthi.rao@scale.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>

* chore: release main (scaleapi#483)

Co-authored-by: Deepthi Rao <deepthi.rao@scale.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>

* chore: release main (scaleapi#487)

Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Nitesh Dhanpal <NiteshDhanpal@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Javed Shaik <javed.shaik@scale.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Alvin Kam <alvin.kam@scale.com>

* chore: release main (scaleapi#492)

Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Nitesh Dhanpal <NiteshDhanpal@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* chore: release main (scaleapi#499)

Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Nitesh Dhanpal <NiteshDhanpal@users.noreply.github.com>
Co-authored-by: Alvin Kam <alvin.kam@scale.com>

* codegen metadata

* feat(tracing): add opt-in commit SHA stamping for SGP spans (scaleapi#505)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* codegen metadata

* chore: release main (scaleapi#506)

Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Cynthia Wang <cynthia.wang@scale.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Rishav Chakravarti <rishav.chakravarti@scale.com>
Co-authored-by: Declan Brady <declan.brady@scale.com>

* ci: guard production-only workflows so they no-op on staging

The promote model keeps staging main and production main SHA-identical, so
every workflow file is shared. Four of them are production-app-specific and
have none of their secrets on staging, where they would run and fail red on
every codegen push -- and permanently red staging CI is what makes a genuinely
red build invisible.

Guarded on github.repository: agentex-tutorials-test (TUTORIAL_* keys),
build-and-push-tutorial-agent (PACKAGE_TOKEN), harness-integration, and
publish-pypi (matching the ts side, whose publish-npm is already guarded).

Entry jobs only -- dependents skip via needs -- except test-summary, which is
if: always() and so needed the condition ANDed.

ci.yml is deliberately left unguarded: it references no secrets and running
the SDK's own lint/test on staging is a useful signal that codegen is sound.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci(bandit): read scan results from file instead of passing them as argv

The 'Format results appropriately from results.json' step passed the entire
results file as a single shell argument:

  jq --argjson scanResults "$(<tmp.json)" ...

Linux caps one argv entry at MAX_ARG_STRLEN (128KB, 32 pages) regardless of the
much larger total ARG_MAX, so once a scan produced more than ~128KB of findings
the step died with 'Argument list too long' (exit 126) and failed the whole job
-- even though 'shell: bash {0}' and the comment above it intend the logging
step to be non-fatal.

--slurpfile reads the file directly, so the payload size stops mattering. It
wraps the file's values in an array, hence the [0]. Verified to produce
byte-identical output to the old form on small inputs, and to handle 20k
findings (3.2MB) where the old form exits non-zero.

This surfaced on the staging reconciliation PR, where bandit's baseline scan
runs against a main branch that has no Python in it -- so every finding in all
792 files landed in results.json at once.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: stainless-app[bot] <142633134+stainless-app[bot]@users.noreply.github.com>
Co-authored-by: Declan Brady <declan.brady@scale.com>
Co-authored-by: Raj Krishnan <raj.krishnan@scale.com>
Co-authored-by: Daniel Miller <daniel.miller@scale.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Prassanna Ravishankar <prassanna.ravishankar@scale.com>
Co-authored-by: Bruce Pannaman <bruce.pannaman@scale.com>
Co-authored-by: Bruce Pannaman <brucey31@users.noreply.github.com>
Co-authored-by: Endre Berki <endre.berki@scale.com>
Co-authored-by: Levi Lentz <levilentz@gmail.com>
Co-authored-by: Stas Moreinis <stas.moreinis@scale.com>
Co-authored-by: Brandon Allen <brandon.allen@scale.com>
Co-authored-by: alvinkam2001 <alvin.kam@scale.com>
Co-authored-by: Devon Peticolas <devon.peticolas@scale.com>
Co-authored-by: Jean Lucas <jeanlpf@hotmail.com>
Co-authored-by: Michael Chou <michael.chou@scale.com>
Co-authored-by: Max Parke <max.parke@scale.com>
Co-authored-by: Matteo Librizzi <matteo.librizzi@scale.com>
Co-authored-by: Stas Moreinis <smoreinis@gmail.com>
Co-authored-by: James Cardenas <james.cardenas@scale.com>
Co-authored-by: Jerome Romualdez <jerome.romualdez@scale.com>
Co-authored-by: Nitesh Dhanpal <NiteshDhanpal@users.noreply.github.com>
Co-authored-by: Vijay Kalmath <158184866+vkalmathscale@users.noreply.github.com>
Co-authored-by: OpenAI <openai@example.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Levi Lentz <levi.lentz@scale.com>
Co-authored-by: Deepthi Rao <deepthi.rao@scale.com>
Co-authored-by: Javed Shaik <javed.shaik@scale.com>
Co-authored-by: Cynthia Wang <cynthia.wang@scale.com>
Co-authored-by: Rishav Chakravarti <rishav.chakravarti@scale.com>
Co-authored-by: stlc-bot <stlc-bot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants