Conversation
Save actual session state before and after inference, and pass separate copies to each metric through an optional context method. Keep the original evaluator interface for existing sync and async metrics. Related: google#4532
|
Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA). View this failed invocation of the CLA check for more information. For the most up to date status, view the checks section at the bottom of the pull request. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Link to Issue or Description of Change
Related: #4532
Custom evaluators can inspect the conversation, but they cannot read session state. This adds
EvaluationContextand anevaluate_with_contextmethod for metrics that need it. The default method callsevaluate_invocationswith its original arguments, so existing sync and async evaluators need no changes.Local inference saves the actual state before the first turn and after the run ends. Each metric receives its own copy, plus the expected final state from the eval case. Later session updates or deletion do not change these snapshots. Old inference results without snapshots keep
Nonefor the actual state fields.Testing Plan
Unit Tests:
pytest tests/unittests/evaluation -q: 960 passed on Python 3.13.15 at the final commit. The 22 new cases cover old and new sync/async evaluators, regular and live inference, reused sessions, state isolation, and saved results after a session update or deletion.The full suite ran through tox on Python 3.10–3.14, with two pytest workers per version. A 60-second per-test timeout let the suite continue past stalled tests.
The evaluation tests pass on all five versions: 960 per version. The full suite does not pass. Two LiveKit tests exceed the timeout, and the import allowlist check fails. All three failures also occur on unmodified
main(9633ab9f) in the same Python 3.14 environment. Some CLI deploy tests fail in the full 3.10 and 3.14 runs, but all 68 tests in that file pass in the isolated baseline check. Their cause is not confirmed.The checks for all changed files pass. The package build and a clean wheel import also pass. Mypy reports the same seven errors on this change and the original base revision, with no new errors.
Manual End-to-End (E2E) Tests:
A local
BaseAgentadds a book to a cart through the real Runner. The test saves the inference result as JSON, deletes the session, then evaluates the saved result with a custom state metric. It needs no model API or credentials.The metric receives an empty cart as the initial state and a cart with one book as the final state. The final state matches the expected state. The result is
PASSEDwith a score of1.0.The check uses
InMemorySessionService,LocalEvalService, and an isolated metric registry. The same setup and public inference/evaluation path are intest_evaluate_uses_snapshots_after_session_change. Run it from the repo after the normal development setup:Console output from the manual run:
Checklist
Additional context
Companion docs: google/adk-docs#2289. The strict docs build and local page check pass.
This is a draft for API feedback. Does the separate context method fit the evaluator interface, and are the snapshot boundaries appropriate? The full suite failures are listed above, so this is not ready to land yet.