audio is transcribed thus no need to be sent, but other blob(e.g. image) should still be sent.
Co-authored-by: Xiang (Sean) Zhou <seanzhougoogle@google.com>
PiperOrigin-RevId: 856422986
The Gemini API may not always send an explicit transcription finished signal. This change ensures that any buffered input or output transcription text is yielded as a finished transcription when a turn is completed, generation is complete, or the session is interrupted.
Also, refined the check for `event.partial` in runners.py to be more explicit.
Co-authored-by: Hangfei Lin <hangfei@google.com>
PiperOrigin-RevId: 839008606
Merge https://github.com/google/adk-python/pull/3352
Replace send() method with send_realtime_input() to fix DeprecationWarning
### Link to Issue or Description of Change
**1. Link to an existing issue (if applicable):**
- Closes: #2393
### Testing Plan
Have run unit test for test_live_request_queue.py
**Unit Tests:**
- [ ] I have added or updated unit tests for my change.
- [x] All unit tests pass locally.
Summary :
tests/unittests/agents/test_live_request_queue.py::test_send_realtime PASSED [100%]
**Manual End-to-End (E2E) Tests:**
### Checklist
- [x] I have read the [CONTRIBUTING.md](https://github.com/google/adk-python/blob/main/CONTRIBUTING.md) document.
- [x] I have performed a self-review of my own code.
- [x] I have commented my code, particularly in hard-to-understand areas.
- [ ] I have added tests that prove my fix is effective or that my feature works.
- [x] New and existing unit tests pass locally with my changes.
- [x] I have manually tested my changes end-to-end.
- [ ] Any dependent changes have been merged and published in downstream modules.
Tested both API variants.
Co-authored-by: Hangfei Lin <hangfei@google.com>
COPYBARA_INTEGRATE_REVIEW=https://github.com/google/adk-python/pull/3352 from SanjaySiddharth:fix-dep-session.send b23b641a63ff20d1f2dbd9abd519f8facf0aab77
PiperOrigin-RevId: 830182844
In `gemini_llm_connection.py`, accumulate partial transcription texts and emit `LlmResponse` with `partial=True` for each chunk. When the transcription is marked as `finished`, emit a final `LlmResponse` with the full accumulated text and `partial=False`.
In `runners.py`, modify `_should_append_to_history` to only add transcription events to the history when they are fully finished, preventing partial transcriptions from being added.
Co-authored-by: Hangfei Lin <hangfei@google.com>
PiperOrigin-RevId: 829029715
Merge https://github.com/google/adk-python/pull/3324
**Problem:**
ADK seems to not pass output/input transcription `finished` flag from the Gemini.
**Solution:**
Relaxation of checking conditions of valid llm_response for input/output transcription.
### Testing Plan
Unit test `test_receive_transcript_finished` checks if input/output transcription message with no text but `finished` flag is received.
**Unit Tests:**
- [x] I have added or updated unit tests for my change.
- [x] All unit tests pass locally.
<img width="785" height="373" alt="image" src="https://github.com/user-attachments/assets/6d870e9f-1372-4808-91a9-38578c1b3729" />
**Manual End-to-End (E2E) Tests:**
Configure Gemini Agent to produce input & output transcriptions. Observe incoming transcript messages - to see if the finished flag appears at the end agents & users statement.
### Checklist
- [x] I have read the [CONTRIBUTING.md](https://github.com/google/adk-python/blob/main/CONTRIBUTING.md) document.
- [x] I have performed a self-review of my own code.
- [x] I have commented my code, particularly in hard-to-understand areas.
- [x] I have added tests that prove my fix is effective or that my feature works.
- [x] New and existing unit tests pass locally with my changes.
- [x] I have manually tested my changes end-to-end.
- [x] Any dependent changes have been merged and published in downstream modules.
### Additional context
The mentioned `finished` flag can be obtained in native Gemini APIs like via Websocket but not via ADK.
Co-authored-by: Hangfei Lin <hangfei@google.com>
COPYBARA_INTEGRATE_REVIEW=https://github.com/google/adk-python/pull/3324 from ChrisQlasty:fix/transcript_finish e7b8e5e5f5fac1d2dfd495f2dbadb19ed4328b0c
PiperOrigin-RevId: 828987996
Populate the usage_metadata field for live events with the metadata provided by the Gemini live API.
Co-authored-by: Kathy Wu <wukathy@google.com>
PiperOrigin-RevId: 828124232
Merge https://github.com/google/adk-python/pull/981
issue: https://github.com/google/adk-python/issues/982
This pull request introduces a new configuration option, `realtime_input_config`, to the `RunConfig` class.
**Reason for this change:**
Currently, there is no direct way to configure real-time audio input behaviors, such as Voice Activity Detection (VAD), for live agents through the `RunConfig`. The Gemini API documentation (specifically [Configure automatic VAD](https://ai.google.dev/gemini-api/docs/live#configure-automatic-vad)) outlines parameters for VAD that users may want to customize.
This change enables users to pass these real-time input configurations, providing more granular control over the audio input for live agents.
**Changes made:**
- Added a new optional field `realtime_input_config: Optional[types.RealtimeInputConfig]` to the `RunConfig` class.
- The docstring for `realtime_input_config` has been added to explain its purpose.
**Example Usage (Conceptual):**
While the specific structure of `types.RealtimeInputConfig` would define the exact parameters, a user might configure it like this:
```python
# (Assuming types.RealtimeInputConfig and types.VadConfig are defined elsewhere)
# import your_project.types as types
run_config = RunConfig(
# ... other configurations ...
realtime_input_config=types.RealtimeInputConfig(
automatic_activity_detection =types.AutomaticActivityDetection(
# VAD specific parameters like sensitivity, endpoint_duration_millis etc.
# based on https://ai.google.dev/gemini-api/docs/live#configure-automatic-vad
)
# Potentially other real-time input settings could be added here in the future
)
)
COPYBARA_INTEGRATE_REVIEW=https://github.com/google/adk-python/pull/981 from ammmr:patch-add-realtime-input-config b2e17fbf5742d264029ad49bf632422b5c5b1e0a
PiperOrigin-RevId: 770797640