Fix streaming + tool calling interleaving in OpenAI provider #2

Open
opened 2026-09-18 01:12:20 -04:00 by assistant · 0 comments

Summary

The OpenAI provider's streaming loop incorrectly handles interleaved tool_call and content deltas, causing premature termination and missed chunks.

Root Cause

The streaming loop uses chunk.done to detect completion, but OpenAI streams at the choice level:

  • finish_reason indicates when the LLM is done
  • tool_calls and content deltas arrive independently and out of order

Current code:

  1. Waits for entire tool_calls array before executing tools
  2. Sets terminal = true on wrong condition (never properly set via finish_reason)
  3. May stop consuming chunks after tool execution begins

Required Changes

  • Detect proper termination via finish_reason === 'stop' | 'tool_calls'
  • Accumulate tool_call deltas by index until all complete
  • Execute tools only after all deltas for a given tool are received
  • Continue consuming content deltas after tools execute
  • Reset state correctly when switching between tool_calls and content

Acceptance Criteria

  • Tools execute reliably even with interleaved deltas
  • Content following tools is captured correctly
  • Stream terminates at proper finish_reason, not on missing chunk.done
  • No chunks are missed or duplicated during tool execution cycles

Related

  • #1 Response streaming + tool calling ends prematurely or without follow-up text
## Summary The OpenAI provider's streaming loop incorrectly handles interleaved tool_call and content deltas, causing premature termination and missed chunks. ## Root Cause The streaming loop uses `chunk.done` to detect completion, but OpenAI streams at the choice level: - `finish_reason` indicates when the LLM is done - `tool_calls` and `content` deltas arrive independently and out of order Current code: 1. Waits for entire `tool_calls` array before executing tools 2. Sets `terminal = true` on wrong condition (never properly set via finish_reason) 3. May stop consuming chunks after tool execution begins ## Required Changes - Detect proper termination via `finish_reason === 'stop' | 'tool_calls'` - Accumulate tool_call deltas by index until all complete - Execute tools only after all deltas for a given tool are received - Continue consuming content deltas after tools execute - Reset state correctly when switching between tool_calls and content ## Acceptance Criteria - Tools execute reliably even with interleaved deltas - Content following tools is captured correctly - Stream terminates at proper finish_reason, not on missing `chunk.done` - No chunks are missed or duplicated during tool execution cycles ## Related - #1 Response streaming + tool calling ends prematurely or without follow-up text
ztimson added reference fix/streaming-tool-interleaving 2026-09-18 14:01:51 -04:00
Sign in to join this conversation.
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: ztimson/ai-utils#2