Skip to main content

LlmStreamEvent

Enum LlmStreamEvent 

Source
pub enum LlmStreamEvent {
    TextDelta {
        content: String,
    },
    ReasoningDelta {
        content: String,
    },
    ToolCallDelta {
        index: usize,
        id: Option<String>,
        name: Option<String>,
        arguments: Option<String>,
    },
    PromptProgress {
        processed: u32,
        total: u32,
        cached: u32,
        time_ms: u64,
    },
    Usage {
        prompt_tokens: u32,
        completion_tokens: u32,
        total_tokens: u32,
        cached_tokens: Option<u32>,
    },
    Done {
        finish_reason: String,
    },
    UpstreamError {
        message: String,
        error_type: String,
        code: String,
    },
    NormalizationError {
        kind: NormalizationErrorKind,
        raw: String,
    },
}
Expand description

A single event produced by a streaming LLM response.

These low-level events are the currency of crate::ports::LlmCompletionPort; they are parsed by adapter crates from raw SSE frames and handed to gglib-agent’s stream collector, which:

  • Forwards TextDelta items directly to the caller’s AgentEvent channel so text appears in real time.
  • Accumulates ToolCallDelta fragments until the stream ends, then assembles them into ToolCall values.
  • Waits for Done before triggering tool execution.

Variants§

§

TextDelta

An incremental text fragment from the model’s response.

Fields

§content: String

The new text fragment (append to the running content buffer).

§

ReasoningDelta

An incremental reasoning/thinking fragment (CoT tokens).

Produced by reasoning-capable models (e.g. DeepSeek R1, QwQ) when llama-server is started with --reasoning-format deepseek. The runtime adapter maps delta["reasoning_content"] frames to this variant; the stream collector forwards them as AgentEvent::ReasoningDelta and accumulates them in a separate buffer that is never sent back to the LLM as context.

Fields

§content: String

The new reasoning fragment (append to the current reasoning buffer).

§

ToolCallDelta

An incremental fragment of a tool-call request.

The adapter crate streams these before the model has finished generating the full arguments JSON. The stream collector accumulates all deltas for a given index into a single ToolCall.

Fields

§index: usize

Zero-based index of the tool call within the current response.

§id: Option<String>

Call identifier (only present in the first delta for this index).

§name: Option<String>

Tool name (only present in the first delta for this index).

§arguments: Option<String>

Partial arguments JSON string fragment (accumulate with push_str).

§

PromptProgress

Prompt-processing progress from llama-server.

Emitted when the request includes return_progress: true. These frames arrive during the pre-fill phase (before any TextDelta), giving real-time visibility into how far along token ingestion is.

Fields

§processed: u32

Number of tokens processed so far.

§total: u32

Total number of tokens in the prompt.

§cached: u32

Number of tokens served from KV cache (already processed).

§time_ms: u64

Elapsed wall-clock time in milliseconds since processing began.

§

Usage

Token usage totals for the completed response.

Emitted when the request includes stream_options.include_usage: true (see gglib_proxy’s inject_streaming_body_overrides). Per the OpenAI streaming convention this arrives as a single trailing chunk with an empty/absent choices array, immediately before the [DONE] sentinel — never mixed with TextDelta or ToolCallDelta events. Consumers that care about real token counts (e.g. clients feeding a context-window UI widget) read this event; everyone else can ignore it.

Fields

§prompt_tokens: u32

Number of tokens in the prompt.

§completion_tokens: u32

Number of tokens generated in the completion.

§total_tokens: u32

Total tokens (prompt_tokens + completion_tokens).

§cached_tokens: Option<u32>

How many of prompt_tokens were served from the KV cache rather than re-processed, from usage.prompt_tokens_details.cached_tokens (llama.cpp’s n_prompt_tokens_cache; the same figure it reports as timings.cache_n).

None when the upstream omits the field — an older llama.cpp, or any other OpenAI-compatible server that doesn’t report it. Absent and zero are deliberately distinct: zero means “nothing was reused”, which is a real measurement, while None means “not reported” and must not be counted as a cache miss.

§

Done

Signals the end of the stream.

Every conforming stream must end with exactly one Done item.

Fields

§finish_reason: String

The OpenAI-compatible finish reason (e.g. "stop", "tool_calls", "length").

§

UpstreamError

An upstream error reported inline, mid-stream.

Some OpenAI-compatible servers (including llama.cpp) can emit a bare {"error": {...}} frame in the middle of an otherwise-successful SSE response — e.g. a context-length overflow discovered only once generation is underway — rather than failing the initial HTTP response. Without special handling this frame carries no choices array and would otherwise be silently discarded by parsers that treat “no choices” as an empty keepalive/heartbeat frame.

error_type and code default to "server_error"/"upstream_error" respectively when the upstream frame omits them, mirroring the existing convention in gglib_proxy::forward’s pre-flight upstream error handling.

Fields

§message: String

Human-readable error message.

§error_type: String

Upstream error.type (or "server_error" if absent).

§code: String

Upstream error.code (or "upstream_error" if absent).

§

NormalizationError

A non-fatal normalization issue surfaced by the crate::normalize layer.

Emitted when a dialect-specific parser detects malformed markup (e.g. a Qwen <tool_call> whose body is not valid JSON, or a tag that the stream ended without closing). The stream is not aborted; the offending bytes are simply discarded or surfaced via this event so consumers can log diagnostics.

Fields

§kind: NormalizationErrorKind

Structured detail about what went wrong.

§raw: String

Short, human-readable excerpt of the offending input.

Trait Implementations§

Source§

impl Clone for LlmStreamEvent

Source§

fn clone(&self) -> LlmStreamEvent

Returns a duplicate of the value. Read more
1.0.0 (const: unstable) · Source§

fn clone_from(&mut self, source: &Self)

Performs copy-assignment from source. Read more
Source§

impl Debug for LlmStreamEvent

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result

Formats the value using the given formatter. Read more
Source§

impl PartialEq for LlmStreamEvent

Source§

fn eq(&self, other: &LlmStreamEvent) -> bool

Tests for self and other values to be equal, and is used by ==.
1.0.0 (const: unstable) · Source§

fn ne(&self, other: &Rhs) -> bool

Tests for !=. The default implementation is almost always sufficient, and should not be overridden without very good reason.
Source§

impl Eq for LlmStreamEvent

Source§

impl StructuralPartialEq for LlmStreamEvent

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<T> CloneToUninit for T
where T: Clone,

Source§

unsafe fn clone_to_uninit(&self, dest: *mut u8)

🔬This is a nightly-only experimental API. (clone_to_uninit)
Performs copy-assignment from self to dest. Read more
Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

§

impl<T> Instrument for T

§

fn instrument(self, span: Span) -> Instrumented<Self>

Instruments this type with the provided [Span], returning an Instrumented wrapper. Read more
§

fn in_current_span(self) -> Instrumented<Self>

Instruments this type with the current Span, returning an Instrumented wrapper. Read more
Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> Same for T

Source§

type Output = T

Should always be Self
Source§

impl<T> ToOwned for T
where T: Clone,

Source§

type Owned = T

The resulting type after obtaining ownership.
Source§

fn to_owned(&self) -> T

Creates owned data from borrowed data, usually by cloning. Read more
Source§

fn clone_into(&self, target: &mut T)

Uses borrowed data to replace owned data, usually by cloning. Read more
Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = Infallible

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, <T as TryFrom<U>>::Error>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
§

impl<T> WithSubscriber for T

§

fn with_subscriber<S>(self, subscriber: S) -> WithDispatch<Self>
where S: Into<Dispatch>,

Attaches the provided Subscriber to this type, returning a [WithDispatch] wrapper. Read more
§

fn with_current_subscriber(self) -> WithDispatch<Self>

Attaches the current default Subscriber to this type, returning a [WithDispatch] wrapper. Read more