pub trait CacheMetricsSink: Send + Sync {
// Required method
fn record(&self, prompt_tokens: u32, cached_tokens: Option<u32>);
}Expand description
A sink for per-request prompt-cache reuse figures.
The recording site — the in-process LLM adapter’s response stream — writes
each completed request’s usage here without knowing where it lands. The
proxy’s dashboard implements it with an in-memory
CacheMetricsStore; a process
with no dashboard passes no sink at all, so recording becomes a no-op. A
future cross-process reporter — a gglib chat run posting its reuse to a
running proxy — would be another implementation behind this same seam,
needing no change to the adapter.
cached_tokens keeps the Option<u32> absent-vs-zero distinction: None
means the upstream didn’t report the field, Some(0) means a real full
re-prefill. Implementations must not collapse the two.
Required Methods§
Implementors§
impl CacheMetricsSink for CacheMetricsStore
The in-process store is the default CacheMetricsSink: the adapter
records straight into the atomics the dashboard snapshots.