Skip to content

feat: keep several prompt cache entries (--prompt-cache-entries) and evict under memory pressure - #210

Merged
solderzzc merged 2 commits into
SharpAI:mainfrom
CodeAndCanvas728:pr/prompt-cache-lru
Oct 5, 2026
Merged

solderzzc merged 2 commits into
SharpAI:mainfrom
CodeAndCanvas728:pr/prompt-cache-lru

Conversation

@CodeAndCanvas728

@CodeAndCanvas728 CodeAndCanvas728 commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Problem

PromptCache holds a single entry. Any request that doesn't share the cached prefix replaces it, so a client that interleaves conversations loses its cache on every switch. Typical cases: an agent that fires a side request (title generation, a summariser, a sub-agent with a different prompt) between two turns of the main conversation, or several chat sessions against one server. Each switch costs a full re-prefill of the long conversation.

Change

  • PromptCache becomes an LRU of up to --prompt-cache-entries N entries (default 1, so default behaviour and memory use are unchanged).
  • save replaces any entry that is a prefix of the new one, so a single growing conversation still occupies one slot.
  • restore (trimmable models) picks the entry with the longest usable match; the per-entry safety checks (excess vs. cached length, wrapped ring buffers) are unchanged, just factored into usableMatch.
  • restoreExactPrefix (hybrid recurrent+attention models) picks the longest exact-prefix entry.
  • A DispatchSource memory-pressure source evicts entries: all but the most recent on .warning, all on .critical. Cached KV is the largest evictable allocation, so it is returned before the OS has to compress or kill. PromptCache.evict(keepMostRecent:) is also the missing way to drop entries while idle.

No byte budget in this PR; that can follow if pressure-based eviction proves too late in practice.

Tests

PromptCacheLRUTests (7 tests): unrelated prompts both hit with 2 entries and not with 1; a hit refreshes recency and the least recent is dropped; an extension replaces its prefix entry; longest match wins; hybrid exact-prefix picks the longest; both evict modes. Existing cache tests pass unchanged. swift test --filter SwiftLMTests: 230 pass.

Measurements

Qwen3.5 35B-A3B (qwen3_5_moe, hybrid GatedDelta + attention, 4-bit), loaded text-only (--no-vision) with --stream-experts --ssd-prefetch --prefill-size 2048, on a 16 GB M2. Sequence: turn 1 (12.7k tokens), an unrelated 2.7k-token request, then turn 2 (14.7k tokens, extending turn 1). Greedy, max_tokens 16.

turn 1 side request turn 2 turn 2 cache
--prompt-cache-entries 4 145 s 37 s 29.8 s hit, 12,681 / 14,657 tokens reused
default (1 entry) 148 s 42 s 209.5 s miss (evicted by the side request)

Memory demand stayed about 8.4–8.7 GB in both runs. The pressure handler fired once (warning level) during the 4-entry run under swap on this machine and kept the newest entry.

Note

On hybrid models, --ctx-size currently disables the hybrid cache (attention layers become RotatingKVCache, which hybridCacheBoundary rejects); that is fixed separately in #207. The runs above did not use --ctx-size.

🤖 Generated with Claude Code

…evict them under memory pressure

PromptCache held a single entry, so a side request between two turns of the
same conversation evicted that conversation's prefix. It is now an LRU of up
to --prompt-cache-entries (default 1, i.e. unchanged). Saving an extension of
an entry replaces it, restore picks the longest usable match, and a memory
pressure source drops entries (all but the newest on warning, all on critical).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ry on eviction, document and test the flag

Follow-up to the multi-entry prompt cache (--prompt-cache-entries).

Stale entries. save() only replaced an entry the new prompt extends exactly,
but the prompt saved for turn N is often not an exact prefix of turn N+1:
templates that re-render history drop or rewrite the generation prompt it
ended with, and edits or regenerations change the last message. The stale
entry then held a slot until it aged out, pushing another conversation out
of the cache. save() now also drops an entry that diverges from the new
prompt in at most PromptCache.supersedeSlack (16) trailing tokens and shares
more than it loses. The rule is stateless, so it needs no conversation
identity. For a cache that can be rewound to any shorter prefix replacing is
bounded: any prompt that could have used the old entry loses at most 16
tokens of reuse. Entries that merely start alike (a shared BOS or system
prompt followed by more than 16 tokens of their own) are kept apart.

Two kinds of state cannot be rewound that way and keep the strict exact-prefix
rule: recurrent (hybrid) snapshots, which are only valid at exactly
tokens.count, and sliding-window ring buffers that have already dropped tokens
(offset beyond the window), which usableMatch rejects for any rewind past one
token. Near-prefix replacement there would turn the old entry's full hit into a
miss (two conversations sharing a long system prompt, each with a short first
message). save() asks the same question usableMatch does, through one shared
helper, so the two cannot drift apart. A ring that has not wrapped yet is
rewound like any attention cache and gets the near-prefix rule. An empty prompt
is no longer saved: nothing can restore it and it only took a slot.

Memory pressure. evict(keepMostRecent:) now returns how many entries it
dropped. The handler logs only when something was dropped and calls
Memory.clearCache() then, from the Task that awaits the eviction rather than
from the GCD handler thread, so the freed KV buffers actually go back to the
system instead of staying in the MLX allocator pool. The comment and the flag
help now state the real behavior: a warning keeps the most recent entry, a
critical event drops all of them, so with the default of one entry only a
critical event empties the cache.

Flag. Document --prompt-cache-entries in the README option table and print the
effective value in the startup Config line. A value below 1 prints a warning at
startup and is raised to 1 (as --gpu-layers does for an invalid value), so the
clamp is visible before any model is loaded.

Tests. Cover the supersede rule (within the slack, one token beyond it, tiny
entries, shared system prompt, ping-pong between two conversations, several
dominated entries at once, exact duplicates, empty prompts, hybrid snapshots),
wrapped and unwrapped ring buffers (shared system prompt with short tails, mixed
full-attention and sliding-window layers, rewind boundaries), the
longest-usable-match pick regardless of recency for restore() and
restoreExactPrefix(), rejection of an unusable candidate falling through to
another entry, tail divergence, hybrid recency refresh, evict() return values,
the capacity floor and the flag parse.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@solderzzc

Copy link
Copy Markdown
Member

Thanks for this. I reviewed it and pushed one follow-up commit (4588b8b) on top of your branch rather than asking for another round trip. AI-assisted (Claude Code); tests run locally on origin/main + this branch (SwiftLMTests 262/0).

  • Stale entries: save() now also drops an entry that diverges from the new prompt in at most 16 trailing tokens and shares more than it loses (re-rendered generation prompts, edited last message), so a conversation no longer pushes another one out of the cache. Hybrid snapshots and wrapped sliding-window rings keep the strict exact-prefix rule, because they can't be rewound.
  • Memory pressure: evict() returns the dropped count; the handler logs only when something was dropped and calls Memory.clearCache() so the freed KV buffers actually return to the system. Note that with the default of 1, a critical event now empties the cache; code comment, help and README say so.
  • Flag: README row, effective value in the startup Config line, values < 1 warn and are raised to 1.
  • Tests: 7 → 36 (supersede rule, ping-pong, shared system prompt, ring buffers, longest-usable-match, rejected candidates, evict counts, flag parse).

Not changed: no byte budget, a short prompt within 16 tokens of a longer one replaces it, and no end-to-end run on a sliding-window model. The PR description is out of date ("unchanged", "7 tests") and could use a refresh before merge.

@solderzzc
solderzzc merged commit 8469a17 into SharpAI:main Oct 5, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants