CLOUDFLARE · LLM INFRASTRUCTURE

LLM Semantic Cache

High-confidence answer reuse before LLM inference.

Worker online
QUESTION
↓
L1 KV EXACT CACHE
↓
BGE-SMALL EMBEDDING
↓
VECTORIZE + SAFETY GUARDS
↓
HIGH-CONFIDENCE MATCH?
YES → CACHED ANSWER
NO → LLM FALLBACK
EXACT CACHE
454 ms
Average · 10 exact repeats
SEMANTIC RECALL
90%
9 / 10 paraphrases
EXACT CACHE COST
~132×
Modeled marginal reduction
FALSE-HIT RATE
0%
0 / 10 hard negatives

Three-tier inference path

01
L1 · Exact KV

Instant reuse for identical normalized questions.

02
L2 · Semantic cache

BGE-small + Vectorize + subject/number/operation guards.

03
L3 · LLM fallback

Generate and store a new answer when confidence is insufficient.