CLOUDFLARE WORKER · VECTORIZE · KV

Cache Statistics

Evaluation results, runtime performance, and cost modeling.

PRODUCTION THRESHOLD
0.88
With semantic guards

Live system statistics

These values represent accumulated requests handled by the deployed Worker, including development and testing traffic.

CACHE HIT RATE
—
TOTAL REQUESTS
—
EXACT HITS · L1
—
SEMANTIC HITS · L2
—
LLM FALLBACKS
—
LLM CALLS AVOIDED
—
Runtime traffic is shown separately from the controlled evaluation above so development/testing requests do not affect the benchmark results.

Response path performance

454 ms
Exact cache
L1 · KV lookup
3.40 s
Semantic cache
Embedding + Vectorize
12.67 s
LLM fallback
Embedding + LLM generation
Values are averages from the final 40-case end-to-end evaluation.

Modeled marginal cache cost

EXACT CACHE
~132× lower
Compared with an LLM request
SEMANTIC CACHE
~6.6× lower
Compared with an LLM request
LLM MISS COST
$0.000066
Modeled average
EXACT HIT COST
$0.00000050
Modeled KV read

Cost figures are modeled estimates using API/list-price assumptions, not direct Cloudflare billing measurements. Exact and semantic cache paths include their respective cache/embedding operations.

Why the production threshold is 0.88

0.88
Production similarity threshold
Used together with semantic guards
The threshold was evaluated against 66 labeled question pairs. At 0.88 with the semantic guards enabled, the evaluation produced a 72% hit rate with 0% false hits on that dataset.
Subject guard Prevents unrelated subject matches.
Number guard Protects questions where numeric values change.
Operation guard Distinguishes operations such as area and perimeter.
Comparison guard Protects comparison questions from mismatched entities.

40-case end-to-end test

10 / 10
Exact repeats
100% exact hits
9 / 10
Semantic paraphrases
90% semantic recall
0 / 10
Hard negatives
0% false hits
40
Total test cases
10 + 10 + 10 + 10