Inference optimization mini-benchmark

Qwen2.5-1.5B-Instruct on llama.cpp (Metal), measuring the production KPIs (TTFT, TPOT, ITL, tokens/sec) across quantization, context length, prefix caching and model load time.

Decode throughput75.8FP16119.0Q8_0176.5Q4_K_Mtokens/sec
Model size (drives cold start + HBM footprint)3.56FP161.89Q8_01.12Q4_K_MGB
Model load time (cold start proxy)0.63FP161.47Q8_00.64Q4_K_Mseconds
Quality gate (exact-match answers)6FP166Q8_06Q4_K_Mcorrect
TTFT vs prompt length (prefill scaling)100w500w2000wFP16Q8_0Q4_K_Mseconds
Prefix-cache TTFT speedup (2000-word prompt)26.4xFP1636.5xQ8_048.2xQ4_K_Mx faster