Model Inference / Latency/Cost
Kind: tension
Record: tension:latency-cost-model-inference
Canonical: Schema
"model-inference" (correctness-core layer) is traded against "Latency/Cost" (performance-core layer) — a principle cannot be scope-separated from a quality, metric, or cost it competes with; resolve by measuring "Latency/Cost" and choosing an explicit operating point.
Listed in Tensions, after Model Evaluation / Metric Completeness and before Retrieval-Augmented Generation (RAG) / Retrieval Quality/Latency.