Model Inference / Latency/Cost

Kind: tension

Record: tension:latency-cost-model-inference

Canonical: Schema

"model-inference" (correctness-core layer) is traded against "Latency/Cost" (performance-core layer) — a principle cannot be scope-separated from a quality, metric, or cost it competes with; resolve by measuring "Latency/Cost" and choosing an explicit operating point.

Listed in Tensions, after Model Evaluation / Metric Completeness and before Retrieval-Augmented Generation (RAG) / Retrieval Quality/Latency.

In tension with

Mechanism

Linked from