# Architecture principles whose category is Scalability / Performance / Optimization

> 28 records

This index as JSON: https://banes-lab.com/json/api/facets/architecture/category/scalability-performance-optimization

## Entries

- [Scalability](https://banes-lab.com/records/architecture/scalability.md): The degree to which a system keeps its throughput and latency as load grows, by adding resources.
- [Horizontal Scaling](https://banes-lab.com/records/architecture/horizontal-scaling.md): A technique for adding capacity by running more stateless instances of a service behind a load balancer.
- [Vertical Scaling](https://banes-lab.com/records/architecture/vertical-scaling.md): A technique for adding capacity by giving one instance more CPU, memory or I/O.
- [Elasticity](https://banes-lab.com/records/architecture/elasticity.md): The degree to which a system's provisioned capacity follows demand up and down automatically.
- [Load Balancing](https://banes-lab.com/records/architecture/load-balancing.md): A mechanism that spreads incoming requests across healthy instances of a service.
- [Sharding](https://banes-lab.com/records/architecture/sharding.md): A design pattern that splits a dataset across separate stores by a shard key, and routes each request to the shard that holds its key.
- [Partitioning](https://banes-lab.com/records/architecture/partitioning.md): A technique for dividing data or work into independent partitions by a key, so each can be processed in parallel.
- [Caching](https://banes-lab.com/records/architecture/caching.md): A design pattern that stores the result of an expensive read or computation under a key derived from every input it depends on, and serves it again while that key matches.
- [Statelessness](https://banes-lab.com/records/architecture/statelessness.md): A design rule that a handler keeps no request state between calls, so any instance can serve any request.
- [Concurrency](https://banes-lab.com/records/architecture/concurrency.md): A conceptual representation of several tasks in progress over overlapping time, with their access to shared state coordinated.
- [Parallelism](https://banes-lab.com/records/architecture/parallelism.md): A technique for running independent units of work at the same time on several cores or workers.
- [Throughput](https://banes-lab.com/records/architecture/throughput.md): The rate at which a system completes requests, messages or items, counted per unit of time.
- [Latency](https://banes-lab.com/records/architecture/latency.md): A measure of the time between a request and its response, usually reported at percentiles.
- [Performance Engineering](https://banes-lab.com/records/architecture/performance-engineering.md): The practice of setting performance budgets, measuring against a representative workload and changing code only where a measurement points.
- [Algorithmic Efficiency](https://banes-lab.com/records/architecture/algorithmic-efficiency.md): A design rule that an algorithm and its data structures are chosen for how their cost grows with input size.
- [Time Complexity](https://banes-lab.com/records/architecture/time-complexity.md): A measure of how an algorithm's running time grows as its input grows.
- [Space Complexity](https://banes-lab.com/records/architecture/space-complexity.md): A measure of how an algorithm's memory use grows as its input grows.
- [Big O Notation](https://banes-lab.com/records/architecture/big-o-notation.md): A method for classifying an algorithm by the upper bound on how its cost grows with input size, ignoring constant factors.
- [Optimization](https://banes-lab.com/records/architecture/optimization.md): The activity of changing code or configuration to reduce the cost of a bottleneck that a measurement has located.
- [Profiling](https://banes-lab.com/records/architecture/profiling.md): A technique for measuring where a running program spends its time and memory under a representative workload.
- [Benchmarking](https://banes-lab.com/records/architecture/benchmarking.md): The activity of timing a fixed workload repeatedly in a controlled environment, so results can be compared across changes.
- [Bottleneck Analysis](https://banes-lab.com/records/architecture/bottleneck-analysis.md): The activity of finding the stage that limits a system's overall throughput or latency, from traces and profiles.
- [Resource Utilization](https://banes-lab.com/records/architecture/resource-utilization.md): A measure of how much of the provisioned CPU, memory, I/O and network capacity is in use.
- [Rate Limiting](https://banes-lab.com/records/architecture/rate-limiting.md): A mechanism that caps how many requests a caller can make in a time window and rejects the excess.
- [Memory Efficiency](https://banes-lab.com/records/architecture/memory-efficiency.md): The degree to which a process handles its input with bounded memory, by streaming or chunking instead of loading it whole.
- [CDN / Edge Caching](https://banes-lab.com/records/architecture/cdn-edge-caching.md): A mechanism that serves cacheable content from servers near the requester, keyed by a fingerprint of the content.
- [Read Replica](https://banes-lab.com/records/architecture/read-replica.md): A technique for routing read queries to replicated copies of a database, so the primary handles only writes.
- [Queuing Theory](https://banes-lab.com/records/architecture/queuing-theory.md): A conceptual representation of a system as queues with arrival and service rates, used to predict waiting time and size capacity.
