# Self-Healing / Recovery / Deployment Safety

> Every principle in this category is listed as a record.

Page: Ontology · Principles
Canonical: https://banes-lab.com/ontology#architecture-category-self-healing-recovery-deployment-safety

Listed in [Ontology · Principles](https://banes-lab.com/api/pages/ontology/principles.md), after [Plugin / Extensibility / IoC](https://banes-lab.com/ontology/principles/architecture-category-plugin-extensibility-ioc.md) and before [Schema / Canonical Data / Semantics](https://banes-lab.com/ontology/principles/architecture-category-schema-canonical-data-semantics.md).

Every principle in this category is listed as a record. Each record carries its kind, its severity, the scopes it applies at and the layer it lives in, then the edge relations that join it to other records, the records that point back at it, the contracts that answer to it and the tensions it takes part in. The descriptors say how it is violated, detected, measured, repaired and enforced. Where the record carries one, an exemplar shows the shape before and after the principle is applied.

Relations diagram

The relations inside this category.

```mermaid
flowchart LR
n_self_healing_architecture["Self-Healing Architecture"]
n_health_checks["Health Checks"]
n_failover["Failover"]
n_redundancy["Redundancy"]
n_replication["Replication"]
n_auto_scaling["Auto-Scaling"]
n_auto_remediation["Auto-Remediation"]
n_rollback["Rollback"]
n_blue_green_deployment["Blue-Green Deployment"]
n_canary_deployment["Canary Deployment"]
n_chaos_engineering["Chaos Engineering"]
n_graceful_shutdown["Graceful Shutdown"]
n_raid_redundancy["RAID Redundancy"]
n_self_healing_architecture --> n_health_checks
n_self_healing_architecture --> n_auto_remediation
n_health_checks --> n_self_healing_architecture
n_failover --> n_redundancy
n_redundancy --> n_failover
n_replication --> n_failover
n_auto_remediation --> n_self_healing_architecture
n_blue_green_deployment --> n_rollback
n_chaos_engineering --> n_self_healing_architecture
n_raid_redundancy --> n_redundancy
```

### Self-Healing Architecture

- Kind: [capability](https://banes-lab.com/records/kind/capability.md)
- Category: [Self-Healing / Recovery / Deployment Safety](https://banes-lab.com/ontology/principles/architecture-category-self-healing-recovery-deployment-safety.md)
- Severity: [contextual](https://banes-lab.com/records/vocabulary/severity-contextual.md)
- Scope: system, infrastructure, runtime
- Aliases: Self-Healing
- Layer: [Correctness Core](https://banes-lab.com/records/layer/correctness-core.md)

Details

Definition
The ability of a system to detect a failed component from its health signals and restore it without human action.

Requires
[Observability](https://banes-lab.com/records/architecture/observability.md), [Health Checks](https://banes-lab.com/records/architecture/health-checks.md), [Automation](https://banes-lab.com/records/lexicon/automation.md)

Reinforces
[Resilience](https://banes-lab.com/records/architecture/resilience.md)

Enables
[Auto-Remediation](https://banes-lab.com/records/architecture/auto-remediation.md)

In tension with
[Unsafe Automation](https://banes-lab.com/records/lexicon/unsafe-automation.md)

Conflicts with
[Manual Runbook Dependency](https://banes-lab.com/records/architecture/manual-runbook-dependency.md)

Referenced by
[Resilience](https://banes-lab.com/records/architecture/resilience.md), [Health Checks](https://banes-lab.com/records/architecture/health-checks.md), [Auto-Remediation](https://banes-lab.com/records/architecture/auto-remediation.md), [Chaos Engineering](https://banes-lab.com/records/architecture/chaos-engineering.md)

Tensions
[Self-Healing Architecture / Unsafe Automation](https://banes-lab.com/records/tension/self-healing-architecture-unsafe-automation.md)

Distinct from
[Auto-Remediation](https://banes-lab.com/records/architecture/auto-remediation.md): A self-healing architecture is the whole system restoring its components, while auto-remediation is one automated response to one known failure.

Distinct from
[Automation](https://banes-lab.com/records/lexicon/automation.md): Self-healing is restoring failed components without human action, while automation is any task run without manual intervention.

Distinct from
[Recovery](https://banes-lab.com/records/lexicon/recovery.md): Self-healing is recovery that the system starts itself, while recovery is returning to correct operation by any means.

Violated by
detectable failure without automated remediation

Detected by
recurring manual recovery steps

Measured by
MTTR, auto-recovery success

Refactored by
Add Health Checks, Add Restart/Remediation Policy

Enforced by
orchestration policy, runbooks

Before

```typescript
process.on("error", error => fooLog.record(error));
```

After

```typescript
supervisor.watch("foo-worker", {
start: startFooWorker,
health: fooWorkerHealth,
restart: { maxAttempts: 5, backoffMs: 1000 },
});
```

How it is checked

Checked by
orchestration policy, runbooks

Population
Every service, replica, release and recovery path in production

Freshness
A verdict stands until the topology, the release or the recovery policy changes, and is renewed by each recovery drill

Refusal
The release gate, disaster-recovery test or orchestration policy blocks a release or topology without a working recovery path

Observation
Health signals, recovery outcomes and drill results recorded against each service

Evidence
None, because the catalog states this check as a class, so a watched run belongs to each system that adopts it

Authoritative side
The recovery policy, which each drill and recovery outcome is compared against

Depends on
[Observability](https://banes-lab.com/records/architecture/observability.md), [Health Checks](https://banes-lab.com/records/architecture/health-checks.md), [Automation](https://banes-lab.com/records/lexicon/automation.md), [Resilience](https://banes-lab.com/records/architecture/resilience.md), [Auto-Remediation](https://banes-lab.com/records/architecture/auto-remediation.md)

Shape it refuses
[Manual Runbook Dependency](https://banes-lab.com/records/architecture/manual-runbook-dependency.md)

### Health Checks

- Kind: [mechanism](https://banes-lab.com/records/kind/mechanism.md)
- Category: [Self-Healing / Recovery / Deployment Safety](https://banes-lab.com/ontology/principles/architecture-category-self-healing-recovery-deployment-safety.md)
- Severity: [contextual](https://banes-lab.com/records/vocabulary/severity-contextual.md)
- Mandatory for: services
- Scope: service, deployment, runtime
- Layer: [Correctness Core](https://banes-lab.com/records/layer/correctness-core.md)

Details

Definition
A mechanism that reports whether an instance and the dependencies it needs are ready to serve, so routing and restarts can act on it.

Requires
[Observable Health Criteria](https://banes-lab.com/records/lexicon/observable-health-criteria.md)

Reinforces
[Self-Healing Architecture](https://banes-lab.com/records/architecture/self-healing-architecture.md), [Load Balancing](https://banes-lab.com/records/architecture/load-balancing.md)

Enables
[Readiness/Liveness Routing](https://banes-lab.com/records/lexicon/readiness-liveness-routing.md)

In tension with
[False Positives](https://banes-lab.com/records/lexicon/false-positives.md)

Conflicts with
[Blind Routing](https://banes-lab.com/records/lexicon/blind-routing.md)

Referenced by
[Service Discovery](https://banes-lab.com/records/architecture/service-discovery.md), [Load Balancing](https://banes-lab.com/records/architecture/load-balancing.md), [Self-Healing Architecture](https://banes-lab.com/records/architecture/self-healing-architecture.md)

Tensions
[Health Checks / False Positives](https://banes-lab.com/records/tension/false-positives-health-checks.md)

Distinct from
[Load Balancing](https://banes-lab.com/records/architecture/load-balancing.md): Health checks report whether an instance can serve, while load balancing spreads requests across the instances that can.

Violated by
traffic routed to unhealthy instance

Detected by
missing or shallow health endpoint

Measured by
health-check accuracy

Refactored by
Add Liveness/Readiness/Dependency Checks

Enforced by
deployment policy

Before

```typescript
app.get("/health", () => "ok");
```

After

```typescript
app.get("/health", async () => {
const fooStoreOk = await fooStore.ping();
const eventBusOk = await eventBus.ping();
return {
state: fooStoreOk && eventBusOk ? "ready" : "blocked",
checks: { fooStore: fooStoreOk, eventBus: eventBusOk },
};
});
```

How it is checked

Checked by
deployment policy

Population
Every service, replica, release and recovery path in production

Freshness
A verdict stands until the topology, the release or the recovery policy changes, and is renewed by each recovery drill

Refusal
The release gate, disaster-recovery test or orchestration policy blocks a release or topology without a working recovery path

Observation
Health signals, recovery outcomes and drill results recorded against each service

Evidence
None, because the catalog states this check as a class, so a watched run belongs to each system that adopts it

Authoritative side
The recovery policy, which each drill and recovery outcome is compared against

Depends on
[Observable Health Criteria](https://banes-lab.com/records/lexicon/observable-health-criteria.md), [Self-Healing Architecture](https://banes-lab.com/records/architecture/self-healing-architecture.md), [Load Balancing](https://banes-lab.com/records/architecture/load-balancing.md), [Readiness/Liveness Routing](https://banes-lab.com/records/lexicon/readiness-liveness-routing.md)

Shape it refuses
[Blind Routing](https://banes-lab.com/records/lexicon/blind-routing.md)

### Failover

- Kind: [mechanism](https://banes-lab.com/records/kind/mechanism.md)
- Category: [Self-Healing / Recovery / Deployment Safety](https://banes-lab.com/ontology/principles/architecture-category-self-healing-recovery-deployment-safety.md)
- Severity: [contextual](https://banes-lab.com/records/vocabulary/severity-contextual.md)
- Scope: service, infrastructure, data
- Layer: [Correctness Core](https://banes-lab.com/records/layer/correctness-core.md)

Details

Definition
A mechanism that switches traffic or reads to a standby instance when the active one fails.

Requires
[Redundancy](https://banes-lab.com/records/architecture/redundancy.md), [Health Detection](https://banes-lab.com/records/lexicon/health-detection.md)

Reinforces
[Availability](https://banes-lab.com/records/lexicon/availability.md)

Enables
[Continuity During Failure](https://banes-lab.com/records/lexicon/continuity-during-failure.md)

In tension with
[Consistency](https://banes-lab.com/records/architecture/consistency.md)

Conflicts with
[Single Instance Dependency](https://banes-lab.com/records/lexicon/single-instance-dependency.md)

Referenced by
[Service Discovery](https://banes-lab.com/records/architecture/service-discovery.md), [Redundancy](https://banes-lab.com/records/architecture/redundancy.md), [Replication](https://banes-lab.com/records/architecture/replication.md)

Tensions
[Failover / Consistency](https://banes-lab.com/records/tension/consistency-failover.md)

Distinct from
[Redundancy](https://banes-lab.com/records/architecture/redundancy.md): Failover is the switch to a standby, while redundancy is keeping the standby there to switch to.

Distinct from
[Replication](https://banes-lab.com/records/architecture/replication.md): Failover switches traffic to another instance, while replication keeps that instance's data current.

Violated by
no alternate instance/path for critical dependency

Detected by
single active dependency with no failover

Measured by
failover time, [availability](https://banes-lab.com/records/lexicon/availability.md)

Refactored by
Add Replica, Add Failover Routing

Enforced by
disaster recovery tests

Before

```typescript
const foo = await primaryFooStore.find(id);
```

After

```typescript
const foo = await failover.read([
primaryFooStore,
secondaryFooStore,
], store => store.find(id));
```

How it is checked

Checked by
disaster recovery tests

Population
Every service, replica, release and recovery path in production

Freshness
A verdict stands until the topology, the release or the recovery policy changes, and is renewed by each recovery drill

Refusal
The release gate, disaster-recovery test or orchestration policy blocks a release or topology without a working recovery path

Observation
Health signals, recovery outcomes and drill results recorded against each service

Evidence
None, because the catalog states this check as a class, so a watched run belongs to each system that adopts it

Authoritative side
The recovery policy, which each drill and recovery outcome is compared against

Depends on
[Redundancy](https://banes-lab.com/records/architecture/redundancy.md), [Health Detection](https://banes-lab.com/records/lexicon/health-detection.md), [Availability](https://banes-lab.com/records/lexicon/availability.md), [Continuity During Failure](https://banes-lab.com/records/lexicon/continuity-during-failure.md)

Shape it refuses
[Single Instance Dependency](https://banes-lab.com/records/lexicon/single-instance-dependency.md)

### Redundancy

- Kind: [mechanism](https://banes-lab.com/records/kind/mechanism.md)
- Category: [Self-Healing / Recovery / Deployment Safety](https://banes-lab.com/ontology/principles/architecture-category-self-healing-recovery-deployment-safety.md)
- Severity: [contextual](https://banes-lab.com/records/vocabulary/severity-contextual.md)
- Scope: infrastructure, service, data
- Layer: [Correctness Core](https://banes-lab.com/records/layer/correctness-core.md)

Details

Definition
A mechanism that keeps spare instances or capacity for a critical component, spread across failure domains.

Requires
[Replication or Alternate Capacity](https://banes-lab.com/records/lexicon/replication-or-alternate-capacity.md)

Reinforces
[Fault Tolerance](https://banes-lab.com/records/architecture/fault-tolerance.md)

Enables
[Failover](https://banes-lab.com/records/architecture/failover.md)

In tension with
[Cost](https://banes-lab.com/records/lexicon/cost.md)

Conflicts with
[Single Point of Failure](https://banes-lab.com/records/lexicon/single-point-of-failure.md)

Referenced by
[Fault Tolerance](https://banes-lab.com/records/architecture/fault-tolerance.md), [Failover](https://banes-lab.com/records/architecture/failover.md), [RAID Redundancy](https://banes-lab.com/records/architecture/raid-redundancy.md)

Tensions
[Redundancy / Cost](https://banes-lab.com/records/tension/cost-redundancy.md)

Violated by
critical singleton dependency

Detected by
SPOF analysis

Measured by
redundancy factor

Refactored by
Add Replica, Add Backup Path

Enforced by
[architecture review](https://banes-lab.com/records/architecture/architecture-review.md)

Before

```typescript
const fooService = deploy({ replicas: 1 });
```

After

```typescript
const fooService = deploy({ replicas: 3, spreadAcross: ["zone-a", "zone-b", "zone-c"] });
```

How it is checked

Checked by
architecture review

Population
Every service, replica, release and recovery path in production

Freshness
A verdict stands until the topology, the release or the recovery policy changes, and is renewed by each recovery drill

Refusal
The release gate, disaster-recovery test or orchestration policy blocks a release or topology without a working recovery path

Observation
Health signals, recovery outcomes and drill results recorded against each service

Evidence
None, because the catalog states this check as a class, so a watched run belongs to each system that adopts it

Authoritative side
The recovery policy, which each drill and recovery outcome is compared against

Depends on
[Replication or Alternate Capacity](https://banes-lab.com/records/lexicon/replication-or-alternate-capacity.md), [Fault Tolerance](https://banes-lab.com/records/architecture/fault-tolerance.md), [Failover](https://banes-lab.com/records/architecture/failover.md)

Shape it refuses
[Single Point of Failure](https://banes-lab.com/records/lexicon/single-point-of-failure.md)

### Replication

- Kind: [mechanism](https://banes-lab.com/records/kind/mechanism.md)
- Category: [Self-Healing / Recovery / Deployment Safety](https://banes-lab.com/ontology/principles/architecture-category-self-healing-recovery-deployment-safety.md)
- Severity: [contextual](https://banes-lab.com/records/vocabulary/severity-contextual.md)
- Scope: database, service, cache
- Layer: [Correctness Core](https://banes-lab.com/records/layer/correctness-core.md)

Details

Definition
A mechanism that keeps copies of data on several nodes under a declared consistency policy, such as a write quorum.

Requires
[Consistency Policy](https://banes-lab.com/records/lexicon/consistency-policy.md)

Reinforces
[Scalability](https://banes-lab.com/records/architecture/scalability.md), [Availability](https://banes-lab.com/records/lexicon/availability.md)

Enables
[Read Scaling](https://banes-lab.com/records/lexicon/read-scaling.md), [Failover](https://banes-lab.com/records/architecture/failover.md)

In tension with
[Consistency Lag](https://banes-lab.com/records/lexicon/consistency-lag.md)

Conflicts with
[Single Copy State](https://banes-lab.com/records/lexicon/single-copy-state.md)

Referenced by
[Read Replica](https://banes-lab.com/records/architecture/read-replica.md)

Tensions
[Replication / Consistency Lag](https://banes-lab.com/records/tension/consistency-lag-replication.md)

Violated by
unreplicated critical state

Detected by
SPOF data stores

Measured by
replication lag, replica count

Refactored by
Add Replica, Define Consistency Model

Enforced by
infrastructure policy

Before

```typescript
await primaryFooStore.save(foo);
```

After

```typescript
await replicatedFooStore.save(foo, { replicas: 3, writeQuorum: 2 });
```

How it is checked

Checked by
infrastructure policy

Population
Every service, replica, release and recovery path in production

Freshness
A verdict stands until the topology, the release or the recovery policy changes, and is renewed by each recovery drill

Refusal
The release gate, disaster-recovery test or orchestration policy blocks a release or topology without a working recovery path

Observation
Health signals, recovery outcomes and drill results recorded against each service

Evidence
None, because the catalog states this check as a class, so a watched run belongs to each system that adopts it

Authoritative side
The recovery policy, which each drill and recovery outcome is compared against

Depends on
[Consistency Policy](https://banes-lab.com/records/lexicon/consistency-policy.md), [Scalability](https://banes-lab.com/records/architecture/scalability.md), [Availability](https://banes-lab.com/records/lexicon/availability.md), [Read Scaling](https://banes-lab.com/records/lexicon/read-scaling.md), [Failover](https://banes-lab.com/records/architecture/failover.md)

Shape it refuses
[Single Copy State](https://banes-lab.com/records/lexicon/single-copy-state.md)

### Auto-Scaling

- Kind: [capability](https://banes-lab.com/records/kind/capability.md)
- Category: [Self-Healing / Recovery / Deployment Safety](https://banes-lab.com/ontology/principles/architecture-category-self-healing-recovery-deployment-safety.md)
- Severity: [contextual](https://banes-lab.com/records/vocabulary/severity-contextual.md)
- Scope: deployment, service, infrastructure
- Aliases: Demand-Based Capacity
- Layer: [Correctness Core](https://banes-lab.com/records/layer/correctness-core.md)

Details

Definition
The ability to change the number of running instances automatically, between set bounds, from a load metric.

Requires
[Horizontal Scalability](https://banes-lab.com/records/lexicon/horizontal-scalability.md), [Metrics](https://banes-lab.com/records/lexicon/metrics.md)

Reinforces
[Elasticity](https://banes-lab.com/records/architecture/elasticity.md)

Enables
none

In tension with
[Cost/Cold Start](https://banes-lab.com/records/lexicon/cost-cold-start.md), [Fixed Capacity](https://banes-lab.com/records/lexicon/fixed-capacity.md)

Conflicts with
none

Referenced by
[Elasticity](https://banes-lab.com/records/architecture/elasticity.md), [Statelessness](https://banes-lab.com/records/architecture/statelessness.md)

Tensions
[Auto-Scaling / Cost/Cold Start](https://banes-lab.com/records/tension/auto-scaling-cost-cold-start.md), [Auto-Scaling / Fixed Capacity](https://banes-lab.com/records/tension/auto-scaling-fixed-capacity.md)

Violated by
manual-only scaling for variable load

Detected by
saturation under load without scale policy

Measured by
scaling latency, saturation rate

Refactored by
Add Scaling Policy, Make Service Stateless

Enforced by
infrastructure-as-code policy

Before

```typescript
deployFooWorkers({ replicas: 4 });
```

After

```typescript
deployFooWorkers({
minReplicas: 2,
maxReplicas: 20,
scaleOn: { queueDepthPerWorker: 100 },
scaleInCooldownSeconds: 300,
});
```

How it is checked

Checked by
infrastructure-as-code policy

Population
Every service, replica, release and recovery path in production

Freshness
A verdict stands until the topology, the release or the recovery policy changes, and is renewed by each recovery drill

Refusal
The release gate, disaster-recovery test or orchestration policy blocks a release or topology without a working recovery path

Observation
Health signals, recovery outcomes and drill results recorded against each service

Evidence
None, because the catalog states this check as a class, so a watched run belongs to each system that adopts it

Authoritative side
The recovery policy, which each drill and recovery outcome is compared against

Depends on
[Horizontal Scalability](https://banes-lab.com/records/lexicon/horizontal-scalability.md), [Metrics](https://banes-lab.com/records/lexicon/metrics.md), [Elasticity](https://banes-lab.com/records/architecture/elasticity.md)

Shape it refuses
Not answered

### Auto-Remediation

- Kind: [capability](https://banes-lab.com/records/kind/capability.md)
- Category: [Self-Healing / Recovery / Deployment Safety](https://banes-lab.com/ontology/principles/architecture-category-self-healing-recovery-deployment-safety.md)
- Severity: [contextual](https://banes-lab.com/records/vocabulary/severity-contextual.md)
- Scope: runtime, service, infrastructure
- Aliases: Autonomous Recovery
- Layer: [Correctness Core](https://banes-lab.com/records/layer/correctness-core.md)

Details

Definition
The ability to run a guarded, automated response, such as a restart, a reset and replay or a compaction, when an alert or a health signal shows a failure whose fix is known.

Requires
[Health Signal](https://banes-lab.com/records/lexicon/health-signal.md), [Remediation Workflow](https://banes-lab.com/records/lexicon/remediation-workflow.md)

Reinforces
[Self-Healing Architecture](https://banes-lab.com/records/architecture/self-healing-architecture.md)

Enables
[Incident Reduction](https://banes-lab.com/records/lexicon/incident-reduction.md), [Reduced Mean Time to Recovery](https://banes-lab.com/records/lexicon/reduced-mean-time-to-recovery.md)

In tension with
[Unsafe Automation](https://banes-lab.com/records/lexicon/unsafe-automation.md), [Manual Remediation](https://banes-lab.com/records/lexicon/manual-remediation.md)

Conflicts with
[Manual Runbook Dependency](https://banes-lab.com/records/architecture/manual-runbook-dependency.md)

Referenced by
[Self-Healing Architecture](https://banes-lab.com/records/architecture/self-healing-architecture.md)

Tensions
[Auto-Remediation / Unsafe Automation](https://banes-lab.com/records/tension/auto-remediation-unsafe-automation.md), [Auto-Remediation / Manual Remediation](https://banes-lab.com/records/tension/auto-remediation-manual-remediation.md)

Distinct from
[Incident Reduction](https://banes-lab.com/records/lexicon/incident-reduction.md): Auto-remediation is running the known fix, while incident reduction is the fall in incidents that reach a responder as a result.

Distinct from
[Reduced Mean Time to Recovery](https://banes-lab.com/records/lexicon/reduced-mean-time-to-recovery.md): Auto-remediation is running the known fix, while reduced mean time to recovery is the shorter outage that results.

Violated by
repeatable failure with no automated response

Detected by
repeated manual runbook actions

Measured by
remediation success, false action rate

Refactored by
Automate Runbook, Add Guardrails

Enforced by
operations policy

Before

```typescript
alert.on("FooDiskFull", pageOnCall);
```

After

```typescript
alert.on("FooDiskFull", async event => {
await fooStorage.compact(event.volumeId);
await fooStorage.verify(event.volumeId);
});
```

How it is checked

Checked by
operations policy

Population
Every service, replica, release and recovery path in production

Freshness
A verdict stands until the topology, the release or the recovery policy changes, and is renewed by each recovery drill

Refusal
The release gate, disaster-recovery test or orchestration policy blocks a release or topology without a working recovery path

Observation
Health signals, recovery outcomes and drill results recorded against each service

Evidence
None, because the catalog states this check as a class, so a watched run belongs to each system that adopts it

Authoritative side
The recovery policy, which each drill and recovery outcome is compared against

Depends on
[Health Signal](https://banes-lab.com/records/lexicon/health-signal.md), [Remediation Workflow](https://banes-lab.com/records/lexicon/remediation-workflow.md), [Self-Healing Architecture](https://banes-lab.com/records/architecture/self-healing-architecture.md), [Incident Reduction](https://banes-lab.com/records/lexicon/incident-reduction.md), [Reduced Mean Time to Recovery](https://banes-lab.com/records/lexicon/reduced-mean-time-to-recovery.md)

Shape it refuses
[Manual Runbook Dependency](https://banes-lab.com/records/architecture/manual-runbook-dependency.md)

### Rollback

- Kind: [mechanism](https://banes-lab.com/records/kind/mechanism.md)
- Category: [Self-Healing / Recovery / Deployment Safety](https://banes-lab.com/ontology/principles/architecture-category-self-healing-recovery-deployment-safety.md)
- Severity: [mandatory](https://banes-lab.com/records/vocabulary/severity-mandatory.md)
- Scope: deployment, release
- Layer: [Correctness Core](https://banes-lab.com/records/layer/correctness-core.md)

Details

Definition
A mechanism that restores the previous versioned release when a deployment fails its verification.

Requires
[Versioned Artifact](https://banes-lab.com/records/lexicon/versioned-artifact.md), [Reversible Deployment](https://banes-lab.com/records/lexicon/reversible-deployment.md)

Reinforces
[Resilience](https://banes-lab.com/records/architecture/resilience.md), [Recovery](https://banes-lab.com/records/lexicon/recovery.md)

Enables
[Fast Failure Recovery](https://banes-lab.com/records/lexicon/fast-failure-recovery.md)

In tension with
[Data Migration Compatibility](https://banes-lab.com/records/lexicon/data-migration-compatibility.md)

Conflicts with
[Irreversible Deployment](https://banes-lab.com/records/lexicon/irreversible-deployment.md), [Irreversible Migration](https://banes-lab.com/records/architecture/irreversible-migration.md)

Referenced by
[Blue-Green Deployment](https://banes-lab.com/records/architecture/blue-green-deployment.md)

Tensions
[Rollback / Data Migration Compatibility](https://banes-lab.com/records/tension/data-migration-compatibility-rollback.md)

Violated by
deployment cannot be reverted

Detected by
no rollback path

Measured by
rollback success time

Refactored by
Add Rollback Plan, Make Migration Backward-Compatible

Enforced by
release gates

Before

```typescript
deploy(fooVersion);
```

After

```typescript
const release = await deploy(fooVersion);
if (!(await release.verify())) await release.rollback(previousFooVersion);
```

How it is checked

Checked by
release gates

Population
Every service, replica, release and recovery path in production

Freshness
A verdict stands until the topology, the release or the recovery policy changes, and is renewed by each recovery drill

Refusal
The release gate, disaster-recovery test or orchestration policy blocks a release or topology without a working recovery path

Observation
Health signals, recovery outcomes and drill results recorded against each service

Evidence
None, because the catalog states this check as a class, so a watched run belongs to each system that adopts it

Authoritative side
The recovery policy, which each drill and recovery outcome is compared against

Depends on
[Versioned Artifact](https://banes-lab.com/records/lexicon/versioned-artifact.md), [Reversible Deployment](https://banes-lab.com/records/lexicon/reversible-deployment.md), [Resilience](https://banes-lab.com/records/architecture/resilience.md), [Recovery](https://banes-lab.com/records/lexicon/recovery.md), [Fast Failure Recovery](https://banes-lab.com/records/lexicon/fast-failure-recovery.md)

Shape it refuses
[Irreversible Deployment](https://banes-lab.com/records/lexicon/irreversible-deployment.md), [Irreversible Migration](https://banes-lab.com/records/architecture/irreversible-migration.md)

### Blue-Green Deployment

- Kind: [pattern](https://banes-lab.com/records/kind/pattern.md)
- Category: [Self-Healing / Recovery / Deployment Safety](https://banes-lab.com/ontology/principles/architecture-category-self-healing-recovery-deployment-safety.md)
- Severity: [contextual](https://banes-lab.com/records/vocabulary/severity-contextual.md)
- Scope: deployment, release
- Layer: [Correctness Core](https://banes-lab.com/records/layer/correctness-core.md)

Details

Definition
A design pattern that deploys a release to an idle copy of production, verifies it, and switches traffic to it at once.

Requires
[Parallel Environments](https://banes-lab.com/records/lexicon/parallel-environments.md)

Reinforces
[Rollback](https://banes-lab.com/records/architecture/rollback.md), [Availability](https://banes-lab.com/records/lexicon/availability.md)

Enables
[Low-Risk Cutover](https://banes-lab.com/records/lexicon/low-risk-cutover.md)

In tension with
[Infrastructure Cost](https://banes-lab.com/records/lexicon/infrastructure-cost.md)

Conflicts with
[In-Place Mutation Only](https://banes-lab.com/records/lexicon/in-place-mutation-only.md)

Tensions
[Blue-Green Deployment / Infrastructure Cost](https://banes-lab.com/records/tension/blue-green-deployment-infrastructure-cost.md)

Violated by
high-risk in-place production deploys

Detected by
no parallel release environment

Measured by
cutover failure rate

Refactored by
Add Blue/Green Environments

Enforced by
deployment pipeline

Before

```typescript
routeAllTraffic(deployFoo("v2"));
```

After

```typescript
const green = await deployFoo("v2");
await verify(green);
await router.switch({ from: "blue", to: "green" });
```

How it is checked

Checked by
deployment pipeline

Population
Every service, replica, release and recovery path in production

Freshness
A verdict stands until the topology, the release or the recovery policy changes, and is renewed by each recovery drill

Refusal
The release gate, disaster-recovery test or orchestration policy blocks a release or topology without a working recovery path

Observation
Health signals, recovery outcomes and drill results recorded against each service

Evidence
None, because the catalog states this check as a class, so a watched run belongs to each system that adopts it

Authoritative side
The recovery policy, which each drill and recovery outcome is compared against

Depends on
[Parallel Environments](https://banes-lab.com/records/lexicon/parallel-environments.md), [Rollback](https://banes-lab.com/records/architecture/rollback.md), [Availability](https://banes-lab.com/records/lexicon/availability.md), [Low-Risk Cutover](https://banes-lab.com/records/lexicon/low-risk-cutover.md)

Shape it refuses
[In-Place Mutation Only](https://banes-lab.com/records/lexicon/in-place-mutation-only.md)

### Canary Deployment

- Kind: [pattern](https://banes-lab.com/records/kind/pattern.md)
- Category: [Self-Healing / Recovery / Deployment Safety](https://banes-lab.com/ontology/principles/architecture-category-self-healing-recovery-deployment-safety.md)
- Severity: [contextual](https://banes-lab.com/records/vocabulary/severity-contextual.md)
- Scope: deployment, release
- Layer: [Correctness Core](https://banes-lab.com/records/layer/correctness-core.md)

Details

Definition
A design pattern that sends a small share of traffic to a new release and widens the share only while its error and latency stay within limits.

Requires
[Traffic Splitting](https://banes-lab.com/records/lexicon/traffic-splitting.md), [Observability](https://banes-lab.com/records/architecture/observability.md)

Reinforces
[Progressive Delivery](https://banes-lab.com/records/lexicon/progressive-delivery.md)

Enables
[Controlled Exposure](https://banes-lab.com/records/lexicon/controlled-exposure.md)

In tension with
[Rollout Complexity](https://banes-lab.com/records/lexicon/rollout-complexity.md)

Conflicts with
[Big-Bang Release](https://banes-lab.com/records/architecture/big-bang-release.md)

Tensions
[Canary Deployment / Rollout Complexity](https://banes-lab.com/records/tension/canary-deployment-rollout-complexity.md)

Violated by
full rollout without health/error guard

Detected by
no staged traffic policy

Measured by
canary error budget, rollback trigger rate

Refactored by
Add Canary Stage, Add Automated Guardrails

Enforced by
deployment pipeline

Before

```typescript
await router.route("foo-v2", 100);
```

After

```typescript
await router.route("foo-v2", 5);
await verifyCanary({ errorRate: 0.01, latencyP95Ms: 200 });
await router.progressiveShift("foo-v2", [25, 50, 100]);
```

How it is checked

Checked by
deployment pipeline

Population
Every service, replica, release and recovery path in production

Freshness
A verdict stands until the topology, the release or the recovery policy changes, and is renewed by each recovery drill

Refusal
The release gate, disaster-recovery test or orchestration policy blocks a release or topology without a working recovery path

Observation
Health signals, recovery outcomes and drill results recorded against each service

Evidence
None, because the catalog states this check as a class, so a watched run belongs to each system that adopts it

Authoritative side
The recovery policy, which each drill and recovery outcome is compared against

Depends on
[Traffic Splitting](https://banes-lab.com/records/lexicon/traffic-splitting.md), [Observability](https://banes-lab.com/records/architecture/observability.md), [Progressive Delivery](https://banes-lab.com/records/lexicon/progressive-delivery.md), [Controlled Exposure](https://banes-lab.com/records/lexicon/controlled-exposure.md)

Shape it refuses
[Big-Bang Release](https://banes-lab.com/records/architecture/big-bang-release.md)

### Chaos Engineering

- Kind: [activity](https://banes-lab.com/records/kind/activity.md)
- Category: [Self-Healing / Recovery / Deployment Safety](https://banes-lab.com/ontology/principles/architecture-category-self-healing-recovery-deployment-safety.md)
- Severity: [contextual](https://banes-lab.com/records/vocabulary/severity-contextual.md)
- Scope: system, resilience, operations
- Layer: [Correctness Core](https://banes-lab.com/records/layer/correctness-core.md)

Details

Definition
The practice of injecting controlled faults into a running system to test a stated hypothesis about how it recovers.

Requires
[Observability](https://banes-lab.com/records/architecture/observability.md)

Reinforces
[Self-Healing Architecture](https://banes-lab.com/records/architecture/self-healing-architecture.md), [Fault Tolerance](https://banes-lab.com/records/architecture/fault-tolerance.md)

Enables
[Empirical Resilience Verification](https://banes-lab.com/records/lexicon/empirical-resilience-verification.md)

In tension with
[Production Risk](https://banes-lab.com/records/lexicon/production-risk.md)

Conflicts with
[Untested Failure Assumptions](https://banes-lab.com/records/lexicon/untested-failure-assumptions.md)

Tensions
[Chaos Engineering / Production Risk](https://banes-lab.com/records/tension/chaos-engineering-production-risk.md)

Violated by
resilience assumed but never exercised

Detected by
no fault-injection testing of recovery paths

Measured by
unverified failure-mode count

Refactored by
Introduce Controlled Fault Injection

Enforced by
resilience review

Before

```typescript
assumeFooSurvivesZoneLoss();
```

After

```typescript
chaos.experiment("foo-zone-loss", {
inject: () => killZone("zone-a"),
hypothesis: () => fooHealth.available(),
});
```

How it is checked

Checked by
resilience review

Population
Every service, replica, release and recovery path in production

Freshness
A verdict stands until the topology, the release or the recovery policy changes, and is renewed by each recovery drill

Refusal
The release gate, disaster-recovery test or orchestration policy blocks a release or topology without a working recovery path

Observation
Health signals, recovery outcomes and drill results recorded against each service

Evidence
None, because the catalog states this check as a class, so a watched run belongs to each system that adopts it

Authoritative side
The recovery policy, which each drill and recovery outcome is compared against

Depends on
[Observability](https://banes-lab.com/records/architecture/observability.md), [Self-Healing Architecture](https://banes-lab.com/records/architecture/self-healing-architecture.md), [Fault Tolerance](https://banes-lab.com/records/architecture/fault-tolerance.md), [Empirical Resilience Verification](https://banes-lab.com/records/lexicon/empirical-resilience-verification.md)

Shape it refuses
[Untested Failure Assumptions](https://banes-lab.com/records/lexicon/untested-failure-assumptions.md)

### Graceful Shutdown

- Kind: [mechanism](https://banes-lab.com/records/kind/mechanism.md)
- Category: [Self-Healing / Recovery / Deployment Safety](https://banes-lab.com/ontology/principles/architecture-category-self-healing-recovery-deployment-safety.md)
- Severity: [contextual](https://banes-lab.com/records/vocabulary/severity-contextual.md)
- Mandatory for: production systems
- Scope: service, runtime, resilience
- Layer: [Correctness Core](https://banes-lab.com/records/layer/correctness-core.md)

Details

Definition
A mechanism that, on a stop signal, stops accepting work, drains in-flight work and closes connections before the process exits.

Requires
[Lifecycle Signals](https://banes-lab.com/records/lexicon/lifecycle-signals.md)

Reinforces
[Reliability](https://banes-lab.com/records/lexicon/reliability.md), [Data Integrity](https://banes-lab.com/records/lexicon/data-integrity.md)

Enables
[In-Flight Work Drain](https://banes-lab.com/records/lexicon/in-flight-work-drain.md), [Connection Cleanup](https://banes-lab.com/records/lexicon/connection-cleanup.md)

In tension with
[Shutdown Latency](https://banes-lab.com/records/lexicon/shutdown-latency.md)

Conflicts with
[Hard Process Kill](https://banes-lab.com/records/lexicon/hard-process-kill.md)

Tensions
[Graceful Shutdown / Shutdown Latency](https://banes-lab.com/records/tension/graceful-shutdown-shutdown-latency.md)

Violated by
processes terminated mid-request with no drain

Detected by
dropped in-flight work on deploy/restart

Measured by
requests lost per restart

Refactored by
Implement Graceful Drain on Shutdown

Enforced by
operations review

Before

```typescript
process.on("SIGTERM", () => process.exit(0));
```

After

```typescript
process.on("SIGTERM", async () => {
server.stopAccepting();
await fooQueue.drain();
await server.close();
process.exit(0);
});
```

How it is checked

Checked by
operations review

Population
Every service, replica, release and recovery path in production

Freshness
A verdict stands until the topology, the release or the recovery policy changes, and is renewed by each recovery drill

Refusal
The release gate, disaster-recovery test or orchestration policy blocks a release or topology without a working recovery path

Observation
Health signals, recovery outcomes and drill results recorded against each service

Evidence
None, because the catalog states this check as a class, so a watched run belongs to each system that adopts it

Authoritative side
The recovery policy, which each drill and recovery outcome is compared against

Depends on
[Lifecycle Signals](https://banes-lab.com/records/lexicon/lifecycle-signals.md), [Reliability](https://banes-lab.com/records/lexicon/reliability.md), [Data Integrity](https://banes-lab.com/records/lexicon/data-integrity.md), [In-Flight Work Drain](https://banes-lab.com/records/lexicon/in-flight-work-drain.md), [Connection Cleanup](https://banes-lab.com/records/lexicon/connection-cleanup.md)

Shape it refuses
[Hard Process Kill](https://banes-lab.com/records/lexicon/hard-process-kill.md)

### RAID Redundancy

- Kind: [technique](https://banes-lab.com/records/kind/technique.md)
- Category: [Self-Healing / Recovery / Deployment Safety](https://banes-lab.com/ontology/principles/architecture-category-self-healing-recovery-deployment-safety.md)
- Severity: [contextual](https://banes-lab.com/records/vocabulary/severity-contextual.md)
- Mandatory for: managed infrastructure
- Scope: storage, redundancy, infrastructure
- Aliases: Redundant Array of Independent Disks
- Layer: [Correctness Core](https://banes-lab.com/records/layer/correctness-core.md)

Details

Definition
A technique for spreading data across several disks with mirroring or parity, so the loss of a disk loses no data.

Requires
[Multiple Physical Disks](https://banes-lab.com/records/lexicon/multiple-physical-disks.md)

Reinforces
[Redundancy](https://banes-lab.com/records/architecture/redundancy.md), [Fault Tolerance](https://banes-lab.com/records/architecture/fault-tolerance.md)

Enables
[Disk-Failure Survival](https://banes-lab.com/records/lexicon/disk-failure-survival.md), [Parity-Based Recovery](https://banes-lab.com/records/lexicon/parity-based-recovery.md)

In tension with
[Write Amplification](https://banes-lab.com/records/lexicon/write-amplification.md)

Conflicts with
[Single-Disk Point of Failure](https://banes-lab.com/records/lexicon/single-disk-point-of-failure.md)

Tensions
[RAID Redundancy / Write Amplification](https://banes-lab.com/records/tension/raid-redundancy-write-amplification.md)

Violated by
durable data written to a single disk with no physical redundancy

Detected by
total data loss when one drive fails

Measured by
tolerated simultaneous disk failures

Refactored by
Place data on a mirrored or parity RAID array (RAID 1/5/10)

Enforced by
storage architecture review

Before

```typescript
const store = new SingleDiskFooStore("/dev/sda");
```

After

```typescript
const store = new FooStore({
volume: raidArray({ level: 10, disks: ["/dev/sda", "/dev/sdb", "/dev/sdc", "/dev/sdd"] }),
});
```

How it is checked

Checked by
storage architecture review

Population
Every service, replica, release and recovery path in production

Freshness
A verdict stands until the topology, the release or the recovery policy changes, and is renewed by each recovery drill

Refusal
The release gate, disaster-recovery test or orchestration policy blocks a release or topology without a working recovery path

Observation
Health signals, recovery outcomes and drill results recorded against each service

Evidence
None, because the catalog states this check as a class, so a watched run belongs to each system that adopts it

Authoritative side
The recovery policy, which each drill and recovery outcome is compared against

Depends on
[Multiple Physical Disks](https://banes-lab.com/records/lexicon/multiple-physical-disks.md), [Redundancy](https://banes-lab.com/records/architecture/redundancy.md), [Fault Tolerance](https://banes-lab.com/records/architecture/fault-tolerance.md), [Disk-Failure Survival](https://banes-lab.com/records/lexicon/disk-failure-survival.md), [Parity-Based Recovery](https://banes-lab.com/records/lexicon/parity-based-recovery.md)

Shape it refuses
[Single-Disk Point of Failure](https://banes-lab.com/records/lexicon/single-disk-point-of-failure.md)

## Links to

- [capability](https://banes-lab.com/records/kind/capability.md)
- [contextual](https://banes-lab.com/records/vocabulary/severity-contextual.md)
- [Correctness Core](https://banes-lab.com/records/layer/correctness-core.md)
- [Observability](https://banes-lab.com/records/architecture/observability.md)
- [Health Checks](https://banes-lab.com/records/architecture/health-checks.md)
- [Automation](https://banes-lab.com/records/lexicon/automation.md)
- [Resilience](https://banes-lab.com/records/architecture/resilience.md)
- [Auto-Remediation](https://banes-lab.com/records/architecture/auto-remediation.md)
- [Unsafe Automation](https://banes-lab.com/records/lexicon/unsafe-automation.md)
- [Manual Runbook Dependency](https://banes-lab.com/records/architecture/manual-runbook-dependency.md)
- [Chaos Engineering](https://banes-lab.com/records/architecture/chaos-engineering.md)
- [Self-Healing Architecture / Unsafe Automation](https://banes-lab.com/records/tension/self-healing-architecture-unsafe-automation.md)
- [Recovery](https://banes-lab.com/records/lexicon/recovery.md)
- [mechanism](https://banes-lab.com/records/kind/mechanism.md)
- [Observable Health Criteria](https://banes-lab.com/records/lexicon/observable-health-criteria.md)
- [Self-Healing Architecture](https://banes-lab.com/records/architecture/self-healing-architecture.md)
- [Load Balancing](https://banes-lab.com/records/architecture/load-balancing.md)
- [Readiness/Liveness Routing](https://banes-lab.com/records/lexicon/readiness-liveness-routing.md)
- [False Positives](https://banes-lab.com/records/lexicon/false-positives.md)
- [Blind Routing](https://banes-lab.com/records/lexicon/blind-routing.md)
- [Service Discovery](https://banes-lab.com/records/architecture/service-discovery.md)
- [Health Checks / False Positives](https://banes-lab.com/records/tension/false-positives-health-checks.md)
- [Redundancy](https://banes-lab.com/records/architecture/redundancy.md)
- [Health Detection](https://banes-lab.com/records/lexicon/health-detection.md)
- [Availability](https://banes-lab.com/records/lexicon/availability.md)
- [Continuity During Failure](https://banes-lab.com/records/lexicon/continuity-during-failure.md)
- [Consistency](https://banes-lab.com/records/architecture/consistency.md)
- [Single Instance Dependency](https://banes-lab.com/records/lexicon/single-instance-dependency.md)
- [Replication](https://banes-lab.com/records/architecture/replication.md)
- [Failover / Consistency](https://banes-lab.com/records/tension/consistency-failover.md)
- [Replication or Alternate Capacity](https://banes-lab.com/records/lexicon/replication-or-alternate-capacity.md)
- [Fault Tolerance](https://banes-lab.com/records/architecture/fault-tolerance.md)
- [Failover](https://banes-lab.com/records/architecture/failover.md)
- [Cost](https://banes-lab.com/records/lexicon/cost.md)
- [Single Point of Failure](https://banes-lab.com/records/lexicon/single-point-of-failure.md)
- [RAID Redundancy](https://banes-lab.com/records/architecture/raid-redundancy.md)
- [Redundancy / Cost](https://banes-lab.com/records/tension/cost-redundancy.md)
- [Architecture Review](https://banes-lab.com/records/architecture/architecture-review.md)
- [Consistency Policy](https://banes-lab.com/records/lexicon/consistency-policy.md)
- [Scalability](https://banes-lab.com/records/architecture/scalability.md)
- [Read Scaling](https://banes-lab.com/records/lexicon/read-scaling.md)
- [Consistency Lag](https://banes-lab.com/records/lexicon/consistency-lag.md)
- [Single Copy State](https://banes-lab.com/records/lexicon/single-copy-state.md)
- [Read Replica](https://banes-lab.com/records/architecture/read-replica.md)
- [Replication / Consistency Lag](https://banes-lab.com/records/tension/consistency-lag-replication.md)
- [Horizontal Scalability](https://banes-lab.com/records/lexicon/horizontal-scalability.md)
- [Metrics](https://banes-lab.com/records/lexicon/metrics.md)
- [Elasticity](https://banes-lab.com/records/architecture/elasticity.md)
- [Cost/Cold Start](https://banes-lab.com/records/lexicon/cost-cold-start.md)
- [Fixed Capacity](https://banes-lab.com/records/lexicon/fixed-capacity.md)
- [Statelessness](https://banes-lab.com/records/architecture/statelessness.md)
- [Auto-Scaling / Cost/Cold Start](https://banes-lab.com/records/tension/auto-scaling-cost-cold-start.md)
- [Auto-Scaling / Fixed Capacity](https://banes-lab.com/records/tension/auto-scaling-fixed-capacity.md)
- [Health Signal](https://banes-lab.com/records/lexicon/health-signal.md)
- [Remediation Workflow](https://banes-lab.com/records/lexicon/remediation-workflow.md)
- [Incident Reduction](https://banes-lab.com/records/lexicon/incident-reduction.md)
- [Reduced Mean Time to Recovery](https://banes-lab.com/records/lexicon/reduced-mean-time-to-recovery.md)
- [Manual Remediation](https://banes-lab.com/records/lexicon/manual-remediation.md)
- [Auto-Remediation / Unsafe Automation](https://banes-lab.com/records/tension/auto-remediation-unsafe-automation.md)
- [Auto-Remediation / Manual Remediation](https://banes-lab.com/records/tension/auto-remediation-manual-remediation.md)
- [mandatory](https://banes-lab.com/records/vocabulary/severity-mandatory.md)
- [Versioned Artifact](https://banes-lab.com/records/lexicon/versioned-artifact.md)
- [Reversible Deployment](https://banes-lab.com/records/lexicon/reversible-deployment.md)
- [Fast Failure Recovery](https://banes-lab.com/records/lexicon/fast-failure-recovery.md)
- [Data Migration Compatibility](https://banes-lab.com/records/lexicon/data-migration-compatibility.md)
- [Irreversible Deployment](https://banes-lab.com/records/lexicon/irreversible-deployment.md)
- [Irreversible Migration](https://banes-lab.com/records/architecture/irreversible-migration.md)
- [Blue-Green Deployment](https://banes-lab.com/records/architecture/blue-green-deployment.md)
- [Rollback / Data Migration Compatibility](https://banes-lab.com/records/tension/data-migration-compatibility-rollback.md)
- [pattern](https://banes-lab.com/records/kind/pattern.md)
- [Parallel Environments](https://banes-lab.com/records/lexicon/parallel-environments.md)
- [Rollback](https://banes-lab.com/records/architecture/rollback.md)
- [Low-Risk Cutover](https://banes-lab.com/records/lexicon/low-risk-cutover.md)
- [Infrastructure Cost](https://banes-lab.com/records/lexicon/infrastructure-cost.md)
- [In-Place Mutation Only](https://banes-lab.com/records/lexicon/in-place-mutation-only.md)
- [Blue-Green Deployment / Infrastructure Cost](https://banes-lab.com/records/tension/blue-green-deployment-infrastructure-cost.md)
- [Traffic Splitting](https://banes-lab.com/records/lexicon/traffic-splitting.md)
- [Progressive Delivery](https://banes-lab.com/records/lexicon/progressive-delivery.md)
- [Controlled Exposure](https://banes-lab.com/records/lexicon/controlled-exposure.md)
- [Rollout Complexity](https://banes-lab.com/records/lexicon/rollout-complexity.md)
- [Big-Bang Release](https://banes-lab.com/records/architecture/big-bang-release.md)
- [Canary Deployment / Rollout Complexity](https://banes-lab.com/records/tension/canary-deployment-rollout-complexity.md)
- [activity](https://banes-lab.com/records/kind/activity.md)
- [Empirical Resilience Verification](https://banes-lab.com/records/lexicon/empirical-resilience-verification.md)
- [Production Risk](https://banes-lab.com/records/lexicon/production-risk.md)
- [Untested Failure Assumptions](https://banes-lab.com/records/lexicon/untested-failure-assumptions.md)
- [Chaos Engineering / Production Risk](https://banes-lab.com/records/tension/chaos-engineering-production-risk.md)
- [Lifecycle Signals](https://banes-lab.com/records/lexicon/lifecycle-signals.md)
- [Reliability](https://banes-lab.com/records/lexicon/reliability.md)
- [Data Integrity](https://banes-lab.com/records/lexicon/data-integrity.md)
- [In-Flight Work Drain](https://banes-lab.com/records/lexicon/in-flight-work-drain.md)
- [Connection Cleanup](https://banes-lab.com/records/lexicon/connection-cleanup.md)
- [Shutdown Latency](https://banes-lab.com/records/lexicon/shutdown-latency.md)
- [Hard Process Kill](https://banes-lab.com/records/lexicon/hard-process-kill.md)
- [Graceful Shutdown / Shutdown Latency](https://banes-lab.com/records/tension/graceful-shutdown-shutdown-latency.md)
- [technique](https://banes-lab.com/records/kind/technique.md)
- [Multiple Physical Disks](https://banes-lab.com/records/lexicon/multiple-physical-disks.md)
- [Disk-Failure Survival](https://banes-lab.com/records/lexicon/disk-failure-survival.md)
- [Parity-Based Recovery](https://banes-lab.com/records/lexicon/parity-based-recovery.md)
- [Write Amplification](https://banes-lab.com/records/lexicon/write-amplification.md)
- [Single-Disk Point of Failure](https://banes-lab.com/records/lexicon/single-disk-point-of-failure.md)
- [RAID Redundancy / Write Amplification](https://banes-lab.com/records/tension/raid-redundancy-write-amplification.md)

## Linked from

- [The layer topology](https://banes-lab.com/ontology/schema/the-layer-topology.md)
- [The membership](https://banes-lab.com/ontology/schema/the-membership.md)
