Cloud World Model on Smithery
    Cloud World Model

    Scenario Library

    Pre-built cloud simulation scenarios.

    Scenario Library

    Resource
    Provider
    AWS EC2 Provider API Throttling
    See how AWS API limits affect a batch of EC2 changes.

    Educational · Guided scenario

    A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.

    beginner
    AWS
    Provider Limits
    API Throttling
    Parallelism
    240 operations~5 min

    What will happen

    See how provider limits, retries, and queueing affect a batch of AWS EC2 changes at the selected concurrency. This is a catalog-backed provider-limit simulation separate from the workspace resource state.

    Simulation details

    CWM models 240 EC2 create, update, or delete requests at the selected concurrency of 8. Some requests may be slowed, queued, or retried. The results show how provider limits, retries, and queueing affect the batch.

    No AWS APIs are called and no EC2 instances are created.

    ECS Fargate Task Scaling
    See a modeled Amazon ECS Fargate service add tasks during a 30→500 RPS ramp, respect its configured fleet ceiling, and account for task startup, bounded queueing, vCPU, and memory billing.

    Predictive · Capacity simulation

    Modeled load and capacity constraints reveal bottlenecks or failure.

    intermediate
    AWS
    ECS
    Fargate
    Tasks
    Scale to Zero
    2 resources~18 min

    Starts at a healthy baseline — enable "Traffic Recovery — watch Fargate scale to zero" in the workspace when you want to run that optional phase.

    What will happen

    See a modeled Amazon ECS Fargate service add tasks during a 30→500 RPS ramp, respect its configured fleet ceiling, and account for task startup, bounded queueing, vCPU, and memory billing.

    Simulation details

    Modeled demand

    Traffic Ramp — watch Fargate tasks start · ramp · steps 0–30 · 30→500 RPS

    Optional phases

    Traffic Recovery — watch Fargate scale to zero · ramp · steps 60–90 · 500→0 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Azure SaaS Web App
    Compare the modeled cost and capacity response of an Azure SaaS stack with Front Door, App Service, PostgreSQL Flexible Server, Monitor, and Defender. The 75% cache assumption and warning-band outcome belong to this configuration, not all Azure SaaS workloads.

    Optimization · Configuration comparison

    The engine compares configuration trade-offs and surfaces the better option.

    beginner
    Azure
    SaaS
    App Service
    Front Door
    PostgreSQL
    Cost
    5 resources~10 min

    Starts at a healthy baseline — enable "Traffic Recovery — see the stack settle back down" in the workspace when you want to run that optional phase.

    What will happen

    Compare the modeled cost and capacity response of an Azure SaaS stack with Front Door, App Service, PostgreSQL Flexible Server, Monitor, and Defender. The 75% cache assumption and warning-band outcome belong to this configuration, not all Azure SaaS workloads.

    Simulation details

    Modeled demand

    Business-hours ramp — watch App Service climb into the warning zone · ramp · steps 0–60 · 200→1,550 RPS

    Optional phases

    Traffic Recovery — see the stack settle back down · ramp · steps 60–90 · 1,550→200 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    GCP Pub/Sub Topic Saturation
    Model a Google Cloud Pub/Sub subscriber fleet under rising event demand. The focus is backlog, subscriber capacity, and delivery pressure—not a live Pub/Sub outage or a guaranteed acknowledgement deadline.

    Predictive · Capacity simulation

    Modeled load and capacity constraints reveal bottlenecks or failure.

    intermediate
    GCP
    Pub/Sub
    Queue
    Failure
    Recovery
    Backlog
    Subscribers
    6 resources~15 min

    Starts at a healthy baseline — enable "Message Burst — overwhelm subscribers" in the workspace when you want to run that optional phase.

    What will happen

    Model a Google Cloud Pub/Sub subscriber fleet under rising event demand. The focus is backlog, subscriber capacity, and delivery pressure—not a live Pub/Sub outage or a guaranteed acknowledgement deadline.

    Simulation details

    Modeled demand

    Steady Publisher Load · Hold · From step 0 · 1,200 RPS

    Optional phases

    Message Burst — overwhelm subscribers · ramp · steps 0–20 · 1,200→5,000 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Azure Service Bus Queue Saturation
    Model an Azure Service Bus order queue as demand rises beyond the configured consumer capacity. This is a queueing and throughput exercise; it does not call Azure or imply that Premium upgrade, auto-forwarding, or consumer scaling happens automatically.

    Predictive · Capacity simulation

    Modeled load and capacity constraints reveal bottlenecks or failure.

    intermediate
    Azure
    Service Bus
    Queue
    Failure
    Recovery
    Backlog
    Consumers
    6 resources~15 min

    Starts at a healthy baseline — enable "Order Burst — overwhelm consumers" in the workspace when you want to run that optional phase.

    What will happen

    Model an Azure Service Bus order queue as demand rises beyond the configured consumer capacity. This is a queueing and throughput exercise; it does not call Azure or imply that Premium upgrade, auto-forwarding, or consumer scaling happens automatically.

    Simulation details

    Modeled demand

    Steady Order Volume · Hold · From step 0 · 1,200 RPS

    Optional phases

    Order Burst — overwhelm consumers · ramp · steps 0–20 · 1,200→5,000 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Launch Day Spike
    Find where a modeled upload server and transcoding worker become capacity-bound during a tenfold launch surge. This is a bounded planning simulation, not a production forecast or a guarantee of an exact concurrency threshold.

    Predictive · Capacity simulation

    Modeled load and capacity constraints reveal bottlenecks or failure.

    beginner
    single-node
    launch
    queue
    upload
    OCI
    Beginner
    3 resources~10 min

    What will happen

    Find where a modeled upload server and transcoding worker become capacity-bound during a tenfold launch surge. This is a bounded planning simulation, not a production forecast or a guarantee of an exact concurrency threshold.

    Simulation details

    Modeled demand

    Launch Day Surge (10×) · ramp · steps 0–30 · 25→250 RPS

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Web App Autoscaling
    See how a load-balanced web tier responds as demand rises, and compare scale-out with the optional later scale-in phase. This is an illustrative capacity walkthrough, not a forecast of a particular AWS deployment.

    Educational · Guided scenario

    A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.

    beginner
    AWS
    Autoscaling
    EC2
    4 resources~10 min

    Starts at a healthy baseline — enable "Traffic Recovery — see scale-in" in the workspace when you want to run that optional phase.

    What will happen

    See how a load-balanced web tier responds as demand rises, and compare scale-out with the optional later scale-in phase. This is an illustrative capacity walkthrough, not a forecast of a particular AWS deployment.

    Simulation details

    Modeled demand

    Traffic Ramp — watch autoscaling fire · ramp · steps 0–40 · 800→4,000 RPS

    Optional phases

    Traffic Recovery — see scale-in · ramp · steps 60–100 · 4,000→800 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Database Failover
    Observe an injected RDS primary overload in a two-AZ application with a standby replica. The card describes a modeled failover path; it is not a live AWS RDS event or a promise of a provider RTO.

    Chaos · Injected failure

    An intentional failure tests how the architecture responds.

    intermediate
    RDS
    High Availability
    Multi-AZ
    6 resources~30 min

    Starts at a healthy baseline — enable "Traffic Recovery" in the workspace when you want to run that optional phase.

    What will happen

    Observe an injected RDS primary overload in a two-AZ application with a standby replica. The card describes a modeled failover path; it is not a live AWS RDS event or a promise of a provider RTO.

    Simulation details

    Modeled demand

    Traffic Ramp — watch RDS Primary reach its limit · ramp · steps 0–25 · 3,000→8,000 RPS

    Optional phases

    Traffic Recovery · ramp · steps 60–80 · 8,000→2,500 RPS · enable in the workspace

    CWM injects severe database overload at RDS Primary · Start step 15 · Duration 35 steps · End step 50.

    AWS Multi-Region Failover — Route 53 Health Checks
    Real incident·AWS · Jul 24, 2026
    Walk through a modeled AWS multi-region DNS failover. Route 53 health checks, DNS caching, and the east fleet’s capacity are the teaching model; the July 2026 incident is context, not proof that this run reproduces its duration.

    Chaos · Injected failure

    An intentional failure tests how the architecture responds.

    intermediate
    AWS
    Route 53
    Multi-Region
    Failover
    DNS
    Resilience
    us-west-2
    us-east-1
    11 resources~15 min

    Route 53 (DNS) · RDS Multi-AZ auto-failover

    Starts at a healthy baseline — enable "Traffic Recovery — watch us-east-1 scale in as us-west-2 recovers" in the workspace when you want to run that optional phase.

    What will happen

    Walk through a modeled AWS multi-region DNS failover. Route 53 health checks, DNS caching, and the east fleet’s capacity are the teaching model; the July 2026 incident is context, not proof that this run reproduces its duration.

    Simulation details

    Modeled demand

    Mid-failover baseline — Route 53 draining west, loading east · Hold · From step 0 · 120 RPS

    East Scale-Out — DNS TTLs expire, us-east-1 absorbs us-west-2 drain · ramp · steps 10–40 · 120→1,500 RPS

    Optional phases

    Traffic Recovery — watch us-east-1 scale in as us-west-2 recovers · ramp · steps 60–90 · 1,500→120 RPS · enable in the workspace

    CWM injects severe availability-zone outage at zone usw2-az1 · Start step 0 · Duration 9999 steps · End step 9999.

    CWM injects severe availability-zone outage at zone usw2-az2 · Start step 0 · Duration 9999 steps · End step 9999.

    Cloud Spanner Outage — When Paxos Can't Help (July 4, 2023)
    Real incident·GCP · Jul 4, 2023
    See how an application responds when all configured Cloud Spanner replicas lose their serving path. CWM models the resulting service interruption and recovery choices; it does not reproduce the real incident’s cause, duration, or customer impact.

    Chaos · Injected failure

    An intentional failure tests how the architecture responds.

    advanced
    GCP
    Cloud Spanner
    Paxos
    Multi-Region
    Outage
    Software Bug
    Circuit Breaker
    Resilience
    6 resources~15 min

    Global LB (anycast) · Paxos consensus (zero RPO)

    What will happen

    See how an application responds when all configured Cloud Spanner replicas lose their serving path. CWM models the resulting service interruption and recovery choices; it does not reproduce the real incident’s cause, duration, or customer impact.

    Simulation details

    Modeled demand

    Pre-outage baseline — 500 RPS hitting Spanner hard · Hold · From step 0 · 500 RPS

    CWM injects severe availability-zone outage at zone usc1-zone-b · Start step 0 · Duration 9999 steps · End step 9999.

    CWM injects severe availability-zone outage at zone use1-zone-c · Start step 0 · Duration 9999 steps · End step 9999.

    CWM injects severe availability-zone outage at zone use4-zone-a · Start step 0 · Duration 9999 steps · End step 9999.

    OCI Multi-Region Failover — Traffic Management Steering + Autonomous Data Guard
    Real incident·OCI · Mar 3, 2026
    Explore how traffic steering and a standby database can help an OCI application during a regional network failure. Promotion and connection changes are modeled choices for the user, not an automatic switchover or a guarantee of zero data loss.

    Chaos · Injected failure

    An intentional failure tests how the architecture responds.

    intermediate
    OCI
    Traffic Management
    Autonomous Database
    Data Guard
    Multi-Region
    HA
    Resilience
    6 resources~14 min

    Traffic Mgmt (DNS) · Data Guard manual SWITCHOVER

    Starts at a healthy baseline — enable "Scale test — confirm Phoenix absorbs full production load" in the workspace when you want to run that optional phase.

    What will happen

    Explore how traffic steering and a standby database can help an OCI application during a regional network failure. Promotion and connection changes are modeled choices for the user, not an automatic switchover or a guarantee of zero data loss.

    Simulation details

    Modeled demand

    us-ashburn-1 degraded — Traffic Management routing to us-phoenix-1 · Hold · From step 0 · 800 RPS

    Optional phases

    Scale test — confirm Phoenix absorbs full production load · ramp · steps 60–90 · 800→4,000 RPS · enable in the workspace

    CWM injects severe availability-zone outage at zone iad-ad-1 · Start step 0 · Duration 9999 steps · End step 9999.

    DigitalOcean Multi-Region HA — Floating IP Failover and the Cross-Region Database Gap
    See how an application can keep serving when its New York site fails and traffic moves to San Francisco. The database still needs manual cross-region recovery, so this exercise shows the difference between routing traffic and restoring data; it does not promise automatic failover or a provider SLA.

    Chaos · Injected failure

    An intentional failure tests how the architecture responds.

    intermediate
    DigitalOcean
    Floating IP
    Multi-Region
    PostgreSQL
    pg_logical
    HA
    Resilience
    5 resources~10 min

    Floating IP (DNS) · pg_logical manual promotion

    Starts at a healthy baseline — enable "Scale test — confirm SFO3 Droplet handles full NYC3 production load" in the workspace when you want to run that optional phase.

    What will happen

    See how an application can keep serving when its New York site fails and traffic moves to San Francisco. The database still needs manual cross-region recovery, so this exercise shows the difference between routing traffic and restoring data; it does not promise automatic failover or a provider SLA.

    Simulation details

    Modeled demand

    NYC3 degraded — DNS update in progress, traffic routing to SFO3 · Hold · From step 0 · 600 RPS

    Optional phases

    Scale test — confirm SFO3 Droplet handles full NYC3 production load · ramp · steps 60–90 · 600→2,500 RPS · enable in the workspace

    CWM injects severe availability-zone outage at zone nyc3-az1 · Start step 0 · Duration 9999 steps · End step 9999.

    AWS Regional Outage — July 2026
    Real incident·AWS · Jul 24, 2026
    Study a bounded, incident-inspired AWS regional routing failure that is structural rather than demand-driven. The historical date and incident basis provide context; the card’s steps and recovery behavior are an illustrative CWM model, not a historical replay.

    Chaos · Injected failure

    An intentional failure tests how the architecture responds.

    intermediate
    AWS
    Outage
    Failure
    us-west-2
    RDS
    ElastiCache
    7 resources~15 min

    Starts at a healthy baseline — enable "Traffic Recovery — watch services stabilize" in the workspace when you want to run that optional phase.

    What will happen

    Study a bounded, incident-inspired AWS regional routing failure that is structural rather than demand-driven. The historical date and incident basis provide context; the card’s steps and recovery behavior are an illustrative CWM model, not a historical replay.

    Simulation details

    Modeled demand

    Steady outage load — structural crisis at low baseline · Hold · From step 0 · 300 RPS

    Optional phases

    Traffic Recovery — watch services stabilize · ramp · steps 60–90 · 300→2,000 RPS · enable in the workspace

    CWM injects severe availability-zone outage at zone us-west-2-regional · Start step 0 · Duration 45 steps · End step 45.

    CWM injects severe availability-zone outage at zone us-west-2-compute · Start step 0 · Duration 45 steps · End step 45.

    CWM injects severe database overload at RDS MySQL Primary · Start step 0 · Duration 40 steps · End step 40.

    Multi-Cloud Hybrid Architecture
    Compare a modeled hybrid path that connects AWS and OCI compute, databases, load balancers, and object storage. The goal is to inspect configuration trade-offs at a steady 1,200 RPS, not to measure a real cross-cloud network.

    Optimization · Configuration comparison

    The engine compares configuration trade-offs and surfaces the better option.

    intermediate
    Multi-Cloud
    AWS
    OCI
    Hybrid
    7 resources~20 min

    What will happen

    Compare a modeled hybrid path that connects AWS and OCI compute, databases, load balancers, and object storage. The goal is to inspect configuration trade-offs at a steady 1,200 RPS, not to measure a real cross-cloud network.

    Simulation details

    Modeled demand

    Steady Baseline · Hold · From step 0 · 1,200 RPS

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    OCI Web Application
    Explore a two-server web application behind an Oracle Cloud Infrastructure (OCI) Load Balancer, with an Autonomous Database and object storage. The scenario starts with a steady modeled 1,400 RPS baseline; it does not contact OCI or predict production performance.

    Educational · Guided scenario

    A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.

    beginner
    OCI
    Web App
    Autonomous DB
    5 resources~12 min

    What will happen

    Explore a two-server web application behind an Oracle Cloud Infrastructure (OCI) Load Balancer, with an Autonomous Database and object storage. The scenario starts with a steady modeled 1,400 RPS baseline; it does not contact OCI or predict production performance.

    Simulation details

    Modeled demand

    Steady Baseline · Hold · From step 0 · 1,400 RPS

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    CDN-Accelerated Web App
    Explore how a modeled CloudFront content delivery network (CDN) can keep a portion of a web surge at the edge while the dynamic path uses an Application Load Balancer, EC2, and RDS. The 75% edge-absorption figure is this scenario’s configured assumption, not a general CloudFront guarantee.

    Educational · Guided scenario

    A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.

    intermediate
    AWS
    CloudFront
    CDN
    EC2
    RDS
    Caching
    5 resources~15 min

    Starts at a healthy baseline — enable "Traffic Recovery" in the workspace when you want to run that optional phase.

    What will happen

    Explore how a modeled CloudFront content delivery network (CDN) can keep a portion of a web surge at the edge while the dynamic path uses an Application Load Balancer, EC2, and RDS. The 75% edge-absorption figure is this scenario’s configured assumption, not a general CloudFront guarantee.

    Simulation details

    Modeled demand

    Traffic Spike — CDN absorbs 75% at the edge · ramp · steps 0–25 · 2,000→8,000 RPS

    Optional phases

    Traffic Recovery · ramp · steps 60–80 · 8,000→2,000 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    GitHub Actions & Pages Database Failover (Unofficial)
    Real incident·GitHub-inspired (unaffiliated) · Aug 26, 2026
    Explore an unaffiliated, illustrative database-primary failover affecting modeled Actions and Pages paths. It is inspired by a status update, not an official GitHub product, exact forensic reconstruction, or set of published impact measurements.

    Chaos · Injected failure

    An intentional failure tests how the architecture responds.

    intermediate
    Azure
    GitHub
    Database Failover
    Real Incident
    Actions
    Pages
    8 resources~10 min

    What will happen

    Explore an unaffiliated, illustrative database-primary failover affecting modeled Actions and Pages paths. It is inspired by a status update, not an official GitHub product, exact forensic reconstruction, or set of published impact measurements.

    Simulation details

    Modeled demand

    Cutover Retry Storm — failover already underway · Hold · steps 0–8 · 620 RPS

    Backlog Drains — replica finishes taking over · ramp · steps 8–20 · 620→120 RPS

    CWM injects severe database overload at Primary Metadata Database · Start step 0 · Duration 30 steps · End step 30.

    LMI Comparison — ECS Fargate Arm
    Compare a modeled ECS Fargate service starting and replacing tasks with the same workload. This shows how ready capacity affects the response; it is not a test of Lambda Managed Instances, RDS Proxy behavior, or universal ECS limits.

    Optimization · Configuration comparison

    The engine compares configuration trade-offs and surfaces the better option.

    advanced
    AWS
    ECS
    Fargate
    LMI Comparison
    Lambda Comparison
    Readiness
    Rolling Deploy
    Warm Capacity
    Research
    Source Validation
    2 resources~10 min

    Starts at a healthy baseline — enable "Traffic Recovery — enable after step 60" in the workspace when you want to run that optional phase.

    What will happen

    Compare a modeled ECS Fargate service starting and replacing tasks with the same workload. This shows how ready capacity affects the response; it is not a test of Lambda Managed Instances, RDS Proxy behavior, or universal ECS limits.

    Simulation details

    Modeled demand

    Idle — establish the warm task baseline · Hold · steps 0–3 · 0 RPS

    Cold burst — watch tasks and application readiness + Warm hold — compare ready capacity and utilization · Hold · steps 3–28 · 2,400 RPS

    Quiesce — invoke rolling deploy here if desired · Hold · steps 28–31 · 0 RPS

    Re-burst — observe replacement and readiness + Post-deployment observation — hold the 2,400-RPS workload · Hold · steps 31–60 · 2,400 RPS

    Optional phases

    Traffic Recovery — enable after step 60 · ramp · steps 60–90 · 2,400→0 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    AWS Middle East Permanent Data Loss
    Real incident·AWS · Incidents beginning March 2026 · Permanent-loss confirmation September 15, 2026
    Explore how an application responds when a Bahrain region and part of a UAE deployment become unavailable while a remote backup remains available. CWM shows the difference between keeping services running and recovering lost data; it is an illustrative exercise, not a replay of a real AWS recovery.

    Chaos · Injected failure

    An intentional failure tests how the architecture responds.

    advanced
    AWS
    Real Incident
    Middle East
    Permanent Data Loss
    Geographic Backup
    8 resources~5 min

    What will happen

    Explore how an application responds when a Bahrain region and part of a UAE deployment become unavailable while a remote backup remains available. CWM shows the difference between keeping services running and recovering lost data; it is an illustrative exercise, not a replay of a real AWS recovery.

    Simulation details

    Modeled demand

    Healthy baseline — illustrative demand · Hold · From step 0 · 100 RPS

    CWM injects severe region outage at aws region me-south-1 (Bahrain) · Start step 10.

    CWM injects moderate network latency at UAE AZ 1 — Illustrative Surviving App · Start step 10.

    CWM injects severe availability-zone outage at zone mec1-az2 · Start step 10.

    CWM injects moderate network latency at UAE AZ 3 — Illustrative Surviving Data · Start step 10.

    CWM injects severe permanent data loss at aws region me-south-1 (Bahrain) · Start step 20 · Terminal data loss; infrastructure recovery cannot restore original data.

    CWM injects severe permanent data loss at zone mec1-az2 · Start step 20 · Terminal data loss; infrastructure recovery cannot restore original data.

    EKS Spot Interruption Migration
    Practice a deterministic Amazon EKS Spot interruption. Demand enters the configured warning band before the notice, and the scenario tests whether four workloads can be rescheduled before the two-minute simulated deadline.

    Chaos · Injected interruption

    A scheduled AWS Spot interruption tests whether affected workloads become ready before the termination deadline.

    advanced
    AWS
    EKS
    Kubernetes
    Spot
    Interruption
    Rescheduling
    3 resources~8 min

    What will happen

    Practice a deterministic Amazon EKS Spot interruption. Demand enters the configured warning band before the notice, and the scenario tests whether four workloads can be rescheduled before the two-minute simulated deadline.

    Simulation details

    Modeled demand

    Steady application traffic · Hold · From step 0 · 400 RPS

    Pre-interruption demand ramp — enter warning · ramp · steps 6–10 · 400→1,100 RPS

    Warning hold — stay above threshold through notice · Hold · steps 10–52 · 1,100 RPS

    Post-migration recovery — return to healthy baseline · ramp · steps 52–55 · 1,100→400 RPS

    Interruption: EKS Spot Worker Fleet · AWS Spot interruption · Start step 10 · 120-second deadline at 3 simulated seconds/step; 4 workloads must be rescheduled.

    CloudFront SPA + API Path Routing — The 404 Trap
    Understand how a global single-page application (SPA) error rewrite can change API status codes when the SPA and API share one CloudFront distribution. This is a configuration walkthrough, not a load failure.

    Educational · Guided scenario

    A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.

    intermediate
    AWS
    CloudFront
    ALB
    S3
    SPA
    API Design
    Architecture
    6 resources~10 min

    What will happen

    Understand how a global single-page application (SPA) error rewrite can change API status codes when the SPA and API share one CloudFront distribution. This is a configuration walkthrough, not a load failure.

    Simulation details

    Modeled demand

    Steady API load — routing bug is live, not a load problem · Hold · From step 0 · 500 RPS

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    AKS Multi-Fault Cascade
    Use this as a manual Azure Kubernetes Service (AKS) multi-fault exercise. The seed starts healthy and waits for the user to inject a zone fault; it does not self-trigger the cascade.

    Chaos · Injected failure

    An intentional failure tests how the architecture responds.

    advanced
    Azure
    AKS
    Kubernetes
    Chaos
    Multi-Fault
    5 resources~30 min

    Starts at a healthy baseline — enable "Ramp to Autoscale — observe recovery behavior after fault injection" in the workspace when you want to run that optional phase.

    What will happen

    Use this as a manual Azure Kubernetes Service (AKS) multi-fault exercise. The seed starts healthy and waits for the user to inject a zone fault; it does not self-trigger the cascade.

    Simulation details

    Modeled demand

    Wave Baseline — correlated faults cascade at steady load · wave · From step 0 · configured traffic

    Optional phases

    Ramp to Autoscale — observe recovery behavior after fault injection · ramp · steps 60–100 · 1,500→8,000 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    API Subdomain Split — SPA on CDN, API on ALB
    See the routing separation that prevents a single-page application (SPA) fallback from rewriting API errors. This is a configuration walkthrough, not a live HTTP or CloudFront test.

    Educational · Guided scenario

    A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.

    beginner
    AWS
    CloudFront
    ALB
    S3
    SPA
    API Design
    Architecture
    6 resources~10 min

    What will happen

    See the routing separation that prevents a single-page application (SPA) fallback from rewriting API errors. This is a configuration walkthrough, not a live HTTP or CloudFront test.

    Simulation details

    Modeled demand

    Steady API baseline — all resources healthy, HTTP semantics intact · Hold · From step 0 · 150 RPS

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Kubernetes App on GKE
    Explore a containerized application on Google Kubernetes Engine (GKE) Autopilot with Memorystore for Redis and Cloud Pub/Sub. The model focuses on how the configured Kubernetes capacity responds to a 2,500-to-15,000 RPS ramp.

    Educational · Guided scenario

    A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.

    intermediate
    GCP
    Kubernetes
    GKE
    Redis
    Pub/Sub
    Cache
    Queue
    5 resources~18 min

    Starts at a healthy baseline — enable "Traffic Recovery — see GKE scale-in" in the workspace when you want to run that optional phase.

    What will happen

    Explore a containerized application on Google Kubernetes Engine (GKE) Autopilot with Memorystore for Redis and Cloud Pub/Sub. The model focuses on how the configured Kubernetes capacity responds to a 2,500-to-15,000 RPS ramp.

    Simulation details

    Modeled demand

    Traffic Ramp — watch GKE autoscale · ramp · steps 0–30 · 2,500→15,000 RPS

    Optional phases

    Traffic Recovery — see GKE scale-in · ramp · steps 60–90 · 15,000→2,500 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Redis Cache Crash & Recovery
    Start from a healthy 1,000 RPS baseline, then manually inject a supported Redis failure to explore how cache misses affect PostgreSQL connection and latency pressure. No cache failure or recovery is scheduled by default.

    Educational · Guided scenario

    A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.

    intermediate
    AWS
    Redis
    Cache
    Failure
    Recovery
    Cache Stampede
    ElastiCache
    5 resources~15 min

    What will happen

    Start from a healthy 1,000 RPS baseline, then manually inject a supported Redis failure to explore how cache misses affect PostgreSQL connection and latency pressure. No cache failure or recovery is scheduled by default.

    Simulation details

    Modeled demand

    Moderate Baseline Traffic · Hold · From step 0 · 1,000 RPS

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Event-Driven Azure Microservices
    See an Azure application handle rising demand through Kubernetes, a cache, and a queue. The simulation shows how the configured capacity and autoscaling respond; it does not contact Azure or predict production performance.

    Educational · Guided scenario

    A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.

    advanced
    Azure
    Kubernetes
    AKS
    Redis
    Service Bus
    Cache
    Queue
    5 resources~20 min

    Starts at a healthy baseline — enable "Traffic Recovery — see AKS scale-in" in the workspace when you want to run that optional phase.

    What will happen

    See an Azure application handle rising demand through Kubernetes, a cache, and a queue. The simulation shows how the configured capacity and autoscaling respond; it does not contact Azure or predict production performance.

    Simulation details

    Modeled demand

    Traffic Ramp — watch AKS autoscale · ramp · steps 0–30 · 2,500→15,000 RPS

    Optional phases

    Traffic Recovery — see AKS scale-in · ramp · steps 60–90 · 15,000→2,500 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    DigitalOcean Kubernetes (DOKS)
    See how a DigitalOcean Kubernetes (DOKS) stack with a Load Balancer, Managed Redis, and Managed Kafka responds to rising demand. The scenario is an illustrative capacity walkthrough, not a claim about every DOKS cluster.

    Educational · Guided scenario

    A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.

    intermediate
    DigitalOcean
    Kubernetes
    DOKS
    Cache
    Queue
    Redis
    Kafka
    4 resources~18 min

    Starts at a healthy baseline — enable "Traffic Recovery — see DOKS scale-in" in the workspace when you want to run that optional phase.

    What will happen

    See how a DigitalOcean Kubernetes (DOKS) stack with a Load Balancer, Managed Redis, and Managed Kafka responds to rising demand. The scenario is an illustrative capacity walkthrough, not a claim about every DOKS cluster.

    Simulation details

    Modeled demand

    Traffic Ramp — watch DOKS autoscale · ramp · steps 0–30 · 2,000→15,000 RPS

    Optional phases

    Traffic Recovery — see DOKS scale-in · ramp · steps 60–90 · 15,000→2,000 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    SQS Queue Backlog Saturation
    Find the modeled point where an Amazon Simple Queue Service (SQS) worker fleet cannot keep up with an event burst. The scenario is a capacity exercise; it does not inject an outage or automatically add consumers.

    Predictive · Capacity simulation

    Modeled load and capacity constraints reveal bottlenecks or failure.

    intermediate
    AWS
    SQS
    Queue
    Failure
    Recovery
    Backlog
    Workers
    6 resources~15 min

    Starts at a healthy baseline — enable "Event Burst — overwhelm workers" in the workspace when you want to run that optional phase.

    What will happen

    Find the modeled point where an Amazon Simple Queue Service (SQS) worker fleet cannot keep up with an event burst. The scenario is a capacity exercise; it does not inject an outage or automatically add consumers.

    Simulation details

    Modeled demand

    Steady Producer Load · Hold · From step 0 · 1,200 RPS

    Optional phases

    Event Burst — overwhelm workers · ramp · steps 0–20 · 1,200→5,000 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Serverless API (Lambda + DynamoDB)
    Model an AWS serverless API with an Application Load Balancer, Lambda, and DynamoDB. A quiet 100 RPS period precedes a 100→2,000 RPS ramp so the simulation can expose the configured cold-start and capacity response.

    Predictive · Capacity simulation

    Modeled load and capacity constraints reveal bottlenecks or failure.

    intermediate
    AWS
    Lambda
    Serverless
    DynamoDB
    Cold Start
    3 resources~15 min

    What will happen

    Model an AWS serverless API with an Application Load Balancer, Lambda, and DynamoDB. A quiet 100 RPS period precedes a 100→2,000 RPS ramp so the simulation can expose the configured cold-start and capacity response.

    Simulation details

    Modeled demand

    Quiet period — Lambda idles · Hold · From step 0 · 100 RPS

    Traffic burst — cold starts fire · ramp · steps 20–35 · 100→2,000 RPS

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    CloudFront Edge Outage
    Real incident·AWS · Jul 16, 2026
    See the modeled origin overload when a CloudFront edge layer stops serving cached content. The July 16, 2026 incident is provenance for the exercise; the simulation is not a replay or a claim about every CloudFront edge.

    Chaos · Injected failure

    An intentional failure tests how the architecture responds.

    intermediate
    AWS
    CloudFront
    CDN
    Outage
    Incident Response
    EC2
    RDS
    5 resources~15 min

    Starts at a healthy baseline — enable "Traffic Drain — reduce load during incident response" in the workspace when you want to run that optional phase.

    What will happen

    See the modeled origin overload when a CloudFront edge layer stops serving cached content. The July 16, 2026 incident is provenance for the exercise; the simulation is not a replay or a claim about every CloudFront edge.

    Simulation details

    Modeled demand

    Steady bypass load — origin absorbing 100% of traffic · Hold · From step 0 · 2,000 RPS

    Optional phases

    Traffic Drain — reduce load during incident response · ramp · steps 60–80 · 2,000→500 RPS · enable in the workspace

    CWM injects severe availability-zone outage at zone us-east-1-regional · Start step 0 · Duration 9999 steps · End step 9999.

    Zombie Infrastructure — OCI
    Real incident·OCI · Aug 4, 2026
    Explore OCI’s smaller but non-zero modeled idle footprint. The ~$90/month amount and the 10 Mbps load-balancer minimum are scenario assumptions; the comparison is not a live OCI quote.

    Optimization · Configuration comparison

    The engine compares configuration trade-offs and surfaces the better option.

    beginner
    OCI
    Cost
    Idle Resources
    Zombie Infrastructure
    FinOps
    8 resources~10 min

    What will happen

    Explore OCI’s smaller but non-zero modeled idle footprint. The ~$90/month amount and the 10 Mbps load-balancer minimum are scenario assumptions; the comparison is not a live OCI quote.

    Simulation details

    Modeled demand

    Zero traffic — the bill keeps running anyway · Hold · From step 0 · 0 RPS

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    LMI Comparison — Ordinary Lambda Approximation
    Compare a Lambda application that must start new execution capacity with one that begins with capacity already available. The scenario shows how startup delay affects a fixed workload. It is a simplified CWM comparison, not a reproduction of AWS Lambda Managed Instances or published benchmark results.

    Optimization · Configuration comparison

    The engine compares configuration trade-offs and surfaces the better option.

    advanced
    AWS
    Lambda
    Ordinary Lambda
    LMI Approximation
    ECS Comparison
    Readiness
    Warm Capacity
    Research
    Source Validation
    2 resources~10 min

    Starts at a healthy baseline — enable "Traffic Recovery — enable after step 60" in the workspace when you want to run that optional phase.

    What will happen

    Compare a Lambda application that must start new execution capacity with one that begins with capacity already available. The scenario shows how startup delay affects a fixed workload. It is a simplified CWM comparison, not a reproduction of AWS Lambda Managed Instances or published benchmark results.

    Simulation details

    Modeled demand

    Idle — establish the cold baseline · Hold · steps 0–3 · 0 RPS

    Cold burst — watch application readiness reject traffic + Warm hold — compare ready capacity and utilization · Hold · steps 3–28 · 2,400 RPS

    Quiesce — clear traffic before the re-burst · Hold · steps 28–31 · 0 RPS

    Re-burst — observe readiness on a second cold start + Post-deployment observation — hold the 2,400-RPS workload · Hold · steps 31–60 · 2,400 RPS

    Optional phases

    Traffic Recovery — enable after step 60 · ramp · steps 60–90 · 2,400→0 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Zombie Infrastructure — AWS
    Real incident·AWS · Aug 4, 2026
    Find the residual AWS charges that remain in this configured zero-traffic topology. The ~$310/month figure is the scenario’s illustrative starting estimate, not a bill or price forecast.

    Optimization · Configuration comparison

    The engine compares configuration trade-offs and surfaces the better option.

    beginner
    AWS
    Cost
    Idle Resources
    Zombie Infrastructure
    FinOps
    13 resources~10 min

    What will happen

    Find the residual AWS charges that remain in this configured zero-traffic topology. The ~$310/month figure is the scenario’s illustrative starting estimate, not a bill or price forecast.

    Simulation details

    Modeled demand

    Zero traffic — the bill keeps running anyway · Hold · From step 0 · 0 RPS

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Zombie Infrastructure — GCP
    Real incident·GCP · Aug 4, 2026
    Inspect modeled GCP residual charges in a decommissioned, zero-traffic project. The ~$120/month amount is an illustrative configuration estimate; the exercise is not a live GCP invoice.

    Optimization · Configuration comparison

    The engine compares configuration trade-offs and surfaces the better option.

    beginner
    GCP
    Cost
    Idle Resources
    Zombie Infrastructure
    FinOps
    10 resources~10 min

    What will happen

    Inspect modeled GCP residual charges in a decommissioned, zero-traffic project. The ~$120/month amount is an illustrative configuration estimate; the exercise is not a live GCP invoice.

    Simulation details

    Modeled demand

    Zero traffic — the bill keeps running anyway · Hold · From step 0 · 0 RPS

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Zombie Infrastructure — Azure
    Real incident·Azure · Aug 4, 2026
    See which Azure resources can continue contributing to a modeled residual bill after application traffic reaches zero. The ~$175/month figure is an illustrative starting estimate, not an Azure invoice.

    Optimization · Configuration comparison

    The engine compares configuration trade-offs and surfaces the better option.

    beginner
    Azure
    Cost
    Idle Resources
    Zombie Infrastructure
    FinOps
    10 resources~10 min

    What will happen

    See which Azure resources can continue contributing to a modeled residual bill after application traffic reaches zero. The ~$175/month figure is an illustrative starting estimate, not an Azure invoice.

    Simulation details

    Modeled demand

    Zero traffic — the bill keeps running anyway · Hold · From step 0 · 0 RPS

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    GCP SaaS Web App
    Compare the modeled cost and capacity response of a Google Cloud SaaS stack with Cloud CDN, App Engine, Cloud SQL, and Cloud Operations Suite. The 80% cache assumption is specific to this scenario, not a promise about real traffic.

    Optimization · Configuration comparison

    The engine compares configuration trade-offs and surfaces the better option.

    beginner
    GCP
    SaaS
    App Engine
    Cloud CDN
    Cloud SQL
    Cost
    4 resources~10 min

    Starts at a healthy baseline — enable "Traffic Recovery — see the stack settle back down" in the workspace when you want to run that optional phase.

    What will happen

    Compare the modeled cost and capacity response of a Google Cloud SaaS stack with Cloud CDN, App Engine, Cloud SQL, and Cloud Operations Suite. The 80% cache assumption is specific to this scenario, not a promise about real traffic.

    Simulation details

    Modeled demand

    Business-hours ramp — watch App Engine climb into the warning zone · ramp · steps 0–60 · 90→680 RPS

    Optional phases

    Traffic Recovery — see the stack settle back down · ramp · steps 60–90 · 680→90 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Inference Autoscaling
    Study how a modeled GKE GPU inference pool spreads fixed node cost across tokens as utilization changes. The T4, utilization, autoscaling, and cost-per-million-token relationships are this configured model, not a forecast for every large-language-model workload.

    Predictive · Capacity simulation

    Modeled load and capacity constraints reveal bottlenecks or failure.

    intermediate
    GCP
    Kubernetes
    GPU
    Inference
    Autoscaling
    Cost
    3 resources~15 min

    Starts at a healthy baseline — enable "Traffic Recovery — see per-token cost climb as GPUs idle" in the workspace when you want to run that optional phase.

    What will happen

    Study how a modeled GKE GPU inference pool spreads fixed node cost across tokens as utilization changes. The T4, utilization, autoscaling, and cost-per-million-token relationships are this configured model, not a forecast for every large-language-model workload.

    Simulation details

    Modeled demand

    Inference ramp — watch cost/M tokens fall as GPU utilization rises · ramp · steps 0–40 · 200→2,000 RPS

    Optional phases

    Traffic Recovery — see per-token cost climb as GPUs idle · ramp · steps 60–90 · 2,000→200 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    GCP Multi-Region Failover — Cloud DNS & Global Load Balancer
    Real incident·GCP · Jun 2, 2019
    Examine a bounded GCP routing-layer failure: Cloud DNS and the Global Load Balancer are modeled as unavailable while the east path remains a possible destination. The 2019 incident supplies provenance; this is not a historical timing replay.

    Chaos · Injected failure

    An intentional failure tests how the architecture responds.

    intermediate
    GCP
    Cloud DNS
    Global Load Balancer
    Multi-Region
    Failover
    GKE
    Resilience
    6 resources~15 min

    Global LB (anycast) · Cloud SQL auto-failover

    Starts at a healthy baseline — enable "East Scale-Out — simulate full central drain to us-east1" in the workspace when you want to run that optional phase.

    What will happen

    Examine a bounded GCP routing-layer failure: Cloud DNS and the Global Load Balancer are modeled as unavailable while the east path remains a possible destination. The 2019 incident supplies provenance; this is not a historical timing replay.

    Simulation details

    Modeled demand

    Mid-failover baseline — Global LB draining central, loading east · Hold · From step 0 · 500 RPS

    Optional phases

    East Scale-Out — simulate full central drain to us-east1 · ramp · steps 60–90 · 500→3,000 RPS · enable in the workspace

    CWM injects severe availability-zone outage at zone usc1-zone-a · Start step 0 · Duration 9999 steps · End step 9999.

    Zombie Infrastructure — DigitalOcean
    Real incident·DigitalOcean · Aug 4, 2026
    Inspect DigitalOcean’s modeled residual charges after traffic reaches zero. The ~$90/month and $10/month load-balancer values are illustrative scenario inputs; they are not a live invoice or universal DigitalOcean pricing claim.

    Optimization · Configuration comparison

    The engine compares configuration trade-offs and surfaces the better option.

    beginner
    DigitalOcean
    Cost
    Idle Resources
    Zombie Infrastructure
    FinOps
    7 resources~10 min

    What will happen

    Inspect DigitalOcean’s modeled residual charges after traffic reaches zero. The ~$90/month and $10/month load-balancer values are illustrative scenario inputs; they are not a live invoice or universal DigitalOcean pricing claim.

    Simulation details

    Modeled demand

    Zero traffic — the bill keeps running anyway · Hold · From step 0 · 0 RPS

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    AWS SaaS Web App
    Compare the modeled cost and capacity response of an AWS SaaS stack with CloudFront, App Runner, Aurora PostgreSQL, and CloudWatch. The 80% cache assumption and warning-band outcome are scenario configuration, not a universal AWS result.

    Optimization · Configuration comparison

    The engine compares configuration trade-offs and surfaces the better option.

    beginner
    AWS
    SaaS
    App Runner
    CloudFront
    Aurora
    Cost
    4 resources~10 min

    Starts at a healthy baseline — enable "Traffic Recovery — see the stack settle back down" in the workspace when you want to run that optional phase.

    What will happen

    Compare the modeled cost and capacity response of an AWS SaaS stack with CloudFront, App Runner, Aurora PostgreSQL, and CloudWatch. The 80% cache assumption and warning-band outcome are scenario configuration, not a universal AWS result.

    Simulation details

    Modeled demand

    Business-hours ramp — watch App Runner climb into the warning zone · ramp · steps 0–60 · 120→850 RPS

    Optional phases

    Traffic Recovery — see the stack settle back down · ramp · steps 60–90 · 850→120 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    DigitalOcean Starter Web App
    Explore a small DigitalOcean web stack: a Load Balancer, one Droplet, and Managed PostgreSQL. It starts at a modeled 400 RPS baseline to show the topology and capacity trade-off; the result is not a promise of a specific monthly bill or production throughput.

    Educational · Guided scenario

    A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.

    beginner
    DigitalOcean
    Beginner
    Web App
    Droplet
    PostgreSQL
    3 resources~10 min

    What will happen

    Explore a small DigitalOcean web stack: a Load Balancer, one Droplet, and Managed PostgreSQL. It starts at a modeled 400 RPS baseline to show the topology and capacity trade-off; the result is not a promise of a specific monthly bill or production throughput.

    Simulation details

    Modeled demand

    Steady Baseline · Hold · From step 0 · 400 RPS

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Microservices with Redis and SQS
    Learn how a Redis cache and an Amazon Simple Queue Service (SQS) queue change the modeled load on an AWS microservices stack as demand rises. This walkthrough shows configuration behavior; it does not claim a universal cache-hit or queue-drain rate.

    Educational · Guided scenario

    A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.

    intermediate
    AWS
    Redis
    SQS
    Microservices
    Cache
    Queue
    6 resources~15 min

    Starts at a healthy baseline — enable "Traffic Recovery" in the workspace when you want to run that optional phase.

    What will happen

    Learn how a Redis cache and an Amazon Simple Queue Service (SQS) queue change the modeled load on an AWS microservices stack as demand rises. This walkthrough shows configuration behavior; it does not claim a universal cache-hit or queue-drain rate.

    Simulation details

    Modeled demand

    Traffic Spike — watch cache absorb the burst · ramp · steps 0–20 · 1,000→5,000 RPS

    Optional phases

    Traffic Recovery · ramp · steps 60–80 · 5,000→1,000 RPS · enable in the workspace

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Idle GPU Infrastructure
    Show the modeled cost of an EKS GPU pool with no inference traffic. With zero requests there are no generated tokens, so cost per million tokens is undefined/infinite while configured GPU capacity continues to carry cost; the result is a waste-detection exercise, not a provider bill.

    Optimization · Configuration comparison

    The engine compares configuration trade-offs and surfaces the better option.

    beginner
    AWS
    Kubernetes
    GPU
    Inference
    Idle
    Zombie
    Cost
    2 resources~10 min

    What will happen

    Show the modeled cost of an EKS GPU pool with no inference traffic. With zero requests there are no generated tokens, so cost per million tokens is undefined/infinite while configured GPU capacity continues to carry cost; the result is a waste-detection exercise, not a provider bill.

    Simulation details

    Modeled demand

    Zero traffic — the GPUs bill anyway · ramp · steps 0–60 · 0 RPS

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    GitHub-Inspired Cascading Retry Storm (Unofficial)
    Real incident·GitHub-inspired (unaffiliated) · Aug 17, 2026
    Investigate a GitHub-inspired, unaffiliated retry cascade with fixed external traffic and modeled internal retry amplification. Compare the configured unprotected and protected dependency policies without treating the result as an official GitHub reconstruction.

    Predictive · Capacity simulation

    Modeled load and capacity constraints reveal bottlenecks or failure.

    advanced
    Azure
    Advanced
    Reliability
    Retry Storm
    Authentication
    AI
    Real Incident
    14 resources~20 min

    What will happen

    Investigate a GitHub-inspired, unaffiliated retry cascade with fixed external traffic and modeled internal retry amplification. Compare the configured unprotected and protected dependency policies without treating the result as an official GitHub reconstruction.

    Simulation details

    Modeled demand

    Healthy 8K RPS baseline and recovery hold · Hold · steps 0–60 · 8,000 RPS

    External traffic stays at 8,000 RPS; when modeled failures trigger the configured dependency retries, effective internal request volume can be higher. Retry attempts are internal, not additional external traffic.

    No failure is scheduled; any degradation should emerge from modeled load and capacity.

    Azure Front Door HA — Active Geo-Replication Through a Front Door Outage
    Real incident·Azure · Oct 9, 2025
    Learn why a healthy multi-region backend does not protect against failure of the global proxy in front of it. The 2025 Front Door incident is context; the Traffic Manager path is a modeled, optional bypass exercise.

    Chaos · Injected failure

    An intentional failure tests how the architecture responds.

    intermediate
    Azure
    Front Door
    Traffic Manager
    Multi-Region
    Active Geo-Replication
    HA
    Resilience
    5 resources~12 min

    Front Door (anycast) · Active Geo-Replication auto-failover

    Starts at a healthy baseline — enable "Traffic Manager failover — direct bypass traffic floods West US origin" in the workspace when you want to run that optional phase.

    What will happen

    Learn why a healthy multi-region backend does not protect against failure of the global proxy in front of it. The 2025 Front Door incident is context; the Traffic Manager path is a modeled, optional bypass exercise.

    Simulation details

    Modeled demand

    Stale-DNS trickle — only clients with cached IPs reaching backends directly · Hold · From step 0 · 50 RPS

    Optional phases

    Traffic Manager failover — direct bypass traffic floods West US origin · ramp · steps 60–90 · 50→3,000 RPS · enable in the workspace

    CWM injects severe availability-zone outage at zone eus-zone-fd · Start step 0 · Duration 9999 steps · End step 9999.