Educational · Guided scenario
A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.
What will happen
See how provider limits, retries, and queueing affect a batch of AWS EC2 changes at the selected concurrency. This is a catalog-backed provider-limit simulation separate from the workspace resource state.
CWM models 240 EC2 create, update, or delete requests at the selected concurrency of 8. Some requests may be slowed, queued, or retried. The results show how provider limits, retries, and queueing affect the batch.
No AWS APIs are called and no EC2 instances are created.
Predictive · Capacity simulation
Modeled load and capacity constraints reveal bottlenecks or failure.
Starts at a healthy baseline — enable "Traffic Recovery — watch Fargate scale to zero" in the workspace when you want to run that optional phase.
What will happen
See a modeled Amazon ECS Fargate service add tasks during a 30→500 RPS ramp, respect its configured fleet ceiling, and account for task startup, bounded queueing, vCPU, and memory billing.
Modeled demand
Traffic Ramp — watch Fargate tasks start · ramp · steps 0–30 · 30→500 RPS
Optional phases
Traffic Recovery — watch Fargate scale to zero · ramp · steps 60–90 · 500→0 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Optimization · Configuration comparison
The engine compares configuration trade-offs and surfaces the better option.
Starts at a healthy baseline — enable "Traffic Recovery — see the stack settle back down" in the workspace when you want to run that optional phase.
What will happen
Compare the modeled cost and capacity response of an Azure SaaS stack with Front Door, App Service, PostgreSQL Flexible Server, Monitor, and Defender. The 75% cache assumption and warning-band outcome belong to this configuration, not all Azure SaaS workloads.
Modeled demand
Business-hours ramp — watch App Service climb into the warning zone · ramp · steps 0–60 · 200→1,550 RPS
Optional phases
Traffic Recovery — see the stack settle back down · ramp · steps 60–90 · 1,550→200 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Predictive · Capacity simulation
Modeled load and capacity constraints reveal bottlenecks or failure.
Starts at a healthy baseline — enable "Message Burst — overwhelm subscribers" in the workspace when you want to run that optional phase.
What will happen
Model a Google Cloud Pub/Sub subscriber fleet under rising event demand. The focus is backlog, subscriber capacity, and delivery pressure—not a live Pub/Sub outage or a guaranteed acknowledgement deadline.
Modeled demand
Steady Publisher Load · Hold · From step 0 · 1,200 RPS
Optional phases
Message Burst — overwhelm subscribers · ramp · steps 0–20 · 1,200→5,000 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Predictive · Capacity simulation
Modeled load and capacity constraints reveal bottlenecks or failure.
Starts at a healthy baseline — enable "Order Burst — overwhelm consumers" in the workspace when you want to run that optional phase.
What will happen
Model an Azure Service Bus order queue as demand rises beyond the configured consumer capacity. This is a queueing and throughput exercise; it does not call Azure or imply that Premium upgrade, auto-forwarding, or consumer scaling happens automatically.
Modeled demand
Steady Order Volume · Hold · From step 0 · 1,200 RPS
Optional phases
Order Burst — overwhelm consumers · ramp · steps 0–20 · 1,200→5,000 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Predictive · Capacity simulation
Modeled load and capacity constraints reveal bottlenecks or failure.
What will happen
Find where a modeled upload server and transcoding worker become capacity-bound during a tenfold launch surge. This is a bounded planning simulation, not a production forecast or a guarantee of an exact concurrency threshold.
Modeled demand
Launch Day Surge (10×) · ramp · steps 0–30 · 25→250 RPS
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Educational · Guided scenario
A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.
Starts at a healthy baseline — enable "Traffic Recovery — see scale-in" in the workspace when you want to run that optional phase.
What will happen
See how a load-balanced web tier responds as demand rises, and compare scale-out with the optional later scale-in phase. This is an illustrative capacity walkthrough, not a forecast of a particular AWS deployment.
Modeled demand
Traffic Ramp — watch autoscaling fire · ramp · steps 0–40 · 800→4,000 RPS
Optional phases
Traffic Recovery — see scale-in · ramp · steps 60–100 · 4,000→800 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Chaos · Injected failure
An intentional failure tests how the architecture responds.
Starts at a healthy baseline — enable "Traffic Recovery" in the workspace when you want to run that optional phase.
What will happen
Observe an injected RDS primary overload in a two-AZ application with a standby replica. The card describes a modeled failover path; it is not a live AWS RDS event or a promise of a provider RTO.
Modeled demand
Traffic Ramp — watch RDS Primary reach its limit · ramp · steps 0–25 · 3,000→8,000 RPS
Optional phases
Traffic Recovery · ramp · steps 60–80 · 8,000→2,500 RPS · enable in the workspace
CWM injects severe database overload at RDS Primary · Start step 15 · Duration 35 steps · End step 50.
Chaos · Injected failure
An intentional failure tests how the architecture responds.
Route 53 (DNS) · RDS Multi-AZ auto-failover
Starts at a healthy baseline — enable "Traffic Recovery — watch us-east-1 scale in as us-west-2 recovers" in the workspace when you want to run that optional phase.
What will happen
Walk through a modeled AWS multi-region DNS failover. Route 53 health checks, DNS caching, and the east fleet’s capacity are the teaching model; the July 2026 incident is context, not proof that this run reproduces its duration.
Modeled demand
Mid-failover baseline — Route 53 draining west, loading east · Hold · From step 0 · 120 RPS
East Scale-Out — DNS TTLs expire, us-east-1 absorbs us-west-2 drain · ramp · steps 10–40 · 120→1,500 RPS
Optional phases
Traffic Recovery — watch us-east-1 scale in as us-west-2 recovers · ramp · steps 60–90 · 1,500→120 RPS · enable in the workspace
CWM injects severe availability-zone outage at zone usw2-az1 · Start step 0 · Duration 9999 steps · End step 9999.
CWM injects severe availability-zone outage at zone usw2-az2 · Start step 0 · Duration 9999 steps · End step 9999.
Chaos · Injected failure
An intentional failure tests how the architecture responds.
Global LB (anycast) · Paxos consensus (zero RPO)
What will happen
See how an application responds when all configured Cloud Spanner replicas lose their serving path. CWM models the resulting service interruption and recovery choices; it does not reproduce the real incident’s cause, duration, or customer impact.
Modeled demand
Pre-outage baseline — 500 RPS hitting Spanner hard · Hold · From step 0 · 500 RPS
CWM injects severe availability-zone outage at zone usc1-zone-b · Start step 0 · Duration 9999 steps · End step 9999.
CWM injects severe availability-zone outage at zone use1-zone-c · Start step 0 · Duration 9999 steps · End step 9999.
CWM injects severe availability-zone outage at zone use4-zone-a · Start step 0 · Duration 9999 steps · End step 9999.
Chaos · Injected failure
An intentional failure tests how the architecture responds.
Traffic Mgmt (DNS) · Data Guard manual SWITCHOVER
Starts at a healthy baseline — enable "Scale test — confirm Phoenix absorbs full production load" in the workspace when you want to run that optional phase.
What will happen
Explore how traffic steering and a standby database can help an OCI application during a regional network failure. Promotion and connection changes are modeled choices for the user, not an automatic switchover or a guarantee of zero data loss.
Modeled demand
us-ashburn-1 degraded — Traffic Management routing to us-phoenix-1 · Hold · From step 0 · 800 RPS
Optional phases
Scale test — confirm Phoenix absorbs full production load · ramp · steps 60–90 · 800→4,000 RPS · enable in the workspace
CWM injects severe availability-zone outage at zone iad-ad-1 · Start step 0 · Duration 9999 steps · End step 9999.
Chaos · Injected failure
An intentional failure tests how the architecture responds.
Floating IP (DNS) · pg_logical manual promotion
Starts at a healthy baseline — enable "Scale test — confirm SFO3 Droplet handles full NYC3 production load" in the workspace when you want to run that optional phase.
What will happen
See how an application can keep serving when its New York site fails and traffic moves to San Francisco. The database still needs manual cross-region recovery, so this exercise shows the difference between routing traffic and restoring data; it does not promise automatic failover or a provider SLA.
Modeled demand
NYC3 degraded — DNS update in progress, traffic routing to SFO3 · Hold · From step 0 · 600 RPS
Optional phases
Scale test — confirm SFO3 Droplet handles full NYC3 production load · ramp · steps 60–90 · 600→2,500 RPS · enable in the workspace
CWM injects severe availability-zone outage at zone nyc3-az1 · Start step 0 · Duration 9999 steps · End step 9999.
Chaos · Injected failure
An intentional failure tests how the architecture responds.
Starts at a healthy baseline — enable "Traffic Recovery — watch services stabilize" in the workspace when you want to run that optional phase.
What will happen
Study a bounded, incident-inspired AWS regional routing failure that is structural rather than demand-driven. The historical date and incident basis provide context; the card’s steps and recovery behavior are an illustrative CWM model, not a historical replay.
Modeled demand
Steady outage load — structural crisis at low baseline · Hold · From step 0 · 300 RPS
Optional phases
Traffic Recovery — watch services stabilize · ramp · steps 60–90 · 300→2,000 RPS · enable in the workspace
CWM injects severe availability-zone outage at zone us-west-2-regional · Start step 0 · Duration 45 steps · End step 45.
CWM injects severe availability-zone outage at zone us-west-2-compute · Start step 0 · Duration 45 steps · End step 45.
CWM injects severe database overload at RDS MySQL Primary · Start step 0 · Duration 40 steps · End step 40.
Optimization · Configuration comparison
The engine compares configuration trade-offs and surfaces the better option.
What will happen
Compare a modeled hybrid path that connects AWS and OCI compute, databases, load balancers, and object storage. The goal is to inspect configuration trade-offs at a steady 1,200 RPS, not to measure a real cross-cloud network.
Modeled demand
Steady Baseline · Hold · From step 0 · 1,200 RPS
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Educational · Guided scenario
A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.
What will happen
Explore a two-server web application behind an Oracle Cloud Infrastructure (OCI) Load Balancer, with an Autonomous Database and object storage. The scenario starts with a steady modeled 1,400 RPS baseline; it does not contact OCI or predict production performance.
Modeled demand
Steady Baseline · Hold · From step 0 · 1,400 RPS
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Educational · Guided scenario
A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.
Starts at a healthy baseline — enable "Traffic Recovery" in the workspace when you want to run that optional phase.
What will happen
Explore how a modeled CloudFront content delivery network (CDN) can keep a portion of a web surge at the edge while the dynamic path uses an Application Load Balancer, EC2, and RDS. The 75% edge-absorption figure is this scenario’s configured assumption, not a general CloudFront guarantee.
Modeled demand
Traffic Spike — CDN absorbs 75% at the edge · ramp · steps 0–25 · 2,000→8,000 RPS
Optional phases
Traffic Recovery · ramp · steps 60–80 · 8,000→2,000 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Chaos · Injected failure
An intentional failure tests how the architecture responds.
What will happen
Explore an unaffiliated, illustrative database-primary failover affecting modeled Actions and Pages paths. It is inspired by a status update, not an official GitHub product, exact forensic reconstruction, or set of published impact measurements.
Modeled demand
Cutover Retry Storm — failover already underway · Hold · steps 0–8 · 620 RPS
Backlog Drains — replica finishes taking over · ramp · steps 8–20 · 620→120 RPS
CWM injects severe database overload at Primary Metadata Database · Start step 0 · Duration 30 steps · End step 30.
Optimization · Configuration comparison
The engine compares configuration trade-offs and surfaces the better option.
Starts at a healthy baseline — enable "Traffic Recovery — enable after step 60" in the workspace when you want to run that optional phase.
What will happen
Compare a modeled ECS Fargate service starting and replacing tasks with the same workload. This shows how ready capacity affects the response; it is not a test of Lambda Managed Instances, RDS Proxy behavior, or universal ECS limits.
Modeled demand
Idle — establish the warm task baseline · Hold · steps 0–3 · 0 RPS
Cold burst — watch tasks and application readiness + Warm hold — compare ready capacity and utilization · Hold · steps 3–28 · 2,400 RPS
Quiesce — invoke rolling deploy here if desired · Hold · steps 28–31 · 0 RPS
Re-burst — observe replacement and readiness + Post-deployment observation — hold the 2,400-RPS workload · Hold · steps 31–60 · 2,400 RPS
Optional phases
Traffic Recovery — enable after step 60 · ramp · steps 60–90 · 2,400→0 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Chaos · Injected failure
An intentional failure tests how the architecture responds.
What will happen
Explore how an application responds when a Bahrain region and part of a UAE deployment become unavailable while a remote backup remains available. CWM shows the difference between keeping services running and recovering lost data; it is an illustrative exercise, not a replay of a real AWS recovery.
Modeled demand
Healthy baseline — illustrative demand · Hold · From step 0 · 100 RPS
CWM injects severe region outage at aws region me-south-1 (Bahrain) · Start step 10.
CWM injects moderate network latency at UAE AZ 1 — Illustrative Surviving App · Start step 10.
CWM injects severe availability-zone outage at zone mec1-az2 · Start step 10.
CWM injects moderate network latency at UAE AZ 3 — Illustrative Surviving Data · Start step 10.
CWM injects severe permanent data loss at aws region me-south-1 (Bahrain) · Start step 20 · Terminal data loss; infrastructure recovery cannot restore original data.
CWM injects severe permanent data loss at zone mec1-az2 · Start step 20 · Terminal data loss; infrastructure recovery cannot restore original data.
Chaos · Injected interruption
A scheduled AWS Spot interruption tests whether affected workloads become ready before the termination deadline.
What will happen
Practice a deterministic Amazon EKS Spot interruption. Demand enters the configured warning band before the notice, and the scenario tests whether four workloads can be rescheduled before the two-minute simulated deadline.
Modeled demand
Steady application traffic · Hold · From step 0 · 400 RPS
Pre-interruption demand ramp — enter warning · ramp · steps 6–10 · 400→1,100 RPS
Warning hold — stay above threshold through notice · Hold · steps 10–52 · 1,100 RPS
Post-migration recovery — return to healthy baseline · ramp · steps 52–55 · 1,100→400 RPS
Interruption: EKS Spot Worker Fleet · AWS Spot interruption · Start step 10 · 120-second deadline at 3 simulated seconds/step; 4 workloads must be rescheduled.
Educational · Guided scenario
A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.
What will happen
Understand how a global single-page application (SPA) error rewrite can change API status codes when the SPA and API share one CloudFront distribution. This is a configuration walkthrough, not a load failure.
Modeled demand
Steady API load — routing bug is live, not a load problem · Hold · From step 0 · 500 RPS
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Chaos · Injected failure
An intentional failure tests how the architecture responds.
Starts at a healthy baseline — enable "Ramp to Autoscale — observe recovery behavior after fault injection" in the workspace when you want to run that optional phase.
What will happen
Use this as a manual Azure Kubernetes Service (AKS) multi-fault exercise. The seed starts healthy and waits for the user to inject a zone fault; it does not self-trigger the cascade.
Modeled demand
Wave Baseline — correlated faults cascade at steady load · wave · From step 0 · configured traffic
Optional phases
Ramp to Autoscale — observe recovery behavior after fault injection · ramp · steps 60–100 · 1,500→8,000 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Educational · Guided scenario
A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.
What will happen
See the routing separation that prevents a single-page application (SPA) fallback from rewriting API errors. This is a configuration walkthrough, not a live HTTP or CloudFront test.
Modeled demand
Steady API baseline — all resources healthy, HTTP semantics intact · Hold · From step 0 · 150 RPS
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Educational · Guided scenario
A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.
Starts at a healthy baseline — enable "Traffic Recovery — see GKE scale-in" in the workspace when you want to run that optional phase.
What will happen
Explore a containerized application on Google Kubernetes Engine (GKE) Autopilot with Memorystore for Redis and Cloud Pub/Sub. The model focuses on how the configured Kubernetes capacity responds to a 2,500-to-15,000 RPS ramp.
Modeled demand
Traffic Ramp — watch GKE autoscale · ramp · steps 0–30 · 2,500→15,000 RPS
Optional phases
Traffic Recovery — see GKE scale-in · ramp · steps 60–90 · 15,000→2,500 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Educational · Guided scenario
A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.
What will happen
Start from a healthy 1,000 RPS baseline, then manually inject a supported Redis failure to explore how cache misses affect PostgreSQL connection and latency pressure. No cache failure or recovery is scheduled by default.
Modeled demand
Moderate Baseline Traffic · Hold · From step 0 · 1,000 RPS
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Educational · Guided scenario
A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.
Starts at a healthy baseline — enable "Traffic Recovery — see AKS scale-in" in the workspace when you want to run that optional phase.
What will happen
See an Azure application handle rising demand through Kubernetes, a cache, and a queue. The simulation shows how the configured capacity and autoscaling respond; it does not contact Azure or predict production performance.
Modeled demand
Traffic Ramp — watch AKS autoscale · ramp · steps 0–30 · 2,500→15,000 RPS
Optional phases
Traffic Recovery — see AKS scale-in · ramp · steps 60–90 · 15,000→2,500 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Educational · Guided scenario
A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.
Starts at a healthy baseline — enable "Traffic Recovery — see DOKS scale-in" in the workspace when you want to run that optional phase.
What will happen
See how a DigitalOcean Kubernetes (DOKS) stack with a Load Balancer, Managed Redis, and Managed Kafka responds to rising demand. The scenario is an illustrative capacity walkthrough, not a claim about every DOKS cluster.
Modeled demand
Traffic Ramp — watch DOKS autoscale · ramp · steps 0–30 · 2,000→15,000 RPS
Optional phases
Traffic Recovery — see DOKS scale-in · ramp · steps 60–90 · 15,000→2,000 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Predictive · Capacity simulation
Modeled load and capacity constraints reveal bottlenecks or failure.
Starts at a healthy baseline — enable "Event Burst — overwhelm workers" in the workspace when you want to run that optional phase.
What will happen
Find the modeled point where an Amazon Simple Queue Service (SQS) worker fleet cannot keep up with an event burst. The scenario is a capacity exercise; it does not inject an outage or automatically add consumers.
Modeled demand
Steady Producer Load · Hold · From step 0 · 1,200 RPS
Optional phases
Event Burst — overwhelm workers · ramp · steps 0–20 · 1,200→5,000 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Predictive · Capacity simulation
Modeled load and capacity constraints reveal bottlenecks or failure.
What will happen
Model an AWS serverless API with an Application Load Balancer, Lambda, and DynamoDB. A quiet 100 RPS period precedes a 100→2,000 RPS ramp so the simulation can expose the configured cold-start and capacity response.
Modeled demand
Quiet period — Lambda idles · Hold · From step 0 · 100 RPS
Traffic burst — cold starts fire · ramp · steps 20–35 · 100→2,000 RPS
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Chaos · Injected failure
An intentional failure tests how the architecture responds.
Starts at a healthy baseline — enable "Traffic Drain — reduce load during incident response" in the workspace when you want to run that optional phase.
What will happen
See the modeled origin overload when a CloudFront edge layer stops serving cached content. The July 16, 2026 incident is provenance for the exercise; the simulation is not a replay or a claim about every CloudFront edge.
Modeled demand
Steady bypass load — origin absorbing 100% of traffic · Hold · From step 0 · 2,000 RPS
Optional phases
Traffic Drain — reduce load during incident response · ramp · steps 60–80 · 2,000→500 RPS · enable in the workspace
CWM injects severe availability-zone outage at zone us-east-1-regional · Start step 0 · Duration 9999 steps · End step 9999.
Optimization · Configuration comparison
The engine compares configuration trade-offs and surfaces the better option.
What will happen
Explore OCI’s smaller but non-zero modeled idle footprint. The ~$90/month amount and the 10 Mbps load-balancer minimum are scenario assumptions; the comparison is not a live OCI quote.
Modeled demand
Zero traffic — the bill keeps running anyway · Hold · From step 0 · 0 RPS
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Optimization · Configuration comparison
The engine compares configuration trade-offs and surfaces the better option.
Starts at a healthy baseline — enable "Traffic Recovery — enable after step 60" in the workspace when you want to run that optional phase.
What will happen
Compare a Lambda application that must start new execution capacity with one that begins with capacity already available. The scenario shows how startup delay affects a fixed workload. It is a simplified CWM comparison, not a reproduction of AWS Lambda Managed Instances or published benchmark results.
Modeled demand
Idle — establish the cold baseline · Hold · steps 0–3 · 0 RPS
Cold burst — watch application readiness reject traffic + Warm hold — compare ready capacity and utilization · Hold · steps 3–28 · 2,400 RPS
Quiesce — clear traffic before the re-burst · Hold · steps 28–31 · 0 RPS
Re-burst — observe readiness on a second cold start + Post-deployment observation — hold the 2,400-RPS workload · Hold · steps 31–60 · 2,400 RPS
Optional phases
Traffic Recovery — enable after step 60 · ramp · steps 60–90 · 2,400→0 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Optimization · Configuration comparison
The engine compares configuration trade-offs and surfaces the better option.
What will happen
Find the residual AWS charges that remain in this configured zero-traffic topology. The ~$310/month figure is the scenario’s illustrative starting estimate, not a bill or price forecast.
Modeled demand
Zero traffic — the bill keeps running anyway · Hold · From step 0 · 0 RPS
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Optimization · Configuration comparison
The engine compares configuration trade-offs and surfaces the better option.
What will happen
Inspect modeled GCP residual charges in a decommissioned, zero-traffic project. The ~$120/month amount is an illustrative configuration estimate; the exercise is not a live GCP invoice.
Modeled demand
Zero traffic — the bill keeps running anyway · Hold · From step 0 · 0 RPS
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Optimization · Configuration comparison
The engine compares configuration trade-offs and surfaces the better option.
What will happen
See which Azure resources can continue contributing to a modeled residual bill after application traffic reaches zero. The ~$175/month figure is an illustrative starting estimate, not an Azure invoice.
Modeled demand
Zero traffic — the bill keeps running anyway · Hold · From step 0 · 0 RPS
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Optimization · Configuration comparison
The engine compares configuration trade-offs and surfaces the better option.
Starts at a healthy baseline — enable "Traffic Recovery — see the stack settle back down" in the workspace when you want to run that optional phase.
What will happen
Compare the modeled cost and capacity response of a Google Cloud SaaS stack with Cloud CDN, App Engine, Cloud SQL, and Cloud Operations Suite. The 80% cache assumption is specific to this scenario, not a promise about real traffic.
Modeled demand
Business-hours ramp — watch App Engine climb into the warning zone · ramp · steps 0–60 · 90→680 RPS
Optional phases
Traffic Recovery — see the stack settle back down · ramp · steps 60–90 · 680→90 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Predictive · Capacity simulation
Modeled load and capacity constraints reveal bottlenecks or failure.
Starts at a healthy baseline — enable "Traffic Recovery — see per-token cost climb as GPUs idle" in the workspace when you want to run that optional phase.
What will happen
Study how a modeled GKE GPU inference pool spreads fixed node cost across tokens as utilization changes. The T4, utilization, autoscaling, and cost-per-million-token relationships are this configured model, not a forecast for every large-language-model workload.
Modeled demand
Inference ramp — watch cost/M tokens fall as GPU utilization rises · ramp · steps 0–40 · 200→2,000 RPS
Optional phases
Traffic Recovery — see per-token cost climb as GPUs idle · ramp · steps 60–90 · 2,000→200 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Chaos · Injected failure
An intentional failure tests how the architecture responds.
Global LB (anycast) · Cloud SQL auto-failover
Starts at a healthy baseline — enable "East Scale-Out — simulate full central drain to us-east1" in the workspace when you want to run that optional phase.
What will happen
Examine a bounded GCP routing-layer failure: Cloud DNS and the Global Load Balancer are modeled as unavailable while the east path remains a possible destination. The 2019 incident supplies provenance; this is not a historical timing replay.
Modeled demand
Mid-failover baseline — Global LB draining central, loading east · Hold · From step 0 · 500 RPS
Optional phases
East Scale-Out — simulate full central drain to us-east1 · ramp · steps 60–90 · 500→3,000 RPS · enable in the workspace
CWM injects severe availability-zone outage at zone usc1-zone-a · Start step 0 · Duration 9999 steps · End step 9999.
Optimization · Configuration comparison
The engine compares configuration trade-offs and surfaces the better option.
What will happen
Inspect DigitalOcean’s modeled residual charges after traffic reaches zero. The ~$90/month and $10/month load-balancer values are illustrative scenario inputs; they are not a live invoice or universal DigitalOcean pricing claim.
Modeled demand
Zero traffic — the bill keeps running anyway · Hold · From step 0 · 0 RPS
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Optimization · Configuration comparison
The engine compares configuration trade-offs and surfaces the better option.
Starts at a healthy baseline — enable "Traffic Recovery — see the stack settle back down" in the workspace when you want to run that optional phase.
What will happen
Compare the modeled cost and capacity response of an AWS SaaS stack with CloudFront, App Runner, Aurora PostgreSQL, and CloudWatch. The 80% cache assumption and warning-band outcome are scenario configuration, not a universal AWS result.
Modeled demand
Business-hours ramp — watch App Runner climb into the warning zone · ramp · steps 0–60 · 120→850 RPS
Optional phases
Traffic Recovery — see the stack settle back down · ramp · steps 60–90 · 850→120 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Educational · Guided scenario
A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.
What will happen
Explore a small DigitalOcean web stack: a Load Balancer, one Droplet, and Managed PostgreSQL. It starts at a modeled 400 RPS baseline to show the topology and capacity trade-off; the result is not a promise of a specific monthly bill or production throughput.
Modeled demand
Steady Baseline · Hold · From step 0 · 400 RPS
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Educational · Guided scenario
A predefined, illustrative walkthrough guides you through modeled behavior; it is not a production forecast.
Starts at a healthy baseline — enable "Traffic Recovery" in the workspace when you want to run that optional phase.
What will happen
Learn how a Redis cache and an Amazon Simple Queue Service (SQS) queue change the modeled load on an AWS microservices stack as demand rises. This walkthrough shows configuration behavior; it does not claim a universal cache-hit or queue-drain rate.
Modeled demand
Traffic Spike — watch cache absorb the burst · ramp · steps 0–20 · 1,000→5,000 RPS
Optional phases
Traffic Recovery · ramp · steps 60–80 · 5,000→1,000 RPS · enable in the workspace
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Optimization · Configuration comparison
The engine compares configuration trade-offs and surfaces the better option.
What will happen
Show the modeled cost of an EKS GPU pool with no inference traffic. With zero requests there are no generated tokens, so cost per million tokens is undefined/infinite while configured GPU capacity continues to carry cost; the result is a waste-detection exercise, not a provider bill.
Modeled demand
Zero traffic — the GPUs bill anyway · ramp · steps 0–60 · 0 RPS
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Predictive · Capacity simulation
Modeled load and capacity constraints reveal bottlenecks or failure.
What will happen
Investigate a GitHub-inspired, unaffiliated retry cascade with fixed external traffic and modeled internal retry amplification. Compare the configured unprotected and protected dependency policies without treating the result as an official GitHub reconstruction.
Modeled demand
Healthy 8K RPS baseline and recovery hold · Hold · steps 0–60 · 8,000 RPS
External traffic stays at 8,000 RPS; when modeled failures trigger the configured dependency retries, effective internal request volume can be higher. Retry attempts are internal, not additional external traffic.
No failure is scheduled; any degradation should emerge from modeled load and capacity.
Chaos · Injected failure
An intentional failure tests how the architecture responds.
Front Door (anycast) · Active Geo-Replication auto-failover
Starts at a healthy baseline — enable "Traffic Manager failover — direct bypass traffic floods West US origin" in the workspace when you want to run that optional phase.
What will happen
Learn why a healthy multi-region backend does not protect against failure of the global proxy in front of it. The 2025 Front Door incident is context; the Traffic Manager path is a modeled, optional bypass exercise.
Modeled demand
Stale-DNS trickle — only clients with cached IPs reaching backends directly · Hold · From step 0 · 50 RPS
Optional phases
Traffic Manager failover — direct bypass traffic floods West US origin · ramp · steps 60–90 · 50→3,000 RPS · enable in the workspace
CWM injects severe availability-zone outage at zone eus-zone-fd · Start step 0 · Duration 9999 steps · End step 9999.