Skip to training content
CertGraph

Stateless web tier that survives an AZ outage

Easy
Resilient Architectures
Architecture Lab · topology construction

Build controls
not started

Click or drag services, connect directional handles, then validate the design rules.

Challenge brief

Halide Health runs a stateless internal portal on one m5.large EC2 instance in eu-central-1a. When that Availability Zone lost power in March, the portal was dark for four hours while an engineer rebuilt it by hand from the release AMI. The platform lead wants the next outage absorbed by the platform: capacity must survive the loss of one Availability Zone, and a failed instance must be replaced with no page. Which design keeps the portal serving through the loss of one Availability Zone with no operator action?

Success criteria

  • 1.Add one Application Load Balancer to route only to healthy targets.
  • 2.Add one Auto Scaling group to own the capacity and launch replacements.
  • 3.Give the Auto Scaling group subnets in two Availability Zones so one zone's loss leaves capacity behind.
  • 4.Connect the load balancer to the Auto Scaling group as its target.
  • 5.Do not leave a standalone EC2 instance on the canvas; the Auto Scaling group has to own every instance.
Locked placement lanes

Selected lanes are stored as placement metadata on newly added services. Canvas coordinates cannot fake placement.

Service palette

Prepared guidance · deterministic simulation

Local

Prewritten hints from this exercise's rules, not live AI.

Use typed validation whenever you want a deterministic check. Suggestions never change the graph without your action or confirmation.

0 services · 0 directed connections·5 design rules