Incident Response
Operational checklists for common incident types.
Use these focused diagnostic checklists as a starting point when an incident is active.
-
appliance 0 over forwarding capacity
In-path appliance over forwarding capacity — inspection load is eating headroom; upgrade the box / reduce inspected traffic / add a parallel path.
-
appliance 0 over inspection capacity
Deep inspection is over its own ceiling (
inspection_capacity_gbps), which is a SEPARATE and much smaller number than the box's forwarding fabric — the box can pass this… -
Site-wide outage — nothing in that AZ serves until it clears.
-
A loop went unchecked and traffic is multiplying — break the loop NOW; pull one leg.
-
Trace the run end to end — check the link LEDs on both ports; dark on both ends confirms the cable.
-
A component inside the host died — inspect the host (walk up, press E) to see which part.
-
Their servers can't see each other — the east/west path between their racks or sites is severed.
-
customer 0 cannot reach internet
Trace this customer's path to the internet: host → switch → gateway → internet box; the break is on that chain.
-
The customer sees degradation but the engine couldn't pin a physical cause — check their hosts (Customers > detail shows placement) and the path from them to the internet.
-
customer 0 Virtual machines unservable
Every host that can run this customer's VMs is full (after failover headroom) — the fleet has aggregate capacity but no single host has room for the next VM. This is…
-
Check the Security lens: attack classes and residual leaking past your scrubbing.
-
Name lookups are failing facility-wide — your servers are up, but customers can't resolve them. This is a macro-scale outage, not a fault in your gear: there is no cable, box, or…
-
Your gateway can't reach the internet — the internet box has NO live path out. Every carrier line on it is dead or unpatched (this is derived reachability, not a power fault).
-
Open the Network lens (Load view): the router's forwarding fabric is over its ceiling — north/south throughput is capped at this box.
-
This host is CPU-overcommitted — offered CPU demand is over the box's physical cores, so every workload on it is being clipped.
-
Walk to the host and read its status LEDs — dark means power, red means the box itself.
-
customer 0 tiers can't route between subnets
Their tiers are in different subnets of one VRF and nothing is routing between them: the fabric has no LIVE L3 first-hop router on the path.
-
The bonded link lost its members — check each member cable and both end ports.
-
This LINK is pegged at its negotiated rate — both devices it joins may still have fabric headroom, the cable between them does not.
-
MAC flapping: switch 0 ↔ gateway 0
Two unbonded parallel cables run between the same switch and gateway — each frame flip-flops between them.
-
Two paths connect the same segment — trace your recent cabling for the redundant leg.
-
customer 0 object store under-replicated
This tenant's object store can't hold its full replica set — fewer than the durability tier's copies are placed because the fleet is short on aggregate SSD.
-
The rack drew more than its breaker rating — check the rack's power draw before resetting.
-
A targeted intrusion campaign ran unmitigated long enough to compromise this tenant — the Intrusion class leaked with no IDS/IPS in path.
-
Open the Network lens (Load view): this switch's northbound trunk and/or backplane fabric is over capacity — it's dropping packets crossing it.
-
Check the switch's power and status LEDs at the rack.
-
Facility cooling is over capacity and heat is compounding — servers across the fleet throttle, then fail (this is facility-wide, not one rack).
-
One line on the internet box is dead — check the port LED and the cable seated in it.
-
Check the Network lens: uplink throughput vs capacity — you're pushing more than the line can carry.