v0.1.20 (Early Access)
Added
Multiple availability zones
- The internet and connectivity screens now label each carrier handoff by its zone. Once a site spans more than one availability zone, each zone has its own internet edge: its own carrier line panel and its own border router. Before, both lenses listed every carrier panel with no way to tell which zone it belonged to, so on a two-zone site you saw two panels named only by their equipment id and had to guess. Now, on any multi-zone site, each carrier panel and each uplink row is prefixed with its zone (for example "AZ-2 · edd-02"), so you always know which zone's lines and routing policy you are editing. Single-zone sites are unchanged, with no extra labelling to clutter them.
- See cross-zone traffic and backbone pressure at a glance. Spreading a database across two zones sends its writes between them over the region's internal backbone. The connectivity lens now carries a BACKBONE readout that shows how much cross-zone traffic is flowing and warns you when that backbone is running out of headroom, the same way a saturated internet pipe is flagged. On a single-zone site it simply reads blank, since there is no cross-zone traffic to show.
- A zone can lean on another zone's internet edge when its own border goes down. With more than one zone in a region, the zones are wired together over that backbone, so if one zone's border router or carrier line fails, its outbound traffic can still leave through a healthy zone's edge instead of going dark. This applies between the larger, region-spanning routers, and it shares the same backbone capacity as cross-zone replication, so heavy failover traffic competes for that pipe.
- The topology map now draws the backbone between your zones. Because the regional backbone is fibre the carrier provides rather than a cable you lay, the network map never showed it, so a multi-zone region looked like two unconnected islands. The map now draws a distinct dashed line between each pair of zones' borders, so you can see the zones are linked, and it lights up hot when that backbone is running full, the same as any other saturated link. It appears only on multi-zone regions and never affects how the game works out what can reach what.
Networking
- See at a glance how each customer's networks route between one another. When a customer runs their machines across more than one of their own networks (say app servers on one and databases on another), something has to route traffic between those networks. The network isolation screen now shows this for every customer: their networks either route locally on a router-capable switch (which keeps working even if the internet router goes down), route through the single border router (a weak point if that is your only router), or can't reach each other at all. Now you can spot a fragile setup before it bites, instead of finding out when the tiers go quiet.
- The topology map now marks which switches are routers. A switch that can route between networks is drawn with a small diamond, the standard map symbol for a router, and a matching key appears in the legend. It makes the shape of your network plain: where traffic can turn the corner between networks on its own, and where it has to travel up to the border router and back. Switches that only pass traffic within a single network are left unmarked, and the marker only appears once you own a router-capable switch.
Diagnosing problems
- A "Diagnose" tool that tells you WHY a customer is unhappy, in one place. The information you needed to work out why a customer was struggling used to be scattered across four different screens (their service health, your capacity, the network, their contract). Now there is a single "Diagnose" button on a customer's page that runs a full checklist in front of you and points at the real root cause. It ticks through every dimension one by one, marking each with a clear pass, warning, fail, or "not relevant": whether their services placed, whether they can reach the internet and each other, whether any of their machines are stranded, and whether every part of their contract (private network, multi-zone, data residency, spreading copies for safety) is actually being met. At the end it names the single deepest problem in plain language, links the codex entry that explains it, and offers a button that takes you straight to the fix (buy a server, open their segment, jump to the network view, or start a repair).
- Diagnose your whole datacenter at once. A "Diagnose" button in the top right of the Operations view runs the same kind of self-test across the entire site: your internet edge, how many servers are online, whether the fleet has room for every contracted machine, how many incidents are open, and how many tenants are healthy versus need attention. Any customer that is not all-clear is listed with their own one-line cause, and you can click straight through to their full diagnosis.
- Diagnose straight from an incident. A customer-rooted incident now carries a "Diagnose" button for the affected tenant, so you can jump from "something is wrong with this customer" to the full checklist without hunting for them in the roster.
- Incident guides now unlock the moment the incident happens, not after you have fixed it. The codex entry that explains a kind of incident (its cause and how to fix it) used to appear only once you had already resolved one. That is exactly backwards: you want to read how to fix something while it is happening. Now the entry unlocks the first time that kind of incident occurs, and a notice names it, so the guidance is there when you actually need it.
Changed
Compute and capacity
- Server CPU load is now real, and busy customers can crowd their neighbours. Before, a customer's machines barely registered any CPU use no matter how much traffic they served, so the figures on the server console and the capacity screen sat near zero and you could pile far more onto a server than it could truly handle with nothing showing the strain. Now CPU use tracks the actual work each machine is doing: it rises and falls with request volume, spikes when a customer gets busy, and varies by the kind of workload (a lightweight web app sips CPU while a heavy compute or database job eats it). When a server gets genuinely overloaded, the tenants sharing it start to feel it, the same noisy-neighbour pressure real hardware has. Overcommitting a server is a real decision with real consequences again, so a rack you have packed tight may now show honest load where it used to read empty.
- Databases, load balancers, Kubernetes and CDNs now show real CPU use too. Previously only plain virtual machines moved the CPU needle, so a database under heavy query load or a load balancer passing thousands of requests a second read as using no CPU at all. Now these services draw CPU in proportion to how hard they are working, and their rows on the server console (and their share of a server's total load) reflect that. Object storage stays at zero, since it is limited by disk and bandwidth rather than CPU.
Service tiers
- The uptime each service tier promises has been rebalanced, and the lower tiers are now genuinely best-effort. Each tier now aims for a clearer, more forgiving uptime figure: Bronze 90%, Silver 95%, Gold 99%, Platinum 99.5% (they were 99%, 99.5%, 99.9% and 99.95%). A tier only owes a service credit once its uptime slips below its own figure, so a Bronze tenant, who pays nothing for a standard guarantee or an uplift, now rides out the ordinary blips as best-effort and earns a refund only if service falls below the 90% floor. The higher tiers still hold higher bars and still pay out sooner when they drop below acceptable.
- Each paid tier now shows the response-time ceiling it is actually held to. Alongside its uptime figure, each paid tier also promises a response-time ceiling, and this is where crowding a server bites: pack a box too tight and every tenant on it slows down, so a premium tenant you overcommit earns a credit even while staying up. The ceilings are Silver under 100ms, Gold under 50ms, Platinum under 30ms (Bronze makes no response-time promise). A healthy request runs in single-digit milliseconds, so these only trip when a server is genuinely overloaded or dropping requests, never for a tenant's own slow workload. The figures shown on the promise were previously stale and far too loose; they now match what the game measures.
Cabling
- A much wider choice of cable colours when you are wiring up. You can already pick a jacket colour while spooling a cable to colour-code a build, but several cable types offered only two colours to choose from, twinax (DAC) and the fibre types especially, so there was little room to keep your own scheme. Every data cable type now offers a broader palette, including five new colours (red, white, pink, cyan and magenta) on top of the ones already there. Copper patch leads get the full range; twinax, active optical and fibre each get a wider realistic set. The default colour for every cable type is unchanged, so existing cables and anything you have already run look exactly as before.
Fixed
Customers
- The outbound-traffic graph on a customer's page now draws. The EGRESS chart on a customer's detail page was always empty, even for busy customers pushing plenty of traffic, because it was looking for its data in the wrong place. It now plots the customer's outbound bandwidth over the last minute, alongside the request-rate and latency graphs beside it.
- Customer machines no longer get stuck half-deployed with nothing telling you why. Some customers ask for their machines to be kept on separate servers, so one server failing can't take all of them down at once. Before, if you didn't have enough separate servers to spread them across, the leftover machines would simply never come online, and no screen anywhere explained it. That was the "only some of the customer's machines placed, and I can't find out why" trap. Now those machines always come online: when there's no separate server free, the game doubles them up rather than leaving the customer short and at risk of walking out. The customer stays served the whole time.
- When a customer's machines aren't placing, the game now names the reason and the fix. Instead of a vague "partially placed", the customer's page (and the matching incident) now tell you which of these is actually happening, each with a plain next step:
- Every server that can run this customer is full. Add more compute, or free some up.
- The customer is walled into its own group of servers (a private network segment or dedicated hosts) and that group is full, even though you have room elsewhere. Add or free a server inside their group. This is the "their network segment has no space" case.
- Their machines are all up and running, but some are doubled onto the same server, so the spread they were promised isn't there. Add a server, or a second availability zone, and they will spread back out on their own. This one is a heads-up about resilience, not an outage, so it never counts against the customer's uptime.
- You have run out of public IP addresses to hand the customer's machines. Free some addresses or add address space, and the waiting machines come up.
- Every kind of service now explains itself when it can't come up, not just plain machines. Load balancers, content delivery, managed Kubernetes and databases used to fail quietly: if one couldn't be placed, it simply read as down with no reason and no incident, while plain virtual machines got a clear explanation. Now all of them go through the same path. If the right kind of server is full, or the customer is walled into a group of servers that is full, or a load balancer has no public address left, it says so and raises an incident with the fix, the same as machines do.
- A managed database no longer refuses to come up just because it can't spread perfectly. A database that keeps several copies for safety wants each copy on separate hardware. Before, if it couldn't get enough separate racks or zones, it placed nothing at all and the customer had no database. Now it comes up anyway on the hardware it can reach and tells you plainly that its copies aren't fully separated yet, so a single failure could take it down. It keeps serving, and adding a rack or zone lets the copies spread back out. You are told the risk rather than left with nothing.
Incidents
- Every incident now names its cause up front. An incident that a customer couldn't be set up used to read only "can't be set up" at the top, with the actual reason buried a section lower. Now the headline itself names why and what to do ("every server is full, add compute"), so you get the cause at a glance. And a security breach now shows its cause in the incident detail instead of the misleading "cause not yet identified" it used to display for a fully-known break-in.
- The warning light on a rack now marks the rack actually worth checking. Every open incident lights a red bar across the top of a rack so you can spot trouble from across the room. But for a whole range of problems, a crowded server, a failed cable, a database that can't keep enough copies for safety, that light always appeared on the rack holding your internet gear, sending you to the wrong end of the room. Now the light lands where the problem really is: a server running hot lights its own rack, a failed cable or bonded link lights the rack it runs into, and a customer whose services have outgrown your hardware lights the rack their machines live on. When a customer's own networks can't route to each other, the light points at their machines, where you would add a router-capable switch, rather than the internet edge. Problems that genuinely live at the internet edge, like an internet outage, a network loop, or a whole zone going dark, still light the edge rack, because that is where you would go to look.
Networking
- The internet link on the topology map now shows the speed it really runs at. The line from your border router out to the carrier was labelled with the router port's own top speed (for example 400G), even when the carrier handoff on the far side only runs at 10G, so the map claimed a much faster link than you actually had. It now shows the slower of the two ends, the speed the connection truly negotiates, the same way every cable elsewhere on the map is measured.
- When a customer's own tiers can't reach each other, the game now names the router as the cause. A customer whose machines are split across more than one of their own networks needs a router between those networks to pass traffic. If the only router on that path is down (or there is no router-capable device there at all), those tiers go silent. Before, this looked like the generic "the customer's pieces can't reach each other" alarm, which pointed you at re-cabling. Now it is called out as its own cause with the right fix: bring the border router back up, or add a switch that can route between networks. Re-cabling was never going to help.
- Plainer wording across the network screens when engineer mode is off. The busy-links warnings, the reliability summary, and a few device labels now use everyday words instead of networking shorthand, so the network views read cleanly whether or not you speak the jargon. For example, a maxed-out internet line now reads "internet line is full, add another line or buy more bandwidth" rather than "uplink saturated", and the crowding tag reads "FULL" instead of "SAT". Turning engineer mode on brings the technical terms back.
Racks and hardware
- Carrier handoff boxes now have their power cables. The rack-mounted carrier box where your internet lines land (the EDD) is an active, powered device with real power sockets moulded onto the back, but it was the one piece of rack gear that sat with nothing plugged into them. Now, like every server, switch and router in your racks, a floor-mounted carrier box comes up with its AC cord run neatly down to the rack's rear power channel. Wall-mounted units in the garage are unchanged, since they run off a wall socket like a home router. This is a visual consistency fix and doesn't change how anything works.
- Power-supply status lights now show which cords are actually plugged in. Every rack device with two power supplies has a little status light for each one, on the back beside its power socket. Those lights were both lit whenever the device had power, even if you had only run one of its two AC cords, so the back of the rack claimed full power redundancy you hadn't wired up. Now each light homes to its own supply: cord one plugged in lights one, both cords lights both, and the light sits over the socket its cord feeds. It reads true across the whole fleet, servers, switches, routers and the carrier handoff boxes alike. This is a visual accuracy fix and doesn't change how anything works.
- Reaching the Data Center through the campaign now gives you the full carrier handoff. When you graduated to the Data Center, the carrier box where your internet lines land was arriving as a smaller two-line unit (about 20 Gb of handoff) rather than the four-line, 100 Gb unit the site is meant to have, and the one every standalone Data Center scenario already ships. That quietly capped how many internet lines you could run into how many border routers, so you had less redundancy and headroom than the site should give. Now the campaign hands you the proper four-line carrier box on move-in, per zone, so you can spread more lines across more routers. Saved games already parked at the Data Center are upgraded to the correct box automatically on load, keeping whatever you had cabled into it.