A Capacity Plan Sizes for Peak Plus Growth at a Target Utilisation, Not for Average Load
A capacity plan is the short piece of arithmetic that turns load numbers into a number of machines, and it is wrong whenever it starts from the average. Three inputs decide it. The first is peak, not mean: a service that averages two hundred requests a second may see nine hundred in its busiest minute, and the busy minute is the one that has to be served. The second is the service time of one request, which with the arrival rate gives the concurrency the system must hold at once; that relationship is the one Little's law states, and a capacity plan is an application of it rather than a separate technique. The third is the target utilisation you intend to run at, and this is where headroom enters. You do not size a fleet to be exactly full at peak, because a queue's waiting time climbs steeply as utilisation approaches one, and because instances fail, deploys take capacity out, and traffic grows between the day you plan and the day you are stuck with the plan. A plan therefore names a utilisation target well below saturation, adds the expected growth over the planning horizon, and states the two assumptions out loud: the peak-to-average ratio and the service time. You can now turn a brief's load numbers into a fleet size, and say which assumption would have to be wrong for the number to be wrong.
This Concept is waiting for its first lesson!
A capacity plan is the short piece of arithmetic that turns load numbers into a number of machines, and it is wrong whenever it starts from the average. Three inputs decide it. The first is peak, not mean: a service that averages two hundred requests a second may see nine hundred in its busiest minute, and the busy minute is the one that has to be served. The second is the service time of one request, which with the arrival rate gives the concurrency the system must hold at once; that relationship is the one Little's law states, and a capacity plan is an application of it rather than a separate technique. The third is the target utilisation you intend to run at, and this is where headroom enters. You do not size a fleet to be exactly full at peak, because a queue's waiting time climbs steeply as utilisation approaches one, and because instances fail, deploys take capacity out, and traffic grows between the day you plan and the day you are stuck with the plan. A plan therefore names a utilisation target well below saturation, adds the expected growth over the planning horizon, and states the two assumptions out loud: the peak-to-average ratio and the service time. You can now turn a brief's load numbers into a fleet size, and say which assumption would have to be wrong for the number to be wrong.
Are you a teacher? Sign in to start contributing.
Sign In