Procurement gets handed a spec. 32 cores, 256GB of RAM, 4TB of disk. It goes to market and asks four vendors to price 32 cores, 256GB of RAM, 4TB of disk.
Nobody stops to ask where the number came from.
That question is the whole job. It is also the one question nobody asks. This is about how to ask it, and what happens to the bill when you do.
The number is almost never a measurement
Start with the cores. 32 of them. Why does it need that many?
That much compute points at a processing-heavy workload, so the first thing to do is check what the machine is actually doing. A SIEM platform grinding through logs might genuinely want the cores. A web front end or a database often does not. High core counts get carried forward without anyone checking which of those they are looking at.
Usually the 32 came from nowhere in particular. Someone bought a server five years ago, sized it for the worst day of its life plus a comfortable margin, and it has run at a fraction of that ever since. The cores were cheap. On commodity kit you buy the box once, rack it, and the spare capacity sits there humming. Headroom you never use costs nothing you can see, once you have already paid for it.
Then the platform moves to the cloud, and the spec travels with it. We have 32 cores, so we must need 32 cores. Out it goes to market as a like for like.
On-prem hides the waste. The cloud bills for it.
On-prem, over-provisioning is a sunk cost you stopped noticing. In the cloud you rent it. Every month. The margin nobody measured becomes a bill nobody can stop paying, and when the like-for-like quote lands high the verdict is that cloud is expensive.
It isn't. The spec was wrong before it ever left the building.
Two things hide inside the word "core", and neither of them gets checked.
The first is the unit. A spec sheet lists physical cores. The cloud rents vCPUs, and the two are not the same thing. Under a typical 1:4 contention ratio, one physical core presents around four virtual ones. Read the number straight across, as if a core were a core, and it can be out by fourfold before anyone has looked at what the workload does.
The second is utilisation. Whatever gets allocated, the workload rarely touches all of it. The figure I see again and again is around 30% CPU. So even once the unit is settled, most of what has been provisioned sits idle.
Here is the part that makes it stick. The person buying does not know there is a difference, and often does not want to. To them there is no physical core and no virtual core. There is 32, and they want 32. Explaining why it matters is hard going. It lands as needless complication on someone who just wants the number. So 32 goes to market with no unit attached to it, and the wrong question gets priced as if it were the right one. The waste was baked in long before anyone opened a pricing tool.
Memory tells the same story, a little more gently, usually nearer 50% used. Storage tracks closest to what's provisioned. When it does drift, though, I've seen it two or three times over. And with hardware prices as volatile as they have been lately, memory especially, paying for capacity you do not touch is a real hit to a budget, for something that was never needed.
What to actually measure
You cannot right-size off a spec sheet. You measure.
I prefer to run an assessment on the live platform for a week or two, across the busiest stretch the business expects, so peak and average show up as facts rather than guesses. Design to the baseline. Build in a way to absorb the peaks on demand when they are rare enough to be worth handling on their own.
How long you watch depends on the business. A retail platform spikes hard during sales and festive periods and sits quiet the rest of the year. There the real question is which load is a genuine once-or-twice-a-year burst worth scaling into, and which is capacity that would otherwise idle for 90% of its life. A real-time data platform under continuous heavy load is the opposite case. The load is real and the spend earns its keep. A week is usually enough to prove it.
Most organisations in the region are still on traditional VMware or physical servers, in-house or sitting in a colocation facility somewhere, often a mix of both. Tools like Live Optics, Azure Migrate and AWS Migration Hub pull the real metrics off that estate. The charts do not lie. They show you peak against average, and used against provisioned, measured over real days on a real workload. That gap is money the customer did not know they were spending.
One caution on the tooling. I don't take a tool's own recommended cloud size at face value. Each was built by a vendor, and the number tends to land somewhere that suits the one who built the tool. Where I can, I run more than one and reconcile them down to a baseline I can defend. It is not always possible. But solid data to back up a conclusion is every technical person's dream.
The metrics are only half of it. The rest is business questions the estate cannot answer for you. What is this platform, and who does it actually serve? What is ancient and what is new? What is due to be retired soon, so it should never be sized for at all? Is the plan to keep running kit in-house, or to get more out of it through a managed partner? Those answers move the sizing as much as any chart does.
Getting them is often the hard part. An assessment can worry an internal team, because change reads as a threat to their roles. Sometimes the knowledge has simply gone: the people who built the platform moved on, and whoever inherited it is wary of touching something that works when they are not sure how to fix it if it breaks. Regulated industries have real questions about what assessment tooling can see and reach. And documentation is often thin, lost to the same churn or quietly held back by someone protecting their own place. None of that is on the spec sheet. All of it changes the answer.
How to spot a guess
Some specs give themselves away.
The clearest tell is heavy CPU and memory piled onto very few machines. Real workloads spread out. A large requirement usually breaks into a group of smaller virtual machines doing distinct jobs. When a big core and memory count sits on one or two boxes, that is a strong sign nobody sized it. They inherited it.
Sizing goes wrong the other way too. It is rarer, and when it happens it is usually because the picture was incomplete, not because anyone got it deliberately wrong. A customer does not always know their own workload. Something an ISV buried in a technical document gets lost or never handed over, and the requirement that mattered never reaches the assessment. That is not the customer being difficult. It is knowledge that fell through a gap.
Getting the virtual cores right is not only a compute question. A lot of software is licensed by the core. Trim a machine to what its role actually needs and the customer saves twice, on the resources and on licences they were about to buy for cores that would never have run anything.
Right-sizing is not cost-cutting
I am not in the business of selling someone a Ferrari when an average car takes them from A to B just as well. Nor the reverse, a business stranded in something too small for the job. The aim is a platform sized to what the workload really uses, with enough overhead that it holds up under real strain and scales when the business needs it to. It should not fall over in the single week it exists to make its money. It should not sit as a nasty line on the budget report for the other fifty-one.
Which leaves an obvious question. If right-sizing so plainly serves the customer, why does the vendor who just types the numbers back tend to win the work?
Why the numbers-back vendor tends to win
Because the person buying often cannot tell the difference, and the process rewards whoever makes their life easiest.
In a formal tender the score goes on a spreadsheet. Points for hitting each line item, points for the price, and the lowest compliant number tends to win. Whether the specification made any sense in the first place is not one of the columns. A procurement lead does not always know a virtual server on a cheap standalone host from one on an enterprise high-availability platform. To them it is a server. That is not a dig. It is not their field, and nobody expects them to size infrastructure any more than they would expect me to run their tender.
The awkward part is that the vendor who quietly right-sizes looks dearer on the page than the one who types the same numbers back, even when the right-sized platform is the cheaper thing to own. Question the spec, ask for more detail, and you become the supplier making it complicated. I have watched a request for more information so we could save the client money land like a request for a favour.
Every so often someone does engage, usually once they realise I am trying to help rather than pick their pocket. In my experience that tends to be the smaller firms, the ones watching their own money. My father has a line for it. Watch the pennies and the pounds look after themselves.
Then there is time. In the UK or Europe a tender gave you three or four weeks. Here you can get five days before it lands on your desk, which is not enough to do the due diligence and get real quotes back. The genuinely useful ones also have a habit of arriving at the least convenient moment, the week before a holiday, which usually says more about what is going on upstream than about the requirement itself. Everyone's deadline is your deadline, and it is always yesterday.
If you want to know what a buyer really cares about, watch where their attention goes. I have used electronic proposal tools that show exactly that, how long a prospect spends on each page. The technical and solution sections, the part that decides whether the thing will actually work, get under a minute. The commercial pages get ten. That gap tells you everything, and it is worth sitting with rather than resenting. The number is what gets read, so the number is where the argument has to be won.
Which is why the most useful sentence I have is not a challenge. It is an offer. I can give you exactly what you asked for, and it will probably be expensive and probably wrong for what you actually need. Or you give us a little time to look at the requirement properly, and we point you at something more suitable and, more often than not, quite a bit cheaper. Put that way, most reasonable people say yes.
When lift-and-shift is the honest answer
Not everything needs re-architecting, and I am not in the business of inventing work.
A straight like-for-like move is genuinely the right call when the usage data is accurate and the client is engaged and willing to act on what the numbers say. Give me a real picture of the workload and a customer who wants the right outcome, and lift-and-shift can be the sensible, honest thing to do.
Where it comes apart is when the same requirement is out with several vendors at once, most of them happy to promise the world and deliver peanuts when it comes to the crunch. I have lost count of the times a conversation has come back around because the first partner the customer picked did not deliver, so they have gone back to the market to run the whole exercise again. Time and money spent twice, to arrive roughly where careful work would have put them the first time.
So if there is one thing to take from all of this, it is an unglamorous one. If you want to do it once and do it right, take the time to get your ducks in a row before you go to market. The reason that is hard is the same reason half of this is hard. The clock was already running before it reached you, usually because someone upstream could not hold the line on expectations and blinked under pressure.
Slow down the start, and the rest tends to look after itself. Rather like the pennies.