If your hosts have, say, dual 4-way CPUs, and you're giving your VMs 8 vCPUs, then a single VM can execute per clock cycle since a VM needs all of it's vCPUs made available to the guest OS. With 8 VMs, that means one VM is executing every 8th clock cycle.
If you "downgrade" those VMs to 2 vCPUs, then 4 VMs can execute per clock cycle, and that VM can now execute every other clock cycle instead of waiting 8 cycles. More work gets done, even though the amount of CPUs has gone down.
Since most VMs are probably executing work that only has 1-2 threads, then there's no loss from lack of parallelism. Remember, a VM needs all cores available, regardless of how many threads need to be executed.
Of course, some workloads do need that many threads, and so balancing number of cores vs. available execution time becomes a little more tricky. but based on what you shared, it sounds like Linode took a closer look and came to the same conclusion.
If your hosts have, say, dual 4-way CPUs, and you're giving your VMs 8 vCPUs, then a single VM can execute per clock cycle since a VM needs all of it's vCPUs made available to the guest OS.
This seems wrong to me. We're talking about virtualization; technologically, it is absolutely feasibly to have only one physical CPU running a VM even if the VM sees multiple virtual CPUs. And it seems like having this capability in virtualization software from the very beginning is really a no-brainer.
In a previous environment of mine, we ran dual socket, 4-way CPUs. VMs were configured with a mix of 1, 2, and 4 vCPU VMs. Our VMs appeared slow, especially on our 4-way systems. However, CPU utilization was low. More digging revealed that our co-stop values were high, meaning that the system couldn't schedule execution time effectively, meaning the VM had to sit in a READY state, which kept CPU utilization low.
Our first fix was to rebalance our cluster, so that 1 and 2 vCPU VMs were relegated to their own set of hosts, and our 4 vCPU VMs executed on their own set. Instantly, out co-stop values dropped, and CPU utilization rates went up...they were now doing work!
The VMware co-scheduler has improved over the years, but I still read (The "Mastering vSphere 5.5" book by Scott Lowe carries a warning on this as well) that carefully balancing vCPUs is a must in a VMware environment. (Again, I don't believe Linode uses VMware, so I can't say with any certainty that KVM or Xen exhibit this behavior.)
So why can't we run 8 vCPUs on one physical one? Because while they're virtual to some extent, they're not completely abstracted. Anytime the hypervisor has to perform a translation between the guest OS and the host, a performance penalty is incured. So while the hypervisor may abstract scheduling, it reveals as much of the physical CPU to the guest VM as possible. Here's a little blurb from an older VMware manual explaining a bit of the difference:
http://pubs.vmware.com/vsphere-4-esx-vcenter/index.jsp?topic...
CPU virtualization =/= emulation
For this reason, (again, at least in a VMware environment) we can't give a VM more vCPUs than exist pCPUs to align them to.
If your hosts have, say, dual 4-way CPUs, and you're giving your VMs 8 vCPUs, then a single VM can execute per clock cycle since a VM needs all of it's vCPUs made available to the guest OS. With 8 VMs, that means one VM is executing every 8th clock cycle.
If you "downgrade" those VMs to 2 vCPUs, then 4 VMs can execute per clock cycle, and that VM can now execute every other clock cycle instead of waiting 8 cycles. More work gets done, even though the amount of CPUs has gone down.
Since most VMs are probably executing work that only has 1-2 threads, then there's no loss from lack of parallelism. Remember, a VM needs all cores available, regardless of how many threads need to be executed.
Of course, some workloads do need that many threads, and so balancing number of cores vs. available execution time becomes a little more tricky. but based on what you shared, it sounds like Linode took a closer look and came to the same conclusion.