I'm seeing a lot of folks here talking about how to-the-minute billing would save them so much money on Amazon, but I think you're not thinking about the problem correctly.
If you want to build for scalability, then it's best to separate the jobs from the machines. Having one job per machine isn't granular enough. It's best to have a farm of machines that can take jobs and process them, and then scale the farm if the jobs aren't processed fast enough. (Incidentally this is what Lambda does for you)
If you do that, then once you have enough jobs to fill a single machine, to-the-minute billing doesn't really save you much money, if any, and you're then on a path to much better future scalability.
As a disclaimer, I work for Google on Cloud Dataproc - a managed Spark and Hadoop service. So, I am passionate and focused on Spark clusters from 3 CPUs/3GB ram to 5k+ CPUs and TBs+ of RAM.
If you are running Spark (or Hadoop, Pig, Hive, Kafka, etc.) jobs than per-minute billing can save you quite a bit. Unless you can seperate your jobs over time (and balance them) to keep n clusters saturated, you're probably paying for an idle Spark/Hadoop cluster to sit around. Moreover, in terms of Cloud Dataproc (only) that's why you can scale clusters up/down, use preemptibles, and custom machine types. You can custom shape your clusters and pay for exactly what you use (disclaimer - there is a 10 minute minimum like most Google Cloud services.)
As a practical example, if you have a 100 node Spark cluster and only use 25 minutes of it, you can stand to save considerably in a given year. Yes, you can possibly rebalance your work internally to saturate a cluster at the optimal n minutes but at that point, you're paying to do the engineering work for it. :)
True but as soon as you have two jobs you can saturate that cluster for an hour.
My point is at only the smallest scales does hourly vs minutely billing make a difference. Yes, it's really nice, but doesn't make as much difference as everyone says.
I think you are correct, but there is a case where saturating the cluster is maybe not the ideal situation.
I think there's a latency trade-off there. So imagine you have a job where latency matters, like a user wants to run a bigquery-style interactive query across 1TB of data, and ideally it should complete within 1 minute. Lets say it takes 10 instance hours (600 instance minutes) to complete this query.
With per-minute billing you could launch 600 instances and complete it in a minute. You could also do the same with per-hour billing, but you would overpay by a factor of 60x or you would launch 10 instances and wait an hour for the query to complete.
Assuming you have a bunch of these queries coming in, you could queue them up, but then the latency will suffer, because you'll have to delay starting a job until you get to that position in the queue.
If you wish to trade some wait time for some cluster efficiency, yes, you can just queue them up and then slowly scale the cluster up and down to keep 100% utilization.
However, it would be nice if you really could scale up and down at per-minute increments and let Google figure out how to get their cluster utilized efficiently and let them take advantage of their different customers with different workloads and preemptible instances.
You're reallllly simplifying a hard problem. "Just pack the bins tighter, no big deal." This is what Omega/Borg/Mesos are all trying to - it's not an easy problem.
Also, you're assuming you have enough work to keep all machines busy all the time....
Actually, what I'm saying is that bin packing is a hard problem so separate the work from the machines, make the bins smaller, and then bin packing becomes easier.
It really depends on how elastic your workload is. Imagine, if amazon had 1 day/week/month minimum for the VMs. You'd run into the same issues (you wouldn't respond to a spike in demand if it was only going to last 2 hours).
I think the point is: Smaller billing windows allow you to be more responsive to your load and be able to shut down faster and save more money.
I may have derailed my comment a bit by mentioning lambda. My point was more about separating work from machines, which Lambda helps with, but you can also do on your own.
If you want to build for scalability, then it's best to separate the jobs from the machines. Having one job per machine isn't granular enough. It's best to have a farm of machines that can take jobs and process them, and then scale the farm if the jobs aren't processed fast enough. (Incidentally this is what Lambda does for you)
If you do that, then once you have enough jobs to fill a single machine, to-the-minute billing doesn't really save you much money, if any, and you're then on a path to much better future scalability.