Guess I'm too ignorant. I need to read up on these. I did know about the persistence feature. I think it's not terrible but also not great, and systems should be designed for being shut down and apps being closed.
> I think it's not terrible but also not great, and systems should be designed for being shut down and apps being closed.
The problem with shutdowns and restarts is the secure bootstrapping problem. The boot process must be within the trusted computing base, so how do you minimize the chance of introducing vulnerabilities? With checkpointing, if you start in a secure state, you're guaranteed to have a secure state after a reboot. This is not the case with other any other form of reboot, particularly ones that are highly configurable and so easy for the user to introduce an insecure configuration.
In any case, many apps are now designed to restore their state on restart, so they are effectively checkpointing themselves, so there's clearly value to checkpointing. In systems with OS-provided checkpointing it's a central shared service and doesn't have to be replicated in every program. That's a significant reduction in overall system code that can go wrong.
It's fallacious to assume that the persistence model of the system can't enter an invalid state and thus cause issues similar to bootstrapping. The threat model also doesn't make sense to me: if an attacker can manipulate the boot process, I feel like they would be able to attack the overall system just fine. Also, there's the bandwidth usage, latency, and whatnot. I think persistence is a strictly less powerful, although certainly convenient, design for an OS.
> The threat model also doesn't make sense to me: if an attacker can manipulate the boot process, I feel like they would be able to attack the overall system just fine.
That's not true actually. These capability systems have the principle of least privilege right down to their core. The checkpointing code is in the kernel which only calls out to the disk driver in user space. The checkpointing code itself is basically just "flush these cached pages to their corresponding locations on disk, then update a boot sector pointer to the new checkpoint", and booting a system is "read these pages pointed to by this disk pointer sequentially into memory and resume".
The attack surface in this system is incomparably small compared to the boot process of a typical OS, which run user-defined scripts and scripts written by completely unknown people from software you downloaded from the internet, often with root or other broad sets of privileges.
I really don't think you can appreciate how this system works without digging into it a little. EROS was built from the design of KeyKOS that ran transactional bank systems back in the 80s. KeyKOS pioneered this kind of checkpointing system, so it saw real industry use in secure systems for years. I recommend at least reading an overview:
EROS is kind of like what you'd get it if you took Smalltalk and tried to push it into the hardware as an operating system, while removing all sources of ambient authority. It lives on as CapROS:
I don't deny that bootstrapping in current systems is ridiculous, but I don't see why it can't be improved. It's not like EROS is a typical OS either. In any case, I'll read up on those OSes.