But for things that really matter, I've tested by configuring the MMU to disable caching for the memory that the realtime code lives in and uses to emulate 0% hitrate. And there's usually still a fair amount of variance on top of that depending on if the memory controller has a small cache, and where the memory controller is in its refresh cycle.
Yeah. And I'm not sure that even that would give you the worst case as far as the cache is concerned. Of course I don't know how these implementations work, but it seems plausible that code that directly uses memory could run faster than code that encounters a cache miss beforehand (or contention, if you're using multiple cores). Moreover there's also the instruction cache, and I'm not sure if you can disable caching for that in a meaningful way?
For soft real time, I don't see a problem. But for hard real time, it seems a bit scary.
You're right! I can think of two cases I've run into where bypassing the cache can be faster compared to a miss.
On some caches the line must be filled before allowing a write(ignoring any write buffer at the interface above the cache) - those basically halve the memory bandwidth when writing to a lot of cache lines. Some systems now have instructions for filling a cache line directly to avoid this. And some CPUs have bit-per-byte validity tracking to avoid this too.
Even on caches with hit-during-fill, a direct read from an address near the last-to-be-filled end of a cacheline can sometimes be a little faster than a cache miss, since the miss will fill the rest of the line first.
> Moreover there's also the instruction cache, and I'm not sure if you can disable caching for that in a meaningful way?
Intels used to boot with their caches disabled, but I haven't worked with them in forever, and never multicore.
I worked with a lot of microcontrollers, and it's not uncommon to be able to disable the instruction cache there.
There are a few things that require the data caches too, like atomic accesses on ARM. Usually we were doing something fairly short though in our realtime code, so it was easy enough to map just the memory it needed as uncacheable.
I almost mentioned page faults, but that's something the kernel has control over. It could just make sure everything is in memory so there aren't any faults. So it's not really an issue I think.
Is there any known upper bound on, say, how long a memory access instruction takes on x86?