Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I advise a talk from Chandler Carruth where he proves longer code can achieve higher performance due to the way computer architectures work.

Unfortunately I no longer remember in which conference he did it, maybe someone else can link it.



In microbenchmarks, on a long pipeline RISC, or similar microarchitectures like NetBurst (P4), I could see that being true. But we're long past that era now. It's the same misguided assumptions that leads to several-KB-long "optimised" memcpy() implementations whose real-world effects on performance are negligible to negative.

If you don't believe me, read what Linus Torvalds has to say about why the Linux kernel is built with -Os by default.


The longer code is typically generated because the compiler will generate vectorized code that provides enormous speedups in case of longer data sets. Take, for example, this code: https://godbolt.org/z/WEx3Gb5jr

At -O2 the assembly it generates is straightforward, and in line with what a human programmer would write. At -O3 it generates vector code that needs a lot more instructions (vector pipeline setup, code to deal with the remaining elements that don't entirely fill up a vector register, etc.) but the main loop takes 4 integers at a time instead of one, so that provides a nice 4x speedup. In order to achieve that it needs 25 instructions to set up the loop / finish the remaining elements, compared to 5 instructions for the -O2 code.

For very short loops the -O2 version will have superior performance, but for runs of data from around 8 integers (wild guess) the -O3 version will begin with pull ahead. So it really depends on the type of data your program is handling, whether it is better to optimize for speed or size.


My recent tests with -Os resulted in distinctly negative effects on performance.

But the main problem with -Os is that it is poorly exercised. The best-exercised modes are -O0 and -O2, so those are the ones to use in production.


> My recent tests with -Os resulted in distinctly negative effects on performance.

Err... yeah, because -Os means "optimize size". Not "speed".

> But the main problem with -Os is that it is poorly exercised

No, the main problem is that you don't understand -Os :) It works as intended.


Evidently you are unaware that use of -Os has on earlier (but quite recent) generations of CPU architecture resulted in notably faster performance.

And, that any compiler feature that is little used will necessarily receive less attention than commonly used features, and be less stable and reliable. Before trying to fix any bug in a program built with -Os, reproducing it first in -O2 will reduce premature balding.


> Before trying to fix any bug in a program built with -Os, reproducing it first in -O2 will reduce premature balding.

I have no idea what you're talking about. Debug your programs in -O0, which means "no optimizations". -Os is, and has always been, optimizing for executable SIZE. It has no guarantees wrt performance.


[flagged]


Please debate in good faith instead of resorting to snide remarks.

Os has never made any guarantees of speed or performance. Any cases where performance increased over O2 are platform specific and are anomalous.

I've never made any statement on edge cases, only guarantees and intent. I have nothing to regret.


He did it on x64.

EDIT: This is the talk, if I remember correctly.

https://isocpp.org/blog/2018/06/cppcon-2017-going-nowhere-fa...


I remember this being talked about in a talk about the Coz profiler. Maybe that was it?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: