> Slug doesn't seem to support any of these, it is just "for a given point, am I in or am I out?", but it doesn't tell you by how much
It does give you an anti-aliased value between 0 and 1 that estimates how much of a pixel is being covered.
But this is a linear estimate based on horizontal and vertical distance to the Bezier curve. It does not look correct at long distances, which is why you shouldn't use it for outlines, drop shadows or the other cool things you can do with (M)SDF. A single pixel outline works fine but is not really legible with modern display resolutions (very thin lines).
For finding the minimum distance between a quadratic Bezier curve and a point would require solving a 3rd degree polynomial, where Slug's algorithm gets away with solving a quadratic equation per pixel. This is makes a big performance difference.
As hobby project, I wrote a font rasterizer algorithm (two algorithms actually) that produces pixel perfect anti-aliased images (like Slug) using a similar but different algorithm. Rather than using Slug's clever root classifier for Bezier curves, I subdivide Bezier curves into monotonic sections where evaluating the winding number is much simpler and can be done in parallel for a group of pixels. It works nicely on GPU with warp/wave/subgroup operations and CPU using SIMD.
At first glance, subdividing the Bezier curves sounds like a bad idea (more Beziers to rasterize) but it opens doors for some parallelism, and most Bezier curves that appear in fonts are monotonic in the first place (so the increase is very modest). This was inspired by this entertaining but not very serious video about font rasterization [0].
The first parallelism optimization is checking against the curve bounding box vs. a rectangular (in uv-space) region of pixels, and this can quickly determine if the Bezier needs to be evaluated in the first place. This can be done per GPU warp.
The second optimization works only for rectilinear transformation (no rotation, skew or perspective). Solving the quadratic equation involves a square root and a division (which alone are >30% of the computation), which can be computed for each row and column of pixels instead of for each pixel (2n instead of n^2).
Both optimizations rely on mathematical invariants of monotonicity, ie. the derivative of the Bezier curve must be non-zero. All Bezier curves can be robustly subdivided into monotonic sections using de Casteljau's algorithm.
My simple benchmarks compare favorably to Slug on the GPU and to "fast" rasterization algorithms on the CPU (which is an order of magnitude faster than "fancy" rasterization algorithms with hinting etc).
Unfortunately there are so many hobby projects and so little time. All I have is messy shaders that draw individual characters and a few benchmarks to see how quickly (and something similar for the CPU). Going from there to a complete text rendering system would be a lot of work. Writing a more detailed article with illustrative code examples is something I'd want to do but haven't gotten around to.
If you want to offer words of encouragement or geek out about rasterization algorithms, I welcome any input.
It's definitely a 80% solution where you occasionally need to drop down to intrinsics (at zero runtime perf cost) for CPU specific instructions.
But just having vector types, arithmetic, swizzling, loads and stores will go a long way for basic tasks.
And with generics you can write code that is type and width agnostic. No need to rewrite your code of you want to go from SSE to AVX512, just change from f32x4 to f32x16 (or use generics) and you are done.
Because there are platform vendors. And SIMD performance very much depends on the use-case, which is a balance of practicalities and specifications and intended deployment targets ..
I also think this is a deployment problem, not a build problem, but okay ..
> also think this is a deployment problem, not a build problem ...
This I agree with, deploying and running code for the correct cpu is a problem with no established solution.
As for actually writing the code, portable_simd is great. You need to adjust simd width and compiler config for the cpu you deploy to and fill in the blanks with intrinsics. Which is much less work per target than writing it all with raw intrinsics if you are deploying to more than one target.
I wrote a multi-target SIMD-using synthesizer, and I don't think I would've had as much success if I were just using someone else's library - its been especially important to have the SIMD instructions for both architectures I'm supporting (ARM and x86) directly available in the code, since a synthesizer necessarily involves a working pipeline from one stage to the other. I wonder how much easier/better the code would have been to write with portable_simd ..
> I wonder how much easier/better the code would have been to write with portable_simd ..
I had a quick glance of the code and it seems to use mostly basic arithmetic instructions on fixed width simd vectors using some helper macros like SIMD_MUL(x, y) for _mm_mul_ps, etc. Some explicit simd intrinsics code for stuff like sin.
That would've been pretty easy to write with portable_simd in Rust or the equivalent C/C++ language extensions (or maybe C++26 std::simd).
You'd just use f32x4 (or add a typedef with attribute in C++) and then use x*y instead of SIMD_MUL.
For basic stuff like this, you should get the exact same generated code.
> They specifies a constant SIMD width so it's non-portable.
This is incorrect, you can use vectors wider than native SIMD width and the compiler will break them down to register size of the target cpu.
In fact it's sometimes better to used wider than native width, in some applications I see 20% better throughput with f32x16 (512 bits) on an AVX2 CPU (256 bits). It is kinda like loop unrolling it.
Except you can't use this in actual code, because either, as is the case in this example with f32x32, you run out of registers and spill all over the place.
Or you aren't using your full vector register or could've gotten better performance by "unrolling" more often for the larger vectors.
If you use f32x16 (the avx-512 wisth), SSE now effectively has 4 registers to work with and will spill when doing anything beyond the most simple stuff.
The default should imo be relative to the native register width, so you can do 1x, 2x or sometimes 4x the native width, depensing on your register preasure.
I can and I do use this is "actual code" and I've got benchmarks to prove that it's got better throughput (for the particular use case, don't extrapolate from there) and the same applies to AVX2 and AVX512: twice the native vector width has ~20% better throughput (ie. using `f32x32` on AVX-512).
I pass in the vector width as a generic parameter like this:
With this I can easily benchmark the same code for any vector width. I can also do some compile time heuristics to choose the vector width based on what's available on the compile target CPU.
> you run out of registers and spill all over the place
As usual when optimizing SIMD code, you should keep an eye on the generated disassembly and the benchmark results and watch for register pressure and the other usual things.
I'm definitely NOT saying that you always get the best perf by using 2x SIMD width, but in this particular case it was so.
This is much much easier to do with portable_simd than if you'd write the same with intrinsics, you can change the SIMD width without having to rewrite all your code (e.g. changing from SSE `_mm_add_ps` to AVX `_mm256_add_ps` etc).
It's still a partial solution, you still need to drop down to intrinsics for some special instructions every now and then (which is easy), but in my projects this accounts for much less than 1% of the lines of code. Not applicable everywhere of course.
> twice the native vector width has ~20% better throughput
Yes, this is what I was saying, but twice the vector width of AVX-512 will perform horrible in SSE, which is why portable SIMD abstractions should make writing code relative to the native vector width simple.
> I pass in the vector width as a generic parameter like this:
My problem is that no portable_simd example code I've seen does this, which causes people to choose one specific N and run with that.
The second part of the problem is how you find the native vector length, so you can instantiate the generic function. IIRC this isn't even exposed in portable_simd and you have to use a seperate crate to get it.
> The second part of the problem is how you find the native vector length, so you can instantiate the generic function. IIRC this isn't even exposed in portable_simd and you have to use a seperate crate to get it.
This is trivial (but not pretty!) to do with something like `#[cfg(target_feature = "avx2")] const SIMD_WIDTH: usize = 8`. You need a few lines of ugly cfg logic to configure this.
A somewhat orthogonal and much more difficult problem is how to select it at runtime. You would either need to have different binaries built with different compiler options, link object files built with different compiler options to same binary, or dynamically link the correct code at runtime.
This is actually one of the (IMO only) cases where intrinsics are more practical: you can use `_mm256_add_ps` from AVX2 intrinsics regardless of whether you've configured your compiler to support AVX2 or not. As long as you check at runtime before calling the code so you don't get illegal instruction exceptions.
> misdirect Ukrainian drones toward NATO airspace (which has happened).
Claims to such effect has been made in public by reputable persons but they are not very credible nor has any supporting evidence been shown in public.
Drones getting lost due to GNSS jamming is happening.
But intentional misdirection towards a target would require GNSS spoofing capabilities over an impractically large area and that the drones would be susceptible to simple spoofing. They have satellite navigation, inertial navigation, terrain following cameras and listen to ground based radio sources like cell towers and are designed to operate in hostile RF environment. Just sending a fake GPS signal (several of them) won't make the drone divert to a specific target or direction.
I'm not buying the misdirection story until more compelling evidence is made public.
I made no claims as to whether misdirection is happening or feasible, only that there are incentives for both sides to do so if it is, while there are no incentives for Russia to take on a much larger enemy while it can barely handle Ukraine, which OP was implying was the obvious explanation.
From what I have read (from unreliable sources), Ukrainian fighter pilots tried this a few years back and stopped due to too many losses.
To engage a slow flying drone with cannons from a fighter jet, they have to get really close while flying close to stall speed. They had some success in shooting down drones but the explosion and the debris caused several lost fighters.
Maybe something like a Embraer Super Tucano turboprop fighter with a gunpod could work. But that has very little other use in the battlefield.
I mean the other way round: rather then combat drones launching missiles, combat drones which use stealth and speed to close to dogfighting distances with aircraft and engage with their own cannons.
It may be that risking some drone airframes for a higher overall kill rate becomes worth it when you're at no risk of losing the pilot, and could be more affordable then the missiles you'd otherwise have to expend (e.g. when you get to things like the Meteor which are very expensive and also use up jet engine components you might otherwise want to put on somewhat more reusable drones).
The example code in this article is using Zig's portable SIMD features. Similar features are available for C/C++ (GCC/Clang extension) [0], (nightly) Rust [1] and C++26 [2].
All of these provide a similar set of features and you can use normal arithmetic operations (+, -, *, etc) for SIMD vectors. Together with templates/generics you can also write code that can deal with any vector width. These get compiled to LLVM vector types and will generally give you pretty good generated code.
This is a very good way of writing basic SIMD code and has the benefit that your code can be compiled to multiple instruction sets. I've been working on a project that can compile down to SSE2, AVX2, AVX-512 and NEON, with just a change of compiler options. Somewhat surprisingly I get the best performance by using 2x the native vector width (ie. f32x16 = 512 bits on 256 bit AVX2), which is kinda like unrolling the loop once.
There are some caveats, though. You will need to keep an eye on the generated assembly code to make sure you're on the happy path. You will inevitably need to drop down to ISA specific intrinsics every now and then (for that fast reciprocal square root with `__mm_rsqrt_ps` etc).
As an example I needed to do a gather load from an array of fp16's on AVX2, which does not do 16 bit loads. Rust's `Simd::gather_select` takes 64 bit usize as the index but AVX2 doesn't do 64 bit indices. But as long as I did all the index arithmetic in 32 bits and cast to usize at the last second, the compiler did what I wanted. But you need to kinda know what is available in the ISA to stay on the happy path. Not really an issue with arithmetic.
I'm sure that an experienced SIMD programmer can get better performance by writing intrinsics manually (say 5-20% better) but I'm already at 3-6x better than the scalar implementation I started with. And you'd have to write (and benchmark) the code for each ISA separately, meaning that you'd spend at least five times more time with it (and have 5x more code to maintain).
FYI you linked to a really old version of the GCC documentation. Google apparently loves those old docs, so they often show up near the top of search results despite being ancient. For posterity, here's the latest version: https://gcc.gnu.org/onlinedocs/gcc-16.1.0/gcc/Vector-Extensi....
Yes, but at typical font sizes this generates a lot of tiny triangles (0 to few pixels) which perfoms badly on GPUs.
reply