Whilst this is an excellent post from vLLM, one of the truly baffling things from either their team or AMDs team, is how much the workstation grade AMD r9700 has been ignored.
Stock vLLM runs so slowly on these cards compared with vLLM forks like Radiance. Going from say 20-30t/s gen, to 150-200t/s
Most of AMD/vLLM work seems to be around their data centre cards, or the AMD AI Halo/Ryzen and ignores the R9700 AI Pro.
> ...one of the truly baffling things from either their team or AMDs team, is how much the workstation grade AMD r9700 has been ignored.
It makes a huge amount of sense after considering AMD's approach to graphics cards from around 2010 to 2025. They just didn't see graphics cards as viable compute platform and many who made the mistake of believing that good specs would translate into in-practice performance got badly burned. I'd have been involved in the AI boom but for an expensive AMD graphics card, I'm not going to forget that for a while.
George Hotz was interesting as a public example, but I think his story probably repeated a few times outside the public eye. People tried to make AMD work and ended up the worse for it.
People who had an interest in using AMD cards to get things done are probably by and large waiting for a new generation of hopefuls to prove this time is different. The mutterings out of AMD are promising, but that isn't persuasive enough given the scale of the failures.
You can get those numbers with https://codeberg.org/ggz14/radiance-vllm-mxfp4. I also can get it on a single R9700 but the 75 ~ 80 t/s is only peak acceptance of very predictable tokens like coding or json, and averages lower for prose. It's still much faster than regular llama.cpp.
Lemonade is one of the worst performing options. Run any modern Linux distro and ask your current LLM to setup llama.cpp with dflash2 for you as an unprivileged container running from a systemd user unit.
Obviously only on a system you do not trust at all.
GCN was such a promising compute architecture, AMD even pioneered stuff like async compute and compute shader heavy rendering pipelines, only to never seriously go beyond that on consumer gear.
I agree with your assessment that the story of supporting the competition, only to get burned, has repeated many times with AMD outside the public eye. It's why I don't put much stock in claims that things work great as long as specific flags are used.
Similar here peak ~250 and down to ~120 as it gets close to 128k (which is where I set DSH compaction) though it can readily do 256k.
I just got DeepSeek Harness (DSH) set up with 2x R9700 and it's rather mind blowing that these can do actual work and quickly. Up until now I've always been evaluating and searching for better hardware/model/tweaks. This is much more than I even hoped for and considered getting extra 3090/4090. Now I can stop looking/tweaking and start using it for all the different things I've yet to discover it's good for. I do plan to also try/use Hermes and Pi. DSH is annoying that every plugin install/remove requires a restart--given that "everything's a plugin".
Thanks! Didn't expect to see this here. Exactly what I needed to run Qwen3.8-27B-Quark-AWQ-MXFP4-native.gguf as well as other experiments on one or 2x R9700's (I hope).
Yeah, as a business AMD should first care about getting their DC grade hardware optimized for inference workloads. It's unfortunate that most of HN discussion has devolved to me-ish.
I'd rather buy two used rtx3090 than a single r9700 AI pro. More VRAM (some wasted due to it being non continuous), more RAM bandwidth, more aggregate compute.
Only if AMD made a card like this with 48G+ I'd consider it.
Also these 20-30t/s jumping to 150-200... Watch out for the massaged numbers coming from vendors.
I believe Intel has claimed something like 1400tok/s (generation! Not prefill) of Qwen3.6-moe on Arc b70.
I was actually very interested in this so I checked the details. Turns out it was 200 simultaneous users running the same 1024 token prompt :D so all the experts got maximum parallelism.
How often are you going to run 200 parallel sessions with a tiny context and same prompt running at 7tok/s.
Based on how much my rtx3090 is getting on a single user (150tok/s) I'm estimating b70 to probably get less than that.
Sadly nvidia is king now.
Also, most of us already have nvidia cards and no inference software supports mixing let's say nvidia, Intel and amd cards in inference of one model.
2x RTX 3090 is not enough for proper use. ideal is 64GB+ so that you get proper cache and 256k context with good enough models. (e.g. Qwen 3.8 27B with MXFP4).
And that is tight already.
You would need 3x RTX 3090 - but then the PCIe bandwidth comes an issue if you really want tensor parallelism with three cards. 3x PCIe5 x16 is not cheap with direct CPU access.
So surprisingly, 2x r9700 starts be a nice deal.
> Also these 20-30t/s jumping to 150-200... Watch out for the massaged numbers coming from vendors.
Well, luckily these are not vendor numbers. Prefill also scales almost linearly with the amount of GPUs.
I don't see how. Even on 32GB you can run Q6_K_XL quant with MTP at 200k context k=q8_0, v=q5_1. So 48GB VRAM is good enough to run Q8 at long context. Also with tings like ninfer and it's various forks I'm seeing people get very good performance out of Qwen3.8 models on all sorts of NVIDIA cards.
You also want prompt/prefix cache. Otherwise there is a lot of duplicate prefill processing if you fork conversation, have a different chat window or anything like that. It therefore also makes subagents much faster.
> I have had direct contact with members of the AMD RTG team and I was disgusted to find that AMD doesn't even provide them with hardware to work on. The developer I was working with had to buy the GPU he was writing drivers for.
An NVIDIA consumer GPU sells for 50+% or more than an equivalent AMD GPU. Because people are buying NVIDIA GPUs to run local models instead of AMD ones.
I did the same thing, I paid 50% more to get an 5070 Ti instead of the equivalent AMD.
This is probably good for gamers, AMD GPUs are not price inflating to the same degree as NVIDIA, because they are bad at LLMs.
> That was the reason for comparing them in the first place: based on performance, they are direct competitors, or at least they are meant to be. However, as things stand today, there is a massive price divide between the two, with the RTX Ti GPU now commanding a premium of more than 50%.
AMD also deprecates cards and throw in wrenches for cards that are hot in used markets. Latest ROCm kind of works on MI50 but requisite files are taken out of just standard Ubuntu installation. They truly don't understand marketing.
They are actually great at LLMs but you need to invest in keeping up with community tuning efforts. But I am fine with most thinking they are bad at LLMs, because I keep buying more of them!
Treating the browser's "disable cookies" feature as a way to reject consent is not real consent. That cripples many legitimate use cases outright; it's neither accessible nor understandable by normal users; it's a technical defence measure, not a way to consciously reject contractual consent.
In contrast, the GDPR demands that you properly ask for consent if you want to process somebody's personal information, inform them why that is necessary, and only process the data if they agree to the processing.
There is clearly a difference here, and IMHO the EU is quite correct here.
I tried to use it, but it didn't remember the "no" answer. Every time I loaded a page, the same confirmation for the same cookie was presented again and again.
Yeah, but that's still way too narrow to capture what the law is about. The GDPR doesn't really care about cookies, or storing data on clients in some way. Instead, it's about end-users giving informed consent to processing their data. Not just by hand-waving away some disclaimer, but actually conscious of the consequences of that action, and why it is necessary to do so.
I know this sounds all lofty and Brussels ivory-tower-ish, but I'm absolutely convinced it's the only sensible way to deal with personal information - even if American companies insist on forcing a new normal of lacking privacy on all of us.
Yeah, I was more talking about ePrivacy cookie banners, which really are about storing data on user devices. The whole thing exists because the already implemented technical solution was deemed inadequate.
It's always interesting watching counties try to lift the birthrate, and doing things that don't actually help improve the conditions that lead to the issue in the first place.
Look here's an idea. If you're under a certain age, and you want to own a home and move to a small village etc.. you are granted 100% protection from the government from your employer and you can work 100% remote. Enshrined into law. Nothing the employer can do about it.
Make it a whole nationalism thing to get the right wing on board.
I know a whole bunch of people who would move to a country town in a heartbeat if they could, heck lots of people did it during the pandemic, the minute that remote work became a possibility.
It's great they're chasing these digital nomads, but I'm not sure that's the solution that's going to work long or even medium term.
I understand that there are a lot of people who tend to believe the events that were given via the presentation, I still have my doubts, as I tend to view OAI as unreliable narrators.
I would love to believe it, sounds really amazing, but it also can be interpreted as a company who needed good marketing for their models security capabilities after their competition's model had it's moment to shine in the security sphere.
For me this is already a better value proposition, as less value add happens in the US. With the US being a perpetrator in trade war against the EU, even this matters. Everything counts, in large amounts...
They found a cheap way to please the kind of customers who care much about their activism and not much about reality. A good business decision with little risk and a potential upside.
To believe your data is any better protected here or there is already unrealistic, and the admiration for the European Union for data privacy among hacker circles is unfounded. But if many people believe something false, you as a business should give them what they ask for and not try to educate them.
Yeah this I think captures my initial feelings in my post better than I could have put it.
Any business that has strict guidelines won't touch this. As there's no real guarantee, and more importantly they have given themselves an out.
Whilst I like fastmail as a product, personally I would have waited to get everything in the EU before I launched this, as you've now got to relaunch it once you solve that last mile problem. Which is hard and costly.
To respond to both of you above. Setting up two of these at once would have increased both the up-front cost (hardware for something like this is well north of a million USD) and the complexity and associated risks. Yes, it's partly a "people want it for their own reasons and we should offer it if it's viable" and partly a "it's good to have service in multiple jurisdictions and experience with them in advance of any further balkanization of the internet".
Re-launching might be costly, but dropping millions more on a second location and splitting our systems more would hav been much riskier (note: many of our customers with 'fastmail.com' addresses have chosen the EU region, we can't segment MX records at a tighter boundary than domain level).
So we do what we can - when you login with username and password, the password is only sent to your region's server (based on a lookup from the username which MUST be global, so it works regardless of which of our servers you hit). If you send a username which doesn't exist, we distribute you to a random region in the same percentage, so you can't use it as an existence oracle.
To think that your data is safe in the EU while the EU is actively pursuing ID checks for social media and pushing Chat control every 6 months is delusional. No, your data is not safe here. If the EU wants it, it will get it.
This cult of the EU privacy needs to stop. The EU wants the same access that the US intelligence has but for some reason, some people don't believe it and defend tooth an nail this idea that things are better here.
Just so you are aware, Europol was lobbying to have access to all text messages/emails in the EU at will without a warrant as part of Chat Control V2. Say what you want about the 5 eyes countries, this is no better.
If tomorrow the EU wants access to your data, Fastmail will give it just like it will give it to the US, to the UK or to Australia.
Anyone who cares about security won't accept "reassurance" anyway. They would use end to end encryption like PGP or similar and not worry about the middlemen.
But I agree with you, it's biggest weakness is that for a real long time the tagline of it was "there are many harnesses, this one is MINE" (That being Mario's)
I have a lot of respect for Mario and his team, but there's things like you've pointed out that deviate from standards, and other issues that I've seen get posted, only to get knocked down by the team as WON'T FIX because, even though the new owners changed the tagline from MINE to YOURS... It's still very much Mario's.
I do like opinionated things. Truly. But I'm also of the opinion that standards exist for a reason.
That said. I like Pi so much that it's my daily driver, and I've created an ecosystem of plugins to do everything I want, having them all tie together and communicate through the shared bus. Pi is really a good harness.
It's just, well. I don't agree with some of the opinions.
If I'm going to add another thing here... Whilst you cannot get everything you need from the openAI API spec, you can get a surprising amount to get a model config. That said. Versions of Pi are still shipping with model configs for certain inference providers. I do hope that gets decoupled at some stage. I see the groundwork being laid.
So the work is being done in the right direction. I applaud the team but I do get the feeling that a lot of this is because people want to contribute, but the team really wants to hand craft this. And that's great
Ignoring the military stuff, I feel like it is saying that this harness, although stamped from a mass produced part, is mine once I take possession of it. An extension of me?
My read sees pi as a starter kit. It's job is to build personal workflows and not to dictate a workflow. Copy on the site calls it minimal and tells us to 'adapt pi to your workflow, not the other way around.' In other words a fleshed out pi is unique to you.
I'm autistic myself. I don't want a "cure" but I would like to have something that could help me reduce some of the harder aspects of life.
Right now I'm on anti anxiety drugs, which whilst mostly affective, have sided effects that kind of suck. I'm lucky my wife and daughter love me very much, and I've always been someone who wanted to push myself a little further to get outside or do new things.. mostly because my daughter is also autistic and I want to be a good example of "small steps" ... But it was so hard with no medication that I eventually had a full blown autistic burnout, where therapy, time off work, and eventually medication helped me get back to work.
Yes I want society to be more accommodating to us, but that doesn't stop the multitude of micro-aggressions I suffer daily just trying to navigate life, relationships, work, etc. If something could help remove those, I'd be forever grateful.
I do not however want the way my mind works to change.
I am very lucky to be in a country where my workplace is required to provide "reasonable accomodations" for me. There's still a long way to go to actually address a lot of the problems that myself and other autistic people have to deal with on a daily basis
The important point for me is that I can see what that world would look like and it doesn't depend on "fixing" me
Same. Norway here, and yeah whilst some work places will be happy to make larger accommodations, most will only do what they are forced too. Working as a SWE has helped me have more accommodating workplace conditions, it was really the last 6 years since work from home did I truly realise that was the only one change a company could give me that made the biggest difference.
Whilst I can take meds to make things better for anxiety, I will never be able to take meds to understand subtext or subtle facial expressions. It can take me weeks to realise that someone meant the opposite. There's no meds that could fix that.... Or if there is. Wow! I'd take it.
It's really nice to hear from someone who's had some of the same experiences as me dealing with autism.
People look at me weird when I say that I have never had it as good as I did during COVID lockdowns, but it was amazing that my entire world was my apartment and I didn't have to deal with anything besides what was within those four walls
I agree with your distinction/example about meds. If there was something that could improve the things I struggle with such as not being able to recognise faces, interpret facial expressions, etc then I'd be very happy to take it
Unfortunately I think these things are intrinsically interconnected with the parts of being autistic I love
Anyway - thanks so much for sharing your perspective and experience. I really appreciated hearing about it
reply