Hacker Newsnew | past | comments | ask | show | jobs | submit | intothemild's commentslogin

Opt-in as in don't connect smart TVs to the internet.

Whilst this is an excellent post from vLLM, one of the truly baffling things from either their team or AMDs team, is how much the workstation grade AMD r9700 has been ignored.

Stock vLLM runs so slowly on these cards compared with vLLM forks like Radiance. Going from say 20-30t/s gen, to 150-200t/s

Most of AMD/vLLM work seems to be around their data centre cards, or the AMD AI Halo/Ryzen and ignores the R9700 AI Pro.

Really wish this would change.


> ...one of the truly baffling things from either their team or AMDs team, is how much the workstation grade AMD r9700 has been ignored.

It makes a huge amount of sense after considering AMD's approach to graphics cards from around 2010 to 2025. They just didn't see graphics cards as viable compute platform and many who made the mistake of believing that good specs would translate into in-practice performance got badly burned. I'd have been involved in the AI boom but for an expensive AMD graphics card, I'm not going to forget that for a while.

George Hotz was interesting as a public example, but I think his story probably repeated a few times outside the public eye. People tried to make AMD work and ended up the worse for it.

People who had an interest in using AMD cards to get things done are probably by and large waiting for a new generation of hopefuls to prove this time is different. The mutterings out of AMD are promising, but that isn't persuasive enough given the scale of the failures.


I have 4x r9700s as my coding daily drivers running qwen 3.8 27b at 80tps each. Zero complaints. Especially at $1200 each.

No way you can get that much each without extremely quants

For a single r9700 you have 637 GB/s and for qwen 3.8 27b q4_k_xl the maximum tg/s is 33 before mtp

Now if you meant 4xr9700 tensor parallelism with mtp, 80 tg/s starts to make sense


You can get those numbers with https://codeberg.org/ggz14/radiance-vllm-mxfp4. I also can get it on a single R9700 but the 75 ~ 80 t/s is only peak acceptance of very predictable tokens like coding or json, and averages lower for prose. It's still much faster than regular llama.cpp.

Can confirm. I've a single R9700 and have maxed out at 45 tokens/second on llama.cpp with Q4 Qwen 3.6 27B (with MTP)

Q4 isn't an extreme quant, and I average 75 toks/s on code, 45 tok/s on prose with MTP.

Please share what operating system and model runtime you use? I have two and don't get close to that with AMD's own Lemonade. Thanks!

Lemonade is one of the worst performing options. Run any modern Linux distro and ask your current LLM to setup llama.cpp with dflash2 for you as an unprivileged container running from a systemd user unit.

Obviously only on a system you do not trust at all.


GCN was such a promising compute architecture, AMD even pioneered stuff like async compute and compute shader heavy rendering pipelines, only to never seriously go beyond that on consumer gear.

I agree with your assessment that the story of supporting the competition, only to get burned, has repeated many times with AMD outside the public eye. It's why I don't put much stock in claims that things work great as long as specific flags are used.


Thankfully, there are still people willing to jump on the R9700 bandwagon and get a vLLM fork working. If you have an RDNA4 card check out https://hub.docker.com/r/stilldeadcode/vllm-radiance

Deadcode is currently working on INT4 right now on his R4D kernel.

The MXFP4 fork is excellent too. Its my daily driver right now. https://codeberg.org/ggz14/radiance-vllm-mxfp4

Also has PARO quant support there too (early stage)

Also speedups in both repos for 4x R9700s


The way it does tensor splitting without all-reduce cost over PCIe bus wasn't something I thought was possible.

What kind of performance are you getting with 4x R9700s--what do you do with all the VRAM (batching, concurrent requests, etc)?


personally? i have 2x gpus.. but i get bursts of ~200tok/s generation, and around 4500-5000tok/s prefil

Yeah the R4D Kernel rules imho.


Similar here peak ~250 and down to ~120 as it gets close to 128k (which is where I set DSH compaction) though it can readily do 256k.

I just got DeepSeek Harness (DSH) set up with 2x R9700 and it's rather mind blowing that these can do actual work and quickly. Up until now I've always been evaluating and searching for better hardware/model/tweaks. This is much more than I even hoped for and considered getting extra 3090/4090. Now I can stop looking/tweaking and start using it for all the different things I've yet to discover it's good for. I do plan to also try/use Hermes and Pi. DSH is annoying that every plugin install/remove requires a restart--given that "everything's a plugin".


Thanks! Didn't expect to see this here. Exactly what I needed to run Qwen3.8-27B-Quark-AWQ-MXFP4-native.gguf as well as other experiments on one or 2x R9700's (I hope).

That MXFP4 is an excellent project. But I have difficulties on reading that README. Is it intentionally generated like that with LLMs?

Feed the setup and run scripts to your LLM.

I don't see a reason why it should AMD doesn't care about lower end prosumers atm.

They might in the future but future is in the future ofc

Edit: to be clear I think it's ridiculous they don't but from a company's stand point it doesn't make much sense


Yeah, as a business AMD should first care about getting their DC grade hardware optimized for inference workloads. It's unfortunate that most of HN discussion has devolved to me-ish.

Then they should stop selling hardware they don’t plan to support. Me-ish when you spend $1,500 on a piece of hardware is completely acceptable.

Never buy hardware based on expectations of future features, especially if there’s no promise from the vendor.

I'd rather buy two used rtx3090 than a single r9700 AI pro. More VRAM (some wasted due to it being non continuous), more RAM bandwidth, more aggregate compute.

Only if AMD made a card like this with 48G+ I'd consider it.

Also these 20-30t/s jumping to 150-200... Watch out for the massaged numbers coming from vendors.

I believe Intel has claimed something like 1400tok/s (generation! Not prefill) of Qwen3.6-moe on Arc b70.

I was actually very interested in this so I checked the details. Turns out it was 200 simultaneous users running the same 1024 token prompt :D so all the experts got maximum parallelism.

How often are you going to run 200 parallel sessions with a tiny context and same prompt running at 7tok/s.

Based on how much my rtx3090 is getting on a single user (150tok/s) I'm estimating b70 to probably get less than that.

Sadly nvidia is king now.

Also, most of us already have nvidia cards and no inference software supports mixing let's say nvidia, Intel and amd cards in inference of one model.


2x RTX 3090 is not enough for proper use. ideal is 64GB+ so that you get proper cache and 256k context with good enough models. (e.g. Qwen 3.8 27B with MXFP4). And that is tight already.

You would need 3x RTX 3090 - but then the PCIe bandwidth comes an issue if you really want tensor parallelism with three cards. 3x PCIe5 x16 is not cheap with direct CPU access.

So surprisingly, 2x r9700 starts be a nice deal.

> Also these 20-30t/s jumping to 150-200... Watch out for the massaged numbers coming from vendors.

Well, luckily these are not vendor numbers. Prefill also scales almost linearly with the amount of GPUs.


> 2x RTX 3090 is not enough for proper use

I don't see how. Even on 32GB you can run Q6_K_XL quant with MTP at 200k context k=q8_0, v=q5_1. So 48GB VRAM is good enough to run Q8 at long context. Also with tings like ninfer and it's various forks I'm seeing people get very good performance out of Qwen3.8 models on all sorts of NVIDIA cards.


You also want prompt/prefix cache. Otherwise there is a lot of duplicate prefill processing if you fork conversation, have a different chat window or anything like that. It therefore also makes subagents much faster.

I'm getting 99% cache hit in DeepSeek Harness?

You don't need PCIe 5.0 x16 since RTX 30 are not PCIe 5.0 to begin with.

Well, that makes them just even slower then

The whole point of dual 3090 over other cards is the nvlink support. At that point the pcie doesn't really matter.

Oops

George Hotz in June 2023:

> I have had direct contact with members of the AMD RTG team and I was disgusted to find that AMD doesn't even provide them with hardware to work on. The developer I was working with had to buy the GPU he was writing drivers for.


I think this has changed since then, their policies toward open source improved (e.g ROCm).

The market says the problem is still there.

An NVIDIA consumer GPU sells for 50+% or more than an equivalent AMD GPU. Because people are buying NVIDIA GPUs to run local models instead of AMD ones.

I did the same thing, I paid 50% more to get an 5070 Ti instead of the equivalent AMD.

This is probably good for gamers, AMD GPUs are not price inflating to the same degree as NVIDIA, because they are bad at LLMs.

> That was the reason for comparing them in the first place: based on performance, they are direct competitors, or at least they are meant to be. However, as things stand today, there is a massive price divide between the two, with the RTX Ti GPU now commanding a premium of more than 50%.

https://www.techspot.com/review/3168-geforce-rtx-5070-vs-rad...


> AMD GPUs are not price inflating to the same degree as NVIDIA, because they are bad at LLMs.

9060 XT 16GB seems to have some great performance with gpt oss 20B and others, and works great with their lemonade-server.

https://lemonade-server.ai/

What exactly do you think is running on a Strix Halo?


wondering when AMD will realize it can charge 2x as much for the same thing, by simply finally writing a fucking driver

The story of “great hardware ruined by poor drivers/software” is so old than its truly shocking its still a thing these days

On Linux the problem is often flipped (historically for Nvidia/AMD).

The AMD drivers are mainline kernel, while Nvidia is still handing out binary blobs.


AMD also deprecates cards and throw in wrenches for cards that are hot in used markets. Latest ROCm kind of works on MI50 but requisite files are taken out of just standard Ubuntu installation. They truly don't understand marketing.

That is a ~9 year old card. I'm not sure any of the newer work AMD has been doing can utilize that hardware at all.

It would have been sold off to grey market or destroyed 3-4yr ago in most companies.

Newer, much cheaper stuff ironically would run better. That card isn't even really designed for AI workloads at all.


> because they are bad at LLMs.

They are actually great at LLMs but you need to invest in keeping up with community tuning efforts. But I am fine with most thinking they are bad at LLMs, because I keep buying more of them!


Honestly a big part of this is AMD's lackluster strategy to GPU software. To say it's lacking is understatement. At least for non-data center gpus.

The cookie banner isn't actually specified in gdpr, it was just how everyone else tried to build the solution to the problem at the last minute.

I remember thinking "ok once this hits an actual web spec, we should see this built into browsers, and sent as headers or something"

Nope


It's a case of industry malicious compliance


i was thinking the same thing. it could be like an actual element of your page.


Browsers already had a way to consent to cookies, since the invention of cookies themselves. But the EU didn't consider that _real_ consent.


Treating the browser's "disable cookies" feature as a way to reject consent is not real consent. That cripples many legitimate use cases outright; it's neither accessible nor understandable by normal users; it's a technical defence measure, not a way to consciously reject contractual consent.

In contrast, the GDPR demands that you properly ask for consent if you want to process somebody's personal information, inform them why that is necessary, and only process the data if they agree to the processing.

There is clearly a difference here, and IMHO the EU is quite correct here.


Correct.. the gdpr isn't the anti cookie law. It's the data privacy law.

If it wasn't cookies it would be something else.


Most browsers in the 00s had a "always ask" option for cookies. Nobody used it, because it's as annoying as gdpr dialogs now are. But it existed.


I tried to use it, but it didn't remember the "no" answer. Every time I loaded a page, the same confirmation for the same cookie was presented again and again.


Yeah, but that's still way too narrow to capture what the law is about. The GDPR doesn't really care about cookies, or storing data on clients in some way. Instead, it's about end-users giving informed consent to processing their data. Not just by hand-waving away some disclaimer, but actually conscious of the consequences of that action, and why it is necessary to do so.

I know this sounds all lofty and Brussels ivory-tower-ish, but I'm absolutely convinced it's the only sensible way to deal with personal information - even if American companies insist on forcing a new normal of lacking privacy on all of us.


Yeah, I was more talking about ePrivacy cookie banners, which really are about storing data on user devices. The whole thing exists because the already implemented technical solution was deemed inadequate.


Correct in principle, completely ineffectual and annoying in practice.


Shaka, when they discovered open models!


omg. so awesome


Shaka, when the bubble burst!


Zinda, his face black, his eyes red…


It's always interesting watching counties try to lift the birthrate, and doing things that don't actually help improve the conditions that lead to the issue in the first place.

Look here's an idea. If you're under a certain age, and you want to own a home and move to a small village etc.. you are granted 100% protection from the government from your employer and you can work 100% remote. Enshrined into law. Nothing the employer can do about it.

Make it a whole nationalism thing to get the right wing on board.

I know a whole bunch of people who would move to a country town in a heartbeat if they could, heck lots of people did it during the pandemic, the minute that remote work became a possibility.

It's great they're chasing these digital nomads, but I'm not sure that's the solution that's going to work long or even medium term.


I understand that there are a lot of people who tend to believe the events that were given via the presentation, I still have my doubts, as I tend to view OAI as unreliable narrators.

I would love to believe it, sounds really amazing, but it also can be interpreted as a company who needed good marketing for their models security capabilities after their competition's model had it's moment to shine in the security sphere.


It's 100% this for me too. If I hear an AI voice, I hang up and try somewhere else.


Wow they completely missed the ball on why people want reassurances that their data stays in the EU


For me this is already a better value proposition, as less value add happens in the US. With the US being a perpetrator in trade war against the EU, even this matters. Everything counts, in large amounts...


I think they just know what they can and can't guarantee.


Agreed. I imagine they’ll announce an EU-only option later once they get more EU infrastructure in place


As long as they are owned by a 5-eyes nation, there's no alternative they can offer.

Instead, look to Exoscale, Proton, Infomaniak, or Scaleway.

Disclaimer: I work for one of 'em.


They found a cheap way to please the kind of customers who care much about their activism and not much about reality. A good business decision with little risk and a potential upside.

To believe your data is any better protected here or there is already unrealistic, and the admiration for the European Union for data privacy among hacker circles is unfounded. But if many people believe something false, you as a business should give them what they ask for and not try to educate them.


Yeah this I think captures my initial feelings in my post better than I could have put it.

Any business that has strict guidelines won't touch this. As there's no real guarantee, and more importantly they have given themselves an out.

Whilst I like fastmail as a product, personally I would have waited to get everything in the EU before I launched this, as you've now got to relaunch it once you solve that last mile problem. Which is hard and costly.


To respond to both of you above. Setting up two of these at once would have increased both the up-front cost (hardware for something like this is well north of a million USD) and the complexity and associated risks. Yes, it's partly a "people want it for their own reasons and we should offer it if it's viable" and partly a "it's good to have service in multiple jurisdictions and experience with them in advance of any further balkanization of the internet".

Re-launching might be costly, but dropping millions more on a second location and splitting our systems more would hav been much riskier (note: many of our customers with 'fastmail.com' addresses have chosen the EU region, we can't segment MX records at a tighter boundary than domain level).

So we do what we can - when you login with username and password, the password is only sent to your region's server (based on a lookup from the username which MUST be global, so it works regardless of which of our servers you hit). If you send a username which doesn't exist, we distribute you to a random region in the same percentage, so you can't use it as an existence oracle.


Absolutely spot on.

To think that your data is safe in the EU while the EU is actively pursuing ID checks for social media and pushing Chat control every 6 months is delusional. No, your data is not safe here. If the EU wants it, it will get it.

This cult of the EU privacy needs to stop. The EU wants the same access that the US intelligence has but for some reason, some people don't believe it and defend tooth an nail this idea that things are better here.

Just so you are aware, Europol was lobbying to have access to all text messages/emails in the EU at will without a warrant as part of Chat Control V2. Say what you want about the 5 eyes countries, this is no better.

If tomorrow the EU wants access to your data, Fastmail will give it just like it will give it to the US, to the UK or to Australia.


they also said in the article it is dependent on them standing up a second EU region. this is simply an announcement of their first.


Anyone who cares about security won't accept "reassurance" anyway. They would use end to end encryption like PGP or similar and not worry about the middlemen.


I really really like pi.

But I agree with you, it's biggest weakness is that for a real long time the tagline of it was "there are many harnesses, this one is MINE" (That being Mario's)

I have a lot of respect for Mario and his team, but there's things like you've pointed out that deviate from standards, and other issues that I've seen get posted, only to get knocked down by the team as WON'T FIX because, even though the new owners changed the tagline from MINE to YOURS... It's still very much Mario's.

I do like opinionated things. Truly. But I'm also of the opinion that standards exist for a reason.

That said. I like Pi so much that it's my daily driver, and I've created an ecosystem of plugins to do everything I want, having them all tie together and communicate through the shared bus. Pi is really a good harness.

It's just, well. I don't agree with some of the opinions.

If I'm going to add another thing here... Whilst you cannot get everything you need from the openAI API spec, you can get a surprising amount to get a model config. That said. Versions of Pi are still shipping with model configs for certain inference providers. I do hope that gets decoupled at some stage. I see the groundwork being laid.

So the work is being done in the right direction. I applaud the team but I do get the feeling that a lot of this is because people want to contribute, but the team really wants to hand craft this. And that's great


I might be misreading this, but I thought it was a reference to "Full Metal Jacket" => https://www.youtube.com/watch?v=YoU2hlDJmFE ?

Ignoring the military stuff, I feel like it is saying that this harness, although stamped from a mass produced part, is mine once I take possession of it. An extension of me?


FMJ references this, I believe - https://en.wikipedia.org/wiki/Rifleman%27s_Creed


My read sees pi as a starter kit. It's job is to build personal workflows and not to dictate a workflow. Copy on the site calls it minimal and tells us to 'adapt pi to your workflow, not the other way around.' In other words a fleshed out pi is unique to you.


Could you share some of the ecosystem/plugins/workflows you're using with Pi?


Sure, I use this, as pretty much my only plugin, I also have one for web-search on top of this. https://github.com/danielcherubini/pi-archimedes

as for workflows, its all skills/agents based, here's my dotagents folder

https://github.com/danielcherubini/dotagents


I'm autistic myself. I don't want a "cure" but I would like to have something that could help me reduce some of the harder aspects of life.

Right now I'm on anti anxiety drugs, which whilst mostly affective, have sided effects that kind of suck. I'm lucky my wife and daughter love me very much, and I've always been someone who wanted to push myself a little further to get outside or do new things.. mostly because my daughter is also autistic and I want to be a good example of "small steps" ... But it was so hard with no medication that I eventually had a full blown autistic burnout, where therapy, time off work, and eventually medication helped me get back to work.

Yes I want society to be more accommodating to us, but that doesn't stop the multitude of micro-aggressions I suffer daily just trying to navigate life, relationships, work, etc. If something could help remove those, I'd be forever grateful.

I do not however want the way my mind works to change.

Keep the positive, remove or dull the negative.


I'm 100% with you

I am very lucky to be in a country where my workplace is required to provide "reasonable accomodations" for me. There's still a long way to go to actually address a lot of the problems that myself and other autistic people have to deal with on a daily basis

The important point for me is that I can see what that world would look like and it doesn't depend on "fixing" me


Same. Norway here, and yeah whilst some work places will be happy to make larger accommodations, most will only do what they are forced too. Working as a SWE has helped me have more accommodating workplace conditions, it was really the last 6 years since work from home did I truly realise that was the only one change a company could give me that made the biggest difference.

Whilst I can take meds to make things better for anxiety, I will never be able to take meds to understand subtext or subtle facial expressions. It can take me weeks to realise that someone meant the opposite. There's no meds that could fix that.... Or if there is. Wow! I'd take it.


It's really nice to hear from someone who's had some of the same experiences as me dealing with autism.

People look at me weird when I say that I have never had it as good as I did during COVID lockdowns, but it was amazing that my entire world was my apartment and I didn't have to deal with anything besides what was within those four walls

I agree with your distinction/example about meds. If there was something that could improve the things I struggle with such as not being able to recognise faces, interpret facial expressions, etc then I'd be very happy to take it

Unfortunately I think these things are intrinsically interconnected with the parts of being autistic I love

Anyway - thanks so much for sharing your perspective and experience. I really appreciated hearing about it


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: