Hacker Newsnew | past | comments | ask | show | jobs | submit | iamniels's commentslogin

I switched to GLM-5.3 flash on high for this reason. Too many "but wait" in the Deepseek-v4 reasoning.


GLM 5.3 Flash was the killer for me. It really feels like we have Claude-approaching models at home.

I just wish they kept parameter count down in order to fit entirely within commonly used RAM sizes


I run a website with 10k unique pages. If I leave the gates open, Meta hits it 200.000 times per day. Every day. What are you paying developers $500k for Mark?


200k req/day = 2.3 req/s

That's bad for static content.

Try adding search, with pagination and 16 filters that can toggled on off. And 1000 tags, give each page 5.


I've simply blocked whole ASNs for Meta, Google, AWS, etc.


This.

Google is an advertising company with an enterprise SaaS branch. I bet they have chosen to focus on running the most efficient "everybody" model instead of running a heavy model for coders.

When using Gemini for other tasks than coding, it is actually pretty good. It grounds well with Google search and gives mostly correct, well written answers to many niche questions.


PPWR is not the only one. Don't forget WEEE and the battery regulation. Only for VAT we have a single entrypoint, for all other regulations, businesses have to submit for each country. Its simply impossible.

On the other hand I estimate that between 50% and 200% of the taxes collected like this are wasted and not applied to their ultimate goals. The average collected amount per submission is so small that you cannot pay the bureaucrats managing the complex submissions from these sums.

An entrepreneur told me that they were being audited regarding the battery regulation. Two consultants came over from EY and audited all submissions in depth. They found out that the business overpaid by €10. The entrepreneur of course didn't mind and told them to keep the €10. The EY consultants told the entrepreneur to resubmit, otherwise the business would be fined. I think the government spent more on that audit by EY than the business will ever bring in on fees.

On top of this we have the regulations like Intrastat and in the future near-real-time submission of digital invoices to our tax authorities.


> On the other hand I estimate that between 50% and 200% of the taxes collected like this are wasted and not applied to their ultimate goals. The average collected amount per submission is so small that you cannot pay the bureaucrats managing the complex submissions from these sums.

The point of this tax is not to raise revenue bit to disincentivize wate. Since the tax weight based, companies have an incentive to reduce packaging.


Customers already pay for the disposal. You could achieve the same thing by requiring sellers to list in checkout how much the disposal of their packaging will cost. That's a 10 minute change for Claude Code, but of course then it would not give more money and power to the useless bureaucrats.


First of all that's a weaker signal to the manufacturer/seller to make changes as it'd only indirectly affect them.

It's also moot - in most (all?) European cities you only pay for waste but not recycling (paper, plastic, glas, aluminum, etc.). Those are free and paid for by exactly the kind of recycling fees we're discussing here.


This is wrong on so many levels. First of all, IT admins receive weekly e-mails from AWS, Google and MS that a certain service will shut down and that it will affect operations. In 90% of the cases, these e-mails are sent to all customers instead of only the affected customers, so admins get alarm fatigue.

Secondly, instead of doing a rug pull, what's wrong with a 2-month brownout period? Only allow the admin to login, put OneDrive in read-only mode. That will alert all users clearly.

Thirdly instead of fending of the questions from slate.com with Orwellian corporate speak, they should acknowledge their mistake and spend a couple of weeks of the highly paid engineers to see whether they can recover the data. Why is it always so hard for MS to say "Sorry, we made a mistake".


> Why is it always so hard for MS to say "Sorry, we made a mistake".

Seems like no corpos ever want to say that. Surely it’s not because of some pride or PR? I’m guessing such an admission might also sometimes have legal risks?


Saying sorry means admitting guilt. If there’s guilt there some level of harm. If there’s harm you can be sued. Avoiding being sued due to causing harm usually means increased costs, because you now need to take care you didn’t take before. Higher cost means you’re less competitive. Being less competitive makes you less attractive investment target. So as always it’s all about money. Saying sorry == lower share price.


The worst thing you can be in America is wrong. Our entire public society is built around, not admitting guilt or wrongdoing in any way shape or form.


I understand why you would like to use an LLM for vision. I do it myself often enough. I don't understand however, why the pill detection and counting is included in this benchmark. That is a task which you would perform with OpenCV right?

In my personal mini benchmark minicpm-v-4.6 scores amazingly well. Its a 0.8B model which runs fine on many consumer hardware.


Generating datasets to train more efficient models is a common use case for VLMs, especially frontier ones. It makes it much cheaper to create that initial dataset and you can abuse the nondeterminism of LLMs to identify data for human review (if they don’t converge, escalate to a human).


Especially the pill counting example. The best model was shown at 81.1% accuracy, which is a terrible rate for pharmacy scenarios. It seems like implementors would be better off instructing the models to use deterministic tools (like OpenCV) until the models are at 99.99% accuracy (or whatever an acceptable error rate is for pharmacy techs).


I think that is because people perceive OpenCV as 'hard to use' and LLMs as easy to use.


OpenCV is no longer hard to use, it just takes longer. Still, a little more complicated than asking LLM to count.

To use an LLM, you just prompt it with an image + text saying "count the pills in this image".

To use OpenCV, ... you just prompt an LLM with an image + text saying "count the pills in this image, using OpenCV instead of eyeballing it".

(I like to throw in "produce intermediary artifacts so I can see the process" for more difficult tasks; this helps the model avoiding making hallucination-prone leaps and gives more opportunities to self-correct. At a cost of extra time and tokens, of course.)

Using OpenCV without an LLM? Nah, not touching that, I don't have free weekends to waste anymore.


I no longer use it but never felt it was particularly complicated, but since the days of resnet there are much faster ways to the goal.


On top of that, the user paid for those tokens, so if there is an owner, it should be the user, not the provider.


This is close to what I need for my company. Currently I'm test driving Open WebUI.

https://github.com/open-webui/open-webui


Stay away, I had bad experience myself, I'm sure there is something better out there. For me I ended up asking Claude to build a custom RAG pipeline, took an afternoon and is 10x faster.


Congratulations with your launch! I suspect the product market fit for tools like these will be huge. Especially for SMBs with just 10s of employees in the office, lacking the budget for SAP consultants. However I think the agent creating an app is adding unnecessary complexity. Users want answers or insight in data, why build an app for that if the agent can provide it directly?


You can reuse the app indefinitely, for cheap, and predictably. I would say these are huge advantages.


Thanks! Agree that for one-off questions, chat is the right interface. We’ve found apps win when the workflow is repeated across a team and they almost become “templates” that seed ideas for other coworkers.


* with no easy and free solution yet.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: