Hacker Newsnew | past | comments | ask | show | jobs | submit | mikert89's commentslogin

theres also a decent chance that AWS itself is vulnerable to AI agents that dont need complicated cloud platforms the way humans do

some ivy league grad with no real world dev experience waved this on

Once Chinese ai can run on 50k priced GPUs and match current models, people will have these running from bunkers. There’s no stopping it

Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated

Disagree.

Examples:

- predict a coinflip: easy to verify, hard to learn

- earn $100: easy to verify, hard to learn

- increase paid subscriptions in an A/B test: easy to verify, hard to learn

I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.


Doesn't "saturated" mean that essentially there won't be any more progress in the benchmarch? Also of note is that two of your points only mean something on an occidental capitalist system.

these just need more compute:

- earn $100: easy to verify, hard to learn

- increase paid subscriptions in an A/B test: easy to verify, hard to learn

but we both know these examples go against the spirit of my point


Perhaps, but I think a bigger problem than lack of compute is the cost of rewards. Games like Chess and Go were solved long before self-driving, partly because it's incredibly cheap to acquire the reward of a bad board game decision, relatively to how expensive it is to acquire the cost of a bad driving decision. With driving, acquiring the reward can cost you $20/hr for human supervisors to generate disengagements, or $100k if you crash, or $30B if you crash the car into a person in a way that causes your company to collapse (e.g., Cruise).

yeah but I think you may be underestimating the amount of capital available for compute. if AGI is possible through some 5 trillion of expenditure on computers, there will be money for it.

also, you are underestimating how short a 10 year time frame is. we are close to self driving, the first neural net image model was in 2013. 13 years is a blink of an eye


You bring up an interesting point. Isn't the reward itself subjective in many domains?

You mean any repeatable benchmark will be saturated.

The problem is that there is a huge perverse incentive. The intelligence is in the training layer not in the model parameters, but the intelligence is really good at remembering things, so if you let it take the test, it can RL it.


this is a short term problem, over 20 years benchmark gaming will be a blip

Then I propose the tomjen-1 benchmark: prove the N vs NP problem formally undecidable.

Yes, but not necessarily under tight budget constraints.

theres no budget constraints for AGI

Yup. Only subjective taste remains.

Nope that will be commodified in short order.

People have, they arent putting them up for sale. Its the same with AI sales systems, if they work, they are worth far more than what they can be sold for as a product.

We are also sort of at a "bespoke factory" stage where building a factory means doing it around your specific codebase, each of which has their own needs and quirks. Just taking one of these wholesale from one company and using it at another would not work.

We'll see if in the future, as people begin software projects this way, if there is more standardization. I suspect that there's just too much going on too fast at the moment to do it any other way, a decent factory for Opus 4.8 looks very different than one good for Astra, I'd guess. And models are just one axis.


Thought I read elsewhere that if these get dirty they are essentially blowing dirty water into your gums, which can cause health problems

tools like this almost need the same level of sanitization you would see at the dentist.

other than that, I am very interested


Qualified product should use antimicrobial plastic. You still need to disinfect it using mild disinfectant like iodophor.

Don’t get it dirty

You're most likely using it in the same room where you poop, probably storing it next to a basin where you wash your dirty hands, and putting it in one of the dirtiest places in/on your body daily. It gets soaking wet hopefully twice daily.

I hate to break it to you, but it will get dirty.


AWS is good at operations, i.e. running something like S3 at massive scale. Or SQS, ddb, etc. High surface area for distributed systems, but low surface area for user experience.

Once you add in product decision making, like how to make the dev experience good on something with alot of user flows (like cognito), the products are shit.


These are temporary visas


I had a similar experience, but I think a lot of other modern philosophy is really pedantic and builds upon a thousand years of writings. So its very hard to follow without proper academic study or understanding what its in response to.

Nietzsche has paragraphs that could be studied and independently thought about without even reading the rest of the book/chapter. You can crack open beyond good and evil and just start reading a random page.

I still would argue his thought is far more groundbreaking than any contemporaries


> paragraphs that could be studied and independently

Yes, but after reading BGE, I felt like I had to read all his other books to really understand what he was saying (and others too, Kant, Spinoza, etc). And then re-read BGE.

Also Walter Kauffman's on him is very good. Working through that now.


yeah a lot of uneducated analysis was off, but it was still very thought provoking and jarring to my earlier world view.

it depends on which chapters/sections though, some are more accessible than others.

one book I enjoyed is:

"The Mask of Enlightenment" by stanley rosen. I read that along Zarathustra


you can run open source models in the privacy of your own home :)


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: