Hacker Newsnew | past | comments | ask | show | jobs | submit | ipsod's commentslogin

That's what I do. Both for this reason, and because I don't want to log into my personal gmail on work devices.

I don't necessarily feel safe, though, and have moved my important stuff off of gmail.


Because it says it more often than kids say "six seven".

My kids stopped saying “six seven” about a year ago. I still have no clue what that was about.

I think it might be like, a praise of inanity.

Gemini Flash is the one I get most excited about, because it's so fast and so good at real-world knowledge, and it's improving so fast - look at how much the benchmarks improved in ~1 month. It's just categorically different than anything else.

Also, I use it every day, and it just got ~10% better at coding, according to the benchmarks. How is that not exciting?


I use Gemini every day and I've noticed any subjective improvement. In many cases it feels worse because it does fewer Google searches than before. As a result I find it hard to get excited about it.

I do think Gemini is underrated on HN though!


I haven't had any issues lately.

3.5 pro doesn't exist yet?

Sorry my bad, I messed up the numbers, 3 and 3.1 pro. All of these model numbers have me confused

IDK if it's smaller, but I know it's way faster. In one test I did, Flash 3.7 high was ~9.4x faster than Luna High.

But, also... Sol crushes Flash 3.7 at writing code in a codebase of any size beyond "tiny".

Flash is my go-to for prototyping, and basically anything that isn't writing production code.


The only company with a proper TPU set-up is bound to have the fast models, now add a market cap like Google to the mix.

They've been my bet to win the AI race for a while. I was starting to doubt, but this 3.6, 3.7, and 3.8 arc has anchored me.

Totally agree. When this current wave of GenAI really started heating up, I guess 2020-2021, my analysis was very straightforward. What are the high level inputs to long-term success? I basically came up with a couple of criteria:

1. Data. Lots of data.

2. Money. Lots of money.

3. Access to necessary hardware.

4. Business alignment/will to do it.

5. Access to talent, current and future.

This is certainly incomplete/naive. In my mind, though, Google was the clear answer.

On a more personal level, I've been deep into the Google ecosystem since I got [email protected] in 2005. (I actually paid 50 cents on ebay to get a very early invite.) There was no question in my mind that Google's AI work would deeply integrate into their whole ecosystem in very powerful and productive ways. (Yes, I can join you to discuss, at length, the various ways that Google's dominance is problematic/scary.)

Having said all that, I'm quite happy that there is, at the moment, a very rich competitive landscape. Indeed, not too long ago, with Gemini Pro 3.1 languishing, I moved most of my deeper thinking work to ChatGPT, which was, for me at least, clearly outperforming Gemini.

While I certainly didn't anticipate it, Google's strategy of making their fast/relatively inexpensive models surprisingly powerful has been a welcomed surprise.


I think you forgot to mention energy efficiency. If you build AI hardware in-house, your only other expense is energy and the producer surplus is greatest for companies producing below the equilibrium market price.

It's the difference between billions in revenue and billions in profits.


Not only do they have access to the hardware - they've been developing it in-house for years.

Luna is way slow. I don't remember an OpenAI model ever being this slow.

edit: I have a subscription; direct call.


Are you using direct or via OpenRouter? I think OpenRouter Luna always uses the `flex` tier, which is quite a bit slower.

Way slow? What are you comparing with?

Its not good at not making mistakes, but what it produces is structurally quite nice, not over-engineered (looking at you Sol) and its personality isn’t annoying (looking at you Claude). A bit like Grok Code, but Grok is a better coder.

Same. Love oneshotting or sanity checks. Which fortunately is a lot of my workflow (lot of long tail stuff fits in one prompt).

> This is of course important for protecting you.

Well, that's very kind of them. I am constantly impressed at the kindness of our governments, and the recent growth of that kindness. I guess that, with all of the power that modern technology is giving them, they're finally getting to live out their heart's desires of being very, very kind.


Do you find the variety helps? I've migrated away from such complexity, and I simply have multiple agents of the same model run the same prompt (usually Sol 5.6 high or max), and generally this gives plenty of adversarial input. I'd be curious to know how much difference it makes to run multiple models.

I frequently find blind spots / edges where one model notices something non-trivial none of the others did. I think the one that surprises me the most often is probably grok, but I wouldn't want grok to be my daily driver. I feel I get benefits but I could also see the argument that it's just a complex token burning furnace lol.

You generally don't even have to convince it, or at least I don't. I just paste the error into the prompt window, and say "you got blocked, try again", and it'll just say, "Oh, that's because ...", then do it.

Web apps are where I have this trouble.

Making a web app secure is literally just finding and patching vulnerabilities, instead of finding and exploiting them. You could have the AI "try to make this app secure", find what it patches, and use it for exploits, and the AI can't know if that's what you're trying to do or not. I don't know how you can get around this. I get around it by not using Anthropic products, at present.


Not to endorse OpenAI's particular guardrails, but unless you're doing something groundbreaking, security best practices should be more than enough for web development.

OpenAI is what I use most. Sol 5.6 still rejects a few requests a day when I'm working on web apps, but, overall, it's not too bad. I wish it'd auto-resume and try again, instead of waiting for me to intervene, but it's rare enough that it's not a huge deal.

It probably doesn't help that I'm using frameworkless PHP - I imagine a lot triggers could be avoided if I was using a framework where secure features were baked in.


With OpenAI you can also apply for the security program, which doesn't require you to be a certified pentester (as per Anthropic).

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: