> Work for business people who want fast results. Agentic coding gets you to something presentable much faster at the cost of code quality. I have never seen a customer or business person care about that.
That's always false. It's like people want their meals delivered fast. They say they don't care about taste or how it is done. Watch when they get sick or don't like it and the drama that happens.
People don't care until they do. They don't know what to care (in this case code quality) or say that because you're not explaining it. People also "gamble" and you take the blame. Long term impacts? Nah doesn't matter. Weeks later and things break -- what did you do?
In my country we have very protective labor laws, but what you describe is one of the few reasons that allow for instant termination: You refuse a direct order. If your boss says that he does not care about long term consequences and he want things done a certain way, you do it that way. Everything else is unprofessional. Your opinion on what you think he wants is legally irrelevant.
I never said. It's not about outright refuse. It's how to deliver better outcomes. It's not a win or lose situation. Don't treat it like that.
> If your boss says that he does not care about long term consequences and he want things done a certain way
As suggested in my previous comment it could be you didn't explain it well. If you say code quality they might not care. Say it will crash randomly and can't fix it then they might. There's some art to it.
> Your opinion on what you think he wants is legally irrelevant.
No it's not. So you're saying if you got told to do something illegal you'd also do it anyway? I doubt it. You need to at least cover your back and document it as such. I'd say termination is better than jail.
> A high score on benchmarks is not as useful because a model overtrained to always answer will give confidently wrong responses.
It's not useful because the benchmarks often measure the wrong thing. They're here yapping about AGI and yet the benchmarks treat it like a trained dog. Fetch this. 100 points.
Each "problem" in these benchmarks likely has more than 1 solution that can be considered correct and even should be graded in many ways. Yet we see in many benchmarks higher effort (or thinking levels) don't help because the benchmark penalizes for doing "more" than what the answers asks for. So what did you ask for?
In human school you often get marks on the process and not just the end result. Thinking tokens have been cut. All we group on is things like cost, turns and time but not the what else.
Will look forward to the "feel" of the model in real testing. But I agree that these benchmarks do get "dealt with" rapidly. That's a shame, but I guess it's the times we live in.
Hype that burned out pretty quickly, it's hard to speak to the size and significance of old hype, I never felt it.
Every time I personally tried Gemini models up until last week they simply couldn't do the long complex tasks I'd being doing with Anthropic models for many months.
reply