Hacker Newsnew | past | comments | ask | show | jobs | submit | maxdo's commentslogin

The hype is almost reverse now . With opus , sonnet and haiku 5.5 , why should I care about Chinese models that are so behind ? Also Jev seems like killed a big chunk of deepseek market too .

I personally used it for secondary research agents and classification . Now classification part is gone .


Yeah, OpenAI models now are really at the frontier of high performing, low-cost models: https://artificialanalysis.ai/#intelligence-comparison-tabs

I agree with the author on a lot points about the joys of using low cost models. My work would pay much more, but I prefer to use cheap models usually. Assuming cost correlates somewhat closely with energy used, I feel good about spending as little as possible, and the cheap models are so capable.


I feel the opposite. An agent sometimes work an entire day . The chance of doing that work twice defeats any motivation to use cheaper , less capable models . Also if you open benchmarks leaderboard for the first time in a while there’s almost 0 Chinese models .

I spent last 2 months in construction stores. A good chunk of people drop a size in claude/chatgpt, and put tiles in their existing picture etc.

You can feel differently but it saves you so much money, time, and bring you so much pleasure. Not the other way around.


I designed my basement renovation using a combination of image modelsand a 3D scanned Houzz Pro FloorPlan, and I gotta say it REALLY helped the contractor deliver on the vision that I had.

Decorating when you can easily swap out pieces and parts in images with realistic dimensions makes rapid iteration on physical things much faster and it's easy to figure out what "speaks" to you. Great use of these tools.


I’m genuinely curious how you use image models for this type of design. I need a way for my wife and I to visualize a remodel project and don’t know where to start or which tools are best for the job.

dump the plan, make sure it has height etc, after dump pictures if needed,

refs of colors. atropic has some nice recipes somewhere online, entire prompts you can feed into cc. i used their workflows feature, otherwise it's painfully slow.

do it in two steps, first generate just the named 3d model, walls, stairs etc inspect for correctness, after proceed to the fun part, materials, colors, etc.

I did it in extra one step, not sure if that was needed, i build a local harness first for extra small room 6 ft by 6ft, and after that expanded into real size. It installed lots of tools.


Just upload the pictures to a chatGPT chat and tell it what you want, My wife uses it for this all the time.

The difference in the replies to this comment are striking.

One is a relatively complicated workflow producing 3D assets, the other is "just give pictures to chatgpt and tell it what you want".


It’s not different . One case you chose tiles , you take a picture of your bathroom , and give it a rough sized of tile + room and it will Give you visual check . Second option if you want it also to calculate materials . This is a 3d path . First one can be done by just chat , another one is coding agent .

And neither shows the outcome as to illustrate/support their point, which is phenomenally ironic considering the context..

I tried a couple of the web-based design-a-floorplan sites and all of them were kinda like using crayons to draw a circuit diagram. Ended up doing a pile of sketches on paper, then when I had what I liked finalised it in Visio and sent it through to an architect who turned it into proper plans which, since they were architect-drawn plans I had to check for f_ckery and then the builder checked it for stuff I'd missed but ended up with a rebuild that was exactly what we wanted.

Well with one bit of f_ckery that both of us missed until after the concrete had been poured because it was too small to see easily on the plan, and a second also too-small bit that I caught once they'd put the framing up and asked them to take it out again.


My wife designed our pool and our patio this exact way. She spent hours on it and it worked perfectly.

I can look out my window and see the exact scene she made in chatGPT.

Really impressive how useful it has become on some very general things


just claude/chatgpt, no specialized app?

yes, it's amazingly/depressingly good at things that would take a $1000+ software, and years of education just a year ago. Now I'm doing this with leftover claude/codex weekly quotas( i do have both due to some testings at work).

I did almost the same but it was completely AI , and started from renovation. it build me material list looked for discounts, saved me ~$2k.

Tricky question how do you know that person did a real drawing, not generated it in codex/claude code?


It doesn't really matter if they're happy with the result does it? Many of us here generate stuff in Claude Code for our employers.

There's still expertise and experience used in prompting, reviewing, refining it - a professional using AI doesn't automatically make it a scam.


You have all the struggles for the price of Anthropic / cursor subscription. I use the first one I code large chunks some PR are 50k LOC and I have at least 2-3 like this a week . It’s a greenfield project .

I still have quotas left I use it for home things build 3d model of my renovation projects, alerts for shopping list etc . And yeah I use cutting edge of cutting edge of models that saves me time and money , only discount monitor saved me ~$2k on my renovation project


Why the hell would you put 50k lines into a single PR?

I mean, why even pretend you’re going to “review” something that large? Just build everything on main.


PR is just an entity to review /do some other LLM processes . Human part check other models review PR's, tests all kind of , security etc. Also human is to checks docs, specs in the pr, db migrations if any , some of the tests related to the PR. We stopped reading the code after opus 4.6. Sometimes for very core parts i skim through files just to make sure if the changes were correct.

I would imagine even LLMs would do a better job reviewing smaller PRs than very large ones.

But to reiterate parent's question: why not just do those things continually at that point? Or on a calendar-based basis?

Yeah exactly, if you have a fully automated SDLC, how do you expect the agent to code review 50k lines properly?

It will take shortcuts and now the entire premise is busted. You now need to build a code review process for large PRs.


This is a common problem. Fully automated SDLC needs to start before the CI/CD.

My suggestion is to have the proper chunking mechanisms and multiple specialised agents. The most important is harness engineering, what we do at dromeas.ai to verify the code that goes to prod is a)have the code mapped before hand for the right agentic context, b)chunks of the right size per model context window c)specialised agents d)deduplication and verification . All before assessing a PR, a commit, a release. Harness engineering is not easy.. Especially when supporting multi model


they hope they can annoy and email every person in the org asking for money , same way as they did it with docker.

looks like a weird abstractions. why do i need this if i have modal.com, e2b, cloudflare that use original docker + some toolings around + way to run + isolations on network level.


Not bad only two major releases behind top tier. Edit : checked its rather 3 generations behind . Oh well

like what , manually refactor code in hours, when agent can do it for you in minutes, or setting up debug/dev env when agent can also do that for you?

Manually refactor code in days if need be, when you have something specific in mind and intend a particular result.

'refactor' isn't just a blender. If I wanted agentic coding I would just do that. I'm quite happy for other people to so completely entangle themselves in agentic coding that they can't get out. Please proceed.

to add: I'm a paying JetBrains customer, but they could lose me if they vibecode themselves into a pile of slop. I'm not sure they're doing that, even if they're dabbling in it. I think there's still expertise there.


why this is a surprise? I have 0 reasons to buy IDE in 2026.

Their main moat was smart refactor, and a few other nice tools to debug. regardless of what you think about coding is done, agents and your harness is the best tool. I heard they doing business accelerator to find a new business ideas. But as a company, hardly imagine anyone buying expensive IDE in 2 years. Why?


> regardless of what you think about coding is done, agents and your harness is the best tool.

I guess you're all in on that Anthropic IPO with the "coding is solved" approach.


There will always be people that write code with pen and paper. Its going to be an art form though rather than something productive.

> why this is a surprise? I have 0 reasons to buy IDE in 2026.

If the first part follows the second, the surprise is that revenue is up while the prevailing opinion appears to be "no-one needs to buy an IDE".


in terms of building sites, gpt is ahead e.g. computer vision etc.

hm, any product they have desperately need more ui to be better. There is an entire ecosystem exists due to their slowness to apply that.

Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: