Hacker Newsnew | past | comments | ask | show | jobs | submit | jackcviers3's commentslogin

On the other hand, the open weights models could crawl and annotate and rl the training data that Anthropic and OAI did in exactly the same way, and take the exact same legal hits. They use distillation because it's cheaper not to do so.


> On the other hand,

What is the other hand here? They could do slow/costly illegal think #1 instead of the fast/cheaper illegal thing #2 that they currently do?


Neither one is illegal though. Scraping the internet is as legal as surfing the internet. And what Anthropic calls "distillation attacks" is really just paying for and using the service Anthropic provides. I would think a judge would have the same view of distillation as they do of scraping, if Anthropic doesn't want a subscriber to have access to the model, they are within their rights to block access, but it's not the user's job to refrain from using the service. If their usage is so different from everyone else's, they should be easy to detect and block. If their usage is so similar to everyone else's that it's difficult to detect, then it shouldn't be of any concern.


Genuine question: is distillation illegal, or just against Anthropic's terms of use?


It seems to be only against ToS; the problem, however, is that “distillation” involves intent. Anthropic seems to be calling for ban on distillation because they cannot reliably identify who does it - if they could, they would’ve just banned all infringing users.

However, if there is no systematic difference between regular use and “distillation use” then they need to establish intent, which is laughably hard if you cannot even distinguish use cases.

Hence the call to make it illegal too (and just against ToS). Because if it is illegal, then they can rely on US to enforce the law. And the US, for a lack of technical solution, can use the good old “China bad” shortcut.


Hmm, why do you think this is true? One reason I'm skeptical of this is because RL envs are often purchased (and are not publicly available), and this might be a sizable component of why models are getting better.


A stopped clock is, after all, correct twice a day.


You need 0.5 to 1.5 acres per person for non-mechanized industrial argriculture. Nowhere with land that is truly arable enough for that is going to _sell_ you 1 acre at a time. In the U.S., you buy at least 40 acres at a time. In the U.S. Midwest, that's going to set you back (on average) $379,000. That's before you buy the equipment you need to be able to farm the land in the first place. Unless you industrialize and grow crops to sell to other people, you will not be able to afford the property taxes on the land to be able to keep it, either.

So, no, you cannot just go out and buy an acre and garden.


What? Yes, you absolutely can buy half an acre. You think in 1880 people went on Zillow to buy land?

You're just going to have to do this the 1880 way.

Knock on the door of a land owner and offer to buy/rent half an acre so you can farm. You'll find takers. My extended family literally has this arrangement with many people who farm different crops during different seasons.

Too hard to to do it that way? Welcome to 1880! Most people weren't land owners on the land they farmed back then and didn't have 'Perfectly arable' plots, and this was pre-fertilizer.

Oh and you'll have to use horse and buggy to get around to find land owners (no evil automobiles from those evil factories full of automation!) who will allow you to farm their land, just like 1880. So good luck.

I don't know how many times I need to explain this to tech doomers: nobody forced people out of subsistence farming. They chose to leave it. It was not a utopia.


For most teams, whether or not you can say no to building something is ambiguous at best, at least if you wish to stay on that team and at that company. It's definitely one of the things that has made me vote with my feet in the past. With agentic coding, the ability to say no is pretty much gone because the perception is that it's just one more parallel thing we can throw an agent at.

The thing I see from agentic adoption that I find lamentable as a software engineer is that timeline expectations have collapsed to absurdity. You can plan a project to do a major migration, do all the estimations on how long something will take, and if you give an answer that says weeks and cite the evidence, product and leadership will now claim it should take days, citing their ai's design.

It's exhausting. Even if you are an expert, you now have lost the implicit trust that came from years of building political capital, shipping efficiently, and delivering value for multiple companies, because a different prompt with different context from the one you provide gave a different answer than what you did.

During delivery, if you read your code produced line-by-line and review for correctness, and put in additional guardrail automations that slow the automated build, and ship 4 times a day with a defect rate of 5.4% with agentic coding, you are compared unfavorably to teams with a change defect rate of 15.7% that ship 13 times per day, because you are too slow.

And you are individually compared with whole team outputs. Even if you deliver at a rate ten times greater than the worst contributor at your company, if you are not outputting code at the rate of an entire team of 5, you are not meeting the expectations of product and leadership anymore.

All of this is to say, yes, people are looking at software engineers as both the bottleneck and unnecessary, even at high technology companies, right now. They are looking at them that way because they have their own agents that are biased to think that the engineering claims are wrong and agents are sycophantic.


That's right, and even more alarming: no one is pushing back, not only the engineers. The leadership team has AI estimates. The product team has AI estimates. Even the engineering team has AI estimates. Everybody says "Ship." So we ship.

What the company lacked was never engineering deliverables; what it really lacked was a prioritization owner who could draw the line. Bad code certainly wasn't the cause of this problem.


It has been a very bad bet that hardware will not evolve to exceed the performance requirements of today's software tomorrow, just as it is a bad bet that tomorrow someone will rewrite today's software to be slower.


Eh, but then as hardware evolves, the software will also follow suit. We’ve had an explosion of compute performance and yet software is crawling for the same tasks we did a decade ago.

Better hardware ensures that software that is “finished” today will run at acceptable levels of performance in the future, and nothing more.

I think we won’t see software performance improve until real constraints are put on the teams writing it and leaders who prioritize performance as a North Star for their product roadmap. Good luck selling that to VCs though.


It seems like a more polite way of handling this in private spaces is just to ask that people take them off - just like we do when a pig farmer walks into our house with their boots on.

I get why people are creeped out by them, but we get filmed or photographed hundreds of times a day in a big city when we are in public spaces. Gatekeeping a potentially useful technology for being filmed in public -- well, everyone is _already_ filmed in public. ATM cameras, stoplight cameras, drone cameras, smartphone cameras, security cameras, doorbell cameras. You are on camera every time you step out of your house. You are on camera every time you open your work computer. Singling out cameras in eyeglasses as "creepy" is kind of worrying about a drop in the ocean. Cameras on self-driving cars. Nanny cams. Closed-circuit cameras. The things are everywhere, and they are always invasions of privacy. Why is the line the "creeper" glasses?

I'd be ok with it if we were for banning all non-consensual recordings in all spaces. But we're very much not.

And if we're not, then having a personal heads-up display that is contextual to your current surroundings or has augmented reality capability is too useful to not use (eventually). I'm bad with names, and good with faces. That use-case alone would be worth it for me, if it were available.


> well, everyone is _already_ filmed in public. ATM cameras, stoplight cameras, drone cameras, smartphone cameras, security cameras, doorbell cameras.

And we probably ought to regulate how all such footage is handled.

> banning all non-consensual recordings in all spaces

It's a false dichotomy. Even if recording is permitted that doesn't mean the systemic invasion of personal privacy needs to be.


Great, let's regulate it! And why are glasses more offensive than cell phone cameras, or go pros, or drones? I genuinely do not understand why people don't worry about the other form factors, but draw the line at the glasses, so help me here. To be clear - I understand why people find being recorded creepy. I don't understand why the glasses form factor is creepy but random cell phone recordings that are shared on the internet all the time without the consent of the recorded people aren't.


tl;dr It's the difference between possessing a camera and actively pointing it at someone.

Think about the practical aspect of it. I have to point my phone at you to record you. It's really quite conspicuous. It's also mildly inconvenient for me so I won't be doing it the vast majority of the time.

Whereas the glasses point wherever you're looking, are expected to be recording constantly, and are expected to do things with the data involving third parties. It's the same as a VR headset except in that case the expectation is that the footage is neither sent anywhere nor even retained, merely presented live to the user as if he were looking at you (and his face is already point in your direction).


I disagree - it’s extremely easy to film on your phone covertly. If someone wants to film you without you noticing they will be able to do so.


"It seems like a more polite way of handling this in private spaces is just to ask that people take them off - just like we do when a pig farmer walks into our house with their boots on."

Just FYI, they do heavily market this towards RX glasses wearers. So, you wouldn't quite be able to just as simply ask someone to take off their glasses and no longer be able to see.


I'm going to guess that someone who can afford smart glasses can afford to have another pair of unsmart glasses. What is it about the _glasses_ that people find creepier than a smartphone that can literally do even more invasive things than the current glasses technology?


It's very obvious when someone is recording you with a smartphone


I mean, I grew up with AOL AIM, Yahoo Messenger, and IRC... yet I switched every time a new tech came out with more of my friends on it. Why do we think discord will be any more sticky than Digg or Slashdot, or any of the above?

People will migrate, some will stay, and it will just be yet another noise machine they have to check in the list of snapchat, instagram, tiktok, reddit, twitter, twitch, discord, group texts, marco polo, tinder, hinge, roblox, minecraft servers, email, whatsapp and telegram, and slack/teams for work.

Absolutely exhausting to be honest.


Kids today are alarmingly bad at technology. This is not a "kids these days" situation, this is absolutely true. They understand "tap on icon, open app, there's a feed and DMs".

I mean it, the tech illiteracy of gen Z/alpha is out of this world, I did not expect a generation that grew up with technology to be so inept, but here we are. But they grew up with a 4x4 grid of app icons, not with a PC.


I don’t think people understand the true level of tech illiteracy of Gen Z. A couple years back I did an internship with the IT guy at my high school, and the vast majority of the problems students had with the Chromebooks we used were, in no specific order:

  - Not understanding that a dead battery means it won’t turn on
  - Trying to use them without an internet connection
  - “The screen won’t work” when trying to non-touchscreen models like a tablet
  - “I can’t see my stuff” when using the guest mode rather than their login, or when they used a PC and they couldn’t see the docs icon on their desktop
That’s not even to mention the abysmal typing skills of most students, so many 15WPM hunt-and-peck typers..

There’s a mountain of issues along those lines we ran into, and it was honestly frightening to watch.


I feel like asking someone working IT about the average technical literacy of the people they work with is similar to asking an EMT about the health of an average person. Not to discredit your experience, but you should account for the fact that a lot of the people you helped were the ones who were already filtered out by their inability to fix trivial problems.

I'm not saying this issue doesn't exist. But I want to reframe it as the low bar for using tech dropping through the floor. Previously, you had to have at least somewhat of an idea for what you're doing, but nowadays most people who don't care about tech are reliant on using the "grandma school of thought" in memorizing basic patterns and relationships without having a bigger model of what's going on. This mostly affects newer generations and older people who only started using technology recently, because this strategy didn't fly in the past. But technical literacy is falling for everyone.

But the absence of the low bar doesn't mean that everyone's chasing it. In high school, I was surrounded by peers who were interested in tech, sometimes being far better than me. The average level of understanding was pretty alright. In university, lots of people did just fine. I know countless people my age who are highly skilled in computer science. We're not in the majority, but there's plenty of us. I'm tired of it always being framed as an issue stemming from some kind of unique lack of personal responsibility and low intelligence related to age, used to apply stereotypes to hundreds of millions of people. Every average user will optimize actually understanding anything out of their brain if given an opportunity, it's just that that opportunity had only appeared fairly recently.


Yeah, I work with kids and it's admittedly a bit disheartening having conversations like

> why don't you make a separate account for your sibling

> I don't know how to make an email

> but you needed an email for your account

> yeah, I just use my school email

By that time my age as a young teen I knew how to make new accounts and research what I didn't know. And I'm not sure of its my place to help them create an email without knowledge from their parents.


Correct. From my personal experience (have kids and nieces/nephew this age), and all think an app is the thing that they scroll in, and any attempt to explain the very basics on internet connectivity, servers, databases, etc, ends up in them basically experiencing blue screen moment and backing away to the safety of the endless scroll.

The most complex concept they can understand is mail/post attachment or capcut, but then this is it. 10 minutes later they will download phone flashlight app that requires Google services for app delivery.

Shocking.

I ended up with refusing to help with anything related to technology in any other way than pointing to help/manual/search engines and asking questions.


Isn't that true of python as well? I would argue that Github's decision to use markdown for formatting, more than any other, is what resulted in its widespread adoption to other use cases. The simple tool to share code ate the world.

I'm continually surprised that Microsoft hasn't completely cornered the market on LLM code generation, given their head start with copilot and ready access to source code on a scale that nobody else really has.


Python has a spec and multiple healthy implementations, and is overwhelmingly more popular than org mode, so I don't really think that's a rebuttal.


The last one is fairly simple to solve. Set up a microphone in any busy location where conversations are occurring. In an agentic loop, send random snippets of audio recordings for transcriptions to be converted to text. Randomly send that to an llm, appending to a conversational context. Then, also hook up a chat interface to discuss topics with the output from the llm. The random background noise and the context output in response serves as a confounding internal dialog to the conversation it is having with the user via the chat interface. It will affect the outputs in response to the user.

If it interrupts the user chain of thought with random questions about what it is hearing in the background, etc. If given tools for web search or generating an image, it might do unprompted things. Of course, this is a trick, but you could argue that any sensory input living sentient beings are also the same sort of trick, I think.

I think the conversation will derail pretty quickly, but it would be interesting to see how uncontrolled input had an impact on the chat.


I'll add to this - if you work on a software project to port an excel spreadsheet to real software that has all those properties, if the spreadsheet is sophisticated enough to warrant the process, the creators won't be able to remember enough details abut how they created it to tell you the requirements necessary to produce the software. You may do all the calculations right, and because they've always had a rounding error that they've worked around somewhere else, your software shows calculations that have driven business decisions for decades were always wrong, and the business will insist that the new software is wrong instead of owning some mistake. It's never pretty, and it always governs something extremely important.


Now, if we could give that excel file to an llm and it creates a design document that explains everything is does, then that would be a great use of an LLM.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: