Hacker Newsnew | past | comments | ask | show | jobs | submit | phoyd's commentslogin

Torrents should really be the preferred method for distributing AI model weights. Why rely on a single point of failure like Hugging Face? BitTorrent was made for exactly this.


In my experience public torrents often die as they grow older. It doesn't help that BitTorrent V1 makes long term seeding annoying, and BitTorrent V2 is almost never used.


I never understood this, is there anything that makes it difficult for the original uploader, the one that supposedly offers the file directly, to offer a torrent instead for the same amount of time?

As far as perennity is concerned it seems strictly better.


Every change to the source is effectively a new torrent. This creates a ton of fragmentation as data is reorganized, remixed, reencoded, and so on.

You can see this with many Linux distros: there is no single Debian torrent that people seed for years because there's always a refreshed version.

Distros are a bad use case for P2P anyway since you depend on upstream as soon as you start upgrading and installing packages.


IPFS has a solution [0] to this problem

[0] https://specs.ipfs.tech/ipns/ipns-record/


Most of IPFS doesn't actually work very well, if you've ever tried to use it


> if you've ever tried to use it

Heh, you got me :) IPFS is one of those things that I love reading and about and thinking about using someday, but somehow never get around to it.


The problems begin with taking 5-10 minutes to locate a file on the network. That's right, when you ask for a file it takes 5-10 minutes. Also if the file isn't in the network at all then it never terminates.

Nobody noticed because everyone just used the central web gateway that cached every file anyone ever accessed.


> Distros are a bad use case for P2P anyway since you depend on upstream as soon as you start upgrading and installing packages.

This is true for any distribution method not just p2p. You can even download a nightly through torrents so what does it matter how the data is transferred if it’s always going to require `apt update`?


Yeah, I just use the "netinstaller" ISOs since it's much smaller and never needs to be updated. If I had a need for air-gapped/offline installs I'd either download a larger ISO or just manually install packages from .deb as needed.


You can trivially have storage deduplication for the files served via torrent, transparent to the protocol. The most trivial version of this that you can do today with pretty much any client is having a single directory containing files serving multiple overlapping torrents.


> files serving multiple overlapping torrents

This sounds wildly complex, especially from a discovery perspective.


I don't see why it would be. It's transparent to other clients just like it is to the protocol. It cannot be more complex than alternatives by construction.


Or use a filesystem with dedupe like ZFS.


but model releases are already non-changeable?


Yes, models are a good use for P2P especially if everyone agrees to share the same torrent and someone (or a cohort) commit to seeding for the long haul.


If they're no longer using that model they may not be willing to continue using their storage for it.


How does this address the previous point? If they are not providing storage, then centralized or decentralized doesn’t make a difference.

Torrent/P2P can only add redundancy, so it’s impossible to have worse availability than a download link?


[flagged]


> Is there a fully in-browser torrent option that has the same UX as a regular file download in Firefox?

Opera did back in the day.


Well, for a torrent to stay healthy, users have to seed after download, so it can't possibly have the same UX as a standard browser file download unless you want to either kill the ecosystem or hide from the user what is consuming upload bandwidth.

That said, for large files, I much prefer the UX of a well-designed torrent client like Transmission to my web browser. If nothing else, the downloads are reliably resumable.


> I much prefer the UX of a well-designed torrent client like Transmission to my web browser. If nothing else, the downloads are reliably resumable.

Brave browser had BitTorrent client built in for a while. I tried it a couple of times as I already use Brave for web browsing on my laptop. It was a very confusing BitTorrent client. I struggled to use it, and wasted time waiting for a download to complete only to not be able to find where the files were and then they disappeared. Using a decent BitTorrent client like you say is much preferable to the one that they had in Brave browser.


yes, but a few dedicated hoarders could keep many models alive, and I'm fairly certain the local LLM community has plenty of people who would.


They don't. That's the conversation topic.

Nor do they need to. 99% of everything is crap, and not worth prescribing except for a random sample so future historians can study our crap.


You can add an existing HTTP download as a "web seed" to a torrent, so they don't actually need to do anything for people to share it as a torrent.


The biggest problem with BTv1 was the lack of per-file checksumming, and swarm merging (i.e. individual files have shared seeding pools across torrents). BTv2 specs the latter, but I think only BiglyBT actually implements it. Having both of those features from the get-go would've gone a LONG way to fixing the dead torrent problem.


A torrent with a webseed is strictly more resilient than a direct download link alone.


You only need one person/organization to commit to seeding. The majority of people do not want to seed at all without some sort of incentive.

If this site represents a coordinated datahoarding effort then there will be at least a few people who will seed indefinitely.


The last guy (kimdotcom) who was working on this (incentive for seeding) seems to be heading to the US: https://www.rnz.co.nz/news/science-and-technology/651123/cou...

It’s interesting he’s no longer getting any media attention any more.


Yeah but that's what provider like huggingface could just do, keep seeding the models so they are still accessible


I always wondered why v2 is never used... you can even search for files by their individual hash with it.


Because history is path-dependent, as engineers keep learning over and over again. It doesn't matter whether Plan9 is theoretically superior to Linux - we're all on Linux and nobody's porting all the apps over.


HF is meant to be a single point of control. AI models and Linux distros aren't usually for normies, so distribution via torrents would make sense, especially to save the provider some bandwidth. Ubuntu has been offering torrent downloads for ages. No mention of torrents on HF. I believe most downloads will soon be account/EULA-walled.


Yeah, I thought they were used for this already. Surprised this is news, but also relieved.


IIRC Mistral used to do it, not sure if that's still the case

EDIT/ Yes they did, that no longer seems to be the case though

https://x.com/MistralAI/status/1833758285167722836


ages ago I tried using IPFS to more or less accomplish this, I imagined it to act more like a weights/training data network fs that everyone would be able to participate in.


Yazi, If, and Ranger file managers use it to display images inline. Euporie allows you to run a graphical Jupyter console session.

Sixel is a very old standard and has been used to display graphics in terminals since the 80s. It fell out of favor when high-resolution bitmapped desktop environments became mainstream (and affordable), but it is still useful if you are doing remote console work and don't want to set up network hoops with ad hoc HTTP servers just to view a remote diagram.


That's literally what this post is about.


Sorry, it was meant to be a reply to a comment.


I am more shocked by the "overnight" aspect. I tried running clang-format on the Chromium source (68,281 .cc files, 21 million lines according to wc):

$ find chromium-149.0.7826.1/ -name ".cc" -exec cat {} + | wc 21640925 55715244 833460441

And that took less than 6 minutes on a single E5-2696 v3 from 2014:

$ time find chromium-149.0.7826.1/ -name *.cc | parallel -j 16 clang-format $x>/dev/null

real 0m5.666s user 1m13.964s sys 0m13.373s

That’s orders of magnitude faster, especially if we assume they’re not running their workloads on potatoes like mine. Is Ruby’s syntax really that much more complicated than C++, or is this a tooling problem?


I don't think the post necessarily means it took multiple hours to format the codebase, I think they're probably just saying they worked on it off-hours and landed it while no one was working so that it didn't run into merge conflicts.


My guess would be tooling. I think the Ruby formatters are written in Ruby. I’d guess the clang one is written in C.


Nah the article says it's rust and calling into a C library for parsing.


I want to drop here the not very well known fact, that the SQL Standard grammar distinguishes between "SQL language identifier" and "regular identifier". According to the rules, a SQL language identifier can not end with an underscore (copied from ISO/IEC 9075-2:1999 "5.4 Names and identifiers":

<SQL language identifier> ::=<SQL language identifier start> [ { <underscore> | <SQL language identifier part> }... ]

<SQL language identifier start> ::= <simple Latin letter> <SQL language identifier part> ::= <simple Latin letter> | <digit>

So, using names with trailing underscore should always be safe.


It_ might_ be_ safe_, but_ it_ is_ definitely_ annoying_.


.NET Core ran on Linux from version 1.0 on and it was released 2016.


I know but it's not really a great way of running and developing apps, that's what I mean. The best tool for it is visual studio (full version) which doesn't run on Linux. And you have to run the bytecode interpreter. Whether it's mono or the official .net doesn't really matter.

I'd rather use something open like python, go etc. And most people do, there's few Linux apps that use .net. Microsoft didn't even use it for visual studio code.


Your arguments are all over the place. The Mono project is practically dead. Official .NET runs on Linux without issues. It uses the modern CoreCLR runtime on Linux. Some other targets, like WASM, still use a runtime based on Mono, but CoreCLR has been used on Linux for nearly a decade. Best IDE for it is JetBrains Rider (by far) which runs on Linux. It uses JIT compiler, not interpreter. For CLI tools, AOT compilation is also supported, which compiles to machine code like Go/C++/Rust do. Also, .NET is open. The only issue is that primary development is done by a single corporation.


> And you have to run the bytecode interpreter. (...) I'd rather use something open like python

Did you know that .net is open and python runs a bytecode interpreter?


Yeah but you are still bound to the cursed company that MS is. Python is developed by passionate, talented devs, .net is developed by "AI" nowadays.

Don't believe me ? Check how many of the .net docs pages have the relevant banner.


Which banner? Pls do show.


> The best tool for it is visual studio (full version) which doesn't run on Linux.

JetBrains Rider is amazing, fyi!


FWIW the dotnet shop I work for has moved entirely to running on Linux Kubernetes in production and issues M series macbooks (with Jetbrains Rider licenses) to developers. The last Windows Server VM was decommissioned in 2019.


Only cellular iPads have GPS IIRC.


Now that you say it, I think our iPad doesn't have GPS indeed. The only app that required it and that we were using was the star map. Since we used the iPad at home it worked for us to just enter our position once manually.


No, all of them do


Giving employees stock options does not solve the "employes should own the company" problem, when there is a market to sell the stock. Shareholders and employees have different goals and interests regarding the company. For example, the employees they may have to decide on the elimination of their own jobs in order to secure the value of their stock holdings (which btw. would make the company not owned by their employees anymore)

A more interesting approach would be a model where being a employee automatically gives you a vote on company wide decisions including. the distribution of profits, just how being a cititzen of a democratic country gives you a vote on the fate of your country (similar to the Mondragon Corporation in Spain)


so in the extreme case, the employees can decide to return 99% profits to themselves and return 1% to shareholders? It sounds cute but will be hard to make it work when incentives are not aligned well. Just like Communism sounds good in theory and everyone with a kind heart will like more equality but in practice, it destroys the entire cake.



I'm also interested. Setting up a passwordless SSH account for some public service sounds like a good way to give your machine away to North Korean hackers, because you forgot to set someting in /etc/sshd to "no".

Is there a usable description somewhere on how to do this safely?


i'd be interested in seeing that. here its ok because it doesnt use sshd at all


So it's using a new stack that hasn't been vetted like OpenSSH? I'd rather use OpenSSH + LibreSSL for this application.



Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: