It was still crap back when it was Wikia. The website itself wasn't so bad yet, but the actual content of most Wikia wikis was very poor quality (often wrong, outdated, and/or poorly sourced; lack of maintenance and amateurish writing/presentation). I quickly learned to not click on Wikia search results whenever I could help it, whereas an independent wiki was a sign that the community cared enough to run their own servers.
The only Wikia sites that had anything worthwhile were the ones that started as independent communities and got bought out, like Memory Alpha or the Minecraft wiki.
OK, but that doesn't change anything. You have a large pool of IPs, each of which only needs to expend a handful of extra milliseconds of work to get unlimited access to the protected resource.
Even if you had to solve a L6 challenge for every request it's faster than the total RTT time of most servers. In other words not a meaningful barrier. And L6 is already a level which severely interferes with human usage of a site.
a $5000 ASIC-based device can hash SHA256 at 200TH/s or more.
And yet many report it works, at least for now, and the excess load due to scraping activity falls precipitously when Anubis or similar solutions are used. Maybe once this sort of challenge is used almost everywhere we'll see concerted attempts to get around it, but for now it is easier for them to just move on to another target.
> a $5000 ASIC-based device can hash SHA256 at 200TH/s or more
Peanuts for the big players, but many (almost all?) running smaller scale scraping operations are going to find $5000 rather prohibitive, and they are unlikely to be able to integrate it as they are probably running a “stock” scraper that they didn't write themselves.
You don’t need to spend $5000 to obtain the hash rate of a $5000 device on a rental basis. You may have heard of this thing called “the cloud”. Obtaining very high hash rates is effectively free, largely as a side effect of the crypto bust.
Not sure why anyone would characterize these scrapers im general as all being fly-by-night operations that don’t have two cents to scrape together.
A significant amount of the effect is everything outside the proof of work. Not the cost of hashing but the need to run the javascript that submits it.
If you've seen anyone post a comparison of crawl rate versus difficulty, I'd love to see it. There's probably some difference but I want to know how much of the overall effect it is.
I never said the approach was so stupid it could never work. I just shared how it’s annoying and how I worked around my annoyance.
An even stupider approach would work exactly as well or better, like a form saying “type the letter y in this box to continue”. The only benefit is as a road bump that makes the site in any way custom. The moment anyone you are defending against so much as looks at the mechanics of solving the challenge it completely falls apart.
Some of the asymmetry might be regained if anubis had thousands of variations of PoW algorithms, each different enough that they must be solved independently.
I wonder if AI might be able to come up with new PoW algorithms in a nightly CI job so every day is a different puzzle...
...That sounds like entropy? As in, the thing computers are bad at (truly random numbers) and /dev/urandom in your kernel already spits out an approximation of?
You can do this on yours. Just have the client and server add an extra "2" after the challenge key or something. A different client which extracts the challenge key and does its own processing will only generate invalid responses.
Cool, so then that invalidates the ASIC problem, right?
My earlier idea was to imagine that each day Anubis picks an entirely different problem-class. Ex: one day it is Sha256, the next it is prime factorization, the next it is twin-prime-finding, the next it is cracking elliptic curves, the next it is some kind of sorting / information theory problem...
All with the goal of adapting constantly so that scrapers have a harder time optimizing for the PoW problem (i.e. with Sha256 ASICs)
The way they're internally implemented doesn't allow pinning an IP. They buy a rotating proxy service from a vendor, and don't get to choose their source IP.
It's not hard to test. Go to a page that demands PoW, change your IP and see what happens. I just did it. Spoiler: kernel.org asks for a new PoW.
If the source IP was an issue, you could do it other ways: for example, make the cookie rotate on every access, and insist there is a single stream of accesses.
Why are you and other defenders of the Anubis approach so fixated on this one specific limitation of a certain type of scraping architecture? It’s hardly an immutable characteristic.
You say “they” as if all scrapers are a monolithic group with the same constraints and goals. Part of the problem is the massive diversity.
Whoa, it's MaskRay! Your blog has been a lifesaver for me every time I've needed to understand some obscure detail about linker behavior. Thanks a ton for all your work in this area, and thanks even more for writing about it.
> If something goes wrong in this process—the overwhelmingly most likely one is that the payer doesn’t have the funds to cover the check (NSF, or “insufficient funds”), but the check being fraudulent or unauthorized is also possible—that wrongness may not be discovered before money “moves” to your bank. And so that payment can be recalled from your bank to the bank the check is drawn on. This will likely result in the bank attempting to recall the money from your account.
> And so by presenting your check, which you think is substantially terminating a transaction, you are actually creating a new credit extension with your bank. They are extremely aware that you just asked them to advance you money, even if you are not aware that you did that. They already partially underwrote this extension of credit; that is why you were not shooed out of the building when you originally asked for a checking account.
The article is about depositing a check as an extension of credit; not spending via check. This creates a creditworthiness requirement that is in practice one of the most common reasons for someone to not have access to a bank account.
I had the same experience reading A Midsummer Night's Dream in high school -- I could follow the story, but it was a slog to get through and I didn't understand what was supposed to be so great about it.
Then I saw an actual performance of the play recently, and it all clicked. It was just as understandable as any contemporary film, and easily one of the most fun and entertaining stories I've seen.
I've had a similar experience, with the same play, but in an absolutely minimalist setting. No sets, the only props were a few chairs, and costumes were the same white slacks and smock worn by all the players.
And it absolutely worked.
This wasn't long after my secondary education --- I'd just started uni. But enough had happened that I got it.
As I've commented recently, even as an excellent reader, there's something that hits quite differently from hearing fiction (or even nonfiction) not just read but acted. Selected Shorts is an excellent source for this. See: <https://news.ycombinator.com/item?id=49163756>.
I think that is one thing about trying to study Shakespeare in class that can be improved. The class shouldn't just be taking turns saying the lines out loud. But you don't need a full theatrical production either or memorize lines. Read the lines with some light acting out of scenes with just the tables, chairs, and a small bit of imagination could be the class room study impactful.
In retrospect, I think I have to apologize to my English teacher. I thought it was too ambitious having the class put on a Shakespeare play. But it was fun and we apparently learned a lot.
I couldn't find the exact point that I was thinking of (when Hamlet is telling Ophelia that men are basically all just assholes), but I think this modern-day version of Hamlet helps to illustrate your point: https://www.youtube.com/watch?v=LdZVR4Ry3jQ
I gave up trying to read plays shortly after school. There's no point to me, they are supposed to be performed. I still think it's worth studying some of the more famous ones in school, but any study should absolutely involve actually going to see the play at some point.
You'd probably be in a better place to understand the scripts now, having seen more performances, for whatever that might be worth. That said, I personally don't read scripts, and have not done so since I was last required to in high school.
I commented elsewhere but my schooling had a pattern of: read > analyze/discuss > watch movie > analyze/discuss. Most of the material has film adaptations and this was fairly sufficient in closing the loop during the course.
Even in history class, when we studied key documents the teacher would often at some point show us an oral presentation of the document- like a clip of (actor portrayed) Abraham Lincoln reading the Gettysburg Address. It’s such a well written speech that’s it’s exceedingly simple to comprehend when read. However, when you hear it orated it becomes evident how powerful it was and monumental moment in American history to justify learning about as it was unifying and motivating for all even in the midst of civil war. (Sorry may not be relatable reference for non Americans was just the first that came to mind)
I think the derailment of this thread into "is Rust's memory safety good or not" is unfortunate and tedious (this debate must have happened hundreds of times by now on this site alone), but I also think it's unfair to lay the blame on Steve for this. Steve left a thoroughly glowing comment praising Zig's work on compiler performance, and comparing to his perspective on the early days of Rust. He mentioned in passing that he's not a Zig user due to preferring to work in memory-safe languages, which was polite, brief, clearly his personal position, and in my opinion an acceptable way to disclose his relationship with Zig without derailing the thread to be about memory safety.
This comment spawned two subthreads. One of them was focused on the differences between Rust and Zig's compilation model, which is directly relevant to the article and illuminating regarding the engineering tradeoffs.
In the other subthread, pron posted paragraphs and paragraphs arguing about what memory safety really means and whether or not Steve is right to have his opinion that Rust is "safe". This tangent had essentially nothing to do with the content of Steve's comment; it (and not Steve's initial comment) was the point where the thread was derailed from the topic of Zig's incremental compilation model. Steve responded politely in this thread to comments and questions directed at him, but did not fan the flames or take the thread further into off-topicness. If the moderators collapsed pron's comment or detached it and pinned it to the bottom of the page, this comment thread would be much better and much more respectful to the Zig project.
I think the RESF trope is just about dead now; it's given way to the Rust Detractor Strike Force showing up to turn unrelated threads into tangential arguments about why Rust is bad.
hint-mostly-unused defers codegen of a crate's functions until compiling a dependent crate where those functions are actually called. Therefore, unused functions will not need to be codegen'd.
The downside is that functions which are called from multiple dependent crates will need to be codegen'd in each of their dependents, so this can increase compile times if the crate is not "mostly unused."
So it's not quite as powerful as full demand-driven compilation, because of how Rust separates the compilation process into separate crates.
The only Wikia sites that had anything worthwhile were the ones that started as independent communities and got bought out, like Memory Alpha or the Minecraft wiki.
reply