Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
erekp
on March 31, 2025
|
parent
|
context
|
favorite
| on:
Ask HN: What are you working on? (March 2025)
how do you exactly fallback to common crawl? isn't the cost to even hold and query common crawl insane?
andrethegiant
on March 31, 2025
|
next
[–]
With AWS Athena, you can query the contents of someone else’s public S3 bucket. You pay per read, but if you craft your query the right way then it’s very inexpensive. Each query I run only scans about 1MB of data.
wfn
on April 1, 2025
|
prev
[–]
Since I was just looking at this accidentally, here are some examples of how to query at a ~cent-per-query cost level (just examples but quite illustrative):
https://commoncrawl.org/blog/index-to-warc-files-and-urls-in...
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: