Hacker Newsnew | past | comments | ask | show | jobs | submit | afiodorov's commentslogin

Queryable climate data for major cities. The main question I'm trying to answer is: where can I maximize the number of comfortable hours I can spend outdoors, whether doing light activity or just sitting around?

DeepSeek can query a processed climate dataset directly, and other agents can use the same interface:

https://climate.fiodorov.es

---

See also:

https://hn.fiodorov.es for RAG search across Hacker News comments


How are you using deepseek to keep your cost low when your traffic explodes ?

So far really wasn't an issue - I put $5 there and forgot about it. If traffic explodes will probably offer bring your own key or add credits/ads. Also just providing MCP for your agent to use is an option.

I used to be a PhD student more than a decade ago, and I published a paper containing a solution to an open problem. Yet shortly after my first publication I became increasingly disillusioned, because I started to think that within my lifetime AI would reach and eventually surpass my ability to solve such problems—and that we only had a decade or two left.

So I started saying that it only made sense to focus on problems whose solutions would be useful immediately. I even emailed my supervisor about it, arguing that our efforts were “pointless” in the sense that AI-related problems were much more pertinent and had to be prioritised.

My supervisor thought I was bonkers. I still have the email, though. Quoting myself from April 2015:

> By 2030-2040 we will have enough computing power to simulate a human brain neuron by neuron. Once we manage to create a human intelligence we will be one little step away from super intelligence: just set the intelligence to modify itself and see the exponential growth in action. Our human intelligence is bounded by a number of biological factors (e.g. size of a skull) and even the smart human who has ever lived will appear to be a primitive ant to a supper intelligence (machine intelligence will also have perfect motivation). There is plenty of literature on this if you are interested in discussing this further. > > What does it have to do with research in pure maths? I can say that research in pure maths which won't come handy in the next 60 years is just wasted effort. The super intelligence will be able to do maths way better than humans. I believe a lot of current efforts should go into researching of artificial intelligence (or areas to do with AI) instead rather than the pure maths. I want to be proven wrong but most mathematicians I interact with are too narrow-minded to counter me and they just laugh about even contemplating the above. Frankly I am myself so perplexed that I take the above seriously, but I do and it's hurting my motivation.

I'm quite curious what my supervisor thinks of that email now.


I think, with all due respect, that your supervisor would be correct to think of that email as an insult to both them and to theoretical mathematicians as a field.

She was not offended but thought I could benefit from therapy. She didn’t share my opinion - she said it’d not bother her if AI could eventually solve problems for her.

I later realised that my motivation depended heavily on believing I was making a contribution that would otherwise go unmade. The prospect of AI doing that work undermined my motivation, but it needn’t undermine hers. I was trying to explain why I was struggling to continue, though I can see how calling the work “pointless” came across as dismissive.


Fantastic! E-mail him and ask if he thought about it at all

I prefer getting rid of old devices that needing to store old cables ad-infinitum.

A RAG search accross all HackerNews comments is back online with a new design: https://hn.fiodorov.es/


I've been embedding all HN comments since 2023 from BigQuery and hosting at https://hn.fiodorov.es

Source is at https://github.com/afiodorov/hn-search


I appreciate the architectural info and details in the GH repo. Cool project.


That's cool - it gave me quite a good answer when I tried it. Does it cost you much to run?

I tried "Who's Gary Marcus" - HN / your thing was considerably more negative about him than Google.


The running costs are very low. Since posting it today we burned 30 cents in DeepSeek inference. Postgres instance though costs me $40 a month on Railway; mostly due to RAM usage during to HNSW incremental update.


That's cool! Some immediate UI feedback after search button is clicked would be nice, I had to press it several times until I noticed some feedback. Maybe just disable it once clicked, my 2 cents


I have a question: what hardware did you use and how long did you need to generate the embeddings ?


Daily updates I do on my m4 mac air: takes about 5 minutes to process roughly 10k fresh comments. Historic backfill was done on an Nvidia GPU rented on vast.ai for a few dollars. If I recall correctly took about an hour or so. It’s mentioned in the README.md on GitHub.


What mechanisms do you have to allow people to remove their comments from your databae


Can users here submit an issue to have data associated with their account removed?


GDPR still holds, so I don’t see why not if that’s what your request is under.

However, it’s out there- and you have no idea where, so there’s not really a moral or feasible way to get rid of it everywhere. (Please don’t nuke the world just to clean your rep.)


The law (at least, in the EU) grants a legal right to privacy, and the motivation behind it is really none of anyone’s business.

Maybe commenters face threats to safety. Maybe commenters didn’t think AI companies profiting off of their non-commercial conversations would ever exist and wouldn’t have put data out there if that was disclosed ahead of time.

Corporations have an unlimited right to bully and threaten to take down embarrassing content and hide their mistakes, they have greatly enhanced leverage over copyright enforcement compared to individuals, but then if individuals do a much less egregious thing to try and take down their content they don’t even get paid for it’s immoral.

This community financially benefits YCombinator and its portfolio companies. Without our contributions, readership, and comments, their ability to hire and recruit founders is diminished. They don’t provide a delete button for profit-motivated reasons, and privacy laws like GDPR guard against that.

(As you might guess, I am personally quite against HN’s policy forbidding most forms of content deletion. Their policy and solution involving manual modifications via the moderation team makes no sense - every other social media platform lets you delete your content)


Finally someone mentioned it. I'm surprised all the "tech enthusiasts" here turn a blind eye when it's their own community, but if it's someone else's then it's atrocious.


Very cool, well done!


Apparently Persian and Russian are close. Which is surprising to say the least. I know people keep getting confused about how Portuguese from Portugal and Russian sound close yet the Persian is new to me.


Idea: Farsi and Russian both have simple list of vowel sounds and no diphtongs. Making it hard/obvious when attempting to speak english, which is rife with them and many different vowel sounds


While Persian has only two diphtongs and 6-8 vowels, Other Languages of Iran are full of them(e.g. Southern Kurdish speakers can pronounce 12+1 vowels and 11 diphtongs). I find it funny if all Iranians are speaking English with the Persian accent.


Yeh they seem to be in the same "major" cluster, although Serbian/Croatian, Romanian, Bulgarian, Turkish, Polish and Czech are all close.

Turkish and Persian seem to be the nearest neighbors.


When I went to Portugal I was struck by how much Portuguese there does sound like Spanish with a Russian accent!


Part of this is the "dark L" sound


I’d guess that the sibilants, consonant clusters, and/or vowel reduction would play a big role.


I thought I was the only one who perceived an audible similarity between Portuguese and Russian.


I am native Russian speaker, and work/visited Portugal. It definitely tricks me when not paying attention, its very similar sounding


I speak neither, and both also sound similar to me depending on the accents of the speakers.


I had that too but it was Brazillian Portuguese where I noticed it.


The characteristics that make pt-PT sound similar to Russian are largely absent in pt-BR.


I've found that building my side projects to be "scalable" is a practical side effect of choosing the most cost-effective hosting.

When a project has little to no traffic, the on-demand pricing of serverless is unbeatable. A static site on S3 or a backend on Lambda with DynamoDB will cost nothing under the AWS free tier. A dedicated server, even a cheap one, is an immediate and fixed $8-10/month liability.

The cost to run a monolith on a VPS only becomes competitive once you have enough users to burn through the very generous free tiers, which for many side projects is a long way off. The primary driver here is minimizing cost and operational overhead from day one.


> A dedicated server, even a cheap one, is an immediate and fixed $8-10/month liability.

Personally, I am more worried about the infinitely-scalable service potentially (liability) sending a huge bill after the fact. This "liability" of $8-10 is predictable, like a Netflix subscription.


RAG search that contains all HN comments since 2023

https://hn.fiodorov.es

I treat it more like a homework exercise for a Coursera course but I like the result.


> It was uncomfortable at first. I had to learn to let go of reading every line of PR code. I still read the tests pretty carefully, but the specs became our source of truth for what was being built and why.

This is exactly right. Our role is shifting from writing implementation details to defining and verifying behavior.

I recently needed to add recursive uploads to a complex S3-to-SFTP Python operator that had a dozen path manipulation flags. My process was:

* Extract the existing behavior into a clear spec (i.e., get the unit tests passing).

* Expand that spec to cover the new recursive functionality.

* Hand the problem and the tests to a coding agent.

I quickly realized I didn't need to understand the old code at all. My entire focus was on whether the new code was faithful to the spec. This is the future: our value will be in demonstrating correctness through verification, while the code itself becomes an implementation detail handled by an agent.


> Our role is shifting from writing implementation details to defining and verifying behavior.

I could argue that our main job was always that - defining and verifying behavior. As in, it was a large part of the job. Time spent on writing implementation details have always been on a downward trend via higher level languages, compilers and other abstractions.


Tell that to all the engineers that want to argue over minutia for days in a PR


> My entire focus was on whether the new code was faithful to the spec

This may be true, but see Postel's Law, that says that the observed behavior of a heavily-used system becomes its public interface and specification, with all its quirks and implementation errors. It may be important to keep testing that the clients using the code are also faithful to the spec, and detect and handle discrepancies.


I believe that's Hyrum's Law.


Claude Plays Pokemon showed that too. AI is bad at deciding when something is "working" - it will go in circles forever. But an AI combined with a human to occasionally course correct is a powerful combo.


If you actually define every inch of behavior, you are pretty much writing code. If there's any line in the PR that you can't instantly grok the meaning of, you probably haven't defined the full breadth of the behavior.


  sign of serious organizational disfunction.
You're not wrong, but it's a "dysfunction" that many successful tech companies have learned to leverage.

The reality is, most engineers spend far less than half their time writing new code. This is where the 80/20 principle comes into play. It's common for 80% of a company's revenue to come from 20% of its features. That core, revenue-generating code is often mature and requires more maintenance than new code. Its stability allows the company to afford what you call "dysfunction": having a large portion of engineers work on speculative features and "big bets" that might never see the light of day.

So, while it looks like a bug from a pure "coding hours" perspective, for many businesses, it's a strategic feature!


I suspect a lot of that organizational dysfunction is related to a couple of things that might be changed by adjusting individual developer coding productivity:

1) aligning the work of multiple developers

2) ensuring that developer attention is focused only on the right problems

3) updating stakeholders on progress of code buildout

4) preventing too much code being produced because of the maintenance burden

If agentic tooling reduces the cost of code ownership, annd allows individual developers to make more changes across a broader scope of a codebase more quickly, all of this organizational overhead also needs to be revisited.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: