Opposite in my experience. I need to limit codex to 500k on medium/low, still run out in 2-3 days with 1 CLI window. CC gives me 4-5 medium/high days with 2-3 CLI windows, and Opus is still great for other regular dumb engineering/refactoring.
On the other hand my head starts to hurt if I read Opus for too long, hopefully they fixed it with 5.5.
This is my experience. After months of hearing how Codex limits were way higher I bumped to the $100/mo plan after hitting my limits a day early on Claude due to some heavy usage + Fable (not normal for me, I often fit nicely in the $200/mo plan). I hit the usage limit in a day with a single agent running on codex and the tiny context window was stifling. Yes, I'm comparing a $100 to a $200 plan but I extrapolated the usage (4x'd it) and it still wasn't close, I got way more done with Opus.
Using Agentsview (which might have it's own issues) I was getting ~$200 of API usage in my 1 week Codex window (paid $100) vs ~$5,000 of API usage in 1 week for Claude (paid $200).
I think that such fine-tuning hinges on what do you consider to be a mistake, which is context dependent. Having task-specific finetuned models goes against the status quo of generalization/centralization where few large companies serve a limited amount of models efficiently - both due to inference efficiency and the need/want to control the model.
Having a human-like LLM ecosystem with deep specialization requires a paradigm change in how LLMs are trained - and held accountable. How do we put trust in a specific finetuned LLM rather than the institution behind it? Is there any better approach than the very inefficient evolutionary?
In Europe I don't think it's commonly done. For example, the standard Dutch MSM STI test [1] covers HIV, Syphilis, Chlamydia, and Gonorrhoea only while Trichomoniasis, HPV 2 and Herpes are rather common. That being said, it's freely available at most hospitals.
Herpes screening is discouraged in the US unless symptoms are present as the widely available IgG tests have terrible specificity. While the false negative rate is practically 0, they have a false positive rate of around 20%. The only high-accuracy confirmatory test (at least in the US) is the Western Blot but there's only a handful of labs in the country that offer it.
The current tests would result in a substantial fraction of the population running around thinking they have a highly stigmatized STI that they in fact do not have with little to zero ways of getting a more accurate test.
Seems like both HSV-1 and HSV-2 seroprevalences are actually rather low in the Netherlands.
As for STI testing: Serological testing is pretty meaningless regarding infection risk and direct viral detection is only done on visible lesions for differential diagnostics AFAIK. You can shed the virus without lesions, HSV-1 can be contracted without direct mucosa contact, and while HSV-1 is associated with oral herpes and HSV-2 with genital herpes, both types can cause lesions anywhere on the body. Sure someone without positive antibody status very likely can't transmit it, but that's a very small minority of people of relevant age.
I believe talking to Americans about "herpes" can be misleading, as it's weirdly stigmatized over there and people seemingly conflate visible "cold sores", genital herpes and antibody status.
Incidentally, genital herpes and HSV-2 in general is more common in the US, which may be a direct outcome of this cultural dynamic. Due to lower incidences of early oral HSV-1 infections, compared to most western Europe, Americans are more likely to get infected with genital herpes of either type, but in particular HSV-2, too. Prior oral HSV-1 infections are somewhat protective.
Most of the stigmatization of HSV in the US is mostly the result of pharmaceutical companies trying to create a market for HSV antiviral drugs. In the 1980s they ran ads that equated HSV with leprosy. Since that was also the time of the AIDS scare the public was probably very susceptible. Even in my childhood I remember ads on TV for the antiviral drugs promising a way of "living a normal life" despite HSV infection.
That's crazy. Herpes can't be cured, so the implication of those ads can only be to "hide" the infection. Additionally, those antivirals aren't effective at preventing outbreaks unless you take them every day, which would be insane unless you are an unfortunate soul who is prone to constant outbreaks.
Assuming a predominantly male population, the suicide rate should be around 20 per 100000 per year. 5 people in a month gives 18x higher rate. Too short period to draw any statistical conclusions but it’s a large blip.
From a life spent on count statistics, one thing I think is interesting is how definitively you can reject perfect randomness from a very small sample like this.
The expected event count from 17,000 people exposed to a month of constant uniform event risk with an annual hazard rate of 20 per 100,000 per year is 0.28. The p-value for observing a count of 5 from a (Poissonian) mean of 0.28 is 1 in 100,000. So it's pretty unlikely by perfect random chance.
The other thing you see in a life spent in count statistics is a lot of statistically extreme events that turn out to have no deep or predictive significance.
Maybe the real reason that small sample sizes don't tell you anything is because there's tons of small samples in the world, and the particular one you're looking at has almost certainly been selected for unusualness.
It's a multiple comparison problem that you don't know is a multiple comparison problem because you yourself didn't do the comparisons. You just got the result of them.
I think you’d have some statistical complications around stress: beyond the inherent pressure of the job itself, it’s highly likely that they were understaffed even before losing a lot of support when the DOGE chainsaws were in “cut first, measure later” mode — I know multiple people in the military who are still running at like half productivity because telework was cancelled so everyone’s day is a couple hours longer commuting into an office where they’re sitting on the floor because there aren’t enough desks — and if they were married to another federal employee they might be dealing with loss of income or childcare (DOGE closed the GSA office which handled that, apparently to pressure parents into resigning) at the same time.
It's wild to think how many indirect deaths Elon Musk and his "chainsaw" must be responsible for. I recently saw an interview and he was asked the question and he said "not a single death"...a true psychopath.
The suicide rate isn't uniformly distributed given gender. There are likely screenings (previous mental health history) and rejection of the most affected age group (70+) that should lead the baseline to be even lower. Meaning 18x is probably on the low end.
I mean, let's be real for a second. It's cybersecurity.
Yes, it does skew male, but you're probably looking at 10x the standard population density of Autistics, and 20x the density of trans women.
I'm sure there are other "sociologically valid" ways you can measure this too, but there is absolutely a correlation between having the type of brain that is able to perform offensive cybersecurity operations and mental illness/social rejection/isolation.
This particular population is highly selected because of the work they do and the types of clearances they hold. I would expect the suicide rate to differ significantly from the general population.
I think that it would be necessary to correct for things like income and education level since engineers in cyber command are likely not representative of the general public.
I agree with you generally but a lot of professions that are well-paid and educated still have high rates. Dentists, for example, are 2x the general population
most medical professions have higher rate. it's usually attributed to being constantly subjected to pain and suffering plus physicians can source pharmaceuticals to exit softly. never heard that it professionals are particularly prone to suicide.
Correct me if I'm wrong, but all of the aforementioned advances were made in the last year? Until very recently few people had access to these tools. Most people still don't know how to use ChatGPT, and very few use tools like CC regularily. If in a few years these frontier tools become commonplace and people upskill we would should see a network effect?
IIRC Taiwan took a page out of Singapore's playbook and went all in on electrical engineering and adjecent fields. It was very much a long-term strategy. Germany probably didn't feel nearly as much pressure, and was already very strong in all industry.
It's from the president's speech. Too lazy to look up the actual text but I guess he meant "pillars", a common metaphor in East Asia. In English axis and pillar are distinct but in East Asia the line is blurry.
For example, the Japanese word 軸 (jiku) is used to mean the "axis" of a graph, but it is also used in business to mean the "core pillar/backbone" of a strategy (e.g., 経営の軸 keiei no jiku, literally "the axis of management," but conceptually "the pillar of management").
The speech was delivered in Korean so this is a choice by a translator. I don’t speak Korean but I asked an LLM and it says …
the phrase used is "대도약" (daedoyak), which literally means "great leap forward" or "great jump forward." This is NOT "대약진" (daeyakjin), which would be the direct translation of China's "Great Leap Forward" (大跃进).
To expand a bit, even saying 대약진 _daeyakjin_ "great leap forward" wouldn't have turned many eyes, because _dae_ is just a common prefix ("great") and _yakjin_ is also a common word meaning "leap forward, push forward, improve". The word simply doesn't have the same connotation of the English phrase "Great Leap Forward", which is almost always used for the infamous Chinese movement.
If a Korean speaker wanted to talk about that Chinese movement, they'd use the full name, 대약진운동 (大跃进运动): the great leap forward movement.
Personal anecdote on ROI - I was at an early stage startup earlier this year where we had some burstable long-running GPU tasks (<100 VMs). Accross GCP and OCI we couldn't get our hands on L40S on-demand, and had to resort to T4s (released 2018). Sometimes even these were unavailable, and we would have a P4 (2016!) fallback. AWS sells A100s (2020) at $4/hr except they don't even have capacity for x1 versions, you have to rent x8.
Opposite in my experience. I need to limit codex to 500k on medium/low, still run out in 2-3 days with 1 CLI window. CC gives me 4-5 medium/high days with 2-3 CLI windows, and Opus is still great for other regular dumb engineering/refactoring.
On the other hand my head starts to hurt if I read Opus for too long, hopefully they fixed it with 5.5.
reply