Sol does not follow instructions well at all. I've caught it multiple times a day now since release going off in incredibly bone-headed directions. It's so easy for it to over-interpret, make wildly out of scope changes, or just completely mis-understand what you're saying. The code it writes also is quite bloated still, and in my still-forming understanding I feel Fable is much more reliably smart. Sol is just more persistent and fast, so it often gets there quicker or after many tries, where Fable will just get it right from the start albeit more slowly.
A friend of mine just got one, ex-Chrome core dev so a fairly sharp guy, his one month review was that it was incredibly capable but had already done two maneuvers that would've led to an accident without intervention.
Sure, it still needs supervision. Today. But it has definitely passed a threshold where it is now safer to supervise it than to drive without it. And it continues to improve quickly. I expect it to work unsupervised within two years (and unlike Elon I have not been saying this every year for the past decade).
The feedback I heard was definitely not that. The mistakes it makes are incredibly hard to predict, and they were lucky that no one was on the side of the road as they could've killed someone if a person had been there and they'd been a half second late.
That's the point. It's actually worse than a dumber system, because it's not even past the margin where it can go a month without a mistake, but you don't have to correct most days. The absolute worst possible scenario for safety.
From the very beginning of self-driving people have been making this claim that supervised systems would cause complacency that would make them perversely less safe. A lot of people believe it, but it's always been just an opinion based on speculation, not data. It's actually an empirical question. The data is in and this claim is definitively disproven. There is no increase in crashes when people use FSD. Not back when it was actually bad, not recently when it was just OK, and still not today when it's actually quite good but not perfect.
I built an iOS simulator simulator, though only for RN. Runs in browser but covers 100% of the API of RN, iOS UI, and the top 1k native libraries basically now. Been an ongoing agentic experiment of mine that's about ready to release.
Kind of fun, you can develop iOS and Android both without a build step and without a Mac even.
I agree we're at diminishing returns, but when brain scanning gets good enough you have a better dataset than the internet, Facebook is deep into that research.
I did initially through some miracle, as I wasn’t overweight. But after that year was up I now do grey market. Finnrick does testing which seems like a decent way to source if you’re looking that way. Not affiliated though and haven’t done a lot of background research on them.
I've been posting about this including here for years now. I wrote a long post about it a while ago here and on Reddit. At time no one was talking about it, and actually my Reddit post was buried behind tons of others which was frustrating at the time given I had basically shared a partial cure.
Now if you search "reddit eds glp-1" or tirzepatide you'll see tons and tons of long threads of people all saying the same thing - it's the only thing that actually helps.
For me it was something of a miracle, I have two overlapping immune issues and it seems to just turn me into a much more normal, functional person. Including fixing my sleep.
Never was overweight beyond maybe ~15lbs btw when I started or took it, and the effects are 100% not because of just fasting or weight loss. I had tried keto and OMAD before, and been at healthy weight my whole life.
There's a definite auto-immune modulating mechanism and it's so strong it seems better than basically most first-class drugs. Even things like prednisone which are like nuclear weapons don't give me relief like Tirzepatide does.
Btw highly recommend Tirzepatide of the three GLP-1 drugs, for me at least it's by far the most effective and least side effects.
I easily burn through 3 $200 plans in less than a week. I am often using 4-6 sessions at once and do run overnight goals though typically 2 at once. Almost never use fast.
Claude plans are more generous now by about 2-3x but Anthropic slowed their tps a month or so ago so you’re not getting the speed. It’s flip flopped, Codex tightened it significantly recently and used to be more generous.
I do split between work, personal and OSS projects, which is why I have the plans.
This has been my experience as well, at least for the last few weeks. Codex 5.5 is the better planner and coder across big projects, but Opus is fine, though my Claude 5 hour window lasts ~2x longer than Codex. So I’ll sometimes use an orchestrator/worker skill to spread the load.
I’ve hired many asian developers anywhere from 1-4k a month.
I get a lot more out of a 200/mo subscription now in a week than I did from them in a month.
Now obviously in today’s world they’d be using a 200/mo subscription themselves. But it’s not like money is nothing, software development doesn’t scale down below 1k/mo for anyone competent even in the poorest areas.
The point the post you replied to is making is that while you get value out of it, and in your case it's not that expensive, it's just simply not the case worldwide
I somehow take the opposite on almost everything here.
4.8 xhigh or max has a slight edge on 5.5 xhigh, for very complex logic perhaps it loses but it's just better in almost every other way, especially code quality. GPT is a slop machine outputs way too much and over-abstracts, plus its communication is so bad in comparison.
Fable was for sure a step above GPT, I tried them both against a few of the same hard tasks and it was not a small difference.