Hacker Newsnew | past | comments | ask | show | jobs | submit | zorked's commentslogin

Well, if Robert O'Callahan reads your post, I am pretty positive he will think he made the right choice.

And they would have to be confident that you wouldn't cheat on it.

People have been using LLMs for two year. It's not just this week's LLM release that is capable something.


A capability isn't binary. There is a massive difference between can produce an impressive demo and can reliably complete the task without a human babysitting it.


Yup, new SOTA models especially with high/xhigh/max reasoning too often overengineer solutions, good for benchmarks that usually measure task completion, bad for normal development where you don't want 'rewrite in rust and 1k LOC unit tests' style solutions when agent does mundane bug fixes.


When it comes to mundane bug fixes the value is in actually finding the cause of the bug, and I find SOTA models way outperform smaller ones here. I don't care about their output - I can write the correct 5 line patch myself once I understand what's wrong.


It's not this week's change. Fable was the step change for programming. And most of truly useful and powerful capabilities arrived in the last eight months.


> Fable was the step change for programming.

AFAIK: Mistral does not even try to compete in this field. There are other use cases for LLMs beside coding. As Mistral AI wrote:

> During the first wave of generative AI, the central question was who could build the most powerful model. Organizations and governments are now asking a different one: how to harness the power of AI for their mission-critical needs without surrendering control over the infrastructure and intelligence loop. Demand for that combination of performance with control, choice and independence is growing internationally, as enterprises and governments weigh the long-term technology dependencies, data governance requirements and deployment choices that come with any AI investment.

> Mistral is the only AI company in the world building the full stack required to answer that question: open-weight models, the infrastructure and the compute capacity they run on, and the products that bring them into production; ensuring that customers are never locked into a single vendor's roadmap, pricing or availability.

> Mistral’s full-stack and open approach also allows organizations to build on it without exposing their most valuable data, workflows and institutional knowledge to anyone outside their own walls. That's what makes Mistral’s stack the sovereign AI layer, meaning retaining control across four dimensions: data that stays inside the organization's boundaries, models that are controllable and customizable, compute that is private and predictable, and systems in production that are fully controllable and auditable.


>> Fable was the step change for programming > AFAIK: Mistral does not even try to compete in this field.

They have released models speficially for programming that ”vibe coding” would be safer.

https://mistral.ai/news/leanstral/


People said this for Opus 4.6 too. Every release the models get RLHF'ed into accomplishing a new task and the people who need to do this task think there was a step change.


To be fair Opus 4.6 was genuinely a really good model when it came out, in fact I'm not sure Opus 5 is even any better.

It's definitely a lot slower, though


Yeah that's kinda my point. I'm not sure if the models have gotten that much smarter, but they're certainly getting more capable. That's not the same thing though.

There are things GPT 6.0 can accomplish for me that 5.3 was not able to. But there are also things it still fails at, and it doesn't seem to be much better at the big picture. It does spam about 100x more tests though and I wonder if just RLHFing it to test everything constantly is carrying it more. 6.0 writes so many tests and spends so much time verifying it's work in python sandboxes. Slow as hell but it tends to get things right the first time more which is good, I guess. I don't love the thought of a 500loc feature adding +4000loc due to tests though.


Yep same with latest Claude models - code isn't really any better than Opus 4.5/4.6, but use 5x as many tokens doing random stuff that's mostly unnecessary.

And yeah still for some reason they often can't understand how to set up any project locally without handholding, which is something you'd think an LLM would actually be good at


If the choice is between Mistral and no AI, I'll take Mistral any day

Even those old llama models were ok for coding

Yes yes they won't be like Claude's fire and forget (until you see how many tokens you burned to write "Hello World")


People fawn over AI brands now like cars and it's silly. OpenAI and Anthropic have been flipping spots for best LLM coder for the last two years and to say one is better feels silly; I've been using them both and they're very similar with different personalities. Recently Grok has become competitive in many aspects, and while I don't have much experience with Gemini it seems to come and go in terms of coding quality.

Saying only anthropic models are competitive frontier coding models is out of touch with the space imo


> Maybe the matching-funds donor is all peachy and clean. Maybe they're the root of all evil.

Money going from the root of all evil to the Internet Archive?

Two cakes.


I could see it being a Zuckerberg "charity" thing motivated by a need for training data, so, yeah.


Is it?

  Then Jesus said to his disciples, “Truly I tell you, it is hard for someone who is rich to enter the kingdom of heaven.  [I]t is easier for a camel to go through the eye of a needle than for someone who is rich to enter the kingdom of God.”
This guy is pretty influential in at least half the globe.


https://news.ycombinator.com/item?id=49582188

That's at least one sector of the Christian world that ignores that line. I don't think they're alone. You also have other etymological canaries like "fortune" meaning both "wealth" and "blessing". The thinking is that good things happen to good people, with "good things" including "coming into money". But, as we see, that seems to be erroneous. Just because good things happen to a person doesn't mean that they're good.


You are reading the term too literally.

(And by the way, you are using a severely outdated translation.)


"IBM Bob? Certainly they have a license deal with Microsoft to release Bob for OS/2"


The website represents this well. My random person was born in the Neolithic and lived to 68. Not bad.


There are entire countries where it's not a thing to use PayPal. Way too many people were scammed out of their money by them. In general you don't give your credit card number to random sites - it mostly goes to Stripe and similar services.


These days stripe is perhaps getting dominant, but there was an age before stripe and many small companies still have custom solutions. And how would Paypal scam you out of your money? I've only ever seen this happen to people who did illegal shit and even then you won't lose much unless you also use paypal to store money like a bank for some weird reason. But then you're definitely more of a merchant than a consumer, so it already sucks because of other reasons as mentioned above. If you use it only for sending payments, getting locked out is a non-issue, because there is no money stored there and I can always revoke their access to my bank account.


Gemini is the best general-purpose model. We hear a lot about the other ones here because we are focusing on coding.


There is an entire graveyward of browsers, almost all that didn't die are now forgotten.


Sure, but I'm sure most of their lead devs are well paid now.

In general C++ work and similar, if not in Chrome development.


cause != effect


There is a confounding variable, however -- the person is most likely a good dev. It stands to reason that they've had a decent career at least since then.


Yes but as a dev you are shaped by the projects you work on.


I think the point isn't that you'll necessarily build a great browser (or LLM), but the experience will benefit you in other ways.


Yup like you go far enought you will pick lots of transferable skills like data cleaning in the case of LLM, DOM parsing case of web browser or welding case of rocket.


If any were built by teenagers, I’d imagine those teenagers ended up with pretty good careers in technology?


For every wunderkind that has a long career there's also plenty who are overlooked or peak early.


Citation desperately needed.



Ken Silverman


https://en.wikipedia.org/wiki/Ken_Silverman

Is this the guy you’re talking about?


I think he's doing fine? Sure, he got out of the video game industry, but that's for young people to burn themselves out.


I don't mean he's destitute. Just that he's no longer exceptional, which is fine.


But you were trying to refute my point of:

> I’d imagine those teenagers ended up with pretty good careers in technology?

The guy you mentioned doesn't seem to have ended up with a bad career in technology?


Thanks for clarifying. I read it as exceptional careers


Source: Sour Grapes


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: