The 370 years thing is doing a lot of work in the headline imo. If people had already figured out it was probably a book cipher, this is impressive, but it's not quite the same as solving something that hundreds of experts had been actively working on
Stuff like this just makes my skin crawl. Sad part is that no matter how much you try to avoid it. There will come a time when everyone will have adopted these technologies
I'm struggling to see the upside of replacing working coreutils when the replacement still has bugs that the old implementation doesn't. Rust being safer is nice, but that doesn't help much if the new rm segfaults
The thing that bothers me here is less that the model cheated and more that it found a way to improve the score that the people running the test didn't intend. That's a pretty nasty failure once you start giving these things more control
I must say I am becoming a huge fan of Deepseek. They keep putting out capable models and actually tell people a lot about how they built them. Even if I don't understand every part of it, I'd rather see companies show their work than give us a few benchmark charts and call it a day
What I find interesting is how AI has changed the cost of maintaining two native apps enough that a decision that made no sense a few years ago is worth revisiting now
The higher price seems less important if it actually gets the job done with fewer tokens. I'm still very worried that this will end up coming back to bite us, by becoming more expensive once they inevitably nerf it. Every major model provider does that now after all
reply