Hacker Newsnew | past | comments | ask | show | jobs | submit | lacker's commentslogin

Recently I showed my 7-yr-old son and a friend of his how they could build their own little video games with the Claude app. Simple stuff, make me a maze game, make the walls move sometimes, make it bigger, make there be a score.

The delight they have at, video games aren't just a thing to play, it's a thing they can design and change, share with each other... I think that dream of a truly malleable experience is still there, I think this "explosion of new forms" you talk about is happening right now with AI, and I think this time it is going to be bigger than ever.


Games that allow custom map is not that new. Minecraft also allows you some basic boolean mechanism that you can build surprisingly complex stuff. I don't know why do you need AI for this

I think the tech companies would be supporters of new rules that would let builders create more housing all throughout San Francisco. We need a big tent, with all pro-housing groups coming together, to change the rules, and start allowing new housing again.

If you own a lot in San Francisco you should be able to build an apartment building on it.


This is like giving you three outputs from md5sum and asking you to guess for which one the input ended in a "q". There's no way to tell unless you break the RNG.


Yeah, it's a pointless exercise. I hope the author is just trolling given that he is knowledgeable in the field.

> Here are three 64-character hex strings. Two are random. One is HMAC-SHA256(secret_key, "anthropic"). You don't have the key. Which one is the HMAC?


I think the point is probably to help convince people that the watermarking doesn’t perceptibly impact quality, which is a concern some people have (whether well founded or not).


"doesn’t perceptibly impact quality" no not perceptibly but it does.


> no not perceptibly but it does.

Similar to how a single particle of dust landing on your shoulder makes you weigh more.

Yes it does - but anyone arguing that is completely missing the point.


Even then, LLMs are based on a lossy compression, so the quality is harmed by design.


Maybe some sort of blockchain could make it so that you don't have to verify the whole proof yourself, to know that it's true. Because some of these proofs could take a long time to verify. Similar to a Merkle tree, but you'd have to make it respect the laws of mathematical proof in some way.

edit: Yes I think this should be possible, using a "recursive SNARK".


I don't think the mathematicians are going to be able to make that work, because journals are already struggling to keep up with their review load, and AI seems like it will make that harder. So a solution that involves "journals will do a lot more effort to review each paper" doesn't seem practical.

It would work better as a bar for hiring, rather than as a bar for publishing.


It will be interesting to see the evolution of journals in the next ten years for sure. Have they outlived their usefulness? Maybe everyone will just upload papers to arXiv, along with a copy of the formal proof.


Just package the proof as a library and put it in some source code repository like github.



My conclusion is the opposite. If benchmarks were meaningless, surely Meta would be able to find some benchmark that shows they are better than Sol and Fable. The fact that they can't do that tells me that benchmarks still do mean something.


Muse 1.1 performed relatively well according to benchmarks, putting it within spitting distance of the premier models. However, based on the results I got from it and the review videos I watched, it wasn’t even close.

Opus 5 is incredible at making games. Almost like a generation better than other models from my experience. You won't see that if you just look at the popular benchmarks..

You have to test each model on your actual use case to see how well it really performs.


Yeah, totally. But... the problem is that I don't have time to test every single model that comes out. So I rely on reports like this to decide, should I even bother testing out Muse?


> Opus 5 is incredible at making games.

This is a bit vague. What sort of games with what technology?


My son was gifted an old Mac from his grandparents. It only supports OSX 10.13. I’ve been able to make several games that he genuinely likes (7 yo) Opus built them on my workstation and then pushed them to his computer and tested them over SSH. It handled all the asset creation or collection from CC0 licensed sources. I believe everything is built on the Godot engine. It’s really amazing to me. I don’t know what it would cost me to get someone to build custom games on a long deprecated computer architecture, but I paid Anthropic $20.


I don't think it is vague in the slightest. Take the most simple examples, how many LLM's have you tested making them? There are stylistic choices pertaining to games that is well beyond a 0/1 reward. Even something as basic as breakout or flappy bird can have wildly different quality between models. Yeah, you could call this animal on a bike benchmarking, but I don't think it is. IMO the problem space occupies an interesting area where you can ignore the pass/fail and focus on the actual level of the model to do something beyond that.

I doubt the OP meant something like creating the whole tech stack for WOW.


You seem to think I was disagreeing somehow.

I was just asking what kinds of games and with which technology.

Neither is stated in the original comment, and the answer obviously isn’t “every kind with every technology”.


Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.


Or they did try to game the benchmarks and just didn’t do it well enough.

Benchmarks are one data point, not the only one, but the easiest one to compare.


Right, but the point is that you can't conclude that a model is necessarily bad because it's not hitting the same scores on benchmarks. I just don't agree with lacker's conclusion, because their logic doesn't seem to consider that. Scoring lower on a benchmark doesn't strictly mean they have a bad model, but it may be the case. Like you said it's one data point, but being the easiest, and obviously most gamed, means you should probably weigh them less heavily.


Personally, I end up throwing away a decent amount of books that the book donation people won't take. Especially old technical books. Textbooks from 2001. These headlines just don't tell me anything useful. "Rare" doesn't mean anything.

Please, give me one example of an interesting and unique book that the AI companies have destroyed.


Oreilley’s 1999 classic: Learning Python.


They destroyed it, there's nowhere to buy it or read it for the public?

Not sure if that was a whoosh?

£5 .. don't think it's destroyed .. https://www.abebooks.co.uk/Learning-Python-Mark-Lutz-OReilly...


I don't think people working at Amazon "know that it is a part of a larger bad", it's one of the most trusted American institutions.

https://www.theargumentmag.com/p/why-everyone-loves-amazon


You really don't think know they are selling endless counterfeit products? Don't know they are taking part in massive return fraud against small sellers? You don't think they know they totally ignore sellers with problems even if their livelihood depends on it?


Yes there are studies, for example last year Pangram's false positives were measured to be under 0.5%.

https://www.pangram.com/blog/third-party-pangram-evals

Personally, at first I thought these sorts of tools were dumb and wouldn't really work, but I think it works because it just isn't designed to be "adversarial". If you want your AI to trick Pangram, you can make an AI to trick Pangram. It just catches people who are cutting and pasting from the AIs without putting any more effort into hiding it.


Any binary classifier can have a FPR under 0.5% if you don't have any restriction on FNR...


While I am quite skeptical of the claims linked above, that link does indeed cover the FNR at the FPR of 0.005, and finds it broadly to be on the same order of magnitude, i.e. also below 0.005.


If the FPR is very low, the FNR rate doesn't really matter if you get a positive result (unless the pre-test probability is very low, which is not the case here)


It is a smell. But it's the EU that smells bad, when it comes to tech regulation. It's the smell of cookie popup warnings.


Nothing in the law requires the pop up. It definitely doesn’t require the obnoxious bullshit that most companies put up (aka the dark pattern to get you to agree to every unreasonable part of their terms just to read the page).

The alternative would be to just stop invasive tracking and add the cookie when it’s actually needed.


Yet somehow all the government/EU-institution pages I visit also choose to track and throw the popup.


Yes, there's a lot of cargo culting in web development.

Many US based companies also do this for US visitors, which is absolutely not required by the GDPR and related regulations, because they don't apply there.

The law states:

> Receive users’ consent before you use any cookies except strictly necessary cookies.

Strictly necessary:

> These cookies are essential for you to browse the website and use its features, such as accessing secure areas of the site. Cookies that allow web shops to hold your items in your cart while you are shopping online are an example of strictly necessary cookies. These cookies will generally be first-party session cookies.

https://gdpr.eu/cookies/

You don't need consent for MOST reasonable uses of cookies. If compliance theatre wasn't such an industry the web would be a lot tidier and we could stop blaming the EU for implementing important privacy and data controls.


You're just acknowledging that intent!=effect, which is a primary criticism of these laws.


You’re shifting the goalposts somewhat, but the thing this misses is that the cookie banners are the least important aspect of EU data regulation. The principle of consent and of minimising data held has actually made a substantial different in European firms, mostly for the better.


I didn't have any initial goalposts. Maybe you are assuming I'm someone else.

I agree with you that cookies banners are used more than legally necessary. They are a consequence of the law nonetheless.


The cookie popup also more often than not doesn't satisfy GDPR, since the option to remove consent disappears with the popup. These dark patterns emerged because the GDPR was used selectively as a club rather than properly enforced. That led to what another comment refers to as "compliance theatre" rather than actual corporate compliance.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: