Hacker Newsnew | past | comments | ask | show | jobs | submit | leumon's commentslogin


dns.adguard-dns.com

Isn't it simply that they are all renting compute from SpaceX (or whatever the company is called that offers these gpus)?

OpenAI does not

Except it's using custom api settings with which gpt-5.6 sol already got 38%.

The annotation on arc-agi-3 is this: > OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations.

With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.


Just to clarify, the 38.3% is on the public set, which is easier. On the private set it’s probably more like 30ish. (This hasn’t been run by ARC, so we can only estimate at the moment.)

next, try: "generate an svg of a human hand". this is a prompt where many models fail imo.

So 89.4% on Terminal Bench 2 but only 19.1% on Tbench 4. Opus 5 is 89.1%/51.8%.

I was thinking the same, obvious suspicion is they benchmaxed it on older bench.

Google has always done quite a lot of benchmaxing for Gemini.

You probably mean 3.5-flash? Pro is still good for a lot of use cases, but it seems it's still officially in the "preview" phase.

Fable 5.1 seems to be the first model who can accurately draw an airbus a320 in 3D space given a set of limited tools (a brush with params color, size hardness and xyz coords): https://youtube.com/shorts/vyHsMqop2yw

how about trying to draw an airbus a320 in 3d space using only one brush tool that can be moved to specific x,y,z coordinates (and its color, size & hardness can be changed). i think fable 5.1 did quite a good job (reasoning high, cost $0,261): https://files.catbox.moe/umx102.png

for comparision, this is fable 5: https://files.catbox.moe/ihl4m1.png


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: