Hacker Newsnew | past | comments | ask | show | jobs | submit | crs_gentleman's commentslogin

We tried this, and it works :) https://arxiv.org/abs/2508.06111

You have to be careful about degenerate / duplicate Qs, as a sibling commenter mentioned.

Recently though, we found that reasoning models have trouble making code-output-prediction tasks (the initial family of verifiable tasks we started with) which other reasoning models can't solve.

We started looking into harder / more agentic tasks (e.g. passing tests, using AISI's Inspect framework) but deprioritised.


Ha! I've considered this, and would really value pointers, if you're able to share?


I haven’t managed to make ends meet yet. It really is uncountable.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: