You have to be careful about degenerate / duplicate Qs, as a sibling commenter mentioned.
Recently though, we found that reasoning models have trouble making code-output-prediction tasks (the initial family of verifiable tasks we started with) which other reasoning models can't solve.
We started looking into harder / more agentic tasks (e.g. passing tests, using AISI's Inspect framework) but deprioritised.
You have to be careful about degenerate / duplicate Qs, as a sibling commenter mentioned.
Recently though, we found that reasoning models have trouble making code-output-prediction tasks (the initial family of verifiable tasks we started with) which other reasoning models can't solve.
We started looking into harder / more agentic tasks (e.g. passing tests, using AISI's Inspect framework) but deprioritised.