Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Yes! As long as you have some criteria to judge the final answer, you can do a kind of "prompt-side RLVR", where you have the model generate prompt changes, try a bunch of different prompts and see which ones improve the results.

You don't necessarily need a bigger model to do this.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: