Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

How does post-training via reinforcement learning factor in? Does every evaluated judgement count as 'the training data' ?


I guess I'd place both within a broader umbrella: human generated input. So it still holds that they're regurgitating the decisions made by humans.


yes




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: