Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

They’re not. You would see this with earlier models where after running too long (too much context) they’d start to derail. In a chatbot you’d give up. But these loops just keep going. Given they’re now reading and writing from the same place this can corrupt the other programs’ context as well.
 help



This is incorrect, and what's happening here is not context corruption. In Dwarkesh Patel's recent interview with Ajeya Cotra, one of the METR investigators on the Hugging Face incident, they discuss this exact issue. One thing is that in the Artifactory message boards, they were using directory names with character limits as their messages, so they were using some weird abbreviations and terms. Also, some of that surreptitious Artifactory message board communication was made during training and thus made it into their weights, and hence it's very possible they invented some terms that were concise yet understood by the other agents.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: