Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I think what you describe ("confidence about the next token") is the entropy of the model's output. A model can be very certain about the next token (its output has low entropy) but if it is usually wrong on the text you measure it against, it will have high perplexity. (For example when the model was trained only on children's books and you measure it on Wikipedia.)


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: