Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Interesting, how would that work? Are there any well-known examples?

Is it: the weights all happen to be where float is sparse, so quantization ends up increasing fidelity? Or is it more of a “worse is better” dropout-type situation?



I suspect it works as regularisation of the network. It usually happens when you train with quantisation instead of post-training quantisation, an I haven't seen that done with LLMs yet.


For image recognition it can sometimes be like that. My gut feeling is that lowering from fp32 to fp16 can get rid of some kind of overfitting or so.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: