Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

How?


Same way you distill any model. Training data efficiency matters only while you train the source model/ensemble. Once you have that you are purely compute bound during distillation.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: