Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I haven't tried it, but colibri is meant to allow running GLM 5.2 the 744B model in 32GB by using prefetches and streaming from fast SSDs.

https://github.com/JustVugg/colibri



Terrible performance on $20,000 worth of GPUs.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: