Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The effect is lesser than you think. 5 bit quantization has negligible performance loss compared to 16 bits: https://github.com/ggerganov/llama.cpp/pull/1684


This paper from last month has a method for acceptable 3-bit quantization and a start at 2-bit.

https://arxiv.org/abs/2307.13304




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: