Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Its the phase where llama.cpp processeses the input, as opposed to the phase where its generating the response.

Speed is particularly important. The response can be streamed word by word, so it doesn't have to be particularly fast, but slow input processing leads to very noticable latency.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: