Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
nodja
12 months ago
|
parent
|
context
|
favorite
| on:
Ask HN: How can ChatGPT serve 700M users when I ca...
Loading here refers to loading from VRAM to the GPUs core cache, loading from VRAM is extremely slow in terms of GPU time that GPU cores end up idle most of the time just waiting for more data to come in.
frabcus
11 months ago
[–]
Thanks, got it! Think I need a deeper article on this - as comment below says you'd then need to load the request specific state in instead.
Consider applying for YC's Fall 2026 batch!
Applications
are open till July 27.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: