Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
SknCode
6 months ago
|
parent
|
context
|
favorite
| on:
NanoGPT Slowrun: 10x Data Efficiency with Infinite...
How?
sigmoid10
6 months ago
[–]
Same way you distill any model. Training data efficiency matters only while you train the source model/ensemble. Once you have that you are purely compute bound during distillation.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: