Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
chatmasta
74 days ago
|
parent
|
context
|
favorite
| on:
Micro-Agent: Beat Frontier Models with Collaborati...
Looks nice (slop article aside), but why is VSR Hybrid only benchmarked on Humanity’s Last Exam and not the other two benchmarks (LiveCodeBench and GPQA-Diamond)? Is this an oversight or are the results too terrible to show?
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: