Hacker Newsnew | past | comments | ask | show | jobs | submit | daveyoung's commentslogin

likely a distilled glm 5.3 that will punch within 20% of that at 2-3x less size. you'll find that capability is typically very jagged on models that are distilled


Two potentials from my pov:

1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/opus.

2. They legitimately shipped a new RL checkpoint over the 7 days, which I find hard to believe.

I am leaning towards 1.


Or 3, they find some bug/regression in their pipeline; maybe they didn't quant parts of a model properly, maybe their inference engine had a bug, maybe some pinned MoE expert wasn't pinned, etc...

That's very plausible to have, identify, and fix in a day; especially when you get community feedback in the wild.


3. Deployment problems unrelated to the weights causing degraded performance


For sure the version accessible from OpenCode had a massive timeout problem the first day or so, which seemed to heavily degrade its task completion rate


2. There was a new checkpoint. Official.


do you have reference to where it was said?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: