Hacker Newsnew | past | comments | ask | show | jobs | submit | Readerium's commentslogin


I have been using it today the whole day and it is definitely better than Luna. Great that bench agrees.

Can you tell something about your tasks? I am pondering both models for cost saving.

Yeah DeepSeek V4.1 beats this by 15 percent (4 points) on terminal bench 4.0

I don't want to be pessimistic, but I can see a single app now showing two different feeds to doomscroll.

Subway surfers and tiktok at the same time. The future is now chud.

Doomscrolls on the left and doomscrolls comments on the right.

Instagram posts on one side and reels on the other side

Perhaps they'll dedicate one screen to advertising so that you can't scroll past them? The horror

Question is what does that button do.

I bet a lot of lawyers are salivating at this question too.


It's not as simple. All our chats are being used by both labs for their future product (unless signed by ZDR). Where should the acknowledgement begin? Who should be acknowledged? The whole world? All the 2B users of AI?

If I know person A is working on problem B.

I am free to work on problem B too. Why should person A be limited to working on it.


Are you free to intercept person A's emails / hack their computer to find their notes on how they're approaching problem B?

Finding Codex session data in the training set that you tie back to these two researchers is like an hour-long task.

You're also free to plagiarise anyone you want. There are no laws against it on most jurisdictions.

Also brain raping* is not illegal in most jurisdictions.

But they're both deeply disturbing.

_________

* https://youtu.be/JlwwVuSUUfc?si=uWl4-LCHAeI7qtb3


But then who is the third person in the call.


It loses to Muse Spark 1.3? Does anyone really believe this index reflects reality?

I'm surprised you feel like you know muse spark 1.3 performance well enough to question the validity of the index based on this benchmark result.

Muse spark 1.3 was only released yesterday.


more like 5.7 not 6

It’s based on a new and larger pre train if I understand correctly, hence the major version.

5.6.1

Its 62 percent when using a neutral harness. https://arcprize.org/blog/astra

Their neutral harness is not very good though, if I read it right, it doesn't preserve the reasoning state between turns. No real harness discards reasoning state like that.

Yeah that must be it. OpenAI doesn't want to disclose internal reasoning, that's why thats typically encrypted_content in OpenAI codex session ledgers etc.; leveraging responses API preserves reasoning server side all the way till a final answer is made; so that's very impressive and to me the score that matters.

Yet it is an impressive number. But yeah when you see a number 99 you have doubts. Thanks for the link

Interesting both this and Sol got approximately a 37% boost with the custom harness.

saturated before (higher degree) AGI-2

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: