Hacker Newsnew | past | comments | ask | show | jobs | submit | j_maffe's commentslogin

I think if you use an LLM just to proofread then it'll not be able to insert a strong enough watermark.

The tasks are the thing to really look at here:

https://github.com/harbor-framework/terminal-bench-science/t...


I am hoping someone with more free time than myself can contribute some things in the RF engineering domain in the 'engineering-sciences' section. There's some problems out there that will definitely stump even a smart LLM.

Looks like most things definitely stump even a "smart LLM"... Best score on this is 30%. Which is what you should assume for tasks you give an LLM if they aren't exactly the same as an existing benchmarked task. They're just not that good for the purposes people seem to think they are. Very limited application space.

> I worry this doesn’t check correctness

Then it's not a valid benchmark. I agree though they're not reliable enough to just put results in a paper.


The Purpose of a System is What it Does.

The Word Purpose was Invented Precisely to Distinguish Between What a System Does and What it Ought to Do.

Less catchy, but damn, I hate that slogan.


No but a provider with more amicable terms can.

Anyone has a link to a report of its capabilities? I can't find a reliable source.

completely vibes based, but ive been using it to port Mindustry game from Java to C# with agents, and its been working for 50 hours (its 15-20 tks so super slow inference). Its done a fantastic work and its almost finished now. Better results than deepseek flash and gpt luna by a mile on this kind of long term work. Less good than gpt sol or opus. We dont know the param count but my guess is 200-300 range.

Just curious, what is the motivation for this conversion?

Its free tokens so i left it running for fun as a experiment

63% at DeepSWE.

https://x.com/davis7/status/2091285712566140986

Wenghi is behind DeepSWE, one of the best benchmarks.


likely a distilled glm 5.3 that will punch within 20% of that at 2-3x less size. you'll find that capability is typically very jagged on models that are distilled

Not if there's a ZDR policy.


I think the idea of having a human-written persistent document describing the operation of the code is a great idea. This document acts as the prompting interface instead of the chat window and changes can still be tracked. Surely something as simple as a skill.md can be made for such a setup, right? I think the pseudocode style is a seperate axis to this setup.


They would rather ban to deterr other governments from taking such measures.


> “It’s no secret, I disagree with the prime minister," he said. “I think targeted assassinations should be carried out in Gaza, taking down 30 to 40 every night,” said the far-right Israeli lawmaker.

> “Not just those who pose an immediate threat, there are people there who are not worthy of life. They shouldn’t live. They’re not even people,” added Ben-Gvir.

Not sure how the full quote is supposed to help.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: