Hacker Newsnew | past | comments | ask | show | jobs | submit | gexla's commentslogin

This may be partially why the US dispatched the LA Rams and the SF 49ers. Even though McDonald's has hundreds of US bases in Australia, the corp lacks a rushing and passing game. I don't see the NFL being successful there, but time will tell. Interesting note, in SE Asia, I watch NFL commentators based in the UK. It's strange seeing NFL talk with a British accent. And TIL. I always thought the Hungry Jack's was a Burger King knockoff, but apparently it's actually Burger King with a different name.

Seems they should be making it more clear they are not associated with Herdr. Just as I use a minimal harness (Pi) I appreciate the minimal features of Herdr. Much of the features this adds, I prefer to bring my own. In many cases, I let my agents go and don't much look at Herdr.

The minimal decisions is partly what gains my trust. If Herdr were to come up with something like this, then I would give it more attention. Since it is not associated with Herdr, I'm no more likely to give it attention than any other the other overwhelming AI coding things that hit Github in too short of a period.


Thanks for mention they are not associated w/ Herdr.

I've banned myself from installing any more software for half a year. Otherwise, I'll never get rid of this anxiety caused by all the available options :)


If everything basically rivals Fable, then why is everything still using it for comparison?


Have you spent at least 10 seconds thinking about it or are you asking just out of spite?


Let me grab my calculator and add up the time I have spent reading about model releases since Fable has been released. It seems they all place themselves relative to Fable. I'm sure that time has added up to far greater than 10 seconds. At some point, it ceased to be a meaningful differentiation. This is especially true when I put the model through real usage.


Interesting how the author compares the situation to smoking. Were any of the execs of the tobacco companies asked to step down because smoking was found to be harmful to health? My guess is they simply continued and adapted based on how they could remain a profitable company with continued shareholder satisfaction.


I don't see how it makes sense to use models outside of Grok or Composer in cursor unless you like the tooling to such degree that you're willing to pay out the nose for access to OpenAI and Anthropic to use with that tooling. And Codex is regularly scoring towards the top among harnesses, so you might as well use the best harness available for the given model. I'm happy with Grok and Composer in Cursor. If Cursor wants to add more models, they should host more open models.


The reason it makes sense is because different models are better at different things. Allowing yourself to be locked in to a particular model by these companies is a bad idea and will lead to worse outcomes for yourself and the market as a whole.


Because I pay $60 per month and not using the included API tokens on other models is wasteful.


You’d use it because until recently Cursor had the best setup for cloud agents and visual verification and a bunch of other features outside of the harness.


Is there a good place for harness comparison/scoring?


https://www.harness-bench.ai/leaderboard.html

This provides a method, but the data looks stale and perhaps a bit thin compared to say, Cursor, or even AntiGravity data.


never heard of nanobot. how reliable is this benchmark?


I can’t speak to the veracity of the benchmarks but it appears their methodology is sound. Nanobot has 47k stars, fwiw https://github.com/HKUDS/nanobot

It has been more of an OpenClaw or Hermes alternative than a coding agent like OpenCode or Pi, so it’s likely to do well given less context bloat.


IMO it's not. It's benchmarking GPT 5.4 and Opus 4.6. It's also missing Claude Code... one of the most popular harnesses (the most?)


No Pi, no Aider either.


GPT gets the most out there, but creates interesting nuance that turn an idea into a more creative exercise. Claude is like a more grounded GPT but may miss the nuance. Gemini is my last pick, but does better than the rest for making something clear and understandable. It gets to be exhausting parsing through Claude outputs and breaking it down into something more easily understandable. One way I use to improve this is to ask the LLM to pull upstream ideas from prior work on the topics. This makes me feel better about the possibility of hallucination, gives me alternate places to look, but the model still may pull things out of context or fall over on the interpretation.

Given all that, I can definitely see how Gemini would be preferred. And good for Google, because I would rather my offering be the top choice for the most people rather than better serving a small subset of users.


My understanding was that it was totally an experiment to create something that works like Cursor's Origin or whatever it's called. Just the idea of it is probably not something you would care about for your own usage.


I don't know about real successful. Since you mentioned Langchain, you could look at https://www.langchain.com/dcode which is a CLI harness build off Langchain deep agents.


Yeah, whatever it is, it's particularly bad at front-end from what I have been seeing.


I remember back in the day when you could apply this to Google search results. When I was freelancing, I had a number of people I worked with who would fire off fast questions, but they would still pay me for my time. Then I would return what was essentially a Google search result shortly after. Yes, I would send my own answer, not some paste from Google.

But the point is. They had search, so why didn't they think to do the search? They had the same tools as I did. And yet regularly they couldn't think of the answer. I still deal with this with AI today. When you ask a question, maybe you're asking for the interpretation of the question by the person on the other side. And it's that interpretation that goes to the LLM. Or otherwise some filter of how that person thinks about the question differently than you do.


There's a couple pieces to this I've noticed.

First of all there's a skill to it. You learn how to phrase things, based on how other people probably phrased it. You learn how google modifies or ignores your queries... (or stops respecting modifiers...)

That comes with practice, and lots of it! I google so much that google regularly throws up the captchas (even when I'm logged in). I remember when Kagi first started, my friend sent me a link to their subscription, and it was like an order of magnitude less search than I did. (They increased it a lot since then.)

So it's skill issue, and skill is a function of volume. Most people barely ever google things. Which is the second thing I've noticed.

I'll often be with people, and they'll be like, "huh I wonder XYZ", and then just trail off, as if there was an implied "I guess we'll never know!" at which point I pull out my phone, bewildered, and simply look it up, and everyone looks at me like I'm an alien. (I do not understand socializing apparently.)


Translating a client question to an actual technical question is the hard part I guess. Google-fu is still a thing, and absolutely a thing you were/are getting paid for it seems.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: