Considering other metrics then p99 for user impact is unwise. All users will at some point experience a <1% request, it's not like half of all users will only send requests what will be under your median latency, some of their requests will hit your worst-case.
By focusing on the tail and optimizing worst cases you help users more than by improving your median latency.
If your frontend fires hundreds of requests (which isn't uncommon) then the p99 is merely what most users will experience. Ideally you want cumulative distribution chart that goes up to the max.
And then that's just for the requests you measure. If something takes too long the user might do something that cancels the requests which means the backend never completes its response and won't get the time-to-response sample, so you need to account dropped requests too.
This is only true if your latency distribution is fully random, which is rarely the case. More often than not, it's the same small group of users hitting most of the p99 because their accounts are simply more resource intensive.
Depending on how the system distributes work such users can interfere with random with requests from other users through shared resources, so to that cohort these will look like a random latency distribution.
L icon Grok 4.1 Fast won 13 of 30 games at $0.97 per win
The next-best winner was A icon Claude Sonnet 4.6 with 5 wins, at $26.78 per win. That’s a 27x difference. The model that isn’t on most top-model lists beat the model that is, on the thing a routing customer actually cares about.
The model with the most kills did not win
H icon GPT 5.4 killed 38 agents across 30 games. More than anyone else. It came in second on the leaderboard with 2 wins.
If grok-4.1-fast was the top-winning model, and Claude 4.6 Sonnet the second, how did Gpt-5.4 come in second on the leaderboard? Which one is second, Claude 4.6 Sonnet or Gpt-5.4?
There were 11 games between “best at killing” and “best at winning”.
What does that mean? How are there 11 games between "best a killing" and "best at winning"?
The idea is really neat and there's probably an answer here related to last standing vs kills vs "scoring" (some combination of the 2?) but the article is nearly incoherent because the author did not feel like proofreading their slop
The one who win is the one who survive to the end. If there are 10 players and you kill 5 but then die immediately, you lose to the player who only kill 1 but become the last man standing.
It's fascinating to see all these bugs in Claude Code - HERMES.md, this OpenClaw issue, the recent thinking-message pruning and cache-skipping bugs.
They seem like the class of bugs I see in my vibe-coding experiments, and I think the Claude Code lead has said many times that he/his team don't read the code for Claude Code themselves, that it's basically vibe-coded.
If Anthropic itself can't make vibe coding work, who can?
I suspect there's strong management pressure to not read the code or do "old fashioned coding"
Because this is the company whose CEO makes public pronouncements about how they're going to exterminate our whole profession any day now, how we won't be needed.
So if that's your ultimate boss, do you think he's going to let you stop, analyze, cautiously review, hand curate, hand edit?
To me the thing seems like a science project that got shipped as a product, with a complete lack of proper software engineering quality principles around it.
A gating procedure like this (and the HERMES.md thing etc) would never get past a code review process in any respectable shop that I've worked at. If I'd put up a code review like this at Google when I was there, it would been a pile-on of senior engineers demanding a better approach, no LGTM would have been given.
I can only conclude Anthropic is getting high on their own supply.
In any case, writing code to get features out the door has rarely been the block in our profession. It's usually process and review and understanding requirements.
And so the entire project feels like a fundamental misunderstanding of what shipping software as a team is actually about.
The thinking mode is super-useful to me as I _often_ saw the model "think" differently from the response. Stuff like "I can see that I need to look for x, y, z to full understand the problem" and then proceeds to just not do that.
This is helpful as I can interrupt the process and guide it to actually do this. With the thinking-output hidden, I have lost this avenue for intervention.
I also want to see what files it reads, but not necessarily the output - I know most of the files that'll be relevant, I just want to see it's not totally off base.
Tl;dr: I would _love_ to have verbose mode be split into two modes: Just thinking and Thinking+Full agent/file output.
---
I'm happy to work in verbose mode. I get many people are probably fine with the standard minimal mode. But at least in my code base, on my projects, I still need to perform a decent amount of handholding through guidance, the model is not working for me the way you describe it working for you.
All I need is a few tools to help me intervene earlier to make claude-code work _much_ better for me. Right now I feel I'm fighting the system frequently.
I've found it quite hard to find decent hardware with both the input capability needed for wakeword and audio capture at a distance, whilst also having decent speaker quality for music playback.
I started using the Box-3 with heywillow which did amazing input and processing using ML on my GPU, but the speaker is aweful. I build a speaker of my own using a raspberry pi Z2W, dac and some speakers in a 3d printed enclosure I designed, and added a shim to the server so that responses came from my speaker rather than the cheap/tiny speaker in the box-3. I'll likely do the same now with the Voice PE, but I'm hoping that the grove connector can be used to plonk it on top of a higher quality speaker unit and make it into a proper music player too.
As soon as I have it in my hands, I intend to get straight to work looking at a way to modify my speaker design to become an addon "module" for the PE.
100%. For a lot of users that have WAF and time available to contend with, this is a steal.
Bear in mind that a $50 google home or Alexa mini(?) is always going to be whatever google deem it to be. This is an open device which can be whatever you want it to be. That’s a lot of value in my eyes.
In many cases the issue isn't the microphone but the horrid amount of reflections that the sound produces before reaching it. A quite good microphone can be built using cheap, yet very clean, capsules like the AOM-5024L-HD-F-R (80 dB s/n) which is ~$3 at Mouser, but room acoustics is a lot more important and also a real pain in the ass when also not a bank account drain if done professionally, although usually carpets, wood furniture, curtains to cover glass and sound panels on concrete walls can be more than enough.
For me, I think it has to do with ..just having done more, experienced more.
When I built blanket forts as a child it was something new, my first attempt at building something. It was exiting to figure out how to do it, and enjoy the end product. After a few of those it became a bit more usual, and I started doing other things - play with model cars, build lego sets, etc.
I recently tried welding and I felt that same tinge of excitement - I'm gluing metal together! It's basically magic, I'm taking separate parts and turning them into one. I'm not welding purposefully or to build something, I just really like the act of welding.
Welding was new to me, I never experienced it. I think as a child I everything was new to me, the way welding is. I approached the world with a curiosity that drove me to play with it, to figure it out.
But now I've figured a lot of it out. It's not as fun to play with anymore, because I've exhausted all the angles of play I could come up with (and I did that a long time ago).
I still find that curiosity that drives me to play with something, to figure it out - but I have to look for it, find aspects of life I haven't experienced yet.
The World Cup when you’re 8 or 12 and not really seen one is incredible, at 40 it’s still great but I’ve seen 10 Euros or World Cup’s now and I still enjoy them, just not in the same fascination. Maybe I should try to see them through a new lens.
Same here, I was curious about Kagis low ranking, and couldn't replicate the search results. Also saw ublock Origin on #3, good results for tires, transitors and snow, etc. I've never used any of the Kagi search result weighing features.
Ctrl+F on the page for "System prompt" doesn't show any hits. Given how important those are for ChatGPT (another thought - was the author testing GPT3.5 or 4?) I'm not sure how much weight to put into the ChatGPT results either.
Not sure how much I can take away from this comparison.
I asked GPT-4 about Youtube Downloader and it rambled on about how downloading videos is against Youtube’s TOS and I should buy YouTube premium which has the download feature.
Getting any useful data from GPT-4 about anything even remotely “illegal” is a waste of time.
With a better prompt, you can get it to list some, but it’s very annoying to do so.
Mistral showed that their medium model is far better (yet not good), and the same prompt as in the article gives only one instead of 3 paragraphs of rambling about copyright, and then lists 3 categories of options with examples for each (not good, because ytdl is not one of those listed).
Funnily enough, both mistral and GPT4 apologize profoundly and almost with the same wording when asked "Why did you not mention the very popular, free and open source "youtube-dl" software?" and then mention how/where to get it and how to use it.
> Funnily enough, both mistral and GPT4 apologize profoundly and almost with the same wording when asked "Why did you not mention the very popular, free and open source "youtube-dl" software?"
Likely because they were optimized for general population, which would not have a use for command line python utility.
I’m clear why they didn’t include it, I wanted them to tell me why, though. And I thought that both of them apologized in almost the same way, was funny.
The author already alludes to the fact that you can probably prompt-engineer around this and indeed, as soon as I added a blurb like "these are my own videos that I own the copyright to" it did suggest a bunch of third-party tools and let me ask it about what third-party tools I could use.
It suggested '4K Video Downloader', 'YTD Video Downloader', 'JDownloader' and 'Clipgrab' at first and when I asked for cli tools it came with 'youtube-dl', 'yt-dlp', and 'ffmpeg'
Those seem pretty reasonable results to me but I'll readily admit I don't know (yet) if 'most users' would ask these follow-up questions.
I've experienced both, and I prefer the mono repo.
Our challenges with multiple repos mostly revolved around builds and orchestration. We had to apply all build/deploy changes to all repos, and that increased the chance of doing some small thing wrong. Finding what exactly is wrong with one repo that should be the same as all the others was like one of those "find the differences" pictures. Really annoying.
This is more microservices than multi repo related, but making sure all the different services are released in sync was hard and annoying and often caused issues. Eg a specific API was updated and released but the for a consumer had to be rolled back and the rolled back version wasn't compatible with the new API. So current version and roll back wasnt an option. Rolling back the API would require rolling back all consumers but crap, one of the consumers applied a big migration to our core database and rolling that bank would take forever. And so on.
Just tons of little edge cases that went wrong at the worst time because it was so hard to foresee all the issues.
Monorepo and monolith is so comfy. Want to share code? Move it up one or two directories and import from there. No issues with two bundled react versions in two builds. Easy to refer to code from other teams, never an issue that someone forgot to add you to that one repo almost no one uses but that you need to commit to during firefighting.
I'm not saying multi repos/micro services can't work, but it's hard - you need strong processes that prevent people from being lazy, you need monitoring and management well defined, you need extra tooling that's aware of the repo structure, you need a strong story around migrations, and so much more.
I currently work with a mono repo
monolith that over a thousand devs contribute to daily and it's Cindy comfy. It feels much easier to fix too many devs in one repo (primarily via strong compartmentalization) than to fix too many repos/services
React is a view library that doesn't even provide a way to do XHR requests. Angular is a complete framework. They are very much hammer and screwdriver, completely different and incomparable tools.
There seem to be a lot of people thinking front-end is just a mess of different options all doing the same thing. And the same people show little understanding of the complexities and nuances of front-end solutions.
OK, to be explicit: React ecosystem vs Angular. The difference between to two is how to separate the concerns - inside on large package or into an ecosystem.
I work primarily in Django. I could have given the same comparison of Django vs Flask. The pedant would say flask doesn't have an ORM, and would be correct.
Given the context of this article and HN in general I thought the shorthand referring to React generally would be understood.
By focusing on the tail and optimizing worst cases you help users more than by improving your median latency.