atm just trying to collect them all. the problem itself is super non-trivial imho, and my feel is that AI empowerment is getting more people to build these systems just because they can. some of the solutions are really interesting, so it could also be a learning resource or we could figure out after some time what happened and if they have really made meaningful impact on those companies
Claude and I have collected here every public mention of “internal agents” which have been popping up a lot recently. This was quite interesting for me to understand how different or similar these approaches are. I have dumped all that I’ve stumbled on x and here but I have prob missed some. I had run deep researchers few times with less luck, as mostly these things are still not well documented publicly. I feel the ownership of agents would become tremendously important in the coming years.
Idk, spending 10x less time dictating a thought to a coding agent and publishing it immediately is incredibly valuable. Take it at face value and don’t overthink the style.
Yes, this is frustrating, but it doesn’t occur in CC. I run the conversation logs through an agent and opencode source, and it identified an issue in the reasoning implementation of opencode for Zai models. Consequently, I ceased my research and opted to use CC instead.
This wqw an interesting postmortem for me. As a regular user, February definitely felt a bit shakier than usual.
I ran into the unicorn error page a few times and had intermittent issues with PR pages and bunch of weirdness. Nothing long-lived, but enough transient failures that my overall feel was it is degrading in general.
was listening to music while coming with groceries and simultaneously juggling stuff to open the doors and change the track with Siri (the only use for Siri I have)
FWIW I work at Steel (not the OP). While we’ve been iterating on the “right shape” for agent tooling, I’ve been building a benchmark harness to measure how different surfaces affect real web task completion: raw API context, CLI-only, opinionated “skills” (structured outputs + artifact capture), and combinations.
If you’ve run agents on the open web, I’d love suggestions for nasty-but-representative workflows to include in the benchmark.
This rings true, as I’ve noticed that with every new model update, I’m leaving behind full workflows I’ve built. The article is really great, and I do admire the system, even if it is overengineered in places, but it already reads like last quarter’s workflow. Now letting Codex 5.3 xhigh chug for 30 minutes on my super long dictated prompt seems to do the trick. And I’m hearing 5.4 is meaningfully better model. Also for fully autonomous scaffolding of new projects towards the first prototype I have my own version of a very simple Ralph loop that gets feed gpt-pro super spec file.