Yeah, what even is the point of a "chronometer" (you don't have a clock elsewhere?), "data stream" (???), "radar visualization" (???), "depth field" (???)
Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked?
As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in performance should be taken with a grain of salt.
A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?
Devin CLI is easily the worst-in-class coding harness I've ever been subjected to using, rife with bugs (up to and including dropping answers to the question tool), so I can't say I'm exactly inspired to try anything else the company produces.
Can't say I've tried the CLI. I've mostly focused on the cloud agents, which I was explicitly looking for. I compared against cursor's cloud agents, ampcode, and hoplite, and came out surprisingly enjoying devin.
I will say that the lack of parity between Devin cloud and Devin desktop is downright embarrassing. It's very clear that the latter is a thinly-reskinned Windsurf. A visually similar UI with vastly different capabilities. Definitely a black mark on the whole thing.
Had the same experience on it when it was going by Windsurf. Had to wire up a skill hooked to terminal runs or else it would hang or freeze and never finish literally every single time
I see issues with other harnesses too but not with the regularity I was getting from this. And the one moat they had with the better UI for per-project multi-agent orch disappeared and now is standardized
Not all that surprising. The original Devin really was just an early attempt at agentic coding before models were really even trained for it. Now that it's a well established pattern and we've figured out what works, I'm not surprised they've morphed into something reasonble.
When they launched Devin it was supposedly at the performance of an engineering intern. Friends who used it found the bad parts of an intern (tons of handholding, review required) but it didn’t learn from mistakes or add throughput.
I think the evolution of the harness and ability to preserve loop context outside the context window has made running these kinds of agentic experiences easier.
sorry so many buzzwords to say, the capabilities to do this kind of work are more accessible and easier to manage, so now it works!
Good to see, and agree they were severely overhyping their product back then.
As others have mentioned, it's really matured a lot and at this point is one of the best cloud-hosted, team-managed coding agents, when factoring overall UX, testing and QA lifecycle via its sandboxes, and its ability to be controlled with an API. We use it quite heavily.
It's coming from a different starting place than Claude Code or Codex are as individually controlled single-developer tools. Devin has been more persistent in pursuing the direction of something that operates more autonomously at the team level, as a peer. And while it might be slightly behind in raw harness ability (maybe?) it's probably ahead on the team-focus.
Yes, perhaps fair. But my point was somewhat narrow. I wasn't saying that team-managed agents are good to go for all cases and that they're better than individual dev-managed ones. Just that they have gotten better and that of those Devin has some of the better UX.
Our experience might also not be typical because we have built infrastructure around making Devin and similar agents work better. And for the record no ties to Devin/Cognition. Just pay them too much as a customer.
It changes the economics of the attack because it requires the attacker to do compute before they can make requests. Depending on how good the fingerprinting is, it can also get very expensive (eg. requiring you to run a full browser, rather than merely computing a few sha256 hashes)
Modern phones are all phablets and it sucks but it's the only format they offer. Smaller phones are nicer to hold and nicer to use for everyday tasks. My thumb can actually reach things. At least AI will no doubt reduce the need for pressing buttons.
Manufacturers want phones to be thin so the only way to fit in the battery life is for the whole phone to be larger. In case you didn't notice it's really really difficult to find a side on photo of the iPhone Duo closed. They don't want to show the thickness.
Note that every time a company comes up with a choice between a smaller phone and a larger phone (up to some limit when you get into phablet territory), the larger phone sells in much, much higher numbers. Apple themselves tried the Mini, it sold at best a few percent of how well the regular phones sell.
While you, personally, might prefer smaller phones, and that's a perfectly valid preference, you are very much in the minority.
The original "phablet" was the Samsung Galaxy Note, considered comically large. At 147mm x 83mm, it's about the same size as the iPhone 18 Pro (150mm x 72mm)
But, and this is probably what you meant, the iPhone 17 has a bigger screen. That's because recent phones got rid of the bezels, so they can fit more screen in the same body. That's why they have taller aspect ratios.
Uh, yes? The first widespread usage of the word "phablet" I can recall was regarding the HTC HD2, which had a whopping 4.3" display and some people were losing their minds saying nobody would want to carry that in their pockets.
Current trend nowadays is watching tiktok, though thankfully I've only seen it at red lights, but that might just be because I'm not staring at other drivers when driving around at 40mph.
I was behind someone just the other week that was actively scrolling tiktok/reels while driving on a 3-lane road during rush hour with what appeared to be at least 2 kids in the back.
So what does this mean for Timmy (16 years old) learning how to drive? Should that also be banned, lest he runs over some kid? Or does he get a pass because he's a human and not a faceless corporation? Maybe he can repay the favor in 60 years when he gets run over by another teenager learning how to drive?
I do want safer vehicles and I do see a path to get there with this technology.
What I don't want is companies training their technologies with legitimate risk to the public, with a clear pathway to profiting off of this risk and public data, with no benefit to the public _unless_ they happen to succeed, and to some degree, unless you can afford their products. Something feels off about that.
Of course, everyone benefits if vehicles are safer, but it does feel like there is an extractive factor regardless. These companies benefit from public resources, are wholly enabled by public resources, but are trying to create a situation in which the technologies they create are more or less required to be used in order to utilize the public resource. Does that make sense? I suppose it's inferred, but, I don't think it's unreasonable to believe any car manufacturer with a working FSD technology would LOVE to have that technology mandated to be on every new vehicle. Imagine leasing that to every manufacturer who didn't develop it themselves so they could continue to sell in the USA.
On one hand, safer vehicles are a MUST because we allow an absurd number of people to die due to unsafe driving every day. On the other hand, I don't feel right about private companies benefiting so much from roads the public built, while training their technology with some degree of risk imposed on the public, and with no clear desire in these companies to give back. They tend to pay paltry sums in taxes, if any. Rivian was at a net loss of over 4 billion for 2024. I understand they're trying to build something that requires massive capital, but, I also understand that these companies are experts at operating at a loss for prolonged periods of time. Amazon is a stellar example.
This definitely cuts both ways, but I'm not sure exactly where I stand. I need more information to decide, but I'm confident that there is some degree of exploitation of public data and resources here. I'm not trying to imply Tesla, Waymo, or Rivian are the only ones guilty of this either. The tragedy of the commons is a well-known phrase because it's such a typical phenomenon. I'm also not suggesting that there's a simple solution. There can't be because the fragmentation of the polities in which these companies operate and train their technology within.
Yes, he would be banned from driving if he ran over a kid. Self-driving models and the companies that create them have not been banned so far, even after they cause accidents.
>Yes, he would be banned from driving if he ran over a kid.
Realistically no, unless he was hilariously reckless.
> SMITH: There’s the case that happened with the little boy on the Upper West Side, Cooper Stock. He and his dad were crossing the street. And a driver was making a turn, and he just ran over the little boy, didn’t see him.
> SMITH: So right now all that is is a summons to the driver for failing to yield. But it does not rise to the level of any kind of manslaughter or homicide charge. There was a study that showed that between 2008 and 2012 there were something like almost 1,300 fatal crashes in New York, and there were like 66 drivers arrested.
>Smith says that New York has some of the narrowest standards for conviction in the country. It’s called the “rule of two” — you need two significant violations of traffic laws in order to bring a charge, including some incredibly reckless or criminally negligent act. Otherwise, it’s just … an accident.
And you had that one example to prove it. That settles it.
I know you want to make your point but stacking up fallacies (same in your "Timmy" comment) isn't helping [1]. Most humans grow up having personal responsibility, then at adult age the legal system puts even more on them. A lot more, on them personally, under threat of deprivation of freedom or monetary loss from their own pockets, or in the case of driving even something as small as losing the license. There are exceptions ("daddy pays" sort of thing).
Now contrast that with corporations, especially large ones, where almost every bit of responsibility is diffused behind a corporate veil that allows every single human to sidestep any personal responsibility whatsoever. Some shareholders' money take the plunge. And starting recently you can offload even more responsibility to the computer/AI.
I might lose my license if I cross a red light. Will a company lose the license to operate all autonomous cars when one crosses a red light?
[1]
> My anti-terrorism AI drone escaped the sandbox and shot up gruez's house. The drone was punished by removing its livery. What else can be done? What next, should we punish soldiers fighting terrorists too?"
> I tested my anti-cancer medication in the drinking water of the neighborhood. gruez got hit the hardest by the side effects, God rest his soul. But what can we do, ban cancer research?
When they say "narrowest", they mean that New York is one of the most difficult places to convict a driver. Right above what you quoted:
> Our neighbors have different vehicular laws than we do. Both Massachusetts and Connecticut have vehicular manslaughter statutes that punish traffic fatalities or serious injury that occurs because of simple negligence. New Jersey doesn’t have the same statute as Massachusetts or Connecticut, but even they have vehicular manslaughter statutes that encompass more behavior than what New York has, which is absolutely nothing other than drunk driving. So around the country — in Iowa, Louisiana, Georgia, Nevada, Kansas, California — all over the country, there are states with vehicular statutes that punish a failure to yield as a traffic fatality, that punish the driver.
So this seems to be a peculiarity of New York, I guess. Also, this is from 12 years ago on a podcast that has been widely criticized for inaccuracy and cherry-picking. I'd be curious if this was even true at the time, and if so, if it's still true or if the laws have changed.
Obviously local laws differ and I'm sure there are many exceptions, but it is very common to have your license suspended if you run over someone with your car.
Self-driving cars also operate at a far larger scale than individual human drivers. Over 300 people died in Boeing 737 crashes, but the entire aviation industry has not been shut down as a result.
the 737-max was grounded for almost two years, Boeing paid a $2bn+ fine and the CEO was fired.
To be fair, actual self-driving car business units have suffered similar fates. Uber's autonomy team was shut down after they killed Elaine Herzberg, and GM's Cruise was functionally destroyed after they contributed to a major bodily injury crash and tried to cover up their errors.
It's the level 2 consumer car companies that have mostly skated under a liability shield. Tesla FSD has been shit for a long time and paid very little price for it.
Civilized countries have the concept of a driving school, where you have to take 20 one hour lessons with an instructor who can override everything as well as theory lessons, followed by an independent state run exam, theory and practice.
First, our current city driving environments were designed for being used by humans, so take into account human strengths and weaknesses.
Second, that kid has spent 16 years in such an environment learning how humans function there. He's got a very good model of how things work.
By the time I was 16 year old Timmy I'd spent a whole lot of time walking next to roads with cars, biking next to or in those roads with cars, and riding in cars on those roads. What was new when I was the driver was mostly learning the physics of cars and driving.
(I suppose nowadays, with the way parents can get in trouble letting kids out of their sight, we might have to worry about a generation of 16 year old Timmys who have not spent any significant time walking next to or biking next to or in roads with cars).
Third, that kid will have a pretty good model of predicting future behavior of other drivers, pedestrians, and cyclists based on figuring out their intent, which can even let them make some pretty good predictions about them even when they are hidden or masked from the kid's sensors.
Fourth, Timmy has a bunch of knowledge learned from listening to other people, reading, movies, etc., that can be applied while driving. For example suppose they live somewhere where there are many cows, but no deer. If they see a cow near the road at a place with no fence, they will have an appropriate model from experience that tells them cows are generally slow moving. Be cautious when passing, maybe giving it some extra space.
If the see a deer even though they have no experience with them they likely know that deer often bolt rapidly into the road right when a car gets close due to hearing stories from people who live in deer heavy regions or from news, books, or TV/movies, so go slow and be very cautious passing that.
Autonomous cars have a big advantage in reaction time, so if a bad situation starts to unfold they are likely better at saving the day than Timmy (or an experienced human driver) would be, but at least the systems so far deployed so far don't seem to be up to Timmy's understanding of other actors and their intent. They also seem unlikely currently to have that widespread non-driving knowledge that can be called upon to understand novel situation involving things that was not in their driver training.
>which could be in the details of how care is provided and its quality
This has "true communism has never been tried" vibes. When enacting a policy, you can't guarantee it will come out optimally, so it makes sense to factor in any potential negative effects, even you think you're not going to make them. After all, I doubt the people campaigning for childcare in quebec were expecting the results to be negative.
I've not read the article, and I expect it goes into this kind of thing, but...
Perhaps it's something like: families that earn enough to have one parent stay at home don't use child care. Therefore children in child care are more likely to be from less affluent families than those who are not in child care. So it's not the child care that causes the difference in outcomes, but just general income scale.
Speaking from my own Western Europe bubble, its the reverse. High-income individuals tend to partner with high-income individuals, both heavily invested in their career, and using more child care to facilitate 2 full time jobs. Parents using less child care tend to be less career focused part timers. Having one parent stay at home is a rare sight that I associate with a low household income.
>Perhaps it's something like: families that earn enough to have one parent stay at home don't use child care. Therefore children in child care are more likely to be from less affluent families than those who are not in child care. So it's not the child care that causes the difference in outcomes, but just general income scale.
No, the OP speculated a variety of reasons, which includes what you've suggested, but I'm specifically focusing on the "which could be in the details of how care is provided and its quality", which is why I quoted that part specifically.
Affordability has little to do with it. The reality is neither parent actually wants to stay at home and watch and raise their own children. So they ship their babies off to full time daycare centers for someone else to do it. Then they both get to go to work instead!
>GrapheneOS is also saying some suspicious stuff on Mastodon in which they are failing the vibe check.
Using "failing the vibe check" as your first argument fails my vibe check.
>They are suddenly very AI obsessed, which is the antithesis of their platform.
1. If you're talking about their post defending the use of AI in development, and that the reviews are still done by humans, that hardly seems "very AI obsessed" to me, unless you subscribe to r/antiai or something.
2. How is the use AI "the antithesis of their platform"?
Presumably whoever is doing the acquisition would look at the financial numbers in addition to the subscriber numbers? Then they'll notice a discrepancy between subscriber count and how much $ is coming in, which makes this whole thing unravel. It's far more likely it's just run of the mill incompetence and/or change of leadership prompting an "efficiency" drive.
You’d think, wouldn’t you? But there are so many promotions and discounts and bundles going on that it’s never a linear relationship. So no one would be able to easily tell the numbers were diverging for at least a few quarters. And as long as they are both headed up, no one actually cares anyway. But if you wanted to juice the growth rate or slow the churn rate just a bit to look better or hit some target that you were just shy of, it’d be easy enough to add a “bug” to forget to bill some certain percentage of trial accounts.
Or maybe they just suck at the coding that handles trial accounts. As rapidly as this stuff was rolled out, and as late as the decisions about actual promotions probably get made, I’d be surprised if it wasn’t buggy.
Yeah, there is no way to back out consumers vs. revenue at basically any big company. They all report "ARPU" or average revenue per user, because there is always a spectrum of pricing, almost never a single price. Between bundle pricing with partners, promo pricing offered for new customers, customers on old plans that haven't upgraded or updates, depending on the type of service and how they manage it there can be hundreds of price points in play for any given customer
That isn't to say there's no way to audit for this kind of situation, they could definitely report out what every single customer is paying and just find the ones that = zero
With the idea being that, like as happened with Frank, multiple of the top-most members Paramount’s leadership will end up in jail for a number of years, at least once there is a neutral admin? It looks like each of Frank’s leaders were sentenced to about 7.5 years in prison each on the 175m deal - do sentence generally scale?
The issue is that "whoever" isn't a single homogeneous entity. I would presume that someone in the organization spotted the discrepancy. Having worked closely with (and sometimes for) large corporations, I cannot presume that C-Suite decision makers were sufficiently aware of the issue. Or perhaps they didn't weigh the issue properly. Or maybe they were consumed by a desire to complete the merger, no matter what. There are quite a few ways this could have gone, internally, and still arrived at this place.
Anecdotally, I've worked at places where the due diligence process wasn't so much about determining if a course of action is good, but justifying a decision that was already made. If the subscriber count and income don't mesh, just use whichever number makes the merger look better.
reply