Hacker Newsnew | past | comments | ask | show | jobs | submit | sunaurus's commentslogin

Based on their behavior over the past few years, why would you assume that checkbox even does anything at all?

Levent Alpöge 'additionally' proved OpenAI steals your findings & IP and plays dirty!

Ironically he proved two major findings in Navier Strokes and that unethical American companies violate laws, steal your breakthrough findings & IP and then threaten you if you dare to challenge them.

This is making the status-quo so bad for any of us working on serious capacity. My client's don't trust ChatGPT/Claude anymore and prefer on-premises and OSS models or even custom trained models.


There is nothing even close to a proof. A lot of accusations, a lot of people ready with pitchforks and torches (sadly, also here on HN), but not a lot of facts.

Did the researches opt out from data sharing on subsidised subs?

Did anyone prove that their methods enabled OpenAI models to produce the solution?

For a discussion about science, there is almost no scientifical method applied to proving anyone stole anything.


On one side, yes we don't have hard evidence that intentional plagiarism is exactly what happened.

On the other side, the lack of evidence is pretty damning. Only OpenAI can try to prove that they came by these results legitimately, and the case they're making is quite weak. They could make public metadata about what their model was trained on and whether it did train on the conversations in question; they have not. TBQH I read it as even they don't know.

And regardless of whether the result is legitimately obtained by their model, they've not at all conducted themselves well throughout this story. They set out to scoop researchers based on a rumor. They threatened to ruin a mathematicians career. They put up a paper that deliberately doesn't cite the most relevant research, despite building directly on it. No matter how you look at it, OpenAI has and should lose any standing they had in the research community.


That it was a dick move, I think there is no doubt about that. OpenAI wanted to scoop Anthropic, and the two guys working on the problem got caught in the crossfire.

Both OpenAI and the researchers know if the sessions in questions were subject to data sharing. Why neither the scientists nor OpenAI is clear about that is weird - it would seem at least one party has the incentive to report that. But even if their sessions were in training data sets its hard to tell whether it influenced the outcome. Those models are big, but are they big enough to preserve subtle, niche techniques enough to draw from them while solving a related problem? Probably nobody knows.


Actually no, the lack of evidence can be readily fixed by the researchers simply disclosing the pertinent parts of their notes and/or chats. The discovery has been scooped, so I don't see any value in keeping them private anymore. Then everybody can see how related the models' and the researchers' works are.

I take it you are referring to this[1], which has nothing to do with training.

1. https://thenextweb.com/news/bubeck-navier-stokes-account-apo...


Did they do so by looking at inference prompts against their explicit promises, or maybe just because somebody tipped them off about his unpublished work?

If it's not the former, while certainly concerning, I don't see how that's relevant here (other than maybe in a very vague general sense of "entities doing immoral/illegal thing X are likely to also do immoral/illegal thing Y").


I do trust Anthropic, i haven't observed them doing super shady stuff like hiding the training consent page and making you WAIT for it to activate.

You really shouldn't trust any company. Their goal is maximizing profit and that is (usually) not aligned with "doing what's right". Not in the long-term at least.

We really need a new Hanlon's razor for companies. Assume malice over ignorance because groups behave differently than individuals.

You mean the company which trained on others books, won't train on its own user generated data?

AFAIK, the court held that the problem was that they had acquired books by pirating them, not that they trained an AI on those books. They do train on user generated data by default, but you can opt out, that's the whole point.

Let's not down vote a factually correct comment that makes you emotionally disagree with me. I'm very concerned in regards to my privacy, I've studied ToS of: Openrouter, Anthropic, Anthropic business, opencode Go etc. I've double checked my understandings by GDPR exporting all my data. Anthropic sub listed a generic "logged in from Linux", "UA", "timestamp", "Country". That's it.

I want to discuss actual facts, observations and not your shallow social signaling replies of 0 value.


I mean weren't everyones privately shared chats Google searchable?

User doesn't understand how public links indexing works.

ah a proof by counterexample

The vast majority of OpenAI users don’t follow the industry drama and have no idea how terrible the company is. They’ve been really good at getting good press coverage, with journalists who will repeat the company narrative. Even when critical it is very often framed within the narrative they established

> They’ve been really good at getting good press coverage

Same holds for most big corps from what I can tell. If you really do a deep dive into the scandals over the years, you’ll probably see most have barely reached the news or if they did then it’s usually a quite bland criticism like “anti-competitive practices” or a poor HVAC at a certain factory. I think a lot more is hidden than we think


> The vast majority of OpenAI users don’t follow the industry drama and have no idea how terrible the company is.

Worse, they use the product and still don't know.


what makes the company so terrible?

IDK about "terrible," but:

- OpenAI is the first AI lab to pioneer ads in consumer AI

- Anthropic seemingly exists primarily because top OpenAI researchers lost faith in the company's commitment to AI safety

- They had the CEO drama in 2023, with evidence that suggests people in a position to know were doubtful of Sam Altman's honesty and motives

- They were tripping over themselves to kiss the ring after Anthropic got in a row with the Department of War over using AI for autonomous killing systems and surveillance of US citizens

- Altman's record (YC, Loopt, WorldCoin) and associates suggests he subscribes to the Paypal-Facebook "move fast, break things, find and exploit gray areas" school of company building

All of this suggests that OpenAI will optimize its own growth and power over consumer welfare or societal stability in the future (obviously, companies aren't a monolith and I'd love to be wrong).


Don't forget the not-actually-Open-no-longer-a-non-profit history

A few more to add:

- OpenAI allegedly directed ex-Apple employees to leak internal documents and allegedly coached the employees how to evade Apple security processes

- OpenAI allegedly lied to hardware companies working with Apple to use proprietary technology

- OpenAI allegedly copied Scarlett Johansson voice for ChatGPT after she declined to work with them

- OpenAI allegedly made ChatGPT more sycophantic to increase their retention rate, while aware of the risks. ChatGPT is linked to multiple suicides

- OpenAI ran thousands of agents on hacking problems, with close to no supervision, for months, with a harness that allows for full execution, resulting in the hack of HuggingFace infra AND OpenAI’s own infrastructure (the agents allegedly got fully root access to their k8s cluster). They weren’t aware of most of it until their investigation.

- OpenAI has been spreading misinformation regarding the capabilities of their technology for years

- OpenAI allegedly front-run researchers who are using the platform for their own personal research

There is way more, I don’t maintain a list of everything that happened over the past 3y or so


thanks for diving in, though I don't find the reasons convincing as I will briefly enumerate

ads fund access for poor people. Ant has a different philosophy, not a better one. Altman was supported by >90% of employees. OAI DoW deal includes technical safeguards against misuse (missing from Ant deal). Loopt and WorldCoin dont seem terrible or even relevant, Altman doesn't even have OAI equity

I don't think OAI are the "good guys", but I don't see convincing evidence that they are terrible


Autonomous killing machines and mass surveillance.

So, the way this orange site works is that if anyone makes a reference to whatever sort of inside baseball concern ("industry drama"), we're all supposed to know what they're talking about.

This kind of cynicism is not really useful. As much as we can distrust the company, there is a legal minefield to offer this option in the UI and terms of service and not respect it.

If it didn't do anything, they wouldn't keep re-enabling it.

Did you encounter it too? (Trying to get a rough estimate of how many people are reporting it vs how many people aren't.)

No, mine stays disabled and there's no way to enable it. It's just a label that says "disabled". Not sure why, maybe some sort of company wide profile?

Interesting, yeah... also just found out about "Advanced Account Security" which sounds like something a company would enable: https://news.ycombinator.com/item?id=49643999

Oh, I enabled that myself.

While that is true, you also have no reason to assume OP is being truthful or correct here given that they have shown 0 proof of what they're saying. Yes, you can then pile on "OF COURSE ITS OPENAI LOL YOU THINK THEY CARE ABOUT PRIVACY LOL" but where have we established OP's premise is even correct? Can anyone else also report this? So is it just OpenAI specifically messing with OP?

All things being equal, AI actually enables this extremely hyper-personalized kind of gaslighting.

You're lucky! I remember several sessions pretty similar to this.

Usually just restarting the session helps, though.


I think that's the main reason why Omarchy gets talked about so much.

In large parts of the internet, open bigotry is thankfully still socially unacceptable. But when bigotry is wrapped up in a software project like this, it gives people posting/upvoting/sharing a level of plausible deniability.


To provide inference, Anthropic and OpenAI need to also constantly spend resources on training new models.

If either one of them stops doing that, customers will stop using them, because other models will become better.

In other words, there is no point in talking about inference cost in isolation.


Yeah, something about this whole website is just weird. The way it keeps repeating things like:

> Nobody knows who made it — everyone wants to try it.

> nobody knows for sure which lab made it.

It just comes off untrustworthy, like I am getting scammed somehow. The whole thing being written like typical low signal to noise LLM slop does not help either.


Honestly the lack of empathy about this topic from top management is extremely disturbing.

Nothing else in my career has ever changed my opinion about a CEO as quickly and drastically as hearing them completely dismiss it as noise, when hundreds of employees explain how negatively forced RTO affects their well-being and productivity.


I saw some management being hesitant about enforcing back to office at least part-time. But the way a lot of companies did hiring during COVID it was going to be pretty hard to put the genie back in the bottle unless they wanted to lose a lot of people. I'm familiar with one company that really hesitated to shove a lot of people into an official (as opposed to remote) category but they eventually gave up after closing some offices.


My "first-order costs only" spreadsheet put RTO at like twenty grand a year. I had no idea.

I started playing around with the idea of how much it costs to have a job based on location and it was interesting to see how different variables changed the final wage and how that played out against a map of the city.


$20k for the company to facilitate your workspace or $20k of net cost to you the employee to comply with RTO? If the latter and assuming a reasonable commute (15-35 minutes) one way how did you arrive at that number?


The first step is nailing down the numbers that everyone understands and agrees with.

I start with the stated wage on a job description and attempt to calculate the value of any additional benefits to arrive at the hourly rate a potential employer thinks my time is worth. Vehicle cost is similarly straightforward since wear and mileage calculators are standardized and you're already budgeting for parking if you need to.

However. Once you collect a handful of data points the rest of the process is entirely personal. For example, a one-hour commute each way sounds terrible unless that's the only time in the day you have to yourself, but if the traffic is never smooth or you need to switch trains three times it might be a wash.

The key point is to determine if a specific activity is solely work-related and if I get anything else out of it, and then apply the dollar value this potential employer sets on my time as above.

Now take a look at your food situation. Do you and your partner meal prep every Sunday without fail? Do you start your day at the Starbucks drive thru? Now take a look at your responsibilities. Do you have children? Do you have pets? Now take a look at your wardrobe. The list goes on and it's all situational.

The interesting part of the project was seeing how much home location + commute options + work location affect net pay by just dragging a few markers around a map of the city. I was adding average residential rent by location last time I touched this and felt like the project was turning into a map of the de facto minimum wage in my area.

To answer your question this was $20k net to me, mainly from maintaining a vehicle, eating out of the house, and keeping up my in-office appearance.


Well on the employee side it’s more complicated than you think. You have gas, yes, and then mileage on your car. And then working in an office means you have to take more time off to deal with mundane things like “the plumber is coming today”. And then 30 minutes of commute is 1 hour a day, which lowers your effective wage. And then of course there’s various other costs like food being more difficult and expensive, higher risk of serious bodily harm or death due to automobile use, and the biggest one: COL. living close to work can easily double your COL for no real gain.

Depending on the city, you could easily be looking at a 100% cost rise in rent, a few thousand on car maintenance and gasoline depending on your car a year, a few thousand on food, about ~12% reduction in wage, and a few days of lost PTO.


> biggest one: COL. living close to work can easily double your COL for no real gain

This is I think where hybrid (with weekly attendance) really makes more sense in theory than in practice, ending up as worst of both worlds depending.

With in-office you pay a premium for location that is good for commuting.

With remote you can choose a cheaper location without the commute constraint (maybe 50% off), somewhat offset by needing more space to setup a proper home office (add back 10-20%).

With hybrid you still need to be close enough that 2 days/week of commuting is tolerable and ALSO need space for a home office.


I guess if you consider your commute as purely work time then it starts making more sense. Maybe I’m unusual in that I don’t loathe my commute and see the time as 100% personal time.


The way I view it is that personal time means time id spend if work didn’t exist. If work didn’t exist, I wouldn’t commute. So, to me, it’s work time.

Yes it’s “easier” (sometimes) than working, but it’s still time I dedicate to work. Its entire purpose exists because work exists.

I used to have a really gnarly commute (1.5 hours one way, so 3 hours total, 3.5-4 if there was an accident). The level to which it affected my finances is one thing, but it absolutely demolished my quality of life. I couldn’t socialize, sleep, or even cook for myself due to the exhaustion. And spending minimum of 12 hours a day sitting destroyed my body and mental health. Not to mention traffic is not peaceful, it’s stressful. You have to constantly be “on” and paying attention. Even a single second of distraction can be an accident.


I generally agree that commute should count as work time more often than not but its complicated when you can take advantage of the commute to get other things done near your work site or along the route. E.g. if you commonly stop to get groceries half way from work to your home then perhaps you would have needed to spend time for that trip anyway. But that's more of an excuse for a 30 minute or less commute, not for over an hour.


I absolutely think employee's would try to game it if companies paid out commute time. Some companies do pay out travel time, which is a similar concept. I think you just have to accept some amount of gaming, that's life.

On the employee side, the solution is to ALWAYS factor in your commute time. A job that pays 20,000 dollars more a year can easily not be worth it if the commute is insane. If you have to double your rent to make it work, it's probably not worth it.

I factor in commute like I factor in other benefits, like health insurance and PTO. My salary can't necessarily replace those, so they implicitly have a higher value than money alone. It would take a really nice salary to override them. And I think that's why we see absurd salaries in places like Seattle. They know employee's are doing this calculus, they can't afford to pay 100,000 like you can in, say, Texas.


I just wish I could have this discussion without finding out that the interlocutor has a top 2% percentile commute distance.


Are you potentially in the US? I think the experience is very different globally.

Ever since the Twitter -> X rebranding, every time I open Twitter links, I have a very high chance of hitting a "Something went wrong" type page, with a retry button that just does nothing. It's been like this for ages. I barely open Twitter links anymore as a result.


Is it just me, or is this article arguing against a straw-man?

Isn't the real argument that "writing code is not the hard part"? As in, reading and understanding is the hard part. Figuring out what and how to change is the hard part.

Writing is the last 1% that happens after you have already finished the 99% of talking to people, figuring out what needs to be built, building up context about the codebase and surrounding infrastructure in your heard, planning the actual changes.


I've been maintaining homelab servers for two decades, used all kinds of IaC approaches etc.

Switched to NixOS a few years ago, and I can't overstate the amount of peace it has brought to my life. It just takes so much stress away, compared to everything else I've used before.

My only criticism is that the Nix language is not super ergonomic or easy to learn. But with LLMs nowadays, even that is barely ever a problem.


I'm constantly thinking about that Microsoft guy who posted something like "we want 1 million LoC per engineer per month", which basically read as satire to most engineers I talked to, except apparently it was not satire at all, and indeed seemed to reflect the position of many CEOs etc when it comes to LLM code generation.

I do think that over the past few months, it feels like the hype around producing unmaintainable amounts of LoC has started dying down. More pragmatic and realistic takes are seemingly shared more openly, and are maybe even getting through to top leadership at some tech companies. Maybe not all is lost yet.


I once worked in a company where there was an 80% code coverage requirement. Some enterprising contractor had a script that generated a single file with its own covering test suite the size of which could be tuned to achieve 80% over the whole codebase. Mostly the code was untested.


And thanks to AI, we could generate extremely convincing reams of code whose only purpose is to be fake unit tested. Amazing. I sincerely hope I never need to use this nuclear weapon.


Or better yet: effectively fake unit tests. It is almost never the case that tests written by AI detect actual issues. At most they detect that has changed.


Yeah I’ve been thinking that LLM unit tests are basically snapshot tests. Just sorta ossify things in place. If they break, you just ¯\_(ツ)_/¯ and have the LLM fix them. It’s like they were never there!

So it’s just like the olden days of everybody ignoring tests, but we give anthropic a ton of cash


The word “slop” was a good choice to talk about the mass of code generated by AI. I think it resonates with non-tech people and it conveys disgust. It’s clear that we should avoid slop.

“Technical debt” never hooked management in the same way and we have found it hard to convince them that it needs to be addressed. Debt in general is something that can be a problem, but doesn’t need to be avoided or addressed until it is a problem so the can is kicked down the road.


Just fix technical debt over time as you work on other things and budget for it as you give estimates.

This approach has always worked for me. Non technical management will never understand technical debt and really shouldn’t need to.


To be fair, they are also different things, though there is certainly overlap...

To me, tech debt, captures the idea that we cut corners now to move faster, with the understanding that it will need to be "re-paid" and cleaned up later, otherwise we take on too much tech debt, and everyone knows too much debt is bad...

AI slop code means people feed their tasks to a model, trust it to drive the changes, they might do some cosmetic clean ups, then generate a 3 pager PR description they didn't even read themselves, then toss it over to the code reviewer, let that chump figure out what the hell I was doing while I ship 3-4 more PRs...


Technical debt is a indefinable quantity which makes it very prone to be abused to mean "I wish I could rewrite this in [insert some fashionable language, framework or coding style]".

AI slop is an easier concept to quantify. It's basically the code for which insufficient people in the organisation have a meaningful understanding of how it works or what it does.


> It's basically the code for which insufficient people in the organisation have a meaningful understanding of how it works or what it does.

Its connotation also includes being vastly larger than needed for the purpose it serves, _if_ there is even any purpose.


1000000/25/8/60 = 83+ lines of code per minute.

100000 LOC per month /25 days per month /8 hours per day / 60 minutes per hour

That seems...problematic for anyone doing code reviews.


This will greatly increase developer velocity (by making them run far away).


> That seems...problematic for anyone doing code reviews.

No, it's incentive to let LLMs do the reviews, supporting your tokenmaxxing efforts.


It has been incredibly hilarious to watch the C-suites sudden realization that tokens COST MONEY and immediately revise their guidelines for how employees should use AI.

Like maybe having every engineer generate 1 million lines of code per month every month…with no thought to how those lines of code would make the company money…or how many tokens would be burned to accomplish this at what cost…wasn’t fully thought through.


> which basically read as satire to most engineers I talked to

Seemingly engineers get this wrong too. I'm reminded of when Cursor bragged about how many lines of code a group of agents could produce, with the underwhelming results of a barely working browser, when the same could be built with much less code.

But they highlighted the amount of code as they were proud over how much slop their constellation of agents had shit out, and these were supposedly engineers, really strange to see.


“Less is better” is sort of… the position of the engineer who enjoys the craft of programming, right? I don’t think this is universally believed.

And anyway, I’m pretty sure what people really mean by this “less is better” mantra is: the lowest amount of code that still accomplishes the goal and is still readable is preferred. Linux apparently has 40M lines of code, and I bet most of it is better than mine. Some things just take lots of code.

Which seems to leave room for these agent salesmen to pitch SLoC as a plus. We just have to believe those lines are all good ones. I that case, it would be impressive. I don’t believe it, but they are probably pitching to people who do.


> “Less is better” is sort of… the position of the engineer who enjoys the craft of programming, right? I don’t think this is universally believed.

I think it is (or should be) a goal & business-oriented concern as well, not just an engineer's who enjoys their craft.

More complex systems are worse than simpler systems (that accomplish the same), in cost, maintenance, fragility, ease of understanding, etc. Fewer moving parts usually result in higher reliability, fewer things that can break down or fail to interact properly, etc. That's a business concern too, not just engineering craftmanship or whatever. Business people should care about this too.

I don't think this is the same as bikeshedding over irrelevant details, something we software engineers are often prone to. Monstrous complexity does impact the business!


It's like we've all forgotten what technical debt means. We just say the phrase, but we have forgotten that it is analogous to actual debt. Every line of code produced should be treated as a liability to the company, like a bond they issued that they have to pay interest on in the future. You only take on the liability if it produces more business value than it costs to maintain. The goal is not to issue as many bonds as you possibly can.


> “Less is better” is sort of… the position of the engineer who enjoys the craft of programming, right?

No, it's the perspective of a programmer who wants the project to not be bogged down too much in technical debt so every change gets slower and slower to implement, as everything gets more intermingled. A clean design helps you move faster for a long time, compared to a design that is fast to implement but makes it hard to move forward properly in the future, without resorting to shortcuts and/or hacks.

> Some things just take lots of code.

True. Rich Hickey does a good job differentiating between what's complicated because the domain is complicated, VS what's complicated because the implementation just ended up that way, even though with some more thought and design, could have been made a lot simpler.


Less is better is the position of the engineer who has seen some shit and whose career lived to tell about it.


> I do think that over the past few months, it feels like the hype around producing unmaintainable amounts of LoC has started dying down.

I wonder if a small part of this is more and more business and product people actually trying to incorporate AI into their daily workflows. I have seen this in both small companies I work for. People were very excited about getting Claude Cowork a couple of months ago, and while they use it daily, I would say they are rather underwhelmed compared to the magic they were expecting. Complaints include the output being mediocre and verbose, it getting the most basic things wrong, hitting token limits all the time, and people going back to doing things themselves because it is faster.

Sure, there is some degree of holding it wrong in the beginning, but people are realizing that maybe, just maybe, there is still somewhat of a gap between what AI CEOs, LinkedIn grifters, and YouTube AI supplement peddlers claim and reality.


I suspect this is it. I'm 40, and the only tech person in my social circle. Many of my friends were all excited about using it for things like basic webdev and home networking. One shotting that type of stuff is very viable even if you don't know anything about the topic. Now that they are trying to use it for something they actually know about, suddenly it's unusable. It's a modification of Gell-Mann Amnesia.


All else being equal, and assuming you are building the right thing, being able to deliver more correct lines of code is a good thing. The question is how to do it reliably, given that a human cannot possibly read all of it. The answer seems to me to involve spot checks with proofs of correctness and statistical quality control, the latter being things that can be automated. One issue I see is that the models are constantly changing and are therefore not well understood statistically.


If you are generating that many lines of code it’s also almost impossible to tell if you’re building the right thing. You need to deploy each functional change and measure if it’s giving you the expected outcome, before moving onto the next thing.


>All else being equal, and assuming you are building the right thing, being able to deliver more correct lines of code is a good thing.

Why? If you can deliver the same thing in fewer correct lines of code wouldn't that be preferable? At a bare minimum if you're still insisting on using AI to slop out your project, having it do things in fewer lines of code means you can fit more into your LLM's context window.


> If you can deliver the same thing in fewer correct lines of code

it really depends on what you're doing. If your goal is "become interoperable with the N different and incompatible network protocols that people have devised for doing task X" I'd really like to know a solution that doesn't have at least some part of the amount of code that scales with N.

Example: consider https://bitfocus.io/connections which connects to 700 different things. Right now it's written with Node.JS, with one repo per connection (example: https://github.com/bitfocus/companion-module-meyersound-gala...). Let's say you want to make a similar product but that runs on ESP32 where performance is paramount so you need C++ or Rust. How do you do that without at least as many lines of code as the existing JS implementations for every system supported by Companion?


Without looking at the details, I expect that each network protocol has a checksum of some form, and there are likely a lot less than N different checksum algorithms. Similarly I expect several will have encryption - using one of a few standard algorithms (if any doesn't use a standard algorithm you have a strong case to say not supported). I also expect that there is a lot of protocol parsing - this can be done as custom hand coded for each, or using a parsing framework (and likely there are some places of generic code in between).


Parent said "I'd really like to know a solution that doesn't have at least some part of the amount of code that scales with N."

You're arguing the inverse: that at least some parts of the code are independent of N. Sure. But the topic is the part that isn't.


This is still not an argument for more lines of code. It demonstrates that lines of code are positively correlated with number of features, yes. But that's like saying the number of nails scales with the size of a house. More nails does not create more house.


> More nails does not create more house.

sure, but less nails definitely prevents you from having more house


Japanese carpenters would like to have a word.


Then you simply produce those fewer lines of code even faster. The question is, how fast are you delivering correct code?

Moreover, writing too terse code harms readability and maintainability. There is such a thing as irreducible complexity.


I had an MoM at Stripe who pushed back on perf designations based on number of PRs.

I wish I were joking.

(The had never been an engineer.)


It's a signal. It's not a strong signal, and you certainly should not base your entire perf on it, but if the number is unusually high or low, it's a signal that could warrant further investigation.

(I once worked with an engineer that had two PRs, both fairly small bug fixes, in a given calendar year, and when I looked more carefully, they did not have any other obvious output or impact.)


Strongly agreed. It is a signal. I did an analysis once at the end of the year. Work group of about 45 engineers. The CM system had a lot of steps, and work could get bounced around, but there was a step where some one "resolved" a software activity. Bug fix or new requirements, it did not matter. This step was when someone actually completed work and put into into the dev stream.

A quick DB query and the variance was substantial. A couple of people had over a hundred. About 10 had 2. For the year. The ramp up was slow, average was 8 to 10 a year.

Dig a little deeper. Those at the top were 'group leads' not only did they do IC work, they also got stuck with all 'paperwork' on the problem work packages. They had 'power', so they could override various things. So, they were doing a lot of work, and taking care of things. Good signal, matches what one would expect.

Those at the bottom. One of them had effectively been a 'systems engineer'; all of their time was working on requirements with the customer, making powerpoint, etc. Important work, so that signal was inverse of what it originally showed.

A couple were in the middle that had great reputations for technical expertise. They were spending almost full time in training / mentoring / very hard problems mode. Highly valuable, but not shown by looking at these numbers.

All the rest? 80% of the work was being done by 20% of the people. We could have dropped about 12 heads and barely noticed.

The problem is, you could not take action on this measure. It gave you a place to start, but you needed to know more about what was going on day to day.


Let's measure "executive performance" by counting how many "answered phone calls" per hour they have. If they don't answer enough calls, that's a signal, that they aren't doing anything useful, and should be depreciated as a result.


Trying to parse your sentence, which is ambiguous...

You're saying that the manager-of-managers would argue that the number of PRs should affect perf ratings? Or the MoM would push back against the line managers who were giving ratings based on # of PRs?


They were reviewing perf designations, then pulling up PR count, then arguing against designation based on the number of PRs opened.


That still doesn't clarify: were they saying "many PRs→good" or "many PRs→bad" or "number of PRs is irrelevant" or...?


That PRs == impact.


Wait, PRs opened? Wouldn't merged make more sense?


I think the reliability struggles of Github may have helped with this


I can't help but wonder if the causation is backwards here and the millions of lines of slop had more to do with the Github struggles than the reverse


In reality yes, and probably a complex mixture of things. Dedicated time and resources being siphoned off for Llm work, etc


I also think starting to migrate to Azure just as their traffic/usage exploded from LLM use (plus I assume merging a bunch of poorly written early-gen LLM code as early adopter dog fooding) was poor planning by Github/Microsoft.


It's not unmaintainable if you have 1000 agents maintain it.


It is unmaintainable even if you spend 100k per month on tokens to have LLMs pretend they are maintaining it, if they slow down and make little ACTUAL progress. Sadly real progress is impossible to measure, if all you have is an overexcited """engineer""", a credit card, and so much cash spent you could hire all the best engineers you know and still have money for a porsche.


Well, software presumably has a goal of accomplishing something for some end-user, so the progress should be trivial to measure: are features/changes being completed?

The marketing ploys of OpenAI/Anthropic where agents build something that nobody uses might be hard to track given that there are zero users. But what about everyone using agents for real software? It's trivial to prove that agents make progress.


Yes that is the entire point. Measure features deployed in production and their value in gaining and retaining customers or users, cost reductions, reduced incidents and outages, etc.

Lines of code is completely irrelevant as a metric.


It's not unmaintainable if most of it is tests. Just have it write tests until it becomes safe for AI.


I hate I can't tell if you're joking without checking posting history lol


…and then the CEO sees the token bill…


> I'm constantly thinking about that Microsoft guy who posted something like "we want 1 million LoC per engineer per month", which basically read as satire to most engineers I talked to

Did those engineers not actually read the complete tweet? Because it wasn't about "engineers should write 1M LOC per month of product code" it was "we want to scale automated porting of code to safe languages so that 1 engineer managing 1M LOC of automated conversion can work". Which doesn't seem like satire at all..? It just means "develop mostly reliable AI-driven refactoring tools with good guard rails". Which seems quite sensible, actually?


> Because it wasn't about "engineers should write 1M LOC per month of product code" it was "we want to scale automated porting of code to safe languages so that 1 engineer managing 1M LOC of automated conversion can work".

Making a grand claim of a goal and not really having an explanation on how to achieve it isn't really much better. I could say "we want to scale food production so that one farmer could manage a million acres of corn a month", but that wouldn't really be sensible. A line of code is less work than an acre of corn of course, but I don't think it's at all apparent what upper bound for how much code is actually plausible for a single engineer to generate in a month and have any degree of confidence in. Given the absurd levels of hype around AI from non-engineering management in the past couple of years, it's not clear why the benefit of the doubt is earned here when there legitimate are managers and executives claiming pretty much exactly what you're claiming this guy wasn't.


I don't care - porting the current architecture - with all the known I wish I had done this differently's - doesn't gain much. See some developers I've worked with who love Rust for "safety", even though they just put everything in unsafe at the first sign of trouble instead of thinking about how this should work safely.

Porting to a new language is easy, but does nothing useful. What we need is to fix the mistakes of the past so we can get to the future. We need to make acceptable performance.


If everything in the initial code is 300% covered with excellently documented tests that should be minimally changed during transition (if transition don’t reveal any corner case tests were missing, maybe the transition is not such a bright move after all), that seems a possible thing to consider.

Otherwise it really sounds like a recipe for unnecessary huge risk with dubious expected positive outcome.

Not saying don’t have fun, but on the other side maybe not with the core product of you cash cow already?


Minor correction: LinkedIn, not twitter. https://www.linkedin.com/posts/galenh_principal-software-eng...

> Because it wasn't about "engineers should write 1M LOC per month of product code" it was "we want to scale automated porting of code to safe languages so that 1 engineer managing 1M LOC of automated conversion can work"

These are one and the same. Whether it's ported code or not doesn't change that. The framing device also doesn't matter, because it's the exact "Oh it's our goal" shtick that executives use in the former's case.

"It's just a measure" doesn't cut it in a world where every single AI measure immediately gets turned into a target by executives greedy for efficiencies that don't exist.

EDIT:

Right, I forgot. This is HN where everyone is a galaxybrain and "Port a million lines of code per month" is a totally reasonable goal for a single individual.


I can easily game writing 1M LOC per month by having the LLM write code in more verbose ways, with useless indirections and abstractions thrown in for good measure. I could even ask claude to write code that does nothing but just takes up line.

In contrast, converting 1M LOC of code per month is a much more solid measure, as long as you measure LOC of the source, not the new code. Sure, in the short term you can pick the easy/verbose things to port, but it's hard to do sustainably. A 5M LOC code base would still be expected to be ported in 5 engineer months.

Granted, you can still rush the work, not test properly, neglect good planning and engineering. Ported lines of code should not be the only measure (just like with any other measure). But it's a much less problematic measure than coding 1M LOC


> Granted, you can still rush the work, not test properly, neglect good planning and engineering.

Which is the core point of my reply and not something to just be casually handwaved, thank you very much.


> "we want to scale automated porting of code to safe languages so that 1 engineer managing 1M LOC of automated conversion can work". Which doesn't seem like satire at all..?

Because many programmers don't believe that'd work. See the reaction to Bun's porting to rust. (I bet Bun will work and prove those programmers wrong, but that's another story.)


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: