The students in the computer lab had no guarantee that they wouldn't, say, download a copy of Napster infected with the CIH virus later. The fact that they were not under imminent threat from some kind of Hollywood-style network worm did not mean they were immune from more realistic attack vectors.
Likewise, one can quite reasonably say there is no credible existential, Hollywood-style threat from AI in the foreseeable future while recognizing far lower-stakes, yet important risks that need to be addressed.
the earth has been 1 decision away from explosion for nearly a hundred years now. do you anticipate a change to this agenda very soon? personally I imagine the same trajectory continuing.
certainly I behave like the article's author and yourself all the time, this isn't meant to be some sort of moral point. watching all these smart people talk so confidently regarding things nobody knows about, it reminds of the confidence ai shows in hallucination. would you say that the singular choice of vasili arkhipov did not prevent annihilation?
> would you say that the singular choice of vasili arkhipov did not prevent annihilation?
He prevented a catastrophe, but not annihilation. We were not at risk of that in 1962 and even if he had decided to go with the others, more decisions would have been needed (not just his) to fully escalate to full scale nuclear war. I will reiterate: We have never been one decision away from full scale nuclear war.
But in case you don't understand why, it's because no one person can actually launch all the missiles. And considering the two major arsenals (US and USSR), there has never been a time when two people could make the same decision (launch) and actually launch all the missiles. The orders still have to go out and acted on, many decisions have to be made in order to have full scale nuclear war and come close to annihilation.
I like the way you frame your opinion as a truth with an obligation to be understood. still I disagree that a decision can only be said to affect its direct successors.
If you think AI will be powerful enough to change the world in huge ways, then it does seem reasonable to assume some level of risk of destroying the world too.
Stating that the probability is zero when we simply don't know what the probabilities are does seem like an extraordinary claim.
Imagine how reassuring it would be to people if we had evidence that there's no existential risk?
> If you think AI will be powerful enough to change the world in huge ways, then it does seem reasonable to assume some level of risk of destroying the world too.
There's an opportunity cost in the doomsday prophesying, though.
The media coverage of "these extremely capable robots might kill us (according to the guys who sell the robots)" comes at the cost of coverage of the real, present issues surrounding the tech oligarchs that we're already facing.
Sure, there are some failure modes that lead to some really bad stuff, but human extinction seems exceedingly unlikely to be one of them, and "10%" is pulled straight out of Dario's derriere.
Far more likely are economic disruptions that impact the tenuous balance between labor and capital and lead to unpredictable societal upheaval.
No it does not seem like an extraordinary claim, unless this is your first time hearing such claims. For the rest of us, we've heard this every decade and it turns out to be entirely untrue.
Nuclear, Overpopulation, Peak Oil, Y2K Bug, Global Warming.
No, you explained why the selection bias doesn't matter, which I agree. The self-report data still is completely worthless in the first place though.
People are really, really good at lying to themselves, let alone to an online poll. They'll tell you that they prefer imperfect or even bad writing as long as it's not AI slop, just like how they'll tell you they like healthier food, they prioritize personality instead of look for potential dates, how they use LLM "only as a spellchecker", and how they use tiktok for educational videos. As long as there is no stake, people will just say what make they feel better.
Self-reporting data for human behavior is just noise.
I mean, fine, it's anecdotal -- but the numbers are also overwhelming (and don't seem to disagree with the tenor here).
And look: you're obviously free to ignore me because you feel that the survey data is "completely worthless" and just slop your way to success -- all I'm doing is trying to explain why you shouldn't expect me (and people like me) to read what you create.
Interesting! Amusingly, if I feed that detector this blog post, it identifies it as confidently robot (97 out of 100 test passages). And running through my last five blog entries, they are all over the map, with three deemed at least "likely robot." Looking further back in time (and taking a somewhat random example), a blog entry from 2008, "Concurrency's Shysters"[0], is also deemed as similarly confidently robot (also 97 out of 100); do you expect this high a false positive rate?
My experience is using Pangram quite often with lots of writing of all flavors (including a bunch of known origin).
As for my own writing, I didn't do this experiment, but one of my co-workers did -- and over 176 posts spanning 22 years, all 176 (well, 177 now with my latest) are 100% human. This is not hugely surprising in that (in addition to me having actually written them!) my voice is very... distinctive. What would be more entertaining would be to try to get an LLM to write like me and fool Pangram that way. I still think that this would be difficult based on the experiences that I've heard, but it wouldn't surprise me if you could pull it off (and I would assuredly find the result entertaining!).
In the dimensions that we use Pangram in the most actionable sense (namely, to audit our own public writing), I am unconcerned about false positives, and leave it to Oxide authors to rework/recast as needed. (Though it sounds like Freddie didn't even need to do that -- he just needed to provide a longer sample.)
Now that you’ve seen it can be brittle (e.g. if a small sample is provided, per this single case), would it be sensible to add a disclaimer to the post? It’s a great ad for the tool (& I’d love for a perfect tool to exist!), so it’ll sell subscriptions & we wanna make sure that some teacher out there doesn’t falsely accuse a kid, or engineer doesn’t think worse of their colleague unfairly, etc.
False negatives are mentioned, but the false positive is what could hurt people.
Well, give it a shot -- you'll likely find that that technique doesn't work nearly as well (at least with Pangram 4) as you think it might. When we had Max on the podcast[0], Adam explicitly asked him about exactly this (after all, you can give an LLM access to Pangram and let it iterate!), and Max reported that someone had attempted to do this -- and ended up burning through $700 in tokens and had a "sad Claude." Another interesting bit: according to Max, newer models are diverging more from human writing not less. I think that that was more anecdotal than quantified, but an interesting comment nonetheless.
99% of college essays and pretty much everything “product” in corporate America is now LLM generated with some marginal oversight. It passes muster for the most part.
This is (obviously?) false, but considering how well-capitalized we are at the moment, you do have me wondering what a quid pro quo would be for; perhaps in this fictional universe Pangram has lucked into some of the PCIe clock buffers that we've been scrambling to secure enough of?
By "quid pro quo" I wasn't suggesting that Pangram's PR people's podcast placement was pay-for-play, just that they traded access for your positioning of their tech and their exec in your content marketing efforts.
That's pretty normal, but the point is that a blog post which is 35% Pangram promotion may not actually be less annoying than the use of AI to help write blog posts.
Yeah, fair -- and definitely not: I am earnestly just a fan of what they built (and I also think it's really important as a way of getting a check against rampant LLM use).
That is honestly the highest possible praise -- thank you. And when this piece was starting to boil inside of me last night (triggered, I'm sorry to report, by an obviously LLM-authored guest blog entry from the Rust Foundation[0]), I messaged one of my colleagues: "Time to do what I do best: bluntly say what lots of people are thinking."
I saw 'the results speak for themselves' but for me this article seemed to have less LLM-ese than some of the more recent obvious LLM prose posted to hackernews. As a reader I think its jarring because you just see a lot of articles purportedly written by different people using a very similar voice. I guess pre-LLM you might see this in a newspaper with very strong editorial oversight. So the phenomena is not completely new but it feels stranger when its not from a single source. It's also kind of sad to see a some people who have written a lot in the past about interesting technical topics in a way that was easy to read to give up their voice and outsource it to an LLM. But given this is basically free labour from the authors it feels a bit ungracious to complain.
My complaint is directed at the Rust Foundation, not the authors, to be clear. Yes, it's free labor, but the Foundation ought to have standards.
I showed Bryan the post while in a meeting room with him, left the room, and a few minutes later heard a very loud, very exasperated primal scream coming from the room!
Wow, thanks for that shout out! When we started Oxide in 2019[0], my kids were 15, 12, and 7, so out of those time-consuming younger years. But when we started Fishworks[1] inside of Sun in 2006, that 15-year-old (now nearly 22) was not yet 3 -- and his brother and sister were still in the future. Point being: I worked plenty hard with young kids.
More generally, I have always believed that the choice between work and parenting is broadly a false dichotomy (especially in our domain and age, which affords quite a bit of flexibility). In my role as a parent, I served on the board of the kids preschool (hello, SOMACC!); I was the president of the elementary school PTA (originally called the Mothers' Club!); I led the kids' Cub Scout Den (hardest parenting job no question); I coached baseball several times over; and I have been (and still am!) the scoutmaster of the kids' Scout troop.
These experiences were all time consuming (and truthfully, there were moments with each of them when I wondered what I had signed up for!) -- but they also made my life richer, and that very much includes my professional life. There are many dimensions of this, but to take one very concrete example: anyone who is familiar with scouting will see an echo in Oxide's values[2] -- but a subtler influence can be found in our hiring materials[3], the open-ended questions of which coming from an adult training that I did in 2017.
I won't profess to the quality of my parenting (it's... really hard), but it's something that I've certainly spent a lot of time on -- and it has paid surprising dividends to my professional life.
"Light on details" is not often a criticism that we get, but I think you might be looking in the wrong place: for the kind of detail that you're looking for, you should be heading into our RFDs -- specifically (for the issue you seem to be most interested in), RFD 26.[0][1]
(I think you meant "like hotcakes", not "like cupcakes"?)
In any event: from the beginning, we have always been targeted at the enterprise buyer who is looking at annual public cloud bills in the tens, hundreds or (in some cases) thousands of millions of dollars per year. We love the enthusiasm that Oxide engenders among the home lab set, but that's not how we have geared the business (for lots of good reasons).
Also, for whatever it's worth: people do want to buy them, it turns out.
I’d really love to know a ballpark price on a minimal install without wasting anyone’s time, have you guys ever discussed real numbers on the podcast? I listen occasionally but definitely haven’t heard every episode.
reply