That is not the clear meaning. They are deliberately picking phrasing where they can lead you to believe they mean A but later you can't pin them down to actually meaning A. That's why they picked intentionally unclear language.
If you think it clearly means anything you are just assuming because it can't "clearly" mean something specific when they go out of their way to use non idiomatic language and they don't give very clear guidance using idiomatic language.
Well I happen to enjoy coffee and $200 AI plans. What if Blue Bottle started watering down it's coffee? Is your answer to stop drinking coffee and make myself tea instead?
Evidence that vendors are being misleading in what they are delivering is important to share, whether or not you personally approve of that product.
It was bought out by employees in 2017 and it seems like they weren't able to manage it successfully. The two that acquired it were well over 70 when they did so, so I guess this isn't very surprising. From the article it sounds like they spent a lot of money trying to 'modernize' the business, which sounds a lot like they brought in management that was more interested in 'doing management. than making money.
I don't understand the point they are trying to make.
It's very often (always?) the case that something general also solves particular problems.
A sorting algorithm is an implementation of min()
A parser also is a syntax checker.
A route planner is a reachability checker.
A computer algebra system is a basic arithmetic calculator.
A general constraint solver is a Soduku hint maker.
It's true that LLM output can be used as an input to another classifier, this is also true of any classifier. The improvement on top of the straight LLM classification is relatively small, and I would argue that working on the prompt or just including in the prompt for the LLM what features might be useful to consider would likely work even better.
Fundamentally I read this article as: We want to build a simpler, dumbed down clone of Mathematica, so we cobbled together the following pieces... We also needed a way to do arithmetic, so we also include a copy of Mathematica to do basic arithmetic.
I took the point as: don't make the LLM the classifier. Use it to turn messy input into useful features, then let a normal model make the actual decision. That gives you thresholds/calibration you can inspect.
What I'm not sure about is how stable those features are when you switch the underlying LLM or model version.
Much like how you shouldn't ask the LLM to solve a (repeated, logical) problem, but you should instead prompt it to generate code that you can inspect/test/fix/reuse.
That isn't what they did here. They took the output of the LLM as one feature, then added 17 other features, then piped it into a crappy model and got a 3% improvement.
I think the reason to not invest your time in PS5 linux is that sony hates it and will patch it away over and over, and nobody really wants it. It's $200, mediocre hardware.
I might be showing a lot of unc energy here, but if you can't afford a raspberry pi or something to play with linux, or a used pc, gaming and having a playstation 5 should be de prioritized for a bit while you get your life together.
You've misunderstood something deeply. PS5 Linux doesn't exist because running Linux on PS5 is something useful or practical. Exists due to the curiosity and fun of running a free OS on a locked-down device. Maybe some nostalgia mixed in as Linux has existed in every other PS (even PSP/PS Vita).
> I might be showing a lot of unc energy here, but if you can't afford a raspberry pi or something to play with linux, or a used pc, gaming and having a playstation 5 should be de prioritized for a bit while you get your life together.
Your comment is framed as though Linux usage and what you deem an unstable lifestyle (lack of monetary means in your judgment) are mutually exclusive. Open-source software is for everybody, including the messy or disenfranchised. Would you frame a kid who got a PS5 for Christmas who doesn't have an allowance or job as a failure because they don't have $100 to spend on a Pi? It sounds like you just look down on people playing video games even when they want to use the very thing you're implicitly deriding to do the thing that you think they should be doing... doesn't seem to make a lot of sense.
Your argument contains a category error. I am arguing based on what is adaptive and in a person's best interests. You are countering that they have the right to be a mess and still have hobbies. That does not answer my point at all; I agree with you that people who are a mess still have the right to enjoy whatever hobbies they want.
Is someone who owns a PS5 but can't put together $100 a failure? Assuming they are older than 16, unqualified yes. They are a failure and also a loser by any colloquial definition. TLC refers to them as a scrub. Their romantic partners' parents are gently encouraging them to end the relationship. There are almost certainly more than 7 empty cans within easy reach of wherever they are sitting right now.
They seem to conclude Luna is a better value, but their analysis is dumb. They just break it down to $/bug found.
However, Luna missed 23 bugs that Astra found, and identified 24 bugs that weren't really bugs. That's horrible. Astra had 96% precision.
The cost to care about here isn't just how much it costs to run the code review, or the cost per true-positive. It's the cost of dealing with this system. A code review system that is right about 2/3 sucks, and one that misses another 1/3 of the bugs is also a lot worse. The Astra code review quoted here would become the foundation of how the team works, the Luna version is at best helpful to find some stuff but does not dramatically increase your confidence. It also will force humans or better AI's to have to run down a lot of false positives, and that is treated as free here.
Actual conclusion: The cost for Astra is low in absolute terms compared to the cost of bugs and human attention, and the added value is far far more than the added cost.
To me, one interesting piece of analysis was whether Luna had benefit on top of Astra -- i.e., running both and synthesizing their findings. But even with that, it raises the question whether running a second Astra pass, or Sol, or even a model from another family (GLM? Fable?) would deliver even more additive benefit.
The nonsense about 'using all the water' and datacenter protests has to be a Chinese psyop. It's so fantastically stupid to protest someone taking an empty field and paying local taxes on it and employing dozens of highly paid people in the middle of nowhere. There's literally no downside. There is effectively no pollution. There are jobs for every level of education, from security guard to the people installing hardware. There's no toxic runoff, no danger of a chemical explosion, and no carcinogens being dumped. This is the only rural industry you can say that about.
Water is a renewable resource. I know you folks in California can't comprehend this, but in the parts of the country where it actually makes sense for humans to live water just falls from the sky multiple times per week. We get so much water the problem is making sure we get rid of it safely, we don't have to fight over who has the most senior claim to it or decide whether we want to have endangered species or almonds more. Our streams don't run dry 4/5 of the year. We don't have to check about water restrictions when we water our lawns because there are never water restrictions and we never have to water our lawns.
There are entire areas miles across in this country where if you drive through it you might throw up from the smell. The ponds full of animal feces make the air un-breathable. Where is the protest over that? You realize that this literally does ruin ground water and poisons surface water? Where are the people coming out to say they have to live 4 miles from oceans of pig feces and when the wind blows their way they can't go outside? Yet the media is able to find the 4 people on earth that want to say a building full of computers is 'loud' and that computers use lots of water? When was the last time you filled the water tank on your computer? Yet people are very quick to believe that somehow a warehouse with a bunch of computers in it somehow destroys water?
When did datacenters start using up all the water? Nobody seemed to worry about this until the US and China were fighting for AI dominance, then suddenly datacenters use water and are so loud people go insane from them. What is the first time someone mentioned datacenters using water and causing pollution?
It's absolute group psychosis. Are you actually so dumb that you don't notice that a thing that has been around for 50 years suddenly is a threat to our survival? Were you born yesterday so you don't remember 2 years ago when nobody had ever realized the existential threat of building datacenters in America? Why didn't anyone notice that datacenters were pumping the wells dry in 2021? Maybe they weren't and still arent...
Also, your last two paragraphs are completely bogus because you overlooked something important: scale.
AI data centers are being built at a massively higher rate than data centers were being built a few years ago, and an AI data center uses 5-10x more energy than a non-AI data center.
It is quite common for something that is not a big problem to become a major problem when scale massively increases. You can't simply dismiss concerns because they weren't concerns a few years ago. You will have to actually do the analysis to determine if they are legitimate concerns now.
Seems to me that most people still have to get up in the morning to go to a bullshit job just to stay alive and consoom things they don't need. I don't see the big change.
Get a better job my guy. Unrelated to AI, why are you living like that it sounds horrible and you can just get another job. I'm happy to help you fix up your resume if you would like.
They are comparing it to 2 and 3 version old flash/fast versions of models but purely for tok/s. Then only comparing it to Mercury 2 on intelligence. This is very misleading and I suspect this model is basically useless.
Even pretty dumb models are useful for running subtasks (especially at this sort of speed). Note that in the coding section they only mention using it as a subagent for a smarter model
If you think it clearly means anything you are just assuming because it can't "clearly" mean something specific when they go out of their way to use non idiomatic language and they don't give very clear guidance using idiomatic language.
reply