From a user experience standpoint this sounds like a clear win. Congratulations.
With AI safety top of mind so much lately, I can't help but notice the announcement does not address this.
With "Chat" mode, there was a user expectation that session had only limited capability to produce unexpected side effects, read sensitive files, etc.
With "Cowork" mode, it seems like more powerful capabilities have been on by default, requiring deep settings and safety understanding to disable if desired.
Merging modes feels like it is removing a simple and easy to understand risk management tool. How does the combined mode help users understand, manage and feel confident about what risks they are accepting?
The dashboard for the Andon Market shows a bank balance of around $7k remaining out of the $100k seed. Rent is due at the end of the month. Last EOM brought a $11k-ish drop. Sales do not appear to be on track to cover the rent that's due in two weeks. What then?
It is pretty funny launch when your longest-term example is losing about $3 for every $1 of revenue. Though I guess the industry standard is to lose money on AI so they're in good company.
We're very explicit that this is for experimentation. And I wouldn't recommend using it to start a business that has expenses (rent+salaries) of tens of thousands per month.
The linked article is intentionally misleading by omission because they left out "in mice" in the university driven article and they certainly know the relevance and consequences of leaving it out.
Partially funded by entity related to manufacturer of daily multivitamins. Study was in people over 60 average age 70 if I recall. I didn't care enough to look, but the question I'd ask is how were the people who died during the study accounted for in the "biological aging testing?"
While that is true, a fact that increases the credibility of the published results is that the study also obtained a negative result.
Besides the multivitamin supplement, they also tested cocoa extract, for which similar effects had been claimed, and for that they did not find any effect.
The fact that the multivitamin supplement had effect is plausible. Presumably a significant number of the participants did not eat a perfect diet that would supply all vitamins in adequate quantities, so for those, taking a multivitamin supplement compensated whatever vitamin deficiency their diet might have had.
Such a positive result does not demonstrate that you need to buy a vitamin supplement, it just demonstrates that it is desirable to eat healthy food. However, there are circumstances when a vitamin supplement may be cheaper or more convenient than buying and eating enough food of an appropriate type.
For example, I take a vitamin supplement from time to time, but that is only because I have a sedentary lifestyle, working at a computer, so I must eat relatively little, otherwise I would gain weight immediately. When you eat little, it is difficult to compose a menu that will provide enough vitamins without also providing too much energy.
Multi vitamin use has been studied A LOT with respect to all cause mortality, where it consistently has no statistical effect (or a slightly negative effect). So, to me, "biological aging" seems a way to hack for effects to advertise. Does the average consumer understand that reduced "biological aging" does NOT mean you will live longer? Because if it did, they would be saying THAT. Centrum (study sponsor) is already creating advertising based on this study.
I have a colleague that recently self-published a book. I can easily tell which parts were LLM driven and which parts represent his own voice. Just like you can tell who's in the next stall in the bathroom at work after hearing just a grunt and a fart. And THAT is a sentence an LLM would not write.
Here's some alternatives. Some are clunky. But, some aren't.
…just like you can tell whose pubes those are on the shared bar of soap without launching a formal investigation.
…just like you can tell who just wanked in the shared bathroom by the specific guilt radiating off them when they finally emerge.
…just like you can tell which of your mates just shitted at the pub by who's suddenly walking like they're auditioning for a period drama.
…just like you can tell which coworker just had a wank on their lunch break by the post-nut serenity that no amount of hand-washing can disguise.
…just like you can tell whose sneeze left that slug trail on the conference room table by the specific way they're not making eye contact with it.
…just like you can identify which flatmate's cum sock you've accidentally stepped on by the vintage of the crunch.
…just like you can tell who just crop-dusted the elevator by the studied intensity with which one person is suddenly reading the inspection certificate.
It's still on you to pick what the LLMs regurgitate. If you don't have a style or taste you will simply make choices that would give you away. And if you already have your own taste and style LLMs don't have much to offer in this regard.
Just as it’s on you to pick the word you want when using Roger’s Thesaurus.
My workflow, when using it for writing, is different than when coding.
When coding, I want an answer that works and is robust.
When writing, I want options.
You pick and choose, run it through again, perhaps use different models, have one agent critique the output of another agent, etc.
This iterative process is much different than asking an LLM to ‘write an article about [insert topic)’ and hope for the best.
In any case, I’ve found the LLMs when properly used greatly benefit prose and knee-jerk comments about how all LLM prose sound the same are a bit outdated… (understandable as few authors are out there admitting they are using AI… there’s a stigma about it. But, trust me, there are some beautiful soulful pieces of prose out there that came out of a properly used LLM… it’s just that the authors aren’t about to admit it.)
One shouldn’t expect the ‘joke’ to have identical tone. (As if that’s even measurable.)
The point was simply that these examples are not trending towards the average or ‘ablating’ things as the article puts it. They seem fairly creative, some are funny, all are gross… and they are the result of very brief prompt… you can ‘sculpt’ the output in ways that go way beyond the boring crap you typically find in AI-generated slop.
The main difference is that the “tests” are predicates over live browser state and are often proposed alongside the plan on the fly, not written upfront by a developer. But conceptually it’s very close: make the expected outcome explicit, try an action, verify, and only move forward if the condition actually holds.
That book used artwork valuation as a performance measure and analyzed it over top artist's lifetimes finding two patterns. The "Young Genius" where an artist has a vision and realizes some innovation and their most valuable works center around that with value tapering off over their life. Picasso. (Who had two peaks but still fit the pattern.) Contrast to the "Old Master." This is someone who keeps refining their craft and their most valuable works and innovations are their late life works. Cézanne.
With AI safety top of mind so much lately, I can't help but notice the announcement does not address this.
With "Chat" mode, there was a user expectation that session had only limited capability to produce unexpected side effects, read sensitive files, etc.
With "Cowork" mode, it seems like more powerful capabilities have been on by default, requiring deep settings and safety understanding to disable if desired.
Merging modes feels like it is removing a simple and easy to understand risk management tool. How does the combined mode help users understand, manage and feel confident about what risks they are accepting?
reply