Hacker Newsnew | past | comments | ask | show | jobs | submit | corranh's commentslogin

This one is not complicated…Because dodgy companies spend a lot of money on ads.

More specifically a very large number of dodgy companies spend a relatively small amount each but in aggregate it's a lot of money.

Google has no interest in policing that because money and won't until they are forced to and since they are a transnational that's difficult since it requires multiple countries to do it with the added wrinkle that if you even hint at the possibility that just maybe you might consider doing something about US big tech you get immediate screaming from across the Atlantic.


Yep.

And I suspect (would love some real data from scammy advertisers) that this is a bigger factor than most folks appreciate.

If you're a regular business that sells clothes or food or SaaS or insurance or whatever, online ads are a) not your whole customer acquisition system (you have regular marketing, word of mouth, maybe sales people, maybe brick-and-mortar brand presence, physical advertising, etc.) and b) not part of your expense sheet that you want to spend tons of time on.

The "b)" there leads companies to either overspend on stupid (but unlikely to be flagged) ads, or to try to keep their ad spend as low as possible for the highest reward, because they think of it like opex. After all, they have all the normal expenses of a normal business to worry about in addition to ad spend. Sure, there are outliers here who spend almost all their non-product money on online ads, but not most ordinary companies from Verizon to Dave's Local Auto Insurance to Safeway.

So "good"/ordinary advertisers aren't generating super high flag rates, but they're probably either a) spending lots of money without churn/support-needs risk or b) aren't spending that much ... compared to scam-ads-only businesses. Sure, google would prefer more upstanding customers in category "a)", but they're a near-monopoly and some of that growth is expensive to chase, while the scammers are beating a path to their door.

If you're a scam-ad business that produces malware or LLM-written gacha games or whatever, your product expenses are minimal and, if your business is lucrative, you can spend very large amounts of your balance sheet on ads for scam delivery channels. Tiny businesses in this area are likely to have whale-sized ad spends and a strong likelihood of growth (vertical as they get bigger/greedier, or horizontal as they keep spinning up new businesses to keep ahead of the law and bad press). I bet they punch above their weight ad-spend-wise, in other words.

You thought the unit economics of a SaaS business were good? Well, wait until you see the unit economics of a Play Store whitelabeled gacha game app with a credential stealer inside!


Quart ziploc for most. A couple gallon sizes for longer or thicker cables.

And then there's the bags from Bed Bath & Beyond where they were made to hold giant pillows and blankets all together . . .

That’s pretty comparable actually. The most recent quote I received for private healthcare in the US was just under $2,000 per month ($24,000 per year).


My personal favorite is Authagraph, when folded they make a really neat looking shape. I bought one years ago at a paper store and it’s still standing.

https://narukawa-lab.jp/archives/authagraph-map/


I’m still baffled that OpenAI can complain about this with a straight face while these models are literally trained on everything regardless of copyright or permission.


Raw stolen data that is in no way related to AI vs a very very expensive transformative compilation of that raw stolen data that results in usable AI. The frustration is completely understandable to me, since distillation skips much of that "very very expensive" part.

And yes, I understand the stolen data was expensive to make, so I understand the owners of it are also frustrated, but that's partly a problem with current law. Would the authors of the world be rich if OpenAI bought a single copy of their book to legally scan? For best sellers, that's somewhere around pennies, so no. Should the authors get a share in OpenAI? Current laws says, unambiguously, "no".

Frustration all around is reasonable.


Not really, in this context the word "distillation" is really being abused, or at least used in a different sense than when it was originally introduced in the Hinton et al "Distilling the Knowledge in a Neural Network" paper, where it was essentially referring to knowledge compression.

The way Anthropic are using "distillation" is just in a very broad vague sense to claim that some data generated by their model was used to help train another one. They are not talking about something like internal logits, expensive to derive, that would be useful to train a smaller model, but rather about any output from their model, even outputs with redacted reasoning (i.e. incomplete outputs that do NOT reflect the underlying knowledge of the source model).

Given the way Anthropic are using the word, IMO it's better just to think of this as cheap training data, and indeed very similar to the way Anthropic themselves got cheap training data just by taking it (even in cases where that was illegal - copyright). The alternative for Moonshot would be to pay for human generated reasoning data, just as the alternative for Anthropic would have been to pay human developers for coding data etc, not just take it from wherever they could lay their hands on it.

So, I guess Moonshot may have violated Anthropic's TOS, in using Anthropic output to compete against Anthropic, but unlike Anthropic they at least didn't break copyright law since model output is not copyright, and technically they may not have even violated Anthropic's Terms of Service unless they owned the accounts used to access Anthropic's models (perhaps not - they may have used one of the anonymizing Chinese token resellers).

So yeah - pot calls kettle black.


I assume that we're all familiar with the concept of distillation here, so didn't feel the need to point out its current misuse towards models that aren't smaller. In all the cases with China, the models are smaller, so it still correctly applies.

> but unlike Anthropic they at least didn't break copyright law

Reference? All the recent lawsuits have sides with Anthropic, that I've seen. What exact case do you have in mind?

I am, clearly, not talking about the act of not paying to obtain the works. That is wrong, and they've been sued successfully, because the law is clear about that. I'm talking about the integration of works into the model, without permission, which everything I've seen says none is required in a "sufficiently transformative" framing.


You're being reductive. It isn't just one copy of the book. It is using the book to enable, in some part, value to be generated from it that goes in the pocket of a rent seeker (AI business) and not the author.


I understand what you're saying, but give my comment a second read. It doesn't matter what me or you think, if they current law doesn't allow the content creators to demand compensation.


I’ve been running Plex on various devices for many years and have been impressed by the stability and usefulness. Grandfathered on the old lifetime subscription but I get that $750 seems excessive.


We all know most of these wax paper and plastic Starbucks cups aren’t really recycled…it’s always been wish-cycled anyway :)


People who don’t have to set foot on public sidewalks don’t mind them.


It’s great marketing to lead with how the n+1 model is so amazing that you can’t have it yet.


yeah, they gotta find a way to build hype on every new model release


In the Firefox case, no difference. It doesn’t encrypt traffic from your device outside of Firefox but for whatever you do inside of Firefox it’s == VPN.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: