American AI companies are offering text and code and other content they illegally scraped from the internet and reselling it packaged as “AI subscriptions” making it impossible for many professionals to compete as they are competing with impossibly cheap, resold stolen goods.
Chinese labs turning the LLMs into open source, making all that money burn that is making so many things unaffordable for us now is literally the best outcome possible.
The issues with LLMs go beyond just IP theft. I would not say PRC making LLMs cheaper is the best outcome (though it is better than nothing). The best outcome would be to make the practice of training on our data without consent illegal, which would simultaneously slow down economic change and make it more organic as well as give PRC companies less capabilities to extract.
> There is no IP theft because LLM outputs aren't protected, just egregious ToS violations
I meant original IP theft that occurs to train LLMs in the first place. But sure that implies that further LLMs based on that LLM are also tainted by that original IP theft.
I can't make heads or tails of your opinion-free comment, made up of only questions.
My best guess is you're suggesting that Anthropic's model outputs are transitively under copyright (as a reproductions of human work under copyright?), but somehow ownership now belongs to Anthropic and not the original owners, and therefore Anthropic has standing against Alibaba? Not only does this go against what Anthropic argued in court against authors and publishers, such jurisprudence would lead to the immediate shutdown all leading LLMs in the US which were all trained on stolen work.
They can license training data. They have trillions, look what they are dumping into it, you seriously think they can't afford to license data.
Obviously it would be easier if they do it from the start, but that was their trick, to do it while people don't notice and get big ASAP. Should they get away with it?
Also, it would solve their Chinese problem, because it would make them violate copyright too. Right now it's more like rules for thee not for me so it's hard to take seriously.
That's the conundrum isn't it? Anyone that posts their datasets would be immediately sued/blocked/boycotted to oblivion due to the obvious and blatant data theft, not to mention IP and copyright issues.
Nvidia's even being sued for providing scripts which automate the downloading of said data from non-Nvidia sources. We certainly don't need copyrights that last nearly a century after the author's death (they literally cannot help the author), so here's hoping that some of the disputes over all this money changing hands can reign in some of the existing copyright sprawl. A stronger public domain would provide more useful training data for everyone, including open source models, and make criminals out of fewer AI researchers.
The TAM that spacex quotes has nothing to do with Space. It is all AI. As in grok. This is a joke, and people that believe it are just providing exit liquidity.
Banking Circle | Data Scientist | London or Copenhagen | Full time
Banking Circle provides account and payment infrastructure to other regulated financial firms. You have probably never heard of us but we have probably moved money on your behalf at some point this month.
As part of this we need to screen payments for sanctions compliance and anti money laundering and fraud concerns. The team hiring builds the ML models that decide if the payment needs to go for manual review.
We’re a diverse mix of people (50% have PhDs) and we work a wide ranging set of problems.
Specifically we are looking for candidates to strengthen our GIS (geolocation) / graph (network) and NLP (genAI) capabilities.
A good candidate will have experience in two of the three areas, a track record of delivering business relevant results and some research experience.
Majority of California based companies employee English only or English and Spanish speakers possibly with some Indian language as well. This leads to lots of problems when you are bilingual or bilingual in other languages such as German in French. Neither Apple nor Microsoft under this sort of language swapping well. Never mind rarer languages like Czech or Greek.
> Majority of California based companies employee English only or English and Spanish speakers possibly with some Indian language as well [...] Never mind rarer languages like Czech or Greek.
That may be generally true, in this case Apple actually has an engineering team in Czechia that works on biometrics and authentication:
Guess what, they’ll do nothing. If Czech market is small enough for them to fix quotation marks, they’re not fixing Czech keyboard.
OTOH, if an American will whine enough on Internet, they may fix it for him. Maybe some other American should use standard Czech quotes as password to get it fixed also.
I'm a little impressed with Google. Recently the assistant started understanding when I speak Portuguese or when my wife switches to it in a text message. I hadn't had that experience before, the assistants would pick one language and mispronounce the other.
Alexa has an experimental bilingual mode but it's nerfed by its general failure to understand well.
Netflix can't even auto-translate subtitles (in the age of genai where we are close to generating entire movies from scratch). Let alone ever imagine that you'd want to see subtitles in two languages at once.
We run into this issue when watching Korean movies/dramas. My wife prefers Japanese subtitles and I prefer English/German.
I haven’t found a way to enable two subtitles in Firefox (via extensions). So in those cases I usually download a release which contains subtitles in both languages and use a script to extract them via ffmpeg and then combine them into a single srt.
Now the issue is that the lines of the different languages don’t always appear/disappear at the same time. This leads to text jumping up and down. I have tried to mitigate it by injecting white space where only one line is visible, but this again fails when the video player breaks long lines or when the location of the subs change to the top (because there is hard-coded text in the image).
I feel like there must be a better way…
Sorry what? While there are right wing idiots in various governments in the EU, the Trump admin is on a completely different level. Also the bosses of big tech are clamouring over each other to s** him off.
I’m not particularly patriotic or bothered about nations in general, but the yanks can go take a hike.
The Shah fell after an ever increasing wave of student protests and violent crack downs.
The regime is doing better to hold on to power by killing more brutally, but there is no guarantee that will suffice to motivate its soldiers to kill the protesters in sufficient numbers to quell the unrest.
Ask Assad how killing his people worked out.
Just to be clear, I’m obviously against killing protesters.
American AI companies are offering text and code and other content they illegally scraped from the internet and reselling it packaged as “AI subscriptions” making it impossible for many professionals to compete as they are competing with impossibly cheap, resold stolen goods.
Chinese labs turning the LLMs into open source, making all that money burn that is making so many things unaffordable for us now is literally the best outcome possible.