Hacker Newsnew | past | comments | ask | show | jobs | submit | heipei's commentslogin

So what is the alternative? It took me a while to get around the data schema / modelling that a database like Cassandra imposed, but when it finally clicked it really clicked for me, so much so that I wish I could have the same performance footprint of consistent hashing / sharding with other databases.

We are using ScyllaDB, but since they discontinued their Open Source version we're stuck on the last supported version, and reading about the operational headaches of Cassandra does not inspire confidence to give it a go.


Drop in with same schematics? It's not going to be easy, I don't have an answer. It's unpopular, but probably I'd architect my software more towards a sharded mongo cluster for write heavy workloads these days, rather than lean on cass/scylla/cockroach.

Mongo bit me a lot pre-wiredtiger and I refused to forgive them for years, but I'm using it again since 8.x and it's really... Boring, Scalable. There are many ways to skin a cat though, and I know not everyone has the luxury of redesigning their apps to fit the persistence layer.


> We are using ScyllaDB, but since they discontinued their Open Source version we're stuck on the last supported version

Do you think you are missing anything by being on scylla oss?


Same issue we had with cockroach. The oss is lagging and you ultimately need to end up paying the company for support. its really just "kinda oss" in that you can look at source code from 2 years ago but to actually use it, might as well be oracle. If your fork it you are just gonna be in a hostile environment (unless you pay them)


Second that question not only to OP/GP. Comparing Scylla and Cassandra - the choice is obvious, but current offering is a complete no-go.

Also Mongo is catching up in the use case we have and it’s even more scary. We are all the time at the exact edge of Mongo scalability and I’m talking about very good, few-years-invested-in setup on NMVe that is really delicate…


i'd say 99% of usecases will be served fine with Postgres, or a managed instance via AWS


We live in wildly different worlds, I think. How is postgres comparable to 6-12 nodes of cassandra? One really giant one with 200k worth of storage? Is there even a multi-master setup yet? every time i’ve looked it’s just around the corner. multigres from supabase seems promising, but im not touching it for at least a couple years…

I’d love to use postgres, and have in various times in the past and love it (it never broke) — but one machine isn’t nearly big enough for really any of my current applications.


You can use a distributed database like yugabytedb or crdb. They are horizontally scalable, support postgres (yugabytedb has much better postgres compatibility). Yugabytedb also supports the Cassandra API.


cocroach is also open source no anymore.

ydb is not clear if easier to maintain at scale.


>multigres from supabase seems promising

I have high hopes for Neki from Planetscale but then again if you are not stuck on Postgres for whatever reason, Vitess with MySQL is about as battle tested as it gets.


Legitimate file hosting services present the biggest total volume of newly discovered phishing pages / unique hostnames. Another (similar) angle is using unrelated legitimate domains which are compromised (think insecure Wordpress) to host phishing sites in subdirectories. A lot of traditional ML scoring and blocking approaches fall flat if the hosting domain is on a very legitimate and hard-to-block domain, such as a government website.


Can we please stop posting here every time Claude or OpenAI are down? These are not critical services like AWS / S3 where downtime means that your application is offline. Nothing is easier than switching to another provider, mid-session even.


No this whole thread is such a nice thing to happen. It's like the lights are out on your building and you're joking about it with your neighbors when you're all just waiting for it to come back. But on a gigantic scale.

I love how people can share moments even though they are not close to each other physically at all. Imagine the cats outside could somehow connect with the cats in the other side of the world. No other animals can do this. Maybe whales, it depends how you define connect.

We are just gifted to live in such times.


On one hand, I fully agree with you. On the other hand, look at all these funny threads! Makes me smile every time.


Neither is AWS


Cloudflare is really good at launching features that facility low-friction deployment of malicious content (such as phishing) on the Internet, piggybacking on their hosting reputation and the fact that you can't easily block their ASN or domains.


I don't know your experience. Once I was toying around and doing a basic auth with registration and so. The weekend was over and couldn't get back to that couple of months. The worker was quarantined and marked as phishing automatically. So I believe they have something in place to prevent those you complain.


Your anecdote just illustrates that their system detects legit uses as abuses, not that they have a system that effectively detects abuses.


But it is not that they have nothing. It was my laziness that I could not setup dev prod env's. When you develop on preview, I don't think they will do much.


Cloudflare is also like a Chinese copycat machine. They mostly copy some successful project and sell it at cheap price.


Be the change you want to see to make the world of your dreams.

And then sell its denizens malice protection services.


This is gonna be amazing for phishing, like most of the features Cloudflare offers (free Turnstile for fresh accounts, CF tunnels, pages.dev, r2.dev).


Depends on what you mean by "local". On your Macbook, large dense models like Qwen 3.6 27B will be slow, sure. On a local workstation with a dedicated RTX card you can get > 100 tps, which is more than good enough to work with it, and faster than cloud models in many cases.


But how smart is it? All the people running local models never seem to mention that they are way dumber than cloud models.

I don't care how many tokens per second of nonsense it can generate.


Qwen 3.6 35b a3b is about as good as sonnet 4.5. It varies but it's at that level.


Not even close.

It may be "about as good" on some very specific task.


It is smart enough that I use for all my coding tasks, and a lot of other mundane tasks.

It is probably not smart enough for "design this whole architecture of this complex system from scratch, make no mistakes", but that is not something I want from a coding tool anyway. I want a model that I can point to a file and tell it to make some changes to the file and related files. Or that I can ask to review a PR with regards to certain aspects.

My suggestion is to simply try it and see what it feels like.


Quantized Gemma 4 26B is as smart or better than GPT 5 in most of my testing. Granted GPT 5 is nearly a year old at this point, but I can run Gemma 4 on a ~6 year old consumer GPU (RTX 3090) and get 140 t/s.


> But how smart is it? All the people running local models never seem to mention that they are way dumber than cloud models.

Well, you aren't going to give it a 20k line sec and have it churn out a full app after 4 hours hours.

But, you can get it to write code for you if you do the design.


Its not going to be as good as Claude, but if you know what you're doing, it may be good enough to get your work done.


This is task dependent.

I find devstral (even though it’s weak generally) much better at writing and documentation than Opus. I’m actually now delegating all documentation to devstral and away from Claude, which makes a mess.


A highly skilled carpenter may be able to 'get work done' by banging nails in with a heavy-bottomed cocktail glass, doesn't mean it's not painful to do so when it is continuously breaking and leaving shards of glass all over the workshop for you to find every day for the rest of your life until you clean up the mess you made using the wrong tool for the job.


More like, a highly-skilled carpenter can work miracles with a $6 hammer from the hardware store, while the pros on the commercial crew are using fancy compressed-air tools.

The carpenter has to get up close and personal with the wood. He can't match the crew's throughput, but maybe that's not what he's trying to do.


I would say the hammer is no AI. Local models are the cheapest XKGYAGH electric nailer on Amazon that "works" but jams up all the time. The $20/mo cloud models are a nice DeWalt that gives an hour of jam-free operation but takes five hours to recharge. And if someone else is paying for it, one can use the heavy duty nail gun with a big generator and compressor on a trailer that can run all day.


If someone comes into the workshop and takes all the tools (hello Donald) then having a cocktail glass to hand might be a bit of a lucky break.

(geddit?)


I'm talking about the common use case that I think hacker news people have:

you get a macbook for work, you run the macbook

they're not going to start giving GPUs to employees to run local models


Same here, I use Qwen 3.6 27b (Q6 quant) with llama.cpp on an RTX 5090 using the pi agent exclusively now. The fact that it's local means that I never have to think about token pricing, quotas, time of day, or data sensitivity. I have limited the GPU from 600W to 450W which means the system stays whisper quiet during inference.

I have become so "lazy" (in a good way), so far that I've started using the model for lots of daily mundane things on top of just coding:

  * "commit this on a branch, push, create a PR and assign $nickname for review"
  * "Use the Stripe CLI to download all open and overdue invoices and reconcile them with this CSV export from our bank account."
  * "Use these Elasticsearch credentials to summarise what kind of operations are causing load at the moment."
  * "Tell me if our codebase already supports X and where it's  implemented."


What context length and kv cache quant (if any) are you using? And MTP?


No KV cache quant, context length 50% of original, MTP absolutely. These are the relevant cmdline attributes. Getting around 100t/s with this setup, even when watt-limited to 450W.

  --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 --presence-penalty 0 --metrics --jinja --chat-template-file chat_template.jinja --chat-template-kwargs '{"preserve_thinking": true}' --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.75 -ngl 99 -c 131072 -fa on -np 1 -hf unsloth/Qwen3.6-27B-MTP-GGUF:Q6_K


Not the person you asked, but I have a 9700 which has the same VRAM, and running Q6 on it with unquantized kv gives me 50k context. Putting -ctv q8_0 ups that to 70k. I normally run Q4 with unquantized kv @ 130k at 50 t/s (mtp 3), with the disclaimer that I'm running PCIe gen4x8, so I'm slightly slowed. I've found that quantizing k leads to broken json on tool calls, which is fairly unrecoverable, but YMMV.


$5B valuation on a $30B run rate for Anthropic?


40B ARR and profitable next month.


It's always profitable next month.


Local AI models are already more than capable enough writing code that surpasses the ability of any bad or even mediocre engineer. That is not something we need to worry about.

In a way, this is less of a cost issue than the fact that some/many engineers do not seem to be willing or able to host things themselves anymore and will happily outsource every part of their stack to managed services, be it CDN, hosting, databases, etc. I don't know why that's not more alarming than the LLMs.


Qwen 3.6 27B is shockingly good, just to add to your point.


Thank goodness for China or Silicon Valley capitalists would be locking us down into an unimaginably awful dystopia. Though they're not done trying.


Serious question: Who actually builds stuff on Cloudflare workers? I mean large software projects / services, and not just side projects where the ability to scale-to-zero is perhaps more important than the scale-to-infinity direction. I feel like Cloudflare keeps pushing workers with its full force yet I fail to see the appeal.


I'm building a commercial SaaS product on Workers. Although I've barely scratched the surface of what Cloudflare offers¹, so far it's been great. The value proposition is effectively the same as serverless in general: You worry about the product, they worry about deployment. Note that Cloudflare Workers is just one (albeit important) star in their constellation of capabilities.

¹https://developers.cloudflare.com/directory/?product-group=D...


Cloudflare Workers solves their "scaling the price to infinity" problem.


Me. I used to deploy everything via Docker swarm but recently migrated everything to workers because wrangler is awesome, it makes blue/green very simple, and it's a lot less of a headache for me to maintain in general. It's a great/flexible product. I also use R1/R2 pretty extensively.



It's always seemed like a solution looking for a problem.


Anything built in Svelte can be deployed to workers easily, and it’s a very good platform.

Just missing compartmentalisation features between prod and dev environments.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: