Hacker Newsnew | past | comments | ask | show | jobs | submit | chaboud's commentslogin

I've been building latency-sensitive LLM systems for a while, and I've come to rely heavily on pre-fill-considerate mechanics like ping-pong overlapped async context construction. For interactive mechanics, the worst case, even if rare, is problematic.

A toy/simplified version lives here: https://github.com/chaboud/goulash

Consideration of mutation rate (a sort of temporal Shannon-ish coding/ordering) lives in there (with some RoPE-friendly structuring). Note: That was a vacation project, not the day job, but similar principles apply even with larger models.


Wait, wait, wait... you're telling me that polars, something made in this decade, is better at modern problems than something made two decades ago?!

I'm shocked! Shocked, I say!

Sometimes you just want something that works with all the things. That's pandas. But, like the bamboo eaters, it wont be long...


"When a measure becomes a target, it ceases to be a good measure."

Goodhart's law strikes again.

https://en.wikipedia.org/wiki/Goodhart%27s_law

However, what is meaningful is whether something is able to create usefully adjacent output, like "let's make Minecraft, but with marching cubes, subdivision surfaces, and global illumination... and behaviorally accurate pandas..." (or something like that).

I have an 11 year old, and most of his game ideas are adjacent to other games he's played. He can make those now, or, at least, enough that he can see where it works and where it doesn't.

Compared to a few years ago, that's pretty cool.


>"When a measure becomes a target, it ceases to be a good measure."

Pelicans?

I heard Gemini was fast so I tested the new one, asked it to clone a popular online game. It took 4 minutes, and worked perfectly first try.

I might need to sit down.


Your observation is game-changing, and it reverses my suggested priority completely.

Codex is fabulous at work, where token use is near limitless and ultra-thorough tool use is welcome. Go ahead and fire off searches for look-alike terms on my 96-core cloud instance.

By contrast, Claude Code's bias to make assumptions of reasonableness about underlying systems has proven to be immensely frustrating over the last month or two, both personally and at work. I've wasted days on "that was my mistake. I've been reporting numbers on the old architecture because I hadn't enabled the new one in the config" both at work and home. It's immensely frustrating.

But here we are. Wrestling with energetic idiots in model form, wrangled by over-specific harnesses that struggle to stay off of deranged side-quests.

What a time to be alive!


Keep in mind that the model is thinking in a token space, itself a compressive representation of language.

(Note: there's still a huge grammar penalty, so, ugh do think small.)


The real breakthrough is going to be thinking in latent space.


Arguably this is already happening: the whole state of the model gets fed through from token to token, and even just shoving a bunch of dashes in between the input tokens and the model's output can improve performance (thinking tokens from the model help a little bit more, but the difference is not as large as you mught expect).


It selects tokens but they expand to embedding vectors which are huge, also in memory and attention requirements, I think?


Listen to them, but don't listen to them.

You are building to solve their problems or open up their capabilities, and if they knew how to do it, they would have already. So listen to your customers for the "why", but use your own judgment for the "what" and the "how" of it all.

And be honest about being differentially valuable. I once had a prospective consultancy customer ask me for a very specific and elaborate piece of software to be built. He'd been thinking about it for years. I dug in on what he was actually trying to achieve, and I realized that small modifications to an existing in-market product would be able to satisfy his actual needs. I connected him with that company, he ended up with exactly what he needed, and they ended up with a high-end expert user as a resource.

Of course, I ended up putting myself out of a lucrative consultancy job, but I don't regret the decision at all. I suspect that most folks on HN feel the same way. Find a way to make a difference, and be honest with your customers (and leads) about what you're bringing to the table. That kind of integrity pays dividends in the long run if you have something truly valuable.


I have heard so many of my co-workers use "load-bearing" over the last couple of months. It's truly comical. Maybe this is a way that we can make "fetch" happen.


Why is that funny. I've had that as part of my professional lexicon for over 20 years


The thing is that some people start to adopt certain "mannerisms" from their LLM of choice. It's not funny in and by itself, but it tends to be unnecessarily pompous words/expressions as well. Relevant: https://www.vice.com/en/article/youre-not-imagining-it-peopl...


I've seen this too and I get it. However, we should not assume that certain phrases are AI tells. And that was my point. There are all kinds of things I see described here as "AI slop" that are things I just do, and have done, for decades.


Not in isolation, but if a variety of telltales is observed it raises the likelihood of slop, cf Naive Bayes as first approximation


I've renovated houses (and I have a bad habit of buying 100 year old hacked up chaos-boxes - my current house was moved from one hill in San Francisco to another 80 years ago, so it's a puzzle box), so yes, I've been using "load-bearing" as a shorthand for decades, reduced to be less jargon-y by adopting "structurally essential" for increasingly international team composition, where English is a second or third language.

However, I've been hearing "load-bearing" at least two orders of magnitude more often over the last few months, particularly after uncorking Claude Code for the team.

I don't think it's a dead give-away of AI usage, and I don't think AI usage is a problem. I just think we can introduce phrases into common use by having them be used by common tools. So let's train the models on obscure/archaic terms and see what happens. Heck, we can just prompt it...


I made the decision to always talk to AI in English, in order to reduce the otherwise inevitable seep-in of AI terminology into my actual speaking and writing patterns. Czech and English are far apart enough that formulations like "load-bearing" don't cross the barrier easily.


adjacent surface seam


I don't think it is. It's really easy to get a persistent, clever, hacky model to dial that down a bit and just come back to chat before stomping off into the woods.

If a model couldn't ever do that in the first place, it'll just get stuck.

I work in the "ZeroOne" space, working on concepts and prototypes for things that don't exist in market yet. Sometimes these models crank hard and immolate tokens while grounding themselves on expensive-to-ingest self-developed frameworks. If the results are well judged and the crank-turn latency is low, I'm okay with the cost as long as the model isn't wasting my time.

But when I want to do more boilerplate work, I turn down the model and thinking level and get more traditional about restraining action. For the really hard stuff, I reach for the models that will start a token bonfire in the back yard.


It's a carrier medium... and it's pretty thick. I would have just committed to "slurry".

Sandy water at the beach that looks a bit brown? Suspension.

Mining run off that looks like soup and coats everything? Slurry.

What's the Half & Half between milk and cream but for tweaky techniques for fabrication? Slurspension?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: