Hacker Newsnew | past | comments | ask | show | jobs | submit | ClassAndBurn's commentslogin

He will always be the Pirate in Muppet Treasure Island to me. Thanks for everything you gave the world


He was easily my favorite part of that movie. The scene where he preaches his way out of getting The Black Spot lives rent-free in my head:

https://clip.cafe/muppet-treasure-island-1996/but-its-not-hi...

(aside: Someone needs to upload a better-quality version of this to YouTube)


I watched that movie so many times as a child I could probably still recite every line.

RIP Long John Silver!


This is my only number!


Any custom harness for a problem shows that harness engineering is going away. Eventually models will introspect problems, then build custom harnesses tailored to that. Then use and modify the ephemeral harness as required.

Sol Ultra style is the path forward. The models are smart enough to self serve their tooling and processes. Given a problem they can figure it out and ask for directions when needed.


How is custom engineering the tooling from scratch for every task the best path forward? There will always be some engineered tools that are better than others even when re-made by the same models - not to mention the cost of starting from scratch every single time just seems wasteful with current token spend.

Does "don't re-invent the wheel" not apply to agents for some reason?


> Does "don't re-invent the wheel" not apply to agents for some reason?

Correct, the purpose of agents is to become the ultimate wheel so we humans can stop needing to reinvent the wheel. For it to do that it needs to know how to reinvent wheels.

And, just to be clear, "don't re-invent the wheel" is just for small teams, as a whole humanity needs to re-invent the wheel all the time to adapt to new situations. For every wheel your team shouldn't reinvent some other teams main job is to reinvent that wheel.


Humans (exhibiting "general intelligence") are tool builders; virtually all our capabilities stem from our ability to create and use tools - often extremely specialized to a task. Why would an artificial general intelligence be any different?


"don't reinvent the wheel" isn't a law of physics and I think is mostly said by people that never designed anything with wheels.


Bespoke tools can be simpler than general purpose ones


I think there is a chasm to cross. the model's training to be aware of the harness it is in, at the same time building probes and observation tool can help it to cross that chasm.

Still seem too soon for a model to have that ability to build a harness on its own, and swap its session to another harness in the same environment. Like a snake shedding its skin, but in this case its harness.


Indeed. And you can make this case about any tooling at all that is model adjacent.


Including by extension all programs, operating systems, or eventually hardware, I suppose.


Yep. That was the entire subtext from the drop in IBM's share price yesterday.


> Sol Ultra style is the path forward

That’s called brute force so not really.


I seriously doubt this, especially in a world in which there's not just one model.

This makes sense if the models some how become unified.


I guess this is kind of an information theory thing. As a model converges AGI would a generalized harness ability appear?


Codex and Claude Code already have the workflow feature which basically does this on demand, so it's getting there.


Except in this case, it isn't yet smart enough. But I agree, building this capability in is coming, and will be really awesome.


It's likely smart enough. It just needs to be told to do it and provided the ability to introspect it. How close could foundational models get to building this harness if explicitly prompted to?

We've only just started training models to use tools. Next, we'll train them to build them. Harness engineering is an ephemeral art.


Let's maybe say not experienced enough / insufficiently RL'ed then -- 5.6 Sol did not reach for a harness solution like this when it got only 13% or son AGI-3 recently. I agree it's interesting to find the point in the prompting when it could 'tip' and do this. I have no instinct for where that point is, except that it must be somewhere, because I bet the Schema Harness was not hardcoded.


Letting the provider decide for the harness is a terrible idea in my eyes. Outsourcing harnessing is giving up control over the AI and equivalent to abandoning your sovereignty. It is a regression to a pre-enlightenment era.


I finally built my side project and the government finally does something to make it pointless! Yay?

https://www.daylightsmearings.com/


There's a life lesson there somewhere.


We'll see if it passes the Senate.


Uncharted island found. Charted.


I think they should name it Islandy McIslandface.


Any non-digital system you describe "as Code" is not a source of truth, but a Source of Hope. The code describes the intended state, something has to reconcile it with reality.

This is the same as having it in unstructured documents. Which means the auditing is still required funny enough.

So yes, this could be done. I'd love to see what run in the CI/CD for a change. When someone works on the wrong thing, or breaks compliance IRL, how do you backport it into this? "Alice is a software engineer, and created this SaaS account with her email when the company was founded. The admin email can not be changed and she has admin even though another role should control that"


They will have to update the openai. Com footer I guess

Latest Advancements

GPT-5

OpenAI o3

OpenAI o4-mini

GPT-4o

GPT-4o mini

Sora


> Deprecating install.md

> Last Friday we announced install.md and it didn't see much adoption.

With thrash like this why would anyone adopt this for something serious?

It's just an .md file, so the overhead is low. The lack on conviction in your design does not inspire confidence though.


Yes, this is ridiculous. Far too much chrun in this ecosystem. This decreases my confidence in this Mintlify company; given this, they seem like the type to randomly rugpull when they feel like it.


It smells of gambling fallacy. Surely, this time theyll win AI and itll finally find that corner case or function flow


Yeah, I can understand where you're coming from. I guess there's balance to the extent that being overly stubborn is also not a great thing. Prioritization at an early stage startup is a tough task.


There is no world outside of a cult or hype bubble where a 6 day turnaround time from announcement to deprecation is acceptable for a standard.

Standards are something that multiple independent groups can build on and should be somewhat stable.

If this was labeled as experimental than maybe, although its crazy to waste time publishing if you are working on something that rapidly iterating. Its not though, because everyone in this AI boom is rushing to be the winner takes all, and is trying their best to signal to the crowd that they are that winner.


Nvidia sees the forest of the trees. The consequences of the US government buying steaks and Intel are that there will be Federal requirements for us companies using Intel. This is entirely about the foundry business. Nvidia is at risk when 100% of the production of its intellectual property occurs in Taiwan. They're more interested than anyone else in diversifying their foundry solutions. Intel has just been a terrible partner and totally disregards its customers. It's only because of the new strategic need for the US to have a foundry business that the government is saving until. NVIDIA is understandably supportive of this.


Open models are going to win long-term. Anthropics' own research has to use OSS models [0]. China is demonstrating how quickly companies can iterate on open models, allowing smaller teams access and augmentation to the abilities of a model without paying the training cost.

My personal prediction is that the US foundational model makers will OSS something close to N-1 for the next 1-3 iterations. The CAPEX for the foundational model creation is too high to justify OSS for the current generation. Unless the US Gov steps up and starts subsidizing power, or Stargate does 10x what it is planned right now.

N-1 model value depreciates insanely fast. Making an OSS release of them and allowing specialized use cases and novel developments allows potential value to be captured and integrated into future model designs. It's medium risk, as you may lose market share. But also high potential value, as the shared discoveries could substantially increase the velocity of next-gen development.

There will be a plethora of small OSS models. Iteration on the OSS releases is going to be biased towards local development, creating more capable and specialized models that work on smaller and smaller devices. In an agentic future, every different agent in a domain may have its own model. Distilled and customized for its use case without significant cost.

Everyone is racing to AGI/SGI. The models along the way are to capture market share and use data for training and evaluations. Once someone hits AGI/SGI, the consumer market is nice to have, but the real value is in novel developments in science, engineering, and every other aspect of the world.

[0] https://www.anthropic.com/research/persona-vectors > We demonstrate these applications on two open-source models, Qwen 2.5-7B-Instruct and Llama-3.1-8B-Instruct.


I'm pretty sure there's no reason that Anthropic has to do research on open models, it's just that they produced their result on open models so that you can reproduce their result on open models without having access to theirs.


> Open models are going to win long-term.

[2 of 3] Assuming we pin down what win means... (which is definitely not easy)... What would it take for this to not be true? There are many ways, including but not limited to:

- publishing open weights helps your competitors catch up

- publishing open weights doesn't improve your own research agenda

- publishing open weights leads to a race dynamic where only the latest and greatest matters; leading to a situation where the resources sunk exceed the gains

- publishing open weights distracts your organization from attaining a sustainable business model / funding stream

- publishing open weights leads to significant negative downstream impacts (there are a variety of uncertain outcomes, such as: deepfakes, security breaches, bioweapon development, unaligned general intelligence, humans losing control [1] [2], and so on)

[1]: "What failure looks like" by Paul Christiano : https://www.alignmentforum.org/posts/HBxe6wdjxK239zajf/what-...

[2]: "An AGI race is a suicide race." - quote from Max Tegmark; article at https://futureoflife.org/statement/agi-manhattan-project-max...


> Once someone hits AGI/SGI

I don't think there will be such a unique event. There is no clear boundary. This is a continuous process. Modells get slightly better than before.

Also, another dimension is the inference cost to run those models. It has to be cheap enough to really take advantage of it.

Also, I wonder, what would be a good target to make profit, to develop new things? There is Isomorphic Labs, which seems like a good target. This company already exists now, and people are working on it. What else?


> I don't think there will be such a unique event.

I guess it depends on your definition of AGI, but if it means human level intelligence then the unique event will be the AI having the ability to act on its own without a "prompt".


> the unique event will be the AI having the ability to act on its own without a "prompt"

That's super easy. The reason they need a prompt is that this is the way we make them useful. We don't need LLMs to generate an endless stream of random "thoughts" otherwise, but if you really wanted to, just hook one up to a webcam and microphone stream in a loop and provide it some storage for "memories".


And the ability to improve itself.


I'm a layman but it seemed to me that the industry is going towards robust foundational models on which we plug tools, databases, and processes to expand their capabilities.

In this setup OSS models could be more than enough and capture the market but I don't see where the value would be to a multitude of specialized models we have to train.


There's no reason that models too large for consumer hardware wouldn't keep a huge edge, is there?


That is fundamentally a big O question.

I have this theory that we simply got over a hump by utilizing a massive processing boost from gpus as opposed to CPUs. That might have been two to three orders of magnitude more processing power.

But that's a one-time success. I don't hardware has any large scale improvements coming, because 3D gaming mostly plumb most of that vector processing hardware development in the last 30 years.

So will software and better training models produce another couple orders of magnitude?

Fundamentally we're talking about nines of of accuracy. What is the processing power required for each line of accuracy? Is it linear? Is it polynomial? Is it exponential?

It just seems strange to me with all the AI knowledge slushing through academia, I haven't seen any basic analysis at that level, which is something that's absolutely going to be necessary for AI applications like self-driving, once you get those insurance companies involved


Could be that you need massive amounts of data from those super expensive production training runs, and it's tough to figure that out from publicly available data and academic computing resources. Maybe the combination of gradual efficiency improvements, bigger compute clusters, and test-time reasoning keeps the cloud models in the lead. Plus, even if it's exponential scaling, wouldn't that still favor the big data centers? That would put local/edge models at a serious disadvantage.


> Open models are going to win long-term.

[1 of 3] For the sake of argument here, I'll grant the premise. If this turns out to be true, it glosses over other key questions, including:

For a frontier lab, what is a rational period of time (according to your organizational mission / charter / shareholder motivations*) to wait before:

1. releasing a new version of an open-weight model; and

2. how much secret sauce do you hold back?

* Take your pick. These don't align perfectly with each other, much less the interests of a nation or world.


> N-1 model value depreciates insanely fast

This implies LLM development isn’t plateaued. Sure the researchers are busting their assess quantizing, adding features like tool calls and structured outputs, etc. But soon enough N-1~=N


To me it depends on 2 factors. Hardware becomes more accessible, and the closed source offerings become more expensive. Right now it's difficult to get enough GPUs to do local inference at production scale, and 2 it's more expensive to run your own GPU's vs closed source models.


> Open models are going to win long-term.

[3 of 3] What would it take for this statement to be false or missing the point?

Maybe we find ourselves in a future where:

- Yes, open models are widely used as base models, but they are also highly customized in various ways (perhaps by industry, person, attitude, or something else). In other words, this would be a blend of open and closed.

- Maybe publishing open weights of a model is more-or-less irrelevant, because it is "table stakes" ... because all the key differentiating advantages have to do with other factors, such as infrastructure, non-LLM computational aspects, regulatory environment, affordable energy, customer base, customer trust, and probably more.

- The future might involve thousands or millions of highly tailored models


This will not be true for future frameworks, though it is likely true for current ones.

Future frameworks will be designed for AI and enablement. There will be a reversal in convention-over-configuration. Explicit referencing and configuration allow models to make fewer assumptions with less training.

All current models are trained on good and bad examples of existing frameworks. This is why asking an LLM to “code like John Carmack” produces better code.. Future frameworks can quickly build out example documentation and provide it within the framework for AI tools to reference directly.


I don't think convention over configuration causes LLMs any problems, GitHub copilot generates code matching rails conventions quite easily for example.


Because there’s enough rails code in the training data to determine the proper conventions :) if you’re making something new without this glut of data, it’s going to be much more difficult for a coding assistant to match a convention it’s never seem before.


The thing is, with some elbow grease, you can write a great plugin for your preferred editor. No need for dubious LLMs results, especially when the difficult part, code intellisense, is already solved with LSP. If you're a shop that has invested in a framework, it would be cheaper and more productive.


true, but the conventions it has seen are the same across all similar domains not just same framework/language, copilot "picks up" the similarity.

What I mean is: if you name your modules consistently, say Operation::Object::Verb or Action::ObjectVerb or ObjectManager.doSomething it's really easy for the LLM to guess the next one, just as it is easy for a human.

Add a new file actions/users/update.rb and start typing "Act" and it may guess "class Actions::Users::Update, and start to fill in the code based on nearby modules, switch to the corresponding unit test and it'll fill it in too.

Source: we have our own in-house conventions and it seems copilot gets them right most of the time, ymmv.


But the new frameworks will never have anywhere near the amount of training data as established frameworks.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: