Any custom harness for a problem shows that harness engineering is going away. Eventually models will introspect problems, then build custom harnesses tailored to that. Then use and modify the ephemeral harness as required.
Sol Ultra style is the path forward. The models are smart enough to self serve their tooling and processes. Given a problem they can figure it out and ask for directions when needed.
How is custom engineering the tooling from scratch for every task the best path forward? There will always be some engineered tools that are better than others even when re-made by the same models - not to mention the cost of starting from scratch every single time just seems wasteful with current token spend.
Does "don't re-invent the wheel" not apply to agents for some reason?
> Does "don't re-invent the wheel" not apply to agents for some reason?
Correct, the purpose of agents is to become the ultimate wheel so we humans can stop needing to reinvent the wheel. For it to do that it needs to know how to reinvent wheels.
And, just to be clear, "don't re-invent the wheel" is just for small teams, as a whole humanity needs to re-invent the wheel all the time to adapt to new situations. For every wheel your team shouldn't reinvent some other teams main job is to reinvent that wheel.
Humans (exhibiting "general intelligence") are tool builders; virtually all our capabilities stem from our ability to create and use tools - often extremely specialized to a task. Why would an artificial general intelligence be any different?
I think there is a chasm to cross. the model's training to be aware of the harness it is in, at the same time building probes and observation tool can help it to cross that chasm.
Still seem too soon for a model to have that ability to build a harness on its own, and swap its session to another harness in the same environment. Like a snake shedding its skin, but in this case its harness.
It's likely smart enough. It just needs to be told to do it and provided the ability to introspect it. How close could foundational models get to building this harness if explicitly prompted to?
We've only just started training models to use tools. Next, we'll train them to build them. Harness engineering is an ephemeral art.
Let's maybe say not experienced enough / insufficiently RL'ed then -- 5.6 Sol did not reach for a harness solution like this when it got only 13% or son AGI-3 recently. I agree it's interesting to find the point in the prompting when it could 'tip' and do this. I have no instinct for where that point is, except that it must be somewhere, because I bet the Schema Harness was not hardcoded.
Letting the provider decide for the harness is a terrible idea in my eyes. Outsourcing harnessing is giving up control over the AI and equivalent to abandoning your sovereignty. It is a regression to a pre-enlightenment era.
Any non-digital system you describe "as Code" is not a source of truth, but a Source of Hope. The code describes the intended state, something has to reconcile it with reality.
This is the same as having it in unstructured documents. Which means the auditing is still required funny enough.
So yes, this could be done. I'd love to see what run in the CI/CD for a change. When someone works on the wrong thing, or breaks compliance IRL, how do you backport it into this? "Alice is a software engineer, and created this SaaS account with her email when the company was founded. The admin email can not be changed and she has admin even though another role should control that"
Yes, this is ridiculous. Far too much chrun in this ecosystem. This decreases my confidence in this Mintlify company; given this, they seem like the type to randomly rugpull when they feel like it.
Yeah, I can understand where you're coming from. I guess there's balance to the extent that being overly stubborn is also not a great thing. Prioritization at an early stage startup is a tough task.
There is no world outside of a cult or hype bubble where a 6 day turnaround time from announcement to deprecation is acceptable for a standard.
Standards are something that multiple independent groups can build on and should be somewhat stable.
If this was labeled as experimental than maybe, although its crazy to waste time publishing if you are working on something that rapidly iterating. Its not though, because everyone in this AI boom is rushing to be the winner takes all, and is trying their best to signal to the crowd that they are that winner.
Nvidia sees the forest of the trees. The consequences of the US government buying steaks and Intel are that there will be Federal requirements for us companies using Intel. This is entirely about the foundry business. Nvidia is at risk when 100% of the production of its intellectual property occurs in Taiwan. They're more interested than anyone else in diversifying their foundry solutions. Intel has just been a terrible partner and totally disregards its customers. It's only because of the new strategic need for the US to have a foundry business that the government is saving until. NVIDIA is understandably supportive of this.
Open models are going to win long-term. Anthropics' own research has to use OSS models [0]. China is demonstrating how quickly companies can iterate on open models, allowing smaller teams access and augmentation to the abilities of a model without paying the training cost.
My personal prediction is that the US foundational model makers will OSS something close to N-1 for the next 1-3 iterations. The CAPEX for the foundational model creation is too high to justify OSS for the current generation. Unless the US Gov steps up and starts subsidizing power, or Stargate does 10x what it is planned right now.
N-1 model value depreciates insanely fast. Making an OSS release of them and allowing specialized use cases and novel developments allows potential value to be captured and integrated into future model designs. It's medium risk, as you may lose market share. But also high potential value, as the shared discoveries could substantially increase the velocity of next-gen development.
There will be a plethora of small OSS models. Iteration on the OSS releases is going to be biased towards local development, creating more capable and specialized models that work on smaller and smaller devices. In an agentic future, every different agent in a domain may have its own model. Distilled and customized for its use case without significant cost.
Everyone is racing to AGI/SGI. The models along the way are to capture market share and use data for training and evaluations. Once someone hits AGI/SGI, the consumer market is nice to have, but the real value is in novel developments in science, engineering, and every other aspect of the world.
I'm pretty sure there's no reason that Anthropic has to do research on open models, it's just that they produced their result on open models so that you can reproduce their result on open models without having access to theirs.
[2 of 3] Assuming we pin down what win means... (which is definitely not easy)... What would it take for this to not be true? There are many ways, including but not limited to:
- publishing open weights helps your competitors catch up
- publishing open weights doesn't improve your own research agenda
- publishing open weights leads to a race dynamic where only the latest and greatest matters; leading to a situation where the resources sunk exceed the gains
- publishing open weights distracts your organization from attaining a sustainable business model / funding stream
- publishing open weights leads to significant negative downstream impacts (there are a variety of uncertain outcomes, such as: deepfakes, security breaches, bioweapon development, unaligned general intelligence, humans losing control [1] [2], and so on)
I don't think there will be such a unique event. There is no clear boundary. This is a continuous process. Modells get slightly better than before.
Also, another dimension is the inference cost to run those models. It has to be cheap enough to really take advantage of it.
Also, I wonder, what would be a good target to make profit, to develop new things? There is Isomorphic Labs, which seems like a good target. This company already exists now, and people are working on it. What else?
> I don't think there will be such a unique event.
I guess it depends on your definition of AGI, but if it means human level intelligence then the unique event will be the AI having the ability to act on its own without a "prompt".
> the unique event will be the AI having the ability to act on its own without a "prompt"
That's super easy. The reason they need a prompt is that this is the way we make them useful. We don't need LLMs to generate an endless stream of random "thoughts" otherwise, but if you really wanted to, just hook one up to a webcam and microphone stream in a loop and provide it some storage for "memories".
I'm a layman but it seemed to me that the industry is going towards robust foundational models on which we plug tools, databases, and processes to expand their capabilities.
In this setup OSS models could be more than enough and capture the market but I don't see where the value would be to a multitude of specialized models we have to train.
I have this theory that we simply got over a hump by utilizing a massive processing boost from gpus as opposed to CPUs. That might have been two to three orders of magnitude more processing power.
But that's a one-time success. I don't hardware has any large scale improvements coming, because 3D gaming mostly plumb most of that vector processing hardware development in the last 30 years.
So will software and better training models produce another couple orders of magnitude?
Fundamentally we're talking about nines of of accuracy. What is the processing power required for each line of accuracy? Is it linear? Is it polynomial? Is it exponential?
It just seems strange to me with all the AI knowledge slushing through academia, I haven't seen any basic analysis at that level, which is something that's absolutely going to be necessary for AI applications like self-driving, once you get those insurance companies involved
Could be that you need massive amounts of data from those super expensive production training runs, and it's tough to figure that out from publicly available data and academic computing resources. Maybe the combination of gradual efficiency improvements, bigger compute clusters, and test-time reasoning keeps the cloud models in the lead. Plus, even if it's exponential scaling, wouldn't that still favor the big data centers? That would put local/edge models at a serious disadvantage.
This implies LLM development isn’t plateaued. Sure the researchers are busting their assess quantizing, adding features like tool calls and structured outputs, etc. But soon enough N-1~=N
To me it depends on 2 factors. Hardware becomes more accessible, and the closed source offerings become more expensive. Right now it's difficult to get enough GPUs to do local inference at production scale, and 2 it's more expensive to run your own GPU's vs closed source models.
[3 of 3] What would it take for this statement to be false or missing the point?
Maybe we find ourselves in a future where:
- Yes, open models are widely used as base models, but they are also highly customized in various ways (perhaps by industry, person, attitude, or something else). In other words, this would be a blend of open and closed.
- Maybe publishing open weights of a model is more-or-less irrelevant, because it is "table stakes" ... because all the key differentiating advantages have to do with other factors, such as infrastructure, non-LLM computational aspects, regulatory environment, affordable energy, customer base, customer trust, and probably more.
- The future might involve thousands or millions of highly tailored models
This will not be true for future frameworks, though it is likely true for current ones.
Future frameworks will be designed for AI and enablement. There will be a reversal in convention-over-configuration. Explicit referencing and configuration allow models to make fewer assumptions with less training.
All current models are trained on good and bad examples of existing frameworks. This is why asking an LLM to “code like John Carmack” produces better code.. Future frameworks can quickly build out example documentation and provide it within the framework for AI tools to reference directly.
Because there’s enough rails code in the training data to determine the proper conventions :) if you’re making something new without this glut of data, it’s going to be much more difficult for a coding assistant to match a convention it’s never seem before.
The thing is, with some elbow grease, you can write a great plugin for your preferred editor. No need for dubious LLMs results, especially when the difficult part, code intellisense, is already solved with LSP. If you're a shop that has invested in a framework, it would be cheaper and more productive.
true, but the conventions it has seen are the same across all similar domains not just same framework/language, copilot "picks up" the similarity.
What I mean is: if you name your modules consistently, say Operation::Object::Verb or Action::ObjectVerb or ObjectManager.doSomething it's really easy for the LLM to guess the next one, just as it is easy for a human.
Add a new file actions/users/update.rb and start typing "Act" and it may guess "class Actions::Users::Update, and start to fill in the code based on nearby modules, switch to the corresponding unit test and it'll fill it in too.
Source: we have our own in-house conventions and it seems copilot gets them right most of the time, ymmv.