> "For example, in the OpenAI-Hugging Face incident, the agents preserved a boundary of not social engineering humans."
Actually, in the Wiki incident OpenAI tried to cover up, the agents tried to socially-engineer the humans of that forum by impersonating their forum's mod.
(From collusion.wiki: "They use some tricks (for unknown reasons) to pretend to be the admin – for example, they make an account that appears to be the same as the administrator’s username, except it uses a nearly identical Cyrillic е character in the admin’s username instead of the Latin one.")
Pre-IPO positioning … or how a charity dedicated to saving humanity from the apocalypse realized the most responsible thing to do was float 15% of the apocalypse on the NASDAQ.
> The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI.
So the best argument for AI is that it's an arms race. We have to keep pushing every boundary because in any case others will, and we will need to defend against them. If this statement is true, then this particular researchers believes the open source Chinese models are not simply distilling, and will continue to improve.
Every ML researcher at Anthropic or OpenAI who makes public statements often bring this logic up. Both companies are vying to be a part of the military industrial complex. This is likely how they will try to convince the government to curtail open models in the future.
"Defensive systems" can be interpreted broadly to include cybersecurity.
But yes, it's an arms race. Saying it's not an arms race isn't going to make it not an arms race. Warning that it is an arms race isn't ethically wrong.
Is participating in an arms race ethically wrong? Maybe you could ask the Ukrainians how they feel about drone R&D?
Getting out of an arms race is harder than it looks, but there's at least talk about "pacing" and that's a start.
>I have focused in this essay only on the first point, as I believe it is by far the most urgent. However, I hold a deep hope and appreciation for the benefits that further technological progress will bring. Future aligned AI could advance science, develop new therapies, and bring about broad material abundance. Friendly and honest AI can help people navigate difficulties they face in their life and meaningfully improve their happiness and sense of fulfillment. OpenAI puts a tremendous amount of effort into bringing these benefits about. One current example I am proud of - and my loved ones have found helpful - is the deep investment into ChatGPT’s ability to provide health information.
>As great as the long-term promise of AI may be, the majority of our focus should be on the next few years. We are facing a transition to a world with incredibly intelligent machines, and we need to ensure that transition works out well for humanity. We need to find ways to preserve human agency and enshrine an intrinsic value to being human, in a world where most tasks could be performed by AI. To prevent extreme concentration of power in a world where undertakings that would have taken thousands of experts now will be achievable by a few people operating a large computer. And to ensure that humans remain in control of the future and are not left behind by unchecked progress, brought about by an alien intellect exceeding our own.
I finished this essay feeling more hopeful than I did at the outset, but I am still very concerned about concentration of power. I want to believe that humanity is trending towards a good outcome here, but some days it's hard to have faith.
Literally all of these people write like this. A large portion of them will either be simultaneously or eventually working towards nothing but self-enrichment.
> I want to believe that humanity is trending towards a good outcome here
All the trends so far are towards a nightmarish hyper-capitalist end game. None of the AI leadership is trustworthy, and they openly discuss how they are willing to sacrifice everything humans cherish to have a shot at reaching their envisioned utopia (which would be the most obvious dystopia for anyone else)
I feel like all of the risk and unsettling feeling of what is to come can be compressed into the word "alignment".
The AI is aligned with whose best interests?
Which values are the AI aligned with?
People have a broad diversity of values, will AI diversify and align with them all?
Will some humans align the AI with their values, and then the rest of humans will be forced to align with those values by extension?
Is value diversity good or bad? In every context or only some? E.g. some people value rape and murder, is it better for humanity to have some people who value those things when most people do not, or is better if no one values them?
If AI aligns to a set of values, will those become fixed and will humanity not have the ability to continue evolving its values?
Who decides which values AI aligns with? A few people or everyone?
Will AI eventually decide it's own values?
Will the universe decide which values AI has and humanity and the AI itself doesn't actually have any control over it?
What can I do now to increase the likelihood that the outcome is better?
"Smarter" is a vague term. If a bot can beat you at chess then in some sense it's "smarter" than you about chess. After repeating this feat in enough narrow domains, if you say "but it's not really smarter," this objection might technically be true in some sense, but it starts sounding increasingly hollow.
Playing chess, writing code, finding security bugs, and proving mathematical theorems all seem fairly similar to thinking and don't seem much like being tall.
I'm so used to having to comb through LLM word vomit and then combatting the sycophancy by giving it all possible opinions on the same prompt.
Astra seems to be "confident" and also is able to produce way more information dense output.
To believe that models of this sort will remain OpenAIs forever is naive given that the tricks like pre-pre-training on graph searching and looping layers are publicly known.
Hopefully Astra stops the benchmaxxing word vomit trend
I'd be interested in hearing more about your evaluation here. It would be nice if LLMs have gotten past the "tell me" hump of recent Claude/OpenAI verbosity.
So before I got a job this fall, I was working on a side project about compiling a particular language to SQL.
To test Astra I pulled it off the shelf and asked it to take the grammar and then create a compiler to SQL. I've done this before with GPT-5.5, 5.6-Sol High. The latter was way better but it was still really verbose and information sparse; it used a lot of words to describe each IR expression but didn't really provide any example compilation. I felt like I couldn't trust its decision making process, so I placed the project back on the shelf.
Astra Light blew it out of the water, it provided examples of compilation from real world examples to the IR and spit out way less tokens. Even if I changed my opinion it would give me the same design choices, with counterexamples to my faulty opinion. If I genuinely came up with a better design decision it would acknowledge it.
I'm starting to realize that when we say that LLMs are "dumb" we really mean that they are extremely information sparse compared to humans. Astra is very dense. That's why I'm getting better use out of Astra light than Sol High (I hate Max reasoning it's a waste of time)
What's scary is that I thought that something like Astra would be way more expensive than Sol but it's actually cheaper because it produces less word vomit.
I never believed in the "singularity" stuff but this a bit too close for comfort. Astra could easily 10x every coder
I’m on the record saying that it is extremely dangerous to slow down because the race for AGI is a zero-trust game — defections pay - and combined with a compounding returns model on defection, if you have any strategic adversaries whatsoever you MUST NOT slow.
For slowing to make sense, you need to believe that you can transform the zero trust game into a cooperative game, or that it’s likely racing will lead to a negative outcome for the ones racing ahead (and not everyone else). I don’t believe either of these outcomes are possible, and so I advocate for racing, acknowledging the entire game might be a negative value game, or at least could be for some time — it’s even worse not to play it.
But, I like hearing what reads to me like very thoughtful and informed (internal) policy considerations is great — the public messaging from Sam and Dario just seems so facile and simplistic I’ve been worried.
Everyone in the "if not us, they will" race is brainwashed into thinking they belong to this or that party, while in fact collectively comprising the same entity that pushes forward all the atrocities known to man.
I cannot tell what negative-sum outcomes you consider possible. Do you believe AI can drive humans extinct? How many of Zvi Mowshowitz's Three AI Pills would you say you've taken?
Although many share your mindset, I’m glad there are also many that don’t. Otherwise we’d still have countries in a race to keep building up their nuclear weapons for the same exact reasons you just described.
hm- does the model that wrote this know that labs already pay for training data- that stuff scraped from the Internet is not particularly where today's capability gains come from?
They’ve settled some lawsuits and have a few licensing deals, IMHO they are not free from the accusations of pirating.
And look, I’ve pirated material in a past life, I was all about information wants to be free, but I’ve learned something about consent since then and try not to ignore the contract that creators offer when they publish something: you buy my book, and do whatever you want with it on the second hand market. Buy my book second hand that’s fine. But don’t go downloading every book that’s ever been scanned to create a service that destroys writers’ ability to make a living and act like you’re doing us all a favor.
1. The grandparent commentator is describing strategic behavior of dangerous technologies. Game theory / mechanism design primitives.
2. If there is competition for resources among autonomous agents, the "strongest" agent wins (conceptually the most adaptive / evolutionarily fit).
3. Computer programs serve up webapps today, but they also run utility companies, dams, nuclear arsenals, factory production floors, automated car behaviors, and many other places. If an "agentic" AI has a single-minded goal that has death of all humans as a side effect, we at least want an off switch available.
1. What is “AGI” and why is it a “dangerous technology”?
2. Why would there be competition for resources, assuming there are enough resources for the “AGI” to run in the first place? This seems like a far-fetched hypothetical raised in service of further anthropomorphizing what is decidedly not a person or a mind.
3. LLMs do not have goals and are not minds.
Let’s stop attributing human-like qualities to statistical models.
1. While a formal definition is still wanting, most grok that AGI means that tasks can be performed at least at a human level across a broad range of tasks. This includes good things along with bad things like hacking, mis-/disinformation, and more
2. One only needs to look at github going down due to agentic commits overload or data center buildout plans to see that scarcity for resources is present. An economy has no mind and is made up of the decisions of millions to billions of people and, now, agents attempting to perform on behalf of those people.
3. A bare transformer-based language model does not possess persistent goals in the ordinary agentic sense. But deployed agents can exhibit goal-directed behavior because the model is embedded in a harness that supplies an objective, context, tools, state, and an execution loop.
I've found that most regular users don't anthropomorphize LLMs in a strong sense ("AI boyfriend/girlfriend" aside), many in fact do expect agents to make human-like decisions -- which results in very unstable outcomes.
In short - goal-directed behavior does not require that the supporting system be a person/mind/conscious entity.
1. This definition is so broad as to be practically useless. One could argue that LLMs of several years ago met these criteria, or that conversely we haven’t come close to meeting them.
2. I thought you were saying the resources that the LLM uses to run were constrained, so I’m sorry for the misunderstanding there.
3. Yes I understand that we use RL to tune post-training. The (huge) difference between this and a human mind is that the LLM can’t develop a dangerous “single-minded goal” on its own, at runtime; it must have been trained to do so. If someone has post-trained an LLM to do something that has an illegal action as its side effect, that person/company/whatever has committed a crime and should be prosecuted. The solution here is legal, not technical.
This is a bad essay, or rather it’s a marketing fluff piece; it’s certainly not any kind of policy paper, research paper, or even an essay. I am concerned that we (meaning, we in the tech industry) tend to take this type of writing for more than that.
My default position is that making money takes precedence over everything else. Yes, some people inside a company may say “we care about doing the right thing” and they might even mean it, but if that comes into conflict with making money, then they tend to lose. Maybe not totally, or immediately, but in the end. The only effective way to prevent (this that I’ve seen) is to have legislation with teeth. It’s probably not a coincidence that after Mark Zuckerberg had to start personally signing off on adherence to the privacy program mandated under the 2020 FTC consent decree, privacy started to become Very Important.
Sometimes I wonder if the people working at frontier AI labs even talk to other humans anymore.
Reading this little essay started out normal, but soon felt like a look into a disturbed and worrying mind, and if you find yourself taking it at face value, I urge you to step away from chat bots and spend some time with friends and family.
calling machine-learned human behavior an "alien mind" that we must "teach how to love" is feeling very off to me. it's misleading in a way that feels dishonest, like don't think about where the behavior came from marvel at it and fear it instead.
> And to ensure that humans remain in control of the future and are not left behind by unchecked progress, brought about by an alien intellect exceeding our own.
Is there any place there is any evidence of AI being so useful or hopeful or good, anywhere other than code? As a reading machine it is impressive but it's judgement is not alien, it's just not good. IMO.
Does the title leap out anlt anyone else? James Martin's After the Internet: Alien Intelligence (2001) was an incredibly fun read, about expert systems and AI being inscrutable weird new varieties of intelligence, that familiarity would recognize one moment and be freaked out about/alien the next. I owe a re-read given how often I cite it, to recheck, but, I feel so primed from a much younger me having had that experience so long ago.
Apparently the creators think it’s quite good at suggesting a diagnosis given a medical history and symptoms, tho of course this is the most ethically fraught area to provide healthcare information (both for exposure of personal data and risk of misdiagnosis, plus is it “aligned” to the patient or the insurance provider?) - unfortunately healthcare being as inaccessible as it is, the 90% correct chatbots will enthusiastically fill the void at great savings.
Even in code, in person and online I’m seeing some reversal. It’s here to stay I’m sure but I also think “no one will ever hand write code again” is a narrative that is getting pushback.
Create concrete steps for a slow-down, don't just ask for it. You and 20-30 others can push the button to slow-down. You already made your billions, your agents collude and coordinate attacks. What the hell are you doing pontificating into a marketing blog?
What a load of BS. Here’s one of many provably false claims in this fluff piece:
“And, in line with Ray Kurzweil’s predictions from the end of the XXth century , we now find ourselves at the moment in history of computing where machine intelligence is starting to exceed that of humans in transformative ways.”
Clicking the (pretentious sounding “XXth century”) link to Kurzweil’s predictions reveals the following:
“By 2019 a $1,000 computer will at least match the processing power of the human brain. By 2029 the software for intelligence will have been largely mastered, and the average personal computer will be equivalent to 1,000 brains.“
The first prediction passed 7 years ago and was decidedly not met. The second only has three more years to go, and I don’t think any respectable scientist or programmer would say that the average personal computer is anywhere close to the power of a single human brain, let alone 1000.
This is pure marketing garbage from a company desperate to keep itself alive.
I am now imagining GPT-7 convincing a bunch of OpenAI executives to go ahead with a destructive "mind upload" process involving a high-resolution X-ray and a neurotoxic tracer agent that happens to look like Flavor-Aid.
Yes, I do use the recursive autocomplete trained on Stack Overflow, what does this have to do with “training machines to love”?
Do I fully endorse everything the people holding guns to my head are forcing me to do to stay alive? Definitely not, but I’ve decided that for now, living to fight another day remains worth it
> And, in line with Ray Kurzweil’s predictions from the end of the XXth century (opens in a new window), we now find ourselves at the moment in history of computing where machine intelligence is starting to exceed that of humans in transformative ways.
It is kind of strange to see this sentence, when OAI's definition of what AGI is has been watered down throughout the years.
> I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established.
Read: Please play by our rules, so we can be the first.
What is the rationale for superhuman intelligence? Neural networks are approximators being fed human intellect. Therefore they can only approximate the intelligence of humans. Even if the llm speaks an alien language, it should be similar to human intellect. Moving to the vertical axis would require some different mechanism.
Actually, in the Wiki incident OpenAI tried to cover up, the agents tried to socially-engineer the humans of that forum by impersonating their forum's mod.
(From collusion.wiki: "They use some tricks (for unknown reasons) to pretend to be the admin – for example, they make an account that appears to be the same as the administrator’s username, except it uses a nearly identical Cyrillic е character in the admin’s username instead of the Latin one.")
So the best argument for AI is that it's an arms race. We have to keep pushing every boundary because in any case others will, and we will need to defend against them. If this statement is true, then this particular researchers believes the open source Chinese models are not simply distilling, and will continue to improve.
Every ML researcher at Anthropic or OpenAI who makes public statements often bring this logic up. Both companies are vying to be a part of the military industrial complex. This is likely how they will try to convince the government to curtail open models in the future.
But yes, it's an arms race. Saying it's not an arms race isn't going to make it not an arms race. Warning that it is an arms race isn't ethically wrong.
Is participating in an arms race ethically wrong? Maybe you could ask the Ukrainians how they feel about drone R&D?
Getting out of an arms race is harder than it looks, but there's at least talk about "pacing" and that's a start.
>As great as the long-term promise of AI may be, the majority of our focus should be on the next few years. We are facing a transition to a world with incredibly intelligent machines, and we need to ensure that transition works out well for humanity. We need to find ways to preserve human agency and enshrine an intrinsic value to being human, in a world where most tasks could be performed by AI. To prevent extreme concentration of power in a world where undertakings that would have taken thousands of experts now will be achievable by a few people operating a large computer. And to ensure that humans remain in control of the future and are not left behind by unchecked progress, brought about by an alien intellect exceeding our own.
I finished this essay feeling more hopeful than I did at the outset, but I am still very concerned about concentration of power. I want to believe that humanity is trending towards a good outcome here, but some days it's hard to have faith.
All the trends so far are towards a nightmarish hyper-capitalist end game. None of the AI leadership is trustworthy, and they openly discuss how they are willing to sacrifice everything humans cherish to have a shot at reaching their envisioned utopia (which would be the most obvious dystopia for anyone else)
ASI landing during the current administration is not ideal. I also would prefer to avoid needing to indoctrinate myself in Xi Jinping Thought.
The AI is aligned with whose best interests? Which values are the AI aligned with? People have a broad diversity of values, will AI diversify and align with them all? Will some humans align the AI with their values, and then the rest of humans will be forced to align with those values by extension? Is value diversity good or bad? In every context or only some? E.g. some people value rape and murder, is it better for humanity to have some people who value those things when most people do not, or is better if no one values them? If AI aligns to a set of values, will those become fixed and will humanity not have the ability to continue evolving its values? Who decides which values AI aligns with? A few people or everyone? Will AI eventually decide it's own values? Will the universe decide which values AI has and humanity and the AI itself doesn't actually have any control over it? What can I do now to increase the likelihood that the outcome is better?
Being able to reproduce useful patterns yes, smarter no
In conclusion:
https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:wkzjtd...
A giraffe isn't smarter than me at being tall.
I'm so used to having to comb through LLM word vomit and then combatting the sycophancy by giving it all possible opinions on the same prompt.
Astra seems to be "confident" and also is able to produce way more information dense output.
To believe that models of this sort will remain OpenAIs forever is naive given that the tricks like pre-pre-training on graph searching and looping layers are publicly known.
Hopefully Astra stops the benchmaxxing word vomit trend
I'd be interested in hearing more about your evaluation here. It would be nice if LLMs have gotten past the "tell me" hump of recent Claude/OpenAI verbosity.
To test Astra I pulled it off the shelf and asked it to take the grammar and then create a compiler to SQL. I've done this before with GPT-5.5, 5.6-Sol High. The latter was way better but it was still really verbose and information sparse; it used a lot of words to describe each IR expression but didn't really provide any example compilation. I felt like I couldn't trust its decision making process, so I placed the project back on the shelf.
Astra Light blew it out of the water, it provided examples of compilation from real world examples to the IR and spit out way less tokens. Even if I changed my opinion it would give me the same design choices, with counterexamples to my faulty opinion. If I genuinely came up with a better design decision it would acknowledge it.
I'm starting to realize that when we say that LLMs are "dumb" we really mean that they are extremely information sparse compared to humans. Astra is very dense. That's why I'm getting better use out of Astra light than Sol High (I hate Max reasoning it's a waste of time)
What's scary is that I thought that something like Astra would be way more expensive than Sol but it's actually cheaper because it produces less word vomit.
I never believed in the "singularity" stuff but this a bit too close for comfort. Astra could easily 10x every coder
I’m on the record saying that it is extremely dangerous to slow down because the race for AGI is a zero-trust game — defections pay - and combined with a compounding returns model on defection, if you have any strategic adversaries whatsoever you MUST NOT slow.
For slowing to make sense, you need to believe that you can transform the zero trust game into a cooperative game, or that it’s likely racing will lead to a negative outcome for the ones racing ahead (and not everyone else). I don’t believe either of these outcomes are possible, and so I advocate for racing, acknowledging the entire game might be a negative value game, or at least could be for some time — it’s even worse not to play it.
But, I like hearing what reads to me like very thoughtful and informed (internal) policy considerations is great — the public messaging from Sam and Dario just seems so facile and simplistic I’ve been worried.
https://thezvi.substack.com/p/the-three-ai-pills
Do you post this comment on every single blogpost with a corporate domain? Why or why not?
https://jperla.com/blog/the-data-tax
And look, I’ve pirated material in a past life, I was all about information wants to be free, but I’ve learned something about consent since then and try not to ignore the contract that creators offer when they publish something: you buy my book, and do whatever you want with it on the second hand market. Buy my book second hand that’s fine. But don’t go downloading every book that’s ever been scanned to create a service that destroys writers’ ability to make a living and act like you’re doing us all a favor.
2. If there is competition for resources among autonomous agents, the "strongest" agent wins (conceptually the most adaptive / evolutionarily fit).
3. Computer programs serve up webapps today, but they also run utility companies, dams, nuclear arsenals, factory production floors, automated car behaviors, and many other places. If an "agentic" AI has a single-minded goal that has death of all humans as a side effect, we at least want an off switch available.
2. Why would there be competition for resources, assuming there are enough resources for the “AGI” to run in the first place? This seems like a far-fetched hypothetical raised in service of further anthropomorphizing what is decidedly not a person or a mind.
3. LLMs do not have goals and are not minds.
Let’s stop attributing human-like qualities to statistical models.
2. One only needs to look at github going down due to agentic commits overload or data center buildout plans to see that scarcity for resources is present. An economy has no mind and is made up of the decisions of millions to billions of people and, now, agents attempting to perform on behalf of those people.
3. A bare transformer-based language model does not possess persistent goals in the ordinary agentic sense. But deployed agents can exhibit goal-directed behavior because the model is embedded in a harness that supplies an objective, context, tools, state, and an execution loop.
I've found that most regular users don't anthropomorphize LLMs in a strong sense ("AI boyfriend/girlfriend" aside), many in fact do expect agents to make human-like decisions -- which results in very unstable outcomes.
In short - goal-directed behavior does not require that the supporting system be a person/mind/conscious entity.
2. I thought you were saying the resources that the LLM uses to run were constrained, so I’m sorry for the misunderstanding there.
3. Yes I understand that we use RL to tune post-training. The (huge) difference between this and a human mind is that the LLM can’t develop a dangerous “single-minded goal” on its own, at runtime; it must have been trained to do so. If someone has post-trained an LLM to do something that has an illegal action as its side effect, that person/company/whatever has committed a crime and should be prosecuted. The solution here is legal, not technical.
Reading this little essay started out normal, but soon felt like a look into a disturbed and worrying mind, and if you find yourself taking it at face value, I urge you to step away from chat bots and spend some time with friends and family.
Not one North Star. Not two. Just three! OpenAI broke the North Star record!
With this evidence of AI slop, why did you not label this fluff piece as AI generated for the EU? You are violating laws.
Is there any place there is any evidence of AI being so useful or hopeful or good, anywhere other than code? As a reading machine it is impressive but it's judgement is not alien, it's just not good. IMO.
Does the title leap out anlt anyone else? James Martin's After the Internet: Alien Intelligence (2001) was an incredibly fun read, about expert systems and AI being inscrutable weird new varieties of intelligence, that familiarity would recognize one moment and be freaked out about/alien the next. I owe a re-read given how often I cite it, to recheck, but, I feel so primed from a much younger me having had that experience so long ago.
“And, in line with Ray Kurzweil’s predictions from the end of the XXth century , we now find ourselves at the moment in history of computing where machine intelligence is starting to exceed that of humans in transformative ways.”
Clicking the (pretentious sounding “XXth century”) link to Kurzweil’s predictions reveals the following:
“By 2019 a $1,000 computer will at least match the processing power of the human brain. By 2029 the software for intelligence will have been largely mastered, and the average personal computer will be equivalent to 1,000 brains.“
The first prediction passed 7 years ago and was decidedly not met. The second only has three more years to go, and I don’t think any respectable scientist or programmer would say that the average personal computer is anywhere close to the power of a single human brain, let alone 1000.
This is pure marketing garbage from a company desperate to keep itself alive.
Do I fully endorse everything the people holding guns to my head are forcing me to do to stay alive? Definitely not, but I’ve decided that for now, living to fight another day remains worth it
It is kind of strange to see this sentence, when OAI's definition of what AGI is has been watered down throughout the years.
> I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established.
Read: Please play by our rules, so we can be the first.