My main gripe here is the lack of transparency around the total experiment and construction. I doubt that they simply pointed their model at these ten specific problems alone and gave the model one shot; therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.
I want to know:
1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up?
2. How many attempts did you give the model at solving these problems?
3. How expensive was the harness, e.g. did the model have access to a job cluster?
Also, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar?
Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
> Also, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar?
A couple small ones that I've seen (example here [0]), but not anything of the magnitude that OpenAI and Anthropic have put out. Likely just related to token limits.
> Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
I think their output has reached a level that precludes this possibility, but I don't have any hard proof.
It's a very important clarification if it took $2000/problem on 20 problem attempts or on 1,000 problem attempts for each successful one. That may be the deciding factor on whether or not it's economically viable to replace a mathematician with a ChatGPT subscription.
Yeah fair I concede that this is somewhat crucial information. The parent seems to write it in a tone that suggests deliberate misleading “lack of transparency” etc.
> claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system’s contribution and the nature of genuine human intellectual work.
AI has no self-awareness. It's a tool. When you assemble a furniture using a screw driver, the torque force interacts with the molecular forces inside the metal and miraculously it transfers the force to the screw though a clever geometry design, communicating the force to the screw to turn it in a certain way.
Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.
A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.
[0]: Not one of the proofs in the linked article, but from OpenAI too.
The human provides the intention and the ability to appreciate the output. Tools do “heavy lifting” all the time, but we still primarily credit the humans who use them precisely because they made the choice to use them.
Provability is just going the way of computation. John Napier had to manually compute logarithm tables over decades and was recognised for his work; now that same work could be performed by a 10 year old with a calculator in an evening.
Your brain also is physical. Electrochemical gradients flow between physical molecular constructs. Isn't it just chemistry? Do you attribute it to physics or some whole-is-greater-than-the-parts idea?
> Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.
When the tool is a 3D printer, or any CNC system really, you bet I attribute a build to it.
I could also attribute the operator; there is no contradiction, it's a free choice, just like saying "I am in Berlin" does not contradict "I am in Germany".
A better analogy would be a manufactured object, say 3d printed for simplicity. The 3d printer is given an input, and an object manifests itself after some time. We say that the creator of the object is the person turning on the machine, sending the data, and collecting the object. Not the machine itself.
I would say „I made this gadget with my 3d printer, but the designer is someone else (I found the model online)“. The intent, the drive, the action comes from the human
Dunno about the parent comment, but I personally interpret the concept as having a hidden representation of self that is continually tended to. This implies statefulness, which models are intentionally not at inference time (*).
(*) Even if we hack around this and just do the usual trick of simply laundering statefulness to a higher level, in this case the context window being fed in, I fail to identify (**) a representation of its own state in these bodies of text that it'd be meticulously maintaining. I further fail to identify how it could be hidden or maintained, considering I control like half of it. The best you could ascribe it is a meticulous maintenance of a persona the user is talking to, but then that doesn't necessarily represent the model's internal state, the same way my own words here aren't doing so either.
You'll sometimes catch models mixing up who's who and how many who-s there even are for example.
(**) I did wish for something hidden though, so maybe it's just concealed? The same way people can encode a lot more of their emotional and mental state than normal into text if they read and write a lot of it, I'm aware of research that suggested the same for LLMs, albeit I cannot cite it. Maybe those phrasing signatures are just alien to me and will never pop out. Either way, I'd expect researchers to stumble upon this during interpretability studies, and either they haven't, they have but it wasn't popsci adopted, or they're keeping awfully tight lipped about it. If you know of anything like this, your turn now, would be happy to learn.
On the token limits etc - one assumes that OpenAI et al are able to “hire expert in field, and let them spend the equivalent of a million dollars of tokens” because they are not actually selling their complete compute 24 hrs a day, so the cost internally is a negligible (ish) electricity bill.
Which is very suggestive - if after everything they are not fully loaded then the next gazillion data centres being built look unlikely to be needed.
I don’t feel the existential dread of mathematicians is correct. It seems to me in fact these results are bringing math mainstream. I now personally look forward to the interpretations and discussions of the significance of such results by human mathematicians.
Now I understand that it’s mostly the super stars benefitting from the increased attention. Folks who are less established don’t share in that glory. But on the other hand it seems like an exciting time to go even deeper for in various specialties of math by deciding where to focus these powerful tools. For every conjecture defeated some seven or eight new ideas open up. Our path through that combination will be set by creative and curious human mathematicians.
Chess found a way to remain and grow with humans, and these other fields will too.
The old way of establishing career credibility is being destroyed, for better or worse. Accomplishments that used to be career-defining are hard to distinguish from AI, and correlate more with access to compute. Think about Bill Gates's math paper he wrote in college. That kind of thing is gone now as a path to credibility. There's still competitions and grades, but the diversity of paths is going away. Maybe new ones will open up. This is a competitive advantage for old people who have credible pre-2025 accomplishments they can point to.
The chess analogy is awful. If you simply want to know the answer to a chess problem, give it to the engine. Chess only lives on because it's a competition between humans to test their skill (just like bicycles, cars, trains didn't eliminate foot races) ... the computer is largely factored out, but not entirely -- people train with the computer, use it to check whether they played correctly, ... and they cheat. A lot. Thus there are more and more sophisticated mechanisms to detect and prevent cheating.
If you translate that to math, then all you get is math competitions, not math as a career. Of course the translation isn't nearly exact ... there's a lot more room for professional mathematicians because the math space is far more vast than the chess space and can't generally be cranked out mechanically (we have proof).
As in chess and go and also coding for the past ~year there are two groups of people: the disappointed and the enthusiastic. The disappointed are sad that they lost their advantage and that the craft they honed for years or decades has rapidly lost its value; the enthusiastic are excited about the future and what computers can bring to their domain and how it will evolve. I’m a bit of both if it comes to programming, more enthusiastic than disappointed, but also more than a bit terrified about the pace of it all. I imagine that’s how Kasparov felt back then, that’s how Lee Sedol felt and now that’s how Terry Tao feels.
The most disappointed folks will simply drop out, but the enthusiastic ones will keep going and with luck make up for the ones who decided to quit. Chess and go certainly went this way.
A fundamental difference being that no one was actually paid to find good moves in chess and go like they are to solve math problems and write code. You're comparing the digital camera and the automobile.
Every time someone makes a comparison to chess I die inside. Chess is a spectator sport primarily funded by a few eccentric billionaires. Players artificially constrain themselves in timed environments knowing that they will never be able to produce better moves than a smartphone because a select few people find it interesting. Only ~30 top professionals actually make enough money to have a full career playing chess, maybe a few hundred more can sustain a meager lifestyle with coaching gigs. I shudder to imagine what will happen to the tens of thousands of non-Fields medalist caliber mathematicians if math goes the way of chess. Perhaps Terence Tao and a few other famous mathematicians will be funded by Peter Thiel to report on how well humanity can keep up with the machines? How do you expect any mathematician to be optimistic about this comparison.
I'm enjoying learning about these hard problems, but this line about credit made me chuckle:
> We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness
Offering to take responsibility for the correctness of a proof written in Lean feels like volunteering to be the fall guy in case someone finds a flaw in basic arithmetic, no?
I'm not an expert at it myself, but my understanding is there are numerous ways to "cheat" in a Lean proof (via `sorry` and similar). They're taking responsibility for fully verifying that none of these cheats were used (and that the theorem statements themselves were all correctly formalized.)
It’s more than you get from free software - you get no proofs, no warranties and any responsibility of its authors are their pure good will. Reminder lean proofs are software!
Not much point to pure math being kept secret, in all honesty. There isn't really industrial value, its only purpose (to them) is showing off their model's capabilities. More realistically they'll just stop paying for it.
Edit: Oh, are you suggesting they just use it to privately improve their models? I imagine a few more correct proofs would have a very marginal benefit, if any. Also, they'll probably just get extracted, meaning it still gets out but OpenAI doesn't get to fancily announce it themselves.
This isn't really a productive way to think about these things, IMO. It's quite possible it would take hundreds of years for any specific group of PhDs to solve them. Or one individual PhD could have the correct flash of insight and solve it in a month. There's absolutely no way to predict this, besides trying to gauge the apparent simplicity of the proof or counterexample (which is likely to be misleading). Until someone actually runs an experiment like this it's not a viable metric.
I understand and I am not trying to deny the impressiveness or the velocity of AI in general. But at some point we have to ask how much do we trust the labs at face value without much transparency of how they got to the results when there is trillions of dollars on the line.
Capabilities of this sort have already been demonstrated by independent parties, and models have consistently gotten better at this. Yes, insinuating that mathematicians and scientists are secretly solving decades-old problems on OpenAI's behalf in as insane conspiracy theory.
Given that OpenAI pays their employees with stock surely a breathtaking number, but not a very meaningful number now that the infrastructure is in place and the models are trained. AI could never get better and it would still be incredibly disruptive.
I want to know:
1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?
Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
A couple small ones that I've seen (example here [0]), but not anything of the magnitude that OpenAI and Anthropic have put out. Likely just related to token limits.
> Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
I think their output has reached a level that precludes this possibility, but I don't have any hard proof.
[0]: https://www.reddit.com/r/math/comments/1uxj3cy/after_openais...
My point here is to not snark. But there should be some level of self skepticism that doesn’t warrant an RCT theatre.
AI has no self-awareness. It's a tool. When you assemble a furniture using a screw driver, the torque force interacts with the molecular forces inside the metal and miraculously it transfers the force to the screw though a clever geometry design, communicating the force to the screw to turn it in a certain way.
Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.
A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.
[0]: Not one of the proofs in the linked article, but from OpenAI too.
Provability is just going the way of computation. John Napier had to manually compute logarithm tables over decades and was recognised for his work; now that same work could be performed by a 10 year old with a calculator in an evening.
When the tool is a 3D printer, or any CNC system really, you bet I attribute a build to it.
I could also attribute the operator; there is no contradiction, it's a free choice, just like saying "I am in Berlin" does not contradict "I am in Germany".
What is your mechanistic model of self awareness that yields this conclusion?
> It's a tool
Does your model suggest that tools can't have self awareness?
(*) Even if we hack around this and just do the usual trick of simply laundering statefulness to a higher level, in this case the context window being fed in, I fail to identify (**) a representation of its own state in these bodies of text that it'd be meticulously maintaining. I further fail to identify how it could be hidden or maintained, considering I control like half of it. The best you could ascribe it is a meticulous maintenance of a persona the user is talking to, but then that doesn't necessarily represent the model's internal state, the same way my own words here aren't doing so either.
You'll sometimes catch models mixing up who's who and how many who-s there even are for example.
(**) I did wish for something hidden though, so maybe it's just concealed? The same way people can encode a lot more of their emotional and mental state than normal into text if they read and write a lot of it, I'm aware of research that suggested the same for LLMs, albeit I cannot cite it. Maybe those phrasing signatures are just alien to me and will never pop out. Either way, I'd expect researchers to stumble upon this during interpretability studies, and either they haven't, they have but it wasn't popsci adopted, or they're keeping awfully tight lipped about it. If you know of anything like this, your turn now, would be happy to learn.
Which is very suggestive - if after everything they are not fully loaded then the next gazillion data centres being built look unlikely to be needed.
Now I understand that it’s mostly the super stars benefitting from the increased attention. Folks who are less established don’t share in that glory. But on the other hand it seems like an exciting time to go even deeper for in various specialties of math by deciding where to focus these powerful tools. For every conjecture defeated some seven or eight new ideas open up. Our path through that combination will be set by creative and curious human mathematicians.
Chess found a way to remain and grow with humans, and these other fields will too.
If you translate that to math, then all you get is math competitions, not math as a career. Of course the translation isn't nearly exact ... there's a lot more room for professional mathematicians because the math space is far more vast than the chess space and can't generally be cranked out mechanically (we have proof).
The most disappointed folks will simply drop out, but the enthusiastic ones will keep going and with luck make up for the ones who decided to quit. Chess and go certainly went this way.
> We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness
Offering to take responsibility for the correctness of a proof written in Lean feels like volunteering to be the fall guy in case someone finds a flaw in basic arithmetic, no?
https://leanprover.zulipchat.com/#narrow/channel/270676-lean...
Edit: Oh, are you suggesting they just use it to privately improve their models? I imagine a few more correct proofs would have a very marginal benefit, if any. Also, they'll probably just get extracted, meaning it still gets out but OpenAI doesn't get to fancily announce it themselves.
https://x.com/polynoamial/status/2083470822258467194