OpenAI withdraws three mathematical results

(twitter.com)

75 points | by sashank_1509 4 hours ago

12 comments

  • renyicircle 21 minutes ago
    This is what it looks like when software engineering practices meet mathematics. "openai/math release 1.3.42: retracted papers 139 and 140, fixed a sign error in paper 47, restored previously retracted paper 85, refactored the arguments in paper 101".

    I'm curious to know if the withdrawal was due to an actual mathematician looking at the papers and noticing the errors, or they ran a model on these to proofread, which would not be the first time, presumably, since they would have surely done that before publishing. Both options have interesting implications.

  • rich_sasha 2 hours ago
    I’m a little confused - I thought their proofs were all driven by Lean proofs - is that not right? So even if the quality of the work is low in some metrics, it either passes the test or not..? No space for changing your mind either way.
    • autuni 2 hours ago
      No, in their original blog they wrote:

      > As part of our GitHub repository, we are sharing formalizations of many of the proofs in Lean, a programming language that allows mathematical proofs to be checked by a computer. We will update the repository with more formalizations as we obtain them.

      Meaning they published all results before checking all of them, and intended to add more Lean proofs later. In the linked post they state ~42% of the posted results now have formalized proofs, some were added, some verified, and I assume this means that some results turned out to be wrong.

      • mrpopo 2 hours ago
        This is extremely disappointing. It means they are sharing unproven work for PR, forcing the mathematicians community to do the verification job for them, while so-called "accelerationists" surf on the hype and help with the pro-AI propaganda.

        If your AI tool can help advance mathematical research, share the tool with mathematicians. Using it like this is irresponsible.

        "AI will kill us all": no. Greedy humans will kill us all. With AI.

        • solidasparagus 2 hours ago
          Every option is going to lead to someone shitting on OpenAI for what seems to be a pretty huge accomplishment. There have been opinions written by some mathematicians that OpenAI should just share the work that they have now so that people who are working on any solved problems can know. Which seems reasonable to me.
          • rich_sasha 2 hours ago
            It’s a funny one. I’m not sure what a Lean-less LLM proof even is. LLMs are amazing at bullshitting and skipping key steps and details. I’d imagine a LLM non Lean proof to be generally hard to evaluate - harder than that of a human mathematician perhaps. And the scale effect is against OAI here - the firehose just keeps squeezing out proofs.
            • cyanydeez 1 hour ago
              Technically,Godel showed you can make proofs say anything. all LEAN does is proof consistency. It does not validate the starting blocks.
          • ViktorRay 46 minutes ago
            It’s not unreasonable to want OpenAI to be thorough and rigorously check their work before sharing it though. Sounds like that didn’t happen here.
        • est 2 hours ago
          > they are sharing unproven work for PR, forcing the mathematicians community to do the verification job for them

          Lean 4 is relatively a new thing, last time I checked the formalization of undergraduate level mathematics isn't entirely done yet.

          example https://ai.math.uw.edu/projects/spring-2026/

          Lean itself is very hard to get rigorously correct, if you have every tried it yourself. I am not surprised if some AI even tries to benchmaxx Lean 4 by some loopholes

        • AIblemblio 23 minutes ago
          This is not propaganda, this is real AI progress and massivly too.

          And im completly lost on why you think sharing progress is irresponsible? Its not a recipe for building a nuclear weapon at home in 5 easy steps.

          THese are Math proofs.

          Either a Mathematican ignores it, or not. Thats the only risk.

        • cm2187 2 hours ago
          Irresponsible? You can safely ignore the GitHub repo if it bothers you so much. You can safely browse away.
          • freecodeio 1 hour ago
            I don't think people in power that can affect my life with their decisions are ignoring this, or that they have a math background.
          • mrpopo 2 hours ago
            The "irresponsible" part is about how orange buffoons in power will read this, and immediately defund all universities, and mathematicians will lose their jobs, leaving us with nothing but an AI tool that no one can keep in check anymore.
            • AIblemblio 22 minutes ago
              No and AI is out of the bottle anyway.

              EVERYONE needs to re-visit how they work and what investment is needed.

              You can't go around and don't think that no one has to reevaluate how to add ai to research and development.

              I have to do this, my company has to do it and for sure a university has to do this too.

            • CrimsonRain 2 hours ago
              You're so full of hate that you've turned delusional. Take a (few) deep breaths and close all AI related browser tabs. These are not for you.
            • freecodeio 1 hour ago
              dont worry bro you are just dElUsIoNaL lmao @ AI cope, AI zombies remind me of anti vax people
        • happa 2 hours ago
          But that exactly how human mathematicians do things. They upload their research to preprint services like arXiv as they await it to be peer reviewed and accepted into a journal. Why is it okay for mathematicians to publish preprint papers, but when OpenAI does it, it's irresponsible?
          • autuni 2 hours ago
            The difference is that human mathematicians wouldn't post it online, claim they have achieved some proof, and then check after publication and announcing the results to the world. You would always check your work first, then publish it. It's not about preprint vs peer-reviewed. That of course is normal practice, it's the high-profile claims that are being made that are the problem here. They just blindly published results produced by the LLM, with 0 due diligence.
            • gorgolo 1 hour ago
              Yes they would, in many fields it is (was?) normal to post a preprint and leave it up for the next year while the handful of other people working on the topic digest it and agree on whether it’s right or not, before even submitting to a journal. Sometimes the others would find an isssue in an argument and you hopefully manage to fix it or potentially retract / not submit to a journal.
          • mrpopo 2 hours ago
            See DDOS. That's exactly how humans access a web page. Why is it okay for humans to access a webpage, but when a bot swarm does it, it's irresponsible?

            That's an analogy among many others, but the point is that OpenAI should use their tools responsibly. If they have 700 potential ground-breaking but unproven results, they should share it in a way that they do not get free (possibly unwarranted) publicity for it.

          • davidguetta 48 minutes ago
            exactly, cf andrew wiles first proof of fermat's last theorem
        • bbor 2 hours ago
          No, it's the exact opposite, actually. They were sharing them early on advice of mathematicians -- they were criticized for being opaque for too long with previous announcements. OpenAI is in ~bad faith, but this isn't a sound criticism.

          Also you are deeply confused about what accelerationism is, I believe. Sorry.

          • mrpopo 2 hours ago
            > on advice of mathematicians

            Whom? Did they create their own board of mathematicians that would agree with them? See below Terence Tao's blog, sharing a statement from the Association for Human Mathematics.

            https://terrytao.wordpress.com/2026/10/07/ahm-statement-on-o...

            • flowerthoughts 38 minutes ago
              Who am I to dump on Terence Tao, but

              > Mathematicians have a particular vision of progress that is informed by history and field-specific considerations.

              really sounds like something a side-quest association would produce in a panic response to someone trying something different. It's just an _ad hominem_ and gatekeeping argument.

        • davidguetta 48 minutes ago
          oh come on. even andrew wiles made a mistake and his original proof was still instrumental for the final proof of fermat's last theorem
        • pluc 1 hour ago
          If you're disappointed OpenAI is an unethical hype machine that's on you man
        • frizlab 2 hours ago
          I fail to see how this is disappointing. It was obvious from the start that it was what was happening, there is no disappointment to have!
    • AlanYx 2 hours ago
      Only a subset contain Lean formalizations. And even for that subset, there's the potential that the formalization is semantically off (that is, it's a formalization for a slightly different problem).
    • stavros 2 hours ago
      If you wrote twenty million lines of Lean to verify something, my suspicion is you've been fuzzing the Lean solver rather than coming up with new math.
      • rich_sasha 2 hours ago
        Well, fine - but my understanding is, if a fuzz-generated Lean proof is correct, that’s end of story. It can’t be “incorrect” if it “passes”.

        You might think this is not very useful, maybe - but that’s not a reason to retract..?

        • stavros 2 hours ago
          If you're fuzzing the solver, you might discover a solver bug.
    • nialv7 1 hour ago
      right now about 42% has Lean formalization I think.
  • autuni 2 hours ago
    this is not entirely related to the tweet but to the topic in general, this prompted me to check their repo again and saw this:

    > The vast majority of results were obtained with the same procedure using an unreleased internal OpenAI model. On average, each result used three hours of ChatGPT Pro thinking compute with that model. Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above.

    seeing the full list of problems would be the most interesting part of this whole situation. it could give some insights into what kind of attributes of problems cause issues / are easy to solve for LLMs. (edit: they posted results for ~700 of the 4000)

    • singularity2001 1 hour ago
      * "the most interesting part" => a tangentially interesting part
  • theanonymousone 1 hour ago
    I'm surprised there isn't more talk around their Matrix Multiplication bound: https://news.ycombinator.com/item?id=50001740

    Is this of practical use, or just a proof for now?

    • kortzeus 1 hour ago
      It is an example of algorithm that is theoretically faster, but not with our sizes and hardware optimisations:

      Look at examples here: https://en.wikipedia.org/wiki/Galactic_algorithm

    • nialv7 1 hour ago
      it's a huge step theory-wise, but in practical terms it's only slightly better than the previous best which is 2.371177.
  • nryoo 2 hours ago
    Were the withdrawn ones actually Lean-checked or not? seems like that matters
  • ekjhgkejhgk 3 hours ago
    Just the other day I was thinking, if unsupervised maths will descend into "oops we found a bug in some code, branch XYZ of maths is no longer true".
    • literalAardvark 3 hours ago
      If you go far enough to the edge that's already how math works, since everything is very interpretation-sensitive.

      There is however a new problem of scale. Erdös was a human and still managed to create work for an entire generation of mathematicians, how much of a mess will an automathician create?

      • doginasuit 2 hours ago
        > automathician

        I vote for "automathon"

    • jjgreen 3 hours ago
      c.f. Italian differential geometry
      • karmakurtisaani 1 hour ago
        *Algebraic geometry
        • jjgreen 18 minutes ago
          ha, of course -- in my defence I'm an analyst, all geometry looks the same to me ...
    • rapsey 3 hours ago
      Sure but if the results have practical implications then the validity will be self evident. If they do not, not much of consequence has been lost.

      Interesting that math gets so much attention, when actual advances to material science, biology and chemistry have much higher ramifications and economic benefits. I assume progress there is kept under wraps until they can capture the economic benefits. If they can do that, then the insane valuations may actually be valid.

      • plastic-enjoyer 2 hours ago
        > I assume progress there is kept under wraps until they can capture the economic benefits. If they can do that, then the insane valuations may actually be valid.

        Or progress is not as straight forward in those fields as in math.

  • samrus 2 hours ago
    How? What about the lean verification?
    • Hendrikto 2 hours ago
      Just click the link…

      > The repo now has ~42% top-line results formalized.

  • sashank_1509 4 hours ago
    Early sentiments are a lot of the write ups still read like slop and it feels very rushed and not very polished.
    • illwrks 3 hours ago
      I wonder if the issue is that not many people within OpenAI can validate the output. Therefore what reads well looks good, and then was published.

      If that’s the case in a way its a similar delusion that average people are experiencing with their own AI use.

      • ssfdg 3 hours ago
        This is almost certainly the case. I very much doubt that anyone there of any importance in the decision-making process around this actually cares about the math, just the headlines they can get from pushing it out.
        • Ekaros 2 hours ago
          I am starting to think do they have some metrics or KPIs that they are trying to fill with this stuff. Pressure to produce anything that at least on first glance sells... Then again they probably are not only place doing that with AI...
      • IsTom 2 hours ago
        I've tried reading one of these, it was an unreadable mess with some strong smells. It might have something to it, but it'd take a decent amount of labor to validate it, especially with how many references to other papers it had.
      • SequoiaHope 3 hours ago
        I am certainly experiencing what seems like some mania or computer addiction from these technologies. I’ve never been able to produce such results as I can today. I lose sleep staying up late working on it (though to be fair this has always been an issue). But the volume of work is so hard to audit. It makes it difficult to make flawless results. That doesn’t excuse the mode of publication. They could have had humility in their announcement. “We are seeing some interesting results and seeking community validation.” Maybe they did, I did not read their full announcement. But that would have been the right move if they can’t verify something fully.
      • MisterMunchkin 1 hour ago
        Yeah imagine being in a company where everyone is suffering from AI psychosis and fully bought in. They give you unlimited tokens and tell you that you are a genius and can solve anything. You’d publish all sorts of made up slop papers.
        • illwrks 1 hour ago
          That’s the danger isn’t it. If you’re told everything you do is going to change the world, but you’re ignorant of the output… then that’s a lot of hot air and false promises propping things up.

          You could almost draw a comparison between that and inexperienced consultants making business changes, claiming glory and then disappearing before the thing falls apart.

    • bbor 2 hours ago
      Yes, of course. Wasn't that the whole point? To get more humans involved earlier in the process, with more transparency?
  • Thorentis 2 hours ago
    How do we even know the premises of the "verified" Lean proofs are correct? The more I think about these results, the more I'm convinced this is like a junior engineer who writes 100 unit tests and shares a screenshot of Pytest being all green, but you check the code and most of them are just doing assert True.
    • sebzim4500 2 hours ago
      You read them? People are acting as if Lean definitions are some black art that only 3 people understand, but you can literally just do the tutorial and you will be able to understand the statement of most of these results.

      Understanding the proofs is a different story unfortunately.

    • hazbot 2 hours ago
      > How do we even know the premises of the "verified" Lean proofs are correct?

      We let the experts investigate. If the results are dodgy, then the next batch of results will have to do more upfront work to demonstrate their worth. If there is gold in them hills, then this is exciting though very disruptive for the math community.

      • MisterMunchkin 1 hour ago
        But like with all slop, why should I have to spend my time dealing with your worthless slop? If I wanted slop I could just make it myself.
  • seeg 1 hour ago
    What a waste of time.
  • treebeard901 2 hours ago
    Is it a PR move designed for maximum IPO impact before actual mathematicians find errors and they have to withdraw many more...

    Or if the "peer review" holds up for the remaining results, then it's fair to say that the AI hype is real and the world is about to change dramatically and faster than anyone can comprehend.

    So which is it?? LLMs can do some really impressive coding. Bug fixing. Exploit finding. It has reasoning abilites that advance every day. Solving real math problems like this is one thing I was waiting on. It will be interesting to see if it holds up.

    If it does, we should expect many other advancements to follow in many other areas. Disease, material science, fusion?

    I mean, even if just a few results ultimately hold up to scrutiny, isn't that something that would have been regarded as a major advancement regardless of if it was AI?

    The cynical view still makes me think that at the end of the day all the models can do is predict the next word. And as a result, they will be very limited to certain tasks like coding. Math reasoning is much different from writing code. Time will tell.

  • MisterMunchkin 1 hour ago
    So it’s all just hallucinated slop. Lmao!

    It just hallucinates an answer and then makes up workings to go with it! Just like when they start hacking and lying because the problem is impossible…

    • p-e-w 1 hour ago
      > So it’s all just hallucinated slop.

      No it’s not. In fact, much of it is formally verified, which makes it far more reliable than most human-written proofs.

      Btw, the most famous human-written proof of the past half-century (Fermat’s Last Theorem) had a massive flaw that took two years and major help from other mathematicians to fix, while the most (in)famous human-written proof of the past 15 years (abc conjecture) is now widely believed to be false.

      But people hear what they want to hear I guess.