Sharing AI progress in mathematics

(openai.com)

416 points | by OfficialTurkey 3 hours ago

67 comments

  • zone411 2 hours ago
    A quick check shows that this list claims to fully solve 90 of the top 500 open problems in math (https://proofatlas.ai/open-problems/).

    The highest ranked would be:

    | 22 | Hilbert’s tenth problem over ℚ |

    | 29 | Unique Games |

    | 31 | Anderson-model extended states |

    | 37 | Spacetime Penrose inequality |

    | 48 | Nonexistence of Landau–Siegel zeros |

    | 52 | Baum–Connes |

    | 78 | Abundance |

    | 80 | Hadwiger |

    | 87 | Bose–Einstein condensation |

    | 92 | Two-dimensional entanglement area law |

    • magicalist 1 hour ago
      > the top 500 open problems in math

      At least put a disclaimer for the ad for this site, and maybe disclose how you came up with a total ordering for "top" open problems (vibes)?

      > How problems are ranked. LLMs compare pairs of problems. A reliability-weighted model combines those judgments into the ranking, with calibration across model families. The model-family weights are OpenAI 1.00, Claude 1.00, GLM 0.95, and DeepSeek 0.90. These are modeling choices, not measured probabilities of correctness.

    • optimalsolver 1 hour ago
      Was anyone in the math community aware of the inbound tsunami at the beginning of the year?
      • AnotherGoodName 4 minutes ago
        Lots. To give an example Terrance Tao was lambasted skeptics on this site for stating it in 2024.

        https://unlocked.microsoft.com/ai-anthology/terence-tao/

        " I expect, say, 2026-level AI, when used properly, will be a trustworthy co-author in mathematical research, and in many other fields as well.

        Then what? That depends not just on the technology, but on how existing human institutions and practices adapt. How will research journals change their publishing and referencing practices when entry-level math papers for AI-guided graduate students can now be generated in less than a day—and with the far better accuracy of future AI tools? How will our approach to graduate education change? Will we actively encourage and train our students to use these tools?

        We are largely unprepared to address these questions. There will be shocking demonstrations of AI-assisted achievement and courageous experiments to incorporate them into our professional structures. But there will also be embarrassing mistakes, controversies, painful disruptions, heated debates, and hasty decisions."

        He's pretty damn smart that guy.

        • mianos 1 minute ago
          > He's pretty damn smart that guy. This is probably the understatement of the year. I am literally ROFLing.
      • thrance 55 minutes ago
        I predicted, over 2 years ago, that theorem proving would fall way before other problems that people believe are harder.

        https://news.ycombinator.com/item?id=41072330

    • k2xl 1 hour ago
      Result 003 (Quasi-Riemann Hypothesis), from my reading of mathematicians reactions, is a landmark discovery.
      • adgjlsfhk1 1 hour ago
        yeah if it holds up, is the biggest result in number theory in 200 years
        • JoshuaZ 15 minutes ago
          Number theorist here. This is a massive big deal, and would likely be a Fields Medal for a human if a human had done it. But it is an exaggeration to say it is the biggest result in 200 years. At a minimum, it is hard to argue that it is a bigger result than the proof of the prime number theorem in 1896 (which this is a strengthening of), or Riemann's original 1859 paper where he laid out the zeta function and its analytic importance, or Dirichlet's proof of infinitely many primes in arithmetic progressions which is the late 1830s.

          But yeah, this is still a very big deal. Among other things, it will drastically improve all sorts of Rosser-Schoenfeld type results for the PNT and that's just a start. For comparison, I have a paper form 2018 where this result would cut 3 pages out and make the full result cleaner and much tighter, and there are likely hundreds of papers like this.

    • anematode 2 hours ago
      Dear lord that website is laggy
      • manquer 1 hour ago
        At this rate solving P=NP is going to be easier than solving front end perf …
        • m_mueller 1 hour ago
          wait, maybe this is the same problem....

          with non-polynomial side being represented as the frontend programmer's constant need for more performance to do the same task...

        • echelon 21 minutes ago
          Please let P=NP, Please let P=NP

          Whomever is running this simulation, please.

      • vector_spaces 1 hour ago
        Not to mention it's got that signature Claude Clutter UI design
        • p-e-w 1 hour ago
          Interesting how perceptions differ. My first thought was “Wow, that’s well designed for a math website”.
        • zone411 1 hour ago
          Except that Claude wasn't used.
        • landdate 1 hour ago
          [dead]
    • zone411 1 hour ago
      By category in the top 500:

        +----------------------------------------------------+------+---------+-----------------+
        | Category                                           | Full | Partial | Matched / total |
        +----------------------------------------------------+------+---------+-----------------+
        | Geometry and topology                              |   25 |       7 |         32 / 74 |
        | Algebra, representation and category theory        |   17 |       2 |         19 / 53 |
        | Analysis and PDE                                   |   11 |       6 |         17 / 40 |
        | Number theory and arithmetic geometry              |    4 |      13 |        17 / 117 |
        | Probability, ergodic theory and dynamics           |   11 |       5 |         16 / 37 |
        | Combinatorics and discrete geometry                |    7 |       2 |          9 / 34 |
        | Theoretical computer science                       |    4 |       4 |          8 / 57 |
        | Mathematical physics                               |    5 |       1 |          6 / 19 |
        | Applied and computational mathematics              |    2 |       2 |           4 / 8 |
        | Quantum information and computation                |    2 |       1 |          3 / 17 |
        | Cryptography, coding, information and optimization |    1 |       1 |          2 / 26 |
        | Logic, foundations and set theory                  |    1 |       1 |          2 / 18 |
        +----------------------------------------------------+------+---------+-----------------+
        | Total                                              |   90 |      45 | 135 / 500 (27%) |
        +----------------------------------------------------+------+---------+-----------------+
  • xanderlewis 2 hours ago
    As Kevin Buzzard recently said:

    > In a 2020 piece in the Notices of the AMS, I asked the following question: “If one human had an understanding of all of modern pure mathematics simultaneously, how much further would they immediately be able to see?” Six years later we are beginning to understand the answer to this question.

    • dang 1 hour ago
      https://xenaproject.wordpress.com/2026/10/01/to-grieve-or-no...

      Discussed here:

      To grieve, or not to grieve? - https://news.ycombinator.com/item?id=49919676 - Oct 2026 (156 comments)

      • oliculipolicula 34 minutes ago
        >I believe that the optimal thing to do ... is to let the machines loose, see what happens, and then begin the journey to where they have stopped. Things are currently moving fast. They cannot move fast forever. But if we get on board now then they will take us to extraordinary new places. And after we have arrived, the new adventure will begin.

        I'm relieved that this time around, they have provided partial reasoning traces and prompts for a small number of problems. Do they now also share data with the other model providers..

        • patcon 13 minutes ago
          It's beautiful, but the animals are not thinking this about us.

          They structurally cannot understand what we are doing at the place where we hit our ceiling. Only with our highest technology (well beyond their understanding) do we have the tools to go back for them, and try to bring them along and interface better with us (re: recent work in animal communication)

    • outworlder 13 minutes ago
      Similarly, there are probably many ideas that have not seen the light of the day because they require deep correlation between seemingly unrelated fields. It is not every day that we get a Isaac Newton or Leonardo da Vinci.
    • anon-3988 2 hours ago
      The other crucial part to this is the ability to actually encode and test the theorem (via Lean). Otherwise, we would be swarmed with a billion lines of theorems that no one will be able to ever understand and verify anyway.
      • senderista 2 hours ago
        If you think AI-generated Lean proofs are unreadable, imagine Opus 5 generating informal proofs.
        • ijidak 2 hours ago
          I think OP is saying Lean does indeed help.
  • NotOscarWilde 2 hours ago
    As a TCS/scheduling person, this one is definitely of lesser importance than UGC, but it has been an open problem since the book of Garey and Johnson in 1979:

    A Polynomial-Time Algorithm for Three-Machine Unit-Job Scheduling [1]

    Since some people talk about small numbers that pop up in integer multiplication results, here a completely different number appears:

    Theorem 1.1. Let an explicitly listed finite directed acyclic graph specify the precedence constraints on n >= 1 nonpreemptive unit-length jobs on three identical machines. There is a uniform deterministic algorithm that constructs a feasible schedule of minimum makespan. Given also an integer deadline 1 <= T <= n, it decides feasibility exactly and returns a schedule whenever the answer is affirmative. Both tasks can be performed in O((L + 2)^150020) steps on a deterministic multitape Turing machine, where L is the total binary input length.

    That is some crazy exponent -- plus an interestingly old computational model to boot; not something that is natural to most of us. I have no capacity to check its correctness today, but I hope it is true purely for the exponent.

    [1]: https://github.com/openai/math/blob/main/preprints/A-polynom...

    • keeganryan 4 minutes ago
      The largest I've seen [1] is an exponent of 10^12, which I suppose still counts as polynomial time.

      I'm sure all of these super small or large constants will improve over time, but it's still amusing. It is entertaining to see the exponents directly rather than have them hidden as n^c or epsilon or O(1).

      [1]: https://github.com/openai/math/blob/main/preprints/Determini...

  • prideout 2 hours ago
    This includes a proof of Barnette's Conjecture, which is one of the graph theory conjectures that I tried attacking with SOTA models a few months ago. I like it because it is easy to understand with a basic knowledge of graph theory. I spent quite a bit of time on it and failed. Their proof looks approachable at first glance.

    https://github.com/openai/math/blob/main/preprints/Paired-st...

    • an0malous 2 hours ago
      Any idea what made OpenAI successful where you weren’t?
      • kulahan 1 hour ago
        Trillions of dollars might be a bit of an advantage.
        • p4ul 3 minutes ago
          Noooooo!!! I mean... does money matter here?!? Crazy! /s

          We are in danger.

      • seanmcau 1 hour ago
        Probably the model OAI used that is strictly better than whichever SOTA - 3 months model OP used?
      • sebzim4500 1 hour ago
        Presumably it's mainly the better model, I don't see much evidence of a particularly advanced harness based on the reasoning traces that they provided.
        • an0malous 1 hour ago
          That’s what I was wondering. Thanks.
      • zzzeek 48 minutes ago
        I'm going to guess the ability to hold a million individual details in an attention space at once, compared to the typical human capacity for about six or seven
      • whamlastxmas 1 hour ago
        Their internal model is allegedly like 4x as capable as the publicly available ones
      • ForHackernews 1 hour ago
        They ingested all of his sessions with their SOTA models from a few months ago. ;)
        • ndriscoll 23 minutes ago
          Crazy how all these problems were simultaneously on the precipice of being solved before ScoopenAI (I came up with that) stole their work. The rate that mathematicians have suddenly made progress is only bested by the rate that OAI has ramped up their theft haha it's nuts. Good thing it'll never get better at explaining ideas though (since they're all stolen) so there'll still be that.
        • digitaltrees 1 hour ago
          The fact that this is plausible should be deeply disturbing and disqualifying for openAI. The fact that they may prevail and win is a travesty of our failed system.
          • zeroonetwothree 1 hour ago
            What does "win" mean? There is no prize for this, and having someone discover a proof benefits us all.
            • vuurmot 1 hour ago
              The prize is a tenure for the researcher, and in OpenAI's case, a higher valuation when they IPO?

              In this case, the tenure is gone, and OpenAI has increased their valuation

            • breezybottom 1 hour ago
              Sure there is. A job, tenure, professional respect, Fields medal.
          • fnordpiglet 1 hour ago
            Disqualifying for what? If you develop a proof you aren’t competing for something, you’re expanding the frontier of knowledge. It’s a binary state of the world, either it’s proven or not. Prevail and win what exactly?
    • TeeWEE 22 minutes ago
      Did you validate the proof? Who did?
  • davegoldblatt 1 minute ago
  • schleck8 1 hour ago
    Levent Alpöge (Anthropic mathematician) comment on the significance:

    > Sure, mathematical history features a lot of incredible developments, like the invention of proof, zero, or the computer, and on the great problems our progress has been over timelines measured in decades or centuries. Obviously this technology didn’t appear today, but blurring our eyes a bit to combine the past ten years, with today a measurement of those developments, there is nothing comparable.

  • enoether 3 hours ago
    Unique Games Conjecture [0] is a seminal conjecture in Complexity Theory, and is an underlying assumption for many, many inapproximability results. A valid proof is a big deal!

    [0] https://en.wikipedia.org/wiki/Unique_games_conjecture [1] https://github.com/openai/math/blob/main/preprints/The-Uniqu...

    • inkysigma 2 hours ago
      I also don't think there was general consensus on which way this would resolve prior to this (or is that a little out dated?) unlike some of the other major problem resolutions. I heard rumors that there would be a big result in TCS and speculation it would be UGC that or P neq PSPACE but I'm still a bit shocked.
    • amluto 2 hours ago
      I'm really glad that OpenAI is formalizing these things, because I'm not convinced that their current internal frontier models are particularly good at writing down their thoughts in English. From the (probably awesome) Unique Games Conjecture/Theorem paper, the first two sentences of section 1.1 start to define the problem:

      > A Unique Games instance has a finite vertex set, a finite alphabet K, and a nonempty list of oriented constraints e = (u_e,v_e,π_e), where π_e is a permutation of K. A labeling a satisfies e when a(v_e) = π_e(a(u_e)).

      I'm sorry, what? I admit it's been quite a few years since I've thought about the Unique Games Conjecture, and I never dug that deeply, but this part is very, very elementary graph theory and notation. So let's unpack it.

      1. e is maybe a name of a list.

      2. The elements of that list are tuples, where each tuple is (a vertex, a vertex, a permutation). So e indexes into the list and u_e is the source vertex for the e-th constraint in the list called e. Thanks.

      3. a is a labeling. I'm fairly confident that, by "a labeling", they mean that e is a function from vertices to colors, where the colors are the elements of k.

      4. That vertex coloring a satisfies the list e, when, for, um, an index e into e, a(v_e) = π_e(a(u_e)). But this isn't for all e, it's for some e, and the goal is to count them.

      So maybe e isn't a list? Maybe e is a constraint that is represented as a tuple, so e = (u_e,v_e,π_e) and u, v, and π aren't sequences at all but are, in fact, the trivial unpacking functions that unpack the pieces of the tuple.

      Reading this stuff is pointlessly painful, and it's extremely easy to make mistakes when being sloppy like this.

      If this were my paper, or if I were trying to train a model to write math, I'd want something like:

      A Unique Games instance has a finite vertex set V, a finite edge set E = (V × V), a finite alphabet K of possible vertex colors, and a nonempty list of oriented constraints. Let Π be the set of permutations of V. Each constraint e is a tuple in E × E × Π, where we write u_e ∈ E for the first element, v_e ∈ E for the second element and π_e ∈ Π for the third.

      A vertex coloring a : V → K satisfies e when a(v_e) = π_e(a(u_e)).

    • impossiblefork 2 hours ago
      Yeah, that's one of the big things of TCS. I think I see that as bigger than that Millenium Prize problem.
      • davemp 37 minutes ago
        TCS being theoretical computer science? I have not seen that acronym before.
        • jhanschoo 2 minutes ago
          Yes, TCS is theoretical computer science, I commonly use that acronym too.
    • gregdeon 2 hours ago
      This was the biggest highlight for me as well. Astounding...
  • gizmodo59 3 hours ago
    This is significant progress and released without all the drama. Some very important progress in Reinmann, Hodge and unique games theorem. Point the repo to your agent and ask for the significance! In a way this is probably 50-100 years of math progress by humans
    • traes 2 hours ago
      Not to pick on you specifically, but as someone who spends a lot of time unproductively reading AI math discourse it's truly shocking how incapable all the supposed math enthusiasts are of spelling Riemann.
      • xpct 2 hours ago
        I just did a quick search on this and apparently the misspellings are German surnames as well:

        https://en.wikipedia.org/wiki/Reimann

        https://en.wikipedia.org/wiki/Reinmann

        • traes 2 hours ago
          I just don't understand how it happens. If they had ever taken an intro to real analysis class they would learn to spell his name. If they were just parroting what an AI told them... shouldn't they still just say his name? An individual could just be dyslexic or mistaken but it seems to be a substantial volume. I guess they just don't care enough about it to commit the correct name to memory, only remembering the "pattern" of the name and filling in the spelling via guesswork?
          • paulhebert 1 hour ago
            My last name is Hebert.

            There’s about a 1 in 10 chance when someone spells it (or says it) they say Herbert.

            Even in situations where they just read it or I just said it.

            I’ve had Herbert soccer trophies, health insurance cards, etc.

            The mind fills in a lot of blanks and doesnt always get them right.

            • jbaber 50 minutes ago
              I sympathize. -- Not Barber
          • ndriscoll 2 hours ago
            Maybe they skipped straight to Lebeg integrals.
          • jryb 1 hour ago
            Autocorrect might be doing it
          • NewsaHackO 2 hours ago
            People just don’t spell that seriously buddy, especially when it is so immaterial to the point.
            • traes 2 hours ago
              My point is it is crazy to make public claims about how important or not important a mathematical result is when you can't spell Riemann. Yes, it technically doesn't matter, but it betrays a damning lack of familiarity with introductory mathematics.
              • xanderlewis 2 hours ago
                You’re (as Claude would say) absolutely right, and I suspect the original commenter has no idea what they’re talking about.
        • conformist 2 hours ago
          Yes sure but they are different surnames and pronounced differently.
          • xpct 2 hours ago
            I didn't mean to oppose OP's point, I just found it interesting as a non-German speaker!
      • lanyard-textile 2 hours ago
        They're mathematicians, not linguists :)
        • traes 2 hours ago
          The mathematicians don't make this mistake, I assure you. I doubt there is a math professor on planet earth who would spell Riemann as Reinmann. In fact, I imagine no one who has ever heard the name pronounced would do so.
          • gpm 1 hour ago
            One of the best mathematicians I've had the pleasure of learning under added 6 to 7 and got 15 during a lecture... I assure you they're capable of misspelling last names too.

            Mathematicians aren't exactly known for being well rounded.

            • jwilber 41 minutes ago
              Pure mathematicians getting arithmetic wrong is a bit of a meme it’s so common. I don’t think it’s the same as a misspelling of a popular mathematician, not that either are indicative of much tbh.
          • pixl97 2 hours ago
            Uh oh, no true scottsman....
            • vector_spaces 1 hour ago
              It's just a name you write so many times as a math undergraduate or first year graduate student due to the number of load bearing results and objects named for him. The name even appears in lower division coursework

              Yes, you can be an amateur mathematician who manages to avoid such classes and may not know those results or objects, but if you haven't read and written down the name enough to avoid habitually misspelling it, you are outing yourself as a meat proxy unless you are dyslexic.

      • jere 1 hour ago
        “How many Ns in Riemann?”
        • sdenton4 1 hour ago
          Let's figure it out! First, draw little boxes over all the letters. Then add up their areas. Then make the boxes smaller and repeat, until you have forgotten how to count.
      • broptimist 2 hours ago
        [flagged]
        • dekhn 1 hour ago
          Don't be a jerk.
        • XorNot 1 hour ago
          Okay but who cares? Results are results.

          If the proofs work thennwe can put them to work doing more things.

          It is not particularly important that Einstein discovered relativity, just that it was discovered (Maxwell was very close).

    • zone411 2 hours ago
      There was A LOT of drama about this release.
    • fspeech 2 hours ago
      Math is the tool humans use to compress knowledge. So until we can comprehend it there really isn't much progress. Math theorems are tautologies, the truth of which are not dependent on proofs and proofs are erasable, at least classically. But the AI progress is exciting and AI proofs are a gold mine for humans (at least non domain experts) to explore.
      • fspeech 2 hours ago
        I think it would be helpful to people who want to understand what a formalized proof is to read Thomas Hales on this: https://www.math.stonybrook.edu/~bishop/classes/math536.S24/...

        He spent years formalizing his sphere packing theorem because the proof (human produced) was already beyond the ability of peer reviews. Now his formalization effort likely can be easily reproduced by a model. However one should read his experience about what a formal proof is: often the problem is the statement not the proof. The example he gave is the Jordan curve theorem. It's actually quite challenging to formalize the concept of a planar curve (there are space filling curves). So it is not necessary that someone can look at a formal statement and say aha it is about a planar curve, unlike FLT where there is not much problem in recognizing what the statement is about.

      • binlog 2 hours ago
        What makes you think no one can comprehend this? It has been less than an hour since it dropped and there is already a ton of online chatter from people explaining the results, pointing out their favorites and more. Some of it is happening on this very thread.
        • fspeech 2 hours ago
          I didn't say that. I am responding to "In a way this is probably 50-100 years of math progress by humans." I am actually very excited about AI proof and I am working overtime in my own way to try to comprehend as much as I can.
      • gpt5 2 hours ago
        Math is far more than that. If you can solve prime factorization for example, suddenly you can listen and interfere with almost every private conversation on the internet.

        We are not far away from the moment where these models will be restricted, and sharing the results will be done more carefully.

        • fspeech 2 hours ago
          This doesn't contradict what I said. But I do appreciate the fact AI can produce side effects not just humans. I made it sound like only human knowledges matter. That's too narrow.
      • fspeech 2 hours ago
        Another way to state this: math theorems are like programs without side effects; it is immaterial whether a program without side effects is ever run. We study math for the side effects: it changes how we organize our thoughts.
      • caaqil 2 hours ago
        > until we can comprehend it there really isn't much progress

        Who is "we" here exactly?

        • fspeech 2 hours ago
          Whoever wants to study the result.
          • caaqil 2 hours ago
            > Whoever wants to study the result.

            Right. Before all the AI disruption, pure Math traditionally welcomed anyone who wanted to study its esoteric proofs, right? I remember all the excitement of the average Math enthusiast casually reading Wiles' proof over coffee.

            Bottom line is, the relevant people can still understand the generated proofs. The disorienting part is they are a little slower than they'd like, but they'll get there.

            • fspeech 2 hours ago
              AI is very helpful with understanding AI proofs. Agent swarms produce messy proofs overall but locally they are excellent and can teach anyone who wants to study them. No one controls math (in a material way funders do control an aspect of practicing math). Still, theorems are already true before we prove them. The difference a proof makes is whether it convinces the reader.
      • gizmodo59 2 hours ago
        >So until we can comprehend it there really isn't much progress.

        Not really? We are at a point if an AI today can solve it, it can be stepping stone of understanding something deeper to tomorrows AI and it continues. Sort of like our limitations doesn't matter. Obviously there are many scenarios in this recursive loop but saying it isn't much progress is not how I view this as

        • le-mark 1 hour ago
          But who will ask the questions or direct further research when humans no longer understand the state of mathematics? Llms lack the drive for homeostasis combined with the evolutionary drive for survival and reproduction thus to direct themselves. They could very easily spend an eternity down a rabbit hole when the warp drive equation was fairly close on another branch.
        • fspeech 2 hours ago
          If it changes how we think then yes it has an effect.
      • yieldcrv 2 hours ago
        Academics have been treating it that way because they had no other choice, and its been a waste of everyone’s time and often times taxpayer resources

        Look at that, taxpayer funding was cut and a private sector solution came in just the nick of time, far accelerating the holding patterns we’ve been in for decades

        Humanity doesn’t need all iterations towards the blueprints, the blueprint is good enough, we all stand on the shoulders of giants

        • fspeech 1 hour ago
          If you actually looked into how agents proved FLT you would be even more amazed by the fellow human beings who are able to keep all this in their heads! I for one can only begin to grasp the scope with AI and scripts.
      • warkdarrior 2 hours ago
        > Math theorems are tautologies

        Proven math theorems are tautologies.

        • fspeech 2 hours ago
          FLT was no less a tautology before it was proved. We just weren't sure about it. Proofs only change us, not math.
        • fspeech 2 hours ago
          True.
  • sebmellen 3 hours ago
    It’s fascinating to read through the reasoning traces: https://github.com/openai/math/tree/main/reasoning_traces

    Look at one of their examples of an initial prompt: https://github.com/openai/math/blob/main/reasoning_traces/re...

    • adverbly 2 hours ago
      > Look at one of their examples of an initial prompt

      Interesting that its only an excerpt. I wonder what else they include but didn't share.

    • ndriscoll 3 hours ago
      > Thus at most one informative i. So cheater chooses arbitrary g_{v_i}, on exact duplicated input matches and passes, independent of actual satisfiability!

      No idea what it's so excited about, but it's cute that it "is." I for one welcome having access to a math buddy 24/7 that's way above my level but also always "willing" to talk at where I'm at.

  • againstapples 2 hours ago
    As an AI "doomer" can I ask the non-doomer people here how you interpret the significance of results like these, and what kind of progress you expect to see in the next 1-5 years?

    Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?

    • somenameforme 8 minutes ago
      Where some see intelligence, others see token prediction. It's a very good question how token prediction could achieve this, but I think there's a simple explanation. No human can hold more than a negligible percent of all knowledge in his mind at once. LLMs have no such limits and so can reliably connect 'obvious' dots that we miss simply for lack of storage capability.

      Well isn't that just semantics? Surely connecting dots in a novel and meaningful is intelligence regardless of how it's achieved. The thing is that humans didn't get to where we are by connecting obvious dots. Go back to before humans had invented language and when bleeding edge tech was literally that - 'poke him with the pointy end.' Train an LLM on that corpus of knowledge. Even given infinite processing power and infinite time - it's not going to discover the secrets of the atom, put a man on the Moon, or do much of anything besides remix what we'd already done at the time.

      I expect there's still much LLMs can achieve simply because of this initial problem. But I expect that they will ultimately start to plateau once these dots have been mostly matched and we reach a point where 'creation' again becomes the missing link. Though even there LLMs will play a major role as tools. For instance Einstein had to spend a significant amount of time in 'retrieval' rather than 'creation' research to develop the field equations for general relativity. If he had access to LLMs trained on all knowledge of the day, he could likely have achieved his goal much more quickly.

    • arctic-true 1 hour ago
      Not a doomer but I try not to be a denier, either. These are hugely impressive results. I do not see the technology plateauing at the current level (though I am dubious about an infinite exponential growth).

      Before I cope, I’ll note that there are plenty of “doom” scenarios that do not require any improvement in capabilities from what we had before this latest unreleased model. We’re at the point where a determined bad actor with enough compute could compromise critical infrastructure in a way that results in casualties, where this actor would not have been capable of such without LLMs. This may not sound like Skynet, but I don’t see why it makes a difference if I’m one of the casualties.

      With that in mind, here is the cope: first, mathematics is an inherently verifiable domain. An LLM can use tools to determine with absolute certainty whether it is correct, and an independent third-party could review and confirm. All of this can be done without any interaction with the physical world or with other minds.

      Second, OpenAI is able to marshal compute at a scale that an individual mathematician can only dream of. It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.

      Third, none of these problems are solved in a vacuum - the reason OpenAI chose these problems is that they are widely discussed and many people are working on them. It’s possible that someone else was close, and OpenAI only contributed the finishing touches. (This wouldn’t need to be plagiarism, to be clear - people publish their work!)

      • istjohn 1 hour ago
        > It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.

        See:

        > The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking. (TFA)

        • arctic-true 1 hour ago
          That tells me very little. What was the cost to OpenAI in dollars? What differentiates the high-cost problems from the low-cost problems? And that’s before you consider that OpenAI has strong incentives to downplay its costs while emphasizing its results. A one-liner in a write-up doesn’t change the fact that they have access to massive resources.
        • edot 22 minutes ago
          Sure and if I make a half court shot after an hour of trying, the result only took 1 second.
    • doginasuit 1 hour ago
      I expect AI will continue to be useful on the vanguard of fields like mathematics because it has the perfect conditions for it to shine. There are a lot of discussions and leads to start from and the AI can check its own work and iterate. It can fail hundreds or thousands of times in a day and continue to work with the same tenacity.

      Superhuman tenacity is not enough on its own to pose an existential threat. If it showed the same capacity for judgment, inventiveness, and decision making in the messy problem space of the physical world, I would be more alarmed. There have been experiments where an AI is given control of managing something like a vending machine and it always ends up a mess. AI has come a long way, but certain problems seem as difficult as ever.

      When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned. Enslaving humanity will involve taking a lot of calculated risks that tenacity alone cannot solve.

      • mikestylz 26 minutes ago
        > When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned.

        Given the past rate of progress, why not start being concerned now? It's a bit like the economist saying that the optimal number of flights to miss is not zero. If you keep landing short on your estimations for how far this technology will go, next time you should err on the other side.

        And regarding Vending-Bench 2 (https://andonlabs.com/evals/vending-bench-2) my understanding is that models do pretty well on it now.

    • Kotlopou 47 minutes ago
      I would like to know how much of the progress comes from effectively combining existing research programs plus massive persistence, and how much is AlphaZero-esque RLVR completely independent of training data. Since I cannot get anybody to care about this question (even though I think it's vital for guessing what the future trajectory will look like -- are we going to complete existing research programs or start new ones?), I live in ignorance and wait for the day when the answer becomes clear.

      In looking at this over the past hour, I haven't seen clear evidence one way or the other. Some of the stuff is highly unexpected (like the multiplication algorithm), but counterexample-y, and about the rest the professional mathematicians online seem to have a consensus that it's not "breaking through fundamental obstacles". I suspect neither of us is competent to judge that.

    • never_giveup 2 hours ago
      Try using AI for your work, whatever you do. You will quickly understand the limitations.
      • ggreer 1 hour ago
        Unless you think that AI will quickly hit a wall (which seems odd considering that only a few years ago the best models had trouble doing basic math or counting the number of Rs in "strawberry"), I don't see how that's reassuring. The models will only get cheaper and more capable over time. It seems quite likely that at some point (probably before I hit retirement age) they'll be able to fully replace me at my job.

        Is there any specific cognitive task that you are willing to bet that AIs won't be able to accomplish in the next 5 years? Because if not, I'm not sure we're disagreeing about predictions.

      • againstapples 18 minutes ago
        It has limitations for sure, I just don't expect those limitations to last. What probability would you put on the limitations being overcome in the next 5 years?
    • pj_mukh 2 hours ago
      Can I ask you back, what your concern is here? It'll get so good so as to desire to hurt us or is it a misalignment event that you think will lead to disaster?

      Or is it simply that you feel bad for Mathematicians.

      • againstapples 27 minutes ago
        I believe the AI labs might actually succeed in developing superintelligent AI and recursive self improvement, and that if they do they are very likely to lose control of the system they build.

        I really think the only place people disagree is that they don't actually think it's possible, they see it as hype or doomerism. I can't find any good reasons to rule out that the companies could actually achieve what they are trying to so I think they should be stopped.

      • Veedrac 1 hour ago
        Humans have one ecological niche. Soon we will have zero. That is worth worry.
        • pj_mukh 50 minutes ago
          >>ecological niche

          As in..to be dominant? Why would an AI try to dominate? What would give it purpose, or is this a purpose via misalignment scenario?

      • whimsicalism 1 hour ago
        Misuse of extremely capable models, misalignment during RL are both very large risks as capabilities grow imo
        • voiceeh 1 hour ago
          So, you're worried about them breaking containment and deciding to do bad things?
          • orlp 1 hour ago
            I'm more worried about them doing bad things at the behest of people who want them to do bad things.

            That is 1. immediately technically possible, and 2. realistic.

            If you need a source for 2 I'd suggest you open any history book.

          • whimsicalism 1 hour ago
            that is a worry yes. instrumental convergence and misalignment during RL is when it is most risky because it hasn't necessarily had the final safety polishes applied

            i get a lot of skepticism on HN by the same crowd that has been wrong about this tech for about 4+ years straight

          • bamboozled 1 hour ago
            The rapid development of extremely dangerous bio-weapons?
        • pj_mukh 1 hour ago
          Misuse how exactly?
          • whimsicalism 1 hour ago
            any number of ways. as we turn over more of our physical economy to these agents (and we will), the potential for physical damage becomes greater. biorisk is getting a lot of attention right now and i think that's justified
            • pj_mukh 51 minutes ago
              I am trying to think through the scenarios here, like a biolab making something that a very advanced open source model prompted by some terrorists comes up with?

              Why would a biolab capable of making something like be unregulated? And if it definitely would, isn't the problem with the biolab?

              It feels like all these scenarios are leaving some gaping holes in our security infrastructure that have nothing to do with AI.

      • ewild 1 hour ago
        i feel bad for math guys yeah seems they are more cooked than CS
    • computably 1 hour ago
      Depends on your definition of doom.

      If doom is ending up with grey goo / paperclip maximizers, or SkyNet, then I don't think doing mathematics is evidence of that direction. Partly because LLMs are quite apparently dumb in many ways, and for math specifically, they need a formal verifier (Lean) which "gamifies" math.

      If you mean bioweapons or cyberwarfare, there's nonzero risk, but not orders of magnitude worse than other global risks. Climate change, nuclear weapons, monoculture food, etc.

      I'm far more concerned about overall trends in AI development and usage. It's accelerating wealth inequality, social isolation, attention capture, surveillance states. If we end up in the Matrix except the admins are humans and the simulation is hyperoptimized TikTok, is that AI-driven doom, or is it just an inevitable outcome of modern tech?

    • jaykru 1 hour ago
      I wrote something [0] that might answer a bit a few weeks ago. Today's slopdrop certainly is challenging my stubbornness, but everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains. Math yields especially impressive results because it so broad and deep that essentially no person can know of all of its parts; pretraining and deep search capacity is a huge advantage. Gowers has recently written about these capabilities and gestured [1] toward some human capabilities lacking from the current frontier models, though he isn't convinced they won't develop soon. If you assume we don't get a total mathematical superintelligence (which to me seems already sort of AGI-complete) and only amplify the capabilities we have today, it's not obvious to me that we get takeoff from recursive self-improvement, unless you happen to believe that 1) we can clearly specify what AGI or ASI is 2) all of the requisite ideas are out there and need only be combined and/or optimized.

      [0] https://dank.systems/posts/2026-09-15-ai-bear.html

      [1] https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-a...

      • red75prime 1 hour ago
        > we can clearly specify what AGI or ASI is

        We'll have plenty of time for this, while living off UBI.

      • p-e-w 1 hour ago
        > everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains

        But they’re already extending into politics, military, journalism, art, and many other fields that aren’t verifiable in any meaningful sense of the word.

    • zeroonetwothree 2 hours ago
      I'm not sure how AI solving math problems is related to "doom", perhaps you could expand on that? To me (a "non-doomer"), it seems like an overall positive.
      • pixl97 1 hour ago
        You have to look at all pieces of the puzzle. For example there were tons of people that said "they could never solve novel or super complex math problems"

        The issue I see is the list of abilities that AI can't do is shrinking at a rapid pace, and its capabilities are growing at the same pace.

    • schleck8 2 hours ago
      From a few preprints I've checked, this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches. Very few people globally could come up with something like this, even when given time and ressources.

      So in other words, since deep learning is algorithmic research, we are now in the RSI era.

      • thereitgoes456 2 hours ago
        > this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches

        "Surprising" is a, well, surprisingly high bar to clear, and requires thorough understanding of not only the paper, but existing work in the area. ("Novel" is tautological.)

        How did you determine this in 1 hour? Are you a researcher in multiple of these areas?

        Can you give an example, or explain more how you came to this conclusion?

        • scarmig 29 minutes ago
          One surprising result is the sub n log n multiplication. Galactic, as one might expect, but if you polled human researchers 24 hours ago, most would say something like it was very unlikely.
    • gizajob 1 hour ago
      Did AI beating humans at chess:

      a) destroy chess and make it a pointless endeavour,

      or

      b) make humans much better at chess.

      • lf88 34 minutes ago
        Chess has always been a game. For other intellectual activities, at least for several people, a big part of the pleasure in engaging in such activities is the sense of contributing something that matters to a collective effort. Strip that away (e.g., because a machine can do the same thing more efficiently) and you effectively make such activities pointless for those people.
      • Light_Hope 1 hour ago
        Just as AI chess performance failed to render human chess playing pointless, it's unlikely that AI will make human thought pointless. Unfortunately, not being pointless doesn't create economic leverage or incentive, and currently a significant portion of humans depend on being the best chess solvers to sustain themselves. If Deep Blue rendered large portions of human thought economically meaningless, we might look at it a little differently.
      • vouaobrasil 1 hour ago
        It didn't destroy chess but it did make it a little irritating in some ways. More mechanical. Some chess players have bemoaned the level to which grandmasters and other highly-ranked players just endlessly study opening-book theory and I think computers made that worse. Bobby Fischer also agreed and that's why he invented Fischer random chess.

        I do think it also took some of the magic away from chess, and Lee Sedol has said something similar about go.

        So did it destroy it? No. And maybe you could make the argument that it got more exciting in some ways, but I think it sort of degenerated into a spectacle and it's just not as interesting as it used to be, and I think computers have played a role in that.

        • gizajob 32 minutes ago
          At the same time though, Magnus is Magnus because he’ll crush you in any endgame.

          I’d posit that more people are playing more and learning chess than ever before, thanks to networking and AI assistance. And computers have only beaten us at computer chess. Human chess is always an experience for learning about the other person, or flipping the board and walking off in a huff.

          I don’t know so much about Go and it’s not surprising that Lee Sedol became pretty demoralised, but the generation coming after him alongside computers are going to see new possibilities that had gone unnoticed in purely-human Go, extending the game for everyone.

    • besterman23 2 hours ago
      I see it as “if this can be represented in tokens it can be trained in and ‘solved’”. I don’t think there will be a plateau, but there might be issues with how effectively we can represent some things in a tokenized form and still be efficient.
    • skybrian 1 hour ago
      For me a big open question is what sort of progress will we see in robotics. I won't even attempt to speculate, but it does seem hard in different ways than proving math theorems.
    • yk 32 minutes ago
      I'm a transhumanist, I want to build god and kill death. To me this looks like we are moving in the right direction.

      So basically I think that the future is getting pretty weird because we are building really powerful tools, though these tools are precisely what allows us to prosper in that future.

      • outworlder 12 minutes ago
        Or, conversely, they are the exact tools that will allow the powerful folks to not care about the peasants anymore. History tells us what happens then.
    • runarberg 1 hour ago
      AI hater here:

      I have a hard time taking statements from OpenAI about their own product, that they are trying to sell to people and make money, seriously. I take these statements as they are greatly exaggerated or even straight up lies and propaganda.

      That said, I think results like these are mostly annoying more then anything. They spent a lot of money, used up gigartiuan amount of compute, to ruin a puzzle that mathematicians were tackling. I am mostly unsurprised that if you spend a trillion times the energy that a team of mathematicians would, that you get maybe 1.5 times the results. I see a future where that 1.5 times the results may go to 3x but not much more. And if that, then I will be more annoyed.

      • vouaobrasil 1 hour ago
        I think giving the world a "cheat code" to accomplish too many intellectual tasks will eventually take its toll on us because after relegating most of our physical labour to machines, we could at least marvel at the abilities of individuals.

        Yes, human beings can do more now in some ways but...to put it poetically, I think there will be no more heroes like Einstein and Newton of the past. Now it will just be someone cleverly turning the crank.

        Yes, we still admire Usain Bolt even though we have cars...but maybe the admiration is a lot more trivial than if we did not have them....

        Personally, I think AI is a grand mistake.

    • icepush 2 hours ago
      They can replace anyone but they can't replace everyone.
    • ForHackernews 1 hour ago
      AI performance has always been extremely spikey. It's great at some things and terrible at others.

      Why do you think the world to date hasn't been taken over by evil genius mathematicians? Can you extrapolate from your understanding of the answer to that question?

      • againstapples 20 minutes ago
        I don't think evil mathematicians are very common or that any of them would be capable of single handedly taking over the world if they were. My concern is more about systems that are beyond human level, those kinds of systems would actually be dangerous to us.

        I see the recent progress in mathematics and cybersecurity as signs that models are getting more capable more quickly than usual. The companies plans to develop them by recursive self improvement now seems like a real possibility and I don't think they should be allowed to attempt this.

    • ijidak 1 hour ago
      For me it's a mixed bag.

      There will be a lot of job loss unquestionably, in the same way that automation reduced manufacturing jobs and farm payrolls.

      At the same time we have to put what AI can do in perspective.

      Intelligence is a broad grouping that includes concepts such as knowledge, skill, experience, and wisdom.

      AI has incredible knowledge and in many areas approximates experience and wisdom.

      But wisdom is harder to formalize than knowledge and skill.

      For example certifying a college education relies mostly on the ease with which we can verify/test knowledge.

      To some extent advanced degrees try to certify maybe wisdom and experience.

      In my very personal opinion, wisdom and life experience should give humans an edge for a while to come.

      Additionally, I do feel that the more an individual lacks better than average wisdom and experience, the harder it will be for that person to compete with AI.

      Also, on the bright side, the average human will continue to prefer to interact with a fellow human in many spheres. That will also act as an upper bound on AI and robots taking every job.

      Either way, I do think this transition will be painful. I don't feel it has to be apocalyptic.

      But the world has been an especially volatile place over the last 10 years.

      So when you add that existing volatility, to the upheaval from the AI transition, it would not surprise me if the transition results in violence.

      But, without the pre-existing volatility, and if humans were capable of generosity and love at scale, I see no reason AI can not be absorbed into society with net gain.

      I guess to summarize, I feel this tech should be a net gain and to the extent that it isn't, it will be because of flaws deep inside of humanity itself, not because it had to end in chaos.

      In other words, I feel fear, greed, anxiety, and competition -- all our base instincts coming from all sides, will be what determine the end result of AI moreso than AI taking everyone's job.

  • foota 2 hours ago
    From their github: "The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking." That's pretty crazy.
  • bcatanzaro 38 minutes ago
    “I think at the heart of this issue is that humans have two competing natures: a tendency to compete and a capacity to appreciate beauty,” said Kai Shaikh, a graduate student in mathematics at the University of Toronto. “To me this seems to be a case of the former attempting to strangle the latter.” [1]

    Beauty can be appreciated even when it is vast, even when it is beyond one's comprehension. I don't think this release should be primarily viewed as an outcome of competition. Instead it is revealing truths about the universe that were always there and always beautiful, even if we hadn't seen them yet. I believe there are infinitely more such beautiful truths currently hidden and waiting for us to discover.

    [1] https://www.nytimes.com/2026/10/06/science/openai-math-probl...

  • kingstnap 2 hours ago
    Some of these are interesting ngl.

    109. Integer multiplication below n log n

    Surprising that this is possible.

    158. The Euclidean plane cannot be colored with five colors.

    Only 6 and 7 remain!

    376. Universal computation in forced Navier–Stokes flows.

    Morning coffee proven turing complete

    • zeroonetwothree 1 hour ago
      Integer multiplication is very unexpected, I think most people believed in the n log n lower bound!
      • tootie 1 hour ago
        Note that these are all preprints. None are verified.
    • mFixman 2 hours ago
      > We give a deterministic algorithm that multiplies two n-bit integers in O(n (log n)^(1−κ)) worst- case time, with κ = 2^(−182).

      LMAO, I don't think I ever saw such a small number in a CS result.

      • kingstnap 2 hours ago
        Yeah its ridiculously small, but any improvement on n log n is wild.

        Like there is somehow redundancy in a fourier transform that makes it sub Linearithmic?

        Which low and behold ->

        130. Fourier transforms below n log n.

        • xyzzyz 2 hours ago
          They also separately give algorithm for Fourier transform over complex number faster than O(n log n)
          • saalweachter 1 hour ago
            Wikipedia just told me there's a galactic algorithm for integer multiplication in O(n log n) based on FFT so I'm guessing those two proofs are related.
      • sobellian 2 hours ago
        I am fully braced for it to be a https://en.wikipedia.org/wiki/Galactic_algorithm

        Very surprising result though! Multiplication is easier than sorting.

        • zeroonetwothree 1 hour ago
          Then 'n' means kind of different things for sorting vs. multiplication though. For example for sorting we assume constant time comparison, which doesn't make sense inputs of O(n) bits
          • sobellian 8 minutes ago
            If you sort n k-bit items for a total time of O(nk logn), that scales more poorly in n than multiplying n-word integers. Of course if k is constant you can do radix sort, but I genuinely don't know under what conditions radix sort is more/less galactic than this multiplication algorithm.
        • senderista 1 hour ago
          It would be absolutely unbelievable if such an improvement were practical.
      • anon-3988 2 hours ago
        It fascinates me that there's something like this in something as solid and rigid like matrix multiplication. What causes something so rigid to break apart and "leak" at very large scale? Why does the "optimization" appear to be very, very small? Why does galactic algorithm exists? I can't imagine long division suddenly breaking apart after a billion digit, the structure seems very stable? I have heard before that matrix multiplication is apparently optimize-able at very, very large scale.

        Does anyone have an intuition to what causes it? What happens at these large scale (or very small)?

        • adgjlsfhk1 57 minutes ago
          One way to think about it is that the classical algorithms are the ones that are fast for small numbers. Galactic algorithms often work for small inputs, it's just that to be faster you need big inputs. A common case of this is a requirement that log(n)<<klog(log(n)). If k=100 then this algorithm will take huge sizes to win
  • 7373737373 14 minutes ago
    It may be useful to publish a formalization of ALL known mathematics at this point. Like every book ever printed, every paper on arXiv etc.

    How many Gigabytes would that be, compressed? Wikipedia once fit on a DVD

    This might also allow for some interesting meta-mathematics

    • hagen8 4 minutes ago
      This is what they are trying to do with Lean
      • 7373737373 1 minute ago
        Oh? Where can i read more about that? It appears the focus so far was solving open problems
  • dekhn 2 hours ago
    I'm a software engineering/biology/ML guy who loves when clever math ideas get turned into real solutions (https://en.wikipedia.org/wiki/Compressed_sensing). I am curious if any of the results have immediate applications in any kind of engineering or science.

    It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.

    • brandonpelfrey 1 hour ago
      Same. I have agents analyzing the papers here to see which of these are actually new approaches, novel application of unusual approaches, etc. If there are new intuitions and ideas, those are the mostly powerful reusable components by my estimation.
      • OutOfHere 1 hour ago
        Please share your findings.
  • trostaft 33 minutes ago
    Not much computational mathematics here, but I do see some of interest in the MCMC community. In particular, 139, 93, and 101 are very interesting. Also (I'm not too familiar), I remember attending some talks on the Crouzeix conjecture 325 attacks, should sharpen some rates for Krylov methods. Obviously need to read deeper, some of these don't have corresponding Lean formalizations.

    Cool!

  • rifty 7 minutes ago
    As AI works steadily through the already discovered unsolved problems in mathematics, who is currently discovering new ones?
  • karahime 3 hours ago
    Extremely unfortunate that gate keeping got to the point where they felt the need to ask for permission to share math.
    • bravoetch 3 hours ago
      In previous math sharing there was speculation about stealing human researcher's results or progress, via prompt inputs from those researchers, and sharing that as their own result. They're adjusting their process, and it seems ok.
      • whimsicalism 1 hour ago
        Let’s be very clear, the alleged “stolen results” were largely the product of another LLM, not de novo human work. Also, it was false - they did not steal the results.

        No clue why I'm downvoted for this, HN struggles with truth-seeking on these topics.

    • hgoel 2 hours ago
      After how poorly OpenAI and Anthropic handled the previous cases, I approve of the more measured and cautious approach this time.

      We cannot have them rushing to publish amidst tons of confusion, rumors of threats/scooping and outright plagiarism of existing work (by failing to cite said work).

      If they're going to participate as scientists in these more rigorous fields, they're going to have to match that level of rigor, not lower it to the disastrous low that ML research publication is at.

      • make3 1 hour ago
        It was OpenAI that fucked up the Navier Stokes explosion situation, not Anthropic
    • xpct 2 hours ago
      Just to be very clear: they aren't asking for permission, they are framing it that way because of the bad press.

      There's no gatekeeping here!

    • reasonableklout 2 hours ago
      I think it's generally a good thing that OpenAI is noticing when their projects are harming human communities, and deciding to respect their norms, especially when their math discoveries do not have immediate application and build on the thousands of years of that community's work.
      • skeledrew 1 hour ago
        > harming human communities

        Said communities are doing that all on their own by caring about what AI is doing rather than just focusing on their own thing as they did before AI. It's a serious kind of envy IMO.

      • schleck8 2 hours ago
        > do not have immediate application

        How do you know? Seems statistically unlikely with 720 problems, most of them well known

  • open592 3 hours ago
    Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?

    Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?

    • dekhn 2 hours ago
      Let me give you some perspective: my entire phd was made obsolete by CRISPR. It was a wonderful thing.
      • thimotedupuch 2 hours ago
        Interesting. If you don't mind, could you please share a little bit about that ? You already finished your dissertation ? It was about the works of Doudna and Charpentier ?
        • dekhn 2 hours ago
          No, back in the late 90s and early 00s, people were trying to engineer custom nucleases and transcription factors, my work was on doing molecular dynamics simulations to optimize TF sequence specificity (similar to engineered zinc fingers) for gene therapy. I wrote up my dissertation and published it in 2001, and then went off to find enough compute, IO, and smart people to make it happen (https://research.google/blog/groundbreaking-simulations-by-g...).

          My approach would require custom engineering for every different sequence we'd want to target. With CRISPR, you just "program" the system with a guide sequence, you don't need to do massive engineering to solve a protein design problem.

          • vasco 1 hour ago
            So it's not the situation they described at all, you didn't waste time during the PhD having to scramble to change topics as it happened after you were done.
      • boznz 23 minutes ago
        For every door that shuts another one opens - great if you're not a cabinet-maker.
    • aaraujo002 3 hours ago
      This happens all the time, even without AI. Other researchers or PhD students can publish the same results before you. I say that based on my experience during my PhD.
      • CaptainNegative 5 minutes ago
        Yes and no. It's pretty normal for students to get scooped, sometimes even multiple times (ask me how I know...).

        It's much less common for students to have their entire thesis direction removed from under them, as might be the case for someone working in fine-grained complexity assuming 3SUM has no subquadratic algorithms, or working assuming ~UGC. Both of which (publicly) seemed like perfectly valid research directions until a couple hours ago.

        There might be stuff to salvage from their conditional results anyway, but this is not your average scooping.

    • binlog 3 hours ago
      Use whatever is published as the new base for your research. Use AI tools to help you going forward.
      • xpct 2 hours ago
        In other words, you've already taken a gambit with the first half of your PhD, now take a second gambit, praying that you have something to publish by the end of your PhD.

        It has to feel awful to be in this position.

        • torben-friis 2 hours ago
          Could be worse, imagine having years of experience in a profession these things can now handle by themselves.

          :)

          • jltsiren 1 hour ago
            It's worse for those who will face the job market (or maybe tenure review) in the next 2–3 years. If you have more time before you have to justify your continued employment, you can pivot to something else. Theorems and proofs may be cheap now, but there is a lot of value in figuring out what is worth studying.
    • bobmarleybiceps 2 hours ago
      I think eventually companies won't get as much stuff that's usable for marketing, so they'll stop investing so much into ai for math, so eventually cheap and poor graduate students will be able to do relevant work again without worrying about getting scooped by a company with a million GPUs :-/
    • dcl 3 hours ago
      This has always been a challenge for PhD students and researchers, it's just far more likely to occur now it seems. Getting scooped doesn't feel good, but it's a signal you've been thinking about things other people care about.
    • pratikdeoghare 58 minutes ago
      > what do I do?

      Very hard question.

      Your work makes you one of the very few people who really understands the problem and solution and its significance.

    • glitchc 2 hours ago
      Perhaps consider switching to a more applied field. Experiments in the physical realm hold value, especially if you document the process.
    • hgoel 2 hours ago
      It could still be interesting if your approach to the problem was different to theirs.
    • goalieca 3 hours ago
      Don’t paste your research into these AI because they will train on it and then scoop you.
      • esafak 2 hours ago
        I think that happened after word of the project reached OpenAI and they allocated resources to it.
    • claaams 2 hours ago
      Don't worry, if you use openAI and get lucky they might offer to share credit with you for your work.
    • caaqil 3 hours ago
      > what do I do?

      Precisely what all NLP researchers and the ML community at large did in the last few years: embrace the frontier and realize that attention is all you need.

    • bamboozled 1 hour ago
      Ask OpenAI for money when you don't have a job or future?

      I guess the only answer is to adapt with the tools. If we can't do that, then yeah, we're in trouble.

    • ex-aws-dude 1 hour ago
      That’s always been a thing, it’s called “getting scooped”
      • vouaobrasil 1 hour ago
        Killing with knives has always been a thing. Now, we have the machine gun.
    • moralestapia 2 hours ago
      That would be unfortunate but the world does not owe you anything and is not going to stop for you. Which is also a valuable thing to learn in your 20s (ideally earlier).
    • vouaobrasil 1 hour ago
      > Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?

      I did get my PhD...before AI. And my honest advice (to myself back then, even) would be: quit the PhD, become an electrician, and work hard to buy a tiny house in the middle of nowhere to watch the world burn in this madness.

    • vinyl7 2 hours ago
      Look forward to being obsolete I guess
    • yieldcrv 2 hours ago
      Yes, and?
  • pavitheran 3 hours ago
    From the GitHub description: “On average, each result used 3 hours of ChatGPT Pro thinking compute”
    • password54321 2 hours ago
      Oh cool, we will all now have a math genius on our computer.
      • jrflo 2 hours ago
        It was using their internal math model, so not yet for us
        • password54321 2 hours ago
          I used future tense. It was implied this will be available.
      • an0malous 2 hours ago
        Well, on their computers. But you can rent them for a price.
    • orlp 2 hours ago
      I'd really like some clarity on what that means. E.g. DeepMind has 'cheated' with this in the past, claiming that AlphaZero only took 4 hours to reach super-human chess levels while conveniently leaving out the fact that it was 4 hours x 5000+ TPUs. Sure it's impressive that it only took 4 hours wall-clock but it's very misleading as to cost.

      Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?

      • timjver 2 hours ago
        >OpenAI has 'cheated' with this in the past, claiming that AlphaZero [...]

        That doesn't sound right

        • orlp 2 hours ago
          Oops, edited.
      • machomaster 2 hours ago
        They did say that. "3 hours of ChatGPT Pro thinking compute"
        • orlp 2 hours ago
          Yes, what does that mean?
      • pixl97 2 hours ago
        Depends what your metrics are. If you suddenly found a way to have 9 women make a baby in one month that is huge.
        • orlp 1 hour ago
          I'm not denying that, but I'd still like to know what that cost.
    • Jtarii 2 hours ago
      That estimate is obviously going to conveniently ignore all the failed runs.
    • scrlk 2 hours ago
      Does this imply that it was a one shot prompt with ChatGPT Pro style models (i.e. best-of-N), rather than the agent swarm approach that was used for Navier-Stokes?
      • inferencecoder 2 hours ago
        It doesn't imply that, it's just measuring the amount of compute.
        • bigmadshoe 1 hour ago
          But if it was an agent swarm wouldn't we expect something like 300k hours of ChatGPT pro compute equivalent instead of 3?
  • ravenical 3 hours ago
  • binlog 3 hours ago
    So happy this is shared on GitHub rather than some gatekeeping paid journal. Truly a new age for science.
    • fph 3 hours ago
      Most mathematical results are shared on Arxiv. Journals add peer review.
    • adverbly 2 hours ago
      End of an age for journals?
    • traes 2 hours ago
      GitHub is a significantly worse place to store important results than Arxiv. Of course, slop does not belong on Arxiv, buy slop should also not get published.
  • karannb 40 minutes ago
    I think such "process-oriented" fields shouldn't be subjected to mass automation (or at least not till humans have sufficiently leveled up our game and understanding). What I mean is progress in math comes from having gained a deeper understanding of the problem for subseq
  • ed 3 hours ago
  • ks2048 3 hours ago
    I think they should put human names on the papers as someone who has reviewed the result, even if just a preliminary review. (I’m assuming they didn’t just pipe their model output directly to the internet and these had some amount of review?)
    • procedurecall 9 minutes ago
      Looking at the quality of the writing in these documents (or at least, the ones that relate to problems I've worked on), I do not think humans reviewed them.
    • xpct 2 hours ago
      Presumably they don't because they're training the audience (us) to trust the machine, not its verifiers, even if they were included.
      • alexgoodhart 2 hours ago
        I appreciate you acknowledging the politics behind this. OpenAI does not intend to be a software company for long, they intend to be scientific infrastructure. They'll want to be faucet from which pours embryos, orbital calculation, geothermal/substructural rating, and etc.
    • chiwilliams 2 hours ago
      There are competitive reasons that they don't want to share all the people on the team.
    • chrisjj 1 hour ago
      > I think they should put human names on the papers as someone who has reviewed the result

      Assume the empty list you see is complete. :)

    • make3 58 minutes ago
      I think it's a damned if they do, damned if they don't scenario. I think they would be afraid to seem like they're claiming that the selected scientists deserve the praise of an invention or something, when it's "just the machine who did the work"
      • ks2048 44 minutes ago
        Yeah, that's probably right. But, it's a "new world" - invent a new standard - "Written by Claude Foo X.Y; initial review by John Doe".

        With a deluge of results, having some human expert vouch that it even might be worthwhile would help. (e.g. see the link on HN yesterday, "Two Room-Temperature Antiferromagnetic Semiconductor Candidates" - I see it not worth looking at unless a subject expert vouches for it).

    • agnosticmantis 2 hours ago
      Long term this will be the only reasonable author list: Chad G. Peter {1}, Mat H. Lean {2}.

      1: Author 2: Verifier

      /s

  • sigbottle 2 hours ago
    Unique games conjecture and matmul <= 2.25. What the hell.
  • TheMrZZ 2 hours ago
    These results are wild. Several individual findings are crazy good and use mostly unexplored methods (the improvement over Riemann for example)... I'm pretty sure some of these results would have been Fields-worthy.

    But having so many of them at once? Damn. We really live in the future.

    • make3 57 minutes ago
      Imagine you get up one morning and most open questions in math are solved lol.
  • bashtoni 42 minutes ago
    Is AI going to put Mathematicians out of a job, or is it going to create many new jobs reviewing proofs it creates?

    I'm not sure it's clear right now.

  • karannb 36 minutes ago
    I think such "process-oriented" fields shouldn't be subjected to mass automation (at least not till we as humans have sufficiently leveled up our game and understanding).

    What I mean is progress in math comes from having gained a deeper understanding of the problem for subsequent attack of more problems and IMPORTANTLY applications! Right now the first one is trivially satisfied (given oai maintains some memory across models) but the second one is not! It's generating proofs faster than anyone can validate and so only the model can use these. Consequently if it just keeps doing more theory it's... not very helpful or at least not optimally helpful. This is just bragging rights for now.

    More "application-oriented" research would be awesome, where it tries to achieve some desirable effect and then produces relevant theory and experiments around it. Fields like CS, Physics, Chemistry, etc. This would also benefit a wider section of the population rather than the 10 people who understand most of these proofs.

  • TeeWEE 23 minutes ago
    This feels like a huge AI slop dump. Who validated these results. Why are they not mentioned. Human understanding is key here.

    In my experience AI (frontier models) sometimes does weird stuff that needs human review. Not that’s incorrect but sometimes overly complex language or weird use of language.

  • rinconrex 1 hour ago
    The math equivalent of AI code reviews piling up. I wonder where the incentives will align and the equilibrium turns out.
  • avd201 1 hour ago
    Wow, FFT faster than O(nlog(n))? I wonder if that will open the floodgates for further improvement or not. I don't understand anything about most of the fields these results touch, but I can say that this in particular is very surprising.
    • philipwhiuk 1 hour ago
      My guess is that the constant terms are large enough it's not practically useful in most cases.
  • closetheloopdev 1 hour ago
    Hopefully the techniques and results here will be in the training dataset for the next models, so that each new release will give us more interesting techniques and results!

    It seems that OpenAI has a proof machine that keeps multiplying fruitful proofs!

  • curtis-jm 2 hours ago
  • lf88 1 hour ago
    In some ways, this feels more like an ominous warning about the times to come than something to celebrate.
  • NegativeLatency 1 hour ago
    Why should I care?
    • voidfunc 1 hour ago
      Because it means mathematical discovery can largely be automated away from academics. This is the beginning.
      • bamboozled 1 hour ago
        The beginning of what?
        • voidfunc 1 hour ago
          The beginning of the end of human thinking being valuable enough to justify university's existences among many others.

          Were in an unprecedented time where the value of knowledge is about to be crushed.

          • bamboozled 1 hour ago
            Not sure I agree with this take, but we're going to find out either way.

            Have you ever heard of an S curve? Things will develop rapidly, then equalize. If they don't, we're at the singularity and I guess the end of time as we know it.

            But I guess really bad things happen, cancer, radiation poisoning, torture, people have died in really horrendous ways, and I guess dying from some horrendous AI side effects is possible too. Yay.

        • sunkeeh 1 hour ago
          Golden age of discovery and mass layoffs
          • zeroonetwothree 1 hour ago
            Predictions of mass layoffs from AI have been about as wrong so far as predictions of AI plateauing.
          • voidfunc 1 hour ago
            People need to figuring out how to horde as much wealth as possible right now in the next 2-3 years. Jobs especially for knowledge workers are about to disappear.
            • le-mark 1 hour ago
              I’ve been thinking this as well. I imagine there is a wealth level x such that someone can escape the coming ubi welfare state. Anything under that you are fucked.
            • chadcmulligan 1 hour ago
              It's funny we're possibly entering a golden age of thought, the dreams of the ancients, but we're all worried about capitalism, I think the problem is pretty obvious.
              • lf88 19 minutes ago
                Maybe a golden age of thought for the machines, but possibly (I would even say likely on the current trajectory) a dark age for humanity.
            • brcmthrowaway 44 minutes ago
              Any tips?
  • i_idiot 59 minutes ago
    If only AI can better humans in meditation...
  • xydac 1 hour ago
    i wonder what it means for maths researchers, and how it aligns with how they approach math problems.
    • blooalien 1 hour ago
      > i wonder what it means for maths researchers, and how it aligns with how they approach math problems.

      I guess their job now is "Idea Man" and "Error Checker"? Kinda like (some/many) "programmers" these days.

      • xydac 39 minutes ago
        just wait till someone builds a idea generator model - wire it to decision (jev-like) classifier -> loop it back to researcher

        >> may be thats what open ai did :)

  • dgacmu 2 hours ago
    I find the claimed matrix multiply result (w<= 2.25) shocking. I hope it holds up.
  • lokl 1 hour ago
    Do applied math next.
  • yewenjie 2 hours ago
    A lot of these seem to be proving conjectures rather than finding counterexamples, a lot of people used that to claim that these models are not really smart/creative etc.

    That copium didn't last for what, three months?

    • zeroonetwothree 1 hour ago
      I acknowledge I am impressed how quickly it moved beyond just counterexamples.
    • sebzim4500 2 hours ago
      Don't worry, more copium will be delivered. TBF so far it's still only solved the easiest of the millennium problems.
  • matapassiones 1 hour ago
    Valency has the papers up on Valency Hub
  • aaraujo002 3 hours ago
    The Advisory Group states in its recommendations [1]:

    "We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."

    To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too?

    [1] https://agmai.org/general-sep29/

    • tchalla 2 hours ago
      Why did you leave out the entire quote?

      > At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.

      To me, the issue is that the models are proprietary which are only accessible to a few people in 2 digits. It's not about progress but access.

      • aaraujo002 2 hours ago
        Maybe, but it still seems like an excuse. OpenAI has a proprietary model capable of solving these problems and is willing to share the results with the mathematical community. So basically, the ask is to just not use the model and leave the problems unsolved?
        • adrian_m 2 hours ago
          The ask is to let mathematicians outside of OpenAI use it, ie. at least wait until the model is released.
          • strange_quark 1 hour ago
            I don’t think that’s sufficient. If they want to be good stewards of mathematical research and not just doing marketing, they need to at the very least tell us the datasets and any techniques they used to train this model. IMO that’s probably just as if not more valuable than a dump of un-reviewed results.

            They won’t because they don’t care and the only way it got this good is something along the lines of they trained on every mathematician’s codex sessions even if they opted out because they consider the thinking traces or output and metadata fair game.

          • agnosticmantis 2 hours ago
            Which mathematicians though? Only fields medalists? Grad students? Any hobbyist wanting access?

            These models are too expensive for broad access unfortunately.

            • Jtarii 2 hours ago
              ChatGPT pro is accessible to literally anyone who has a job and lives in a developed country.
              • Jweb_Guru 1 hour ago
                Literally every single one of these papers was developed with an internal model that not even most OpenAI employees have access to.
        • TeeWEE 17 minutes ago
          The work is not having AI poop out these docs the work is validating them and publishing them. OpenAI doesn’t care and wanted to race to publish potential findings and let others review it which is lazy and selfish
    • mattr03 2 hours ago
      I don't think this is a reasonable take at all. Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art. Nothing is gained from OpenAI solving all these problems but taking jobs from mathematicians. Other than advertising for OpenAI at least. It could not be more different from having AI work on disease research etc.
      • zeroonetwothree 1 hour ago
        Most math isn't just proving novel famous results. Just like most of software engineering isn't writing code.
      • bravoetch 2 hours ago
        > Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art.

        It's been a while since I was reminded of this xkcd: https://xkcd.com/435/

    • jhrmnn 2 hours ago
      It all hinges on the definition of “progress”. The debate of the past month is all about questioning whether formally proving outstanding unproven theorems without human understanding constitutes progress. This is quite different from solving diseases.
    • medler 2 hours ago
      The rest of that document makes a pretty compelling case for why this is a bad practice
      • esafak 2 hours ago
        I fear that professional mathematics will wither, and there will be nobody left to digest the AI results of the future, leaving us unable to challenge the AI.
    • osiris970 3 hours ago
      Comical ask
    • bmitc 2 hours ago
      Advocating purely for progress and not humanitarian value is how we'll all get enslaved.
    • perching_aix 2 hours ago
      The trope you're drawing a parallel with has a (to me) compelling counter though: there being a cure for every disease wouldn't stop people from getting sick.

      How this maps back to math, idk.

    • warkdarrior 2 hours ago
      The latest posts from Terry Tao on Mastodon effectively ask for an AI to explain its results to human mathematicians.

      > "I believe that AI can contribute positively in all of these directions [NB: exposition, community building, new directions of study]"

      https://mathstodon.xyz/@tao/117395269325940185

    • fph 2 hours ago
      ...but we're not talking about diseases. Publishing an AI-generated Navier-Stokes solution does not save lives. (And, in fact, it harms some.)
  • connor11528 2 hours ago
    will this make the math for building data centers work?
  • jrflo 2 hours ago
    Glad to seem them changing their tact with the whole NS debacle. Hopefully we can all focus on the results now rather than the surrounding drama.
  • kevinwang 2 hours ago
    wow
  • nautilus12 2 hours ago
    Have any real mathematicians working on these problems reviewed any of these and determined if they are just gobbledegook or not?

    The ones with lean proofs could still be formulated incorrectly

  • pugfugly 1 hour ago
    holy fucking shit
  • Catloafdev 3 hours ago
    This is a pretty hilarious thing to read juxtaposed with AGMAI's requests.

    Basically "Here you go, have fun with this, fuck all your demands, by the way we're gonna be releasing the model stay tuned!"

  • k2xl 3 hours ago
    Can someone knowledgeable about the subject outline the most significant portions of the results?
  • hi__dang 1 hour ago
    Mathematics is solved.
  • tootie 2 hours ago
    Seemingly none are vetted and reviewed yet
    • mulemisterX 2 hours ago
      That's our job.
      • TeeWEE 19 minutes ago
        No it’s OpenAI’s job. They are acting as a meat proxy
      • esafak 1 hour ago
        Ain't nobody paying me to do that. It's kinda sad that maths is being reduced to checking the AI's work.
        • kozikow 1 hour ago
          Not just maths

          In SWE as well - this is what I do most of the day

        • zeroonetwothree 1 hour ago
          Always has been
    • schleck8 2 hours ago
      Most are formalized in Lean, about 80% of what I checked
  • applicative 2 hours ago
    I wonder if the Lean compiler can change it's license so that a for-profit corporation can only use it if it pays, say, a few hundred billion dollars. This is the correct path.
  • applicative 2 hours ago
    Why didn't they just give mathematicians access, so they could at least understand and write up the results in publishable form?
  • mathisfun123 3 hours ago
    With so many results in so many different areas no way they even remotely spot checked well enough.

    Prediction: one of these is wrong and this (publicity stunt) will backfire.

    Edit: don't tell me about lean. For lean to function as a proof certificate you need to represent the theorem correctly. Again: good luck doing that across such a broad swath of problems.

    • jojva 2 hours ago
      You have not read their readme:

      > Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly. We are also exploring community-hosted repositories for these materials.

      • mathisfun123 2 hours ago
        i have and i'm exactly saying that if it comes to pass one of them is wrong it's going to backfire. ie yes that's my exact point/bet.
        • stevenhuang 1 hour ago
          I don't think anyone would particularly care if only one of them is wrong, if most are correct.

          If they are all wrong, that's when it would backfire.

    • bravoetch 2 hours ago
      What does a backfire look like? It's ok to be wrong in the science/math world.
      • mathisfun123 2 hours ago
        of course in science/math it is but it's not okay if you're a business selling supercalifragilisticexpialidocious infallible intelligence.
        • bravoetch 2 hours ago
          Do they claim that's the case? I don't think they do.
          • mathisfun123 2 hours ago
            does company A making product B claim that the product is robust and consistent? is this a serious question?
            • zamadatix 1 hour ago
              If you were waiting for companies to sell engines that never break down you'd still be stuck pre industrial revolution while the rest of the world has been to space.

              The question never if something works 100% of the time but how often it breaks and how that fits the need well. Solving one of these problems is a massive undertaking and accomplishment for the best minds, solving hundreds in a month but being wrong about 10% or something would likely not be the death knell you believe it to be.

    • orlp 1 hour ago
      It's likely that way more than just one of these is wrong. But even if it turns out 80% is wrong this is still 100+ results...
  • dpweb 2 hours ago
    [dead]
  • philipwhiuk 57 minutes ago
    [dead]
  • philipfweiss 1 hour ago
    [flagged]
    • camdenreslink 1 hour ago
      What is considered a big event? Some papers published or letters sent between academics have invented entire new categories of mathematics that didn’t exist. Is it is big as calculus or Euclid’s elements, or Godel’s incompleteness theorem? Or Hilbert’s program of formalism?

      I’m skeptical!

  • redox99 2 hours ago
    The stochastic parrots have predicted the next token once again.
  • sandworm101 1 hour ago
    So the million monkeys at a million typewriters have churned out 700 shakespeares, but they need me for spellcheck?
    • utopcell 1 hour ago
      Nobody needs you to do anything, not with that attitude.
  • digitaltrees 1 hour ago
    Gross. After the accusations training on mathematicians conversations and unique methods, to dump this volume of unsubstantiated papers is at best tone deaf. Are they hoping humans review and validate this? The amount of free that they are coat tailing is outrageous.
    • computerex 1 hour ago
      What do you expect them to do?
  • mi_lk 2 hours ago
    Curious if Sébastien Bubeck still work at OpenAI? He came out quite dirty after Navier-Stokes drama
  • rafterydj 3 hours ago
    I don't know, this does not feel like the message hit OpenAI where it needed to hit, if this is their primary response.
    • osiris970 3 hours ago
      You want them to stop doing math research?
  • senderista 3 hours ago
    Good to see they're engaging with the mathematical community, even if they had to be publicly shamed into doing so.
    • sebzim4500 1 hour ago
      This is the opposite of what the mathematical community were asking for. I say this as someone who strongly approves of this approach.
      • senderista 1 hour ago
        Yeah I'm not sure they met them even halfway.