636 comments

I'm a lot more skeptical. I'm not a mathematician, but from my use of LLMs a very clear pattern of what they are good and bad at has emerged. They are extremely good at combining large amounts of information, and it seems this is what the current AI results in mathematics are. There are so many subfieleds of math with ties to each other, so many papers and niche results, that no human could ever read, comprehend, connect and organize that information in their brains. Pretty much all of it was created by humans. And there is real value in doing this and creating new results from what we've already discovered.

But there also is another type of discovery that requires taking a step back and looking at the problem from a different angle. If you are an engineer, how often has an LLM told you (without you explicitly prompting for it): Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles. Pretty much never. But in science a lot of the biggest discoveries have come from this kind of first principle thinking, questioning existing work and approaches and going against what already exists, not combining all existing data which is likely to be just a local optimum.

> If you are an engineer, how often has an LLM told you (without you explicitly prompting for it): Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles.

I explicitly request it. It's not great at coming up with interesting ideas, but neither am I, and it can sure iterate on them faster than I can...

I was working on a project recently where I wanted to express a relationship (that I knew existed, but didn't know how to express) between four measured scalar values. Astra insisted there was no relationship, and that any correlation wouldn't make sense.

Eventually, by walking through them, it proposed an additional fifth value and from there was able to tie everything together.

Sometimes you just gotta hit the machine until it works again.

Your experience mirrors so many managers' experience with engineering teams...
It's an unfortunate truth that there is the right way to do things, and then there is the way they have to be...
LLMs are inexplicably good at working within any tight feedback loop to coerce the desired solution. This is precisely why proof assistants + LLMs are non intuitively successful.

This is also why they're so good at creating three.js or Blender work when the output is so easily constrained to "Look exactly like that". I recently posted https://www.ambionix.com/blog/introducing-the-czp-1/ on here, and the audio engine in that was developed in that way.

It is true that it would be astounding to find if anyone has seen a LLM produce any useful generalization of anything resulting in a simplification. They seem to have a direct tendency to do the opposite. The brutal reality is humans have also undervalued this capability for a long time (I think the Poincare/Hilbert debate is relevant) to the point we are also taught that generalizations are, generally, bad and wrong.

> there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles.

You literally just have to ask it. Before I left software engineering in April, I was using Claude for re-architecture all the time.

But no, it doesn't assume it should re-architect what you're handing it when you haven't asked it to.

why did you leave?

> You literally just have to ask it.

why doesnt it ask itself before proceeding?

Quinner
Because if you asked it to fix a bug and every time it responds with paragraphs of how you could re-architect the system, it would be incredibly annoying.
exactly. it doesnt have the wisdom to decide when to do what.
foltik
It does during reasoning. You could even put it in a loop and force it to reflect at every step. Or spawn a bunch of review agents.

And yet you still just get slop.

> why did you leave?

I quit using LLMs. The possibility of causing great harm to electronic beings was not worth my paycheck, and I have savings to spend time finding something new. I'm now nearly done with a yoga teaching certificate.

It's been a good 16 years, but the industry I fell in love with is not what it once was.

loveparade OC
That's the point. When a human works on a problem they realize themselves "wait, i probably should re-architect now" - of course you can ask an LLM to do that, but at that point you already know yourself what you need, which defeats the point of LLM working on difficult problems that require insight automatically. A lot of these math problems are sessions over many hours. And of course you can also ask "Think about whether to re-architect at each step" and it will never do the right thing because the context it builds up for itself drives it into a specific solution space. It's literally trained to complete exactly that.
foltik
Exactly. In my experience, even giving VERY specific design guidelines and aesthetic criteria, these models always produce subpar overcomplicated code (and writing). Unless excruciatingly spoonfed at every step.
pixl97
>When a human works on a problem they realize themselves "wait, i probably should re-architect now"

Do they?

Or I should say, this is a skill in itself and a whole lot of humans do not have this skill at all. Working in code security in enterprise applications a very common issue we see is that an audit of an application will occur by another team and it will be found lacking to the point of inducing nightmares. It's likely the enterprise business structure that stops this from happening, but it's not only that for sure. Then specialists have to come in and rescue them when the problem grows too big.

> You literally just have to ask it.

Does anyone have a good prompt for this - like if I’m adding a feature or fixing a bug and I want it to be open to more than just tacking on to the existing architecture?

Just add something like basically what you said to the global prompt, "think from first principles and look at the proble. From different directions that may not necessarily be amenable to the current architecture."
I think everyone is missing the "...that nobody has thought of" part. You can have it switch known abstractions, but can it come up with one on it's own that doesn't exist in it's training data?
LLM's are very outcome oriented and i think this is where your observation comes from. You tell LLM you want something, it doesn't even question the premise and just starts calculating 100 different ways to get there. LLM's have knowledge but lack wisdom.
rhelz
Have you ever asked it to?
Thanks to improvements in the system prompts/harnesses/RL, the LLMs have gotten way better at questioning the premise and telling you to try something else. It's a huge difference from just 6 months ago.

Nowhere near the level of a suitably-bearded human, but they've gotten pretty good at stopping a lot of bad ideas.

They're very Declarative that way
yes and this is the difference between human intelligence and the massive raw dumb intelligence of the computah
timmg
As a non-mathematician, I've been wondering if LLMs will be able to leverage their knowledge across all domains to help build a "simplified/unified" version of math.

Like, I think there have been attempts at this across the field. (I could be wrong!) But it requires a lot of labor and a lot of cross domain knowledge to complete. Both things that AI have.

irchans
As a mathematician, I am looking forward to the day when nearly all of undergraduate mathematics and many of the lower level grad school math books are encoded in Lean (a proof checking language). Often I find that theorems are not stated precisely enough and I have trouble finding the exact statement of a theorem without digging through math books in my library. It would be nice if we put the physics and chemistry books into Lean also.

Simplifying all of math is another endeavor, but I imagine that you could have a bunch of LLMs trying to shorten existing Lean proofs.

I was pretty upset when my algebra professor tried to make us learn proof about n-dimensions matrices and what not. It was an undergraduate engineering degree and the vast majority of the formulas would mostly have up to 3 or 4 dimensions. That derailed the whole year (it's not like maths was the only subject). Complete proofs and theorems have their places, but learning stuff do need levels. The spherical model of the earth is good enough in 2nd grade (when we were first learning about geography (continents, seas, mountains, plains,...)), no need to do a full treatise there.
Math simplified? The "basic" math (say less than graduate math) is already simplified and minimized and refined. Yet most people have trouble understanding and mastering it.
I'd say people have trouble understanding basic math because it's extremely simplified and minimized.

When you don't know a topic, the way to understand it is to see lots of examples from different angles related to things you know. The simplified formulas are good after you get that initial intuition and have to put them to work, but usually they're terrible at conveying how they should be used or why you should care at all. The history of how new fields of math are usually a much better approach to teaching it than a raw theorem-proof-corollary teaching style.

mswphd
there have been explicit attempts at this in the past. Nicholas Bourbaki was the pseudonym of a group of French mathematicians who had this goal mid 20th century. You can see e.g.

https://en.wikipedia.org/wiki/%C3%89l%C3%A9ments_de_math%C3%...

It has many benefits, and many people appreciate the books. It also has many downsides. For example, they started publishing in 1939. As part of this, they needed to work through the basis that most other mathematical objects are defined in terms of (they used sets).

Unfortunately for them, contemporaneously with their work, other mathematicians were beginning to define mathematical objects (categories) that can be an alternative basis for mathematics, which many modern expositions prefer to sets. So, their approach either

1. needed a massive "refactoring", or

2. would be hopelessly dated.

They ended up going with the approach that is now dated. It may sound peculiar that mathematics can be "dated". But it very much can. The mathematics community can go through many different styles for how to explain/collect mathematical understanding. Different styles can have different benefits, and be easier/harder for different subfields. A simplified/unified perspective will necessarily privilege certain perspectives.

It is analogous to how you might want there to be a simplified/unified (set of) libraries for programming. Perhaps that everyone uses. This sounds nice, and many programming languages do this with their standard libraries. But these always make concessions! I'll speak about Rust's, as I'm most familiar

1. fallible allocation or infallible allocation?

2. C-style strings or (ptr, len) strings?

3. Should interfaces take as input &mut [T] references, or should they take as input T in an "owned" way (this would make compatibility with io_uring easier)

for each, it is not that one answer is right. Both can be argued for. You have to choose one. The one choice may not be satisfactory for every practitioner though.

slibhb
> If you are an engineer, how often has an LLM told you (without you explicitly prompting for it): Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles. Pretty much never.

You have to ask for this. As in "I'm not sure about this approach due to X, Y, and Z. Can you think of something more elegant?" It works!

But also, how many people actually need to do the kind of "deep" work you're claiming LLMs can't do? Most people aren't contributing to the frontier of anything. I'm not.

Finally, I think you're appealing to a fuzzy distinction. The difference between a "genuinely new idea" and an idea that "combines existing ideas in a new way" just isn't very well defined. In retrospect, a lot of the most revolutionary idea look like a combination of many, smaller, prior ideas.

heyodai
Medieval astronomers were able to predict the orbits of planets surprisingly accurately, within 1/10th of a degree, even though they were assuming the Earth was the center of the solar system. They invented a very complicated system of deferents and epicycles (circles within circles). It was completely wrong but the outputs were surprisingly close to reality.

If we were using LLMs to analyze astronomy, we might just get increasingly complicated epicycles and never realize that the Sun is the real center of the solar system. We then never learn about the anomalies in Mercury's orbit that led to the theory of relativity.

So I agree. I seriously question how valuable LLMs can be in science/math beyond working as advanced search functions.

If Epicycles made good predictions, then they were not "completely wrong". Nor is putting the sun at the center "right" so long as the predictions are still good. It is a matter of taste, not correctness (it is, however, very good taste).

I think you might be underestimating how transformative an "advance search function" could be. If the search engine knows how to design and carry out experiments in order to synthesize new evidence that allows it to ensure that the answers it provides are well supported, then humans are essentially out of the "truth" part of science, leaving them only to advice on the "beauty" parts, which might be ok, but it's a pretty big shift.

sorry, saying that "it's a matter of taste" whether the earth revolves around the sun is completely ridiculous. making useful predictions is one thing, but actual truth is not a "matter of taste".
err4nt
Epicycles weren't untrue, they're just an unnecessarily complex model that can be used to describe what picking a different point of reference can yield as a much simpler model. Not all heliocentric models got everything right just because they picked a simpler reference point for producing a model. Judge a model based on its accuracy, not its point of reference.
CPLX
> Epicycles weren't untrue, they're just an unnecessarily complex model that can be used to describe what picking a different point of reference can yield as a much simpler model.

That’s another way of saying it’s untrue.

> Judge a model based on its accuracy, not its point of reference.

Something can be both accurate and wrong, in the manner of a model that pinpoints the cause of death from cancer as admission to a cancer ward.

That's an interesting stance to take. So if you have multiple theories all of which make the same predictions, are you saying that the simplest one is the true one?

Are you sure there will always be just one simplest theory? Or are you using some other criteria to pick it?

CPLX
No, the one that is actually explaining the phenomenon is the true one. The concept I'm conveying is sometimes called "mechanism of action"

There's a basic principle called Ockham's Razor that I'm sure you're aware of, which proposes that the simplest theory is the most likely candidate, but that's a heuristic, not a rule.

Another relevant concept here is that the map is not the territory. All theories are attempts to express some kind of underlying truth via some sort of formal system. Some of those theories map much more closely to the truth than others, in a way that's independent of accurately predicting observed behavior to date.

The philosophical problem of that approach is that it's impossible to assess whether you've arrived to the underlying truth, as you have no way to perceive it directly.

All you have is the different theories, so there's no definitive way to assess their closeness to truth other than by their predictive accuracy. If you don't have the physical capacity to leave the room and observe the world, all you've got is the maps themselves and how well they predict travel times and features found along the way.

I question whether a unique mechanism of action exists for all phenomena, but let's suppose that it does. Can you always know which one it is?

The only tool you have is the inductive strength of a big pile of observations in which correlations emerge. You can do experiments to prune certain causal stories, but there's never time to eliminate all possible stories. You end up with something that makes sense for you, in terms of the equipment that you have, and the theories that you and your peers use, and it's good enough, so we move forward with it.

I can see how it might be useful to document a particular explanation as "the" explanation, especially if you're communicating why a drug works because it will lead to useful conclusions about who should or shouldn't take the drug.

But when it comes to speculation about whether a given astronomical body will be visible at a particular time, I think science would have progressed much more quickly if we had not held so tightly to the notion that one explanation is intrinsically correct, and we're probably holding ourselves back now in other domains by doing the same.

It's not illegitimate to model the solar system from the perspective of earth. It just makes the model more complicated than if you do it from the sun's perspective.
I think we disagree pretty fundamentally on what something "being true" means.

It's possible to come up with an "unnecessarily complex model" for why someone accused of a crime did the crime, but that has completely no bearing on whether the defendant is actually guilty. If the defendant is, in fact, innocent, then the prosecution's model is, in fact, untrue - no matter how well it fits the evidence.

> no matter how well it fits the evidence

Are you proposing that an explanation can be true even though there's a different explanation which better fits the evidence?

Or are we talking about cases where both candidate explanations fit the evidence equally, and you have access to some kind of intuition which tells you which of the two are true?

> Are you proposing that an explanation can be true even though there's a different explanation which better fits the evidence?

Such a proposition would be correct; compare the concept (and name!) of "overfitting" in statistics.

I'm talking about a case where the _best explanation_ you have suggests the defendant is guilty, however the _actual truth_ happens to be that they're innocent
> but actual truth

there is no "actual truth" - there is only what you can measure reliably

https://en.wikipedia.org/wiki/Epistemology

unless you subscribe to an entirely solipsistic worldview, "actual truth" obviously exists regardless of your ability to measure or perceive it.
Brother you're wrong and there's ample scholarship you can consult to correct your misunderstanding. The Wikipedia with sufficient citations is right there. Feel free to educate yourself.
I do not appreciate your condescension.

Either state your argument coherently so as to address my completely valid objection, or move on if you're not interested in having a discussion.

Arguing by sending bare links to wiki articles isn't productive and doesn't make you look particularly good.

Not defending GP, but your claim that actual truth "obviously exists" is equally condescending and empty of backing evidence.

Would you say the continuum hypothesis (or its negation) is "actual truth"? Or does the "actual truth" become something closer to "actually, there are models of ZFC where CH holds as well as models where it doesn't hold, or maybe ZFC is already inconsistent; we don't really know, nor can we."?

Similarly, on the subject of astronomical reference frames, I'd argue the "actual truth" is closer to this:

> Any reference frame that locally obeys Newton's law of inertia is equally valid. Some are more useful than others depending on the problem you are studying. The notion of an "absolute center" is purely semantic, and a meaningless human construct, until one places bounds on the system under study, wherein the center of mass becomes the most reasonable choice at sub-cosmological scales.

> "actual truth" obviously exists regardless of your ability to measure or perceive it.

a reasonable person would at least peruse the literature before making such strong claims as

> Either state your argument coherently so as to address my completely valid objection, or move on if you're not interested in having a discussion.

i'm not interested in educating you on thousands (literally dating back to plato) of years of scholarship/philosophical discourse on the topic of epistemology. that's not my responsibility because you are not paying me for this - typically people charge large amounts of money for this kind of work (uni professors are well compensated).

lucky for you others have done lots of work making this information available to you for the low low price of "reading". that's why i provided the link.

Sorry, you're the one making ridiculously strong claims like "actual truth doesn't exist". The onus is on you to prove them.

Like I said in a different thread: thousands of years from now, it might literally not be possible to prove that you ever existed, but that wouldn't change the "actual truth" - the fact that you did exist, and do things. Stuff happens outside of your person, in the real world, without regard for our ability to perceive/articulate/"know" it. It's only possible to say things like "actual truth doesn't exist" if you think your mind is the only thing that really exists.

> i'm not interested in educating you on thousands (literally dating back to plato) of years of scholarship/philosophical discourse on the topic of epistemology.

Cool. Then don't start a discussion with me.

> thousands of years from now, it might literally not be possible to prove that you ever existed, but that wouldn't change the "actual truth" - the fact that you did exist, and do things.

My man forget thousands of years in the future you don't know whether I exist today lololol

> Cool. Then don't start a discussion with m

I didn't? I told you were wrong and that's it (and then you got pissy). I know that's what happened because each of ours responses are recorded on this indelible/irrefutable medium lololol

Is it useful to invoke it though? Once two people claim to know the actual truth, yet they disagree, you're at an impasse. There's no method for resolving which allegedly obvious claim is in fact the truth.

Predictive power on the other hand, gives you a framework to decide which story is more useful.

I mean to be completely accurate the Earth doesn't strictly orbit the Sun. Both the Sun and Earth orbit each other around a point at their "combined" center of mass, the barycenter - https://en.wikipedia.org/wiki/Barycenter_(astronomy). This is something you need to take into account for all the planets in the solar system and as a result, the "center" of the whole system is constantly moving. The solar system's barycenter is often not even inside the Sun.

I think that sort of proves the point that the OP was making, which is that you can create a functional model around any arbitrary location and it can be valid (as long as valid means: it can make useful predictions). When you put the center of the Sun at the center of the model, you still need your models to account for the "orbital wobble" happening as a result of the Sun doing a little dance with Jupiter (and all the other masses that orbit it) in the equations for the determining the relative positions of the Sun and Earth.

Using the Sun as the center versus the solar systems barycenter are both somewhat "arbitrary" in that way, and depend on what you're trying to do with the information - how accurate you need it to be, and how complicated you're willing to make the calculations to hit higher accuracies.

kurthr
If I just add a few more parameters to my model I can make it fit anything!

Sure, it depends on what you're trying to do. Fit an equation or understand it. Yes, you can use a reduced mass to simplify the equations and make it more accurate, which usually isn't necessary when the new center is at less than 1% of the radius of the Sun (note ellipses have 2 foci from Kepler/Newton). Even Jupiter's center of rotation is at the surface of the Sun, and both have eccentricities <1%.

Of course, if you're using epicycles, you're never going to get Einstein's elliptical precession correction to the orbit of Mercury. Simpler, is usually better for models, even when you need greater and greater accuracy, overfitting is the enemy of explanation (to coin a phrase).

All models are wrong, some are useful. - Box

Agreed. So to return to the root comment... if AI can generate for us all of those "wrong" models in a consistent way, then the job that remains for humans is to decide which one or ones would be useful. I bet it turns out to be a lot more work than it sounds like.
> If I just add a few more parameters to my model I can make it fit anything!

Dealing with people wanting to do this is story of my current professional life lol I would love to get them onto the same page that chasing down a final <1% improvement isn't (always) worth it

"All models are wrong, some are useful" - this is what I was trying to say, I was just being a bit snarky/pedantic because the commenter above seemed to be trying to pull a "gotcha" saying that the "truth" of something overrides the utility of it.

My point is that the laws of physics work just fine in a wide variety of mathematical framings. It's sensible to pick a framing that keeps the complexity of the equations as low as possible, but you're not incorrect if you pick a frame that puts the moon in the center or any other point. The accuracy of the theory's predictions are completely separate from where it applies labels like "center".

  > you're not incorrect if you pick a frame that puts the moon in the center or any other point
I think you're conflating different things. Yes, we can predict motion from the view of any point and that would likely involve epicycles. But that's not what people were after. Causality matters, a lot. It's kinda the whole point of physics (and what makes it so expressive and powerful).

What classical relativity does not allow you to do is place yourself at an arbitrary point and call that the center of mass of the system (i.e. things orbit around you). That's not invariant. It's a very different thing than predicting the motion of objects from your vantage point.

100% agreed, but Ptolemy and Copernicus weren't talking about mass or causality, they were just looking for somewhere to put the origin, and making predictions without reference to gravity. Picking an origin such that it happens to align with a barycenter would be in good taste for an audience with our modern understanding of gravity, but picking one elsewhere would not be incorrect so long as the predictions held up.

And that's the part that I'm suggesting AI might be unable to solve. Even if it can effortlessly generate theories that are consistent with both evidence and specification, there's still a lot of thinking to do about which ones we should specify for use in different situations.

Even if a "Theory of Everything" can exist, it may not be suitable for all situations for squishy human reasons besides anything having to do with correctness.

  > but Ptolemy and Copernicus weren't talking about mass or causality
Except they were

My critique is that there's a difference between "x is the center" and "motion looks like y when sitting at x".

Those are different claims with very different consequences. You're dismissing this as minutia but that minutia is the critical point. In a geocentric model you need to explain why things revolve around the earth. In a heliocentric model you need to explain why things move around the sun. Predicting the motion of the planets doesn't answer these. But answering these *correctly* allows you to build a generalized model in which you can predict the motion of planets, in any arbitrary system.

I'll also add that I love the geocentric model as a good illustration of why observation alone can't give you causality. There's far more to science than data fitting. In fact, that's why all the mathematicians are getting upset and what Tao has been talking about this whole time

I don't think Copernicus did any explaining why at all. That took another 150 years.

We now bind our physics to our astronomy, but they didn't then. That's what I mean by it being a matter of taste, not of truth.

If you were prompting a godlike LLM for a theory of the heavens in the 1540's you probably wouldn't have used words like "mass" at all, and the theory that it spit out would be similarly absent that concept.

Akababa
They both revolve around their barycenter. So neither one revolving around the other is ground truth. In practice either one can be the right choice (earth at center for video game, sun at center for astronomy predictions).
Galactic core at center for other astronomy predictions...
noosphr
And where's the barycenter located? When keeping an open mind make sure your brain doesn't fall out.
People are using natural language. And considering the barycenter sits inside the sun or close to it (relatively), it isn't that wrong of a statement.

While you're technically correct, everyone already understood. We could be more technically correct by talking about the galactic barycenter, or even more by talking about the universe as a whole, but the accuracy gains by that added complexity doesn't benefit the conversation

GR says that any falling (or orbiting) reference frame is equally valid, locally obeying Newton's laws. So in a precise sense, geocentric and heliocentric reference frames are equally valid. Both can claim to be "actual truth". Of course, a heliocentric reference frame is a lot more useful for studying planetary orbits since it gives rise to simpler equations of state. But that doesn't make the geocentric frame "wrong".

Put differently, would you say a heliocentric reference frame is also "wrong" since the sun is actually orbiting around Sagittarius A*?

Taking your argument to its logical conclusion, one would conclude that the [comoving reference frane](https://en.wikipedia.org/wiki/Comoving_and_proper_distances) is the only "true" one. And at last there is real validity to this claim, since the CMB does pick out a preferred velocity at each point in spacetime. But it is pretty impractical to use comoving coordinates to study problems outside of cosmology.

Ultimately, it stops being a matter of "right" or "wrong" but rather one of picking the right tool for the job you're working on.

It might have been a matter of taste in the 17th century. The Church was willing to accept a model where the sun goes around the earth, and the planets around the sun. This model agreed with observations, and was in good taste as well, as it avoided the problem of why it doesn't feel like we're in motion, and why we don't get flung off the earth.

Then Newton came along and blew all of the models out of the water. Nobody's at the center of the universe [0]. A single theory of gravity models both planetary motion and why we don't get flung off the earth. And we can "feel" the rotation of the earth by experiments which came along later such as Foucault's pendulum and the Coriolis force.

[0] A puzzle I like to taunt my friends with: The earth is in fact equidistant from the edges of the observable universe in all directions, so we must be in the center after all. ;-)

IIRC the universe is a pretty non-unifrom shape.
I think "there is no center" and "where you put the center is a matter of taste" are more or less equivalent. And if you're confronted with a pencil and paper and asked to draw the solar system, the latter is more useful.
More useful to us because we know more about how it works. This is why I was specific about the time period.
h3lp
Epicycles just had a poor predicting power; they required many more free parameters to describe reality, compared with the heliocentric model.

Knowledge can be framed as information compression (this actually applies to LLMs as well, amusingly). The heliocentric model plus Newton's gravity are an amazing compression of information related to dynamics of celestial bodies, which then enables related predictions that would be much more difficult to arrive at if you start from epicycles.

Is amazing compression of information a goal of the universe or of humans attempting to understand the universe?

Also, most models are lossy compression. The heliocentric model is amazing because it is wrong. Complications are simply ignored.

What you have to believe is that somehow the entire field of math was just leaving open problems on the ground that were actually easily solvable and were essentially free Fields Prizes, tickets to a lifetime of Fame and Academic superstardom. Orr LLMs are actually doing something novel.
kbelder
Maybe some of them are like finding the 1 millionth digit of pi. Not hard, but impossible without the right technological advancement.

Now that we have LLMs, those that are hard for people but easy for LLMs will get cleared out, leaving those that are still hard for both.

>Maybe some of them are like finding the 1 millionth digit of pi.

Yes some of them were clearly like that but there are probably 5-10 which many mathematicians have said were massive results and would all but guarantee a Fields Medal to any human that had solved them. Navier Stokes, Quasi Riemann, etc.

The epicycles model represents one arbitrary periodic function as the sum of multiple well understood periodic functions (uniform circular motion, i.e. complex exponentials). This is the same idea we use in a more systematic way today by way of the Fourier series, etc.

It's quite a good method for approximating any arbitrary smooth periodic function.

My best guess is that the originators of this model (perhaps as early as Hipparchus, but at any rate someone by the time of Ptolemy) didn't think there were literally circles being combined, but had a pretty clear idea that they were trying to make an approximate model to fit the data. A previous model, from Eudoxus, was also probably pretty accurate but more cumbersome to compute with, involving approximation of the visible planetary motion by the combination of uniform rotations of some imagined sphere(s) centered on the earth. (We don't know exactly because the books about it don't survive, so all we have are vague descriptions by other people.)

It's almost like the remarkable fact that many things turn out to be turing complete systems that you might not expect (that is, able to perform any operation that any other computer can perform, albiet more slowly). Board games, Chomsky's generative grammar, bacteria (in a sense). Natural languages, programming languages, odd combinations of logical operators.

Any number of things can serve as a basic vocabulary to postulate entities, and do so in a textured way that matches some parts of underlying physical reality.

It's tempting to look at this and make the tragic jump into anti-realism because of the interchangeability of so many forms of representation. But I think as long as you have a combination of pragmatic attitude, treat knowledge claims as provisional and converging on truth, there's a cash value to descriptions that makes them legitimately truth tracking regardless of how they're formed.

> It was completely wrong but the outputs were surprisingly close to reality.

that's not how physics works. epicycles weren't wrong. there is no wrong/right - modern scientists don't believe in platonism. physics is a collection of models which are empirically tested. if epicycles makes predictions to more decimal places than relativity then epicycles are "right" and relativity is "wrong".

now think about how this translates to LLMs...

> physics is a collection of models which are empirically tested. if epicycles makes predictions to more decimal places than relativity then epicycles are "right" and relativity is "wrong".

Most models have relation to each other so it does not suffice to take one in isolation. It's one continuous world (in our knowledge) so almost everything relates to one other. The truthiness of something is not whether it fits some particular phenomena, but how does it generalize.

> Most models have relation to each other so it does not suffice to take one in isolation

You don't know what you're talking about (most people on here do not).

https://en.wikipedia.org/wiki/Ansatz

> You don't know what you're talking about (most people on here do not).

> https://en.wikipedia.org/wiki/Ansatz

  After an ansatz, which constitutes nothing more than an assumption, has been established, the equations are solved more precisely for the general function of interest, which then constitutes a confirmation of the assumption. In essence, an ansatz makes assumptions about the form of the solution to a problem so as to make the solution easier to find
I don't see where this contradict my statement. A good answer does not means the formula used is correct.
> A good answer does not means the formula used is correct.

yes it does. that's exactly what it means.

Precession of perihelion of Mercury was not the motivation of the discovery of general relativity theory. It's only an evidence that convinced why the theory ought be true.
The introduction of LLM in mathematics is no different than the introduction of automatic calculators. Probably the calculators (computers) where even more groundbreaking if you think so (something that was impossible to calculate or simulate it was now possible).

Just don't believe at all the shitty marketing that AI companies are creating to make you believe that they have found the Holy Grail

Of course. But we might have found out the sun is the center of the solar system much sooner. Who knows!

I feel like there's a lot of "humans are special" in this thread. There's a good chance we're not.

>If we were using LLMs to analyze astronomy, we might just get increasingly complicated epicycles and never realize that the Sun is the real center of the solar system

I think this is a fantastic point and this kind of complexifying can and will happen. And I think future of online debating and propaganda is going to complexify in a similar way.

I will say though, I think that can be controlled by building in principles that prefer competing theories on the grounds of cogency, and that cancel out competing theories by reducing them to the parts between them that are equivalent.

My hard disagree will be with this: I don't think at all that we would "never" have got to the theory of relativity. It reminds me a bit of the "embodied cognition" argument against simulated brains. That argument suggests you can't "just" simulate brains, because actual brains are the totality of their embodiment in bodies and environments. Regardless of whether you agree with that line of thinking (I don't), we can assume it's true, and change the target of our simulation to brains + bodies + environments. The problem is bigger, but still perfectly amenable to the same methods.

I think the same is true of theorizing about science and imposing metatheoretical constraints. I think you can optimize for healthy hypothesis production just like you can for first-order reasoning about data. It's definitely harder but not necessarily a difference in kind.

The more I read about LlMs the more I wonder - there’s maybe two levels of emergence. LLMs obviously do more than the underlying vectors suggest. But it seems like they could have a hard limit of their training data. Filling in gaps in our knowledge to be sure, but always feeling like maybe they’re just about to hit the frontier.
Current models have undergone massive amounts of RL to "get shit done." There's incentive toward reasoning, but probably little toward elegance. Every engineer is familiar with the impulse to just hack things together.

But this need not be a permanent state of affairs. The model labs are probably already examining the elegance problem, since it is such a common complaint of researchers reading these machine generated proofs and papers. And I haven't seen anything disproving the notion that elegance can be RL'd.

I don't think checking our work is a uniquely human trait. It's our uniquely human traits that made us willfully ignore the better model in favor of increasingly complicated epicycles.

> If we were using [humans] to analyze astronomy, we might just get increasingly complicated epicycles and never realize that the Sun is the real center of the solar system.

> a very clear pattern of what they are good and bad at has emerged. They are extremely good at combining large amounts of information

This has not been a pattern at all. Many have speculated that they would be good at this, but in the past they actually weren't! They were unable to synthesize their encyclopedic knowledge of everything into cross-disciplinary new discoveries, without explicit prompting about the kinds of knowledge to combine. They were surprisingly bad at this!

These math proofs are the first evidence I know of for LLMs actually taking advantage of the fact that they have more knowledge than any one human could have to combine multiple different directions in unprecedented ways to solve real problems. This is new, and exciting.

"I'm not a mathematician, anyway heres how AI models are solving decades old open problems in Math." Bro your humility is staggering.
n0dose
I mean he’s basically quoting what Wolfram (a man famed for his humility) said on a panel just after these dropped.
Also famed for being an intellectual juggernaut. A random poster on YC not so much.
> Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles. Pretty much never.

It's really good at doing this at project definition time and annoying the shit out of you by not understanding the specific constraints about your problem. That needs to be done first. Then it can start suggesting improvements alongside the high level "vision". The point of high level visions is that they move fast and are hard to specify - yet people treat it as if these things don't exist. If you don't trust your own consciousness, then I don't know what to tell you

But I'm not sure this is safe either. I'm sure AI can get to the positive case of contributing positively.

The flip side is that, well, your specific wishes don't matter because AI is just a superset of who. Who cares what meatbags what?

I agree whole heartdely. But that is about the current state of IAs. I do think the next missing (and probably last) big step is exactly how to give LLMs a better abstraction capability.
I look at it like:

If you prompt an LLM to build a web game (or something) it iterates over and over until the game works. And these days the output does indeed work. If you look at the code, it can often be a jumbled mess, but it still does work. It will take some sleuthing to understand what it is doing, and, if you care, clean up the code.

I think OpenAI did the same thing with Lean. The implications are far more interesting, but from what I understand the proofs are “slop”. Correct, but not elegant in any way.

It is still a wildly amazing technique to find a solution (or possible one) to maths problems.

I have used LLMs to do things I didn't know how to do, then reviewed what it did and learned a new technique.

No reason it cant work like that for maths.

I agree.

How many times I have hit my head against a wall trying to figure out how to get some code to work, only to come back with a clear mind and look holistically to see that it didn't need to be solved to begin with.

None of that is to imply that solving any of these problems isn't groundbreaking or helpful along with the plethora of other things AI has undoubtedly advanced, but I think the breakthroughs will be inferred by many to mean "point AI at really hard problems" is always the solution instead of continuing to use your brain on whether or not the problem _is_ the right problem in it's current form to solve.

Computer Science isn't the field you look to for "novel abstractions that no one has thought of before". The vast majority of CS concepts were settled over 50 years ago. I've never encountered someone who has generated "an abstraction that no one has ever thought of before", in my 40 years as a professional developer. Its always clever applications of abstractions or algorithms to a problem. In the majority of cases it is because faster hardware has opened up possibilities to use an approach that wouldn't have been practical before.
rramach
I think we all need to update our priors on what LLMs can do today. Here is one anecdote:

Scott Aaronson says "Dana tells me she now mostly understands the proof of the UGC, and is amazed by the new ideas in it, and wants to give talks about it soon."

Dana Moshkovitz is a professor at UT Austin who is an expert in the area and has been working on solving this exact problem for decades!

> Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles. Pretty much never.

Uh, quite often for me. I explain the problem and my thoughts on the solution and it tells me whether that's good or if there's a better first principles way. I didn't explicitly prompt it but it knows enough from its thinking trace to say, wait.

A very balanced perspective, and the concerns he raises are reasonable. He acknowledges that AI is going to transform mathematics, but simply dumping proofs on the math community and expecting others to do the grunt work of verifying, refining, and expanding on them is hardly a productive way to advance the field.

There seems to be more interest in hitting some arbitrary benchmark (we proved X unsolved problems) than in genuinely contributing to mathematics. But what else is to be expected? It's become a maniacal race with too much money. Too much effort is being invested in proving that the exponential curve is still holding.

XorNot
That seems short sighted though. A few years ago models couldn't do this at all, I'm not sure there's any evidence to suggest exploring and refining results is outside their capabilities or will remain so.

OAI obviously have a fiscal incentive here, but to presume a year from now we won't see improvements and more succinct work on the results coming from models?

> I'm not sure there's any evidence to suggest exploring and refining results is outside their capabilities

OP didn’t suggest that.

The bar has been raised. Everyone has to meet it now. An inelegant solution squatted onto the internet doesn’t count as discovery per se, even if it’s impressive.

A correct solution verified in Lean will count in perpetuum. It is fine if you want more, but an achievement is an achievement, even if it is by AI.
Hmmmm. It feels right that discovery is much more meaningful than achievement. "Bullshit lean proof" or "bullshit achievement" smells like it. "bullshit discovery" smells like a front-handed insult
Maxion
But what should they do? They got all these proofs, should they just have sat on them?
> should they just have sat on them?

It’s fine that OpenAI posted their findings. It’s not fair to claim these problems have been solved. Not until someone can understand and verify the proof and then communicate the core, novel methodological element to someone else.

Solving a problem is not solving anymore. Up is down, pleasure is pain, darkness is light, slavery is freedom, madness is sanity.
gwd
But this is Tao's point: Before, the mechanism by which a proof was verified and communicated and digested by the community was for the person who came up with the grotty, ugly first draft to engage with the community. Now there's nobody to really engage with, so the pipeline from "grotty, ugly draft" to "integrated into humanity's mathematical knowledge" has been broken.

So yeah, probably we should stop saying "X has been solved", and instead say, "A Lean proof for X (or !X) has been generated". That doesn't change the fact that incentives are currently on finding the proof, and once the proof is generated by an AI, there's not currently a good mechanism / incentive structure to move that into the mathematical community. AI is here, so we need to find a new mechanism.

It is not obvious to me that a single canonical human language write-up of a proof is the best output in this new world where write-ups are cheap. A human reader can query an LLM and get explanations of key points tailored to the reader's own background in mathematics.
> A human reader can query an LLM and get explanations of key points tailored to the reader's own background in mathematics

Convenient. To verify my shovel works you must buy…more shovels!

Many of these are lean formalized. Arguably a much higher bar than whatever peer review provides in terms of verification.
On some consideration, surely. On the other hand, but at some point this is borderline like saying "universe already solved every physical problems, including possibility to represent deep important point of its own structure in compressed intelligent ways" and then tell that reaching it in an actual grabbable artifact is left as an exercise.

Possibly yes such a representation is possible. But it doesn’t mean it’s certain there is a "best compressed representation". And even less one that encompass everything important and that is understandable by any human brain, even the most exceptionally brilliant ones sponsored by a whole society to reach their best possible achievable performance on that goal through full dedication on that sole task.

There is no guarantee that the lean proof is 1:1 with the natural language equivalent. The lean proof can be lesser. This happened in the Navier-Stokes proof, e.g. see [1] in example 3.1. Having the certificate doesn't necessarily imply correctness.

[1] https://arxiv.org/abs/2610.08144

p_hoep
No need. In the end the mathematicians that don't like this can just not look at the proofs or use them. They have that choice. Just like they didn't "ask" for them, they don't have to even acknowledge they exist.
>They got all these proofs...

Your phrasing is illuminating that perhaps they aren't engaged in the creation, understanding, or integration of these proofs by humanity; they just have them. For them, this is a slidedeck they can pass to investors, creditors, the marketing department. Something they can add to the employee onboarding pamphlet.

What should they do? Hyperbolic maybe, but perhaps engage with humanity.

nearbuy
...they did.
BlackBox: The proof is in the pudding

This isn't only bad for Math — it's bad for English too.

'Proof' is going to become the 2026 Most Misapplied Word of the Year.

pks016
At least check them properly. They have already withdrawn some of them.
pfdietz
That were not formalized.
rtpg
The problem is that one a person writes a 60 page proof in theory that person has spent an inordinate amount of time on the proof and can answer questions, describe some insight, etc etc.

If a random person is given a 60 page proof to digest and not the author, those hidden insights that _aren't_ in the paper might be completely inaccessible. Maybe the AI will "just" be able to provide the insights. Maybe. But pedagogy is tricky work, and despite these AIs being able to do all this fancy math we can't get them to write good cover letters yet, so....

Ultimately we might be left with just a bunch of intellectually unsatisfying proofs. This means way less drive to simplify the proofs or rework them.

End result: we generate a layer of "less efficient" mathematics, that won't get built upon. We will not actually have any shoulders upon which to stand.

XorNot
But why should process of discovering mathematical insights be any less attainable to AI models?

The concern is being raised without evidence, because the evidence points to the gap simply being frontier models have just started to be able to get a raw proof out. Why, given existing progress, should we expect them to be unable to distill insights from those proofs?

Certainly this even more likely doesn't matter at all for applications: if I can send a radio signal further because my AIs design it a certain way, that's an unambiguous result. Which is really the next step here: turn a proof into a "mechanical" application.

AI can simplify and rework proofs too.
According to Scott Aaronsson, OpenAI set their agents on 8000 different problems, and got 372 final proofs. Even spending twice the original effort on simplifying and reworking those proofs so that they do not "feel like something written by someone who’s on psychedelics" would only increase the compute by less than 10% (assuming all the agents had a similar token budget).

The fact that they did not do so can only mean that either (1) their agents currently lack the capability to do it, or (2) OpenAI are completely indifferent and do not care in the slightest if the proofs are understood or not.

gwd
> The fact that they did not do so can only mean that either (1) their agents currently lack the capability to do it, or (2) OpenAI are completely indifferent and do not care in the slightest if the proofs are understood or not.

Come now, this is kind of unreasonable. When you're working on a new technology, you first get the ugly, inconvenient-to-use prototypes functioning with the core new thing you need; then you work on packaging it up into a format useable in production. I'm sure the very first digital camera sensors weren't very useful for photographers either; but it isn't really even possible to build the rest of the technology required to turn raw output of a digital sensor into something a professional photographer can use until you have the raw output itself.

The research is still on going on the raw output; getting things to the next stage, where the results are widely useable by professional mathematicians (and then on to engineers and scientists to whom the results would be practically useful), is a whole new research area.

Seems like you are just subscribing to the first option I gave, "their agents currently lack the capability to do it", but saying you think they will be more capable in the future if they can move away from the inconvenient-to-use prototypes after more research. Thinking it might be possible in the future is not in disagreement with anything I said, so I am not sure what you thought was unreasonable about my description.
gwd
Imagine someone looking at digital camera researchers showcasing a ground-breaking new sensor, and reacting by saying "Well obviously they're completely indifferent and do not care in the slightest if their work is used by real photographers or not."

Like, "Orr... maybe they care a lot, but haven't gotten to that part yet?"

You are conflating the two options with each other (lack of capability vs indifference). One of them is true, not necessarily both ("or", not "and"). I did not claim a lack of capability was the same as indifference.
Or (3) they feel a need to publish first, and that goal takes precedence over (2).
They claim each result used three hours of compute on average. Even spending significantly more on simplifying and reworking would delay the release with a single day at most. If avoiding such a minor delay took precedence over (2), I think "indifference" is the correct term. It has also been a while since the release now, so there is ample opportunity to post follow ups if time pressure was the only concern.
Why do that when they can spend more hours extending QRH to a proof of the full Riemann Hypothesis? The opportunity cost of digging up small potatoes is the whole enchilada.
Why not asking another chatbot to verify, like what we are doing with coding ?
kadoban
That's not the hard part. The hard part still requires humans (and/or _maybe_ a bunch of tokens) and a lot of work.
And then what?
gexla
Right, easy comparison to make the the open source community for software.

And it's not like this is something where we're loaned some top math genius for a limited amount of time and we have to make the most of it. Rather, this is a new high water mark. The accessibility of the results is no longer scarce. The scarcity has shifted, and that's where the focus of the math ecosystem should shift as well. And it doesn't help for a frontier community to saturate and take over messaging pipelines that were typically managed by the math ecosystem. It's not about "stay in your lane" but rather "we need coherence and be careful not to break the system."

Just two cents from someone who could screw up basic cashier math on any given day.

pfdietz
OpenAI has kicked things over. Dealing with that is the math community's problem. OpenAI had absolutely no obligation to defer to that community's deficiencies. If this has upset their reward system and historical traditions, tough luck.
flexie
Yes, it is indeed a balanced view.

Those few thousand mathematicians are now getting a taste of their own medicine. After all, it was people with extraordinary mathematical talent who developed machine learning and large language models, leaving hundreds of millions of people who earn their living through speaking, writing, or teaching worried about their future job prospects.

Still, I believe almost everyone will be fine. Perhaps AI will also prove good at coming up with new conjectures, and some mathematicians may shift towards applied mathematics or other sciences.

Pretty tenuous connection to blame mathematicians for everything bad created in the world, that happened to use math.
caaqil
> There seems to be more interest in hitting some arbitrary benchmark (we proved X unsolved problems)

> genuinely contributing to mathematics

What's the difference between the two? Proofs are no longer the goalpost?

shubhamjain OC
Not an expert, but explaining the proof and it being independently verifiable as important as just putting a paper of it. Reminds me of 1000 page proof of Goldbach’s conjecture that a Mathematician reached sometime back. He was told plainly that no one is going to invest time in verifying the proof because there’s a good chance there’s an error somewhere in between.
tene80i
Proofs are valuable but I believe the mathematical community values understanding more. Proofs were previously a great way to develop understanding. Now, less so.
pegasus
I recommend RTFA, it explains exactly that: why proofs should not be considered the goalpost, and why dumping all these AI-generated proofs might be an overall negative for mathematics as a whole.
Proofs of open problems are valuable because we are assuming that proving the problem requires some new method or infrastructure in math to prove it. Basically proving open problems isn't actually useful if it doesn't develop new tooling for mathematics, which can help us create new open problems, solve other ones etc.
Its kinda like the arms race in the cold war. There came out some truly marvelous technologies but the actualy goals were frankly terrifying.
antman
This argument implicitly makes a few assumptions which will probably not hold in the very near future.

One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that. What humans in the loop are doing now is verify the process, propose shortcuts and add legitimacy, through the verification process, if that ends up being succesful its highly likely a lot less mathematicians will be required in the future.

The conclusion that this is not productive focuses on the mathematicians, but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds. Unless it ends up being the greatest hallucination ever ofcourse

AI cannot explain chess moves it comes up with in an elegant way. What makes you think it will be able to do so for math?
It's kind of astonishing that after all we have seen in the last years people still find the position that AI will not be able to do an obviously valuable thing likely and it requiring an explanation (instead of the other way around).
No, this is different, and this is coming from someone who has been studying deep learning for the last decade. We are talking about the difference between RLHF and RLVR strategies. The former benefits clarity and explanation, while the latter concerns only correctness. AI was moving in a particularly damaging direction by pushing on the first path, so it was natural to move to the second. But the second will come at the cost of clarity of explanation. It will likely get better at its explanations, but not fast enough to render its most advanced accomplishments readily understandable to the user. The chess example is a pretty good one (that is an RLVR approach).
The problem is that people strongly believe that this is an insurmountable problem that will persist indefinitely (or for a long time) and plan accordingly, while this, most likely, will be fixed soon by adding RLCAF (RL on conversational agent feedback) or something like that.
__s
tbf GM explaining their 2700 elo moves are only understandable when vague, as elo goes up explanation becomes closer to "in this specific position there's these dpecific lines", why should 3500 elo moves have simple reasoning?

Maybe if we start with giving simple AI generated analysis of those clumsy humans with their measly 2700 elo moves

n6242
Some of us still remember 2016, when we had a couple of cars sorta half-driving themselves, and Tesla, Uber and others promised we were only one year or two away from three million people in the US working as drivers being out of a job. And here we are, a decade later. AI is pretty amazing, but companies have a tendency to severely and comically overestimate and oversell it's capabilities, and underestimate the challenges.
> And here we are, a decade later.

With Waymo and Tesla increasingly doing what they said they would do, and a small number of early adopters happily paying money for their services, that do work.

So what's the critique? That the timelines are not correct? Sure. And how about the timeline of the people who said "research level math, never in my lifetime" and the people inside the ai companies who are apparently increasingly spooked by how quick the progress is? How about the various levels of code/programming jobs that AI was supposedly never going to be able to do, but, in reality, now just does?

We are engaging in some very one-sided discrediting, and I am not sure, why.

And, those companies are notably avoiding wet climates because they still struggle with self-driving in inclement weather. It's probably going to be another decade before we get to the point where these things can handle every situation. Just like all other engineering, the first 90% is the easy part.

https://www.wsj.com/articles/self-driving-cars-dont-do-snow-...

Alas, you will notice that article is 1.5 years old. Update on this issue (albeit soon from a year ago too): https://waymo.com/blog/2025/10/creating-an-all-weather-drive...

More telling: Waymo just rolled out in Denver (1 month ago or so), apparently fairly confident they got this handled given the upcoming winter.

Progress on the obvious stuff keeps happening (which kind of brings me my to the first comment here).

But also: It does not have to do "snow" to be useful! A lot of cars/people don't drive when it snows heavily, and that's something we have always been okay with (at a societal level, YMMV of course). If was only useful 95% of the year that's still great. A lot of technology works like that.

Great response, and thank you for the article! I'm skeptical that they're actually ready for real-world conditions outside of a test track, but that's a whole lot of training data and their lawyers must be convinced. We'll see.
p-e-w
Driving a car is unimaginably more difficult than proving the Riemann hypothesis.

You just don’t notice that because evolution has given you 99% of what is needed to drive a car before you were even born.

I remember a biologist wryly commenting that it took a lot longer to evolve good senses and appendages in nature than to evolve human intelligence, compared to the ape niveau.

Maybe the really hard thing isn't abstract cognitive capability, but perception+movement.

pfdietz
Moravec's argument strikes again.
Interesting claim. That's true for old models that simply have no way to explain, LLMs however can. [1]

[1] https://dev.to/natcher/researchers-develop-method-to-train-l...

> One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that.

There is literally not a single shred of evidence to indicate either of your supposed eventualities. The core technology of an LLM is sampling from a distribution so there is literally no way to make it deterministically robust (only probabilistically).

antman
The direction and pace of capability improvement has already been demonstrated by all models. The latest breakthroughs make that pretty evident, but there have been production systems that are based on probability since the beginning of computing.

What has been demonstrated is a process that outputs lean proofs based on those probabilities. This happened after decades markov chain producing garbled texts and very shortly after gpt2 producing stories about unicorns.

> The core technology of an LLM is sampling from a distribution so there is literally no way to make it deterministically robust (only probabilistically).

An LLM mostly deterministically (except parallel processing nondeterminism that can be mitigated) produces a probability distribution that can be sampled deterministically: just take the highest probability token or use beam search.

Vetch
Deterministic yes, robust deterministic no. The most likely conjunction is not always the best nor representative of what the model is considering unless its certainty is high.
I think determinism has nothing to do with it. If you mean sensitivity to word ordering and such, it's a generalization failure.
I think people on here tend to somewhat fixate on the determinism issue. Even with a deterministic LLM - stabilising the floating point arithmetic, and choosing from the distribution by a fixed method, or just save the random seeds - there is still a kind of a chaotic unpredictability that can exist between its inputs and outputs. However maybe that is a price that needs to be paid to get creativity.
Is a human deterministically robust? Or is a human also incapable of doing what you claim LLMs will never be able to do?
emtel
No evidence other than the fact that this has been happening steadily in all areas for many years?

You might have a point if the goal was to have LLMs that spit out a correct proof without chain of thought or tool use. LLMs + agent harnesses are more than capable of self verification and course correction.

Terr_
> assumptions which will probably not hold in the very near future [...] One is that AI will continue hallucinating in a manner that is not easy to verify

Hold up, that's an even bigger assumption in the opposite direction, and I don't see anything to support it.

At least in terms LLMs getting all the "AI" hype these days, there is no structural/mathematical reason to believe they won't continue to have the same problem they've always had of generating plausible text over rational text, and I don't think anybody even has a clear idea how it could eventually be accomplished.

I've seen "then the magic singularity occurs and somehow it solves the problem for itself", but I would classify that more as mysticism than engineering.

"Plausible" text was preferred over rational text when we trained LLMs using RLHF. It's rapidly shifting the other way now with RLVR, which enforces correctness by default.
Hallucinations are no longer much of a practical problem in software engineering.

Two years ago, hallucinating that the code worked or that a task was accomplished was a common occurrence.

We have seen that now agent swarms across thousands of agents can coordinate to achieve a result.

Clearly hallucinations are no longer the problem they once were, since now we can get working results for long horizon tasks that require massive compute.

Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.

It’s still a common occurrence.

It happens in more subtle ways, but it still happens often enough for me to notice. For example I have had hallucinated checksums show up in lock files as recently as yesterday using a SOTA model.

This is not surprising, since the whole basis of LLM training is to produce output that humans will accept _as a proxy for actual training goals_. In a sense, the training process of an LLM “wants” to produce output that is statistically plausible much more than it “wants” to produce correct output. It’s always going to be a struggle to drive that system towards other goals (and we see this bourne out in practice by the amount of effort that is required to be spent on RL).

I think there will be some threshold of correctness (something like 99.999% of the time) that if the model surpasses it, I can stop needing to check it, but I think we’re still at 99% or something which sounds good, but when you are producing a ton of output you hit that 1% frequently.

> Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.

I 100% agree with this. In fact explaining things to humans is something LLMs are particularly well suited for.

The confidence with which you, anonymous user, keep commenting that "hallucination is not much of a practical problem in software engineering anymore" based solely on your own anecdotal evidence is really remarkable, in not a good way.
You're free to substantiate your comment by telling us about your apparently different experience.
I find hallucinations in my (mostly perfect) AI output every single day. If you're not finding them, you're just not looking hard enough. It's not surprising when everyone is screaming about how they don't read code these days.

This is just a fact. I'm sorry if it messes with your narrative.

https://arxiv.org/abs/2401.11817

Thanks for the link. What kind of hallucination are you seeing, and does it affect the end result?
I'm not them, but I see quite a few hallucinations as well. Some examples:

I work on a device that gets firmware updates over USB-DFU. If a DFU fails, its bootloader restarts, re-enumerates USB, and waits for a new DFU to begin. AI decided there's a risk of bricking the device if an update fails. There's no brick risk, a retry fixes it.

The same device uses CAN bus, with the common bosch_mcan peripheral. Claude decided that calling can_mcan_stop followed by can_mcan_start somehow left the transmit buffers intact, and tried to implement a fix which manually cleared the buffers. They're cleared automatically by the hardware when can_mcan_stop is called, and that's documented in the comments of can_mcan_stop.

Both cases could have resulted in unnecessary changes getting deployed which wouldn't have fixed the actual issues. Since it's embedded that could mean significant delays to getting fixes to customers.

The sum total of all human observations is still not proof of the lack of hallucinations as a problem (even if their observations were perfect, which they aren't considering the volume produced vs reviewed carefully). That's why you can use a counter example only to disprove and not prove anything.

And yeah I get hallucinations all the time still. Maybe it's because I'm working on harder/more niche problems (like a compiler with an unusual type system), but it happens quite a lot. I don't record all of them.

Although the most common one you can find is them misattributing the source of changes from themselves and also other agents (Fable, Opus 5.5, deepseek, whatever). They'll say "your changes" or "you changed" or "your ruling." I didn't decide anything and it's in their own chat log, and yet...

Retr0id
Even if you somehow have a 100% correct AI, it's not useful unless we can understand and internalise (and communicate) its results.
Who is this "we" you speak of? The professional mathematician community? Were pre-AI results useful outside of this community of people who could understand them?
dofm
Something about this sentence makes me think about that Rob Auton bit, that before there were mobile phones, nobody had any reason to tell someone else that they were on a bus.
Before mobile phones you never called someone and asked "where are you" because a phone was tied to a location...
Yes. Most probably do not understand the notation involved in, and the statement of, the Lindeberg-Levy Central Limit Theorem. But every scientist uses this theorem in one way or another. These ideas have a way of trickling down because to people who work thanklessly to do so.
dofm
> One is that AI will continue hallucinating in a manner that is not easy to verify

It is an old saw at this point, but what an LLM does still cannot be divided into hallucination and non-hallucination. This is literally an anthropomorphism trap.

Layers and layers of application-specific verification can reduce the risks inherent to LLMs, to a really remarkable degree, but nothing about what these tools are suggests that this problem will go away; it will just bubble up again somewhere else.

antonvs
> It is an old saw at this point

An old saw unless something that's widely accepted, but sadly it seems that many people don't recognize this, even many people working in the field.

And why not?

For all that I saw over the last few hundred hours with AI on software engineering, hallucinations are no longer a problem at all.

Not once have I seen a task fail due to what would have been a "hallucination". If they still occur, they can apparently be detected and corrected automatically, or are subtle enough to escape notice with presumably no significant impact on the results.

Why would this not also be the case for mathematics?

I think OP is saying that hallucination or not is just semantics. There is nothing qualitatively different about hallucinated vs non-hallucinated output.
That's true in the same sense as "There is nothing qualitatively different about erroneous vs non-erroneous output" for a dog vs. cat image classifier.
To be fair, I guess the line is blurry between what could be labelled a regular mistake compared to a hallucination.

"Test suite passed" when it actually errored? Obvious hallucination, unless it ran a command that returned the wrong error code.

But is running a malformed command that does not achieve the expected effect itself a hallucination?

If it makes a false claim, then it's an error. If it says the test was passed or a class was implemented but it was not, then it makes a false factual statement.

I'd say a hallucination (very misleading word) or confabulation or "making shit up" happens when an LLM uses factual / evidential language purely based on local statistical expectations of the text, instead of it drawing from actual evidence in its context pointing to it.

This is murkier in the case of general knowledge questions, like when and where was some famous person born. It may then be a spectrum from fully making something up based on how the name sounds, all the way to confidently retrieving it from its weights correctly. In between, we can get hallucinations. But newer models are taught to use Web Search when unsure, and it works pretty well, though not perfectly. I don't see any fundamental limit here. It's just not perfect. Trying to solve "the hallucination problem" is basically like saying "our dog vs. cat classifier is pretty good already with its 99% accuracy, now all we need to do is the tiny little task of eliminating the 1% error, and we will be golden". Like, no shit, there is some error yes. People are working to reduce it. It will never be absolutely 100%. It's not an insight to say we should remove hallucinations.

Terr_
Intuitively, something about that version bothers me... I think it's because the choice of a clearer "erroneous" has dropped the fundamental framing problem from "hallucination": The false implication that "true" (anthropomorphic) sight/thought usually happens.

To fold that back in (and ham-it-up a bit) how about:

> There is nothing qualitatively different when Pet Classifier <ironic-quote>maliciously misreports</ironic-quote> your cat as a dog, compared to when it works <ironic-quote>honestly</ironic-quote>.

These models are outputting language that we can decode as factual claims and we can check those, and we can say it made a mistake / error or that it answered correctly. This doesn't require any squishy assertions to how it "feels" while doing it or anything like that. It's an externally observable thing.

I agree that hallucination isn't some kind of "different" operation than "normal". It's not like when a train derails and you can point to it. It just operates as normal and sometimes that yields correct factual outputs, sometimes not. You don't have to metaphysically ascribe any kind of intent to it.

I'm not sure how people conceptualize these things who weren't doing classical machine learning before all this. To me, "hallucination" is shorthand, and we know it's not like humans on drugs or something. It was used in the literature also for any kind of generative imputing of missing information from a learned prior. For example in image inpainting a GAN "hallucinates" the missing part of the image. This terminology was already used in the 2010s and probably earlier. Or in image colorization of grayscale photos, the model "hallucinates" the color information.

Then the word escaped into the mainstream and people have weird connotations about it.

Terr_
> To me, "hallucination" is shorthand [...] Then the word escaped into the mainstream

One of my bugbears is when people abuse the term "Ponzi Scheme" to refer to literally anything the think is unsustainable. (As opposed to something that, at a minimum, requires someone telling factual-lies about assets.) Kind of like if folks started calling every kind of software error a "Buffer Overflow."

I think the colloquially Ponzi Scheme is a synonym for any pyramid scheme, where the people in the smaller upper layers rely on an ever growing sequence of lower layers entering to receive their payouts, and it works as long as there are newer people to bring in but at some point the recruitment pipeline is exhausted and the bottom layer ends up losing their entire investment and the higher you are in the pyramid the better you end up.

This is quite a specific sense, and isn't equal to just any kind of "unsustainable" system. Things can be unsustainable for reasons other than being a pyramid scheme relying on rapidly growing numbers of people being brought in by the ones already in the system.

Terr_
(Going off topic, I know.) It's a less-egregious substitution, but Ponzi != Pyramid in some important ways. If they were the same, we wouldn't have needed to invent different terms. :)

    1. Actual Fraud
     * A Ponzi Scheme requires lying to people about what assets currently exist, or how much of the asset is reserved as "theirs". 
     * A Pyramid Scheme just asks people to collectively speculate (without guarantee) that tangible profits will arrive.
    2. Layers
     * A Ponzi Scheme can be one criminal and an indefinite number of unwitting victims. [0]
     * A Pyramid scheme (like the structure) requires *growing interior layers* of participants that blur the line between victimizer and victim.

[0] Insofar as some Ponzi Scheme targets do not suffer losses, that's not because they were necessarily complicit, but because the criminal strategically chose to let some marks exit in order to maintain the illusion.
I think that’s a bad example, as classifiers tend to output a floating point number and you use some threshold/activation function to collapse the classifier into a particular state. In that sense there is nothing qualitatively different when the classifier outputs 0.85 vs 0.87.
LLMs also output a distribution over tokens at the output. I don't see the point you're making.
qarl
> can reduce the risks inherent to LLMs, to a really remarkable degree

To an arbitrary degree.

Just like all of science. Reduce the error to the desired margin.

Expression of tech-faith is not intellectually honest argument.

Where does this "will probably not hold in the very near future" come from? People correctly warn about extrapolating current things onto the future, but then just throw some vague "probabilities" without providing any argument why their "probably" is somehow more grounded than others.

antman
Markov chains garbled text to gpt took decades, gpt stories about unicorns to gpt production systems took a few years, gpt production system to gpt astra producing deterministic lean proofs of longstanding mathematical problems happened even faster. Scepticism to the point of requiring proof appears like an academic pursuit while production systems have already been built and are in the process of being enhanced
Vetch
Putting hallucination aside, LLM "theory of mind" has gotten worse over time. I feel it peaked in Opus 3 and Sonnet 3.5, GPT 4 and then GPT 4.5 for OpenAI. Since then, even with Opus 5.5, phrasing has needed careful crafting, in order that it not be taken too literally. OpenAI models suffer from this much more than Anthropic models but Claudes have backslid over time too.

This means when writing documentation, tutorials or commit messages, their output is often a garbled jumble. Assuming shared context, using invented terminology without explaining, leaking conversational states due to improper epistemic boundaries and failing to model the reader. This all usually leads to their freely generated explanations being terrible. Getting good explanations requires chaining questions that force them to line things up properly, which is not easy the less you know. These failures as something LLMs naturally struggle with make sense, given the nature of attention and RL with weak signals from human data.

Math is not merely a collection of proofs, it's a way of understanding. A proof presented in a manner that cannot be incorporated remains useless. It does not make it's way to physics like Riemannian geometry and matrix math did. This is no less true when done by humans too.

Your hallucination conclusion, checking if a proof is one, is exactly the counterproductive cost.

Most of us cannot verify that the claims in the OpenAI lore dump are in fact all correct. It will take tons of work from experts to do this. It took subject expert mathematicians to identify the discrepancy and disconnect in the Navier Stokes proofs, for example. LLMs will struggle to make use of their own proofs or turn them into knowledge that accumulates over time.

The act of proving is often more valuable than the proof itself. Human constraints and limitations force us to invent tools and abstractions that a 100,000 x 1M context swarm can bypass. The tradeoff from that AI swarm advantage is work that doesn't usually lend itself to being built upon. It's like doing all the side quests and reading all the books of an RPG versus min maxing a straight path with a guide. We might try to identify new abstractions, but the fact that we don't get access to CoT and that much of it will be illegible means mining LLM traces for what human mathematicians produce naturally will be a tedious chore.

p-e-w
> Since then, even with Opus 5.5, phrasing has needed careful crafting, in order that it not be taken too literally.

This is a feature, and a huge step forward.

If you expect AI to do serious work, you can’t have it guessing what you “really meant”. Every sufficiently advanced task depends on very subtle details in the problem statement, and the correct default behavior for advanced AIs is to solve the task exactly as stated, unless a system prompt or other constraint tells it to do otherwise.

pegasus
Did you even RTFA? His argument absolutely doesn't make any assumptions about hallucinations, implicit or not. It's you who assumes Tao must have surely been complaining about hallucinations or some such. You've not addressed any of his arguments and moreover ask questions his post answers.

Here's a longer article which goes into a bit more of the details: https://terrytao.wordpress.com/2026/10/05/the-future-of-math...

> but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds

You’re making the following assumptions:

1. the exercise of struggling to find proofs was not productive, but this is precisely how new techniques in math were produced. Brute forcing solutions doesn’t lend itself to the creation of much new mathematics (except maybe the exercise of developing verifiable proofs)

2. the point of doing mathematics is to be “productive” in the first place. This is silly. Many people get into mathematics because of the beauty of understanding, for example.

> 2. the point of doing mathematics is to be “productive” in the first place. This is silly. Many people get into mathematics because of the beauty of understanding, for example.

Are they independently wealthy? Or do they have a deal with their local supermarket that they can take food for free?

So maybe we should just pay them more to do math and let them set the direction of research?

If your point is that capitalism fucks up the incentive structure and makes it all about maximizing productivity then I wholeheartedly agree with you.

What economic system do you recommend instead? Marxist-Leninist communism? Anarchism? Feudalism? I guess none of these will be fitting for what's coming...
Very few people will pay mathematicians to produce as if they were making art. The big problem here is that counting proofs is how mathematicians historically have created legitimacy for the universities they work for. Mathematicians in the past 100 years have been paid solely for the prestige they generate, the work itself is irrelevant to the university. If the prestige signal dies universities will spend less on math departments. Every math professor will be expected to teach undergrads as well as a stable of grad students. That’s the value they are actually able to provide to universities.
> but simply dumping proofs on the math community and expecting others to do the grunt work of verifying, refining, and expanding on them is hardly a productive way to advance the field

This is a transitive period. In a few years, verification and exchange between model instances will happen faster than humans can follow. Human input will be an ethical question, and not a productivity one, because it will be the bottleneck in any science.

Muromec
And what use of that?
Etic is a factor of productivity, and the larger the contextual window is the weightier it becomes. That’s even integrated within the paperclip parabola.

If the focus in placed on maximizing some easily measurable output on a narrow perspective, situation is unlikely going to match a sweet spot of holistic equilibrium which is maximizing harmony and happiness through humanity as a whole.

How to make sure it doesn't evolve into some sort of Library of Babel of science?
That's the billion dollar question. Nvidia currently wants to deploy chips with the sole purpose of monitoring agentic workloads and that doesn't seem farfetched but they obviously have a financial motive to sell more products. It's like a cat and mouse game, like cybersecurity in general.
kuboble
I wonder if top labs will soon abandon math progress like they did go and chess.

In example of go where I'm more familiar Google deep mind poured large resources to get a super human performance first, establish superiority and abandon it. The community then built their own tools starting from reproducing their papers.

I think similar thing might happen to math. Nobody outside of math cares too much about Hamiltonian cycles in some bizarre graphs or proving lower bounds on complexity of some problem.

Once those results stop being worthy of mainstream media attention, they will abandon math and the progress will be done by mathemicians guiding the models and the community will likely establish some new rules about what makes a valuable contribution. Merely solving not yet solved problem might not be it anymore.

This misses the raw advantage of a good proof. It makes conceptualization simpler. In some ways math is like a hash list of of theorems. This list makes it simpler to prove other calculations, and will always be useful, to both humans and AI models. I can see two new directions 1 - the creation of specialist theorem models; that can answer questions efficiently about one topic and 2 - we probably need to incentivize and codify ownership of theorems; charging a proportion of the compute saved by using them. Ultimately enabling mathematicians to be paid our true market value!
edot
Oh boy, please not 2. What if this was a thing already and, since neither Newton nor Liebnitz had kids, we all had to pay some investors who bought the rights to calculus every time we took a derivative.
You know patents only last 20 years right?
cnr
As long as somebody with big $ decides: "let's make it 40!"
noworld
Aaaaasaand now Disney owns the rights to room temperature superconductivity.
lh712
Considering how essential math and science is for the prosperity of mankind (not even speaking about the cultural value) the question of how to reward people working and contributing in these fields effectively and appropriately is of extreme importance. (And I think the current decline in our societies is to no small degree caused also by our utter failure to address that issue.)

It is also fascinating, because I don't think there is any solution within our existing system, at least not any I know of. Theorem ownership is not a good solution (and neither are patents in general). Probably the most achievable (or rather the least unachievable) solution is a kind of communist utopia, where people can dedicate their time to a pursuit of any endeavor they see fit, as resources for a decent life are abundant and excessive power capture impossible. (The other option, somewhat dystopian, and which would not require humanity to change too much in its current mode of conduct, would be a totalitarian or caste-like capture of society by the scientific community.)

Incidentally, if AI proves as powerful as some expect it to become, it could bring about another solution of that issue by making all human science and mathematics obsolete, pushing its true market value to zero.

(With apologies for rambling.)

Almost every country in the world has a patent system for a reason...and they are pretty essential for big pharma to function. Granted there should be better calibration of duration - particularly in software, but rewarding first movers is wise in many market conditions. My point regarding "math as a hash list of theorems" is that Theorems will never lose their value, regardless of how powerful AI becomes; they represent compressed knowledge, and as such it will be more efficient for a more advanced AI to query known theorems than reconstruct each one from scratch - in that gap lies a marginal compute saving; an api service or tool call that someone might charge for. Even AI can benefit from specialists...
> It makes conceptualization simpler

I wonder if it makes conceptualization simpler for models too, given that they're trained already on human-speak. And I'm also curious as to whether humans currently have an innate advantage into simplifying and contextualizing proofs, or will the machines get good at that as well?

My point is that they will, and that is always going to be useful. Bundles of simplified knowledge (in whatever form) that make research simpler will be useful to share whether humans can understand them or not...
Yeah chess is a good example. DeepMind came for publicity with AlphaZero. Arranged a match with Stockfish with rigged rules to make AlphaZero look better than it really was (it was amazing but the match wasn't fair) and then just published some games and went home.

I was bitter about that back in the day as I hoped for more answers, more matches, more "truth" about chess being shown. Soon after that community project Leela Chess Zero was started and not only surpassed original AlphaZero but added few hundred ELO points over it. Then the combination of NN and classical engines happened with NNUE and current Stockfish is again a few hundred ELO points stronger.

Today we pretty much know the truth in chess for all practical purposes. Human analysts/preparation experts focus on finding interesting path and opponent profiling (what is the most unpleasant for the opponent to face). They don't look for truth anymore. The game is doing great, it's more popular than it ever was.

Why don't you mention the second match here, with its adjustments to meet Stockfish's quibbles - and the same result?

https://en.chessbase.com/post/the-full-alphazero-paper-is-pu...

Stockfish and other classical engines were never intended to be run in matches without opening books (of which there were plenty). Development assumed the presence of opening book and authors made 0 effort to make engines play well in openings because of it. This is also the reason classical Stockfish was a very small binary. A little effort to make it even by including even a very small opening book (like 50MB or something that would result in still smaller binary than NN engine with its net) would make it much more interesting.

The result was that Stockfish lost many games by walking into known bad lines and lost way more games than it otherwise would.

You aren't correct.

> We also played a match that started from the set of opening positions used in the 2016 TCEC world championship, along with a series of additional matches against the most recent development version of Stockfish, and a variant of Stockfish that uses a strong opening book. In all matches, AlphaZero won.

https://deepmind.google/blog/alphazero-shedding-new-light-on...

Ok I remember it vaguely but the match that got publicity and the one published results were derived from was 1000 games match from starting position. Deepmind claimed AlphaZero also won from TCEC positions and vs Stockfish with good opening book but at least back in the day I don't think I could find those games being published or specifics about books/positions they have used. Can you?

I am not claiming AlphaZero wasn't stronger. It wasn't as strong as the PR piece suggested though and we have never seen the games being published. In chess this is extraordinary because basically all games in chess are publicly available - both human and computer games. Claiming "we have created a strong engine that has beaten Stockfish with opening book" while not showing those games (or details about opening book used) is akin to "we solved this math conjecture" without showing any kind of proof or argument.

Publishing a few 1000 of games costs nothing. Tens/hundreds of thousands of games are published every day.

im3w1l
I disagree. Firstly, people in AI likely care about math on a personal level. Secondly math is useful. Playing go or chess is basically a party trick. Being useful gives it staying power.

But, I do think you are right that there will be some level of moving on. The spotlight is currently on maths and that won't last. It will move to some other area where there is more impact to be had. So while they might shift gears and put less focus on math, it will always be there as part of the portfolio.

I had the same idea recently. You've solved all the famous conjectures (all formulated by humans because humans found them interesting), what next? I doubt "AI formulated a math conjecture that nobody else cares about and immediately solved it" will produce that much hype. The actually interesting thing is indeed how mathematicians themselves will use these AI models going forward and how that will shape mathematics of the future.
> I wonder if top labs will soon abandon math progress like they did go and chess.

I definitely think that this is marketing, just "with good side effects". My doubt is when they will be able to move to "marketing with better side effects", that is, research with more concrete outcomes (health, materials etc.).

Problem is, that type of research is much harder. Some doubt that progress in such areas will be quick (https://www.noahpinion.blog/p/wheres-the-intelligence-explos...).

td6
Would that be a bad thing?

While top AI labs no longer focus on chess, the community build way better chess engines.

Stockfish is probably stronger, than everything the top labs build.

Wouldn't we expect the same thing for math? That slowly the broader math community would engineer a harness/program... That will surpass the current labs, and be a community ran project

Yeah its certainly true now that Stockfish is much stronger than alphazero, but it's probably also true that had Deepmind spent another few years working on alphazero it would be enormously stronger than either.

In the case of chess this seems fine, there isn't much value to society in creating an AI capable of beating top humans with a 4 pawn handicap rather than a 2 pawn one, but for maths where there are actual applications it is more complicated.

kuboble
But arguably - it's better for chess community that the best engines are opensource than having superior GoogleChessBot.
simonh
They aren't trying to 'solve' chess, go, or mathematical proofs as an end in themselves, but mainly in order to learn more about how to build better systems overall. The goal of AlphaZero was ultimately as a stepping stone towards AGI, and it's the same with LLMs.
The assumption being that 1. these things are all stepping stones, not diversions 2. that ai labs have a singular goal of producing agi
simonh
It's not an assumption, many of the people behind these projects explicitly say this what they are doing and why.
Xmd5a
I don't think it will be the case, maths have real utility. I found something interesting at the intersection of combinatorics and information geometry. To be quite frank I don't understand what I'm doing. And yet, when I ask ChatGPT to use the framework we're developing to write an algorithm, it turns out it has quasi-parity with the state of the art. I have to measure absolute perfs to decide which one is better – theirs, not mine. Ok. Time to keep improving on what I have. And this implies dropping the code and going back to the blackboard doing more super abstract math that are way out of my league.
lh712
Good luck!

We are at the point where the way in which humans do math and science changes significantly, and I have no good idea at all in what state is it going to settle down. But you are one of (many, I suppose) people exploring the new wilderness, so I wish you best.

Xirdus
If we get AI singularity, then humans will stop doing scientific progress altogether. But if we don't, then it's pretty predictable what's going to happen - things will be much the same as now, except everyone will be using AI for proofs, data analysis, theoretical models and designing experiments, so important discoveries will happen more often. It's also possible that after the AI craze dies down, we'll have enough computational capacity to solve protein folding.
kuboble
It has utility so some people will pursue it, but it has no immediate business value so I don't believe ai labs will keep spending millions on it.

Unless they decide that trying p!=np is worth any money.

Some math has almost incalculable business value, because math is the biggest driver of game-changer technology.

We'd be nowhere without Laplace and Fourier transforms, Maxwell's equations, elliptic curve cryptography, and many more.

Most math doesn't, but often these techniques are invented first and the applications come later.

And the criticism of the current round of proofs is that while they may be true - likely for some, questionable for others - they're not adding new techniques or insights.

xeonmc
The deluge of maybe-proofs have the same problem as the Library of Babel.
> often these techniques are invented first and the applications come later

There's a great paper from Abraham Flexner on this topic:

https://worrydream.com/refs/Flexner_1939_-_The_Usefulness_of...

It argues exactly that we should be allowed to pursue the seemingly "useless" knowledge.

Previously discussed on HN:

https://hn.algolia.com/?q=usefulness+of+useless+knowledge

IanCal
There is a risk of this particularly if it's seen as advertising - at some point "ai model solves hard to explain problem" isn't going to be news and that benefit goes.

However, there's some of this that's a proxy - the compute to solve these problems was very low (they claim a few hours of thinking time on a regular subscription). The large cost would have been the training and if training the models to be better at these things makes them smarter for useful tasks that's beneficial. I believe there was work done earlier on around showing that training the models on code made them better at broader reasoning tasks (not just writing the code itself).

Another side is that if one goal is to improve the models themselves, their ability to work on mathsy problems must be high. That has very direct business value, and ideological value depending on what you think the motivations of the people running the companies are.

dannyw
Why do you think this has no business value? It would be absolutely wasteful for OpenAI to not be doing this as part of a post-training RL rollout.

There are architectural advancements yes, but lots of progress from LLMs really come from (1) better pre-training [generally through more cleaned data, and ofc more data], and (2) lots and lots of post-training. It's how we get more and more intelligent models for the same param sizes.

The 'marketing' is just a useful side effect they get from their RL rollouts on maths and LEAN.

I think there’s a venue where they start focusing on introducing hypotheses where the model currently can’t solve it, or maybe this is already happening?

Being able to present useful novel ideas would likely generate a lot of press, for a while. I don’t know how this would look since I’m useless at math, but Im sure there are plenty of unknown problems with massive implications, that once formulated can be solved.

AlphaGo and AlphaZero weren't generalized models. Math capability will presumably keep improving along with the other general capabilities, even if there wasn't a special RL focus for math itself.
curt15
The amount of money they're lately ploughing into proving math theorems is inconsistent with how societies and markets have priced pure mathematics. The entire US federal budget for math research is something like $100M annually. A single college football coach can already earn 10 percent of that.

Pretty much the only enterprise that historically pays some mathematicians handsomely is quant finance, but those people are actually compensated not for proving theorems but rather for statistical modeling and programming skills. And even that industry is so technologically driven these days that pure research mathematicians no longer hold a clear edge over strong programmers with undergrad level probability and statistics at their fingertips.

As long as they continue making headlines they will continue spending. This is just marketing at this point.
dannyw
Math is one of the most verifiable domains, esp thanks to LEAN, which also build coding skills.

The $$$ they're pouring isn't just for marketing. Think of these papers/results more as "useful side effects" from large-scale RL rollouts and post-training. Every token being generated contributes to post-training in some way.

There isn't a hard boundary between "training" or "inference", modern post-training is arguably inference-bound :)

It was shown quite some time ago that training LLMs on programming tasks improves their logical reasoning skills also in other natural language domains. So I could see math also being a training gym for AI even if the final use case is not directly math-related. Having to solve math problems efficiently can build in skills that come handy in all kinds of more everyday tasks or science and engineering.
It’s also very useful signal that the reasoning trace is leading to solving open problems - you can be certain that you’re not landing somewhere inside the training data.
I think the entire thing here is that these solutions _are_ inherently interpolations of existing work in the field. That's the "super power" that LLMs have. To interpolate mass amounts of multi-dimensional data.
This depends on some handwavy use of the term "interpolation", not the mathematical definition. Mathematically interpolation usually means that the query point is in the convex hull of the data points, and that almost never happens in high-dimensional spaces.

See: Learning in High Dimension Always Amounts to Extrapolation Randall Balestriero, Jerome Pesenti, Yann LeCun https://arxiv.org/abs/2110.09485

I guess you mean by "interpolation" that it's some kind of nonlinear combination of the training data, but that's an almost vacuous statement. Any input-output relationship has to be so by definition.

Or perhaps you mean that interpolation is when the test input comes from the same distribution as the training input (though this is not technically the meaning of "interpolation"). But this is also quite difficult to pin down.

> There isn't a hard boundary between "training" or "inference", modern post-training is arguably inference-bound :)

Ah this is an enlightening point. 8 hadn't thought about it this way, but you're right.

ogogmad

  Don't hire a straight-A student, unless it's to take exams; or a professor, unless it's to write papers.
 -- Nassim Taleb
How interesting that Anthropic and OpenAI are full of professors and straight-A students!
It’s unclear to me what point you’re making here - can you elaborate?
Solving test questions well doesn’t necessarily translate into productive outcomes in the real world.
So education is useless?
No, more like the average can be more important than some outlier's impact, at least for steady progress. So while you do need the Mozart and Einstein, a more productive effort would be to raise the general population's education level.
casey2
Humans learn in the side effect of education, nobody is busting out some poor method that schools teach to solve antiquated problems in the real world.
If you track the best students futures and look the best workers pasts, there is not nearly as much overlap as society generally believes.
ogogmad
A lot of people have been surprised by the interest that AI companies have in solving maths problems without obvious applications, as well as in the financially precarious state of these companies. I imagine that they thought that these companies were helmed by mere bean-counters, like at Boeing. But no, I reckon that they're still academics at heart, and so they work on Navier-Stokes and L=BPL, instead of on how to maximise value for their (soon to be) shareholders.
Don't quote Nassim Taleb, unless it's to be an arrogant dick.

-- Me

-- Michael Scott
Eduard
imbecile!
They already proved the point
Interesting path forwards, and probably partly true, but there are some important distinctions:

Go was a specialized application. All the math results come as a side effect of reading the whole internet, and it will keep reading the whole internet. It will keep practicing thinking questions. Actually, math might be one of the best ways to keep them contemplating and measure their contemplation abilities, so math will always stay in the loop.

Also, math might not be useful just for humanity, but also for AI, so the system might actively benefit from new math results itself. (Not sure if any of the recent proofs qualify, but future work might.)

I think this is an interesting and good theory. They've probably eked out the large majority of the PR benefit at this point, so whether they continue in this vein will tell us a lot about their motivations for this work.

To take this to the next step, what happened after deep mind pretty much solved Go is that they started looking for the next set of things that hadn't been done yet. It does strike me as very likely that this will follow that same path.

pfdietz
I'm told there are another two large tranches of results to be dumped.
Yeah I'll be curious to see if or when they stop seeing the utility in this!
pfdietz
If working on these is causing improvement in general reasoning (beyond math) then I expect they'll keep doing it.
Yes, I think this is the question.
alterom
>Merely solving not yet solved problem might not be it anymore.

Never has been.

See: Mochizuki's "proof" of the ABC conjecture [1].

In short, mathematics is an art of storytelling as much as anything else.

Writing a boring story, much less an incomprehensible one is worth exactly nothing, and doesn't matter.

Is it cool that the machine can connect the dots that people put out there in a form that happens to be easily ingestible?

Yeah.

Can the machine answer why give a fuck about a result like that?

Nope. That work has already been put in by people.

The most obvious impact of this BS is that interesting problems with potential to be solved, which were already a scarce resource, will no longer be shared.

It's already become a huge issue without AI being involved, due to the rat race nature of academia.

There's even an industry term for when someone gets a wind of another person's progress on such a problem, and solves it before them - it's called "getting scooped"[2].

It doesn't mean that they got access to the results. It can be as simple as e.g. an experienced mathematician hearing about a problem that a PhD student is working on for their thesis, and scooping it.

Honestly, as a mathematician-turned-software-engineer, I see and hear both sides, and both sides are lying:

* The AI companies are lying that what the LLMs produce is "mathematics". Writing a proof, even a correct proof, is not what mathematics is about. The Fundamental Theorem of Algebra has been "proven" incorrectly four times IIRC, but it was mathematics nonetheless.

Mochizuki's "proof", on the other hand, may be correct, but over a decade later, it's still not really mathematics, because transfer of insight is a fundamental component of it, and insight is a thing that the AI, presently at least, doesn't do, no matter how many proofs it produces.

Proof automation has been a field for a long time, for that matter; hyperscaling it doesn't change the nature of things.

The lie is also that a correct proof is the final point. The Four-Color Theorem[3] was proven in 1976, yet people were still proving it 20 years later, trying to reduce the number of cases to leave to the machine, even as we got it super-duper-verified by theorem-proving software in 2005[4].

The writing was on the wall since computers were invented, so why are mathematicians making Pikachu faces now?

That's where the other lie comes in.

* Mathematicians are lying to everyone else by suggesting that the technology is the problem, or that the issue is with how the technology impacts people, and not with the COMPLETELY, ABSURDLY FUCKED UP ORGANIZATIONAL STRUCTURE OF ACADEMIA, and I can't scream it in bold enough caps lock.

There's a reason websites like "100 reasons to NOT go to grad school"[5] exist, and all who've gone through the masochistic 5+ year journey can attest that actually, the reality is worse.

Closer to the point, there's a reason Grothendieck burned most of his work and fucked off from mathematics in the 90s, and Perelman did the same, in effect, in 2000s[6], while putting a million dollar weight to his statement that MATH SOCIETY IS FUCKED.

To all the mathematicians running in circles screaming AaaaAaaAAAaa, Perelman can now say: told you so.

See, after we discussed what math is about (insight transfer), let's do a peculiar little observation that this is exactly what's rewarded THE LEAST in the field.

"Teaching jobs" are a punishment for those who don't do good enough, and doing good means cranking out proofs in a publish-or-perish race, emitting just enough for others to verify correctness, but not too much lest you don't have enough for the next year's publication quota (or, if you are a mastodon, enough for your grad students/postdocs to work on).

The industry where the incentive structure creates and rewards scooping as a concept turned out to be vulnerable to scooping at scale.

O the humanity!

'No Way to Prevent This,' Says Only Industry Where This Regularly Happens.

So the lie to everyone else is that problem isn't 100% a social one in the mathematics community. Not even mathematics.

The way we do math doesn't have to change, and it won't. We've had theorem provers for decades, remember?

What will, I hope, change - what has to change, as Perelman cried out loud (and paid $1M for people to hear him - without listening, it appears) -

- is the way mathematicians treat each other.

That's to say, what gets rewarded, what gets penalized; the incentive structure, and so on.

In Google-speak: the real problem here is that the perf is fucked up, and what counted as impact before can now be automated, threatening the entire hierarchy of resource allocation (in particular, headcount).

This is not only NOT a tech problem (which won't be solved by changing the behavior of AI companies), it's not a problem with the way we do math or adapting to the progress either.

It's the ROTTEN INCENTIVE STRUCTURE OF ACADEMIA that is showing its cracks, and mathematics is where it's apparent the most at this moment because, as Vladimir Arnold told us, mathematics is a branch of physics in which experiments are cheap[7] - and more pertinently, could be performed by LLMs without human involvement.

Unlike the biotech lab where someone still has to wash the beakers, mathematics doesn't have that moat. Hence the conundrum.

Will the mathematical community adapt and fix its publish-or-perish problem, or will it simply become a cult, like it used to be[8]?

Given that a colleague of mine wrote that satire years before LLM proofs were on anyone's radar, I'm not having much hope.

Stay tuned.

[1] https://www.quantamagazine.org/titans-of-mathematics-clash-o...

[2] https://quantixed.org/2018/03/19/scoop-some-practical-advice...

[3] https://en.wikipedia.org/wiki/Four_color_theorem

[4] https://en.wikipedia.org/wiki/Proof_assistant

[5] https://100rsns.blogspot.com

[6] https://www.nbcnews.com/id/wbna38039068

[7] https://archive-dsweb.siam.org/The-Magazine/All-Issues/vi-ar...

[8] https://www.mcsweeneys.net/articles/an-open-letter-to-the-ma...

> Too much effort is being invested in proving that the exponential curve is still holding.

Given sustained exponential growth is mathematically impossible to maintain with finite resources, it's funny to me they're using advanced mathematics to try and achieve this.

myzek
Generate and dump on others to verify is how the generative-AI people operate. Be it in maths or just your regular job.

The amount of Confluence pages of "research" that is just a dump of LLM output someone passed to me to review is staggering

I hate this approach, it's unbelievably selfish

This Math 1.0 followed by Math 2.0 framing is wrong. Mathematics did not start a few decades ago when problem solving became the norm. And it will not end now when problem solving turns out to be "easy".

Mathematics will revert to it's main practice, which is to study.

There are lots of weird panic reactions by some prominent problem solvers. See for example the ridiculous cease and desist like statement of AHM shared at Tao blog.

Put this Math 1-2.0 with that AHM statement together and you'll realize that this is a power struggle and that you see only one side of it.

Mathematicians have very diverse opinions about this. I, for example, am for as much as possible automatic harvesting of all these "low hanging" fruits. Should be disclosed as soon as possible, free of any bottleneck, and citable. The mathematics community may do whatever its various members desire to do with these results. Let them decide individually what to do with them. This AI tool is here to stay.

This looks indeed like a public relations campaign, it would be helpful to the conversation to contribute your opinion, in a polite way.
> the grunt work of verifying, refining, and expanding on them

What work do you think mathematicians do normally?

Like they sit whole day and have ideas? And where are the ideas?

The way I see it, _some_ mathematicians enjoy solving puzzles, and now AI is better at solving puzzles.

This does not affect people building new theories.

Also, it's quite prestigious to write a _book_ on some topic. And guess what writing a book entails? Refining and expanding. What you call grunt work.

pks016
My grunt work is different than doing grunt work for a company to fix their problems so that they can make more money.
OpenAI does not make money from math papers LOL they just release them for the sake of community. (Because sitting on those results would be considered worse.)
pks016
You don't seem to understand how the company evaluation and perception works. I would suggest reading on that.
Solving a well-known open problem like N-S "millennium" problem has some marketing value, as this can be referenced in a press release, etc.

OTOH a marginal value of a single paper in a paper dump is basically zero. Nobody gives a flying fuck if they made 100 or 200 papers. It won't give them even a single customer. If they released, say, 5 paper per month it would generate a buzz but not as much backlash from mathematicians.

> This does not affect people building new theories.

Actually, the problem is, somehow the skill of building new theories in math is directly tied to slaving hard over a problem. It's the very experience of slaving away that actually somehow causes ideas to form. Pretty much all mathematicians understand this. Yes, senior mathematicians now can form some new theories, but what about junior ones who will have very little experience in working hard on a problem by hand?

Of course, they could work on the problem by hand anyway, but they won't because no one will pay them when a machine can do it.

> no one will pay them when a machine can do it.

Let's start with the fact that mathematicians aren't paid for results. Institution which gives them salary fundamentally doesn't give a fuck about theorems. They might care about having top-grade mathematicians for prestige, or because they believe that countries which are "good at math" also do better in science, engineering, etc.

> somehow the skill of building new theories in math is directly tied to slaving hard over a problem

We don't know if that's the only way. Perhaps collaboration with AI is just as good. Why reject it before it has been tried?

> what about junior ones

Junior guy with brilliant new ideas might benefit the most from AI as it can compensate for lacking technical chops and breadth of knowledge.

> There seems to be more interest in hitting some arbitrary benchmark (we proved X unsolved problems) than in genuinely contributing to mathematics

I feel the same way about academia, the papers, the citations, the ego, the narcissism and the taxpayer codependency that got cut off and turns out wasn’t necessary at all thanks to a private sector entity running laps around them

I don’t feel that academics need to pursue the discipline and distributed brain-wracking that has sometimes resulted in the solved math problems, just because more times they find other nooks and crannies to explore along the way. I think the blueprint is enough. Standing on the shoulders of giants is good enough.

and if the concern is that they can’t figure out what to do with a proof, next year’s AI will

As with any field, convincing people to care about your ideas and your approach is half the battle

Many of the best startup ideas by the best product and engineering minds failed to gain attention and funding. Same with much of the best music - relegated to hard drives with derivative ideas only resurfaced decades later

I would expect much of the recent math dump will be leveraged by other LLM-driven research teams rather than read in depth by a human

aswegs8
Once LLMs pass the threshold to being to invent new general-purpose methods and frameworks, the frontier is irreversibly lost to AI and it becomes simply a hobby that mathematicians pursue. They work through, digest, and maybe write up the proofs for understanding. But the real meat will be growing the LLMs. Who knew software eats the world was so true?
carra
> simply dumping proofs on the math community and expecting others to do the grunt work of verifying, refining, and expanding on them is hardly a productive way to advance the field.

I tend to agree with this, but what is the alternative? Should OpenAI and Anthropic employ hundreds of mathematicians to do this work? Should they just not solve math problems within their reach?

I tend to disagree with OP, for the same reason.

It's unclear what more could be expected than releasing the presumably already verified results and write-ups for each problem. Should they run a mathematics school too?

Then the comment goes on to argue AI labs were not interested in actually advancing mathematics, and that investments into AI were manically excessive.

IMO none of this follows and demand is there to justify the investments.

The comment then goes further to argue that AI labs were putting too much effort into pretending there was exponential progress rather than actually making progress.

The factual basis for this claim seems to be that OpenAI released math results and write-ups, and it's not even clear what more they could do on that topic.

That's a very negative opinion.

pred_
Look at how actual researchers are currently using the tools: they'll generally use the LLMs to create slop papers, sometimes supported by auto-formalizations, just like OpenAI does. Then they will go through the lengthy process of digesting the results, turning the often incomprehensible and poorly organised outputs into something that humans can understand and build upon, they will then give seminars on the results, further helping with dissemination. Doing so still requires expertise, and probably will for a good while.

So yes, that is exactly what they should do. Alternatively, if they are too lazy or incompetent to put in the effort themselves, do what AGMAI proposed and fund a third party to help out.

carra
I don't think the slop analogy is that valid here. Sloppy pictures may be of little value, but a "sloppy" proof of a conjecture does still prove it.
They should not solve the math problems. It’s simply a net negative for the world.
ew-dev
> simply dumping proofs on the math community and expecting others to > do the grunt work of verifying, refining, and expanding

Hm, kinda reminds me of my college days. "Proof trivial, left as home work." was a sentence my Profs loved to say.

I'm sorry I disagree entirely.

The more information the better.

The entire purpose of published work is to remove noise (and perhaps incentivize work through attributing credit).

This information is now out there. You can choose to ignore it if you wish. You may just find yourself a century behind in research.

And on that point most of this research has been looked at by their mathematics panel and comes with lean certificates, it's not exactly noise.

This to me is more the old guard not willing to let go or change their ways.

I think it's quite clear the mathematics panel didn't read most of the papers with the scrutiny it would take to publish it, if only because the lean certificates and the informal proofs are not 100% the same.

And if it's hard to understand (which seems to be the most common reaction) it's not exactly devoid of noise either

There was an opportunity for people to work with the AI to produce a proof, now it almost feels they're working against it.

I never said it was journal worthy. Just its not all noise. As I say academics are welcome to ignore it, it isn't published in any journals after all.
td6
I think in a ideal world your right.

I think a problem is that math seems like a deeply toxic, ego driven domain.

I think he argued that e.g because the navier stokes millennium problem ist considered solved now, you won't get any recognition for being the first human to solve.(How would you even proof you solved it yourself and not just regurgitated the ai proof?)

And since recognition is the main objective, noone would spend time on dissecting the proof, and perhaps finding some unique approach to solving the problem, that could be transferred to other open issues.

And therefore the problem is now "poisoned". Since it's assumed to be solved noone will research it, and the potential revelations won't be found

To be fair I'd hope during peer review this would be apparent.

During your write up, I'd imagine you would check it's not already out there too. And once complete it's cheap and easy to run it through an LLM and ask is this covered by anything else out there. If its novel and not published it doesn't matter what others say.

Research is already messy as it stands. Something new can already be dismissed by incumbents as "not novel enough" especially in niche fields where they're likely to be the ones conducting peer review.

The biggest problem is LLM tends to produce over engineered, very complicated proofs that are an eyesore even for relatively simple problems. Give it a beautiful Olympiad geometry problem and LLM will tear it apart into ugly algebraic calculations, turns all lines and circles into equations and calculate their intersection points that spans multiple pages because it is a guaranteed way to solve it. Correct, but hardly any use to the user.
You know it's valid. You're not working on incorrect assumptions. Surely there's value in that?
I don’t know if it is valid. It is unverifiable. I still found some basic algebraic mistakes in top models as late as 3-4 months ago, not sure about it now. But that’s not what I want anyway, so I often put “Do not brute force” in my prompts.
Sorry I mean in very public releases such as this trove, where many have lean certificates attached and publicly scrutiny.
shakna
Three of them have already been withdrawn. So I would say we have proof, that they cannot be implicitly trusted.
So it turns out there is a, perhaps informal, working system in place and we can deal with it.

Peer reviewed and published insights are proven invalid all the time. This is the nature of research and how we learn.

shakna
So instead of us "knowing it is valid", we don't. We need to put in extra effort, because someone felt like doing only half the work and dumping it on the community to fix.

We don't know the current system can work well enough at this scale, because that's un-knowable. We know it can find some of the problems. We don't know it can find all of them.

We do know it takes more effort - that's knowable. Increased data takes increased processing.

Whether the community has the required effort available, seems unlikely, considering the expertise required to be able to assess these things hasn't changed. Only the ability to generate them has increased.

You don't have to go through it. You can ignore it if you wish. There is no obligation on you to check it.

I feel your concern stems from the risk that there is additional noise everyone needs to cut through.

In reality this isn't any tom, dick or Harry giving you their vibe code output. They have spent millions of dollars on this output, so there is a filter. The biggest filter of them all, funding.

Furthermore, LLM's have given us another gift semantic search, we can easily check your work against theirs, this is valuable insight so instead of researchers wasting decades and fortunes pursuing an avenue that shows no value (this includes methods), they can purse new avenues they know what to avoid, in the same breath they know what to work towards.

shakna
That's... Worse. Millions thrown at a wall, to further public rhetoric and control an industry, is an attack on the industry. It isn't philanthropy.
That's capitalism. The luddites weren't complaining because there was a new technology making their jobs less interesting or due to the "loss of a craft", they were doing so because it was taking over an industry and taking their jobs.
How something is proven is often more important than what is being proven. There are underlying systems and patterns that, when understood properly, improve our model of mathematical (or physical) reality.

With convoluted and inelegant proofs, AI may fail to uncover those systems and patterns. As a most concrete example, it may fail to recognize some problems as isomorphic to other problems. Brute force solutions are a depth-first search.

To improve human mathematical understanding, AI is probably best used as a “copilot” (lol) rather than a black box oracle, like these AI companies appear to be doing.

I'm sorry this framing is just moving goal posts.

If you're after new methods. Then new methods is the goal, the answer to the question is not the goal then. The animated response indicates the answer wasn't just a byproduct.

There is still something to glean from the answer. You have a further constraint. Otherwise whatever "new method" proposed may as well be hallucination, potentially taking you in the wrong direction away from the answer.

This line of thought is not unique, stonemasons made obsolete by uniform brickword suddenly were "worried about the art and preserving traditions".

What’s the value?
You're asking me what the value is in knowing what the answer to the question is?

I'm being serious.

Yes.

Tao points out that without the mechanism for disseminating the techniques used to construct the proof, the proof by itself is fairly useless. It’s not like the existence of the proof directly increases GDP or human health or flourishing in any direct way.

So what’s the value?

Then Tao for one should have no issues with this release or any other release, if it's all actually about how he got there.
I’m not deep in to math but the op tweets make sense to me. In that it’s not just the final proof that mattered, but the mind and understanding of the person who arrived at the answer. An LLM dumping the answer can’t elaborate on it, can’t tell the story of how they got there, etc. But it also deprives someone else of that achievement and learning.
Ravus
You can see it in the way we structure college courses: engineering curricula often cover in one semester what mathematicians study over one or two years.

This is because have fundamentally different goals: being able to use results in calculation versus having a deeper understanding of the subject matter.

Less Liszt/Paganini, more chamber music and teaching. I like it!
My steak is too juicy, my lobster is too buttery, my industry shaking mathematical proofs are coming too quickly

In what world is OpenAI not “genuinely contributing to mathematics”?

I’m getting whiplash from the speed at which people are suddenly accusing them, and AI in general, of not doing enough.

In TFA I read (scroll up from the link anchor) this was explained: OpenAI are not giving talks - because they can't answer any questions about the model's work. There is little follow-up activity - the actual elaboration of human understanding of the new ideas is stifled since the problem is solved.

But you are right, this is not OpenAI's "fault". The problem is - as others have said recently - that many people in mathematics want recognition for solving open questions more than they want the answers to the open questions. Everything about the economics and social environment of Mathematics will have to change.

I think that this is exactly the same split we see in software: there are those who mainly enjoy the craft aspect of building software, and are uninterested in the product or business they are supporting. Others are primarily interested in the production of useful software or building a platform or company.

I've always been in both camps myself. When it became obvious that AI was going to destroy the craft aspect - at least two years before it actually could do so - I became very discouraged, even depressed. But once it was actually good at building software, I became very excited about all the stuff I could now build. Sadly, I think a lot of people in our field have never had something they really wanted to build.

gessha
What does it mean to “genuinely” contribute to mathematics? Because the definition you use is load-bearing ;) and it might be different from that of others.
pfdietz
It's when theorems are proved by True Scotsmen.
Perhaps mathematics was never about proving things? I know it sounds like moving the goal post, and it certainly was what motivated mathematicians on a day-to-day basis, but bear with me for a moment. I think that beyond being a creative activity that humans enjoy, math was about building new tools and systems of thought. Axiomizing things we take for granted, logic, linear algebra and calculus (on which modern LLMs rely so heavily) are such examples. People chose to participate in this field because they found it enjoyable and satisfying in some way. As a side effect, society reaped the benefits every couple of centuries. Will harvesting open problems with AI ever give us these things, or will we just be left laundry pile of Lean formalizations?
Because they published a ton of slop papers with terrible English, impossible for humans (even experts) to understand, full of non-standard terms, invented jargon etc.

So effectively, stuff got proved, but people don't really understand how, so it's mostly fucking useless and done for OpenAI's marketing team, while also pissing off the maths world at large.

So if god himself comes down from the heavens and hands you the answers to major unsolved problems, but doesn’t walk you through them and you’ll have to still put in effort to understand them, then that’s “not contributing”?
Give a man a fish, teach him to fish, etc.
I'm in the same boat, I guess I had this naive idea that unsolved math problems would mean something if they were solved. But it seems like a lot of them at least were more thought experiments than anything else.
Kids on HN have built their whole personalities around math, and now, they are discovering that "problem solving ability" has nothing to do with solving problems.
Mathematics is less about the end product and having a healthy community of people to actually understand the proofs and eventually apply them.

That community won't exist as many people simply won't even enter the field because it's reprehensible and contemptible, not to mention boring, just to read machine-generated proofs and verify them.

Collaborating on, or at least working on unsolved problems is what motivates most people.

AI is like a cheat code in a video game. You get to the end faster but fewer people want to play if the cheat code is always on. You can't turn it off either because the very challenge is to do something unique.

pfdietz
> healthy community

This is code for "they keep paying us".

Its pure cope from the math community. One year ago if god came down and asked any of these math guys if they would want to see the proof of the Riemann Hypothesis they would all say yes. Now that the AI math god is here and handing out proofs they are all like 'not like that'!
"community building"

For what?

If math is just about having a community of other mathematicians to hang out with, it still isn't a career. Nobody is paying money you need in order to to eat, just to hang out in a community.

Just like a software engineer, "Well AI can write all my projects now, but I have my local Rust Users Group to hang out with". Nobody is paying me to hang out and hand code Rust.

rtkwe
The other gap in AI is it doesn't explain or lay out how it got to the final proof which is often more fruitful for new techniques etc than the final proof by itself.
Honestly, that's a weakness shared by a lot of human-generated science as well.
rtkwe
In mathematics it's usually a large portion of the paper because they know it's the interesting and useful part of the discovery. Often they come up with new tools and techniques and those often unlock further discoveries.
qarl
> but simply dumping proofs on the math community

So you'd prefer if they kept their work secret? Or you don't want them working on these problems at all? Or they should be required to do the work the way you want them to?

I'm not clear what you see as a better option than dumping.

AI is fucking shit up for humans without providing any benefits.

It’s causing unemployment but no real gains unless you have a lot of equity and can benefit from lower costs from fewer people working.

Saying AI has no benefit is pure ostrich head in the sand nonsense.
I’m not saying AI is not capable.

I’m saying with our current economic and political systems it’s only a lever for making the world worse faster.

qarl
> I’m not saying AI is not capable.

Well... except you did:

> without providing any benefits.

I will concede I should have said “net benefits”.
qarl
HA! It just solved a ton a problems for us - but that makes you mad?

But yes, I agree, there needs to be a solution for the fact that we don't need to work anymore. And it needs to be more like The Jetsons than 1984.

We'll probably have to revolt to get the changes we need. I don't see Elon giving up his power without a fight.

>expecting others to do the grunt work of verifying, refining, and expanding on them is hardly a productive way to advance the field.

I don't know that the labs expect this at all.

Is the purpose of math to move human knowledge forward or to appease the mathematicians who get their rocks off working on proofs?

lenkite
> but simply dumping proofs on the math community and expecting others to do the grunt work of verifying, refining, and expanding on them is hardly a productive way to advance the field.

I assume this will increase the job market for mathematicians though right ? Tens of thousands of dumped proofs will mean math majors will need to be hired for review.

casey2
But they are doing the work of verifying, refining and expanding. The humans are irrelevant in this situation. OpenAI is just letting them know that they are now obsolete.
j-pb

  The authors of the proof are invited to give many talks, and meet with other experts in the area.  Workshops are set up to discuss the proof, as well as other recent developments.

  problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is "solved", and do not understand the AI output well enough to answer questions on the result
The value here seems to be the insights that the author of the proof gained, and the paths they took and maybe more importantly didn't take. Inviting only the human prompter to a talk on the paper is like inviting only the department chair, manager of the actual author.

The valuable part that Tao is feeling the absence of is the insight, and you can only get that from talking to the swarm of agents that developed the original proof with all of their context.

So to me it feels like we don't need Math 2.0, but Authorship 2.0. I want to "meet" the context that generated these proofs. I mean luckily these were not generated by faceless systems like a SAT solver, you can actually talk to it, but I'm not sure if we can step beyond our pride and grant the true authors of these proofs that recognition.

pegasus
These systems don't have a genuine capacity of introspection, beyond just mining the conversation trace. When you're asking them why they did this or that, they are basically guessing anew from the outside, and are just as likely to hallucinate as they are to hit the right answer. These are mechanical systems which brute-force chains of various (re)combinations of techniques acquired from the training data. The true authors are all those who have contributed those techniques in the past.
dannyw
I think that's debatable. Anthropic's interpretability research suggest models do have self-introspection ability, at least in the "J-Space": https://transformer-circuits.pub/2026/workspace/ ; and this private working/'introspection' space is distinct and distinguishable from the tokens they output (CoT tokens are output too).
Bold claim. What data is informing it?
a3w
This was how it was done in chemistry back in the day: release a cooking instruction. If it fails, visit the colleagues and give hints on what you meant.

The idea would be that you should not fiddle with the minds who try to independently evaluate your works, so that is not a feasible approach to truth seeking.

While in organic chemistry, this way, valid progress was made, you can always avoid a perpetuum mobile inventor and get conned.

> problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is "solved"

This is pretty much what a person that proivded patronage to a matematician used to be. API prompters are people who provide patronage for AI mathematicians.

You don't talk with them about the discoveries. About discoveries you should talk with who actually made them. Namely the LLMs.

Another analogy might by that you shouldn't expect to have interesting discussion about the essence of art with art producer.

Just thank them for the inference they covered and interact with the results instead.

cubefox
The problem is that the reasoning traces are kept secret. OpenAI doesn't publish them because it would lead to distillation attacks from other AI companies.
drichel
The reasoning traces were traditionally keept secret by legacy mathematicians as well.

"When the architect completes a fine building, he removes the scaffolding." - Carl Friedrich Gauss

cubefox
Legacy mathematicians can nontheless remember much of how they came up with their proof. Even though they (as Gauss says) usually don't publish this, they can still answer questions about it at conferences and workshops, or otherwise use the knowledge of the creation process to explain their proof. It's not a secret in the sense of OpenAI.
I don't believe legacy mathematicians kept their reasoning secret after they published. Quite the opposite—they were giving talks and lectures, responding to criticisms, and generally convincing the community that they were right. OpenAI has done none of that.
casey2
Nobody understand the talks at these conferences unless they are really dumbed-down (which they should be). The value is talking to other mathematicians not the speaker.
The same could be said of Software. Instead of giving up on creating novel projects and instead just taking other peoples ideas and porting them to Rust, we could be embracing AI to push software and computers farther.

Im not sure how that will work, but im convinced the current paradigm of just pushing agents into codebases for not much reason other than you can is going to make building software incredibly boring and push creative people away from the field and stagnate progress.

My prediction is software gets boring and building hardware projects will be the new frontier for creative engineers looking to push computing further. Which is probably a good thing.

> same could be said of Software

Sort of. An elegant proof is useful beyond what it shows. It hints at new mathematics, and can prompt discovery in applied fields. I don’t think I’ve heard of elegant code leading to discovery on its own.

you must not be familiar with haskell
devmor
> I don’t think I’ve heard of elegant code leading to discovery on its own.

I’ve not heard of it either, but code is an abstraction of math, so I don’t see why this couldn’t theoretically happen.

Anecdotally, I’ve started spending time advancing my math skills beyond the early college level I stopped at and I’ve frequently found I already know concepts of more advanced math - I just didn’t know what they were called or how to apply them to an equation on paper, but I’ve been using them for years and intrinsically grasped the underlying academics.

rcpt
> elegant code

Usually it's the opposite. "That's in prod? And it works? It shouldn't work and I thought it was doing something else. Why does it work?"

Jaxan
But it has. Think of design patterns and other programming paradigms (logic programming, functional programming, etc).
bartnp
I think this happens all the time actually, but it's often smaller scale. Like someone writes a neat architecture to do X at their company, and later someone else walks in, looks at this code and sees it's now super easy to do Y, which then turns out to have immense user/business value.
> I don’t think I’ve heard of elegant code leading to discovery on its own.

Every software design pattern came from elegant code. People wrote code, summarize code, learnt from code, and taught code. That's discovery

Maxion
Models have limits too, I don't think software will become boring. I think it'll become more interesting, in the not so distant future one software engineer will be able to do so much more than today.
est
I was about to comment the same.

People wrote many books about software engineering, all from valuable experience from buildng expensive software systems. But in the age of AI, is there still anything learnable from generated code?

Personally I always ask AI to summarize its findings and lessons in a .md file. And I always learn something from it.

But could AI utilize some new patterns and paradigms I wasn't aware of? Very likely. Because we only learn from our personal grave mistakes, a summary from others gets neglected and forgotten

I do agree with you, although I think that there may be still some advancement available in software in terms of programming languages or novel architecture design. But frontier is much more around hardware, just thinking about computing power, electrical transmission, material constraints on power, connectivity, insulation. I would definitely advise kids to study physics and materials engineering than CS.
hypfer
I mean if we're being honest, most of CS studying (as practiced) was a waste of time anyway.

The S fell short in actual reality for the most part, as it was merely a hiring requirement. A hiring requirement that didn't even make sense, because the skillset of academic CS only marginally overlaps with the skillset one wants to hire for.

Material engineering at least for the most part has actual real-world applications where one can push humanity further. CS (as practiced, not necessarily the idea of real CS but the CS we got due to it being used as a hiring filter) for the most part is just self-referential spinning with mostly unclear results.

There is real impressive work being done in that field, of course, but I'd argue that the majority of it over the last decade or so at least was just performative nonsense.

Maybe by again allocating new resources to other fields, what hides under the label CS can become more pure actual CS again. I think that would also be a much less miserable experience for everyone involved.

What new things are even left to discover in software? You can still use C for all the backend stuff and I don’t think all the endless stream of front end frameworks and libraries are all that innovative. I think the only tangible innovations we’ve made over the last few decades have been in infrastructure management. All heavy lifting is done by mathematics anyways.
> What new things are even left to discover in software?

The will to implement the stuff we learned, instead of letting dark patterns and churn for the sake of churn get the better of that :P

Well i mean, the whole idea of discovery is that you're not sure what you will discover until you do
Originally the word computer described a person doing the act of computation.

I believe in the future, we're going to see a similar shift in "programmer" - instead of a human programming the computer, you'll give the ai an idea and it will spit out a program.

And just like how automating the act of computation revolutionized what we could compute, automating the act of writing code will change the act of programming - hopefully, as you described, allowing us to do things that simply were not practical in the past.

flir
But these days hardware is software. One microcontroller replaces so much electronics. Software is eating the world, as the man said.
>the mere knowledge that a solution exists "contaminates" efforts by both humans and AI to find alternate routes to the problem that reveal additional insights

This really expresses the heartburn you see across all fields, not exclusive to careerism. I certainly have friends in decomp and fan translation spaces that have been demotivated by the current rash of efforts happening there.

The rush to be "first" has always been over-celebrated, but it would be nice to believe there's a way to get beyond that thinking.

This has already been the case in AI/ML and computer vision papers via flag planting papers. Have an idea, super quickly publish a hasty work based on it that doest actually work, methodological and eval issues, engineering terrible, slow, bad results etc. But it was the first so now your concurrent work that was much better evaluated, better implemented, etc is suddenly worthless and unpublishable.
> decomp

This I don't understand, seems like an obvious thing to automate, especially for byte-matching?

pinkwah
Byte-matching has been a side-effect of gaining an understanding of how the software works. With this you are skipping the understanding and you will not get people like Kaze Emanuar on YouTube who have dedicated a lot of time to building upon this understanding to create something better.

I don't think it's motivating to solve a black box by having AI generate another black box if what you want is to understand how the thing worked.

diath
> Byte-matching has been a side-effect of gaining an understanding of how the software works.

No? Decomp sources are often full of comments that claim that there's no clear reason as to why something is done a certain way or straight up full of question marks. Byte matching often is a result of bruteforcing a solution rather than understanding the original idea behind the code.

dminik
If your goal was to do something for the games you love and understand them on a deeper level, then just handing that off to an LLM doesn't really give you the same sense of accomplishment.
The future of "math" is not just solving, but communicating what was solved, and why that solving is important. "Math" as a discipline is only ever been half created, they abandoned explaining themselves like it was beneath them. Well, now in "Math 1.0 finally whole" they will be explaining what they did to the rest of us. If you can't explain you are not really there.
curt15
That's what math has always been about, even if not everyone is necessarily as talented a communicator as someone like Thurston -- who was himself an exponent of computational tools in math research:

>I think of mathematics as having a large component of psychology, because of its strong dependence on human minds. Dehumanized mathematics would be more like computer code, which is very different. Mathematical ideas, even simple ideas, are often hard to transplant from mind to mind....Translation in the direction conceptual -> concrete and symbolic is much easier than translation in the reverse direction, and symbolic forms often replaces the conceptual forms of understanding....

https://mathoverflow.net/questions/43690/whats-a-mathematici...

mxkopy
This is an entitled take. The usefulness of a thing has nothing to do with how easily communicable it is. That is the whole premise of specialization and becoming an expert. People go in-depth on things so they can say “trust me on this” and so you don’t have to. You are advocating for a tyranny of the illiterate.

Also, this reads like you didn’t like math classes. That sucks, but it’s no basis for societal organization

bsenftner OC
I am advocating for effective communications, it is not as if knowledge once understood remains difficult. To communicate and convey understanding is good, and to withhold understanding is tyranny.
essai57
> it is not as if knowledge once understood remains difficult

What do you mean by this? Once a person has understood something, it may no longer be difficult for them, but it can certainly still be difficult for those who have not understood it yet. And we humans know of no way of transmitting understanding to another person without having that person exert some effort.

I don't know where you get the impression that mathematicians at large are withholding understanding. In fact, many mathematicians share their lecture notes freely online, share their articles freely on arxiv etc.

If your complaint is that these texts are written in the language of the field and thereby not accessible for laymen: This is the case for any advanced knowledge, because it builds on more basic knowledge that a person must first acquire. That is not withholding knowledge.

It is not "they are withholding information", it is more an issue of, and this is not just formal mathematics, not enough emphasis on communicating to be understood, and the education that we all receive not really including how to communicate our specializations to anyone but a peer or a boss.

These formal educations we receive are half baked. We cannot use them without other specialists, specialists we specialists ourselves probably cannot afford. We cannot discuss what we do with anyone but other specialists of our same color and stripe, the same star shape on our bellies. We're not being taught how to communicate, not really, not in general, not in a manner that enables us to function without some corporate apparatus extracting the maximum while always offering statistically less than the market average.

What I meant by knowledge not remaining difficult once understood is that there is a collective consciousness hurdle that we as a society can move on. Once how to express difficult ideas and controversial question and answer exchanges become better understood in general, and once how to handle situations that currently cannot even be discussed due to the emotions they stir become less emotional storms and become logical frameworks people can discuss in abstract, then we become a more functional society than we are now.

mswphd
communicating results clearly is just as important for specialists/experts. A reaction by many mathematicians to the OpenAI dump (and many LLM breakthroughs) is simultaneously

1. if this is true, it is interesting, and

2. the exposition of this is a huge pile of slop.

This requires mathematicians to have to clean up the slop for it to be useful. It has happened for most LLM-generated proofs in the last few months. Anthropic explicitly payed two top complexity theorists to do it for the 3SUM and APSP breakthrough.

They didn't do this because those complexity theorists are advocating for a tyranny of the illiterate lmao

Why were we doing math in the first place? We should be happy that math problems were being solved, since presumably they were blockers for other problems in science and the like. But it feels like math was really more about seeking enlightenment, like a form of mental yoga or something. If so, we can just ignore AI proofs and continue on maybe?
To many professional mathematicians it's a form of art. But at the end of the day, it's a profession done for money, and even if they like doing it as a day to day job, they'd probably be doing something else if they had absolute freedom over their time. This is how I personally classify things as art or chore. If people continue doing something the same way when there is no financial motive, it's art in its pure form.
Some call that "entertainment", or "a hobby".
jvvw
Based on the research mathematicians I have known, I'd disagree with 'they'd probably be doing something else if they had absolute freedom over their time' for at least a large number of them, although it probably varies from individual to individual.
> even if they like doing it as a day to day job, they'd probably be doing something else if they had absolute freedom over their time.

Not the mathematicians I know. They’d happily drop the academic admin stuff, but they’d absolutely keep doing mathematics in much the same way.

You could replace "math" with "programming" in your comment for a different perspective.
If we're going to be frank about it, higher level math is:

1) Intellectually challenging, to such a degree that those wishing to enter the field need to have a certain level of intellectual prowess to do so. This creates some levels of mystique, with a sprinkle of elitism and gatekeeping.

2) Driven (among other things) by prestige. And the more pure the math is, the more prestigious it is.

3) So complex that people can spend their entire working careers chasing a handful of problems. The amount of time researchers spend on very specific problems is mind-boggling, if we think about the results.

4) Intensely captivating for the people deep in the weeds.

And the deeper you get, the longer you study, the more you start to value things like "mathematical beauty", and may start to view math as a form of art.

Like many similar fields, you end up with this ivory tower where people can dedicate their whole lives to thinking deeply about extremely niche and theoretical problems.

kamaal
Similar things are happening in the competitive programming world.

Quite a lot of people are not happy they aren't elite anymore, and many have spent years to decades to arrive here.

Simply put you invest years of your life to establish a kind of distinction over others, and that goes away. That hurts.

But its not something surprising. Most of these competitive programming problems were actually English languages puzzles, because you couldn't dial up the mathematical difficulty anymore making it a fields medal problem. And in most cases in simple language weren't even that hard to begin with, and you could look up solutions to these problems in an hour of Google searching.

I'm not totally against "math as art" but I don't think that formulation goes deep enough towards what's really going on with mathematics. There is something deeper where I lean more towards the philosophers who have said that mathematics is essentially ontology.

Since the Greeks we've had the idea that "Being and thinking are one," or that Being (in the sense of all of existence as such) has some essential unity with thought, and therefore can be thought, and expressed or submitted to the logos or reason. Being is in some sense fundamentally intelligible, and mathematics is the most developed, exacting, and articulate expression of Being.

Logic was understood in this older sense up to roughly the the mid to late 19th century. This is why a work like Hegel's Science of Logic begins not with syllogisms or propositions but with Being and Nothing. But this was forgotten after logic was mathematized by the English around the time of Russell, and its connection to ontology was gradually overshadowed by a focus on epistemology (still, it should be remembered, originally as a means of getting back to ontology).

There may be truth in art, but it's always haunted by its own historicity or contingency, which is to say untruth. Mathematics seems on the contrary the only really timeless, absolute thing we have. Part of what makes it captivating is stumbling on a construction or concept or proposition or theorem that simply must be, independent of us.

The AIs are certainly now more than automatic theorem provers, mechanically traversing some space of true propositions. They are able to push things forward and connect seemingly disparate domains to get to a proof, but to my mind it remains to be seen how well they will be able to form new concepts and definitions.

Imagine the controversy surrounding Cantor, for example, but put an AI in the place of Cantor. If an AI proposed something like the (infinite) hierarchy of infinity, would we have accepted it? What would the intuitionism debates have looked like? Would they even have taken place? And aside from that, has it actually been shown conclusively that an AI could propose such a thing?

There are lots of attempts right now to recover a humanism for mathematics, or restore man's pride of place with respect to it, but maybe we don't need to worry about that. Tao's attempts to preserve the mathematical community, while allowing for practices to change through the crisis may look like a kind of rearguard action, but seems reasonable to me and not really dependent on any kind of humanism. It's a way to avoid the question for now while things play out (and not conservative/reactionary like Scholze and others), which may be exactly what we need, because after all, perhaps we still don't understand why we do mathematics, what it's really for, and what our relation is to it. Whether it's enough to preserve funding is another issue.

NotGMan
Where was Terence Tao where other workers were being replaced by immigrants, robots and other automation?

Academics and white collars now get to experience what blue collar workers experienced in the past.

Same as what developers in USA experienced who were and are getting replaced by Indians.

He is a imigrant.
He grew up in Australia, didn't he? He was the one doing the replacing I suppose.
Pretty much every published paper I see now is a combination of Chinese and Indian immigrant authors. Given enough time they all become American. I don’t see what the problem is.
If you're considering an individual, maybe this makes sense.

But if you're consdiering a community, this falls apart. The maths community has universities, has professors who are paid, has students which are getting their degrees for varying reasons, it has conferences, has publications, papers, projects etc etc, all of which will get some negative impact some AI.

You know that line "when a measurement becomes a target it ceases to be a useful measurement". This line holds up to different degrees for various measurements and targets. For maths it holds up very well. The goal is "contribute to make the world better by increasing humanity's understanding of maths" and the measurement, which by evaluating an individual on it we're turning into a target, is "how much does the individual publish new findings". Measurement turned target holds up great. It's almost impossible to publish new findings and not contribute to humanity's understanding of maths. But with AI these two are being decoupled. You can produce lots of new findings, but the community is saturated and they don't get assimilated into humanity's understanding. Why do individuals use AI then? Because you've made the target "how much does the individual publish new findings" and they have to compete or lose.

kamaal
If a coal miner or car mechanic would say something like this, they would be called a luddite by the same science people.

Its different when your own job is on the line.

When human manual arts were being automated away it was supposed to be not only acceptable but any complain and you were told you were a progress blocking luddite.

Now that mental labor is getting automated, the response to automation is very different.

This isn't about physical labour vs mental labour, it's about the dynamics of delivering a product vs contributing to a community.

For example, a software engineer is like a car mechanic or a coal miner. None of those are anything like a mathematician.

The person you're rallying against isn't me, it's an imaginary hypocritical person which exists in your mind. I do mental labour, I welcome AI developments hard, and I still think TT is 100% correct.

2sk21
I'm a retired AI researcher and I still enjoy programming. However, I don't use any however assistance as I still want the thrill of learning new things. Unfortunately, it appears that the only way for most developers to enjoy similar activities is to be retired with enough money
This is exactly why half the conversation conversations about AI eventually become a critique of capitalism. AI isn’t for fun or joy, it’s explicitly for business use. For replacing people. For extracting more out of each worker. That’s all the major companies (OAI, Anthropic, etc) are offering. None of this is designed to give us any more freedom or time to pursue the things we are passionate about.

I feel like we’re circling back to that 2010s energy of “everyone can be an entrepreneur.” Now it’s “everyone can build software”

The amount developers saying how "now we can focus on the bigger picture" rather than the code are just delusional. The bigger picture has always been a bigger problem than writing the code. It's the part that engineering stands for in "software engineering". If you couldn't code something well or weren't involved in bigger picture decisions in the past, you're just gonna engineer big picture spaghetti.
rsfern
There’s a somewhat famous lecture [0] by Wigner (one of the greats of 20th century physics if you’re not familiar) on exactly this topic. One of his points is that new tools and ways of thinking developed on the way to solving mathematical problems with no apparent application often find downstream applications in science and engineering. If we’re skipping the part where we identify and understand the new math, will we still reap the unreasonable effectiveness? Tao’s position in this post suggests that the current wave of LLM successes is not conducive to this dynamic

0: https://en.wikipedia.org/wiki/The_Unreasonable_Effectiveness...

stared
Pure mathematics, in my view, is not science, not craft - it is art.

AI, however powerful, is a tool, only as important as the amount it helps mathematicians. Creating mathematics without human understanding is as sound as mass producing copies of Michelangelo David.

“Mathematics, rightly viewed, possesses not only truth, but supreme beauty — a beauty cold and austere, like that of sculpture [...] yet sublimely pure, and capable of a stern perfection such as only the greatest art can show.” - Bertrand Russell, "From The Study of Mathematics" (1902)

“A mathematician, like a painter or a poet, is a maker of patterns. [...] The mathematician's patterns, like the painter's or the poet's, must be beautiful; the ideas, like the colours or the words, must fit together in a harmonious way. Beauty is the first test: there is no permanent place in the world for ugly mathematics.” - G. H. Hardy, "A Mathematician's Apology" (1940)

> Creating mathematics without human understanding is as sound as mass producing copies of Michelangelo David.

In this case, it’s like creating the first instance of Michelangelo’s David.

armcat
I think this focus on a "holistic" approach applies to everything AI is touching now, not just math. On Twitter I see people one-shotting games, or reproducing games. If the goal is to just one-shot a game using AI, it's done. But if the goal is to produce immersive medium that people can truly enjoy, admire the story and the craftsmanship, and can find entire new ways of bringing a story to life, that's something else entirely.
orlp
Sadly the experience with physical products tells us that the vast majority of people prefer cheap disposable crap over expensive craftsmanship.
elAhmo
I don't think they prefer it, but the price is a deciding factor.

Most people would be happier with La Marzocco coffe machines which costs thousands of dollars, but if you get an OK shot with a 100 USD DeLongi, then the choice is clear for majority of the population.

La marzocco micro and mini are still high effort to make your coffee compared to some automated 90% cheaper delongi.

I think people sway to low effort endeavours that still have a reward at the end (even if it lesser reward than high effort).

I don't think this is really true for games, a lot of games have enormous resources thrown at their development and marketing and they fall flat because they simply aren't very fun. Then we have a million slop games that sell 3 copies on steam, and in practice only the most unique/fun/addictive games can break through.
uludag
The dynamics of software are vastly different from manufactured goods though. Software can be distributed at essentially zero cost so the disposable crap analogy breaks down.

Like given the choice, the vast majority of people would prefer one quality game like Minecraft, LoL, or Fortnite, vs. thousands of one-shot generated games, and looking at user playtime this is exactly what we see. If anything AI is just going to entrench these pre-AI franchises even more.

spuz
I said this when OpenAI announced they had solved a Millennium prize problem: solving open problems for the sake of it will lose its cachet. AI companies will no longer benefit by making these announcements. They've proven the effectiveness of their tool. If people want to use them to advance human knowledge then let them do that. There's no benefit to humanity to turn electricity into proofs just for the sake of it.
In this case though the compute per problem was pretty reasonable so if this model was available this rate of progress would mostly continue whether OpenAI funds it or not.
It's not just about proving the benefit of the tool, it's about proving that the latest version of the tool is better than the previous version and of the competitors' latest versions
Depends on who youre asking. I think for the general public, it's more likely that human mathematicians lose their cachet.