In late 2023, a Reddit user known as "evilscientist311" conducted a comprehensive analysis of Cleo's activity patterns and interactions with other Stack Exchange accounts.
I think there are some interesting parallels between what the Cleo account did, and what the large AI companies are doing now with some of the proofs for previously unsolved theorems.
There may be some value in the answers, but for some the value is in the understanding of how to get from the problem to the solution.
I wouldn't call myself a mathematician, but I encountered a few problems in my CS classes where I could work through the proof by hand, see that an algorithm works, and even mathematically prove it, but yet still didn't feel like I understood why it worked, or how the algorithm was invented/derived. But for many, the fact that it did work was enough.
Joe McCann and his audience discovered that Cleo was a ploy at ragebaiting the math forum in order to drive engagement for more useful proofs. A real-world Cunningham's Law social experiment.
So maybe OpenAI's ridiculous approach to mathematics has value in how it's ragebaiting us to scrutinize and/or come up with better proofs!
AI capabilities keep getting bumpier and bumpier. There's no reason to believe it won't stall out on a front that it is already quite bad at (correctly formalizing in human language)
The point of understanding is to get better at proving new things. Now that humans are obsolete for proving things, human understanding is unnecessary.
> for some the value is in the understanding of how to get from the problem to the solution
All of the value is in this understanding.
So many people are still in denial about it, but "AI" really is just a natural language search engine. AI helps organize what has already been written by humans and was technically self-evident from our work.
If we allow ourselves to continue down this anti-intellectual path and stop caring about why proofs work, we've basically entered another dark age.
I'm not saying AI can't possibly help at all with building intuition, but clearly there are a lot of idiots in the class who just gotta raise their hand first to blurt out whatever trash they come up with on impulse. Those people need to be told to sit the fuck down already. Their uncontrolled ADHD is ruining everything. There's nothing intelligent about mindlessly rushing to compute an answer. That is literally what the last few years have felt like for everyone else.
> AI helps organize what has already been written by humans and was technically self-evident from our work.
I am not 100% sure about the "self-evident" part here.
To your point: yes, AI doesn't come up with new axioms. But then we enter the topic of "math is invented or discovered?". Each of the discoveries are logical consequences of the axioms, but that doesn't mean they were "self-evident", especially for humans.
If everything was already self-evident, AI would just be a useless tautological machine. However, I think it is more than just that
Well, it does make the solutions less impressive, right? Because it's much easier to compute the inverse of an integral than an integral. What's impressive here is more in composing problems that have nice answers and can't be solved by computer algebra systems, then the sock puppet trick to conceal that you ha the answer all along.
Reshetnikov didn't confess to them being reverse engineered, though. He confessed to them being conjecture arrived at after working on the problem for a bit by tweaking similar integrals in symbolic math software / numerical approximations. Every post was a gamble because it could have been wrong.
Of course, that could be a lie, but he came clean about other things that make sense retrospectively.
> Every post was a gamble because it could have been wrong
Symbolic differentiation is a deterministic algorithm. There is no reason for his posts to be a gamble. I'm sure there was a lots of interesting trial and error in developing good question and answer pairs, but once he posted the question he knew what the answer was already.
As I understand, he was specifically targeting problems that symbolic math software had given up on. He was only using symbolic software to solve the similarly shaped problems first, which allowed him to be fairly confident in his conjecture, as he became quite keen on how to slip his way from the output of symbolic math software from one solvable formula to a closely related formula that broke the software.
The symbolic math software couldn't perform the integration problems he posted. Integration lacks a deterministic algorithm and therefore permits constructing challenges for stack overflow and computer algebra systems. However, there is a deterministic algorithm to go the other way.
In other words, he was posting challenges to go from X -> Y (hard). However, going from Y -> X is easy. Therefore, since he had both X and Y before posting, he could verify the solutions perfectly well. The reason why constructing (X, Y) together is easier than going from Y->X is because you can always tweak Y a bit and then work backwards to see what X falls out.
There was certainly creativity in finding (X, Y) pairs where the symbolic software couldn't go from X -> Y. However, again this is more about trial and error iteration than producing brilliant insights from scratch, which is what it appeared Cleo was doing on Stack Overflow. Again, this is just a consequence of X -> Y being hard but Y -> X being easy. It was a wonderful parlor trick that took a good deal of effort to set up.
(It looks like OP's most recent post was deleted, but I spent some time writing this up so I'll post it anyway. The deleted post contended that Cleo never confessed to "cheating" by differentiating and then reversing the direction.)
I think you are misunderstanding something that is just implicit in the story. IE he doesn't need to "confess" this, it's a fundamental part of how the trick works.
I'll take one more stab at explaining what's going on. (Out of curiosity, are you familiar with calculus? I don't want to assume that you're not, but your comment reads as if you're not very familiar with it, so I'm going to explain things a bit better this time.)
Let's start with how the beautiful parlor trick looked to everyone else.
Random user: Asks how to integrate ABC expression.
(This is difficult, because integrate(ABC) has no general algorithm. It often requires many subtle tricks to perform a given integration, and there is no guarantee that there even is an elementary answer for integrate(ABC)).
Cleo: Answers integrate(ABC) = XYZ, with no notes.
(Wow! This obviously must have required many subtle tricks, but they are not provided!)
Importantly, anyone can verify that Cleo is right, because it turns out that UndoIntegration(XYZ) => ABC is easy, and there is a deterministic algorithm to do it. So, we all can tell that Cleo's answer is correct. But how could she have done this, since the integration direction is difficult???
-----
How the parlor trick really works.
First, as you can see, if Cleo takes any random XYZ, she can easily run UndoIntegration(XYZ) => ABC, and now she knows for free that Integrate(ABC) => XYZ. The neat thing is that doing this doesn't require figuring out any of the subtle steps required to run the Integrate operation either, which is convenient since Cleo doesn't plan to post them anyway.
So then, how does Cleo find a good ABC and XYZ without being a genius who is smarter than a computer? The most important thing is to find a pair such that ComputerAlgebraSystem_Integrate(ABC) doesn't work. Since integration is generally done by a bag of tricks, there are always holes you can find. So, you can basically just do this:
1. Start with a candidate XYZ_1.
2. Run UndoIntegration(XYZ_1) => ABC_1. (Remember, this is easy.)
3. Check if ComputerAlgebraSystem_Integrate(ABC_1) works. (Also easy to check, although the computer program has to work hard.)
4. If the computer is stumped, good. We can make a StackOverflow post.
5. Otherwise, try a different XYZ_2 and go back to step 1.
This is still an interesting game, but at no point does Cleo need to come up with a crazy bag of integration tricks here, the way that everyone assumes she did when she runs the parlor trick in the forum. This is the whole point of the trick, and the reason she hid her identity.
Apologies, I had figured out what distinction you were getting at right after writing my reply, so didn't feel the need to keep my misunderstanding up. But thanks for the detailed explanation!
Edit: Actually, I'm confused again. To be clear, it sounded to me from the discussion in the Joe McCann video that there was no tweaking the answer to get the question, i.e. that doing anything other than verification in reverse would have been against their ethic, that the answer must truly follow the question for it not to be cheating. The integral is fixed in place, and then legitimately solved by a highly competent Reshetnikov who is good at this, afforded plenty of time by scheming in advanced (as opposed to the illusion of only taking 3 hours), but is too lazy to formalize their work or is interested to see a 'clean' solution unbiased by their own approach to the problem or by the software that aided their work. And then it's verified trivially using differentiation (something symbolic math software almost never fails at), as opposed to my original misunderstanding that there still would have been uncertainty. Right?
No, although he was certainly an integral enthusiast, the video does not claim that he was solving them from scratch. Rather, it says he was starting from integrals with known answers and tweaking them slightly to see if he could break the CAS. At that point, although he did apparently try to solve the resulting problems himself, he already knew roughly what the answer would look like (by comparing to the previous answer, and also to the previous solution path.)
Specifically step 5 in my previous post is more work than it sounds, he was doing some calculations by hand, but it's like he's starting 90% of the way there and trying to do the last 10% by hand to fool the CAS. And remember, he can try as many variants of the tweak+10% as he wants until he finds something that works.
Treat it as an oracle problem. Posting exact closed forms without derivations forced the community to build the actual algorithmic bridges. High-effort trolling, higher-tier result.
The Wikipedia page says that in the early 2000s, he emigrated from his native Uzbekistan. Where too? Of course almost everyone would have the US as their first guess. And they'd be right. What a boon for the Americans to be the default destination for geniuses like this!
But now... What have you guys done? Would the guy even be allowed to come if he wanted to?
In late 2023, a Reddit user known as "evilscientist311" conducted a comprehensive analysis of Cleo's activity patterns and interactions with other Stack Exchange accounts.
However, I and others were on the scent at least as far back as April, 2023. See for example: https://x.com/TheDavidSJ/status/1650957407902658571
There may be some value in the answers, but for some the value is in the understanding of how to get from the problem to the solution.
I wouldn't call myself a mathematician, but I encountered a few problems in my CS classes where I could work through the proof by hand, see that an algorithm works, and even mathematically prove it, but yet still didn't feel like I understood why it worked, or how the algorithm was invented/derived. But for many, the fact that it did work was enough.
So maybe OpenAI's ridiculous approach to mathematics has value in how it's ragebaiting us to scrutinize and/or come up with better proofs!
All of the value is in this understanding.
So many people are still in denial about it, but "AI" really is just a natural language search engine. AI helps organize what has already been written by humans and was technically self-evident from our work.
If we allow ourselves to continue down this anti-intellectual path and stop caring about why proofs work, we've basically entered another dark age.
I'm not saying AI can't possibly help at all with building intuition, but clearly there are a lot of idiots in the class who just gotta raise their hand first to blurt out whatever trash they come up with on impulse. Those people need to be told to sit the fuck down already. Their uncontrolled ADHD is ruining everything. There's nothing intelligent about mindlessly rushing to compute an answer. That is literally what the last few years have felt like for everyone else.
I am not 100% sure about the "self-evident" part here.
To your point: yes, AI doesn't come up with new axioms. But then we enter the topic of "math is invented or discovered?". Each of the discoveries are logical consequences of the axioms, but that doesn't mean they were "self-evident", especially for humans.
If everything was already self-evident, AI would just be a useless tautological machine. However, I think it is more than just that
Doesn’t necessarily make the solutions less impressive, but it’s kind of a relevant detail that is left out in the Wikipedia summary.
Of course, that could be a lie, but he came clean about other things that make sense retrospectively.
Symbolic differentiation is a deterministic algorithm. There is no reason for his posts to be a gamble. I'm sure there was a lots of interesting trial and error in developing good question and answer pairs, but once he posted the question he knew what the answer was already.
In other words, he was posting challenges to go from X -> Y (hard). However, going from Y -> X is easy. Therefore, since he had both X and Y before posting, he could verify the solutions perfectly well. The reason why constructing (X, Y) together is easier than going from Y->X is because you can always tweak Y a bit and then work backwards to see what X falls out.
There was certainly creativity in finding (X, Y) pairs where the symbolic software couldn't go from X -> Y. However, again this is more about trial and error iteration than producing brilliant insights from scratch, which is what it appeared Cleo was doing on Stack Overflow. Again, this is just a consequence of X -> Y being hard but Y -> X being easy. It was a wonderful parlor trick that took a good deal of effort to set up.
I think you are misunderstanding something that is just implicit in the story. IE he doesn't need to "confess" this, it's a fundamental part of how the trick works.
I'll take one more stab at explaining what's going on. (Out of curiosity, are you familiar with calculus? I don't want to assume that you're not, but your comment reads as if you're not very familiar with it, so I'm going to explain things a bit better this time.)
Let's start with how the beautiful parlor trick looked to everyone else.
Random user: Asks how to integrate ABC expression.
(This is difficult, because integrate(ABC) has no general algorithm. It often requires many subtle tricks to perform a given integration, and there is no guarantee that there even is an elementary answer for integrate(ABC)).
Cleo: Answers integrate(ABC) = XYZ, with no notes.
(Wow! This obviously must have required many subtle tricks, but they are not provided!) Importantly, anyone can verify that Cleo is right, because it turns out that UndoIntegration(XYZ) => ABC is easy, and there is a deterministic algorithm to do it. So, we all can tell that Cleo's answer is correct. But how could she have done this, since the integration direction is difficult???
-----
How the parlor trick really works.
First, as you can see, if Cleo takes any random XYZ, she can easily run UndoIntegration(XYZ) => ABC, and now she knows for free that Integrate(ABC) => XYZ. The neat thing is that doing this doesn't require figuring out any of the subtle steps required to run the Integrate operation either, which is convenient since Cleo doesn't plan to post them anyway.
So then, how does Cleo find a good ABC and XYZ without being a genius who is smarter than a computer? The most important thing is to find a pair such that ComputerAlgebraSystem_Integrate(ABC) doesn't work. Since integration is generally done by a bag of tricks, there are always holes you can find. So, you can basically just do this:
1. Start with a candidate XYZ_1.
2. Run UndoIntegration(XYZ_1) => ABC_1. (Remember, this is easy.)
3. Check if ComputerAlgebraSystem_Integrate(ABC_1) works. (Also easy to check, although the computer program has to work hard.)
4. If the computer is stumped, good. We can make a StackOverflow post.
5. Otherwise, try a different XYZ_2 and go back to step 1.
This is still an interesting game, but at no point does Cleo need to come up with a crazy bag of integration tricks here, the way that everyone assumes she did when she runs the parlor trick in the forum. This is the whole point of the trick, and the reason she hid her identity.
Apologies, I had figured out what distinction you were getting at right after writing my reply, so didn't feel the need to keep my misunderstanding up. But thanks for the detailed explanation!
Edit: Actually, I'm confused again. To be clear, it sounded to me from the discussion in the Joe McCann video that there was no tweaking the answer to get the question, i.e. that doing anything other than verification in reverse would have been against their ethic, that the answer must truly follow the question for it not to be cheating. The integral is fixed in place, and then legitimately solved by a highly competent Reshetnikov who is good at this, afforded plenty of time by scheming in advanced (as opposed to the illusion of only taking 3 hours), but is too lazy to formalize their work or is interested to see a 'clean' solution unbiased by their own approach to the problem or by the software that aided their work. And then it's verified trivially using differentiation (something symbolic math software almost never fails at), as opposed to my original misunderstanding that there still would have been uncertainty. Right?
No, although he was certainly an integral enthusiast, the video does not claim that he was solving them from scratch. Rather, it says he was starting from integrals with known answers and tweaking them slightly to see if he could break the CAS. At that point, although he did apparently try to solve the resulting problems himself, he already knew roughly what the answer would look like (by comparing to the previous answer, and also to the previous solution path.)
Specifically step 5 in my previous post is more work than it sounds, he was doing some calculations by hand, but it's like he's starting 90% of the way there and trying to do the last 10% by hand to fool the CAS. And remember, he can try as many variants of the tweak+10% as he wants until he finds something that works.
[0] https://www.youtube.com/watch?v=7gQ9DnSYsXg&t=14s
Also, didn't Cleo post the problems themselves?
But now... What have you guys done? Would the guy even be allowed to come if he wanted to?