The Millennium Problems are a set of the most important open problems in mathematics. so far only 2 have been solved: the Poincare Conjecture by Grigori Perelman in 2010 and today OpenAI released this https://openai.com/index/navier-stokes-solution/

The only problem is, that their AI didn’t even come up with the solution itself. Mathematicians working on the same problem recently made big strides in solving the Navier Stokes equations, and their chat logs were scraped and used in training data before they could release the proof themselves.

Here is the unpaywalled statements of the Mathematician: https://mastodon.social/@tristanbuckmaster/117233413705701198

Western AI companies have been tackling a bunch of open problems and techbro chuds cannot shut up about it, this is just Marketing and the fact they have to steal real people’s work to do so just proves it. They really want that IPO moneyyy capitalist-laugh

  • searabbit@piefed.social
    link
    fedilink
    English
    arrow-up
    15
    ·
    1 day ago

    I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.

    Wow I had suspected this would eventually happen, but for it to be this blatant this quickly is alarming. They feel so emboldened due to the complete lack of consequences for stealing.

  • db0@lemmy.dbzer0.com
    link
    fedilink
    arrow-up
    8
    ·
    1 day ago

    They also spend the equivalent of something like 15% of that university’s mathematician research budget for the past 50 years to do this one proof, using plagiarized data to do so.

  • Hirom@beehaw.org
    link
    fedilink
    arrow-up
    5
    ·
    edit-2
    24 hours ago

    while unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models⁠.

    So either they don’t track what goes in their training data, or they do but their LLM is unable to precisely credit sources. Maybe a bit of both. That’s convenient, this way they claim great discoveries without crediting all prior work they’re relying on.

    LLMs are a great plausible deniability generator. These allow one to plagiarize while saying with a straight face no one knows what source material the result are based on.