Responsible Release of AI-Generated Mathematics

September 29, 2026

Back to main page


At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.

1. Background

The mathematical community has long-standing norms concerning the dissemination and peer review of results. These norms have been essential for the reliability, trustworthiness, and effectiveness of mathematical work. One of the most important scholarly norms in mathematics is that the authors of a paper should understand the mathematical argument of that paper, have verified its correctness themselves, and take full responsibility for the content. Moreover, in the mathematical community, authors of works that significantly advance the field regularly give seminars at other institutions and talks at conferences, explaining their new developments and answering questions from colleagues. This all works towards developing the deepest possible human understanding of mathematics, which is one of the crowning glories of millennia of human development.

However, it is now the case that AI can output mathematical arguments in situations without the human who prompted it being able to understand the arguments, verify them, or take responsibility for them. We believe that human understanding of mathematics remains of paramount importance. How, in this new era, can we work towards a new paradigm that includes human understanding of mathematics as part of responsible scholarly output?

We asked the mathematical community for feedback about what it would mean for AI labs to responsibly release mathematical results and received over 600 replies (the survey asked about a specific situation in which OpenAI announced the existence of many results without giving details). Informed by these responses, we arrived at a set of recommendations, supported by a clear plurality of respondents, aimed at any AI lab whose models are likely to have a significant impact on mathematics. The overarching principles underlying these recommendations are the following.

  1. If AI labs produce significant mathematical results, they should responsibly release the results, as outlined in this document, as soon as possible.
  2. AI labs that release substantial mathematical output without immediate accompanying human understanding must take responsibility for ensuring that human understanding will follow. In particular, AI labs should provide significant support, including funding, to help develop this understanding.
  3. The development of human understanding must remain organic and community led. It should not be directed by AI labs, even when the labs have produced the results.

We present our recommendations themselves in Section 2, and in Section 3 we make some comments about access to powerful model.

2. Responsible release of results generated by AI labs

We recommend two possible courses of action, depending on the level of human understanding that accompanies a result.

2.A. Papers that a human understands

Papers for which there is a mathematician responsible who fully understands the content should follow the academic mathematical community’s traditional norms: the mathematician(s) concerned should post a preprint, submit a paper for peer review at a journal, and give talks to explain the work to other mathematicians.

2.B. Papers that are not yet understood by anybody

The recommendations below are for labs that have AI mathematical output that is not understood by the people who prompted the AI systems. They are split into two parts. The first part is a set of proposed technical norms for the release of AI-generated mathematics. The second part is a recommendation that AI labs provide support for the additional mathematical activities that are needed for humans to be able to understand and assimilate their AI-generated mathematical output and identify possible applications of it.

Step I: Initial release

1. With the help of LLMs, it is easy to make substantial improvements to the initial written version of an AI-generated result. The following actions should be carried out by the AI labs rather than left to mathematicians afterwards.

(a) The literature should be scoured for any ideas that are related to the ideas in the proofs of the results released. Even if the AI lab’s model discovered those ideas independently, it should follow standard mathematical practice and cite the papers in which the ideas were first introduced.

(b) The model, or some other model, should be prompted to produce a version of each proof that is written up in a style that follows the conventions of a traditional mathematical paper. They should contain friendly introductions and precise theorem statements and proofs. They
should not be full of wordy reasoning and non-standard terminology that renders them virtually incomprehensible.

The current abilities of LLMs may not be able to reproduce the level of attribution or quality of exposition that we expect of mathematicians, and thus more important work from mathematicians after release may be needed to reach this standard. If so, this should be supported as described in Step II. However, inability to reach a high standard does not absolve AI labs of the responsibility to do the best they can with their models on the two points above.

2. When results are announced, they should be deposited in a timely manner in appropriate scholarly repositories. These should not be controlled by any AI lab and should guarantee certain standards, including that submissions have a persistent citable identifier and that subsequent modifications are appropriately recorded. For AI-generated results, it would be particularly useful if the repository
allows for comments on papers. We strongly recommend that AI labs refrain from treating the release of mathematical results as marketing vehicles to promote their models, ignoring the substantial negative externalities that this practice inflicts on the mathematical community.

3. For each result released, the AI lab should make public the name of the model, the prompts used, a (summarized) chain of thought, the time taken, and the estimated cost of computation. Releasing additional material that sheds light on the scientific process that produced the results, such as the initial LLM outputs before they were cleaned up (as recommended in (1) above), is strongly encouraged.

4. As far as possible, a proof released by an AI lab should be formalized. Formalization artifacts should meet community standards, including having copyright headers, a challenge file for comparator, and a formalization.yaml. It would be helpful to include machine-readable metadata correlating the natural language and formal artifacts. Where formalization would lead to unacceptable delays, the formalization status of the paper should be clearly stated. For example, perhaps it has been formalized modulo standard results that are accepted by the community

5. Each time a solution to a problem is released, it should be clearly documented how exactly AI came to be used on that particular problem. If many results are released at once, then in addition to the results themselves a further document should be written and made public that references all of the released results and explains how many other problems of comparable difficulty the models tried and failed to solve, as well as how the problems were chosen.

Step II: Supporting human understanding

One of our principles is that AI labs have a responsibility to provide support, including funding, for the development of human understanding of the AI mathematical output that they release. Here we give more detail about what such support might be used for, though the ultimate decision about which particular efforts are supported should be taken neither by this group nor by AI labs, but by separate, already existing nonprofit institutions that have established processes for distributing funding. The scope of the appropriate level of support should depend on the importance and complexity of the released AI output, as these features determine the amount of work that will be needed to support human understanding of the output.

A few examples of activities that could be supported are the following.

  1. For results with highly complex proofs or results that could open many new avenues of investigation, conferences or summer schools dedicated to understanding them could be critical for analysis and understanding of the proofs so they can be applied in future mathematical developments.
  2. For other results, a more targeted workshop or a longer-term working group may be necessary. For long-term work, postdocs or students might be supporting the work and thus need to be funded.
  3. For some results, it would be useful to support experts to write books or long expository articles.

It should be understood that the acceptance of such support by the mathematical community would not be conferring legitimacy on the practices of the AI labs. Rather, it would be asking the AI labs to make an appropriate contribution to fulfilling their responsibility towards the development of human understanding of their AI output.

3. Ensuring broad access

The use of proprietary internal models by AI labs to do mathematical research risks creating a two-tier system where labs outrun the rest of the field, effectively alienating the mathematical community from its own discipline.


As for publicly available models, unequal access due to economic and other factors risks establishing a multi-tier hierarchy across the mathematical world and greatly exacerbating existing inequalities that arise from institutional wealth and technological privilege.

We therefore advise the AI labs to grant the global mathematical community broad, equitable access to their publicly available models: this is a way both to accelerate progress and to ensure wide access to it. Mathematics advances primarily through collective verification, conceptual synthesis, and shared intuition, none of which can thrive when some instruments generating frontier results are accessible only to a select few.