Mathematicians Seek Proof OpenAI Did Not Train on Their Unpublished Work
A group of mathematicians has raised formal concerns about whether OpenAI used their unpublished or proprietary mathematical work in training its models, and is demanding transparency about training data provenance. The dispute centers on the difficulty of detecting whether specialized, non-public mathematical content appeared in training corpora — a problem that is technically hard to audit with current tools. This is part of a broader pattern of domain experts questioning whether their intellectual output was ingested without consent or compensation. For developers and researchers building on OpenAI models, this highlights ongoing legal and ethical uncertainty around training data sourcing that could affect model licensing and liability in the future. It also underscores the demand for more rigorous data documentation practices from frontier AI labs.
Read original source ↗Part of the 2026-09-11 briefing→