Concerns have emerged in the AI community following the revelation that OpenAI provided financial support to a project assessing AI’s mathematical capabilities? Epoch AI, a nonprofit, disclosed only recently that OpenAI funded the development of FrontierMath, a benchmark designed to challenge AI systems with complex mathematical problems?
Behind the Scenes
Epoch AI, primarily financed by Open Philanthropy, announced on December 20 that OpenAI had contributed to FrontierMath? This benchmark was part of OpenAI’s testing for its anticipated AI model, o3? However, many contributors to FrontierMath were unaware of OpenAI’s involvement until this public disclosure? The lack of transparency raised concerns about the benchmark’s perceived objectivity?
Community Reactions
On forums like LessWrong, users expressed unease about the secretive nature of OpenAI’s involvement, fearing it could compromise the benchmark�s credibility? One contributor, under the alias ‘Meemi,’ criticized the opacity surrounding OpenAI’s funding, emphasizing the need for transparency for contributors making choices about their work?
Addressing the Issues
Tamay Besiroglu, Associate Director of Epoch AI, acknowledged these missteps? Besiroglu admitted that disclosing the partnership was not feasible until the o3 launch but recognized the need for more negotiation flexibility to inform contributors earlier? He stated that while OpenAI had access to the problems in FrontierMath, there exists a verbal agreement not to use the problems to train AI systems � a practice akin to ‘teaching to the test?’
Maintaining Benchmark Integrity
As a precaution, Epoch AI established a ‘separate holdout set’ to ensure the independence and accuracy of benchmarking results? Besiroglu assured the community that OpenAI respected this decision, supporting the benchmark’s integrity?
Awaiting Verification
Despite these efforts, Epoch AI’s lead mathematician, Elliot Glazer, noted difficulty in verifying OpenAI’s results independently? While Glazer believes in the legitimacy of OpenAI’s FrontierMath results, full confidence will only come with their own evaluation results?
This development highlights the intricate balance in creating empirical benchmarks for AI, emphasizing the importance of transparency and the potential conflict of interest when securing development resources?

