arxiv:2502.10197

MathConstruct: Challenging LLM Reasoning with Constructive Proofs

Published on Feb 14

Authors:

Nikola Jovanović ,

Abstract

While Large Language Models (LLMs) demonstrate impressive performance in mathematics, existing math benchmarks come with significant limitations. Many focus on problems with fixed ground-truth answers, and are often saturated due to problem simplicity or the viability of guessing or memorization. Crucially, they capture only a narrow subset of relevant math problems. To address this research gap, we introduce \mc, a new benchmark of 126 challenging problems sourced from various math competitions, which targets constructive proofs, a widely encountered problem type requiring the construction of mathematical objects with specific properties. These proofs are particularly suitable for LLM evaluation, as solution correctness can be easily verified. Our automated verifiers also enable MathConstruct to generate problem variations, used to evaluate robustness. State-of-the-art LLMs solve only 54% of MathConstruct problems, highlighting its complexity and importance for LLM evaluation.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment

No model linking this paper

Cite arxiv.org/abs/2502.10197 in a model README.md to link it from this page.

No dataset linking this paper

Cite arxiv.org/abs/2502.10197 in a dataset README.md to link it from this page.

No Space linking this paper

Cite arxiv.org/abs/2502.10197 in a Space README.md to link it from this page.

No Collection including this paper

Add this paper to a collection to link it from this page.