LLMs struggle with math reasoning, because they can't conjecture (arxiv.org)

🤖 AI Summary
Researchers argue that a critical, previously neglected step—conjecturing—is necessary before autoformalisation (turning informal math into formal statements), and that ignoring it leads to overoptimistic assessments of LLM mathematical reasoning. They introduce ConjectureBench, augment existing benchmarks, and redesign evaluation and metrics to treat conjecture generation as an independent task rather than assuming the target conjecture is provided. The paper shows many problems cannot be formalised without first producing an explicit answer, bound, or claim, and finds that popular models’ autoformalisation scores (tested on GPT-4.1 and DeepSeek‑V3.1) drop substantially once conjecturing is evaluated explicitly. To address the gap they propose Lean‑FIRe, an inference‑time method that boosts conjecturing and end‑to‑end autoformalisation. With it they report the first successful full autoformalisation on 13 PutnamBench problems with GPT‑4.1 (and 7 with DeepSeek‑V3.1), demonstrating that LLMs often have the latent knowledge to propose correct conjectures but fail when conjecturing isn’t handled as a separate subtask. The work has practical implications for benchmark design, model training objectives, and formal‑math pipelines: improving conjecture generation and its integration is key to real progress in machine-assisted formal reasoning.
Loading comments...
loading comments...