A research team gave an AI pipeline 51,110 mathematics papers and asked it to look for unresolved problems in combinatorics.
The pipeline came back with a lot of work.
This paper, from researchers affiliated with Carnegie Mellon University and Anysphere, describes a system called FAR: Find, Attempt, and Recommend. A mathematician sets a research direction. FAR searches the literature, recovers open questions, attempts solutions, and narrows the results into a queue for expert review. That queue begins with a steep funnel. FAR identified 5,245 combinatorics papers, extracted 6,453 candidate conjectures or open problems, and checked 4,717 as apparently well posed and still open under the evidence it found. Each checked problem then received one attempt.
The system classified 1,050 outputs as claimed new resolutions. Three automated judges examined each claim, passing 598. A final grading stage recommended 77 as substantial enough for expert attention.
But those numbers can run away from us if we read them too quickly. The paper does not report 598 verified discoveries or 77 new theorems.
The authors manually reviewed 15 of the 77 recommended artifacts, choosing the ones they found particularly interesting. They judged all 15 mathematically correct. But one had already been solved a few months before their run through a route the pipeline's searches missed.
The proof held. The novelty claim did not.
And that correction points to the operating lesson. Faster generation creates a larger field of plausible work. It also creates more claims that need someone qualified to check the mathematics, search for prior results, decide whether the contribution matters, and take responsibility for reporting it.
So the authors designed FAR around that constraint. The system keeps each attempted resolution connected to its source paper and original statement. It uses cheaper models to label and extract, stronger models to attempt and judge, and a final recommendation stage to focus scarce expert attention. The output is a research queue rather than a truth machine.
Verification bottleneck
Verification is becoming the scarce institutional function.
- A literature-to-attempt pipeline can generate candidate resolutions faster than specialists can examine every proof and novelty claim.
- Mathematicians still have to verify correctness, prior work, significance, and the final public record.
- Watch for source-bound artifacts, random audits of rejected and recommended items, independent reproduction, correction trails, and clear labels separating machine judgment from expert review.
This unreviewed part of the experiment matters. The 15 checked artifacts were selected by author interest rather than random sampling, so their success cannot establish the accuracy of the full 77-item recommendation set. The study is also one arXiv preprint, one combinatorics pilot, and one model pipeline. This review did not run the authors' code or independently reproduce the results.
Still, the shape of the problem travels. Science already has more papers than any person can read. If AI adds another layer of candidate proofs, counterexamples, experiments, and interpretations, institutions will need better ways to decide what deserves expensive human attention.
Opportunities
A builder could create a source-bound candidate packet that travels with each proposed result. It would include the original question, source context, search trail, attempted solution, automated critiques, unresolved objections, expert owner, and final disposition.
There is also room for dedicated novelty checking. A useful service could search discipline-specific indexes, preprint servers, proceedings, and citation networks before a candidate is described as new. The missed prior resolution in this pilot shows why novelty deserves its own evidence trail.
Research groups could also build specialist-routing queues. Instead of handing one general reviewer a pile of polished outputs, the system could match each artifact to the narrow expertise required, preserve the review decision, and record whether the work was corrected, rejected, merged with prior work, or prepared for publication.
The value sits in making expert attention easier to aim and easier to audit.
The pipeline can fill the tray. A qualified person still has to decide what survives.
