
AI-guided equation search reports strong benchmark results
An arXiv preprint details an LLM-guided equation-search method achieving 95% exact recovery on 100 Feynman problems, alongside tests on 240 scientific problems…

An arXiv preprint details an LLM-guided equation-search method achieving 95% exact recovery on 100 Feynman problems, alongside tests on 240 scientific problems…
1 editorial report · 1 verified social mention. The most authoritative report leads while later evidence completes the story.
The official benchmark card for Ling-3.0-flash-Fin is a useful reminder that the unit being tested is rarely just “the model.” The release says most runs used temperature 1, top_p 0.95 and the highest available reasoning effort. FinFIRST and FinSearchComp Verified used a common ReAct scaffold with Web Search, Visit and Python. SpreadsheetBench used Claude Code 2.1.173 with LibreOffice 25.8.7, Search disabled, 120 or 300 maximum interaction turns and a three-hour task timeout; Ling used temperature 0.6 there. The chart also mixes evidence types. Some results come from official or externally published scores, while others are internal runs. FinSearchComp Verified is an internal 145-question set with expert-revised answers and a GPT-5 judge. FinCRAFT is internal. FinFIRST is announced as “com
Open mention