Articles
Cost-Effective Automated Judging of Natural-Language Mathematical Proofs⭐9
Cheap open-weight LLMs judge math proofs with accuracy matching frontier models at up to 100× lower cost, though ensemble voting offers no advantage.
Cheap open-weight LLMs judge math proofs with accuracy matching frontier models at up to 100× lower cost, though ensemble voting offers no advantage.