#AI Benchmarks

1개의 포스트

[논문리뷰] Towards Robust Mathematical Reasoning

기존 수학 벤치마크들의 포화 상태와 단답형 답변 위주의 한계를 극복하기 위해, 논문은 국제 수학 올림피아드(IMO) 수준의 견고한 수학적 추론 능력을 평가하는 새로운 벤치마크 스위트인 IMO-Bench 를 제안합니다.

#Review #Mathematical Reasoning #Large Language Models (LLMs)#AI Benchmarks #International Mathematical Olympiad (IMO)#Proof Verification #Automatic Grading #Robustness

2025년 11월 9일