[논문리뷰] Skaling: Chinchilla's Exponents Meet Kaplan's Coupling본 논문은 기존의 Additive Chinchilla law가 데이터 부족(data-scarce) 및 과잉 학습(overtraining) 극단 영역에서 Systematic prediction bias를 유발한다는 점을 지적한다.#Review#Neural Scaling Laws#Chinchilla#Kaplan#Loss Surface#Compute Allocation#Extrapolation#Sparse Profiling2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Understanding Reasoning from Pretraining to Post-Training본 논문은 LLM 훈련 과정에서 Pretraining 단계의 선택(모델 크기, 데이터 등)이 이후 RL 효율성에 미치는 정량적 관계를 규명하고자 한다.#Review#Reinforcement Learning#Pretraining#Scaling Law#LLM#Reasoning#Compute Allocation#Policy Evolution2026년 7월 19일댓글 수 로딩 중
[논문리뷰] ConceptMoE: Adaptive Token-to-Concept Compression for Implicit Compute Allocation대규모 언어 모델(LLMs)이 모든 토큰에 균일하게 연산을 할당하여 비효율적인 연산 자원 사용을 초래하는 문제를 해결하는 것이 목표입니다.#Review#MoE#LLMs#Adaptive Compression#Token Merging#Compute Allocation#Efficiency#Vision-Language Models#Continual Training2026년 1월 29일댓글 수 로딩 중