[논문리뷰] Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs본 논문은 구조적 압축과 4-bit 양자화가 순차적으로 적용된 대규모 언어 모델(LLM)에서 발생하는 성능 저하 문제를 해결하고자 합니다. 기존의 QAT 방식은 학습 과정에서 수렴이 느리고, 최적 지점을 지나치게 되면 모델 성능이 급격히 붕괴하는 불안정성을 보입니다.#Review#Large Language Models#Model Compression#Quantization-Aware Training#Knowledge Distillation#MXFP4#Model Healing2026년 8월 24일댓글 수 로딩 중
[논문리뷰] Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs본 논문은 VLM을 모바일 및 자원 제약 환경에 배포할 때 발생하는 막대한 메모리 및 컴퓨팅 자원 요구 문제를 해결하고자 합니다.#Review#Vision-Language Models#Quantization-Aware Training#Model Compression#Numerical Formats#Mobile Inference#Arm CPU2026년 8월 23일댓글 수 로딩 중
[논문리뷰] Gemma 4 Technical Report본 논문은 최신 LLM 생태계에서 요구되는 강력한 multimodal 이해도, 복잡한 추론 능력, 그리고 컴퓨팅 효율성을 동시에 달성하기 위해 Gemma 4 모델 제품군을 제안합니다.#Review#Multimodal#Mixture-of-Experts#Reasoning Trace#Speculative Decoding#Quantization-Aware Training#Long-context#Encoder-free2026년 7월 7일댓글 수 로딩 중
[논문리뷰] SLA2: Sparse-Linear Attention with Learnable Routing and QAT본 논문은 기존 Sparse-Linear Attention (SLA)의 한계, 즉 주의 가중치 크기에 기반한 휴리스틱 기반의 어텐션 분할 과 희소 및 선형 어텐션 출력 간의 불일치 를 해결하는 것을 목표로 합니다.#Review#Sparse-Linear Attention#Diffusion Models#Video Generation#Learnable Routing#Quantization-Aware Training#Attention Acceleration#Model Optimization2026년 2월 18일댓글 수 로딩 중
[논문리뷰] Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking본 논문은 텍스트, 이미지, 문서 이미지, 비디오 등 다양한 양식의 데이터를 통합 하여 고정밀 멀티모달 검색을 수행하는 Qwen3-VL-Embedding 및 Qwen3-VL-Reranker 모델 시리즈를 소개합니다.#Review#Multimodal Retrieval#Multimodal Ranking#Foundation Models#Embedding Models#Reranking Models#Contrastive Learning#Knowledge Distillation#Matryoshka Representation Learning#Quantization-Aware Training2026년 1월 11일댓글 수 로딩 중