[논문리뷰] SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs본 논문은 LLM의 Multi-task 성능 향상을 위한 두 주류 방법론인 SFT와 RL이 Multi-task 환경에서 보이는 극명한 거동 차이를 규명하고자 합니다. 기존 연구들은 SFT와 RL의 단일 태스크 학습 특성을 비교했으나, Multi-stage 학습 상황에서의 근본적인 차이는 명확히 설명하지 못했습니다.#Review#Large Language Models#Supervised Fine-Tuning#Reinforcement Learning#Multi-Task Learning#Gradient Interference#Parallel-RL2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors본 연구는 오토레그레시브 모델이 long-horizon rollout 과정에서 누적하는 오차(rollout error)를 배포 시점에 측정할 수 없는 문제를 해결합니다 .#Review#Bidirectional Diffusion#Round-Trip Consistency#Self-Supervised Uncertainty Quantification#Rollout Error#Latent Dynamics#Scientific Machine Learning2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression본 논문은 대규모 컨텍스트를 처리하는 LLM의 추론 비용을 줄이기 위한 Hard Prompt Compression 기법이, 정보 조각을 독립적으로 선택함에 따라 필연적으로 발생하는 구조적 결함인 Referential Dangling 문제를 제기합니다 .#Review#Prompt Compression#Hard Prompt Compression#Referential Dangling#Long-Context#LLM Inference#Information Retrieval#Contextual Dependency2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning본 연구는 기존 오디오 추론 모델이 결과 중심의 보상(Outcome-based rewards)이나 고정된 루브릭 보상(Process-based rewards) 체계 하에서 겪는 한계를 해결하고자 합니다.#Review#Audio Reasoning#Reinforcement Learning#Evolving Rubrics#Large Audio Language Models#Process-based Reward#Audio-grounded Supervision2026년 8월 9일댓글 수 로딩 중
[논문리뷰] PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say본 논문은 LLM 기반 에이전트가 외부 툴을 통해 작업 수행 시 과도한 민감 정보를 획득하는 문제를 해결하기 위해 PrivacyPeek을 제안합니다 .#Review#LLM-based Agents#Privacy Leakage#Data Acquisition#Benchmark#Contextual Integrity#Tool-use2026년 8월 9일댓글 수 로딩 중
[논문리뷰] OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction본 논문은 기존 감정 지능 연구가 특정 태스크에만 치중된 Specialist 모델 위주로 구성되어 있어, 감정의 다면적 속성을 포괄하지 못하고 태스크 간 시너지(Synergy)를 활용하지 못하는 문제를 해결하고자 합니다.#Review#Multimodal Large Language Models#Affective Computing#Reinforcement Learning#Reasoning Trajectories#Emotion Perception#Multi-task Learning2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Modular TTT: Rethinking Test-Time Training as Composable Modules본 논문은 기존 TTT 연구들이 각 모델 변형을 개별적으로 hard-coding하여 TTT 디자인 공간을 체계적으로 이해하기 어렵다는 문제에서 출발합니다 .#Review#Test-Time Training#Modular Framework#Sequence Modeling#Directed Acyclic Graph#Inner Learner#Fast Weights2026년 8월 9일댓글 수 로딩 중
[논문리뷰] FATE: Frame-Level Audio-Visual Temporal Embedding본 논문은 오디오-비주얼 모델에서 '무엇(what)'에 해당하는 의미적 정보와 '언제(when)'에 해당하는 시간적 정렬 정보가 효과적으로 결합되지 못하는 문제를 해결합니다.#Review#Audio-Visual Embedding#Temporal Alignment#Contrastive Learning#Event Localization#Cross-Modal Retrieval#Frame-Level Representation2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss본 연구는 자원 제약이 심한 환경에서 대규모 언어 모델(LLM)을 배포하기 위한 효율적인 지식 증류(Knowledge Distillation) 파이프라인 구축을 목표로 합니다.#Review#Knowledge Distillation#Offline Distillation#Fused Chunked KL Loss#Long-context Healing#LLM Compression#Sparse Top-K Logits2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Douyin Multimodal Embedding Model Technical Report본 논문은 대규모 산업용 다중 모달 플랫폼에서 효율적인 검색과 세밀한 의미론적 식별이라는 두 가지 상충하는 요구사항을 동시에 만족시키기 위해 DME를 제안합니다.#Review#Multimodal Embedding#Retrieval#Semantic Sufficiency#Latent Reasoning#Cross-Conditional Reconstruction#Bi-Encoder#Representation Learning2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events본 연구는 LLM 기반 에이전트가 시간이 지남에 따라 경험하는 삶의 사건에 대해 인간과 유사한 심리학적 성격 변화를 보이는지 규명하고자 합니다.#Review#LLM Agents#Personality Evolution#Big Five Personality#Life Events#BFI-Adapt#Psychometric Evaluation2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Characterizing the Quality Profile of AI-Generated C++ in Production본 논문은 대규모 엔터프라이즈 환경에서 AI가 생성한 C++ 코드의 품질, 성능, 유지보수 특성을 실증적으로 분석합니다 . 기존 연구들은 대부분 통제된 환경이나 벤치마크 데이터셋 위주로 수행되어, 실제 프로덕션 환경의 복잡한 코드 리뷰 및 배포 과정을 반영하지 못한다는 한계가 있었습니다.#Review#AI-generated code#Software quality taxonomy#Authoring-time provenance#Static analysis#Code review#Empirical software engineering2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence본 논문은 기존의 VLM들이 embodied agent의 반복적이고 동적인 실행 환경(observation–reasoning–action cycle)을 통합적으로 처리하지 못하고 특정 작업에 편향되어 있다는 점을 핵심 문제로 정의합니다 .#Review#Vision-Language Model#Embodied Intelligence#Capability Specialist#GRPO#TIES#Model Consolidation#Action Guidance2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding본 논문은 MLLM의 창의적 능력을 체계적으로 평가하기 위한 객관적인 지표가 부족하다는 점을 해결하고자 한다. 기존의 많은 멀티모달 벤치마크는 명시적인 정답이나 직접적인 보상 신호가 존재하는 작업 위주로 구성되어 있어, 창의성에 필수적인 novelty와 appropriateness를 평가하는 데 한계가 있다.#Review#Multimodal Models#Creativity#Cross-Concept Understanding#Multimodal Evaluation#Chengyu2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning본 논문은 단순한 multimodal 환경 수의 확장이 반드시 에이전트 학습에 이득이 되지 않으며, 오히려 환경 간 성능 저하를 초래할 수 있다는 문제 제기에서 시작합니다.#Review#Multimodal Agent#Environment Distribution#Curriculum Learning#Ability-aware Selection#Negative Transfer#Hierarchical Difficulty2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle본 논문은 현대의 멀티모달 파이프라인에서 시각적 콘텐츠의 권한이 생성자에게 귀속되지 않는 구조적 불균형 문제를 해결하고자 한다. 기존의 법적·제도적 대응은 무단 사용이 발생한 이후의 사후 조치에 국한되어 있으며, 콘텐츠가 AI 파이프라인으로 유입되기 전 단계의 기술적 개입이 절실하다.#Review#Adversarial Machine Learning#Proactive Protection#Visual Content Lifecycle#Adversarial Examples#Deep Learning Security#Generative AI#Content Provenance2026년 8월 9일댓글 수 로딩 중
[논문리뷰] Addressable Memory for Video World Models본 논문은 Autoregressive 비디오 세계 모델에서 긴 시간의 생성 과정 중 발생하는 visual persistence 저하 문제를 해결합니다.#Review#Video World Models#Addressable Memory#KV Cache Compression#RoPE#Temporal Coherence#Episodic Recall#WorldTrace2026년 8월 9일댓글 수 로딩 중
[hermes-agent] Hermes Agent: 10배 빠른 프로젝트 그룹화 최적화 분석불필요한 git subprocess 호출을 제거하고 캐싱 전략을 개선하여 프로젝트 트리 빌드 속도를 3.3초에서 0.3초로 단단히 개선한 사례입니다.#Python#Performance#Git#Optimization#Software Engineering2026년 8월 9일댓글 수 로딩 중
[sglang] SGLang, Sol-Attn 도입으로 비디오 생성 속도 1.23배 향상SGLang에 Sol-Attn 희소 어텐션 백엔드를 도입하여 비디오 생성 속도를 크게 향상시킨 기술 블로그 글입니다.#SGLang#AI#비디오 생성#최적화#Sol-Attn#희소 어텐션2026년 8월 9일댓글 수 로딩 중
[ultralytics] RT-DETR FLOPs 프로파일링 성능 최적화 및 안정화RT-DETR 모델의 FLOPs 계산 시 발생하는 비효율적인 stride-based proxy 방식을 개선하여 프로파일링 속도를 최대 5배 이상 향상시켰습니다.#RT-DETR#FLOPs#Performance Optimization#Deep Learning#THOP2026년 8월 9일댓글 수 로딩 중