<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://bits-bytes-nn.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://bits-bytes-nn.github.io/" rel="alternate" type="text/html" /><updated>2026-08-19T23:26:20+00:00</updated><id>https://bits-bytes-nn.github.io/feed.xml</id><title type="html">Bits, Bytes and Neural Networks</title><subtitle>A tech blog focusing on AI/ML paper reviews and latest research trends</subtitle><author><name>Jonas Kim</name></author><entry xml:lang="en"><title type="html">Writing Context Into Your Data — AI-Ready Data, Semantic Layers, Knowledge Graphs, Ontologies</title><link href="https://bits-bytes-nn.github.io/insights/data-architecture/2026/07/27/ai-ready-data-semantic-layer-knowledge-graph-en.html" rel="alternate" type="text/html" title="Writing Context Into Your Data — AI-Ready Data, Semantic Layers, Knowledge Graphs, Ontologies" /><published>2026-07-27T12:00:00+00:00</published><updated>2026-07-27T12:00:00+00:00</updated><id>https://bits-bytes-nn.github.io/insights/data-architecture/2026/07/27/ai-ready-data-semantic-layer-knowledge-graph-en</id><author><name>Jonas Kim</name></author><category term="Insights" /><category term="Data-Architecture" /><category term="AI-Ready-Data" /><category term="Semantic-Layer" /><category term="Knowledge-Graph" /><category term="Ontology" /><category term="GraphRAG" /><category term="Agentic-AI" /><category term="Model-Context-Protocol" /><category term="Data-Governance" /><summary type="html"><![CDATA["AI-ready data" isn't clean data — it's context written in machine-executable form. Why the same schema takes text-to-SQL accuracy from 16.7% to 54.2%.]]></summary></entry><entry xml:lang="ko"><title type="html">데이터에 맥락을 새겨 넣는 법 — AI Ready Data, Semantic Layer, Knowledge Graph, Ontology</title><link href="https://bits-bytes-nn.github.io/insights/data-architecture/2026/07/27/ai-ready-data-semantic-layer-knowledge-graph.html" rel="alternate" type="text/html" title="데이터에 맥락을 새겨 넣는 법 — AI Ready Data, Semantic Layer, Knowledge Graph, Ontology" /><published>2026-07-27T12:00:00+00:00</published><updated>2026-07-27T12:00:00+00:00</updated><id>https://bits-bytes-nn.github.io/insights/data-architecture/2026/07/27/ai-ready-data-semantic-layer-knowledge-graph</id><author><name>Jonas Kim</name></author><category term="Insights" /><category term="Data-Architecture" /><category term="AI-Ready-Data" /><category term="Semantic-Layer" /><category term="Knowledge-Graph" /><category term="Ontology" /><category term="GraphRAG" /><category term="Agentic-AI" /><category term="Model-Context-Protocol" /><category term="Data-Governance" /><summary type="html"><![CDATA['AI-ready 데이터'는 깨끗한 데이터가 아니라 맥락을 기계가 실행할 수 있는 형태로 새겨 넣은 데이터입니다. 같은 스키마에서 text-to-SQL 정확도가 16.7%에서 54.2%로 오른 이유를 짚습니다.]]></summary></entry><entry><title type="html">Amazon Bedrock AgentCore를 하네스로 읽다</title><link href="https://bits-bytes-nn.github.io/insights/agentic-ai/2026/04/12/agentcore-harness-engineering-analysis.html" rel="alternate" type="text/html" title="Amazon Bedrock AgentCore를 하네스로 읽다" /><published>2026-04-12T12:00:00+00:00</published><updated>2026-04-12T12:00:00+00:00</updated><id>https://bits-bytes-nn.github.io/insights/agentic-ai/2026/04/12/agentcore-harness-engineering-analysis</id><author><name>Jonas Kim</name></author><category term="Insights" /><category term="Agentic-AI" /><category term="AgentCore" /><category term="AWS-Bedrock" /><category term="Harness-Engineering" /><category term="Agentic-Infrastructure" /><category term="Model-Context-Protocol" /><category term="Cedar-Policy" /><category term="Managed-RAG" /><category term="Agent-Registry" /><category term="Agentic-AI" /><summary type="html"><![CDATA[에이전트의 '나머지 전부'를 AWS는 어떻게 제품화했나. Amazon Bedrock AgentCore를 Build·Deploy·Assess 세 층으로 열어 하네스 체크리스트의 어디를 채우고 어디를 비워 두었는지 짚습니다.]]></summary></entry><entry xml:lang="en"><title type="html">From Prompts to Harnesses — Four Years of AI Agentic Patterns</title><link href="https://bits-bytes-nn.github.io/insights/agentic-ai/2026/04/05/evolution-of-ai-agentic-patterns-en.html" rel="alternate" type="text/html" title="From Prompts to Harnesses — Four Years of AI Agentic Patterns" /><published>2026-04-05T12:00:00+00:00</published><updated>2026-04-05T12:00:00+00:00</updated><id>https://bits-bytes-nn.github.io/insights/agentic-ai/2026/04/05/evolution-of-ai-agentic-patterns-en</id><author><name>Jonas Kim</name></author><category term="Insights" /><category term="Agentic-AI" /><category term="Prompt-Engineering" /><category term="Context-Engineering" /><category term="Harness-Engineering" /><category term="Agentic-Patterns" /><category term="LLM-Architecture" /><category term="Vibe-Coding" /><category term="Agentic-AI" /><summary type="html"><![CDATA[Engineering rigor never disappeared — it relocated. Four years, three paradigm shifts from prompts to context to harnesses, traced by why each era failed.]]></summary></entry><entry xml:lang="ko"><title type="html">프롬프트에서 하네스까지 — AI 에이전틱 패턴 4년의 기록</title><link href="https://bits-bytes-nn.github.io/insights/agentic-ai/2026/04/05/evolution-of-ai-agentic-patterns.html" rel="alternate" type="text/html" title="프롬프트에서 하네스까지 — AI 에이전틱 패턴 4년의 기록" /><published>2026-04-05T12:00:00+00:00</published><updated>2026-04-05T12:00:00+00:00</updated><id>https://bits-bytes-nn.github.io/insights/agentic-ai/2026/04/05/evolution-of-ai-agentic-patterns</id><author><name>Jonas Kim</name></author><category term="Insights" /><category term="Agentic-AI" /><category term="Prompt-Engineering" /><category term="Context-Engineering" /><category term="Harness-Engineering" /><category term="Agentic-Patterns" /><category term="LLM-Architecture" /><category term="Vibe-Coding" /><category term="Agentic-AI" /><summary type="html"><![CDATA[엔지니어링의 엄밀함은 사라지지 않았습니다, 이동했을 뿐입니다. 프롬프트에서 컨텍스트로, 컨텍스트에서 하네스로 패러다임이 세 번 바뀐 2022-2026년을 각 시대가 왜 실패했는지로 추적합니다.]]></summary></entry><entry xml:lang="en"><title type="html">Claude Code Architecture Analysis</title><link href="https://bits-bytes-nn.github.io/insights/agentic-ai/2026/03/31/claude-code-architecture-analysis.html" rel="alternate" type="text/html" title="Claude Code Architecture Analysis" /><published>2026-03-31T12:00:01+00:00</published><updated>2026-03-31T12:00:01+00:00</updated><id>https://bits-bytes-nn.github.io/insights/agentic-ai/2026/03/31/claude-code-architecture-analysis</id><author><name>Jonas Kim</name></author><category term="Insights" /><category term="Agentic-AI" /><category term="Claude-Code" /><category term="Agentic-Architecture" /><category term="Context-Compaction" /><category term="Multi-Agent-Orchestration" /><category term="Security-Architecture" /><category term="Agentic-AI" /><summary type="html"><![CDATA[An npm source map leak exposed all 4,600+ files of Claude Code's core engine. Its 8-layer security and 4-tier message compaction, read as architecture.]]></summary></entry><entry xml:lang="ko"><title type="html">Claude Code 내부 아키텍처 분석</title><link href="https://bits-bytes-nn.github.io/insights/agentic-ai/2026/03/31/claude-code-source-map-leak-analysis.html" rel="alternate" type="text/html" title="Claude Code 내부 아키텍처 분석" /><published>2026-03-31T12:00:00+00:00</published><updated>2026-03-31T12:00:00+00:00</updated><id>https://bits-bytes-nn.github.io/insights/agentic-ai/2026/03/31/claude-code-source-map-leak-analysis</id><author><name>Jonas Kim</name></author><category term="Insights" /><category term="Agentic-AI" /><category term="Claude-Code" /><category term="Agentic-Architecture" /><category term="Context-Compaction" /><category term="Multi-Agent-Orchestration" /><category term="Security-Architecture" /><category term="Agentic-AI" /><summary type="html"><![CDATA[npm Source Map 유출로 Claude Code의 비공개 코어 엔진 4,600여 파일이 드러났습니다. 8계층 보안과 4단 메시지 압축, 에이전틱 루프를 아키텍처로 읽습니다.]]></summary></entry><entry><title type="html">DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models</title><link href="https://bits-bytes-nn.github.io/paper%20reviews/language-models/2025/12/02/deepseek-v3.2-pushing-the-frontier-of-open-large-language-models.html" rel="alternate" type="text/html" title="DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models" /><published>2025-12-02T09:25:14+00:00</published><updated>2025-12-02T09:25:14+00:00</updated><id>https://bits-bytes-nn.github.io/paper%20reviews/language-models/2025/12/02/deepseek-v3.2--pushing-the-frontier-of-open-large-language-models</id><author><name>DeepSeek AI</name></author><category term="Paper Reviews" /><category term="Language-Models" /><category term="DeepSeek-Sparse-Attention" /><category term="Scalable-Reinforcement-Learning-Framework" /><category term="Large-Scale-Agentic-Task-Synthesis-Pipeline" /><category term="Group-Relative-Policy-Optimization-Scaling" /><category term="Unbiased-KL-Estimate-for-RL" /><category term="Off-Policy-Sequence-Masking" /><category term="Keep-Routing-for-Mixture-of-Experts" /><category term="Thinking-Context-Management-for-Tool-Use" /><category term="Long-Chain-of-Thought-Cold-Start-Integration" /><category term="Multi-Stage-Agentic-Environment-Synthesis" /><category term="Mixture-of-Experts" /><category term="Reasoning-Models" /><category term="DeepSeek" /><summary type="html"><![CDATA[최근 몇 년간 대규모 언어 모델의 발전은 눈부신 성과를 이루었지만, 오픈소스 모델과 상용 모델 간의 성능 격차가 점점 벌어지고 있는 현실에 직면하고 있습니다. GPT-5, Claude-4.5-Sonnet, Gemini 3.0과 같은 비공개 상용 모델들이 복잡한 추론 작업에서 급속도로 성능을 향상시키는 동안, Qwen3, GLM, MiniMax-M2 등의 오픈소스 모델들은 상대적으로 뒤처지고 있는 상황입니다. 이러한 격차는 단순한 성능 차이를 넘어 오픈소스 커뮤니티의 기술 혁신 능력에 대한 근본적인 의문을 제기합니다…]]></summary></entry><entry><title type="html">Kimi K2: Open Agentic Intelligence</title><link href="https://bits-bytes-nn.github.io/paper%20reviews/language-models/2025/07/28/kimi-k2-open-agentic-intelligence.html" rel="alternate" type="text/html" title="Kimi K2: Open Agentic Intelligence" /><published>2025-07-28T05:35:43+00:00</published><updated>2025-07-28T05:35:43+00:00</updated><id>https://bits-bytes-nn.github.io/paper%20reviews/language-models/2025/07/28/kimi-k2--open-agentic-intelligence</id><author><name>Moonshot AI</name></author><category term="Paper Reviews" /><category term="Language-Models" /><category term="MuonClip-Optimizer" /><category term="QK-Clip-Attention-Stabilization" /><category term="Large-Scale-Agentic-Data-Synthesis" /><category term="Multi-Stage-Reinforcement-Learning-with-Self-Critique" /><category term="Mixture-of-Experts-Sparsity-Scaling-Law" /><category term="Synthetic-Data-Rephrasing-for-Token-Efficiency" /><category term="Verifiable-Rewards-Reinforcement-Learning" /><category term="Agentic-Intelligence-Framework" /><category term="Computational-Efficiency-in-Large-Language-Models" /><category term="Efficient-Mixture-of-Experts-Architecture" /><category term="Mixture-of-Experts" /><summary type="html"><![CDATA[대규모 언어 모델(LLM)의 발전은 인공지능 분야에서 혁명적인 변화를 예고하고 있습니다. 그러나 기존 모델들은 정적인 데이터 모방에 그치며, 실제 환경에서 자율적으로 추론하고 행동하는 능력에 한계를 보였습니다. 특히 도구 사용, 소프트웨어 개발, 복잡한 다단계 추론과 같은 에이전틱 인텔리전스 영역에서 기존 모델들의 성능은 매우 제한적이었습니다. 연구팀은 이러한 한계를 극복하고, 모델이 단순한 응답 생성을 넘어 실제 환경과 상호작용하며 학습하고 적응할 수 있는 새로운 접근법의 필요성을 인식했습니다.]]></summary></entry><entry><title type="html">Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities</title><link href="https://bits-bytes-nn.github.io/paper%20reviews/multimodal-learning/2025/07/07/gemini-2.5-pushing-the-frontier-with-advanced-reasoning-multimodality-long-context-and-next-generation-agentic-capabilities.html" rel="alternate" type="text/html" title="Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities" /><published>2025-07-07T17:36:04+00:00</published><updated>2025-07-07T17:36:04+00:00</updated><id>https://bits-bytes-nn.github.io/paper%20reviews/multimodal-learning/2025/07/07/gemini-2.5--pushing-the-frontier-with-advanced-reasoning--multimodality--long-context--and-next-generation-agentic-capabilities</id><author><name>Google DeepMind</name></author><category term="Paper Reviews" /><category term="Multimodal-Learning" /><category term="Sparse-Mixture-of-Experts-Architecture" /><category term="Long-Context-Reasoning" /><category term="Dynamic-Thinking-Budget" /><category term="Multimodal-Tool-Use" /><category term="Natively-Multimodal-Transformer" /><category term="Reinforcement-Learning-from-Human-and-Critic-Feedback" /><category term="Agentic-Workflow-Generation" /><category term="Thinking-Mode-Fusion" /><category term="Advanced-Reasoning-Capabilities" /><category term="Automated-Red-Teaming" /><category term="Reasoning-Models" /><category term="Multimodal-Models" /><summary type="html"><![CDATA[인공지능 기술의 급속한 발전과 함께 대규모 언어 모델의 능력을 확장하고 개선하는 것은 현대 AI 연구의 핵심 과제로 자리 잡았습니다. Google의 Gemini 팀은 기존 AI 모델들이 가진 한계를 극복하고, 더욱 복잡하고 다양한 작업을 수행할 수 있는 멀티모달 AI 시스템을 개발하고자 했습니다. 특히 코딩, 추론, 긴 컨텍스트 처리, 멀티모달 이해와 같은 영역에서 기존 모델들의 성능적 제약을 뛰어넘는 것이 주요 연구 동기였습니다.]]></summary></entry><entry><title type="html">Qwen3 Technical Report</title><link href="https://bits-bytes-nn.github.io/paper%20reviews/language-models/2025/05/14/qwen3-technical-report.html" rel="alternate" type="text/html" title="Qwen3 Technical Report" /><published>2025-05-14T13:41:34+00:00</published><updated>2025-05-14T13:41:34+00:00</updated><id>https://bits-bytes-nn.github.io/paper%20reviews/language-models/2025/05/14/qwen3-technical-report</id><author><name>Alibaba Group</name></author><category term="Paper Reviews" /><category term="Language-Models" /><category term="Thinking-Budget-Mechanism" /><category term="Strong-to-Weak-Distillation" /><category term="Dynamic-Mode-Switching" /><category term="Mixture-of-Experts-Architecture" /><category term="Fine-Grained-Expert-Segmentation" /><category term="Global-Batch-Load-Balancing-Loss" /><category term="Long-Chain-of-Thought-Cold-Start" /><category term="Reasoning-Reinforcement-Learning" /><category term="Thinking-Mode-Fusion" /><category term="Multi-Stage-Post-Training-Recipe" /><category term="Mixture-of-Experts" /><category term="Reasoning-Models" /><summary type="html"><![CDATA[인공지능 분야에서 대규모 언어 모델의 발전은 끊임없이 인간의 지능에 근접하려는 도전의 연장선상에 있습니다. GPT-4o, Claude 3.7, Gemini 2.5와 같은 최신 모델들은 인공 일반 지능(AGI)을 향한 중요한 이정표를 제시하고 있지만, 여전히 계산 효율성, 다국어 지원, 추론 능력 등에서 한계를 보여왔습니다. 특히 기존 모델들은 복잡한 추론 작업과 다양한 언어 환경에서 일관된 성능을 유지하는 데 어려움을 겪어왔습니다.]]></summary></entry><entry><title type="html">Gemma 3 Technical Report</title><link href="https://bits-bytes-nn.github.io/paper%20reviews/multimodal-learning/2025/03/25/gemma-3-technical-report.html" rel="alternate" type="text/html" title="Gemma 3 Technical Report" /><published>2025-03-25T15:52:34+00:00</published><updated>2025-03-25T15:52:34+00:00</updated><id>https://bits-bytes-nn.github.io/paper%20reviews/multimodal-learning/2025/03/25/gemma-3-technical-report</id><author><name>Google DeepMind</name></author><category term="Paper Reviews" /><category term="Multimodal-Learning" /><category term="Alternating-Local-Global-Attention" /><category term="Long-Context-Adaptation" /><category term="Pan-&amp;-Scan-Image-Processing" /><category term="Vision-Encoder-Token-Condensation" /><category term="Multimodal-Knowledge-Distillation" /><category term="RoPE-Positional-Embedding-Extension" /><category term="Grouped-Query-Attention" /><category term="Quantization-Aware-Training" /><category term="Sliding-Window-Attention-Optimization" /><category term="Efficient-Long-Context-Attention-Mechanism" /><category term="Multimodal-Models" /><summary type="html"><![CDATA[대규모 언어 모델의 발전은 AI 기술의 핵심 동력으로 자리 잡았지만, 기존 모델들은 여전히 심각한 한계를 가지고 있었습니다. 특히 긴 컨텍스트 처리의 메모리 비효율성, 제한된 멀티모달 능력, 그리고 다국어 성능의 불균형은 AI 시스템의 실용성을 크게 제한하는 주요 문제였습니다. Google DeepMind 연구팀은 이러한 근본적인 기술적 제약을 극복하고, 더욱 접근성 높고 효율적인 AI 모델을 개발하고자 Gemma 3 프로젝트를 시작했습니다.]]></summary></entry><entry><title type="html">DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning</title><link href="https://bits-bytes-nn.github.io/paper%20reviews/language-models/2025/01/22/deepseek-r1-incentivizing-reasoning-capability-in-llms-via-reinforcement-learning.html" rel="alternate" type="text/html" title="DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning" /><published>2025-01-22T15:19:35+00:00</published><updated>2025-01-22T15:19:35+00:00</updated><id>https://bits-bytes-nn.github.io/paper%20reviews/language-models/2025/01/22/deepseek-r1--incentivizing-reasoning-capability-in-llms-via-reinforcement-learning</id><author><name>DeepSeek AI</name></author><category term="Paper Reviews" /><category term="Language-Models" /><category term="Large-Scale-Reinforcement-Learning-on-Base-Model" /><category term="Group-Relative-Policy-Optimization" /><category term="Reasoning-Oriented-Reinforcement-Learning" /><category term="Reinforcement-Learning-with-Cold-Start" /><category term="Distillation-of-Reasoning-Capability" /><category term="Multi-Stage-Reinforcement-Learning-with-Self-Critique" /><category term="Verifiable-Rewards-Reinforcement-Learning" /><category term="Rejection-Sampling-and-Supervised-Fine-Tuning" /><category term="Iterative-Reinforcement-Learning" /><category term="Unified-Paradigm-for-Reinforcement-Learning" /><category term="Reasoning-Models" /><category term="Alignment" /><category term="DeepSeek" /><summary type="html"><![CDATA[대규모 언어 모델(LLM)의 발전은 인공지능 분야에서 혁명적인 변화를 가져왔지만, 추론 능력의 근본적인 한계는 여전히 중요한 도전 과제로 남아있었습니다. 기존의 지도 학습 미세 조정 방법은 모델에 외부에서 정의된 지식을 주입하는 데 집중했지만, 모델 스스로 복잡한 문제 해결 전략을 자율적으로 개발하는 능력은 제한적이었습니다. 특히 수학, 코딩, 과학적 추론과 같은 고도의 논리적 사고를 요구하는 영역에서 언어 모델들은 일관된 성능을 보이지 못했습니다.]]></summary></entry><entry><title type="html">Zep: A Temporal Knowledge Graph Architecture for Agent Memory</title><link href="https://bits-bytes-nn.github.io/paper%20reviews/retrieval-augmented-generation/2025/01/20/zep-a-temporal-knowledge-graph-architecture-for-agent-memory.html" rel="alternate" type="text/html" title="Zep: A Temporal Knowledge Graph Architecture for Agent Memory" /><published>2025-01-20T16:52:48+00:00</published><updated>2025-01-20T16:52:48+00:00</updated><id>https://bits-bytes-nn.github.io/paper%20reviews/retrieval-augmented-generation/2025/01/20/zep-a-temporal-knowledge-graph-architecture-for-agent-memory</id><author><name>Preston Rasmussen et al.</name></author><category term="Paper Reviews" /><category term="Retrieval-Augmented-Generation" /><category term="Temporally-Aware-Knowledge-Graph-Engine" /><category term="Bi-Temporal-Knowledge-Graph-Modeling" /><category term="Hierarchical-Knowledge-Graph-Construction" /><category term="Dynamic-Edge-Invalidation-for-Temporal-Reasoning" /><category term="Episodic-and-Semantic-Memory-Subgraphs" /><category term="Community-Detection-with-Label-Propagation" /><category term="Hybrid-Search-with-Breadth-First-Graph-Traversal" /><category term="Non-Lossy-Knowledge-Graph-Updates" /><category term="Multi-Hop-Entity-and-Relationship-Extraction" /><category term="Graph-Based-Memory-Retrieval-for-LLM-Agents" /><category term="Retrieval-Augmented-Generation" /><category term="Knowledge-Graph" /><summary type="html"><![CDATA[대규모 언어 모델(LLM) 기반 대화형 에이전트는 산업과 연구 커뮤니티에서 광범위하게 활용되고 있지만, 근본적인 한계를 마주하고 있습니다. LLM의 컨텍스트 윈도우 크기 제약, 긴 맥락에 대한 효과적인 정보 활용의 어려움, 그리고 사전 학습 이후의 새로운 정보나 도메인 특화 지식에 대한 무지가 그것입니다. 이러한 제약을 극복하기 위해 검색 증강 생성(RAG) 기법이 등장했으나, 기존 RAG 접근법은 주로 정적인 코퍼스를 다루도록 설계되어 있어 끊임없이 진화하는 대화 데이터와 비즈니스 정보를 효과적으로 처리하지 못합니다. 에이전트가…]]></summary></entry><entry><title type="html">DeepSeek-V3 Technical Report</title><link href="https://bits-bytes-nn.github.io/paper%20reviews/language-models/2024/12/27/deepseek-v3-technical-report.html" rel="alternate" type="text/html" title="DeepSeek-V3 Technical Report" /><published>2024-12-27T04:03:16+00:00</published><updated>2024-12-27T04:03:16+00:00</updated><id>https://bits-bytes-nn.github.io/paper%20reviews/language-models/2024/12/27/deepseek-v3-technical-report</id><author><name>DeepSeek AI</name></author><category term="Paper Reviews" /><category term="Language-Models" /><category term="Auxiliary-Loss-Free-Load-Balancing" /><category term="Multi-Token-Prediction" /><category term="Multi-Head-Latent-Attention" /><category term="DeepSeekMoE-Architecture" /><category term="FP8-Mixed-Precision-Training" /><category term="Efficient-Cross-Node-All-to-All-Communication" /><category term="Node-Limited-Routing" /><category term="Computation-Communication-Overlap" /><category term="Tile-Wise-Fine-Grained-Quantization" /><category term="Speculative-Decoding" /><category term="Mixture-of-Experts" /><category term="DeepSeek" /><summary type="html"><![CDATA[대규모 언어 모델(LLM) 분야는 최근 몇 년간 급속한 발전을 거듭하고 있으며, 인공 일반 지능(AGI)을 향한 중요한 이정표를 계속해서 세우고 있습니다. 그러나 기존 모델들은 여전히 계산 효율성, 훈련 비용, 추론 성능 측면에서 상당한 한계를 보이고 있었습니다. 특히 클로즈드소스 모델들에 비해 오픈소스 모델들의 성능 격차가 큰 문제였으며, 이는 연구진들이 DeepSeek-V3를 개발하게 된 핵심 동기가 되었습니다.]]></summary></entry><entry><title type="html">Tulu 3: Pushing Frontiers in Open Language Model Post-Training</title><link href="https://bits-bytes-nn.github.io/paper%20reviews/language-models/2024/11/22/tulu-3-pushing-frontiers-in-open-language-model-post-training.html" rel="alternate" type="text/html" title="Tulu 3: Pushing Frontiers in Open Language Model Post-Training" /><published>2024-11-22T18:44:04+00:00</published><updated>2024-11-22T18:44:04+00:00</updated><id>https://bits-bytes-nn.github.io/paper%20reviews/language-models/2024/11/22/tulu-3--pushing-frontiers-in-open-language-model-post-training</id><author><name>Allen Institute for AI</name></author><category term="Paper Reviews" /><category term="Language-Models" /><category term="Reinforcement-Learning-with-Verifiable-Rewards" /><category term="Multi-Stage-Post-Training-Recipe" /><category term="Direct-Preference-Optimization" /><category term="Persona-Driven-Data-Synthesis" /><category term="Skill-Specific-Synthetic-Data-Generation" /><category term="Prompt-Decontamination" /><category term="Open-Language-Model-Evaluation-System" /><category term="Length-Normalized-Preference-Optimization" /><category term="Asynchronous-Reinforcement-Learning-Infrastructure" /><category term="Skill-Targeted-Model-Training" /><category term="Alignment" /><summary type="html"><![CDATA[대규모 언어 모델의 발전은 인공지능 분야에서 혁명적인 변화를 가져왔지만, 사후 훈련 방법론에서 오픈소스와 폐쇄형 접근법 사이에는 여전히 큰 격차가 존재했습니다. 기존의 폐쇄형 모델들은 훈련 데이터와 방법론을 투명하게 공개하지 않아 연구자들의 접근을 제한했고, 오픈소스 모델들은 성능과 정교함에서 뒤처져 있었습니다. 특히 지시 따르기, 수학적 추론, 코딩과 같은 핵심 기술 영역에서 오픈소스 모델들의 성능은 상업용 모델에 비해 현저히 낮았습니다.]]></summary></entry><entry><title type="html">Pixtral 12B</title><link href="https://bits-bytes-nn.github.io/paper%20reviews/multimodal-learning/2024/10/09/pixtral-12b.html" rel="alternate" type="text/html" title="Pixtral 12B" /><published>2024-10-09T17:16:22+00:00</published><updated>2024-10-09T17:16:22+00:00</updated><id>https://bits-bytes-nn.github.io/paper%20reviews/multimodal-learning/2024/10/09/pixtral-12b</id><author><name>Mistral AI</name></author><category term="Paper Reviews" /><category term="Multimodal-Learning" /><category term="RoPE-2D-Positional-Encoding" /><category term="Block-Diagonal-Attention-Masking" /><category term="Flexible-Vision-Encoder-Architecture" /><category term="Natively-Multimodal-Transformer-Architecture" /><category term="Multi-Turn-Instruction-Tuning" /><category term="Sequence-Packing-Optimization" /><category term="Variable-Image-Resolution-Processing" /><category term="Break-Tokens-for-Image-Tokenization" /><category term="Standardized-Multimodal-Evaluation" /><category term="MM-MT-Bench-Benchmark" /><category term="Multimodal-Models" /><summary type="html"><![CDATA[현대 인공지능 연구에서 멀티모달 언어 모델의 발전은 매우 중요한 과제로 대두되고 있습니다. 기존의 대부분 멀티모달 모델들은 이미지 이해 능력과 텍스트 처리 능력 사이에 심각한 성능 불균형을 보였으며, 특히 오픈소스 모델들은 상업용 클로즈드 모델들에 비해 현저히 낮은 성능을 나타냈습니다. 연구진은 이러한 한계를 극복하고, 실제 사용자 환경에서 유용하게 활용될 수 있는 멀티모달 AI 시스템의 필요성을 깊이 인식했습니다.]]></summary></entry><entry><title type="html">LightRAG: Simple and Fast Retrieval-Augmented Generation</title><link href="https://bits-bytes-nn.github.io/paper%20reviews/retrieval-augmented-generation/2024/10/08/lightrag-simple-and-fast-retrieval-augmented-generation.html" rel="alternate" type="text/html" title="LightRAG: Simple and Fast Retrieval-Augmented Generation" /><published>2024-10-08T08:00:12+00:00</published><updated>2024-10-08T08:00:12+00:00</updated><id>https://bits-bytes-nn.github.io/paper%20reviews/retrieval-augmented-generation/2024/10/08/lightrag--simple-and-fast-retrieval-augmented-generation</id><author><name>Beijing University of Posts and Telecommunications</name></author><category term="Paper Reviews" /><category term="Retrieval-Augmented-Generation" /><category term="Graph-Based-Text-Indexing" /><category term="Dual-Level-Retrieval-Paradigm" /><category term="Low-Level-Entity-Retrieval" /><category term="High-Level-Relationship-Retrieval" /><category term="Graph-Enhanced-Entity-and-Relationship-Extraction" /><category term="LLM-Profiling-for-Key-Value-Pair-Generation" /><category term="Incremental-Knowledge-Base-Updates" /><category term="Graph-Vector-Hybrid-Retrieval" /><category term="Multi-Hop-Subgraph-Information-Extraction" /><category term="Deduplication-for-Graph-Optimization" /><category term="Retrieval-Augmented-Generation" /><category term="Knowledge-Graph" /><summary type="html"><![CDATA[검색 증강 생성(RAG) 기술은 대규모 언어 모델이 외부 지식 소스를 활용하여 더욱 정확하고 맥락에 적합한 응답을 생성할 수 있도록 하는 중요한 기술입니다. 그러나 기존 RAG 시스템들은 두 가지 근본적인 한계를 가지고 있습니다. 첫째, 많은 방법들이 평면적 데이터 표현에 의존하고 있어 엔터티 간의 복잡한 관계를 기반으로 정보를 이해하고 검색하는 능력이 제한됩니다. 둘째, 다양한 엔터티와 그 상호 관계에 걸쳐 일관성을 유지하는 맥락 인식 능력이 부족하여, 사용자의 질의에 완전히 대응하지 못하는 단편적인 응답이 생성됩니다. 예를…]]></summary></entry><entry><title type="html">Gemma 2: Improving Open Language Models at a Practical Size</title><link href="https://bits-bytes-nn.github.io/paper%20reviews/language-models/2024/07/31/gemma-2-improving-open-language-models-at-a-practical-size.html" rel="alternate" type="text/html" title="Gemma 2: Improving Open Language Models at a Practical Size" /><published>2024-07-31T19:13:07+00:00</published><updated>2024-07-31T19:13:07+00:00</updated><id>https://bits-bytes-nn.github.io/paper%20reviews/language-models/2024/07/31/gemma-2--improving-open-language-models-at-a-practical-size</id><author><name>Google DeepMind</name></author><category term="Paper Reviews" /><category term="Language-Models" /><category term="Knowledge-Distillation-for-Small-Language-Models" /><category term="Interleaving-Local-Global-Attention" /><category term="Grouped-Query-Attention" /><category term="Logit-Soft-Capping" /><category term="RMSNorm-Stabilization" /><category term="Multi-Stage-Post-Training-Recipe" /><category term="Sliding-Window-Attention-Optimization" /><category term="Instruction-Fine-Tuning-with-Direct-Preference-Optimization" /><category term="Model-Merging-through-Weight-Averaging" /><category term="Responsible-Open-Model-Development" /><summary type="html"><![CDATA[대규모 언어 모델(LLM)의 발전은 최근 인공지능 분야에서 가장 주목받는 연구 영역 중 하나입니다. 기존의 대규모 모델들은 놀라운 성능을 보여주었지만, 대부분 계산 비용이 매우 높고 접근성이 제한적이었습니다. 특히 소규모 모델들의 성능 개선은 주로 훈련 길이 증가에 의존해 왔으며, 이는 데이터셋 크기에 대해 로그적으로만 확장되는 한계를 가지고 있었습니다. Chinchilla 논문에서 제시된 스케일링 법칙에 따르면, 최신 소형 모델들이 최첨단 성능을 1-2% 개선하기 위해서는 최대 15조 토큰이 필요하다는 점이 이러한 접근법의…]]></summary></entry><entry><title type="html">The Llama 3 Herd of Models</title><link href="https://bits-bytes-nn.github.io/paper%20reviews/language-models/2024/07/31/the-llama-3-herd-of-models.html" rel="alternate" type="text/html" title="The Llama 3 Herd of Models" /><published>2024-07-31T17:54:27+00:00</published><updated>2024-07-31T17:54:27+00:00</updated><id>https://bits-bytes-nn.github.io/paper%20reviews/language-models/2024/07/31/the-llama-3-herd-of-models</id><author><name>Meta AI</name></author><category term="Paper Reviews" /><category term="Language-Models" /><category term="Direct-Preference-Optimization" /><category term="Scaling-Laws-for-Large-Language-Models" /><category term="Multimodal-Knowledge-Distillation" /><category term="Efficient-Long-Context-Attention-Mechanism" /><category term="Multilingual-Multimodal-Understanding" /><category term="Safety-Alignment" /><category term="Synthetic-Data-Generation-for-Mathematics" /><category term="Tool-Use-Emergence" /><category term="Rejection-Sampling-Fine-Tuning" /><category term="Multi-Stage-Post-Training-Recipe" /><category term="Llama" /><summary type="html"><![CDATA[대규모 언어 모델의 발전은 인공지능 분야에서 가장 중요한 연구 주제 중 하나로 자리 잡았습니다. 기존의 언어 모델들은 여러 가지 한계점을 가지고 있었는데, 특히 데이터 품질, 모델 규모, 그리고 다국어 및 다중 모달 능력에서 제한적이었습니다. Meta AI 연구팀은 이러한 한계를 극복하고 더욱 강력하고 유연한 언어 모델을 개발하고자 Llama 3 프로젝트를 시작했습니다.]]></summary></entry></feed>