<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title>Transformer - Tag - 枕石的个人博客</title><link>https://blog.zsmgc.love/tags/transformer/</link><description>Transformer - Tag - 枕石的个人博客</description><generator>Hugo -- gohugo.io</generator><language>zh-CN</language><managingEditor>3499129952@qq.com (枕石)</managingEditor><webMaster>3499129952@qq.com (枕石)</webMaster><lastBuildDate>Fri, 10 Jul 2026 10:00:00 +0800</lastBuildDate><atom:link href="https://blog.zsmgc.love/tags/transformer/" rel="self" type="application/rss+xml"/><item><title>Qwen3 推理过程：从输入文本到下一个 Token</title><link>https://blog.zsmgc.love/posts/qwen3_inference_architecture/</link><pubDate>Fri, 10 Jul 2026 10:00:00 +0800</pubDate><author>枕石</author><guid>https://blog.zsmgc.love/posts/qwen3_inference_architecture/</guid><description>&lt;p>前面在分析 FlashAttention、Paged KV Cache 和 Paged Attention 时，我一直在研究 Qwen3 推理过程中的局部模块：Attention 是怎么计算的、KV Cache 为什么要保存、Decode 为什么需要分页访问历史 KV。把这些模块单独拆开以后，反而容易丢掉一个更基础的问题：&lt;strong>一个输入句子究竟是怎样经过 Qwen3，最后变成下一个 Token 的？&lt;/strong>&lt;/p></description></item></channel></rss>