<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://chaojiezhang.me/feed.xml" rel="self" type="application/atom+xml"/><link href="https://chaojiezhang.me/" rel="alternate" type="text/html" hreflang="en"/><updated>2026-10-06T06:34:05+00:00</updated><id>https://chaojiezhang.me/feed.xml</id><title type="html">blank</title><subtitle>Chaojie Zhang — plasma/accelerator physicist at UCLA working on plasma wakefield acceleration, ultrafast laser-matter interaction, and AI/ML-driven modeling. </subtitle><entry><title type="html">理解 Transformer</title><link href="https://chaojiezhang.me/posts/2026/09/transformer-experimentalist/" rel="alternate" type="text/html" title="理解 Transformer"/><published>2026-09-18T00:00:00+00:00</published><updated>2026-09-18T00:00:00+00:00</updated><id>https://chaojiezhang.me/posts/2026/09/transformer-experimentalist</id><content type="html" xml:base="https://chaojiezhang.me/posts/2026/09/transformer-experimentalist/"><![CDATA[<p>本文介绍 Transformer 从输入文本到输出预测的基本计算过程，包括 token embedding、注意力、多头结构、残差连接、encoder–decoder，以及训练和推理的区别。阅读需要矩阵乘法和梯度的基础知识。</p> <p>文中以《Attention Is All You Need》的原始架构为主，并在相关位置说明它与 GPT-2 的区别。除特别说明外，省略 batch 维度，以单条序列为单位讨论。</p> <h2 id="1-词表大小序列长度与隐藏维度">1. 词表大小、序列长度与隐藏维度</h2> <p>计算中需要区分三个量：</p> <table> <thead> <tr> <th>符号</th> <th>含义</th> <th>决定什么</th> </tr> </thead> <tbody> <tr> <td>\(N_{\mathrm{vocab}}\)</td> <td>词表里有多少种 token</td> <td>embedding 表的行数</td> </tr> <tr> <td>\(S\)</td> <td>当前序列有多少个 token</td> <td>当前状态矩阵的行数</td> </tr> <tr> <td>\(d\)</td> <td>每个位置的向量维度</td> <td>当前状态矩阵的列数</td> </tr> </tbody> </table> <p>“100k 上下文”说的是可容纳的 token 数量上限。实际输入只有 1,000 个 token 时，当前序列长度就是 1,000。词表里的其他 token 不会因此各占一个注意力位置。</p> <p>分词器有不同设计。BPE 等子词方法会从较小单位出发，把频繁出现的组合合并为 token。词表不必收录每一个完整单词：一个没见过的长词，仍可能被拆成已有片段。采用完整字节覆盖的方法可以进一步用字节表示输入；其他方法可能需要未知 token。</p> <p>词表大小通常通过实验和预算取舍。词表小，文字可能切得更碎；词表大，embedding 和词表输出层更大，一些罕见 token 也更难获得充分训练。在模型训练完成后，词表和编号映射通常保持固定。</p> <h2 id="2-token-embedding">2. Token embedding</h2> <p>每个 token 有一个编号。编号用于查表，其大小和相邻关系不表示语义关系。设 embedding 表为：</p> \[E\in\mathbb{R}^{N_{\mathrm{vocab}}\times d}.\] <p>根据当前序列的 token 编号，从 \(E\) 中取出对应的 \(S\) 行，按序列顺序排列，得到：</p> \[X_{\mathrm{token}}\in\mathbb{R}^{S\times d}.\] <p>同一个 token 出现多次，就会重复取出同一行，放到不同的序列位置。向量的数值通过训练确定，单个分量通常没有明确的语义标签，信息可以由多个分量共同表示。</p> <p>从头训练时，embedding 权重通常随机初始化，然后与整个网络一起优化。训练好的模型直接加载这张表。同一个 token 的初始向量相同，但后续会根据上下文产生不同状态。</p> <p>这里需要区分两类量：<strong>参数是模型保存的可训练数值；状态是给定输入后计算出的中间结果。</strong> \(E\) 是参数矩阵，查表得到的 \(X_{\mathrm{token}}\) 是当前输入的表示。</p> <h2 id="3-位置编码">3. 位置编码</h2> <p>仅有 token embedding 还没有显式表示顺序。原论文将每个位置的编码与该位置的 token embedding 相加。设当前序列的位置编码矩阵为 \(P\)，则：</p> \[X=\sqrt{d}\,X_{\mathrm{token}}+P,\qquad X,P\in\mathbb{R}^{S\times d}.\] <p>这里的 \(\sqrt{d}\) 是原论文采用的 embedding 缩放系数。相加按元素进行，结果的尺寸不变。token 编号决定从 embedding 表取哪一行；位置编号决定使用哪个位置向量。词表大小与可表示的位置数量没有必要相等。</p> <p>同一个 token 出现在不同位置时，token embedding 相同，位置编码不同。相加后，网络便能利用位置差异区分不同的排列顺序。</p> <p>原始 Transformer 使用固定的正弦、余弦位置编码，不同分量具有不同频率。GPT-2 使用可学习的绝对位置表。前者的位置向量由公式生成，后者的位置向量参与训练。[1, 2]</p> <p>下文用 \(X\) 表示送入某个注意力模块的状态：第一层接收上述输入，后续层接收上一层处理后的结果。为具体说明尺寸，采用原论文 base 配置：\(d=512\)，8 个注意力头，每头 64 维。</p> <h2 id="4-查询键和值">4. 查询、键和值</h2> <p>一个注意力头有三个可训练的投影矩阵：</p> \[W_{Q},W_{K},W_{V}\in\mathbb{R}^{512\times64}.\] <p>它们分别作用于同一个输入矩阵：</p> \[Q=XW_{Q},\qquad K=XW_{K},\qquad V=XW_{V}.\] <p>因此，\(Q\)、\(K\)、\(V\) 都是 \(S\times64\) 的矩阵。它们的每一行对应一个序列位置：</p> <table> <thead> <tr> <th>矩阵</th> <th>含义</th> <th>作用</th> </tr> </thead> <tbody> <tr> <td>\(Q\)</td> <td>查询（query）</td> <td>提供接收位置的匹配特征</td> </tr> <tr> <td>\(K\)</td> <td>键（key）</td> <td>提供源位置的匹配特征</td> </tr> <tr> <td>\(V\)</td> <td>值（value）</td> <td>提供后续加权求和的内容</td> </tr> </tbody> </table> <p>“某个位置的 query”指 \(Q\) 中对应的那一行，是一个向量。同一个头中，所有位置都使用同一套 \(W_{Q}\)。矩阵乘法将每一行分别投影，再把结果按行排列；每个位置并没有独立的投影参数。\(K\) 和 \(V\) 同理。</p> <p>\(K\) 用于计算匹配分数，\(V\) 用于后续的加权求和。分别使用两个可训练投影，使模型可以独立调整匹配所需的特征和需要传递的内容。</p> <h2 id="5-注意力权重与加权求和">5. 注意力权重与加权求和</h2> <p>先计算所有接收位置与源位置之间的匹配分数：</p> \[QK^{\mathsf{T}}\in\mathbb{R}^{S\times S}.\] <p>上标 \(\mathsf{T}\) 表示转置。尺寸关系为：</p> \[(S\times64)(64\times S)=S\times S.\] <p>这个分数矩阵的每一行对应一个接收位置，每一列对应一个源位置。\(Q\) 的某一行与 \(K\) 的每一行分别做点积，得到 \(S\) 个分数。<strong>点积内部对向量分量求和，不同源位置的分数则分别保留。</strong></p> <p>设遮挡矩阵为 \(M\)：允许读取的位置取 0，禁止读取的位置取负无穷。注意力权重为：</p> \[A=\operatorname{softmax}_{\mathrm{row}} \left(\frac{QK^{\mathsf{T}}}{\sqrt{64}}+M\right).\] <p>缩放用于控制点积分数的尺度。softmax 按行执行，将每行分数转换为非负、总和为 1 的权重；被遮挡的位置权重为零。这里假设每行至少有一个允许读取的位置。</p> <p>随后用权重对值向量进行混合：</p> \[H=AV,\qquad (S\times S)(S\times64)=S\times64.\] <p>\(H\) 是这个头的输出。它的每一行，是 \(V\) 中各行按对应注意力权重形成的加权和。</p> <figure> <img src="/assets/img/posts/transformer/attention.svg" alt="输入矩阵分别投影为查询、键和值；查询与键计算匹配分数，经缩放、遮挡和 softmax 得到注意力矩阵，再与值矩阵相乘得到输出。" style="width:100%;height:auto;" loading="lazy"/> <figcaption>图 1：单个注意力头的计算过程。匹配分数决定权重，值矩阵提供被加权求和的内容。</figcaption> </figure> <p>因此，不同位置的信息通过两步影响输出：其他位置的键向量影响读取权重，其他位置的值向量参与输出的加权求和。</p> <p>这也解释了同一个 token 如何在不同语境下产生不同表示。即使某个位置的初始表示相同，周围位置的键和值不同，注意力输出也可以不同。新状态继续进入后续层，但不会直接写回 embedding 表；训练时更新参数表的是反向传播得到的梯度。</p> <p>将输入依赖关系显式写出：</p> \[H=A(X)V(X).\] <p>给定 \(A\) 后，加权求和对 \(V\) 是线性的；整体注意力操作则是非线性的，因为 \(A\) 也由输入计算得到。注意力权重描述这一层的信息读取，不能单独解释模型的最终判断。</p> <h2 id="6-多头注意力">6. 多头注意力</h2> <p>每个头都读取完整的输入矩阵，再通过各自的投影生成 \(Q\)、\(K\)、\(V\)。每个 \(512\times64\) 投影矩阵都能将全部 512 个输入分量混合到 64 个输出分量中，并不是只读取输入的一部分。</p> <p>多个头的投影矩阵可以并排放成一个大矩阵，一次乘法算完，再将结果拆成各头。这与分别投影在数学上等价。</p> <p>随后，各头分别计算自己的注意力权重矩阵。这样，同一个接收位置可以通过不同的头，使用不同权重读取信息。如果把所有头的通道合起来只计算一次点积和 softmax，就只得到一套位置权重。</p> <figure> <img src="/assets/img/posts/transformer/multihead.svg" alt="完整输入分别进入八个注意力头，每头输出为序列长度乘64；拼接后恢复512维，再通过输出投影混合不同头的信息。" style="width:100%;height:auto;" loading="lazy"/> <figcaption>图 2：各头读取完整输入，独立计算注意力；合并后再混合不同头的输出分量。</figcaption> </figure> <p>设八个头的输出为 \(H_{1},\ldots,H_{8}\)。沿列方向拼接，再经过注意力输出投影：</p> \[O=\operatorname{Concat}(H_{1},\ldots,H_{8})W_{O}, \qquad W_{O}\in\mathbb{R}^{512\times512}.\] <p>拼接结果和 \(O\) 的尺寸都是 \(S\times512\)。\(W_{O}\) 允许不同头的输出相互混合。标准多头注意力会使用全部头的输出，各头在共同损失下训练，其功能可能互补，也可能重叠。</p> <p>同一层各头的注意力计算不等待其他头的结果；下一层则可以基于上一层已经整合的信息继续处理。<strong>头数增加并行读取方式，深度增加有先后依赖的处理轮次。</strong> 若隐藏维度固定，更多头通常意味着每头更窄，并不自动带来更强能力。</p> <h2 id="7-残差连接与-mlp">7. 残差连接与 MLP</h2> <p>注意力输出 \(O\) 与输入 \(X\) 尺寸相同。原论文先将它们相加，再做层归一化：</p> \[Z=\operatorname{LayerNorm}(X+O).\] <p>在归一化之前，残差连接 \(X+O\) 提供了一条系数为 1 的直接通路。当更新量 \(O\) 接近零时，相加结果接近 \(X\)。这并不表示经过 LayerNorm 后的整个模块严格等于恒等变换。</p> <p>用标量形式说明残差本身的梯度：</p> \[z=x+f(x),\qquad \frac{dz}{dx}=1+f'(x).\] <p>导数中的 1 对应直接通路。这有助于梯度传播，但完整网络还包含归一化等操作，不能据此保证梯度始终稳定。其他架构也可以采用缩放或门控残差。</p> <p>接下来，MLP 对每一行使用同一套参数：</p> \[U=\operatorname{ReLU}(ZW_{1}+b_{1}),\qquad F=UW_{2}+b_{2}.\] <table> <thead> <tr> <th>参数或结果</th> <th>尺寸</th> </tr> </thead> <tbody> <tr> <td>\(Z\)</td> <td>\(S\times512\)</td> </tr> <tr> <td>\(W_{1}\)</td> <td>\(512\times2048\)</td> </tr> <tr> <td>\(b_{1}\)</td> <td>\(2048\)，广播到每一行</td> </tr> <tr> <td>\(U\)</td> <td>\(S\times2048\)</td> </tr> <tr> <td>\(W_{2}\)</td> <td>\(2048\times512\)</td> </tr> <tr> <td>\(b_{2}\)</td> <td>\(512\)，广播到每一行</td> </tr> <tr> <td>\(F\)</td> <td>\(S\times512\)</td> </tr> </tbody> </table> <p>偏置的大小不随输入长度变化。ReLU 提供非线性，否则两个线性变换可以合成一个。MLP 在每个位置内部加工已有信息，本身不跨位置通信。然后再做一次残差与归一化：</p> \[X_{\mathrm{new}}=\operatorname{LayerNorm}(Z+F).\] <p>这就是原论文 encoder block 的主干。原论文在残差相加前等位置还使用 dropout；为简化表达，这里没有把它写进每条公式。</p> <h2 id="8-原始-transformerencoder-与-decoder">8. 原始 Transformer：encoder 与 decoder</h2> <p>前面介绍的注意力和 MLP 构成了原始 Transformer 的基本模块。完整模型由 encoder 和 decoder 两部分组成，各有六个串联的 block，同一部分内各层结构相同、参数独立。下图根据原论文 Figure 1 重绘，数据从上向下流动。[1]</p> <p>本文采用的 base 配置为 512 维、8 个头。GPT-2 small 则是 768 维、12 个头和 12 个 decoder-only block；它的 MLP 激活、位置表示和归一化顺序也不同。[2]</p> <figure> <img src="/assets/img/posts/transformer/architecture.svg" alt="原始 Transformer 的 encoder 和 decoder 各六层。Encoder 最终输出供每层 decoder 的交叉注意力读取。各子层均有残差连接和层归一化，decoder 最终输出再投影到词表。" style="width:100%;height:auto;" loading="lazy"/> <figcaption>图 3：原始架构的简化示意。虚线为残差；每个大框串联六次，各层参数独立。</figcaption> </figure> <p>为区分两条序列，记原文长度为 \(S_{\mathrm{src}}\)，当前译文前缀长度为 \(S_{\mathrm{tgt}}\)。原文经过六层 encoder，得到：</p> \[C\in\mathbb{R}^{S_{\mathrm{src}}\times512}.\] <p>它保留原文的各个位置，并没有把整句话压成单个向量。Encoder 自注意力允许读取原文的前后位置，但不读取补齐长度用的 padding。</p> <p>Decoder 的每个 block 则依次包含因果自注意力、交叉注意力和 MLP，每个子层后都有残差相加与归一化。三种注意力的区别如下：</p> <table> <thead> <tr> <th>模块</th> <th>\(Q\) 的输入来源</th> <th>\(K,V\) 的输入来源</th> <th>可读取范围</th> </tr> </thead> <tbody> <tr> <td>Encoder 自注意力</td> <td>当前原文状态</td> <td>当前原文状态</td> <td>完整原文</td> </tr> <tr> <td>Decoder 自注意力</td> <td>当前译文状态</td> <td>当前译文状态</td> <td>自己及之前的译文位置</td> </tr> <tr> <td>Decoder 交叉注意力</td> <td>decoder 自注意力及归一化后的状态</td> <td>encoder 最终输出 \(C\)</td> <td>完整原文</td> </tr> </tbody> </table> <p>具体地，记 decoder 经过因果自注意力及归一化后的状态为 \(B\)。一个交叉注意力头计算：</p> \[Q=BW_{Q},\qquad K=CW_{K},\qquad V=CW_{V}.\] <p>这里的投影参数属于该交叉注意力模块，与其他注意力模块的参数独立。各矩阵尺寸为：</p> \[Q:S_{\mathrm{tgt}}\times64,\qquad K,V:S_{\mathrm{src}}\times64.\] <p>因此：</p> \[A:S_{\mathrm{tgt}}\times S_{\mathrm{src}}, \qquad AV:S_{\mathrm{tgt}}\times64.\] <p>注意力矩阵不必是方阵。输出的每一行仍对应一个译文位置，却已经读取了原文信息。多头合并和输出投影后，残差加回的是 \(B\)，不是 \(C\)。</p> <p>每个 decoder block 都读取同一个最终 \(C\)，但使用各自的投影参数；不是第几层 decoder 对接第几层 encoder。</p> <h2 id="9-decoder-only-模型如何读取输入">9. Decoder-only 模型如何读取输入</h2> <p>Decoder-only 模型可以把任务说明、原文和译文前缀按顺序放进同一条序列。生成译文时，原文已经位于前面，因果自注意力可以读取它。</p> <p><strong>Self-attention 的 self 指同一组序列表示，不是一个 token 只能读取自己。</strong> 同一条序列中不同位置之间的读取仍是自注意力；当查询来自 decoder，而键和值来自独立的 encoder 表示时，才是上一节介绍的交叉注意力。</p> <p>因此，GPT-2 这类 decoder-only 模型不需要通过额外的交叉注意力接收文字输入。能否做好翻译还取决于训练数据与训练方式，架构提供可行的信息通道，并不自动赋予能力。</p> <h2 id="10-推理预测下一个-token">10. 推理：预测下一个 token</h2> <p>回到原始 encoder–decoder 模型。Decoder 最后一层的输出经词表投影和 softmax，为每个位置产生一个“下一个 token”的分布：</p> \[\begin{aligned} D&amp;\in\mathbb{R}^{S_{\mathrm{tgt}}\times512},\\ W_{\mathrm{vocab}}&amp;\in\mathbb{R}^{512\times N_{\mathrm{vocab}}},\\ \Pi&amp;=\operatorname{softmax}_{\mathrm{row}}(DW_{\mathrm{vocab}}) \in\mathbb{R}^{S_{\mathrm{tgt}}\times N_{\mathrm{vocab}}}. \end{aligned}\] <p>\(D\) 是 decoder 的最终状态，\(W_{\mathrm{vocab}}\) 是词表输出投影，\(\Pi\) 是预测概率矩阵。词表输出投影与注意力内部的 \(W_{O}\) 作用不同；前者输出词表分数，后者混合多头结果。原论文还让词表输出投影与 embedding 共享权重。</p> <p>推理从起始标记开始。给定当前前缀，尚未生成的下一个 token 紧跟最后一个位置，因此使用 \(\Pi\) 的最后一行。前面的状态仍参与注意力计算，为最后一个位置提供信息。</p> <p>选出新 token 后加入前缀，继续预测，直到结束标记或长度限制。最直接的实现每次重算完整前缀；实际推理通常缓存各层已有位置的 \(K\)、\(V\)。因果遮挡保证旧位置不依赖后来追加的 token，因此其缓存可以复用。原文固定时，encoder 输出 \(C\) 也可复用。</p> <p>选择 token 可以采用贪心、采样或束搜索。原论文的翻译实验采用束搜索，保留多个候选前缀；上述说明只追踪其中一条。</p> <h2 id="11-训练输入与目标错开一位">11. 训练：输入与目标错开一位</h2> <p>训练时有完整的原文和正确译文。设正确译文包含 \(T\) 个 token，依次记为 \(y_{1},\ldots,y_{T}\)。这里的下标只表示译文位置。</p> <p>Decoder 输入与监督目标排列如下：</p> <table> <thead> <tr> <th>位置</th> <th>Decoder 输入</th> <th>该位置的监督目标</th> </tr> </thead> <tbody> <tr> <td>第 0 位</td> <td>起始标记</td> <td>\(y_{1}\)</td> </tr> <tr> <td>第 1 位</td> <td>\(y_{1}\)</td> <td>\(y_{2}\)</td> </tr> <tr> <td>…</td> <td>…</td> <td>…</td> </tr> <tr> <td>第 \(T\) 位</td> <td>\(y_{T}\)</td> <td>结束标记</td> </tr> </tbody> </table> <p>这就是右移一位的输入。训练时 decoder 使用正确前缀，即使模型在前一个位置预测错了，也不在这次计算中用错误预测替换输入。这叫 teacher forcing。</p> <p>因果遮挡让每个位置不能读取后面的答案。因此，所有位置的预测可以在一次前向计算中并行完成。得到概率矩阵后，不必先抽样生成 token，直接计算监督目标的损失。</p> <p>设某个位置的正确目标为 \(y\)，模型分配给它的概率为 \(p(y)\)。普通交叉熵损失为：</p> \[\ell=-\ln p(y).\] <p>正确目标的概率越高，损失越小。即使它还不是概率最高的候选，提高它的概率也会降低损失。原论文还使用了 label smoothing，对监督分布进行平滑处理。</p> <p>把有效目标位置的损失汇总为 \(L\)，通过链式法则反向传播，再更新参数。对任意一个可训练参数 \(w\)，最简单的梯度下降更新为：</p> \[w_{\mathrm{new}}=w_{\mathrm{old}}-\eta\frac{\partial L}{\partial w}.\] <p>\(\eta\) 是学习率。一次前向、反向计算期间参数不变，算完梯度后才更新。Embedding、注意力投影、MLP、归一化中的可训练参数和输出层都可参与；\(Q\)、\(K\)、\(V\) 是中间结果，不作为独立参数更新。</p> <p>训练优化的是所有样本上的共同目标。Encoder 不需要单独的目标表示 \(C\)：最终预测损失的梯度会通过交叉注意力传回 encoder。各个注意力头也不需要预先指定任务。</p> <p>不同样本的梯度可能相反。Batch 平均能减小单条样本带来的波动，并适合硬件并行；学习率决定每步走多远。两者是不同的控制量，不存在“越小越好”的普遍规则。原论文实际使用 Adam，而不是这里为解释方便采用的普通梯度下降。</p> <p>训练时正确前缀已知，能并行计算多个位置；生成时未来 token 未知，需要逐个推进。推理时一次错误会改变后续前缀，因此，除了验证损失，还需要让模型生成完整译文来检查效果。</p> <h2 id="12-模型容量与训练数据">12. 模型容量与训练数据</h2> <p>参数多意味着可用容量更大，不保证已经学到了更多。数据不足时，大模型可能记住训练样本而泛化不佳；模型过小，也可能无法表达所需规律。模型大小、数据量和计算量要一起考虑，不能只数参数。</p> <p>本文介绍基本计算过程。KV cache 的实现、长序列的计算成本、分布式训练和数值精度等工程问题，留待后续讨论。</p> <h2 id="参考资料与说明">参考资料与说明</h2> <ol> <li>Vaswani et al., <a href="https://arxiv.org/abs/1706.03762">Attention Is All You Need</a>, 2017。架构、尺寸和训练细节参照原论文；图 3 为重新绘制的简化示意。</li> <li>OpenAI, <a href="https://github.com/openai/gpt-2/blob/master/src/model.py">GPT-2 官方实现：model.py</a>。用于核对 GPT-2 与原始 Transformer 的区别。</li> <li>3Blue1Brown, <a href="https://www.3blue1brown.com/lessons/gpt/">Transformers, the tech behind LLMs</a>。Transformer 的可视化介绍。</li> </ol> <p>本文根据与 AI 助手的学习问答整理。</p>]]></content><author><name></name></author><category term="Machine Learning"/><category term="Transformer"/><category term="Learning Notes"/><summary type="html"><![CDATA[介绍 Transformer 的基本计算过程：token embedding、注意力、多头结构、encoder–decoder，以及训练与推理的区别。]]></summary></entry><entry><title type="html">The Fourth Generation: Where the X-Ray Light Sources Stand in 2026</title><link href="https://chaojiezhang.me/posts/2026/07/fourth-generation-light-sources/" rel="alternate" type="text/html" title="The Fourth Generation: Where the X-Ray Light Sources Stand in 2026"/><published>2026-07-16T00:00:00+00:00</published><updated>2026-07-16T00:00:00+00:00</updated><id>https://chaojiezhang.me/posts/2026/07/fourth-generation-light-sources</id><content type="html" xml:base="https://chaojiezhang.me/posts/2026/07/fourth-generation-light-sources/"><![CDATA[<p>Most of what we know about the atomic structure of matter — the shape of a protein, the way a battery electrode degrades, how a catalyst grips a molecule — we learned by shining X-rays on it. The X-rays come from a few dozen machines around the world, and those machines are in the middle of the largest transition in their history.</p> <p>Seven of them now run on a design that did not exist in usable form fifteen years ago, and four of the seven switched on in the last two years. Several more of the best-known facilities are shutting down, having their interiors torn out, and being rebuilt inside their own tunnels before the decade is out — the first of them went dark last summer. On the free-electron laser side, one machine has just reached a repetition rate nearly a thousand times higher than the previous generation, and another has, for the first time, made an X-ray laser lase inside a resonant cavity.</p> <p>Everything below is based on publicly available information: commissioning papers, design reports, budget documents, and laboratory announcements. At the end I offer an assessment of where plasma acceleration fits. That part is opinion, and since plasma acceleration is what I work on, I have tried to rely on other people’s evaluations rather than my own.</p> <hr/> <h3 id="what-these-machines-are-and-what-brightness-means">What These Machines Are, and What “Brightness” Means</h3> <p>A synchrotron light source is a ring, a few hundred meters to a couple of kilometers around, in which electrons circulate at very nearly the speed of light. Magnets bend them, and any charged particle forced to turn radiates. At these energies the radiation comes out as X-rays, in a narrow forward cone, and it is piped down “beamlines” to experiments arranged around the ring like spokes. A large facility runs about thirty to seventy of them at once, which is the ring’s great advantage: many experiments, all day, every day.</p> <p>The figure of merit is <strong>brightness</strong>: how many photons per second, emitted from how small a spot, into how narrow a cone, within how narrow a band of color. All four matter. You need many photons because most of them miss. You need a small spot and a narrow cone because that is what lets you focus the beam onto something tiny, and because it is what makes the light <em>coherent</em> — able to interfere with itself, which is what turns a shadow into an image with atomic detail.</p> <p>Spot size multiplied by angular spread is a single quantity called <strong>emittance</strong>, and it belongs to the electron beam, not the light. The X-rays inherit it. So the entire enterprise of building a better light source reduces, to first approximation, to one question: how do you make the circulating electron beam smaller and less divergent?</p> <p>There is a floor. Light itself has a minimum emittance set by diffraction — roughly the wavelength divided by 4π. A beam at that floor is called <strong>diffraction-limited</strong>, and its light is fully coherent. Getting close to it is the whole point of the current generation, which is why these machines are called diffraction-limited storage rings.</p> <hr/> <h3 id="one-equation-explains-most-of-the-landscape">One Equation Explains Most of the Landscape</h3> <p>Here is the scaling that organizes everything that follows. A ring’s natural emittance goes as</p> \[\varepsilon_x \propto \gamma^2 \theta^3\] <p>where \(\gamma\) is the electron energy in units of its rest mass and \(\theta\) is the angle through which each individual bending magnet turns the beam.</p> <p>The important term is the cube. Emittance grows because bending spreads the beam: each time an electron radiates a photon inside a bending magnet it recoils slightly, and the recoil puts it onto a slightly different orbit. The larger the angle a single magnet bends through, the larger that orbit error becomes, and the accumulated errors spread the beam sideways. Make each individual bend angle small and the growth drops — very sharply, because of that exponent.</p> <p>So: instead of a few magnets each bending through a large angle, use many magnets each bending through a small one. Split each bend into seven pieces and each piece turns the beam through one-seventh the angle, which reduces the emittance by a factor of \(7^3 = 343\). That idea, proposed in 1993 and called the <strong>multi-bend achromat</strong>, is the design every fourth-generation ring is built on. The price is that everything else gets harder — the beam must be focused far more strongly, which means smaller magnets packed tighter, vacuum chambers narrowed from tens of millimeters to twenty or less, and an unforgiving sensitivity to alignment.</p> <p>The cube has a second consequence, less often stated. The magnets must collectively bend the beam through a full circle, so \(\theta\) is fixed by how many magnets you can fit, which is fixed by circumference. <strong>At a given beam energy, circumference is destiny:</strong></p> <table> <thead> <tr> <th>Ring</th> <th>Energy</th> <th>Circumference</th> <th>Natural emittance</th> </tr> </thead> <tbody> <tr> <td>ESRF-EBS (France)</td> <td>6 GeV</td> <td>844 m</td> <td>~135 pm·rad</td> </tr> <tr> <td>APS-U (USA)</td> <td>6 GeV</td> <td>1104 m</td> <td>42 pm design, 33 pm measured</td> </tr> <tr> <td>HEPS (China)</td> <td>6 GeV</td> <td>1360 m</td> <td>34.2 pm</td> </tr> <tr> <td>PETRA-IV (Germany)</td> <td>6 GeV</td> <td>2304 m</td> <td>20 pm target</td> </tr> </tbody> </table> <p>Four machines at the same energy, ordered by size. (A picometer-radian is \(10^{-12}\) meter-radians; for scale, the diffraction limit at 10 keV — a typical hard X-ray — is about 10 pm·rad.) PETRA-IV expects to lead this list not because of a physics insight but because it occupies a 2.3-kilometer tunnel built decades ago for a particle collider.</p> <p>That is worth holding onto. Which facility has the world’s brightest hard X-rays is a question answered substantially by civil engineering choices made in a previous era, and the answer will change when someone else finishes a bigger tunnel.</p> <hr/> <h3 id="the-machines-that-are-running">The Machines That Are Running</h3> <table> <thead> <tr> <th>Facility</th> <th>Location</th> <th>Energy</th> <th>Circumference</th> <th>Emittance</th> <th>Users since</th> </tr> </thead> <tbody> <tr> <td>MAX IV</td> <td>Lund, Sweden</td> <td>3 GeV</td> <td>528 m</td> <td>~330 pm</td> <td>2016</td> </tr> <tr> <td>ESRF-EBS</td> <td>Grenoble, France</td> <td>6 GeV</td> <td>844 m</td> <td>~135 pm</td> <td>2020</td> </tr> <tr> <td>SIRIUS</td> <td>Campinas, Brazil</td> <td>3 GeV</td> <td>518 m</td> <td>~250 pm</td> <td>2021</td> </tr> <tr> <td>NanoTerasu</td> <td>Sendai, Japan</td> <td>3 GeV</td> <td>349 m</td> <td>1.14 nm</td> <td>2024</td> </tr> <tr> <td>APS-U</td> <td>Argonne, USA</td> <td>6 GeV</td> <td>1104 m</td> <td>33 pm</td> <td>2024</td> </tr> <tr> <td>SLS 2.0</td> <td>Villigen, Switzerland</td> <td>2.7 GeV</td> <td>288 m</td> <td>~135 pm</td> <td>2025</td> </tr> <tr> <td>HEPS</td> <td>Beijing, China</td> <td>6 GeV</td> <td>1360 m</td> <td>34.2 pm</td> <td>2026</td> </tr> </tbody> </table> <p><strong>MAX IV</strong> got there first, in 2016, and everything since is downstream of what it proved: that a multi-bend achromat ring can actually be built and operated. Its injector is worth a note of its own. A single 3 GeV linear accelerator serves as continuous top-up injector to two storage rings and, separately, as the driver for a short-pulse facility — with two different electron sources feeding it, one optimized for filling the rings and one for the short-pulse beam.</p> <p><strong>ESRF-EBS</strong> was the first machine to rebuild an existing hard X-ray ring inside its own tunnel, and the first to use the “hybrid” variant of the multi-bend achromat that most later designs adopted. The “hybrid” part is a deliberate compromise: rather than keeping the coupling between an electron’s energy and its position — the dispersion — small everywhere, the lattice lets it rise in two controlled bumps within each cell and puts the strong correcting sextupole magnets there, where a little dispersion makes them far more effective. That compromise is what made such a tight lattice practical to operate, and it is the reason the other upgrades believe they can work.</p> <p><strong>APS-U</strong> is the largest recent success in the U.S. program. The ring went dark in April 2023, restarted in 2024, and was verified that August as the brightest storage ring in the world. It received final project approval in January 2026 at a total cost of $815 million, delivered on budget and ahead of schedule, with nine new beamlines. The <a href="https://www.energy.gov/documents/fy-2027-basic-energy-sciences-budget-request">FY2027 Department of Energy budget justification</a> reports a measured horizontal emittance of 33 pm·rad; the design goal was 42.</p> <p><strong>NanoTerasu</strong> is the instructive outlier. It is a green-field machine in Sendai, only 349 meters around, with an emittance of 1.14 nanometer-radians — thirty times larger than APS-U. That is not a shortfall; it is a choice. NanoTerasu is aimed at soft and tender X-rays around 1–3 keV, where the diffraction limit is much less demanding, and it is positioned explicitly as the complement to SPring-8’s hard X-rays. It reached stable operation at 400 milliamps in 2025, doubling its output. It has also signalled plans to extend its injector linac and add a soft X-ray free-electron laser later.</p> <p><strong>SLS 2.0</strong> achieved the fastest turnaround anyone has managed: ring replaced inside the existing building between October 2023 and December 2024, <a href="https://www.psi.ch/en/news/psi-stories/sls-2-0-how-to-start-up-a-particle-accelerator">beam back in January 2025</a>, first experiments in August 2025, regular operation in 2026. The rebuild raised the beam energy from 2.4 to 2.7 GeV and improved the emittance by a factor of 40; brightness rose by well over an order of magnitude from the ring alone, and by up to three at beamlines whose undulators were replaced too. Its bending magnets are permanent magnets, not electromagnets, which means the ring’s energy is now fixed for good. That is a deliberate trade of flexibility for stability and electricity, and more machines will make it.</p> <p><strong>HEPS</strong> is the first high-energy fourth-generation ring built from scratch rather than retrofitted, with 34.2 pm·rad emittance; it has been in trial operation since December 2025 and begins regular user operation this year.</p> <hr/> <h3 id="the-machines-being-built-rebuilt-or-waiting">The Machines Being Built, Rebuilt, or Waiting</h3> <table> <thead> <tr> <th>Project</th> <th>Location</th> <th>Energy</th> <th>Circumference</th> <th>Injector</th> <th>Status</th> </tr> </thead> <tbody> <tr> <td>ALS-U</td> <td>Berkeley, USA</td> <td>2 GeV</td> <td>197 m</td> <td>Linac + booster, plus a new accumulator ring</td> <td>Under construction; dark period from Oct 2027 at the earliest (~22 months); completion ~FY2030</td> </tr> <tr> <td>PETRA-IV</td> <td>Hamburg, Germany</td> <td>6 GeV</td> <td>2304 m</td> <td>Upgraded linac + new booster</td> <td><a href="https://www.wissenschaftsrat.de/download/2026/3152-26.pdf">Funding recommended Mar 2026</a>; final approval anticipated within 2026</td> </tr> <tr> <td>Diamond-II</td> <td>Oxfordshire, UK</td> <td>3.5 GeV</td> <td>561 m</td> <td>Existing booster, upgraded to 3.5 GeV</td> <td>Approved; dark period from Dec 2027 (~18 months); completion Mar 2030</td> </tr> <tr> <td>SOLEIL-II</td> <td>Saint-Aubin, France</td> <td>2.75 GeV</td> <td>354 m</td> <td>Upgraded linac + rebuilt booster</td> <td>Under construction</td> </tr> <tr> <td>Elettra 2.0</td> <td>Trieste, Italy</td> <td>2.4 GeV</td> <td>259 m</td> <td>Existing linac + full-energy booster</td> <td>Under construction; dark since July 2025; user operation targeted Jan 2027</td> </tr> <tr> <td>ALBA-II</td> <td>Barcelona, Spain</td> <td>3 GeV</td> <td>269 m</td> <td>Existing linac + full-energy booster</td> <td>Design</td> </tr> <tr> <td>HALF</td> <td>Hefei, China</td> <td>2.2 GeV</td> <td>480 m</td> <td>Full-energy linac</td> <td>Under construction; completion Sept 2028</td> </tr> <tr> <td>Korea-4GSR</td> <td>Ochang, South Korea</td> <td>4 GeV</td> <td>800 m</td> <td>Linac + full-energy booster</td> <td>Under construction; completion 2029</td> </tr> <tr> <td>SAPS</td> <td>Guangdong, China</td> <td>3.5 GeV</td> <td>810 m</td> <td>Under study</td> <td>Design</td> </tr> <tr> <td>SPring-8-II</td> <td>Hyōgo, Japan</td> <td>6 GeV</td> <td>1436 m</td> <td>Injection from the SACLA linac</td> <td>Funded; shutdown summer 2027; user operation resumes FY2029</td> </tr> <tr> <td>NSLS-II-U</td> <td>Upton, USA</td> <td>3 GeV</td> <td>792 m</td> <td>—</td> <td>Concept</td> </tr> <tr> <td>SSRL-X</td> <td>Menlo Park, USA</td> <td>—</td> <td>—</td> <td>—</td> <td>Published lattice studies</td> </tr> </tbody> </table> <p>Two things stand out in this table.</p> <p><strong>Most of the European and American entries are rebuilds.</strong> Diamond, SOLEIL, Elettra, ALBA, ALS, PETRA and SPring-8 are each replacing the ring inside the existing tunnel — Elettra went dark in July 2025, the first of the wave, and <a href="https://www.riken.jp/en/news_pubs/research_news/rr/20260317_1/">SPring-8 follows in 2027</a> — and each rebuild takes the beam away from its user community for roughly a year and a half to two years. On present schedules several of those dark periods overlap in the late 2020s, so the facilities that remain in operation through the transition will carry a larger share of the world’s demand for beamtime.</p> <p><strong>Published schedules are snapshots, not commitments.</strong> ALS-U illustrates the point. Construction was approved in November 2022; since then <a href="https://als.lbl.gov/joint-als-als-u-statement-on-dark-time-delay/">its planned shutdown has moved</a> from October 2025 to June 2026 to no earlier than October 2027, and project completion is now expected around FY2030, under a <a href="https://research.lbl.gov/2026/02/18/looking-forward-to-the-upgraded-als/">cost and schedule revised in 2026</a>. The hardware is well along — the accumulator ring’s magnet assemblies are installed and storage-ring magnet production is under way. First-of-a-kind machines are difficult to schedule precisely, and the dates in this article should be read accordingly.</p> <p>The last two rows are earlier-stage proposals. <strong>NSLS-II-U</strong> was assessed as “absolutely central” by <a href="https://science.osti.gov/-/media/bes/besac/pdf/Reports/Report-to-BESAC-on-New-and-Upgraded-National-User-Facilities-2024-05-28Final.pdf">an advisory subcommittee in May 2024</a>: it would be the world’s brightest source between 1 and 10 keV, ten to twenty times brighter than any existing U.S. source, using a magnet design that replaces long electromagnets with strings of permanent magnets and thereby cuts the ring’s electricity consumption substantially. The same report recommended further engineering study, drawing on the teams from other recent upgrades, before the design is frozen. <strong>SSRL-X</strong> exists as <a href="https://arxiv.org/abs/2311.13667">published lattice studies</a>, including one option that would reuse a decommissioned collider tunnel — which is, per the equation above, the PETRA-IV move. Both await a formal project start.</p> <hr/> <h3 id="free-electron-lasers-the-other-half-of-the-field">Free-Electron Lasers: The Other Half of the Field</h3> <p>A storage ring’s electrons radiate independently, like a crowd of people each striking a bell at random. The light adds up in intensity but not in phase.</p> <p>A free-electron laser does something different. Send a very high-quality electron beam down a long straight line of alternating magnets — an <strong>undulator</strong> — and the light the electrons emit acts back on the electrons themselves, sorting them into thin sheets spaced one wavelength apart. Once sorted, they radiate in step, and the power grows exponentially rather than linearly. The result is a pulse perhaps a billion times brighter at its peak than anything a ring can produce, and a few femtoseconds long instead of tens of picoseconds.</p> <p>The trade is arithmetic. A ring serves dozens of beamlines simultaneously. An X-ray FEL serves two or three, and those driven by normal-conducting copper linacs fire only around a hundred times a second. Peak brightness is spectacular; the <em>average</em> number of photons delivered to the world per year has, until recently, been comparable.</p> <table> <thead> <tr> <th>Facility</th> <th>Location</th> <th>Energy</th> <th>Technology</th> <th>Status</th> </tr> </thead> <tbody> <tr> <td>LCLS</td> <td>SLAC, USA</td> <td>up to 15 GeV</td> <td>copper linac, 120 Hz</td> <td>Operating</td> </tr> <tr> <td>LCLS-II</td> <td>SLAC, USA</td> <td>4 GeV</td> <td>superconducting, continuous</td> <td>First light 2023; 93 kHz reached Dec 2025; 1 MHz target</td> </tr> <tr> <td>LCLS-II-HE</td> <td>SLAC, USA</td> <td>8 GeV</td> <td>superconducting</td> <td>Approved Sep 2024, $716M; completion expected FY2028</td> </tr> <tr> <td>European XFEL</td> <td>Hamburg, Germany</td> <td>17.5 GeV</td> <td>superconducting, burst mode</td> <td>Operating; 3 undulator lines</td> </tr> <tr> <td>SACLA</td> <td>Hyōgo, Japan</td> <td>8 GeV</td> <td>copper linac</td> <td>Operating</td> </tr> <tr> <td>SwissFEL</td> <td>Villigen, Switzerland</td> <td>5.8 GeV</td> <td>copper linac</td> <td>Operating</td> </tr> <tr> <td>PAL-XFEL</td> <td>Pohang, Korea</td> <td>10 GeV</td> <td>copper linac</td> <td>Operating</td> </tr> <tr> <td>FLASH / FERMI</td> <td>Hamburg / Trieste</td> <td>soft X-ray</td> <td>superconducting / seeded</td> <td>Operating</td> </tr> <tr> <td>SXFEL</td> <td>Shanghai, China</td> <td>1.5 GeV</td> <td>copper linac</td> <td>Operating</td> </tr> <tr> <td>SHINE</td> <td>Shanghai, China</td> <td>8 GeV</td> <td>superconducting, continuous</td> <td>3.1 km; 0.4–25 keV at 1 MHz; first beam targeted 2026</td> </tr> <tr> <td>DCLS</td> <td>Dalian, China</td> <td>0.3 GeV</td> <td>copper linac, seeded</td> <td>Operating</td> </tr> </tbody> </table> <p>A copper accelerator can only fire in short pulses, because it would melt otherwise. A superconducting one can run continuously, and that lifts the repetition rate from 120 per second to a million. <strong>LCLS-II</strong> <a href="https://www6.slac.stanford.edu/news/2025-12-09-new-world-record-lcls-approaches-100000-pulses-second-path-million">reached 93 kilohertz in December 2025</a> on the way to a megahertz. <strong>LCLS-II-HE</strong> will double its energy to 8 GeV, extending megahertz operation to hard X-rays of 5 to 13 keV; it was <a href="https://www6.slac.stanford.edu/news/2024-09-27-new-upgrade-will-supercharge-atomic-vision-worlds-most-powerful-x-ray-laser">approved in September 2024</a> at $716 million and should finish in FY2028. <strong>SHINE</strong>, in Shanghai, is an 8 GeV continuous-wave machine in 3.1 kilometers of tunnel 29 meters underground, aiming at its first electron beam this year.</p> <p>When both are running there will be two continuous megahertz X-ray lasers in the world, and the average-flux comparison with storage rings will no longer be close.</p> <hr/> <h3 id="an-x-ray-cavity-finally">An X-Ray Cavity, Finally</h3> <p>An ordinary FEL starts from nothing — the electrons’ own random noise, amplified. That works, but it means the output is a different jagged spectrum every shot: broad, spiky, and fluctuating by roughly 100 percent divided by the square root of the number of independent spikes.</p> <p>The fix is the one every laser uses: put the light in a cavity and let it go around many times, so each pass is seeded by the last. For X-rays this is hard, because ordinary mirrors do not reflect them. Diamond crystals do, by Bragg diffraction, but only within an extremely narrow band of colors — roughly one part in \(10^5\), compared with the one part in \(10^3\) that an ordinary FEL produces.</p> <p>That narrowness is what a cavity buys, and it is worth being precise about what the numbers mean. The case for the proposed LCLS-X quotes a 100-to-1000-fold gain in <em>average spectral brightness</em> from cavity-based sources. That is not more photons. An X-ray oscillator’s gain per pass is low — tens of percent — and its peak power is <em>below</em> that of a conventional FEL. The entire factor comes from the denominator: brightness is quoted per 0.1 percent of bandwidth, and the cavity cuts the bandwidth by a hundred to a thousand. What you actually buy is coherence along the pulse and a collapse of the shot-to-shot fluctuation from 100 percent to a few percent. For many experiments the stability is worth more than the brightness.</p> <p>The catch is geometric and unforgiving: the light’s round trip through the cavity must take exactly as long as the gap between electron bunches. That requires bunches arriving at megahertz rates, which requires a superconducting accelerator, which is why this concept sits behind LCLS-II-HE and SHINE in every plan rather than beside them.</p> <p>It is no longer hypothetical. In January 2026, a team reported lasing from a diamond cavity 132.8 meters around at the European XFEL, at 6.952 keV, matched to that machine’s 2.23 MHz bunch spacing, with the light building up bunch by bunch (<a href="https://doi.org/10.1038/s41586-025-10025-x"><em>Nature</em> <strong>650</strong>, 93</a>). The optics were proven at SLAC first: a SLAC–Argonne collaboration stored hard X-ray pulses through dozens of round trips in a 14-meter diamond cavity at LCLS (<a href="https://doi.org/10.1038/s41566-023-01267-0"><em>Nature Photonics</em> <strong>17</strong>, 878</a>, 2023), and the same team’s cavity-FEL project with RIKEN aims to demonstrate two-pass gain at 9.831 keV (<a href="https://doi.org/10.1103/PhysRevAccelBeams.27.110701"><em>Phys. Rev. Accel. Beams</em> <strong>27</strong>, 110701</a>) using a workaround for their copper accelerator: two bunches, 218 nanoseconds apart.</p> <p>In the space of a few years, the X-ray cavity went from a design study to a measurement.</p> <hr/> <h3 id="brightness-is-not-the-only-axis">Brightness Is Not the Only Axis</h3> <p>The usual summary of the storage ring landscape is coverage: one source optimized for soft X-rays, one for the middle of the spectrum, one for hard X-rays.</p> <p>That summary describes one axis — average brightness against photon energy — and, as the table near the top of this article suggests, the leader on that axis will keep changing as larger tunnels are finished.</p> <p>Other axes matter too. Consider time. A storage ring’s bunches are typically tens of picoseconds long, because the bunch length is set by an equilibrium between the radiation the electrons emit and the radio-frequency field that restores their energy. Shorter X-ray pulses can be produced at a ring, but every method has a price: slicing a short piece out of a long bunch with a laser sacrifices most of the flux; filling patterns with isolated or specially prepared bunches sacrifice repetition rate; deflecting-cavity schemes make trade-offs of their own. The window between roughly one hundred femtoseconds and a few picoseconds — faster than a ring bunch, slower than an FEL pulse — is therefore comparatively underserved, and emittance reduction does not change that. To my mind this is a larger capability gap than anything on the photon-energy axis.</p> <p>Several other axes need no tunnel at all. Electron sources are a shared upstream bottleneck for rings and FELs alike — the U.S. accelerator research program emphasizes “significant improvements in very high brightness and high current electron sources.” Detectors and data systems determine what a beamline can actually measure, and advisory reports have repeatedly identified them as a priority alongside the accelerators themselves. And automation is moving quickly: among the achievements highlighted in the FY2027 U.S. budget document is a machine-learning system that tunes an FEL’s electron beam, improving emittance by a factor of two and doing it ten times faster than manual tuning.</p> <hr/> <h3 id="where-plasma-acceleration-fits">Where Plasma Acceleration Fits</h3> <p><strong>The idea.</strong> A conventional accelerator pushes electrons with radio waves inside metal cavities, and it is limited by electrical breakdown: push past roughly 30 to 100 million volts per meter and the metal arcs. A plasma cannot break down, because it is already broken down — it is a gas whose electrons have been stripped off. Drive a laser pulse or a dense electron bunch through it and the plasma electrons are shoved aside and snap back, forming a wave that trails the driver like a boat’s wake. That wave holds electric fields of tens to a hundred <em>billion</em> volts per meter, a thousand times what metal allows. The same energy gain, in a thousandth of the length.</p> <p><strong>What has actually been demonstrated.</strong> Free-electron lasing driven by plasma-accelerated beams has now been shown four ways: <a href="https://doi.org/10.1038/s41586-021-03678-x">laser-driven at 27 nanometers</a> by a group in China in 2021; <a href="https://doi.org/10.1038/s41586-022-04589-1">beam-driven at Frascati</a> in 2022; <a href="https://doi.org/10.1038/s41566-022-01104-w">laser-driven and seeded</a> by a French-German collaboration in 2023; and <a href="https://newscenter.lbl.gov/2025/07/29/researchers-make-key-gains-in-unlocking-the-promise-of-compact-x-ray-free-electron-lasers/">laser-driven at Berkeley</a> in 2025. All four are below 1 GeV, and all four are at ultraviolet or longer wavelengths. None is an X-ray FEL. The gap between what has been demonstrated and what a light source needs is the central fact of this subject, and it is a large gap.</p> <p>The beam physics has moved quickly. <a href="https://doi.org/10.1038/s41467-024-50320-1">Preserving the beam’s emittance</a> through a plasma stage has been demonstrated. So has preserving its energy spread below one percent, and actively compressing it afterward. So has extracting more energy from the wave than the driver put in per unit charge — a ratio called the transformer ratio, whose classical ceiling of 2 for a symmetric driver has been <a href="https://doi.org/10.1103/PhysRevLett.121.064801">exceeded in plasma</a> by shaping the driver’s current profile. And in April 2026 a laser-plasma-driven FEL ran continuously for more than eight hours (<a href="https://doi.org/10.1103/z2d3-bhyt"><em>Phys. Rev. Accel. Beams</em> <strong>29</strong>, 041301</a>).</p> <p>That last result addresses the objection that actually matters. Gradient was never the problem. A user facility is judged on uptime and on mean time between failures; a plasma source that performs beautifully on a good afternoon is not yet a component of one.</p> <p><strong>Where it is being built.</strong> EuPRAXIA is the first plasma accelerator project placed on the <a href="https://roadmap2021.esfri.eu/projects-and-landmarks/browse-the-catalogue/eupraxia/">European roadmap for research infrastructure</a>, and its beam-driven half at Frascati has <a href="https://www.eupraxia-project.eu/major-boost-to-european-plasma-accelerator-facility.html">roughly €120 million committed</a> from the Italian government, the Latium region, and INFN. Its architecture is deliberately modest: a conventional photoinjector and X-band linac deliver up to 500 MeV at up to 100 pulses per second, a plasma stage boosts that to 1 GeV, and the result drives an FEL. First pilot users are targeted for 2028. The second site, laser-driven and aiming at 1 to 5 GeV, has been chosen: <a href="https://www.eli-laser.eu/news/eupraxia-selects-eli-for-second-laser-driven-accelerator-site/">ELI Beamlines</a>, outside Prague.</p> <p>Elsewhere, DESY has published <a href="https://bib-pubdb1.desy.de/record/615183/files/PIP4_CDR.pdf">a conceptual design</a> for a laser-plasma injector that would deliver 6 GeV bunches to top up PETRA-IV, driven by the petawatt upgrade of its KALDERA laser, noting that such an injector could eventually replace the conventional chain — though an injector, whether a full-energy linac or a linac plus booster, accounts for only a modest fraction of a facility’s total cost. KIT is building a storage ring, cSTART, <a href="https://publikationen.bibliothek.kit.edu/1000173412">designed to be fed by a laser-plasma injector</a>. The UK’s CLARA accelerator reached its design energy of 250 MeV in 2025, with plasma-wakefield experiments — a plasma-driven FEL demonstration among the stated aims — in its exploitation program.</p> <p><strong>The official assessment.</strong> The <a href="https://science.osti.gov/-/media/bes/besac/pdf/Reports/Report-to-BESAC-on-New-and-Upgraded-National-User-Facilities-2024-05-28Final.pdf">May 2024 advisory report</a> that evaluated candidate technologies for a future U.S. light source described plasma acceleration as potentially a thousand times more efficient than current linear accelerators, noted that its development is rapid — parenthetically, “including in China” — and concluded that in 10 to 15 years it may be mature enough for such a facility.</p> <p>A quieter signal points the same way from a different direction. The FY2027 U.S. budget document lists, among the year’s achievements in basic energy sciences, a compact free-electron laser result in which a national laboratory and an American company produced high-brightness electron beams from a plasma and achieved nearly a thousandfold amplification through an undulator — and the document itself observes that compact plasma-based FELs offer opportunity for scientific discovery and for industrial applications including microelectronics fabrication. The commercial side has moved accordingly: a Palo Alto company founded in 2021 is developing an accelerator-driven FEL to replace the tin-droplet plasma sources used in extreme-ultraviolet lithography, with one machine intended to feed up to twenty chip-printing tools. It <a href="https://www.xlight.com/news">raised a Series B in July 2025</a>, received a <a href="https://www.nist.gov/news-events/news/2025/12/department-commerce-and-nist-announce-chips-research-and-development-letter">CHIPS research letter of intent</a> from the Commerce Department in December 2025, and plans a prototype in Albany from 2028. I have written about that direction <a href="https://chaojiezhang.me/posts/2025/08/accelerators-moores-law/">before</a>. For high-average-power FELs, the pull may currently be industrial rather than scientific.</p> <p><strong>My reading.</strong> Plasma acceleration will not replace a storage ring. A ring’s value is fifty beamlines running at once at high average flux, and nothing about a plasma stage helps with that. It will not replace a superconducting FEL linac in this decade either; LCLS-II-HE and SHINE will define that frontier and cavity-based sources will extend it.</p> <p>What it offers is narrower and worth stating plainly. It offers energy per meter, which matters only where the alternative is a tunnel that cannot be afforded or sited. It offers a beam whose quality is set by how it is born inside the plasma rather than inherited from whatever drove the wave, which matters where the available driver is not good enough on its own. Those are real. Whether they are worth their cost depends on what the alternative costs, and the alternative — superconducting linacs, better electron guns, better undulators, better detectors — is getting cheaper and more reliable every year.</p> <hr/> <h3 id="outlook">Outlook</h3> <p>The picture in one paragraph. The fourth generation of storage rings has arrived and is spreading, mostly by rebuilding existing machines inside their own tunnels at a cost of roughly two dark years each; the title of brightest hard X-ray source will pass to whoever finishes the largest tunnel, and will keep passing; the era of continuous, megahertz-class X-ray lasers is opening in two places; X-ray cavities stopped being hypothetical in January; and plasma acceleration is a decade or more from a user facility.</p> <p>Those last two clauses are not in tension. A technology fifteen years from usefulness can be entirely worth the intervening fifteen years, and this one has produced a great deal of physics along the way that had nothing to do with light sources.</p> <p>The milestones that would shorten that timescale are well defined: a plasma-driven FEL at soft X-ray rather than ultraviolet wavelengths; a plasma source that meets a facility’s availability requirements over months rather than hours; a high transformer ratio demonstrated at GeV energies rather than at tens of MeV. All three are the subject of active work.</p> <hr/> <h3 id="sources">Sources</h3> <p><em>All numbers and schedules are from the public record: commissioning papers, design reports, advisory-committee and budget documents, and laboratory announcements. Corrections are welcome. The load-bearing sources:</em></p> <p><strong>Reports and budget documents</strong></p> <ul> <li><a href="https://science.osti.gov/-/media/bes/besac/pdf/Reports/Report-to-BESAC-on-New-and-Upgraded-National-User-Facilities-2024-05-28Final.pdf">Report to BESAC on New and Upgraded National User Facilities</a> — Basic Energy Sciences Advisory Committee subcommittee, May 2024. The NSLS-II-U assessment and the plasma-technology outlook quoted above.</li> <li><a href="https://www.energy.gov/documents/fy-2027-basic-energy-sciences-budget-request">FY 2027 Congressional Budget Justification, DOE Basic Energy Sciences</a> — APS-U completion and measured emittance, LCLS-II-HE and ALS-U status, the machine-learning tuning result, and the compact plasma-FEL highlight.</li> <li><a href="https://www.wissenschaftsrat.de/download/2026/3152-26.pdf">Wissenschaftsrat statement on PETRA IV</a> — German Science Council, March 2026.</li> </ul> <p><strong>Storage rings</strong></p> <ul> <li>MAX IV — <a href="https://www.maxiv.lu.se/beamlines-accelerators/accelerators/guns-and-linear-accelerator">the linac and its two guns</a> (MAX IV Laboratory).</li> <li>SLS 2.0 — <a href="https://doi.org/10.1080/08940886.2024.2312059">Willmott, <em>Synchrotron Radiation News</em> (2024)</a>; <a href="https://www.psi.ch/en/news/psi-stories/sls-2-0-how-to-start-up-a-particle-accelerator">PSI, “SLS 2.0: how to start up a particle accelerator” (2025)</a>.</li> <li>ALS-U — <a href="https://als.lbl.gov/joint-als-als-u-statement-on-dark-time-delay/">joint statement on the dark-time delay (2023)</a>; <a href="https://research.lbl.gov/2026/02/18/looking-forward-to-the-upgraded-als/">“Looking forward to the upgraded ALS” (February 2026)</a>.</li> <li>NanoTerasu — <a href="https://doi.org/10.1103/PhysRevAccelBeams.28.020701">accelerator commissioning, <em>Phys. Rev. Accel. Beams</em> <strong>28</strong>, 020701 (2025)</a>.</li> <li>HEPS — <a href="https://doi.org/10.1007/s43673-026-00184-y">design and commissioning overview, <em>AAPPS Bulletin</em> (2026)</a>.</li> <li>SPring-8-II — <a href="https://www.spring8.or.jp/pdf/en/ann_rep/24/SPring-8-II.pdf">FY2024 annual report chapter (PDF)</a>; <a href="https://www.riken.jp/en/news_pubs/research_news/rr/20260317_1/">RIKEN research news, March 2026</a>.</li> <li>Diamond-II — <a href="https://indico.global/event/5643/">project overview talk</a>.</li> <li>SOLEIL II — <a href="https://meow.elettra.eu/81/pdf/MOPB074.pdf">“Entrance in the Construction Phase” (IPAC proceedings)</a>.</li> <li>SSRL-X — <a href="https://arxiv.org/abs/2311.13667">lattice-options study, arXiv:2311.13667</a>.</li> </ul> <p><strong>Free-electron lasers and cavities</strong></p> <ul> <li>LCLS-II at 93 kHz — <a href="https://www6.slac.stanford.edu/news/2025-12-09-new-world-record-lcls-approaches-100000-pulses-second-path-million">SLAC news, December 9, 2025</a>.</li> <li>LCLS-II-HE approval — <a href="https://www6.slac.stanford.edu/news/2024-09-27-new-upgrade-will-supercharge-atomic-vision-worlds-most-powerful-x-ray-laser">SLAC news, September 27, 2024</a>.</li> <li>SHINE — <a href="https://www.sixthtone.com/news/1018387">Sixth Tone</a>.</li> <li>Cavity lasing at the European XFEL — <a href="https://doi.org/10.1038/s41586-025-10025-x"><em>Nature</em> <strong>650</strong>, 93 (2026)</a>.</li> <li>X-ray cavity storage at LCLS — <a href="https://doi.org/10.1038/s41566-023-01267-0">Margraf et al., <em>Nature Photonics</em> <strong>17</strong>, 878 (2023)</a>.</li> <li>The cavity-based FEL project — <a href="https://doi.org/10.1103/PhysRevAccelBeams.27.110701"><em>Phys. Rev. Accel. Beams</em> <strong>27</strong>, 110701 (2024)</a>.</li> <li>X-ray oscillator output properties — <a href="https://arxiv.org/abs/1903.09317">arXiv:1903.09317</a>.</li> </ul> <p><strong>Plasma acceleration</strong></p> <ul> <li>Wang et al., <a href="https://doi.org/10.1038/s41586-021-03678-x"><em>Nature</em> <strong>595</strong>, 516 (2021)</a> — laser-driven lasing at 27 nm.</li> <li>Pompili et al., <a href="https://doi.org/10.1038/s41586-022-04589-1"><em>Nature</em> <strong>605</strong>, 659 (2022)</a> — beam-driven plasma FEL.</li> <li>Labat et al., <a href="https://doi.org/10.1038/s41566-022-01104-w"><em>Nature Photonics</em> <strong>17</strong>, 150 (2023)</a> — seeded laser-plasma FEL.</li> <li>BELLA lasing — <a href="https://newscenter.lbl.gov/2025/07/29/researchers-make-key-gains-in-unlocking-the-promise-of-compact-x-ray-free-electron-lasers/">Berkeley Lab news, July 2025</a>.</li> <li>Eight hours of continuous operation — <a href="https://doi.org/10.1103/z2d3-bhyt"><em>Phys. Rev. Accel. Beams</em> <strong>29</strong>, 041301 (2026)</a>.</li> <li>Emittance preservation — <a href="https://doi.org/10.1038/s41467-024-50320-1"><em>Nature Communications</em> <strong>15</strong> (2024)</a>.</li> <li>Transformer ratio above 2 — <a href="https://doi.org/10.1103/PhysRevLett.121.064801">Loisch et al., <em>Phys. Rev. Lett.</em> <strong>121</strong>, 064801 (2018)</a>.</li> <li>EuPRAXIA — <a href="https://roadmap2021.esfri.eu/projects-and-landmarks/browse-the-catalogue/eupraxia/">ESFRI Roadmap 2021</a>; <a href="https://www.eupraxia-project.eu/major-boost-to-european-plasma-accelerator-facility.html">funding announcement</a>; <a href="https://www.eli-laser.eu/news/eupraxia-selects-eli-for-second-laser-driven-accelerator-site/">ELI Beamlines chosen as second site</a>.</li> <li>Laser-plasma injector for PETRA IV — <a href="https://bib-pubdb1.desy.de/record/615183/files/PIP4_CDR.pdf">conceptual design report (DESY)</a>.</li> <li>cSTART — <a href="https://publikationen.bibliothek.kit.edu/1000173412">laser-plasma injector for an electron storage ring (KIT)</a>.</li> <li>xLight — <a href="https://www.xlight.com/news">company news</a>; <a href="https://www.nist.gov/news-events/news/2025/12/department-commerce-and-nist-announce-chips-research-and-development-letter">NIST CHIPS letter of intent, December 2025</a>.</li> </ul>]]></content><author><name></name></author><category term="Particle Accelerators"/><category term="Synchrotron Light Sources"/><category term="Free-Electron Laser"/><category term="X-ray Laser"/><category term="Laser-Plasma Acceleration"/><category term="Advanced Accelerator Concepts"/><summary type="html"><![CDATA[A survey of the synchrotrons and free-electron lasers that produce the world’s brightest X-rays — what is running, what is being rebuilt, what is only proposed — and an assessment of where plasma acceleration fits.]]></summary></entry><entry><title type="html">Laser-Driven Proton Therapy: Strong Physics, Hard Road to the Clinic</title><link href="https://chaojiezhang.me/posts/2026/07/laser-proton-therapy/" rel="alternate" type="text/html" title="Laser-Driven Proton Therapy: Strong Physics, Hard Road to the Clinic"/><published>2026-07-11T00:00:00+00:00</published><updated>2026-07-11T00:00:00+00:00</updated><id>https://chaojiezhang.me/posts/2026/07/laser-proton-therapy</id><content type="html" xml:base="https://chaojiezhang.me/posts/2026/07/laser-proton-therapy/"><![CDATA[<h3 id="the-promise-of-a-small-cheap-proton-source">The Promise of a Small, Cheap Proton Source</h3> <p>Proton therapy has a real advantage over ordinary X-ray treatment. A proton beam deposits most of its energy at a chosen depth and then stops — a feature called the Bragg peak — so it can strike a tumor while sparing the healthy tissue before and behind it. The difficulty has always been the machine. Producing protons with enough energy for treatment, up to about 200 MeV to reach roughly 25 cm deep in human body, has traditionally required a cyclotron or synchrotron with hundred-ton magnets, a heavily shielded room, and, until recently, a budget in the hundreds of millions of dollars.</p> <p>About twenty years ago, laser-driven ion acceleration seemed to offer a shortcut. When an intense (petawatt) laser pulse strikes a thin foil, it generates an enormous electric field on the rear side — on the order of a trillion volts per meter, tens of thousands of times stronger than the radio-frequency cavities of a conventional accelerator. In principle, such a field could push protons to treatment energy over several tens of microns rather than many meters. The promise was irresistible: a proton source small enough and cheap enough to bring a scarce treatment to far more patients.</p> <p>It is a beautiful idea, and strong groups have pursued it for close to two decades. This post is the proton companion to my earlier one on <a href="/posts/2025/08/laser-driven-radiotherapy/">VHEE electrons</a>. The physics, in my opinion, has largely delivered — yet the case for laser-driven protons as a hospital product runs into trouble that has little to do with the physics at all. To be clear from the outset, this is not a criticism of the science, which is great and worth doing for its own sake. My concern is narrower, and it is only about one goal: turning the method into a hospital machine.</p> <hr/> <h3 id="the-physics-has-almost-delivered">The Physics Has (Almost) Delivered</h3> <p>The progress here is genuine and deserves acknowledgment. TNSA, the workhorse mechanism, routinely produces protons of several to tens MeV, and the maximum energy has climbed steadily. Recent experiments have reached nearly 150 MeV by chaining two acceleration stages at extreme laser intensity. New facilities are being built specifically to reach 100–200 MeV with usable beams at a selected energy, and the beams themselves keep improving — more reproducible, carrying more charge, with a narrower spread in energy.</p> <p>Many of the problems once thought impossible have been solved. If the only question were whether a laser can make a proton beam good enough for therapy, the answer would be trending toward yes.</p> <p>But that was never the hard question.</p> <hr/> <h3 id="the-gradient-was-never-the-real-problem">The Gradient Was Never the Real Problem</h3> <p>A trillion-volt-per-meter field is spectacular, yet the field strength — the “gradient” — was never what made proton therapy machines large or expensive. A clinical system is far more than an accelerating gap. It is an instrument that must place a planned dose onto a three-dimensional tumor, safely and reproducibly, every day, at a price that competes. The laser’s giant field solves the problem of making the acceleration part compact, but compact acceleration was never the part that made these machines big or costly in the first place.</p> <p>The fair way to judge laser proton therapy, then, is against the tests that any real product must pass. Three stand out. Does ultra-fast, or “FLASH,” delivery give the laser a unique medical advantage? Can its beam actually be shaped onto a whole tumor? And does the complete system cost less than what hospitals can already buy? On all three the answer, I think, is likely no — and the first is the most interesting, because it fails on physics grounds alone.</p> <hr/> <h3 id="does-flash-change-the-picture-a-look-from-basic-physics">Does FLASH Change the Picture? A Look from Basic Physics</h3> <p>The recent excitement around FLASH radiotherapy — the finding that delivering an entire dose in a fraction of a second appears to spare healthy tissue — looks tailor-made for lasers, which excel at delivering a very high dose rate in an extremely short burst. On closer inspection, though, it probably does not help, and the reason is worth working through.</p> <p>The FLASH effect is defined by delivering the full dose, typically at least 4–8 Gy, at an average rate above about 40 Gy/s, with the whole exposure finished in under roughly 200 ms. The benefit appears to level off somewhere around 100–1000 Gy/s. What matters in that definition is an average rate and a total time; nothing in it refers to structure below a microsecond.</p> <p>That points to the real question: can beam timing finer than a microsecond matter at all? There are only three ways the timing of a beam can reach the biology.</p> <ol> <li><strong>Inside a single ion track.</strong> The energy here is deposited in femtoseconds, but this happens for each particle as it passes, regardless of how the overall pulse is shaped. A single track already reaches a local dose rate of about 10¹² Gy/s, whether the beam is continuous or arrives in bunches. The laser’s femtosecond pulse coincides with this timescale but cannot intensify it.</li> <li><strong>Between tracks.</strong> This is the only channel a beam can genuinely influence, and it works through chemistry. The reactive fragments left in water — hydroxyl radicals (·OH) and free electrons among them — live and react over about a nanosecond to a microsecond. A hydroxyl radical travels only some 3.6 nm in a nanosecond, roughly the size of one small cluster of fragments, and these clusters blur together within about a microsecond. The chemistry, in other words, runs on its own clock, from nanoseconds to microseconds. Deliver faster than that and the effect saturates; compressing the pulse from nanoseconds to picoseconds to femtoseconds likely changes nothing.</li> <li><strong>Heating and other bulk effects.</strong> These would require a real surge of energy, but 10 Gy is only 10 J/kg — a temperature rise of about two thousandths of a degree, no matter how quickly it arrives. There is no shock wave and no thermal effect.</li> </ol> <p>The biology, then, has no clock faster than about a nanosecond for a beam to drive. The distinction that matters here is between two things that are easy to conflate. The <em>duration</em> of an individual pulse — femtoseconds versus nanoseconds — is wasted resolution, since nothing in the chemistry responds to it. The <em>dose delivered per pulse</em> is a different variable, and here a laser is genuinely unique: it could in principle place an entire treatment’s worth of dose inside a single radical lifetime, a dose per chemistry window perhaps 10⁵ times what a cyclotron delivers, and no conventional machine can come close.</p> <p>Whether that regime actually does anything is the real open question. If the FLASH effect is governed by transient oxygen depletion, which unfolds over milliseconds, then only the average rate and total time matter, a cyclotron already reaches them, and the laser’s extreme dose-per-pulse buys nothing. If instead it is governed by radical–radical recombination, which depends on the instantaneous density of reactive fragments, then dose-per-pulse could matter and lasers would have a real, though narrow, opening. The evidence so far leans toward saturation once the mean rate clears roughly 40–100 Gy/s — rates conventional machines can meet — but the matter is not settled. My reading is that FLASH is unlikely to be the advantage that rescues laser proton therapy, while conceding it is the one place a laser-specific benefit cannot yet be ruled out.</p> <hr/> <h3 id="delivery-a-pulsed-beam-is-the-wrong-tool-for-scanning">Delivery: A Pulsed Beam Is the Wrong Tool for Scanning</h3> <p>Even setting the biology aside, delivery poses a problem of its own. Modern proton therapy relies on pencil-beam scanning. The beam’s energy sets its depth — the Bragg peak — while magnets sweep it from side to side, and the tumor is built up from perhaps 10–40 energy layers and thousands of individually weighted spots. This spot-by-spot painting is exactly what lets the dose conform to the tumor’s shape, which is the whole point of using protons.</p> <p>Scanning of this kind wants a beam that is nearly continuous, quick to change energy, and fast to steer — which is what a cyclotron provides. A laser source is different in kind: it delivers a large charge in a single short burst, today at only a few shots per second. No one seriously proposes rastering such a source spot by spot; the realistic plans use its broad energy spread and high per-shot charge to fill a volume at once, or capture and transport a large bunch for delivery in a few shots. The catch is that these routes rely on broad-field or “shoot-through” (single-energy) delivery, which gives up some of the conformity that motivated pencil-beam scanning to begin with. In fairness, the FLASH frontier is pushing conventional machines the same way: scanned FLASH is hard even for a cyclotron, because switching energy layer by layer is slow, so transmission delivery is a shared compromise rather than a laser-specific flaw.</p> <p>That fairness cuts only so far. Shared compromise or not, the pulsed source still offers less conformity than modern scanned therapy, and its charge is limited within the therapeutic energy window once the beam has been energy-selected; higher repetition rates would help but remain distant at therapeutic energies. The very quality that makes a laser a superb physics source makes it an awkward treatment source.</p> <hr/> <h3 id="the-cost-problem">The Cost Problem</h3> <p>The original case rested on size and cost, but while the laser community advanced, so did the competition. Commercial single-room cyclotron systems now install for roughly $30–40M, much reduced from the $150M+ of the old multi-room centers. The decisive detail is that the accelerator itself accounts for only a fraction of that — very roughly half. The rest — the shielded room, the gantry, the imaging, the patient couch, the quality-control systems — does not shrink much when a laser replaces the cyclotron. Even a free and perfect laser source would cut the total cost only by around half.</p> <p>And the laser is far from free. A single-shot multi-petawatt system runs on the order of $15–20M, and such costs are falling. A version firing at the repetition rate a clinic needs, with the improved pumping and cooling that implies, is a much larger undertaking, tens of millions and rising. It also brings its own shielding burden — the intense interaction sprays secondary electrons, neutrons, hard X-rays and an electromagnetic pulse — that is arguably worse per useful proton than a cyclotron’s. Proponents point out that a single laser could in principle feed several treatment rooms, spreading its cost; that helps, but the per-room vault, gantry, imaging and target handling remain, and those are where much of the money goes. Add a target “factory” able to deliver fresh, ultrathin foils many times a second, reliably, for years, and the whole-system arithmetic does not favor the laser today.</p> <hr/> <h3 id="outlook">Outlook</h3> <p>Laser-driven ion acceleration will remain one of the most exciting tools in physics — and it will probably not become a competitive proton-therapy product in five to ten years. Someone will very likely treat a patient with a laser-driven proton beam, and that will be a genuine milestone worth celebrating. But demonstrating that something can work and persuading a hospital to buy the next one are entirely different achievements. The machine that treats the first patient will still, I suspect, struggle to win the following purchase against a cyclotron judged on reliability, dose shaping, and cost.</p> <p>I would be glad to be proven wrong, and the routes are clear enough. A demonstrated way to produce a single-energy, high-charge beam — through radiation-pressure acceleration, sometimes called “light sail,” or related schemes — that tames both the energy spread and the charge at a high repetition rate. A sharp fall in the cost of high-power lasers, carried along by the fusion-energy supply chain. Firm evidence that the FLASH benefit truly requires the ultrafast pulse only a laser can provide. Or a treatment paradigm built around the pulsed beam’s own strengths, rather than forced to imitate a cyclotron — volumetric single-shot delivery, or spatially fractionated “minibeam” schemes — that turns the time structure from a liability into a feature. My argument judges the laser against how proton therapy is practiced today; a genuinely new way to deliver dose could rewrite the comparison.</p> <p>The effort belongs where the laser genuinely excels: ultrafast radiation biology and chemistry at dose rates nothing else can reach, compact pulsed sources of neutrons and isotopes, and warm-dense-matter research. These are not consolation prizes, but frontiers where a femtosecond, high-charge burst is exactly the right instrument, and where laser-driven ion beams may prove genuinely transformative. Whatever caution belongs here is only about framing. A great scientific tool need not become a hospital machine to be worth building; it is already unmatched at what it does best.</p>]]></content><author><name></name></author><category term="Laser-Plasma Acceleration"/><category term="Proton Therapy"/><category term="FLASH Therapy"/><category term="Medical Physics"/><category term="TNSA"/><category term="Radiotherapy"/><summary type="html"><![CDATA[A follow-up to my post on VHEE electrons: why laser-driven proton acceleration is great physics but, in my view, unlikely to become a competitive medical product — and where its real strengths are.]]></summary></entry><entry><title type="html">The 10 TeV Horizon: A Perspective on the Future of High Energy Physics</title><link href="https://chaojiezhang.me/posts/2026/01/future-colliders-perspective/" rel="alternate" type="text/html" title="The 10 TeV Horizon: A Perspective on the Future of High Energy Physics"/><published>2026-01-04T00:00:00+00:00</published><updated>2026-01-04T00:00:00+00:00</updated><id>https://chaojiezhang.me/posts/2026/01/future-colliders-perspective</id><content type="html" xml:base="https://chaojiezhang.me/posts/2026/01/future-colliders-perspective/"><![CDATA[<h3 id="the-10-tev-horizon-a-perspective-on-the-future-of-high-energy-physics">The 10 TeV Horizon: A Perspective on the Future of High Energy Physics</h3> <p>As the High-Luminosity LHC era matures, the global high-energy physics (HEP) community is engaging in a deep strategic dialogue about the future. The recent US P5 Report helped crystalize a shared vision: the need to reach the <strong>10 TeV parton center-of-mass energy scale</strong>.</p> <p>This target represents the energy frontier required to explore physics beyond the Standard Model. However, reaching this scale requires navigating a complex landscape of technological readiness, economic constraints, and facility longevity. Coming from the Advanced Accelerator Concepts (AAC) community, I offer a perspective on how these different approaches stack up and where our technologies might best contribute.</p> <hr/> <h3>The Physics Context: Defining the 10 TeV Goal</h3> <p>To evaluate future machines, it helps to look at the “effective” energy available for particle discovery.</p> <ul> <li><strong>Proton Colliders (100 TeV):</strong> Because protons are composite particles, the collision energy is shared among constituent quarks and gluons. A 100 TeV proton collider (like the proposed FCC-hh or SPPC) effectively probes the 10-20 TeV energy scale at the constituent level.</li> <li><strong>Lepton Colliders (10 TeV):</strong> Electrons and muons are fundamental particles, meaning the full beam energy is available for the interaction. Therefore, a 10 TeV lepton collider is generally considered the physics equivalent of a 100 TeV proton machine.</li> </ul> <p><strong>The Challenge:</strong> While 10 TeV is the clear target for the energy frontier, achieving this with leptons presents significantly different challenges than with protons.</p> <hr/> <h3>Strategic Considerations for Linear Colliders</h3> <p>Linear electron-positron colliders (like the ILC or CLIC) have long been studied as precision Higgs factories. However, when viewed through the lens of the 10 TeV long-term goal, they face distinct strategic hurdles compared to their circular counterparts.</p> <p><strong>Infrastructure and Upgradability</strong> A key advantage of circular collider proposals (like FCC-ee) is the “tunnel value.” Excavating a 90-100 km tunnel is a massive investment, but it creates a permanent infrastructure asset. Once the electron machine completes its mission, the same tunnel can host a 100 TeV proton machine (FCC-hh), guaranteeing the facility’s relevance for decades. In contrast, a linear tunnel is optimized for a single pass. Upgrading a linear collider from a Higgs factory (~250 GeV) to the 10 TeV frontier is technically prohibitive due to the immense length required (potentially &gt;100 km) and the physics of beamstrahlung (radiative energy loss), which becomes manageable only at lower energies or with novel particle species.</p> <p><strong>The “Middle Energy” Dilemma</strong> This leaves linear electron colliders in a difficult position: they are excellent precision machines at the 100s GeV scale, but extending them to the 10 TeV energy frontier—where the physics community aims to go next — requires a paradigm shift in technology rather than just scaling up.</p> <p><strong>The “Narrow Window” for Viability</strong> This leaves a very narrow path for linear colliders. To be justifiable, the footprint and cost of a linear Higgs factory must be radically reduced — perhaps to a small fraction (e.g., &lt;20%) of a circular counterpart like FCC-ee or CEPC. Without such a disruptive cost advantage, a linear machine represents a significant strategic risk. As a single-purpose facility, it lacks the “infrastructure insurance” of a circular tunnel; if no new physics is found, the community is left with a dead-end facility rather than a stepping stone to the 10 TeV frontier.</p> <hr/> <h3>The Global Landscape</h3> <p>Different regions are adopting different strategies to address these challenges.</p> <h4>Europe: The Infrastructure Approach (FCC)</h4> <p>CERN’s Future Circular Collider (FCC) proposal prioritizes infrastructure longevity. By starting with an e-e+ machine and planning for a proton machine later, it maximizes the return on the civil engineering investment. This is a robust, albeit expensive, path that relies on proven technology and CERN’s existing institutional strength.</p> <h4>China: The Window of Opportunity (CEPC)</h4> <p>The Circular Electron Positron Collider (CEPC) shares the same logic as the FCC — a circular Higgs factory followed by a super proton collider (SPPC). The primary challenge here is timing. For CEPC to secure its position as the world’s first Higgs factory, it needs to move to construction ahead of the FCC to maximize its strategic value to the global community.</p> <h4>USA: The Technology Leap (Muon Collider)</h4> <p>The US community is exploring a different path. Rather than committing to a massive new tunnel, interest is growing in the <strong>Muon Collider</strong>.</p> <ul> <li><strong>The Logic:</strong> Muons, being heavy leptons, suppress synchrotron radiation. This allows for a <strong>circular</strong> 10 TeV machine that could essentially fit within the existing Fermilab site.</li> <li><strong>The Status:</strong> This is an R&amp;D intensive “moonshot.” It requires mastering muon cooling (compressing the beam phase space). If successful, it offers a compact, cost-effective route to 10 TeV that bypasses the size constraints of proton rings and linear electron accelerators.</li> </ul> <hr/> <h3>The Evolution of Advanced Accelerator Concepts (AAC)</h3> <p>For years, the AAC community (including Plasma Wakefield Acceleration) focused heavily on designing a compact linear collider. While we have achieved remarkable gradients (10-100 GeV energy gain per meter), the stringent luminosity (collision rate) requirements of a particle physics collider remain a significant challenge.</p> <p><strong>A Pivot to Brightness</strong> However, this challenge has clarified where plasma acceleration truly shines. The unique strength of plasma-based schemes is not necessarily high average power (required for collider luminosity), but rather the ability to produce ultra-short beams with <strong>extreme peak brightness</strong>.</p> <p><strong>The Future: Photon Science</strong> This realization is driving a strategic evolution in our field. Rather than competing directly with conventional RF technology for collider luminosity, AAC is poised to revolutionize <strong>future light sources</strong>.</p> <p>As we demonstrated in our recent paper, <em>“<a href="https://www.nature.com/articles/s41467-025-65742-8">Plasma wakefield accelerator simultaneously boosts electron beam energy and brightness</a>,”</em> plasma accelerators can act as an <strong>“Energy and Brightness Dual Transformer.”</strong> By simultaneously increasing energy and peak brightness in a single stage, this technology enables a new class of <strong>compact, next-generation Free Electron Lasers (FELs)</strong>. This leverages plasma accelerators’ unique potential in chasing the <strong>Brightness Frontier</strong> and opens new doors for <strong>photon science</strong> and advanced materials research, playing to the inherent strengths of plasma technology while the high-energy physics community pursues the Muon and Proton paths for the <strong>Energy Frontier</strong>.</p>]]></content><author><name></name></author><category term="Future Colliders"/><category term="High Energy Physics"/><category term="Muon Collider"/><category term="FCC"/><category term="Advanced Accelerator Concepts"/><summary type="html"><![CDATA[As the community looks toward the 10 TeV energy frontier, a look at the strategic trade-offs between linear and circular approaches, and why the future of plasma acceleration may lie in high-brightness photon science.]]></summary></entry><entry><title type="html">A Verdict on the Annual Modulation: Settling a 20-Year Dark Matter Debate</title><link href="https://chaojiezhang.me/posts/2025/09/dark-matter-debate-settled/" rel="alternate" type="text/html" title="A Verdict on the Annual Modulation: Settling a 20-Year Dark Matter Debate"/><published>2025-09-23T00:00:00+00:00</published><updated>2025-09-23T00:00:00+00:00</updated><id>https://chaojiezhang.me/posts/2025/09/dark-matter-debate-settled</id><content type="html" xml:base="https://chaojiezhang.me/posts/2025/09/dark-matter-debate-settled/"><![CDATA[<h3 id="a-verdict-on-the-annual-modulation-settling-a-20-year-dark-matter-debate">A Verdict on the Annual Modulation: Settling a 20-Year Dark Matter Debate</h3> <p>For over two decades, one of the most tantalizing and controversial results in particle physics has been the signal from the DAMA/LIBRA experiment. Deep under the Gran Sasso mountain in Italy, its sodium iodide detectors have recorded a persistent annual modulation—a subtle, cyclical rise and fall in events that peaks each June and troughs in December. This is exactly the signature one would expect from the Earth moving through a galactic “wind” of WIMP dark matter.</p> <p>The problem? No other experiment has seen it. This discrepancy, known as the <strong>DAMA anomaly</strong>, created a deep rift in the field. The results from DAMA/LIBRA contradict other direct detection dark matter experiments under the most conventional scenarios. Was DAMA/LIBRA seeing the first definitive proof of dark matter, or was it an incredibly subtle, long-lived systematic effect?</p> <p>A new combined analysis from two independent experiments, <strong>COSINE-100</strong> and <strong>ANAIS-112</strong>, has now provided the strongest answer yet: the signal is likely not from dark matter.</p> <hr/> <h3 id="a-direct-confrontation-with-the-anomaly">A Direct Confrontation with the Anomaly</h3> <p>To test the DAMA/LIBRA claim, you need an identical setup. This was the explicit goal of the COSINE-100 experiment, located in South Korea, and ANAIS-112 in Spain. Crucially, both experiments use the exact same thallium-doped sodium iodide (NaI(Tl)) crystal detectors as DAMA/LIBRA. This allows for a direct, apples-to-apples comparison by removing ambiguities that arise from using different target materials. If the WIMP wind is real, they should see the same modulation.</p> <p>Independently, neither experiment saw a signal that supported DAMA/LIBRA’s claim. But to increase their statistical power, the collaborations took the next logical step: combining their data.</p> <hr/> <h3 id="the-combined-result-a-null-finding">The Combined Result: A Null Finding</h3> <p>In a new paper, researchers present a combined analysis of the first three years of data from both experiments. By carefully modeling and subtracting the known background events from each detector, they could search the combined “residual” data for any hint of modulation.</p> <p>The result is unambiguous.</p> <ul> <li>The combined data is fully consistent with a <strong>null hypothesis</strong>—that is, no modulation at all.</li> <li>The best-fit modulation amplitude they found was \(-0.0002 \pm 0.0026\) cpd/kg/keV in the 1-6 keV energy range. This is statistically indistinguishable from zero and fundamentally incompatible with DAMA/LIBRA’s reported amplitude of \(0.0105 \pm 0.0011\) cpd/kg/keV in the same range.</li> </ul> <p>The paper goes a step further by performing a simple combination of the newly released, much larger 6-year datasets from both experiments. This larger dataset, totaling nearly 1,000 kg-years of exposure, excludes the DAMA/LIBRA signal with even higher confidence: <strong>4.7σ</strong> in the 1-6 keV region and <strong>3.5σ</strong> in the 2-6 keV region.</p> <p>While this result does not identify the source of the DAMA/LIBRA signal, it strongly challenges the interpretation of the DAMA/LIBRA modulation in terms of galactic dark matter. This finding helps to resolve a long-standing puzzle and allows the field to move forward, focusing on other detection strategies as the search for dark matter continues.</p> <hr/> <h3 id="references">References</h3> <ol> <li>Carlin, N., et al. (2025). “Combined Annual Modulation Dark Matter Search with COSINE-100 and ANAIS-112.” <em>Physical Review Letters</em> 135.12: 121002.</li> <li>Bernabei, R., et al. (2020). “The DAMA project: Achievements, implications and perspectives.” <em>Progress in Particle and Nuclear Physics</em> 114: 103810.</li> </ol>]]></content><author><name></name></author><category term="Dark Matter"/><category term="Particle Physics"/><category term="DAMA"/><category term="COSINE-100"/><category term="ANAIS-112"/><category term="WIMP"/><summary type="html"><![CDATA[A new combined analysis from the COSINE-100 and ANAIS-112 experiments provides the strongest evidence yet against the long-standing DAMA/LIBRA dark matter claim.]]></summary></entry><entry><title type="html">The Kernel Trick: A Guide to High-Dimensional Feature Spaces</title><link href="https://chaojiezhang.me/posts/2025/08/kernel-trick-explained/" rel="alternate" type="text/html" title="The Kernel Trick: A Guide to High-Dimensional Feature Spaces"/><published>2025-08-21T00:00:00+00:00</published><updated>2025-08-21T00:00:00+00:00</updated><id>https://chaojiezhang.me/posts/2025/08/kernel-trick-explained</id><content type="html" xml:base="https://chaojiezhang.me/posts/2025/08/kernel-trick-explained/"><![CDATA[<h3 id="the-kernel-trick-a-guide-to-high-dimensional-feature-spaces">The Kernel Trick: A Guide to High-Dimensional Feature Spaces</h3> <p>Linear models are foundational in machine learning due to their simplicity and computational efficiency. They operate by learning a linear decision boundary to separate data. However, many real-world datasets are not linearly separable, limiting the direct application of these models.</p> <p>The kernel trick is a method that enables linear classifiers, such as Support Vector Machines (SVMs), to learn non-linear decision boundaries. It achieves this by implicitly mapping data to high-dimensional feature spaces where the data becomes linearly separable, without incurring the computational cost associated with explicitly performing this mapping.</p> <hr/> <h3 id="the-problem-non-linear-separability">The Problem: Non-Linear Separability</h3> <p>A linear classifier is restricted to learning hyperplane decision boundaries. For datasets where the class separation is curved or non-linear, such as in the case of concentric circles, a linear model cannot achieve effective separation. This is the problem of <strong>non-linearly separable data</strong>.</p> <hr/> <h3 id="the-brute-force-solution-explicit-feature-mapping">The “Brute Force” Solution: Explicit Feature Mapping</h3> <p>A direct approach to address non-linear data is to construct a new set of features through a process called <strong>feature mapping</strong>. A mapping function, denoted by \(\phi(x)\), transforms the original data into a new, higher-dimensional feature space. The objective is for the data to become linearly separable in this new space.</p> <p>Consider a one-dimensional dataset that is not linearly separable:</p> <p>Class B (-2) — Class A (-1) — Class A (1) — Class B (2)</p> <p>No single point can separate Class A from Class B. A feature map can be defined to add a second dimension, for instance, \(\phi(x) = (x_1, x_2) = (x, x^2)\). Applying this transformation yields new coordinates:</p> <ul> <li>Class B at \(x=-2\) maps to \((-2, 4)\)</li> <li>Class A at \(x=-1\) maps to \((-1, 1)\)</li> <li>Class A at \(x=1\) maps to \((1, 1)\)</li> <li>Class B at \(x=2\) maps to \((2, 4)\)</li> </ul> <p>In the resulting 2D space, the classes are now linearly separable by a horizontal line, such as \(x_2=2.5\).</p> <p>This demonstrates that the problem is solvable in a higher-dimensional space.</p> <h4 id="the-curse-of-dimensionality">The Curse of Dimensionality</h4> <p>The primary issue with explicit feature mapping is its lack of scalability. For complex decision boundaries, the required number of new features can grow exponentially, leading to several challenges known as the <strong>curse of dimensionality</strong>. The computation and storage of these high-dimensional vectors become prohibitively expensive, and the risk of model overfitting increases significantly.</p> <hr/> <h3 id="the-efficient-solution-the-kernel-trick">The Efficient Solution: The Kernel Trick</h3> <p>An analysis of certain algorithms, including the dual formulation of the SVM, reveals that the feature vectors are only used in the form of <strong>dot products</strong> (or inner products), \(\phi(x)^T \phi(z)\). The algorithm’s execution depends only on these scalar dot product values, not on the individual coordinates of the vectors.</p> <p>A <strong>kernel function</strong>, denoted \(K(x, z)\), is a computationally efficient function that takes two low-dimensional vectors, \(x\) and \(z\), as input and returns the value of the dot product between their high-dimensional representations, \(\phi(x)\) and \(\phi(z)\).</p> \[\underbrace{K(x, z)}_{\text{computationally efficient}} = \underbrace{\phi(x)^T \phi(z)}_{\text{computationally expensive}}\] <p>The kernel trick consists of replacing the expensive high-dimensional dot product with its corresponding kernel function. This allows the algorithm to gain the full benefit of operating in the high-dimensional feature space without ever explicitly computing or storing the vectors \(\phi(x)\). Consequently, the mapping function \(\phi(x)\) does not need to be known.</p> <hr/> <h3 id="how-to-use-kernels-in-practice">How to Use Kernels in Practice</h3> <p>The practical application of kernel methods involves selecting an appropriate kernel function and tuning its parameters, rather than designing new kernels for each problem.</p> <h4 id="standard-kernel-functions">Standard Kernel Functions</h4> <ol> <li><strong>Linear Kernel</strong> <ul> <li><strong>Formula</strong>: \(K(x, z) = x^T z\)</li> <li><strong>Use Case</strong>: This is the standard dot product and results in a linear classifier. It is computationally efficient and serves as a good baseline, particularly for datasets with a large number of features.</li> </ul> </li> <li><strong>Polynomial Kernel</strong> <ul> <li><strong>Formula</strong>: \(K(x, z) = (\gamma x^T z + c)^d\)</li> <li><strong>Hyperparameters</strong>: The degree \(d\), the slope \(\gamma\), and the constant \(c\).</li> <li><strong>Use Case</strong>: This kernel is effective for problems with a known polynomial structure in their decision boundary.</li> </ul> </li> <li><strong>Gaussian (RBF) Kernel</strong> <ul> <li><strong>Formula</strong>: \(K(x, z) = \exp\left(-\gamma \mid x-z\ \mid ^2\right)\)</li> <li><strong>Hyperparameter</strong>: The bandwidth \(\gamma\) (gamma), where \(\gamma = 1/(2\sigma²)\).</li> <li><strong>Use Case</strong>: The Radial Basis Function (RBF) kernel is a common default choice due to its flexibility. It corresponds to an infinite-dimensional feature space and can learn complex, non-linear decision boundaries. The \(\gamma\) parameter controls the smoothness of the boundary; smaller values result in smoother boundaries, while larger values can lead to more complex boundaries that fit the data more tightly.</li> </ul> </li> </ol> <h4 id="the-practical-workflow">The Practical Workflow</h4> <p>For a typical machine learning project utilizing a kernelized algorithm like an SVM:</p> <ol> <li><strong>Select a kernel.</strong> The RBF kernel is a standard and effective starting point for most problems.</li> <li><strong>Tune hyperparameters using cross-validation.</strong> For an RBF SVM, this involves finding the optimal combination of the kernel’s \(\gamma\) and the SVM’s regularization parameter \(C\). These parameters control model complexity and the trade-off between classification accuracy and margin size.</li> <li><strong>Train the final model.</strong> After identifying the optimal hyperparameters, train the model on the entire training dataset.</li> </ol> <p>The RBF kernel is generally the recommended first choice unless there is specific domain knowledge suggesting that a different kernel structure is more appropriate for the data.</p> <hr/> <h3 id="theoretical-connection-basis-and-representation">Theoretical Connection: Basis and Representation</h3> <p>The analogy between kernel methods and the Fourier Transform is insightful. Both techniques operate by changing the basis of the problem space.</p> <p>The Fourier Transform projects a signal from the time domain onto a basis of sine and cosine functions to simplify analysis in the frequency domain. Similarly, a kernel function implicitly projects data onto a basis of high-dimensional feature functions, where a linear separation becomes possible.</p> <p>The theoretical foundation for this is <strong>Mercer’s Theorem</strong>. It states that any valid kernel function (specifically, any continuous, symmetric, positive semi-definite function) corresponds to a dot product in some high-dimensional feature space. This theorem provides the mathematical guarantee for the validity of the kernel trick.</p> <p>By changing the data’s representation, these methods can transform a computationally difficult problem into a more tractable one.</p>]]></content><author><name></name></author><category term="Machine Learning"/><category term="Kernel Methods"/><category term="SVM"/><category term="Support Vector Machine"/><category term="Kernel Trick"/><category term="Data Science"/><category term="Tutorial"/><summary type="html"><![CDATA[A technical explanation of how the kernel trick allows linear models to solve non-linear problems by implicitly operating in high-dimensional feature spaces.]]></summary></entry><entry><title type="html">A Scalpel of Electrons: The Promise and Clinical Reality of Laser-Driven Radiotherapy</title><link href="https://chaojiezhang.me/posts/2025/08/laser-driven-radiotherapy/" rel="alternate" type="text/html" title="A Scalpel of Electrons: The Promise and Clinical Reality of Laser-Driven Radiotherapy"/><published>2025-08-16T00:00:00+00:00</published><updated>2025-08-16T00:00:00+00:00</updated><id>https://chaojiezhang.me/posts/2025/08/laser-driven-radiotherapy</id><content type="html" xml:base="https://chaojiezhang.me/posts/2025/08/laser-driven-radiotherapy/"><![CDATA[<h3 id="the-quest-for-a-more-precise-beam-in-cancer-treatment">The Quest for a More Precise Beam in Cancer Treatment</h3> <p>For decades, the workhorse of radiation oncology has been the medical linear accelerator (linac), using high-energy X-rays (photons) to destroy tumors. While incredibly effective, the goal has always been to find new ways to maximize damage to cancer cells while further sparing the healthy tissue around them.</p> <p>One of the most exciting frontiers in this quest comes from the world of high-energy physics: using compact, powerful lasers to accelerate electrons to very high energies (VHEE) for therapy. This technology, known as Laser-Wakefield Acceleration (LWFA), promises to shrink room-sized accelerators down to a tabletop. After years of development, these systems have evolved from laboratory curiosities into stable, high-performance machines. But as the technology matures, it faces a tougher question: Is there a widespread clinical problem for it to solve?</p> <hr/> <h3 id="from-lab-curiosity-to-engineering-marvel">From Lab Curiosity to Engineering Marvel</h3> <p>Early-generation LWFAs were known for instabilities, making them unsuitable for the precision demands of medicine. However, the state-of-the-art has taken a big leap forward.</p> <p>Most recent LWFAs now demonstrate remarkable progress. Researchers have achieved significant improvements in the stability of the beam’s energy and direction, bringing them closer to the consistency required for medical applications. Furthermore, the core promise of overall compactness has been realized (due to the development of compact laser systems), with the accelerator component itself being dramatically smaller than conventional linear accelerators. These advances prove that many fundamental physics challenges have been overcome, resulting in a legitimate, high-performance technology capable of delivering a high-energy electron beam.</p> <hr/> <h3 id="the-clinical-question-why-vhee">The Clinical Question: Why VHEE?</h3> <p>The central challenge for VHEE emerges when we compare its dose profile to the existing standard of care. For a deep-seated tumor, the dose delivered by a high-energy electron beam doesn’t look dramatically different from that of a standard X-ray beam. Both can reach the target effectively. So, why bother with the complexity of VHEE?</p> <p>The answer lies in the subtle but critical details of the beam’s behavior.</p> <ol> <li> <p><strong>A Sharper Edge (Lateral Penumbra):</strong> Due to their high inertia, VHEE beams scatter less to the side than photons. This creates an incredibly sharp beam edge, which in principle allows clinicians to “paint” the dose onto a tumor that is wrapped around a critical organ with less collateral damage.</p> </li> <li> <p><strong>Magnetic Steering:</strong> Unlike photons, electrons are charged particles. This means they can be steered and focused with magnets. This unique capability allows for novel treatment approaches, such as concentrating the dose into small, highly resistant pockets within a tumor—a feat impossible with X-rays.</p> </li> </ol> <p>These advantages suggest VHEE isn’t meant to replace the workhorse photon linac, but rather to serve as a specialized tool for the most difficult 5-10% of cancer cases where surgical precision is paramount.</p> <hr/> <h3 id="the-flash-reality-check">The FLASH Reality Check</h3> <p>Much of the recent excitement around new accelerator technologies is tied to <strong>FLASH radiotherapy</strong>, a technique that delivers the entire radiation dose in a fraction of a second. This ultra-fast delivery has shown potential for sparing healthy tissue.</p> <p>Here, the physics of accelerators presents a stark reality. The threshold to achieve the FLASH effect is widely considered to be delivering an entire therapeutic dose at an average rate greater than <strong>40 Gy/s</strong>, with the full irradiation lasting less than <strong>200 milliseconds</strong>. This requires a massive, sustained flux of particles (to estimate order of magnitude, ~10 nC of 100 MeV electrons deliver 1 Gy (1 Joule of energy deposited in 1 kG of mass)).</p> <ul> <li><strong>LWFA’s Limitation:</strong> While an LWFA pulse has an immense <em>instantaneous</em> dose rate, its low repetition rate is a bottleneck. To achieve the average dose rates required for FLASH, the system would need to pulse thousands of times per second—orders of magnitude faster than current capabilities.</li> <li><strong>Conventional Linacs:</strong> In contrast, advanced linacs are already capable of producing the high average currents needed to meet the FLASH threshold.</li> </ul> <p>For the specific application of FLASH therapy, significant technological advances in laser repetition rates are needed before LWFA becomes a viable platform.</p> <hr/> <h3 id="outlook-charting-a-path-to-the-clinic">Outlook: Charting a Path to the Clinic</h3> <p>LWFA-VHEE technology represents a monumental engineering achievement. It has matured from a physics experiment into a platform capable of producing high-quality electron beams with medically relevant properties.</p> <p>However, the path from a technically proven system to a widely adopted clinical tool is complex. For the vast majority of cases, conventional photon therapy remains an effective and reliable standard of care. For the niche cases requiring higher conformity, VHEE technology must be evaluated alongside other advanced modalities like proton therapy.</p> <p>The ultimate role of these new accelerator technologies in medicine is still being defined. The journey involves overcoming not just technical hurdles, but also regulatory and economic challenges. Dedicated teams of scientists and engineers continue to push the boundaries of what is possible, and the evolution of these systems will determine their place in the future of cancer treatment.</p>]]></content><author><name></name></author><category term="Radiotherapy"/><category term="Particle Accelerators"/><category term="VHEE"/><category term="FLASH Therapy"/><category term="Medical Physics"/><summary type="html"><![CDATA[An analysis of laser-wakefield accelerators for cancer therapy, exploring how this remarkable technology has overcome physics barriers but now faces the challenge of clinical necessity.]]></summary></entry><entry><title type="html">The Next Light: Can Particle Accelerators Power the Future of Moore’s Law?</title><link href="https://chaojiezhang.me/posts/2025/08/accelerators-moores-law/" rel="alternate" type="text/html" title="The Next Light: Can Particle Accelerators Power the Future of Moore’s Law?"/><published>2025-08-14T00:00:00+00:00</published><updated>2025-08-14T00:00:00+00:00</updated><id>https://chaojiezhang.me/posts/2025/08/accelerators-moores-law</id><content type="html" xml:base="https://chaojiezhang.me/posts/2025/08/accelerators-moores-law/"><![CDATA[<h3 id="the-need-for-a-brighter-future-in-chipmaking">The Need for a Brighter Future in Chipmaking</h3> <p>The engine of the digital age has long been Moore’s Law—the relentless doubling of transistors on a microchip. This progress is a direct result of advances in photolithography, the process of using light to etch circuits onto silicon wafers. To create today’s most advanced chips, the industry relies on Extreme Ultraviolet (EUV) light with a wavelength of 13.5 nanometers.</p> <p>The current state-of-the-art EUV source technology, while a monumental engineering achievement, is approaching fundamental limits in the power required for the next generation of high-volume manufacturing. This has spurred a search for a new kind of light source, with a promising candidate emerging from the world of high-energy physics: the particle accelerator.</p> <hr/> <h3 id="the-incumbent-laser-produced-plasma-lpp">The Incumbent: Laser-Produced Plasma (LPP)</h3> <p>Today’s EUV lithography is enabled by ASML’s Laser-Produced Plasma (LPP) sources. In this system, a high-power CO2 laser pulverizes thousands of molten tin droplets per second, creating an intensely hot plasma that radiates the required EUV light. This technology has successfully brought EUV sources to ~500 W of average power.</p> <p>However, the industry roadmap calls for source power to scale beyond 1 kW to improve manufacturing throughput and reduce defects. Scaling LPP technology faces significant challenges, primarily related to managing the tin debris generated by the plasma, which can contaminate and damage the priceless collection optics within the system.</p> <hr/> <h3 id="the-challenger-accelerator-driven-free-electron-lasers-fels">The Challenger: Accelerator-Driven Free-Electron Lasers (FELs)</h3> <p>An accelerator-based approach offers a fundamentally different path. In a Free-Electron Laser (FEL), a high-energy beam of electrons is passed through a series of alternating magnets called an undulator. This forces the electrons to “wiggle” and radiate powerful, laser-like light.</p> <p>An FEL-based source has several key advantages:</p> <ul> <li><strong>High Power:</strong> FELs are scalable to kilowatt-level average power if powered by high rep. rate e- beams.</li> <li><strong>Clean Operation:</strong> The process occurs in an ultra-high vacuum, producing no debris.</li> <li><strong>Tunability:</strong> The output wavelength can be adjusted, offering a potential path to “Beyond EUV” lithography at even shorter wavelengths.</li> </ul> <p>The primary drawback of a single-pass FEL is its low intrinsic efficiency, typically converting less than 0.1% of the electron beam’s energy into EUV light. The central challenge, therefore, is to engineer a system that dramatically improves this efficiency and/or increase the rep. rate of the FEL pulses.</p> <hr/> <h3 id="architectures-for-a-high-efficiency-fel-source">Architectures for a High-Efficiency FEL Source</h3> <p>To be commercially viable, an accelerator-based EUV source must be both powerful and efficient. Two main architectures have been proposed in the scientific literature to achieve this. Both place the FEL undulator inside an optical cavity to create a high-gain Regenerative Amplifier FEL (RAFEL), but they differ in how they manage the electron beam.</p> <p><strong>1. The Energy Recovery Linac (ERL) Approach:</strong> An ERL is an exceptionally efficient type of linear accelerator. A high-current, MHz-rate beam of electrons passes through the undulator <em>once</em> to create light. The spent beam is then looped back through the main accelerator out of phase, where it deposits its vast remaining energy back into the accelerating field to power the next bunch. This “energy recovery” process can achieve over 99% efficiency, drastically reducing the system’s power consumption. A recent paper by He et al. (<em>Phys. Rev. Accel. Beams</em>, 2025) details a compact ERL-RAFEL design capable of producing over 2 kW of EUV power.</p> <p><strong>2. The Storage Ring Approach:</strong> This architecture uses a circular accelerator, or storage ring, to act as a repetition rate multiplier. A less demanding, “slow” injector (e.g., kHz) fills the ring with many electron bunches, which then circulate at MHz frequencies. The RAFEL is placed in a straight section of the ring. On every turn, each bunch passes through the RAFEL, amplifying the light stored in the optical cavity. The key challenge is managing the “heat” imparted to the beam by the FEL interaction, which must be balanced by the natural “cooling” effect of synchrotron radiation in the ring’s magnets. This sets a limit on the maximum power that can be extracted.</p> <hr/> <h3 id="outlook">Outlook</h3> <p>The engineering task of building an accelerator complex with the reliability and uptime required for 24/7 high-volume manufacturing is immense. However, both the ERL and Storage Ring concepts represent sound, intensively discussed physical bases for a next-generation EUV source. They directly address the core challenges of power scaling and efficiency, offering plausible and compelling answers to what might light the way for the future of Moore’s Law.</p>]]></content><author><name></name></author><category term="EUV Lithography"/><category term="Particle Accelerators"/><category term="Free-Electron Laser"/><category term="Energy Recovery Linac"/><category term="ASML"/><summary type="html"><![CDATA[An overview of the push for next-generation EUV light sources, comparing the incumbent laser-produced plasma technology against emerging accelerator-based Free-Electron Laser concepts.]]></summary></entry><entry><title type="html">Reconstructing Electron Bunch Current Profiles with Conditional Diffusion Models</title><link href="https://chaojiezhang.me/posts/2025/07/diffusion-bunch-reconstruction/" rel="alternate" type="text/html" title="Reconstructing Electron Bunch Current Profiles with Conditional Diffusion Models"/><published>2025-07-31T00:00:00+00:00</published><updated>2025-07-31T00:00:00+00:00</updated><id>https://chaojiezhang.me/posts/2025/07/diffusion-bunch-reconstruction</id><content type="html" xml:base="https://chaojiezhang.me/posts/2025/07/diffusion-bunch-reconstruction/"><![CDATA[<h3 id="introduction-the-diagnostic-challenge-in-plasma-wakefield-accelerators">Introduction: The Diagnostic Challenge in Plasma Wakefield Accelerators</h3> <p>The accurate characterization of the longitudinal current profile of electron bunches within Plasma Wakefield Accelerators (PWFAs) is critical for optimizing performance. Downstream diagnostics, such as spectrometers, provide integrated information about the bunch after it has exited the plasma. Reconstructing the initial in-plasma current profile (\(dQ/d\xi\)) from these final measurements constitutes a challenging, ill-posed inverse problem.</p> <hr/> <h3 id="the-inverse-problem-and-its-ambiguity">The Inverse Problem and Its Ambiguity</h3> <p>Two primary diagnostics are commonly used:</p> <ol> <li><strong>Electron Energy Spectrometer:</strong> Measures the final energy distribution (\(dQ/dE\)). This observable is the result of the complex, non-linear wakefield loading integrated over the full acceleration distance. Different initial current profiles can lead to similar final energy spectra, creating significant ambiguity.</li> <li><strong>COTR Spectrometer:</strong> Measures the spectrum of Coherent Optical Transition Radiation, which is proportional to the squared amplitude of the bunch form factor (\(\lvert b(k) \rvert^2\)). While directly related to the bunch’s Fourier transform, this measurement lacks the phase information necessary for unique profile reconstruction.</li> </ol> <p>Each diagnostic alone is insufficient to uniquely solve the inverse problem. Our work proposes a method to fuse the information from both diagnostics using a conditional generative model to overcome these limitations.</p> <hr/> <h3 id="methodology-a-conditional-diffusion-model">Methodology: A Conditional Diffusion Model</h3> <p>We propose using a conditional denoising diffusion probabilistic model (DDPM) to solve this inverse problem. Diffusion models are a class of generative models that learn to reverse a fixed Markovian process that gradually adds Gaussian noise to data.</p> <p>The reverse process is where the learning occurs. A neural network, typically a <strong>U-Net</strong>, is trained to denoise the data at each step <code class="language-plaintext highlighter-rouge">t</code> by predicting the noise that was added. The key to solving inverse problems is that this denoising process can be <strong>conditioned</strong> on external data <code class="language-plaintext highlighter-rouge">y</code>—in our case, the measured energy and COTR spectra. The network learns to approximate the score of the conditional distribution, \(\nabla_{x_t} \log p(x_t \mid y)\), guiding the generation process from random noise toward a high-fidelity solution that is consistent with the specific experimental measurements.</p> <hr/> <h3 id="advantages-over-feed-forward-mlp-networks">Advantages Over Feed-Forward MLP Networks</h3> <p>While a simple Multi-Layer Perceptron (MLP) can be trained to directly map spectra to a current profile, it is fundamentally ill-suited for this type of ill-posed problem.</p> <ul> <li> <p><strong>Handling the One-to-Many Mapping:</strong> An MLP is, by definition, a single-valued function. When trained with a standard loss function like Mean Squared Error (MSE) on an ambiguous (one-to-many) dataset, the network minimizes its average error by learning to predict the <strong>conditional mean</strong> of all possible solutions. This often results in an “averaged,” overly smoothed output that may not represent any single physically plausible solution.</p> </li> <li> <p><strong>Sampling from the Posterior Distribution:</strong> A diffusion model, in contrast, is a generative process. Given a conditioning signal <code class="language-plaintext highlighter-rouge">y</code>, it learns to draw samples from the true posterior distribution <code class="language-plaintext highlighter-rouge">p(x|y)</code>. By initializing the reverse denoising process with different random noise seeds, we can generate multiple distinct, high-fidelity samples, each representing a valid solution consistent with the measured spectra.</p> </li> <li> <p><strong>Intrinsic Uncertainty Quantification:</strong> This sampling capability is the diffusion model’s most significant scientific advantage. The generated ensemble of profiles provides a direct estimate of the posterior distribution. The mean of the ensemble can be interpreted as the expected solution, while the variance across the ensemble serves as a robust, pixel-wise <strong>uncertainty quantification</strong>. This allows us to determine precisely which features of the reconstructed profile are well-constrained by the diagnostics and which remain ambiguous—a capability a standard MLP completely lacks.</p> </li> </ul> <hr/> <h3 id="implementation-and-outlook">Implementation and Outlook</h3> <p>To ensure the model learns the underlying physics rather than experimental biases, the training data will be generated using Latin Hypercube Sampling (LHS) to create an unbiased distribution of simulated current profiles and their corresponding diagnostic signals. This work represents a novel approach to accelerator diagnostics, offering not only a path to high-fidelity profile reconstruction but also a principled framework for quantifying the inherent uncertainties of the solution.</p>]]></content><author><name></name></author><category term="Plasma Wakefield Accelerator"/><category term="Machine Learning"/><category term="Inverse Problem"/><category term="Diffusion Model"/><category term="Diagnostics"/><summary type="html"><![CDATA[This work proposes using a conditional diffusion model to solve the ill-posed inverse problem of reconstructing electron bunch current profiles from downstream diagnostics in plasma wakefield accelerators.]]></summary></entry><entry><title type="html">Laser frequency combs: a precision revolution</title><link href="https://chaojiezhang.me/posts/2025/07/laser-frequency-combs/" rel="alternate" type="text/html" title="Laser frequency combs: a precision revolution"/><published>2025-07-07T00:00:00+00:00</published><updated>2025-07-07T00:00:00+00:00</updated><id>https://chaojiezhang.me/posts/2025/07/laser-frequency-combs</id><content type="html" xml:base="https://chaojiezhang.me/posts/2025/07/laser-frequency-combs/"><![CDATA[<p><strong>The 2005 Nobel Prize-winning technology that transformed optical metrology from laboratory curiosity to commercial reality continues to revolutionize precision measurement, enabling adavnces from ultra-stable atomic clocks to chip-scale sensors for autonomous vehicles.</strong></p> <p>In 2000, a breakthrough at JILA transformed one of physics’ most complex measurement challenges into an elegant solution using a single laser system. <a href="https://www.nobelprize.org/prizes/physics/2005/hall/facts/">John Hall</a> and <a href="https://www.nobelprize.org/prizes/physics/2005/hansch/facts/">Theodor Hänsch</a> demonstrated that a mode-locked femtosecond laser could directly bridge the gap between optical and microwave frequencies—eliminating the need for elaborate 15-stage frequency multiplication chains that had dominated optical metrology for decades. This <strong>frequency comb technique</strong> earned them the <a href="https://www.nobelprize.org/prizes/physics/2005/press-release/">2005 Nobel Prize in Physics</a> and launched a technological revolution that continues to reshape precision science and commercial applications today.</p> <p>The impact extends far beyond the laboratory. Modern atomic clocks now achieve <strong>precision at the 19th decimal place</strong>, enabling tests of fundamental physics that were previously impossible. The global optical frequency comb market, valued at $542 million in 2024, is projected to reach $1.4 billion by 2032, driven by applications ranging from autonomous vehicle LiDAR to environmental monitoring systems that can detect greenhouse gases over 113-kilometer open-air paths.</p> <h2 id="a-40-year-journey-from-concept-to-nobel-recognition">A 40-year journey from concept to Nobel recognition</h2> <p>The frequency comb concept emerged from Hänsch’s visionary work in the late 1970s at Stanford University, where he recognized that synchronized femtosecond pulses could serve as precision frequency rulers. However, the technology remained impractical for decades due to technical limitations. Early picosecond dye lasers lacked the bandwidth and stability needed for widespread adoption, while traditional harmonic frequency chains—requiring complex multi-stage setups—remained the only viable approach for optical frequency metrology.</p> <p>The breakthrough came through a convergence of enabling technologies in the 1990s. <strong>Titanium-doped sapphire lasers</strong> provided the femtosecond pulse generation capabilities, while the development of photonic crystal fibers at Bell Labs in 2000 enabled the crucial octave-spanning spectra needed for self-referencing. This self-referencing technique, developed independently by Hall’s team at JILA and Hänsch’s group at the Max Planck Institute, eliminated the need for external optical calibration by allowing the frequency comb to determine its own absolute frequencies.</p> <p>The technical elegance lies in the frequency comb equation: \(f_n = n \times f_{rep} + f_0\), where just two microwave-domain parameters (repetition rate and carrier-envelope offset frequency) define the entire optical spectrum. This seemingly simple relationship enabled scientists to <strong>count optical oscillations at \(10^{15}\) Hz</strong> with unprecedented precision, revolutionizing optical frequency metrology overnight.</p> <h2 id="nobel-recognition-validates-transformative-impact">Nobel recognition validates transformative impact</h2> <p>The <a href="https://www.nobelprize.org/prizes/physics/2005/press-release/">2005 Nobel Committee</a> recognized Hall and Hänsch’s contributions as revolutionary, noting that their technique “made it possible to measure frequencies with an accuracy of fifteen digits.” The Committee emphasized that the development enabled “studies of the stability of the constants of nature over time and to develop extremely accurate clocks and improved GPS technology.”</p> <p>The recognition came just five years after the critical 2000 demonstrations, reflecting the immediate impact of the technology. Two seminal papers published that year established the practical foundation: Jones et al. demonstrated carrier-envelope phase control in femtosecond lasers, while <a href="https://link.aps.org/doi/10.1103/PhysRevLett.85.2264">Holzwarth et al.</a> realized an optical frequency synthesizer with uncertainty below \(5.1\times10^{-16}\). These achievements directly replaced complex frequency chains with a single laser system, transforming optical metrology from an esoteric laboratory technique into an accessible precision tool.</p> <p>The Nobel Committee’s citation specifically highlighted how the technique <strong>“revolutionized the art of counting the frequency of light”</strong> and created the “long missing clockwork for optical atomic clocks.” This recognition validated frequency combs as foundational technology for precision science, spurring commercial development and expanding applications across multiple fields.</p> <h2 id="current-applications-span-precision-science-to-commercial-markets">Current applications span precision science to commercial markets</h2> <p>Today’s frequency comb applications extend far beyond the original metrology focus. <strong>Optical atomic clocks</strong> now serve as the world’s most precise timekeeping devices, with accuracy improvements enabling next-generation GPS systems that could achieve centimeter-level precision. The technology has become essential for fundamental physics experiments, including tests of general relativity and searches for dark matter through precision spectroscopy.</p> <p>In telecommunications, microresonator-based frequency combs provide multiple wavelength channels for high-capacity data transmission, while enabling interconnection of thousands of computers in cloud computing facilities. The technology has revolutionized chemical analysis through <strong>dual-comb spectroscopy</strong>, enabling real-time detection of greenhouse gases, industrial emissions monitoring, and portable medical diagnostics.</p> <p>Commercial systems are now available from multiple vendors, including <a href="https://www.menlosystems.com/">Menlo Systems</a>, <a href="https://www.toptica.com/">Toptica Photonics</a>, and <a href="https://www.nktphotonics.com/">NKT Photonics</a>, with turnkey solutions deployed in research laboratories and industrial facilities worldwide. The technology has moved from specialized scientific instrumentation to practical sensing applications, with companies like <a href="https://longpathtech.com/">LongPath Technologies</a> developing frequency comb-based systems for environmental monitoring and oil field emission detection.</p> <h2 id="chip-scale-integration-unlocks-new-frontiers">Chip-scale integration unlocks new frontiers</h2> <p>The field’s most exciting developments center on <strong>chip-scale integration</strong>, which promises to transform frequency combs from laboratory instruments into ubiquitous sensing devices. Recent breakthroughs in silicon nitride microresonators have achieved greater than 50% pump-to-comb conversion efficiency, while battery-operated soliton microcomb systems now consume only 98 mW of electrical power.</p> <p>Novel comb architectures are expanding capabilities beyond traditional limitations. <a href="https://www.nist.gov/news-events/news/2024/03/researchers-develop-new-type-frequency-comb-promises-further-boost-accuracy">NIST’s parametrically driven dual-pump laser systems</a>, demonstrated in March 2024, offer enhanced accuracy for precision measurements. <strong>Electro-optic combs</strong> based on lithium niobate platforms achieve 30% conversion efficiency with 132 nm optical spans, while quantum cascade lasers enable mid-infrared frequency combs for molecular spectroscopy applications.</p> <p>The integration with quantum technologies represents a particularly promising frontier. Researchers have demonstrated <strong>quantum frequency combs</strong> generating entangled photon pairs across more than 500 frequency channels, enabling high-dimensional quantum communications and enhanced sensing capabilities. These developments position frequency combs as essential tools for the emerging quantum technology ecosystem.</p> <h2 id="future-outlook-from-specialty-instruments-to-everywhere-devices">Future outlook: from specialty instruments to everywhere devices</h2> <p>Looking toward 2030, frequency combs are poised to become as ubiquitous as GPS systems are today. Chip-scale integration will enable deployment in autonomous vehicles for precision LiDAR, environmental monitoring networks for real-time pollution tracking, and point-of-care medical devices for disease diagnosis. The technology’s integration with artificial intelligence and machine learning promises autonomous optimization and self-calibrating systems.</p> <p>The convergence of frequency combs with quantum technologies will likely drive the next wave of scientific discoveries and commercial applications. From quantum computers requiring precise frequency control to gravitational wave detectors with enhanced sensitivity, frequency combs will remain at the forefront of precision measurement science.</p> <h2 id="conclusion">Conclusion</h2> <p>The frequency comb revolution exemplifies how fundamental physics research can yield transformative practical applications. Hall and Hänsch’s breakthrough solved a decades-old measurement problem while simultaneously creating entirely new technological possibilities. As the technology continues evolving from Nobel Prize-winning scientific achievement to commercial reality, frequency combs demonstrate the continuing power of precision measurement to drive both scientific discovery and technological innovation.</p> <p>The field’s trajectory from laboratory curiosity to billion-dollar market illustrates how “precision breeds discovery”—by improving measurement accuracy by orders of magnitude, frequency combs continue enabling scientific breakthroughs that reshape our understanding of the universe while creating practical tools that benefit society at large.</p> <h2 id="references">References</h2> <ul> <li><a href="https://www.nobelprize.org/prizes/physics/2005/press-release/">Nobel Prize in Physics 2005</a></li> <li><a href="https://www.nist.gov/topics/physics/optical-frequency-combs">NIST Optical Frequency Combs Overview</a></li> <li><a href="https://www.science.org/doi/10.1126/science.aay3676">Science Magazine: Optical frequency combs review</a></li> <li><a href="https://www.nature.com/articles/s41566-024-01525-9">Nature Photonics: Dual-comb spectroscopy over 100 km</a></li> </ul>]]></content><author><name></name></author><category term="Laser Frequency Combs"/><category term="Precision Measurement"/><category term="Optical Physics"/><summary type="html"><![CDATA[How laser frequency combs evolved from precision metrology tools to ubiquitous sensing devices powering quantum technologies and autonomous systems.]]></summary></entry></feed>