参考文献
参考文献
核心论文
Vaswani, A., et al. “Attention Is All You Need.” NeurIPS, 2017.
Dean, J., et al. “Large Scale Distributed Deep Networks.” NeurIPS, 2012.
Rajbhandari, S., et al. “ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.” SC, 2020.
Shoeybi, M., et al. “Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.” arXiv:1909.08053, 2019.
Narayanan, D., et al. “Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.” SC, 2021.
Korthikanti, V., et al. “Reducing Activation Recomputation in Large Transformer Models.” MLSys, 2023.
Shazeer, N., et al. “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.” ICLR, 2017.
Fedus, W., et al. “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.” JMLR, 2022.
Kwon, W., et al. “Efficient Memory Management for Large Language Model Serving with PagedAttention.” SOSP, 2023.
Leviathan, Y., et al. “Fast Inference from Transformers via Speculative Decoding.” ICML, 2023.
分布式系统
Zaharia, M., et al. “Resilient Distributed Datasets.” NSDI, 2012.
Verma, A., et al. “Large-scale cluster management at Google with Borg.” EuroSys, 2015.
Burns, B., et al. “Borg, Omega, and Kubernetes.” Communications of the ACM, 2016.
硬件与网络
NVIDIA. “NVIDIA H100 Tensor Core GPU Architecture Whitepaper.” 2024.
NVIDIA. “NVIDIA GB200 NVL72 System Architecture.” 2025.
Zahavi, E., et al. “Fat-Tree Topology for Data Center Networks.” IEEE, 2014.
推理优化
Frantar, E., et al. “GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.” ICLR, 2023.
Lin, J., et al. “AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.” MLSys, 2024.
Zheng, L., et al. “SGLang: Scheduling Language Model Programs.” arXiv:2312.07104, 2023.
开源模型
Touvron, H., et al. “LLaMA: Open and Efficient Foundation Language Models.” arXiv:2302.13971, 2023.
Touvron, H., et al. “Llama 2: Open Foundation and Fine-Tuned Chat Models.” arXiv:2307.09288, 2023.
Jiang, A. Q., et al. “Mistral 7B.” arXiv:2310.06825, 2023.
Jiang, A. Q., et al. “Mixtral of Experts.” arXiv:2401.04088, 2024.
Bai, J., et al. “Qwen Technical Report.” arXiv:2309.16609, 2023.
DeepSeek-AI. “DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.” arXiv:2405.04434, 2024.
DeepSeek-AI. “DeepSeek-V3 Technical Report.” arXiv:2412.19437, 2024.
书籍
Huyen, C. AI Engineering. O’Reilly Media, 2025.
Patterson, D. A., & Hennessy, J. L. Computer Architecture: A Quantitative Approach. 6th ed., 2019.
Kleppmann, M. Designing Data-Intensive Applications. O’Reilly Media, 2017.
博客与技术报告
Anthropic. “Effective Context Engineering for AI Agents.” Anthropic Engineering Blog, 2025.
OpenAI. “Harness Engineering.” OpenAI Blog, 2025.
Google. “TPU v5e Cloud TPU Architecture.” Google Cloud Blog, 2024.
Meta. “Building Meta’s GenAI Infrastructure.” Meta Engineering Blog, 2024.
本参考文献列表持续更新中。完整论文索引请访问:GitHub