Mac Anderson

Part VII: Back matter

Sources

The works cited, in order of first citation.

Mac Anderson119 sources
View markdown

119 works are cited in this book. They appear below in the order of their first citation. Each chapter also lists its own sources at its end.

  1. Cottier, B., You, J., Martemianova, N., & Owen, D. (2024). How far behind are open models? Epoch AI. https://epoch.ai/blog/open-models-report
  2. Stanford Institute for Human-Centered Artificial Intelligence (2025). AI Index Report 2025, Chapter 2: Technical Performance. https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance
  3. OpenAI (2025). gpt-oss-120b and gpt-oss-20b Model Card. arXiv. https://arxiv.org/abs/2508.10925
  4. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv; ICLR 2024. https://arxiv.org/abs/2310.01798
  5. Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM Evaluators Recognize and Favor Their Own Generations. arXiv; NeurIPS 2024. https://arxiv.org/abs/2404.13076
  6. Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). STaR: Bootstrapping Reasoning With Reasoning. arXiv. https://arxiv.org/abs/2203.14465
  7. DeepSeek-AI (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv. https://arxiv.org/abs/2501.12948
  8. Baker, B., Huizinga, J., Gao, L., et al. (2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. arXiv. https://arxiv.org/abs/2503.11926
  9. Denison, C., MacDiarmid, M., Barez, F., et al. (2024). Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models. arXiv. https://arxiv.org/abs/2406.10162
  10. Blum, A., & Hardt, M. (2015). The Ladder: A Reliable Leaderboard for Machine Learning Competitions. ICML 2015, PMLR 37, 1006–1014. https://arxiv.org/abs/1502.04585
  11. Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., & Roth, A. (2015). The reusable holdout: Preserving validity in adaptive data analysis. Science, 349(6248), 636–638. https://doi.org/10.1126/science.aaa9375
  12. Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv; ICML 2025. https://arxiv.org/abs/2412.21139
  13. Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). SWE-smith: Scaling Data for Software Engineering Agents. arXiv. https://arxiv.org/abs/2504.21798
  14. Zeng, L., Li, Y., Xiao, Y., et al. (2025). Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs. arXiv. https://arxiv.org/abs/2506.19290
  15. Cottier, B., Snodin, B., Owen, D., & Adamczewski, T. (2025). LLM inference prices have fallen rapidly but unequally across tasks. Epoch AI. https://epoch.ai/data-insights/llm-inference-price-trends
  16. Wei, Y., Duchenne, O., Copet, J., et al. (2025). SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution. arXiv; NeurIPS 2025. https://arxiv.org/abs/2502.18449
  17. Mistral AI & All Hands AI (2025). Devstral. https://mistral.ai/news/devstral
  18. Qwen Team (2025). Qwen3 Technical Report. arXiv. https://arxiv.org/abs/2505.09388
  19. Grattafiori, A., et al. (2024). The Llama 3 Herd of Models. arXiv. https://arxiv.org/abs/2407.21783
  20. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv; ICLR 2024. https://arxiv.org/abs/2310.06770
  21. Anthropic (2026). Hooks reference. Claude Code documentation. https://code.claude.com/docs/en/hooks
  22. Barr, E. T., Harman, M., McMinn, P., Shahbaz, M., & Yoo, S. (2015). The Oracle Problem in Software Testing: A Survey. IEEE Transactions on Software Engineering, 41(5), 507–525. https://doi.org/10.1109/TSE.2014.2372785
  23. Luo, Q., Hariri, F., Eloussi, L., & Marinov, D. (2014). An Empirical Analysis of Flaky Tests. FSE 2014, 643–653. https://doi.org/10.1145/2635868.2635920
  24. Micco, J. (2016). Flaky Tests at Google and How We Mitigate Them. Google Testing Blog. https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html
  25. Lamb, C., & Zacchiroli, S. (2022). Reproducible Builds: Increasing the Integrity of Software Supply Chains. IEEE Software, 39(2), 62–70. https://arxiv.org/abs/2104.06020
  26. Just, R., Jalali, D., & Ernst, M. D. (2014). Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs. ISSTA 2014, 437–440. https://doi.org/10.1145/2610384.2628055
  27. Qi, Z., Long, F., Achour, S., & Rinard, M. (2015). An Analysis of Patch Plausibility and Correctness for Generate-and-Validate Patch Generation Systems. ISSTA 2015, 24–36. https://doi.org/10.1145/2771783.2771791
  28. Smith, E. K., Barr, E. T., Le Goues, C., & Brun, Y. (2015). Is the Cure Worse Than the Disease? Overfitting in Automated Program Repair. ESEC/FSE 2015, 532–543. https://doi.org/10.1145/2786805.2786825
  29. Inozemtseva, L., & Holmes, R. (2014). Coverage Is Not Strongly Correlated with Test Suite Effectiveness. ICSE 2014, 435–445. https://doi.org/10.1145/2568225.2568271
  30. DeMillo, R. A., Lipton, R. J., & Sayward, F. G. (1978). Hints on Test Data Selection: Help for the Practicing Programmer. IEEE Computer, 11(4), 34–41. https://doi.org/10.1109/C-M.1978.218136
  31. Petrović, G., & Ivanković, M. (2018). State of Mutation Testing at Google. ICSE-SEIP 2018, 163–171. https://doi.org/10.1145/3183519.3183521
  32. Claessen, K., & Hughes, J. (2000). QuickCheck: A Lightweight Tool for Random Testing of Haskell Programs. ICFP 2000, 268–279. https://doi.org/10.1145/351240.351266
  33. Segura, S., Fraser, G., Sanchez, A. B., & Ruiz-Cortés, A. (2016). A Survey on Metamorphic Testing. IEEE Transactions on Software Engineering, 42(9), 805–824. https://doi.org/10.1109/TSE.2016.2532875
  34. McKeeman, W. M. (1998). Differential Testing for Software. Digital Technical Journal, 10(1), 100–107.
  35. Ziegler, A., Kalliamvakou, E., Li, X. A., Rice, A., Rifkin, D., Simister, S., Sittampalam, G., & Aftandilian, E. (2024). Measuring GitHub Copilot's Impact on Productivity. Communications of the ACM, 67(3), 54–63. https://doi.org/10.1145/3633453
  36. Von Arx, S., Chan, L., & Barnes, E. (2025). Recent Frontier Models Are Reward Hacking. METR. https://metr.org/blog/2025-06-05-recent-reward-hacking/
  37. MacDiarmid, M., Wright, B., Uesato, J., et al. (2025). Natural Emergent Misalignment from Reward Hacking in Production RL. arXiv. https://arxiv.org/abs/2511.18397
  38. Skalse, J., Howe, N. H. R., Krasheninnikov, D., & Krueger, D. (2022). Defining and Characterizing Reward Hacking. NeurIPS 2022. https://arxiv.org/abs/2209.13085
  39. Gao, L., Schulman, J., & Hilton, J. (2022). Scaling Laws for Reward Model Overoptimization. arXiv; ICML 2023. https://arxiv.org/abs/2210.10760
  40. Thompson, K. (1984). Reflections on Trusting Trust. Communications of the ACM, 27(8), 761–763. https://doi.org/10.1145/358198.358210
  41. Agache, A., Brooker, M., Florescu, A., Iordache, A., Liguori, A., Neugebauer, R., Piwonka, P., & Popa, D.-M. (2020). Firecracker: Lightweight Virtualization for Serverless Applications. NSDI 2020, 419–434. https://www.usenix.org/conference/nsdi20/presentation/agache
  42. Young, E. G., Zhu, P., Caraza-Harter, T., Arpaci-Dusseau, A. C., & Arpaci-Dusseau, R. H. (2019). The True Cost of Containing: A gVisor Case Study. HotCloud 2019. https://www.usenix.org/conference/hotcloud19/presentation/young
  43. Chen, X., Lin, M., Schärli, N., & Zhou, D. (2023). Teaching Large Language Models to Self-Debug. arXiv; ICLR 2024. https://arxiv.org/abs/2304.05128
  44. Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks. https://arxiv.org/abs/2306.05685
  45. Wang, P., Li, L., Chen, L., et al. (2023). Large Language Models are not Fair Evaluators. arXiv; ACL 2024. https://arxiv.org/abs/2305.17926
  46. Stechly, K., Marquez, M., & Kambhampati, S. (2023). GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems. arXiv. https://arxiv.org/abs/2310.12397
  47. Valmeekam, K., Marquez, M., & Kambhampati, S. (2023). Can Large Language Models Really Improve by Self-critiquing Their Own Plans? arXiv. https://arxiv.org/abs/2310.08118
  48. Tyen, G., Mansoor, H., Cărbune, V., Chen, P., & Mak, T. (2024). LLMs cannot find reasoning errors, but can correct them given the error location. Findings of ACL 2024. https://arxiv.org/abs/2311.08516
  49. Kamoi, R., Zhang, Y., Zhang, N., Han, J., & Zhang, R. (2024). When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs. Transactions of the ACL, 12. https://arxiv.org/abs/2406.01297
  50. Xu, W., Zhu, G., Zhao, X., Pan, L., Li, L., & Wang, W. Y. (2024). Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement. ACL 2024. https://arxiv.org/abs/2402.11436
  51. Meli, M., McNiece, M. R., & Reaves, B. (2019). How Bad Can It Git? Characterizing Secret Leakage in Public GitHub Repositories. NDSS 2019. https://doi.org/10.14722/ndss.2019.23418
  52. Anthropic (2026). Plugin manifest reference. Claude Code documentation. https://code.claude.com/docs/en/plugins-reference
  53. Nagappan, N., Maximilien, E. M., Bhat, T., & Williams, L. (2008). Realizing quality improvement through test driven development: results and experiences of four industrial teams. Empirical Software Engineering, 13(3), 289–302. https://doi.org/10.1007/s10664-008-9062-z
  54. OpenAI (2024). Introducing SWE-bench Verified. https://openai.com/index/introducing-swe-bench-verified/
  55. Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., & Mirhoseini, A. (2024). Large Language Monkeys: Scaling Inference Compute with Repeated Sampling. arXiv. https://arxiv.org/abs/2407.21787
  56. Touvron, H., et al. (2023). Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv. https://arxiv.org/abs/2307.09288
  57. Just, R., Jalali, D., Inozemtseva, L., Ernst, M. D., Holmes, R., & Fraser, G. (2014). Are Mutants a Valid Substitute for Real Faults in Software Testing? FSE 2014, 654–665. https://doi.org/10.1145/2635868.2635929
  58. Badertdinov, I., Golubev, A., Nekrashevich, M., et al. (2025). SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents. arXiv; NeurIPS 2025. https://arxiv.org/abs/2505.20411
  59. Zhao, A., Wu, Y., Yue, Y., et al. (2025). Absolute Zero: Reinforced Self-play Reasoning with Zero Data. arXiv. https://arxiv.org/abs/2505.03335
  60. Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the Knowledge in a Neural Network. NeurIPS 2014 Deep Learning Workshop. https://arxiv.org/abs/1503.02531
  61. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. https://arxiv.org/abs/2305.18290
  62. Andrychowicz, M., Wolski, F., Ray, A., et al. (2017). Hindsight Experience Replay. NeurIPS 2017. https://arxiv.org/abs/1707.01495
  63. Anthony, T., Tian, Z., & Barber, D. (2017). Thinking Fast and Slow with Deep Learning and Tree Search. NeurIPS 2017. https://arxiv.org/abs/1705.08439
  64. Silver, D., Schrittwieser, J., Simonyan, K., et al. (2017). Mastering the game of Go without human knowledge. Nature, 550, 354–359. https://doi.org/10.1038/nature24270
  65. Gulcehre, C., Le Paine, T., Srinivasan, S., et al. (2023). Reinforced Self-Training (ReST) for Language Modeling. arXiv. https://arxiv.org/abs/2308.08998
  66. Singh, A., Co-Reyes, J. D., Agarwal, R., et al. (2023). Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models. arXiv; TMLR 2024. https://arxiv.org/abs/2312.06585
  67. Li, Y., Choi, D., Chung, J., et al. (2022). Competition-level code generation with AlphaCode. Science, 378(6624), 1092–1097. https://doi.org/10.1126/science.abq1158
  68. Singhal, P., Goyal, T., Xu, J., & Durrett, G. (2023). A Long Way to Go: Investigating Length Correlations in RLHF. arXiv; COLM 2024. https://arxiv.org/abs/2310.03716
  69. Lambert, N., Morrison, J., Pyatkin, V., et al. (2024). Tülu 3: Pushing Frontiers in Open Language Model Post-Training. arXiv. https://arxiv.org/abs/2411.15124
  70. Agentica & Together AI (2025). DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL. https://www.together.ai/blog/deepswe
  71. Golubev, A., Trofimova, M., Polezhaev, S., et al. (2025). Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning. arXiv. https://arxiv.org/abs/2508.03501
  72. Yu, Q., Zhang, Z., Zhu, R., et al. (2025). DAPO: An Open-Source LLM Reinforcement Learning System at Scale. arXiv. https://arxiv.org/abs/2503.14476
  73. Qwen Team (2025). Qwen3-Coder: Agentic Coding in the World. https://qwenlm.github.io/blog/qwen3-coder/
  74. Yue, Y., Chen, Z., Lu, R., et al. (2025). Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? arXiv; NeurIPS 2025. https://arxiv.org/abs/2504.13837
  75. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv; ICLR 2022. https://arxiv.org/abs/2106.09685
  76. Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient Finetuning of Quantized LLMs. NeurIPS 2023. https://arxiv.org/abs/2305.14314
  77. Biderman, D., Portes, J., Gonzalez Ortiz, J. J., et al. (2024). LoRA Learns Less and Forgets Less. Transactions on Machine Learning Research. https://arxiv.org/abs/2405.09673
  78. Schulman, J., & Thinking Machines Lab (2025). LoRA Without Regret. https://thinkingmachines.ai/blog/lora/
  79. McCloskey, M., & Cohen, N. J. (1989). Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem. Psychology of Learning and Motivation, 24, 109–165. https://doi.org/10.1016/S0079-7421(08)60536-8
  80. Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., & Zhang, Y. (2023). An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning. arXiv. https://arxiv.org/abs/2308.08747
  81. Ibrahim, A., Thérien, B., Gupta, K., et al. (2024). Simple and Scalable Strategies to Continually Pre-train Large Language Models. arXiv; TMLR. https://arxiv.org/abs/2403.08763
  82. Wortsman, M., Ilharco, G., Gadre, S. Y., et al. (2022). Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. ICML 2022. https://arxiv.org/abs/2203.05482
  83. Yadav, P., Tam, D., Choshen, L., Raffel, C., & Bansal, M. (2023). TIES-Merging: Resolving Interference When Merging Models. NeurIPS 2023. https://arxiv.org/abs/2306.01708
  84. Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755–759. https://doi.org/10.1038/s41586-024-07566-y
  85. Gerstgrasser, M., Schaeffer, R., Dey, A., et al. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv. https://arxiv.org/abs/2404.01413
  86. Shao, R., Li, S. S., Xin, R., et al. (2025). Spurious Rewards: Rethinking Training Signals in RLVR. arXiv. https://arxiv.org/abs/2506.10947
  87. Zhou, C., Liu, P., Xu, P., et al. (2023). LIMA: Less Is More for Alignment. NeurIPS 2023. https://arxiv.org/abs/2305.11206
  88. Muennighoff, N., Yang, Z., Shi, W., et al. (2025). s1: Simple test-time scaling. arXiv. https://arxiv.org/abs/2501.19393
  89. Ye, Y., Huang, Z., Xiao, Y., Chern, E., Xia, S., & Liu, P. (2025). LIMO: Less is More for Reasoning. arXiv (v1, February 2025); COLM 2025. https://arxiv.org/abs/2502.03387
  90. Chen, B., Shu, C., Shareghi, E., Collier, N., Narasimhan, K., & Yao, S. (2023). FireAct: Toward Language Agent Fine-tuning. arXiv. https://arxiv.org/abs/2310.05915
  91. Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., & Tang, J. (2023). AgentTuning: Enabling Generalized Agent Abilities for LLMs. arXiv. https://arxiv.org/abs/2310.12823
  92. Zhang, B., Liu, Z., Cherry, C., & Firat, O. (2024). When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method. ICLR 2024. https://arxiv.org/abs/2402.17193
  93. Chen, L., Li, S., Yan, J., et al. (2023). AlpaGasus: Training a Better Alpaca with Fewer Data. arXiv; ICLR 2024. https://arxiv.org/abs/2307.08701
  94. Humble, J., & Farley, D. (2010). Continuous Delivery: Reliable Software Releases through Build, Test, and Deployment Automation. Addison-Wesley.
  95. Deng, C., Zhao, Y., Tang, X., Gerstein, M., & Cohan, A. (2024). Investigating Data Contamination in Modern Benchmarks for Large Language Models. NAACL 2024. https://arxiv.org/abs/2311.09783
  96. Aleithan, R., Xue, H., Mohajer, M. M., Nnorom, E., Uddin, G., & Wang, S. (2024). SWE-Bench+: Enhanced Coding Benchmark for LLMs. arXiv. https://arxiv.org/abs/2410.06992
  97. Zan, D., Huang, Z., Liu, W., et al. (2025). Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving. arXiv. https://arxiv.org/abs/2504.02605
  98. Kwon, W., Li, Z., Zhuang, S., et al. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention. SOSP 2023. https://arxiv.org/abs/2309.06180
  99. Sheng, Y., Cao, S., Li, D., et al. (2023). S-LoRA: Serving Thousands of Concurrent LoRA Adapters. arXiv; MLSys 2024. https://arxiv.org/abs/2311.03285
  100. Carlini, N., Tramèr, F., Wallace, E., et al. (2021). Extracting Training Data from Large Language Models. USENIX Security 2021. https://arxiv.org/abs/2012.07805
  101. Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramèr, F., & Zhang, C. (2023). Quantifying Memorization Across Neural Language Models. ICLR 2023. https://arxiv.org/abs/2202.07646
  102. Sculley, D., Holt, G., Golovin, D., et al. (2015). Hidden Technical Debt in Machine Learning Systems. NeurIPS 2015.
  103. Cobbe, K., Kosaraju, V., Bavarian, M., et al. (2021). Training Verifiers to Solve Math Word Problems. arXiv. https://arxiv.org/abs/2110.14168
  104. Le, H., Wang, Y., Gotmare, A. D., Savarese, S., & Hoi, S. C. H. (2022). CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning. NeurIPS 2022. https://arxiv.org/abs/2207.01780
  105. Gehring, J., Zheng, K., Copet, J., Mella, V., Cohen, T., & Synnaeve, G. (2024). RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning. arXiv. https://arxiv.org/abs/2410.02089
  106. Jain, N., Singh, J., Shetty, M., Zheng, L., Sen, K., & Stoica, I. (2025). R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents. arXiv. https://arxiv.org/abs/2504.07164
  107. Miserendino, S., Wang, M., Patwardhan, T., & Heidecke, J. (2025). SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? arXiv. https://arxiv.org/abs/2502.12115
  108. Bader, J., Scott, A., Pradel, M., & Chandra, S. (2019). Getafix: Learning to Fix Bugs Automatically. Proceedings of the ACM on Programming Languages, 3(OOPSLA), Article 159. https://doi.org/10.1145/3360585
  109. Marginean, A., Bader, J., Chandra, S., Harman, M., Jia, Y., Mao, K., Mols, A., & Scott, A. (2019). SapFix: Automated End-to-End Repair at Scale. ICSE-SEIP 2019, 269–278. https://doi.org/10.1109/ICSE-SEIP.2019.00039
  110. Maniatis, P., & Tarlow, D. (2023). Large sequence models for software development activities. Google Research Blog. https://research.google/blog/large-sequence-models-for-software-development-activities/
  111. Tabachnyk, M., & Nikolov, S. (2022). ML-Enhanced Code Completion Improves Developer Productivity. Google Research Blog. https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/
  112. Singhal, M., Carelli, R., Segato, G., Kumar, V., & Catasta, M. (2024). Building LLMs for Code Repair. Replit. https://replit.com/blog/code-repair
  113. Bojarski, M., Del Testa, D., Dworakowski, D., et al. (2016). End to End Learning for Self-Driving Cars. arXiv. https://arxiv.org/abs/1604.07316
  114. Ross, S., Gordon, G. J., & Bagnell, J. A. (2011). A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS 2011. https://arxiv.org/abs/1011.0686
  115. Hindle, A., Barr, E. T., Su, Z., Gabel, M., & Devanbu, P. (2012). On the Naturalness of Software. ICSE 2012, 837–847. https://doi.org/10.1109/ICSE.2012.6227135
  116. Allamanis, M., Barr, E. T., Devanbu, P., & Sutton, C. (2018). A Survey of Machine Learning for Big Code and Naturalness. ACM Computing Surveys, 51(4), Article 81. https://arxiv.org/abs/1709.06182
  117. Kimi Team (2025). Kimi K2: Open Agentic Intelligence. arXiv. https://arxiv.org/abs/2507.20534
  118. Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. NeurIPS 2024. https://arxiv.org/abs/2405.15793
  119. Wang, X., Li, B., Song, Y., et al. (2024). OpenHands: An Open Platform for AI Software Developers as Generalist Agents. arXiv; ICLR 2025. https://arxiv.org/abs/2407.16741

Cite this

Anderson, M. (2026). Sources. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Sources). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/sources

BibTeX
@incollection{anderson2026continuousdeliveryof,
  author    = {Anderson, Mac},
  title     = {Sources},
  booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
  year      = {2026},
  url       = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/sources}
}

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.