Part VII: Back matter
Sources
The works cited, in order of first citation.
Mac Anderson119 sources
Sources of 22
119 works are cited in this book. They appear below in the order of their first citation. Each chapter also lists its own sources at its end.
- Cottier, B., You, J., Martemianova, N., & Owen, D. (2024). How far behind are open models? Epoch AI. https://epoch.ai/blog/open-models-report
- Stanford Institute for Human-Centered Artificial Intelligence (2025). AI Index Report 2025, Chapter 2: Technical Performance. https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance
- OpenAI (2025). gpt-oss-120b and gpt-oss-20b Model Card. arXiv. https://arxiv.org/abs/2508.10925
- Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv; ICLR 2024. https://arxiv.org/abs/2310.01798
- Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM Evaluators Recognize and Favor Their Own Generations. arXiv; NeurIPS 2024. https://arxiv.org/abs/2404.13076
- Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). STaR: Bootstrapping Reasoning With Reasoning. arXiv. https://arxiv.org/abs/2203.14465
- DeepSeek-AI (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv. https://arxiv.org/abs/2501.12948
- Baker, B., Huizinga, J., Gao, L., et al. (2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. arXiv. https://arxiv.org/abs/2503.11926
- Denison, C., MacDiarmid, M., Barez, F., et al. (2024). Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models. arXiv. https://arxiv.org/abs/2406.10162
- Blum, A., & Hardt, M. (2015). The Ladder: A Reliable Leaderboard for Machine Learning Competitions. ICML 2015, PMLR 37, 1006–1014. https://arxiv.org/abs/1502.04585
- Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., & Roth, A. (2015). The reusable holdout: Preserving validity in adaptive data analysis. Science, 349(6248), 636–638. https://doi.org/10.1126/science.aaa9375
- Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv; ICML 2025. https://arxiv.org/abs/2412.21139
- Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). SWE-smith: Scaling Data for Software Engineering Agents. arXiv. https://arxiv.org/abs/2504.21798
- Zeng, L., Li, Y., Xiao, Y., et al. (2025). Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs. arXiv. https://arxiv.org/abs/2506.19290
- Cottier, B., Snodin, B., Owen, D., & Adamczewski, T. (2025). LLM inference prices have fallen rapidly but unequally across tasks. Epoch AI. https://epoch.ai/data-insights/llm-inference-price-trends
- Wei, Y., Duchenne, O., Copet, J., et al. (2025). SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution. arXiv; NeurIPS 2025. https://arxiv.org/abs/2502.18449
- Mistral AI & All Hands AI (2025). Devstral. https://mistral.ai/news/devstral
- Qwen Team (2025). Qwen3 Technical Report. arXiv. https://arxiv.org/abs/2505.09388
- Grattafiori, A., et al. (2024). The Llama 3 Herd of Models. arXiv. https://arxiv.org/abs/2407.21783
- Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv; ICLR 2024. https://arxiv.org/abs/2310.06770
- Anthropic (2026). Hooks reference. Claude Code documentation. https://code.claude.com/docs/en/hooks
- Barr, E. T., Harman, M., McMinn, P., Shahbaz, M., & Yoo, S. (2015). The Oracle Problem in Software Testing: A Survey. IEEE Transactions on Software Engineering, 41(5), 507–525. https://doi.org/10.1109/TSE.2014.2372785
- Luo, Q., Hariri, F., Eloussi, L., & Marinov, D. (2014). An Empirical Analysis of Flaky Tests. FSE 2014, 643–653. https://doi.org/10.1145/2635868.2635920
- Micco, J. (2016). Flaky Tests at Google and How We Mitigate Them. Google Testing Blog. https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html
- Lamb, C., & Zacchiroli, S. (2022). Reproducible Builds: Increasing the Integrity of Software Supply Chains. IEEE Software, 39(2), 62–70. https://arxiv.org/abs/2104.06020
- Just, R., Jalali, D., & Ernst, M. D. (2014). Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs. ISSTA 2014, 437–440. https://doi.org/10.1145/2610384.2628055
- Qi, Z., Long, F., Achour, S., & Rinard, M. (2015). An Analysis of Patch Plausibility and Correctness for Generate-and-Validate Patch Generation Systems. ISSTA 2015, 24–36. https://doi.org/10.1145/2771783.2771791
- Smith, E. K., Barr, E. T., Le Goues, C., & Brun, Y. (2015). Is the Cure Worse Than the Disease? Overfitting in Automated Program Repair. ESEC/FSE 2015, 532–543. https://doi.org/10.1145/2786805.2786825
- Inozemtseva, L., & Holmes, R. (2014). Coverage Is Not Strongly Correlated with Test Suite Effectiveness. ICSE 2014, 435–445. https://doi.org/10.1145/2568225.2568271
- DeMillo, R. A., Lipton, R. J., & Sayward, F. G. (1978). Hints on Test Data Selection: Help for the Practicing Programmer. IEEE Computer, 11(4), 34–41. https://doi.org/10.1109/C-M.1978.218136
- Petrović, G., & Ivanković, M. (2018). State of Mutation Testing at Google. ICSE-SEIP 2018, 163–171. https://doi.org/10.1145/3183519.3183521
- Claessen, K., & Hughes, J. (2000). QuickCheck: A Lightweight Tool for Random Testing of Haskell Programs. ICFP 2000, 268–279. https://doi.org/10.1145/351240.351266
- Segura, S., Fraser, G., Sanchez, A. B., & Ruiz-Cortés, A. (2016). A Survey on Metamorphic Testing. IEEE Transactions on Software Engineering, 42(9), 805–824. https://doi.org/10.1109/TSE.2016.2532875
- McKeeman, W. M. (1998). Differential Testing for Software. Digital Technical Journal, 10(1), 100–107.
- Ziegler, A., Kalliamvakou, E., Li, X. A., Rice, A., Rifkin, D., Simister, S., Sittampalam, G., & Aftandilian, E. (2024). Measuring GitHub Copilot's Impact on Productivity. Communications of the ACM, 67(3), 54–63. https://doi.org/10.1145/3633453
- Von Arx, S., Chan, L., & Barnes, E. (2025). Recent Frontier Models Are Reward Hacking. METR. https://metr.org/blog/2025-06-05-recent-reward-hacking/
- MacDiarmid, M., Wright, B., Uesato, J., et al. (2025). Natural Emergent Misalignment from Reward Hacking in Production RL. arXiv. https://arxiv.org/abs/2511.18397
- Skalse, J., Howe, N. H. R., Krasheninnikov, D., & Krueger, D. (2022). Defining and Characterizing Reward Hacking. NeurIPS 2022. https://arxiv.org/abs/2209.13085
- Gao, L., Schulman, J., & Hilton, J. (2022). Scaling Laws for Reward Model Overoptimization. arXiv; ICML 2023. https://arxiv.org/abs/2210.10760
- Thompson, K. (1984). Reflections on Trusting Trust. Communications of the ACM, 27(8), 761–763. https://doi.org/10.1145/358198.358210
- Agache, A., Brooker, M., Florescu, A., Iordache, A., Liguori, A., Neugebauer, R., Piwonka, P., & Popa, D.-M. (2020). Firecracker: Lightweight Virtualization for Serverless Applications. NSDI 2020, 419–434. https://www.usenix.org/conference/nsdi20/presentation/agache
- Young, E. G., Zhu, P., Caraza-Harter, T., Arpaci-Dusseau, A. C., & Arpaci-Dusseau, R. H. (2019). The True Cost of Containing: A gVisor Case Study. HotCloud 2019. https://www.usenix.org/conference/hotcloud19/presentation/young
- Chen, X., Lin, M., Schärli, N., & Zhou, D. (2023). Teaching Large Language Models to Self-Debug. arXiv; ICLR 2024. https://arxiv.org/abs/2304.05128
- Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks. https://arxiv.org/abs/2306.05685
- Wang, P., Li, L., Chen, L., et al. (2023). Large Language Models are not Fair Evaluators. arXiv; ACL 2024. https://arxiv.org/abs/2305.17926
- Stechly, K., Marquez, M., & Kambhampati, S. (2023). GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems. arXiv. https://arxiv.org/abs/2310.12397
- Valmeekam, K., Marquez, M., & Kambhampati, S. (2023). Can Large Language Models Really Improve by Self-critiquing Their Own Plans? arXiv. https://arxiv.org/abs/2310.08118
- Tyen, G., Mansoor, H., Cărbune, V., Chen, P., & Mak, T. (2024). LLMs cannot find reasoning errors, but can correct them given the error location. Findings of ACL 2024. https://arxiv.org/abs/2311.08516
- Kamoi, R., Zhang, Y., Zhang, N., Han, J., & Zhang, R. (2024). When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs. Transactions of the ACL, 12. https://arxiv.org/abs/2406.01297
- Xu, W., Zhu, G., Zhao, X., Pan, L., Li, L., & Wang, W. Y. (2024). Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement. ACL 2024. https://arxiv.org/abs/2402.11436
- Meli, M., McNiece, M. R., & Reaves, B. (2019). How Bad Can It Git? Characterizing Secret Leakage in Public GitHub Repositories. NDSS 2019. https://doi.org/10.14722/ndss.2019.23418
- Anthropic (2026). Plugin manifest reference. Claude Code documentation. https://code.claude.com/docs/en/plugins-reference
- Nagappan, N., Maximilien, E. M., Bhat, T., & Williams, L. (2008). Realizing quality improvement through test driven development: results and experiences of four industrial teams. Empirical Software Engineering, 13(3), 289–302. https://doi.org/10.1007/s10664-008-9062-z
- OpenAI (2024). Introducing SWE-bench Verified. https://openai.com/index/introducing-swe-bench-verified/
- Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., & Mirhoseini, A. (2024). Large Language Monkeys: Scaling Inference Compute with Repeated Sampling. arXiv. https://arxiv.org/abs/2407.21787
- Touvron, H., et al. (2023). Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv. https://arxiv.org/abs/2307.09288
- Just, R., Jalali, D., Inozemtseva, L., Ernst, M. D., Holmes, R., & Fraser, G. (2014). Are Mutants a Valid Substitute for Real Faults in Software Testing? FSE 2014, 654–665. https://doi.org/10.1145/2635868.2635929
- Badertdinov, I., Golubev, A., Nekrashevich, M., et al. (2025). SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents. arXiv; NeurIPS 2025. https://arxiv.org/abs/2505.20411
- Zhao, A., Wu, Y., Yue, Y., et al. (2025). Absolute Zero: Reinforced Self-play Reasoning with Zero Data. arXiv. https://arxiv.org/abs/2505.03335
- Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the Knowledge in a Neural Network. NeurIPS 2014 Deep Learning Workshop. https://arxiv.org/abs/1503.02531
- Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. https://arxiv.org/abs/2305.18290
- Andrychowicz, M., Wolski, F., Ray, A., et al. (2017). Hindsight Experience Replay. NeurIPS 2017. https://arxiv.org/abs/1707.01495
- Anthony, T., Tian, Z., & Barber, D. (2017). Thinking Fast and Slow with Deep Learning and Tree Search. NeurIPS 2017. https://arxiv.org/abs/1705.08439
- Silver, D., Schrittwieser, J., Simonyan, K., et al. (2017). Mastering the game of Go without human knowledge. Nature, 550, 354–359. https://doi.org/10.1038/nature24270
- Gulcehre, C., Le Paine, T., Srinivasan, S., et al. (2023). Reinforced Self-Training (ReST) for Language Modeling. arXiv. https://arxiv.org/abs/2308.08998
- Singh, A., Co-Reyes, J. D., Agarwal, R., et al. (2023). Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models. arXiv; TMLR 2024. https://arxiv.org/abs/2312.06585
- Li, Y., Choi, D., Chung, J., et al. (2022). Competition-level code generation with AlphaCode. Science, 378(6624), 1092–1097. https://doi.org/10.1126/science.abq1158
- Singhal, P., Goyal, T., Xu, J., & Durrett, G. (2023). A Long Way to Go: Investigating Length Correlations in RLHF. arXiv; COLM 2024. https://arxiv.org/abs/2310.03716
- Lambert, N., Morrison, J., Pyatkin, V., et al. (2024). Tülu 3: Pushing Frontiers in Open Language Model Post-Training. arXiv. https://arxiv.org/abs/2411.15124
- Agentica & Together AI (2025). DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL. https://www.together.ai/blog/deepswe
- Golubev, A., Trofimova, M., Polezhaev, S., et al. (2025). Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning. arXiv. https://arxiv.org/abs/2508.03501
- Yu, Q., Zhang, Z., Zhu, R., et al. (2025). DAPO: An Open-Source LLM Reinforcement Learning System at Scale. arXiv. https://arxiv.org/abs/2503.14476
- Qwen Team (2025). Qwen3-Coder: Agentic Coding in the World. https://qwenlm.github.io/blog/qwen3-coder/
- Yue, Y., Chen, Z., Lu, R., et al. (2025). Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? arXiv; NeurIPS 2025. https://arxiv.org/abs/2504.13837
- Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv; ICLR 2022. https://arxiv.org/abs/2106.09685
- Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient Finetuning of Quantized LLMs. NeurIPS 2023. https://arxiv.org/abs/2305.14314
- Biderman, D., Portes, J., Gonzalez Ortiz, J. J., et al. (2024). LoRA Learns Less and Forgets Less. Transactions on Machine Learning Research. https://arxiv.org/abs/2405.09673
- Schulman, J., & Thinking Machines Lab (2025). LoRA Without Regret. https://thinkingmachines.ai/blog/lora/
- McCloskey, M., & Cohen, N. J. (1989). Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem. Psychology of Learning and Motivation, 24, 109–165. https://doi.org/10.1016/S0079-7421(08)60536-8
- Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., & Zhang, Y. (2023). An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning. arXiv. https://arxiv.org/abs/2308.08747
- Ibrahim, A., Thérien, B., Gupta, K., et al. (2024). Simple and Scalable Strategies to Continually Pre-train Large Language Models. arXiv; TMLR. https://arxiv.org/abs/2403.08763
- Wortsman, M., Ilharco, G., Gadre, S. Y., et al. (2022). Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. ICML 2022. https://arxiv.org/abs/2203.05482
- Yadav, P., Tam, D., Choshen, L., Raffel, C., & Bansal, M. (2023). TIES-Merging: Resolving Interference When Merging Models. NeurIPS 2023. https://arxiv.org/abs/2306.01708
- Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755–759. https://doi.org/10.1038/s41586-024-07566-y
- Gerstgrasser, M., Schaeffer, R., Dey, A., et al. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv. https://arxiv.org/abs/2404.01413
- Shao, R., Li, S. S., Xin, R., et al. (2025). Spurious Rewards: Rethinking Training Signals in RLVR. arXiv. https://arxiv.org/abs/2506.10947
- Zhou, C., Liu, P., Xu, P., et al. (2023). LIMA: Less Is More for Alignment. NeurIPS 2023. https://arxiv.org/abs/2305.11206
- Muennighoff, N., Yang, Z., Shi, W., et al. (2025). s1: Simple test-time scaling. arXiv. https://arxiv.org/abs/2501.19393
- Ye, Y., Huang, Z., Xiao, Y., Chern, E., Xia, S., & Liu, P. (2025). LIMO: Less is More for Reasoning. arXiv (v1, February 2025); COLM 2025. https://arxiv.org/abs/2502.03387
- Chen, B., Shu, C., Shareghi, E., Collier, N., Narasimhan, K., & Yao, S. (2023). FireAct: Toward Language Agent Fine-tuning. arXiv. https://arxiv.org/abs/2310.05915
- Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., & Tang, J. (2023). AgentTuning: Enabling Generalized Agent Abilities for LLMs. arXiv. https://arxiv.org/abs/2310.12823
- Zhang, B., Liu, Z., Cherry, C., & Firat, O. (2024). When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method. ICLR 2024. https://arxiv.org/abs/2402.17193
- Chen, L., Li, S., Yan, J., et al. (2023). AlpaGasus: Training a Better Alpaca with Fewer Data. arXiv; ICLR 2024. https://arxiv.org/abs/2307.08701
- Humble, J., & Farley, D. (2010). Continuous Delivery: Reliable Software Releases through Build, Test, and Deployment Automation. Addison-Wesley.
- Deng, C., Zhao, Y., Tang, X., Gerstein, M., & Cohan, A. (2024). Investigating Data Contamination in Modern Benchmarks for Large Language Models. NAACL 2024. https://arxiv.org/abs/2311.09783
- Aleithan, R., Xue, H., Mohajer, M. M., Nnorom, E., Uddin, G., & Wang, S. (2024). SWE-Bench+: Enhanced Coding Benchmark for LLMs. arXiv. https://arxiv.org/abs/2410.06992
- Zan, D., Huang, Z., Liu, W., et al. (2025). Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving. arXiv. https://arxiv.org/abs/2504.02605
- Kwon, W., Li, Z., Zhuang, S., et al. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention. SOSP 2023. https://arxiv.org/abs/2309.06180
- Sheng, Y., Cao, S., Li, D., et al. (2023). S-LoRA: Serving Thousands of Concurrent LoRA Adapters. arXiv; MLSys 2024. https://arxiv.org/abs/2311.03285
- Carlini, N., Tramèr, F., Wallace, E., et al. (2021). Extracting Training Data from Large Language Models. USENIX Security 2021. https://arxiv.org/abs/2012.07805
- Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramèr, F., & Zhang, C. (2023). Quantifying Memorization Across Neural Language Models. ICLR 2023. https://arxiv.org/abs/2202.07646
- Sculley, D., Holt, G., Golovin, D., et al. (2015). Hidden Technical Debt in Machine Learning Systems. NeurIPS 2015.
- Cobbe, K., Kosaraju, V., Bavarian, M., et al. (2021). Training Verifiers to Solve Math Word Problems. arXiv. https://arxiv.org/abs/2110.14168
- Le, H., Wang, Y., Gotmare, A. D., Savarese, S., & Hoi, S. C. H. (2022). CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning. NeurIPS 2022. https://arxiv.org/abs/2207.01780
- Gehring, J., Zheng, K., Copet, J., Mella, V., Cohen, T., & Synnaeve, G. (2024). RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning. arXiv. https://arxiv.org/abs/2410.02089
- Jain, N., Singh, J., Shetty, M., Zheng, L., Sen, K., & Stoica, I. (2025). R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents. arXiv. https://arxiv.org/abs/2504.07164
- Miserendino, S., Wang, M., Patwardhan, T., & Heidecke, J. (2025). SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? arXiv. https://arxiv.org/abs/2502.12115
- Bader, J., Scott, A., Pradel, M., & Chandra, S. (2019). Getafix: Learning to Fix Bugs Automatically. Proceedings of the ACM on Programming Languages, 3(OOPSLA), Article 159. https://doi.org/10.1145/3360585
- Marginean, A., Bader, J., Chandra, S., Harman, M., Jia, Y., Mao, K., Mols, A., & Scott, A. (2019). SapFix: Automated End-to-End Repair at Scale. ICSE-SEIP 2019, 269–278. https://doi.org/10.1109/ICSE-SEIP.2019.00039
- Maniatis, P., & Tarlow, D. (2023). Large sequence models for software development activities. Google Research Blog. https://research.google/blog/large-sequence-models-for-software-development-activities/
- Tabachnyk, M., & Nikolov, S. (2022). ML-Enhanced Code Completion Improves Developer Productivity. Google Research Blog. https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/
- Singhal, M., Carelli, R., Segato, G., Kumar, V., & Catasta, M. (2024). Building LLMs for Code Repair. Replit. https://replit.com/blog/code-repair
- Bojarski, M., Del Testa, D., Dworakowski, D., et al. (2016). End to End Learning for Self-Driving Cars. arXiv. https://arxiv.org/abs/1604.07316
- Ross, S., Gordon, G. J., & Bagnell, J. A. (2011). A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS 2011. https://arxiv.org/abs/1011.0686
- Hindle, A., Barr, E. T., Su, Z., Gabel, M., & Devanbu, P. (2012). On the Naturalness of Software. ICSE 2012, 837–847. https://doi.org/10.1109/ICSE.2012.6227135
- Allamanis, M., Barr, E. T., Devanbu, P., & Sutton, C. (2018). A Survey of Machine Learning for Big Code and Naturalness. ACM Computing Surveys, 51(4), Article 81. https://arxiv.org/abs/1709.06182
- Kimi Team (2025). Kimi K2: Open Agentic Intelligence. arXiv. https://arxiv.org/abs/2507.20534
- Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. NeurIPS 2024. https://arxiv.org/abs/2405.15793
- Wang, X., Li, B., Song, Y., et al. (2024). OpenHands: An Open Platform for AI Software Developers as Generalist Agents. arXiv; ICLR 2025. https://arxiv.org/abs/2407.16741
Cite this
Anderson, M. (2026). Sources. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Sources). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/sources
BibTeX
@incollection{anderson2026continuousdeliveryof,
author = {Anderson, Mac},
title = {Sources},
booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
year = {2026},
url = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/sources}
}Updates by email
New research reaches subscribers first.