← back

CSCI 8710: Modern Software Development Methodologies

Course Description

This graduate Software Engineering course introduces students to advanced software engineering methodologies, modern software development paradigms, and intelligent techniques for building, analyzing, and maintaining software systems. The course is intended for graduate students in computer science and related fields who are interested in modern software engineering methods and intelligent software development tools. While prior exposure to software engineering is helpful, students from other departments with basic programming experience are encouraged to enroll and will be introduced to the necessary foundations during the course.

The course covers advanced methodologies for designing, developing, analyzing, and maintaining modern software systems. Topics include advanced object-oriented software development, component-based software engineering, software architecture and design patterns, software quality assurance and testing, and software evolution and maintenance. Building on these foundations, the course introduces program analysis techniques such as abstract syntax trees, program parsing, semantic analysis and name binding, and data-flow analysis, which provide the technical basis for understanding and constructing intelligent software engineering tools.

A major focus of the course is LLM4Code and intelligent coding tasks. Students study how large language models and code language models complement and extend traditional software engineering techniques to support code completion, code generation, code search, code summarization, automated testing, program repair, vulnerability detection, and developer-assistance tools. Through research paper discussions and hands-on projects, students examine both the capabilities and limitations of intelligent coding systems and explore how program analysis, software design principles, and LLM-based approaches can be integrated to build practical software engineering tools.

Student Learning Outcomes

Major Topics

The major topics central to this course include advanced software engineering, program analysis, IDE-integrated development tools, and LLM4Code. Students connect software engineering foundations with intelligent coding techniques through the following themes.

Course Evaluation

Student performance will be evaluated through homework, quizzes, programming assignments, research paper activities, class participation, and a term project. These components assess students' mastery of advanced software engineering concepts, their ability to apply program analysis and intelligent coding techniques, and their capacity to evaluate current research in software engineering and LLM4Code.

Evaluation Component Weight
Homework, Quizzes, and In-Class Exercises 15%
Midterm Examination 20%
Final Examination 20%
Research Paper Presentations and Scholarly Participation 30%
Term Project 15%
Total 100%

Homework, quizzes, and in-class exercises reinforce lecture topics and assess students' understanding of software engineering concepts, program analysis techniques, and intelligent software engineering methods. Programming assignments may involve source-code parsing, AST-based analysis, static analysis, def-use or data-flow analysis, IDE extension development, or LLM-based coding assistance.

The research paper presentation requires students to analyze and present selected papers from recent software engineering conferences and journals. Assessment considers technical accuracy, clarity of presentation, quality of slides, understanding of the research problem and methodology, and the ability to lead scholarly discussion. Class participation includes active engagement in discussions, submission of paper discussion questions, and constructive interaction with peer presentations.

The term project may be completed individually or in small teams. Projects should address a topic related to advanced software engineering, program analysis, LLM4Code, or intelligent software engineering tools. Typical deliverables include a project proposal, progress presentation, final presentation or demonstration, and a final research-style report containing the problem statement, motivating example, related work, proposed approach, evaluation plan or results, and conclusion.

Grading Scale

The final course grade will be assigned according to the following scale.

Percentage Letter Grade
97–100A+
94–96A
90–93A-
87–89B+
84–86B
80–83B-
77–79C+
74–76C
70–73C-
67–69D+
64–66D
60–63D-

Bibliography

The bibliography focuses on LLMs for code, intelligent coding, code generation, program repair, testing, code search, code representation, and related software engineering tasks.

Code Models and Representations

  1. [ICSE, 2019] Jian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun, Kaixuan Wang, and Xudong Liu. 2019. A Novel Neural Source Code Representation Based on Abstract Syntax Tree. In 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE). IEEE, 783–794.
  2. [TOSEM, 2020] Wenhan Wang, Ge Li, Sijie Shen, Xin Xia, and Zhi Jin. 2020. Modular Tree Network for Source Code Representation Learning. ACM Transactions on Software Engineering and Methodology (TOSEM) 29, 4 (2020), 1–23.
  3. [ICSE, 2021] Jinfeng Lin, Yalin Liu, Qingkai Zeng, Meng Jiang, and Jane Cleland-Huang. 2021. Traceability Transformed: Generating More Accurate Links with Pre-Trained BERT Models. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE). IEEE, 324–335.
  4. [ICSE, 2022] Changan Niu, Chuanyi Li, Vincent Ng, Jidong Ge, Liguo Huang, and Bin Luo. 2022. SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code Representations. In Proceedings of the 44th International Conference on Software Engineering. 2006–2018.
  5. [TSE, 2022] Julian Von der Mosel, Alexander Trautsch, and Steffen Herbold. 2022. On the Validity of Pre-Trained Transformers for Natural Language Processing in the Software Engineering Domain. IEEE Transactions on Software Engineering 49, 4 (2022), 1487–1507.
  6. [FSE, 2023] Yali Du and Zhongxing Yu. 2023. Pre-Training Code Representation with Semantic Flow Graph for Effective Bug Localization. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on The Foundations of Software Engineering. 579–591.

Code Generation and Completion

  1. [TSE, 2021] Matteo Ciniselli, Nathan Cooper, Luca Pascarella, Antonio Mastropaolo, Emad Aghajani, Denys Poshyvanyk, Massimiliano Di Penta, and Gabriele Bavota. 2021. An Empirical Study On the Usage of Transformer Models for Code Completion. IEEE Transactions on Software Engineering 48, 12 (2021), 4818–4837.
  2. [ICSE, 2022] Maliheh Izadi, Roberta Gismondi, and Georgios Gousios. 2022. CodeFill: Multi-Token Code Completion by Jointly Learning from Structure and Naming Sequences. In Proceedings of the 44th International Conference on Software Engineering. 401–412.
  3. [ICSE, 2022] Naman Jain, Skanda Vaidyanath, Arun Iyer, Nagarajan Natarajan, Suresh Parthasarathy, Sriram Rajamani, and Rahul Sharma. 2022. Jigsaw: Large Language Models Meet Program Synthesis. In Proceedings of the 44th International Conference on Software Engineering. 1219–1231.
  4. [FSE, 2023] Shangwen Wang, Mingyang Geng, Bo Lin, Zhensu Sun, Ming Wen, Yepang Liu, Li Li, Tegawendé F Bissyandé, and Xiaoguang Mao. 2023. Natural Language to Code: How Far Are We? In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 375–387.
  5. [TSE, 2023] Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q Feldman, et al. 2023. MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation. IEEE Transactions on Software Engineering (2023).

Program Repair, Debugging, and Fault Localization

  1. [FSE, 2023] Emily First, Markus Rabe, Talia Ringer, and Yuriy Brun. 2023. Baldur: Whole-proof Generation and Repair with Large Language Models. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on The Foundations of Software Engineering. 1229–1241.
  2. [FSE, 2023] Yuxiang Wei, Chunqiu Steven Xia, and Lingming Zhang. 2023. Copiloting the Copilots: Fusing Large Language Models with Completion Engines for Automated Program Repair. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 172–184.
  3. [FSE, 2023] Weishi Wang, Yue Wang, Shafiq Joty, and Steven CH Hoi. 2023. RAP-Gen: Retrieval-Augmented Patch Generation with CodeT5 for Automatic Program Repair. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 146–158.
  4. [ICSE, 2023] Chunqiu Steven Xia, Yuxiang Wei, and Lingming Zhang. 2023. Automated Program Repair In the Era of Large Pre-Trained Language Models. In Proceedings of the 45th International Conference on Software Engineering (ICSE '23). https://doi.org/10.1109/ICSE48619.2023.00129
  5. [ICSE, 2023] Zhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury, and Shin Hwei Tan. 2023. Automated Repair of Programs from Large Language Models. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 1469–1481.
  6. [ICSE, 2023] Zongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang, Dong Chen, Shuai Wang, and Cuiyun Gao. 2023. CCTest: Testing and Repairing Code Completion Systems. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 1238–1250.
  7. [ICSE, 2024] Aidan ZH Yang, Claire Le Goues, Ruben Martins, and Vincent Hellendoorn. 2024. Large Language Models for Test-Free Fault Localization. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering. 1–12.

Testing and Code Quality

  1. [TOSEM, 2015] Gordon Fraser, Matt Staats, Phil McMinn, Andrea Arcuri, and Frank Padberg. 2015. Does Automated Unit Test Generation Really Help Software Testers? A Controlled Empirical Study. ACM Transactions on Software Engineering and Methodology (TOSEM) 24, 4 (2015), 1–49.
  2. [FSE, 2022] Ali Reza Ibrahimzada, Yigit Varli, Dilara Tekinoglu, and Reyhaneh Jabbarvand. 2022. Perfect Is The Enemy of Test Oracle. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on The Foundations of Software Engineering. 70–81.
  3. [ICSE, 2023] Zhe Liu, Chunyang Chen, Junjie Wang, Xing Che, Yuekai Huang, Jun Hu, and Qing Wang. 2023. Fill In the Blank: Context-Aware Automated Text Input Generation for Mobile GUI Testing. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 1355–1367.
  4. [TSE, 2023] Max Schäfer, Sarah Nadi, Aryaz Eghbali, and Frank Tip. 2023. An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation. IEEE Transactions on Software Engineering (2023).

Code Search, Summarization, Review, and API Support

  1. [FSE, 2022] Yao Wan, Shijie Zhang, Hongyu Zhang, Yulei Sui, Guandong Xu, Dezhong Yao, Hai Jin, and Lichao Sun. 2022. You See What I Want You to See: Poisoning Vulnerabilities In Neural Code Search. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on The Foundations of Software Engineering (ESEC/FSE 2022). Association for Computing Machinery, New York, NY, USA, 1233–1245.
  2. [ICSE, 2022] Zhensu Sun, Li Li, Yan Liu, Xiaoning Du, and Li Li. 2022. On the Importance of Building High-Quality Training Datasets for Neural Code Search. In Proceedings of the 44th International Conference on Software Engineering. 1609–1620.
  3. [ICSE, 2022] Rosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella, Denys Poshyvanyk, and Gabriele Bavota. 2022. Using Pre-Trained Models to Boost Code Review Automation. In Proceedings of the 44th International Conference on Software Engineering. 2291–2302.
  4. [TOSEM, 2023] Zhihao Li, Chuanyi Li, Ze Tang, Wanhong Huang, Jidong Ge, Bin Luo, Vincent Ng, Ting Wang, Yucheng Hu, and Xiaopeng Zhang. 2023. PTM-APIRec: Leveraging Pre-Trained Models of Source Code In API Recommendation. ACM Transactions on Software Engineering and Methodology (2023).
  5. [TOSEM, 2024] Guodong Fan, Shizhan Chen, Cuiyun Gao, Jianmao Xiao, Tao Zhang, and Zhiyong Feng. 2024. Rapid: Zero-Shot Domain Adaptation for Code Search with Pre-Trained Models. ACM Transactions on Software Engineering and Methodology (2024).

Security and Robustness

  1. [ICSE, 2024] Benjamin Steenhoek, Hongyang Gao, and Wei Le. 2024. Dataflow Analysis-Inspired Deep Learning for Efficient Vulnerability Detection. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering. 1–13.

Other Relevant SE/AI Papers

  1. [FSE, 2018] Vincent J Hellendoorn, Christian Bird, Earl T Barr, and Miltiadis Allamanis. 2018. Deep Learning Type Inference. In Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on The Foundations of Software Engineering. 152–162.
  2. [ICSE, 2024] Junjielong Xu, Ziang Cui, Yuan Zhao, Xu Zhang, Shilin He, Pinjia He, Liqun Li, Yu Kang, Qingwei Lin, Yingnong Dang, et al. 2024. UniLog: Automatic Logging via LLM and in-Context Learning. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering. 1–12.