Chi tiết tin tức
A- A A+ | Trong Tăng tương phản Giảm tương phản

Trang thông tin luận án của Nghiên cứu sinh Nguyễn Thị Tuyển

THÔNG TIN VỀ LUẬN ÁN TIẾN SĨ

1. Họ và tên nghiên cứu sinh: Nguyễn Thị Tuyển

2. Giới tính: Nữ

3. Ngày sinh: 26/03/1981

4. Nơi sinh: Phường Linh Sơn, tỉnh Thái Nguyên

5. Quyết định công nhận nghiên cứu sinh: 1006/QĐ- ĐHCNTT&TT ngày 30 tháng 11 năm 2022 của hiệu trưởng Trường ĐH CNTT và Truyền Thông.

6. Các thay đổi trong quá trình đào tạo (nếu có): Tên đề tài luận án theo quyết định là “Dự đoán chức năng protein sử dụng mô hình học máy”, đổi thành “Dự đoán chức năng protein sử dụng mô hình học sâu”.

7. Tên đề tài luận án: Dự đoán chức năng protein sử dụng mô hình học sâu

8. Chuyên ngành: Khoa học máy tính

9. Mã số: 9 48 01 01

10. Cán bộ hướng dẫn khoa học:

  1. PGS.TS. Lê Nguyễn Quốc Khánh
  2. PGS.TS. Nguyễn Văn Núi

11. Tóm tắt các kết quả mới của luận án.

Luận án tập trung nghiên cứu bài toán dự đoán chức năng protein, khai thác thông tin từ trình tự và cấu trúc protein. Trên cơ sở kế thừa và phát triển các công trình đã công bố, các đóng góp khoa học của luận án được xây dựng theo một tiến trình nâng cấp về phương pháp, thể hiện ở cách biểu diễn dữ liệu, thiết kế kiến trúc và phát triển các mô hình dự đoán.

Trên cơ sở đó, luận án đã đạt được các kết quả mới sau:

  1. Làm rõ vai trò của các phương pháp biểu diễn trình tự protein dựa trên xử lý ngôn ngữ tự nhiên và mô hình ngôn ngữ protein trong dự đoán chức năng protein. Thông qua thực nghiệm, luận án chứng minh hiệu quả của các biểu diễn học được so với các đặc trưng thủ công, đồng thời cho thấy các biểu diễn ngữ cảnh sâu từ mô hình ngôn ngữ protein góp phần nâng cao khả năng dự đoán chức năng protein.
  2. Đề xuất các mô hình học sâu lai CNN–BiLSTM, cho phép khai thác đồng thời các đặc trưng cục bộ và thông tin ngữ cảnh dài hạn của chuỗi protein. Sự kết hợp này giúp nâng cao khả năng biểu diễn và cải thiện hiệu năng dự đoán trong các bài toán dự đoán chức năng protein dạng nhị phân.
  3. Đề xuất mô hình học sâu MEARN kết hợp biểu diễn ngữ cảnh sâu từ mô hình ngôn ngữ protein với khối tự chú ý đa đầu và cơ chế học phần dư trong khối perceptron nhiều lớp, nhằm nâng cao khả năng khai thác và biến đổi biểu diễn đặc trưng cho bài toán dự đoán epitope tế bào B trong điều kiện dữ liệu lớn và mất cân bằng nghiêm trọng.
  4. Đề xuất các mô hình STRUCTSEQ2GO và ESM-STRUCT2GO hợp nhất thông tin trình tự và cấu trúc protein, kết hợp mô hình ngôn ngữ protein với biểu diễn cấu trúc cho bài toán dự đoán chức năng protein dạng đa nhãn theo Gene Ontology (GO), qua đó mở rộng hướng tiếp cận từ bài toán nhị phân sang bài toán đa nhãn với mức độ biểu diễn thông tin toàn diện hơn.

12. Giá trị khoa học, khả năng ứng dụng trong thực tiễn và những vấn đề cần tiếp tục nghiên cứu.

12.1. Giá trị khoa học và khả năng ứng dụng của kết quả nghiên cứu

Ý nghĩa khoa học:

  1. Góp phần khẳng định hiệu quả của các phương pháp biểu diễn trình tự dựa trên xử lý ngôn ngữ tự nhiên và mô hình ngôn ngữ protein trong bài toán dự đoán chức năng protein.
  2. Cung cấp bằng chứng thực nghiệm về hiệu quả của các kiến trúc học sâu lai, khối tự chú ý đa đầu, cơ chế học phần dư và phương pháp hợp nhất thông tin trình tự–cấu trúc trong việc cải thiện khả năng biểu diễn và hiệu năng dự đoán chức năng protein ở cả bài toán nhị phân và đa nhãn.
  3. Góp phần xây dựng cơ sở phương pháp luận và kỹ thuật cho việc phát triển các mô hình học sâu dự đoán chức năng protein có khả năng khai thác đồng thời thông tin trình tự, biểu diễn ngữ cảnh sâu từ mô hình ngôn ngữ protein và thông tin cấu trúc protein.

Về mặt thực tiễn:

Luận án cung cấp các mô hình dự đoán chức năng protein có thể ứng dụng trong nhận diện protein checkpoint miễn dịch, dự đoán epitope tế bào B và chú giải chức năng protein theo Gene Ontology. Các mô hình này có khả năng hỗ trợ sàng lọc và xác định các protein hoặc vị trí chức năng tiềm năng trên quy mô lớn, góp phần giảm thời gian và chi phí cho các bước kiểm chứng thực nghiệm. Kết quả nghiên cứu có thể hỗ trợ các nghiên cứu trong sinh học phân tử, sinh tin học và y sinh học, đồng thời tạo tiền đề cho các ứng dụng liên quan đến phát hiện mục tiêu sinh học, nghiên cứu vaccine, phát triển thuốc và chú giải chức năng cho các protein chưa được xác định đầy đủ.

12.2. Những vấn đề cần tiếp tục nghiên cứu và phát triển

Từ những kết quả đạt được, luận án định hướng một số hướng nghiên cứu và phát triển tiếp theo như sau:

  1. Tiếp tục nâng cao hiệu năng dự đoán trên nhánh BPO thông qua việc phát triển các chiến lược huấn luyện phù hợp với phân bố nhãn mất cân bằng, nâng cao khả năng học đối với các thuật ngữ hiếm và khai thác hiệu quả hơn mối quan hệ phân cấp giữa các thuật ngữ GO.
  2. Nghiên cứu các cơ chế khai thác thông tin cấu trúc protein theo hướng linh hoạt hơn, có xét đến mức độ tin cậy của cấu trúc dự đoán, nhằm nâng cao độ ổn định và độ tin cậy của mô hình.
  3. Mở rộng đánh giá các mô hình trên các bộ dữ liệu độc lập và đa dạng hơn về loài và chức năng protein nhằm kiểm chứng đầy đủ hơn khả năng khái quát hóa của mô hình trong các điều kiện thực tế.
  4. Phân tích sâu hơn ý nghĩa sinh học của các kết quả dự đoán nhằm nâng cao tính giải thích và giá trị ứng dụng của các mô hình trong lĩnh vực tin sinh học.
  5. Tối ưu hóa các mô hình theo hướng gọn nhẹ hơn nhưng vẫn duy trì hiệu năng cao, nhằm đáp ứng yêu cầu tính toán khi mở rộng sang các bộ dữ liệu có quy mô lớn hơn.

13. Các công trình công bố liên quan đến luận án:

[CB1] Thi-Tuyen Nguyen, Van-Nui Nguyen, Thi-Xuan Tran, Nguyen-Quoc-Khanh Le
(2024). A Machine Learning Approach for Predictive Insights into Tight Junction
Protein Functions. Proceedings of the 13th International Conference on Information Technology and Its Applications (CITA 2024), pp. 37–47. Vietnam-Korea University of Information and Communication Technology. ISBN: 978-604-80-9774-5. Available at: https://elib.vku.udn.vn/handle/123456789/4005.

[CB2] Nguyen Quoc Khanh Le, Van-Nui Nguyen, Thi-Tuyen Nguyen, Thi-Xuan Tran,
Trang-Thi Ho, Van-Lam Ho (2024). Deep Learning-Based Identification of
Rab Proteins: A Convolutional Neural Network Approach with Evolutionary
Information Integration. Lecture Notes in Data Engineering and Communications Technologies. Springer Nature Switzerland AG, pp. 177–187. DOI:
10.1007/978-3-031-75596-5_17 (Scopus Q3).

[CB3] Thi-Tuyen Nguyen, Van-Nui Nguyen, Thi-Xuan Tran, Nguyen-Quoc-Khanh Le
(2025). Improved Linear B-Cell Epitope Prediction Using CNN and BiLSTM.
Advances in Information and Communication Technology. Springer, Lecture
Notes in Networks and Systems, 1205, pp. 466–475. DOI: 10.1007/978-3-031-
80943-9_50 (Scopus Q4).

[CB4] Thi-Tuyen Nguyen, Van-Nui Nguyen, Thi-Xuan Tran, Nguyen-Quoc-Khanh
Le (2026). Integrating Protein Language Models and Deep Learning for Immune Checkpoint Protein Prediction. Proceedings of the 4th International
Conference on Advances in Information and Communication Technology (ICTA
2025). Springer, Lecture Notes in Networks and Systems, pp. 281–291. DOI:
10.1007/978-3-031-81662-5_29 (Scopus Q4).

[CB5] Thi-Tuyen Nguyen, Thi-Xuan Tran, Thi-Tuyen Ho, Hai-Thanh Tran, Nguyen Quoc-Khanh Le, Van-Nui Nguyen (2025). MEARN: Enhancing B-cell Epitope Prediction with ESM-Derived Protein Embeddings and Attention-Residual Neural Network. Journal of Computational Biology. Under review.

[CB6] Thi-Tuyen Nguyen, Wenqing Zheng, Van-Nui Nguyen, Nguyen Quoc Khanh
Le, Matthew Chin Heng Chua (2025, online; 2026, issue). A unified graphbased approach for protein function prediction using AlphaFold structures and
sequence features. Computational Biology and Chemistry, 120, 108609. DOI:
10.1016/j.compbiolchem.2025.108609. (SCIE Q2, IF:3.1).

[CB7] Thi-Tuyen Nguyen, Zhuocheng Jiang, Van-Nui Nguyen, Nguyen Quoc Khanh Le, Matthew Chin Heng Chua (2025). Integrating ESM-2 and Graph Neural Networks with AlphaFold-2 Structures for Enhanced Protein Function Prediction.
ACS Omega, 10(33), 38103–38111. DOI: 10.1021/acsomega.5c05484. (SCIE
Q1, IF: 4.3).

INFORMATION ON DOCTORAL THESIS

1. Full name: Nguyen Thi Tuyen

2. Sex: Female

3. Date of birth: 26/03/1981

4. Place of birth: Linh Son Ward, Thai Nguyen Province

5. Admission decision number: Decision No. 1006/QĐ-ĐHCNTT&TT dated November 30, 2022, issued by the Rector of the University of Information and Communication Technology

6. Changes in academic process: The title of the doctoral dissertation under the original decision was “Protein Function Prediction Using Machine Learning Models” and has been changed to “Protein Function Prediction Using Deep Learning Models.”

7. Official thesis title: Protein Function Prediction Using Deep Learning Models

8. Major: Computer Science

9. Code: 9 48 01 01

10. Supervisors:

1: Assoc. Prof. Dr. Le Nguyen Quoc Khanh

2: Assoc. Prof. Dr. Nguyen Van Nui

11. Summary of the Novel Results of the Dissertation

The dissertation focuses on the problem of protein function prediction by exploiting information from protein sequences and structures. Building upon and further developing the published works, the scientific contributions of the dissertation are established through a progressive methodological advancement, reflected in data representation, architectural design, and the development of predictive models.

On this basis, the dissertation has achieved the following novel results:

  1. Clarifying the role of protein sequence representation methods based on natural language processing and protein language models in protein function prediction. Through experiments, the dissertation demonstrates the effectiveness of learned representations compared with handcrafted features, while also showing that deep contextual representations derived from protein language models contribute to improving protein function prediction capability.
  2. Proposing hybrid CNN–BiLSTM deep learning models that enable the simultaneous exploitation of local features and long-range contextual information from protein sequences. This combination improves representation capability and prediction performance in binary protein function prediction tasks.
  3. Proposing the MEARN deep learning model, which combines deep contextual representations from protein language models with a multi-head self-attention block and residual learning within a multilayer perceptron block, with the aim of improving the exploitation and transformation of feature representations for B-cell epitope prediction under conditions of large-scale and severely imbalanced data.
  4. Proposing the STRUCTSEQ2GO and ESM-STRUCT2GO models for integrating protein sequence and structural information by combining protein language models with structural representations for multi-label protein function prediction according to Gene Ontology (GO), thereby extending the approach from binary prediction tasks to multi-label prediction tasks with a more comprehensive level of information representation.

12. Scientific Significance, Practical Applicability, and Outstanding Issues Requiring Further Research

12.1. Scientific Significance and Practical Applicability of the Research Findings

Scientific Significance:

  1. Contributes to confirming the effectiveness of sequence representation methods based on natural language processing and protein language models in protein function prediction.
  2. Provides empirical evidence of the effectiveness of hybrid deep learning architectures, multi-head self-attention blocks, residual learning mechanisms, and sequence–structure information integration methods in improving representation capability and protein function prediction performance in both binary and multi-label tasks.
  3. Contributes to establishing a methodological and technical foundation for the development of deep learning models for protein function prediction that are capable of simultaneously exploiting sequence information, deep contextual representations from protein language models, and protein structural information.

Practical Applicability:

The dissertation provides protein function prediction models that can be applied to the identification of immune checkpoint proteins, B-cell epitope prediction, and protein function annotation according to Gene Ontology. These models can support the large-scale screening and identification of potential proteins or functional sites, thereby helping to reduce the time and cost required for experimental validation. The research findings can support studies in molecular biology, bioinformatics, and biomedicine, while also providing a foundation for applications related to biological target discovery, vaccine research, drug development, and functional annotation of proteins that have not yet been fully characterized.

12.2. Outstanding Issues Requiring Further Research

Based on the results achieved, the dissertation identifies several directions for future research and development as follows:

  1. Continue improving prediction performance for the BPO branch by developing training strategies suitable for imbalanced label distributions, enhancing the learning capability for rare terms, and exploiting the hierarchical relationships among GO terms more effectively.
  2. Investigate more flexible mechanisms for exploiting protein structural information, taking into account the confidence level of predicted structures, in order to improve the stability and reliability of the models.
  3. Extend the evaluation of the models to independent datasets that are more diverse in terms of species and protein functions in order to more comprehensively verify their generalization capability under practical conditions.
  4. Further analyze the biological significance of the prediction results in order to improve the interpretability and application value of the models in the field of bioinformatics.
  5. Optimize the models toward more lightweight architectures while maintaining high performance, in order to meet computational requirements when scaling to larger datasets.

13. Thesis-related publications

[CB1] Thi-Tuyen Nguyen, Van-Nui Nguyen, Thi-Xuan Tran, Nguyen-Quoc-Khanh Le
(2024). A Machine Learning Approach for Predictive Insights into Tight Junction
Protein Functions. Proceedings of the 13th International Conference on Information Technology and Its Applications (CITA 2024), pp. 37–47. Vietnam-Korea University of Information and Communication Technology. ISBN: 978-604-80-9774-5. Available at: https://elib.vku.udn.vn/handle/123456789/4005.

[CB2] Nguyen Quoc Khanh Le, Van-Nui Nguyen, Thi-Tuyen Nguyen, Thi-Xuan Tran,
Trang-Thi Ho, Van-Lam Ho (2024). Deep Learning-Based Identification of
Rab Proteins: A Convolutional Neural Network Approach with Evolutionary
Information Integration. Lecture Notes in Data Engineering and Communications Technologies. Springer Nature Switzerland AG, pp. 177–187. DOI:
10.1007/978-3-031-75596-5_17 (Scopus Q3).

[CB3] Thi-Tuyen Nguyen, Van-Nui Nguyen, Thi-Xuan Tran, Nguyen-Quoc-Khanh Le
(2025). Improved Linear B-Cell Epitope Prediction Using CNN and BiLSTM.
Advances in Information and Communication Technology. Springer, Lecture
Notes in Networks and Systems, 1205, pp. 466–475. DOI: 10.1007/978-3-031-
80943-9_50 (Scopus Q4).

[CB4] Thi-Tuyen Nguyen, Van-Nui Nguyen, Thi-Xuan Tran, Nguyen-Quoc-Khanh
Le (2026). Integrating Protein Language Models and Deep Learning for Immune Checkpoint Protein Prediction. Proceedings of the 4th International
Conference on Advances in Information and Communication Technology (ICTA
2025). Springer, Lecture Notes in Networks and Systems, pp. 281–291. DOI:
10.1007/978-3-031-81662-5_29 (Scopus Q4).

[CB5] Thi-Tuyen Nguyen, Thi-Xuan Tran, Thi-Tuyen Ho, Hai-Thanh Tran, Nguyen Quoc-Khanh Le, Van-Nui Nguyen (2025). MEARN: Enhancing B-cell Epitope Prediction with ESM-Derived Protein Embeddings and Attention-Residual Neural Network. Journal of Computational Biology. Under review.

[CB6] Thi-Tuyen Nguyen, Wenqing Zheng, Van-Nui Nguyen, Nguyen Quoc Khanh
Le, Matthew Chin Heng Chua (2025, online; 2026, issue). A unified graphbased approach for protein function prediction using AlphaFold structures and
sequence features. Computational Biology and Chemistry, 120, 108609. DOI:
10.1016/j.compbiolchem.2025.108609. (SCIE Q2, IF:3.1).

[CB7] Thi-Tuyen Nguyen, Zhuocheng Jiang, Van-Nui Nguyen, Nguyen Quoc Khanh Le, Matthew Chin Heng Chua (2025). Integrating ESM-2 and Graph Neural Networks with AlphaFold-2 Structures for Enhanced Protein Function Prediction.
ACS Omega, 10(33), 38103–38111. DOI: 10.1021/acsomega.5c05484. (SCIE
Q1, IF: 4.3).

Thích

Bài viết liên quan