{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"},{"sourceId":7392775,"sourceType":"datasetVersion","datasetId":4297782}],"dockerImageVersionId":30635,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Đây là notebook dịch từ [🧠 Exploring EEG: A Beginner's Guide](https://www.kaggle.com/code/yorkyong/exploring-eeg-a-beginner-s-guide)","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding: 30px;color:white;margin:10;font-size:60%;text-align:left;display:fill;border-radius:10px;background-color:#FFFFFF;overflow:hidden;background-color:#FFCE30\"><b><span style='color:#FFFFFF'>1 |</span></b> <b>Giới thiệu</b></div>\n\n👋 Chào mừng bạn đến với \"🧠Khám Phá EEG: Hướng Dẫn Dành Cho Người Mới Bắt Đầu\"! \n\nNếu bạn đam mê những kỳ diệu của bộ não con người và các mẫu sóng não phức tạp, nhưng lại cảm thấy thế giới phân tích Điện não đồ (EEG) có phần khó hiểu, thì bạn đã đến đúng nơi. \n\nNotebook này được thiết kế dành cho những người mới bắt đầu như tôi và bạn, với mục tiêu giải mã sự phức tạp của dữ liệu EEG và làm cho hành trình học hỏi của bạn vừa thú vị vừa bổ ích.\n\n### <b><span style='color:#FFCE30'> 1.1 |</span> Mục Đích Của Notebook</b>\nTrong notebook này, chúng ta sẽ bắt đầu hành trình khám phá vào lĩnh vực phân tích dữ liệu EEG. Mục tiêu của chúng tôi là cung cấp một hướng dẫn rõ ràng, từng bước về cách hiểu và phân tích các tín hiệu EEG, những tín hiệu này rất quan trọng trong việc phát hiện và phân loại các hoạt động của não bộ, chẳng hạn như các cơn động kinh. Chúng tôi nhắm đến việc:\n\n* Phân tách các khái niệm phức tạp thành những phần dễ hiểu.\n* Minh họa từng bước với các ví dụ mã thực tế.\n* Tham chiếu các notebook và thảo luận công khai để nâng cao trải nghiệm học hỏi của bạn.\n\n\n### <b><span style='color:#FFCE30'> 1.2 |</span> Mục Tiêu Học Tập</b>\nCuối notebook này, bạn sẽ có một nền tảng cơ bản về:\n\n* Các tín hiệu EEG và tầm quan trọng của chúng trong nghiên cứu y học và thần kinh học.\n* Cách tiền xử lý và phân tích dữ liệu EEG.\n* Thực hành mã cơ bản để xây dựng một mô hình học máy cho phân loại dữ liệu EEG.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding: 30px;color:white;margin:10;font-size:60%;text-align:left;display:fill;border-radius:10px;background-color:#FFFFFF;overflow:hidden;background-color:#FFCE30\"><b><span style='color:#FFFFFF'>2 |</span></b> <b>TÀI LIỆU THAM KHẢO & LỜI CẢM ƠN</b></div>\n\nNotebook này sẽ không thể hoàn thành nếu không có những đóng góp và những hiểu biết quý giá từ cộng đồng Kaggle. Tôi đã tận dụng một số tài nguyên để biên soạn lộ trình học hiệu quả nhất cho chúng ta:\n\n* https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-8\n* https://www.kaggle.com/code/mvvppp/hms-eda-and-domain-journey\n* https://www.kaggle.com/code/ksooklall/hms-banana-montage\n* https://www.kaggle.com/code/mpwolke/seizures-classification-parquet\n\n\nHãy thoải mái khám phá các tài nguyên này cùng với notebook này để nâng cao sự hiểu biết của bạn.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding: 30px;color:white;margin:10;font-size:60%;text-align:left;display:fill;border-radius:10px;background-color:#FFFFFF;overflow:hidden;background-color:#FFCE30\"><b><span style='color:#FFFFFF'>3 |</span></b> <b>TẢI CÁC THƯ VIỆN</b></div>","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd, numpy as np\nfrom glob import glob\nimport matplotlib.pyplot as plt\nVER = 1","metadata":{"execution":{"iopub.status.busy":"2025-02-13T15:16:56.841691Z","iopub.execute_input":"2025-02-13T15:16:56.842569Z","iopub.status.idle":"2025-02-13T15:16:56.846576Z","shell.execute_reply.started":"2025-02-13T15:16:56.842537Z","shell.execute_reply":"2025-02-13T15:16:56.845598Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"padding: 30px;color:white;margin:10;font-size:60%;text-align:left;display:fill;border-radius:10px;background-color:#FFFFFF;overflow:hidden;background-color:#FFCE30\"><b><span style='color:#FFFFFF'>4 |</span></b> <b>GIỚI THIỆU VỀ EEG VÀ PHÁT HIỆN CÁC CƠN ĐỘNG KINH</b></div>\n\n<b><span style='color:#FFCE30'> 4.1 |</span> Điện não đồ (EEG) - Cửa Sổ Bước Vào Hoạt Động Của Não</b>\n\n* Điện não đồ, hay còn gọi là EEG, là phương pháp không xâm lấn được các chuyên gia y tế sử dụng để ghi lại hoạt động điện của não bộ. \n* Điều này được thực hiện bằng cách đặt các điện cực dọc theo da đầu. \n* EEG là công cụ quan trọng trong việc chẩn đoán các rối loạn thần kinh, đặc biệt là động kinh, một căn bệnh đặc trưng bởi các cơn động kinh tái phát.\n\n<img src=\"https://www.researchgate.net/profile/Sebastian-Nagel-4/publication/338423585/figure/fig1/AS:844668573073409@1578396089381/Sketch-of-how-to-record-an-Electroencephalogram-An-EEG-allows-measuring-the-electrical.png\" alt=\"EEG\" width=\"600\" height=\"400\">\n\n","metadata":{"execution":{"iopub.status.busy":"2024-01-14T15:14:51.032997Z","iopub.execute_input":"2024-01-14T15:14:51.033378Z","iopub.status.idle":"2024-01-14T15:14:51.038874Z","shell.execute_reply.started":"2024-01-14T15:14:51.033348Z","shell.execute_reply":"2024-01-14T15:14:51.037454Z"}}},{"cell_type":"code","source":"# Định nghĩa đường dẫn chứa dữ liệu của cuộc thi\nBASE_PATH = '/kaggle/input/hms-harmful-brain-activity-classification/'\n\n# Tạo một DataFrame chứa danh sách tất cả các tệp .parquet trong thư mục BASE_PATH (bao gồm cả thư mục con)\ndf = pd.DataFrame({'path': glob(BASE_PATH + '**/*.parquet')})\n\n# Trích xuất loại test từ đường dẫn của từng tệp\n# Cụ thể:\n# - Chia đường dẫn theo dấu '/' và lấy phần tử thứ hai từ cuối lên (tên thư mục chứa tệp).\n# - Chia tên thư mục đó theo dấu '_' và lấy phần cuối cùng.\n# Ví dụ: \"/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/1000913311.parquet\"\n# - Tên thư mục chứa file: \"train_eegs\"\n# - Sau khi tách \"_\", lấy phần cuối: \"eegs\"\ndf['test_type'] = df['path'].str.split('/').str.get(-2).str.split('_').str.get(-1)\n\n\n# Trích xuất ID của tệp từ tên file\n# - Lấy phần cuối của đường dẫn (tên file, ví dụ: \"1000913311.parquet\").\n# - Chia theo dấu '.' và lấy phần đầu tiên (bỏ đi phần mở rộng .parquet).\n# Ví dụ: \"1000913311.parquet\" -> \"1000913311\"\ndf['id'] = df['path'].str.split('/').str.get(-1).str.split('.').str.get(0)\n\n\n# Đọc dữ liệu EEG từ một tệp cụ thể (ID 1000913311) trong thư mục train_eegs\ndf_eeg = pd.read_parquet(BASE_PATH + 'train_eegs/1000913311.parquet')\ndf_eeg.head()","metadata":{"execution":{"iopub.status.busy":"2025-02-13T15:17:00.344093Z","iopub.execute_input":"2025-02-13T15:17:00.344937Z","iopub.status.idle":"2025-02-13T15:17:00.591055Z","shell.execute_reply.started":"2025-02-13T15:17:00.344875Z","shell.execute_reply":"2025-02-13T15:17:00.590145Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Xác định số lượng kênh (channels)\n# Giả định rằng mỗi hàng (row) trong df_eeg là một điểm theo thời gian (time point),\n# và mỗi cột (column) là một kênh tín hiệu EEG (channel).\nn_channels = df_eeg.shape[1]\nn_channels","metadata":{"execution":{"iopub.status.busy":"2025-02-13T15:17:03.080278Z","iopub.execute_input":"2025-02-13T15:17:03.081085Z","iopub.status.idle":"2025-02-13T15:17:03.088581Z","shell.execute_reply.started":"2025-02-13T15:17:03.081046Z","shell.execute_reply":"2025-02-13T15:17:03.087627Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"* Các tiêu đề trong tập dữ liệu (Fp1, F3, C3, P3, F7, T3, T5, O1, Fz, Cz, Pz, Fp2, F4, C4, P4, F8, T4, T6, O2, EKG) là các nhãn vị trí điện cực tiêu chuẩn được sử dụng trong điện não đồ (EEG). \n* Những nhãn này tương ứng với các vị trí cụ thể trên da đầu, nơi các điện cực EEG được đặt để ghi lại hoạt động của não. \n* Dưới đây là mô tả ngắn gọn về ý nghĩa của chúng:\n\n1. **Fp1, Fp2:** Điện cực trán trước, nằm trên trán, bên trái và bên phải.\n2. **F3, F4:** Điện cực vùng trán, nằm ở phần trán bên trái và bên phải.\n3. **C3, C4:** Điện cực trung tâm, đặt phía trên hai bán cầu não trái và phải.\n4. **P3, P4:** Điện cực vùng đỉnh, nằm ở phần trên phía sau của đầu, bên trái và bên phải.\n5. **O1, O2:** Điện cực vùng chẩm, đặt ở phía sau đầu, gần vỏ não thị giác.\n6. **T3, T4, T5, T6:** Điện cực vùng thái dương, nằm ở hai bên đầu gần tai. Chúng thường được sử dụng để theo dõi các chức năng thính giác.\n7. **F7, F8:** Điện cực vùng trán-thái dương, đặt ở phía trước của thùy thái dương.\n8. **Fz, Cz, Pz:** Điện cực đường giữa, đặt ở các vị trí trán (Fz), trung tâm (Cz) và đỉnh đầu (Pz) dọc theo đường giữa của đầu.\n9. **EKG:** Điện cực điện tâm đồ (ECG), ghi lại hoạt động điện của tim. Nó không liên quan trực tiếp đến hoạt động của não nhưng có thể quan trọng trong một số phân tích EEG.\n\n\n<img src=\"https://www.researchgate.net/profile/Danny-Plass-Oude-Bos/publication/237777779/figure/fig3/AS:669556259434497@1536646060035/10-20-system-of-electrode-placement.png\" alt=\"10-20-system-of-electrode-placement\" width=\"300\" height=\"150\">","metadata":{}},{"cell_type":"markdown","source":"<b><span style='color:#FFCE30'> 4.2 |</span> Động Kinh và Ảnh Hưởng Của Nó</b>\n* Động kinh là những rối loạn điện đột ngột, không kiểm soát trong não, có thể gây ra thay đổi về hành vi, cảm giác, cử động và mức độ ý thức.\n* Việc phát hiện và phân loại chính xác các cơn động kinh là rất quan trọng để đưa ra phương pháp điều trị và chăm sóc phù hợp, đặc biệt đối với bệnh nhân nặng.\n\n<b><span style='color:#FFCE30'> 4.3 |</span> Thách Thức Của Việc Phân Tích EEG Thủ Công</b>\n\n* Truyền thống, việc phân tích dữ liệu EEG dựa vào quan sát trực tiếp của các nhà thần kinh học được đào tạo. \n* Quá trình này không chỉ tốn nhiều thời gian, đòi hỏi công sức mà còn dễ mắc lỗi do mệt mỏi và sự chủ quan trong đánh giá.\n\n<img src=\"https://slideplayer.com/slide/12925171/78/images/2/Manual+Interpretation+of+EEGs.jpg\" alt=\"Manual Interpretation of EEG\" width=\"700\" height=\"300\">\nNguồn: Automated Identification of Abnormal Adult EEG, S. López, G. Suarez, D. Jungreis, I. Obeid and J. Picone, Neural Engineering Data Consortium, Temple University\n","metadata":{}},{"cell_type":"markdown","source":"<b><span style='color:#FFCE30'> 4.4 |</span> Vai Trò Của Khoa Học Dữ Liệu Trong Phân Tích EEG</b>\n\n* Tự Động Hóa Việc Giải Thích EEG\n\nSự phát triển của học máy và khoa học dữ liệu mang đến cơ hội tự động hóa quá trình giải thích dữ liệu EEG. Bằng cách xây dựng các thuật toán có khả năng phát hiện và phân loại các mẫu tín hiệu EEG khác nhau, chúng ta có thể hỗ trợ các nhà thần kinh học trong việc đưa ra chẩn đoán nhanh hơn và chính xác hơn.\n\n* Cách Tiếp Cận Khoa Học Dữ Liệu\n\nCác nhà khoa học dữ liệu tiếp cận thách thức này bằng cách tiền xử lý dữ liệu EEG, bao gồm lọc nhiễu và trích xuất các đặc trưng quan trọng. Học máy sau đó được áp dụng để phân tích và tìm ra các mẫu có ý nghĩa từ dữ liệu EEG, giúp cải thiện độ chính xác trong chẩn đoán và hỗ trợ nghiên cứu y học.\n\n<img src=\"https://www.researchgate.net/profile/Huiguang-He/publication/336336651/figure/fig1/AS:834361356197888@1575938657076/The-flow-chart-of-EEG-emotion-classification-with-similarity-learning-network.png\" alt=\"flowchart for EEG classification\" width=\"700\" height=\"300\">\n","metadata":{}},{"cell_type":"markdown","source":"<b><span style='color:#FFCE30'> 4.5 |</span> Hiểu Về Các Mẫu Sóng EEG</b>\n\nTrong phân tích EEG để phát hiện động kinh, một số mẫu sóng đặc biệt quan trọng:\n\n1. **Seizure (SZ) - Cơn động kinh:** Được đặc trưng bởi hoạt động dao động bất thường, dấu hiệu của một cơn động kinh.\n2. **Generalized Periodic Discharges (GPD) - Sóng xả chu kỳ tổng quát:** Các mẫu sóng có thể xuất hiện trong nhiều dạng bệnh não khác nhau.\n3. **Lateralized Periodic Discharges (LPD) - Sóng xả chu kỳ khu trú:** Thường liên quan đến các tổn thương khu trú trong não.\n4. **Lateralized Rhythmic Delta Activity (LRDA) - Sóng delta nhịp điệu khu trú:** Có thể quan sát thấy trong các rối loạn chức năng khu trú của não.\n5. **Generalized Rhythmic Delta Activity (GRDA) - Sóng delta nhịp điệu tổng quát:** Thường liên quan đến rối loạn chức năng não lan tỏa.\n6. **\"Các mẫu sóng khác\":** Bất kỳ loại hoạt động nào không thuộc các danh mục trên.\n\n<b><span style='color:#FFCE30'> 4.6 |</span> Giải Thích Dữ Liệu EEG Phức Tạp</b>\n\nViệc giải thích dữ liệu EEG có thể rất phức tạp, đặc biệt trong những trường hợp ngoại lệ, khi ngay cả các chuyên gia thần kinh cũng có thể không đồng thuận về một phân loại nhất định. Đây là lúc các mô hình học máy phát huy tác dụng, cung cấp một lớp phân tích bổ sung để hỗ trợ đánh giá chính xác hơn.\n\n<img src=\"https://www.neurology.org/cms/10.1212/WNL.0000000000207127/asset/bd84c182-712c-41ab-8742-cecf9d49a322/assets/images/large/5ff2.jpg\" alt=\"flowchart for EEG classification\" width=\"700\" height=\"300\">\n\nNguồn: Development of Expert-Level Classification of Seizures and Rhythmic and Periodic Patterns During EEG Interpretation https://www.neurology.org/doi/10.1212/WNL.0000000000207127\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding: 30px;color:white;margin:10;font-size:60%;text-align:left;display:fill;border-radius:10px;background-color:#FFFFFF;overflow:hidden;background-color:#FFCE30\"><b><span style='color:#FFFFFF'>5 |</span></b> <b>TẢI DỮ LIỆU HUẤN LUYỆN</b></div>","metadata":{"execution":{"iopub.status.busy":"2024-01-14T15:16:07.035029Z","iopub.execute_input":"2024-01-14T15:16:07.035699Z","iopub.status.idle":"2024-01-14T15:16:07.040799Z","shell.execute_reply.started":"2024-01-14T15:16:07.035656Z","shell.execute_reply":"2024-01-14T15:16:07.039623Z"}}},{"cell_type":"code","source":"df = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\n\n# Xác định các cột mục tiêu (TARGETS) bằng cách lấy 6 cột cuối cùng của DataFrame\nTARGETS = df.columns[-6:]\n\n# Kích thước của tập dữ liệu\nprint('Train shape:', df.shape )\n\n# Ddanh sách các cột mục tiêu\nprint('Targets', list(TARGETS))\n\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2025-02-13T15:17:13.317547Z","iopub.execute_input":"2025-02-13T15:17:13.317839Z","iopub.status.idle":"2025-02-13T15:17:13.551263Z","shell.execute_reply.started":"2025-02-13T15:17:13.317818Z","shell.execute_reply":"2025-02-13T15:17:13.550417Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"padding: 30px;color:white;margin:10;font-size:60%;text-align:left;display:fill;border-radius:10px;background-color:#FFFFFF;overflow:hidden;background-color:#FFCE30\"><b><span style='color:#FFFFFF'>6 |</span></b> <b>TẠO DỮ LIỆU HUẤN LUYỆN EEG KHÔNG TRÙNG</b></div>\n\nDựa theo notebook của Chris Deotte: https://www.kaggle.com/code/cdeotte/catboost-starter-lb-0-8\n\nThảo luận ban đầu tại đây: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/467021\n\nChúng ta thực hiện các bước sau vì:\n\n* **Khớp Định Dạng Dữ Liệu Huấn Luyện Với Dữ Liệu Kiểm Tra:** Cuộc thi quy định rằng dữ liệu kiểm tra không chứa nhiều đoạn tín hiệu từ cùng một eeg_id. TĐể làm cho dữ liệu huấn luyện tương tự với dữ liệu kiểm tra, chúng ta cũng chỉ sử dụng một đoạn tín hiệu cho mỗi eeg_id trong dữ liệu huấn luyện.\n\n* **Loại Bỏ Sự Dư Thừa:** Phương pháp này đảm bảo rằng dữ liệu huấn luyện không có thông tin trùng lặp hoặc chồng lấp, giúp mô hình học máy trở nên chính xác hơn và tổng quát hóa tốt hơn.\n\n* **Tính Nhất Quán Trong Dữ Liệu:** Việc chuẩn hóa cách xử lý các đoạn EEG trong tập huấn luyện giúp đảm bảo mô hình học từ dữ liệu có định dạng tương tự với dữ liệu nó sẽ được kiểm tra.\n\n* **Chuẩn Bị Dữ Liệu Cho Học Máy:** Việc chuẩn hóa các biến mục tiêu và bao gồm các đặc trưng quan trọng như patient_id và expert_consensus giúp chuẩn bị tập dữ liệu cho việc xây dựng mô hình học máy hiệu quả.","metadata":{"execution":{"iopub.status.busy":"2024-01-14T15:17:26.898038Z","iopub.execute_input":"2024-01-14T15:17:26.898867Z","iopub.status.idle":"2024-01-14T15:17:26.903917Z","shell.execute_reply.started":"2024-01-14T15:17:26.898818Z","shell.execute_reply":"2024-01-14T15:17:26.902565Z"}}},{"cell_type":"code","source":"# Tạo một phân đoạn EEG duy nhất cho mỗi eeg_id:\n# - Nhóm dữ liệu (groupby) theo eeg_id, mỗi eeg_id đại diện cho một bản ghi EEG riêng biệt.\n# - Chọn spectrogram_id đầu tiên và thời điểm bắt đầu sớm nhất (min) spectrogram_label_offset_seconds cho mỗi eeg_id.\n# - Kết quả giúp xác định điểm bắt đầu của mỗi phân đoạn EEG.\ntrain = df.groupby('eeg_id')[['spectrogram_id','spectrogram_label_offset_seconds']].agg(\n    {'spectrogram_id':'first','spectrogram_label_offset_seconds':'min'})\n\n# Đổi tên các cột để dễ đọc hơn:\n# - 'spec_id': Chứa giá trị spectrogram_id đầu tiên của mỗi EEG.\n# - 'min': Thời điểm bắt đầu sớm nhất của nhãn EEG.\ntrain.columns = ['spec_id','min']\n\n\n# Tìm thời điểm kết thúc của mỗi phân đoạn EEG:\n# - Nhóm dữ liệu theo eeg_id, tìm giá trị lớn nhất (max) của spectrogram_label_offset_seconds.\n# - Giá trị này biểu thị điểm cuối của phân đoạn EEG.\ntmp = df.groupby('eeg_id')[['spectrogram_id','spectrogram_label_offset_seconds']].agg(\n    {'spectrogram_label_offset_seconds':'max'})\n\n\n# Thêm giá trị thời điểm kết thúc vào DataFrame train.\ntrain['max'] = tmp\n\n# Thêm thông tin bệnh nhân vào train DataFrame:\n# - Nhóm dữ liệu theo eeg_id, lấy giá trị đầu tiên của patient_id.\n# - Điều này giúp liên kết mỗi EEG với một bệnh nhân cụ thể.\ntmp = df.groupby('eeg_id')[['patient_id']].agg('first')\n\n# Gán patient_id vào train\ntrain['patient_id'] = tmp\n\n# Tính tổng số lần xuất hiện của mỗi mục tiêu (target labels) theo eeg_id:\n# - Nhóm dữ liệu theo eeg_id, tính tổng giá trị của các nhãn mục tiêu (TARGETS).\n# - Điều này giống như \"đếm phiếu bầu\" cho từng loại nhãn (seizure, LPD, GPD, ...).\ntmp = df.groupby('eeg_id')[TARGETS].agg('sum') \n\n\n# Thêm giá trị tổng của từng nhãn mục tiêu vào train DataFrame.\nfor t in TARGETS:\n    train[t] = tmp[t].values\n\n# Chuẩn hóa dữ liệu nhãn mục tiêu:\n# - Chuyển đổi số lượng phiếu bầu thành xác suất bằng cách chuẩn hóa tổng của chúng về 1.\n# - Đây là bước quan trọng để sử dụng dữ liệu trong bài toán phân loại.\ny_data = train[TARGETS].values\ny_data = y_data / y_data.sum(axis=1,keepdims=True)\n\n\n# Gán giá trị chuẩn hóa vào train DataFrame.\ntrain[TARGETS] = y_data\n\n# Thêm thông tin đánh giá từ chuyên gia:\n# - Với mỗi eeg_id, lấy giá trị đầu tiên của expert_consensus (nhận định của chuyên gia về phân đoạn EEG này).\ntmp = df.groupby('eeg_id')[['expert_consensus']].agg('first')\n\n# Thêm vào train\ntrain['target'] = tmp\n\n# Reset index để eeg_id trở thành một cột thay vì chỉ số index.\ntrain = train.reset_index()\n\n# Kích thước của tập train sau khi xử lý.\nprint('Train non-overlapp eeg_id shape:', train.shape )\n\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:39:15.419607Z","iopub.execute_input":"2024-01-17T13:39:15.41999Z","iopub.status.idle":"2024-01-17T13:39:15.515608Z","shell.execute_reply.started":"2024-01-17T13:39:15.419957Z","shell.execute_reply":"2024-01-17T13:39:15.514644Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"padding: 30px;color:white;margin:10;font-size:60%;text-align:left;display:fill;border-radius:10px;background-color:#FFFFFF;overflow:hidden;background-color:#FFCE30\"><b><span style='color:#FFFFFF'>7 |</span></b> <b>KỸ THUẬT XÂY DỰNG ĐẶC TRƯNG</b></div>","metadata":{}},{"cell_type":"markdown","source":"<b><span style='color:#FFCE30'> 7.1 |</span> Cửa sổ thời gian 10 phút và 20 giây</b>\n\n* Đoạn mã dưới đây đọc dữ liệu phổ (spectrogram) một cách hiệu quả từ một tệp kết hợp duy nhất, dựa trên biến được thiết lập. Chúng tôi sử dụng bộ dữ liệu của Chris Deotte để tiết kiệm thời gian: https://www.kaggle.com/datasets/cdeotte/brain-spectrograms\n* Sau đó, nó thực hiện kỹ thuật xây dựng đặc trưng bằng cách tính giá trị trung bình và nhỏ nhất trên hai cửa sổ thời gian khác nhau cho mỗi tần số trong phổ tín hiệu.\nQuá trình này tạo ra 1.600 đặc trưng (400 đặc trưng × 4 phép tính) cho mỗi EEG ID.\n* Các đặc trưng mới này giúp mô hình hiểu và phân loại dữ liệu EEG tốt hơn.\n* Cách tiếp cận này được thiết kế nhằm cải thiện hiệu suất mô hình bằng cách cung cấp thêm thông tin chi tiết được trích xuất từ dữ liệu phổ tín hiệu EEG.","metadata":{}},{"cell_type":"code","source":"READ_SPEC_FILES = False  # Nếu READ_SPEC_FILES = False, code sẽ đọc tệp kết hợp thay vì các tệp riêng lẻ.\nFEATURE_ENGINEER = True  # Bật chế độ trích xuất đặc trưng (feature engineering).","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:39:15.516627Z","iopub.execute_input":"2024-01-17T13:39:15.516955Z","iopub.status.idle":"2024-01-17T13:39:15.52139Z","shell.execute_reply.started":"2024-01-17T13:39:15.516918Z","shell.execute_reply":"2024-01-17T13:39:15.520335Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\n# Định nghĩa đường dẫn chứa dữ liệu spectrogram\nPATH = '/kaggle/input/hms-harmful-brain-activity-classification/train_spectrograms/'\n\n# Lấy danh sách tất cả các tệp trong thư mục này\nfiles = os.listdir(PATH)\n\n# Số lượng tệp .parquet có trong thư mục\nprint(f'There are {len(files)} spectrogram parquets')\n\n\nif READ_SPEC_FILES:    \n    spectrograms = {}\n    for i, f in enumerate(files):\n        if i % 100 == 0: print(i, ', ', end='')  # In tiến trình đọc dữ liệu sau mỗi 100 tệp\n        tmp = pd.read_parquet(f'{PATH}{f}')  # Đọc tệp .parquet\n        name = int(f.split('.')[0])  # Lấy ID từ tên file (loại bỏ phần mở rộng .parquet)\n        spectrograms[name] = tmp.iloc[:, 1:].values  # Lưu dữ liệu spectrogram (bỏ cột đầu tiên)\nelse:\n    spectrograms = np.load('/kaggle/input/brain-spectrograms/specs.npy', allow_pickle=True).item() ","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:39:15.522641Z","iopub.execute_input":"2024-01-17T13:39:15.523043Z","iopub.status.idle":"2024-01-17T13:40:17.405646Z","shell.execute_reply.started":"2024-01-17T13:39:15.523005Z","shell.execute_reply":"2024-01-17T13:40:17.404521Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%time\n# ENGINEER FEATURES\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Đoạn code này tạo ra các đặc trưng từ dữ liệu spectrogram để sử dụng trong mô hình.\n# Các đặc trưng được tính bằng cách lấy giá trị trung bình (mean) và giá trị nhỏ nhất (min) theo thời gian \n# trên mỗi trong số 400 tần số của spectrogram.\n# Hai loại cửa sổ thời gian được sử dụng để tính toán:\n# - Cửa sổ 10 phút (_mean_10m, _min_10m).\n# - Cửa sổ 20 giây (_mean_20s, _min_20s).\n# Quá trình này tạo ra tổng cộng 1600 đặc trưng (400 tần số × 4 phép tính) cho mỗi EEG ID.\n\n# Trích xuất danh sách các cột spectrogram (bỏ cột đầu tiên)\nSPEC_COLS = pd.read_parquet(f'{PATH}1000086677.parquet').columns[1:]\n\nFEATURES = [f'{c}_mean_10m' for c in SPEC_COLS]  # Giá trị trung bình trên 10 phút\nFEATURES += [f'{c}_min_10m' for c in SPEC_COLS]   # Giá trị nhỏ nhất trên 10 phút\nFEATURES += [f'{c}_mean_20s' for c in SPEC_COLS]  # Giá trị trung bình trên 20 giây\nFEATURES += [f'{c}_min_20s' for c in SPEC_COLS]   # Giá trị nhỏ nhất trên 20 giây\n\nprint(f'Chúng ta đang tạo {len(FEATURES)} đặc trưng cho {len(train)} mẫu dữ liệu... ', end='')\n\n# Một ma trận dữ liệu `data` được khởi tạo để lưu trữ các đặc trưng mới cho mỗi `eeg_id` trong DataFrame `train`.\n# Đối với mỗi dòng trong `train`, đoạn code sẽ tính toán giá trị trung bình (mean) và giá trị nhỏ nhất (min) \n# trong các cửa sổ thời gian được chỉ định (10 phút và 20 giây).\n# Các giá trị đã tính toán này sau đó được lưu vào ma trận dữ liệu.\n# Cuối cùng, ma trận này được thêm vào DataFrame `train` dưới dạng các cột mới.\n\nif FEATURE_ENGINEER:\n    # Khởi tạo ma trận dữ liệu rỗng để lưu đặc trưng\n    data = np.zeros((len(train), len(FEATURES)))  \n    \n    # Duyệt từng mẫu dữ liệu trong train\n    for k in range(len(train)):\n        if k % 100 == 0: print(k, ', ', end='')  # Cứ mỗi 100 mẫu in số lượng đã xử lý\n        \n        row = train.iloc[k]  # Lấy một dòng từ DataFrame train\n        r = int((row['min'] + row['max']) // 4)  # Xác định vị trí trung tâm để trích xuất đặc trưng\n        \n        # 🚀 Tính toán đặc trưng trên cửa sổ 10 phút\n        x = np.nanmean(spectrograms[row.spec_id][r:r+300, :], axis=0)  # Trung bình 10 phút\n        data[k, :400] = x  # Lưu vào 400 cột đầu\n        \n        x = np.nanmin(spectrograms[row.spec_id][r:r+300, :], axis=0)  # Giá trị nhỏ nhất 10 phút\n        data[k, 400:800] = x  # Lưu vào cột 400 - 799\n\n        # 🚀 Tính toán đặc trưng trên cửa sổ 20 giây\n        x = np.nanmean(spectrograms[row.spec_id][r+145:r+155, :], axis=0)  # Trung bình 20 giây\n        data[k, 800:1200] = x  # Lưu vào cột 800 - 1199\n\n        x = np.nanmin(spectrograms[row.spec_id][r+145:r+155, :], axis=0)  # Giá trị nhỏ nhất 20 giây\n        data[k, 1200:1600] = x  # Lưu vào cột 1200 - 1599\n\n    # Gán các đặc trưng vừa tính vào DataFrame train\n    train[FEATURES] = data\nelse:\n    train = pd.read_parquet('/kaggle/input/brain-spectrograms/train.pqt')\n\nprint('New train shape:',train.shape)","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:40:17.410051Z","iopub.execute_input":"2024-01-17T13:40:17.410383Z","iopub.status.idle":"2024-01-17T13:40:35.676525Z","shell.execute_reply.started":"2024-01-17T13:40:17.410355Z","shell.execute_reply":"2024-01-17T13:40:35.675443Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<b><span style='color:#FFCE30'> 7.2 |</span>  Phân Tích Dải Tần Số</b>\n\n#### Trích Xuất Đặc Trưng Dải Tần Số:\n\n* Hàm extract_frequency_band_features được thiết kế để xử lý một đoạn dữ liệu EEG. EEG là một tín hiệu phức tạp phản ánh hoạt động điện của não.\n* Hàm này chia tín hiệu EEG thành các dải tần số khác nhau: Delta, Theta, Alpha, Beta, và Gamma. Những dải tần này có ý nghĩa quan trọng trong nghiên cứu thần kinh vì chúng liên quan đến các trạng thái và hoạt động khác nhau của não.\n\n![](https://ars.els-cdn.com/content/image/3-s2.0-B9780128044902000026-f02-01-9780128044902.jpg)\n\n\n1. **Delta (0.5 – 4 Hz):**\nSóng Delta là sóng não chậm nhất và thường liên quan đến giấc ngủ sâu và quá trình phục hồi của cơ thể. Chúng xuất hiện nhiều nhất khi ngủ không mơ và đóng vai trò quan trọng trong việc chữa lành và tái tạo tế bào.\n2. **Theta (4 – 8 Hz):**\nSóng Theta xuất hiện trong giấc ngủ nông, thiền định sâu và giấc ngủ REM (Rapid Eye Movement - chuyển động mắt nhanh). Chúng có liên quan đến sự sáng tạo, trực giác, mơ mộng và hoạt động của tiềm thức.\n3. **Alpha (8 – 12 Hz):**\nSóng Alpha xuất hiện khi cơ thể và tâm trí đang ở trạng thái thư giãn nhưng vẫn tỉnh táo. Chúng thường thấy trong trạng thái thức tỉnh nhưng thoải mái, hỗ trợ điều phối tinh thần, bình tĩnh, cảnh giác và học tập.\n4. **Beta (12 – 30 Hz):**\nSóng Beta chiếm ưu thế trong trạng thái tỉnh táo bình thường khi chúng ta tập trung vào các nhiệm vụ nhận thức hoặc thế giới bên ngoài. Chúng liên quan đến tư duy năng động, sự lo lắng hoặc tập trung cao độ.\n5. **Gamma (30 – 45 Hz):**\nSóng Gamma liên quan đến hoạt động nhận thức cao cấp và quá trình củng cố thông tin. Chúng đóng vai trò quan trọng trong học tập, trí nhớ và xử lý thông tin. Sóng Gamma được xem là dải tần nhanh nhất và liên quan đến quá trình xử lý đồng thời thông tin từ nhiều vùng não khác nhau.\n\n\n\n\n* Với mỗi dải tần số, hàm áp dụng bộ lọc thông dải (bandpass filter) để cô lập tín hiệu của dải tần đó. Sau đó, nó tính toán các đặc trưng thống kê (trung bình, độ lệch chuẩn, giá trị lớn nhất và nhỏ nhất) cho mỗi dải, giúp nắm bắt đặc điểm của tín hiệu EEG trong các dải tần số khác nhau.\n* Việc sử dụng np.nanmean, np.nanstd, np.nanmax, và np.nanmin đảm bảo rằng các phép tính không bị ảnh hưởng bởi các giá trị NaN (Not a Number) có thể xuất hiện do mất tín hiệu hoặc nhiễu.\n\n#### Tổng Hợp Đặc Trưng và Ứng Dụng PCA:\n\n* Tập lệnh chính khởi tạo mô hình Phân Tích Thành Phần Chính (PCA - Principal Component Analysis) với mục tiêu giảm số chiều của tập hợp đặc trưng đã trích xuất. PCA là một kỹ thuật phổ biến giúp chuyển đổi tập dữ liệu có số chiều cao thành không gian có số chiều thấp hơn, đồng thời giữ lại phần lớn phương sai trong dữ liệu.\n* Tập lệnh lặp qua từng dòng trong tập dữ liệu huấn luyện, trích xuất các đoạn EEG và áp dụng hàm extract_frequency_band_features lên từng kênh trong các đoạn tín hiệu này. Các đặc trưng trích xuất từ tất cả các kênh sau đó được tổng hợp lại.\n* Tuy nhiên, trước khi áp dụng PCA, mọi giá trị NaN trong dữ liệu tổng hợp (data_original) được xử lý bằng phương pháp thay thế trung bình (mean imputation). Bước này đảm bảo rằng thuật toán PCA không gặp lỗi do dữ liệu bị thiếu.\n* Sau khi hoàn tất quá trình xử lý NaN, PCA được áp dụng để chuyển đổi các đặc trưng vào không gian thành phần chính, và các đặc trưng đã biến đổi này được thêm lại vào tập dữ liệu huấn luyện (train DataFrame).\n* Quá trình này cuối cùng giúp tạo ra tập hợp đặc trưng vừa súc tích vừa mang nhiều thông tin, hỗ trợ hiệu quả trong các nhiệm vụ phân loại (classification) hoặc phát hiện bất thường (anomaly detection) trên dữ liệu EEG.","metadata":{"execution":{"iopub.status.busy":"2024-01-17T03:03:50.993523Z","iopub.execute_input":"2024-01-17T03:03:50.993878Z","iopub.status.idle":"2024-01-17T03:03:51.029962Z","shell.execute_reply.started":"2024-01-17T03:03:50.993848Z","shell.execute_reply":"2024-01-17T03:03:51.028783Z"}}},{"cell_type":"code","source":"from scipy import signal\nfrom sklearn.decomposition import PCA","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:40:35.677918Z","iopub.execute_input":"2024-01-17T13:40:35.678281Z","iopub.status.idle":"2024-01-17T13:40:36.51601Z","shell.execute_reply.started":"2024-01-17T13:40:35.67825Z","shell.execute_reply":"2024-01-17T13:40:36.515134Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np\nfrom scipy import signal\n\ndef extract_frequency_band_features(segment):\n    \"\"\"\n    Hàm này trích xuất các đặc trưng của dải tần số EEG từ một đoạn tín hiệu EEG.\n\n    Đặc trưng được trích xuất từ các dải tần số phổ biến của EEG:\n    - Delta: 0.5 - 4 Hz\n    - Theta: 4 - 8 Hz\n    - Alpha: 8 - 12 Hz\n    - Beta: 12 - 30 Hz\n    - Gamma: 30 - 45 Hz\n    \n    Đối với mỗi dải tần, các đặc trưng sau được tính toán:\n    - Trung bình (mean)\n    - Độ lệch chuẩn (standard deviation)\n    - Giá trị lớn nhất (max)\n    - Giá trị nhỏ nhất (min)\n\n    Đầu vào:\n    - segment: Một đoạn tín hiệu EEG (numpy array).\n\n    Đầu ra:\n    - band_features: Danh sách chứa đặc trưng của tất cả các dải tần EEG.\n    \"\"\"\n\n    # Định nghĩa các dải tần số EEG\n    eeg_bands = {\n        'Delta': (0.5, 4),\n        'Theta': (4, 8),\n        'Alpha': (8, 12),\n        'Beta': (12, 30),\n        'Gamma': (30, 45)\n    }\n    \n    band_features = []  # Danh sách lưu đặc trưng của từng dải tần\n\n    for band in eeg_bands:\n        low, high = eeg_bands[band]  # Lấy giá trị tần số thấp nhất và cao nhất của dải tần hiện tại\n\n        # Áp dụng bộ lọc thông dải (band-pass filter) để chỉ giữ lại tín hiệu trong khoảng tần số cần thiết\n        band_pass_filter = signal.butter(\n            3, [low, high], btype='bandpass', fs=200, output='sos'\n        )  # Bộ lọc bậc 3, tần số lấy mẫu 200Hz\n\n        # Lọc tín hiệu EEG theo dải tần số hiện tại\n        filtered = signal.sosfilt(band_pass_filter, segment)\n\n        # Trích xuất các đặc trưng thống kê từ tín hiệu đã lọc\n        band_features.extend([\n            np.nanmean(filtered),  # Giá trị trung bình của tín hiệu trong dải tần\n            np.nanstd(filtered),   # Độ lệch chuẩn của tín hiệu trong dải tần\n            np.nanmax(filtered),   # Giá trị lớn nhất của tín hiệu trong dải tần\n            np.nanmin(filtered)    # Giá trị nhỏ nhất của tín hiệu trong dải tần\n        ])\n\n    return band_features  # Trả về danh sách chứa đặc trưng của tất cả dải tần","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:40:36.519187Z","iopub.execute_input":"2024-01-17T13:40:36.519926Z","iopub.status.idle":"2024-01-17T13:40:36.526465Z","shell.execute_reply.started":"2024-01-17T13:40:36.519894Z","shell.execute_reply":"2024-01-17T13:40:36.52533Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import time\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.decomposition import PCA\nimport numpy as np\n\n# Khởi tạo mô hình PCA, giữ lại 95% phương sai của dữ liệu\npca = PCA(n_components=0.95)\nprint(\"PCA model initialized.\")\n\n# Khởi tạo ma trận dữ liệu gốc để lưu trữ đặc trưng\nnum_rows = len(train)  # Số lượng mẫu trong tập dữ liệu\nnum_features = 20 * n_channels  # 20 đặc trưng cho mỗi kênh EEG\n\n# Tạo ma trận dữ liệu với giá trị ban đầu là 0\ndata_original = np.zeros((num_rows, num_features))\n\nprint(\"Starting feature extraction and PCA processing...\")\n\n# Bắt đầu đo thời gian thực thi\nstart_time = time.time()\n\n# Vòng lặp để trích xuất đặc trưng từ từng EEG sample\nfor k in range(num_rows):\n    if k % 1000 == 0:\n        print(f\"Processing row {k} of {num_rows}...\")  # In tiến trình mỗi 1000 mẫu\n\n    row = train.iloc[k]  # Lấy một dòng dữ liệu EEG\n    r = int((row['min'] + row['max']) // 4)  # Xác định vị trí trung tâm để lấy tín hiệu EEG\n\n    # Lấy một đoạn tín hiệu EEG từ spectrograms\n    eeg_segment = spectrograms[row.spec_id][r:r+300, :]\n\n    # Danh sách lưu trữ đặc trưng của tất cả kênh EEG\n    all_channel_features = []\n    \n    # Duyệt qua từng kênh EEG để trích xuất đặc trưng\n    for i in range(n_channels):\n        channel_features = extract_frequency_band_features(eeg_segment[:, i])  # Gọi hàm trích xuất đặc trưng\n        all_channel_features.extend(channel_features)  # Gộp tất cả đặc trưng lại\n    \n    # Lưu đặc trưng vào ma trận dữ liệu\n    data_original[k, :] = all_channel_features\n\nprint(\"Data matrix constructed\")\n\n# Xử lý giá trị NaN trong ma trận dữ liệu\nimputer = SimpleImputer(strategy='mean')  # Thay thế giá trị NaN bằng giá trị trung bình của mỗi cột\ndata_imputed = imputer.fit_transform(data_original)  # Áp dụng Imputer để xử lý dữ liệu\n\nprint(f\"NaN values handled. Imputed data matrix shape: {data_imputed.shape}\")\n\n# Huấn luyện mô hình PCA trên dữ liệu đã được xử lý NaN\npca.fit(data_imputed)\nprint(\"PCA fitting completed.\")\n\n# Biến đổi dữ liệu sang không gian mới với số chiều thấp hơn\ndata_pca = pca.transform(data_imputed)\n\n# Đặt tên cho các đặc trưng PCA\npca_feature_columns = [f'pca_feature_{i}' for i in range(data_pca.shape[1])]\n\n# Thêm các đặc trưng PCA vào DataFrame train\ntrain[pca_feature_columns] = data_pca\n\n# Đo thời gian thực thi toàn bộ quy trình\ntotal_time = time.time() - start_time\nprint(f\"Total processing time: {total_time:.2f} seconds.\")","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:40:36.528124Z","iopub.execute_input":"2024-01-17T13:40:36.528463Z","iopub.status.idle":"2024-01-17T13:40:36.540152Z","shell.execute_reply.started":"2024-01-17T13:40:36.528426Z","shell.execute_reply":"2024-01-17T13:40:36.538243Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:40:36.541454Z","iopub.execute_input":"2024-01-17T13:40:36.541814Z","iopub.status.idle":"2024-01-17T13:40:36.590096Z","shell.execute_reply.started":"2024-01-17T13:40:36.541776Z","shell.execute_reply":"2024-01-17T13:40:36.589007Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\n\n# Danh sách các cột KHÔNG thực hiện chuẩn hóa (các cột ID và nhãn mục tiêu)\nexcluded_columns = [\n    'eeg_id', 'spec_id', 'min', 'max', 'patient_id',  # Thông tin EEG\n    'seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote',  # Phiếu chẩn đoán\n    'target'  # Nhãn mục tiêu\n]\n\n# Lưu lại các cột bị loại trừ để giữ nguyên dữ liệu này\nexcluded_data = train[excluded_columns]\n\n# Tạo DataFrame chỉ chứa các cột cần được chuẩn hóa (bỏ các cột bị loại trừ)\nfeatures = train.drop(columns=excluded_columns)\n\n# Khởi tạo StandardScaler để chuẩn hóa dữ liệu\nscaler = StandardScaler()\n\n# Áp dụng StandardScaler lên dữ liệu đặc trưng\nfeatures_scaled = scaler.fit_transform(features)\n\n# Chuyển đổi dữ liệu đã chuẩn hóa thành DataFrame với cùng tên cột\nfeatures_scaled_df = pd.DataFrame(features_scaled, columns=features.columns)\n\n# Kết hợp lại dữ liệu đã chuẩn hóa với các cột bị loại trừ (giữ nguyên index)\ntrain_scaled_df = pd.concat([excluded_data.reset_index(drop=True), features_scaled_df], axis=1)\n\n# Hiển thị DataFrame sau khi chuẩn hóa\ntrain_scaled_df","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:40:36.591581Z","iopub.execute_input":"2024-01-17T13:40:36.592454Z","iopub.status.idle":"2024-01-17T13:40:36.598136Z","shell.execute_reply.started":"2024-01-17T13:40:36.59241Z","shell.execute_reply":"2024-01-17T13:40:36.597088Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_scaled_df.info()","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:40:36.599942Z","iopub.execute_input":"2024-01-17T13:40:36.600704Z","iopub.status.idle":"2024-01-17T13:40:36.619112Z","shell.execute_reply.started":"2024-01-17T13:40:36.600663Z","shell.execute_reply":"2024-01-17T13:40:36.617906Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"padding: 30px;color:white;margin:10;font-size:60%;text-align:left;display:fill;border-radius:10px;background-color:#FFFFFF;overflow:hidden;background-color:#FFCE30\"><b><span style='color:#FFFFFF'>8 |</span></b> <b>HUẤN LUYỆN MÔ HÌNH</b></div>\n\n* Công trình gốc sử dụng CatBoost, trong phiên bản này, chúng ta sẽ thử XGBoost để xem sự khác biệt về hiệu suất mô hình.","metadata":{"execution":{"iopub.status.busy":"2024-01-14T15:48:13.199779Z","iopub.execute_input":"2024-01-14T15:48:13.200329Z","iopub.status.idle":"2024-01-14T15:48:13.206828Z","shell.execute_reply.started":"2024-01-14T15:48:13.200282Z","shell.execute_reply":"2024-01-14T15:48:13.205408Z"}}},{"cell_type":"code","source":"# import xgboost as xgb\n# import gc\n# from sklearn.model_selection import KFold, GroupKFold\n\n# print('XGBoost version', xgb.__version__)","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:40:36.621382Z","iopub.execute_input":"2024-01-17T13:40:36.622373Z","iopub.status.idle":"2024-01-17T13:40:36.820884Z","shell.execute_reply.started":"2024-01-17T13:40:36.622328Z","shell.execute_reply":"2024-01-17T13:40:36.819697Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# all_oof = []\n# all_true = []\n# TARS = {'Seizure':0, 'LPD':1, 'GPD':2, 'LRDA':3, 'GRDA':4, 'Other':5}\n\n# gkf = GroupKFold(n_splits=5)\n# for i, (train_index, valid_index) in enumerate(gkf.split(train , train .target, train .patient_id)):   \n    \n#     print('#'*25)\n#     print(f'### Fold {i+1}')\n#     print(f'### train size {len(train_index)}, valid size {len(valid_index)}')\n#     print('#'*25)\n    \n#     model = xgb.XGBClassifier(\n#         objective='multi:softprob', \n#         num_class=len(TARS),\n#         learning_rate = 0.1, \n                      \n# #         tree_method='gpu_hist',  #skip GPU acceleration\n#     )\n    \n#     # Prepare training and validation data\n#     X_train = train.loc[train_index, FEATURES]\n#     y_train = train.loc[train_index, 'target'].map(TARS)\n#     X_valid = train.loc[valid_index, FEATURES]\n#     y_valid = train.loc[valid_index, 'target'].map(TARS)\n    \n#     model.fit(X_train, y_train, \n#               eval_set=[(X_valid, y_valid)], \n#               verbose=True, \n#               early_stopping_rounds=10)\n#     model.save_model(f'XGB_v{VER}_f{i}.model')\n    \n#     oof = model.predict_proba(X_valid)\n#     all_oof.append(oof)\n#     all_true.append(train.loc[valid_index, TARGETS].values)\n    \n#     del X_train, y_train, X_valid, y_valid, oof\n#     gc.collect()\n    \n# all_oof = np.concatenate(all_oof)\n# all_true = np.concatenate(all_true)","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:45:00.748977Z","iopub.execute_input":"2024-01-17T13:45:00.749394Z","iopub.status.idle":"2024-01-17T14:06:02.331533Z","shell.execute_reply.started":"2024-01-17T13:45:00.749361Z","shell.execute_reply":"2024-01-17T14:06:02.330233Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"padding: 30px;color:white;margin:10;font-size:60%;text-align:left;display:fill;border-radius:10px;background-color:#FFFFFF;overflow:hidden;background-color:#FFCE30\"><b><span style='color:#FFFFFF'>9 |</span></b> <b>ĐIỀU CHỈNH SIÊU THAM SỐ</b></div>\n\n### <b><span style='color:#FFCE30'> 9.1 |</span> Nhập Thư Viện và Thiết Lập Optuna</b>\n* Đầu tiên, bạn nhập các thư viện cần thiết:\n1. optuna để tối ưu hóa siêu tham số.\n2. xgboost để xây dựng mô hình học máy.\n3. log_loss từ scikit-learn làm thước đo đánh giá hiệu suất.\n4. GroupKFold để thực hiện cross-validation.\n* Lệnh optuna.create_study(direction='minimize') tạo một phiên tối ưu hóa mới. direction='minimize' nghĩa là Optuna sẽ tìm cách giảm thiểu giá trị của hàm mục tiêu, trong trường hợp này là log loss.\n\n### <b><span style='color:#FFCE30'> 9.2 |</span> Xác Định Hàm Mục Tiêu</b>\n* Hàm mục tiêu chính là hàm mà Optuna sẽ tối ưu hóa. Hàm này nhận một đối tượng trial, đối tượng này dùng để đề xuất giá trị cho các siêu tham số.\n* Bên trong hàm, ta xác định không gian siêu tham số, trong đó Optuna sẽ thử nghiệm nhiều tổ hợp khác nhau:\n1. lambda, alpha: Các tham số điều chuẩn (Regularization parameters).\n2. colsample_bytree, subsample: Tỷ lệ chọn cột và hàng.\n3. learning_rate: Tốc độ học, giúp ngăn mô hình bị overfitting.\n4. n_estimators: Số lượng cây quyết định.\n5. max_depth: Độ sâu tối đa của cây.\n6. min_child_weight: Tổng trọng số tối thiểu của các mẫu trong một nút con.\n\n### <b><span style='color:#FFCE30'> 9.3 |</span> Vòng Lặp Cross-Validation</b>\n\n* THàm này sử dụng GroupKFold để chia dữ liệu. Kỹ thuật này phù hợp khi dữ liệu có nhóm (ví dụ: mã bệnh nhân - patient_id) cần được giữ nguyên trong tập huấn luyện hoặc tập kiểm tra.\n* Trong mỗi vòng lặp cross-validation, hàm thực hiện:\n1. Chia dữ liệu thành tập huấn luyện và tập kiểm tra.\n2. Huấn luyện mô hình XGBoost với các siêu tham số do Optuna đề xuất.\n3. Tính log loss trên tập kiểm tra.\n4. Trả về giá trị log loss trung bình trên tất cả các fold. Optuna sử dụng giá trị này để xác định bộ siêu tham số tối ưu nhất.\n\n### <b><span style='color:#FFCE30'> 9.4 |</span> Chạy Quá Trình Tối Ưu Hóa Với Optuna</b>\n\n* study.optimize(objective, n_trials=100) yêu cầu Optuna chạy 100 thử nghiệm (n_trials=100) để tìm ra bộ siêu tham số tối ưu..\n* Bắt đầu với số lượng thử nghiệm nhỏ trước khi tăng dần, để quản lý thời gian tốt hơn.\n* Khi quá trình tối ưu hóa hoàn tất, các siêu tham số tốt nhất sẽ được in ra.","metadata":{}},{"cell_type":"code","source":"# import optuna\n# from sklearn.metrics import log_loss\n\n\n# def objective(trial):\n#     # Hyperparameters to be tuned by Optuna\n#     param = {\n#         'objective': 'multi:softprob',\n#         'num_class': len(TARS),\n#         'tree_method': 'gpu_hist',  # use 'gpu_hist' for GPU\n#         'lambda': trial.suggest_loguniform('lambda', 1e-4, 10.0),\n#         'alpha': trial.suggest_loguniform('alpha', 1e-4, 10.0),\n#         'colsample_bytree': trial.suggest_categorical('colsample_bytree', [0.5, 0.6, 0.7, 0.8, 0.9, 1.0]),\n#         'subsample': trial.suggest_categorical('subsample', [0.6, 0.7, 0.8, 0.9, 1.0]),\n#         'learning_rate': trial.suggest_categorical('learning_rate', [0.008, 0.01, 0.02, 0.05, 0.1]),\n#         'n_estimators': 1000,\n#         'max_depth': trial.suggest_categorical('max_depth', [5, 7, 9, 11, 13]),\n#         'min_child_weight': trial.suggest_int('min_child_weight', 1, 300),\n#     }\n\n#     gkf = GroupKFold(n_splits=5)\n#     cv_scores = []\n\n#     for train_index, valid_index in gkf.split(train, train.target, train.patient_id):\n#         X_train, X_valid = train.loc[train_index, FEATURES], train.loc[valid_index, FEATURES]\n#         y_train, y_valid = train.loc[train_index, 'target'].map(TARS), train.loc[valid_index, 'target'].map(TARS)\n\n#         model = xgb.XGBClassifier(**param)\n#         model.fit(X_train, y_train, eval_set=[(X_valid, y_valid)], verbose=False, early_stopping_rounds=10)\n#         preds = model.predict_proba(X_valid)\n#         cv_scores.append(log_loss(y_valid, preds))\n\n#     return np.mean(cv_scores)\n\n# study = optuna.create_study(direction='minimize')\n# study.optimize(objective, n_trials=10)  # Increase n_trials for more extensive search\n\n# print('Number of finished trials:', len(study.trials))\n# print('Best trial:', study.best_trial.params)","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:40:37.295221Z","iopub.status.idle":"2024-01-17T13:40:37.295669Z","shell.execute_reply.started":"2024-01-17T13:40:37.295467Z","shell.execute_reply":"2024-01-17T13:40:37.295489Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"* [I 2024-01-15 06:06:55,585] A new study created in memory with name: no-name-21c5e987-7941-4f05-8e92-2903b9a7e304\n* [I 2024-01-15 06:20:36,414] Trial 0 finished with value: 1.0734795198486586 and parameters: {'lambda': 8.578460041884545, 'alpha': 1.0420327876364774, 'colsample_bytree': 1.0, 'subsample': 1.0, 'learning_rate': 0.02, 'max_depth': 7, 'min_child_weight': 30}. Best is trial 0 with value: 1.0734795198486586.\n* [I 2024-01-15 06:39:35,894] Trial 1 finished with value: 1.0792154540612213 and parameters: {'lambda': 0.0029055528938438085, 'alpha': 0.03244985452963714, 'colsample_bytree': 0.5, 'subsample': 1.0, 'learning_rate': 0.01, 'max_depth': 11, 'min_child_weight': 77}. Best is trial 0 with value: 1.0734795198486586.\n* [I 2024-01-15 06:42:19,332] Trial 2 finished with value: 1.107279543251576 and parameters: {'lambda': 1.5919274718287213, 'alpha': 0.042136459342788604, 'colsample_bytree': 0.7, 'subsample': 0.7, 'learning_rate': 0.1, 'max_depth': 5, 'min_child_weight': 152}. Best is trial 0 with value: 1.0734795198486586.\n* [I 2024-01-15 06:57:56,577] Trial 3 finished with value: 1.107956784549986 and parameters: {'lambda': 6.873357111288177, 'alpha': 0.05538406375764404, 'colsample_bytree': 0.5, 'subsample': 0.8, 'learning_rate': 0.008, 'max_depth': 7, 'min_child_weight': 108}. Best is trial 0 with value: 1.0734795198486586.\n* [I 2024-01-15 07:12:36,921] Trial 4 finished with value: 1.1453593669298952 and parameters: {'lambda': 0.0012348624625841455, 'alpha': 6.698933350539058, 'colsample_bytree': 0.8, 'subsample': 0.9, 'learning_rate': 0.01, 'max_depth': 9, 'min_child_weight': 258}. Best is trial 0 with value: 1.0734795198486586.\n* [I 2024-01-15 07:27:45,543] Trial 5 finished with value: 1.1234561427497631 and parameters: {'lambda': 0.0779496745099949, 'alpha': 0.001997110519034328, 'colsample_bytree': 0.5, 'subsample': 0.7, 'learning_rate': 0.008, 'max_depth': 11, 'min_child_weight': 145}. Best is trial 0 with value: 1.0734795198486586.\n* [I 2024-01-15 07:30:45,104] Trial 6 finished with value: 1.139466297500727 and parameters: {'lambda': 0.011643281906929, 'alpha': 0.06334100511005662, 'colsample_bytree': 0.6, 'subsample': 0.7, 'learning_rate': 0.1, 'max_depth': 9, 'min_child_weight': 274}. Best is trial 0 with value: 1.0734795198486586.\n* [I 2024-01-15 07:33:47,478] Trial 7 finished with value: 1.074140292040741 and parameters: {'lambda': 5.487810272015954, 'alpha': 2.266845998962579, 'colsample_bytree': 0.6, 'subsample': 0.8, 'learning_rate': 0.1, 'max_depth': 7, 'min_child_weight': 51}. Best is trial 0 with value: 1.0734795198486586.\n* [I 2024-01-15 07:48:19,863] Trial 8 finished with value: 1.1705891200455967 and parameters: {'lambda': 0.03333036846711228, 'alpha': 0.0004482362892025373, 'colsample_bytree': 0.9, 'subsample': 0.8, 'learning_rate': 0.008, 'max_depth': 13, 'min_child_weight': 294}. Best is trial 0 with value: 1.0734795198486586.\n* [I 2024-01-15 07:53:41,442] Trial 9 finished with value: 1.115599617296683 and parameters: {'lambda': 0.000284724944614318, 'alpha': 0.020480207664040264, 'colsample_bytree': 0.6, 'subsample': 0.9, 'learning_rate': 0.05, 'max_depth': 9, 'min_child_weight': 248}. Best is trial 0 with value: 1.0734795198486586.\nNumber of finished trials: 10\n\n* **Best trial: {'lambda': 8.578460041884545, 'alpha': 1.0420327876364774, 'colsample_bytree': 1.0, 'subsample': 1.0, 'learning_rate': 0.02, 'max_depth': 7, 'min_child_weight': 30}**","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding: 30px;color:white;margin:10;font-size:60%;text-align:left;display:fill;border-radius:10px;background-color:#FFFFFF;overflow:hidden;background-color:#FFCE30\"><b><span style='color:#FFFFFF'>10 |</span></b> <b>TẦM QUAN TRỌNG CỦA CÁC ĐẶC TRƯNG</b></div>","metadata":{"execution":{"iopub.status.busy":"2024-01-14T16:08:35.106237Z","iopub.execute_input":"2024-01-14T16:08:35.106804Z","iopub.status.idle":"2024-01-14T16:08:35.112735Z","shell.execute_reply.started":"2024-01-14T16:08:35.106757Z","shell.execute_reply":"2024-01-14T16:08:35.111659Z"}}},{"cell_type":"code","source":"# TOP = 30\n\n# # Assuming 'model' is your trained model\n# feature_importance = model.feature_importances_\n\n# # Get the feature names from 'train'\n# feature_names = train.columns\n\n# # Sort the feature importances and get the indices of the sorted array\n# sorted_idx = np.argsort(feature_importance)\n\n# # Plot only the top 'TOP' features\n# fig = plt.figure(figsize=(10, 8))\n# plt.barh(np.arange(len(sorted_idx))[-TOP:], feature_importance[sorted_idx][-TOP:], align='center')\n# plt.yticks(np.arange(len(sorted_idx))[-TOP:], feature_names[sorted_idx][-TOP:])\n# plt.title(f'Feature Importance - Top {TOP}')\n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-17T14:12:19.822466Z","iopub.execute_input":"2024-01-17T14:12:19.823028Z","iopub.status.idle":"2024-01-17T14:12:20.406006Z","shell.execute_reply.started":"2024-01-17T14:12:19.822988Z","shell.execute_reply":"2024-01-17T14:12:20.404813Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"padding: 30px;color:white;margin:10;font-size:60%;text-align:left;display:fill;border-radius:10px;background-color:#FFFFFF;overflow:hidden;background-color:#FFCE30\"><b><span style='color:#FFFFFF'>11 |</span></b> <b>DỰ ĐOÁN TRÊN TẬP KIỂM TRA</b></div>","metadata":{"execution":{"iopub.status.busy":"2024-01-14T16:17:09.800209Z","iopub.execute_input":"2024-01-14T16:17:09.802007Z","iopub.status.idle":"2024-01-14T16:17:09.809141Z","shell.execute_reply.started":"2024-01-14T16:17:09.801957Z","shell.execute_reply":"2024-01-14T16:17:09.807551Z"}}},{"cell_type":"code","source":"# test = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/test.csv')\n# print('Test shape',test.shape)\n# test.head()","metadata":{"execution":{"iopub.status.busy":"2024-01-17T14:12:32.465945Z","iopub.execute_input":"2024-01-17T14:12:32.466394Z","iopub.status.idle":"2024-01-17T14:12:32.485369Z","shell.execute_reply.started":"2024-01-17T14:12:32.466357Z","shell.execute_reply":"2024-01-17T14:12:32.48421Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# PATH2 = '/kaggle/input/hms-harmful-brain-activity-classification/test_spectrograms/'\n# spec = pd.read_parquet(f'{PATH2}{s}.parquet')\n# spec","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:40:37.300846Z","iopub.status.idle":"2024-01-17T13:40:37.301226Z","shell.execute_reply.started":"2024-01-17T13:40:37.301049Z","shell.execute_reply":"2024-01-17T13:40:37.301068Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# %%time\n# # READ ALL TEST SPECTROGRAMS\n# PATH2 = '/kaggle/input/hms-harmful-brain-activity-classification/test_spectrograms/'\n# files = os.listdir(PATH2)\n# print(f'There are {len(files)} spectrogram parquets')\n\n# spectrograms = {}\n# for i,f in enumerate(files):\n#     if i%100==0: print(i,', ',end='')\n#     tmp = pd.read_parquet(f'{PATH2}{f}')\n#     name = int(f.split('.')[0])\n#     spectrograms_test[name] = tmp.iloc[:,1:].values","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:40:37.302384Z","iopub.status.idle":"2024-01-17T13:40:37.302745Z","shell.execute_reply.started":"2024-01-17T13:40:37.302572Z","shell.execute_reply":"2024-01-17T13:40:37.302589Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# %time\n# # ENGINEER FEATURES\n# import warnings\n# warnings.filterwarnings('ignore')\n\n# # The code generates features from the spectrogram data for use in a model \n# # The features are derived by calculating the mean and minimum values over time for each of the 400 spectrogram frequencies.\n# # Two types of windows are used for these calculations:\n# # A 10-minute window (_mean_10m, _min_10m).\n# # A 20-second window (_mean_20s, _min_20s).\n# # This process results in 1600 features (400 features × 4 calculations) for each EEG ID.\n\n# SPEC_COLS = pd.read_parquet(f'{PATH}1000086677.parquet').columns[1:]\n# FEATURES = [f'{c}_mean_10m' for c in SPEC_COLS]\n# FEATURES += [f'{c}_min_10m' for c in SPEC_COLS]\n# FEATURES += [f'{c}_mean_20s' for c in SPEC_COLS]\n# FEATURES += [f'{c}_min_20s' for c in SPEC_COLS]\n# print(f'We are creating {len(FEATURES)} features for {len(test)} rows... ',end='')\n\n\n# # A data matrix data is initialized to store the new features for each eeg_id in the train DataFrame.\n# # For each row in train, the code calculates the mean and minimum values within the specified 10-minute and 20-second windows.\n# # These calculated values are then stored in the data matrix.\n# # Finally, the matrix is added to the train DataFrame as new columns.\n\n# data = np.zeros((len(test),len(FEATURES)))\n# for k in range(len(test)):\n#     if k%100==0: print(k,', ',end='')\n#     row = test.iloc[k]\n            \n#     # 10 MINUTE WINDOW FEATURES\n#     x = np.nanmean( spec.iloc[:,1:].values, axis=0)\n#     data[k,:400] = x\n#     x = np.nanmin( spec.iloc[:,1:].values, axis=0)\n#     data[k,400:800] = x\n\n#     # 20 SECOND WINDOW FEATURES\n#     x = np.nanmean( spec.iloc[145:155,1:].values, axis=0)\n#     data[k,800:1200] = x\n#     x = np.nanmin( spec.iloc[145:155,1:].values, axis=0)\n#     data[k,1200:1600] = x\n\n#     test[FEATURES] = data\n\n    \n# print()\n# print('New test shape:',test.shape)","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:40:37.303729Z","iopub.status.idle":"2024-01-17T13:40:37.304126Z","shell.execute_reply.started":"2024-01-17T13:40:37.303944Z","shell.execute_reply":"2024-01-17T13:40:37.303963Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# from sklearn.impute import SimpleImputer\n\n# # Initialize a PCA model\n# pca = PCA(n_components=0.95)\n# print(\"PCA model initialized.\")\n\n# # Initialize an array for original features\n# num_rows = len(test)\n# num_features = 20 * n_channels  # 20 features per channel\n# data_original = np.zeros((num_rows, num_features))\n\n# print(\"Starting feature extraction and PCA processing...\")\n# start_time = time.time()\n\n# for k in range(num_rows):\n#     if k % 1000 == 0:\n#         print(f\"Processing row {k} of {num_rows}...\")\n\n#     row = train.iloc[k]\n#     eeg_segment = spectrograms_test[853520][r:r+300, :]\n\n#     # Apply the feature extraction function to each EEG channel\n#     all_channel_features = []\n#     for i in range(n_channels):\n#         channel_features = extract_frequency_band_features(eeg_segment[:, i])\n#         all_channel_features.extend(channel_features)\n    \n#     data_original[k, :] = all_channel_features\n\n# print(\"Data matrix constructed\")\n\n# # Impute NaN values in the data matrix\n# imputer = SimpleImputer(strategy='mean')\n# data_imputed = imputer.fit_transform(data_original)\n\n# print(f\"NaN values handled. Imputed data matrix shape: {data_imputed.shape}\")\n\n# # Apply PCA on the imputed data\n# pca.fit(data_imputed)\n# print(\"PCA fitting completed.\")\n\n# # Transform data using PCA\n# data_pca = pca.transform(data_imputed)\n\n# # Add PCA features to DataFrame\n# pca_feature_columns = [f'pca_feature_{i}' for i in range(data_pca.shape[1])]\n# test[pca_feature_columns] = data_pca\n\n# # Measure total processing time\n# total_time = time.time() - start_time\n# print(f\"Total processing time: {total_time:.2f} seconds.\")\n\n# test.head()","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:40:37.305186Z","iopub.status.idle":"2024-01-17T13:40:37.305554Z","shell.execute_reply.started":"2024-01-17T13:40:37.305379Z","shell.execute_reply":"2024-01-17T13:40:37.305396Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# # Columns to be excluded from scaling\n# excluded_columns = ['eeg_id', 'spectrogram_id', 'patient_id']\n\n# # Save the columns to be excluded\n# excluded_data = test[excluded_columns]\n\n# # DataFrame with only the columns to be scaled\n# features = test.drop(columns=excluded_columns)\n\n# # Initialize the StandardScaler\n# scaler = StandardScaler()\n\n# # Fit the scaler to the features and transform them\n# features_scaled = scaler.fit_transform(features)\n\n# # Create a DataFrame from the scaled features\n# features_scaled_df = pd.DataFrame(features_scaled, columns=features.columns)\n\n# # Concatenate the scaled features with the excluded columns\n# test_scaled_df = pd.concat([excluded_data.reset_index(drop=True),features_scaled_df,], axis=1)\n# test_scaled_df \n","metadata":{"execution":{"iopub.status.busy":"2024-01-17T13:40:37.308779Z","iopub.status.idle":"2024-01-17T13:40:37.309167Z","shell.execute_reply.started":"2024-01-17T13:40:37.308993Z","shell.execute_reply":"2024-01-17T13:40:37.309011Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# # FEATURE ENGINEER TEST\n# PATH2 = '/kaggle/input/hms-harmful-brain-activity-classification/test_spectrograms/'\n# data = np.zeros((len(test),len(FEATURES)))\n    \n# for k in range(len(test)):\n#     row = test.iloc[k]\n#     s = int( row.spectrogram_id )\n#     spec = pd.read_parquet(f'{PATH2}{s}.parquet')\n    \n#     # 10 MINUTE WINDOW FEATURES\n#     x = np.nanmean( spec.iloc[:,1:].values, axis=0)\n#     data[k,:400] = x\n#     x = np.nanmin( spec.iloc[:,1:].values, axis=0)\n#     data[k,400:800] = x\n\n#     # 20 SECOND WINDOW FEATURES\n#     x = np.nanmean( spec.iloc[145:155,1:].values, axis=0)\n#     data[k,800:1200] = x\n#     x = np.nanmin( spec.iloc[145:155,1:].values, axis=0)\n#     data[k,1200:1600] = x\n\n# test[FEATURES] = data\n# print('New test shape',test.shape)","metadata":{"execution":{"iopub.status.busy":"2024-01-17T14:12:38.149388Z","iopub.execute_input":"2024-01-17T14:12:38.149775Z","iopub.status.idle":"2024-01-17T14:12:38.894043Z","shell.execute_reply.started":"2024-01-17T14:12:38.149746Z","shell.execute_reply":"2024-01-17T14:12:38.893069Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# # INFER XGBOOST ON TEST\n# preds = []\n\n# for i in range(5):\n#     print(i, ', ', end='')\n    \n#     # Load the XGBoost model\n#     model = xgb.XGBClassifier()\n#     model.load_model(f'XGB_v{VER}_f{i}.model')\n    \n#     # Make predictions\n#     pred = model.predict_proba(test[FEATURES])\n#     preds.append(pred)\n\n# # Average the predictions from each fold\n# pred = np.mean(preds, axis=0)\n# print()\n# print('Test preds shape', pred.shape)","metadata":{"execution":{"iopub.status.busy":"2024-01-17T14:12:50.63105Z","iopub.execute_input":"2024-01-17T14:12:50.631491Z","iopub.status.idle":"2024-01-17T14:12:51.908914Z","shell.execute_reply.started":"2024-01-17T14:12:50.631455Z","shell.execute_reply":"2024-01-17T14:12:51.907722Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <div style=\"padding: 30px;color:white;margin:10;font-size:60%;text-align:left;display:fill;border-radius:10px;background-color:#FFFFFF;overflow:hidden;background-color:#FFCE30\"><b><span style='color:#FFFFFF'>12 |</span></b> <b>NỘP KẾT QUẢ</b></div>","metadata":{}},{"cell_type":"code","source":"# sub = pd.DataFrame({'eeg_id':test.eeg_id.values})\n# sub[TARGETS] = pred\n# sub.to_csv('submission.csv',index=False)\n# print('Submission shape',sub.shape)\n# sub.head()","metadata":{"execution":{"iopub.status.busy":"2024-01-17T14:12:56.955648Z","iopub.execute_input":"2024-01-17T14:12:56.956094Z","iopub.status.idle":"2024-01-17T14:12:56.979024Z","shell.execute_reply.started":"2024-01-17T14:12:56.956057Z","shell.execute_reply":"2024-01-17T14:12:56.978179Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# # SANITY CHECK TO CONFIRM PREDICTIONS SUM TO ONE\n# sub.iloc[:,-6:].sum(axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-01-17T14:13:01.825709Z","iopub.execute_input":"2024-01-17T14:13:01.826129Z","iopub.status.idle":"2024-01-17T14:13:01.837955Z","shell.execute_reply.started":"2024-01-17T14:13:01.826094Z","shell.execute_reply":"2024-01-17T14:13:01.83659Z"},"trusted":true},"outputs":[],"execution_count":null}]}