{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Setup","metadata":{}},{"cell_type":"code","source":"from fastai.vision.all import *\nimport timm","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-03-03T07:53:07.429724Z","iopub.execute_input":"2024-03-03T07:53:07.430071Z","iopub.status.idle":"2024-03-03T07:53:18.994726Z","shell.execute_reply.started":"2024-03-03T07:53:07.430044Z","shell.execute_reply":"2024-03-03T07:53:18.993921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path = Path('/kaggle/input/hms-harmful-brain-activity-classification')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-03-03T07:53:21.021852Z","iopub.execute_input":"2024-03-03T07:53:21.022217Z","iopub.status.idle":"2024-03-03T07:53:21.026588Z","shell.execute_reply.started":"2024-03-03T07:53:21.022189Z","shell.execute_reply":"2024-03-03T07:53:21.025722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path.ls()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-03-03T07:53:21.786033Z","iopub.execute_input":"2024-03-03T07:53:21.786478Z","iopub.status.idle":"2024-03-03T07:53:21.794482Z","shell.execute_reply.started":"2024-03-03T07:53:21.786450Z","shell.execute_reply":"2024-03-03T07:53:21.793589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Background\n\nAs part of Lesson 6 of the fastai course, I walked through Jeremy Howard's Live Coding YouTube videos where he worked on the Paddy Doctor Kaggle Competition---a competition where you had to predict the disease classification given different images of rice (some diseased, some not). I wrote an [8-part blog post series](https://vishalbakshi.github.io/blog/index.html#category=paddy%20doctor) with notes and code that correspond with each of the 6 Live Coding videos on this topic, plus a couple more with my own experiments with model training and submission.\n\nJeremy has distilled this Live Coding content into the Kaggle Notebook series, [Road to the Top](https://www.kaggle.com/code/jhoward/first-steps-road-to-the-top-part-1). I'll be following his approach in this series as best as I can apply it to this competition. \n\nI'm fresh off finishing my first live competition--which I wrote about in [this blog post](https://vishalbakshi.github.io/blog/posts/2024-02-29-first-live-competition/)--where I didn't perform well. My goal this year is to finish in the top 50% of the final leaderboard in a live competition (playground or otherwise) so I'm taking another stab at it with this one.","metadata":{}},{"cell_type":"markdown","source":"## Quick and Dirty Submission","metadata":{}},{"cell_type":"markdown","source":"I'm first going to focus on just getting a submission in successfully. I'll start by looking at the sample submission file and the test data.","metadata":{}},{"cell_type":"code","source":"ss = pd.read_csv(path/\"sample_submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:24.642398Z","iopub.execute_input":"2024-03-03T07:53:24.642801Z","iopub.status.idle":"2024-03-03T07:53:24.655243Z","shell.execute_reply.started":"2024-03-03T07:53:24.642771Z","shell.execute_reply":"2024-03-03T07:53:24.654177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ss.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:24.929915Z","iopub.execute_input":"2024-03-03T07:53:24.930281Z","iopub.status.idle":"2024-03-03T07:53:24.948771Z","shell.execute_reply.started":"2024-03-03T07:53:24.930253Z","shell.execute_reply":"2024-03-03T07:53:24.947810Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note: the `_vote` columns have to all add up to 1. ","metadata":{}},{"cell_type":"code","source":"test_df = pd.read_csv(path/\"test.csv\")\ntest_df","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:25.509747Z","iopub.execute_input":"2024-03-03T07:53:25.510102Z","iopub.status.idle":"2024-03-03T07:53:25.523343Z","shell.execute_reply.started":"2024-03-03T07:53:25.510074Z","shell.execute_reply":"2024-03-03T07:53:25.522581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(path/'test_eegs').ls()","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:25.790172Z","iopub.execute_input":"2024-03-03T07:53:25.791086Z","iopub.status.idle":"2024-03-03T07:53:25.798433Z","shell.execute_reply.started":"2024-03-03T07:53:25.791054Z","shell.execute_reply":"2024-03-03T07:53:25.797469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.read_parquet(path/'test_eegs'/'3911565283.parquet').head(3)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:26.107586Z","iopub.execute_input":"2024-03-03T07:53:26.108421Z","iopub.status.idle":"2024-03-03T07:53:26.309364Z","shell.execute_reply.started":"2024-03-03T07:53:26.108388Z","shell.execute_reply":"2024-03-03T07:53:26.308389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(path/'test_spectrograms').ls()","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:26.312155Z","iopub.execute_input":"2024-03-03T07:53:26.312870Z","iopub.status.idle":"2024-03-03T07:53:26.321245Z","shell.execute_reply.started":"2024-03-03T07:53:26.312833Z","shell.execute_reply":"2024-03-03T07:53:26.320392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.read_parquet(path/'test_spectrograms'/'853520.parquet').head(3)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:26.448732Z","iopub.execute_input":"2024-03-03T07:53:26.449075Z","iopub.status.idle":"2024-03-03T07:53:26.511626Z","shell.execute_reply.started":"2024-03-03T07:53:26.449033Z","shell.execute_reply":"2024-03-03T07:53:26.510789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, I'll look at the different forms of training data available to us, starting with the `train_eegs` parquet files.","metadata":{}},{"cell_type":"code","source":"(path/'train_eegs').ls()[:5]","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:26.943494Z","iopub.execute_input":"2024-03-03T07:53:26.943854Z","iopub.status.idle":"2024-03-03T07:53:27.215843Z","shell.execute_reply.started":"2024-03-03T07:53:26.943827Z","shell.execute_reply":"2024-03-03T07:53:27.214861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_eeg = pd.read_parquet(path/'train_eegs'/'2208063991.parquet')","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:27.273643Z","iopub.execute_input":"2024-03-03T07:53:27.274026Z","iopub.status.idle":"2024-03-03T07:53:27.291867Z","shell.execute_reply.started":"2024-03-03T07:53:27.273999Z","shell.execute_reply":"2024-03-03T07:53:27.290891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_eeg.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:27.629397Z","iopub.execute_input":"2024-03-03T07:53:27.629764Z","iopub.status.idle":"2024-03-03T07:53:27.652459Z","shell.execute_reply.started":"2024-03-03T07:53:27.629736Z","shell.execute_reply":"2024-03-03T07:53:27.651555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_eeg.shape","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:27.884662Z","iopub.execute_input":"2024-03-03T07:53:27.885488Z","iopub.status.idle":"2024-03-03T07:53:27.891250Z","shell.execute_reply.started":"2024-03-03T07:53:27.885458Z","shell.execute_reply":"2024-03-03T07:53:27.890264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_eeg.columns","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:28.124661Z","iopub.execute_input":"2024-03-03T07:53:28.125098Z","iopub.status.idle":"2024-03-03T07:53:28.131503Z","shell.execute_reply.started":"2024-03-03T07:53:28.125061Z","shell.execute_reply":"2024-03-03T07:53:28.130561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Obviously there's a lot to unpack here, but I'm probably not going to use this data for my first submission.","metadata":{}},{"cell_type":"markdown","source":"`train.csv` contains metadata for these files. I'll look at it to see what this particular EEG is classified as:","metadata":{}},{"cell_type":"code","source":"df_train = pd.read_csv(path/\"train.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:29.101100Z","iopub.execute_input":"2024-03-03T07:53:29.101751Z","iopub.status.idle":"2024-03-03T07:53:29.329352Z","shell.execute_reply.started":"2024-03-03T07:53:29.101716Z","shell.execute_reply":"2024-03-03T07:53:29.328414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:29.331034Z","iopub.execute_input":"2024-03-03T07:53:29.331349Z","iopub.status.idle":"2024-03-03T07:53:29.346579Z","shell.execute_reply.started":"2024-03-03T07:53:29.331324Z","shell.execute_reply":"2024-03-03T07:53:29.345721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are multiple subsets (by time) for each EEG and spectrogram, each with their own set of `_vote` values. Before I train the data, I'll group by `spectrogram_id` or `eeg_id`, average the each of the `_vote` columns, and pick the one with the maximum votes as the dependent variable to make it a single classification problem for my first attempt. My thinking is validated by [Chris Deotte's approach.](https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43?scriptVersionId=159911317)","metadata":{}},{"cell_type":"code","source":"df_train.query('eeg_id == 2208063991').loc[:, df_train.columns.str.endswith('vote')]","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:29.836575Z","iopub.execute_input":"2024-03-03T07:53:29.837217Z","iopub.status.idle":"2024-03-03T07:53:29.855379Z","shell.execute_reply.started":"2024-03-03T07:53:29.837187Z","shell.execute_reply":"2024-03-03T07:53:29.854492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's look at the spectrogram corresponding to this EEG:","metadata":{}},{"cell_type":"code","source":"df_train.query('eeg_id == 2208063991')['spectrogram_id']","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:30.382346Z","iopub.execute_input":"2024-03-03T07:53:30.382722Z","iopub.status.idle":"2024-03-03T07:53:30.395656Z","shell.execute_reply.started":"2024-03-03T07:53:30.382669Z","shell.execute_reply":"2024-03-03T07:53:30.394813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(path/'train_spectrograms').ls()[:3]","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:30.636809Z","iopub.execute_input":"2024-03-03T07:53:30.637447Z","iopub.status.idle":"2024-03-03T07:53:30.881443Z","shell.execute_reply.started":"2024-03-03T07:53:30.637415Z","shell.execute_reply":"2024-03-03T07:53:30.880546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_trn_sgram = pd.read_parquet(path/'train_spectrograms'/'1111500860.parquet')","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:30.944631Z","iopub.execute_input":"2024-03-03T07:53:30.945169Z","iopub.status.idle":"2024-03-03T07:53:30.988567Z","shell.execute_reply.started":"2024-03-03T07:53:30.945141Z","shell.execute_reply":"2024-03-03T07:53:30.987660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_trn_sgram.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:31.316943Z","iopub.execute_input":"2024-03-03T07:53:31.317565Z","iopub.status.idle":"2024-03-03T07:53:31.338167Z","shell.execute_reply.started":"2024-03-03T07:53:31.317536Z","shell.execute_reply":"2024-03-03T07:53:31.337166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'll use the following function from [this notebook](https://www.kaggle.com/code/clehmann10/plot-spectrograms) that was referenced in [this notebook](https://www.kaggle.com/code/oles04/introduction-eda) to visualize the spectrograms:","metadata":{}},{"cell_type":"code","source":"def plot_spectrogram(spectrogram_path):\n    sample_spect = pd.read_parquet(spectrogram_path)\n    \n    split_spect = {\n        \"LL\": sample_spect.filter(regex='^LL', axis=1),\n        \"RL\": sample_spect.filter(regex='^RL', axis=1),\n        \"RP\": sample_spect.filter(regex='^RP', axis=1),\n        \"LP\": sample_spect.filter(regex='^LP', axis=1),\n    }\n    \n    fig, axes = plt.subplots(nrows=2, ncols=2, figsize=(15, 12))\n    axes = axes.flatten()\n    label_interval = 5\n    for i, split_name in enumerate(split_spect.keys()):\n        ax = axes[i]\n        img = ax.imshow(np.log(split_spect[split_name]).T, cmap='viridis', aspect='auto', origin='lower')  # You can choose any colormap (cmap) that suits your preferences\n        cbar = fig.colorbar(img, ax=ax)\n        cbar.set_label('Log(Value)')\n        ax.set_title(split_name)\n        ax.set_ylabel(\"Frequency (Hz)\")\n        ax.set_xlabel(\"Time\")\n\n        ax.set_yticks(np.arange(len(split_spect[split_name].columns)))\n        ax.set_yticklabels([column_name[3:] for column_name in split_spect[split_name].columns])\n        frequencies = [column_name[3:] for column_name in split_spect[split_name].columns]\n        ax.set_yticks(np.arange(0, len(split_spect[split_name].columns), label_interval))\n        ax.set_yticklabels(frequencies[::label_interval])\n    plt.tight_layout()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:32.049884Z","iopub.execute_input":"2024-03-03T07:53:32.050815Z","iopub.status.idle":"2024-03-03T07:53:32.064717Z","shell.execute_reply.started":"2024-03-03T07:53:32.050764Z","shell.execute_reply":"2024-03-03T07:53:32.063404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sgram_path = path/'train_spectrograms'/'1111500860.parquet'\nplot_spectrogram(sgram_path);","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:32.472940Z","iopub.execute_input":"2024-03-03T07:53:32.473729Z","iopub.status.idle":"2024-03-03T07:53:35.320450Z","shell.execute_reply.started":"2024-03-03T07:53:32.473691Z","shell.execute_reply":"2024-03-03T07:53:35.319550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Awesome! However, this still gives me 4 inputs for each observation. I'll certainly explore how to incorporate 4 inputs in a future attempt, but for now, I just want to have one input and one output. I don't know if it's technically correct or appropriate to take the average of these four readings (I think they come from 4 different parts of the brain) but I'm going to do that for this first attempt to simplify the approach.","metadata":{}},{"cell_type":"markdown","source":"### Average Spectrogram","metadata":{}},{"cell_type":"markdown","source":"Here's the bit in that function's code that splits the parquet file into the four region's readings:","metadata":{}},{"cell_type":"code","source":"sample_spect = pd.read_parquet(sgram_path)\n    \nsplit_spect = {\n    \"LL\": sample_spect.filter(regex='^LL', axis=1),\n    \"RL\": sample_spect.filter(regex='^RL', axis=1),\n    \"RP\": sample_spect.filter(regex='^RP', axis=1),\n    \"LP\": sample_spect.filter(regex='^LP', axis=1),\n}","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:35.322433Z","iopub.execute_input":"2024-03-03T07:53:35.322885Z","iopub.status.idle":"2024-03-03T07:53:35.361485Z","shell.execute_reply.started":"2024-03-03T07:53:35.322850Z","shell.execute_reply":"2024-03-03T07:53:35.360796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"split_spect['LL'].shape, split_spect['RL'].shape, split_spect['RP'].shape, split_spect['LP'].shape","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:35.362456Z","iopub.execute_input":"2024-03-03T07:53:35.362751Z","iopub.status.idle":"2024-03-03T07:53:35.369092Z","shell.execute_reply.started":"2024-03-03T07:53:35.362727Z","shell.execute_reply":"2024-03-03T07:53:35.368150Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems like the 4 regions all have similarly named columns.","metadata":{}},{"cell_type":"code","source":"split_spect['LL'].columns[:5], \\\nsplit_spect['RL'].columns[:5], \\\nsplit_spect['RP'].columns[:5], \\\nsplit_spect['LP'].columns[:5]","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:35.370615Z","iopub.execute_input":"2024-03-03T07:53:35.370915Z","iopub.status.idle":"2024-03-03T07:53:35.380187Z","shell.execute_reply.started":"2024-03-03T07:53:35.370892Z","shell.execute_reply":"2024-03-03T07:53:35.379263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here's what I'll do---remove the prefix `XX_` from each column (`XX` being the region acronym), concatenate the four `DataFrame`s and average by index to get a 456 row x 100 column result.","metadata":{}},{"cell_type":"code","source":"def remove_prefix(df):\n    df.columns = df.columns.str.replace(r'^[A-Z]+_', '', regex=True)\n    return df\n\nsgram_avg_df = pd.concat([remove_prefix(df) for df in split_spect.values()])","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:35.941658Z","iopub.execute_input":"2024-03-03T07:53:35.942489Z","iopub.status.idle":"2024-03-03T07:53:35.949975Z","shell.execute_reply.started":"2024-03-03T07:53:35.942457Z","shell.execute_reply":"2024-03-03T07:53:35.948807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sgram_avg_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:36.377686Z","iopub.execute_input":"2024-03-03T07:53:36.378546Z","iopub.status.idle":"2024-03-03T07:53:36.384257Z","shell.execute_reply.started":"2024-03-03T07:53:36.378514Z","shell.execute_reply":"2024-03-03T07:53:36.383243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sgram_avg_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:36.877612Z","iopub.execute_input":"2024-03-03T07:53:36.878491Z","iopub.status.idle":"2024-03-03T07:53:36.898118Z","shell.execute_reply.started":"2024-03-03T07:53:36.878460Z","shell.execute_reply":"2024-03-03T07:53:36.897123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sgram_avg_df = sgram_avg_df.groupby(sgram_avg_df.index).mean()","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:37.284788Z","iopub.execute_input":"2024-03-03T07:53:37.285431Z","iopub.status.idle":"2024-03-03T07:53:37.294849Z","shell.execute_reply.started":"2024-03-03T07:53:37.285401Z","shell.execute_reply":"2024-03-03T07:53:37.293921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sgram_avg_df.shape","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:37.831426Z","iopub.execute_input":"2024-03-03T07:53:37.831795Z","iopub.status.idle":"2024-03-03T07:53:37.837718Z","shell.execute_reply.started":"2024-03-03T07:53:37.831768Z","shell.execute_reply":"2024-03-03T07:53:37.836737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sgram_avg_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:38.077668Z","iopub.execute_input":"2024-03-03T07:53:38.078055Z","iopub.status.idle":"2024-03-03T07:53:38.098767Z","shell.execute_reply.started":"2024-03-03T07:53:38.078029Z","shell.execute_reply":"2024-03-03T07:53:38.097721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks good! Let's see how it looks---","metadata":{}},{"cell_type":"code","source":"fig,ax = plt.subplots(1, figsize=(2, 2))\nfig.subplots_adjust(left=0,right=1,bottom=0,top=1)\n\nax.axis('off')\nax.imshow(np.log(sgram_avg_df).T, cmap='viridis', aspect='auto', origin='lower')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:39.117653Z","iopub.execute_input":"2024-03-03T07:53:39.117995Z","iopub.status.idle":"2024-03-03T07:53:39.214731Z","shell.execute_reply.started":"2024-03-03T07:53:39.117972Z","shell.execute_reply":"2024-03-03T07:53:39.213332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cool, at least it works. I'll wrap this all into a function.","metadata":{}},{"cell_type":"code","source":"def avg_sgram(sgram_path):\n    # read the parquet file and separate into four DataFrames\n    sample_spect = pd.read_parquet(sgram_path)\n    \n    split_spect = {\n        \"LL\": sample_spect.filter(regex='^LL', axis=1),\n        \"RL\": sample_spect.filter(regex='^RL', axis=1),\n        \"RP\": sample_spect.filter(regex='^RP', axis=1),\n        \"LP\": sample_spect.filter(regex='^LP', axis=1),\n    }\n\n    # concanate the four DataFrames with column prefixes removed\n    sgram_avg_df = pd.concat([remove_prefix(df) for df in split_spect.values()])\n    sgram_avg_df = sgram_avg_df.groupby(sgram_avg_df.index).mean()\n    \n    return sgram_avg_df","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:40.009167Z","iopub.execute_input":"2024-03-03T07:53:40.009530Z","iopub.status.idle":"2024-03-03T07:53:40.016245Z","shell.execute_reply.started":"2024-03-03T07:53:40.009502Z","shell.execute_reply":"2024-03-03T07:53:40.015194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And test it out:","metadata":{}},{"cell_type":"code","source":"fig,ax = plt.subplots(1, figsize=(2, 2))\nfig.subplots_adjust(left=0,right=1,bottom=0,top=1)\n\nax.axis('off')\nax.imshow(np.log(avg_sgram(sgram_path)).T, cmap='viridis', aspect='auto', origin='lower')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:41.029244Z","iopub.execute_input":"2024-03-03T07:53:41.029930Z","iopub.status.idle":"2024-03-03T07:53:41.131742Z","shell.execute_reply.started":"2024-03-03T07:53:41.029896Z","shell.execute_reply":"2024-03-03T07:53:41.130816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Maximum Averaged Target","metadata":{}},{"cell_type":"markdown","source":"With the input taken care of, I'll now turn to getting a single dependent variable. There are multiple rows for each `eeg_id` so I'll group by that column and take the sum of the `_vote` columns. Then, whichever `_vote` column has the maximum votes will be chosen as the target.","metadata":{}},{"cell_type":"markdown","source":"Let's first see how many unique `eeg_id` values there are---this should be the length of the final `DataFrame`.","metadata":{}},{"cell_type":"code","source":"len(df_train.eeg_id.unique())","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:43.037572Z","iopub.execute_input":"2024-03-03T07:53:43.038406Z","iopub.status.idle":"2024-03-03T07:53:43.047284Z","shell.execute_reply.started":"2024-03-03T07:53:43.038375Z","shell.execute_reply":"2024-03-03T07:53:43.046360Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.query('eeg_id == 568657')","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:43.613628Z","iopub.execute_input":"2024-03-03T07:53:43.613998Z","iopub.status.idle":"2024-03-03T07:53:43.635294Z","shell.execute_reply.started":"2024-03-03T07:53:43.613969Z","shell.execute_reply":"2024-03-03T07:53:43.634210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols = ['eeg_id', 'spectrogram_id', 'seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']\nagg_funcs = {c: 'sum' for c in cols if 'vote' in c}\n\nunique_df = df_train[cols].groupby(['eeg_id', 'spectrogram_id'], as_index=False).agg(agg_funcs)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:43.953404Z","iopub.execute_input":"2024-03-03T07:53:43.954104Z","iopub.status.idle":"2024-03-03T07:53:43.985204Z","shell.execute_reply.started":"2024-03-03T07:53:43.954073Z","shell.execute_reply":"2024-03-03T07:53:43.984417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:44.385522Z","iopub.execute_input":"2024-03-03T07:53:44.386401Z","iopub.status.idle":"2024-03-03T07:53:44.397410Z","shell.execute_reply.started":"2024-03-03T07:53:44.386365Z","shell.execute_reply":"2024-03-03T07:53:44.396381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(unique_df)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:44.809353Z","iopub.execute_input":"2024-03-03T07:53:44.810232Z","iopub.status.idle":"2024-03-03T07:53:44.815853Z","shell.execute_reply.started":"2024-03-03T07:53:44.810201Z","shell.execute_reply":"2024-03-03T07:53:44.814917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'll use the `idxmax` method (which is like the PyTorch `argmax`):","metadata":{}},{"cell_type":"code","source":"unique_df['target'] = unique_df[[c for c in cols if 'vote' in c]].idxmax(axis=1)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:45.905422Z","iopub.execute_input":"2024-03-03T07:53:45.906258Z","iopub.status.idle":"2024-03-03T07:53:45.918698Z","shell.execute_reply.started":"2024-03-03T07:53:45.906227Z","shell.execute_reply":"2024-03-03T07:53:45.917586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_df.target.unique()","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:46.445503Z","iopub.execute_input":"2024-03-03T07:53:46.445870Z","iopub.status.idle":"2024-03-03T07:53:46.453341Z","shell.execute_reply.started":"2024-03-03T07:53:46.445842Z","shell.execute_reply":"2024-03-03T07:53:46.452384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:46.921621Z","iopub.execute_input":"2024-03-03T07:53:46.922586Z","iopub.status.idle":"2024-03-03T07:53:46.933712Z","shell.execute_reply.started":"2024-03-03T07:53:46.922551Z","shell.execute_reply":"2024-03-03T07:53:46.932743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Creating a `DataBlock`","metadata":{}},{"cell_type":"markdown","source":"So far so good. Now comes the fun part! I'll start by adding `.parquet` to the `spectrogram_id` column since the `DataBlock` I'm creating will want a path object including the filename.","metadata":{}},{"cell_type":"code","source":"unique_df['spectrogram_id'] = 'train_spectrograms/' + unique_df['spectrogram_id'].astype(str) + '.parquet'","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:48.214058Z","iopub.execute_input":"2024-03-03T07:53:48.214646Z","iopub.status.idle":"2024-03-03T07:53:48.233561Z","shell.execute_reply.started":"2024-03-03T07:53:48.214614Z","shell.execute_reply":"2024-03-03T07:53:48.232730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In order to create a `DataBlock`, I need to define a function to assign to its parameter `get_x` which will take the path to the training spectogram parquet file and convert it to an image. \n\nMy `avg_sgram` function currently creates a matplotlib plot, but I need something that creates a `PILImage` object. I'll modify this function as follows:","metadata":{}},{"cell_type":"code","source":"import io\n\ndef avg_sgram(sgram_path):\n    # read the parquet file and separate into four DataFrames\n    sample_spect = pd.read_parquet(sgram_path)\n    \n    split_spect = {\n        \"LL\": sample_spect.filter(regex='^LL', axis=1),\n        \"RL\": sample_spect.filter(regex='^RL', axis=1),\n        \"RP\": sample_spect.filter(regex='^RP', axis=1),\n        \"LP\": sample_spect.filter(regex='^LP', axis=1),\n    }\n\n    # concanate the four DataFrames with column prefixes removed\n    sgram_avg_df = pd.concat([remove_prefix(df) for df in split_spect.values()])\n    sgram_avg_df = sgram_avg_df.groupby(sgram_avg_df.index).mean()\n    \n    img_buf = io.BytesIO()\n    fig.savefig(img_buf, format='png')\n    im = PILImage.create(img_buf)\n    \n    return im","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:48.933851Z","iopub.execute_input":"2024-03-03T07:53:48.934218Z","iopub.status.idle":"2024-03-03T07:53:48.941729Z","shell.execute_reply.started":"2024-03-03T07:53:48.934188Z","shell.execute_reply":"2024-03-03T07:53:48.940719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"im = avg_sgram(sgram_path)\nim","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:49.317221Z","iopub.execute_input":"2024-03-03T07:53:49.318029Z","iopub.status.idle":"2024-03-03T07:53:49.397626Z","shell.execute_reply.started":"2024-03-03T07:53:49.317995Z","shell.execute_reply":"2024-03-03T07:53:49.396809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"type(im)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:49.745368Z","iopub.execute_input":"2024-03-03T07:53:49.745968Z","iopub.status.idle":"2024-03-03T07:53:49.752010Z","shell.execute_reply.started":"2024-03-03T07:53:49.745936Z","shell.execute_reply":"2024-03-03T07:53:49.751065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, I'll add an `is_valid` column which will be `True` if the row is included in the validation set.","metadata":{}},{"cell_type":"code","source":"valid_bool = [False for _ in range(int(0.8 * len(unique_df)))]\ntrain_bool = [True for _ in range(len(unique_df) - int(0.8 * len(unique_df)))]\nis_valid_bool = pd.Series(valid_bool + train_bool).sample(frac=1).reset_index(drop=True)\nlen(is_valid_bool)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:50.825302Z","iopub.execute_input":"2024-03-03T07:53:50.826047Z","iopub.status.idle":"2024-03-03T07:53:50.839509Z","shell.execute_reply.started":"2024-03-03T07:53:50.826014Z","shell.execute_reply":"2024-03-03T07:53:50.838528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_df[\"is_valid\"] = is_valid_bool","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:51.388979Z","iopub.execute_input":"2024-03-03T07:53:51.389704Z","iopub.status.idle":"2024-03-03T07:53:51.394343Z","shell.execute_reply.started":"2024-03-03T07:53:51.389664Z","shell.execute_reply":"2024-03-03T07:53:51.393336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:51.937282Z","iopub.execute_input":"2024-03-03T07:53:51.938119Z","iopub.status.idle":"2024-03-03T07:53:51.950409Z","shell.execute_reply.started":"2024-03-03T07:53:51.938086Z","shell.execute_reply":"2024-03-03T07:53:51.949415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I have everything I need to create a `DataBlock`:\n\n`get_x=ColReader('spectrogram_id', pref=path/'train_spectrograms')` creates a filepath to the spectrogram parquet file.\n\n`get_y=ColReader('target')` gets the class name from the `target` column.\n\n`TransformBlock(type_tfms=avg_sgram, batch_tfms=IntToFloatTensor)` passes a filepath to `avg_sgram` which returns a `PILImage` object ","metadata":{}},{"cell_type":"code","source":"dblock = DataBlock(\n            blocks=(TransformBlock(type_tfms=avg_sgram, batch_tfms=IntToFloatTensor), CategoryBlock),\n            splitter=ColSplitter(),\n            get_x=ColReader('spectrogram_id', pref=path),\n            get_y=ColReader('target'))","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:53.119821Z","iopub.execute_input":"2024-03-03T07:53:53.120471Z","iopub.status.idle":"2024-03-03T07:53:53.127271Z","shell.execute_reply.started":"2024-03-03T07:53:53.120438Z","shell.execute_reply":"2024-03-03T07:53:53.126387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls = dblock.dataloaders(unique_df)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:53.818393Z","iopub.execute_input":"2024-03-03T07:53:53.819074Z","iopub.status.idle":"2024-03-03T07:53:55.389699Z","shell.execute_reply.started":"2024-03-03T07:53:53.819043Z","shell.execute_reply":"2024-03-03T07:53:55.388823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls.show_batch(nrows=1, ncols=3)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:53:55.391297Z","iopub.execute_input":"2024-03-03T07:53:55.391591Z","iopub.status.idle":"2024-03-03T07:54:00.422396Z","shell.execute_reply.started":"2024-03-03T07:53:55.391565Z","shell.execute_reply":"2024-03-03T07:54:00.421479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check vocab\ndls.vocab","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:54:00.424174Z","iopub.execute_input":"2024-03-03T07:54:00.424770Z","iopub.status.idle":"2024-03-03T07:54:00.430834Z","shell.execute_reply.started":"2024-03-03T07:54:00.424734Z","shell.execute_reply":"2024-03-03T07:54:00.429889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Training an Image Classifier","metadata":{}},{"cell_type":"markdown","source":"With a `DataLoaders` object working as expected, now I can train a model. For this submission, I'll train a resnet.","metadata":{}},{"cell_type":"code","source":"# model = timm.create_model(model_name='resnet18', pretrained=False)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:54:03.377526Z","iopub.execute_input":"2024-03-03T07:54:03.378249Z","iopub.status.idle":"2024-03-03T07:54:03.382775Z","shell.execute_reply.started":"2024-03-03T07:54:03.378214Z","shell.execute_reply":"2024-03-03T07:54:03.381713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# model.load_state_dict(torch.load('/kaggle/input/resnet18/resnet18.pth'))","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:54:03.836726Z","iopub.execute_input":"2024-03-03T07:54:03.837647Z","iopub.status.idle":"2024-03-03T07:54:03.841957Z","shell.execute_reply.started":"2024-03-03T07:54:03.837608Z","shell.execute_reply":"2024-03-03T07:54:03.840824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn = vision_learner(dls, timm.models.resnet18, metrics=accuracy, pretrained=False).to_fp16()","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:54:04.832521Z","iopub.execute_input":"2024-03-03T07:54:04.833309Z","iopub.status.idle":"2024-03-03T07:54:05.192303Z","shell.execute_reply.started":"2024-03-03T07:54:04.833279Z","shell.execute_reply":"2024-03-03T07:54:05.191456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.fine_tune(1)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T07:54:08.497204Z","iopub.execute_input":"2024-03-03T07:54:08.497538Z","iopub.status.idle":"2024-03-03T08:10:29.434075Z","shell.execute_reply.started":"2024-03-03T07:54:08.497506Z","shell.execute_reply":"2024-03-03T08:10:29.432870Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Okay great, the model at least runs without error.","metadata":{}},{"cell_type":"markdown","source":"### Creating a Submission File","metadata":{}},{"cell_type":"markdown","source":"I need to add `.parquet` to the `spectrogram_id` column of the test `DataFrame`:","metadata":{}},{"cell_type":"code","source":"test_df['spectrogram_id'] = 'test_spectrograms/' + test_df['spectrogram_id'].astype(str) + '.parquet'","metadata":{"execution":{"iopub.status.busy":"2024-03-03T08:10:42.839813Z","iopub.execute_input":"2024-03-03T08:10:42.840786Z","iopub.status.idle":"2024-03-03T08:10:42.848371Z","shell.execute_reply.started":"2024-03-03T08:10:42.840749Z","shell.execute_reply":"2024-03-03T08:10:42.847288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df","metadata":{"execution":{"iopub.status.busy":"2024-03-03T08:10:43.949004Z","iopub.execute_input":"2024-03-03T08:10:43.949389Z","iopub.status.idle":"2024-03-03T08:10:43.960085Z","shell.execute_reply.started":"2024-03-03T08:10:43.949360Z","shell.execute_reply":"2024-03-03T08:10:43.958947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now I can create a `test_dl`:","metadata":{}},{"cell_type":"code","source":"tst_dl = learn.dls.test_dl(test_df)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T08:10:58.084884Z","iopub.execute_input":"2024-03-03T08:10:58.085249Z","iopub.status.idle":"2024-03-03T08:10:58.093224Z","shell.execute_reply.started":"2024-03-03T08:10:58.085221Z","shell.execute_reply":"2024-03-03T08:10:58.092266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tst_dl.show_batch()","metadata":{"execution":{"iopub.status.busy":"2024-03-03T08:11:00.742621Z","iopub.execute_input":"2024-03-03T08:11:00.743384Z","iopub.status.idle":"2024-03-03T08:11:00.918547Z","shell.execute_reply.started":"2024-03-03T08:11:00.743351Z","shell.execute_reply":"2024-03-03T08:11:00.917255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"probs,_,idxs = learn.get_preds(dl=tst_dl, with_decoded=True)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T08:11:06.960764Z","iopub.execute_input":"2024-03-03T08:11:06.961405Z","iopub.status.idle":"2024-03-03T08:11:07.326488Z","shell.execute_reply.started":"2024-03-03T08:11:06.961377Z","shell.execute_reply":"2024-03-03T08:11:07.325512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I referenced [Sonu Jha's notebook](https://www.kaggle.com/code/sonujha090/hms-hbac-fastai-starter/notebook) for the first line of this next step:","metadata":{}},{"cell_type":"code","source":"probs_df = pd.DataFrame(probs, columns=dls.vocab)\nprobs_df['eeg_id'] = test_df['eeg_id']\nprobs_df = probs_df[['eeg_id', 'seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']]\nprobs_df","metadata":{"execution":{"iopub.status.busy":"2024-03-03T08:11:10.490130Z","iopub.execute_input":"2024-03-03T08:11:10.491091Z","iopub.status.idle":"2024-03-03T08:11:10.507714Z","shell.execute_reply.started":"2024-03-03T08:11:10.491056Z","shell.execute_reply":"2024-03-03T08:11:10.506718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"probs_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-03-03T08:11:13.169490Z","iopub.execute_input":"2024-03-03T08:11:13.169898Z","iopub.status.idle":"2024-03-03T08:11:13.177890Z","shell.execute_reply.started":"2024-03-03T08:11:13.169868Z","shell.execute_reply":"2024-03-03T08:11:13.177041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!head submission.csv","metadata":{"execution":{"iopub.status.busy":"2024-03-03T08:11:14.568805Z","iopub.execute_input":"2024-03-03T08:11:14.569504Z","iopub.status.idle":"2024-03-03T08:11:15.534740Z","shell.execute_reply.started":"2024-03-03T08:11:14.569476Z","shell.execute_reply":"2024-03-03T08:11:15.533694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Next Steps","metadata":{}},{"cell_type":"markdown","source":"The first task I need to figure out is how to load a pretrained model's weights with internet access disabled. That might take awhile, but if I can't figure that out I can't actually compete in this competition. I now understand why code competitions are tougher.\n\nI have some hope though thanks to the Kaggle community---I'll be referencing [this training notebook](https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-training#Inference-Out-Of-Fold) and [this inference notebook](https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-inference) by Tawara. It seems like you are allowed to train a model in one notebook (that you don't use for submission), export it, and load it into a separate inference notebook (that you **do** use for submission) as long as it uses a freely/publicly available architecture like from `timm`. \n\nThe second task I want to work on next is to decrease the time it takes to train a model so I can iterate faster. I think in order to do that, I should first create a temporary folder with all of the training and test images, similar to done in [Awsaf et. al's notebook](https://www.kaggle.com/code/awsaf49/hms-hbac-kerascv-starter-notebook) referenced in [Sonu Jha's notebook](https://www.kaggle.com/code/sonujha090/hms-hbac-fastai-starter/notebook).\n\nAfter that, I'll generally follow Jeremy's approach in the Road to the Top series, submitting along the way to see if each step improves my score, starting with two initial steps:\n\n- train a resnet34 and improve the accuracy (more epochs, maybe a different learning rate).\n- try a quick and dirty RandomForestRegressor.\n\nDo a series of vision model experiments, submitting after each step to see if how it changes the score:\n\n- train different small vision models. \n- try an ensemble of small vision models.\n- consider if I want to use data augmentations (I don't think it will work well, but will see).\n- train larger vision models based on the smaller ones that did well.\n- try an ensemble of large vision models.\n- try multi-category classification.\n- try multiple inputs or different inputs than an average of the four regions.\n\nDo a series of RandomForestRegressor experiments, submitting after each step to see if how it changes the score:\n\n- remove low importance features.\n- remove redundant features.\n- remove out-of-domain features.\n- try different hyperparameters.\n- try a neural net.\n- try an ensemble with a Random Forest and neural net.\n- train a RF with NN embeddings and ensemble the two.\n\nHopefully I have enough time, with one month remaining in the competition, to try all of this out!\n","metadata":{}}]}