{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"14/2 Update: I deleted the noisy channels, and looked at the resulting models. Using Chris Deotte EfficientNet model, the score is worse with the \"bad\" channels deleted. I'm not sure what is going on, but I suspect pyprep may not deal well with epilepsy traces. \n\nI'm trying to work out if particular channels are not suitable for eeg analysis. Here I have connected the data to **pyprep** via the **mne** system. I count the number of bad channels that are detected by pyprep. This takes quite a long time to run, but gives a clear picture on which channels are perhaps suspect. My plan is to see whether deleting bad channels improves time (or perhaps also spectrogram) based model performance. \n\nIf you look at the pyprep system examples, they don't seem to work anymore. But I think pyprep itself still works. ","metadata":{}},{"cell_type":"code","source":"!pip install pyprep","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"###############################################################################\n# \nimport os\nimport pathlib\n\nimport mne\nimport numpy as np\nimport scipy.io as sio\nimport matplotlib.pyplot as plt\nimport pandas as pd\nimport gc\n\nfrom pyprep.prep_pipeline import PrepPipeline\n\nimport os\nimport contextlib\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"###############################################################################\n# Load data and prepare it\n# ------------------------\n\nid = pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/train.csv\")\neeg_path = '/kaggle/input/hms-harmful-brain-activity-classification/train_eegs/'\n\n\ndf = pd.read_parquet(eeg_path+str(1000913311)+'.parquet')\ndf = df.drop('EKG',axis=1)\nfeats = df.columns\n\neeg_unique = id['eeg_id'].drop_duplicates()\n\nbad_count = dict()\nfor feat in feats:\n    bad_count[feat] = 0 \n        \nfor eeg in eeg_unique:\n    print(eeg)\n    filename = eeg_path + str(eeg) +'.parquet'\n\n    sample_rate = 200\n    no_channels = len(feats)\n\n    info = mne.create_info(\n        ch_names=list(feats), ch_types=[\"eeg\"]*no_channels, sfreq=sample_rate\n    )\n    data = np.transpose(df.to_numpy())\n    raw = mne.io.RawArray(data, info)\n    # Make a copy of the data\n    raw_copy = raw.copy()\n    # Fit prep\n    prep_params = {\n        \"ref_chs\": \"eeg\",\n        \"reref_chs\": \"eeg\",\n        \"line_freqs\": np.arange(60, sample_rate / 2, 60),\n    }\n    montage_kind = \"standard_1005\"\n    montage = mne.channels.make_standard_montage(montage_kind)\n\n    with open(os.devnull, \"w\") as file:\n        with contextlib.redirect_stdout(file):\n            prep = PrepPipeline(raw_copy, prep_params, montage)\n            prep.fit()\n\n    print(\"interpolated channels\")\n    print(prep.interpolated_channels)\n    \n    for channel in prep.interpolated_channels:\n        bad_count[channel] = bad_count[channel] + 1\n        \n    for key in bad_count.keys():\n        print(key,bad_count[key])\n\n    gc.collect()\n\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}}]}