{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":59093,"databundleVersionId":7469972,"sourceType":"competition"},{"sourceId":7392775,"sourceType":"datasetVersion","datasetId":4297782},{"sourceId":7447509,"sourceType":"datasetVersion","datasetId":4334995}],"dockerImageVersionId":30635,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Neural networks like input data to have a Gaussian distribution (i.e. normal bell shaped histogram with mean=0 and std=1). Therefore we transform inputs to transform the distribution.\n\nThe most common transformation is standardization which is new data = (old data - mean) / std. When the original data has a skewed distribution (i.e. the histogram has a tail extending on only one side), then we use a log transform to remove the skew. So first we log transform and next we standardize.\n\nLog transforms can only handle positive numbers (i.e x>0), so we must shift, clip, and/or flip all the data to be data > 0 before we perform log transform.\n\nAnother trick is using Gauss Rank Transform [here](\"https://medium.com/rapids-ai/gauss-rank-transformation-is-100x-faster-with-rapids-and-cupy-7c947e3397da\"). Or using a Quantile Transform [here](\"https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.QuantileTransformer.html\"). Both can transform any distribution into a nearly perfect Gaussian distribution afterward.\n\nBy @cdeotte in this discuss [Magic Formula to Convert EEG to Spectrograms!](\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760\")\n\n","metadata":{}},{"cell_type":"markdown","source":"# Motivation\n```python\nnp.clip(img,np.exp(-4),np.exp(8)) # improved LB score by 0.10 introduced by Chris\n```\n\n# Back to Basics","metadata":{}},{"cell_type":"markdown","source":"![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F2e5159c656d78753e70132867945a7a8%2Fposter.jpeg?generation=1706062656796456&alt=media)\n- **Source - bing**","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nspectrograms = np.load('/kaggle/input/brain-spectrograms/specs.npy',allow_pickle=True).item()\neeg_spectrograms = np.load('/kaggle/input/brain-eeg-spectrograms/eeg_specs.npy',allow_pickle=True).item()","metadata":{"execution":{"iopub.status.busy":"2024-01-24T01:09:53.312581Z","iopub.execute_input":"2024-01-24T01:09:53.313073Z","iopub.status.idle":"2024-01-24T01:12:23.985624Z","shell.execute_reply.started":"2024-01-24T01:09:53.313039Z","shell.execute_reply":"2024-01-24T01:12:23.983423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(spectrograms), len(eeg_spectrograms)","metadata":{"execution":{"iopub.status.busy":"2024-01-24T03:15:53.411102Z","iopub.execute_input":"2024-01-24T03:15:53.411588Z","iopub.status.idle":"2024-01-24T03:15:53.420254Z","shell.execute_reply.started":"2024-01-24T03:15:53.411553Z","shell.execute_reply":"2024-01-24T03:15:53.419077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_spectrograms = spectrograms[next(iter(spectrograms))].reshape(-1)\nsample_eeg_spectrograms = eeg_spectrograms[next(iter(eeg_spectrograms))].reshape(-1).reshape(-1)","metadata":{"execution":{"iopub.status.busy":"2024-01-24T03:15:54.281178Z","iopub.execute_input":"2024-01-24T03:15:54.281623Z","iopub.status.idle":"2024-01-24T03:15:54.289891Z","shell.execute_reply.started":"2024-01-24T03:15:54.281590Z","shell.execute_reply":"2024-01-24T03:15:54.288289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12, 6))\nsns.histplot(sample_spectrograms, kde=True, bins=50, color='green')\nplt.title(f'Distribution of the Sample Spectrogram Values, Mean: {round(np.nanmean(sample_spectrograms),1)} Std: {round(np.nanstd(sample_spectrograms),1)}')\nplt.xlabel('Spectrogram Value')\nplt.ylabel('Freq')\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-24T03:15:54.718388Z","iopub.execute_input":"2024-01-24T03:15:54.719225Z","iopub.status.idle":"2024-01-24T03:15:56.740578Z","shell.execute_reply.started":"2024-01-24T03:15:54.719183Z","shell.execute_reply":"2024-01-24T03:15:56.739385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12, 6))\nsns.histplot(sample_eeg_spectrograms, kde=True, bins=50, color='blue')\nplt.title(f\"Distribution of the Sample EEG's Spectrogram Values, Mean: {round(np.nanmean(sample_eeg_spectrograms),1)} Std: {round(np.nanstd(sample_eeg_spectrograms),1)}\")\nplt.xlabel(\"EEG's Spectrogram Value\")\nplt.ylabel('Freq')\nplt.grid(True)\n\"mean\", np.mean(sample_eeg_spectrograms), \"std\", np.std(sample_eeg_spectrograms)","metadata":{"execution":{"iopub.status.busy":"2024-01-24T03:15:56.743409Z","iopub.execute_input":"2024-01-24T03:15:56.743834Z","iopub.status.idle":"2024-01-24T03:15:58.089819Z","shell.execute_reply.started":"2024-01-24T03:15:56.743797Z","shell.execute_reply":"2024-01-24T03:15:58.087810Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Spectrogram: Its left Skewed Distribution\n# EEG's Spectrogram: Already in Gaussian Distribution as Chris already normalised","metadata":{}},{"cell_type":"code","source":"sample_spectrograms_log = np.log(sample_spectrograms)\nplt.figure(figsize=(12, 6))\nsns.histplot(sample_spectrograms_log, kde=True, bins=50, color='green')\nplt.title(f'Distribution of the Sample Spectrogram Log Values, Mean: {round(np.nanmean(sample_spectrograms_log),1)} Std: {round(np.nanstd(sample_spectrograms_log),1)}')\nplt.xlabel('Spectrogram Value')\nplt.ylabel('Freq')\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-24T03:15:58.091632Z","iopub.execute_input":"2024-01-24T03:15:58.092127Z","iopub.status.idle":"2024-01-24T03:16:00.239840Z","shell.execute_reply.started":"2024-01-24T03:15:58.092087Z","shell.execute_reply":"2024-01-24T03:16:00.238634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Not bad, similar to Guesian Distribution (mean is not 0) - lets normalise","metadata":{}},{"cell_type":"code","source":"ep = 1e-6\nm = np.nanmean(sample_spectrograms_log)\ns = np.nanstd(sample_spectrograms_log)\nsample_spectrograms_log_norm = (sample_spectrograms_log-m)/(s+ep)\nsample_spectrograms_log_norm = np.nan_to_num(sample_spectrograms_log_norm, nan=0.0)\nplt.figure(figsize=(12, 6))\nsns.histplot(sample_spectrograms_log_norm, kde=True, bins=50, color='green')\nplt.title(f'Distribution of the Sample Spectrogram Log Norm Values, Mean: {round(np.mean(sample_spectrograms_log_norm),2)} Std: {round(np.std(sample_spectrograms_log_norm),2)}')\nplt.xlabel('Spectrogram Value')\nplt.ylabel('Freq')\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-24T03:16:00.242641Z","iopub.execute_input":"2024-01-24T03:16:00.243056Z","iopub.status.idle":"2024-01-24T03:16:02.350695Z","shell.execute_reply.started":"2024-01-24T03:16:00.243020Z","shell.execute_reply":"2024-01-24T03:16:02.349344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Mean and Std are now Gaussian distribution","metadata":{}},{"cell_type":"markdown","source":"# Lets reverse engineering the -> np.clip(img,np.exp(-4),np.exp(8))\n> As we seen log values range from -4 to 10 but mean ~ 0.69 and std ~ 1.7 i.e in norm distribution we have values from -2 to 6 but mean = 0 and std = 1\n\n# Outlers ( 3 sigma )\n![](https://news.mit.edu/sites/default/files/styles/news_article__image_gallery/public/images/201202/20120208160239-1_0.jpg?itok=1X1a_HCs)\n\n### Outliers range ( 0.69 - 3 * 1.7 ~ -4.41 and 0.69 + 3 * 1.7 ~ 5.79 ) close to Chris clip range -4 to 8","metadata":{}},{"cell_type":"code","source":"sample_spectrograms_log_norm_clip = np.clip(sample_spectrograms_log, -4, 8)\nplt.figure(figsize=(12, 6))\nsns.histplot(sample_spectrograms_log_norm_clip, kde=True, bins=50, color='orange')\nplt.title(f'Distribution of the Sample Spectrogram Log Christ Clip Values, Mean: {round(np.mean(sample_spectrograms_log_norm_clip),2)} Std: {round(np.std(sample_spectrograms_log_norm_clip),2)}')\nplt.xlabel('Spectrogram Value')\nplt.ylabel('Freq')\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-24T03:16:02.352805Z","iopub.execute_input":"2024-01-24T03:16:02.353759Z","iopub.status.idle":"2024-01-24T03:16:04.443210Z","shell.execute_reply.started":"2024-01-24T03:16:02.353710Z","shell.execute_reply":"2024-01-24T03:16:04.441963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_spectrograms_log_norm_clip = np.clip(sample_spectrograms_log, -4.41, 5.79)\nplt.figure(figsize=(12, 6))\nsns.histplot(sample_spectrograms_log_norm_clip, kde=True, bins=50, color='red')\nplt.title(f'Distribution of the Sample Spectrogram Log 3-Sigma Clip Values, Mean: {round(np.mean(sample_spectrograms_log_norm_clip),2)} Std: {round(np.std(sample_spectrograms_log_norm_clip),2)}')\nplt.xlabel('Spectrogram Value')\nplt.ylabel('Freq')\nplt.grid(True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-24T03:16:04.445286Z","iopub.execute_input":"2024-01-24T03:16:04.446316Z","iopub.status.idle":"2024-01-24T03:16:07.302685Z","shell.execute_reply.started":"2024-01-24T03:16:04.446264Z","shell.execute_reply":"2024-01-24T03:16:07.301387Z"},"trusted":true},"execution_count":null,"outputs":[]}]}