{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<img src=\"https://media.wired.com/photos/5c6c53bc0135c22d73b8ccd4/master/pass/ligo.jpg\">","metadata":{}},{"cell_type":"markdown","source":"I am using spectrogram images generated by @yasufuminakama which can be found here for <a href=\"https://www.kaggle.com/yasufuminakama/g2net-spectrogram-generation-train\">train</a> and here for <a href=\"https://www.kaggle.com/yasufuminakama/g2net-spectrogram-generation-test\">test</a>","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom matplotlib import pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nimport tensorflow as tf\nimport keras\nimport keras.layers as L\nimport math\nfrom keras.utils import Sequence\nfrom keras.preprocessing import image\nfrom random import shuffle\nfrom sklearn.model_selection import train_test_split","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:09:29.073059Z","iopub.execute_input":"2021-07-07T08:09:29.07356Z","iopub.status.idle":"2021-07-07T08:09:38.035751Z","shell.execute_reply.started":"2021-07-07T08:09:29.073456Z","shell.execute_reply":"2021-07-07T08:09:38.03489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# About the Competition🚩\n<p style=\"font-size:15px\">It's been said that teamwork makes the dream work. This couldn't be truer for the breakthrough discovery of gravitational waves (GW), signals from colliding binary black holes in 2015. It required the collaboration of experts in physics, mathematics, information science, and computing. GW signals have led researchers to observe a new population of massive, stellar-origin black holes, to unlock the mysteries of neutron star mergers, and to measure the expansion of the Universe. These signals are unimaginably tiny ripples in the fabric of space-time and even though the global network of GW detectors are some of the most sensitive instruments on the planet, the signals are buried in detector noise. Analysis of GW data and the detection of these signals is a crucial mission for the growing global network of increasingly sensitive GW detectors. These challenges in data analysis and noise characterization could be solved with the help of data science.\n</p><p>\nIn this competition, you’ll aim to detect GW signals from the mergers of binary black holes. Specifically, you'll build a model to analyze simulated GW time-series data from a network of Earth-based detectors.\n</p><p>\nSubmissions are evaluated on area under the ROC curve between the predicted probability and the observed target.\n</p>","metadata":{}},{"cell_type":"markdown","source":"# Data Description","metadata":{}},{"cell_type":"markdown","source":"<div style=\"font-size:15px\">\n We are given 2 npy files and 2 csv files:-\n<ul>\n    <li><code>train:</code> the training set files, one npy file per observation; labels are provided in a files shown below\n</li>\n    <li><code>test:</code> the test set files; you must predict the probability that the observation contains a gravitational wave</li>\n    <li><code>training_labels.csv:</code> target values of whether the associated signal contains a gravitational wave</li>\n    <li><code>sample_submission.csv:</code> a sample submission file in the correct format\n</li>\n</ul>    \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Train Dataset","metadata":{}},{"cell_type":"markdown","source":"# EDA","metadata":{}},{"cell_type":"code","source":"train_labels = pd.read_csv('../input/g2net-gravitational-wave-detection/training_labels.csv')\nsample_submission = pd.read_csv('../input/g2net-gravitational-wave-detection/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:09:38.037292Z","iopub.execute_input":"2021-07-07T08:09:38.037867Z","iopub.status.idle":"2021-07-07T08:09:38.692509Z","shell.execute_reply.started":"2021-07-07T08:09:38.037821Z","shell.execute_reply":"2021-07-07T08:09:38.691339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets take a look at train_labels","metadata":{}},{"cell_type":"code","source":"train_labels.head()","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:09:38.694147Z","iopub.execute_input":"2021-07-07T08:09:38.694494Z","iopub.status.idle":"2021-07-07T08:09:38.720611Z","shell.execute_reply.started":"2021-07-07T08:09:38.694463Z","shell.execute_reply":"2021-07-07T08:09:38.719487Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(x='target',data=train_labels,palette='Set2')","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:09:38.722193Z","iopub.execute_input":"2021-07-07T08:09:38.722623Z","iopub.status.idle":"2021-07-07T08:09:38.918539Z","shell.execute_reply.started":"2021-07-07T08:09:38.722581Z","shell.execute_reply":"2021-07-07T08:09:38.917467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path = list(train_labels['id'])\nfor i in range(len(path)):\n    path[i] = '../input/g2net-gravitational-wave-detection/train/' +path[i][0]+'/'+path[i][1]+'/'+path[i][2]+'/' + path[i] + '.npy'","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:09:39.731079Z","iopub.execute_input":"2021-07-07T08:09:39.731469Z","iopub.status.idle":"2021-07-07T08:09:40.510201Z","shell.execute_reply.started":"2021-07-07T08:09:39.731436Z","shell.execute_reply":"2021-07-07T08:09:40.509058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def id2path(idx,is_train=True):\n    path = '../input/g2net-gravitational-wave-detection'\n    if is_train:\n        path += '/train/'+idx[0]+'/'+idx[1]+'/'+idx[2]+'/'+idx+'.npy'\n    else:\n        path += '/test/'+idx[0]+'/'+idx[1]+'/'+idx[2]+'/'+idx+'.npy'\n    return path","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:09:41.471628Z","iopub.execute_input":"2021-07-07T08:09:41.471979Z","iopub.status.idle":"2021-07-07T08:09:41.480352Z","shell.execute_reply.started":"2021-07-07T08:09:41.471948Z","shell.execute_reply":"2021-07-07T08:09:41.479147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install -q nnAudio","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:09:41.945187Z","iopub.execute_input":"2021-07-07T08:09:41.945819Z","iopub.status.idle":"2021-07-07T08:09:50.867076Z","shell.execute_reply.started":"2021-07-07T08:09:41.945773Z","shell.execute_reply":"2021-07-07T08:09:50.8654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch\nfrom nnAudio.Spectrogram import CQT1992v2\ndef increase_dimension(idx,is_train,transform=CQT1992v2(sr=2048, fmin=20, fmax=1024, hop_length=64)): # in order to use efficientnet we need 3 dimension images\n    waves = np.load(id2path(idx,is_train))\n    waves = np.hstack(waves)\n    waves = waves / np.max(waves)\n    waves = torch.from_numpy(waves).float()\n    image = transform(waves)\n    image = np.array(image)\n    image = np.transpose(image,(1,2,0))\n    return image","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:09:50.86919Z","iopub.execute_input":"2021-07-07T08:09:50.869547Z","iopub.status.idle":"2021-07-07T08:09:52.055861Z","shell.execute_reply.started":"2021-07-07T08:09:50.869509Z","shell.execute_reply":"2021-07-07T08:09:52.054853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"lets take a look at example file","metadata":{}},{"cell_type":"code","source":"example = np.load(path[0])","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:09:52.057537Z","iopub.execute_input":"2021-07-07T08:09:52.05784Z","iopub.status.idle":"2021-07-07T08:09:52.0733Z","shell.execute_reply.started":"2021-07-07T08:09:52.057803Z","shell.execute_reply":"2021-07-07T08:09:52.072262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig,a =  plt.subplots(3,1)\na[0].plot(example[1],color='green')\na[1].plot(example[1],color='red')\na[2].plot(example[1],color='yellow')\nfig.suptitle('Target 1', fontsize=16)\nplt.show()\nplt.imshow(increase_dimension(train_labels['id'][0],is_train=True))\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-07-07T08:09:52.074751Z","iopub.execute_input":"2021-07-07T08:09:52.075102Z","iopub.status.idle":"2021-07-07T08:09:52.639507Z","shell.execute_reply.started":"2021-07-07T08:09:52.075071Z","shell.execute_reply":"2021-07-07T08:09:52.638368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"example = np.load(path[1])\nfig,a =  plt.subplots(3,1)\na[0].plot(example[1],color='green')\na[1].plot(example[1],color='red')\na[2].plot(example[1],color='yellow')\nfig.suptitle('Target 0', fontsize=16)\nplt.show()\nplt.imshow(increase_dimension(train_labels['id'][1],is_train=True))\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-07-07T08:09:52.641017Z","iopub.execute_input":"2021-07-07T08:09:52.641344Z","iopub.status.idle":"2021-07-07T08:09:53.120115Z","shell.execute_reply.started":"2021-07-07T08:09:52.64131Z","shell.execute_reply":"2021-07-07T08:09:53.118882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"example = np.load(path[4])\nfig,a =  plt.subplots(3,1)\na[0].plot(example[1],color='green')\na[1].plot(example[1],color='red')\na[2].plot(example[1],color='yellow')\nfig.suptitle('Target 1', fontsize=16)\nplt.show()\nplt.imshow(increase_dimension(train_labels['id'][4],is_train=True))\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-07-07T08:09:53.121969Z","iopub.execute_input":"2021-07-07T08:09:53.122394Z","iopub.status.idle":"2021-07-07T08:09:53.604927Z","shell.execute_reply.started":"2021-07-07T08:09:53.122348Z","shell.execute_reply":"2021-07-07T08:09:53.603856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"example = np.load(path[2])\nfig,a =  plt.subplots(3,1)\na[0].plot(example[1],color='green')\na[1].plot(example[1],color='red')\na[2].plot(example[1],color='yellow')\nfig.suptitle('Target 0', fontsize=16)\nplt.show()\nplt.imshow(increase_dimension(train_labels['id'][2],is_train=True))\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-07-07T08:09:53.606205Z","iopub.execute_input":"2021-07-07T08:09:53.606543Z","iopub.status.idle":"2021-07-07T08:09:54.070537Z","shell.execute_reply.started":"2021-07-07T08:09:53.606512Z","shell.execute_reply":"2021-07-07T08:09:54.069118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model","metadata":{}},{"cell_type":"code","source":"class Dataset(Sequence):\n    def __init__(self,idx,y=None,batch_size=256,shuffle=True):\n        self.idx = idx\n        self.batch_size = batch_size\n        self.shuffle = shuffle\n        if y is not None:\n            self.is_train=True\n        else:\n            self.is_train=False\n        self.y = y\n    def __len__(self):\n        return math.ceil(len(self.idx)/self.batch_size)\n    def __getitem__(self,ids):\n        batch_ids = self.idx[ids * self.batch_size:(ids + 1) * self.batch_size]\n        if self.y is not None:\n            batch_y = self.y[ids * self.batch_size: (ids + 1) * self.batch_size]\n            \n        list_x = np.array([increase_dimension(x,self.is_train) for x in batch_ids])\n        batch_X = np.stack(list_x)\n        if self.is_train:\n            return batch_X, batch_y\n        else:\n            return batch_X\n    \n    def on_epoch_end(self):\n        if self.shuffle and self.is_train:\n            ids_y = list(zip(self.idx, self.y))\n            shuffle(ids_y)\n            self.idx, self.y = list(zip(*ids_y))","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:09:54.072314Z","iopub.execute_input":"2021-07-07T08:09:54.072618Z","iopub.status.idle":"2021-07-07T08:09:54.346295Z","shell.execute_reply.started":"2021-07-07T08:09:54.072586Z","shell.execute_reply":"2021-07-07T08:09:54.345013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_idx =  train_labels['id'].values\ny = train_labels['target'].values\ntest_idx = sample_submission['id'].values","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:09:54.348642Z","iopub.execute_input":"2021-07-07T08:09:54.349545Z","iopub.status.idle":"2021-07-07T08:09:54.360369Z","shell.execute_reply.started":"2021-07-07T08:09:54.349498Z","shell.execute_reply":"2021-07-07T08:09:54.359386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_train,x_valid,y_train,y_valid = train_test_split(train_idx,y,test_size=0.05,random_state=42,stratify=y)","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:09:54.527759Z","iopub.execute_input":"2021-07-07T08:09:54.528111Z","iopub.status.idle":"2021-07-07T08:09:55.037031Z","shell.execute_reply.started":"2021-07-07T08:09:54.52808Z","shell.execute_reply":"2021-07-07T08:09:55.036192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dataset = Dataset(x_train,y_train)\nvalid_dataset = Dataset(x_valid,y_valid)\ntest_dataset = Dataset(test_idx)","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:09:55.05213Z","iopub.execute_input":"2021-07-07T08:09:55.052525Z","iopub.status.idle":"2021-07-07T08:09:55.056933Z","shell.execute_reply.started":"2021-07-07T08:09:55.052492Z","shell.execute_reply":"2021-07-07T08:09:55.055882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install -U efficientnet","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:09:56.922314Z","iopub.execute_input":"2021-07-07T08:09:56.922832Z","iopub.status.idle":"2021-07-07T08:10:04.339461Z","shell.execute_reply.started":"2021-07-07T08:09:56.922767Z","shell.execute_reply":"2021-07-07T08:10:04.338274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import efficientnet.keras as efn","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:10:04.343919Z","iopub.execute_input":"2021-07-07T08:10:04.344308Z","iopub.status.idle":"2021-07-07T08:10:04.725778Z","shell.execute_reply.started":"2021-07-07T08:10:04.344265Z","shell.execute_reply":"2021-07-07T08:10:04.724669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since efficient net requires image shape to be at least 32X32 I am first using a Conv layer to increase channels to 6 then using reshape to increase image size if someone has a better way then please suggest","metadata":{}},{"cell_type":"code","source":"model = tf.keras.Sequential([L.InputLayer(input_shape=(69,193,1)),L.Conv2D(3,3,activation='relu',padding='same'),efn.EfficientNetB0(include_top=False,input_shape=(),weights='imagenet'),\n        L.GlobalAveragePooling2D(),\n        L.Dense(32,activation='relu'),\n        L.Dense(1, activation='sigmoid')])\n\nmodel.summary()\nmodel.compile(optimizer=keras.optimizers.Adam(learning_rate=0.001),\n              loss='binary_crossentropy', metrics=[keras.metrics.AUC()])\n","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:10:04.727383Z","iopub.execute_input":"2021-07-07T08:10:04.72769Z","iopub.status.idle":"2021-07-07T08:10:08.265251Z","shell.execute_reply.started":"2021-07-07T08:10:04.727659Z","shell.execute_reply":"2021-07-07T08:10:08.264024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.fit(train_dataset,epochs=2,validation_data=valid_dataset)","metadata":{"execution":{"iopub.status.busy":"2021-07-07T08:10:08.26851Z","iopub.execute_input":"2021-07-07T08:10:08.268827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds = model.predict(test_dataset)\npreds = preds.reshape(-1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame({'id':sample_submission['id'],'target':preds})","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv',index=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2><center>If you found this notebook useful please upvote</center></h2>","metadata":{}},{"cell_type":"markdown","source":"<h2><center>Work in Progress ... ⏳</center></h2>","metadata":{}}]}