{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"text-align:center; font-size:300%; background-color:#000000; color: #ffffff; padding:5%;\">Gravitational Waves</h1>","metadata":{"execution":{"iopub.status.busy":"2021-08-06T15:50:55.199298Z","iopub.execute_input":"2021-08-06T15:50:55.199707Z","iopub.status.idle":"2021-08-06T15:50:55.205288Z","shell.execute_reply.started":"2021-08-06T15:50:55.19967Z","shell.execute_reply":"2021-08-06T15:50:55.204188Z"}}},{"cell_type":"markdown","source":"<h3 style=\"text-align:center; font-style:italic;\">\"Bundle of woods makes them strong\", unlike black holes who don't agree with the statement!</h3>","metadata":{}},{"cell_type":"markdown","source":"<img style=\"width:70%; margin:0 auto; border-radius:15px;\" src=\"https://i.ibb.co/4gLcXM7/d41586-019-00573-4-16480554.jpg\">","metadata":{}},{"cell_type":"markdown","source":"<p style=\"text-align: center; font-size: 100%;\">Gravitational waves are basically the ripples or disturbances in the curvature of\nspacetime, generated by the violent masses in the universe, that propogate as waves from their source with the\nspeed of light. They were first discovered in 2015, due to collision of two black holes as shown in the figure\nabove. These waves are basically the signals generated by collision and can be detected\nthrough different detectors like LIGO Hanford, LIGO Livingston, and Virgo used in the dataset of this competetion.\nEach data example basically contains the three waves detected by the three detectors.</p>","metadata":{}},{"cell_type":"markdown","source":"<div>\n    <h2 style=\"font-size:140%;\">Contributers (Kaggle Profile)</h2>\n    <a style=\"padding:1.5% 2.4%; display:inline-block;background-color: #0a0037; color:white;font-size:85%; border-radius:10px;\" href=\"https://www.kaggle.com/mohammadasimbluemoon\">Mohammad Asim</a>\n    <a style=\"padding:1.5%; display:inline-block;background-color: #2d2c2c;; color:white;font-size:85%; border-radius:10px;\" href=\"https://www.kaggle.com/khhassan\">Muhammad Hassan</a>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div>\n    <h2 style=\"font-size:200%;\">Purpose and Description of Notebook</h2>\n    <p style=\"text-align:justify;\">The problem with the given dataset <a href=\"https://www.kaggle.com/c/g2net-gravitational-wave-detection\">G2Net Gravitational Wave Detection</a> is that it contains gravitational signals detected by the detectors\n    and does not convey enough information about the variation of the signals in the dataset.\n    So, to visualize the dataset, this should be converted into the image form. This notebook basically generates\n    the image data from the given gravitational waves (GW) detected by the detectors and performs the EDA\n    of the given training data.</p>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<hr>","metadata":{}},{"cell_type":"markdown","source":"**Installing the required package 'pycbc' used to convert the data to time series format as you will see after.**","metadata":{}},{"cell_type":"code","source":"!pip -q install pycbc","metadata":{"execution":{"iopub.status.busy":"2021-08-08T15:10:29.026891Z","iopub.execute_input":"2021-08-08T15:10:29.027381Z","iopub.status.idle":"2021-08-08T15:11:03.800077Z","shell.execute_reply.started":"2021-08-08T15:10:29.027236Z","shell.execute_reply":"2021-08-08T15:11:03.798795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Importing the required Libraries\n**These are some of the librarries needed throughout the notebook.**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport random\nimport pylab\nimport pycbc.types\nfrom PIL import Image\nfrom multiprocessing import Pool\nfrom multiprocessing import cpu_count\nimport os\nimport glob\nfrom tqdm.notebook import tqdm","metadata":{"execution":{"iopub.status.busy":"2021-08-08T15:11:41.602419Z","iopub.execute_input":"2021-08-08T15:11:41.602826Z","iopub.status.idle":"2021-08-08T15:11:41.609554Z","shell.execute_reply.started":"2021-08-08T15:11:41.602788Z","shell.execute_reply":"2021-08-08T15:11:41.608191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Reading the Training and Submission Data\n**Reading the training data to get an overview of it.**\n* The following block will give the data frame of the training data.","metadata":{}},{"cell_type":"code","source":"submission_df = pd.read_csv('../input/g2net-gravitational-wave-detection/sample_submission.csv')\nsubmission_df.head()\ntrain_df = pd.read_csv('../input/g2net-gravitational-wave-detection/training_labels.csv')\nprint(\"No. of Training Examples : \", len(train_df))\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-08T15:11:43.898595Z","iopub.execute_input":"2021-08-08T15:11:43.898986Z","iopub.status.idle":"2021-08-08T15:11:44.642604Z","shell.execute_reply.started":"2021-08-08T15:11:43.898954Z","shell.execute_reply":"2021-08-08T15:11:44.641076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* The following block will plot the histogram of the training data.","metadata":{}},{"cell_type":"code","source":"train_df.hist(bins=3)\nplt.title(\"Distribution of Training Data\")\nplt.xlabel('target')\nplt.ylabel('id')","metadata":{"execution":{"iopub.status.busy":"2021-08-08T15:11:47.130393Z","iopub.execute_input":"2021-08-08T15:11:47.130797Z","iopub.status.idle":"2021-08-08T15:11:47.453017Z","shell.execute_reply.started":"2021-08-08T15:11:47.130764Z","shell.execute_reply":"2021-08-08T15:11:47.451686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualizing Training Data\n**The following block will read the given \".npy\" file and visualize the training data.**\n* This block will first of all read the file path using \"convert_id_to_path\" function.\n* After that, it will give the path to the \"visualize_signals\" function to convert the data and plot the signals for easy visualization.\n* Also, it will give the shape of the running data example.","metadata":{}},{"cell_type":"code","source":"main_path = '../input/g2net-gravitational-wave-detection/'\ntrain_folder = 'train'\ntest_folder = 'test'\ndef convert_id_to_path(_id):\n    return f\"{main_path}/{train_folder}/{_id[0]}/{_id[1]}/{_id[2]}/{_id}.npy\"\ndef visualize_signals(_id, target, colors=(\"black\",\"red\",\"green\"), signal_names=(\"LIGO Hanford\",\"LIGO Livingston\",\"Virgo\")):\n    path = convert_id_to_path(_id)\n    data = np.load(path) #This function returns the input array (from a disk file with 'npy' extension).\n    x = data\n    print(data.shape)\n    plt.figure(figsize=(16,7))\n    for i in range(3):\n        plt.subplot(4,1,i+1)\n        plt.plot(data[i], color=colors[i])\n        plt.legend(signal_names[i], fontsize=12, loc=\"lower right\")\n        \n        plt.subplot(4,1,4)\n        plt.plot(data[i], color=colors[i])\n    plt.subplot(4,1,4)\n    plt.legend(signal_names, fontsize=12, loc=\"lower right\")\n    plt.suptitle(f\"id: {_id} target: {target}\", fontsize=16)\n    return x","metadata":{"execution":{"iopub.status.busy":"2021-08-08T15:11:49.745883Z","iopub.execute_input":"2021-08-08T15:11:49.746314Z","iopub.status.idle":"2021-08-08T15:11:49.75831Z","shell.execute_reply.started":"2021-08-08T15:11:49.746273Z","shell.execute_reply":"2021-08-08T15:11:49.757058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**This block will basically call the above function and plot the above signals based on the given index.**","metadata":{}},{"cell_type":"code","source":"for i in range(2):   #We are viusalizing the first two examples only. You can visualize...\n                     #...more by changing the value in the range function                    \n    _id = train_df.iloc[i][\"id\"]  #This line will get the '_id/index' from the training dataframe.\n    target = train_df.iloc[i][\"target\"] #This line will get the 'target/label' from the training dataframe.\n    data = visualize_signals(_id, target) \n","metadata":{"execution":{"iopub.status.busy":"2021-08-08T15:11:53.810328Z","iopub.execute_input":"2021-08-08T15:11:53.810703Z","iopub.status.idle":"2021-08-08T15:11:55.324353Z","shell.execute_reply.started":"2021-08-08T15:11:53.810671Z","shell.execute_reply":"2021-08-08T15:11:55.32301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Appyling the Q_Transform on the signals\n**To visualize the signals and variation in the data correctly, we have to apply the q_transform to the given signal. So, the following block will do the following:**\n* It will first convert the data into the time series instances.\n* After that, it will normalize/whitten the data.\n* Then, it will take the q_transform of the time series instances.\n* After that, it will normalize it to 0-255 scale.\n* At the end, it will save the data into the image form.\n","metadata":{}},{"cell_type":"code","source":"def get_constant_q_transform(file_names):\n    esp=1e-6\n    normalize=True\n    \n    for i in tqdm(range(len(file_names)),  desc='Progress'):\n        example_id = file_names[i].split('/')[-1].split('.')[0]\n        # load the specific 2s sample\n        data = np.load(file_names[i])\n        channels = []\n        for i in range(3):\n            # convert the data to a TimeSeries instance\n            ts = pycbc.types.TimeSeries(data[i, :], epoch=0, delta_t=1.0/2048) #this is the use of...\n                                                                               #...the installed library. \n            # whiten the data (i.e. normalize the noise power at different frequencies)\n            ts = ts.whiten(0.125, 0.125)\n            # calculate the qtransform\n            time, freq, power = ts.qtransform(.002, logfsteps=100, qrange=(10, 10), frange=(20, 512))\n            # normalize and scale to 0-255\n            if normalize:\n                mean = power.mean()\n                std = power.std()\n                power = (power - mean) / (std + esp)\n                _min, _max = power.min(), power.max()\n                power[power < _min] = _min\n                power[power > _max] = _max\n                power = 255 * (power - _min) / (_max - _min)\n                power = power.astype(np.uint8)\n            channels.append(power)\n        vstacked = np.stack(channels, axis = -1)\n        im = Image.fromarray(vstacked)\n        im.save(f\"train_cqt/{example_id}.png\")\n\n","metadata":{"execution":{"iopub.status.busy":"2021-08-08T15:11:58.547448Z","iopub.execute_input":"2021-08-08T15:11:58.547913Z","iopub.status.idle":"2021-08-08T15:11:58.560403Z","shell.execute_reply.started":"2021-08-08T15:11:58.547864Z","shell.execute_reply":"2021-08-08T15:11:58.559474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plotting Q_Transform\n**The following blocks will plot the q_transform and the time series instances that we have generated from the given training data.**\n* The follwing block will generate the function that will plot the q_transform.","metadata":{}},{"cell_type":"code","source":"import cv2\ndef q_transform_plot(ts, name):\n    esp=1e-6\n    ts = ts.whiten(0.125, 0.125)\n    time, freq, power = ts.qtransform(.002, logfsteps=100, qrange=(10, 10), frange=(20, 512))\n    mean = power.mean()\n    std = power.std()\n    power = (power - mean) / (std + esp)\n    _min, _max = power.min(), power.max()\n    power[power < _min] = _min\n    power[power > _max] = _max\n    power = 255 * (power - _min) / (_max - _min)\n    power = power.astype(np.uint8)\n    fig, ax = plt.subplots(1, 1, figsize=(15, 7))\n    ax.imshow(np.array(power))\n    plt.title('Q-Transformed ('+name+\")\")\n    return power","metadata":{"execution":{"iopub.status.busy":"2021-08-08T15:12:01.741001Z","iopub.execute_input":"2021-08-08T15:12:01.741398Z","iopub.status.idle":"2021-08-08T15:12:01.90654Z","shell.execute_reply.started":"2021-08-08T15:12:01.741359Z","shell.execute_reply":"2021-08-08T15:12:01.905307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following block will plot the three graphs:\n* Time series data before whitening of every detector in data.\n* Time series data after whitening of every detector in data.\n* Q_transform of every detector in data.","metadata":{}},{"cell_type":"code","source":"names = ['LIGO Hanford', 'LIGO Livingston', 'Virgo']\ncolor = ['r', 'g', 'b']\npath = convert_id_to_path(train_df.iloc[0]['id'])\ndata = np.load(path)\nfor i in range(data.shape[0]):\n    ts = pycbc.types.TimeSeries(data[i, :], epoch=0, delta_t=1.0/2048) \n    plt.figure(figsize=(16,3))\n    plt.subplot(2,2,1)\n    plt.plot(ts, color[i])\n    plt.legend(names[i], fontsize=12, loc=\"lower right\")\n    plt.title(\"Before Whitening (\"+names[i]+\")\")\n    plt.subplot(2,2,2)\n    plt.plot(ts.whiten(0.125, 0.125), color[i])\n    plt.legend(names[i], fontsize=12, loc=\"lower right\")\n    plt.title(\"After Applying Whitening (\"+names[i]+\")\")\n    q_transform_plot(ts, names[i])","metadata":{"execution":{"iopub.status.busy":"2021-08-08T15:12:04.549164Z","iopub.execute_input":"2021-08-08T15:12:04.549622Z","iopub.status.idle":"2021-08-08T15:12:09.770438Z","shell.execute_reply.started":"2021-08-08T15:12:04.549588Z","shell.execute_reply":"2021-08-08T15:12:09.769169Z"},"trusted":true},"execution_count":null,"outputs":[]}]}