{"cells":[{"metadata":{"_uuid":"e154a47bf09b8770980486e87786317a1b3038e1"},"cell_type":"markdown","source":"<img src = \"https://storage.googleapis.com/kaggle-media/competitions/Enet%20Centre/logo_vyska.png\" width = 300 hieght = 300/>\n\n<h6><center><font size=\"6\">End to End Machine Learning for VSB Power Line Fault Detection</center></font></h6>\n<br>\n<b>Faults in electric transmission lines can lead to a destructive phenomenon called partial discharge. If left alone, partial discharges can damage equipment to the point that it stops functioning entirely. Your challenge is to detect partial discharges so that repairs can be made before any lasting harm occurs. Each signal contains 800,000 measurements of a power line's voltage, taken over 20 milliseconds. As the underlying electric grid operates at 50 Hz, this means each signal covers a single complete grid cycle. The grid itself operates on a 3-phase power scheme, and all three phases are measured simultaneously.\n<br><br>\nThe Challenge is to detect partial discharge patterns in signals acquired from these power lines with a new meter designed at the ENET Centre at VSB. Effective classifiers using this data will make it possible to continuously monitor power lines for faults.</b>\n\n<pre><b>Inspired from Bruno Aquino's Kernel:\nhttps://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694</b></pre>\n\n<pre><b>\n<a id = \"#\">Content</a>\n- <a href = \"#1\"> Import Library</a>\n- <a href = \"#2\"> Define Matthews Correlation</a>\n- <a href = \"#3\"> Create Attention Layer</a>\n- <a href = \"#4\"> Import Data</a>\n- <a href = \"#5\"> Define Transformations</a>\n- <a href = \"#6\"> Prepare Training Data</a>\n- <a href = \"#7\"> Data Exploration</a>\n    - <a href = \"#71\">Distribution of mean values per rows</a>\n    - <a href = \"#72\">Distribution of mean values per columns</a>\n    - <a href = \"#73\">Distribution of S.D values per rows</a>\n    - <a href = \"#74\">Distribution of S.D values per columns</a>\n    - <a href = \"#75\">Distribution of Skewness values per rows</a>\n    - <a href = \"#76\">Distribution of Skewness values per columns</a>\n    - <a href = \"#77\">Distribution of Kurtosis values per rows</a>\n    - <a href = \"#78\">Distribution of Kurtosis values per columns</a>\n- <a href = \"#8\">Applying TSNE Transformation</a>\n- <a href = \"#9\">Applying PCA Transformation></a>\n- <a href = \"#19\">Applying ISO Transformation</a>\n- <a href = \"#10\">Applying LLE Transformation></a>\n- <a href = \"#11\">Building the RNN Model Architecture</a>\n- <a href =\"#12\">Applying Stratified K Fold Technique</a>\n- <a href =\"#13\">Identify the Best Threshold</a>\n- <a href = \"#14\">Create Submission File</a>\n</b></pre>","execution_count":null},{"metadata":{"_uuid":"9882a9623b6a1b939d465a51ea381e56d9dc4f9c"},"cell_type":"markdown","source":"<a id=1><pre><b>Import Library</b></pre></a>","execution_count":null},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport pyarrow.parquet as pq # Used to read the data\nimport os \nimport numpy as np\nfrom keras.layers import * # Keras is the most friendly Neural Network library, this Kernel use a lot of layers classes\nfrom keras.models import Model\nfrom tqdm import tqdm # Processing time measurement\nfrom sklearn.model_selection import train_test_split \nfrom keras import backend as K # The backend give us access to tensorflow operations and allow us to create the Attention class\nfrom keras import optimizers # Allow us to access the Adam class to modify some parameters\nfrom sklearn.model_selection import GridSearchCV, StratifiedKFold,RepeatedStratifiedKFold # Used to use Kfold to train our model\nfrom keras.callbacks import * # This object helps the model to train in a smarter way, avoiding overfitting\nimport tensorflow as tf","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6e6379386e44afc69bee8895a52da22199e888fb","trusted":true},"cell_type":"code","source":"# select how many folds will be created\nN_SPLITS = 5\n# it is just a constant with the measurements data size\nsample_size = 800000","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"38d0c7c3b97be5cd3e5b72135f5a7de6933de953"},"cell_type":"markdown","source":"<a id=2><pre><b>Define Matthews Correlation</b></pre></a>","execution_count":null},{"metadata":{"_uuid":"c3340ee96becb5ca8f075d9c44b7df383ddba5ee","trusted":true},"cell_type":"code","source":"# It is the official metric used in this competition\n# below is the declaration of a function used inside the keras model, calculation with K (keras backend / thensorflow)\ndef matthews_correlation(y_true, y_pred):\n    '''Calculates the Matthews correlation coefficient measure for quality\n    of binary classification problems.\n    '''\n    y_pred = tf.convert_to_tensor(y_pred, np.float32)\n    y_true = tf.convert_to_tensor(y_true, np.float32)\n    \n    y_pred_pos = K.round(K.clip(y_pred, 0, 1))\n    y_pred_neg = 1 - y_pred_pos\n\n    y_pos = K.round(K.clip(y_true, 0, 1))\n    y_neg = 1 - y_pos\n\n    tp = K.sum(y_pos * y_pred_pos)\n    tn = K.sum(y_neg * y_pred_neg)\n\n    fp = K.sum(y_neg * y_pred_pos)\n    fn = K.sum(y_pos * y_pred_neg)\n\n    numerator = (tp * tn - fp * fn)\n    denominator = K.sqrt((tp + fp) * (tp + fn) * (tn + fp) * (tn + fn))\n\n    return numerator / (denominator + K.epsilon())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"240ff24f0e5c2d976131af097f4becb883a11f95"},"cell_type":"markdown","source":"<pre><a id = 3><b>Create Attention Layer</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"eda7ea366117d1ce8e5fce69e5bba333821d8b48","trusted":true},"cell_type":"code","source":"# https://www.kaggle.com/suicaokhoailang/lstm-attention-baseline-0-652-lb\n\nclass Attention(Layer):\n    def __init__(self, step_dim,\n                 W_regularizer=None, b_regularizer=None,\n                 W_constraint=None, b_constraint=None,\n                 bias=True, **kwargs):\n        self.supports_masking = True\n        self.init = initializers.get('glorot_uniform')\n\n        self.W_regularizer = regularizers.get(W_regularizer)\n        self.b_regularizer = regularizers.get(b_regularizer)\n\n        self.W_constraint = constraints.get(W_constraint)\n        self.b_constraint = constraints.get(b_constraint)\n\n        self.bias = bias\n        self.step_dim = step_dim\n        self.features_dim = 0\n        super(Attention, self).__init__(**kwargs)\n\n    def build(self, input_shape):\n        assert len(input_shape) == 3\n\n        self.W = self.add_weight((input_shape[-1],),\n                                 initializer=self.init,\n                                 name='{}_W'.format(self.name),\n                                 regularizer=self.W_regularizer,\n                                 constraint=self.W_constraint)\n        self.features_dim = input_shape[-1]\n\n        if self.bias:\n            self.b = self.add_weight((input_shape[1],),\n                                     initializer='zero',\n                                     name='{}_b'.format(self.name),\n                                     regularizer=self.b_regularizer,\n                                     constraint=self.b_constraint)\n        else:\n            self.b = None\n\n        self.built = True\n\n    def compute_mask(self, input, input_mask=None):\n        return None\n\n    def call(self, x, mask=None):\n        features_dim = self.features_dim\n        step_dim = self.step_dim\n\n        eij = K.reshape(K.dot(K.reshape(x, (-1, features_dim)),\n                        K.reshape(self.W, (features_dim, 1))), (-1, step_dim))\n\n        if self.bias:\n            eij += self.b\n\n        eij = K.tanh(eij)\n\n        a = K.exp(eij)\n\n        if mask is not None:\n            a *= K.cast(mask, K.floatx())\n\n        a /= K.cast(K.sum(a, axis=1, keepdims=True) + K.epsilon(), K.floatx())\n\n        a = K.expand_dims(a)\n        weighted_input = x * a\n        return K.sum(weighted_input, axis=1)\n\n    def compute_output_shape(self, input_shape):\n        return input_shape[0],  self.features_dim","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0c94824ab8a2f0e7ee38811a6effecd7e2b6477d"},"cell_type":"markdown","source":"<pre><a id = 4><b>Import Data</b></a></pre>","execution_count":null},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"# just load train data\ndf_train = pd.read_csv('../input/metadata_train.csv')\n# set index, it makes the data access much faster\ndf_train = df_train.set_index(['id_measurement', 'phase'])\ndf_train.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"26df6c7fbfecd537404866faec13d1238ae3ebc6","trusted":true},"cell_type":"code","source":"# in other notebook I have extracted the min and max values from the train data, the measurements\nmax_num = 127\nmin_num = -128","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"87d74c8547a580e74dd945600b361a96ce8aa0d8"},"cell_type":"markdown","source":"<pre><a id = 5><b>Define Transformations</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"7b0717b14bcfcba1f48d33c8161ae51c778687af","trusted":true},"cell_type":"code","source":"# This function standardize the data from (-128 to 127) to (-1 to 1)\n# Theoretically it helps in the NN Model training, but I didn't tested without it\ndef min_max_transformation(ts, min_data, max_data, range_needed=(-1,1)):\n    if min_data < 0:\n        ts_std = (ts + abs(min_data)) / (max_data + abs(min_data))\n    else:\n        ts_std = (ts - min_data) / (max_data - min_data)\n    if range_needed[0] < 0:    \n        return ts_std * (range_needed[1] + abs(range_needed[0])) + range_needed[0]\n    else:\n        return ts_std * (range_needed[1] - range_needed[0]) + range_needed[0]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c6137bbbe75c3a1509a5f98e08805dbbd492aa37","trusted":true},"cell_type":"code","source":"def transform_ts(ts, n_dim=160, min_max=(-1,1)):\n    # convert data into -1 to 1\n    ts_std = min_max_transformation(ts, min_data=min_num, max_data=max_num)\n    # bucket or chunk size, 5000 in this case (800000 / 160)\n    bucket_size = int(sample_size / n_dim)\n    # new_ts will be the container of the new data\n    new_ts = []\n    # this for iteract any chunk/bucket until reach the whole sample_size (800000)\n    for i in range(0, sample_size, bucket_size):\n        # cut each bucket to ts_range\n        ts_range = ts_std[i:i + bucket_size]\n        # calculate each feature\n        mean = ts_range.mean()\n        std = ts_range.std() # standard deviation\n        std_top = mean + std # I have to test it more, but is is like a band\n        std_bot = mean - std\n        # I think that the percentiles are very important, it is like a distribuiton analysis from eath chunk\n        percentil_calc = np.percentile(ts_range, [0, 1, 25, 50, 75, 99, 100]) \n        max_range = percentil_calc[-1] - percentil_calc[0] # this is the amplitude of the chunk\n        relative_percentile = percentil_calc - mean # maybe it could heap to understand the asymmetry\n        new_ts.append(np.concatenate([np.asarray([mean, std, std_top, std_bot, max_range]),percentil_calc, relative_percentile]))\n    return np.asarray(new_ts)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"48f1bb0248e8ec5a8e20e48061f71692491d2c5e"},"cell_type":"markdown","source":"<pre><a id = 6><b>Prepare Training Data</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"7460e718a605803f1d9e4fbec61750a0deb02a47","trusted":true},"cell_type":"code","source":"# this function take a piece of data and convert using transform_ts(), but it does to each of the 3 phases\n# if we would try to do in one time, could exceed the RAM Memmory\ndef prepare_data(start, end):\n\n    praq_train = pq.read_pandas('../input/train.parquet', columns=[str(i) for i in range(start, end)]).to_pandas()\n    X = []\n    y = []\n   \n    for id_measurement in tqdm(df_train.index.levels[0].unique()[int(start/3):int(end/3)]):\n        X_signal = []\n        # for each phase of the signal\n        for phase in [0,1,2]:\n            # extract from df_train both signal_id and target to compose the new data sets\n            signal_id, target = df_train.loc[id_measurement].loc[phase]\n            # but just append the target one time, to not triplicate it\n            if phase == 0:\n                y.append(target)\n            # extract and transform data into sets of features\n            X_signal.append(transform_ts(praq_train[str(signal_id)]))\n        # concatenate all the 3 phases in one matrix\n        X_signal = np.concatenate(X_signal, axis=1)\n        # add the data to X\n        X.append(X_signal)\n    X = np.asarray(X)\n    y = np.asarray(y)\n    return X, y","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"52dc826ab9ee1dd56c9fb29bd5c1b2d26b5928bf","trusted":true},"cell_type":"code","source":"# this code is very simple, divide the total size of the df_train into two sets and process it\nX = []\ny = []\ndef load_all():\n    total_size = len(df_train)\n    for ini, end in [(0, int(total_size/2)), (int(total_size/2), total_size)]:\n        X_temp, y_temp = prepare_data(ini, end)\n        X.append(X_temp)\n        y.append(y_temp)\nload_all()\nX = np.concatenate(X)\ny = np.concatenate(y)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"51ad0e25b00536de6170168499923d82ae1d735f","trusted":true},"cell_type":"code","source":"print(X.shape, y.shape)\n# save data into file, a numpy specific format\nnp.save(\"X.npy\",X)\nnp.save(\"y.npy\",y)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"29a9c9c03a2d3f91e5cc3bc2c10dd787e4dc3b59"},"cell_type":"markdown","source":"<pre><a id = 7><b>Data Exploration</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"c0e45b630dff00cde5fa737bd611a4ae65bf2b2f"},"cell_type":"markdown","source":"<pre><b> I used Panel to convert the X and Y into Data Frames</b></pre>","execution_count":null},{"metadata":{"_uuid":"609e87b0474f5ce27cb610eddc6845c30c500a3a","trusted":true},"cell_type":"code","source":"warnings.filterwarnings('ignore')\n\npan = pd.Panel(X)\ndf = pan.swapaxes(1, 2).to_frame()\ndf.index = df.index.droplevel('major')\ndf.index = df.index+1","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"622cc5e7213351c5108f234f02456c633d973d60","trusted":true},"cell_type":"code","source":"df.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7fd8406261c01dd8036110228bfdeced82a8f444","trusted":true},"cell_type":"code","source":"df_y = pd.DataFrame(y.reshape(1,-1))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5fbe52acd76558b792df7398f954ae22a1d6f5c9","trusted":true},"cell_type":"code","source":"df_y.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c963d60ecea97c0c313e39564e5b7edec306f528"},"cell_type":"markdown","source":"<pre><b>Lets check the Explorations for Mean, Standard Deviation, Skewness and Kurtosis on the Data generated</b></pre>\n","execution_count":null},{"metadata":{"_uuid":"57094aa26ea59c6127306b25f6de72e908f908cc"},"cell_type":"markdown","source":"<pre><a id = 71><b>Distribution of mean values per row in the training set</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"af7c80ed71411c83a966ea3c08971e313f3cdb75","trusted":true},"cell_type":"code","source":"import matplotlib\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\nplt.style.use('seaborn')\nfrom scipy.stats import norm, skew\n\nplt.figure(figsize=(16,6))\nplt.title(\"Distribution of mean values per row in the training set\")\nsns.distplot(df.mean(axis=0),color=\"red\", kde=True,bins=120, label='train')\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0b00108257564d9b6d57bf46c978220ce4a9458a"},"cell_type":"markdown","source":"<pre><a id = 72><b>Distribution of mean values per column in the training set</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"2016ee1a75b80fc6f593383f0e21067f1a306030","trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16,6))\nplt.title(\"Distribution of mean values per column in the train and test set\")\nsns.distplot(df.mean(axis=1),color=\"red\", kde=True,bins=120, label='train')\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a198c93e39a8a7147a4163d99d702ae662c1d513"},"cell_type":"markdown","source":"<pre><a id = 73><b>Distribution of S.D values per row in the training set</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"4958f96df5b85552f8ff4aa88558a706af9aac90","trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16,6))\nplt.title(\"Distribution of Standard Deviation values per row in the training set\")\nsns.distplot(df.std(axis=0),color=\"blue\", kde=True,bins=120, label='train')\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b58b6285d65d73d6f96dda5825a324458eb184e4"},"cell_type":"markdown","source":"<pre><a id = 74><b>Distribution of S.D values per column in the training set</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"c6bcfcf055384013002c23a643b5dcb3f7553563","trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16,6))\nplt.title(\"Distribution of Standard Deviation values per column in the training set\")\nsns.distplot(df.std(axis=1),color=\"blue\", kde=True,bins=120, label='train')\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"01fe57b117c2b31a1d03c9f5a962bf6e7ed9a185"},"cell_type":"markdown","source":"<pre><a id = 75><b>Distribution of Skewness values per rows in the training set</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"eaa3a5f32349ab59881275625e97ceac27aa3404","scrolled":false,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16,6))\nplt.title(\"Skewness Distribution per row in the training set\")\nsns.distplot(df.skew(axis=0),color=\"red\", kde=True,bins=120, label='train')\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3332aa4d7c13a0f3ad20dec1b03ea037d99120e0"},"cell_type":"markdown","source":"<pre><a id = 76><b>Distribution of Skewness values per column in the training set</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"4ad17336c6dcd97ef758c856b11b5a0f3f298aef","trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16,6))\nplt.title(\"Skewness Distribution per column in the training set\")\nsns.distplot(df.skew(axis=1),color=\"red\", kde=True,bins=120, label='train')\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ec36b1b447d6151b19477d24c300c7a051d44c13"},"cell_type":"markdown","source":"<pre><b>What we understand now, is the reason behind the BIModal behavior. Since the power line fault is detected in a very few cases, there is a presence of class imbalance in the dataset, which is valid for this kind of data.</pre></b>","execution_count":null},{"metadata":{"_uuid":"d53b150cfbdf4ef5a9a1b19753852ddc74ee89d1"},"cell_type":"markdown","source":"<pre><a id = 77><b>Distribution of Kurtosis values per rows in the training set</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"ad3fc481219acf4a90bfebe38f875097a0ebb385","trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16,6))\nplt.title(\"Kurtosis Distribution per row in the training set\")\nsns.distplot(df.kurtosis(axis=0),color=\"red\", kde=True,bins=120, label='train')\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cfa744283f2eb7a951f2b0a080b080b61e4f881d"},"cell_type":"markdown","source":"<pre><b>The distribution is Platokurtic.</b></pre>","execution_count":null},{"metadata":{"_uuid":"e4583426314b10fb67a3136713f651127acdc962"},"cell_type":"markdown","source":"<pre><a id = 78><b>Distribution of Kurtosis values per columns in the training set</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"79a2a5041118388e7c520581311e688d70e2b021","trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16,6))\nplt.title(\"Kurtosis Distribution per row in the training set\")\nsns.distplot(df.kurtosis(axis=1),color=\"red\", kde=True,bins=120, label='train')\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a3f228594dcf4f764927c5535f2d0ac880ecf264"},"cell_type":"markdown","source":"<pre><b>Now Lets perform TSNE to see if the distribution is well seperated?</b></pre>","execution_count":null},{"metadata":{"_uuid":"f435440a31a125077364422a820b3e5b21f04cdb","scrolled":true,"trusted":true},"cell_type":"code","source":"df.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6577dc56c7e9d3ea74ce7c9c0abf253a80bea00e","trusted":true},"cell_type":"code","source":"df_y.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c1d23b64100b4a63bc0a33e8194f8dec381c291f"},"cell_type":"markdown","source":"<pre><b>Please note that since y's shape is different than X, i cant really use Y's for the color bar.  I am using a customized function for the color bar.</b></pre>","execution_count":null},{"metadata":{"_uuid":"f499ca2af17f037765df1128b015523264a50290"},"cell_type":"markdown","source":"<pre><a id = 8><b>Applying TSNE Transformation</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"a2164a45e3d69b06c7211c9db1992a79b8ff7279","trusted":true},"cell_type":"code","source":"#Now let's use t-SNE to reduce dimensionality down to 2D so we can plot the dataset:\nfrom sklearn.manifold import TSNE\n\ntsne = TSNE(n_components=2, random_state=42, verbose = 2)\nTSNE_X = tsne.fit_transform(df)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"476c6d75046232ae66360901fb0ffd1e027c980e","trusted":true},"cell_type":"code","source":"import matplotlib as mpl\nmin, max = (-40, 30)\nstep = 10\n\n# Setting up a colormap that's a simple transtion\nmymap = mpl.colors.LinearSegmentedColormap.from_list('mycolors',['green','yellow','red'])\n\n# Using contourf to provide my colorbar info, then clearing the figure\nZ = [[0,0],[0,0]]\nlevels = range(min,max+step,step)\nCS3 = plt.contourf(Z, levels, cmap=mymap)\nplt.clf()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"78fefd9f89cac8942ecc558aa5f4d72df9ed9d55","trusted":true},"cell_type":"code","source":"plt.figure(figsize=(13,10))\ncm = plt.cm.get_cmap('RdYlGn_r')\ncolors=[cm(1*i) for i in [75,145,255]]\nxy = range(3)\ncolorlist=[colors[x] for x in xy]\nplt.scatter(TSNE_X[:, 0], TSNE_X[:, 1], c=colorlist, cmap=cm)\nplt.colorbar(CS3)\nplt.axis('off')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0c095ae18c3135dcf598fb7218fdfb6031e5349b"},"cell_type":"markdown","source":"<pre><a id = 19><b>Applying ISO Transformation</b></a></pre>","execution_count":null},{"metadata":{"trusted":true,"_uuid":"3b496c77034a3ef167fe57d2d04ed5a2333e3b4d"},"cell_type":"code","source":"from sklearn.manifold import Isomap\n\nisomap = Isomap(n_components=200)\ndf_reduced = isomap.fit_transform(df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4c58be4694f802c718df5773c5b0e2e8ad3605a2"},"cell_type":"code","source":"plt.scatter(df_reduced[:, 0], df_reduced[:, 1], c=colorlist, cmap=plt.cm.hot)\nplt.axis('off')\nplt.colorbar(CS3)\nplt.grid(True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"70ff90ffc9f61b20a4060858ae915c026ea504b0"},"cell_type":"markdown","source":"<pre><b>Now Lets perform PCA to see if the distribution is well seperated?</pre></b>","execution_count":null},{"metadata":{"_uuid":"f815c9c21d02ea4e13ca25bb8166c45c87973442"},"cell_type":"markdown","source":"<pre><a id = 9><b>Applying PCA Transformation</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"9c31a1ca7a0a6e9c967b332e2e2c952f3f4b39aa","trusted":true},"cell_type":"code","source":"from sklearn.decomposition import PCA\n\nPCA_train_x = PCA(n_components=2, random_state=42).fit_transform(df)\nplt.scatter(PCA_train_x[:, 0], PCA_train_x[:, 1], c=colorlist, cmap=cm)\nplt.axis('off')\nplt.colorbar(CS3)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5d61a79961ed69f5ec02f4f40adfeafa129982dd","trusted":true},"cell_type":"code","source":"from sklearn.decomposition import KernelPCA\n\nlin_pca = KernelPCA(n_components = 2, kernel=\"linear\", fit_inverse_transform=True)\nrbf_pca = KernelPCA(n_components = 2, kernel=\"rbf\", gamma=0.0433, fit_inverse_transform=True)\nsig_pca = KernelPCA(n_components = 2, kernel=\"sigmoid\", gamma=0.001, coef0=1, fit_inverse_transform=True)\n\nplt.figure(figsize=(11, 4))\nfor subplot, pca, title in ((131, lin_pca, \"Linear kernel\"), (132, rbf_pca, \"RBF kernel, $\\gamma=0.04$\"), (133, sig_pca, \"Sigmoid kernel, $\\gamma=10^{-3}, r=1$\")):\n    X_reduced = pca.fit_transform(df)\n    if subplot == 132:\n        X_reduced_rbf = X_reduced\n    \n    plt.subplot(subplot)\n    plt.title(title, fontsize=14)\n    plt.scatter(X_reduced[:, 0], X_reduced[:, 1], c=colorlist, cmap=plt.cm.hot)\n    plt.xlabel(\"$z_1$\", fontsize=18)\n    if subplot == 131:\n        plt.ylabel(\"$z_2$\", fontsize=18, rotation=0)\n    plt.grid(True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"33ef8994a342182a4c472ae03971559ecc1808eb"},"cell_type":"markdown","source":"<pre><b>After seeing the PCA model, I am eager to use Local Linear Embedding, as I do see the curve is resembling similar to a swiss roll.</b></pre>","execution_count":null},{"metadata":{"_uuid":"23f6855927df81c9d52ead8b99eb949f3cd76e1f"},"cell_type":"markdown","source":"<pre><a id = 10><b>Applying LLE Transformation</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"5133adf8a8a3fa6fb8d575ab92e44aa03f1b9126","trusted":false},"cell_type":"code","source":"from sklearn.manifold import LocallyLinearEmbedding\n\nlle = LocallyLinearEmbedding(n_components=2, n_neighbors=10, random_state=42)\nlle_X = lle.fit_transform(df)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"03823498ba0a3af149632bcb907ac0cc300be62c","trusted":false},"cell_type":"code","source":"plt.title(\"Unrolling PCA graph using LLE\", fontsize=14)\nplt.scatter(lle_X[:, 0], lle_X[:, 1], c= colorlist, cmap=plt.cm.hot)\nplt.xlabel(\"$z_1$\", fontsize=18)\nplt.ylabel(\"$z_2$\", fontsize=18)\nplt.colorbar(CS3)\nplt.grid(True)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"20c3ca659de29ac0ed899bcacd65231109e24667"},"cell_type":"markdown","source":"<pre><a id = 11><b>Building the RNN Model Architecture</b></a></pre>\n","execution_count":null},{"metadata":{"_uuid":"289bc7d1ab8048a60025801b457f8df1d848acbc","trusted":false},"cell_type":"code","source":"# This is NN LSTM Model creation\ndef model_lstm(input_shape):\n    # The shape was explained above, must have this order\n    inp = Input(shape=(input_shape[1], input_shape[2],))\n    \n    x = Bidirectional(CuDNNLSTM(128, return_sequences=True))(inp)\n    x = Bidirectional(CuDNNLSTM(64, return_sequences=True))(x)\n    \n    x = Attention(input_shape[1])(x)\n    \n    x = Dense(64, activation=\"relu\")(x)\n    x = Dense(1, activation=\"sigmoid\")(x)\n    model = Model(inputs=inp, outputs=x)\n    model.compile(loss='binary_crossentropy', optimizer='adam', metrics=[matthews_correlation])\n    \n    return model","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0b0c205f394c57d65b3d5704d06603508dd87c7e"},"cell_type":"markdown","source":"<pre><a id = 12><b>Applying Stratified K Fold Technique</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"8d6f4ca319c383b1b4f671a37c5a324136e7a466","trusted":false},"cell_type":"code","source":"# First, create a set of indexes of the 5 folds\nsplits = list(StratifiedKFold(n_splits=N_SPLITS,random_state=42).split(X, y))\npreds_val = []\ny_val = []\n# Then, iteract with each fold\n# If you dont know, enumerate(['a', 'b', 'c']) returns [(0, 'a'), (1, 'b'), (2, 'c')]\nfor idx, (train_idx, val_idx) in enumerate(splits):\n    K.clear_session() # I dont know what it do, but I imagine that it \"clear session\" :)\n    print(\"Beginning fold {}\".format(idx+1))\n    train_X, train_y, val_X, val_y = X[train_idx], y[train_idx], X[val_idx], y[val_idx]\n    model = model_lstm(train_X.shape)\n    #es = EarlyStopping(monitor='val_matthews_correlation', verbose=2, patience=50, mode='max')\n    ckpt = ModelCheckpoint('weights_{}.h5'.format(idx), save_best_only=True, save_weights_only=True, verbose=1, monitor='val_matthews_correlation', mode='max')\n    # Train, train, train\n    model.fit(train_X, train_y, batch_size=128, epochs=50, validation_data=[val_X, val_y], callbacks=[ckpt])\n    # loads the best weights saved by the checkpoint\n    model.load_weights('weights_{}.h5'.format(idx))\n    # Add the predictions of the validation to the list preds_val\n    preds_val.append(model.predict(val_X, batch_size=512))\n    # and the val true y\n    y_val.append(val_y)\n\n# concatenates all and prints the shape    \npreds_val = np.concatenate(preds_val)[...,0]\ny_val = np.concatenate(y_val)\npreds_val.shape, y_val.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d28151fd0be9fd9762f3f55e307d82f89bfbd291","trusted":false},"cell_type":"code","source":"import tensorflow as tf\n\ndef threshold_search(y_true, y_proba):\n    best_threshold = 0\n    best_score = 0\n    for threshold in tqdm([i * 0.01 for i in range(100)]):\n        score = K.eval(matthews_correlation(y_true.astype(np.float64), (y_proba > threshold).astype(np.float64)))\n        if score > best_score:\n            best_threshold = threshold\n            best_score = score\n    search_result = {'threshold': best_threshold, 'matthews_correlation': best_score}\n    return search_result","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"92cb38e80553c4c4bc43bb1e39c6bd993b65b275"},"cell_type":"markdown","source":"<pre><a id = 13><b>Identify the Best Threshold</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"c252624065ee8551ae88a2c10b343e4fee3ecc80","trusted":false},"cell_type":"code","source":"best_threshold = threshold_search(y_val, preds_val)['threshold']","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8fc5b3dceb6f9a2b3882af4342db975c4f487062","trusted":false},"cell_type":"code","source":"best_threshold","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ae9bd3fa9d8c0781c0708846bb7f2a9f9e6cbd3c","trusted":false},"cell_type":"code","source":"meta_test = pd.read_csv('../input/metadata_test.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3eb186d032f79c99ffba05dd1a7fabb77e13cec5","trusted":false},"cell_type":"code","source":"meta_test = meta_test.set_index(['signal_id'])\nmeta_test.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6f8e94387f625bff0a9a6289e1ee038908bc5856","trusted":false},"cell_type":"code","source":"first_sig = meta_test.index[0]\nn_parts = 10\nmax_line = len(meta_test)\npart_size = int(max_line / n_parts)\nlast_part = max_line % n_parts\nprint(first_sig, n_parts, max_line, part_size, last_part, n_parts * part_size + last_part)\n# Here we create a list of lists with start index and end index for each of the 10 parts and one for the last partial part\nstart_end = [[x, x+part_size] for x in range(first_sig, max_line + first_sig, part_size)]\nstart_end = start_end[:-1] + [[start_end[-1][0], start_end[-1][0] + last_part]]\nprint(start_end)\nX_test = []\n# now, very like we did above with the train data, we convert the test data part by part\n# transforming the 3 phases 800000 measurement in matrix (160,57)\nfor start, end in start_end:\n    subset_test = pq.read_pandas('../input/test.parquet', columns=[str(i) for i in range(start, end)]).to_pandas()\n    for i in tqdm(subset_test.columns):\n        id_measurement, phase = meta_test.loc[int(i)]\n        subset_test_col = subset_test[i]\n        subset_trans = transform_ts(subset_test_col)\n        X_test.append([i, id_measurement, phase, subset_trans])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"af9aa6b2b8f8a2beda1a02ff998e3072fcad8d06","trusted":false},"cell_type":"code","source":"X_test_input = np.asarray([np.concatenate([X_test[i][3],X_test[i+1][3], X_test[i+2][3]], axis=1) for i in range(0,len(X_test), 3)])\nnp.save(\"X_test.npy\",X_test_input)\nX_test_input.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"222d3d99e51f4ef6629a6c47864440721c9b4be2"},"cell_type":"markdown","source":"<pre><a id = 14><b>Create Submission File</b></a></pre>","execution_count":null},{"metadata":{"_uuid":"cfd265d3e07c4cc1679d2c4d55fe7de631c813e7","trusted":false},"cell_type":"code","source":"submission = pd.read_csv('../input/sample_submission.csv')\nprint(len(submission))\nsubmission.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2f7342296138f6bfd3e9cedd029e1035de3b98fc","trusted":false},"cell_type":"code","source":"preds_test = []\nfor i in range(N_SPLITS):\n    model.load_weights('weights_{}.h5'.format(i))\n    pred = model.predict(X_test_input, batch_size=300, verbose=1)\n    pred_3 = []\n    for pred_scalar in pred:\n        for i in range(3):\n            pred_3.append(pred_scalar)\n    preds_test.append(pred_3)\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4b212616b85a0c704ad8dd11d2ea88c818a4d2a2","trusted":false},"cell_type":"code","source":"preds_test = (np.squeeze(np.mean(preds_test, axis=0)) > best_threshold).astype(np.int)\npreds_test.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"347a76deca88ff3834c2b372391ada6b30a6ae08","trusted":false},"cell_type":"code","source":"submission['target'] = preds_test\nsubmission.to_csv('submission.csv', index=False)\nsubmission.head()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}