{"cells":[{"metadata":{"_uuid":"7982834001d894783090561a22e00f211cb75732"},"cell_type":"markdown","source":"This kernel uses 1D convolutions on signals from power lines to identify partial faults"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport pyarrow.parquet as pq\nimport os\n\nos.listdir('../input')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7a0a816882adf498a88e6a85429d8ec124af4b10"},"cell_type":"markdown","source":"Read the parquet file. The full length of each signal is 800000. We will halve it to 400000 readings to create the pipeline."},{"metadata":{"trusted":true,"_uuid":"bd92583ca1db17938c34c8699836b3e17aad346f"},"cell_type":"code","source":"subset_train = pq.read_pandas('../input/train.parquet').to_pandas() #, columns=[str(i) for i in range(10)]).to_pandas()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a7164fff4b7c9f0bcf494f7f3fdf164e3b21bc6d"},"cell_type":"code","source":"subset_train = subset_train.iloc[:400000,:]\nsubset_train.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2973ddcfff4f06ac98ccde46d0b5a5157afe0683"},"cell_type":"markdown","source":"Now read the metadata file."},{"metadata":{"trusted":true,"_uuid":"6de250a9d4bf7faec4d39160610ef2fcde2c952d"},"cell_type":"code","source":"metadata_train = pd.read_csv('../input/metadata_train.csv')\nmetadata_train.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f2359163fe48f20610ed7c426efe94d7239252b7"},"cell_type":"code","source":"metadata_train.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6f5a84750f1ca51a4fca03e4d0ec03cfaa2bc2dd"},"cell_type":"markdown","source":"Import plotting libraries and create some basic plots."},{"metadata":{"trusted":true,"_uuid":"8aaad8abc0bc4e95ab5c592cd3605994b3c5d1f0"},"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\n\n#plt.hist(metadata_train['target'])\nsns.countplot(metadata_train['target'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d07f9ed35254ca6577def6c94de901129617e1f9"},"cell_type":"markdown","source":"This is a plot of the target values. As expected, a faulty power line is a kind of rare event. Let's visualize some negative and positive (faulty) signals."},{"metadata":{"trusted":true,"_uuid":"ab20ac984b83a05d25eaf79f49dcf981c35356ed"},"cell_type":"code","source":"fig = plt.figure(figsize=(10,8))\n\nplt.subplot(431)\nplt.plot(subset_train['0'])\nplt.subplot(432)\nplt.boxplot(subset_train['0'])\nplt.subplot(433)\nplt.hist(subset_train['0'])\n    \nplt.subplot(434)\nplt.plot(subset_train['1'])\nplt.subplot(435)\nplt.boxplot(subset_train['1'])\nplt.subplot(436)\nplt.hist(subset_train['1'])\n\nplt.subplot(437)\nplt.plot(subset_train['3'])\nplt.subplot(438)\nplt.boxplot(subset_train['3'])\nplt.subplot(439)\nplt.hist(subset_train['3'])\n\nplt.subplot(4,3,10)\nplt.plot(subset_train['4'])\nplt.subplot(4,3,11)\nplt.boxplot(subset_train['4'])\nplt.subplot(4,3,12)\nplt.hist(subset_train['4'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"28b29bfcc7af20c19e15cec6c78afb59449f5fbc"},"cell_type":"markdown","source":"At least from these couple of plots, we notice that the faulty signals (last 2) have relatively more outliers than the non-faulty ones. We will analyze this further with more data.\n\nLet's separate the positive and negative signals for further analysis. I'm going to reduce the sample sizes to make sure we don't run out of memory limits."},{"metadata":{"trusted":true,"_uuid":"a5fb73acf543de684313260c9358acf252e9cbdf"},"cell_type":"code","source":"import numpy as np\n\n# Temporarily reduce data size to build the pipeline\nsmall_subset_train = subset_train.iloc[:25000,:]\nsmall_subset_train = small_subset_train.transpose()\nsmall_subset_train.index = small_subset_train.index.astype(np.int32)\ntrain_dataset = metadata_train.join(small_subset_train, how='right')\n\n# Uncomment the following to train on the full dataset\n#subset_train = subset_train.transpose()\n#subset_train.index = subset_train.index.astype(np.int32)\n#train_dataset = metadata_train.join(subset_train, how='right')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a42f12fb8b9a3223674cc16cc5b4441d90023ee6"},"cell_type":"code","source":"positive_samples = train_dataset[train_dataset['target']==1]\npositive_samples = positive_samples.iloc[:,3:]\npositive_samples.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"24379b4d531df931d7c15e004d657425e9a7c7de"},"cell_type":"code","source":"positive_samples.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3a44931440815cb30e88b057c84e7213b356b025"},"cell_type":"markdown","source":"Now let's visualize the positive (faulty) signals using a boxplot for several of them."},{"metadata":{"trusted":true,"_uuid":"8b942a1a76c8475c9c6be3d84c8d8687ee16e3a8"},"cell_type":"code","source":"plt.figure(figsize=(10,4))\nplt.subplot(151)\nplt.boxplot(positive_samples.iloc[3,1:])\nplt.subplot(152)\nplt.boxplot(positive_samples.iloc[4,1:])\nplt.subplot(153)\nplt.boxplot(positive_samples.iloc[5,1:])\nplt.subplot(154)\nplt.boxplot(positive_samples.iloc[201,1:])\nplt.subplot(155)\nplt.boxplot(positive_samples.iloc[202,1:])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7a2d03a69c9b3d7f8f264abb020460b741aa8764"},"cell_type":"markdown","source":"We see that the data values differ a lot. Let's normalize the data first, this will also be needed for training some type of models later."},{"metadata":{"trusted":true,"_uuid":"29e59231f9412c9d820b391293a880ebaa9f1689"},"cell_type":"code","source":"# Normalize the data set\nfrom sklearn.preprocessing import StandardScaler\ny_train_pos = positive_samples.iloc[:, 0]\nX_train_pos = positive_samples.iloc[:, 1:]\nscaler = StandardScaler()\nscaler.fit(X_train_pos.T)\nX_train_pos = scaler.transform(X_train_pos.T).T","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"08dce41d4527ebe95444f46919d949684333b080"},"cell_type":"markdown","source":"Let's visualize the boxplots again using this normalized data."},{"metadata":{"trusted":true,"_uuid":"07ddaf401fb44e18f67fd1e6a901baae3ca83d01"},"cell_type":"code","source":"plt.figure(figsize=(10,4))\nplt.subplot(151)\nplt.boxplot(X_train_pos[0,:])\nplt.subplot(152)\nplt.boxplot(X_train_pos[1,:])\nplt.subplot(153)\nplt.boxplot(X_train_pos[2,:])\nplt.subplot(154)\nplt.boxplot(X_train_pos[3,:])\nplt.subplot(155)\nplt.boxplot(X_train_pos[4,:])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"08d130bb14bf8a096fdd3810b7e298a4b56f98cb"},"cell_type":"markdown","source":"Again we notice that there are a lot of outliers in the positive (faulty) signals.\n\nNow let's extract the negative (non-faulty) samples and visualize the same boxplots, and see if we can notice any apparent difference."},{"metadata":{"trusted":true,"_uuid":"f7c99280ac588f718325b75a41fda68f62475966"},"cell_type":"code","source":"negative_samples = train_dataset[train_dataset['target']==0]\nnegative_samples = negative_samples.iloc[:,3:]\nnegative_samples.info(), negative_samples.head()\n\ny_train_neg = negative_samples.iloc[:, 0]\nX_train_neg = negative_samples.iloc[:, 1:]\nscaler.fit(X_train_neg.T)\nX_train_neg = scaler.transform(X_train_neg.T).T\n\nplt.figure(figsize=(10,4))\nplt.subplot(151)\nplt.boxplot(X_train_neg[0,:])\nplt.subplot(152)\nplt.boxplot(X_train_neg[1,:])\nplt.subplot(153)\nplt.boxplot(X_train_neg[2,:])\nplt.subplot(154)\nplt.boxplot(X_train_neg[3,:])\nplt.subplot(155)\nplt.boxplot(X_train_neg[4,:])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d9bc3cef7d6e57668a441a11715d05af5c0e3436"},"cell_type":"markdown","source":"The negative (non-faulty) signals have much fewer outliers, and their magnitudes also seem to be very low. Seems like the number of outliers could be a promising feature.\n\nNow let's create the test/train split for training a Conv1D model."},{"metadata":{"trusted":true,"_uuid":"3945a11048622a9a88e5335d62bfcb1be870bc3e"},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX_train_pos, X_valid_pos, y_train_pos, y_valid_pos = train_test_split(X_train_pos, y_train_pos, \n                                                                    test_size=0.2,\n                                                                    random_state = 0,\n                                                                    shuffle=True)\n\nX_train_neg, X_valid_neg, y_train_neg, y_valid_neg = train_test_split(X_train_neg, y_train_neg, \n                                                                    test_size=0.2,\n                                                                    random_state = 0,\n                                                                    shuffle=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"28373fafd9cf52acb37e49f7aafc8a6abf57630e"},"cell_type":"code","source":"X_train_pos.shape, X_train_neg.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4496e1c8d6c09d1c95a03768c75706339856e7e1"},"cell_type":"markdown","source":"As we know, the positive samples are fewer, so we will only select a subset of negative samples for training."},{"metadata":{"trusted":true,"_uuid":"feb7e65bd79d359efaeda750ef519ecf6dca5d3c"},"cell_type":"code","source":"# Combine positive and negative samples for training...\ndef combine_positive_and_negative_samples(pos_samples, neg_samples, y_pos, y_neg):\n    X_combined = np.concatenate((pos_samples, neg_samples)) \n                                                    # don't select all negative samples, to\n                                                    # keep the samples balanced\n    y_combined = np.concatenate((y_pos, y_neg))\n    #X_train_combined.shape, y_train_combined.shape\n    combined_samples = np.hstack((X_combined, y_combined.reshape(y_combined.shape[0],1)))\n    np.random.shuffle(combined_samples)\n    return combined_samples\n\n# Only use 500 negative samples, to create a balanced dataset with the positive samples...\ntrain_samples = combine_positive_and_negative_samples(X_train_pos, X_train_neg[:500, :], y_train_pos, y_train_neg[:500])\nX_train = train_samples[:,:-1]\ny_train = train_samples[:,-1]\nX_train.shape, y_train.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1e2248e93cb07a9c0fb2466099283dfc9bdfb7bb"},"cell_type":"code","source":"# Create the validation set\n#X_valid_combined = np.concatenate((X_valid_pos, X_valid_neg[:500,:])) # don't select all negative samples, to\n                                                  # keep the samples balanced\n#y_valid_combined = np.concatenate((y_valid_pos, y_valid_neg[:500]))\n#X_valid_combined.shape, y_valid_combined.shape\n#validation_samples = np.hstack((X_valid_combined, y_valid_combined.reshape(y_valid_combined.shape[0],1)))\n#np.random.shuffle(validation_samples)\n\nvalidation_samples = combine_positive_and_negative_samples(X_valid_pos, X_valid_neg[:500,:], y_valid_pos, y_valid_neg[:500])\nX_valid = validation_samples[:,:-1]\ny_valid = validation_samples[:,-1]\nX_valid.shape, y_valid.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4383b24325024d54a179ee1a064ab3d2b94fc336"},"cell_type":"markdown","source":"A 1-D ConvNet would be an interesting model to try out on this signal. Earlier we saw that there are a lot of outliers in fauty signals. Since the actual signal value differs at different times, the outliers are relative to this mean signal value. A 1-D ConvNet can analyze the signal in various windows of increasing lengths and create high-level features out of that to classify on."},{"metadata":{"trusted":true,"_uuid":"dc8c83b24a0834051a811a1c1bb494397470d40b"},"cell_type":"code","source":"# Reshape training and validation data for keras input layer\nX_train = X_train.reshape(-1, X_train.shape[1], 1)\nX_valid = X_valid.reshape(-1, X_valid.shape[1], 1)\n\nX_train.shape, X_valid.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e0a4db51023b35cde4428e8ab827f49e4192ba74"},"cell_type":"code","source":"import tensorflow as tf\nimport tensorflow.keras as keras\nfrom tensorflow.keras.layers import Dense, Dropout, Flatten\nfrom tensorflow.keras import Input, layers\nfrom tensorflow.keras import backend as K\n\n#Conv1D Model\ndrop_out_rate = 0.2\nlr = 0.0005\ninput_tensor = Input(shape=([X_train_pos.shape[1],1]))\n\nx = layers.Conv1D(8, 11, padding='valid', activation='relu', strides=1)(input_tensor)\nx = layers.MaxPooling1D(2)(x)\nx = layers.Dropout(drop_out_rate)(x)\nx = layers.Conv1D(16, 7, padding='valid', activation='relu', strides=1)(x)\nx = layers.MaxPooling1D(2)(x)\nx = layers.Dropout(drop_out_rate)(x)\nx = layers.Conv1D(32, 5, padding='valid', activation='relu', strides=1)(x)\nx = layers.MaxPooling1D(2)(x)\nx = layers.Dropout(drop_out_rate)(x)\nx = layers.Conv1D(64, 5, padding='valid', activation='relu', strides=1)(x)\nx = layers.MaxPooling1D(2)(x)\nx = layers.Dropout(drop_out_rate)(x)\nx = layers.Conv1D(128, 3, padding='valid', activation='relu', strides=1)(x)\nx = layers.MaxPooling1D(2)(x)\nx = layers.Flatten()(x)\nx = layers.Dense(256, activation='relu')(x)\nx = layers.Dropout(drop_out_rate)(x)\nx = layers.Dense(128, activation='relu')(x)\nx = layers.Dropout(drop_out_rate)(x)\noutput_tensor = layers.Dense(1, activation='sigmoid')(x)\n\nmodel = tf.keras.Model(input_tensor, output_tensor)\n\nmodel.compile(loss=keras.losses.binary_crossentropy,\n             optimizer=keras.optimizers.Adam(lr = lr),\n             metrics=['accuracy'])\n\nmodel.summary()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bf6747eeba31c182575c9eb519913cebe19b5278"},"cell_type":"code","source":"from keras.callbacks import ModelCheckpoint\n\nweights_file=\"best_weights.hdf5\"\ncheckpoint = ModelCheckpoint(weights_file, monitor='val_acc', verbose=1, save_best_only=True, mode='max')\ncallbacks = [checkpoint]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"96e1e695b4ae8a8c2ef548e3448ae206863d9947"},"cell_type":"code","source":"batch_size = 100\nhistory = model.fit(X_train, y_train, validation_data=[X_valid, y_valid],\n          batch_size=batch_size, \n          epochs=25,\n          verbose=1, callbacks=callbacks)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8394d9d307d009502826412f94d485f1cd904830"},"cell_type":"code","source":"# plot history for accuracy\nplt.plot(history.history['acc'])\nplt.plot(history.history['val_acc'])\nplt.title('model accuracy')\nplt.ylabel('accuracy')\nplt.xlabel('epoch')\nplt.legend(['train', 'test'], loc='upper left')\nplt.show()\n# plot history for loss\nplt.plot(history.history['loss'])\nplt.plot(history.history['val_loss'])\nplt.title('model loss')\nplt.ylabel('loss')\nplt.xlabel('epoch')\nplt.legend(['train', 'test'], loc='upper left')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b34128f356f6d8d893887c04149ea0e18d8d4773"},"cell_type":"code","source":"os.listdir('./')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b23648520f869b56c027655eca6005677f21fc9e"},"cell_type":"code","source":"model.load_weights(\"best_weights.hdf5\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c2f3a811166d924c6230872fc166bd88e842a4ba"},"cell_type":"code","source":"from sklearn.metrics import roc_curve\nfrom sklearn.metrics import auc\n\ny_pred = model.predict(X_valid).ravel()\nfpr, tpr, thresholds = roc_curve(y_valid, y_pred)\nauc = auc(fpr, tpr)\n\nplt.figure(1)\nplt.plot([0, 1], [0, 1], 'k--')\nplt.plot(fpr, tpr, label='Conv1D (area = {:.3f})'.format(auc))\n#plt.plot(fpr_rf, tpr_rf, label='RF (area = {:.3f})'.format(auc_rf))\nplt.xlabel('False positive rate')\nplt.ylabel('True positive rate')\nplt.title('ROC curve')\nplt.legend(loc='best')\nplt.show()\n# Zoom in view of the upper left corner.\nplt.figure(2)\nplt.xlim(0, 0.5)\nplt.ylim(0.6, 1)\nplt.plot([0, 1], [0, 1], 'k--')\nplt.plot(fpr, tpr, label='Conv1D (area = {:.3f})'.format(auc))\n#plt.plot(fpr_rf, tpr_rf, label='RF (area = {:.3f})'.format(auc_rf))\nplt.xlabel('False positive rate')\nplt.ylabel('True positive rate')\nplt.title('ROC curve (zoomed in at top left)')\nplt.legend(loc='best')\nplt.show()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}