{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"This is my first kernel on Kaggle. I am very excited to contribute to this competition and going back to my days of studying biomechatonics where I first met the concept of feature extraction to classify signals from sensors on human body. \n\nI made a fast literature review on feature extraction and signal classification and prepared following experiment for my first submissions.\n\nI will be updating explanatory walktrough alongside model improvements.\n\nRunning and submitting from full notebook on interactive session was problemmatic due to memory troubles with the test set. I will work my way around it meanwhile you can download to use or fork to improve. Open for feedbacks to improve my first kernel.\n\nHave fun everyone!"},{"metadata":{"_uuid":"efe99fc9f6a1fa26d7fb8d05974514278fd0ec0a"},"cell_type":"markdown","source":"Many thanks to following kernels:\n- For shortening the signals with a simple feature extraction thanks to: https://www.kaggle.com/ashishpatel26/transfer-learning-in-basic-nn\n- For signal denoising and fft: https://www.kaggle.com/theoviel/fast-fourier-transform-denoising"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"import keras\nimport keras.backend as K\nfrom keras.layers import LSTM,Dropout,Dense,TimeDistributed,Conv1D,MaxPooling1D,Flatten\nfrom keras.models import Sequential\nimport tensorflow as tf\nimport gc\nfrom numba import jit\nfrom IPython.display import display, clear_output\nfrom tqdm import tqdm\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns\nimport sys\nsns.set_style(\"whitegrid\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"508ddbf86299fef8a425c2fa09e2cfe375a08bee"},"cell_type":"markdown","source":"## 0. Info"},{"metadata":{"_uuid":"b08145bc2ec474e929454c7d493731725133bffb"},"cell_type":"markdown","source":"**Signals:**\n\n- 800.000 measurement points for 8712 signals.\n- The signals are **three-phased** so there are 2904 distinct signaling instances.\n- Three phase signals:\n    - Sums to zero.\n    - When one fails other continue to carry the current.\n    - Can be rectified to be converted to DC current.\n    - Ripples in rectification can be seen on failure.\n   "},{"metadata":{"_uuid":"4fe933082e21e354b71c839bf1acbc4544a14347"},"cell_type":"markdown","source":"![](https://upload.wikimedia.org/wikipedia/commons/thumb/4/48/3-phase_flow.gif/357px-3-phase_flow.gif)"},{"metadata":{"_uuid":"90ef4c9350a8b972c257fe1c9794728abb1d0912"},"cell_type":"markdown","source":"### What is partial discharge?"},{"metadata":{"_uuid":"83199678dc23010ddb96c2628f5b56b2009cb465"},"cell_type":"markdown","source":" - Typical situation of PD: imagine there is an internal cavity/void or **impurity in insulation**. \n - When **High Voltage** is applied on conductor, a field is also induced on the cavity. Further, when the field increases, this **defect breaks down** and **discharges** different forms of energy which result in partial discharge.\n - This phenomenon is damaging over a long period of time. It is not event that occurs suddenly. "},{"metadata":{"_uuid":"0828866956a682075544fb7e032db7c68776c91b"},"cell_type":"markdown","source":"### Classical Modes of Detection\n- Partial Discharges can be detected by **measuring the emissions** they give off: Ultrasonic Sound, Transient Earth Voltages (TEV and UHF energy).\n- Is it possible to enhance the modes of detection by **better feature extraction** for the classifiers?\n- **Intel Mobile ODT** challenge on 2017 was about topping **classical image processing** methods by automatic feature extaction using pre-trained CNN models and **transfer learning**.\n- **Two possible approaches**:\n    - FE on signals and feeding them into NNs for classification.\n    - Using NNs further as feature extractors and then use shallow classifiers (XGBoost) for binary classification\n"},{"metadata":{"_uuid":"623885326f62cdc140f914e26e7d7d48ee971636"},"cell_type":"markdown","source":"### **TASK:** Classify long-term failure of covered conductors based on signal characteristics:\n- Extract features from time series data for classification.\n- Use **CNN** for further FE and **LSTM** to get temporal dependencies and perform time series classification on the top layer."},{"metadata":{"_uuid":"f195687039a2f89dbad20ec8bb132f34d2dd5e55"},"cell_type":"markdown","source":"## 1. Load Data"},{"metadata":{"trusted":true,"_uuid":"e06765b00810004c271944be1360e0b62fa64134"},"cell_type":"code","source":"import pyarrow.parquet as pq\nimport pandas as pd\nimport numpy as np","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"935c2ba4d25cae753e04645a168decb84b6ce964"},"cell_type":"code","source":"%%time \ntrain_set = pq.read_pandas('../input/train.parquet').to_pandas()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e217109aeec443c5cb64fdc716ef433f306c436c"},"cell_type":"code","source":"%%time\nmeta_train = pd.read_csv('../input/metadata_train.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"809b6250e1643fa4a4dae5230dc07772c9b56084"},"cell_type":"markdown","source":"## 2. Process and Minimize Data"},{"metadata":{"trusted":true,"_uuid":"36f9662e947b286fa6505a4c8183daff1b8c1570"},"cell_type":"code","source":"@jit('float32(float32[:,:], int32)')\ndef feature_extractor(x, n_part=1000):\n    lenght = len(x)\n    pool = np.int32(np.ceil(lenght/n_part))\n    output = np.zeros((n_part,))\n    for j, i in enumerate(range(0,lenght, pool)):\n        if i+pool < lenght:\n            k = x[i:i+pool]\n        else:\n            k = x[i:]\n        output[j] = np.max(k, axis=0) - np.min(k, axis=0)\n    return output","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"83ce36dd1a85aaad8b446e452cbc7e19b882164b"},"cell_type":"code","source":"x_train = []\ny_train = []\nfor i in tqdm(meta_train.signal_id):\n    idx = meta_train.loc[meta_train.signal_id==i, 'signal_id'].values.tolist()\n    y_train.append(meta_train.loc[meta_train.signal_id==i, 'target'].values)\n    x_train.append(abs(feature_extractor(train_set.iloc[:, idx].values, n_part=400)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4cc6eef2ae63a2c79d08c5fe1ce8670449fb807a"},"cell_type":"code","source":"del train_set; gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cad66118494f25f70c9ace7b85bea500f1c680c8"},"cell_type":"code","source":"y_train = np.array(y_train).reshape(-1,)\nX_train = np.array(x_train).reshape(-1,x_train[0].shape[0])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f634073305f25a13dbea2713f3987a98cd7a98f5"},"cell_type":"markdown","source":"## 3. Build Primitive CNN + LSTM Model"},{"metadata":{"_uuid":"0dea1e4b86a1b6d9a6a516ae97eeefa60581b71e"},"cell_type":"markdown","source":"* CNN is for feature extraction and LSTM is for capturing time dependency."},{"metadata":{"trusted":true,"_uuid":"5250bf32e062836cbf10394e5f9df4d6b695380e"},"cell_type":"code","source":"def keras_auc(y_true, y_pred):\n    auc = tf.metrics.auc(y_true, y_pred)[1]\n    K.get_session().run(tf.local_variables_initializer())\n    return auc","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"034ea6d8cf8bd8a09eb5943ebb3faa01cd991e9d"},"cell_type":"code","source":"n_signals = 1 #So far each instance is one signal. We will diversify them in next step\nn_outputs = 1 #Binary Classification","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1c05ecc6abef0b822e422950a5824ea8c55ba861"},"cell_type":"code","source":"#Build the model\nverbose, epochs, batch_size = True, 15, 16\nn_steps, n_length = 40, 10\nX_train = X_train.reshape((X_train.shape[0], n_steps, n_length, n_signals))\n# define model\nmodel = Sequential()\nmodel.add(TimeDistributed(Conv1D(filters=64, kernel_size=3, activation='relu'), input_shape=(None,n_length,n_signals)))\nmodel.add(TimeDistributed(Conv1D(filters=64, kernel_size=3, activation='relu')))\nmodel.add(TimeDistributed(Dropout(0.5)))\nmodel.add(TimeDistributed(MaxPooling1D(pool_size=2)))\nmodel.add(TimeDistributed(Flatten()))\nmodel.add(LSTM(100))\nmodel.add(Dropout(0.5))\nmodel.add(Dense(100, activation='relu'))\nmodel.add(Dense(n_outputs, activation='sigmoid'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e48f5b3c1e9cffb34ef0f395f668a205b9cd2508"},"cell_type":"code","source":"model.compile(loss='binary_crossentropy', optimizer='adam', metrics=[keras_auc])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7092f9b2feb4b98b47a72738b956a38f178e0075"},"cell_type":"code","source":"model.fit(X_train, y_train, epochs=epochs, batch_size=batch_size, verbose=verbose)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"470dee0dc3d9edaaad403eb14275308cb84fb9f8"},"cell_type":"code","source":"model.save_weights('model1.hdf5')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"18ba795d40eabef956f7c66bf5cdfbd4733b08d9"},"cell_type":"code","source":"#%%time\n#test_set = pq.read_pandas('../input/test.parquet').to_pandas()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f0b9e956922420544e5526499647fa6a810a619f"},"cell_type":"code","source":"#%%time\n#meta_test = pd.read_csv('../input/metadata_test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6aec10ef8cd423326af7871140ed62285255c67a"},"cell_type":"code","source":"#x_test = []\n#for i in tqdm(meta_test.signal_id.values):\n#    idx=i-8712\n#    clear_output(wait=True)\n#    x_test.append(abs(feature_extractor(test_set.iloc[:, idx].values, n_part=400)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a7e9b508c6cd1531b1a1574a2397871d4b9ebc32"},"cell_type":"code","source":"#del test_set; gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9fd4b08be3425a2d6b194adb9dff35c4fd03f01e"},"cell_type":"code","source":"#X_test = x_test.reshape((x_test.shape[0], n_steps, n_length, n_signals))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"337e04c8799ba03e528cebc92485615b322b44bd"},"cell_type":"code","source":"#preds = model.predict(X_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"72a34b08ab91a2a7239d42eb6c69b37ecd08aa7b"},"cell_type":"code","source":"#threshpreds = (preds>0.5)*1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"28fa72d1dcdf0139173d1ee9f2b0ddf21ce53b3c"},"cell_type":"code","source":"#sub = pd.read_csv('../input/sample_submission.csv')\n#sub.target = threshpreds","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b4065303092c127e26350dc645827efe23636dde"},"cell_type":"code","source":"#sub.to_csv('first_sub.csv',index=False)\n#Gave me an LB score of 0.450","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a1f27cd1e603a8fe5563c5c826de7fb8ea7eb967"},"cell_type":"markdown","source":"## 4. Further processing of Signals to Diversify the Model"},{"metadata":{"_uuid":"2601daed31d00d2fb9c6d1699c6fba95ddaaf507"},"cell_type":"markdown","source":"We only fed what we were given and only feature engineering was to reduce signal lengths because 800000 time steps is troublesome with LSTM. Signal classification applications takes more than one modalities or channels in real life. So I will try to increase the feature size in terms of channel depth rather than feature count.\n\nProposed data structure: **n_instances x n_timesteps x  n_channels**"},{"metadata":{"_uuid":"9e6a98f5ecfc0e3becff82a504be2e7292685c31"},"cell_type":"markdown","source":"### 4.a - Filter & Transform Signals"},{"metadata":{"trusted":true,"_uuid":"7f687572ed8f1d315c45213f78fc40abe198657b"},"cell_type":"code","source":"#Both numpy and scipy has utilities for FFT which is an endlessly useful algorithm\nfrom numpy.fft import *\nfrom scipy import fftpack","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"47f7e4eb85d0e1b80a2b76e7d3cb4b1042a5a633"},"cell_type":"code","source":"%%time \ntrain_set = pq.read_pandas('../input/train.parquet').to_pandas()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"11c755d8e57368c280c0f0a029008b78027e1f87"},"cell_type":"code","source":"#FFT to filter out HF components and get main signal profile\ndef low_pass(s, threshold=1e4):\n    fourier = rfft(s)\n    frequencies = rfftfreq(s.size, d=2e-2/s.size)\n    fourier[frequencies > threshold] = 0\n    return irfft(fourier)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"39be5cedb33a36e82bd9c391776b18ecbf6bca69"},"cell_type":"code","source":"def phase_indices(signal_num):\n    phase1 = 3*signal_num\n    phase2 = 3*signal_num + 1\n    phase3 = 3*signal_num + 2\n    return phase1,phase2,phase3","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a68235428f0a1fc22cb99c8067e5bd2d026f1547"},"cell_type":"code","source":"s_id = 14\np1,p2,p3 = phase_indices(s_id)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"04e65ef8f770d1ed0fb00b374702c8b898b8a339"},"cell_type":"code","source":"plt.figure(figsize=(10,5))\nplt.title('Signal %d / Target:%d'%(s_id,meta_train[meta_train.id_measurement==s_id].target.unique()[0]))\nplt.plot(train_set.iloc[:,p1])\nplt.plot(train_set.iloc[:,p2])\nplt.plot(train_set.iloc[:,p3])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"be8b5b0f9808be748fef478a2090914dc6d92ac6"},"cell_type":"code","source":"lf_signal_1 = low_pass(train_set.iloc[:,p1])\nlf_signal_2 = low_pass(train_set.iloc[:,p2])\nlf_signal_3 = low_pass(train_set.iloc[:,p3])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b4e6bcd6b78a7083708b43d5f47c0bc8692920f7"},"cell_type":"code","source":"plt.figure(figsize=(10,5))\nplt.title('De-noised Signal %d / Target:%d'%(s_id,meta_train[meta_train.id_measurement==s_id].target.unique()[0]))\nplt.plot(lf_signal_1)\nplt.plot(lf_signal_2)\nplt.plot(lf_signal_3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"47481a5628b7063c2039e6e7640f51ef8a240187"},"cell_type":"code","source":"plt.figure(figsize=(10,5))\nplt.title('Signal %d AbsVal / Target: %d'%(s_id,meta_train[meta_train.id_measurement==s_id].target.unique()[0]))\nplt.plot(np.abs(lf_signal_1))\nplt.plot(np.abs(lf_signal_2))\nplt.plot(np.abs(lf_signal_3))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e17b09b88409c0cf738836cacc952550a9a2518c"},"cell_type":"code","source":"plt.figure(figsize=(10,5))\nplt.title('Signal %d / Target: %d'%(s_id,meta_train[meta_train.id_measurement==s_id].target.unique()[0]))\nplt.plot(lf_signal_1)\nplt.plot(lf_signal_2)\nplt.plot(lf_signal_3)\nplt.plot((np.abs(lf_signal_1)+np.abs(lf_signal_2)+np.abs(lf_signal_3)))\nplt.legend(['phase 1','phase 2','phase 3','DC Component'],loc=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c0b762e3058accf3f8de4d0a1e0fe8fa07f8845c"},"cell_type":"code","source":"###Filter out low frequencies from the signal to get HF characteristics\ndef high_pass(s, threshold=1e7):\n    fourier = rfft(s)\n    frequencies = rfftfreq(s.size, d=2e-2/s.size)\n    fourier[frequencies < threshold] = 0\n    return irfft(fourier)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5c74a2a35e4d6a858af94ae220cfa14ea3942413"},"cell_type":"code","source":"hf_signal_1 = high_pass(train_set.iloc[:,p1])\nhf_signal_2 = high_pass(train_set.iloc[:,p2])\nhf_signal_3 = high_pass(train_set.iloc[:,p3])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5a6f0e8175af41d523e91a5c76ca895fce94360d"},"cell_type":"code","source":"plt.figure(figsize=(10,5))\nplt.title('Signal %d / Target:%d'%(s_id,meta_train[meta_train.id_measurement==s_id].target.unique()[0]))\nplt.plot(hf_signal_1)\n#plt.plot(hf_signal_2)\n#plt.plot(hf_signal_3)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7a0387ec172763edcd3685194e5276accc63fca9"},"cell_type":"markdown","source":"As seen above we can decouple the signals into their high and low frequency components using FFT. So the number of signal channels that are fed into the model can may very well be diversified using different decouplings in time and frequency domains. \n\nThe mode of decoupling so far was filtering what we have to get new signals. Following part is about playing around the frequency domain to create features. "},{"metadata":{"_uuid":"80acc5554c7f86b1e031f255f8dcd90e8cf8c94e"},"cell_type":"markdown","source":"### 4.b - Spectogram Features"},{"metadata":{"trusted":true,"_uuid":"26757e0293dce7977d097cf6e933f177bcb0267b"},"cell_type":"code","source":"signal = train_set.iloc[:,p1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9067c754a762f671f210748af4be6a0fa0443cd1"},"cell_type":"code","source":"%%time\nx = signal\nX = fftpack.fft(x,n=400)\nfreqs = fftpack.fftfreq(n=400,d=2e-2/x.size) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"023dc77fb968a6cc05bb85462ae94262183de69b"},"cell_type":"code","source":"plt.plot(x)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"26dd8cdcaed41376c998414404724ff1da9e6e6a"},"cell_type":"code","source":"fig, ax = plt.subplots()\nax.set_title('Full Spectrum with Scipy')\nax.set_xlabel('Frequency in Hertz [Hz]')\nax.set_ylabel('Frequency Domain (Spectrum) Magnitude')\nax.stem(freqs[1:], np.abs(X)[1:])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"11bf4c4dd27eb0925132822db73a9082869737eb"},"cell_type":"code","source":"%%time\nx = high_pass(train_set.iloc[:,p1])\nX = fftpack.fft(x,n=400)\nfreqs = fftpack.fftfreq(n=400,d=2e-2/x.size) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e27e040f311a24f35aef6d956bb16af04fe216e4"},"cell_type":"code","source":"plt.plot(x)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b2bb4747bc4e35699ace4151530993d1345cd9d5"},"cell_type":"code","source":"fig, ax = plt.subplots()\nax.set_title('High Frequency Spectrum')\nax.set_xlabel('Frequency in Hertz [Hz]')\nax.set_ylabel('Frequency Domain (Spectrum) Magnitude')\nax.stem(freqs[1:], np.abs(X)[1:])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"855e9fd8cc209d5c181f451fc048a93b19aa6d94"},"cell_type":"code","source":"%%time\nx = low_pass(signal)\nX = fftpack.fft(x,n=400)\nfreqs = fftpack.fftfreq(n=400,d=2e-2/x.size) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"945fc0540eef72d0da58bfabf914264e1a58847a"},"cell_type":"code","source":"plt.plot(x)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6bc15753fa47ab9117ccd1a6a94b95d6a0dd4a02"},"cell_type":"code","source":"fig, ax = plt.subplots()\nax.set_title('Low Frequency Spectrum')\nax.set_xlabel('Frequency in Hertz [Hz]')\nax.set_ylabel('Frequency Domain (Spectrum) Magnitude')\nax.stem(freqs[1:], np.abs(X)[1:])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"04bdfdc27c5fafa08f972a72d78f0640daf14784"},"cell_type":"markdown","source":"### 4.c - Frequencies vs. Time\n"},{"metadata":{"trusted":true,"_uuid":"5cd1e3aa881bd66d9a6f89032666658ce2847f51"},"cell_type":"code","source":"p1,p2,p3 = phase_indices(100)\nsignal = train_set.iloc[:,p1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c8a8e7aef58c9a52f2bf7f8b3d6fc60c9ca0f09d"},"cell_type":"code","source":"from scipy import signal as sgn\nM = 1024\nrate = 1/(2e-2/signal.size)\n\nfreqs, times, Sx = sgn.spectrogram(signal.values, fs=rate, window='hanning',\n                                      nperseg=1024, noverlap=M - 100,\n                                      detrend='constant', scaling='spectrum')\n\nf, ax = plt.subplots(figsize=(10, 5))\nax.set_ylabel('Frequency [Hz]')\nax.set_xlabel('Time(s)')\nax.set_title('Spectogram')\nax.pcolormesh(times, freqs, np.log10(Sx), cmap='viridis')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"909fa8bad93291ba0c3f88bc8bde788fb241438f"},"cell_type":"markdown","source":"Now we have the temporal behavior of frequencies that compose our signal in terms of their magnitudes. Aggregations can be made on both time and frequency axes by selecting appropriate window sizes."},{"metadata":{"_uuid":"992088be6daace02127ec10ac93b4055d6503c31"},"cell_type":"markdown","source":"## 5. Add Features"},{"metadata":{"_uuid":"193bd5f0560ab4f93ce202065c3f4a002b744744"},"cell_type":"markdown","source":"Apart from the signals themselves we have: \n- Low Pass Filtered Signal (800000)\n- High Pass Filtered Signal (800000)\n- DC component from adding absolute values of 3-phases. (800000)\n- Frequency Magnitude Spectrum of full signal (400)\n- Frequency Magnitude Spectrum of Low Freq signal (400)\n- Frequency Magnitude Spectrum of High Freq signal (400)\n- Frequency Spectrum "},{"metadata":{"_uuid":"a88a5a709e99b43ad120192e752fc23c6874d20b"},"cell_type":"markdown","source":"For initial improvement of our model we will use diversification of signals from **4.a** and create **4 channels**:\n- **Signal itself**\n- **LF** component\n- **HF** component\n- **DC** component from three-phase merge)\n\nThe justification for DC component is that all three phase signals are rectified to form a DC behavior with a small amount of ripple. I naively assumed that a partial discharce occuring in any of the 3 signals, will result in corruption at the resulting voltage behavior. While composing up the features, each instance that belong to a triple will have the same \"DC component.\""},{"metadata":{"trusted":true,"_uuid":"70265fd797e470aa5e54663c2aadae8f43565dd9"},"cell_type":"code","source":"x_train_lp = []\nx_train_hp = []\nx_train_dc = []\nfor i in meta_train.signal_id:\n    idx = meta_train.loc[meta_train.signal_id==i, 'signal_id'].values.tolist()\n    clear_output(wait=True)\n    display(idx)\n    hp = high_pass(train_set.iloc[:, idx[0]])\n    lp = low_pass(train_set.iloc[:, idx[0]])\n    meas_id = meta_train.id_measurement[meta_train.signal_id==idx].values[0]\n    p1,p2,p3=phase_indices(meas_id)\n    lf_signal_1,lf_signal_2,lf_signal_3 = low_pass(train_set.iloc[:,p1]), low_pass(train_set.iloc[:,p2]), low_pass(train_set.iloc[:,p3])\n    dc = np.abs(lf_signal_1)+np.abs(lf_signal_2)+np.abs(lf_signal_3)\n    x_train_lp.append(abs(feature_extractor(lp, n_part=400)))\n    x_train_hp.append(abs(feature_extractor(hp, n_part=400)))\n    x_train_dc.append(abs(feature_extractor(dc, n_part=400)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e46111dc607821b99c9b3e65cf6063f71673a42f"},"cell_type":"code","source":"del train_set; gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"14924ecbcfbfd2c7044d6c6fdde666e9861c1772"},"cell_type":"code","source":"#x_test_lp = []\n#x_test_hp = []\n#x_test_dc = []\n#for i in tqdm(meta_test.signal_id):\n#    idx = idx=i-8712\n#    clear_output(wait=True)\n#    #display(idx)\n#    hp = high_pass(test_set.iloc[:, idx])\n#    lp = low_pass(test_set.iloc[:, idx])\n#    meas_id = meta_test.id_measurement[meta_test.signal_id==i].values[0]\n#    p1,p2,p3=phase_indices(meas_id)\n#    lf_signal_1,lf_signal_2,lf_signal_3 = low_pass(test_set.iloc[:,p1-8712]), low_pass(test_set.iloc[:,p2-8712]), low_pass(test_set.iloc[:,p3-8712])\n#    dc = np.abs(lf_signal_1)+np.abs(lf_signal_2)+np.abs(lf_signal_3)\n#    x_test_lp.append(abs(feature_extractor(lp, n_part=400)))\n#    x_test_hp.append(abs(feature_extractor(hp, n_part=400)))\n#    x_test_dc.append(abs(feature_extractor(dc, n_part=400)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cc3e12c50c8d0605026b723194b872ccd6f84069"},"cell_type":"code","source":"x_train = np.array(x_train).reshape(-1,x_train[0].shape[0])\nx_train_lp = np.array(x_train).reshape(-1,x_train_lp[0].shape[0])\nx_train_hp = np.array(x_train).reshape(-1,x_train_hp[0].shape[0])\nx_train_dc = np.array(x_train).reshape(-1,x_train_dc[0].shape[0])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"97d24a21c443595875c24d930942c005534434a3"},"cell_type":"code","source":"#x_test = np.array(x_test).reshape(-1,x_test[0].shape[0])\n#x_test_lp = np.array(x_test).reshape(-1,x_test_lp[0].shape[0])\n#x_test_hp = np.array(x_test).reshape(-1,x_test_hp[0].shape[0])\n#x_test_dc = np.array(x_test).reshape(-1,x_test_dc[0].shape[0])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2c5cdada44633512db4c90aa46ae9f3f8e59b7b0"},"cell_type":"code","source":"train = np.dstack((x_train,x_train_lp,x_train_hp,x_train_dc))\n#test = np.dstack((x_test,x_test_lp,x_test_hp,x_test_dc))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6065f7fdad52237dc0f4ebaeb50310a5842d7326"},"cell_type":"code","source":"y_train = np.array(y_train).reshape(-1,)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b943c536c2dcb5484695ef64af16a323c6811667"},"cell_type":"code","source":"verbose, epochs, batch_size = True, 15, 16\nn_signals,n_steps, n_length = 4,40, 10\ntrain = train.reshape((train.shape[0], n_steps, n_length, n_signals))\n# define model\nmodel = Sequential()\nmodel.add(TimeDistributed(Conv1D(filters=64, kernel_size=3, activation='relu'), input_shape=(None,n_length,n_signals)))\nmodel.add(TimeDistributed(Conv1D(filters=64, kernel_size=3, activation='relu')))\nmodel.add(TimeDistributed(Dropout(0.5)))\nmodel.add(TimeDistributed(MaxPooling1D(pool_size=2)))\nmodel.add(TimeDistributed(Flatten()))\nmodel.add(LSTM(100))\nmodel.add(Dropout(0.5))\nmodel.add(Dense(100, activation='relu'))\nmodel.add(Dense(n_outputs, activation='sigmoid'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"db18c624f053718fe0c45968300f25a185973f32"},"cell_type":"code","source":"model.compile(loss='binary_crossentropy', optimizer='adam', metrics=[keras_auc])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"53f2c0bd22234e11c941c7b94955c23e3839d509"},"cell_type":"code","source":"# fit network\nmodel.fit(train, y_train, epochs=epochs, batch_size=batch_size, verbose=verbose)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"64070e81d596caebeb85f35ba375fc32dd50c313"},"cell_type":"code","source":"model.save_weights('model2.hdf5')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5453b1751be24359f1aab05d480fbc250916a643"},"cell_type":"code","source":"#X_test = test.reshape((test.shape[0], n_steps, n_length, n_signals))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c0805178ae75af58f4adb66ad29780a777bdab8f"},"cell_type":"code","source":"#preds = model.predict(X_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a145327bc88ce5db35424e4f63295e914dd47d63"},"cell_type":"code","source":"#threshpreds = (preds>0.5)*1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"66244cae3ebd17c315e61778bbf4c670afdf32de"},"cell_type":"code","source":"#sub = pd.read_csv('data/sample_submission.csv')\n#sub.target = threshpreds","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c6abe96e37f9c59ac1694f3f6df6dd74969b0f85"},"cell_type":"code","source":"#sub.to_csv('submissions/second_sub.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"517451abf6dc8d05e94d75c1fea124901df73894"},"cell_type":"markdown","source":"### Second submission with 4 channels scores higher on LB. Increased number of epochs and playing with threshold may yield better results. I got 0.513 with re-training the model for 3+ times. "},{"metadata":{"_uuid":"2aaa70b4f064ad49b6111959975b02e1c71516e5"},"cell_type":"markdown","source":"# TODO#1: Turn Spectral Analysis into Features\n# TODO#2: Enhcance the experiment with cross_val, adaptive learning rate, early stopping, ensembling etc."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}