{
  "id": 86616,
  "title": "2nd Place Solution",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/86616",
  "author_name": "HELLO TANG",
  "post_date": "2019-03-25T11:38:47.441000",
  "votes": 30,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Congratulations to all the winners! \nAnd thanks a lot to VSB/Enet Center and Kaggle for this exciting competition.</p>\n\n<h2>Here are the details about my single model:</h2>\n\n<p>Preprocessing:\nBefore extracting features, I used DWT for noise reduction, and I think this is helpful for the stability of the model.<a href=\"https://www.kaggle.com/jackvial/dwt-signal-denoising\">you can find here</a></p>\n\n<h3>Features:</h3>\n\n<p>I tried a lot of different methods, such as RNN taking 80, 100, 160, 250 time steps, basic features, energy features, peak features, and also tried to automatically extract features from CNN models, different RNN models. Unfortunately, none of these attempts have brought me a big improvement. Finally, the addition of the peak feature on the basic RNN architecture can effectively alleviate the over-fitting. This is the problem that I think is the most important problem to solve. Extracting the peak features using python's signal library, all this features can be found in <a href=\"https://ieeexplore.ieee.org/abstract/document/7909221\">here</a>. Here is the code:</p>\n\n<p>```python</p>\n\n<h1>Extract peak features</h1>\n\n<p>from scipy.signal import find_peaks, peak_widths, peak_prominences</p>\n\n<p>def remove_false_peak(signal, p1, p2, maxDistance=10):\n    peak_diff = np.diff(p2)\n    if len(peak_diff) == 0:\n        return p1\n    ticks = []\n    for i, d in enumerate(peak_diff):\n        ratio = signal[p2[i+1]]/signal[p2[i]]\n        if d &lt; maxDistance and -0.25 &gt; ratio and ratio &gt; -4:\n            ticks.append((p2[i], p2[i+1]))\n    mask = np.array([True]*len(p1))\n    for i, j in ticks:\n        mask = mask &amp; ((p1 &lt; i) | (p1 &gt; 500+j))\n    return p1[mask]</p>\n\n<p>def get_peaks(signal):\n    p1_1, _ = find_peaks(signal, height=[5, 100])\n    p1_2, _ = find_peaks(-signal, height=[5, 100])\n    p1 = np.union1d(p1_1, p1_2)\n    n_peaks, _ = find_peaks(-signal, height=[10, 100])\n    p_peaks, _ = find_peaks(signal, height=[10, 100])\n    p2 = np.union1d(n_peaks, p_peaks)\n    p = remove_false_peak(signal, p1, p2, maxDistance=10)\n    return np.intersect1d(p1_1, p), np.intersect1d(p1_2, p)</p>\n\n<p>def extract_peak_feature(signal):\n    p_peaks, n_peaks = get_peaks(signal)</p>\n\n<pre><code>num_p, num_n = len(p_peaks), len(n_peaks)\n\nsig_peak_width = np.concatenate(\n    [peak_widths(signal, p_peaks)[0], peak_widths(-signal, n_peaks)[0]])\nsig_peak_height = abs(signal[np.concatenate([p_peaks, n_peaks])])\n\nif num_n or num_p:\n    height_mean = sig_peak_height.mean()\n    height_max = sig_peak_height.max()\n    height_min = sig_peak_height.min()\n    height_median = np.median(sig_peak_height)\n\n    width_mean = sig_peak_width.mean()\n    width_max = sig_peak_width.max()\n    width_min = sig_peak_width.min()\n    width_median = np.median(sig_peak_width)\n\n    return np.array([height_mean, height_max, height_min, height_median,\n                     width_mean, width_max, width_min, width_median, num_p, num_n])\nelse:\n    return np.zeros(10)\n</code></pre>\n\n<p>```</p>\n\n<p>At the same time, the added peak feature has a large number of outliers, which I convert to a missing value. Then perform the missing value processing(dividing the data into groups based on the attribute with the largest correlation coefficient of the missing value, and then calculating the average value of each group. Just put these averages in the missing values.)</p>\n\n<h3>Training:</h3>\n\n<p>epochs: 25\nCheckpoint monitor='val_loss'</p>\n\n<h2>My final solution was a ensemble of three models:</h2>\n\n<p>My single Model\n<a href=\"https://www.kaggle.com/tarunpaparaju/vsb-competition-attention-bilstm-with-features?scriptVersionId=10690570\">VSB Competition : Stacked Attention Capsule BiLSTM</a>\n<a href=\"https://www.kaggleusercontent.com/kf/10818864/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..rLDkCCqNGYG5hfrj7oMt9A.ucfuA1j7MlivrTFzzAvVq7SpDSojFzTdXNVHqC5T7q0Vc4AjG5OP-2Pi0EngziSLz6FHxHoY4lqxvYj02gOtfMh9dGMJnCitLHWZ4JrZX10kzWvvLYhAmbtfm6Mk2ej46868zJzHFQ9RKnvcUjjBNQ.abKNk9CPW8feEC28o41osg/__results__.html\">Handmade features</a></p>\n\n<p>If you think there is something incorrect or that could be improved, please leave your comments! And thank you everybody for the great kernels and discussions.(From 6th but it is exactly what I want to say)</p>",
  "messages": [
    {
      "id": 499909,
      "postDate": "2019-03-25T11:38:47.440Z",
      "content": "<p>Congratulations to all the winners! \nAnd thanks a lot to VSB/Enet Center and Kaggle for this exciting competition.</p>\n\n<h2>Here are the details about my single model:</h2>\n\n<p>Preprocessing:\nBefore extracting features, I used DWT for noise reduction, and I think this is helpful for the stability of the model.<a href=\"https://www.kaggle.com/jackvial/dwt-signal-denoising\">you can find here</a></p>\n\n<h3>Features:</h3>\n\n<p>I tried a lot of different methods, such as RNN taking 80, 100, 160, 250 time steps, basic features, energy features, peak features, and also tried to automatically extract features from CNN models, different RNN models. Unfortunately, none of these attempts have brought me a big improvement. Finally, the addition of the peak feature on the basic RNN architecture can effectively alleviate the over-fitting. This is the problem that I think is the most important problem to solve. Extracting the peak features using python's signal library, all this features can be found in <a href=\"https://ieeexplore.ieee.org/abstract/document/7909221\">here</a>. Here is the code:</p>\n\n<p>```python</p>\n\n<h1>Extract peak features</h1>\n\n<p>from scipy.signal import find_peaks, peak_widths, peak_prominences</p>\n\n<p>def remove_false_peak(signal, p1, p2, maxDistance=10):\n    peak_diff = np.diff(p2)\n    if len(peak_diff) == 0:\n        return p1\n    ticks = []\n    for i, d in enumerate(peak_diff):\n        ratio = signal[p2[i+1]]/signal[p2[i]]\n        if d &lt; maxDistance and -0.25 &gt; ratio and ratio &gt; -4:\n            ticks.append((p2[i], p2[i+1]))\n    mask = np.array([True]*len(p1))\n    for i, j in ticks:\n        mask = mask &amp; ((p1 &lt; i) | (p1 &gt; 500+j))\n    return p1[mask]</p>\n\n<p>def get_peaks(signal):\n    p1_1, _ = find_peaks(signal, height=[5, 100])\n    p1_2, _ = find_peaks(-signal, height=[5, 100])\n    p1 = np.union1d(p1_1, p1_2)\n    n_peaks, _ = find_peaks(-signal, height=[10, 100])\n    p_peaks, _ = find_peaks(signal, height=[10, 100])\n    p2 = np.union1d(n_peaks, p_peaks)\n    p = remove_false_peak(signal, p1, p2, maxDistance=10)\n    return np.intersect1d(p1_1, p), np.intersect1d(p1_2, p)</p>\n\n<p>def extract_peak_feature(signal):\n    p_peaks, n_peaks = get_peaks(signal)</p>\n\n<pre><code>num_p, num_n = len(p_peaks), len(n_peaks)\n\nsig_peak_width = np.concatenate(\n    [peak_widths(signal, p_peaks)[0], peak_widths(-signal, n_peaks)[0]])\nsig_peak_height = abs(signal[np.concatenate([p_peaks, n_peaks])])\n\nif num_n or num_p:\n    height_mean = sig_peak_height.mean()\n    height_max = sig_peak_height.max()\n    height_min = sig_peak_height.min()\n    height_median = np.median(sig_peak_height)\n\n    width_mean = sig_peak_width.mean()\n    width_max = sig_peak_width.max()\n    width_min = sig_peak_width.min()\n    width_median = np.median(sig_peak_width)\n\n    return np.array([height_mean, height_max, height_min, height_median,\n                     width_mean, width_max, width_min, width_median, num_p, num_n])\nelse:\n    return np.zeros(10)\n</code></pre>\n\n<p>```</p>\n\n<p>At the same time, the added peak feature has a large number of outliers, which I convert to a missing value. Then perform the missing value processing(dividing the data into groups based on the attribute with the largest correlation coefficient of the missing value, and then calculating the average value of each group. Just put these averages in the missing values.)</p>\n\n<h3>Training:</h3>\n\n<p>epochs: 25\nCheckpoint monitor='val_loss'</p>\n\n<h2>My final solution was a ensemble of three models:</h2>\n\n<p>My single Model\n<a href=\"https://www.kaggle.com/tarunpaparaju/vsb-competition-attention-bilstm-with-features?scriptVersionId=10690570\">VSB Competition : Stacked Attention Capsule BiLSTM</a>\n<a href=\"https://www.kaggleusercontent.com/kf/10818864/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..rLDkCCqNGYG5hfrj7oMt9A.ucfuA1j7MlivrTFzzAvVq7SpDSojFzTdXNVHqC5T7q0Vc4AjG5OP-2Pi0EngziSLz6FHxHoY4lqxvYj02gOtfMh9dGMJnCitLHWZ4JrZX10kzWvvLYhAmbtfm6Mk2ej46868zJzHFQ9RKnvcUjjBNQ.abKNk9CPW8feEC28o41osg/__results__.html\">Handmade features</a></p>\n\n<p>If you think there is something incorrect or that could be improved, please leave your comments! And thank you everybody for the great kernels and discussions.(From 6th but it is exactly what I want to say)</p>",
      "rawMarkdown": "Congratulations to all the winners! \nAnd thanks a lot to VSB/Enet Center and Kaggle for this exciting competition.\n\n## Here are the details about my single model:\nPreprocessing:\nBefore extracting features, I used DWT for noise reduction, and I think this is helpful for the stability of the model.[you can find here](https://www.kaggle.com/jackvial/dwt-signal-denoising)\n\n### Features:\nI tried a lot of different methods, such as RNN taking 80, 100, 160, 250 time steps, basic features, energy features, peak features, and also tried to automatically extract features from CNN models, different RNN models. Unfortunately, none of these attempts have brought me a big improvement. Finally, the addition of the peak feature on the basic RNN architecture can effectively alleviate the over-fitting. This is the problem that I think is the most important problem to solve. Extracting the peak features using python's signal library, all this features can be found in [here](https://ieeexplore.ieee.org/abstract/document/7909221). Here is the code:\n\n```python\n# Extract peak features\nfrom scipy.signal import find_peaks, peak_widths, peak_prominences\n\ndef remove_false_peak(signal, p1, p2, maxDistance=10):\n    peak_diff = np.diff(p2)\n    if len(peak_diff) == 0:\n        return p1\n    ticks = []\n    for i, d in enumerate(peak_diff):\n        ratio = signal[p2[i+1]]/signal[p2[i]]\n        if d &lt; maxDistance and -0.25 &gt; ratio and ratio &gt; -4:\n            ticks.append((p2[i], p2[i+1]))\n    mask = np.array([True]*len(p1))\n    for i, j in ticks:\n        mask = mask &amp; ((p1 &lt; i) | (p1 &gt; 500+j))\n    return p1[mask]\n\n\ndef get_peaks(signal):\n    p1_1, _ = find_peaks(signal, height=[5, 100])\n    p1_2, _ = find_peaks(-signal, height=[5, 100])\n    p1 = np.union1d(p1_1, p1_2)\n    n_peaks, _ = find_peaks(-signal, height=[10, 100])\n    p_peaks, _ = find_peaks(signal, height=[10, 100])\n    p2 = np.union1d(n_peaks, p_peaks)\n    p = remove_false_peak(signal, p1, p2, maxDistance=10)\n    return np.intersect1d(p1_1, p), np.intersect1d(p1_2, p)\n\n\ndef extract_peak_feature(signal):\n    p_peaks, n_peaks = get_peaks(signal)\n\n    num_p, num_n = len(p_peaks), len(n_peaks)\n\n    sig_peak_width = np.concatenate(\n        [peak_widths(signal, p_peaks)[0], peak_widths(-signal, n_peaks)[0]])\n    sig_peak_height = abs(signal[np.concatenate([p_peaks, n_peaks])])\n\n    if num_n or num_p:\n        height_mean = sig_peak_height.mean()\n        height_max = sig_peak_height.max()\n        height_min = sig_peak_height.min()\n        height_median = np.median(sig_peak_height)\n\n        width_mean = sig_peak_width.mean()\n        width_max = sig_peak_width.max()\n        width_min = sig_peak_width.min()\n        width_median = np.median(sig_peak_width)\n\n        return np.array([height_mean, height_max, height_min, height_median,\n                         width_mean, width_max, width_min, width_median, num_p, num_n])\n    else:\n        return np.zeros(10)\n```\n\nAt the same time, the added peak feature has a large number of outliers, which I convert to a missing value. Then perform the missing value processing(dividing the data into groups based on the attribute with the largest correlation coefficient of the missing value, and then calculating the average value of each group. Just put these averages in the missing values.)\n\n### Training:\nepochs: 25\nCheckpoint monitor='val_loss'\n\n## My final solution was a ensemble of three models:\nMy single Model\n[VSB Competition : Stacked Attention Capsule BiLSTM](https://www.kaggle.com/tarunpaparaju/vsb-competition-attention-bilstm-with-features?scriptVersionId=10690570)\n[Handmade features](https://www.kaggleusercontent.com/kf/10818864/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..rLDkCCqNGYG5hfrj7oMt9A.ucfuA1j7MlivrTFzzAvVq7SpDSojFzTdXNVHqC5T7q0Vc4AjG5OP-2Pi0EngziSLz6FHxHoY4lqxvYj02gOtfMh9dGMJnCitLHWZ4JrZX10kzWvvLYhAmbtfm6Mk2ej46868zJzHFQ9RKnvcUjjBNQ.abKNk9CPW8feEC28o41osg/__results__.html)\n\nIf you think there is something incorrect or that could be improved, please leave your comments! And thank you everybody for the great kernels and discussions.(From 6th but it is exactly what I want to say)",
      "votes": 30
    },
    {
      "id": 501394,
      "postDate": "2019-03-27T09:51:51.093Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations",
      "votes": 1
    },
    {
      "id": 500281,
      "postDate": "2019-03-25T19:50:49.107Z",
      "content": "<p>Congratulations! I also tried to put peak features into LSTM, but I got 0.541 in public - in private this model obtained 0.692 and of course I didn't choose it. ;) I have few questions. \n1. What RNN architecture did you use? Is it the same as from Bruno kernel? \n2. Didn't you use high pass filter before getting peaks?\n3. Why did you set height=[5, 100]? If I remember correctly peaks with amplitude above 20 are probably corona discharge (I don't remember where I read about it).\n4. What's public and private score of your single model?</p>",
      "rawMarkdown": "Congratulations! I also tried to put peak features into LSTM, but I got 0.541 in public - in private this model obtained 0.692 and of course I didn't choose it. ;) I have few questions. \n1. What RNN architecture did you use? Is it the same as from Bruno kernel? \n2. Didn't you use high pass filter before getting peaks?\n3. Why did you set height=[5, 100]? If I remember correctly peaks with amplitude above 20 are probably corona discharge (I don't remember where I read about it).\n4. What's public and private score of your single model?\n",
      "votes": 1,
      "replies": [
        {
          "id": 500512,
          "postDate": "2019-03-26T06:31:00.610Z",
          "content": "<ol>\n<li>yes</li>\n<li>DWT before getting peaks</li>\n<li>[5,100] is the range of peaks I extracted, amplitude above 10(manual selection, because it seems that I don't see the specific value in the paper) was used to remove the false peak</li>\n<li>public 0.61-0.68 private 0.63-0.7</li>\n</ol>",
          "rawMarkdown": "1. yes\n2. DWT before getting peaks\n3. [5,100] is the range of peaks I extracted, amplitude above 10(manual selection, because it seems that I don't see the specific value in the paper) was used to remove the false peak\n4. public 0.61-0.68 private 0.63-0.7",
          "votes": 2
        }
      ]
    },
    {
      "id": 499971,
      "postDate": "2019-03-25T12:56:14.750Z",
      "content": "<p>Thanks for sharing! And congratulations for 2nd place!</p>",
      "rawMarkdown": "Thanks for sharing! And congratulations for 2nd place!",
      "votes": 1
    },
    {
      "id": 909361,
      "postDate": "2020-06-30T14:36:45.620Z",
      "content": "<p>The Handmade features link provided above is down. Can anyone share the updated link if there is any?</p>",
      "rawMarkdown": "The Handmade features link provided above is down. Can anyone share the updated link if there is any?"
    },
    {
      "id": 501584,
      "postDate": "2019-03-27T13:31:40.827Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 500624,
      "postDate": "2019-03-26T10:07:24.657Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 500682,
          "postDate": "2019-03-26T12:17:31.550Z",
          "content": "<p>I found the model trained with 'val loss'(25 epochs) had a similar public score compared to the same model trained with 'val MCC'(40 epochs), but 'val_loss' gave me a better cv/lb relationship</p>",
          "rawMarkdown": "I found the model trained with 'val loss'(25 epochs) had a similar public score compared to the same model trained with 'val MCC'(40 epochs), but 'val_loss' gave me a better cv/lb relationship",
          "votes": 2
        }
      ]
    },
    {
      "id": 724859,
      "postDate": "2020-01-21T15:07:59.677Z",
      "content": "<p>Thank you for sharing! Congratulation!</p>",
      "rawMarkdown": "Thank you for sharing! Congratulation!"
    }
  ],
  "comments": [
    {
      "id": 501394,
      "author_name": "coderwangson",
      "author_url": "",
      "post_date": "2019-03-27T09:51:51.093000",
      "content": "<p>Congratulations</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 500281,
      "author_name": "Rafał Żyła",
      "author_url": "",
      "post_date": "2019-03-25T19:50:49.107000",
      "content": "<p>Congratulations! I also tried to put peak features into LSTM, but I got 0.541 in public - in private this model obtained 0.692 and of course I didn't choose it. ;) I have few questions. \n1. What RNN architecture did you use? Is it the same as from Bruno kernel? \n2. Didn't you use high pass filter before getting peaks?\n3. Why did you set height=[5, 100]? If I remember correctly peaks with amplitude above 20 are probably corona discharge (I don't remember where I read about it).\n4. What's public and private score of your single model?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 500512,
          "author_name": "HELLO TANG",
          "author_url": "",
          "post_date": "2019-03-26T06:31:00.610000",
          "content": "<ol>\n<li>yes</li>\n<li>DWT before getting peaks</li>\n<li>[5,100] is the range of peaks I extracted, amplitude above 10(manual selection, because it seems that I don't see the specific value in the paper) was used to remove the false peak</li>\n<li>public 0.61-0.68 private 0.63-0.7</li>\n</ol>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 499971,
      "author_name": "dhaqui the kaggler",
      "author_url": "",
      "post_date": "2019-03-25T12:56:14.750000",
      "content": "<p>Thanks for sharing! And congratulations for 2nd place!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 909361,
      "author_name": "Suhas Aithal",
      "author_url": "",
      "post_date": "2020-06-30T14:36:45.620000",
      "content": "<p>The Handmade features link provided above is down. Can anyone share the updated link if there is any?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 501584,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-27T13:31:40.827000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 500624,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-26T10:07:24.657000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 500682,
          "author_name": "HELLO TANG",
          "author_url": "",
          "post_date": "2019-03-26T12:17:31.550000",
          "content": "<p>I found the model trained with 'val loss'(25 epochs) had a similar public score compared to the same model trained with 'val MCC'(40 epochs), but 'val_loss' gave me a better cv/lb relationship</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 724859,
      "author_name": "JohnQ",
      "author_url": "",
      "post_date": "2020-01-21T15:07:59.677000",
      "content": "<p>Thank you for sharing! Congratulation!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "499909": "Congratulations to all the winners! \nAnd thanks a lot to VSB/Enet Center and Kaggle for this exciting competition.\n\n## Here are the details about my single model:\nPreprocessing:\nBefore extracting features, I used DWT for noise reduction, and I think this is helpful for the stability of the model.[you can find here](https://www.kaggle.com/jackvial/dwt-signal-denoising)\n\n### Features:\nI tried a lot of different methods, such as RNN taking 80, 100, 160, 250 time steps, basic features, energy features, peak features, and also tried to automatically extract features from CNN models, different RNN models. Unfortunately, none of these attempts have brought me a big improvement. Finally, the addition of the peak feature on the basic RNN architecture can effectively alleviate the over-fitting. This is the problem that I think is the most important problem to solve. Extracting the peak features using python's signal library, all this features can be found in [here](https://ieeexplore.ieee.org/abstract/document/7909221). Here is the code:\n\n```python\n# Extract peak features\nfrom scipy.signal import find_peaks, peak_widths, peak_prominences\n\ndef remove_false_peak(signal, p1, p2, maxDistance=10):\n    peak_diff = np.diff(p2)\n    if len(peak_diff) == 0:\n        return p1\n    ticks = []\n    for i, d in enumerate(peak_diff):\n        ratio = signal[p2[i+1]]/signal[p2[i]]\n        if d &lt; maxDistance and -0.25 &gt; ratio and ratio &gt; -4:\n            ticks.append((p2[i], p2[i+1]))\n    mask = np.array([True]*len(p1))\n    for i, j in ticks:\n        mask = mask &amp; ((p1 &lt; i) | (p1 &gt; 500+j))\n    return p1[mask]\n\n\ndef get_peaks(signal):\n    p1_1, _ = find_peaks(signal, height=[5, 100])\n    p1_2, _ = find_peaks(-signal, height=[5, 100])\n    p1 = np.union1d(p1_1, p1_2)\n    n_peaks, _ = find_peaks(-signal, height=[10, 100])\n    p_peaks, _ = find_peaks(signal, height=[10, 100])\n    p2 = np.union1d(n_peaks, p_peaks)\n    p = remove_false_peak(signal, p1, p2, maxDistance=10)\n    return np.intersect1d(p1_1, p), np.intersect1d(p1_2, p)\n\n\ndef extract_peak_feature(signal):\n    p_peaks, n_peaks = get_peaks(signal)\n\n    num_p, num_n = len(p_peaks), len(n_peaks)\n\n    sig_peak_width = np.concatenate(\n        [peak_widths(signal, p_peaks)[0], peak_widths(-signal, n_peaks)[0]])\n    sig_peak_height = abs(signal[np.concatenate([p_peaks, n_peaks])])\n\n    if num_n or num_p:\n        height_mean = sig_peak_height.mean()\n        height_max = sig_peak_height.max()\n        height_min = sig_peak_height.min()\n        height_median = np.median(sig_peak_height)\n\n        width_mean = sig_peak_width.mean()\n        width_max = sig_peak_width.max()\n        width_min = sig_peak_width.min()\n        width_median = np.median(sig_peak_width)\n\n        return np.array([height_mean, height_max, height_min, height_median,\n                         width_mean, width_max, width_min, width_median, num_p, num_n])\n    else:\n        return np.zeros(10)\n```\n\nAt the same time, the added peak feature has a large number of outliers, which I convert to a missing value. Then perform the missing value processing(dividing the data into groups based on the attribute with the largest correlation coefficient of the missing value, and then calculating the average value of each group. Just put these averages in the missing values.)\n\n### Training:\nepochs: 25\nCheckpoint monitor='val_loss'\n\n## My final solution was a ensemble of three models:\nMy single Model\n[VSB Competition : Stacked Attention Capsule BiLSTM](https://www.kaggle.com/tarunpaparaju/vsb-competition-attention-bilstm-with-features?scriptVersionId=10690570)\n[Handmade features](https://www.kaggleusercontent.com/kf/10818864/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..rLDkCCqNGYG5hfrj7oMt9A.ucfuA1j7MlivrTFzzAvVq7SpDSojFzTdXNVHqC5T7q0Vc4AjG5OP-2Pi0EngziSLz6FHxHoY4lqxvYj02gOtfMh9dGMJnCitLHWZ4JrZX10kzWvvLYhAmbtfm6Mk2ej46868zJzHFQ9RKnvcUjjBNQ.abKNk9CPW8feEC28o41osg/__results__.html)\n\nIf you think there is something incorrect or that could be improved, please leave your comments! And thank you everybody for the great kernels and discussions.(From 6th but it is exactly what I want to say)",
    "501394": "Congratulations",
    "500281": "Congratulations! I also tried to put peak features into LSTM, but I got 0.541 in public - in private this model obtained 0.692 and of course I didn't choose it. ;) I have few questions. \n1. What RNN architecture did you use? Is it the same as from Bruno kernel? \n2. Didn't you use high pass filter before getting peaks?\n3. Why did you set height=[5, 100]? If I remember correctly peaks with amplitude above 20 are probably corona discharge (I don't remember where I read about it).\n4. What's public and private score of your single model?\n",
    "499971": "Thanks for sharing! And congratulations for 2nd place!",
    "909361": "The Handmade features link provided above is down. Can anyone share the updated link if there is any?",
    "501584": "",
    "500624": "",
    "724859": "Thank you for sharing! Congratulation!"
  }
}