{
  "id": 85170,
  "title": "6th Place Solution Overview",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/85170",
  "author_name": "Vinicius",
  "post_date": "2019-03-22T03:21:48.024000",
  "votes": 61,
  "comment_count": 39,
  "views": 0,
  "content": "<p>Thanks to VSB/Enet Centre and Kaggle for this great competition. I had a lot of fun with that in the last months!</p>\n\n<p>I'm going to brief the main points I believe have helped in my final score. I'm not a Pro (yet!), so if you think there is something incorrect or that could be improved, please leave your comments!</p>\n\n<p><strong>My final solution was a ensemble of three main branchs:</strong></p>\n\n<ol>\n<li>Gradient Boosting Trees (Lightgbm)</li>\n<li>Recurrent Neural Network (GRU and LSTM)</li>\n<li>Convolutional Neural Network (Custom CNN, Resnet50, DenseNet101)</li>\n</ol>\n\n<p><strong>Preprocessing:</strong>\nThe 3 signal phases were used as just one sample. Each signal was aligned (using the 50Hz phase of the fourrier transform) to start where the signal crosses the axis from the negative to positive (𝜋/2) and the last quarter of the signal was removed.\nIn each model, I used a few denoised versions (varying thresholds) of this aligned/cropped signal:\n1. Wavelet\n    - Following the MaxHalford's repository: <a href=\"https://github.com/MaxHalford/kaggle-vsb-power/blob/master/scripts/extract_solo_features.py\">extract_solo_features.py</a>\n2. Fourrier / IQR\n    - Eliminate low frequencies with Fourrier transform and points with interquartile range</p>\n\n<p>I ended up with 4 signal version (Wavelet/Fourrier with 2 threshold levels)</p>\n\n<p>To RNN and CNN, I undersampled the 600000 size signal to 300000 taking the position with maximum absolute value at each pair of points:\n<code>und_signal = np.where(np.abs(signal[::2])&gt;np.abs(signal[1::2]), signal[::2], signal[1::2])</code></p>\n\n<p><strong>Inputs:</strong>\n1. GBT:\n    - The features were min, max, mean, std, skew and kurtosis of:\n        - Peaks count, height, width, prominences: <a href=\"https://github.com/MaxHalford/kaggle-vsb-power/blob/master/scripts/extract_solo_features.py\">extract_solo_features.py</a>\n        - Entropy &amp; Fractal: <a href=\"https://www.kaggle.com/tarunpaparaju/vsb-competition-attention-bilstm-with-features\">vsb-competition-attention-bilstm-with-features</a>\n        - Slope of lines connecting positive/negative peaks to the next negative/positive peaks\n        - Ratio of positive peaks to the minimum of the next maxDistance (defined in Vantuch's Thesis) points.\n        - Ratio of negative peaks to the maximum of the next maxDistance points.\n    - R² and weights of some polynomial regressions fit in positive/negative peaks</p>\n\n<ol>\n<li><p>RNN:</p>\n\n<ul><li>Preprocessed signal reshaped in 40 columns</li></ul></li>\n<li><p>CNN:</p>\n\n<ul><li>Preprocessed signal reshaped in 200 columns</li></ul></li>\n</ol>\n\n<p><strong>LB Prediction</strong>\nAt some point in the competition, I built a Lasso to predict the LB score of a submission. I knew that it would create problems with overfitting so I tried to avoid it with:\n- Threshold selection: I fixed all thresholds at the value 0.5\n- Pseudo-Labelling: I used it in all models and I reduced the sample weights of test examples witch had low weight in the trained Lasso\n- Ensemble: I should have ended up with 24 models (6 models x 4 preprocessed signals) but I got just 16 (lack of time...!)</p>\n\n<p>I was not able to create a stable stacking so i used a very elaborate strategy to ensemble the models: Take the arithmetic mean of the predictions...</p>\n\n<p>The last but not the least, I wanted to thank you everybody for the great kernels and discussions. I learned a lot and got my first gold medal!</p>\n\n<p>That's all folks!</p>",
  "messages": [
    {
      "id": 496307,
      "postDate": "2019-03-22T03:21:48.023Z",
      "content": "<p>Thanks to VSB/Enet Centre and Kaggle for this great competition. I had a lot of fun with that in the last months!</p>\n\n<p>I'm going to brief the main points I believe have helped in my final score. I'm not a Pro (yet!), so if you think there is something incorrect or that could be improved, please leave your comments!</p>\n\n<p><strong>My final solution was a ensemble of three main branchs:</strong></p>\n\n<ol>\n<li>Gradient Boosting Trees (Lightgbm)</li>\n<li>Recurrent Neural Network (GRU and LSTM)</li>\n<li>Convolutional Neural Network (Custom CNN, Resnet50, DenseNet101)</li>\n</ol>\n\n<p><strong>Preprocessing:</strong>\nThe 3 signal phases were used as just one sample. Each signal was aligned (using the 50Hz phase of the fourrier transform) to start where the signal crosses the axis from the negative to positive (𝜋/2) and the last quarter of the signal was removed.\nIn each model, I used a few denoised versions (varying thresholds) of this aligned/cropped signal:\n1. Wavelet\n    - Following the MaxHalford's repository: <a href=\"https://github.com/MaxHalford/kaggle-vsb-power/blob/master/scripts/extract_solo_features.py\">extract_solo_features.py</a>\n2. Fourrier / IQR\n    - Eliminate low frequencies with Fourrier transform and points with interquartile range</p>\n\n<p>I ended up with 4 signal version (Wavelet/Fourrier with 2 threshold levels)</p>\n\n<p>To RNN and CNN, I undersampled the 600000 size signal to 300000 taking the position with maximum absolute value at each pair of points:\n<code>und_signal = np.where(np.abs(signal[::2])&gt;np.abs(signal[1::2]), signal[::2], signal[1::2])</code></p>\n\n<p><strong>Inputs:</strong>\n1. GBT:\n    - The features were min, max, mean, std, skew and kurtosis of:\n        - Peaks count, height, width, prominences: <a href=\"https://github.com/MaxHalford/kaggle-vsb-power/blob/master/scripts/extract_solo_features.py\">extract_solo_features.py</a>\n        - Entropy &amp; Fractal: <a href=\"https://www.kaggle.com/tarunpaparaju/vsb-competition-attention-bilstm-with-features\">vsb-competition-attention-bilstm-with-features</a>\n        - Slope of lines connecting positive/negative peaks to the next negative/positive peaks\n        - Ratio of positive peaks to the minimum of the next maxDistance (defined in Vantuch's Thesis) points.\n        - Ratio of negative peaks to the maximum of the next maxDistance points.\n    - R² and weights of some polynomial regressions fit in positive/negative peaks</p>\n\n<ol>\n<li><p>RNN:</p>\n\n<ul><li>Preprocessed signal reshaped in 40 columns</li></ul></li>\n<li><p>CNN:</p>\n\n<ul><li>Preprocessed signal reshaped in 200 columns</li></ul></li>\n</ol>\n\n<p><strong>LB Prediction</strong>\nAt some point in the competition, I built a Lasso to predict the LB score of a submission. I knew that it would create problems with overfitting so I tried to avoid it with:\n- Threshold selection: I fixed all thresholds at the value 0.5\n- Pseudo-Labelling: I used it in all models and I reduced the sample weights of test examples witch had low weight in the trained Lasso\n- Ensemble: I should have ended up with 24 models (6 models x 4 preprocessed signals) but I got just 16 (lack of time...!)</p>\n\n<p>I was not able to create a stable stacking so i used a very elaborate strategy to ensemble the models: Take the arithmetic mean of the predictions...</p>\n\n<p>The last but not the least, I wanted to thank you everybody for the great kernels and discussions. I learned a lot and got my first gold medal!</p>\n\n<p>That's all folks!</p>",
      "rawMarkdown": "Thanks to VSB/Enet Centre and Kaggle for this great competition. I had a lot of fun with that in the last months!\n\nI'm going to brief the main points I believe have helped in my final score. I'm not a Pro (yet!), so if you think there is something incorrect or that could be improved, please leave your comments!\n\n**My final solution was a ensemble of three main branchs:**\n\n1. Gradient Boosting Trees (Lightgbm)\n2. Recurrent Neural Network (GRU and LSTM)\n3. Convolutional Neural Network (Custom CNN, Resnet50, DenseNet101)\n\n**Preprocessing:**\nThe 3 signal phases were used as just one sample. Each signal was aligned (using the 50Hz phase of the fourrier transform) to start where the signal crosses the axis from the negative to positive (𝜋/2) and the last quarter of the signal was removed.\nIn each model, I used a few denoised versions (varying thresholds) of this aligned/cropped signal:\n1. Wavelet\n\t- Following the MaxHalford's repository: [extract\\_solo\\_features.py](https://github.com/MaxHalford/kaggle-vsb-power/blob/master/scripts/extract_solo_features.py)\n2. Fourrier / IQR\n\t- Eliminate low frequencies with Fourrier transform and points with interquartile range\n\nI ended up with 4 signal version (Wavelet/Fourrier with 2 threshold levels)\n\nTo RNN and CNN, I undersampled the 600000 size signal to 300000 taking the position with maximum absolute value at each pair of points:\n```und_signal = np.where(np.abs(signal[::2])&gt;np.abs(signal[1::2]), signal[::2], signal[1::2])```\n\n\n**Inputs:**\n1. GBT:\n\t- The features were min, max, mean, std, skew and kurtosis of:\n\t\t- Peaks count, height, width, prominences: [extract\\_solo\\_features.py](https://github.com/MaxHalford/kaggle-vsb-power/blob/master/scripts/extract_solo_features.py)\n\t\t- Entropy &amp; Fractal: [vsb-competition-attention-bilstm-with-features](https://www.kaggle.com/tarunpaparaju/vsb-competition-attention-bilstm-with-features)\n\t\t- Slope of lines connecting positive/negative peaks to the next negative/positive peaks\n\t\t- Ratio of positive peaks to the minimum of the next maxDistance (defined in Vantuch's Thesis) points.\n\t\t- Ratio of negative peaks to the maximum of the next maxDistance points.\n\t- R² and weights of some polynomial regressions fit in positive/negative peaks\n\n2. RNN:\n\t- Preprocessed signal reshaped in 40 columns\n\n3. CNN:\n\t- Preprocessed signal reshaped in 200 columns\n\n**LB Prediction**\nAt some point in the competition, I built a Lasso to predict the LB score of a submission. I knew that it would create problems with overfitting so I tried to avoid it with:\n- Threshold selection: I fixed all thresholds at the value 0.5\n- Pseudo-Labelling: I used it in all models and I reduced the sample weights of test examples witch had low weight in the trained Lasso\n- Ensemble: I should have ended up with 24 models (6 models x 4 preprocessed signals) but I got just 16 (lack of time...!)\n\nI was not able to create a stable stacking so i used a very elaborate strategy to ensemble the models: Take the arithmetic mean of the predictions...\n\nThe last but not the least, I wanted to thank you everybody for the great kernels and discussions. I learned a lot and got my first gold medal!\n\nThat's all folks!\n",
      "votes": 60
    },
    {
      "id": 500367,
      "postDate": "2019-03-25T22:53:08.560Z",
      "content": "<p>Congratulations, great job!</p>",
      "rawMarkdown": "Congratulations, great job!",
      "votes": 1
    },
    {
      "id": 499824,
      "postDate": "2019-03-25T08:44:37.967Z",
      "content": "<p>Congratulations and thanks for sharing..</p>",
      "rawMarkdown": "Congratulations and thanks for sharing..",
      "votes": 1
    },
    {
      "id": 497772,
      "postDate": "2019-03-23T22:18:00.023Z",
      "content": "<p>Congrats on the gold medal and thank you very much for sharing the description of your solution!</p>",
      "rawMarkdown": "Congrats on the gold medal and thank you very much for sharing the description of your solution!",
      "votes": 1
    },
    {
      "id": 496792,
      "postDate": "2019-03-22T15:29:51.310Z",
      "content": "<p>Hi again <a href=\"/vhessel\">@vhessel</a>, \nI came back to understand your solution more. I am not sure that I understand on this point</p>\n\n<blockquote>\n  <p>RNN: Preprocessed signal reshaped in 40 columns\n  CNN: Preprocessed signal reshaped in 200 columns</p>\n</blockquote>\n\n<p>Before this line, you mentioned that you undersampled the signal data into 300,000.</p>\n\n<p>Do I understand correctly that the final shape that you feed into RNN is something like</p>\n\n<p>8712 x 40 x 7500  (40*7500 = 300,000), and then feed this into RNN?\n(so you did not group the 3 phases to predict together?)</p>",
      "rawMarkdown": "Hi again @vhessel, \nI came back to understand your solution more. I am not sure that I understand on this point\n\n&gt; RNN: Preprocessed signal reshaped in 40 columns\nCNN: Preprocessed signal reshaped in 200 columns\n\nBefore this line, you mentioned that you undersampled the signal data into 300,000.\n\nDo I understand correctly that the final shape that you feed into RNN is something like\n\n8712 x 40 x 7500  (40*7500 = 300,000), and then feed this into RNN?\n(so you did not group the 3 phases to predict together?)",
      "votes": 1,
      "replies": [
        {
          "id": 496800,
          "postDate": "2019-03-22T15:40:03.987Z",
          "content": "<p>Almost that...\nI reshaped the 300,000 points signal to 7500x40 and concatenate the 3 phases (axis=-1)\nSo the input of RNN was a 2904x7500x120 matrix and the input to CNN was a 2904x1500x200x3 matrix</p>",
          "rawMarkdown": "Almost that...\nI reshaped the 300,000 points signal to 7500x40 and concatenate the 3 phases (axis=-1)\nSo the input of RNN was a 2904x7500x120 matrix and the input to CNN was a 2904x1500x200x3 matrix",
          "votes": 1
        },
        {
          "id": 496804,
          "postDate": "2019-03-22T15:44:51.953Z",
          "content": "<p>Thanks so much!! never imagine that RNN would be able to handle such a long-range 7500 time-step data! This is so valuable to know its potential.</p>\n\n<p>May I ask a couple more questions? \n- How can you decide to reshape to 40/200 in cases of RNN/CNN (why different numbers?)\n- How long does it take to train  2904x7500x120 ? My biggest processed data is 2904 x 640 x 69 and it took a while to finish.</p>",
          "rawMarkdown": "Thanks so much!! never imagine that RNN would be able to handle such a long-range 7500 time-step data! This is so valuable to know its potential.\n\nMay I ask a couple more questions? \n- How can you decide to reshape to 40/200 in cases of RNN/CNN (why different numbers?)\n- How long does it take to train  2904x7500x120 ? My biggest processed data is 2904 x 640 x 69 and it took a while to finish.",
          "votes": 1
        },
        {
          "id": 496812,
          "postDate": "2019-03-22T15:58:19.220Z",
          "content": "<p>I'm not a RNN expert, so I don't know if it is a too much long sequence to that kind of model but the data was a bit sparse (after denoising), maybe it has helped.\n- The rationale of CNN was easier: a kernel matrix multiplying 2 subsequent lines is using information of 2 points, in the signal, separated by 200 steps (lag-200). So, the number of columns control the distance of the information used in the convolution (at least in the first layer)\nThe RNN decision was a bit arbitrary. I chose to divide each 𝜋/2 arc of the signal in 10 parts. I was using 10000x40 before crop the signal\n- The full RNN training was taking about 3-4 hours: 5-Folds x 20 Epochs x 2 minutes. I used a GTX1080TI in this competition.</p>",
          "rawMarkdown": "I'm not a RNN expert, so I don't know if it is a too much long sequence to that kind of model but the data was a bit sparse (after denoising), maybe it has helped.\n- The rationale of CNN was easier: a kernel matrix multiplying 2 subsequent lines is using information of 2 points, in the signal, separated by 200 steps (lag-200). So, the number of columns control the distance of the information used in the convolution (at least in the first layer)\nThe RNN decision was a bit arbitrary. I chose to divide each 𝜋/2 arc of the signal in 10 parts. I was using 10000x40 before crop the signal\n- The full RNN training was taking about 3-4 hours: 5-Folds x 20 Epochs x 2 minutes. I used a GTX1080TI in this competition.",
          "votes": 1
        },
        {
          "id": 496816,
          "postDate": "2019-03-22T16:07:30Z",
          "content": "<p>Thank you so much again! I will try to emulate your RNN/CNN approach next week, and hope to get a good score too ;)</p>",
          "rawMarkdown": "Thank you so much again! I will try to emulate your RNN/CNN approach next week, and hope to get a good score too ;)",
          "votes": 1
        }
      ]
    },
    {
      "id": 496588,
      "postDate": "2019-03-22T11:19:05.980Z",
      "content": "<p>Congratulations, great work ! Glad my kernel came in handy for you :)</p>",
      "rawMarkdown": "Congratulations, great work ! Glad my kernel came in handy for you :)",
      "votes": 1,
      "replies": [
        {
          "id": 496594,
          "postDate": "2019-03-22T11:26:25.223Z",
          "content": "<p>Yeah! Thanks for the kernel, <a href=\"/tarunpaparaju\">@tarunpaparaju</a> !</p>",
          "rawMarkdown": "Yeah! Thanks for the kernel, @tarunpaparaju !",
          "votes": 1
        },
        {
          "id": 496597,
          "postDate": "2019-03-22T11:27:37.587Z",
          "content": "<p>My pleasure :)</p>",
          "rawMarkdown": "My pleasure :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 496456,
      "postDate": "2019-03-22T08:22:41.340Z",
      "content": "<p>good job!</p>",
      "rawMarkdown": "good job!",
      "votes": 1
    },
    {
      "id": 496452,
      "postDate": "2019-03-22T08:20:17.030Z",
      "content": "<p>Congrats and thanks for sharing. I like the diversity in your selection(Gradient Boosting, RNNs, and CNNs).</p>",
      "rawMarkdown": "Congrats and thanks for sharing. I like the diversity in your selection(Gradient Boosting, RNNs, and CNNs).",
      "votes": 1
    },
    {
      "id": 496435,
      "postDate": "2019-03-22T07:46:16.897Z",
      "content": "<p>Congratulations! Nice work!</p>",
      "rawMarkdown": "Congratulations! Nice work!",
      "votes": 1
    },
    {
      "id": 496425,
      "postDate": "2019-03-22T07:22:56.037Z",
      "content": "<p>Congratz! Any idea on how your 3 models perform without ensembling ? I'm curious!</p>",
      "rawMarkdown": "Congratz! Any idea on how your 3 models perform without ensembling ? I'm curious!",
      "votes": 1,
      "replies": [
        {
          "id": 496592,
          "postDate": "2019-03-22T11:25:15.160Z",
          "content": "<p>Thanks! With the limitation of 2 submits per day, i didn't tested the results of my final single models. On local CV I've got RNN &gt; GBT &gt; CNN. I will submit the results and tell you later</p>",
          "rawMarkdown": "Thanks! With the limitation of 2 submits per day, i didn't tested the results of my final single models. On local CV I've got RNN &gt; GBT &gt; CNN. I will submit the results and tell you later",
          "votes": 2
        }
      ]
    },
    {
      "id": 496342,
      "postDate": "2019-03-22T04:10:24.023Z",
      "content": "<p>Congrats <a href=\"/vhessel\">@vhessel</a> on your 1st gold medal and thanks for sharing you solution overview.</p>\n\n<p>On the CV strategy, I have a <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85143\">discussion post here</a> if you want to add your thoughts.</p>",
      "rawMarkdown": "Congrats @vhessel on your 1st gold medal and thanks for sharing you solution overview.\n\nOn the CV strategy, I have a [discussion post here](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85143) if you want to add your thoughts.",
      "votes": 1
    },
    {
      "id": 496316,
      "postDate": "2019-03-22T03:32:53.237Z",
      "content": "<p>You did a superb job here, congratulation!!</p>\n\n<p>Some questions:\n- In the final solution, you aligned &amp; cropped the signal, but did you try unaligned/uncropped signal ? Do they give very different results?\n- And how much did the denoising help? Unfortunately for me, I haven’t tried this at all! (I thought of the DL principle of giving it raw data :p )\n- How could you psuedo label data?</p>\n\n<p>Lastly, This is the first time I heard of LB prediction ... truly nice to hear! But I could still not understand how it work, does it relate to your pseudo labelling? (is that just linear regression lasso from 20337 example to [0,1] ?? )</p>\n\n<p>Congratulation again and sorry to ask many questions !</p>",
      "rawMarkdown": "You did a superb job here, congratulation!!\n\nSome questions:\n- In the final solution, you aligned &amp; cropped the signal, but did you try unaligned/uncropped signal ? Do they give very different results?\n- And how much did the denoising help? Unfortunately for me, I haven’t tried this at all! (I thought of the DL principle of giving it raw data :p )\n- How could you psuedo label data?\n\n\nLastly, This is the first time I heard of LB prediction ... truly nice to hear! But I could still not understand how it work, does it relate to your pseudo labelling? (is that just linear regression lasso from 20337 example to [0,1] ?? )\n\nCongratulation again and sorry to ask many questions !\n",
      "votes": 1,
      "replies": [
        {
          "id": 496333,
          "postDate": "2019-03-22T03:56:39.753Z",
          "content": "<p>Hi, <a href=\"/ratthachat\">@ratthachat</a>. Thank you and no problem about the questions!!!\n- I used the whole signal, aligned by the phase 0 signal almost until the end. Then, I noticed that cropping and aligning each signal gave me a boost on local CV to the GBT model. Since I did not have enought time to fine tuning and experimenting with each model, I changed all of them.\n- For the CNN models, I was not able to get even decent results without denoising. I started using raw data with RNN, but after the experiments with CNN, I changed all my functions of preprocessing.\n- I used the ensemble of predictions as targets to the test set. Then I trained the models with the train+test data, with increasing sample weights for the test data, and a proportion of 50/50 in the training batches.</p>\n\n<p>To be honest, it was the first time I tried LB prediction. I heard that from Heng CherKeng: <a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/63984\">common tricks in kaggle image segmentation problems</a>. What I did was to train a Lasso using my submissions to predict (regression problem) the Public Score. So, input was a submission and output was the public LB</p>",
          "rawMarkdown": "Hi, @ratthachat. Thank you and no problem about the questions!!!\n- I used the whole signal, aligned by the phase 0 signal almost until the end. Then, I noticed that cropping and aligning each signal gave me a boost on local CV to the GBT model. Since I did not have enought time to fine tuning and experimenting with each model, I changed all of them.\n- For the CNN models, I was not able to get even decent results without denoising. I started using raw data with RNN, but after the experiments with CNN, I changed all my functions of preprocessing.\n- I used the ensemble of predictions as targets to the test set. Then I trained the models with the train+test data, with increasing sample weights for the test data, and a proportion of 50/50 in the training batches.\n\nTo be honest, it was the first time I tried LB prediction. I heard that from Heng CherKeng: [common tricks in kaggle image segmentation problems](https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/63984). What I did was to train a Lasso using my submissions to predict (regression problem) the Public Score. So, input was a submission and output was the public LB",
          "votes": 4
        },
        {
          "id": 496341,
          "postDate": "2019-03-22T04:09:23.667Z",
          "content": "<p>Thanks <a href=\"/vhessel\">@vhessel</a> for your explanation!  Could it be possible to use the lasso LB prediction instead of CV ? (or we have to use them together to make sense?)</p>\n\n<p>May I ask one more question : Based on your CV strategy mentioned below with Russ, what is your best single model CV ? (and how does it result on LB?, i.e. does your CV / LB correlate well?)</p>",
          "rawMarkdown": "Thanks @vhessel for your explanation!  Could it be possible to use the lasso LB prediction instead of CV ? (or we have to use them together to make sense?)\n\nMay I ask one more question : Based on your CV strategy mentioned below with Russ, what is your best single model CV ? (and how does it result on LB?, i.e. does your CV / LB correlate well?)",
          "votes": 1
        },
        {
          "id": 496356,
          "postDate": "2019-03-22T04:21:34.713Z",
          "content": "<p>I think the problem to use the LB prediction is overfitting. The model can just rely on your past submissions (it's good to have a number of low correlated models), limited data (number of submits) and the public samples. I used it in conjunction with my local CV. For example, my best solution got:\n- Local CV ranging from 0.69x to : 0.73x (models of the final ensemble)\n- LB Prediction: 0.75x\n- Public LB: 0.77x\n- Private LB: 0.69x</p>",
          "rawMarkdown": "I think the problem to use the LB prediction is overfitting. The model can just rely on your past submissions (it's good to have a number of low correlated models), limited data (number of submits) and the public samples. I used it in conjunction with my local CV. For example, my best solution got:\n- Local CV ranging from 0.69x to : 0.73x (models of the final ensemble)\n- LB Prediction: 0.75x\n- Public LB: 0.77x\n- Private LB: 0.69x",
          "votes": 1
        },
        {
          "id": 496362,
          "postDate": "2019-03-22T04:33:25.570Z",
          "content": "<p>Many thanks and noted. Hope to apply the insights to the next challenging competition! </p>",
          "rawMarkdown": "Many thanks and noted. Hope to apply the insights to the next challenging competition! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 496314,
      "postDate": "2019-03-22T03:31:58.837Z",
      "content": "<p>Congratulations and great work Vinicius!   What kind of local cross validation scheme did you use?</p>",
      "rawMarkdown": "Congratulations and great work Vinicius!   What kind of local cross validation scheme did you use?",
      "votes": 1,
      "replies": [
        {
          "id": 496321,
          "postDate": "2019-03-22T03:40:16.253Z",
          "content": "<p>Many thanks, Russ. I tried some multilabel crossvalidation schemes based on clustering, number of peaks, statistics from fft but at the end I did a stratified classification on targets with the validation set using 50% data from training examples and 50% from pseudo labeled</p>",
          "rawMarkdown": "Many thanks, Russ. I tried some multilabel crossvalidation schemes based on clustering, number of peaks, statistics from fft but at the end I did a stratified classification on targets with the validation set using 50% data from training examples and 50% from pseudo labeled",
          "votes": 3
        }
      ]
    },
    {
      "id": 497530,
      "postDate": "2019-03-23T16:37:38.860Z",
      "content": "<p>good job!</p>",
      "rawMarkdown": "good job!",
      "votes": 2
    },
    {
      "id": 496405,
      "postDate": "2019-03-22T06:33:31.560Z",
      "content": "<p>Congratulations on the strong finish ! Your ensemble is amazing by itself.</p>\n\n<p>Can I ask what was the reasoning behind cutting the last quarter of the signal ? And why 2/1 downsampling by just choosing the maximum ?</p>",
      "rawMarkdown": "Congratulations on the strong finish ! Your ensemble is amazing by itself.\n\nCan I ask what was the reasoning behind cutting the last quarter of the signal ? And why 2/1 downsampling by just choosing the maximum ?",
      "votes": 2,
      "replies": [
        {
          "id": 496589,
          "postDate": "2019-03-22T11:21:34.067Z",
          "content": "<p>Hi <a href=\"/nyleve\">@nyleve</a>. Tomas Vantuch commented about the regions where PD patterns tend to occur: <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/75771#460739\">Problem description</a> <br>\nThe downsampling was just because the models were taking too long to be trained. I was trying a lot of configurations and wanted use the original signal at the end but...\nI chose the maximum trying to reduce the resample effect in the peaks</p>",
          "rawMarkdown": "Hi @nyleve. Tomas Vantuch commented about the regions where PD patterns tend to occur: [Problem description](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/75771#460739)  \nThe downsampling was just because the models were taking too long to be trained. I was trying a lot of configurations and wanted use the original signal at the end but...\nI chose the maximum trying to reduce the resample effect in the peaks",
          "votes": 2
        }
      ]
    },
    {
      "id": 496365,
      "postDate": "2019-03-22T04:41:14.333Z",
      "content": "<blockquote>\n  <p>I was not able to create a stable stacking so i used a very elaborate strategy to ensemble the models: Take the arithmetic mean of the predictions…</p>\n</blockquote>\n\n<p>Interesting. What do you mean by <em>arithmetic mean</em>. Is that a mean of the probability, or a mean of yes/no?</p>\n\n<p>Our team tried both, and it seems a voting (mean of yes/no) works better than mean of probability. </p>\n\n<p>Thanks</p>",
      "rawMarkdown": "&gt;I was not able to create a stable stacking so i used a very elaborate strategy to ensemble the models: Take the arithmetic mean of the predictions…\n\nInteresting. What do you mean by *arithmetic mean*. Is that a mean of the probability, or a mean of yes/no?\n\nOur team tried both, and it seems a voting (mean of yes/no) works better than mean of probability. \n\nThanks",
      "votes": 2,
      "replies": [
        {
          "id": 496368,
          "postDate": "2019-03-22T04:45:20.953Z",
          "content": "<p>Hi, <a href=\"/huyunwei\">@huyunwei</a> . I took the mean of probability and then rounded the values to get the target. I tried voting ensemble just in the beginning... I should have tried at end.</p>",
          "rawMarkdown": "Hi, @huyunwei . I took the mean of probability and then rounded the values to get the target. I tried voting ensemble just in the beginning... I should have tried at end."
        },
        {
          "id": 496418,
          "postDate": "2019-03-22T07:17:50.207Z",
          "content": "<p>Congrats man. You are one of the very few who kept and improved the LB position after the competition ended. You fully deserve your medal. a lot of lessons can be learned based on your shared experience. thanks for sharing !</p>\n\n<p>However, I have a question.</p>\n\n<p>Can you explain how did you choose which techniques to use inside the ensemble ? \nAt some point I was considering to make an example RNN + other</p>\n\n<p>Also, did you experiment an ensemble only with GBT + RNN, and if you did, do you have a comparison with current ensemble that uses also CNN ?</p>\n\n<p>Thanks !</p>",
          "rawMarkdown": "Congrats man. You are one of the very few who kept and improved the LB position after the competition ended. You fully deserve your medal. a lot of lessons can be learned based on your shared experience. thanks for sharing !\n\nHowever, I have a question.\n\nCan you explain how did you choose which techniques to use inside the ensemble ? \nAt some point I was considering to make an example RNN + other\n\nAlso, did you experiment an ensemble only with GBT + RNN, and if you did, do you have a comparison with current ensemble that uses also CNN ?\n\nThanks !",
          "votes": 2
        },
        {
          "id": 496583,
          "postDate": "2019-03-22T11:10:32.430Z",
          "content": "<p>Thanks, <a href=\"/nicupetridean\">@nicupetridean</a>.\nI started the competition using GBT and changed to RNN at the time when the 0.694 public kernel was released. I also tried some few experiments with CNN but i didn't get good results because, I guess, was using raw data (have to confirm it) and just back to try this model in the beginning of march.\nI was not using pseudo-labelling when experimenting with GBT+RNN so the ensemble was instable. I will submit some single model and partial ensemble predictions and update here de results.</p>",
          "rawMarkdown": "Thanks, @nicupetridean.\nI started the competition using GBT and changed to RNN at the time when the 0.694 public kernel was released. I also tried some few experiments with CNN but i didn't get good results because, I guess, was using raw data (have to confirm it) and just back to try this model in the beginning of march.\nI was not using pseudo-labelling when experimenting with GBT+RNN so the ensemble was instable. I will submit some single model and partial ensemble predictions and update here de results.",
          "votes": 2
        },
        {
          "id": 496641,
          "postDate": "2019-03-22T12:25:03.683Z",
          "content": "<p>Thanks for the details ! Looking forward to see the results of partial ensemble predictions !</p>",
          "rawMarkdown": "Thanks for the details ! Looking forward to see the results of partial ensemble predictions !",
          "votes": 2
        },
        {
          "id": 496689,
          "postDate": "2019-03-22T13:24:18.553Z",
          "content": "<p>I have a public/private LB of 0.776/0.686 to the GBT+RNN ensemble against 0.771/0.696 in the final one</p>",
          "rawMarkdown": "I have a public/private LB of 0.776/0.686 to the GBT+RNN ensemble against 0.771/0.696 in the final one",
          "votes": 3
        },
        {
          "id": 496858,
          "postDate": "2019-03-22T17:09:25.650Z",
          "content": "<p>Close enough. Nice.</p>",
          "rawMarkdown": "Close enough. Nice."
        }
      ]
    },
    {
      "id": 499628,
      "postDate": "2019-03-25T02:47:28.423Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 496602,
      "postDate": "2019-03-22T11:35:41.593Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 499765,
      "postDate": "2019-03-25T06:52:14.233Z",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 497178,
      "postDate": "2019-03-23T04:14:14.440Z",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!",
      "votes": 2
    },
    {
      "id": 496352,
      "postDate": "2019-03-22T04:16:55.360Z",
      "content": "<p>Congratulations and Thanks for sharing. </p>",
      "rawMarkdown": "Congratulations and Thanks for sharing. ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 500367,
      "author_name": "Roman Ilechko",
      "author_url": "",
      "post_date": "2019-03-25T22:53:08.560000",
      "content": "<p>Congratulations, great job!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 499824,
      "author_name": "Puneet Mehta",
      "author_url": "",
      "post_date": "2019-03-25T08:44:37.967000",
      "content": "<p>Congratulations and thanks for sharing..</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 497772,
      "author_name": "WholeGrain",
      "author_url": "",
      "post_date": "2019-03-23T22:18:00.023000",
      "content": "<p>Congrats on the gold medal and thank you very much for sharing the description of your solution!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 496792,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-03-22T15:29:51.310000",
      "content": "<p>Hi again <a href=\"/vhessel\">@vhessel</a>, \nI came back to understand your solution more. I am not sure that I understand on this point</p>\n\n<blockquote>\n  <p>RNN: Preprocessed signal reshaped in 40 columns\n  CNN: Preprocessed signal reshaped in 200 columns</p>\n</blockquote>\n\n<p>Before this line, you mentioned that you undersampled the signal data into 300,000.</p>\n\n<p>Do I understand correctly that the final shape that you feed into RNN is something like</p>\n\n<p>8712 x 40 x 7500  (40*7500 = 300,000), and then feed this into RNN?\n(so you did not group the 3 phases to predict together?)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 496800,
          "author_name": "Vinicius",
          "author_url": "",
          "post_date": "2019-03-22T15:40:03.987000",
          "content": "<p>Almost that...\nI reshaped the 300,000 points signal to 7500x40 and concatenate the 3 phases (axis=-1)\nSo the input of RNN was a 2904x7500x120 matrix and the input to CNN was a 2904x1500x200x3 matrix</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 496804,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-03-22T15:44:51.953000",
          "content": "<p>Thanks so much!! never imagine that RNN would be able to handle such a long-range 7500 time-step data! This is so valuable to know its potential.</p>\n\n<p>May I ask a couple more questions? \n- How can you decide to reshape to 40/200 in cases of RNN/CNN (why different numbers?)\n- How long does it take to train  2904x7500x120 ? My biggest processed data is 2904 x 640 x 69 and it took a while to finish.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 496812,
          "author_name": "Vinicius",
          "author_url": "",
          "post_date": "2019-03-22T15:58:19.220000",
          "content": "<p>I'm not a RNN expert, so I don't know if it is a too much long sequence to that kind of model but the data was a bit sparse (after denoising), maybe it has helped.\n- The rationale of CNN was easier: a kernel matrix multiplying 2 subsequent lines is using information of 2 points, in the signal, separated by 200 steps (lag-200). So, the number of columns control the distance of the information used in the convolution (at least in the first layer)\nThe RNN decision was a bit arbitrary. I chose to divide each 𝜋/2 arc of the signal in 10 parts. I was using 10000x40 before crop the signal\n- The full RNN training was taking about 3-4 hours: 5-Folds x 20 Epochs x 2 minutes. I used a GTX1080TI in this competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 496816,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-03-22T16:07:30",
          "content": "<p>Thank you so much again! I will try to emulate your RNN/CNN approach next week, and hope to get a good score too ;)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 496588,
      "author_name": "Tarun Paparaju",
      "author_url": "",
      "post_date": "2019-03-22T11:19:05.980000",
      "content": "<p>Congratulations, great work ! Glad my kernel came in handy for you :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 496594,
          "author_name": "Vinicius",
          "author_url": "",
          "post_date": "2019-03-22T11:26:25.223000",
          "content": "<p>Yeah! Thanks for the kernel, <a href=\"/tarunpaparaju\">@tarunpaparaju</a> !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 496597,
          "author_name": "Tarun Paparaju",
          "author_url": "",
          "post_date": "2019-03-22T11:27:37.587000",
          "content": "<p>My pleasure :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 496456,
      "author_name": "Stanislav Blinov",
      "author_url": "",
      "post_date": "2019-03-22T08:22:41.340000",
      "content": "<p>good job!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 496452,
      "author_name": "jamie_omoya",
      "author_url": "",
      "post_date": "2019-03-22T08:20:17.030000",
      "content": "<p>Congrats and thanks for sharing. I like the diversity in your selection(Gradient Boosting, RNNs, and CNNs).</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 496435,
      "author_name": "Anton ",
      "author_url": "",
      "post_date": "2019-03-22T07:46:16.897000",
      "content": "<p>Congratulations! Nice work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 496425,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2019-03-22T07:22:56.037000",
      "content": "<p>Congratz! Any idea on how your 3 models perform without ensembling ? I'm curious!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 496592,
          "author_name": "Vinicius",
          "author_url": "",
          "post_date": "2019-03-22T11:25:15.160000",
          "content": "<p>Thanks! With the limitation of 2 submits per day, i didn't tested the results of my final single models. On local CV I've got RNN &gt; GBT &gt; CNN. I will submit the results and tell you later</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 496342,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-03-22T04:10:24.023000",
      "content": "<p>Congrats <a href=\"/vhessel\">@vhessel</a> on your 1st gold medal and thanks for sharing you solution overview.</p>\n\n<p>On the CV strategy, I have a <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85143\">discussion post here</a> if you want to add your thoughts.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 496316,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-03-22T03:32:53.237000",
      "content": "<p>You did a superb job here, congratulation!!</p>\n\n<p>Some questions:\n- In the final solution, you aligned &amp; cropped the signal, but did you try unaligned/uncropped signal ? Do they give very different results?\n- And how much did the denoising help? Unfortunately for me, I haven’t tried this at all! (I thought of the DL principle of giving it raw data :p )\n- How could you psuedo label data?</p>\n\n<p>Lastly, This is the first time I heard of LB prediction ... truly nice to hear! But I could still not understand how it work, does it relate to your pseudo labelling? (is that just linear regression lasso from 20337 example to [0,1] ?? )</p>\n\n<p>Congratulation again and sorry to ask many questions !</p>",
      "votes": 1,
      "replies": [
        {
          "id": 496333,
          "author_name": "Vinicius",
          "author_url": "",
          "post_date": "2019-03-22T03:56:39.753000",
          "content": "<p>Hi, <a href=\"/ratthachat\">@ratthachat</a>. Thank you and no problem about the questions!!!\n- I used the whole signal, aligned by the phase 0 signal almost until the end. Then, I noticed that cropping and aligning each signal gave me a boost on local CV to the GBT model. Since I did not have enought time to fine tuning and experimenting with each model, I changed all of them.\n- For the CNN models, I was not able to get even decent results without denoising. I started using raw data with RNN, but after the experiments with CNN, I changed all my functions of preprocessing.\n- I used the ensemble of predictions as targets to the test set. Then I trained the models with the train+test data, with increasing sample weights for the test data, and a proportion of 50/50 in the training batches.</p>\n\n<p>To be honest, it was the first time I tried LB prediction. I heard that from Heng CherKeng: <a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/63984\">common tricks in kaggle image segmentation problems</a>. What I did was to train a Lasso using my submissions to predict (regression problem) the Public Score. So, input was a submission and output was the public LB</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 496341,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-03-22T04:09:23.667000",
          "content": "<p>Thanks <a href=\"/vhessel\">@vhessel</a> for your explanation!  Could it be possible to use the lasso LB prediction instead of CV ? (or we have to use them together to make sense?)</p>\n\n<p>May I ask one more question : Based on your CV strategy mentioned below with Russ, what is your best single model CV ? (and how does it result on LB?, i.e. does your CV / LB correlate well?)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 496356,
          "author_name": "Vinicius",
          "author_url": "",
          "post_date": "2019-03-22T04:21:34.713000",
          "content": "<p>I think the problem to use the LB prediction is overfitting. The model can just rely on your past submissions (it's good to have a number of low correlated models), limited data (number of submits) and the public samples. I used it in conjunction with my local CV. For example, my best solution got:\n- Local CV ranging from 0.69x to : 0.73x (models of the final ensemble)\n- LB Prediction: 0.75x\n- Public LB: 0.77x\n- Private LB: 0.69x</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 496362,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-03-22T04:33:25.570000",
          "content": "<p>Many thanks and noted. Hope to apply the insights to the next challenging competition! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 496314,
      "author_name": "Russ Wolfinger",
      "author_url": "",
      "post_date": "2019-03-22T03:31:58.837000",
      "content": "<p>Congratulations and great work Vinicius!   What kind of local cross validation scheme did you use?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 496321,
          "author_name": "Vinicius",
          "author_url": "",
          "post_date": "2019-03-22T03:40:16.253000",
          "content": "<p>Many thanks, Russ. I tried some multilabel crossvalidation schemes based on clustering, number of peaks, statistics from fft but at the end I did a stratified classification on targets with the validation set using 50% data from training examples and 50% from pseudo labeled</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 497530,
      "author_name": "Andrii Bohachuk",
      "author_url": "",
      "post_date": "2019-03-23T16:37:38.860000",
      "content": "<p>good job!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 496405,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2019-03-22T06:33:31.560000",
      "content": "<p>Congratulations on the strong finish ! Your ensemble is amazing by itself.</p>\n\n<p>Can I ask what was the reasoning behind cutting the last quarter of the signal ? And why 2/1 downsampling by just choosing the maximum ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 496589,
          "author_name": "Vinicius",
          "author_url": "",
          "post_date": "2019-03-22T11:21:34.067000",
          "content": "<p>Hi <a href=\"/nyleve\">@nyleve</a>. Tomas Vantuch commented about the regions where PD patterns tend to occur: <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/75771#460739\">Problem description</a> <br>\nThe downsampling was just because the models were taking too long to be trained. I was trying a lot of configurations and wanted use the original signal at the end but...\nI chose the maximum trying to reduce the resample effect in the peaks</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 496365,
      "author_name": "Yunwei Hu",
      "author_url": "",
      "post_date": "2019-03-22T04:41:14.333000",
      "content": "<blockquote>\n  <p>I was not able to create a stable stacking so i used a very elaborate strategy to ensemble the models: Take the arithmetic mean of the predictions…</p>\n</blockquote>\n\n<p>Interesting. What do you mean by <em>arithmetic mean</em>. Is that a mean of the probability, or a mean of yes/no?</p>\n\n<p>Our team tried both, and it seems a voting (mean of yes/no) works better than mean of probability. </p>\n\n<p>Thanks</p>",
      "votes": 2,
      "replies": [
        {
          "id": 496368,
          "author_name": "Vinicius",
          "author_url": "",
          "post_date": "2019-03-22T04:45:20.953000",
          "content": "<p>Hi, <a href=\"/huyunwei\">@huyunwei</a> . I took the mean of probability and then rounded the values to get the target. I tried voting ensemble just in the beginning... I should have tried at end.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 496418,
          "author_name": "Nicu",
          "author_url": "",
          "post_date": "2019-03-22T07:17:50.207000",
          "content": "<p>Congrats man. You are one of the very few who kept and improved the LB position after the competition ended. You fully deserve your medal. a lot of lessons can be learned based on your shared experience. thanks for sharing !</p>\n\n<p>However, I have a question.</p>\n\n<p>Can you explain how did you choose which techniques to use inside the ensemble ? \nAt some point I was considering to make an example RNN + other</p>\n\n<p>Also, did you experiment an ensemble only with GBT + RNN, and if you did, do you have a comparison with current ensemble that uses also CNN ?</p>\n\n<p>Thanks !</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 496583,
          "author_name": "Vinicius",
          "author_url": "",
          "post_date": "2019-03-22T11:10:32.430000",
          "content": "<p>Thanks, <a href=\"/nicupetridean\">@nicupetridean</a>.\nI started the competition using GBT and changed to RNN at the time when the 0.694 public kernel was released. I also tried some few experiments with CNN but i didn't get good results because, I guess, was using raw data (have to confirm it) and just back to try this model in the beginning of march.\nI was not using pseudo-labelling when experimenting with GBT+RNN so the ensemble was instable. I will submit some single model and partial ensemble predictions and update here de results.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 496641,
          "author_name": "Nicu",
          "author_url": "",
          "post_date": "2019-03-22T12:25:03.683000",
          "content": "<p>Thanks for the details ! Looking forward to see the results of partial ensemble predictions !</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 496689,
          "author_name": "Vinicius",
          "author_url": "",
          "post_date": "2019-03-22T13:24:18.553000",
          "content": "<p>I have a public/private LB of 0.776/0.686 to the GBT+RNN ensemble against 0.771/0.696 in the final one</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 496858,
          "author_name": "Nicu",
          "author_url": "",
          "post_date": "2019-03-22T17:09:25.650000",
          "content": "<p>Close enough. Nice.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 499628,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-25T02:47:28.423000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 496602,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-22T11:35:41.593000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 499765,
      "author_name": "Kim",
      "author_url": "",
      "post_date": "2019-03-25T06:52:14.233000",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 497178,
      "author_name": "David J. Slate",
      "author_url": "",
      "post_date": "2019-03-23T04:14:14.440000",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 496352,
      "author_name": "huiqin",
      "author_url": "",
      "post_date": "2019-03-22T04:16:55.360000",
      "content": "<p>Congratulations and Thanks for sharing. </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "496307": "Thanks to VSB/Enet Centre and Kaggle for this great competition. I had a lot of fun with that in the last months!\n\nI'm going to brief the main points I believe have helped in my final score. I'm not a Pro (yet!), so if you think there is something incorrect or that could be improved, please leave your comments!\n\n**My final solution was a ensemble of three main branchs:**\n\n1. Gradient Boosting Trees (Lightgbm)\n2. Recurrent Neural Network (GRU and LSTM)\n3. Convolutional Neural Network (Custom CNN, Resnet50, DenseNet101)\n\n**Preprocessing:**\nThe 3 signal phases were used as just one sample. Each signal was aligned (using the 50Hz phase of the fourrier transform) to start where the signal crosses the axis from the negative to positive (𝜋/2) and the last quarter of the signal was removed.\nIn each model, I used a few denoised versions (varying thresholds) of this aligned/cropped signal:\n1. Wavelet\n\t- Following the MaxHalford's repository: [extract\\_solo\\_features.py](https://github.com/MaxHalford/kaggle-vsb-power/blob/master/scripts/extract_solo_features.py)\n2. Fourrier / IQR\n\t- Eliminate low frequencies with Fourrier transform and points with interquartile range\n\nI ended up with 4 signal version (Wavelet/Fourrier with 2 threshold levels)\n\nTo RNN and CNN, I undersampled the 600000 size signal to 300000 taking the position with maximum absolute value at each pair of points:\n```und_signal = np.where(np.abs(signal[::2])&gt;np.abs(signal[1::2]), signal[::2], signal[1::2])```\n\n\n**Inputs:**\n1. GBT:\n\t- The features were min, max, mean, std, skew and kurtosis of:\n\t\t- Peaks count, height, width, prominences: [extract\\_solo\\_features.py](https://github.com/MaxHalford/kaggle-vsb-power/blob/master/scripts/extract_solo_features.py)\n\t\t- Entropy &amp; Fractal: [vsb-competition-attention-bilstm-with-features](https://www.kaggle.com/tarunpaparaju/vsb-competition-attention-bilstm-with-features)\n\t\t- Slope of lines connecting positive/negative peaks to the next negative/positive peaks\n\t\t- Ratio of positive peaks to the minimum of the next maxDistance (defined in Vantuch's Thesis) points.\n\t\t- Ratio of negative peaks to the maximum of the next maxDistance points.\n\t- R² and weights of some polynomial regressions fit in positive/negative peaks\n\n2. RNN:\n\t- Preprocessed signal reshaped in 40 columns\n\n3. CNN:\n\t- Preprocessed signal reshaped in 200 columns\n\n**LB Prediction**\nAt some point in the competition, I built a Lasso to predict the LB score of a submission. I knew that it would create problems with overfitting so I tried to avoid it with:\n- Threshold selection: I fixed all thresholds at the value 0.5\n- Pseudo-Labelling: I used it in all models and I reduced the sample weights of test examples witch had low weight in the trained Lasso\n- Ensemble: I should have ended up with 24 models (6 models x 4 preprocessed signals) but I got just 16 (lack of time...!)\n\nI was not able to create a stable stacking so i used a very elaborate strategy to ensemble the models: Take the arithmetic mean of the predictions...\n\nThe last but not the least, I wanted to thank you everybody for the great kernels and discussions. I learned a lot and got my first gold medal!\n\nThat's all folks!\n",
    "500367": "Congratulations, great job!",
    "499824": "Congratulations and thanks for sharing..",
    "497772": "Congrats on the gold medal and thank you very much for sharing the description of your solution!",
    "496792": "Hi again @vhessel, \nI came back to understand your solution more. I am not sure that I understand on this point\n\n&gt; RNN: Preprocessed signal reshaped in 40 columns\nCNN: Preprocessed signal reshaped in 200 columns\n\nBefore this line, you mentioned that you undersampled the signal data into 300,000.\n\nDo I understand correctly that the final shape that you feed into RNN is something like\n\n8712 x 40 x 7500  (40*7500 = 300,000), and then feed this into RNN?\n(so you did not group the 3 phases to predict together?)",
    "496588": "Congratulations, great work ! Glad my kernel came in handy for you :)",
    "496456": "good job!",
    "496452": "Congrats and thanks for sharing. I like the diversity in your selection(Gradient Boosting, RNNs, and CNNs).",
    "496435": "Congratulations! Nice work!",
    "496425": "Congratz! Any idea on how your 3 models perform without ensembling ? I'm curious!",
    "496342": "Congrats @vhessel on your 1st gold medal and thanks for sharing you solution overview.\n\nOn the CV strategy, I have a [discussion post here](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85143) if you want to add your thoughts.",
    "496316": "You did a superb job here, congratulation!!\n\nSome questions:\n- In the final solution, you aligned &amp; cropped the signal, but did you try unaligned/uncropped signal ? Do they give very different results?\n- And how much did the denoising help? Unfortunately for me, I haven’t tried this at all! (I thought of the DL principle of giving it raw data :p )\n- How could you psuedo label data?\n\n\nLastly, This is the first time I heard of LB prediction ... truly nice to hear! But I could still not understand how it work, does it relate to your pseudo labelling? (is that just linear regression lasso from 20337 example to [0,1] ?? )\n\nCongratulation again and sorry to ask many questions !\n",
    "496314": "Congratulations and great work Vinicius!   What kind of local cross validation scheme did you use?",
    "497530": "good job!",
    "496405": "Congratulations on the strong finish ! Your ensemble is amazing by itself.\n\nCan I ask what was the reasoning behind cutting the last quarter of the signal ? And why 2/1 downsampling by just choosing the maximum ?",
    "496365": "&gt;I was not able to create a stable stacking so i used a very elaborate strategy to ensemble the models: Take the arithmetic mean of the predictions…\n\nInteresting. What do you mean by *arithmetic mean*. Is that a mean of the probability, or a mean of yes/no?\n\nOur team tried both, and it seems a voting (mean of yes/no) works better than mean of probability. \n\nThanks",
    "499628": "",
    "496602": "",
    "499765": "Congrats and thanks for sharing!",
    "497178": "Congratulations and thanks for sharing!",
    "496352": "Congratulations and Thanks for sharing. "
  }
}