{
  "id": 87038,
  "title": "Overview of 1st place solution",
  "url": "/competitions/vsb-power-line-fault-detection/writeups/mark4h-overview-of-1st-place-solution",
  "author_name": "",
  "post_date": "2019-03-28T11:19:11.675285Z",
  "votes": 77,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Congratulation to all the winners and thanks to VSB for hosting the competition.</p>\n\n<p>I'll provide an overview of the approach I took here. If you want to take a closer look I have uploaded the <a href=\"https://www.kaggle.com/mark4h/vsb-1st-place-solution\">winning kernel</a>.</p>\n\n<p>Like everyone has been discussing, it was challenging to cope with the peculiarities of the dataset. I didn't use any specific CV technique to prevent over-fitting, instead I spent most of the time trying to understand the data. In the end my best model was a simple LightGBM model, with some preprocessing of the data and a small number (9 in total) of features based on the patterns common to the train and test data. </p>\n\n<p>Preprocessing:</p>\n\n<pre><code>• I spent some time trying to denoise the traces but in the end came to the conclusion that it was just as likely to remove signal as it was noise\n• Instead I focused on finding a way to pick a threshold for each trace to separate the peaks (signal) from the noise\n• First I found the local maxima in the data with a window size of 51. I picked the window sized based on the data.\n• I then used a knee point detection to find the noise floor. I ordered the peaks from highest to lowest and then looked for the point where the height of the peaks flattened off.\n</code></pre>\n\n<p>Features:</p>\n\n<p>For each peak identified in the preprocessing I then calculated a few features, of which the first 2 where the most important:</p>\n\n<pre><code>1. The absolute height of the peak\n2. The RMSE between the peak and a sawtooth shaped template. I picked this because I noticed this sort of shape was common in traces marked as faulty\n3. The absolute ratio of the peak to the next data point\n4. The absolute ratio of the peak to the previous data point\n5. The distance to the maximum, of opposite polarity to the peak, within a window of 5 either side of the peak\n</code></pre>\n\n<p>For the features I fed into the LightGBM model I then aggregated the peak features for each measurement ID. The model was trained on the measurement ID instead of the signal ID because there were a number of examples in the training data where it looked as if all 3 signal traces had been marked as faulty even though only one had any obvious signs of a fault. Some of the features were calculated for only certain portions of the phase of each signal so I phase resolved the position of each of the peaks. The features I then calculated, of which the first 5 were the most important, for each measurement ID were:</p>\n\n<pre><code>1. The total number of peaks\n2. The number of peaks in quarters 0 and 2\n3. The number of peaks in quarters 1 and 3\n4. The mean “sawtooth” RMSE value in quarters 0 and 2\n5. The std height in quarters 0 and 2\n6. The mean height in quarters 0 and 2\n7. The mean value of the ratio with the previous data point feature in quarters 0 and 2\n8. The mean value of the ratio with the next data point feature in quarters 0 and 2\n9. The mean value of the absolute distance to the opposite polarity maximum\n</code></pre>\n\n<p>Finally, I trained a LightGBM model using 5 fold CV but repeated 25 times using a different seed for the CV split each time.</p>\n\n<p>It would be great to get peoples thoughts on this approach.</p>\n\n<p>mark4h</p>",
  "messages": [
    {
      "id": "502270",
      "postDate": "03/28/2019 11:19:11",
      "content": "<p>Congratulation to all the winners and thanks to VSB for hosting the competition.</p>\n\n<p>I'll provide an overview of the approach I took here. If you want to take a closer look I have uploaded the <a href=\"https://www.kaggle.com/mark4h/vsb-1st-place-solution\">winning kernel</a>.</p>\n\n<p>Like everyone has been discussing, it was challenging to cope with the peculiarities of the dataset. I didn't use any specific CV technique to prevent over-fitting, instead I spent most of the time trying to understand the data. In the end my best model was a simple LightGBM model, with some preprocessing of the data and a small number (9 in total) of features based on the patterns common to the train and test data. </p>\n\n<p>Preprocessing:</p>\n\n<pre><code>• I spent some time trying to denoise the traces but in the end came to the conclusion that it was just as likely to remove signal as it was noise\n• Instead I focused on finding a way to pick a threshold for each trace to separate the peaks (signal) from the noise\n• First I found the local maxima in the data with a window size of 51. I picked the window sized based on the data.\n• I then used a knee point detection to find the noise floor. I ordered the peaks from highest to lowest and then looked for the point where the height of the peaks flattened off.\n</code></pre>\n\n<p>Features:</p>\n\n<p>For each peak identified in the preprocessing I then calculated a few features, of which the first 2 where the most important:</p>\n\n<pre><code>1. The absolute height of the peak\n2. The RMSE between the peak and a sawtooth shaped template. I picked this because I noticed this sort of shape was common in traces marked as faulty\n3. The absolute ratio of the peak to the next data point\n4. The absolute ratio of the peak to the previous data point\n5. The distance to the maximum, of opposite polarity to the peak, within a window of 5 either side of the peak\n</code></pre>\n\n<p>For the features I fed into the LightGBM model I then aggregated the peak features for each measurement ID. The model was trained on the measurement ID instead of the signal ID because there were a number of examples in the training data where it looked as if all 3 signal traces had been marked as faulty even though only one had any obvious signs of a fault. Some of the features were calculated for only certain portions of the phase of each signal so I phase resolved the position of each of the peaks. The features I then calculated, of which the first 5 were the most important, for each measurement ID were:</p>\n\n<pre><code>1. The total number of peaks\n2. The number of peaks in quarters 0 and 2\n3. The number of peaks in quarters 1 and 3\n4. The mean “sawtooth” RMSE value in quarters 0 and 2\n5. The std height in quarters 0 and 2\n6. The mean height in quarters 0 and 2\n7. The mean value of the ratio with the previous data point feature in quarters 0 and 2\n8. The mean value of the ratio with the next data point feature in quarters 0 and 2\n9. The mean value of the absolute distance to the opposite polarity maximum\n</code></pre>\n\n<p>Finally, I trained a LightGBM model using 5 fold CV but repeated 25 times using a different seed for the CV split each time.</p>\n\n<p>It would be great to get peoples thoughts on this approach.</p>\n\n<p>mark4h</p>",
      "rawMarkdown": "Congratulation to all the winners and thanks to VSB for hosting the competition.\n\nI'll provide an overview of the approach I took here. If you want to take a closer look I have uploaded the [winning kernel](https://www.kaggle.com/mark4h/vsb-1st-place-solution).\n\nLike everyone has been discussing, it was challenging to cope with the peculiarities of the dataset. I didn't use any specific CV technique to prevent over-fitting, instead I spent most of the time trying to understand the data. In the end my best model was a simple LightGBM model, with some preprocessing of the data and a small number (9 in total) of features based on the patterns common to the train and test data. \n\nPreprocessing:\n\n    • I spent some time trying to denoise the traces but in the end came to the conclusion that it was just as likely to remove signal as it was noise\n    • Instead I focused on finding a way to pick a threshold for each trace to separate the peaks (signal) from the noise\n    • First I found the local maxima in the data with a window size of 51. I picked the window sized based on the data.\n    • I then used a knee point detection to find the noise floor. I ordered the peaks from highest to lowest and then looked for the point where the height of the peaks flattened off.\n\nFeatures:\n\nFor each peak identified in the preprocessing I then calculated a few features, of which the first 2 where the most important:\n\n    1. The absolute height of the peak\n    2. The RMSE between the peak and a sawtooth shaped template. I picked this because I noticed this sort of shape was common in traces marked as faulty\n    3. The absolute ratio of the peak to the next data point\n    4. The absolute ratio of the peak to the previous data point\n    5. The distance to the maximum, of opposite polarity to the peak, within a window of 5 either side of the peak\n\nFor the features I fed into the LightGBM model I then aggregated the peak features for each measurement ID. The model was trained on the measurement ID instead of the signal ID because there were a number of examples in the training data where it looked as if all 3 signal traces had been marked as faulty even though only one had any obvious signs of a fault. Some of the features were calculated for only certain portions of the phase of each signal so I phase resolved the position of each of the peaks. The features I then calculated, of which the first 5 were the most important, for each measurement ID were:\n\n    1. The total number of peaks\n    2. The number of peaks in quarters 0 and 2\n    3. The number of peaks in quarters 1 and 3\n    4. The mean “sawtooth” RMSE value in quarters 0 and 2\n    5. The std height in quarters 0 and 2\n    6. The mean height in quarters 0 and 2\n    7. The mean value of the ratio with the previous data point feature in quarters 0 and 2\n    8. The mean value of the ratio with the next data point feature in quarters 0 and 2\n    9. The mean value of the absolute distance to the opposite polarity maximum\n\nFinally, I trained a LightGBM model using 5 fold CV but repeated 25 times using a different seed for the CV split each time.\n\nIt would be great to get peoples thoughts on this approach.\n\nmark4h",
      "votes": null
    },
    {
      "id": "502328",
      "postDate": "03/28/2019 13:00:02",
      "content": "<blockquote>\n  <p>instead I spent most of the time trying to understand the data. </p>\n</blockquote>\n\n<p>Winning quote of this competition! Thanks for sharing, and of course big congratulation for you!</p>",
      "rawMarkdown": "&gt; instead I spent most of the time trying to understand the data. \n\nWinning quote of this competition! Thanks for sharing, and of course big congratulation for you!",
      "votes": null
    },
    {
      "id": "502345",
      "postDate": "03/28/2019 13:25:32",
      "content": "<p>Love the simplicity of your solution. Great Kernel thanks for sharing !</p>",
      "rawMarkdown": "Love the simplicity of your solution. Great Kernel thanks for sharing !",
      "votes": null
    },
    {
      "id": "502603",
      "postDate": "03/28/2019 20:04:02",
      "content": "<p>Congratulations <a href=\"/mark4h\">@mark4h</a> ! And Thank you a lot for sharing your marvelous solution !\nI'm really surprised by the elegance and originality of your solution !</p>",
      "rawMarkdown": "Congratulations @mark4h ! And Thank you a lot for sharing your marvelous solution !\nI'm really surprised by the elegance and originality of your solution !",
      "votes": null
    },
    {
      "id": "502624",
      "postDate": "03/28/2019 20:51:22",
      "content": "<p>So you don't use neural networks, right? And blending with other models. Cool!</p>",
      "rawMarkdown": "So you don't use neural networks, right? And blending with other models. Cool!",
      "votes": null
    },
    {
      "id": "503026",
      "postDate": "03/29/2019 11:55:50",
      "content": "<p>Thanks everyone.</p>\n\n<p><a href=\"/sergeyzlobin\">@sergeyzlobin</a>, that’s right, no neural network, no blending, just a single LightGBM model.</p>",
      "rawMarkdown": "Thanks everyone.\n\n@sergeyzlobin, that’s right, no neural network, no blending, just a single LightGBM model.",
      "votes": null
    },
    {
      "id": "503827",
      "postDate": "03/30/2019 16:02:48",
      "content": "<p>Amazing simplicity! Only a master can make his/her handcraft look that easy.</p>",
      "rawMarkdown": "Amazing simplicity! Only a master can make his/her handcraft look that easy.",
      "votes": null
    },
    {
      "id": "504346",
      "postDate": "03/31/2019 12:58:46",
      "content": "<p>Simple, but elegant way!  Thank you for sharing your solution. I have a question.\n I'd like to know the differences among the avg and the std of the scores by the models with different random seeds.</p>",
      "rawMarkdown": "Simple, but elegant way!  Thank you for sharing your solution. I have a question.\n I'd like to know the differences among the avg and the std of the scores by the models with different random seeds.",
      "votes": null
    },
    {
      "id": "518960",
      "postDate": "04/18/2019 06:33:04",
      "content": "<p>Nice work</p>",
      "rawMarkdown": "Nice work",
      "votes": null
    },
    {
      "id": "520169",
      "postDate": "04/20/2019 10:20:03",
      "content": "<p>Congratulations <a href=\"/mark4h\">@mark4h</a>, and thanks for sharing your super elegant solution.</p>",
      "rawMarkdown": "Congratulations @mark4h, and thanks for sharing your super elegant solution.",
      "votes": null
    },
    {
      "id": "521846",
      "postDate": "04/23/2019 14:29:23",
      "content": "<p>Thanks for sharing, great approach!</p>",
      "rawMarkdown": "Thanks for sharing, great approach!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 502328,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "03/28/2019 13:00:02",
      "content": "<blockquote>\n  <p>instead I spent most of the time trying to understand the data. </p>\n</blockquote>\n\n<p>Winning quote of this competition! Thanks for sharing, and of course big congratulation for you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 502345,
      "author_name": "areveillon",
      "author_url": "",
      "post_date": "03/28/2019 13:25:32",
      "content": "<p>Love the simplicity of your solution. Great Kernel thanks for sharing !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 502603,
      "author_name": "yukinkgwa",
      "author_url": "",
      "post_date": "03/28/2019 20:04:02",
      "content": "<p>Congratulations <a href=\"/mark4h\">@mark4h</a> ! And Thank you a lot for sharing your marvelous solution !\nI'm really surprised by the elegance and originality of your solution !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 502624,
      "author_name": "sergeyzlobin",
      "author_url": "",
      "post_date": "03/28/2019 20:51:22",
      "content": "<p>So you don't use neural networks, right? And blending with other models. Cool!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 503026,
      "author_name": "mark4h",
      "author_url": "",
      "post_date": "03/29/2019 11:55:50",
      "content": "<p>Thanks everyone.</p>\n\n<p><a href=\"/sergeyzlobin\">@sergeyzlobin</a>, that’s right, no neural network, no blending, just a single LightGBM model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 503827,
      "author_name": "wkirgsn",
      "author_url": "",
      "post_date": "03/30/2019 16:02:48",
      "content": "<p>Amazing simplicity! Only a master can make his/her handcraft look that easy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 504346,
      "author_name": "yatzhash",
      "author_url": "",
      "post_date": "03/31/2019 12:58:46",
      "content": "<p>Simple, but elegant way!  Thank you for sharing your solution. I have a question.\n I'd like to know the differences among the avg and the std of the scores by the models with different random seeds.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 518960,
      "author_name": "wangwangsuibinbin",
      "author_url": "",
      "post_date": "04/18/2019 06:33:04",
      "content": "<p>Nice work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 520169,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "04/20/2019 10:20:03",
      "content": "<p>Congratulations <a href=\"/mark4h\">@mark4h</a>, and thanks for sharing your super elegant solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 521846,
      "author_name": "alanmurphy1",
      "author_url": "",
      "post_date": "04/23/2019 14:29:23",
      "content": "<p>Thanks for sharing, great approach!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "502270": "Congratulation to all the winners and thanks to VSB for hosting the competition.\n\nI'll provide an overview of the approach I took here. If you want to take a closer look I have uploaded the [winning kernel](https://www.kaggle.com/mark4h/vsb-1st-place-solution).\n\nLike everyone has been discussing, it was challenging to cope with the peculiarities of the dataset. I didn't use any specific CV technique to prevent over-fitting, instead I spent most of the time trying to understand the data. In the end my best model was a simple LightGBM model, with some preprocessing of the data and a small number (9 in total) of features based on the patterns common to the train and test data. \n\nPreprocessing:\n\n    • I spent some time trying to denoise the traces but in the end came to the conclusion that it was just as likely to remove signal as it was noise\n    • Instead I focused on finding a way to pick a threshold for each trace to separate the peaks (signal) from the noise\n    • First I found the local maxima in the data with a window size of 51. I picked the window sized based on the data.\n    • I then used a knee point detection to find the noise floor. I ordered the peaks from highest to lowest and then looked for the point where the height of the peaks flattened off.\n\nFeatures:\n\nFor each peak identified in the preprocessing I then calculated a few features, of which the first 2 where the most important:\n\n    1. The absolute height of the peak\n    2. The RMSE between the peak and a sawtooth shaped template. I picked this because I noticed this sort of shape was common in traces marked as faulty\n    3. The absolute ratio of the peak to the next data point\n    4. The absolute ratio of the peak to the previous data point\n    5. The distance to the maximum, of opposite polarity to the peak, within a window of 5 either side of the peak\n\nFor the features I fed into the LightGBM model I then aggregated the peak features for each measurement ID. The model was trained on the measurement ID instead of the signal ID because there were a number of examples in the training data where it looked as if all 3 signal traces had been marked as faulty even though only one had any obvious signs of a fault. Some of the features were calculated for only certain portions of the phase of each signal so I phase resolved the position of each of the peaks. The features I then calculated, of which the first 5 were the most important, for each measurement ID were:\n\n    1. The total number of peaks\n    2. The number of peaks in quarters 0 and 2\n    3. The number of peaks in quarters 1 and 3\n    4. The mean “sawtooth” RMSE value in quarters 0 and 2\n    5. The std height in quarters 0 and 2\n    6. The mean height in quarters 0 and 2\n    7. The mean value of the ratio with the previous data point feature in quarters 0 and 2\n    8. The mean value of the ratio with the next data point feature in quarters 0 and 2\n    9. The mean value of the absolute distance to the opposite polarity maximum\n\nFinally, I trained a LightGBM model using 5 fold CV but repeated 25 times using a different seed for the CV split each time.\n\nIt would be great to get peoples thoughts on this approach.\n\nmark4h",
    "502328": "&gt; instead I spent most of the time trying to understand the data. \n\nWinning quote of this competition! Thanks for sharing, and of course big congratulation for you!",
    "502345": "Love the simplicity of your solution. Great Kernel thanks for sharing !",
    "502603": "Congratulations @mark4h ! And Thank you a lot for sharing your marvelous solution !\nI'm really surprised by the elegance and originality of your solution !",
    "502624": "So you don't use neural networks, right? And blending with other models. Cool!",
    "503026": "Thanks everyone.\n\n@sergeyzlobin, that’s right, no neural network, no blending, just a single LightGBM model.",
    "503827": "Amazing simplicity! Only a master can make his/her handcraft look that easy.",
    "504346": "Simple, but elegant way!  Thank you for sharing your solution. I have a question.\n I'd like to know the differences among the avg and the std of the scores by the models with different random seeds.",
    "518960": "Nice work",
    "520169": "Congratulations @mark4h, and thanks for sharing your super elegant solution.",
    "521846": "Thanks for sharing, great approach!"
  },
  "source": "meta"
}