{
  "id": 76302,
  "title": "A (currently) non DL starter code [0.362]",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/76302",
  "author_name": "Rand Xie",
  "post_date": "2018-12-31T20:48:36.609000",
  "votes": 30,
  "comment_count": 8,
  "views": 0,
  "content": "<p>It is amazing to see deep learning based approaches have achieved very good results. Before exploring the DL side, I would love to try and push forward the limit of (manual) feature engineering, specifically the signal processing based feature engineering techniques (compute peak statistics after wavelet de-noising, signal entropy, ...).</p>\n\n<p>In many previous competitions, I am more of a consumer than producer. As also promised to myself, I will share more of my thoughts and solutions in the future competitions. In this competition, I would like to share my starter code written in Pytorch, which is compatible to both signal processing workflow and deep learning workflow. Please find the code here: <a href=\"https://github.com/randxie/Kaggle-VSB-Baseline\">https://github.com/randxie/Kaggle-VSB-Baseline</a></p>\n\n<p>Key Features are:</p>\n\n<ol>\n<li>Use Pytorch dataloader to do feature extraction, which leverages the build-in multiprocessing capability. It also makes an easy extension to deep learning workflow.</li>\n<li>Add new feature extractor only requires inheriting the AbstractExtractor then write your <strong>process_fn</strong> and <strong>extract_fn</strong>. Before computation, feature extractor will use the first signal to infer feature dimension so that you don't need to pre-specify that. The pipeline will also automatically cache the computed features for you.</li>\n<li>Use json to configure parameters in the pipeline. It helps you to document your experiments.</li>\n<li>Use random seed to control data augmentation so that results are reproducible.</li>\n</ol>\n\n<p>Thoughts:</p>\n\n<p>The biggest challenge I have is to design a stable validation scheme. Currently, I used group KFold to make sure signals from same group of power lines are in the same CV folds. However, the validation score is still way off from the LB scores. I suspect there are some feature distribution shift or label distribution shift. Please also share your validation journey.</p>\n\n<p>Finally, happy new year and enjoy the competition!! </p>",
  "messages": [
    {
      "id": 448348,
      "postDate": "2018-12-31T20:48:36.610Z",
      "content": "<p>It is amazing to see deep learning based approaches have achieved very good results. Before exploring the DL side, I would love to try and push forward the limit of (manual) feature engineering, specifically the signal processing based feature engineering techniques (compute peak statistics after wavelet de-noising, signal entropy, ...).</p>\n\n<p>In many previous competitions, I am more of a consumer than producer. As also promised to myself, I will share more of my thoughts and solutions in the future competitions. In this competition, I would like to share my starter code written in Pytorch, which is compatible to both signal processing workflow and deep learning workflow. Please find the code here: <a href=\"https://github.com/randxie/Kaggle-VSB-Baseline\">https://github.com/randxie/Kaggle-VSB-Baseline</a></p>\n\n<p>Key Features are:</p>\n\n<ol>\n<li>Use Pytorch dataloader to do feature extraction, which leverages the build-in multiprocessing capability. It also makes an easy extension to deep learning workflow.</li>\n<li>Add new feature extractor only requires inheriting the AbstractExtractor then write your <strong>process_fn</strong> and <strong>extract_fn</strong>. Before computation, feature extractor will use the first signal to infer feature dimension so that you don't need to pre-specify that. The pipeline will also automatically cache the computed features for you.</li>\n<li>Use json to configure parameters in the pipeline. It helps you to document your experiments.</li>\n<li>Use random seed to control data augmentation so that results are reproducible.</li>\n</ol>\n\n<p>Thoughts:</p>\n\n<p>The biggest challenge I have is to design a stable validation scheme. Currently, I used group KFold to make sure signals from same group of power lines are in the same CV folds. However, the validation score is still way off from the LB scores. I suspect there are some feature distribution shift or label distribution shift. Please also share your validation journey.</p>\n\n<p>Finally, happy new year and enjoy the competition!! </p>",
      "rawMarkdown": "It is amazing to see deep learning based approaches have achieved very good results. Before exploring the DL side, I would love to try and push forward the limit of (manual) feature engineering, specifically the signal processing based feature engineering techniques (compute peak statistics after wavelet de-noising, signal entropy, ...).\n\nIn many previous competitions, I am more of a consumer than producer. As also promised to myself, I will share more of my thoughts and solutions in the future competitions. In this competition, I would like to share my starter code written in Pytorch, which is compatible to both signal processing workflow and deep learning workflow. Please find the code here: https://github.com/randxie/Kaggle-VSB-Baseline\n\nKey Features are:\n\n1. Use Pytorch dataloader to do feature extraction, which leverages the build-in multiprocessing capability. It also makes an easy extension to deep learning workflow.\n2. Add new feature extractor only requires inheriting the AbstractExtractor then write your **process_fn** and **extract_fn**. Before computation, feature extractor will use the first signal to infer feature dimension so that you don't need to pre-specify that. The pipeline will also automatically cache the computed features for you.\n3. Use json to configure parameters in the pipeline. It helps you to document your experiments.\n4. Use random seed to control data augmentation so that results are reproducible.\n\nThoughts:\n\nThe biggest challenge I have is to design a stable validation scheme. Currently, I used group KFold to make sure signals from same group of power lines are in the same CV folds. However, the validation score is still way off from the LB scores. I suspect there are some feature distribution shift or label distribution shift. Please also share your validation journey.\n\nFinally, happy new year and enjoy the competition!! \n",
      "votes": 30
    },
    {
      "id": 448421,
      "postDate": "2019-01-01T05:22:06.740Z",
      "content": "<blockquote>\n  <p>The biggest challenge I have is to design a stable validation scheme.</p>\n</blockquote>\n\n<p>That's the biggest challenge in <em>every</em> Kaggle competition. Figure that one out, and you are 90% on your way to Gold. :)</p>",
      "rawMarkdown": "&gt;The biggest challenge I have is to design a stable validation scheme.\n\nThat's the biggest challenge in _every_ Kaggle competition. Figure that one out, and you are 90% on your way to Gold. :)",
      "votes": 3,
      "replies": [
        {
          "id": 448763,
          "postDate": "2019-01-02T03:19:21.133Z",
          "content": "<p>That's true. In this competition, the unbalance label makes it more challenging to build a good validation scheme. </p>",
          "rawMarkdown": "That's true. In this competition, the unbalance label makes it more challenging to build a good validation scheme. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 469044,
      "postDate": "2019-02-10T11:23:17.407Z",
      "content": "<blockquote>\n  <p>However, the validation score is still way off from the LB scores.</p>\n</blockquote>\n\n<p>I saw the following code snippet in src/common.py</p>\n\n<p><code>SAMPLING_FREQ = 80000/0.02 # 80,000 data points taken over 20 ms</code></p>\n\n<p>However, according to the data page, \"Each signal contains 800,000 measurements of a power line's voltage, taken over 20 milliseconds.\"\nI am just wondering if you can check this, just in case!</p>",
      "rawMarkdown": "&gt;However, the validation score is still way off from the LB scores.\n\nI saw the following code snippet in src/common.py\n\n`SAMPLING_FREQ = 80000/0.02 # 80,000 data points taken over 20 ms`\n\nHowever, according to the data page, \"Each signal contains 800,000 measurements of a power line's voltage, taken over 20 milliseconds.\"\nI am just wondering if you can check this, just in case!",
      "votes": 1,
      "replies": [
        {
          "id": 471955,
          "postDate": "2019-02-15T06:23:05.030Z",
          "content": "<p>This may be caused by unbalanced distribution of the train &amp; test dataset. I see the auc between train &amp; test is above 0.9. You can refer to :<a href=\"https://www.kaggle.com/its7171/train-vs-test-analisys\">https://www.kaggle.com/its7171/train-vs-test-analisys</a>.</p>",
          "rawMarkdown": "This may be caused by unbalanced distribution of the train &amp; test dataset. I see the auc between train &amp; test is above 0.9. You can refer to :https://www.kaggle.com/its7171/train-vs-test-analisys."
        },
        {
          "id": 472673,
          "postDate": "2019-02-16T12:56:36.107Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 464755,
      "postDate": "2019-02-01T12:39:19.130Z",
      "content": "<p>Great Code and documentation as always ! Thanks for sharing :)</p>",
      "rawMarkdown": "Great Code and documentation as always ! Thanks for sharing :)",
      "votes": 1
    },
    {
      "id": 471954,
      "postDate": "2019-02-15T06:22:44.603Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 449654,
      "postDate": "2019-01-03T14:14:56.103Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 448421,
      "author_name": "Bojan Tunguz",
      "author_url": "",
      "post_date": "2019-01-01T05:22:06.740000",
      "content": "<blockquote>\n  <p>The biggest challenge I have is to design a stable validation scheme.</p>\n</blockquote>\n\n<p>That's the biggest challenge in <em>every</em> Kaggle competition. Figure that one out, and you are 90% on your way to Gold. :)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 448763,
          "author_name": "Rand Xie",
          "author_url": "",
          "post_date": "2019-01-02T03:19:21.133000",
          "content": "<p>That's true. In this competition, the unbalance label makes it more challenging to build a good validation scheme. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 469044,
      "author_name": "Mohammad Azam Khan",
      "author_url": "",
      "post_date": "2019-02-10T11:23:17.407000",
      "content": "<blockquote>\n  <p>However, the validation score is still way off from the LB scores.</p>\n</blockquote>\n\n<p>I saw the following code snippet in src/common.py</p>\n\n<p><code>SAMPLING_FREQ = 80000/0.02 # 80,000 data points taken over 20 ms</code></p>\n\n<p>However, according to the data page, \"Each signal contains 800,000 measurements of a power line's voltage, taken over 20 milliseconds.\"\nI am just wondering if you can check this, just in case!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 471955,
          "author_name": "LongYin/杰少",
          "author_url": "",
          "post_date": "2019-02-15T06:23:05.030000",
          "content": "<p>This may be caused by unbalanced distribution of the train &amp; test dataset. I see the auc between train &amp; test is above 0.9. You can refer to :<a href=\"https://www.kaggle.com/its7171/train-vs-test-analisys\">https://www.kaggle.com/its7171/train-vs-test-analisys</a>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 472673,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-16T12:56:36.107000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 464755,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2019-02-01T12:39:19.130000",
      "content": "<p>Great Code and documentation as always ! Thanks for sharing :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 471954,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-15T06:22:44.603000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 449654,
      "author_name": "xie",
      "author_url": "",
      "post_date": "2019-01-03T14:14:56.103000",
      "content": "<p>Thanks!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "448348": "It is amazing to see deep learning based approaches have achieved very good results. Before exploring the DL side, I would love to try and push forward the limit of (manual) feature engineering, specifically the signal processing based feature engineering techniques (compute peak statistics after wavelet de-noising, signal entropy, ...).\n\nIn many previous competitions, I am more of a consumer than producer. As also promised to myself, I will share more of my thoughts and solutions in the future competitions. In this competition, I would like to share my starter code written in Pytorch, which is compatible to both signal processing workflow and deep learning workflow. Please find the code here: https://github.com/randxie/Kaggle-VSB-Baseline\n\nKey Features are:\n\n1. Use Pytorch dataloader to do feature extraction, which leverages the build-in multiprocessing capability. It also makes an easy extension to deep learning workflow.\n2. Add new feature extractor only requires inheriting the AbstractExtractor then write your **process_fn** and **extract_fn**. Before computation, feature extractor will use the first signal to infer feature dimension so that you don't need to pre-specify that. The pipeline will also automatically cache the computed features for you.\n3. Use json to configure parameters in the pipeline. It helps you to document your experiments.\n4. Use random seed to control data augmentation so that results are reproducible.\n\nThoughts:\n\nThe biggest challenge I have is to design a stable validation scheme. Currently, I used group KFold to make sure signals from same group of power lines are in the same CV folds. However, the validation score is still way off from the LB scores. I suspect there are some feature distribution shift or label distribution shift. Please also share your validation journey.\n\nFinally, happy new year and enjoy the competition!! \n",
    "448421": "&gt;The biggest challenge I have is to design a stable validation scheme.\n\nThat's the biggest challenge in _every_ Kaggle competition. Figure that one out, and you are 90% on your way to Gold. :)",
    "469044": "&gt;However, the validation score is still way off from the LB scores.\n\nI saw the following code snippet in src/common.py\n\n`SAMPLING_FREQ = 80000/0.02 # 80,000 data points taken over 20 ms`\n\nHowever, according to the data page, \"Each signal contains 800,000 measurements of a power line's voltage, taken over 20 milliseconds.\"\nI am just wondering if you can check this, just in case!",
    "464755": "Great Code and documentation as always ! Thanks for sharing :)",
    "471954": "",
    "449654": "Thanks!"
  }
}