{
  "id": 87150,
  "title": "Denoised Dataset and Performance Improvement Wrap-up",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/87150",
  "author_name": "",
  "post_date": "2019-03-29T05:19:17.276651300Z",
  "votes": 10,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>Introduction</h1>\n\n<p>Regarding VSB partial fault detection, it is found by many top-score participants that the denoised version can efficiently increase predictive performance <strong>for neural networks</strong> (Note : not tree models) in private test set. So our team (credit : @putalay) think that by making this dataset available should benefit and save the time for our community.</p>\n\n<p>We apply train/test parquet files with the method described by Jack <a href=\"/jackvial\">@jackvial</a> : <a href=\"https://www.kaggle.com/jackvial/dwt-signal-denoising\">https://www.kaggle.com/jackvial/dwt-signal-denoising</a></p>\n\n<h2>Dataset URL</h2>\n\n<p>if you are interested, please import this data into your kernel : \n<a href=\"https://www.kaggle.com/thaikeras/vsb-wavelet-denoised/\">https://www.kaggle.com/thaikeras/vsb-wavelet-denoised/</a></p>\n\n<p>You can use this data directly in place of the original train.parq / test.parq</p>\n\n<h2>Effects and Notes on Performance Improvement</h2>\n\n<p>1) by applying this data version to our top public kernels (e.g. <a href=\"/braquino\">@braquino</a> Bruno's and <a href=\"/tarunpaparaju\">@tarunpaparaju</a> Tarun's), we immediately got the privateLB results around 0.640 - 0.670 ... <strong>So, in the first place, people who believe in this denoising method already won medals.</strong></p>\n\n<p>2) However, it is certainly not easy to decide and believe in this method since direct application will reduce public LB score to 0.630-0.670</p>\n\n<p>3) In order to also get a good public score (as well as private scores) we need to do more, for examples by some insights of our participants : </p>\n\n<ul>\n<li><p>Justify the validation of the method (using adversarial validation) as mentioned by <a href=\"/yukinkgwa\">@yukinkgwa</a> <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85258\">here</a></p></li>\n<li><p>Apply ensembles :  as used by many top scorers</p></li>\n<li><p>Apply justified complicated pre-processing techniques like <a href=\"/vhessel\">@vhessel</a> <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85170\">here</a> </p></li>\n<li><p>Use a good probed data in addition to a training data as <a href=\"/bigswimatom\">@bigswimatom</a> in <a href=\"https://www.kaggle.com/bigswimatom/5-fold-lstm-attention-stateful-metrics-with-exp\">https://www.kaggle.com/bigswimatom/5-fold-lstm-attention-stateful-metrics-with-exp</a>\n(I empirically found that stateful metric doesn't help much as the probed data)</p></li>\n<li><p>or if you are able to find a killer feature like <a href=\"/tw1994\">@tw1994</a> Tang's <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/86616\">here</a></p></li>\n</ul>\n\n<p>4) This denoised data seems cannot improve the performance of the tree based method. (I try to use it in <a href=\"https://www.kaggle.com/qinhui1999/handmade-features-0685-private-score\">https://www.kaggle.com/qinhui1999/handmade-features-0685-private-score</a> )\nBut <a href=\"/qinhui1999\">@qinhui1999</a> Hui’s pseudo labelling method could also improve the publicLB in the DL kernel.</p>\n\n<p>Final note to <a href=\"/sheriytm\">@sheriytm</a> my friend, I tried re-implemented many methods shared by our top participants, but the single most important and effective factor (exclude more complicated techniques mentioned above) seems to be the denoised dataset here, so I share it here as promised to you :) .</p>",
  "messages": [
    {
      "id": "502815",
      "postDate": "03/29/2019 05:19:17",
      "content": "<h1>Introduction</h1>\n\n<p>Regarding VSB partial fault detection, it is found by many top-score participants that the denoised version can efficiently increase predictive performance <strong>for neural networks</strong> (Note : not tree models) in private test set. So our team (credit : @putalay) think that by making this dataset available should benefit and save the time for our community.</p>\n\n<p>We apply train/test parquet files with the method described by Jack <a href=\"/jackvial\">@jackvial</a> : <a href=\"https://www.kaggle.com/jackvial/dwt-signal-denoising\">https://www.kaggle.com/jackvial/dwt-signal-denoising</a></p>\n\n<h2>Dataset URL</h2>\n\n<p>if you are interested, please import this data into your kernel : \n<a href=\"https://www.kaggle.com/thaikeras/vsb-wavelet-denoised/\">https://www.kaggle.com/thaikeras/vsb-wavelet-denoised/</a></p>\n\n<p>You can use this data directly in place of the original train.parq / test.parq</p>\n\n<h2>Effects and Notes on Performance Improvement</h2>\n\n<p>1) by applying this data version to our top public kernels (e.g. <a href=\"/braquino\">@braquino</a> Bruno's and <a href=\"/tarunpaparaju\">@tarunpaparaju</a> Tarun's), we immediately got the privateLB results around 0.640 - 0.670 ... <strong>So, in the first place, people who believe in this denoising method already won medals.</strong></p>\n\n<p>2) However, it is certainly not easy to decide and believe in this method since direct application will reduce public LB score to 0.630-0.670</p>\n\n<p>3) In order to also get a good public score (as well as private scores) we need to do more, for examples by some insights of our participants : </p>\n\n<ul>\n<li><p>Justify the validation of the method (using adversarial validation) as mentioned by <a href=\"/yukinkgwa\">@yukinkgwa</a> <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85258\">here</a></p></li>\n<li><p>Apply ensembles :  as used by many top scorers</p></li>\n<li><p>Apply justified complicated pre-processing techniques like <a href=\"/vhessel\">@vhessel</a> <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85170\">here</a> </p></li>\n<li><p>Use a good probed data in addition to a training data as <a href=\"/bigswimatom\">@bigswimatom</a> in <a href=\"https://www.kaggle.com/bigswimatom/5-fold-lstm-attention-stateful-metrics-with-exp\">https://www.kaggle.com/bigswimatom/5-fold-lstm-attention-stateful-metrics-with-exp</a>\n(I empirically found that stateful metric doesn't help much as the probed data)</p></li>\n<li><p>or if you are able to find a killer feature like <a href=\"/tw1994\">@tw1994</a> Tang's <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/86616\">here</a></p></li>\n</ul>\n\n<p>4) This denoised data seems cannot improve the performance of the tree based method. (I try to use it in <a href=\"https://www.kaggle.com/qinhui1999/handmade-features-0685-private-score\">https://www.kaggle.com/qinhui1999/handmade-features-0685-private-score</a> )\nBut <a href=\"/qinhui1999\">@qinhui1999</a> Hui’s pseudo labelling method could also improve the publicLB in the DL kernel.</p>\n\n<p>Final note to <a href=\"/sheriytm\">@sheriytm</a> my friend, I tried re-implemented many methods shared by our top participants, but the single most important and effective factor (exclude more complicated techniques mentioned above) seems to be the denoised dataset here, so I share it here as promised to you :) .</p>",
      "rawMarkdown": "# Introduction\n\nRegarding VSB partial fault detection, it is found by many top-score participants that the denoised version can efficiently increase predictive performance **for neural networks** (Note : not tree models) in private test set. So our team (credit : @putalay) think that by making this dataset available should benefit and save the time for our community.\n\nWe apply train/test parquet files with the method described by Jack @jackvial : https://www.kaggle.com/jackvial/dwt-signal-denoising\n\n## Dataset URL\nif you are interested, please import this data into your kernel : \nhttps://www.kaggle.com/thaikeras/vsb-wavelet-denoised/\n\nYou can use this data directly in place of the original train.parq / test.parq\n\n## Effects and Notes on Performance Improvement\n\n1) by applying this data version to our top public kernels (e.g. @braquino Bruno's and @tarunpaparaju Tarun's), we immediately got the privateLB results around 0.640 - 0.670 ... **So, in the first place, people who believe in this denoising method already won medals.**\n\n2) However, it is certainly not easy to decide and believe in this method since direct application will reduce public LB score to 0.630-0.670\n\n3) In order to also get a good public score (as well as private scores) we need to do more, for examples by some insights of our participants : \n\n- Justify the validation of the method (using adversarial validation) as mentioned by @yukinkgwa [here](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85258)\n\n- Apply ensembles :  as used by many top scorers\n\n- Apply justified complicated pre-processing techniques like @vhessel [here](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85170) \n\n- Use a good probed data in addition to a training data as @bigswimatom in https://www.kaggle.com/bigswimatom/5-fold-lstm-attention-stateful-metrics-with-exp\n(I empirically found that stateful metric doesn't help much as the probed data)\n\n- or if you are able to find a killer feature like @tw1994 Tang's [here](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/86616)\n\n\n4) This denoised data seems cannot improve the performance of the tree based method. (I try to use it in https://www.kaggle.com/qinhui1999/handmade-features-0685-private-score )\nBut @qinhui1999 Hui’s pseudo labelling method could also improve the publicLB in the DL kernel.\n\nFinal note to @sheriytm my friend, I tried re-implemented many methods shared by our top participants, but the single most important and effective factor (exclude more complicated techniques mentioned above) seems to be the denoised dataset here, so I share it here as promised to you :) .",
      "votes": null
    },
    {
      "id": "502884",
      "postDate": "03/29/2019 07:38:41",
      "content": "<p>Well done <a href=\"/ratthachat\">@ratthachat</a> and thank you very much for generously sharing your findings. I find the problem addressed in this competion fascinating. My goal is to find time later on to build a simple model that can produce nearly as good a result as the top solutions. When and if I do that, I will be posting my updates here.</p>",
      "rawMarkdown": "Well done @ratthachat and thank you very much for generously sharing your findings. I find the problem addressed in this competion fascinating. My goal is to find time later on to build a simple model that can produce nearly as good a result as the top solutions. When and if I do that, I will be posting my updates here.",
      "votes": null
    },
    {
      "id": "502932",
      "postDate": "03/29/2019 09:20:35",
      "content": "<p>Thank you very much for citing my kernel. </p>",
      "rawMarkdown": "Thank you very much for citing my kernel.",
      "votes": null
    },
    {
      "id": "503069",
      "postDate": "03/29/2019 12:39:30",
      "content": "<p>Great work, <a href=\"/ratthachat\">@ratthachat</a> ! Thanks for citing my kernel !</p>",
      "rawMarkdown": "Great work, @ratthachat ! Thanks for citing my kernel !",
      "votes": null
    },
    {
      "id": "504122",
      "postDate": "03/31/2019 02:47:38",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a>, I can confirm your observations that denoised data does not predictive performance of tree bsed models. I plugged you data denoised to <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146\">my best model</a> on the private LB i.e. <a href=\"https://www.kaggle.com/sheriytm/5-fold-lstm-atten-fully-commented-original-cpu?scriptVersionId=10248597\">my public Catboost kernel</a> and it did not improve the private LB score of 0.62059 but scored the highest of all my models on th public LB with 0.71467. I will try it with my NN model when I resume experimentation on this data.</p>",
      "rawMarkdown": "ratthachat, I can confirm your observations that denoised data does not predictive performance of tree bsed models. I plugged you data denoised to [my best model](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146) on the private LB i.e. [my public Catboost kernel](https://www.kaggle.com/sheriytm/5-fold-lstm-atten-fully-commented-original-cpu?scriptVersionId=10248597) and it did not improve the private LB score of 0.62059 but scored the highest of all my models on th public LB with 0.71467. I will try it with my NN model when I resume experimentation on this data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 502884,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "03/29/2019 07:38:41",
      "content": "<p>Well done <a href=\"/ratthachat\">@ratthachat</a> and thank you very much for generously sharing your findings. I find the problem addressed in this competion fascinating. My goal is to find time later on to build a simple model that can produce nearly as good a result as the top solutions. When and if I do that, I will be posting my updates here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 502932,
      "author_name": "bigswimatom",
      "author_url": "",
      "post_date": "03/29/2019 09:20:35",
      "content": "<p>Thank you very much for citing my kernel. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 503069,
      "author_name": "tarunpaparaju",
      "author_url": "",
      "post_date": "03/29/2019 12:39:30",
      "content": "<p>Great work, <a href=\"/ratthachat\">@ratthachat</a> ! Thanks for citing my kernel !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 504122,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "03/31/2019 02:47:38",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a>, I can confirm your observations that denoised data does not predictive performance of tree bsed models. I plugged you data denoised to <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146\">my best model</a> on the private LB i.e. <a href=\"https://www.kaggle.com/sheriytm/5-fold-lstm-atten-fully-commented-original-cpu?scriptVersionId=10248597\">my public Catboost kernel</a> and it did not improve the private LB score of 0.62059 but scored the highest of all my models on th public LB with 0.71467. I will try it with my NN model when I resume experimentation on this data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "502815": "# Introduction\n\nRegarding VSB partial fault detection, it is found by many top-score participants that the denoised version can efficiently increase predictive performance **for neural networks** (Note : not tree models) in private test set. So our team (credit : @putalay) think that by making this dataset available should benefit and save the time for our community.\n\nWe apply train/test parquet files with the method described by Jack @jackvial : https://www.kaggle.com/jackvial/dwt-signal-denoising\n\n## Dataset URL\nif you are interested, please import this data into your kernel : \nhttps://www.kaggle.com/thaikeras/vsb-wavelet-denoised/\n\nYou can use this data directly in place of the original train.parq / test.parq\n\n## Effects and Notes on Performance Improvement\n\n1) by applying this data version to our top public kernels (e.g. @braquino Bruno's and @tarunpaparaju Tarun's), we immediately got the privateLB results around 0.640 - 0.670 ... **So, in the first place, people who believe in this denoising method already won medals.**\n\n2) However, it is certainly not easy to decide and believe in this method since direct application will reduce public LB score to 0.630-0.670\n\n3) In order to also get a good public score (as well as private scores) we need to do more, for examples by some insights of our participants : \n\n- Justify the validation of the method (using adversarial validation) as mentioned by @yukinkgwa [here](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85258)\n\n- Apply ensembles :  as used by many top scorers\n\n- Apply justified complicated pre-processing techniques like @vhessel [here](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85170) \n\n- Use a good probed data in addition to a training data as @bigswimatom in https://www.kaggle.com/bigswimatom/5-fold-lstm-attention-stateful-metrics-with-exp\n(I empirically found that stateful metric doesn't help much as the probed data)\n\n- or if you are able to find a killer feature like @tw1994 Tang's [here](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/86616)\n\n\n4) This denoised data seems cannot improve the performance of the tree based method. (I try to use it in https://www.kaggle.com/qinhui1999/handmade-features-0685-private-score )\nBut @qinhui1999 Hui’s pseudo labelling method could also improve the publicLB in the DL kernel.\n\nFinal note to @sheriytm my friend, I tried re-implemented many methods shared by our top participants, but the single most important and effective factor (exclude more complicated techniques mentioned above) seems to be the denoised dataset here, so I share it here as promised to you :) .",
    "502884": "Well done @ratthachat and thank you very much for generously sharing your findings. I find the problem addressed in this competion fascinating. My goal is to find time later on to build a simple model that can produce nearly as good a result as the top solutions. When and if I do that, I will be posting my updates here.",
    "502932": "Thank you very much for citing my kernel.",
    "503069": "Great work, @ratthachat ! Thanks for citing my kernel !",
    "504122": "ratthachat, I can confirm your observations that denoised data does not predictive performance of tree bsed models. I plugged you data denoised to [my best model](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146) on the private LB i.e. [my public Catboost kernel](https://www.kaggle.com/sheriytm/5-fold-lstm-atten-fully-commented-original-cpu?scriptVersionId=10248597) and it did not improve the private LB score of 0.62059 but scored the highest of all my models on th public LB with 0.71467. I will try it with my NN model when I resume experimentation on this data."
  },
  "source": "meta"
}