{
  "id": 85156,
  "title": "My useful tricks in this hard competition",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/85156",
  "author_name": "",
  "post_date": "2019-03-22T01:46:31.524712200Z",
  "votes": 12,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Congratulations to all the winners.  This competition is ver hard . We tried many ways to improve our scores, but most of them are overfitting.  There are still some useful tricks .\n<br>\n1.  pseudo label .   We use the pseudo label to retrain our models in the last two weeks.  It really improve the cv and lb scores (including private scores).  <br></p>\n\n<ol>\n<li><p>data augment.   In the last two days,  I see the data augment <a href=\"https://www.kaggle.com/jiweiliu/lgb-2-leaves-augment\">kernel</a> . Then I used it in our models.  It works and improve our cv and lb scores. <br></p></li>\n<li><p>Catboost model is better than LGB and RNN.   I used this <a href=\"https://www.kaggle.com/junkoda/handmade-features?scriptVersionId=10818864\">kernel</a>() and train  it with  pseudo label and data augment.  Finally I got cv0.765,  0.690 (lb) and 0.679 (private).  To my regret, we did not choose this model.  : ( \n<br></p></li>\n</ol>\n\n<p>I learned a lot of tricks and lesson from this competition. Thanks Kaggle.  Thanks every sharings in this competition.</p>",
  "messages": [
    {
      "id": "496235",
      "postDate": "03/22/2019 01:46:31",
      "content": "<p>Congratulations to all the winners.  This competition is ver hard . We tried many ways to improve our scores, but most of them are overfitting.  There are still some useful tricks .\n<br>\n1.  pseudo label .   We use the pseudo label to retrain our models in the last two weeks.  It really improve the cv and lb scores (including private scores).  <br></p>\n\n<ol>\n<li><p>data augment.   In the last two days,  I see the data augment <a href=\"https://www.kaggle.com/jiweiliu/lgb-2-leaves-augment\">kernel</a> . Then I used it in our models.  It works and improve our cv and lb scores. <br></p></li>\n<li><p>Catboost model is better than LGB and RNN.   I used this <a href=\"https://www.kaggle.com/junkoda/handmade-features?scriptVersionId=10818864\">kernel</a>() and train  it with  pseudo label and data augment.  Finally I got cv0.765,  0.690 (lb) and 0.679 (private).  To my regret, we did not choose this model.  : ( \n<br></p></li>\n</ol>\n\n<p>I learned a lot of tricks and lesson from this competition. Thanks Kaggle.  Thanks every sharings in this competition.</p>",
      "rawMarkdown": "Congratulations to all the winners.  This competition is ver hard . We tried many ways to improve our scores, but most of them are overfitting.  There are still some useful tricks .\n<br>\n1.  pseudo label .   We use the pseudo label to retrain our models in the last two weeks.  It really improve the cv and lb scores (including private scores).  <br>\n\n2.  data augment.   In the last two days,  I see the data augment [kernel](https://www.kaggle.com/jiweiliu/lgb-2-leaves-augment) . Then I used it in our models.  It works and improve our cv and lb scores. <br>\n\n3.  Catboost model is better than LGB and RNN.   I used this [kernel](https://www.kaggle.com/junkoda/handmade-features?scriptVersionId=10818864)() and train  it with  pseudo label and data augment.  Finally I got cv0.765,  0.690 (lb) and 0.679 (private).  To my regret, we did not choose this model.  : ( \n<br>\n\nI learned a lot of tricks and lesson from this competition. Thanks Kaggle.  Thanks every sharings in this competition.",
      "votes": null
    },
    {
      "id": "496291",
      "postDate": "03/22/2019 02:59:42",
      "content": "<p>Thanks qinhui1999 for sharing your experience!</p>\n\n<p>Could you please elaborate how you make pseudo labels and data augmentation specifically in this problem?  I tried also both of them but did not see obvious benefit.</p>\n\n<p>And what is your CV strategy? Is it correlated well to both private / public ?</p>\n\n<p>Lastly, would you mind sharing your thought on another issue here:\n<a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85167\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85167</a></p>",
      "rawMarkdown": "Thanks qinhui1999 for sharing your experience!\n\nCould you please elaborate how you make pseudo labels and data augmentation specifically in this problem?  I tried also both of them but did not see obvious benefit.\n\nAnd what is your CV strategy? Is it correlated well to both private / public ?\n\nLastly, would you mind sharing your thought on another issue here:\nhttps://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85167",
      "votes": null
    },
    {
      "id": "496308",
      "postDate": "03/22/2019 03:22:39",
      "content": "<p>I also used pseudo-labeling (just add test set to training with label = best current model prediction) and augmentation (random phase shift and random phase ordering).   For local CV, I repped every model with 5 seeds x 5 different kfold splits (k=8 mostly), and averaged all reps for test set predictions.</p>",
      "rawMarkdown": "I also used pseudo-labeling (just add test set to training with label = best current model prediction) and augmentation (random phase shift and random phase ordering).   For local CV, I repped every model with 5 seeds x 5 different kfold splits (k=8 mostly), and averaged all reps for test set predictions.",
      "votes": null
    },
    {
      "id": "496317",
      "postDate": "03/22/2019 03:34:07",
      "content": "<p>Thanks Russ for sharing! Did you use neural networks ?</p>",
      "rawMarkdown": "Thanks Russ for sharing! Did you use neural networks ?",
      "votes": null
    },
    {
      "id": "496324",
      "postDate": "03/22/2019 03:46:02",
      "content": "<p>For pseudo labels, we firstly generate some good performace test predictions (Through weigted voting, I got 0.744 lb score ). Then I use this 0.744 test prediction as the test labels and concate it with train data.  Finally I feed the concated datas into the model for training.   <br></p>\n\n<p>For data augmentation, in fact I had tried mixup data augmentation and it did not work.  After I see the data augment <a href=\"https://www.kaggle.com/jiweiliu/lgb-2-leaves-augment\">kernel</a>, I tried to use it and it worked.  The key of this data augmentation is shuffling the datas according to the y target.  I used it only for the train data, not for the validating data and test data. <br></p>\n\n<p>For cv strategy, I use the straightforward KFold 10folds.  Because the skf has a better score in lb than kfold.  Using the sfk 10 folds, pseudo labels and shuffle data augmentation, <br>\nmy rnn model lb score is from 0.692 to 0.741,  private score is from 0.636 to 0.642.<br>\nmy catboost models lb score is from 0.655 to 0.690, private score is from 0.647 to 0.679<br></p>\n\n<p>Sure, you are welcome to use my thought on your post.</p>",
      "rawMarkdown": "For pseudo labels, we firstly generate some good performace test predictions (Through weigted voting, I got 0.744 lb score ). Then I use this 0.744 test prediction as the test labels and concate it with train data.  Finally I feed the concated datas into the model for training.   <br>\n\nFor data augmentation, in fact I had tried mixup data augmentation and it did not work.  After I see the data augment [kernel](https://www.kaggle.com/jiweiliu/lgb-2-leaves-augment), I tried to use it and it worked.  The key of this data augmentation is shuffling the datas according to the y target.  I used it only for the train data, not for the validating data and test data. <br>\n\nFor cv strategy, I use the straightforward KFold 10folds.  Because the skf has a better score in lb than kfold.  Using the sfk 10 folds, pseudo labels and shuffle data augmentation, <br>\nmy rnn model lb score is from 0.692 to 0.741,  private score is from 0.636 to 0.642.<br>\nmy catboost models lb score is from 0.655 to 0.690, private score is from 0.647 to 0.679<br>\n\nSure, you are welcome to use my thought on your post.",
      "votes": null
    },
    {
      "id": "496328",
      "postDate": "03/22/2019 03:50:35",
      "content": "<p>Yes, plus some boosted tree models for diversity.</p>",
      "rawMarkdown": "Yes, plus some boosted tree models for diversity.",
      "votes": null
    },
    {
      "id": "496351",
      "postDate": "03/22/2019 04:15:58",
      "content": "<p>Thanks so much again!</p>",
      "rawMarkdown": "Thanks so much again!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 496291,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "03/22/2019 02:59:42",
      "content": "<p>Thanks qinhui1999 for sharing your experience!</p>\n\n<p>Could you please elaborate how you make pseudo labels and data augmentation specifically in this problem?  I tried also both of them but did not see obvious benefit.</p>\n\n<p>And what is your CV strategy? Is it correlated well to both private / public ?</p>\n\n<p>Lastly, would you mind sharing your thought on another issue here:\n<a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85167\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85167</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 496324,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "03/22/2019 03:46:02",
          "content": "<p>For pseudo labels, we firstly generate some good performace test predictions (Through weigted voting, I got 0.744 lb score ). Then I use this 0.744 test prediction as the test labels and concate it with train data.  Finally I feed the concated datas into the model for training.   <br></p>\n\n<p>For data augmentation, in fact I had tried mixup data augmentation and it did not work.  After I see the data augment <a href=\"https://www.kaggle.com/jiweiliu/lgb-2-leaves-augment\">kernel</a>, I tried to use it and it worked.  The key of this data augmentation is shuffling the datas according to the y target.  I used it only for the train data, not for the validating data and test data. <br></p>\n\n<p>For cv strategy, I use the straightforward KFold 10folds.  Because the skf has a better score in lb than kfold.  Using the sfk 10 folds, pseudo labels and shuffle data augmentation, <br>\nmy rnn model lb score is from 0.692 to 0.741,  private score is from 0.636 to 0.642.<br>\nmy catboost models lb score is from 0.655 to 0.690, private score is from 0.647 to 0.679<br></p>\n\n<p>Sure, you are welcome to use my thought on your post.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496351,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "03/22/2019 04:15:58",
          "content": "<p>Thanks so much again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 496308,
      "author_name": "sasrdw",
      "author_url": "",
      "post_date": "03/22/2019 03:22:39",
      "content": "<p>I also used pseudo-labeling (just add test set to training with label = best current model prediction) and augmentation (random phase shift and random phase ordering).   For local CV, I repped every model with 5 seeds x 5 different kfold splits (k=8 mostly), and averaged all reps for test set predictions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 496317,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "03/22/2019 03:34:07",
          "content": "<p>Thanks Russ for sharing! Did you use neural networks ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496328,
          "author_name": "sasrdw",
          "author_url": "",
          "post_date": "03/22/2019 03:50:35",
          "content": "<p>Yes, plus some boosted tree models for diversity.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "496235": "Congratulations to all the winners.  This competition is ver hard . We tried many ways to improve our scores, but most of them are overfitting.  There are still some useful tricks .\n<br>\n1.  pseudo label .   We use the pseudo label to retrain our models in the last two weeks.  It really improve the cv and lb scores (including private scores).  <br>\n\n2.  data augment.   In the last two days,  I see the data augment [kernel](https://www.kaggle.com/jiweiliu/lgb-2-leaves-augment) . Then I used it in our models.  It works and improve our cv and lb scores. <br>\n\n3.  Catboost model is better than LGB and RNN.   I used this [kernel](https://www.kaggle.com/junkoda/handmade-features?scriptVersionId=10818864)() and train  it with  pseudo label and data augment.  Finally I got cv0.765,  0.690 (lb) and 0.679 (private).  To my regret, we did not choose this model.  : ( \n<br>\n\nI learned a lot of tricks and lesson from this competition. Thanks Kaggle.  Thanks every sharings in this competition.",
    "496291": "Thanks qinhui1999 for sharing your experience!\n\nCould you please elaborate how you make pseudo labels and data augmentation specifically in this problem?  I tried also both of them but did not see obvious benefit.\n\nAnd what is your CV strategy? Is it correlated well to both private / public ?\n\nLastly, would you mind sharing your thought on another issue here:\nhttps://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85167",
    "496308": "I also used pseudo-labeling (just add test set to training with label = best current model prediction) and augmentation (random phase shift and random phase ordering).   For local CV, I repped every model with 5 seeds x 5 different kfold splits (k=8 mostly), and averaged all reps for test set predictions.",
    "496317": "Thanks Russ for sharing! Did you use neural networks ?",
    "496324": "For pseudo labels, we firstly generate some good performace test predictions (Through weigted voting, I got 0.744 lb score ). Then I use this 0.744 test prediction as the test labels and concate it with train data.  Finally I feed the concated datas into the model for training.   <br>\n\nFor data augmentation, in fact I had tried mixup data augmentation and it did not work.  After I see the data augment [kernel](https://www.kaggle.com/jiweiliu/lgb-2-leaves-augment), I tried to use it and it worked.  The key of this data augmentation is shuffling the datas according to the y target.  I used it only for the train data, not for the validating data and test data. <br>\n\nFor cv strategy, I use the straightforward KFold 10folds.  Because the skf has a better score in lb than kfold.  Using the sfk 10 folds, pseudo labels and shuffle data augmentation, <br>\nmy rnn model lb score is from 0.692 to 0.741,  private score is from 0.636 to 0.642.<br>\nmy catboost models lb score is from 0.655 to 0.690, private score is from 0.647 to 0.679<br>\n\nSure, you are welcome to use my thought on your post.",
    "496328": "Yes, plus some boosted tree models for diversity.",
    "496351": "Thanks so much again!"
  },
  "source": "meta"
}