{
  "id": 85331,
  "title": "What I tried to reduce high variance of models (also want to know yours)",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/85331",
  "author_name": "yuusha_aaaaa",
  "post_date": "2019-03-23T08:17:37.333000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>In this competition, I was suffering from the high variance of models (especially RNN/CNN). <br>\nTo reduce the variances, I tried some approaches.  </p>\n\n<h3>worked</h3>\n\n<ul>\n<li>hard voting of k-fold models</li>\n<li>small window size to extract feature</li>\n<li>ensemble of locally good and public LB</li>\n<li>original random seed different from public kernels <br>\n(Of course, I didn't tune the random seed)\nThis helps me a lot to judge whether the kernel models really work or get a better public LB score by luck.</li>\n</ul>\n\n<h3>worked partially</h3>\n\n<ul>\n<li>soft voting / naive bayes / svm / logistic regression of cv models <br>\n(inferior to hard voting)</li>\n<li>mix-up \n(worked for private LB, not worked)</li>\n<li>bagging + under-sampling of class 0 <br>\nworked, generate the most robust model, but lower LB scores than those wo bagging</li>\n<li><a href=\"https://www.kaggle.com/yatzhash/smote-to-learn-from-a-few-anomaly-sample/edit\">SMOTH (worked for private LB, but worsen public LB)</a></li>\n</ul>\n\n<h3>not worked at all</h3>\n\n<ul>\n<li>fine-tuning pre-trained model with ImageNet. <br>\nThe signals are convert into images by plotting them. The models are inferior to the non-image models. <br>\nThis plotting approach is introduced in <a href=\"https://www.sciencedirect.com/science/article/pii/S2213158219300348\">this eeg paper</a>\nI can't find any pre-trained models by signal datasets.   </li>\n<li>cosine loss <br>\nnot worked at all, training didn't progress. It is possible my implementation of the cosine loss isn't correct. <br>\nabout cosine loss, please refer to <a href=\"https://arxiv.org/abs/1901.09054\">this paper</a></li>\n<li>multiple dropouts <br>\nintroduced in <a href=\"https://www.ijcai.org/proceedings/2017/0318.pdf\">Deep Neural Networks for High Dimension, Low Sample Size Data</a> <br>\nIt caused over-fitting to local data and public LB</li>\n<li>threshold tuning <br>\ncaused over-fitting</li>\n<li>data augmentation by splitting a signal into 2 parts and swapping them <br>\nnot worked at all, causing over-fitting to augmented samples  </li>\n</ul>\n\n<p><br>\nI also want to know what approaches of yours worked and didn't to reduce the model variances. </p>",
  "messages": [
    {
      "id": 497265,
      "postDate": "2019-03-23T08:17:37.333Z",
      "content": "<p>In this competition, I was suffering from the high variance of models (especially RNN/CNN). <br>\nTo reduce the variances, I tried some approaches.  </p>\n\n<h3>worked</h3>\n\n<ul>\n<li>hard voting of k-fold models</li>\n<li>small window size to extract feature</li>\n<li>ensemble of locally good and public LB</li>\n<li>original random seed different from public kernels <br>\n(Of course, I didn't tune the random seed)\nThis helps me a lot to judge whether the kernel models really work or get a better public LB score by luck.</li>\n</ul>\n\n<h3>worked partially</h3>\n\n<ul>\n<li>soft voting / naive bayes / svm / logistic regression of cv models <br>\n(inferior to hard voting)</li>\n<li>mix-up \n(worked for private LB, not worked)</li>\n<li>bagging + under-sampling of class 0 <br>\nworked, generate the most robust model, but lower LB scores than those wo bagging</li>\n<li><a href=\"https://www.kaggle.com/yatzhash/smote-to-learn-from-a-few-anomaly-sample/edit\">SMOTH (worked for private LB, but worsen public LB)</a></li>\n</ul>\n\n<h3>not worked at all</h3>\n\n<ul>\n<li>fine-tuning pre-trained model with ImageNet. <br>\nThe signals are convert into images by plotting them. The models are inferior to the non-image models. <br>\nThis plotting approach is introduced in <a href=\"https://www.sciencedirect.com/science/article/pii/S2213158219300348\">this eeg paper</a>\nI can't find any pre-trained models by signal datasets.   </li>\n<li>cosine loss <br>\nnot worked at all, training didn't progress. It is possible my implementation of the cosine loss isn't correct. <br>\nabout cosine loss, please refer to <a href=\"https://arxiv.org/abs/1901.09054\">this paper</a></li>\n<li>multiple dropouts <br>\nintroduced in <a href=\"https://www.ijcai.org/proceedings/2017/0318.pdf\">Deep Neural Networks for High Dimension, Low Sample Size Data</a> <br>\nIt caused over-fitting to local data and public LB</li>\n<li>threshold tuning <br>\ncaused over-fitting</li>\n<li>data augmentation by splitting a signal into 2 parts and swapping them <br>\nnot worked at all, causing over-fitting to augmented samples  </li>\n</ul>\n\n<p><br>\nI also want to know what approaches of yours worked and didn't to reduce the model variances. </p>",
      "rawMarkdown": "\nIn this competition, I was suffering from the high variance of models (especially RNN/CNN).  \nTo reduce the variances, I tried some approaches.  \n\n### worked\n\n- hard voting of k-fold models\n- small window size to extract feature\n- ensemble of locally good and public LB\n- original random seed different from public kernels  \n    (Of course, I didn't tune the random seed)\n    This helps me a lot to judge whether the kernel models really work or get a better public LB score by luck.\n\n### worked partially \n\n- soft voting / naive bayes / svm / logistic regression of cv models  \n    (inferior to hard voting)\n- mix-up \n    (worked for private LB, not worked)\n- bagging + under-sampling of class 0  \n    worked, generate the most robust model, but lower LB scores than those wo bagging\n- [SMOTH (worked for private LB, but worsen public LB)](https://www.kaggle.com/yatzhash/smote-to-learn-from-a-few-anomaly-sample/edit)\n\n### not worked at all\n- fine-tuning pre-trained model with ImageNet.  \n  The signals are convert into images by plotting them. The models are inferior to the non-image models.  \n  This plotting approach is introduced in [this eeg paper](https://www.sciencedirect.com/science/article/pii/S2213158219300348)\n  I can't find any pre-trained models by signal datasets.   \n- cosine loss  \n  not worked at all, training didn't progress. It is possible my implementation of the cosine loss isn't correct.  \n  about cosine loss, please refer to [this paper](https://arxiv.org/abs/1901.09054)\n- multiple dropouts  \n  introduced in [Deep Neural Networks for High Dimension, Low Sample Size Data](https://www.ijcai.org/proceedings/2017/0318.pdf)  \n  It caused over-fitting to local data and public LB\n- threshold tuning  \n    caused over-fitting\n- data augmentation by splitting a signal into 2 parts and swapping them  \n    not worked at all, causing over-fitting to augmented samples  \n\n<br>\nI also want to know what approaches of yours worked and didn't to reduce the model variances. ",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "497265": "\nIn this competition, I was suffering from the high variance of models (especially RNN/CNN).  \nTo reduce the variances, I tried some approaches.  \n\n### worked\n\n- hard voting of k-fold models\n- small window size to extract feature\n- ensemble of locally good and public LB\n- original random seed different from public kernels  \n    (Of course, I didn't tune the random seed)\n    This helps me a lot to judge whether the kernel models really work or get a better public LB score by luck.\n\n### worked partially \n\n- soft voting / naive bayes / svm / logistic regression of cv models  \n    (inferior to hard voting)\n- mix-up \n    (worked for private LB, not worked)\n- bagging + under-sampling of class 0  \n    worked, generate the most robust model, but lower LB scores than those wo bagging\n- [SMOTH (worked for private LB, but worsen public LB)](https://www.kaggle.com/yatzhash/smote-to-learn-from-a-few-anomaly-sample/edit)\n\n### not worked at all\n- fine-tuning pre-trained model with ImageNet.  \n  The signals are convert into images by plotting them. The models are inferior to the non-image models.  \n  This plotting approach is introduced in [this eeg paper](https://www.sciencedirect.com/science/article/pii/S2213158219300348)\n  I can't find any pre-trained models by signal datasets.   \n- cosine loss  \n  not worked at all, training didn't progress. It is possible my implementation of the cosine loss isn't correct.  \n  about cosine loss, please refer to [this paper](https://arxiv.org/abs/1901.09054)\n- multiple dropouts  \n  introduced in [Deep Neural Networks for High Dimension, Low Sample Size Data](https://www.ijcai.org/proceedings/2017/0318.pdf)  \n  It caused over-fitting to local data and public LB\n- threshold tuning  \n    caused over-fitting\n- data augmentation by splitting a signal into 2 parts and swapping them  \n    not worked at all, causing over-fitting to augmented samples  \n\n<br>\nI also want to know what approaches of yours worked and didn't to reduce the model variances. "
  }
}