{
  "id": 85220,
  "title": "single silver models (hard voting, mixup, SEResNet)",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/85220",
  "author_name": "yuusha_aaaaa",
  "post_date": "2019-03-22T09:59:15.872000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I share two of my single \"possibly\" silver models.</p>\n\n<p>To my disappointment, I don't select these model as final model, \nso I missed the silver medal because I underestimated the local scores. <br>\nHowever,  these models contributed to my final ensemble ones and brought me the bronze medal. <br>\nThese 2 models are inspired by <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge\">TensorFlow Speech Recognition Challenge</a>  I joined in past.   </p>\n\n<ol>\n<li><p><strong>bigru + attention +  mix-up + hard voting  (private LB 0.668 (same as 38th) / public LB 0.568)</strong>\nIt's similar model with <a href=\"https://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694\">Bruno Aquino's kenel</a>, but 2 main different points.  </p>\n\n<ul><li><em>hard voting of each cv model's classification (threshold=0.5)</em>\nInstead of tuning the threshold with the averaged probabilities  of those by cv models, \nthe classes of the samples are decided by majority of the cv models with the fixed threshold 0.5. I guess this hard voting bring the model robustness.   </li>\n<li><em>mix-up</em>\nThis augmentation method is introduced by the paper <a href=\"https://arxiv.org/abs/1710.09412\">mixup: Beyond Empirical Risk Minimization</a>. <br>\nThe augmentation generates a new sample from 2 training samples based of the adding. <br>\n<code>math\n\\tilde{x} = λx_i +(1−λ)x_j\n\\tilde{y} = λy_i +(1−λ)y_j\n</code>\nIf  \\tilde{y} &gt;= 0.5 \\, I labeled the new sample as class 1. <br>\nI used 3 as the mix-up ratio, but I don't have enough time to tune the ratio. \nPlease see the blog post [the detail](<a href=\"http://www.inference.vc/mixup-data-dependent-data-\">http://www.inference.vc/mixup-data-dependent-data-</a> \naugmentation/) about mix-up. <br>\nThe mix-up improved private LB to 0.668 from 0.599 but worsened public LB to 0.568 from \n0.632. Therefore, it requires more careful evaluation to know whether mix-up really works.  </li></ul>\n\n<p>In addition, I implement it with PyTorch instead of Keras (I don't know what effect that brings the model performance. ) <br>\n<br></p></li>\n<li><p><strong>SE-ResNet 50 + window (chunk) 40 step 20 summary feature +  hard voting (private LB 0.663 (same as 53th) / public LB 0.587)</strong>\nThis model composed of 2 specific features.  </p>\n\n<ul><li><p><em>2D-1D SE-Resnet 50</em>\nI learned cnn works well for signal classification tasks in the above speech recognition \ncompetition.  That drives me to adopt cnn models for the power line signals. <br>\nInstead of simple cnn models, I used SE-ResNet 50 expecting it to find more deep features. <br>\n(I also tried more simple CNN models with attention, but they are inferior to SE-ResNet in both public/private LB. )  </p>\n\n<p>To support signal inputs, I changed 3 points from original 2D SE-ResNet for image.</p>\n\n<ul><li>input concatenated features from 3 phase in height axis.</li>\n<li>2D 1st conv with kernel size = (n feature(height),  your kernel size for time step axis) . It weaves the features into 1D time steps and the channels. </li>\n<li>following 1D conv blocks </li></ul></li></ul>\n\n<p>These steps are inspired by <a href=\"https://arxiv.org/abs/1703.05051\">Deep Learning With Convolutional Neural Networks for EEG Decoding and Visualization by Schirrmeister etc.</a> and  <a href=\"https://arxiv.org/abs/1611.01942\">DeepSense: A Unified Deep Learning Framework for Time-Series Mobile Sensing Data Processing by Yao etc.\narXiv:1611.01942v2</a>.  </p>\n\n<ul><li><em>smaller window size to extract summarized feature</em>\nIn the public kernels, the features are extracted with window size 5000. <br>\nHowever, I want to leave feature extracting to cnn as much as possible . <br>\nTherefore the inputted features are extracted with more smaller window (40) and step (20) size. <br>\nI also tried to train the model with the raw 3 phase signals or the features from smaller window \nsizes, but I cannot do due to oom.  </li></ul>\n\n<p>The SE-ResNet also seems over-fitting to privte LB and it leaves improvements for the model parameters.  </p></li>\n</ol>\n\n<p>In conclusion,  the both of them seem to include some hints but requires more improvements \nto obtain the high scores in both private and public LB.\nI'll share my codes in near feature.</p>",
  "messages": [
    {
      "id": 496534,
      "postDate": "2019-03-22T09:59:15.873Z",
      "content": "<p>I share two of my single \"possibly\" silver models.</p>\n\n<p>To my disappointment, I don't select these model as final model, \nso I missed the silver medal because I underestimated the local scores. <br>\nHowever,  these models contributed to my final ensemble ones and brought me the bronze medal. <br>\nThese 2 models are inspired by <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge\">TensorFlow Speech Recognition Challenge</a>  I joined in past.   </p>\n\n<ol>\n<li><p><strong>bigru + attention +  mix-up + hard voting  (private LB 0.668 (same as 38th) / public LB 0.568)</strong>\nIt's similar model with <a href=\"https://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694\">Bruno Aquino's kenel</a>, but 2 main different points.  </p>\n\n<ul><li><em>hard voting of each cv model's classification (threshold=0.5)</em>\nInstead of tuning the threshold with the averaged probabilities  of those by cv models, \nthe classes of the samples are decided by majority of the cv models with the fixed threshold 0.5. I guess this hard voting bring the model robustness.   </li>\n<li><em>mix-up</em>\nThis augmentation method is introduced by the paper <a href=\"https://arxiv.org/abs/1710.09412\">mixup: Beyond Empirical Risk Minimization</a>. <br>\nThe augmentation generates a new sample from 2 training samples based of the adding. <br>\n<code>math\n\\tilde{x} = λx_i +(1−λ)x_j\n\\tilde{y} = λy_i +(1−λ)y_j\n</code>\nIf  \\tilde{y} &gt;= 0.5 \\, I labeled the new sample as class 1. <br>\nI used 3 as the mix-up ratio, but I don't have enough time to tune the ratio. \nPlease see the blog post [the detail](<a href=\"http://www.inference.vc/mixup-data-dependent-data-\">http://www.inference.vc/mixup-data-dependent-data-</a> \naugmentation/) about mix-up. <br>\nThe mix-up improved private LB to 0.668 from 0.599 but worsened public LB to 0.568 from \n0.632. Therefore, it requires more careful evaluation to know whether mix-up really works.  </li></ul>\n\n<p>In addition, I implement it with PyTorch instead of Keras (I don't know what effect that brings the model performance. ) <br>\n<br></p></li>\n<li><p><strong>SE-ResNet 50 + window (chunk) 40 step 20 summary feature +  hard voting (private LB 0.663 (same as 53th) / public LB 0.587)</strong>\nThis model composed of 2 specific features.  </p>\n\n<ul><li><p><em>2D-1D SE-Resnet 50</em>\nI learned cnn works well for signal classification tasks in the above speech recognition \ncompetition.  That drives me to adopt cnn models for the power line signals. <br>\nInstead of simple cnn models, I used SE-ResNet 50 expecting it to find more deep features. <br>\n(I also tried more simple CNN models with attention, but they are inferior to SE-ResNet in both public/private LB. )  </p>\n\n<p>To support signal inputs, I changed 3 points from original 2D SE-ResNet for image.</p>\n\n<ul><li>input concatenated features from 3 phase in height axis.</li>\n<li>2D 1st conv with kernel size = (n feature(height),  your kernel size for time step axis) . It weaves the features into 1D time steps and the channels. </li>\n<li>following 1D conv blocks </li></ul></li></ul>\n\n<p>These steps are inspired by <a href=\"https://arxiv.org/abs/1703.05051\">Deep Learning With Convolutional Neural Networks for EEG Decoding and Visualization by Schirrmeister etc.</a> and  <a href=\"https://arxiv.org/abs/1611.01942\">DeepSense: A Unified Deep Learning Framework for Time-Series Mobile Sensing Data Processing by Yao etc.\narXiv:1611.01942v2</a>.  </p>\n\n<ul><li><em>smaller window size to extract summarized feature</em>\nIn the public kernels, the features are extracted with window size 5000. <br>\nHowever, I want to leave feature extracting to cnn as much as possible . <br>\nTherefore the inputted features are extracted with more smaller window (40) and step (20) size. <br>\nI also tried to train the model with the raw 3 phase signals or the features from smaller window \nsizes, but I cannot do due to oom.  </li></ul>\n\n<p>The SE-ResNet also seems over-fitting to privte LB and it leaves improvements for the model parameters.  </p></li>\n</ol>\n\n<p>In conclusion,  the both of them seem to include some hints but requires more improvements \nto obtain the high scores in both private and public LB.\nI'll share my codes in near feature.</p>",
      "rawMarkdown": "I share two of my single \"possibly\" silver models.\n\nTo my disappointment, I don't select these model as final model, \nso I missed the silver medal because I underestimated the local scores.    \nHowever,  these models contributed to my final ensemble ones and brought me the bronze medal.   \nThese 2 models are inspired by [TensorFlow Speech Recognition Challenge](https://www.kaggle.com/c/tensorflow-speech-recognition-challenge)  I joined in past.   \n\n\n1.   **bigru + attention +  mix-up + hard voting  (private LB 0.668 (same as 38th) / public LB 0.568)**\n    It's similar model with [Bruno Aquino's kenel](https://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694), but 2 main different points.  \n   \n  -  *hard voting of each cv model's classification (threshold=0.5)*\n     Instead of tuning the threshold with the averaged probabilities  of those by cv models, \n     the classes of the samples are decided by majority of the cv models with the fixed threshold 0.5. I guess this hard voting bring the model robustness.   \n  - *mix-up*\n    This augmentation method is introduced by the paper [mixup: Beyond Empirical Risk Minimization](https://arxiv.org/abs/1710.09412).  \n    The augmentation generates a new sample from 2 training samples based of the adding.  \n    ```math\n       \\tilde{x} = λx_i +(1−λ)x_j\n       \\tilde{y} = λy_i +(1−λ)y_j\n   ```\n    If  \\\\tilde{y} &gt;= 0.5 \\\\, I labeled the new sample as class 1.  \n    I used 3 as the mix-up ratio, but I don't have enough time to tune the ratio. \n Please see the blog post [the detail](http://www.inference.vc/mixup-data-dependent-data- \n augmentation/) about mix-up.  \n The mix-up improved private LB to 0.668 from 0.599 but worsened public LB to 0.568 from \n 0.632. Therefore, it requires more careful evaluation to know whether mix-up really works.  \n  \n  In addition, I implement it with PyTorch instead of Keras (I don't know what effect that brings the model performance. )  \n<br>\n\n2.  **SE-ResNet 50 + window (chunk) 40 step 20 summary feature +  hard voting (private LB 0.663 (same as 53th) / public LB 0.587)**\n  This model composed of 2 specific features.  \n\n  - *2D-1D SE-Resnet 50*\n     I learned cnn works well for signal classification tasks in the above speech recognition \n     competition.  That drives me to adopt cnn models for the power line signals.  \n     Instead of simple cnn models, I used SE-ResNet 50 expecting it to find more deep features.  \n    (I also tried more simple CNN models with attention, but they are inferior to SE-ResNet in both public/private LB. )  \n    \n      To support signal inputs, I changed 3 points from original 2D SE-ResNet for image.\n    \n      - input concatenated features from 3 phase in height axis.\n      - 2D 1st conv with kernel size = (n feature(height),  your kernel size for time step axis) . It weaves the features into 1D time steps and the channels. \n     - following 1D conv blocks \n \n    These steps are inspired by [Deep Learning With Convolutional Neural Networks for EEG Decoding and Visualization by Schirrmeister etc.](https://arxiv.org/abs/1703.05051) and  [DeepSense: A Unified Deep Learning Framework for Time-Series Mobile Sensing Data Processing by Yao etc.\narXiv:1611.01942v2](https://arxiv.org/abs/1611.01942).  \n    \n  - *smaller window size to extract summarized feature*\n    In the public kernels, the features are extracted with window size 5000.    \n    However, I want to leave feature extracting to cnn as much as possible .  \n   Therefore the inputted features are extracted with more smaller window (40) and step (20) size.  \n    I also tried to train the model with the raw 3 phase signals or the features from smaller window \n   sizes, but I cannot do due to oom.  \n  \n   The SE-ResNet also seems over-fitting to privte LB and it leaves improvements for the model parameters.  \n\n\nIn conclusion,  the both of them seem to include some hints but requires more improvements \nto obtain the high scores in both private and public LB.\nI'll share my codes in near feature.\n\n\n\n\n\n\n\n",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "496534": "I share two of my single \"possibly\" silver models.\n\nTo my disappointment, I don't select these model as final model, \nso I missed the silver medal because I underestimated the local scores.    \nHowever,  these models contributed to my final ensemble ones and brought me the bronze medal.   \nThese 2 models are inspired by [TensorFlow Speech Recognition Challenge](https://www.kaggle.com/c/tensorflow-speech-recognition-challenge)  I joined in past.   \n\n\n1.   **bigru + attention +  mix-up + hard voting  (private LB 0.668 (same as 38th) / public LB 0.568)**\n    It's similar model with [Bruno Aquino's kenel](https://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694), but 2 main different points.  \n   \n  -  *hard voting of each cv model's classification (threshold=0.5)*\n     Instead of tuning the threshold with the averaged probabilities  of those by cv models, \n     the classes of the samples are decided by majority of the cv models with the fixed threshold 0.5. I guess this hard voting bring the model robustness.   \n  - *mix-up*\n    This augmentation method is introduced by the paper [mixup: Beyond Empirical Risk Minimization](https://arxiv.org/abs/1710.09412).  \n    The augmentation generates a new sample from 2 training samples based of the adding.  \n    ```math\n       \\tilde{x} = λx_i +(1−λ)x_j\n       \\tilde{y} = λy_i +(1−λ)y_j\n   ```\n    If  \\\\tilde{y} &gt;= 0.5 \\\\, I labeled the new sample as class 1.  \n    I used 3 as the mix-up ratio, but I don't have enough time to tune the ratio. \n Please see the blog post [the detail](http://www.inference.vc/mixup-data-dependent-data- \n augmentation/) about mix-up.  \n The mix-up improved private LB to 0.668 from 0.599 but worsened public LB to 0.568 from \n 0.632. Therefore, it requires more careful evaluation to know whether mix-up really works.  \n  \n  In addition, I implement it with PyTorch instead of Keras (I don't know what effect that brings the model performance. )  \n<br>\n\n2.  **SE-ResNet 50 + window (chunk) 40 step 20 summary feature +  hard voting (private LB 0.663 (same as 53th) / public LB 0.587)**\n  This model composed of 2 specific features.  \n\n  - *2D-1D SE-Resnet 50*\n     I learned cnn works well for signal classification tasks in the above speech recognition \n     competition.  That drives me to adopt cnn models for the power line signals.  \n     Instead of simple cnn models, I used SE-ResNet 50 expecting it to find more deep features.  \n    (I also tried more simple CNN models with attention, but they are inferior to SE-ResNet in both public/private LB. )  \n    \n      To support signal inputs, I changed 3 points from original 2D SE-ResNet for image.\n    \n      - input concatenated features from 3 phase in height axis.\n      - 2D 1st conv with kernel size = (n feature(height),  your kernel size for time step axis) . It weaves the features into 1D time steps and the channels. \n     - following 1D conv blocks \n \n    These steps are inspired by [Deep Learning With Convolutional Neural Networks for EEG Decoding and Visualization by Schirrmeister etc.](https://arxiv.org/abs/1703.05051) and  [DeepSense: A Unified Deep Learning Framework for Time-Series Mobile Sensing Data Processing by Yao etc.\narXiv:1611.01942v2](https://arxiv.org/abs/1611.01942).  \n    \n  - *smaller window size to extract summarized feature*\n    In the public kernels, the features are extracted with window size 5000.    \n    However, I want to leave feature extracting to cnn as much as possible .  \n   Therefore the inputted features are extracted with more smaller window (40) and step (20) size.  \n    I also tried to train the model with the raw 3 phase signals or the features from smaller window \n   sizes, but I cannot do due to oom.  \n  \n   The SE-ResNet also seems over-fitting to privte LB and it leaves improvements for the model parameters.  \n\n\nIn conclusion,  the both of them seem to include some hints but requires more improvements \nto obtain the high scores in both private and public LB.\nI'll share my codes in near feature.\n\n\n\n\n\n\n\n"
  }
}