{
  "id": 134929,
  "title": "Worth Seeing Skills(Classification)",
  "url": "/competitions/bengaliai-cv19/discussion/134929",
  "author_name": "",
  "post_date": "2020-03-11T06:39:15.538546500Z",
  "votes": 15,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi, Im new to Kaggle competitions, and late for this match. But I will boost my grade as high as I can. \nHope everyone have a nice grade.</p>\n\n<h2>Augmentation</h2>\n\n<p>ref: <br>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132492\">[list of all mixup variant]</a>  </p>\n\n<ul>\n<li><p>cutout <br>\n\"Improved Regularization of Convolutional Neural Networks with Cutout\" - Terrance DeVries, arvix 2017\n<a href=\"https://arxiv.org/abs/1708.04552\">https://arxiv.org/abs/1708.04552</a></p></li>\n<li><p>mixup <br>\n\"mixup: Beyond Empirical Risk Minimization\" - Hongyi Zhang, arvix 2017\n<a href=\"https://arxiv.org/abs/1710.09412\">https://arxiv.org/abs/1710.09412</a>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128115#763163\">[code]</a>  </p></li>\n<li><p>manifold mixup <br>\n\"Manifold Mixup: Better Representations by Interpolating Hidden States\" - Vikas Verma, arvix 2018\n<a href=\"https://arxiv.org/abs/1806.05236\">https://arxiv.org/abs/1806.05236</a></p></li>\n<li><p>cutmix <br>\n\"CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features\" - Sangdoo Yun, iccv 2019\n<a href=\"https://arxiv.org/abs/1905.04899\">https://arxiv.org/abs/1905.04899</a>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/126504\">[cutmix is all you need code]</a>  </p></li>\n<li><p>augmix <br>\n\"AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty\" - Dan Hendrycks, arvix 2019\n<a href=\"https://arxiv.org/abs/1912.02781\">https://arxiv.org/abs/1912.02781</a></p></li>\n<li><p>gridmask <br>\n\"GridMask Data Augmentation\" - Pengguang Chen, arvix 2020\n<a href=\"https://arxiv.org/abs/2001.04086\">https://arxiv.org/abs/2001.04086</a></p></li>\n<li><p>dropblock <br>\n\"DropBlock: A regularization method for convolutional networks\" - Golnaz Ghiasi, nips 2019\n<a href=\"https://github.com/sujatasaini/Kuzushiji-DropBlock\">[code]</a>  </p></li>\n<li><p>fmix \n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/133322\">[paper+code]</a></p></li>\n<li><p>batchboost: regularization for stabilizing training with resistance to underfitting &amp; overfitting\nMaciej A. Czyzewski</p></li>\n<li><p>MaxUp: A Simple Way to Improve Generalization of Neural Network Training <br>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757\">[ref]</a> <br>\n```python\nMethod    Top-1 error Top-5 error\nVanilla (He et al., 2016a)    76.3    -\nDropout (Srivastava et al., 2014)    76.8    93.4\nDropPath (Larsson et al., 2017)    77.1    93.5\nManifold Mixup (Verma et al., 2019)    77.5    93.8\nAutoAugment (Cubuk et al., 2019a)    77.6    93.8\nMixup (Zhang et al., 2018)    77.9    93.9\nDropBlock (Ghiasi et al., 2018)    78.3    94.1\nCutMix (Yun et al., 2019)    78.6    94.0\nMaxUp+CutMix    78.9    94.2</p></li>\n</ul>\n\n<p>Table 1: Summary of top1 and top5 accuracies on the validation set of ImageNet for ResNet-50.\n```</p>\n\n<ul>\n<li><p>report on experiments on \"No Augmentation\" <br>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132898\">[ref from Heng]</a>  </p></li>\n<li><p>(2020)  SuperMix <br>\n<a href=\"https://github.com/alldbi/SuperMix\">[code]</a></p></li>\n</ul>\n\n<h2>Model</h2>\n\n<ul>\n<li>Se-Resnext50/101</li>\n<li>EfficientNet-b2</li>\n<li>GhostNet</li>\n<li>[TPAMI2020]Res2Net <a href=\"https://arxiv.org/abs/1904.01169\">[paper]</a>  <a href=\"https://mmcheng.net/res2net/\">[pretrained]</a>  </li>\n</ul>\n\n<h2>Tricks</h2>\n\n<ul>\n<li><p><a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/127976\">[worth seeing posts and tricks]</a>  </p></li>\n<li><p>Half precision</p></li>\n<li><p>channel_1 to channel_3\n```python\ninput1ch = Input(shape=(224, 224, 1))\ninput3ch = Conv2D(3, (1, 1))(input1ch)    #1 -&gt; 3 channels</p></li>\n</ul>\n\n<h1>or</h1>\n\n<p>def forward(self, x):\n    h = x.repeat(1,3,1,1)\n    h = self.base_model.features(h)</p>\n\n<p>```</p>\n\n<ul>\n<li><p>pretrained <br>\n```python</p>\n\n<ol><li><p>add a 1 to 3 conversion conv2d layer (or many layer)</p></li>\n<li><p>freeze all imagenet pretrained weights. Now only the conversion conv2d and the last fc logit layer are trainable.</p></li>\n<li><p>start to train. this force the conversion conv2d to produce activation signal to match the magnitude that is suitable for the pretrained weights.</p></li>\n<li><p>when there can be not more improvement, unfreeze everything and do the training as usual.\n```</p></li></ol></li>\n<li><p>OHEM <br>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">[OHEM+label smoothing]</a>  </p></li>\n<li><p>Label Smoothing <br>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128115#763163\">[Label Smoothing]</a>  </p></li>\n<li><p>CamDrop: A New Explanation of Dropout and A Guided Regularization Method for Deep Neural Networks</p></li>\n</ul>\n\n<h2>Infer</h2>\n\n<ul>\n<li>K-fold</li>\n<li>StritifiedKFold </li>\n<li>TTA <br>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/126648\">[TTA-Multiple_models]</a>  </li>\n<li><p>Ensemble   </p></li>\n<li><p>new validation (unseen)\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134434\">[discussion]</a></p></li>\n</ul>",
  "messages": [
    {
      "id": "768754",
      "postDate": "03/11/2020 06:39:15",
      "content": "<p>Hi, Im new to Kaggle competitions, and late for this match. But I will boost my grade as high as I can. \nHope everyone have a nice grade.</p>\n\n<h2>Augmentation</h2>\n\n<p>ref: <br>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132492\">[list of all mixup variant]</a>  </p>\n\n<ul>\n<li><p>cutout <br>\n\"Improved Regularization of Convolutional Neural Networks with Cutout\" - Terrance DeVries, arvix 2017\n<a href=\"https://arxiv.org/abs/1708.04552\">https://arxiv.org/abs/1708.04552</a></p></li>\n<li><p>mixup <br>\n\"mixup: Beyond Empirical Risk Minimization\" - Hongyi Zhang, arvix 2017\n<a href=\"https://arxiv.org/abs/1710.09412\">https://arxiv.org/abs/1710.09412</a>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128115#763163\">[code]</a>  </p></li>\n<li><p>manifold mixup <br>\n\"Manifold Mixup: Better Representations by Interpolating Hidden States\" - Vikas Verma, arvix 2018\n<a href=\"https://arxiv.org/abs/1806.05236\">https://arxiv.org/abs/1806.05236</a></p></li>\n<li><p>cutmix <br>\n\"CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features\" - Sangdoo Yun, iccv 2019\n<a href=\"https://arxiv.org/abs/1905.04899\">https://arxiv.org/abs/1905.04899</a>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/126504\">[cutmix is all you need code]</a>  </p></li>\n<li><p>augmix <br>\n\"AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty\" - Dan Hendrycks, arvix 2019\n<a href=\"https://arxiv.org/abs/1912.02781\">https://arxiv.org/abs/1912.02781</a></p></li>\n<li><p>gridmask <br>\n\"GridMask Data Augmentation\" - Pengguang Chen, arvix 2020\n<a href=\"https://arxiv.org/abs/2001.04086\">https://arxiv.org/abs/2001.04086</a></p></li>\n<li><p>dropblock <br>\n\"DropBlock: A regularization method for convolutional networks\" - Golnaz Ghiasi, nips 2019\n<a href=\"https://github.com/sujatasaini/Kuzushiji-DropBlock\">[code]</a>  </p></li>\n<li><p>fmix \n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/133322\">[paper+code]</a></p></li>\n<li><p>batchboost: regularization for stabilizing training with resistance to underfitting &amp; overfitting\nMaciej A. Czyzewski</p></li>\n<li><p>MaxUp: A Simple Way to Improve Generalization of Neural Network Training <br>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757\">[ref]</a> <br>\n```python\nMethod    Top-1 error Top-5 error\nVanilla (He et al., 2016a)    76.3    -\nDropout (Srivastava et al., 2014)    76.8    93.4\nDropPath (Larsson et al., 2017)    77.1    93.5\nManifold Mixup (Verma et al., 2019)    77.5    93.8\nAutoAugment (Cubuk et al., 2019a)    77.6    93.8\nMixup (Zhang et al., 2018)    77.9    93.9\nDropBlock (Ghiasi et al., 2018)    78.3    94.1\nCutMix (Yun et al., 2019)    78.6    94.0\nMaxUp+CutMix    78.9    94.2</p></li>\n</ul>\n\n<p>Table 1: Summary of top1 and top5 accuracies on the validation set of ImageNet for ResNet-50.\n```</p>\n\n<ul>\n<li><p>report on experiments on \"No Augmentation\" <br>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132898\">[ref from Heng]</a>  </p></li>\n<li><p>(2020)  SuperMix <br>\n<a href=\"https://github.com/alldbi/SuperMix\">[code]</a></p></li>\n</ul>\n\n<h2>Model</h2>\n\n<ul>\n<li>Se-Resnext50/101</li>\n<li>EfficientNet-b2</li>\n<li>GhostNet</li>\n<li>[TPAMI2020]Res2Net <a href=\"https://arxiv.org/abs/1904.01169\">[paper]</a>  <a href=\"https://mmcheng.net/res2net/\">[pretrained]</a>  </li>\n</ul>\n\n<h2>Tricks</h2>\n\n<ul>\n<li><p><a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/127976\">[worth seeing posts and tricks]</a>  </p></li>\n<li><p>Half precision</p></li>\n<li><p>channel_1 to channel_3\n```python\ninput1ch = Input(shape=(224, 224, 1))\ninput3ch = Conv2D(3, (1, 1))(input1ch)    #1 -&gt; 3 channels</p></li>\n</ul>\n\n<h1>or</h1>\n\n<p>def forward(self, x):\n    h = x.repeat(1,3,1,1)\n    h = self.base_model.features(h)</p>\n\n<p>```</p>\n\n<ul>\n<li><p>pretrained <br>\n```python</p>\n\n<ol><li><p>add a 1 to 3 conversion conv2d layer (or many layer)</p></li>\n<li><p>freeze all imagenet pretrained weights. Now only the conversion conv2d and the last fc logit layer are trainable.</p></li>\n<li><p>start to train. this force the conversion conv2d to produce activation signal to match the magnitude that is suitable for the pretrained weights.</p></li>\n<li><p>when there can be not more improvement, unfreeze everything and do the training as usual.\n```</p></li></ol></li>\n<li><p>OHEM <br>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">[OHEM+label smoothing]</a>  </p></li>\n<li><p>Label Smoothing <br>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128115#763163\">[Label Smoothing]</a>  </p></li>\n<li><p>CamDrop: A New Explanation of Dropout and A Guided Regularization Method for Deep Neural Networks</p></li>\n</ul>\n\n<h2>Infer</h2>\n\n<ul>\n<li>K-fold</li>\n<li>StritifiedKFold </li>\n<li>TTA <br>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/126648\">[TTA-Multiple_models]</a>  </li>\n<li><p>Ensemble   </p></li>\n<li><p>new validation (unseen)\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134434\">[discussion]</a></p></li>\n</ul>",
      "rawMarkdown": "Hi, Im new to Kaggle competitions, and late for this match. But I will boost my grade as high as I can. \nHope everyone have a nice grade.\n\n## Augmentation\n\nref:  \n[[list of all mixup variant]](https://www.kaggle.com/c/bengaliai-cv19/discussion/132492)  \n\n\n- cutout  \n\"Improved Regularization of Convolutional Neural Networks with Cutout\" - Terrance DeVries, arvix 2017\nhttps://arxiv.org/abs/1708.04552\n\n- mixup  \n\"mixup: Beyond Empirical Risk Minimization\" - Hongyi Zhang, arvix 2017\nhttps://arxiv.org/abs/1710.09412\n[[code]](https://www.kaggle.com/c/bengaliai-cv19/discussion/128115#763163)  \n\n\n- manifold mixup  \n\"Manifold Mixup: Better Representations by Interpolating Hidden States\" - Vikas Verma, arvix 2018\nhttps://arxiv.org/abs/1806.05236\n\n\n- cutmix  \n\"CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features\" - Sangdoo Yun, iccv 2019\nhttps://arxiv.org/abs/1905.04899\n[[cutmix is all you need code]](https://www.kaggle.com/c/bengaliai-cv19/discussion/126504)  \n\n- augmix  \n\"AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty\" - Dan Hendrycks, arvix 2019\nhttps://arxiv.org/abs/1912.02781\n\n\n- gridmask  \n\"GridMask Data Augmentation\" - Pengguang Chen, arvix 2020\nhttps://arxiv.org/abs/2001.04086\n\n\n- dropblock  \n\"DropBlock: A regularization method for convolutional networks\" - Golnaz Ghiasi, nips 2019\n[[code]](https://github.com/sujatasaini/Kuzushiji-DropBlock)  \n\n\n- fmix \n[[paper+code]](https://www.kaggle.com/c/bengaliai-cv19/discussion/133322)\n\n- batchboost: regularization for stabilizing training with resistance to underfitting &amp; overfitting\nMaciej A. Czyzewski\n\n\n- MaxUp: A Simple Way to Improve Generalization of Neural Network Training  \n[[ref]](https://www.kaggle.com/c/bengaliai-cv19/discussion/123757)   \n```python\nMethod    Top-1 error Top-5 error\nVanilla (He et al., 2016a)    76.3    -\nDropout (Srivastava et al., 2014)    76.8    93.4\nDropPath (Larsson et al., 2017)    77.1    93.5\nManifold Mixup (Verma et al., 2019)    77.5    93.8\nAutoAugment (Cubuk et al., 2019a)    77.6    93.8\nMixup (Zhang et al., 2018)    77.9    93.9\nDropBlock (Ghiasi et al., 2018)    78.3    94.1\nCutMix (Yun et al., 2019)    78.6    94.0\nMaxUp+CutMix    78.9    94.2\n\nTable 1: Summary of top1 and top5 accuracies on the validation set of ImageNet for ResNet-50.\n```\n\n\n- report on experiments on \"No Augmentation\"  \n[[ref from Heng]](https://www.kaggle.com/c/bengaliai-cv19/discussion/132898)  \n\n\n\n- (2020)  SuperMix  \n[[code]](https://github.com/alldbi/SuperMix)\n\n\n## Model  \n- Se-Resnext50/101\n- EfficientNet-b2\n- GhostNet\n- [TPAMI2020]Res2Net [[paper]](https://arxiv.org/abs/1904.01169)  [[pretrained]](https://mmcheng.net/res2net/)  \n\n\n\n\n## Tricks\n\n- [[worth seeing posts and tricks]](https://www.kaggle.com/c/bengaliai-cv19/discussion/127976)  \n\n\n- Half precision\n\n- channel_1 to channel_3\n```python\ninput1ch = Input(shape=(224, 224, 1))\ninput3ch = Conv2D(3, (1, 1))(input1ch)    #1 -&gt; 3 channels\n\n# or\ndef forward(self, x):\n    h = x.repeat(1,3,1,1)\n    h = self.base_model.features(h)\n\n```\n\n- pretrained  \n```python\n1.  add a 1 to 3 conversion conv2d layer (or many layer)\n\n2. freeze all imagenet pretrained weights. Now only the conversion conv2d and the last fc logit layer are trainable.\n\n3. start to train. this force the conversion conv2d to produce activation signal to match the magnitude that is suitable for the pretrained weights.\n\n4. when there can be not more improvement, unfreeze everything and do the training as usual.\n```\n\n\n- OHEM    \n[[OHEM+label smoothing]](https://www.kaggle.com/c/bengaliai-cv19/discussion/128637)  \n- Label Smoothing  \n[[Label Smoothing]](https://www.kaggle.com/c/bengaliai-cv19/discussion/128115#763163)  \n\n\n- CamDrop: A New Explanation of Dropout and A Guided Regularization Method for Deep Neural Networks\n\n\n\n\n\n\n\n## Infer  \n- K-fold\n- StritifiedKFold \n- TTA   \n [[TTA-Multiple_models]](https://www.kaggle.com/c/bengaliai-cv19/discussion/126648)  \n- Ensemble   \n\n\n- new validation (unseen)\n[[discussion]](https://www.kaggle.com/c/bengaliai-cv19/discussion/134434)",
      "votes": null
    },
    {
      "id": "768973",
      "postDate": "03/11/2020 11:53:31",
      "content": "<p>update Res2Net (TPAMI2020)\n<a href=\"https://arxiv.org/abs/1904.01169\">https://arxiv.org/abs/1904.01169</a></p>",
      "rawMarkdown": "update Res2Net (TPAMI2020)\nhttps://arxiv.org/abs/1904.01169",
      "votes": null
    },
    {
      "id": "769293",
      "postDate": "03/11/2020 18:30:26",
      "content": "<p>Thanks for sharing. Good work. </p>",
      "rawMarkdown": "Thanks for sharing. Good work.",
      "votes": null
    },
    {
      "id": "769573",
      "postDate": "03/12/2020 02:47:45",
      "content": "<p>update [(Submitted on 10 Mar 2020)] <strong>SuperMix</strong> <br>\n<a href=\"https://github.com/alldbi/SuperMix\">[code]</a></p>\n\n<blockquote>\n  <p>In this paper, we propose a supervised mixing augmentation method, termed SuperMix, which exploits the knowledge of a teacher to mix images based on their salient regions. SuperMix optimizes a mixing objective that considers: i) forcing the class of input images to appear in the mixed image, ii) preserving the local structure of images, and iii) reducing the risk of suppressing important features. To make the mixing suitable for large-scale applications, we develop an optimization technique, 65× faster than gradient descent on the same problem. We validate the effectiveness of SuperMix through extensive evaluations and ablation studies on two tasks of object classification and knowledge distillation. On the classification task, SuperMix provides the same performance as the advanced augmentation methods, such as AutoAugment. On the distillation task, SuperMix sets a new state-of-the-art with a significantly simplified distillation method. Particularly, in six out of eight teacher-student setups from the same architectures, the students trained on the mixed data surpass their teachers with a notable margin.  </p>\n</blockquote>",
      "rawMarkdown": "update [(Submitted on 10 Mar 2020)] **SuperMix**  \n[[code]](https://github.com/alldbi/SuperMix)\n&gt; In this paper, we propose a supervised mixing augmentation method, termed SuperMix, which exploits the knowledge of a teacher to mix images based on their salient regions. SuperMix optimizes a mixing objective that considers: i) forcing the class of input images to appear in the mixed image, ii) preserving the local structure of images, and iii) reducing the risk of suppressing important features. To make the mixing suitable for large-scale applications, we develop an optimization technique, 65× faster than gradient descent on the same problem. We validate the effectiveness of SuperMix through extensive evaluations and ablation studies on two tasks of object classification and knowledge distillation. On the classification task, SuperMix provides the same performance as the advanced augmentation methods, such as AutoAugment. On the distillation task, SuperMix sets a new state-of-the-art with a significantly simplified distillation method. Particularly, in six out of eight teacher-student setups from the same architectures, the students trained on the mixed data surpass their teachers with a notable margin.",
      "votes": null
    },
    {
      "id": "823930",
      "postDate": "04/28/2020 03:12:40",
      "content": "<p>in the link below there is an implementation of dropblock2D,i want to implement a dropblock1D for time series data, does anyone have a good implement way ?\n<a href=\"https://github.com/DHZS/tf-dropblock/blob/master/nets/dropblock.py\">https://github.com/DHZS/tf-dropblock/blob/master/nets/dropblock.py</a></p>",
      "rawMarkdown": "in the link below there is an implementation of dropblock2D,i want to implement a dropblock1D for time series data, does anyone have a good implement way ?\nhttps://github.com/DHZS/tf-dropblock/blob/master/nets/dropblock.py",
      "votes": null
    },
    {
      "id": "945971",
      "postDate": "07/26/2020 08:52:59",
      "content": "<p>How can you apply fmix to target detection? How to deal with the target boxes? Thank you</p>",
      "rawMarkdown": "How can you apply fmix to target detection? How to deal with the target boxes? Thank you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 768973,
      "author_name": "bryce1010",
      "author_url": "",
      "post_date": "03/11/2020 11:53:31",
      "content": "<p>update Res2Net (TPAMI2020)\n<a href=\"https://arxiv.org/abs/1904.01169\">https://arxiv.org/abs/1904.01169</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 769293,
      "author_name": "akashshingha850",
      "author_url": "",
      "post_date": "03/11/2020 18:30:26",
      "content": "<p>Thanks for sharing. Good work. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 769573,
      "author_name": "bryce1010",
      "author_url": "",
      "post_date": "03/12/2020 02:47:45",
      "content": "<p>update [(Submitted on 10 Mar 2020)] <strong>SuperMix</strong> <br>\n<a href=\"https://github.com/alldbi/SuperMix\">[code]</a></p>\n\n<blockquote>\n  <p>In this paper, we propose a supervised mixing augmentation method, termed SuperMix, which exploits the knowledge of a teacher to mix images based on their salient regions. SuperMix optimizes a mixing objective that considers: i) forcing the class of input images to appear in the mixed image, ii) preserving the local structure of images, and iii) reducing the risk of suppressing important features. To make the mixing suitable for large-scale applications, we develop an optimization technique, 65× faster than gradient descent on the same problem. We validate the effectiveness of SuperMix through extensive evaluations and ablation studies on two tasks of object classification and knowledge distillation. On the classification task, SuperMix provides the same performance as the advanced augmentation methods, such as AutoAugment. On the distillation task, SuperMix sets a new state-of-the-art with a significantly simplified distillation method. Particularly, in six out of eight teacher-student setups from the same architectures, the students trained on the mixed data surpass their teachers with a notable margin.  </p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 823930,
      "author_name": "skgone123",
      "author_url": "",
      "post_date": "04/28/2020 03:12:40",
      "content": "<p>in the link below there is an implementation of dropblock2D,i want to implement a dropblock1D for time series data, does anyone have a good implement way ?\n<a href=\"https://github.com/DHZS/tf-dropblock/blob/master/nets/dropblock.py\">https://github.com/DHZS/tf-dropblock/blob/master/nets/dropblock.py</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 945971,
      "author_name": "irinali",
      "author_url": "",
      "post_date": "07/26/2020 08:52:59",
      "content": "<p>How can you apply fmix to target detection? How to deal with the target boxes? Thank you</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "768754": "Hi, Im new to Kaggle competitions, and late for this match. But I will boost my grade as high as I can. \nHope everyone have a nice grade.\n\n## Augmentation\n\nref:  \n[[list of all mixup variant]](https://www.kaggle.com/c/bengaliai-cv19/discussion/132492)  \n\n\n- cutout  \n\"Improved Regularization of Convolutional Neural Networks with Cutout\" - Terrance DeVries, arvix 2017\nhttps://arxiv.org/abs/1708.04552\n\n- mixup  \n\"mixup: Beyond Empirical Risk Minimization\" - Hongyi Zhang, arvix 2017\nhttps://arxiv.org/abs/1710.09412\n[[code]](https://www.kaggle.com/c/bengaliai-cv19/discussion/128115#763163)  \n\n\n- manifold mixup  \n\"Manifold Mixup: Better Representations by Interpolating Hidden States\" - Vikas Verma, arvix 2018\nhttps://arxiv.org/abs/1806.05236\n\n\n- cutmix  \n\"CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features\" - Sangdoo Yun, iccv 2019\nhttps://arxiv.org/abs/1905.04899\n[[cutmix is all you need code]](https://www.kaggle.com/c/bengaliai-cv19/discussion/126504)  \n\n- augmix  \n\"AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty\" - Dan Hendrycks, arvix 2019\nhttps://arxiv.org/abs/1912.02781\n\n\n- gridmask  \n\"GridMask Data Augmentation\" - Pengguang Chen, arvix 2020\nhttps://arxiv.org/abs/2001.04086\n\n\n- dropblock  \n\"DropBlock: A regularization method for convolutional networks\" - Golnaz Ghiasi, nips 2019\n[[code]](https://github.com/sujatasaini/Kuzushiji-DropBlock)  \n\n\n- fmix \n[[paper+code]](https://www.kaggle.com/c/bengaliai-cv19/discussion/133322)\n\n- batchboost: regularization for stabilizing training with resistance to underfitting &amp; overfitting\nMaciej A. Czyzewski\n\n\n- MaxUp: A Simple Way to Improve Generalization of Neural Network Training  \n[[ref]](https://www.kaggle.com/c/bengaliai-cv19/discussion/123757)   \n```python\nMethod    Top-1 error Top-5 error\nVanilla (He et al., 2016a)    76.3    -\nDropout (Srivastava et al., 2014)    76.8    93.4\nDropPath (Larsson et al., 2017)    77.1    93.5\nManifold Mixup (Verma et al., 2019)    77.5    93.8\nAutoAugment (Cubuk et al., 2019a)    77.6    93.8\nMixup (Zhang et al., 2018)    77.9    93.9\nDropBlock (Ghiasi et al., 2018)    78.3    94.1\nCutMix (Yun et al., 2019)    78.6    94.0\nMaxUp+CutMix    78.9    94.2\n\nTable 1: Summary of top1 and top5 accuracies on the validation set of ImageNet for ResNet-50.\n```\n\n\n- report on experiments on \"No Augmentation\"  \n[[ref from Heng]](https://www.kaggle.com/c/bengaliai-cv19/discussion/132898)  \n\n\n\n- (2020)  SuperMix  \n[[code]](https://github.com/alldbi/SuperMix)\n\n\n## Model  \n- Se-Resnext50/101\n- EfficientNet-b2\n- GhostNet\n- [TPAMI2020]Res2Net [[paper]](https://arxiv.org/abs/1904.01169)  [[pretrained]](https://mmcheng.net/res2net/)  \n\n\n\n\n## Tricks\n\n- [[worth seeing posts and tricks]](https://www.kaggle.com/c/bengaliai-cv19/discussion/127976)  \n\n\n- Half precision\n\n- channel_1 to channel_3\n```python\ninput1ch = Input(shape=(224, 224, 1))\ninput3ch = Conv2D(3, (1, 1))(input1ch)    #1 -&gt; 3 channels\n\n# or\ndef forward(self, x):\n    h = x.repeat(1,3,1,1)\n    h = self.base_model.features(h)\n\n```\n\n- pretrained  \n```python\n1.  add a 1 to 3 conversion conv2d layer (or many layer)\n\n2. freeze all imagenet pretrained weights. Now only the conversion conv2d and the last fc logit layer are trainable.\n\n3. start to train. this force the conversion conv2d to produce activation signal to match the magnitude that is suitable for the pretrained weights.\n\n4. when there can be not more improvement, unfreeze everything and do the training as usual.\n```\n\n\n- OHEM    \n[[OHEM+label smoothing]](https://www.kaggle.com/c/bengaliai-cv19/discussion/128637)  \n- Label Smoothing  \n[[Label Smoothing]](https://www.kaggle.com/c/bengaliai-cv19/discussion/128115#763163)  \n\n\n- CamDrop: A New Explanation of Dropout and A Guided Regularization Method for Deep Neural Networks\n\n\n\n\n\n\n\n## Infer  \n- K-fold\n- StritifiedKFold \n- TTA   \n [[TTA-Multiple_models]](https://www.kaggle.com/c/bengaliai-cv19/discussion/126648)  \n- Ensemble   \n\n\n- new validation (unseen)\n[[discussion]](https://www.kaggle.com/c/bengaliai-cv19/discussion/134434)",
    "768973": "update Res2Net (TPAMI2020)\nhttps://arxiv.org/abs/1904.01169",
    "769293": "Thanks for sharing. Good work.",
    "769573": "update [(Submitted on 10 Mar 2020)] **SuperMix**  \n[[code]](https://github.com/alldbi/SuperMix)\n&gt; In this paper, we propose a supervised mixing augmentation method, termed SuperMix, which exploits the knowledge of a teacher to mix images based on their salient regions. SuperMix optimizes a mixing objective that considers: i) forcing the class of input images to appear in the mixed image, ii) preserving the local structure of images, and iii) reducing the risk of suppressing important features. To make the mixing suitable for large-scale applications, we develop an optimization technique, 65× faster than gradient descent on the same problem. We validate the effectiveness of SuperMix through extensive evaluations and ablation studies on two tasks of object classification and knowledge distillation. On the classification task, SuperMix provides the same performance as the advanced augmentation methods, such as AutoAugment. On the distillation task, SuperMix sets a new state-of-the-art with a significantly simplified distillation method. Particularly, in six out of eight teacher-student setups from the same architectures, the students trained on the mixed data surpass their teachers with a notable margin.",
    "823930": "in the link below there is an implementation of dropblock2D,i want to implement a dropblock1D for time series data, does anyone have a good implement way ?\nhttps://github.com/DHZS/tf-dropblock/blob/master/nets/dropblock.py",
    "945971": "How can you apply fmix to target detection? How to deal with the target boxes? Thank you"
  },
  "source": "meta"
}