{
  "id": 101366,
  "title": "Classification Ensemble: From 0.713 to [0.787 PB]",
  "url": "/competitions/aptos2019-blindness-detection/discussion/101366",
  "author_name": "QuantScientist",
  "post_date": "2019-07-25T09:26:54.933000",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello, \nThere are not so many Classification kernels scoring so high. </p>\n\n<p><strong>The kernel is here:</strong>\n<a href=\"https://www.kaggle.com/solomonk/pytorch-aptos-ensemble-8-x-tta-0-787?scriptVersionId=17769269\">https://www.kaggle.com/solomonk/pytorch-aptos-ensemble-8-x-tta-0-787?scriptVersionId=17769269</a></p>\n\n<p><a href=\"https://www.kaggle.com/solomonk/ensemble-8-x-tta-2-models-008?scriptVersionId=17544570\">https://www.kaggle.com/solomonk/ensemble-8-x-tta-2-models-008?scriptVersionId=17544570</a>\n(actual kernel with public score) </p>\n\n<p><strong>The pre-trained pytorch models are here:</strong>\n<a href=\"https://www.kaggle.com/solomonk/pth-best/settings\">https://www.kaggle.com/solomonk/pth-best/settings</a></p>\n\n<p><strong>Techniques used:</strong>\n1. ImageNetPolicy augmentation (<a href=\"https://github.com/DeepVoltaire/AutoAugment\">https://github.com/DeepVoltaire/AutoAugment</a>)\n2. TTA x 8 \n3. An Ensemble of 3 seresnext50 models \n4. Image size of 512 <br>\n5. Use of MixUpSoftmaxLoss (    \"Reference: <a href=\"https://github.com/fastai/fastai/blob/master/fastai/callbacks/mixup.py#L6\">https://github.com/fastai/fastai/blob/master/fastai/callbacks/mixup.py#L6</a>\") </p>\n\n<p>`\nimport torch.nn as nn\nclass MixUpSoftmaxLoss(nn.Module):</p>\n\n<pre><code>def __init__(self, crit, reduction='mean'):\n    super().__init__()\n    self.crit = crit\n    setattr(self.crit, 'reduction', 'none')\n    self.reduction = reduction\n\ndef forward(self, output, target):\n    if len(target.size()) == 2:\n        loss1 = self.crit(output, target[:, 0].long())\n        loss2 = self.crit(output, target[:, 1].long())\n        lambda_ = target[:, 2]\n        d = (loss1 * lambda_ + loss2 * (1-lambda_)).mean()\n    else:\n        # This handles the cases without MixUp for backward compatibility\n        d = self.crit(output, target)\n    if self.reduction == 'mean':\n        return d.mean()\n    elif self.reduction == 'sum':\n        return d.sum()\n    return d`\n</code></pre>\n\n<p>Let me know if you have any questions. \nGood luck. </p>",
  "messages": [
    {
      "id": 583991,
      "postDate": "2019-07-25T09:26:54.933Z",
      "content": "<p>Hello, \nThere are not so many Classification kernels scoring so high. </p>\n\n<p><strong>The kernel is here:</strong>\n<a href=\"https://www.kaggle.com/solomonk/pytorch-aptos-ensemble-8-x-tta-0-787?scriptVersionId=17769269\">https://www.kaggle.com/solomonk/pytorch-aptos-ensemble-8-x-tta-0-787?scriptVersionId=17769269</a></p>\n\n<p><a href=\"https://www.kaggle.com/solomonk/ensemble-8-x-tta-2-models-008?scriptVersionId=17544570\">https://www.kaggle.com/solomonk/ensemble-8-x-tta-2-models-008?scriptVersionId=17544570</a>\n(actual kernel with public score) </p>\n\n<p><strong>The pre-trained pytorch models are here:</strong>\n<a href=\"https://www.kaggle.com/solomonk/pth-best/settings\">https://www.kaggle.com/solomonk/pth-best/settings</a></p>\n\n<p><strong>Techniques used:</strong>\n1. ImageNetPolicy augmentation (<a href=\"https://github.com/DeepVoltaire/AutoAugment\">https://github.com/DeepVoltaire/AutoAugment</a>)\n2. TTA x 8 \n3. An Ensemble of 3 seresnext50 models \n4. Image size of 512 <br>\n5. Use of MixUpSoftmaxLoss (    \"Reference: <a href=\"https://github.com/fastai/fastai/blob/master/fastai/callbacks/mixup.py#L6\">https://github.com/fastai/fastai/blob/master/fastai/callbacks/mixup.py#L6</a>\") </p>\n\n<p>`\nimport torch.nn as nn\nclass MixUpSoftmaxLoss(nn.Module):</p>\n\n<pre><code>def __init__(self, crit, reduction='mean'):\n    super().__init__()\n    self.crit = crit\n    setattr(self.crit, 'reduction', 'none')\n    self.reduction = reduction\n\ndef forward(self, output, target):\n    if len(target.size()) == 2:\n        loss1 = self.crit(output, target[:, 0].long())\n        loss2 = self.crit(output, target[:, 1].long())\n        lambda_ = target[:, 2]\n        d = (loss1 * lambda_ + loss2 * (1-lambda_)).mean()\n    else:\n        # This handles the cases without MixUp for backward compatibility\n        d = self.crit(output, target)\n    if self.reduction == 'mean':\n        return d.mean()\n    elif self.reduction == 'sum':\n        return d.sum()\n    return d`\n</code></pre>\n\n<p>Let me know if you have any questions. \nGood luck. </p>",
      "rawMarkdown": "Hello, \nThere are not so many Classification kernels scoring so high. \n\n**The kernel is here:**\nhttps://www.kaggle.com/solomonk/pytorch-aptos-ensemble-8-x-tta-0-787?scriptVersionId=17769269\n\nhttps://www.kaggle.com/solomonk/ensemble-8-x-tta-2-models-008?scriptVersionId=17544570\n(actual kernel with public score) \n\n**The pre-trained pytorch models are here:**\nhttps://www.kaggle.com/solomonk/pth-best/settings\n\n**Techniques used:**\n1. ImageNetPolicy augmentation (https://github.com/DeepVoltaire/AutoAugment)\n2. TTA x 8 \n3. An Ensemble of 3 seresnext50 models \n4. Image size of 512  \n5. Use of MixUpSoftmaxLoss (    \"Reference: https://github.com/fastai/fastai/blob/master/fastai/callbacks/mixup.py#L6\") \n\n`\nimport torch.nn as nn\nclass MixUpSoftmaxLoss(nn.Module):\n\n    def __init__(self, crit, reduction='mean'):\n        super().__init__()\n        self.crit = crit\n        setattr(self.crit, 'reduction', 'none')\n        self.reduction = reduction\n\n    def forward(self, output, target):\n        if len(target.size()) == 2:\n            loss1 = self.crit(output, target[:, 0].long())\n            loss2 = self.crit(output, target[:, 1].long())\n            lambda_ = target[:, 2]\n            d = (loss1 * lambda_ + loss2 * (1-lambda_)).mean()\n        else:\n            # This handles the cases without MixUp for backward compatibility\n            d = self.crit(output, target)\n        if self.reduction == 'mean':\n            return d.mean()\n        elif self.reduction == 'sum':\n            return d.sum()\n        return d`\n\nLet me know if you have any questions. \nGood luck. \n",
      "votes": 9
    },
    {
      "id": 584000,
      "postDate": "2019-07-25T09:41:21.293Z",
      "content": "<p>Mixup is better in pb?</p>",
      "rawMarkdown": "Mixup is better in pb?",
      "replies": [
        {
          "id": 584002,
          "postDate": "2019-07-25T09:45:59.310Z",
          "content": "<p>For this specific CNN network, I found MixUp loss to be better than BCE, even without using MixUp during training. </p>",
          "rawMarkdown": "For this specific CNN network, I found MixUp loss to be better than BCE, even without using MixUp during training. "
        },
        {
          "id": 584009,
          "postDate": "2019-07-25T09:54:26.683Z",
          "content": "<p>OK，thanks. I find alpha parameter is Hard to adjust\n&nbsp;</p>",
          "rawMarkdown": "OK，thanks. I find alpha parameter is Hard to adjust\n&nbsp;"
        }
      ]
    },
    {
      "id": 585105,
      "postDate": "2019-07-27T01:00:58.377Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 584000,
      "author_name": "JieXuan Zhou",
      "author_url": "",
      "post_date": "2019-07-25T09:41:21.293000",
      "content": "<p>Mixup is better in pb?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 584002,
          "author_name": "QuantScientist",
          "author_url": "",
          "post_date": "2019-07-25T09:45:59.310000",
          "content": "<p>For this specific CNN network, I found MixUp loss to be better than BCE, even without using MixUp during training. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 584009,
          "author_name": "JieXuan Zhou",
          "author_url": "",
          "post_date": "2019-07-25T09:54:26.683000",
          "content": "<p>OK，thanks. I find alpha parameter is Hard to adjust\n&nbsp;</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 585105,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-27T01:00:58.377000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "583991": "Hello, \nThere are not so many Classification kernels scoring so high. \n\n**The kernel is here:**\nhttps://www.kaggle.com/solomonk/pytorch-aptos-ensemble-8-x-tta-0-787?scriptVersionId=17769269\n\nhttps://www.kaggle.com/solomonk/ensemble-8-x-tta-2-models-008?scriptVersionId=17544570\n(actual kernel with public score) \n\n**The pre-trained pytorch models are here:**\nhttps://www.kaggle.com/solomonk/pth-best/settings\n\n**Techniques used:**\n1. ImageNetPolicy augmentation (https://github.com/DeepVoltaire/AutoAugment)\n2. TTA x 8 \n3. An Ensemble of 3 seresnext50 models \n4. Image size of 512  \n5. Use of MixUpSoftmaxLoss (    \"Reference: https://github.com/fastai/fastai/blob/master/fastai/callbacks/mixup.py#L6\") \n\n`\nimport torch.nn as nn\nclass MixUpSoftmaxLoss(nn.Module):\n\n    def __init__(self, crit, reduction='mean'):\n        super().__init__()\n        self.crit = crit\n        setattr(self.crit, 'reduction', 'none')\n        self.reduction = reduction\n\n    def forward(self, output, target):\n        if len(target.size()) == 2:\n            loss1 = self.crit(output, target[:, 0].long())\n            loss2 = self.crit(output, target[:, 1].long())\n            lambda_ = target[:, 2]\n            d = (loss1 * lambda_ + loss2 * (1-lambda_)).mean()\n        else:\n            # This handles the cases without MixUp for backward compatibility\n            d = self.crit(output, target)\n        if self.reduction == 'mean':\n            return d.mean()\n        elif self.reduction == 'sum':\n            return d.sum()\n        return d`\n\nLet me know if you have any questions. \nGood luck. \n",
    "584000": "Mixup is better in pb?",
    "585105": ""
  }
}