{
  "id": 77254,
  "title": "92nd place pytorch solution",
  "url": "/competitions/human-protein-atlas-image-classification/writeups/yurii-rebryk-92nd-place-pytorch-solution",
  "author_name": "",
  "post_date": "2019-01-11T23:23:58.167Z",
  "votes": 31,
  "comment_count": 11,
  "views": 0,
  "content": "<p><strong>1. Data</strong></p>\n\n<ul>\n<li>Channels: RGB</li>\n<li>Oversampling</li>\n<li>External data: <a href=\"http://v18.proteinatlas.org\">http://v18.proteinatlas.org</a></li>\n</ul>\n\n<p><strong>2. Augmentation</strong></p>\n\n<ul>\n<li>Resize, Rotate, RandomRotate90, HorizontalFlip, RandomBrightnessContrast, Normalize</li>\n</ul>\n\n<p><strong>3. Model design</strong></p>\n\n<ul>\n<li>Backbone: Resnet50 pretrained on ImageNet</li>\n<li>Head: 2 linear layers with batch normalization and dropout</li>\n</ul>\n\n<p><strong>4. Loss</strong></p>\n\n<ul>\n<li>Binary Cross Entropy</li>\n</ul>\n\n<p><strong>5. Training</strong></p>\n\n<ul>\n<li>5-fold CV</li>\n<li>Optimizer: Adam</li>\n<li>Different learning rates for different layers</li>\n<li>Head fine-tuning with frozen backbone (1 epoch)</li>\n<li>Scheduler: Cyclical Learning Rates</li>\n</ul>\n\n<p>Stage 1:</p>\n\n<ul>\n<li>Image size: 256</li>\n<li>Batch size: 128</li>\n<li>Epochs: 16</li>\n</ul>\n\n<p>Stage 2:</p>\n\n<ul>\n<li>Image size: 512</li>\n<li>Batch size: 32</li>\n<li>Epochs: 6</li>\n</ul>\n\n<p><strong>6. Prediciton</strong></p>\n\n<ul>\n<li>TTA: 8</li>\n<li>TTA augmentation: Resize, Rotate, RandomRotate90, HorizontalFlip, Normalize</li>\n<li>The mean of the predictions</li>\n<li>Threshold: 0.2</li>\n</ul>\n\n<p><strong>7. Result</strong></p>\n\n<ul>\n<li>Training takes ~35 hours on Tesla v100</li>\n<li>Public LB: 0.595</li>\n<li>Private LB: 0.523</li>\n</ul>\n\n<p><strong>8. Observations</strong></p>\n\n<ul>\n<li>Mixed precision works poorly</li>\n<li>External data helps a lot</li>\n<li>BCE Loss with oversampling is much better than Focal Loss</li>\n<li>Resnet50 outperforms Resnet18 and Resnet34</li>\n<li>5 folds improve score by 0.024</li>\n<li>TTA helps too</li>\n</ul>\n\n<p>GitHub link: <a href=\"https://github.com/rebryk/kaggle/tree/master/human-protein\">https://github.com/rebryk/kaggle/tree/master/human-protein</a></p>",
  "messages": [
    {
      "id": "453903",
      "postDate": "01/11/2019 00:11:49",
      "content": "<p><strong>1. Data</strong></p>\n\n<ul>\n<li>Channels: RGB</li>\n<li>Oversampling</li>\n<li>External data: <a href=\"http://v18.proteinatlas.org\">http://v18.proteinatlas.org</a></li>\n</ul>\n\n<p><strong>2. Augmentation</strong></p>\n\n<ul>\n<li>Resize, Rotate, RandomRotate90, HorizontalFlip, RandomBrightnessContrast, Normalize</li>\n</ul>\n\n<p><strong>3. Model design</strong></p>\n\n<ul>\n<li>Backbone: Resnet50 pretrained on ImageNet</li>\n<li>Head: 2 linear layers with batch normalization and dropout</li>\n</ul>\n\n<p><strong>4. Loss</strong></p>\n\n<ul>\n<li>Binary Cross Entropy</li>\n</ul>\n\n<p><strong>5. Training</strong></p>\n\n<ul>\n<li>5-fold CV</li>\n<li>Optimizer: Adam</li>\n<li>Different learning rates for different layers</li>\n<li>Head fine-tuning with frozen backbone (1 epoch)</li>\n<li>Scheduler: Cyclical Learning Rates</li>\n</ul>\n\n<p>Stage 1:</p>\n\n<ul>\n<li>Image size: 256</li>\n<li>Batch size: 128</li>\n<li>Epochs: 16</li>\n</ul>\n\n<p>Stage 2:</p>\n\n<ul>\n<li>Image size: 512</li>\n<li>Batch size: 32</li>\n<li>Epochs: 6</li>\n</ul>\n\n<p><strong>6. Prediciton</strong></p>\n\n<ul>\n<li>TTA: 8</li>\n<li>TTA augmentation: Resize, Rotate, RandomRotate90, HorizontalFlip, Normalize</li>\n<li>The mean of the predictions</li>\n<li>Threshold: 0.2</li>\n</ul>\n\n<p><strong>7. Result</strong></p>\n\n<ul>\n<li>Training takes ~35 hours on Tesla v100</li>\n<li>Public LB: 0.595</li>\n<li>Private LB: 0.523</li>\n</ul>\n\n<p><strong>8. Observations</strong></p>\n\n<ul>\n<li>Mixed precision works poorly</li>\n<li>External data helps a lot</li>\n<li>BCE Loss with oversampling is much better than Focal Loss</li>\n<li>Resnet50 outperforms Resnet18 and Resnet34</li>\n<li>5 folds improve score by 0.024</li>\n<li>TTA helps too</li>\n</ul>\n\n<p>GitHub link: <a href=\"https://github.com/rebryk/kaggle/tree/master/human-protein\">https://github.com/rebryk/kaggle/tree/master/human-protein</a></p>",
      "rawMarkdown": "**1. Data**\n\n- Channels: RGB\n- Oversampling\n- External data: http://v18.proteinatlas.org\n\n**2. Augmentation**\n\n- Resize, Rotate, RandomRotate90, HorizontalFlip, RandomBrightnessContrast, Normalize\n\n**3. Model design**\n\n- Backbone: Resnet50 pretrained on ImageNet\n- Head: 2 linear layers with batch normalization and dropout\n\n**4. Loss**\n\n- Binary Cross Entropy\n\n**5. Training**\n\n- 5-fold CV\n- Optimizer: Adam\n- Different learning rates for different layers\n- Head fine-tuning with frozen backbone (1 epoch)\n- Scheduler: Cyclical Learning Rates\n\nStage 1:\n\n- Image size: 256\n- Batch size: 128\n- Epochs: 16\n\nStage 2:\n\n- Image size: 512\n- Batch size: 32\n- Epochs: 6\n\n**6. Prediciton**\n\n- TTA: 8\n- TTA augmentation: Resize, Rotate, RandomRotate90, HorizontalFlip, Normalize\n- The mean of the predictions\n- Threshold: 0.2\n\n**7. Result**\n\n- Training takes ~35 hours on Tesla v100\n- Public LB: 0.595\n- Private LB: 0.523\n\n**8. Observations**\n\n- Mixed precision works poorly\n- External data helps a lot\n- BCE Loss with oversampling is much better than Focal Loss\n- Resnet50 outperforms Resnet18 and Resnet34\n- 5 folds improve score by 0.024\n- TTA helps too\n\nGitHub link: https://github.com/rebryk/kaggle/tree/master/human-protein",
      "votes": null
    },
    {
      "id": "453971",
      "postDate": "01/11/2019 01:52:18",
      "content": "<p>Thanks for sharing this ! </p>\n\n<p>For the loss you could also use Focal Loss with over sampling. In our case we used Under and Over Sample + Focal Loss and it gaves very good results. FL is good to put high loss on sample where your prediction is very far from the actual label.</p>",
      "rawMarkdown": "Thanks for sharing this ! \n\nFor the loss you could also use Focal Loss with over sampling. In our case we used Under and Over Sample + Focal Loss and it gaves very good results. FL is good to put high loss on sample where your prediction is very far from the actual label.",
      "votes": null
    },
    {
      "id": "453989",
      "postDate": "01/11/2019 02:21:11",
      "content": "<p>Congratulations !. Can you tell us more about over sampling? Did you use smoot?</p>",
      "rawMarkdown": "Congratulations !. Can you tell us more about over sampling? Did you use smoot?",
      "votes": null
    },
    {
      "id": "454013",
      "postDate": "01/11/2019 02:57:32",
      "content": "<p>Thanks for Sharing this is amazing..</p>",
      "rawMarkdown": "Thanks for Sharing this is amazing..",
      "votes": null
    },
    {
      "id": "454229",
      "postDate": "01/11/2019 09:19:26",
      "content": "<p>Hi, thanks for sharing. In the head part, why do you use \"2 linear layers\"?</p>",
      "rawMarkdown": "Hi, thanks for sharing. In the head part, why do you use \"2 linear layers\"?",
      "votes": null
    },
    {
      "id": "454237",
      "postDate": "01/11/2019 09:35:04",
      "content": "<p>Sure! I use log weights and choose label with the maximum weight.<br>\nHere is my code:</p>\n\n<pre><code>def parse_target(target: str) -&gt; np.ndarray:\n    y = np.zeros(28, dtype=np.int)\n    indices = [int(it) for it in target.split()]\n    y[indices] = 1\n    return y\n\ndef get_sampler(df: pd.DataFrame, alpha: float = 0.5) -&gt; Sampler:\n    y = np.array([parse_target(target) for target in df.Target])\n    class_weights = np.round(np.log(alpha * y.sum() / y.sum(axis=0)), 2)\n    class_weights[class_weights &lt; 1.0] = 1.0\n\n    weights = np.zeros(len(df))\n    for i, target in enumerate(y):\n        weights[i] = class_weights[target == 1].max()\n\n    return WeightedRandomSampler(weights, len(df))\n</code></pre>",
      "rawMarkdown": "Sure! I use log weights and choose label with the maximum weight.<br>\nHere is my code:\n    \n    def parse_target(target: str) -&gt; np.ndarray:\n        y = np.zeros(28, dtype=np.int)\n        indices = [int(it) for it in target.split()]\n        y[indices] = 1\n        return y\n\n    def get_sampler(df: pd.DataFrame, alpha: float = 0.5) -&gt; Sampler:\n        y = np.array([parse_target(target) for target in df.Target])\n        class_weights = np.round(np.log(alpha * y.sum() / y.sum(axis=0)), 2)\n        class_weights[class_weights &lt; 1.0] = 1.0\n\n        weights = np.zeros(len(df))\n        for i, target in enumerate(y):\n            weights[i] = class_weights[target == 1].max()\n\n        return WeightedRandomSampler(weights, len(df))",
      "votes": null
    },
    {
      "id": "454244",
      "postDate": "01/11/2019 09:53:04",
      "content": "<p>Good question. I started with this and didn't change it later.</p>",
      "rawMarkdown": "Good question. I started with this and didn't change it later.",
      "votes": null
    },
    {
      "id": "454246",
      "postDate": "01/11/2019 09:56:12",
      "content": "<p>Could you please share your code snippets? </p>",
      "rawMarkdown": "Could you please share your code snippets?",
      "votes": null
    },
    {
      "id": "454247",
      "postDate": "01/11/2019 09:56:35",
      "content": "<p>You are welcome!</p>",
      "rawMarkdown": "You are welcome!",
      "votes": null
    },
    {
      "id": "454861",
      "postDate": "01/12/2019 12:00:26",
      "content": "<p>Hi, thanks for sharing. What's the effectof \"soft_f1\"(challenge/utils.py)?</p>",
      "rawMarkdown": "Hi, thanks for sharing. What's the effectof \"soft_f1\"(challenge/utils.py)?",
      "votes": null
    },
    {
      "id": "455899",
      "postDate": "01/14/2019 19:57:57",
      "content": "<p>Some people used this to find thresholds. I tried too. It didn't help me.</p>",
      "rawMarkdown": "Some people used this to find thresholds. I tried too. It didn't help me.",
      "votes": null
    },
    {
      "id": "456869",
      "postDate": "01/16/2019 17:16:09",
      "content": "<p>hi thanks for posting your solution...\ncould you let me know how did u do preprocessing for external data before using it..i see many have done before using that</p>",
      "rawMarkdown": "hi thanks for posting your solution...\ncould you let me know how did u do preprocessing for external data before using it..i see many have done before using that",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 453971,
      "author_name": "areveillon",
      "author_url": "",
      "post_date": "01/11/2019 01:52:18",
      "content": "<p>Thanks for sharing this ! </p>\n\n<p>For the loss you could also use Focal Loss with over sampling. In our case we used Under and Over Sample + Focal Loss and it gaves very good results. FL is good to put high loss on sample where your prediction is very far from the actual label.</p>",
      "votes": null,
      "replies": [
        {
          "id": 454246,
          "author_name": "rebryk",
          "author_url": "",
          "post_date": "01/11/2019 09:56:12",
          "content": "<p>Could you please share your code snippets? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 453989,
      "author_name": "abualabed",
      "author_url": "",
      "post_date": "01/11/2019 02:21:11",
      "content": "<p>Congratulations !. Can you tell us more about over sampling? Did you use smoot?</p>",
      "votes": null,
      "replies": [
        {
          "id": 454237,
          "author_name": "rebryk",
          "author_url": "",
          "post_date": "01/11/2019 09:35:04",
          "content": "<p>Sure! I use log weights and choose label with the maximum weight.<br>\nHere is my code:</p>\n\n<pre><code>def parse_target(target: str) -&gt; np.ndarray:\n    y = np.zeros(28, dtype=np.int)\n    indices = [int(it) for it in target.split()]\n    y[indices] = 1\n    return y\n\ndef get_sampler(df: pd.DataFrame, alpha: float = 0.5) -&gt; Sampler:\n    y = np.array([parse_target(target) for target in df.Target])\n    class_weights = np.round(np.log(alpha * y.sum() / y.sum(axis=0)), 2)\n    class_weights[class_weights &lt; 1.0] = 1.0\n\n    weights = np.zeros(len(df))\n    for i, target in enumerate(y):\n        weights[i] = class_weights[target == 1].max()\n\n    return WeightedRandomSampler(weights, len(df))\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 454013,
      "author_name": "viswanathravindran",
      "author_url": "",
      "post_date": "01/11/2019 02:57:32",
      "content": "<p>Thanks for Sharing this is amazing..</p>",
      "votes": null,
      "replies": [
        {
          "id": 454247,
          "author_name": "rebryk",
          "author_url": "",
          "post_date": "01/11/2019 09:56:35",
          "content": "<p>You are welcome!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 454229,
      "author_name": "zjucor",
      "author_url": "",
      "post_date": "01/11/2019 09:19:26",
      "content": "<p>Hi, thanks for sharing. In the head part, why do you use \"2 linear layers\"?</p>",
      "votes": null,
      "replies": [
        {
          "id": 454244,
          "author_name": "rebryk",
          "author_url": "",
          "post_date": "01/11/2019 09:53:04",
          "content": "<p>Good question. I started with this and didn't change it later.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 454861,
      "author_name": "dldmw579",
      "author_url": "",
      "post_date": "01/12/2019 12:00:26",
      "content": "<p>Hi, thanks for sharing. What's the effectof \"soft_f1\"(challenge/utils.py)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 455899,
          "author_name": "rebryk",
          "author_url": "",
          "post_date": "01/14/2019 19:57:57",
          "content": "<p>Some people used this to find thresholds. I tried too. It didn't help me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 456869,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "01/16/2019 17:16:09",
      "content": "<p>hi thanks for posting your solution...\ncould you let me know how did u do preprocessing for external data before using it..i see many have done before using that</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "453903": "**1. Data**\n\n- Channels: RGB\n- Oversampling\n- External data: http://v18.proteinatlas.org\n\n**2. Augmentation**\n\n- Resize, Rotate, RandomRotate90, HorizontalFlip, RandomBrightnessContrast, Normalize\n\n**3. Model design**\n\n- Backbone: Resnet50 pretrained on ImageNet\n- Head: 2 linear layers with batch normalization and dropout\n\n**4. Loss**\n\n- Binary Cross Entropy\n\n**5. Training**\n\n- 5-fold CV\n- Optimizer: Adam\n- Different learning rates for different layers\n- Head fine-tuning with frozen backbone (1 epoch)\n- Scheduler: Cyclical Learning Rates\n\nStage 1:\n\n- Image size: 256\n- Batch size: 128\n- Epochs: 16\n\nStage 2:\n\n- Image size: 512\n- Batch size: 32\n- Epochs: 6\n\n**6. Prediciton**\n\n- TTA: 8\n- TTA augmentation: Resize, Rotate, RandomRotate90, HorizontalFlip, Normalize\n- The mean of the predictions\n- Threshold: 0.2\n\n**7. Result**\n\n- Training takes ~35 hours on Tesla v100\n- Public LB: 0.595\n- Private LB: 0.523\n\n**8. Observations**\n\n- Mixed precision works poorly\n- External data helps a lot\n- BCE Loss with oversampling is much better than Focal Loss\n- Resnet50 outperforms Resnet18 and Resnet34\n- 5 folds improve score by 0.024\n- TTA helps too\n\nGitHub link: https://github.com/rebryk/kaggle/tree/master/human-protein",
    "453971": "Thanks for sharing this ! \n\nFor the loss you could also use Focal Loss with over sampling. In our case we used Under and Over Sample + Focal Loss and it gaves very good results. FL is good to put high loss on sample where your prediction is very far from the actual label.",
    "453989": "Congratulations !. Can you tell us more about over sampling? Did you use smoot?",
    "454013": "Thanks for Sharing this is amazing..",
    "454229": "Hi, thanks for sharing. In the head part, why do you use \"2 linear layers\"?",
    "454237": "Sure! I use log weights and choose label with the maximum weight.<br>\nHere is my code:\n    \n    def parse_target(target: str) -&gt; np.ndarray:\n        y = np.zeros(28, dtype=np.int)\n        indices = [int(it) for it in target.split()]\n        y[indices] = 1\n        return y\n\n    def get_sampler(df: pd.DataFrame, alpha: float = 0.5) -&gt; Sampler:\n        y = np.array([parse_target(target) for target in df.Target])\n        class_weights = np.round(np.log(alpha * y.sum() / y.sum(axis=0)), 2)\n        class_weights[class_weights &lt; 1.0] = 1.0\n\n        weights = np.zeros(len(df))\n        for i, target in enumerate(y):\n            weights[i] = class_weights[target == 1].max()\n\n        return WeightedRandomSampler(weights, len(df))",
    "454244": "Good question. I started with this and didn't change it later.",
    "454246": "Could you please share your code snippets?",
    "454247": "You are welcome!",
    "454861": "Hi, thanks for sharing. What's the effectof \"soft_f1\"(challenge/utils.py)?",
    "455899": "Some people used this to find thresholds. I tried too. It didn't help me.",
    "456869": "hi thanks for posting your solution...\ncould you let me know how did u do preprocessing for external data before using it..i see many have done before using that"
  },
  "source": "meta"
}