{
  "id": 77285,
  "title": "30 place solution writeup",
  "url": "/competitions/human-protein-atlas-image-classification/writeups/draw-me-a-star-30-place-solution-writeup",
  "author_name": "",
  "post_date": "2019-01-11T23:01:25.853Z",
  "votes": 23,
  "comment_count": 10,
  "views": 0,
  "content": "<p>First of all, thanks to Kaggle and HPA team for this interesting competition!. This is the most difficult and memorable competition since I start my journey as a Kaggle competitor. I learned a lot from this</p>\n\n<p>Overview, our solution is so basic. I see there are some interesting points inside it. I hope they will be useful for you.  </p>\n\n<ol>\n<li>Data <br>\nWe use RGBY and addition HPAv18. </li>\n<li>Preprocessing <br>\n<ul><li>We were only able to train with 512x512 images. To save the disk space and loading time, we resize the HPAv18 image to 512x512.  </li>\n<li>Upsampling rare classes: Any class which has a number of image less than 450, we upsample up to 450 with random rotation from (0, 360) </li></ul></li>\n<li>Model <br>\nWe use following model for final submission: \n<ul><li>SEResnext50 ( public: 0.604, private: 0.534)</li>\n<li>SEResnet50 (public: 0.599, private: 0.534)</li>\n<li>InceptionV3 (public: 0.585, private: 0.521) </li>\n<li>Resnet34 (public: 0.580, private: 0.508) </li></ul></li>\n<li>Loss <br>\nIn our experiment, weighted BCE loss performs best. Other combinations such as: FocalLoss, F1 \ndoes not work. We stuck at public 0.577LB for a long time when using it. <br>\nAfter that, we try a tricky loss method: Train the model with weighted BCE loss with even epoch, \nand F1 loss with odd epoch. Then, we get 0.04 addition for each model and ensemble. </li>\n<li>Augmentation \n<ul><li>Rotation (0, 360) </li>\n<li>HFlip, VFlip </li>\n<li>Random rotate 90</li>\n<li>Affine   </li></ul></li>\n<li>Threshold <br>\nFixed threshold: 0.2 </li>\n<li>Ensemble <br>\nWeighted average: (1 * resnet34 + 1 * inceptionv3 + 1 * seresnet50 + 3 * seresnext50) / 6  </li>\n</ol>\n\n<p>Cheers,</p>\n\n<p>Edit: We use 3TTA: Flip, VFlip and Normal</p>",
  "messages": [
    {
      "id": "454078",
      "postDate": "01/11/2019 05:20:45",
      "content": "<p>First of all, thanks to Kaggle and HPA team for this interesting competition!. This is the most difficult and memorable competition since I start my journey as a Kaggle competitor. I learned a lot from this</p>\n\n<p>Overview, our solution is so basic. I see there are some interesting points inside it. I hope they will be useful for you.  </p>\n\n<ol>\n<li>Data <br>\nWe use RGBY and addition HPAv18. </li>\n<li>Preprocessing <br>\n<ul><li>We were only able to train with 512x512 images. To save the disk space and loading time, we resize the HPAv18 image to 512x512.  </li>\n<li>Upsampling rare classes: Any class which has a number of image less than 450, we upsample up to 450 with random rotation from (0, 360) </li></ul></li>\n<li>Model <br>\nWe use following model for final submission: \n<ul><li>SEResnext50 ( public: 0.604, private: 0.534)</li>\n<li>SEResnet50 (public: 0.599, private: 0.534)</li>\n<li>InceptionV3 (public: 0.585, private: 0.521) </li>\n<li>Resnet34 (public: 0.580, private: 0.508) </li></ul></li>\n<li>Loss <br>\nIn our experiment, weighted BCE loss performs best. Other combinations such as: FocalLoss, F1 \ndoes not work. We stuck at public 0.577LB for a long time when using it. <br>\nAfter that, we try a tricky loss method: Train the model with weighted BCE loss with even epoch, \nand F1 loss with odd epoch. Then, we get 0.04 addition for each model and ensemble. </li>\n<li>Augmentation \n<ul><li>Rotation (0, 360) </li>\n<li>HFlip, VFlip </li>\n<li>Random rotate 90</li>\n<li>Affine   </li></ul></li>\n<li>Threshold <br>\nFixed threshold: 0.2 </li>\n<li>Ensemble <br>\nWeighted average: (1 * resnet34 + 1 * inceptionv3 + 1 * seresnet50 + 3 * seresnext50) / 6  </li>\n</ol>\n\n<p>Cheers,</p>\n\n<p>Edit: We use 3TTA: Flip, VFlip and Normal</p>",
      "rawMarkdown": "First of all, thanks to Kaggle and HPA team for this interesting competition!. This is the most difficult and memorable competition since I start my journey as a Kaggle competitor. I learned a lot from this\n\nOverview, our solution is so basic. I see there are some interesting points inside it. I hope they will be useful for you.  \n\n1. Data  \nWe use RGBY and addition HPAv18. \n2. Preprocessing  \n - We were only able to train with 512x512 images. To save the disk space and loading time, we resize the HPAv18 image to 512x512.  \n - Upsampling rare classes: Any class which has a number of image less than 450, we upsample up to 450 with random rotation from (0, 360) \n3. Model  \n We use following model for final submission: \n - SEResnext50 ( public: 0.604, private: 0.534)\n - SEResnet50 (public: 0.599, private: 0.534)\n - InceptionV3 (public: 0.585, private: 0.521) \n - Resnet34 (public: 0.580, private: 0.508) \n4. Loss  \n In our experiment, weighted BCE loss performs best. Other combinations such as: FocalLoss, F1 \n does not work. We stuck at public 0.577LB for a long time when using it.  \n After that, we try a tricky loss method: Train the model with weighted BCE loss with even epoch, \n and F1 loss with odd epoch. Then, we get 0.04 addition for each model and ensemble. \n5. Augmentation \n - Rotation (0, 360) \n - HFlip, VFlip \n - Random rotate 90\n - Affine   \n6. Threshold  \n Fixed threshold: 0.2 \n7.  Ensemble  \n Weighted average: (1 * resnet34 + 1 * inceptionv3 + 1 * seresnet50 + 3 * seresnext50) / 6  \n\nCheers,\n\nEdit: We use 3TTA: Flip, VFlip and Normal",
      "votes": null
    },
    {
      "id": "454086",
      "postDate": "01/11/2019 05:32:20",
      "content": "<p>Thanks for sharing.</p>\n\n<p>Do SEResnext50 have pretrained weights and model generator for pytorch?</p>",
      "rawMarkdown": "Thanks for sharing.\n\nDo SEResnext50 have pretrained weights and model generator for pytorch?",
      "votes": null
    },
    {
      "id": "454087",
      "postDate": "01/11/2019 05:35:36",
      "content": "<p>Yes, we used pytorch with pretrained SEResnext50</p>",
      "rawMarkdown": "Yes, we used pytorch with pretrained SEResnext50",
      "votes": null
    },
    {
      "id": "454151",
      "postDate": "01/11/2019 07:10:14",
      "content": "<p>Congratulations and thanks!</p>",
      "rawMarkdown": "Congratulations and thanks!",
      "votes": null
    },
    {
      "id": "454216",
      "postDate": "01/11/2019 08:59:16",
      "content": "<p>I want to ask a stupid question, does public mean the score that the system gave? And what does private mean?</p>",
      "rawMarkdown": "I want to ask a stupid question, does public mean the score that the system gave? And what does private mean?",
      "votes": null
    },
    {
      "id": "454223",
      "postDate": "01/11/2019 09:15:13",
      "content": "<p>Testset is divided into two parts called public and private. We dont know which test case belongs to which test set, but we do know how many percentages of each test set. When you submit, the system gives you a public score that is evaluated on the public test set until the deadline. The private score will be used for the final ranking.</p>",
      "rawMarkdown": "Testset is divided into two parts called public and private. We dont know which test case belongs to which test set, but we do know how many percentages of each test set. When you submit, the system gives you a public score that is evaluated on the public test set until the deadline. The private score will be used for the final ranking.",
      "votes": null
    },
    {
      "id": "454260",
      "postDate": "01/11/2019 10:12:13",
      "content": "<p>In weighted BCE, how are weights calculated? The inverse of the frequency of the class? or something different.</p>",
      "rawMarkdown": "In weighted BCE, how are weights calculated? The inverse of the frequency of the class? or something different.",
      "votes": null
    },
    {
      "id": "454368",
      "postDate": "01/11/2019 14:31:44",
      "content": "<p>Exactly, the inverse of the frequency of the classes (after upsampling).</p>",
      "rawMarkdown": "Exactly, the inverse of the frequency of the classes (after upsampling).",
      "votes": null
    },
    {
      "id": "454459",
      "postDate": "01/11/2019 17:00:45",
      "content": "<p>How did you create your CV splits? Was it randomly sampled, or did you use some other method?</p>",
      "rawMarkdown": "How did you create your CV splits? Was it randomly sampled, or did you use some other method?",
      "votes": null
    },
    {
      "id": "454639",
      "postDate": "01/11/2019 23:01:12",
      "content": "<p>We use stratified kfold which is mentioned in the discussion somewhere. </p>",
      "rawMarkdown": "We use stratified kfold which is mentioned in the discussion somewhere.",
      "votes": null
    },
    {
      "id": "454952",
      "postDate": "01/12/2019 15:52:34",
      "content": "<p>So we can't know our private score until the competition is closed, right?</p>",
      "rawMarkdown": "So we can't know our private score until the competition is closed, right?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 454086,
      "author_name": "ildoonet",
      "author_url": "",
      "post_date": "01/11/2019 05:32:20",
      "content": "<p>Thanks for sharing.</p>\n\n<p>Do SEResnext50 have pretrained weights and model generator for pytorch?</p>",
      "votes": null,
      "replies": [
        {
          "id": 454087,
          "author_name": "backaggle",
          "author_url": "",
          "post_date": "01/11/2019 05:35:36",
          "content": "<p>Yes, we used pytorch with pretrained SEResnext50</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 454151,
      "author_name": "sgalib",
      "author_url": "",
      "post_date": "01/11/2019 07:10:14",
      "content": "<p>Congratulations and thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454216,
      "author_name": "theredarrow",
      "author_url": "",
      "post_date": "01/11/2019 08:59:16",
      "content": "<p>I want to ask a stupid question, does public mean the score that the system gave? And what does private mean?</p>",
      "votes": null,
      "replies": [
        {
          "id": 454223,
          "author_name": "backaggle",
          "author_url": "",
          "post_date": "01/11/2019 09:15:13",
          "content": "<p>Testset is divided into two parts called public and private. We dont know which test case belongs to which test set, but we do know how many percentages of each test set. When you submit, the system gives you a public score that is evaluated on the public test set until the deadline. The private score will be used for the final ranking.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 454952,
          "author_name": "theredarrow",
          "author_url": "",
          "post_date": "01/12/2019 15:52:34",
          "content": "<p>So we can't know our private score until the competition is closed, right?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 454260,
      "author_name": "sourajmishra",
      "author_url": "",
      "post_date": "01/11/2019 10:12:13",
      "content": "<p>In weighted BCE, how are weights calculated? The inverse of the frequency of the class? or something different.</p>",
      "votes": null,
      "replies": [
        {
          "id": 454368,
          "author_name": "hainamnguyen",
          "author_url": "",
          "post_date": "01/11/2019 14:31:44",
          "content": "<p>Exactly, the inverse of the frequency of the classes (after upsampling).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 454459,
      "author_name": "sushifan",
      "author_url": "",
      "post_date": "01/11/2019 17:00:45",
      "content": "<p>How did you create your CV splits? Was it randomly sampled, or did you use some other method?</p>",
      "votes": null,
      "replies": [
        {
          "id": 454639,
          "author_name": "backaggle",
          "author_url": "",
          "post_date": "01/11/2019 23:01:12",
          "content": "<p>We use stratified kfold which is mentioned in the discussion somewhere. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "454078": "First of all, thanks to Kaggle and HPA team for this interesting competition!. This is the most difficult and memorable competition since I start my journey as a Kaggle competitor. I learned a lot from this\n\nOverview, our solution is so basic. I see there are some interesting points inside it. I hope they will be useful for you.  \n\n1. Data  \nWe use RGBY and addition HPAv18. \n2. Preprocessing  \n - We were only able to train with 512x512 images. To save the disk space and loading time, we resize the HPAv18 image to 512x512.  \n - Upsampling rare classes: Any class which has a number of image less than 450, we upsample up to 450 with random rotation from (0, 360) \n3. Model  \n We use following model for final submission: \n - SEResnext50 ( public: 0.604, private: 0.534)\n - SEResnet50 (public: 0.599, private: 0.534)\n - InceptionV3 (public: 0.585, private: 0.521) \n - Resnet34 (public: 0.580, private: 0.508) \n4. Loss  \n In our experiment, weighted BCE loss performs best. Other combinations such as: FocalLoss, F1 \n does not work. We stuck at public 0.577LB for a long time when using it.  \n After that, we try a tricky loss method: Train the model with weighted BCE loss with even epoch, \n and F1 loss with odd epoch. Then, we get 0.04 addition for each model and ensemble. \n5. Augmentation \n - Rotation (0, 360) \n - HFlip, VFlip \n - Random rotate 90\n - Affine   \n6. Threshold  \n Fixed threshold: 0.2 \n7.  Ensemble  \n Weighted average: (1 * resnet34 + 1 * inceptionv3 + 1 * seresnet50 + 3 * seresnext50) / 6  \n\nCheers,\n\nEdit: We use 3TTA: Flip, VFlip and Normal",
    "454086": "Thanks for sharing.\n\nDo SEResnext50 have pretrained weights and model generator for pytorch?",
    "454087": "Yes, we used pytorch with pretrained SEResnext50",
    "454151": "Congratulations and thanks!",
    "454216": "I want to ask a stupid question, does public mean the score that the system gave? And what does private mean?",
    "454223": "Testset is divided into two parts called public and private. We dont know which test case belongs to which test set, but we do know how many percentages of each test set. When you submit, the system gives you a public score that is evaluated on the public test set until the deadline. The private score will be used for the final ranking.",
    "454260": "In weighted BCE, how are weights calculated? The inverse of the frequency of the class? or something different.",
    "454368": "Exactly, the inverse of the frequency of the classes (after upsampling).",
    "454459": "How did you create your CV splits? Was it randomly sampled, or did you use some other method?",
    "454639": "We use stratified kfold which is mentioned in the discussion somewhere.",
    "454952": "So we can't know our private score until the competition is closed, right?"
  },
  "source": "meta"
}