{
  "id": 77256,
  "title": "33th Place Algorithm on Private LB",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/77256",
  "author_name": "Ildoo Kim",
  "post_date": "2019-01-11T00:52:40.709000",
  "votes": 30,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I participate this competition as a kind of test. I wanted to develop a algorithm, which generate reasonable models without lots of human intervention.</p>\n\n<p>Since data preprocessing and augmentations are very labor intensive works, I focused on cross-validated, various-architecture, robust single models and ensembled models from them, with very clean and basic script.</p>\n\n<p>My single models are like 0.55~0.57 on public LB. Unfortunately, unlike my expectations, ensemble from LOTS of models wasn't that much effective(0.601 on public LB). I learned that basic single model should perform better than this. If I have a more time(I participated late), I will definately try to improve my single models, with changing input size, data augmentations, various network architecture, ROI cropping and ensemble with other known hand crafted features for this dataset.</p>\n\n<p>Here is the github repository so I hope this is helpful for many people. \n<a href=\"https://github.com/ildoonet/kaggle-human-protein-atlas-image-classification\">https://github.com/ildoonet/kaggle-human-protein-atlas-image-classification</a></p>\n\n<h3>Models</h3>\n\n<ul>\n<li>vgg16</li>\n<li>resnet50, resnet101, ...</li>\n<li>densenet121, densenet169 *</li>\n<li>inception v3, inception v4 *</li>\n<li>se152</li>\n<li>polynet</li>\n<li>NASNet, PNASNet</li>\n</ul>\n\n<h3>Implementations</h3>\n\n<ul>\n<li>Data Loader for External Datas and Merger</li>\n<li>Basic data augmentations\n<ul><li>Rotation, Flip *</li>\n<li>Channel drops</li></ul></li>\n<li>16 Test-Time Augmentation</li>\n<li>5-Folds Cross Validation</li>\n<li>Simple Threshold Search Algorithm</li>\n<li>Ensembles\n<ul><li>Test-Time Augmentation Averaging *</li>\n<li>Majority Voting *</li>\n<li>Fully-Connected Neural Network</li>\n<li>logits -&gt; output</li>\n<li>logits + features -&gt; output</li>\n<li>XGBoost</li></ul></li>\n<li>Loss\n<ul><li>Soft F1 Loss *</li>\n<li>Binary Cross Entropy *</li>\n<li>Focal Loss</li>\n<li>MultiLabelMarginLoss</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": 453935,
      "postDate": "2019-01-11T00:52:40.710Z",
      "content": "<p>I participate this competition as a kind of test. I wanted to develop a algorithm, which generate reasonable models without lots of human intervention.</p>\n\n<p>Since data preprocessing and augmentations are very labor intensive works, I focused on cross-validated, various-architecture, robust single models and ensembled models from them, with very clean and basic script.</p>\n\n<p>My single models are like 0.55~0.57 on public LB. Unfortunately, unlike my expectations, ensemble from LOTS of models wasn't that much effective(0.601 on public LB). I learned that basic single model should perform better than this. If I have a more time(I participated late), I will definately try to improve my single models, with changing input size, data augmentations, various network architecture, ROI cropping and ensemble with other known hand crafted features for this dataset.</p>\n\n<p>Here is the github repository so I hope this is helpful for many people. \n<a href=\"https://github.com/ildoonet/kaggle-human-protein-atlas-image-classification\">https://github.com/ildoonet/kaggle-human-protein-atlas-image-classification</a></p>\n\n<h3>Models</h3>\n\n<ul>\n<li>vgg16</li>\n<li>resnet50, resnet101, ...</li>\n<li>densenet121, densenet169 *</li>\n<li>inception v3, inception v4 *</li>\n<li>se152</li>\n<li>polynet</li>\n<li>NASNet, PNASNet</li>\n</ul>\n\n<h3>Implementations</h3>\n\n<ul>\n<li>Data Loader for External Datas and Merger</li>\n<li>Basic data augmentations\n<ul><li>Rotation, Flip *</li>\n<li>Channel drops</li></ul></li>\n<li>16 Test-Time Augmentation</li>\n<li>5-Folds Cross Validation</li>\n<li>Simple Threshold Search Algorithm</li>\n<li>Ensembles\n<ul><li>Test-Time Augmentation Averaging *</li>\n<li>Majority Voting *</li>\n<li>Fully-Connected Neural Network</li>\n<li>logits -&gt; output</li>\n<li>logits + features -&gt; output</li>\n<li>XGBoost</li></ul></li>\n<li>Loss\n<ul><li>Soft F1 Loss *</li>\n<li>Binary Cross Entropy *</li>\n<li>Focal Loss</li>\n<li>MultiLabelMarginLoss</li></ul></li>\n</ul>",
      "rawMarkdown": "I participate this competition as a kind of test. I wanted to develop a algorithm, which generate reasonable models without lots of human intervention.\n\nSince data preprocessing and augmentations are very labor intensive works, I focused on cross-validated, various-architecture, robust single models and ensembled models from them, with very clean and basic script.\n\nMy single models are like 0.55~0.57 on public LB. Unfortunately, unlike my expectations, ensemble from LOTS of models wasn't that much effective(0.601 on public LB). I learned that basic single model should perform better than this. If I have a more time(I participated late), I will definately try to improve my single models, with changing input size, data augmentations, various network architecture, ROI cropping and ensemble with other known hand crafted features for this dataset.\n\nHere is the github repository so I hope this is helpful for many people. \nhttps://github.com/ildoonet/kaggle-human-protein-atlas-image-classification\n\n### Models\n\n- vgg16\n- resnet50, resnet101, ...\n- densenet121, densenet169 *\n- inception v3, inception v4 *\n- se152\n- polynet\n- NASNet, PNASNet\n\n### Implementations\n\n- Data Loader for External Datas and Merger\n- Basic data augmentations\n  - Rotation, Flip *\n  - Channel drops\n- 16 Test-Time Augmentation\n- 5-Folds Cross Validation\n- Simple Threshold Search Algorithm\n- Ensembles\n  - Test-Time Augmentation Averaging *\n  - Majority Voting *\n  - Fully-Connected Neural Network\n    - logits -&gt; output\n    - logits + features -&gt; output\n  - XGBoost\n- Loss\n  - Soft F1 Loss *\n  - Binary Cross Entropy *\n  - Focal Loss\n  - MultiLabelMarginLoss\n",
      "votes": 29
    },
    {
      "id": 454019,
      "postDate": "2019-01-11T03:06:23.180Z",
      "content": "<p>Congratulation surprised to see VGG-16 as one of the things that worked, I would probably better understand if I go through the code...</p>",
      "rawMarkdown": "Congratulation surprised to see VGG-16 as one of the things that worked, I would probably better understand if I go through the code...",
      "votes": 1
    },
    {
      "id": 453991,
      "postDate": "2019-01-11T02:23:43.450Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 1
    },
    {
      "id": 454290,
      "postDate": "2019-01-11T11:23:30.980Z",
      "content": "<p>Congratulations！I'm curious about each result of these ensemble methods. For me, average is the best...</p>",
      "rawMarkdown": "Congratulations！I'm curious about each result of these ensemble methods. For me, average is the best..."
    },
    {
      "id": 454240,
      "postDate": "2019-01-11T09:45:37.873Z",
      "content": "<p>I think 0.03-0.05 LB boost after ensemble is quite good....</p>",
      "rawMarkdown": "I think 0.03-0.05 LB boost after ensemble is quite good...."
    },
    {
      "id": 454878,
      "postDate": "2019-01-12T12:25:14.580Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 454153,
      "postDate": "2019-01-11T07:11:45.547Z",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 454019,
      "author_name": "Vishy",
      "author_url": "",
      "post_date": "2019-01-11T03:06:23.180000",
      "content": "<p>Congratulation surprised to see VGG-16 as one of the things that worked, I would probably better understand if I go through the code...</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 453991,
      "author_name": "Hilal Shaath",
      "author_url": "",
      "post_date": "2019-01-11T02:23:43.450000",
      "content": "<p>Congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 454290,
      "author_name": "zjuyang",
      "author_url": "",
      "post_date": "2019-01-11T11:23:30.980000",
      "content": "<p>Congratulations！I'm curious about each result of these ensemble methods. For me, average is the best...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 454240,
      "author_name": "good good study",
      "author_url": "",
      "post_date": "2019-01-11T09:45:37.873000",
      "content": "<p>I think 0.03-0.05 LB boost after ensemble is quite good....</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 454878,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-12T12:25:14.580000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 454153,
      "author_name": "Shai",
      "author_url": "",
      "post_date": "2019-01-11T07:11:45.547000",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "453935": "I participate this competition as a kind of test. I wanted to develop a algorithm, which generate reasonable models without lots of human intervention.\n\nSince data preprocessing and augmentations are very labor intensive works, I focused on cross-validated, various-architecture, robust single models and ensembled models from them, with very clean and basic script.\n\nMy single models are like 0.55~0.57 on public LB. Unfortunately, unlike my expectations, ensemble from LOTS of models wasn't that much effective(0.601 on public LB). I learned that basic single model should perform better than this. If I have a more time(I participated late), I will definately try to improve my single models, with changing input size, data augmentations, various network architecture, ROI cropping and ensemble with other known hand crafted features for this dataset.\n\nHere is the github repository so I hope this is helpful for many people. \nhttps://github.com/ildoonet/kaggle-human-protein-atlas-image-classification\n\n### Models\n\n- vgg16\n- resnet50, resnet101, ...\n- densenet121, densenet169 *\n- inception v3, inception v4 *\n- se152\n- polynet\n- NASNet, PNASNet\n\n### Implementations\n\n- Data Loader for External Datas and Merger\n- Basic data augmentations\n  - Rotation, Flip *\n  - Channel drops\n- 16 Test-Time Augmentation\n- 5-Folds Cross Validation\n- Simple Threshold Search Algorithm\n- Ensembles\n  - Test-Time Augmentation Averaging *\n  - Majority Voting *\n  - Fully-Connected Neural Network\n    - logits -&gt; output\n    - logits + features -&gt; output\n  - XGBoost\n- Loss\n  - Soft F1 Loss *\n  - Binary Cross Entropy *\n  - Focal Loss\n  - MultiLabelMarginLoss\n",
    "454019": "Congratulation surprised to see VGG-16 as one of the things that worked, I would probably better understand if I go through the code...",
    "453991": "Congratulations!",
    "454290": "Congratulations！I'm curious about each result of these ensemble methods. For me, average is the best...",
    "454240": "I think 0.03-0.05 LB boost after ensemble is quite good....",
    "454878": "",
    "454153": "Congratulations and thanks for sharing!"
  }
}