{
  "id": 77327,
  "title": "My solution (with github)",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/77327",
  "author_name": "Femi",
  "post_date": "2019-01-11T15:15:20.066000",
  "votes": 28,
  "comment_count": 17,
  "views": 0,
  "content": "<p>First of all, thanks to Kaggle and HPA team for this meaningful competition! Thanks @Spytensor for sharing his pytorch baseline code. I started the competition by his code. Thanks @Allunia, @Iafoss, <a href=\"/jschnab\">@jschnab</a>, <a href=\"/kwentar\">@kwentar</a> for sharing their nice exploration and visualization kernels. Thanks <a href=\"/zhangmiao\">@zhangmiao</a>, <a href=\"/interneuron\">@interneuron</a>, <a href=\"/pascal1129\">@pascal1129</a>, @maw501for, ... answered my question.</p>\n\n<p>I am a new  person in the field of data science and this is my first time to participate in competitions. I am very glad that I have chosen kaggle to enter the world of cv. Because there are too many people sharing their ideas here. I have learned too much in the past month. Because of the lack of knowledge, I don't have much trick to strengthen my model. You will find my model is very simple, maybe I got this place by luck haha.</p>\n\n<h2>Here is my solution:</h2>\n\n<ol>\n<li><strong>data</strong>: kaggle + external data 512x512 RGBY</li>\n<li><strong>model</strong>: Densenet121 (conv0 3-&gt;4 channels, linear layers with batch normalization and dropout)</li>\n<li><strong>augmentation</strong>: horizontal flip, vertical flip, rotate, shear, lighter/darker</li>\n<li><strong>normalize</strong>: kaggle data and external data with different mean and std</li>\n<li><strong>optimizer</strong>: SGD</li>\n<li><strong>loss</strong>: BCE (no weight)</li>\n<li><strong>oversampling</strong>: each image sampled by 'WeightedRandomSampler' function in pytorch. their weight is calculated by math.log(#all_target / #class-i).</li>\n<li><strong>cv</strong>: 5 folds</li>\n<li><strong>threshold search</strong>: I use the code supplied by Brian</li>\n<li><strong>TTA</strong>: 2-horizontal flip, 2-vertical flip</li>\n</ol>\n\n<h2>I want to share my understanding to strategies:</h2>\n\n<ol>\n<li><p>TTA and oversampling is really helpful, and most importantly 5-fold cv improve score by 0.02. \n<br></p></li>\n<li><p>Threshold is crucial, different threshold have a great influence on the LB score. The threshold I searched by valid set lowered my score at first. So I use the constant threshold (0.15) for a long time. The searched threshold for each class varies from 0.1 to 0.9 and different, and I cannot find any relationship between rare class and its threshold. I just want to give up to use searched threshold until I found that using smaller threshold would increase the score (not always). So I just multiply each searched threshold by a factor(~0.5). That mean I want predicted more target each class and get high recall. Although doing so may lower my f1 score, it does improve my public/private LB score. I think the reason may be that the wrong classification can be eliminated by tta and ensemble, and finally get more #TP and high score. My final submission ensemble 5-fold x 4-tta x 3-factor(0.3, 0.4, 0.5).\n<code>note</code>: there are 5 version of threshold corresponding 5-fold cv. When using a model to predict, I choose the corresponding threshold. You can find them in my github repository <code>data.py</code> file(fold 12-16)\n<br></p></li>\n<li><p>I think it's important to handle external data correctly. For me, the red, green and blue images are directly extracted from the corresponding channels of original jpg images ,the yellow image is combined with R and Y channel in original image. Finally save all of them into 512x512 gray png images separately. Because kaggle and external data have different distributions, so I normalized them with two different mean and std.\n<br></p></li>\n<li><p>Put all images into SSD faster than HDD in training. and I want to know is there any other way to accelerate training? Because I train 1 fold for 30 hours (GTX 1070, 24 epochs)</p></li>\n</ol>\n\n<p>Here is the github repository and I hope this is helpful for you.\n<a href=\"https://github.com/feifei9099/kaggle_human_protein\">https://github.com/feifei9099/kaggle_human_protein</a>\n<br></p>\n\n<p><strong>update</strong>: I am sorry that my grades have been cancelled. Because this is my first time to participate in the competition, there are some rules that I am not very clear. I promise to follow the rules of kaggle next time.</p>",
  "messages": [
    {
      "id": 454393,
      "postDate": "2019-01-11T15:15:20.067Z",
      "content": "<p>First of all, thanks to Kaggle and HPA team for this meaningful competition! Thanks @Spytensor for sharing his pytorch baseline code. I started the competition by his code. Thanks @Allunia, @Iafoss, <a href=\"/jschnab\">@jschnab</a>, <a href=\"/kwentar\">@kwentar</a> for sharing their nice exploration and visualization kernels. Thanks <a href=\"/zhangmiao\">@zhangmiao</a>, <a href=\"/interneuron\">@interneuron</a>, <a href=\"/pascal1129\">@pascal1129</a>, @maw501for, ... answered my question.</p>\n\n<p>I am a new  person in the field of data science and this is my first time to participate in competitions. I am very glad that I have chosen kaggle to enter the world of cv. Because there are too many people sharing their ideas here. I have learned too much in the past month. Because of the lack of knowledge, I don't have much trick to strengthen my model. You will find my model is very simple, maybe I got this place by luck haha.</p>\n\n<h2>Here is my solution:</h2>\n\n<ol>\n<li><strong>data</strong>: kaggle + external data 512x512 RGBY</li>\n<li><strong>model</strong>: Densenet121 (conv0 3-&gt;4 channels, linear layers with batch normalization and dropout)</li>\n<li><strong>augmentation</strong>: horizontal flip, vertical flip, rotate, shear, lighter/darker</li>\n<li><strong>normalize</strong>: kaggle data and external data with different mean and std</li>\n<li><strong>optimizer</strong>: SGD</li>\n<li><strong>loss</strong>: BCE (no weight)</li>\n<li><strong>oversampling</strong>: each image sampled by 'WeightedRandomSampler' function in pytorch. their weight is calculated by math.log(#all_target / #class-i).</li>\n<li><strong>cv</strong>: 5 folds</li>\n<li><strong>threshold search</strong>: I use the code supplied by Brian</li>\n<li><strong>TTA</strong>: 2-horizontal flip, 2-vertical flip</li>\n</ol>\n\n<h2>I want to share my understanding to strategies:</h2>\n\n<ol>\n<li><p>TTA and oversampling is really helpful, and most importantly 5-fold cv improve score by 0.02. \n<br></p></li>\n<li><p>Threshold is crucial, different threshold have a great influence on the LB score. The threshold I searched by valid set lowered my score at first. So I use the constant threshold (0.15) for a long time. The searched threshold for each class varies from 0.1 to 0.9 and different, and I cannot find any relationship between rare class and its threshold. I just want to give up to use searched threshold until I found that using smaller threshold would increase the score (not always). So I just multiply each searched threshold by a factor(~0.5). That mean I want predicted more target each class and get high recall. Although doing so may lower my f1 score, it does improve my public/private LB score. I think the reason may be that the wrong classification can be eliminated by tta and ensemble, and finally get more #TP and high score. My final submission ensemble 5-fold x 4-tta x 3-factor(0.3, 0.4, 0.5).\n<code>note</code>: there are 5 version of threshold corresponding 5-fold cv. When using a model to predict, I choose the corresponding threshold. You can find them in my github repository <code>data.py</code> file(fold 12-16)\n<br></p></li>\n<li><p>I think it's important to handle external data correctly. For me, the red, green and blue images are directly extracted from the corresponding channels of original jpg images ,the yellow image is combined with R and Y channel in original image. Finally save all of them into 512x512 gray png images separately. Because kaggle and external data have different distributions, so I normalized them with two different mean and std.\n<br></p></li>\n<li><p>Put all images into SSD faster than HDD in training. and I want to know is there any other way to accelerate training? Because I train 1 fold for 30 hours (GTX 1070, 24 epochs)</p></li>\n</ol>\n\n<p>Here is the github repository and I hope this is helpful for you.\n<a href=\"https://github.com/feifei9099/kaggle_human_protein\">https://github.com/feifei9099/kaggle_human_protein</a>\n<br></p>\n\n<p><strong>update</strong>: I am sorry that my grades have been cancelled. Because this is my first time to participate in the competition, there are some rules that I am not very clear. I promise to follow the rules of kaggle next time.</p>",
      "rawMarkdown": "First of all, thanks to Kaggle and HPA team for this meaningful competition! Thanks @Spytensor for sharing his pytorch baseline code. I started the competition by his code. Thanks @Allunia, @Iafoss, @jschnab, @kwentar for sharing their nice exploration and visualization kernels. Thanks @zhangmiao, @interneuron, @pascal1129, @maw501for, ... answered my question.\n\nI am a new  person in the field of data science and this is my first time to participate in competitions. I am very glad that I have chosen kaggle to enter the world of cv. Because there are too many people sharing their ideas here. I have learned too much in the past month. Because of the lack of knowledge, I don't have much trick to strengthen my model. You will find my model is very simple, maybe I got this place by luck haha.\n\n## Here is my solution:\n1. **data**: kaggle + external data 512x512 RGBY\n2. **model**: Densenet121 (conv0 3-&gt;4 channels, linear layers with batch normalization and dropout)\n3. **augmentation**: horizontal flip, vertical flip, rotate, shear, lighter/darker\n4. **normalize**: kaggle data and external data with different mean and std\n5. **optimizer**: SGD\n6. **loss**: BCE (no weight)\n7. **oversampling**: each image sampled by 'WeightedRandomSampler' function in pytorch. their weight is calculated by math.log(\\#all_target / \\#class-i).\n8. **cv**: 5 folds\n9. **threshold search**: I use the code supplied by Brian\n10. **TTA**: 2-horizontal flip, 2-vertical flip\n\n## I want to share my understanding to strategies:\n1. TTA and oversampling is really helpful, and most importantly 5-fold cv improve score by 0.02. \n<br>\n\n2. Threshold is crucial, different threshold have a great influence on the LB score. The threshold I searched by valid set lowered my score at first. So I use the constant threshold (0.15) for a long time. The searched threshold for each class varies from 0.1 to 0.9 and different, and I cannot find any relationship between rare class and its threshold. I just want to give up to use searched threshold until I found that using smaller threshold would increase the score (not always). So I just multiply each searched threshold by a factor(~0.5). That mean I want predicted more target each class and get high recall. Although doing so may lower my f1 score, it does improve my public/private LB score. I think the reason may be that the wrong classification can be eliminated by tta and ensemble, and finally get more \\#TP and high score. My final submission ensemble 5-fold x 4-tta x 3-factor(0.3, 0.4, 0.5).\n`note`: there are 5 version of threshold corresponding 5-fold cv. When using a model to predict, I choose the corresponding threshold. You can find them in my github repository `data.py` file(fold 12-16)\n<br>\n\n3. I think it's important to handle external data correctly. For me, the red, green and blue images are directly extracted from the corresponding channels of original jpg images ,the yellow image is combined with R and Y channel in original image. Finally save all of them into 512x512 gray png images separately. Because kaggle and external data have different distributions, so I normalized them with two different mean and std.\n<br>\n\n4. Put all images into SSD faster than HDD in training. and I want to know is there any other way to accelerate training? Because I train 1 fold for 30 hours (GTX 1070, 24 epochs)\n\n\nHere is the github repository and I hope this is helpful for you.\nhttps://github.com/feifei9099/kaggle_human_protein\n<br>\n\n**update**: I am sorry that my grades have been cancelled. Because this is my first time to participate in the competition, there are some rules that I am not very clear. I promise to follow the rules of kaggle next time.",
      "votes": 27
    },
    {
      "id": 455456,
      "postDate": "2019-01-14T00:13:16.383Z",
      "content": "<p>Thanks for sharing your solution. It is astonishing that you could achieve that high rank with such a a low number of ensembles and without to much tricks and tweaks. Sorry for what happened to your grades, but why?  I am too knew in kaggle, and don't know the rules like you</p>",
      "rawMarkdown": "Thanks for sharing your solution. It is astonishing that you could achieve that high rank with such a a low number of ensembles and without to much tricks and tweaks. Sorry for what happened to your grades, but why?  I am too knew in kaggle, and don't know the rules like you"
    },
    {
      "id": 455111,
      "postDate": "2019-01-13T02:10:52.397Z",
      "content": "<p>Thanks for sharing! Great work and such a nice and clear github README.\nSorry for what happened to your medal. Wish you good luck!</p>",
      "rawMarkdown": "Thanks for sharing! Great work and such a nice and clear github README.\nSorry for what happened to your medal. Wish you good luck!",
      "replies": [
        {
          "id": 455138,
          "postDate": "2019-01-13T03:57:27.497Z",
          "content": "<p>Thanks :-)</p>",
          "rawMarkdown": "Thanks :-)"
        }
      ]
    },
    {
      "id": 454719,
      "postDate": "2019-01-12T05:00:02.177Z",
      "content": "<p>Congrats!And if I can use your collections for top solutions in this competition to add it in my blog?</p>",
      "rawMarkdown": "Congrats!And if I can use your collections for top solutions in this competition to add it in my blog?",
      "replies": [
        {
          "id": 454922,
          "postDate": "2019-01-12T14:02:38.140Z",
          "content": "<p>Of course, this is my pleasure. I have put top solutions at the forefront of my README.md. And I will keep to update it</p>",
          "rawMarkdown": "Of course, this is my pleasure. I have put top solutions at the forefront of my README.md. And I will keep to update it"
        }
      ]
    },
    {
      "id": 454687,
      "postDate": "2019-01-12T02:13:30.190Z",
      "content": "<p>thanks for sharing. so you are doing voting ensemble on different folds? Have you tried averaging probability of all folds?\nI was always doing probability averaging  during contest, this kind of ensemble can never boost me 0.02 on LB.</p>",
      "rawMarkdown": "thanks for sharing. so you are doing voting ensemble on different folds? Have you tried averaging probability of all folds?\nI was always doing probability averaging  during contest, this kind of ensemble can never boost me 0.02 on LB.",
      "replies": [
        {
          "id": 454694,
          "postDate": "2019-01-12T03:11:26.757Z",
          "content": "<p>I use average mean ensemble of final target. I have never tried averaging probability.</p>",
          "rawMarkdown": "I use average mean ensemble of final target. I have never tried averaging probability."
        }
      ]
    },
    {
      "id": 454496,
      "postDate": "2019-01-11T18:39:49.023Z",
      "content": "<p>Great job, Femi! \nDense121 was one of my better performers, it worked strangely well. ibnDense121 was even better but densenets became incredibly slow for me as the image size increased. </p>\n\n<p>I am also very new and this was my first competition I really worked on so I have much to learn. Thanks for sharing your solution, extremely informative!</p>",
      "rawMarkdown": "Great job, Femi! \nDense121 was one of my better performers, it worked strangely well. ibnDense121 was even better but densenets became incredibly slow for me as the image size increased. \n\nI am also very new and this was my first competition I really worked on so I have much to learn. Thanks for sharing your solution, extremely informative!",
      "replies": [
        {
          "id": 454678,
          "postDate": "2019-01-12T01:17:51.933Z",
          "content": "<p>Thanks, slow training has also plagued me. 256x256 took me 7 hours, 512x512 took me 30 hours. It seems that the time spent exponentially increases with the size of the image. I want to know if there is any way to speed up the training.\nI also have a lot of things to learn and work hard together :-)</p>",
          "rawMarkdown": "Thanks, slow training has also plagued me. 256x256 took me 7 hours, 512x512 took me 30 hours. It seems that the time spent exponentially increases with the size of the image. I want to know if there is any way to speed up the training.\nI also have a lot of things to learn and work hard together :-)"
        },
        {
          "id": 454679,
          "postDate": "2019-01-12T01:25:57.803Z",
          "content": "<p>512x512 has 4 times as much data as 256x256 so 28h seems reasonable.\nI also wonder how some of the top winners managed to train 14 models or so. You can't make any mistakes... </p>",
          "rawMarkdown": "512x512 has 4 times as much data as 256x256 so 28h seems reasonable.\nI also wonder how some of the top winners managed to train 14 models or so. You can't make any mistakes... "
        },
        {
          "id": 454698,
          "postDate": "2019-01-12T03:15:34.013Z",
          "content": "<p>thanks!</p>",
          "rawMarkdown": "thanks!"
        }
      ]
    },
    {
      "id": 454464,
      "postDate": "2019-01-11T17:08:30.637Z",
      "content": "<p><a href=\"/femichen\">@femichen</a> Thanks for sharing and great job being 28th with a single 5 fold model! Most of the write ups so far have many more models.</p>",
      "rawMarkdown": "@femichen Thanks for sharing and great job being 28th with a single 5 fold model! Most of the write ups so far have many more models.",
      "replies": [
        {
          "id": 454675,
          "postDate": "2019-01-12T01:01:31.013Z",
          "content": "<p>thanks, I think the key to my success is finding the right threshold.I will try more models in the next competition.</p>",
          "rawMarkdown": "thanks, I think the key to my success is finding the right threshold.I will try more models in the next competition."
        }
      ]
    },
    {
      "id": 454411,
      "postDate": "2019-01-11T15:46:57.220Z",
      "content": "<p>Congrats and thanks  for your solution! It seems from your and other top-rated solutions , that preprocessing external data was extremely important. </p>",
      "rawMarkdown": "Congrats and thanks  for your solution! It seems from your and other top-rated solutions , that preprocessing external data was extremely important. ",
      "replies": [
        {
          "id": 454418,
          "postDate": "2019-01-11T16:04:09.567Z",
          "content": "<p>Thanks, you are right. And I am curious how the gold medal winner uses external data. :-)</p>",
          "rawMarkdown": "Thanks, you are right. And I am curious how the gold medal winner uses external data. :-)"
        }
      ]
    },
    {
      "id": 454989,
      "postDate": "2019-01-12T17:34:58.943Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 455116,
          "postDate": "2019-01-13T02:40:21.090Z",
          "content": "<p>hah~, thanks! I am sorry that I have not tried stand-alone oversampling. My understanding of stand-alone oversampling is duplication of low frequency classes like code here. Am I right?\n<code>\ntrain_df_orig=train_df.copy() <br>\nlows = [15,15,15,8,9,10,8,9,10,8,9,10,17,20,24,26,15,27,15,20,24,17,8,15,27,27,27]\nfor i in lows:\n    target = str(i)\n    indicies = train_df_orig.loc[train_df_orig['Target'] == target].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.startswith(target+\" \")].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.endswith(\" \"+target)].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.contains(\" \"+target+\" \")].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n</code>\nSuppose my understanding is correct. I think such oversampling needs to decide which rare classes need to be oversampling and how many times they are duplicated. In other words, there are too many hyperparameters to be adjusted. You need to find the right parameters to get him to work. However, log weighted helped us decide all of this.\nI don't think log weighted is the most correct way, but I think it is the reasonable way. Hope to get your reply.</p>",
          "rawMarkdown": "hah~, thanks! I am sorry that I have not tried stand-alone oversampling. My understanding of stand-alone oversampling is duplication of low frequency classes like code here. Am I right?\n```\ntrain_df_orig=train_df.copy()    \nlows = [15,15,15,8,9,10,8,9,10,8,9,10,17,20,24,26,15,27,15,20,24,17,8,15,27,27,27]\nfor i in lows:\n    target = str(i)\n    indicies = train_df_orig.loc[train_df_orig['Target'] == target].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.startswith(target+\" \")].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.endswith(\" \"+target)].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.contains(\" \"+target+\" \")].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n```\nSuppose my understanding is correct. I think such oversampling needs to decide which rare classes need to be oversampling and how many times they are duplicated. In other words, there are too many hyperparameters to be adjusted. You need to find the right parameters to get him to work. However, log weighted helped us decide all of this.\nI don't think log weighted is the most correct way, but I think it is the reasonable way. Hope to get your reply."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 455456,
      "author_name": "Haider Alwasiti",
      "author_url": "",
      "post_date": "2019-01-14T00:13:16.383000",
      "content": "<p>Thanks for sharing your solution. It is astonishing that you could achieve that high rank with such a a low number of ensembles and without to much tricks and tweaks. Sorry for what happened to your grades, but why?  I am too knew in kaggle, and don't know the rules like you</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 455111,
      "author_name": "zhangboshen",
      "author_url": "",
      "post_date": "2019-01-13T02:10:52.397000",
      "content": "<p>Thanks for sharing! Great work and such a nice and clear github README.\nSorry for what happened to your medal. Wish you good luck!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 455138,
          "author_name": "Femi",
          "author_url": "",
          "post_date": "2019-01-13T03:57:27.497000",
          "content": "<p>Thanks :-)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 454719,
      "author_name": "Spytensor",
      "author_url": "",
      "post_date": "2019-01-12T05:00:02.177000",
      "content": "<p>Congrats!And if I can use your collections for top solutions in this competition to add it in my blog?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 454922,
          "author_name": "Femi",
          "author_url": "",
          "post_date": "2019-01-12T14:02:38.140000",
          "content": "<p>Of course, this is my pleasure. I have put top solutions at the forefront of my README.md. And I will keep to update it</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 454687,
      "author_name": "good good study",
      "author_url": "",
      "post_date": "2019-01-12T02:13:30.190000",
      "content": "<p>thanks for sharing. so you are doing voting ensemble on different folds? Have you tried averaging probability of all folds?\nI was always doing probability averaging  during contest, this kind of ensemble can never boost me 0.02 on LB.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 454694,
          "author_name": "Femi",
          "author_url": "",
          "post_date": "2019-01-12T03:11:26.757000",
          "content": "<p>I use average mean ensemble of final target. I have never tried averaging probability.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 454496,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-01-11T18:39:49.023000",
      "content": "<p>Great job, Femi! \nDense121 was one of my better performers, it worked strangely well. ibnDense121 was even better but densenets became incredibly slow for me as the image size increased. </p>\n\n<p>I am also very new and this was my first competition I really worked on so I have much to learn. Thanks for sharing your solution, extremely informative!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 454678,
          "author_name": "Femi",
          "author_url": "",
          "post_date": "2019-01-12T01:17:51.933000",
          "content": "<p>Thanks, slow training has also plagued me. 256x256 took me 7 hours, 512x512 took me 30 hours. It seems that the time spent exponentially increases with the size of the image. I want to know if there is any way to speed up the training.\nI also have a lot of things to learn and work hard together :-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 454679,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2019-01-12T01:25:57.803000",
          "content": "<p>512x512 has 4 times as much data as 256x256 so 28h seems reasonable.\nI also wonder how some of the top winners managed to train 14 models or so. You can't make any mistakes... </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 454698,
          "author_name": "Femi",
          "author_url": "",
          "post_date": "2019-01-12T03:15:34.013000",
          "content": "<p>thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 454464,
      "author_name": "David Wagner",
      "author_url": "",
      "post_date": "2019-01-11T17:08:30.637000",
      "content": "<p><a href=\"/femichen\">@femichen</a> Thanks for sharing and great job being 28th with a single 5 fold model! Most of the write ups so far have many more models.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 454675,
          "author_name": "Femi",
          "author_url": "",
          "post_date": "2019-01-12T01:01:31.013000",
          "content": "<p>thanks, I think the key to my success is finding the right threshold.I will try more models in the next competition.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 454411,
      "author_name": "Vaagn Minasian",
      "author_url": "",
      "post_date": "2019-01-11T15:46:57.220000",
      "content": "<p>Congrats and thanks  for your solution! It seems from your and other top-rated solutions , that preprocessing external data was extremely important. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 454418,
          "author_name": "Femi",
          "author_url": "",
          "post_date": "2019-01-11T16:04:09.567000",
          "content": "<p>Thanks, you are right. And I am curious how the gold medal winner uses external data. :-)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 454989,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-12T17:34:58.943000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 455116,
          "author_name": "Femi",
          "author_url": "",
          "post_date": "2019-01-13T02:40:21.090000",
          "content": "<p>hah~, thanks! I am sorry that I have not tried stand-alone oversampling. My understanding of stand-alone oversampling is duplication of low frequency classes like code here. Am I right?\n<code>\ntrain_df_orig=train_df.copy() <br>\nlows = [15,15,15,8,9,10,8,9,10,8,9,10,17,20,24,26,15,27,15,20,24,17,8,15,27,27,27]\nfor i in lows:\n    target = str(i)\n    indicies = train_df_orig.loc[train_df_orig['Target'] == target].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.startswith(target+\" \")].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.endswith(\" \"+target)].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n    indicies = train_df_orig.loc[train_df_orig['Target'].str.contains(\" \"+target+\" \")].index\n    train_df = pd.concat([train_df,train_df_orig.loc[indicies]], ignore_index=True)\n</code>\nSuppose my understanding is correct. I think such oversampling needs to decide which rare classes need to be oversampling and how many times they are duplicated. In other words, there are too many hyperparameters to be adjusted. You need to find the right parameters to get him to work. However, log weighted helped us decide all of this.\nI don't think log weighted is the most correct way, but I think it is the reasonable way. Hope to get your reply.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "454393": "First of all, thanks to Kaggle and HPA team for this meaningful competition! Thanks @Spytensor for sharing his pytorch baseline code. I started the competition by his code. Thanks @Allunia, @Iafoss, @jschnab, @kwentar for sharing their nice exploration and visualization kernels. Thanks @zhangmiao, @interneuron, @pascal1129, @maw501for, ... answered my question.\n\nI am a new  person in the field of data science and this is my first time to participate in competitions. I am very glad that I have chosen kaggle to enter the world of cv. Because there are too many people sharing their ideas here. I have learned too much in the past month. Because of the lack of knowledge, I don't have much trick to strengthen my model. You will find my model is very simple, maybe I got this place by luck haha.\n\n## Here is my solution:\n1. **data**: kaggle + external data 512x512 RGBY\n2. **model**: Densenet121 (conv0 3-&gt;4 channels, linear layers with batch normalization and dropout)\n3. **augmentation**: horizontal flip, vertical flip, rotate, shear, lighter/darker\n4. **normalize**: kaggle data and external data with different mean and std\n5. **optimizer**: SGD\n6. **loss**: BCE (no weight)\n7. **oversampling**: each image sampled by 'WeightedRandomSampler' function in pytorch. their weight is calculated by math.log(\\#all_target / \\#class-i).\n8. **cv**: 5 folds\n9. **threshold search**: I use the code supplied by Brian\n10. **TTA**: 2-horizontal flip, 2-vertical flip\n\n## I want to share my understanding to strategies:\n1. TTA and oversampling is really helpful, and most importantly 5-fold cv improve score by 0.02. \n<br>\n\n2. Threshold is crucial, different threshold have a great influence on the LB score. The threshold I searched by valid set lowered my score at first. So I use the constant threshold (0.15) for a long time. The searched threshold for each class varies from 0.1 to 0.9 and different, and I cannot find any relationship between rare class and its threshold. I just want to give up to use searched threshold until I found that using smaller threshold would increase the score (not always). So I just multiply each searched threshold by a factor(~0.5). That mean I want predicted more target each class and get high recall. Although doing so may lower my f1 score, it does improve my public/private LB score. I think the reason may be that the wrong classification can be eliminated by tta and ensemble, and finally get more \\#TP and high score. My final submission ensemble 5-fold x 4-tta x 3-factor(0.3, 0.4, 0.5).\n`note`: there are 5 version of threshold corresponding 5-fold cv. When using a model to predict, I choose the corresponding threshold. You can find them in my github repository `data.py` file(fold 12-16)\n<br>\n\n3. I think it's important to handle external data correctly. For me, the red, green and blue images are directly extracted from the corresponding channels of original jpg images ,the yellow image is combined with R and Y channel in original image. Finally save all of them into 512x512 gray png images separately. Because kaggle and external data have different distributions, so I normalized them with two different mean and std.\n<br>\n\n4. Put all images into SSD faster than HDD in training. and I want to know is there any other way to accelerate training? Because I train 1 fold for 30 hours (GTX 1070, 24 epochs)\n\n\nHere is the github repository and I hope this is helpful for you.\nhttps://github.com/feifei9099/kaggle_human_protein\n<br>\n\n**update**: I am sorry that my grades have been cancelled. Because this is my first time to participate in the competition, there are some rules that I am not very clear. I promise to follow the rules of kaggle next time.",
    "455456": "Thanks for sharing your solution. It is astonishing that you could achieve that high rank with such a a low number of ensembles and without to much tricks and tweaks. Sorry for what happened to your grades, but why?  I am too knew in kaggle, and don't know the rules like you",
    "455111": "Thanks for sharing! Great work and such a nice and clear github README.\nSorry for what happened to your medal. Wish you good luck!",
    "454719": "Congrats!And if I can use your collections for top solutions in this competition to add it in my blog?",
    "454687": "thanks for sharing. so you are doing voting ensemble on different folds? Have you tried averaging probability of all folds?\nI was always doing probability averaging  during contest, this kind of ensemble can never boost me 0.02 on LB.",
    "454496": "Great job, Femi! \nDense121 was one of my better performers, it worked strangely well. ibnDense121 was even better but densenets became incredibly slow for me as the image size increased. \n\nI am also very new and this was my first competition I really worked on so I have much to learn. Thanks for sharing your solution, extremely informative!",
    "454464": "@femichen Thanks for sharing and great job being 28th with a single 5 fold model! Most of the write ups so far have many more models.",
    "454411": "Congrats and thanks  for your solution! It seems from your and other top-rated solutions , that preprocessing external data was extremely important. ",
    "454989": ""
  }
}