{
  "id": 74146,
  "title": "Baseline Performance ?",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/74146",
  "author_name": "",
  "post_date": "2018-12-09T07:32:28.591354300Z",
  "votes": 3,
  "comment_count": 22,
  "views": 0,
  "content": "<p>Before moving to external data and class-imbalance handling, I want to check my baseline whether the performance makes sense (no serious bugs).  Hope someone can give some comments.</p>\n\n<p>Data split:  using trent-b MultilabelStratifiedShuffleSplit\nimage size:  512*512 (4 channel)\naugmentaiton: fliplr, flipud, transpose\npretrained: ImageNet pretrained model.\nloss funciton: BCE\n1. resnet34  threshold=0.5, local validation: 0.68 ===&gt; LB: 0.42\n2. resnet34  thershold=0.5, better_lr_schedule,  local validation: 0.71 ===&gt; LB: 0.437\n3. resnet34  thershold=(fitting local validation), local validation: 0.723  ===&gt; LB: 0.44</p>\n\n<p>Reading from some kernel, I found that baseline performance is about 0.46. I am wondering the score after LB update ? Is 0.44 LB score reasonable ? Thanks.</p>\n\n<p>update: with threshold selection, the average score of above setting in tensorflow get ~0.45</p>",
  "messages": [
    {
      "id": "435977",
      "postDate": "12/09/2018 07:32:28",
      "content": "<p>Before moving to external data and class-imbalance handling, I want to check my baseline whether the performance makes sense (no serious bugs).  Hope someone can give some comments.</p>\n\n<p>Data split:  using trent-b MultilabelStratifiedShuffleSplit\nimage size:  512*512 (4 channel)\naugmentaiton: fliplr, flipud, transpose\npretrained: ImageNet pretrained model.\nloss funciton: BCE\n1. resnet34  threshold=0.5, local validation: 0.68 ===&gt; LB: 0.42\n2. resnet34  thershold=0.5, better_lr_schedule,  local validation: 0.71 ===&gt; LB: 0.437\n3. resnet34  thershold=(fitting local validation), local validation: 0.723  ===&gt; LB: 0.44</p>\n\n<p>Reading from some kernel, I found that baseline performance is about 0.46. I am wondering the score after LB update ? Is 0.44 LB score reasonable ? Thanks.</p>\n\n<p>update: with threshold selection, the average score of above setting in tensorflow get ~0.45</p>",
      "rawMarkdown": "Before moving to external data and class-imbalance handling, I want to check my baseline whether the performance makes sense (no serious bugs).  Hope someone can give some comments.\n\nData split:  using trent-b MultilabelStratifiedShuffleSplit\nimage size:  512*512 (4 channel)\naugmentaiton: fliplr, flipud, transpose\npretrained: ImageNet pretrained model.\nloss funciton: BCE\n1. resnet34  threshold=0.5, local validation: 0.68 ===&gt; LB: 0.42\n2. resnet34  thershold=0.5, better_lr_schedule,  local validation: 0.71 ===&gt; LB: 0.437\n3. resnet34  thershold=(fitting local validation), local validation: 0.723  ===&gt; LB: 0.44\n\nReading from some kernel, I found that baseline performance is about 0.46. I am wondering the score after LB update ? Is 0.44 LB score reasonable ? Thanks.\n\nupdate: with threshold selection, the average score of above setting in tensorflow get ~0.45",
      "votes": null
    },
    {
      "id": "435979",
      "postDate": "12/09/2018 07:42:32",
      "content": "<ul>\n<li><p>The baseline posted in here: <br>\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72812#latest-434873\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72812#latest-434873</a> <br>\nHe used <code>bninception</code> as the backbone. According to him, it can archive <code>0.461</code>.  </p></li>\n<li><p>The <code>fastai</code> baseline: <br>\n<a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\">https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb</a> <br>\nHe used <code>resnet34</code>. It was <code>0.460</code> before removing leak. Now it turns out <code>0.45</code> after removing leak.  </p></li>\n<li><p>My baseline (also <code>resnet34</code> ) <br>\nLook like I have a similar baseline as you. But I only get <code>0.42x</code> .</p></li>\n</ul>\n\n<p>So, I think your <code>0.44 LB</code> is reasonable compared to others.</p>",
      "rawMarkdown": "* The baseline posted in here:  \nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72812#latest-434873  \nHe used `bninception` as the backbone. According to him, it can archive `0.461`.  \n\n* The `fastai` baseline:  \nhttps://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb  \nHe used `resnet34`. It was `0.460` before removing leak. Now it turns out `0.45` after removing leak.  \n\n* My baseline (also `resnet34` )  \nLook like I have a similar baseline as you. But I only get `0.42x` .\n\nSo, I think your `0.44 LB` is reasonable compared to others.",
      "votes": null
    },
    {
      "id": "435982",
      "postDate": "12/09/2018 07:45:49",
      "content": "<p>Thanks ! \nI was experimenting threshold selection these days, 042-&gt;0.44 mostly come from there.</p>",
      "rawMarkdown": "Thanks ! \nI was experimenting threshold selection these days, 042-&gt;0.44 mostly come from there.",
      "votes": null
    },
    {
      "id": "436024",
      "postDate": "12/09/2018 10:44:25",
      "content": "<p>resnet34(modified) thershold=(fitting local validation), local validation: 0.613 ===&gt; LB: 0.468\nresnet50 thershold=(fitting local validation), local validation: 0.723 ===&gt; LB: 0.47</p>",
      "rawMarkdown": "resnet34(modified) thershold=(fitting local validation), local validation: 0.613 ===&gt; LB: 0.468\nresnet50 thershold=(fitting local validation), local validation: 0.723 ===&gt; LB: 0.47",
      "votes": null
    },
    {
      "id": "436028",
      "postDate": "12/09/2018 11:07:20",
      "content": "<p>Thanks for sharing. Amazing baseline performance.\nMaybe my lr schedule led to overfitting...\nMay I ask what's is your dropout rate ?</p>",
      "rawMarkdown": "Thanks for sharing. Amazing baseline performance.\nMaybe my lr schedule led to overfitting...\nMay I ask what's is your dropout rate ?",
      "votes": null
    },
    {
      "id": "436032",
      "postDate": "12/09/2018 11:40:06",
      "content": "<p>No dropout. The gap between cv and lb puzzles me. I did a lot of experiments on this. </p>",
      "rawMarkdown": "No dropout. The gap between cv and lb puzzles me. I did a lot of experiments on this.",
      "votes": null
    },
    {
      "id": "436047",
      "postDate": "12/09/2018 12:25:53",
      "content": "<p>Same here. Trying to figure out better way to evaluate performance :(</p>",
      "rawMarkdown": "Same here. Trying to figure out better way to evaluate performance :(",
      "votes": null
    },
    {
      "id": "436649",
      "postDate": "12/10/2018 17:42:02",
      "content": "<p>There are some duplicate or very similar images in the train dataset. These might explain the gap between cv and lb. \n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534</a></p>",
      "rawMarkdown": "There are some duplicate or very similar images in the train dataset. These might explain the gap between cv and lb. \nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534",
      "votes": null
    },
    {
      "id": "436715",
      "postDate": "12/10/2018 20:04:55",
      "content": "<p>I believe the the fastai baseline used 256x256.</p>",
      "rawMarkdown": "I believe the the fastai baseline used 256x256.",
      "votes": null
    },
    {
      "id": "436790",
      "postDate": "12/11/2018 00:14:12",
      "content": "<p>Ok. I will change the notebook to see the results.</p>",
      "rawMarkdown": "Ok. I will change the notebook to see the results.",
      "votes": null
    },
    {
      "id": "436856",
      "postDate": "12/11/2018 03:10:38",
      "content": "<p>I found the fastai baseline kernel is hard to reproduce on other DL framework. Even I use pytorch to reproduce, it is hard to get LB 0.45 with 256 size of pretrained resnet34.</p>",
      "rawMarkdown": "I found the fastai baseline kernel is hard to reproduce on other DL framework. Even I use pytorch to reproduce, it is hard to get LB 0.45 with 256 size of pretrained resnet34.",
      "votes": null
    },
    {
      "id": "436864",
      "postDate": "12/11/2018 03:17:46",
      "content": "<p>I currently using Tensorflow. With different threshold (fitting with equal threshold among classes), score vary from 0.42 to 0.45.\nDid you use the similar least square threshold fitting as the kernel in your pytorch baseline ?</p>",
      "rawMarkdown": "I currently using Tensorflow. With different threshold (fitting with equal threshold among classes), score vary from 0.42 to 0.45.\nDid you use the similar least square threshold fitting as the kernel in your pytorch baseline ?",
      "votes": null
    },
    {
      "id": "436910",
      "postDate": "12/11/2018 05:04:55",
      "content": "<p>@Appian, Actually considering these duplicates doesn't help much since they do not include rare classes if I remember correctly. Right now I'm doing validation with excluding these duplicates from validation set, but the mismatch between public LB and val is exactly the same. Partially it may be explained if some rare classes are not present in the public LB and automatically get zero score in contrast to validation. Another strange thing is that for one of my models I got 0.8+ val that gave only 0.42 public LB, while another approach gives 0.47+ public LB with only 0.69 val (and visual analysis of the val and test probability distributions for each label shows than the separation of classes is much worse that in the first approach).</p>\n\n<p>Regarding the baseline, the <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\">kernel</a> I have posted gives 0.460 (0.454 after LB update) public LB on 256x256 (training took only ~6 hours on K80). I'd say it varies from run to run in 0.445-0.46 range based on different versions of the kernel. On 512x512 images some people mention that they got ~0.5+ using this kernel.</p>",
      "rawMarkdown": "Appian, Actually considering these duplicates doesn't help much since they do not include rare classes if I remember correctly. Right now I'm doing validation with excluding these duplicates from validation set, but the mismatch between public LB and val is exactly the same. Partially it may be explained if some rare classes are not present in the public LB and automatically get zero score in contrast to validation. Another strange thing is that for one of my models I got 0.8+ val that gave only 0.42 public LB, while another approach gives 0.47+ public LB with only 0.69 val (and visual analysis of the val and test probability distributions for each label shows than the separation of classes is much worse that in the first approach).\n\nRegarding the baseline, the [kernel][1] I have posted gives 0.460 (0.454 after LB update) public LB on 256x256 (training took only ~6 hours on K80). I'd say it varies from run to run in 0.445-0.46 range based on different versions of the kernel. On 512x512 images some people mention that they got ~0.5+ using this kernel.\n\n\n  [1]: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb",
      "votes": null
    },
    {
      "id": "436955",
      "postDate": "12/11/2018 06:38:51",
      "content": "<p>Thank you @lafoss for sharing your experiment. Excluding duplicates was exactly I was going to do and nice to know the results you got.</p>\n\n<p>Here is what I got recently. I use label distribution of training data for thresholding for now.</p>\n\n<p>local cv: 0.7259\npublic LB: 0.416</p>\n\n<ul>\n<li>00 0.814 | 01 0.853 | 02 0.785 | 03 0.773</li>\n<li>04 0.817 | 05 0.726 | 06 0.668 | 07 0.827</li>\n<li>08 0.545 | 09 1.000 | 10 1.000 | 11 0.840</li>\n<li>12 0.674 | 13 0.704 | 14 0.883 | 15 1.000</li>\n<li>16 0.443 | 17 0.571 | 18 0.547 | 19 0.700 </li>\n<li>20 0.500 | 21 0.701 | 22 0.640 | 23 0.808</li>\n<li>24 0.769 | 25 0.688 | 26 0.545 | 27 0.500</li>\n</ul>\n\n<p>And just like your result, much lower local cv sometimes gives better public LB.</p>",
      "rawMarkdown": "Thank you @lafoss for sharing your experiment. Excluding duplicates was exactly I was going to do and nice to know the results you got.\n\nHere is what I got recently. I use label distribution of training data for thresholding for now.\n\nlocal cv: 0.7259\npublic LB: 0.416\n\n- 00 0.814 | 01 0.853 | 02 0.785 | 03 0.773\n- 04 0.817 | 05 0.726 | 06 0.668 | 07 0.827\n- 08 0.545 | 09 1.000 | 10 1.000 | 11 0.840\n- 12 0.674 | 13 0.704 | 14 0.883 | 15 1.000\n- 16 0.443 | 17 0.571 | 18 0.547 | 19 0.700 \n- 20 0.500 | 21 0.701 | 22 0.640 | 23 0.808\n- 24 0.769 | 25 0.688 | 26 0.545 | 27 0.500\n\nAnd just like your result, much lower local cv sometimes gives better public LB.",
      "votes": null
    },
    {
      "id": "442551",
      "postDate": "12/20/2018 05:29:43",
      "content": "<p>InceptionV3 (modified) threshold=0.2, local validation: 0.773 ==&gt; LB: 0.493</p>",
      "rawMarkdown": "InceptionV3 (modified) threshold=0.2, local validation: 0.773 ==&gt; LB: 0.493",
      "votes": null
    },
    {
      "id": "442567",
      "postDate": "12/20/2018 06:17:48",
      "content": "<p>.515 -&gt; .386 for me with InceptionV3. Any tips? I am not using the AuxLogits layer, and have tried swapping the avg_pool layer out for something that would accept variable image size input, but haven't had much luck with 512 on InceptionV3 yet. I am thinking about changing input to accept 4 channels similar to the fastai baseline kernel that got .46. </p>",
      "rawMarkdown": ".515 -&gt; .386 for me with InceptionV3. Any tips? I am not using the AuxLogits layer, and have tried swapping the avg_pool layer out for something that would accept variable image size input, but haven't had much luck with 512 on InceptionV3 yet. I am thinking about changing input to accept 4 channels similar to the fastai baseline kernel that got .46.",
      "votes": null
    },
    {
      "id": "442582",
      "postDate": "12/20/2018 06:49:47",
      "content": "<p>try TTA</p>",
      "rawMarkdown": "try TTA",
      "votes": null
    },
    {
      "id": "442590",
      "postDate": "12/20/2018 07:02:50",
      "content": "<p>Thanks. I have tried TTA, but only with my rotation / reflection augmentation. I have seen others using Brightness, Blur, Noise, and Zoom. Not so sure about zoom, but I'm guessing I would benefit from these additional augmentations. </p>",
      "rawMarkdown": "Thanks. I have tried TTA, but only with my rotation / reflection augmentation. I have seen others using Brightness, Blur, Noise, and Zoom. Not so sure about zoom, but I'm guessing I would benefit from these additional augmentations.",
      "votes": null
    },
    {
      "id": "442596",
      "postDate": "12/20/2018 07:12:37",
      "content": "<p>In tta (test time) you usually do inference on the  image several times and average the results or something similar. The idea is that the neural network is better in detection of objects in some orientations then others. For tta you usually stick to 90 degrees rotation and flips. As for augmentation during training.... Zoom is tricky, for example, as you might cause a rare kind of protein to be cropped or not appear in the zoomed image. </p>",
      "rawMarkdown": "In tta (test time) you usually do inference on the  image several times and average the results or something similar. The idea is that the neural network is better in detection of objects in some orientations then others. For tta you usually stick to 90 degrees rotation and flips. As for augmentation during training.... Zoom is tricky, for example, as you might cause a rare kind of protein to be cropped or not appear in the zoomed image.",
      "votes": null
    },
    {
      "id": "443120",
      "postDate": "12/21/2018 03:49:12",
      "content": "<p>You can try my <a href=\"https://www.kaggle.com/mathormad/inceptionv3-baseline-lb-0-379\">public kernel</a>, train more epochs on your local machine, apply cross-validation (using 6-fold and average the results) -&gt; you can get around 0.47 on LB.</p>\n\n<p>Try to implement some <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70478\">ideas</a> here and you can get 0.517 on LB.  </p>",
      "rawMarkdown": "You can try my [public kernel][1], train more epochs on your local machine, apply cross-validation (using 6-fold and average the results) -&gt; you can get around 0.47 on LB.\n\nTry to implement some [ideas][2] here and you can get 0.517 on LB.  \n\n\n  [1]: https://www.kaggle.com/mathormad/inceptionv3-baseline-lb-0-379\n  [2]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70478",
      "votes": null
    },
    {
      "id": "443134",
      "postDate": "12/21/2018 04:35:13",
      "content": "<p>Thanks! I'm not sure how I missed your kernel. That is some clean code! I am working on attention now, hopefully can figure it out. In that thread he suggested using red, blue, yellow without green as negative samples to help train the network attention. That seems like an interesting, easy to implement idea. </p>",
      "rawMarkdown": "Thanks! I'm not sure how I missed your kernel. That is some clean code! I am working on attention now, hopefully can figure it out. In that thread he suggested using red, blue, yellow without green as negative samples to help train the network attention. That seems like an interesting, easy to implement idea.",
      "votes": null
    },
    {
      "id": "443336",
      "postDate": "12/21/2018 12:58:05",
      "content": "<p>Hi, what's the meaning of \"apply cross-validation\"? (sorry i am a newcomer)</p>",
      "rawMarkdown": "Hi, what's the meaning of \"apply cross-validation\"? (sorry i am a newcomer)",
      "votes": null
    },
    {
      "id": "443410",
      "postDate": "12/21/2018 15:18:07",
      "content": "<p>@Wang-Xinliang  I believe he is referring to k-fold</p>",
      "rawMarkdown": "Wang-Xinliang  I believe he is referring to k-fold",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 435979,
      "author_name": "backaggle",
      "author_url": "",
      "post_date": "12/09/2018 07:42:32",
      "content": "<ul>\n<li><p>The baseline posted in here: <br>\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72812#latest-434873\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72812#latest-434873</a> <br>\nHe used <code>bninception</code> as the backbone. According to him, it can archive <code>0.461</code>.  </p></li>\n<li><p>The <code>fastai</code> baseline: <br>\n<a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\">https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb</a> <br>\nHe used <code>resnet34</code>. It was <code>0.460</code> before removing leak. Now it turns out <code>0.45</code> after removing leak.  </p></li>\n<li><p>My baseline (also <code>resnet34</code> ) <br>\nLook like I have a similar baseline as you. But I only get <code>0.42x</code> .</p></li>\n</ul>\n\n<p>So, I think your <code>0.44 LB</code> is reasonable compared to others.</p>",
      "votes": null,
      "replies": [
        {
          "id": 435982,
          "author_name": "tkuanlun",
          "author_url": "",
          "post_date": "12/09/2018 07:45:49",
          "content": "<p>Thanks ! \nI was experimenting threshold selection these days, 042-&gt;0.44 mostly come from there.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436715,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "12/10/2018 20:04:55",
          "content": "<p>I believe the the fastai baseline used 256x256.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436790,
          "author_name": "tkuanlun",
          "author_url": "",
          "post_date": "12/11/2018 00:14:12",
          "content": "<p>Ok. I will change the notebook to see the results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436856,
          "author_name": "zjucor",
          "author_url": "",
          "post_date": "12/11/2018 03:10:38",
          "content": "<p>I found the fastai baseline kernel is hard to reproduce on other DL framework. Even I use pytorch to reproduce, it is hard to get LB 0.45 with 256 size of pretrained resnet34.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436864,
          "author_name": "tkuanlun",
          "author_url": "",
          "post_date": "12/11/2018 03:17:46",
          "content": "<p>I currently using Tensorflow. With different threshold (fitting with equal threshold among classes), score vary from 0.42 to 0.45.\nDid you use the similar least square threshold fitting as the kernel in your pytorch baseline ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 436024,
      "author_name": "hdzheng",
      "author_url": "",
      "post_date": "12/09/2018 10:44:25",
      "content": "<p>resnet34(modified) thershold=(fitting local validation), local validation: 0.613 ===&gt; LB: 0.468\nresnet50 thershold=(fitting local validation), local validation: 0.723 ===&gt; LB: 0.47</p>",
      "votes": null,
      "replies": [
        {
          "id": 436028,
          "author_name": "tkuanlun",
          "author_url": "",
          "post_date": "12/09/2018 11:07:20",
          "content": "<p>Thanks for sharing. Amazing baseline performance.\nMaybe my lr schedule led to overfitting...\nMay I ask what's is your dropout rate ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436032,
          "author_name": "hdzheng",
          "author_url": "",
          "post_date": "12/09/2018 11:40:06",
          "content": "<p>No dropout. The gap between cv and lb puzzles me. I did a lot of experiments on this. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436047,
          "author_name": "tkuanlun",
          "author_url": "",
          "post_date": "12/09/2018 12:25:53",
          "content": "<p>Same here. Trying to figure out better way to evaluate performance :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436649,
          "author_name": "appian",
          "author_url": "",
          "post_date": "12/10/2018 17:42:02",
          "content": "<p>There are some duplicate or very similar images in the train dataset. These might explain the gap between cv and lb. \n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436910,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "12/11/2018 05:04:55",
          "content": "<p>@Appian, Actually considering these duplicates doesn't help much since they do not include rare classes if I remember correctly. Right now I'm doing validation with excluding these duplicates from validation set, but the mismatch between public LB and val is exactly the same. Partially it may be explained if some rare classes are not present in the public LB and automatically get zero score in contrast to validation. Another strange thing is that for one of my models I got 0.8+ val that gave only 0.42 public LB, while another approach gives 0.47+ public LB with only 0.69 val (and visual analysis of the val and test probability distributions for each label shows than the separation of classes is much worse that in the first approach).</p>\n\n<p>Regarding the baseline, the <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\">kernel</a> I have posted gives 0.460 (0.454 after LB update) public LB on 256x256 (training took only ~6 hours on K80). I'd say it varies from run to run in 0.445-0.46 range based on different versions of the kernel. On 512x512 images some people mention that they got ~0.5+ using this kernel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436955,
          "author_name": "appian",
          "author_url": "",
          "post_date": "12/11/2018 06:38:51",
          "content": "<p>Thank you @lafoss for sharing your experiment. Excluding duplicates was exactly I was going to do and nice to know the results you got.</p>\n\n<p>Here is what I got recently. I use label distribution of training data for thresholding for now.</p>\n\n<p>local cv: 0.7259\npublic LB: 0.416</p>\n\n<ul>\n<li>00 0.814 | 01 0.853 | 02 0.785 | 03 0.773</li>\n<li>04 0.817 | 05 0.726 | 06 0.668 | 07 0.827</li>\n<li>08 0.545 | 09 1.000 | 10 1.000 | 11 0.840</li>\n<li>12 0.674 | 13 0.704 | 14 0.883 | 15 1.000</li>\n<li>16 0.443 | 17 0.571 | 18 0.547 | 19 0.700 </li>\n<li>20 0.500 | 21 0.701 | 22 0.640 | 23 0.808</li>\n<li>24 0.769 | 25 0.688 | 26 0.545 | 27 0.500</li>\n</ul>\n\n<p>And just like your result, much lower local cv sometimes gives better public LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442551,
      "author_name": "mathormad",
      "author_url": "",
      "post_date": "12/20/2018 05:29:43",
      "content": "<p>InceptionV3 (modified) threshold=0.2, local validation: 0.773 ==&gt; LB: 0.493</p>",
      "votes": null,
      "replies": [
        {
          "id": 442567,
          "author_name": "davidwagnerkc",
          "author_url": "",
          "post_date": "12/20/2018 06:17:48",
          "content": "<p>.515 -&gt; .386 for me with InceptionV3. Any tips? I am not using the AuxLogits layer, and have tried swapping the avg_pool layer out for something that would accept variable image size input, but haven't had much luck with 512 on InceptionV3 yet. I am thinking about changing input to accept 4 channels similar to the fastai baseline kernel that got .46. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442582,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "12/20/2018 06:49:47",
          "content": "<p>try TTA</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442590,
          "author_name": "davidwagnerkc",
          "author_url": "",
          "post_date": "12/20/2018 07:02:50",
          "content": "<p>Thanks. I have tried TTA, but only with my rotation / reflection augmentation. I have seen others using Brightness, Blur, Noise, and Zoom. Not so sure about zoom, but I'm guessing I would benefit from these additional augmentations. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442596,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "12/20/2018 07:12:37",
          "content": "<p>In tta (test time) you usually do inference on the  image several times and average the results or something similar. The idea is that the neural network is better in detection of objects in some orientations then others. For tta you usually stick to 90 degrees rotation and flips. As for augmentation during training.... Zoom is tricky, for example, as you might cause a rare kind of protein to be cropped or not appear in the zoomed image. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443120,
          "author_name": "mathormad",
          "author_url": "",
          "post_date": "12/21/2018 03:49:12",
          "content": "<p>You can try my <a href=\"https://www.kaggle.com/mathormad/inceptionv3-baseline-lb-0-379\">public kernel</a>, train more epochs on your local machine, apply cross-validation (using 6-fold and average the results) -&gt; you can get around 0.47 on LB.</p>\n\n<p>Try to implement some <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70478\">ideas</a> here and you can get 0.517 on LB.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443134,
          "author_name": "davidwagnerkc",
          "author_url": "",
          "post_date": "12/21/2018 04:35:13",
          "content": "<p>Thanks! I'm not sure how I missed your kernel. That is some clean code! I am working on attention now, hopefully can figure it out. In that thread he suggested using red, blue, yellow without green as negative samples to help train the network attention. That seems like an interesting, easy to implement idea. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443336,
          "author_name": "dldmw579",
          "author_url": "",
          "post_date": "12/21/2018 12:58:05",
          "content": "<p>Hi, what's the meaning of \"apply cross-validation\"? (sorry i am a newcomer)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443410,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "12/21/2018 15:18:07",
          "content": "<p>@Wang-Xinliang  I believe he is referring to k-fold</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "435977": "Before moving to external data and class-imbalance handling, I want to check my baseline whether the performance makes sense (no serious bugs).  Hope someone can give some comments.\n\nData split:  using trent-b MultilabelStratifiedShuffleSplit\nimage size:  512*512 (4 channel)\naugmentaiton: fliplr, flipud, transpose\npretrained: ImageNet pretrained model.\nloss funciton: BCE\n1. resnet34  threshold=0.5, local validation: 0.68 ===&gt; LB: 0.42\n2. resnet34  thershold=0.5, better_lr_schedule,  local validation: 0.71 ===&gt; LB: 0.437\n3. resnet34  thershold=(fitting local validation), local validation: 0.723  ===&gt; LB: 0.44\n\nReading from some kernel, I found that baseline performance is about 0.46. I am wondering the score after LB update ? Is 0.44 LB score reasonable ? Thanks.\n\nupdate: with threshold selection, the average score of above setting in tensorflow get ~0.45",
    "435979": "* The baseline posted in here:  \nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72812#latest-434873  \nHe used `bninception` as the backbone. According to him, it can archive `0.461`.  \n\n* The `fastai` baseline:  \nhttps://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb  \nHe used `resnet34`. It was `0.460` before removing leak. Now it turns out `0.45` after removing leak.  \n\n* My baseline (also `resnet34` )  \nLook like I have a similar baseline as you. But I only get `0.42x` .\n\nSo, I think your `0.44 LB` is reasonable compared to others.",
    "435982": "Thanks ! \nI was experimenting threshold selection these days, 042-&gt;0.44 mostly come from there.",
    "436024": "resnet34(modified) thershold=(fitting local validation), local validation: 0.613 ===&gt; LB: 0.468\nresnet50 thershold=(fitting local validation), local validation: 0.723 ===&gt; LB: 0.47",
    "436028": "Thanks for sharing. Amazing baseline performance.\nMaybe my lr schedule led to overfitting...\nMay I ask what's is your dropout rate ?",
    "436032": "No dropout. The gap between cv and lb puzzles me. I did a lot of experiments on this.",
    "436047": "Same here. Trying to figure out better way to evaluate performance :(",
    "436649": "There are some duplicate or very similar images in the train dataset. These might explain the gap between cv and lb. \nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534",
    "436715": "I believe the the fastai baseline used 256x256.",
    "436790": "Ok. I will change the notebook to see the results.",
    "436856": "I found the fastai baseline kernel is hard to reproduce on other DL framework. Even I use pytorch to reproduce, it is hard to get LB 0.45 with 256 size of pretrained resnet34.",
    "436864": "I currently using Tensorflow. With different threshold (fitting with equal threshold among classes), score vary from 0.42 to 0.45.\nDid you use the similar least square threshold fitting as the kernel in your pytorch baseline ?",
    "436910": "Appian, Actually considering these duplicates doesn't help much since they do not include rare classes if I remember correctly. Right now I'm doing validation with excluding these duplicates from validation set, but the mismatch between public LB and val is exactly the same. Partially it may be explained if some rare classes are not present in the public LB and automatically get zero score in contrast to validation. Another strange thing is that for one of my models I got 0.8+ val that gave only 0.42 public LB, while another approach gives 0.47+ public LB with only 0.69 val (and visual analysis of the val and test probability distributions for each label shows than the separation of classes is much worse that in the first approach).\n\nRegarding the baseline, the [kernel][1] I have posted gives 0.460 (0.454 after LB update) public LB on 256x256 (training took only ~6 hours on K80). I'd say it varies from run to run in 0.445-0.46 range based on different versions of the kernel. On 512x512 images some people mention that they got ~0.5+ using this kernel.\n\n\n  [1]: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb",
    "436955": "Thank you @lafoss for sharing your experiment. Excluding duplicates was exactly I was going to do and nice to know the results you got.\n\nHere is what I got recently. I use label distribution of training data for thresholding for now.\n\nlocal cv: 0.7259\npublic LB: 0.416\n\n- 00 0.814 | 01 0.853 | 02 0.785 | 03 0.773\n- 04 0.817 | 05 0.726 | 06 0.668 | 07 0.827\n- 08 0.545 | 09 1.000 | 10 1.000 | 11 0.840\n- 12 0.674 | 13 0.704 | 14 0.883 | 15 1.000\n- 16 0.443 | 17 0.571 | 18 0.547 | 19 0.700 \n- 20 0.500 | 21 0.701 | 22 0.640 | 23 0.808\n- 24 0.769 | 25 0.688 | 26 0.545 | 27 0.500\n\nAnd just like your result, much lower local cv sometimes gives better public LB.",
    "442551": "InceptionV3 (modified) threshold=0.2, local validation: 0.773 ==&gt; LB: 0.493",
    "442567": ".515 -&gt; .386 for me with InceptionV3. Any tips? I am not using the AuxLogits layer, and have tried swapping the avg_pool layer out for something that would accept variable image size input, but haven't had much luck with 512 on InceptionV3 yet. I am thinking about changing input to accept 4 channels similar to the fastai baseline kernel that got .46.",
    "442582": "try TTA",
    "442590": "Thanks. I have tried TTA, but only with my rotation / reflection augmentation. I have seen others using Brightness, Blur, Noise, and Zoom. Not so sure about zoom, but I'm guessing I would benefit from these additional augmentations.",
    "442596": "In tta (test time) you usually do inference on the  image several times and average the results or something similar. The idea is that the neural network is better in detection of objects in some orientations then others. For tta you usually stick to 90 degrees rotation and flips. As for augmentation during training.... Zoom is tricky, for example, as you might cause a rare kind of protein to be cropped or not appear in the zoomed image.",
    "443120": "You can try my [public kernel][1], train more epochs on your local machine, apply cross-validation (using 6-fold and average the results) -&gt; you can get around 0.47 on LB.\n\nTry to implement some [ideas][2] here and you can get 0.517 on LB.  \n\n\n  [1]: https://www.kaggle.com/mathormad/inceptionv3-baseline-lb-0-379\n  [2]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70478",
    "443134": "Thanks! I'm not sure how I missed your kernel. That is some clean code! I am working on attention now, hopefully can figure it out. In that thread he suggested using red, blue, yellow without green as negative samples to help train the network attention. That seems like an interesting, easy to implement idea.",
    "443336": "Hi, what's the meaning of \"apply cross-validation\"? (sorry i am a newcomer)",
    "443410": "Wang-Xinliang  I believe he is referring to k-fold"
  },
  "source": "meta"
}