{
  "id": 115787,
  "title": "[lb 0.6709] yet another pytorch starter-kit",
  "url": "/competitions/understanding_cloud_organization/discussion/115787",
  "author_name": "hengck23",
  "post_date": "2019-11-05T07:51:55.732000",
  "votes": 96,
  "comment_count": 193,
  "views": 0,
  "content": "<p>what this is not <em>NOT</em>\n- a click and run solution. do not expect just to run the scripts as it is and get the results</p>\n\n<p>what this is \n- the code base does indeed produce the results as mentioned, however, you have to find your hyper parameters and made some modification according to the discussion. you probably have to fill in some simple code yourself like making train splits, etc</p>\n\n<p>why another starter kit?\n- it is sometimes difficult to discuss without reference code.\n- training log is important to analyse model behavior. it is probably the most important reference for beginner to learn how to train a model. the provision of such log is to compare your version with the reference version to see how change of hyper-parameters affect results.\n- another tool is the visualization  of results.\n- this starter kit provides good logging and visualization. reference training log is also provided.</p>\n\n<p><strong>7 days is enough to make good results for this challenge due to the small dataset and noise in the annotation. have fun, work hard and good luck!</strong></p>\n\n<p>please see</p>\n\n<p><a href=\"https://drive.google.com/open?id=1FTAjX-rUDbvb2G8wWyC0YEvhPd5gN-B4\">https://drive.google.com/open?id=1FTAjX-rUDbvb2G8wWyC0YEvhPd5gN-B4</a></p>\n\n<hr>\n\n<p>pre-relase version  (2019-11-05):\n - dirty code, not all function are finished coded yet\n - but you can make  a submission of LB 0.644 using resnet34-fpn and trained with 3 to  4 hrs</p>\n\n<hr>\n\n<p>version.1: (2019-11-10)\n - LB 0.6586 (resnet34-fpn x3 fold, single fold lb 0.656)</p>\n\n<p>version.1: (2019-11-10a)\n- added partial code for resnet34-unet(single fold lb 0.654 @ label threshold 0.6, 0.657 @ 0.7)\n- ensemble of 1xresnet34-unet + 3xfpn : lb 0.6671</p>\n\n<p>note: \n  - you can improve the code by addition different input size in ensemble. Also add double flip = flip lr + up =rotate 180 as additional TTA. you should be about to get lb 0.670~</p>",
  "messages": [
    {
      "id": 665639,
      "postDate": "2019-11-05T07:51:55.733Z",
      "content": "<p>what this is not <em>NOT</em>\n- a click and run solution. do not expect just to run the scripts as it is and get the results</p>\n\n<p>what this is \n- the code base does indeed produce the results as mentioned, however, you have to find your hyper parameters and made some modification according to the discussion. you probably have to fill in some simple code yourself like making train splits, etc</p>\n\n<p>why another starter kit?\n- it is sometimes difficult to discuss without reference code.\n- training log is important to analyse model behavior. it is probably the most important reference for beginner to learn how to train a model. the provision of such log is to compare your version with the reference version to see how change of hyper-parameters affect results.\n- another tool is the visualization  of results.\n- this starter kit provides good logging and visualization. reference training log is also provided.</p>\n\n<p><strong>7 days is enough to make good results for this challenge due to the small dataset and noise in the annotation. have fun, work hard and good luck!</strong></p>\n\n<p>please see</p>\n\n<p><a href=\"https://drive.google.com/open?id=1FTAjX-rUDbvb2G8wWyC0YEvhPd5gN-B4\">https://drive.google.com/open?id=1FTAjX-rUDbvb2G8wWyC0YEvhPd5gN-B4</a></p>\n\n<hr>\n\n<p>pre-relase version  (2019-11-05):\n - dirty code, not all function are finished coded yet\n - but you can make  a submission of LB 0.644 using resnet34-fpn and trained with 3 to  4 hrs</p>\n\n<hr>\n\n<p>version.1: (2019-11-10)\n - LB 0.6586 (resnet34-fpn x3 fold, single fold lb 0.656)</p>\n\n<p>version.1: (2019-11-10a)\n- added partial code for resnet34-unet(single fold lb 0.654 @ label threshold 0.6, 0.657 @ 0.7)\n- ensemble of 1xresnet34-unet + 3xfpn : lb 0.6671</p>\n\n<p>note: \n  - you can improve the code by addition different input size in ensemble. Also add double flip = flip lr + up =rotate 180 as additional TTA. you should be about to get lb 0.670~</p>",
      "rawMarkdown": "what this is not *NOT*\n- a click and run solution. do not expect just to run the scripts as it is and get the results\n\nwhat this is \n- the code base does indeed produce the results as mentioned, however, you have to find your hyper parameters and made some modification according to the discussion. you probably have to fill in some simple code yourself like making train splits, etc\n\nwhy another starter kit?\n- it is sometimes difficult to discuss without reference code.\n- training log is important to analyse model behavior. it is probably the most important reference for beginner to learn how to train a model. the provision of such log is to compare your version with the reference version to see how change of hyper-parameters affect results.\n- another tool is the visualization  of results.\n- this starter kit provides good logging and visualization. reference training log is also provided.\n\n\n**7 days is enough to make good results for this challenge due to the small dataset and noise in the annotation. have fun, work hard and good luck!**\n\nplease see\n\nhttps://drive.google.com/open?id=1FTAjX-rUDbvb2G8wWyC0YEvhPd5gN-B4\n\n\n----\n\npre-relase version  (2019-11-05):\n - dirty code, not all function are finished coded yet\n - but you can make  a submission of LB 0.644 using resnet34-fpn and trained with 3 to  4 hrs\n\n\n----\n\nversion.1: (2019-11-10)\n - LB 0.6586 (resnet34-fpn x3 fold, single fold lb 0.656)\n\nversion.1: (2019-11-10a)\n- added partial code for resnet34-unet(single fold lb 0.654 @ label threshold 0.6, 0.657 @ 0.7)\n- ensemble of 1xresnet34-unet + 3xfpn : lb 0.6671\n\nnote: \n  - you can improve the code by addition different input size in ensemble. Also add double flip = flip lr + up =rotate 180 as additional TTA. you should be about to get lb 0.670~",
      "votes": 96
    },
    {
      "id": 668036,
      "postDate": "2019-11-07T22:26:00.563Z",
      "content": "<p>here is one data leak!</p>\n\n<ul>\n<li>each image has at least one label .... you can use adaptive threshold to ensure each image has at least one type of cloud or you can design a loss function for that (e.g. ranking loss)</li>\n</ul>",
      "rawMarkdown": "here is one data leak!\n\n- each image has at least one label .... you can use adaptive threshold to ensure each image has at least one type of cloud or you can design a loss function for that (e.g. ranking loss)",
      "votes": 9,
      "replies": [
        {
          "id": 668901,
          "postDate": "2019-11-09T03:51:29.447Z",
          "content": "<p>this could be helpful for postprocess\n.6609 model, found 100 images negative, force them to predict mask, improve .6614</p>",
          "rawMarkdown": "this could be helpful for postprocess\n.6609 model, found 100 images negative, force them to predict mask, improve .6614"
        },
        {
          "id": 668902,
          "postDate": "2019-11-09T03:56:03.343Z",
          "content": "<p>Training a softmax classification model for class selection may work</p>",
          "rawMarkdown": "Training a softmax classification model for class selection may work"
        }
      ]
    },
    {
      "id": 669537,
      "postDate": "2019-11-10T06:19:34.493Z",
      "content": "<p><a href=\"https://events.ecmwf.int/event/118/contributions/538/attachments/139/244/OCBWF-Bony.pdf\">https://events.ecmwf.int/event/118/contributions/538/attachments/139/244/OCBWF-Bony.pdf</a>\n<a href=\"https://rmets.onlinelibrary.wiley.com/doi/epdf/10.1002/qj.3662\">https://rmets.onlinelibrary.wiley.com/doi/epdf/10.1002/qj.3662</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe371d138ff40a6d50a9d09d41343d903%2FSelection_053.png?generation=1573366766042602&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd4d6206c17b3abb64bb2fa195e4ecf41%2FSelection_051.png?generation=1573366770599581&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F8527544c55b481b58e7a635cdffd1ed8%2FSelection_052.png?generation=1573366771716980&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "https://events.ecmwf.int/event/118/contributions/538/attachments/139/244/OCBWF-Bony.pdf\nhttps://rmets.onlinelibrary.wiley.com/doi/epdf/10.1002/qj.3662\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe371d138ff40a6d50a9d09d41343d903%2FSelection_053.png?generation=1573366766042602&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd4d6206c17b3abb64bb2fa195e4ecf41%2FSelection_051.png?generation=1573366770599581&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F8527544c55b481b58e7a635cdffd1ed8%2FSelection_052.png?generation=1573366771716980&amp;alt=media)\n",
      "votes": 8
    },
    {
      "id": 671393,
      "postDate": "2019-11-12T16:25:41.047Z",
      "content": "<p>0.6621 single model!</p>\n\n<p>i tried several segmentation head: unet decoder head, fpn head, ASPP head. I think ASPP is the best.</p>\n\n<p>```</p>\n\n<p>class Net(nn.Module):\n    def load_pretrain(self, skip=['logit.'], is_print=True):\n        load_pretrain(self, skip, pretrain_file=PRETRAIN_FILE, conversion=CONVERSION, is_print=is_print)</p>\n\n<pre><code>def __init__(self, num_class=4):\n    super(Net, self).__init__()\n\n    e = ResNet34()\n    self.block0 = e.block0\n    self.block1 = e.block1\n    self.block2 = e.block2\n    self.block3 = e.block3\n    self.block4 = e.block4\n    e = None  #dropped\n\n    self.jpu = JointPyramidUpsample([512,256,128],128)\n    #self.aspp = ASPP(512, 128, rate=[6,12,18], dropout_rate=0.1)\n    self.aspp = ASPP(512, 128, rate=[4,8,12], dropout_rate=0.1)\n    self.logit = nn.Conv2d(128,num_class,kernel_size=1)\n\n\ndef forward(self, x):\n    batch_size,C,H,W = x.shape\n\n    x0 = self.block0(x)\n    x1 = self.block1(x0)\n    x2 = self.block2(x1)\n    x3 = self.block3(x2)\n    x4 = self.block4(x3)\n\n    x = self.jpu([x4,x3,x2])\n    x = self.aspp(x)\n    logit = self.logit(x)\n\n    #---\n    probability_mask  = torch.sigmoid(logit)\n    probability_label = F.adaptive_max_pool2d(probability_mask,1).view(batch_size,-1)\n    return probability_label, probability_mask\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "0.6621 single model!\n\ni tried several segmentation head: unet decoder head, fpn head, ASPP head. I think ASPP is the best.\n\n```\n\nclass Net(nn.Module):\n    def load_pretrain(self, skip=['logit.'], is_print=True):\n        load_pretrain(self, skip, pretrain_file=PRETRAIN_FILE, conversion=CONVERSION, is_print=is_print)\n\n    def __init__(self, num_class=4):\n        super(Net, self).__init__()\n\n        e = ResNet34()\n        self.block0 = e.block0\n        self.block1 = e.block1\n        self.block2 = e.block2\n        self.block3 = e.block3\n        self.block4 = e.block4\n        e = None  #dropped\n\n        self.jpu = JointPyramidUpsample([512,256,128],128)\n        #self.aspp = ASPP(512, 128, rate=[6,12,18], dropout_rate=0.1)\n        self.aspp = ASPP(512, 128, rate=[4,8,12], dropout_rate=0.1)\n        self.logit = nn.Conv2d(128,num_class,kernel_size=1)\n\n\n    def forward(self, x):\n        batch_size,C,H,W = x.shape\n\n        x0 = self.block0(x)\n        x1 = self.block1(x0)\n        x2 = self.block2(x1)\n        x3 = self.block3(x2)\n        x4 = self.block4(x3)\n\n        x = self.jpu([x4,x3,x2])\n        x = self.aspp(x)\n        logit = self.logit(x)\n\n        #---\n        probability_mask  = torch.sigmoid(logit)\n        probability_label = F.adaptive_max_pool2d(probability_mask,1).view(batch_size,-1)\n        return probability_label, probability_mask\n\n\n```",
      "votes": 7,
      "replies": [
        {
          "id": 671395,
          "postDate": "2019-11-12T16:27:44.930Z",
          "content": "<p>submission details:</p>\n\n<p>```</p>\n\n<p>submitting .... @ ['null', 'flip_lr', 'flip_ud', 'flip_both']\ninitial_checkpoint  = /root/share/project/kaggle/2019/cloud/result/run3/resnet34-jpu-aspp-576x384-fold_a2/checkpoint/00012000_model.pth\nthreshold_label = [0.7, 0.7, 0.7, 0.7]\nthreshold_mask  = [0.3, 0.3, 0.3, 0.3]\nthreshold_mask  = [1, 1, 1, 1]</p>\n\n<p>** net setting **\n 3698 / 3698   0 hr 02 min\n 0 min 02 sec\ntest submission .... @ ['null', 'flip_lr', 'flip_ud', 'flip_both']\ninitial_checkpoint=/root/share/project/kaggle/2019/cloud/result/run3/resnet34-jpu-aspp-576x384-fold_a2/checkpoint/00012000_model.pth</p>\n\n<p>compare with LB probing ... \n        num_image =  3698(3698) </p>\n\n<pre><code>    pos0 =  1303(1864)  0.699\n    pos1 =  1410(1508)  0.935\n    pos2 =  1084(1982)  0.547\n    pos3 =  2305(2382)  0.968\n\n    neg0 =  2395(1834)  1.306\n    neg1 =  2288(1940)  1.179\n    neg2 =  2614(2638)  0.991\n</code></pre>\n\n<h2>        neg3 =  1393(2017)  0.691</h2>\n\n<pre><code>    all_zero =   104 (?)\n</code></pre>\n\n<p>```</p>\n\n<p>validation</p>\n\n<p>```\ntest_dataset : \n    len = 300</p>\n\n<pre><code>mode    = train\nsplit   = ['by_random1/valid_fold_a2_300.npy']\ncsv     = ['train.csv']\nfolder  = {'image': '1050x700', 'mask': '525x350'}\nnum_image = 300\n            Fish   neg0, pos0 =   142  (0.473),    158  (0.527)\n          Flower   neg1, pos1 =   179  (0.597),    121  (0.403)\n          Gravel   neg2, pos2 =   144  (0.480),    156  (0.520)\n           Sugar   neg3, pos3 =   101  (0.337),    199  (0.663)\n</code></pre>\n\n<p>submitting .... @ ['null', 'flip_lr', 'flip_ud', 'flip_both']\ninitial_checkpoint  = /root/share/project/kaggle/2019/cloud/result/run3/resnet34-jpu-aspp-576x384-fold_a2/checkpoint/00012000_model.pth\nthreshold_label = [0.7, 0.7, 0.7, 0.7]\nthreshold_mask  = [0.3, 0.3, 0.3, 0.3]\nthreshold_mask  = [1, 1, 1, 1]</p>\n\n<p>** net setting **\n** all threshold **</p>\n\n<pre><code>             |   truth  |  predict |              |              |          \n</code></pre>\n\n<h2>                 | neg  pos | neg  pos | tn     tp    | dn     dp    | kaggle  </h2>\n\n<p>0      Fish     | 142  158 | 196  104 | 0.930  0.595 | 0.000  0.652 | 0.656 (0.761) <br>\n 1    Flower     | 179  121 | 188  112 | 0.911  0.793 | 0.000  0.776 | 0.790 (0.863) <br>\n 2    Gravel     | 144  156 | 213   87 | 0.931  0.494 | 0.000  0.703 | 0.618 (0.696) <br>\n 3     Sugar     | 101  199 | 113  187 | 0.743  0.809 | 0.000  0.686 | 0.622 (0.785)  </p>\n\n<p>kaggle (classification only) = 0.67159 (0.77636)</p>\n\n<p>** segmentation only **</p>\n\n<pre><code>             |   truth  |  predict |              |              |          \n</code></pre>\n\n<h2>                 | neg  pos | neg  pos | tn     tp    | dn     dp    | kaggle  </h2>\n\n<p>0      Fish     | 142  158 |   0  300 | 0.000  1.000 | 0.613  0.457 | 0.534 (0.504) <br>\n 1    Flower     | 179  121 |   0  300 | 0.000  1.000 | 0.754  0.665 | 0.718 (0.408) <br>\n 2    Gravel     | 144  156 |   0  300 | 0.000  1.000 | 0.611  0.477 | 0.539 (0.536) <br>\n 3     Sugar     | 101  199 |   0  300 | 0.000  1.000 | 0.297  0.616 | 0.503 (0.644)  </p>\n\n<p>kaggle (classification only) = 0.57352 (0.52299)</p>\n\n<p>```</p>",
          "rawMarkdown": "submission details:\n\n```\n\nsubmitting .... @ ['null', 'flip_lr', 'flip_ud', 'flip_both']\ninitial_checkpoint  = /root/share/project/kaggle/2019/cloud/result/run3/resnet34-jpu-aspp-576x384-fold_a2/checkpoint/00012000_model.pth\nthreshold_label = [0.7, 0.7, 0.7, 0.7]\nthreshold_mask  = [0.3, 0.3, 0.3, 0.3]\nthreshold_mask  = [1, 1, 1, 1]\n\n** net setting **\n 3698 / 3698   0 hr 02 min\n 0 min 02 sec\ntest submission .... @ ['null', 'flip_lr', 'flip_ud', 'flip_both']\ninitial_checkpoint=/root/share/project/kaggle/2019/cloud/result/run3/resnet34-jpu-aspp-576x384-fold_a2/checkpoint/00012000_model.pth\n\ncompare with LB probing ... \n\t\tnum_image =  3698(3698) \n\n\t\tpos0 =  1303(1864)  0.699\n\t\tpos1 =  1410(1508)  0.935\n\t\tpos2 =  1084(1982)  0.547\n\t\tpos3 =  2305(2382)  0.968\n\n\t\tneg0 =  2395(1834)  1.306\n\t\tneg1 =  2288(1940)  1.179\n\t\tneg2 =  2614(2638)  0.991\n\t\tneg3 =  1393(2017)  0.691\n--------------------------------------------------\n\n\t\tall_zero =   104 (?)\n\n\n```\n\nvalidation\n\n```\ntest_dataset : \n\tlen = 300\n\n\tmode    = train\n\tsplit   = ['by_random1/valid_fold_a2_300.npy']\n\tcsv     = ['train.csv']\n\tfolder  = {'image': '1050x700', 'mask': '525x350'}\n\tnum_image = 300\n\t            Fish   neg0, pos0 =   142  (0.473),    158  (0.527)\n\t          Flower   neg1, pos1 =   179  (0.597),    121  (0.403)\n\t          Gravel   neg2, pos2 =   144  (0.480),    156  (0.520)\n\t           Sugar   neg3, pos3 =   101  (0.337),    199  (0.663)\n\n\nsubmitting .... @ ['null', 'flip_lr', 'flip_ud', 'flip_both']\ninitial_checkpoint  = /root/share/project/kaggle/2019/cloud/result/run3/resnet34-jpu-aspp-576x384-fold_a2/checkpoint/00012000_model.pth\nthreshold_label = [0.7, 0.7, 0.7, 0.7]\nthreshold_mask  = [0.3, 0.3, 0.3, 0.3]\nthreshold_mask  = [1, 1, 1, 1]\n\n** net setting **\n** all threshold **\n\n                 |   truth  |  predict |              |              |          \n                 | neg  pos | neg  pos | tn     tp    | dn     dp    | kaggle  \n----------------------------------------------------------------------------------------\n 0      Fish     | 142  158 | 196  104 | 0.930  0.595 | 0.000  0.652 | 0.656 (0.761)  \n 1    Flower     | 179  121 | 188  112 | 0.911  0.793 | 0.000  0.776 | 0.790 (0.863)  \n 2    Gravel     | 144  156 | 213   87 | 0.931  0.494 | 0.000  0.703 | 0.618 (0.696)  \n 3     Sugar     | 101  199 | 113  187 | 0.743  0.809 | 0.000  0.686 | 0.622 (0.785)  \n\nkaggle (classification only) = 0.67159 (0.77636)\n\n\n** segmentation only **\n\n                 |   truth  |  predict |              |              |          \n                 | neg  pos | neg  pos | tn     tp    | dn     dp    | kaggle  \n----------------------------------------------------------------------------------------\n 0      Fish     | 142  158 |   0  300 | 0.000  1.000 | 0.613  0.457 | 0.534 (0.504)  \n 1    Flower     | 179  121 |   0  300 | 0.000  1.000 | 0.754  0.665 | 0.718 (0.408)  \n 2    Gravel     | 144  156 |   0  300 | 0.000  1.000 | 0.611  0.477 | 0.539 (0.536)  \n 3     Sugar     | 101  199 |   0  300 | 0.000  1.000 | 0.297  0.616 | 0.503 (0.644)  \n\nkaggle (classification only) = 0.57352 (0.52299)\n\n```",
          "votes": 1
        },
        {
          "id": 671426,
          "postDate": "2019-11-12T17:21:05.560Z",
          "content": "<p>We got with UNet a single fold 0.6678. So it means that the decoder type has nothing to do with a high score?</p>",
          "rawMarkdown": "We got with UNet a single fold 0.6678. So it means that the decoder type has nothing to do with a high score?",
          "votes": 2
        },
        {
          "id": 671652,
          "postDate": "2019-11-13T01:56:51.070Z",
          "content": "<p>\"We got with UNet a single fold 0.6678\"</p>\n\n<p>0.6678 is very good score and i am surprised. </p>\n\n<p>in my experiments, the score are quite sensitive. because of the negative images, the metric itself is non smooth. The BCE is smooth however. in my experiment, the ASPP head has a much BCE lower loss than others. so i conclude it is a better decoder type.</p>",
          "rawMarkdown": "\"We got with UNet a single fold 0.6678\"\n\n0.6678 is very good score and i am surprised. \n\nin my experiments, the score are quite sensitive. because of the negative images, the metric itself is non smooth. The BCE is smooth however. in my experiment, the ASPP head has a much BCE lower loss than others. so i conclude it is a better decoder type.\n\n",
          "votes": 2
        },
        {
          "id": 671950,
          "postDate": "2019-11-13T11:10:05.607Z",
          "content": "<p>Your findings look good <a href=\"/hengck23\">@hengck23</a>! What metric do you use for loss?</p>",
          "rawMarkdown": "Your findings look good @hengck23! What metric do you use for loss?"
        },
        {
          "id": 671955,
          "postDate": "2019-11-13T11:19:01.117Z",
          "content": "<p>only BCE. you can check the code for more detail</p>",
          "rawMarkdown": "only BCE. you can check the code for more detail",
          "votes": 3
        },
        {
          "id": 671961,
          "postDate": "2019-11-13T11:26:46.627Z",
          "content": "<p>Interesting. Thanks! Will do.</p>",
          "rawMarkdown": "Interesting. Thanks! Will do."
        },
        {
          "id": 674442,
          "postDate": "2019-11-16T14:15:01.717Z",
          "content": "<p>JPU valid loss looks actually better than UNet. I accidently submitted a resnet34-JPU with a wrong label threshold 0.55 -&gt; LB 0.658. I think with a higher threshold it can achieve a good LB. The problem seems to be it predicts many masks in the test set -&gt; you have to search again some thresholds...</p>",
          "rawMarkdown": "JPU valid loss looks actually better than UNet. I accidently submitted a resnet34-JPU with a wrong label threshold 0.55 -&gt; LB 0.658. I think with a higher threshold it can achieve a good LB. The problem seems to be it predicts many masks in the test set -&gt; you have to search again some thresholds..."
        }
      ]
    },
    {
      "id": 670024,
      "postDate": "2019-11-10T22:55:16.470Z",
      "content": "<p>The competition is going to the final week. Please stop sharing. It's not nice to see solutions keeping posted like this.</p>",
      "rawMarkdown": "The competition is going to the final week. Please stop sharing. It's not nice to see solutions keeping posted like this.",
      "votes": 6,
      "replies": [
        {
          "id": 670914,
          "postDate": "2019-11-12T03:30:59.567Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F07717d002f9a96888b2439d175227cbc%2Fguideline.png?generation=1573529402516418&amp;alt=media\" alt=\"\"></p>\n\n<p>kaggle guideline is stop publishing at the last week. it should be ok to put code publicly before this message comes out.</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F07717d002f9a96888b2439d175227cbc%2Fguideline.png?generation=1573529402516418&amp;alt=media)\n\n\nkaggle guideline is stop publishing at the last week. it should be ok to put code publicly before this message comes out.",
          "votes": 16
        }
      ]
    },
    {
      "id": 665655,
      "postDate": "2019-11-05T08:32:46.167Z",
      "content": "<p>Please, be carefull )</p>",
      "rawMarkdown": "Please, be carefull )",
      "votes": 6
    },
    {
      "id": 670074,
      "postDate": "2019-11-11T00:51:58.773Z",
      "content": "<p>lesson from the steel competition is that  results is going to be sensitive to the image label threshold.</p>\n\n<p>one way to deal with this is to select proper high threshold, (like 0.70).</p>\n\n<p>another way is to modify the loss e.g loss = weight_pos*loss_pos +  weight_neg*loss_neg. By adjusting the weights, you can shift the threshold back to 0.50.</p>\n\n<p>to visualize the effects on classification, you can observe the roc curve of the weighted and un-weighted loss model. indicate tpr, fpr and threshold on your roc curve.</p>\n\n<p>for segmentation, just draw the probability heatmap in range (0,1). you can see the weighted version will be dilated (or eroded) version of the unweighted one.</p>",
      "rawMarkdown": "lesson from the steel competition is that  results is going to be sensitive to the image label threshold.\n\none way to deal with this is to select proper high threshold, (like 0.70).\n\nanother way is to modify the loss e.g loss = weight\\_pos*loss\\_pos +  weight\\_neg*loss\\_neg. By adjusting the weights, you can shift the threshold back to 0.50.\n\nto visualize the effects on classification, you can observe the roc curve of the weighted and un-weighted loss model. indicate tpr, fpr and threshold on your roc curve.\n\nfor segmentation, just draw the probability heatmap in range (0,1). you can see the weighted version will be dilated (or eroded) version of the unweighted one.\n",
      "votes": 4
    },
    {
      "id": 666095,
      "postDate": "2019-11-05T18:14:38.337Z",
      "content": "<p>daily tracking results:</p>\n\n<p>entry date : 2019-11-05</p>\n\n<p>```\nday-1, 2019-11-05: LB 0.644   (non-tuned resnet34-fpn) <br>\nday-3, 2019-11-08: LB 0.6586  (resnet34-fpn x3 fold) . for single fold : lb 0.656\nday-5, 2019-11-10: LB 0.6577  (resnet34-unet single). \n             ensemble of 1 x unet + 3 x fpn : lb 0.667  (finetune threshold, etc)</p>\n\n<p>day-6, 2019-11-11: LB 0.6709 (added a few misc models to ensemble, e.g. different input size, jpu+aspp (joint-pyramid-upsampling))</p>\n\n<p>day-8, 2019-11-13: LB 0.6713 (corrected a bug. wrongly resize image to widthxheight instead of heightxwidth). perform threshold search on LB to use up free submission slots, try 0.60 0.65, 0.67 as image label threshold. the lb score can ranges from  0.6664,0.6707, 0.6705. nevethelss, high threshold is preferred</p>\n\n<p>day-10 2019-11-14: LB 0.6721 (corrected the resize bug again! seems that this bug occurs in different part of my code). add 2 more models to my previous ensemble. i am still using the same models except for for more fold</p>\n\n<p>```</p>",
      "rawMarkdown": "daily tracking results:\n\nentry date : 2019-11-05\n\n```\nday-1, 2019-11-05: LB 0.644   (non-tuned resnet34-fpn)  \nday-3, 2019-11-08: LB 0.6586  (resnet34-fpn x3 fold) . for single fold : lb 0.656\nday-5, 2019-11-10: LB 0.6577  (resnet34-unet single). \n             ensemble of 1 x unet + 3 x fpn : lb 0.667  (finetune threshold, etc)\n\n\nday-6, 2019-11-11: LB 0.6709 (added a few misc models to ensemble, e.g. different input size, jpu+aspp (joint-pyramid-upsampling))\n\n\nday-8, 2019-11-13: LB 0.6713 (corrected a bug. wrongly resize image to widthxheight instead of heightxwidth). perform threshold search on LB to use up free submission slots, try 0.60 0.65, 0.67 as image label threshold. the lb score can ranges from  0.6664,0.6707, 0.6705. nevethelss, high threshold is preferred\n\nday-10 2019-11-14: LB 0.6721 (corrected the resize bug again! seems that this bug occurs in different part of my code). add 2 more models to my previous ensemble. i am still using the same models except for for more fold\n\n\n```",
      "votes": 3
    },
    {
      "id": 671545,
      "postDate": "2019-11-12T21:24:11.957Z",
      "content": "<p>You still are sharing top tips and even partial code for a near gold solutions. Can you please just stop and wait until it ends and share, I’m glad to discuss after that, not now. </p>",
      "rawMarkdown": "You still are sharing top tips and even partial code for a near gold solutions. Can you please just stop and wait until it ends and share, I’m glad to discuss after that, not now. ",
      "votes": 2,
      "replies": [
        {
          "id": 671558,
          "postDate": "2019-11-12T21:42:09.380Z",
          "content": "<p>This competition is really just \"ensemble hell\". I don't think <a href=\"/hengck23\">@hengck23</a> is spoiling much of anything by saying \"1xresnet34-unet + 3xfpn + TTA\" and publishing some scaffolding code.</p>",
          "rawMarkdown": "This competition is really just \"ensemble hell\". I don't think @hengck23 is spoiling much of anything by saying \"1xresnet34-unet + 3xfpn + TTA\" and publishing some scaffolding code.",
          "votes": -1
        },
        {
          "id": 671561,
          "postDate": "2019-11-12T21:47:03.923Z",
          "content": "<p>This is the final week. Imagine if I share some tips to directly reach 0.671 on LB? If he is still sharing, who knows what he will share next? What’s the point of sharing here? Please just STOP.  Why not wait for 6 days and then he can share everything. But he rarely shared after the competition ends. That’s really strange. </p>",
          "rawMarkdown": "This is the final week. Imagine if I share some tips to directly reach 0.671 on LB? If he is still sharing, who knows what he will share next? What’s the point of sharing here? Please just STOP.  Why not wait for 6 days and then he can share everything. But he rarely shared after the competition ends. That’s really strange. ",
          "votes": 1
        },
        {
          "id": 671646,
          "postDate": "2019-11-13T01:42:28.397Z",
          "content": "<p>Seems like the best way to get my first medal is to avoid join competition with Heng 😂\nBut still thankful for his share, I learned a lot.</p>",
          "rawMarkdown": "Seems like the best way to get my first medal is to avoid join competition with Heng 😂\nBut still thankful for his share, I learned a lot."
        },
        {
          "id": 671649,
          "postDate": "2019-11-13T01:52:27.137Z",
          "content": "<p>In the past segment competition, this can easily get you a medal, but now the community is getting familiar with segment task, it's harder, so it push me to learn more knowledge rather than copy my past code.</p>",
          "rawMarkdown": "In the past segment competition, this can easily get you a medal, but now the community is getting familiar with segment task, it's harder, so it push me to learn more knowledge rather than copy my past code.",
          "votes": 4
        }
      ]
    },
    {
      "id": 666792,
      "postDate": "2019-11-06T14:03:28.513Z",
      "content": "<p>my plan  (to be updated as competition proceeds ...)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbd4150e743b7ac4b6eb9a1da4c03229d%2FSelection_053.png?generation=1573049004668282&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "my plan  (to be updated as competition proceeds ...)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbd4150e743b7ac4b6eb9a1da4c03229d%2FSelection_053.png?generation=1573049004668282&amp;alt=media)\n\n",
      "votes": 3
    },
    {
      "id": 676064,
      "postDate": "2019-11-19T00:17:46.960Z",
      "content": "<p>closing remarks:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F851c2335980981005516610ceb5348dc%2FSelection_061.png?generation=1574122453993841&amp;alt=media\" alt=\"\"></p>\n\n<p>best private score this code can get is LB 0.66441\nmethod: \n1. just train with all samples, use 1x  resnet34-unet (384x256), 1x resnet34-jpu-psp(576x384), 1x resnet34-\nfpn(1050x700)</p>\n\n<p>attached: resnet34-jpu-psp</p>\n\n<p>if you have suggestion on  how to improve starter kit and how to improve discussion process (e.g. how to select threshold, how to interpret train log)  to help kagglers, please leave your comments here. thanks!</p>",
      "rawMarkdown": "closing remarks:\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F851c2335980981005516610ceb5348dc%2FSelection_061.png?generation=1574122453993841&amp;alt=media)\n\n\nbest private score this code can get is LB 0.66441\nmethod: \n1. just train with all samples, use 1x  resnet34-unet (384x256), 1x resnet34-jpu-psp(576x384), 1x resnet34-\nfpn(1050x700)\n\nattached: resnet34-jpu-psp\n\n\nif you have suggestion on  how to improve starter kit and how to improve discussion process (e.g. how to select threshold, how to interpret train log)  to help kagglers, please leave your comments here. thanks!",
      "votes": 4,
      "replies": [
        {
          "id": 676099,
          "postDate": "2019-11-19T00:44:43.187Z",
          "content": "<p>Thanks for your notes. I haven't had a chance to read all the material here yet. How did you end up merging threshold and min_size for multiple models? Did you average the values? Use the ones from your best model?</p>",
          "rawMarkdown": "Thanks for your notes. I haven't had a chance to read all the material here yet. How did you end up merging threshold and min_size for multiple models? Did you average the values? Use the ones from your best model?"
        },
        {
          "id": 676102,
          "postDate": "2019-11-19T00:50:23.470Z",
          "content": "<p>use average results in ensemble. no classifier, no min size. use max pixel probability to threshold against 0.65 to get image level label</p>",
          "rawMarkdown": "use average results in ensemble. no classifier, no min size. use max pixel probability to threshold against 0.65 to get image level label",
          "votes": 1
        },
        {
          "id": 676109,
          "postDate": "2019-11-19T00:57:42.483Z",
          "content": "<p>You maybe underestimating your code...the best LB score that we got from your code is <code>0.66520</code>, although it was not the one we selected :(\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F93d269dd7c0012a7d850e24ae90a970c%2Fheng.png?generation=1574124988572045&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "You maybe underestimating your code...the best LB score that we got from your code is `0.66520`, although it was not the one we selected :(\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F93d269dd7c0012a7d850e24ae90a970c%2Fheng.png?generation=1574124988572045&amp;alt=media)\n"
        },
        {
          "id": 676111,
          "postDate": "2019-11-19T01:00:50.133Z",
          "content": "<p>that is very good news!\nthis is the reason that i share my code (and ideas) before the competition ends.</p>\n\n<p>previous experiences always show that others can take my code (and ideas) to higher levels.</p>\n\n<p>good work and thanks!</p>",
          "rawMarkdown": "that is very good news!\nthis is the reason that i share my code (and ideas) before the competition ends.\n\nprevious experiences always show that others can take my code (and ideas) to higher levels.\n\ngood work and thanks!",
          "votes": 3
        },
        {
          "id": 676112,
          "postDate": "2019-11-19T01:03:07.743Z",
          "content": "<p>Thank you for your amazing contribution. It wouldn't have been possible without your ideas/codes. We will write a post summarizing our experience and learning from your codes/ideas</p>",
          "rawMarkdown": "Thank you for your amazing contribution. It wouldn't have been possible without your ideas/codes. We will write a post summarizing our experience and learning from your codes/ideas"
        }
      ]
    },
    {
      "id": 675646,
      "postDate": "2019-11-18T11:13:30.467Z",
      "content": "<p>2 new arvix paper today:</p>\n\n<p><a href=\"https://arxiv.org/pdf/1911.06357.pdf\">https://arxiv.org/pdf/1911.06357.pdf</a>\n\"Give me (un)certainty - An exploration of parameters\nthat affect segmentation uncertainty\"</p>\n\n<p><a href=\"https://arxiv.org/pdf/1911.06667.pdf\">https://arxiv.org/pdf/1911.06667.pdf</a>\n\"CenterMask:Real-Time Anchor-Free Instance Segmentation\"</p>",
      "rawMarkdown": "2 new arvix paper today:\n\nhttps://arxiv.org/pdf/1911.06357.pdf\n\"Give me (un)certainty - An exploration of parameters\nthat affect segmentation uncertainty\"\n\nhttps://arxiv.org/pdf/1911.06667.pdf\n\"CenterMask:Real-Time Anchor-Free Instance Segmentation\"",
      "votes": 4
    },
    {
      "id": 670923,
      "postDate": "2019-11-12T03:55:55.943Z",
      "content": "<p>Thanks, I'm getting better at reading your code😹 😹 </p>",
      "rawMarkdown": "Thanks, I'm getting better at reading your code😹 😹 ",
      "votes": 4,
      "replies": [
        {
          "id": 671239,
          "postDate": "2019-11-12T12:18:55.493Z",
          "content": "<p>Hey , mate , I  also   read  your   code  before 😂 😂 </p>",
          "rawMarkdown": "Hey , mate , I  also   read  your   code  before 😂 😂 "
        },
        {
          "id": 671692,
          "postDate": "2019-11-13T04:02:25.780Z",
          "content": "<p>Oh, That's my pleasure:)</p>",
          "rawMarkdown": "Oh, That's my pleasure:)"
        }
      ]
    },
    {
      "id": 672685,
      "postDate": "2019-11-14T04:18:13.163Z",
      "content": "<p><a href=\"/joshvarty\">@joshvarty</a> </p>\n\n<p>Kaggle RSNA challenge just ended today, and it gives me idea on how to use the satellite image  pair. for example instead of predicting single cloud image, predict pair of image together!!</p>\n\n<p>if you have timestamp or location stamp, you can predict them together. </p>\n\n<p>we assume there is some correlation between images taken are different time or location. e.g. if it is fish in latitude A, it is likely to be sugar in latitude B? e.g. fish today flower tomorrow?</p>\n\n<p>```</p>\n\n<p>input = concate[  satellite1 image, satellite2 image ] \n[  satellite1 mask, satellite2 mask] = net(input)</p>\n\n<p>other combinations:\ninput = concate[  first hour, 2nd hour, ....] \ninput = concate[  location 1, location 2, ....] \n```</p>\n\n<p><a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117235#latest-672655\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117235#latest-672655</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117228#latest-672613\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117228#latest-672613</a></p>",
      "rawMarkdown": "@joshvarty \n\nKaggle RSNA challenge just ended today, and it gives me idea on how to use the satellite image  pair. for example instead of predicting single cloud image, predict pair of image together!!\n\n\nif you have timestamp or location stamp, you can predict them together. \n\nwe assume there is some correlation between images taken are different time or location. e.g. if it is fish in latitude A, it is likely to be sugar in latitude B? e.g. fish today flower tomorrow?\n\n```\n\ninput = concate[  satellite1 image, satellite2 image ] \n[  satellite1 mask, satellite2 mask] = net(input)\n\nother combinations:\ninput = concate[  first hour, 2nd hour, ....] \ninput = concate[  location 1, location 2, ....] \n```\n\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117235#latest-672655\n\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117228#latest-672613",
      "votes": 2,
      "replies": [
        {
          "id": 672691,
          "postDate": "2019-11-14T04:23:00.987Z",
          "content": "<p>on a side note, TTA is what we usually do. but who knows if there is a better combination than just averaging or temperature sharpening?</p>\n\n<p>```\n mask = net(input)\n mask_tta1= net(input_tta1)\n mask_tta2= net(input_tta2)\n...\nmask_final = fuse_net([mask , mask_tta1, mask_tta2 ...])</p>\n\n<p>```</p>",
          "rawMarkdown": "on a side note, TTA is what we usually do. but who knows if there is a better combination than just averaging or temperature sharpening?\n\n```\n mask = net(input)\n mask_tta1= net(input_tta1)\n mask_tta2= net(input_tta2)\n...\nmask_final = fuse_net([mask , mask_tta1, mask_tta2 ...])\n\n\n```"
        }
      ]
    },
    {
      "id": 670198,
      "postDate": "2019-11-11T06:55:30.203Z",
      "content": "<p>overfitting</p>\n\n<p>```\nresnet34-unet</p>\n\n<p>constant learning rate=0.01\nbatch_size=16,  iter_accum=2 (100 iter is about 0.3 epoch)</p>\n\n<p>train_dataset : \n    len = 5246</p>\n\n<pre><code>mode    = train\nsplit   = ['by_random1/train_fold_a2_5246.npy']\ncsv     = ['train.csv']\nfolder  = {'image': '1050x700', 'mask': '525x350'}\nnum_image = 5246\n            Fish   neg0, pos0 =  2623  (0.500),   2623  (0.500)\n          Flower   neg1, pos1 =  3002  (0.572),   2244  (0.428)\n          Gravel   neg2, pos2 =  2463  (0.470),   2783  (0.530)\n           Sugar   neg3, pos3 =  1694  (0.323),   3552  (0.677)\n</code></pre>\n\n<p>valid_dataset : \n    len = 300</p>\n\n<pre><code>mode    = train\nsplit   = ['by_random1/valid_fold_a2_300.npy']\ncsv     = ['train.csv']\nfolder  = {'image': '1050x700', 'mask': '525x350'}\nnum_image = 300\n            Fish   neg0, pos0 =   142  (0.473),    158  (0.527)\n          Flower   neg1, pos1 =   179  (0.597),    121  (0.403)\n          Gravel   neg2, pos2 =   144  (0.480),    156  (0.520)\n           Sugar   neg3, pos3 =   101  (0.337),    199  (0.663)\n</code></pre>\n\n<p>```</p>\n\n<p>legend is wrong: </p>\n\n<p>orange=validation kaggle score\ngray=train loss\nblue=validation loss</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe194a1b8fad8eb11ae4aca98ef7ce392%2FSelection_060.png?generation=1573455219875682&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "overfitting\n\n```\nresnet34-unet\n\nconstant learning rate=0.01\nbatch_size=16,  iter_accum=2 (100 iter is about 0.3 epoch)\n\n\ntrain_dataset : \n\tlen = 5246\n\n\tmode    = train\n\tsplit   = ['by_random1/train_fold_a2_5246.npy']\n\tcsv     = ['train.csv']\n\tfolder  = {'image': '1050x700', 'mask': '525x350'}\n\tnum_image = 5246\n\t            Fish   neg0, pos0 =  2623  (0.500),   2623  (0.500)\n\t          Flower   neg1, pos1 =  3002  (0.572),   2244  (0.428)\n\t          Gravel   neg2, pos2 =  2463  (0.470),   2783  (0.530)\n\t           Sugar   neg3, pos3 =  1694  (0.323),   3552  (0.677)\n\nvalid_dataset : \n\tlen = 300\n\n\tmode    = train\n\tsplit   = ['by_random1/valid_fold_a2_300.npy']\n\tcsv     = ['train.csv']\n\tfolder  = {'image': '1050x700', 'mask': '525x350'}\n\tnum_image = 300\n\t            Fish   neg0, pos0 =   142  (0.473),    158  (0.527)\n\t          Flower   neg1, pos1 =   179  (0.597),    121  (0.403)\n\t          Gravel   neg2, pos2 =   144  (0.480),    156  (0.520)\n\t           Sugar   neg3, pos3 =   101  (0.337),    199  (0.663)\n```\n\nlegend is wrong: \n\norange=validation kaggle score\ngray=train loss\nblue=validation loss\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe194a1b8fad8eb11ae4aca98ef7ce392%2FSelection_060.png?generation=1573455219875682&amp;alt=media)\n",
      "votes": 2,
      "replies": [
        {
          "id": 670230,
          "postDate": "2019-11-11T07:43:21.553Z",
          "content": "<p>the reason why it can overfit at constant high learning rate is as follows:</p>\n\n<ul>\n<li>the cloud problem is a texture classification problem.</li>\n<li>we are looking at the texture at \"some large scale\", we ignore differences lower than \"this scale\".</li>\n<li>but model at good are picking differences at small scale. these are nuisance feature as they do not generalize and can cause overfitting</li>\n</ul>\n\n<p>hence to build a good model, scale is important here i think. </p>",
          "rawMarkdown": "the reason why it can overfit at constant high learning rate is as follows:\n\n- the cloud problem is a texture classification problem.\n- we are looking at the texture at \"some large scale\", we ignore differences lower than \"this scale\".\n- but model at good are picking differences at small scale. these are nuisance feature as they do not generalize and can cause overfitting\n\nhence to build a good model, scale is important here i think. \n ",
          "votes": 1
        },
        {
          "id": 671712,
          "postDate": "2019-11-13T04:29:55.850Z",
          "content": "<p>one thing to mention here is that resnet34 unet is trained with augmentation. </p>\n\n<p>unlike overfitting resnet18 unet where no augmentation is used. also, with augmentation, resnet18 is less likely to overfit. hence good augmentation also prevent overfitting</p>",
          "rawMarkdown": "one thing to mention here is that resnet34 unet is trained with augmentation. \n\nunlike overfitting resnet18 unet where no augmentation is used. also, with augmentation, resnet18 is less likely to overfit. hence good augmentation also prevent overfitting",
          "votes": 2
        },
        {
          "id": 671723,
          "postDate": "2019-11-13T04:43:00.513Z",
          "content": "<p>to find augmentation that prevent overfits, you can try the below:\n(note this note necessary improve loss as it can lead to underfitting?)</p>\n\n<ol>\n<li><p>for a trained model, perform new test  augmentation. it can be: just add gaussian noise,  just scale, just zero out random part, etc ....</p></li>\n<li><p>measure the effect,change of loss, score, etc of these augmentation. you can sorted them in decreasing \"change of loss\".</p></li>\n<li><p>if the \"change of loss\" is within certain limits, they are probability safe to added for training.</p></li>\n</ol>\n\n<p>rather than \" just add gaussian noise,  just scale, just zero out random part, etc ....\", some adversarial loss should be better?</p>",
          "rawMarkdown": "to find augmentation that prevent overfits, you can try the below:\n(note this note necessary improve loss as it can lead to underfitting?)\n\n1. for a trained model, perform new test  augmentation. it can be: just add gaussian noise,  just scale, just zero out random part, etc ....\n\n2. measure the effect,change of loss, score, etc of these augmentation. you can sorted them in decreasing \"change of loss\".\n\n3. if the \"change of loss\" is within certain limits, they are probability safe to added for training.\n\nrather than \" just add gaussian noise,  just scale, just zero out random part, etc ....\", some adversarial loss should be better?\n\n ",
          "votes": 2
        }
      ]
    },
    {
      "id": 669929,
      "postDate": "2019-11-10T17:47:07.353Z",
      "content": "<p>I wonder if the jigsaw puzzle trick can be used here? The black slant is a good clue to stich back the big image </p>",
      "rawMarkdown": "I wonder if the jigsaw puzzle trick can be used here? The black slant is a good clue to stich back the big image ",
      "votes": 1,
      "replies": [
        {
          "id": 669931,
          "postDate": "2019-11-10T17:56:10.643Z",
          "rawMarkdown": "",
          "votes": 1
        },
        {
          "id": 669987,
          "postDate": "2019-11-10T20:53:26.310Z",
          "content": "<p>My understanding is its not purely 3 locations that were imaged though, right? I thought it was several tiles from that specific region. So then you could jigsaw them together. You know one tile came above the other and you could take the top half of the bottom one and the bottom half of the top one and then you've got a new combination of the two spliced together images</p>",
          "rawMarkdown": "My understanding is its not purely 3 locations that were imaged though, right? I thought it was several tiles from that specific region. So then you could jigsaw them together. You know one tile came above the other and you could take the top half of the bottom one and the bottom half of the top one and then you've got a new combination of the two spliced together images"
        },
        {
          "id": 670070,
          "postDate": "2019-11-11T00:33:53.737Z",
          "content": "<p>besides, you can fill in the missing black region from one satellite to another </p>",
          "rawMarkdown": "besides, you can fill in the missing black region from one satellite to another "
        }
      ]
    },
    {
      "id": 671913,
      "postDate": "2019-11-13T10:10:00.180Z",
      "content": "<p>In your training log, there is a <code>kaggle</code> columns which prints out two values? what do they indicate? Higher values means better?</p>",
      "rawMarkdown": "In your training log, there is a `kaggle` columns which prints out two values? what do they indicate? Higher values means better?",
      "votes": 1,
      "replies": [
        {
          "id": 671919,
          "postDate": "2019-11-13T10:16:22.763Z",
          "content": "<p>higher value is better.</p>\n\n<p>the first column is the kaggle score you would get on the leader board (classification + segmentation)</p>\n\n<p>the second column is only classification. (segmentation is assumed have perfect dice=1). it represents the upper limit if you try to freeze your classification results and improve segmentation only (e.g. via ensemble, post-processing)</p>",
          "rawMarkdown": "higher value is better.\n\nthe first column is the kaggle score you would get on the leader board (classification + segmentation)\n\nthe second column is only classification. (segmentation is assumed have perfect dice=1). it represents the upper limit if you try to freeze your classification results and improve segmentation only (e.g. via ensemble, post-processing)",
          "votes": 2
        }
      ]
    },
    {
      "id": 667136,
      "postDate": "2019-11-06T21:22:45.910Z",
      "content": "<p>Good luck!</p>",
      "rawMarkdown": "Good luck!",
      "votes": 1
    },
    {
      "id": 670846,
      "postDate": "2019-11-12T00:01:22.447Z",
      "content": "<p>i wonder what do you get if you train only on images like these:</p>\n\n<p>train image contains only positive region only. maybe less noise?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fb8f393144067e469ee65ef984966951b%2F__results___28_0.png?generation=1573533605694538&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "i wonder what do you get if you train only on images like these:\n\ntrain image contains only positive region only. maybe less noise?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fb8f393144067e469ee65ef984966951b%2F__results___28_0.png?generation=1573533605694538&amp;alt=media)\n",
      "votes": 2,
      "replies": [
        {
          "id": 671235,
          "postDate": "2019-11-12T12:12:59.150Z",
          "content": "<p>Well model might learn that the places that are not black are always positive</p>",
          "rawMarkdown": "Well model might learn that the places that are not black are always positive",
          "votes": 2
        },
        {
          "id": 671659,
          "postDate": "2019-11-13T02:24:03.083Z",
          "content": "<p>how about using 2 dataset in stages, e.g. original  and masked-out version. use them only at pretrain, train, finetune ... there are many possible combinations and possibly  different results</p>\n\n<p>how about  training  with  original  + masked-out version, in this case the masked out is a form of augmentation</p>",
          "rawMarkdown": "how about using 2 dataset in stages, e.g. original  and masked-out version. use them only at pretrain, train, finetune ... there are many possible combinations and possibly  different results\n\nhow about  training  with  original  + masked-out version, in this case the masked out is a form of augmentation",
          "votes": 2
        }
      ]
    },
    {
      "id": 669247,
      "postDate": "2019-11-09T18:47:27.600Z",
      "content": "<p>how i overfit my model\n- use resnet18-unet\n- disable all augmentation\n- train for long epoch .... close 100 ...until loss = NAN</p>\n\n<p>i am surprised that even resnet18 can overfits. below shows how overfitting on train images looks like.\n(it also means we can clean up and create new \"correct label\" from these  train results just before  overfitting occurs.  however, clean label may not benefit the challenge as the test ground truth are also noisy)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F7e8a1370185a67ad02f937c6be034480%2F00031.png?generation=1573325055726255&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F24740b97fd2fcda9bbf457f633018ab9%2F00015.png?generation=1573325032499993&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ffd39a9c3b9181d36b230598620b065a1%2F00000.png?generation=1573324976914170&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "how i overfit my model\n- use resnet18-unet\n- disable all augmentation\n- train for long epoch .... close 100 ...until loss = NAN\n\ni am surprised that even resnet18 can overfits. below shows how overfitting on train images looks like.\n(it also means we can clean up and create new \"correct label\" from these  train results just before  overfitting occurs.  however, clean label may not benefit the challenge as the test ground truth are also noisy)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F7e8a1370185a67ad02f937c6be034480%2F00031.png?generation=1573325055726255&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F24740b97fd2fcda9bbf457f633018ab9%2F00015.png?generation=1573325032499993&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ffd39a9c3b9181d36b230598620b065a1%2F00000.png?generation=1573324976914170&amp;alt=media)\n",
      "replies": [
        {
          "id": 669256,
          "postDate": "2019-11-09T19:06:27.423Z",
          "content": "<p>overfitting reveals the inlier and outlier</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F2352e0b2c0b22c74b5efae2a2b83e77c%2F00023.png?generation=1573326264546928&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "overfitting reveals the inlier and outlier\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F2352e0b2c0b22c74b5efae2a2b83e77c%2F00023.png?generation=1573326264546928&amp;alt=media)\n\n"
        },
        {
          "id": 669263,
          "postDate": "2019-11-09T19:20:11.547Z",
          "content": "<p>A few examples of pairs of images for the same day/region:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2Fcf24962bc79e1b97507a7dc27b26504e%2Fim5.png?generation=1573327168294730&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2F9645162abf3ff12a974c17461b13a156%2Fim0.png?generation=1573327099727112&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2Fd9301cf3e1a0305837ded969cefd95c2%2Fim1.png?generation=1573327112040475&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2F7300c9859825ba029e0c9f2a04bc5bbb%2Fim2.png?generation=1573327123327558&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2Fcb7c5285aa0b26fdad39907d65473f37%2Fim3.png?generation=1573327136181854&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2F436733a096828758aaa10b0ac4cbbe11%2Fim4.png?generation=1573327149580874&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "A few examples of pairs of images for the same day/region:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2Fcf24962bc79e1b97507a7dc27b26504e%2Fim5.png?generation=1573327168294730&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2F9645162abf3ff12a974c17461b13a156%2Fim0.png?generation=1573327099727112&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2Fd9301cf3e1a0305837ded969cefd95c2%2Fim1.png?generation=1573327112040475&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2F7300c9859825ba029e0c9f2a04bc5bbb%2Fim2.png?generation=1573327123327558&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2Fcb7c5285aa0b26fdad39907d65473f37%2Fim3.png?generation=1573327136181854&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2F436733a096828758aaa10b0ac4cbbe11%2Fim4.png?generation=1573327149580874&amp;alt=media)\n"
        },
        {
          "id": 669394,
          "postDate": "2019-11-10T01:40:00.470Z",
          "content": "<p>thanks. this explain</p>\n\n<ol>\n<li>the presence of label noise\n2.why label smoothing is not required\n3.why disable augmentation overfit .with almost zero loss... the network may have learned the black band</li>\n</ol>\n\n<p>i wonder if an annotator only label one satellite or both. if a annotator only deal with single satellite, training individual classifier may work?</p>",
          "rawMarkdown": "\n\nthanks. this explain\n\n1. the presence of label noise\n2.why label smoothing is not required\n3.why disable augmentation overfit .with almost zero loss... the network may have learned the black band\n\ni wonder if an annotator only label one satellite or both. if a annotator only deal with single satellite, training individual classifier may work?"
        },
        {
          "id": 669990,
          "postDate": "2019-11-10T21:05:34.830Z",
          "content": "<blockquote>\n  <p>i wonder if an annotator only label one satellite or both. if a annotator only deal with single satellite, training individual classifier may work?</p>\n</blockquote>\n\n<p>Annotators label images from both satellites. You can see how they annotated them here:</p>\n\n<p><a href=\"https://www.zooniverse.org/projects/raspstephan/sugar-flower-fish-or-gravel/classify\">https://www.zooniverse.org/projects/raspstephan/sugar-flower-fish-or-gravel/classify</a></p>",
          "rawMarkdown": "&gt; i wonder if an annotator only label one satellite or both. if a annotator only deal with single satellite, training individual classifier may work?\n\nAnnotators label images from both satellites. You can see how they annotated them here:\n\nhttps://www.zooniverse.org/projects/raspstephan/sugar-flower-fish-or-gravel/classify"
        },
        {
          "id": 670075,
          "postDate": "2019-11-11T00:57:20.063Z",
          "content": "<p>@Miguel Pinto</p>\n\n<p>how you get the date and domain of the images? is it provided?</p>\n\n<p>have you download the same images from the website for the website, you may have larger image and more information (like cloud precipitation) for that.</p>",
          "rawMarkdown": "@Miguel Pinto\n \nhow you get the date and domain of the images? is it provided?\n\nhave you download the same images from the website for the website, you may have larger image and more information (like cloud precipitation) for that.\n"
        },
        {
          "id": 670382,
          "postDate": "2019-11-11T12:04:04.003Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> I downloaded all images from worldview (identical to the competition data) for the regions and period described in the paper about the data. Then I used perceptual hashing to match downloaded images to the identical ones provided in the competition data. I didn't download full-size images from worldview, 350 x 525 is faster to download and is enough to find the pairs. </p>",
          "rawMarkdown": "@hengck23 I downloaded all images from worldview (identical to the competition data) for the regions and period described in the paper about the data. Then I used perceptual hashing to match downloaded images to the identical ones provided in the competition data. I didn't download full-size images from worldview, 350 x 525 is faster to download and is enough to find the pairs. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 668561,
      "postDate": "2019-11-08T15:05:25.203Z",
      "content": "<p>some submission statistics</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F509471fed377077574518f8c124b140a%2FSelection_058.png?generation=1573225729280934&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "some submission statistics\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F509471fed377077574518f8c124b140a%2FSelection_058.png?generation=1573225729280934&amp;alt=media)\n",
      "replies": [
        {
          "id": 669918,
          "postDate": "2019-11-10T17:16:37.710Z",
          "content": "<p>how lb 0.6671 looks like:( there is a mistake in the spreadsheet. pos recall should be 0.64647, )</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ff437c2c01db64abc8dce18b34c7b758a%2FSelection_057.png?generation=1573406195419154&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "how lb 0.6671 looks like:( there is a mistake in the spreadsheet. pos recall should be 0.64647, )\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ff437c2c01db64abc8dce18b34c7b758a%2FSelection_057.png?generation=1573406195419154&amp;alt=media)\n"
        },
        {
          "id": 669937,
          "postDate": "2019-11-10T18:14:16.643Z",
          "content": "<blockquote>\n  <p>how lb 0.6671 looks like:</p>\n</blockquote>\n\n<p>Publishing a silver level result with 8 days to go sounds more like a “finisher-kit” than a “starter-kit” ?</p>",
          "rawMarkdown": "&gt; how lb 0.6671 looks like:\n\nPublishing a silver level result with 8 days to go sounds more like a “finisher-kit” than a “starter-kit” ?",
          "votes": 10
        },
        {
          "id": 669958,
          "postDate": "2019-11-10T19:08:26.697Z",
          "content": "<p>His single best model score is 0.657. There are public kernels which score same range. So I wouldn't say this is a \"finisher-kit\". What impressive is that he can get 0.6671 from such lower score models...</p>",
          "rawMarkdown": "His single best model score is 0.657. There are public kernels which score same range. So I wouldn't say this is a \"finisher-kit\". What impressive is that he can get 0.6671 from such lower score models...",
          "votes": -5
        },
        {
          "id": 669988,
          "postDate": "2019-11-10T20:57:57.140Z",
          "content": "<p>it's not \"starter-kit\" ...</p>",
          "rawMarkdown": "it's not \"starter-kit\" ...",
          "votes": 1
        },
        {
          "id": 670013,
          "postDate": "2019-11-10T22:21:28.200Z",
          "content": "<p>what's the point here? \ntips are cool\nstarter codes are cool\na complete silver medal solution is not cool</p>",
          "rawMarkdown": "what's the point here? \ntips are cool\nstarter codes are cool\na complete silver medal solution is not cool",
          "votes": 1
        },
        {
          "id": 670066,
          "postDate": "2019-11-11T00:23:12.210Z",
          "content": "<p>\"Publishing a silver level result with 8 days to go sounds more like a “finisher-kit” than a “starter-kit” ?\"  </p>\n\n<p>you can't download my code and run as it is.  </p>\n\n<p>some files are intentionally left missing, but you can create them yourself, e.g. split file. in particular, the training iteration don't automatically stop. you will have to find the hyper-parameters yourself. </p>\n\n<p>What i would like to provide is a possible direction for new kagglers to follow and not a \"click and run solution\"</p>",
          "rawMarkdown": "\"Publishing a silver level result with 8 days to go sounds more like a “finisher-kit” than a “starter-kit” ?\"  \n\nyou can't download my code and run as it is.  \n\nsome files are intentionally left missing, but you can create them yourself, e.g. split file. in particular, the training iteration don't automatically stop. you will have to find the hyper-parameters yourself. \n\nWhat i would like to provide is a possible direction for new kagglers to follow and not a \"click and run solution\"",
          "votes": 10
        },
        {
          "id": 670134,
          "postDate": "2019-11-11T04:03:39.140Z",
          "content": "<p>I   think  the  “finisher-kit”   here  means  that  you  will  get   higer  score .</p>",
          "rawMarkdown": "I   think  the  “finisher-kit”   here  means  that  you  will  get   higer  score .",
          "votes": 3
        },
        {
          "id": 670155,
          "postDate": "2019-11-11T05:09:44.153Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa77f2aff2081e2851ab566e81f818f31%2FSelection_056.png?generation=1573549207608697&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": " ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa77f2aff2081e2851ab566e81f818f31%2FSelection_056.png?generation=1573549207608697&amp;alt=media)\n",
          "votes": 1
        },
        {
          "id": 670234,
          "postDate": "2019-11-11T07:49:11.387Z",
          "content": "<p>take a look at 'steel' competition private lb, lots of silver medal on public lb dropped far away, I image the same thing happen here</p>",
          "rawMarkdown": "take a look at 'steel' competition private lb, lots of silver medal on public lb dropped far away, I image the same thing happen here"
        },
        {
          "id": 670241,
          "postDate": "2019-11-11T07:52:42.470Z",
          "content": "<p>i think the shakeup will be much smaller for cloud here.</p>\n\n<p>there is a lot of difference when you can see and cannot see both public+private test data.</p>",
          "rawMarkdown": "i think the shakeup will be much smaller for cloud here.\n\nthere is a lot of difference when you can see and cannot see both public+private test data.\n\n",
          "votes": 1
        },
        {
          "id": 670248,
          "postDate": "2019-11-11T08:05:34.993Z",
          "content": "<p>Am I missing something? What is the dice column in your image? Not sure where that is coming from. Seems very high. </p>",
          "rawMarkdown": "Am I missing something? What is the dice column in your image? Not sure where that is coming from. Seems very high. "
        },
        {
          "id": 670251,
          "postDate": "2019-11-11T08:11:00.723Z",
          "content": "<p>dice of the positive classified image. it is value of pos_dice in the formula below:</p>\n\n<p>```\nkaggle_score = avergae(\n     tnr *num_true_negative _image +\\\n     tpr *num_true_positive_image*pos_dice +\\\n    (1-tnr)*num_true_negative_image*neg_dice +\\\n)</p>\n\n<p>```</p>",
          "rawMarkdown": "dice of the positive classified image. it is value of pos\\_dice in the formula below:\n\n```\nkaggle_score = avergae(\n     tnr *num_true_negative _image +\\\n     tpr *num_true_positive_image*pos_dice +\\\n    (1-tnr)*num_true_negative_image*neg_dice +\\\n)\n\n\n```\n\n"
        },
        {
          "id": 670695,
          "postDate": "2019-11-11T18:46:16.737Z",
          "content": "<p>bestfitting score is a good guide. he is unusually unshaken and he don't overfits.</p>",
          "rawMarkdown": "bestfitting score is a good guide. he is unusually unshaken and he don't overfits."
        },
        {
          "id": 670766,
          "postDate": "2019-11-11T20:14:12.603Z",
          "content": "<blockquote>\n  <p>bestfitting score is a good guide. he is unusually unshaken and he don't overfits.</p>\n</blockquote>\n\n<p>This sounds like top teams are overfitting and we might expect shakeup again?</p>",
          "rawMarkdown": "&gt; bestfitting score is a good guide. he is unusually unshaken and he don't overfits.\n\nThis sounds like top teams are overfitting and we might expect shakeup again?"
        },
        {
          "id": 670963,
          "postDate": "2019-11-12T05:00:39.923Z",
          "content": "<p>@Bibek</p>\n\n<p>there will not be major  shakeup. </p>",
          "rawMarkdown": "@Bibek\n\nthere will not be major  shakeup. "
        }
      ]
    },
    {
      "id": 672162,
      "postDate": "2019-11-13T15:15:09.460Z",
      "content": "<p><a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/70421\">https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/70421</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc800e5ea0c5d05dc4d44e4a9a9ad114f%2FSelection_051.png?generation=1573658106274069&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/70421\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc800e5ea0c5d05dc4d44e4a9a9ad114f%2FSelection_051.png?generation=1573658106274069&amp;alt=media)\n",
      "votes": 1,
      "replies": [
        {
          "id": 672165,
          "postDate": "2019-11-13T15:18:00.787Z",
          "content": "<p>this is how they handle intersection of box (different from our cases)</p>\n\n<p>\"the intersection of boxes was used as opposed to the average. We tried to mimic this intersection process by taking the average of multiple intersections of boxes across models. This gave ~10-15% (!) improvement on stage 1 public LB. As an alternative to this, we simply resized the boxes by multiplying length/width by a fraction. We found 87.5% for each was a good reduction and worked a bit better than doing the intersection. It was also much easier to implement. We discovered this early on, and didn't submit anything without resizing after that\"</p>\n\n<p>... should we expand our box in ground truth ???</p>\n\n<p>it my code and experiment, it actually used 0.3 for threshold which works better. this is essentially a mask dilation </p>",
          "rawMarkdown": "this is how they handle intersection of box (different from our cases)\n\n\"the intersection of boxes was used as opposed to the average. We tried to mimic this intersection process by taking the average of multiple intersections of boxes across models. This gave ~10-15% (!) improvement on stage 1 public LB. As an alternative to this, we simply resized the boxes by multiplying length/width by a fraction. We found 87.5% for each was a good reduction and worked a bit better than doing the intersection. It was also much easier to implement. We discovered this early on, and didn't submit anything without resizing after that\"\n\n... should we expand our box in ground truth ???\n\nit my code and experiment, it actually used 0.3 for threshold which works better. this is essentially a mask dilation "
        }
      ]
    },
    {
      "id": 668913,
      "postDate": "2019-11-09T04:43:07.590Z",
      "content": "<p>when zero is not actually zero?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F02f6f6dc79fd8f87852ec5100df6d89d%2FSelection_061.png?generation=1573274564563490&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "when zero is not actually zero?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F02f6f6dc79fd8f87852ec5100df6d89d%2FSelection_061.png?generation=1573274564563490&amp;alt=media)\n",
      "votes": 1,
      "replies": [
        {
          "id": 669059,
          "postDate": "2019-11-09T12:09:17.997Z",
          "content": "<p>According to the data section, the labels are actually the union of annotators and not the intersection. </p>\n\n<blockquote>\n  <p>\"Ground truth was determined by the union of the areas marked by all labelers for that image, after removing any black band area from the areas.\" </p>\n</blockquote>",
          "rawMarkdown": "According to the data section, the labels are actually the union of annotators and not the intersection. \n&gt; \"Ground truth was determined by the union of the areas marked by all labelers for that image, after removing any black band area from the areas.\" ",
          "votes": 3
        },
        {
          "id": 669067,
          "postDate": "2019-11-09T12:31:14.610Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 669221,
          "postDate": "2019-11-09T17:53:58.933Z",
          "content": "<p>@Miguel Pinto</p>\n\n<p>you are right. hence it should be \"one is actually not one?\".</p>",
          "rawMarkdown": "@Miguel Pinto\n\nyou are right. hence it should be \"one is actually not one?\"."
        },
        {
          "id": 669299,
          "postDate": "2019-11-09T20:53:54.807Z",
          "content": "<p>Yes. Given any mask that is not shaped as a rectangle, you can split it into the separate annotators' masks and make more training data. Or as you say, you can relabel the 1's to be the correct values.</p>",
          "rawMarkdown": "Yes. Given any mask that is not shaped as a rectangle, you can split it into the separate annotators' masks and make more training data. Or as you say, you can relabel the 1's to be the correct values.",
          "votes": 2
        }
      ]
    },
    {
      "id": 665939,
      "postDate": "2019-11-05T14:55:46.177Z",
      "content": "<p>at first look, the images are randomly reflected or rotated (maybe to prevent the jigsaw puzzle assembling). you can realign all images so that all the \"black slant\" are in the same direction.</p>",
      "rawMarkdown": "at first look, the images are randomly reflected or rotated (maybe to prevent the jigsaw puzzle assembling). you can realign all images so that all the \"black slant\" are in the same direction.",
      "votes": 1,
      "replies": [
        {
          "id": 665943,
          "postDate": "2019-11-05T15:00:16.787Z",
          "content": "<p>any helpful resources that can demonstrate how to design network with higher image sizes to predict small size masks? thanks in advance <a href=\"/hengck23\">@hengck23</a> </p>",
          "rawMarkdown": "any helpful resources that can demonstrate how to design network with higher image sizes to predict small size masks? thanks in advance @hengck23 ",
          "votes": 1
        },
        {
          "id": 665986,
          "postDate": "2019-11-05T16:00:32.727Z",
          "content": "<p>simply:</p>\n\n<p>```\ninput --&gt; encode(scale by half)  --&gt; encode .....  --&gt;decode (scale by two)  ---&gt; decode</p>\n\n<p>if num of decoder  is less than num of encoder, mask size will be small \n```</p>",
          "rawMarkdown": "simply:\n\n```\ninput --&gt; encode(scale by half)  --&gt; encode .....  --&gt;decode (scale by two)  ---&gt; decode\n\nif num of decoder  is less than num of encoder, mask size will be small \n```",
          "votes": 2
        },
        {
          "id": 666161,
          "postDate": "2019-11-05T20:06:40.883Z",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> You can do it with Qubvel's segmentation models. Each backbone has five 2x encoders. So the original image gets reduced by a factor of <code>2^5 = 32</code>. Now regarding decoders, the default is <code>decoder_filters = [256, 128, 64, 32, 16]</code> where <code>len(decoder_filters)</code> is the number of decoder filters and the numbers are how many convolutional maps each has. This says that there are five 2x decoders, so the reduced image becomes a mask similar to original size. If instead you use:</p>\n\n<pre><code>model = Unet('resnet18', decoder_filters=[256, 128, 64])\n</code></pre>\n\n<p>Then you only have three decoder filters. So if original is <code>640x960</code>, it gets encoded into <code>20x30</code>, and it only gets decoded to <code>160x240</code>. Hence the network outputs masks of size <code>160x240</code> and accepts images of size <code>640x960</code>.</p>",
          "rawMarkdown": "@mobassir You can do it with Qubvel's segmentation models. Each backbone has five 2x encoders. So the original image gets reduced by a factor of `2^5 = 32`. Now regarding decoders, the default is `decoder_filters = [256, 128, 64, 32, 16]` where `len(decoder_filters)` is the number of decoder filters and the numbers are how many convolutional maps each has. This says that there are five 2x decoders, so the reduced image becomes a mask similar to original size. If instead you use:\n\n    model = Unet('resnet18', decoder_filters=[256, 128, 64])\n\nThen you only have three decoder filters. So if original is `640x960`, it gets encoded into `20x30`, and it only gets decoded to `160x240`. Hence the network outputs masks of size `160x240` and accepts images of size `640x960`.",
          "votes": 8
        },
        {
          "id": 666358,
          "postDate": "2019-11-06T03:06:32.447Z",
          "content": "<p>thank you a lot for this nice explanation <a href=\"/cdeotte\">@cdeotte</a>  </p>",
          "rawMarkdown": "thank you a lot for this nice explanation @cdeotte  ",
          "votes": 1
        },
        {
          "id": 667957,
          "postDate": "2019-11-07T21:01:04.103Z",
          "content": "<p>Edit: Sorry, I misunderstood.</p>",
          "rawMarkdown": "Edit: Sorry, I misunderstood."
        }
      ]
    },
    {
      "id": 1023955,
      "postDate": "2020-09-23T14:49:01.277Z",
      "content": "<p>can any one will explain about this table please?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4444375%2F394d70364686cd678fbfa129736bd49f%2Fdicevsthreshold.png?generation=1600872408682015&amp;alt=media\" alt=\"\"></p>\n<p>any documents in support of this please?  i am trying very hard to find out but didn't find any documents. please someone help me in this regard.</p>",
      "rawMarkdown": "can any one will explain about this table please?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4444375%2F394d70364686cd678fbfa129736bd49f%2Fdicevsthreshold.png?generation=1600872408682015&alt=media)\n\nany documents in support of this please?  i am trying very hard to find out but didn't find any documents. please someone help me in this regard."
    },
    {
      "id": 683084,
      "postDate": "2019-11-28T04:22:51.257Z",
      "content": "<p>my net </p>\n\n<p>class DinkNet101(nn.Module):\n    def <strong>init</strong>(self, num_classes=4):\n        super(DinkNet101, self).<strong>init</strong>()</p>\n\n<pre><code>    filters = [256, 512, 1024, 2048]\n    resnet = models.resnet101(pretrained=True)\n    self.firstconv = resnet.conv1\n    self.firstbn = resnet.bn1\n    self.firstrelu = resnet.relu\n    self.firstmaxpool = resnet.maxpool\n    self.encoder1 = resnet.layer1\n    self.encoder2 = resnet.layer2\n    self.encoder3 = resnet.layer3\n    self.encoder4 = resnet.layer4\n\n    self.dblock = Dblock_more_dilate(2048)\n\n    self.decoder4 = DecoderBlock(filters[3], filters[2])\n    self.decoder3 = DecoderBlock(filters[2], filters[1])\n    self.decoder2 = DecoderBlock(filters[1], filters[0])\n    self.decoder1 = DecoderBlock(filters[0], filters[0])\n\n    self.finaldeconv1 = nn.ConvTranspose2d(filters[0], 32, 4, 2, 1)\n    self.finalrelu1 = nonlinearity\n    self.finalconv2 = nn.Conv2d(32, 32, 3, padding=1)\n    self.finalrelu2 = nonlinearity\n    self.finalconv3 = nn.Conv2d(32, num_classes, 3, padding=1)\n\ndef forward(self, x):\n    # Encoder\n    x = self.firstconv(x)\n    x = self.firstbn(x)\n    x = self.firstrelu(x)\n    x = self.firstmaxpool(x)\n    e1 = self.encoder1(x)\n    e2 = self.encoder2(e1)\n    e3 = self.encoder3(e2)\n    e4 = self.encoder4(e3)\n\n    # Center\n    e4 = self.dblock(e4)\n\n    # Decoder\n    d4 = self.decoder4(e4) + e3\n    d3 = self.decoder3(d4) + e2\n    d2 = self.decoder2(d3) + e1\n    d1 = self.decoder1(d2)\n    out = self.finaldeconv1(d1)\n    out = self.finalrelu1(out)\n    out = self.finalconv2(out)\n    out = self.finalrelu2(out)\n    out = self.finalconv3(out)\n\n    return F.sigmoid(out)\n</code></pre>",
      "rawMarkdown": "my net \n\nclass DinkNet101(nn.Module):\n    def __init__(self, num_classes=4):\n        super(DinkNet101, self).__init__()\n\n        filters = [256, 512, 1024, 2048]\n        resnet = models.resnet101(pretrained=True)\n        self.firstconv = resnet.conv1\n        self.firstbn = resnet.bn1\n        self.firstrelu = resnet.relu\n        self.firstmaxpool = resnet.maxpool\n        self.encoder1 = resnet.layer1\n        self.encoder2 = resnet.layer2\n        self.encoder3 = resnet.layer3\n        self.encoder4 = resnet.layer4\n        \n        self.dblock = Dblock_more_dilate(2048)\n\n        self.decoder4 = DecoderBlock(filters[3], filters[2])\n        self.decoder3 = DecoderBlock(filters[2], filters[1])\n        self.decoder2 = DecoderBlock(filters[1], filters[0])\n        self.decoder1 = DecoderBlock(filters[0], filters[0])\n\n        self.finaldeconv1 = nn.ConvTranspose2d(filters[0], 32, 4, 2, 1)\n        self.finalrelu1 = nonlinearity\n        self.finalconv2 = nn.Conv2d(32, 32, 3, padding=1)\n        self.finalrelu2 = nonlinearity\n        self.finalconv3 = nn.Conv2d(32, num_classes, 3, padding=1)\n\n    def forward(self, x):\n        # Encoder\n        x = self.firstconv(x)\n        x = self.firstbn(x)\n        x = self.firstrelu(x)\n        x = self.firstmaxpool(x)\n        e1 = self.encoder1(x)\n        e2 = self.encoder2(e1)\n        e3 = self.encoder3(e2)\n        e4 = self.encoder4(e3)\n        \n        # Center\n        e4 = self.dblock(e4)\n\n        # Decoder\n        d4 = self.decoder4(e4) + e3\n        d3 = self.decoder3(d4) + e2\n        d2 = self.decoder2(d3) + e1\n        d1 = self.decoder1(d2)\n        out = self.finaldeconv1(d1)\n        out = self.finalrelu1(out)\n        out = self.finalconv2(out)\n        out = self.finalrelu2(out)\n        out = self.finalconv3(out)\n\n        return F.sigmoid(out)"
    },
    {
      "id": 675090,
      "postDate": "2019-11-17T15:35:59.207Z",
      "rawMarkdown": "\n",
      "replies": [
        {
          "id": 675096,
          "postDate": "2019-11-17T15:42:19.850Z",
          "content": "<p>don't think it is a bug.\n\"segmentation only\" uses threshold_label = -1</p>",
          "rawMarkdown": "don't think it is a bug.\n\"segmentation only\" uses threshold\\_label = -1\n"
        },
        {
          "id": 675107,
          "postDate": "2019-11-17T15:55:32.600Z",
          "content": "<p>i just check and i do have the same results as you if i use size threshold (i never use it before)\nlet me check if it is a bug\n<code>\nthreshold_size  = [ 30000, 30000, 30000, 30000]\n</code></p>",
          "rawMarkdown": "i just check and i do have the same results as you if i use size threshold (i never use it before)\nlet me check if it is a bug\n```\nthreshold_size  = [ 30000, 30000, 30000, 30000]\n```"
        },
        {
          "id": 675133,
          "postDate": "2019-11-17T16:42:47.787Z",
          "content": "<p>i check that the code should have no bug (but i am not 100% sure).\ni made a submission. using large area threshold like 30000 don't give good results</p>",
          "rawMarkdown": "i check that the code should have no bug (but i am not 100% sure).\ni made a submission. using large area threshold like 30000 don't give good results"
        }
      ]
    },
    {
      "id": 674993,
      "postDate": "2019-11-17T12:06:27.133Z",
      "content": "<p>I trained a model(different from yours) using your codes/pipeline and got a validation score of <code>0.659xx</code> but when I submit it, I get around <code>0.58xx</code>on LB. I have attached the screenshot...don't know what could be wrong?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F9cabc472aaa647c0c3d6e94e1f58ac4c%2Fval_score.bmp?generation=1573992298447004&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I trained a model(different from yours) using your codes/pipeline and got a validation score of `0.659xx` but when I submit it, I get around `0.58xx`on LB. I have attached the screenshot...don't know what could be wrong?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F9cabc472aaa647c0c3d6e94e1f58ac4c%2Fval_score.bmp?generation=1573992298447004&amp;alt=media)\n",
      "replies": [
        {
          "id": 674995,
          "postDate": "2019-11-17T12:09:49.807Z",
          "content": "<p>This is the training log...I used 0.658 model</p>",
          "rawMarkdown": "This is the training log...I used 0.658 model"
        },
        {
          "id": 675003,
          "postDate": "2019-11-17T12:39:38.003Z",
          "content": "<p>i suspect threshold problem.</p>\n\n<p>when you submit in 'test' mode, you should see the results of</p>\n\n<p><code>\n        print('initial_checkpoint=%s'%initial_checkpoint)\n        text = summarise_submission_csv(df)\n        log.write('\\n')\n        log.write('%s'%(text))\n</code></p>\n\n<p>you check the number of predicted masks, etc similar to what i have below</p>\n\n<p>```</p>\n\n<p>compare with LB probing ... \n        num_image =  3698(3698) </p>\n\n<pre><code>    pos0 =  1303(1864)  xxx\n    pos1 =  1410(1508)  xxx\n    pos2 =  1084(1982)  xxx\n    pos3 =  2305(2382)  xxx\n\n    neg0 =  xxx(1834)  1.306\n    neg1 =  xxx(1940)  1.179\n    neg2 =  xxx(2638)  0.991\n</code></pre>\n\n<h2>        neg3 =  xxx(2017)  0.691</h2>\n\n<pre><code>    all_zero =   xxx (?)\n</code></pre>\n\n<p>```\nadjust threshold until you get good number of predicted mask in the LB test data</p>\n\n<p>you can also see:\n<a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/117310#latest-674964\">https://www.kaggle.com/c/understanding_cloud_organization/discussion/117310#latest-674964</a></p>",
          "rawMarkdown": "i suspect threshold problem.\n\nwhen you submit in 'test' mode, you should see the results of\n\n```\n        print('initial_checkpoint=%s'%initial_checkpoint)\n        text = summarise_submission_csv(df)\n        log.write('\\n')\n        log.write('%s'%(text))\n```\n\nyou check the number of predicted masks, etc similar to what i have below\n\n```\n\ncompare with LB probing ... \n        num_image =  3698(3698) \n\n        pos0 =  1303(1864)  xxx\n        pos1 =  1410(1508)  xxx\n        pos2 =  1084(1982)  xxx\n        pos3 =  2305(2382)  xxx\n\n        neg0 =  xxx(1834)  1.306\n        neg1 =  xxx(1940)  1.179\n        neg2 =  xxx(2638)  0.991\n        neg3 =  xxx(2017)  0.691\n--------------------------------------------------\n\n        all_zero =   xxx (?)\n```\nadjust threshold until you get good number of predicted mask in the LB test data\n\n\nyou can also see:\nhttps://www.kaggle.com/c/understanding_cloud_organization/discussion/117310#latest-674964",
          "votes": 2
        }
      ]
    },
    {
      "id": 674509,
      "postDate": "2019-11-16T15:53:35.030Z",
      "content": "<p>'compute_metric_label' is not defined ???</p>",
      "rawMarkdown": "'compute_metric_label' is not defined ???\n",
      "replies": [
        {
          "id": 674511,
          "postDate": "2019-11-16T15:55:29.467Z",
          "content": "<p>kaggle.py?</p>",
          "rawMarkdown": "kaggle.py?",
          "votes": 2
        },
        {
          "id": 674515,
          "postDate": "2019-11-16T16:04:29.963Z",
          "content": "<p>no\nsubmmit.py</p>",
          "rawMarkdown": "no\nsubmmit.py"
        },
        {
          "id": 674574,
          "postDate": "2019-11-16T18:14:44.900Z",
          "content": "<p>The function should be found in kaggle.py. please the latest version of the code</p>",
          "rawMarkdown": "The function should be found in kaggle.py. please the latest version of the code",
          "votes": 2
        }
      ]
    },
    {
      "id": 674345,
      "postDate": "2019-11-16T10:02:52.960Z",
      "content": "<p>hi heng, do you feel strange about andrew's threshold_mask and your threshold_mask ? the threshold_mask &gt; 0.5 always have good result with andrew's model(i trained with FPN resnet34), but your model always need to set the threshold_mask &lt; 0.5. and you can see andrew's experimental result at <a href=\"https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools\">here</a>.</p>",
      "rawMarkdown": "hi heng, do you feel strange about andrew's threshold_mask and your threshold_mask ? the threshold_mask &gt; 0.5 always have good result with andrew's model(i trained with FPN resnet34), but your model always need to set the threshold_mask &lt; 0.5. and you can see andrew's experimental result at [here](https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools).",
      "replies": [
        {
          "id": 674350,
          "postDate": "2019-11-16T10:07:28.900Z",
          "content": "<p>and you all use sigmod with model's output.</p>",
          "rawMarkdown": "and you all use sigmod with model's output."
        },
        {
          "id": 674353,
          "postDate": "2019-11-16T10:16:24.223Z",
          "content": "<p>i also use  threshold &gt;0.5 for image label. \nthreshold can be less than 0.5 for pixel label. this is to encourage dilation.</p>\n\n<p>did you train with loss_mask only or ( loss_label + loss_mask )?</p>",
          "rawMarkdown": "i also use  threshold &gt;0.5 for image label. \nthreshold can be less than 0.5 for pixel label. this is to encourage dilation.\n\ndid you train with loss\\_mask only or ( loss\\_label + loss\\_mask )?",
          "votes": 1
        },
        {
          "id": 674415,
          "postDate": "2019-11-16T13:18:12.767Z",
          "content": "<p>“i also use threshold &gt;0.5 for image label.\nthreshold can be less than 0.5 for pixel label. this is to encourage dilation.”</p>\n\n<p>log of andrew's model(i trained with FPN resnet34) is lost, i will try it again. and if we do not consider the difference between Unet and FPN,  i think your submit mode of segmentation only is equal to Andrew's submit mode.  the threshold mask &gt; 0.5 with andrew's model will be better (see picture), and the threshold mask &gt; 0.5 with your model will be worse (see log).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F08a9a48a429996b20cd7872e839ebd7d%2Fs.png?generation=1573910110677859&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "“i also use threshold &gt;0.5 for image label.\nthreshold can be less than 0.5 for pixel label. this is to encourage dilation.”\n\nlog of andrew's model(i trained with FPN resnet34) is lost, i will try it again. and if we do not consider the difference between Unet and FPN,  i think your submit mode of segmentation only is equal to Andrew's submit mode.  the threshold mask &gt; 0.5 with andrew's model will be better (see picture), and the threshold mask &gt; 0.5 with your model will be worse (see log).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F08a9a48a429996b20cd7872e839ebd7d%2Fs.png?generation=1573910110677859&amp;alt=media)\n"
        },
        {
          "id": 674417,
          "postDate": "2019-11-16T13:19:48.740Z",
          "content": "<p>\"did you train with loss_mask only or ( loss_label + loss_mask )?\"</p>\n\n<p>andrew' model is loss_mask only,  your model is ((loss_mask )/iter_accum).backward().</p>",
          "rawMarkdown": "\"did you train with loss_mask only or ( loss_label + loss_mask )?\"\n\nandrew' model is loss_mask only,  your model is ((loss_mask )/iter_accum).backward()."
        },
        {
          "id": 674436,
          "postDate": "2019-11-16T14:04:09.263Z",
          "content": "<p>do note that my model do not use size post processing. It uses max pixel probability instead.</p>\n\n<p>\"i think your submit mode of segmentation only is equal to Andrew's submit mode\"\ni think that too. that is why i mentioned before that the notebook kernel are actually good enough for LB 0.67 if they are properly trained, post processed and ensembled.</p>",
          "rawMarkdown": "do note that my model do not use size post processing. It uses max pixel probability instead.\n\n\"i think your submit mode of segmentation only is equal to Andrew's submit mode\"\ni think that too. that is why i mentioned before that the notebook kernel are actually good enough for LB 0.67 if they are properly trained, post processed and ensembled.",
          "votes": 2
        },
        {
          "id": 674444,
          "postDate": "2019-11-16T14:18:54.780Z",
          "content": "<p>thank u, heng, wait for me to traine andrew's model(fnp resnet) again, maybe after finished competion, because i am working on your code now.</p>",
          "rawMarkdown": "thank u, heng, wait for me to traine andrew's model(fnp resnet) again, maybe after finished competion, because i am working on your code now."
        },
        {
          "id": 674449,
          "postDate": "2019-11-16T14:26:23.973Z",
          "content": "<p>do also note that his fpn is not exactly the same as mine. the decoder head is differ slightly in number of convs, etc. but i think that does affect results. but do note that my fpn is tuned for size 1050x700 and not other size.</p>\n\n<p>i have tried many models, different sizes, etc in the past week. i carried about 80 experiments. most of my models will have 0.65 in validation, +/-0.05. </p>",
          "rawMarkdown": "do also note that his fpn is not exactly the same as mine. the decoder head is differ slightly in number of convs, etc. but i think that does affect results. but do note that my fpn is tuned for size 1050x700 and not other size.\n\ni have tried many models, different sizes, etc in the past week. i carried about 80 experiments. most of my models will have 0.65 in validation, +/-0.05. ",
          "votes": 1
        },
        {
          "id": 674468,
          "postDate": "2019-11-16T14:44:54.140Z",
          "content": "<p>thank u.</p>",
          "rawMarkdown": "thank u."
        },
        {
          "id": 674710,
          "postDate": "2019-11-16T23:54:24.060Z",
          "content": "<p>after i check andrew's code and results in detail, indeed our results are close but using very different parameters.</p>\n\n<p>you may want to check the correlation of the results of the two models. if there is diversity, there is a good chance for ensemble.</p>\n\n<p>but it is tricky for ensemble since both model use different parameter.</p>\n\n<p>```\neg \nmodel1 needs param1\nmodel2 needs param2</p>\n\n<p>but parm1+2 may/may not be good for model1+2?\nor param1 is better?\nor param2 is better?\n```</p>",
          "rawMarkdown": "after i check andrew's code and results in detail, indeed our results are close but using very different parameters.\n\nyou may want to check the correlation of the results of the two models. if there is diversity, there is a good chance for ensemble.\n\nbut it is tricky for ensemble since both model use different parameter.\n\n```\neg \nmodel1 needs param1\nmodel2 needs param2\n\nbut parm1+2 may/may not be good for model1+2?\nor param1 is better?\nor param2 is better?\n```"
        },
        {
          "id": 675093,
          "postDate": "2019-11-17T15:39:27.223Z",
          "content": "<p>thank you for your reply, i will do experiment about different between the two models after competion.</p>",
          "rawMarkdown": "thank you for your reply, i will do experiment about different between the two models after competion."
        }
      ]
    },
    {
      "id": 674098,
      "postDate": "2019-11-15T22:15:53.193Z",
      "content": "<p><a href=\"/hengck23\">@hengck23</a>  this is great stuff for learning for sure! if you can follow this up with like additional details around </p>\n\n<blockquote>\n  <p>\"training log is important to analyse model behavior. it is probably the most important reference for beginner to learn how to train a model. the provision of such log is to compare your version with the reference version to see how change of hyper-parameters affect results.\"</p>\n</blockquote>\n\n<p>with examples for intuition building / reading you are platinum! </p>",
      "rawMarkdown": "@hengck23  this is great stuff for learning for sure! if you can follow this up with like additional details around \n\n&gt; \"training log is important to analyse model behavior. it is probably the most important reference for beginner to learn how to train a model. the provision of such log is to compare your version with the reference version to see how change of hyper-parameters affect results.\"\n\nwith examples for intuition building / reading you are platinum! "
    },
    {
      "id": 673776,
      "postDate": "2019-11-15T13:26:22.197Z",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> are you using for the ensembling your sharpening method?</p>",
      "rawMarkdown": "@hengck23 are you using for the ensembling your sharpening method?",
      "replies": [
        {
          "id": 673785,
          "postDate": "2019-11-15T13:32:12.887Z",
          "content": "<p>i do not use sharpening method now. i haven't test sharpening yet.</p>\n\n<p>sharpening is useful for classification and when threshold is near 0.50 without sharpening.</p>\n\n<p>now, i use label probability = max pixel probability. this is already a high very and maybe sharpening is difficult to work.</p>\n\n<p>i am thinking of using joint label loss and pixel loss instead</p>",
          "rawMarkdown": "i do not use sharpening method now. i haven't test sharpening yet.\n\nsharpening is useful for classification and when threshold is near 0.50 without sharpening.\n\nnow, i use label probability = max pixel probability. this is already a high very and maybe sharpening is difficult to work.\n\ni am thinking of using joint label loss and pixel loss instead",
          "votes": 2
        },
        {
          "id": 674140,
          "postDate": "2019-11-15T23:41:22.063Z",
          "content": "<p>I try to draw sharpen function,\nit work like a equalizer which will seperate small value but push large value togater.\nIt will help when your predictions are hard to seperate with simple threshold selection.\nProbably won't help in this competition.</p>\n\n<p>.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1008248%2F2a34ff421e5adeeea8c1a6ba9fb06fc1%2Fsharpenpng?generation=1573860262086355&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "I try to draw sharpen function,\nit work like a equalizer which will seperate small value but push large value togater.\nIt will help when your predictions are hard to seperate with simple threshold selection.\nProbably won't help in this competition.\n\n.![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1008248%2F2a34ff421e5adeeea8c1a6ba9fb06fc1%2Fsharpenpng?generation=1573860262086355&amp;alt=media)\n"
        },
        {
          "id": 674197,
          "postDate": "2019-11-16T01:52:34.757Z",
          "content": "<p>you should think of it as:</p>\n\n<ol>\n<li>assume i have 2 classifier predictions p1 and p2</li>\n<li>if i use normal ensemble average: (p1+p2)/2 &lt;0.5.\nnow draw the region in the plot of p1,p2 that satisfy inequality above. it should be triangle</li>\n<li>instead, if we use (p1^0.5 + p2^0.5 )/2 &lt;0.5. the inequality will change slightly, favoring larger or smaller values, depend on the power constant. the region  will have a slight curve boundary</li>\n<li>in the extreme case,  (p1^999 + p2^999 )/2 &lt;0.5 is like majority voting</li>\n</ol>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fb70ac1c15fc77dc78c7553f68a6b08d6%2FSelection_057.png?generation=1573889846574797&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "you should think of it as:\n\n1. assume i have 2 classifier predictions p1 and p2\n2. if i use normal ensemble average: (p1+p2)/2 &lt;0.5.\nnow draw the region in the plot of p1,p2 that satisfy inequality above. it should be triangle\n3. instead, if we use (p1^0.5 + p2^0.5 )/2 &lt;0.5. the inequality will change slightly, favoring larger or smaller values, depend on the power constant. the region  will have a slight curve boundary\n4. in the extreme case,  (p1^999 + p2^999 )/2 &lt;0.5 is like majority voting\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fb70ac1c15fc77dc78c7553f68a6b08d6%2FSelection_057.png?generation=1573889846574797&amp;alt=media)\n",
          "votes": 5
        },
        {
          "id": 674498,
          "postDate": "2019-11-16T15:31:00.060Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>  Thanks your explanation.\nDoes the optimal threshold also shift? as we transform the probability into new range.\nSo it shouldn't be 0.5 anymore,\nI remember that I need to find a new threshold in the steel competition.\nThanks</p>",
          "rawMarkdown": "@hengck23  Thanks your explanation.\nDoes the optimal threshold also shift? as we transform the probability into new range.\nSo it shouldn't be 0.5 anymore,\nI remember that I need to find a new threshold in the steel competition.\nThanks"
        }
      ]
    },
    {
      "id": 673752,
      "postDate": "2019-11-15T12:45:48.167Z",
      "content": "<p>hi heng, I ran your code, [dn0,1,2,3] always zero, normal?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F3ff73219a66bd3c20c7549fa44fa42d1%2FQQ20191115204236.png?generation=1573821978090365&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "hi heng, I ran your code, [dn0,1,2,3] always zero, normal?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F3ff73219a66bd3c20c7549fa44fa42d1%2FQQ20191115204236.png?generation=1573821978090365&amp;alt=media)\n\n",
      "replies": [
        {
          "id": 673768,
          "postDate": "2019-11-15T13:21:07.763Z",
          "content": "<p><a href=\"/caoqineng\">@caoqineng</a> </p>\n\n<p>it is normal.</p>\n\n<p>it means that for false positive at image level (i.e. we are unable to to reject non-mask image), the segmentation is unable to correct it too at pixel level</p>",
          "rawMarkdown": "@caoqineng \n\nit is normal.\n\nit means that for false positive at image level (i.e. we are unable to to reject non-mask image), the segmentation is unable to correct it too at pixel level",
          "votes": 1
        },
        {
          "id": 673792,
          "postDate": "2019-11-15T13:47:15.400Z",
          "content": "<p>Thank u. but i have another question, Why is the experimental result different from yours? I simply modified your code about split data part in class CloudDataset. The following is the result of my plotting position and plotting result. I think the part I modified should be correct. Can you give me some advice?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F650167b8221d766c5a141f7e968b5881%2F1.png?generation=1573825493877471&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F051e3074bea76b2ac9cd887c227424d4%2F2.png?generation=1573825593039824&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thank u. but i have another question, Why is the experimental result different from yours? I simply modified your code about split data part in class CloudDataset. The following is the result of my plotting position and plotting result. I think the part I modified should be correct. Can you give me some advice?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F650167b8221d766c5a141f7e968b5881%2F1.png?generation=1573825493877471&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F051e3074bea76b2ac9cd887c227424d4%2F2.png?generation=1573825593039824&amp;alt=media)\n\n"
        },
        {
          "id": 673795,
          "postDate": "2019-11-15T13:53:31.550Z",
          "content": "<p>can you attach the whole py file here?</p>",
          "rawMarkdown": "can you attach the whole py file here?",
          "votes": 1
        },
        {
          "id": 673799,
          "postDate": "2019-11-15T14:00:01.357Z",
          "content": "<p>i think i know. i read mask png file and image png file of difference size</p>",
          "rawMarkdown": "i think i know. i read mask png file and image png file of difference size",
          "votes": 1
        },
        {
          "id": 673806,
          "postDate": "2019-11-15T14:07:57.720Z",
          "content": "<p>there is a test code  to output result results from the dataloader and dataset, please see\n```\ndataset.py</p>\n\n<pre><code>#run_check_dataset()\n#run_check_dataloader()\n#run_check_augment()\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "there is a test code  to output result results from the dataloader and dataset, please see\n```\ndataset.py\n\n    #run_check_dataset()\n    #run_check_dataloader()\n    #run_check_augment()\n\n\n```",
          "votes": 1
        },
        {
          "id": 673835,
          "postDate": "2019-11-15T14:40:20.730Z",
          "content": "<p>yes, i sorted out the file. at first, i ran three methods run_make_train_split(), run_dump_mask_to_png(),run_dump_image_to_png() in kaggle.py. then i changed the code in dataset.py and train_a2.py. the modification in dataset.py and train_a2.py is marked likes below image.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F5cfb8b28687439a83bb6eb36a5423ef4%2F3.png?generation=1573828568794753&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "yes, i sorted out the file. at first, i ran three methods run_make_train_split(), run_dump_mask_to_png(),run_dump_image_to_png() in kaggle.py. then i changed the code in dataset.py and train_a2.py. the modification in dataset.py and train_a2.py is marked likes below image.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F5cfb8b28687439a83bb6eb36a5423ef4%2F3.png?generation=1573828568794753&amp;alt=media)\n\n"
        },
        {
          "id": 673847,
          "postDate": "2019-11-15T14:53:53.183Z",
          "content": "<p>but you will resize the probability_mask likes truth_mask before criterion().\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F0d18fe1232990726641c6d1c9236f1f9%2FQQ20191115225205.png?generation=1573829630886915&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "but you will resize the probability_mask likes truth_mask before criterion().\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F0d18fe1232990726641c6d1c9236f1f9%2FQQ20191115225205.png?generation=1573829630886915&amp;alt=media)\n\n"
        },
        {
          "id": 673853,
          "postDate": "2019-11-15T15:01:46.120Z",
          "content": "<p>i check your files. they seems correct. \nbut my loss is better (actually you should see good results within 12 K training iteration and don't have to train so long)</p>\n\n<p>how about using more training data? your validation set is 1000 images, try reducing it to 300?</p>\n\n<p>attached, my logfile</p>",
          "rawMarkdown": "i check your files. they seems correct. \nbut my loss is better (actually you should see good results within 12 K training iteration and don't have to train so long)\n\nhow about using more training data? your validation set is 1000 images, try reducing it to 300?\n\n\nattached, my logfile"
        },
        {
          "id": 673856,
          "postDate": "2019-11-15T15:06:53.413Z",
          "content": "<p>ran run_check_augment(), looks a bit strange.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F4343629c7694c7ae85cd583b451a2dd6%2Fbefore.png?generation=1573830369233901&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F7e4dc626e34e70084d29e0b576843ba7%2Fafter.png?generation=1573830381899427&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "ran run_check_augment(), looks a bit strange.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F4343629c7694c7ae85cd583b451a2dd6%2Fbefore.png?generation=1573830369233901&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F7e4dc626e34e70084d29e0b576843ba7%2Fafter.png?generation=1573830381899427&amp;alt=media)\n"
        },
        {
          "id": 673857,
          "postDate": "2019-11-15T15:06:54.590Z",
          "content": "<p>\"but you will resize the probabilitymask likes truthmask before criterion().\"</p>\n\n<p>yes, that is correct. it is more flexible to code in this way. so the model can output any size it wants</p>",
          "rawMarkdown": "\"but you will resize the probabilitymask likes truthmask before criterion().\"\n\nyes, that is correct. it is more flexible to code in this way. so the model can output any size it wants",
          "votes": 1
        },
        {
          "id": 673858,
          "postDate": "2019-11-15T15:10:18.140Z",
          "content": "<p>\"ran runcheckaugment(), looks a bit strange.\"</p>\n\n<p>edit here\n```\ndef run_check_augment():\n    # 'image': '1050x700', 'mask': '525x350'\n    def augment(image, label, mask, infor):\n        if 0:\n            #if np.random.rand()&lt;0.5:  image, mask = do_flip_ud(image, mask)\n            if np.random.rand()&lt;0.5:  image, mask = do_flip_lr(image, mask)</p>\n\n<p>````</p>\n\n<p>what you see is random grid shuffle augmentation which i don't use.\n(<a href=\"https://albumentations.readthedocs.io/en/latest/api/augmentations.html#albumentations.augmentations.transforms.RandomGridShuffle\">https://albumentations.readthedocs.io/en/latest/api/augmentations.html#albumentations.augmentations.transforms.RandomGridShuffle</a>)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F4ac05dcff647be6c9847912e0f957503%2FSelection_055.png?generation=1573830900187980&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "\"ran runcheckaugment(), looks a bit strange.\"\n\nedit here\n```\ndef run_check_augment():\n    # 'image': '1050x700', 'mask': '525x350'\n    def augment(image, label, mask, infor):\n        if 0:\n            #if np.random.rand()&lt;0.5:  image, mask = do_flip_ud(image, mask)\n            if np.random.rand()&lt;0.5:  image, mask = do_flip_lr(image, mask)\n\n\n````\n\nwhat you see is random grid shuffle augmentation which i don't use.\n(https://albumentations.readthedocs.io/en/latest/api/augmentations.html#albumentations.augmentations.transforms.RandomGridShuffle)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F4ac05dcff647be6c9847912e0f957503%2FSelection_055.png?generation=1573830900187980&amp;alt=media)\n"
        },
        {
          "id": 673867,
          "postDate": "2019-11-15T15:21:59.620Z",
          "content": "<p>Thank you for checking my file,  otherwise i have to check it  all night，i thought dice&gt;0.650 is the correct result before your check.</p>",
          "rawMarkdown": "Thank you for checking my file,  otherwise i have to check it  all night，i thought dice&gt;0.650 is the correct result before your check."
        },
        {
          "id": 673869,
          "postDate": "2019-11-15T15:22:22.703Z",
          "content": "<p>\"what you see is random grid shuffle augmentation which i don't use.\n(<a href=\"https://albumentations.readthedocs.io/en/latest/api/augmentations.html#albumentations.augmentations.transforms.RandomGridShuffle\">https://albumentations.readthedocs.io/en/latest/api/augmentations.html#albumentations.augmentations.transforms.RandomGridShuffle</a>)\"</p>\n\n<p>thank u.</p>",
          "rawMarkdown": "\"what you see is random grid shuffle augmentation which i don't use.\n(https://albumentations.readthedocs.io/en/latest/api/augmentations.html#albumentations.augmentations.transforms.RandomGridShuffle)\"\n\nthank u."
        },
        {
          "id": 673870,
          "postDate": "2019-11-15T15:24:49.950Z",
          "content": "<p>\"i thought dice&gt;0.650 is the correct result before your check.\"</p>\n\n<p>at the submit.py, using TTA and adjusting threshold you should better results than when you see at the train logfile. i suggest your try to try submit.py with the same validation set. (set mode to \"valid\")</p>",
          "rawMarkdown": "\"i thought dice&gt;0.650 is the correct result before your check.\"\n\nat the submit.py, using TTA and adjusting threshold you should better results than when you see at the train logfile. i suggest your try to try submit.py with the same validation set. (set mode to \"valid\")",
          "votes": 1
        },
        {
          "id": 673875,
          "postDate": "2019-11-15T15:32:49.503Z",
          "content": "<p>really appreciate your help. i will try it right away.</p>",
          "rawMarkdown": "really appreciate your help. i will try it right away."
        }
      ]
    },
    {
      "id": 673230,
      "postDate": "2019-11-14T17:35:29.107Z",
      "content": "<p><a href=\"/hengck23\">@hengck23</a>  thanks your sharing, my experiment also show that JPU + ASPP converge faster than other model\nMay I ask that what norm layer you use in JPU and ASPP? because with small batch size (4~6), BN won't behave well isn't it?\nThanks</p>",
      "rawMarkdown": "@hengck23  thanks your sharing, my experiment also show that JPU + ASPP converge faster than other model\nMay I ask that what norm layer you use in JPU and ASPP? because with small batch size (4~6), BN won't behave well isn't it?\nThanks",
      "replies": [
        {
          "id": 673233,
          "postDate": "2019-11-14T17:38:48.837Z",
          "content": "<p>\"JPU + ASPP converge faster than other model\" </p>\n\n<p>you are very right. they results are also much better.  (interesting, they require lower threshold because of better BCE loss)</p>\n\n<p>\"because with small batch size (4~6), BN won't behave well isn't it?\"\nnot true. it really depends on dataset. for this cloud dataset i can train with batch =6.</p>\n\n<p>please see the attached file (model.py) i post before for complete description of my model</p>",
          "rawMarkdown": "\"JPU + ASPP converge faster than other model\" \n\nyou are very right. they results are also much better.  (interesting, they require lower threshold because of better BCE loss)\n\n\n\"because with small batch size (4~6), BN won't behave well isn't it?\"\nnot true. it really depends on dataset. for this cloud dataset i can train with batch =6.\n\nplease see the attached file (model.py) i post before for complete description of my model",
          "votes": 2
        },
        {
          "id": 673372,
          "postDate": "2019-11-14T22:38:18.980Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>  Thanks for your reply,\nYes, I notice that you only use Group norm in decoder part, so I conduct some experiement and have some conclusion, that if I swap decoder batch norm to group norm (ex: Unet), then small batch (4) with accumulate gradient works (suprisely), otherwise, cannot make it work.\nDon't know that this also apply in JPU and ASPP as well, im going to try this later\nThanks</p>",
          "rawMarkdown": "@hengck23  Thanks for your reply,\nYes, I notice that you only use Group norm in decoder part, so I conduct some experiement and have some conclusion, that if I swap decoder batch norm to group norm (ex: Unet), then small batch (4) with accumulate gradient works (suprisely), otherwise, cannot make it work.\nDon't know that this also apply in JPU and ASPP as well, im going to try this later\nThanks"
        },
        {
          "id": 673379,
          "postDate": "2019-11-14T22:48:42.960Z",
          "content": "<p>I just follow the paper and the code in their github repo. Group norm is used in fpn.</p>\n\n<p>For jpu and aspp i used batchnorm. The cloud images are quite consistent in terms of intensity values</p>\n\n<p>Things sometimes work and don’t work. So i do encourage you to explore different methods </p>",
          "rawMarkdown": "I just follow the paper and the code in their github repo. Group norm is used in fpn.\n\nFor jpu and aspp i used batchnorm. The cloud images are quite consistent in terms of intensity values\n\nThings sometimes work and don’t work. So i do encourage you to explore different methods ",
          "votes": 1
        },
        {
          "id": 673388,
          "postDate": "2019-11-14T23:07:32.947Z",
          "content": "<p>A side note for aspp atrous spatial kernel. To know what it is going on, you can output the response map for each rate to see what they are detecting.   </p>\n\n<p>I also tried jpu and aspp separately. It seems that the combinations is required. There are some other jpu architecture in the paper like pspnet + jpu that i am trying now</p>",
          "rawMarkdown": "A side note for aspp atrous spatial kernel. To know what it is going on, you can output the response map for each rate to see what they are detecting.   \n\nI also tried jpu and aspp separately. It seems that the combinations is required. There are some other jpu architecture in the paper like pspnet + jpu that i am trying now",
          "votes": 1
        },
        {
          "id": 673831,
          "postDate": "2019-11-15T14:36:43.010Z",
          "content": "<p>completed \"jpu + pspnet\" . I think it is good. it is the only network that i tried and it don't overfit easily</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa8e78d09c8f068eda49203b1eda36cb4%2FSelection_054.png?generation=1573828595753343&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "completed \"jpu + pspnet\" . I think it is good. it is the only network that i tried and it don't overfit easily\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa8e78d09c8f068eda49203b1eda36cb4%2FSelection_054.png?generation=1573828595753343&amp;alt=media)\n",
          "votes": 1
        },
        {
          "id": 674502,
          "postDate": "2019-11-16T15:34:39.037Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>  Thanks for your reply,\nI also agree that the threshold is lower than the default setting (0.7, 0.3)\nHere is the 5 fold with resnet 50 JPU ASPP, threshold found by hyperopt:\n-------- 0 -------------\n[[11/15/2019 11:55:35 PM]] max_threshold: 0.6499999999999999\n[[11/15/2019 11:55:35 PM]] threshold: 0.3\n[[11/15/2019 11:55:35 PM]] best dice: 0.6662194166950253</p>\n\n<p>[[11/16/2019 03:38:42 AM]] max_threshold: 0.5499999999999999\n[[11/16/2019 03:38:42 AM]] threshold: 0.25\n[[11/16/2019 03:38:42 AM]] best dice: 0.6269578569840203</p>\n\n<p>[[11/16/2019 09:26:30 AM]] max_threshold: 0.7\n[[11/16/2019 09:26:30 AM]] threshold: 0.35\n[[11/16/2019 09:26:30 AM]] best dice: 0.6324725536992442</p>\n\n<p>[[11/16/2019 11:38:55 AM]] max_threshold: 0.7499999999999998\n[[11/16/2019 11:38:55 AM]] threshold: 0.3\n[[11/16/2019 11:38:55 AM]] best dice: 0.6536089501721206</p>\n\n<p>[[11/16/2019 01:01:42 PM]] max_threshold: 0.6499999999999999\n[[11/16/2019 01:01:42 PM]] threshold: 0.25\n[[11/16/2019 01:01:42 PM]] best dice: 0.6198696383341208</p>\n\n<p>-------- 1 -------------\n[[11/15/2019 11:56:34 PM]] max_threshold: 0.7499999999999998\n[[11/15/2019 11:56:34 PM]] threshold: 0.35\n[[11/15/2019 11:56:34 PM]] best dice: 0.767333773896778</p>\n\n<p>[[11/16/2019 03:39:44 AM]] max_threshold: 0.7\n[[11/16/2019 03:39:44 AM]] threshold: 0.35\n[[11/16/2019 03:39:44 AM]] best dice: 0.7741650036825072</p>\n\n<p>[[11/16/2019 09:27:04 AM]] max_threshold: 0.6499999999999999\n[[11/16/2019 09:27:04 AM]] threshold: 0.35\n[[11/16/2019 09:27:04 AM]] best dice: 0.7717044965056499</p>\n\n<p>[[11/16/2019 11:39:28 AM]] max_threshold: 0.7\n[[11/16/2019 11:39:28 AM]] threshold: 0.25\n[[11/16/2019 11:39:28 AM]] best dice: 0.7702028346574968</p>\n\n<p>[[11/16/2019 01:02:16 PM]] max_threshold: 0.7\n[[11/16/2019 01:02:16 PM]] threshold: 0.44999999999999996\n[[11/16/2019 01:02:16 PM]] best dice: 0.7569014497993047</p>\n\n<p>-------- 2 -------------</p>\n\n<p>[[11/15/2019 11:57:37 PM]] max_threshold: 0.5999999999999999\n[[11/15/2019 11:57:37 PM]] threshold: 0.3\n[[11/15/2019 11:57:37 PM]] best dice: 0.6211370617538596</p>\n\n<p>[[11/16/2019 03:40:51 AM]] max_threshold: 0.49999999999999994\n[[11/16/2019 03:40:51 AM]] threshold: 0.3\n[[11/16/2019 03:40:51 AM]] best dice: 0.6190079927221759</p>\n\n<p>[[11/16/2019 09:27:40 AM]] max_threshold: 0.6499999999999999\n[[11/16/2019 09:27:40 AM]] threshold: 0.25\n[[11/16/2019 09:27:40 AM]] best dice: 0.6250109995914948</p>\n\n<p>[[11/16/2019 11:40:06 AM]] max_threshold: 0.5999999999999999\n[[11/16/2019 11:40:06 AM]] threshold: 0.39999999999999997\n[[11/16/2019 11:40:06 AM]] best dice: 0.6094671966791182</p>\n\n<p>[[11/16/2019 01:02:53 PM]] max_threshold: 0.49999999999999994\n[[11/16/2019 01:02:53 PM]] threshold: 0.25\n[[11/16/2019 01:02:53 PM]] best dice: 0.6099291254010601</p>\n\n<p>-------- 3 -------------</p>\n\n<p>[[11/15/2019 11:58:44 PM]] max_threshold: 0.5999999999999999\n[[11/15/2019 11:58:44 PM]] threshold: 0.3\n[[11/15/2019 11:58:44 PM]] best dice: 0.5958070768192975</p>\n\n<p>[[11/16/2019 03:42:02 AM]] max_threshold: 0.6499999999999999\n[[11/16/2019 03:42:02 AM]] threshold: 0.39999999999999997\n[[11/16/2019 03:42:02 AM]] best dice: 0.596168368005109</p>\n\n<p>[[11/16/2019 09:28:20 AM]] max_threshold: 0.5999999999999999\n[[11/16/2019 09:28:20 AM]] threshold: 0.3\n[[11/16/2019 09:28:20 AM]] best dice: 0.5984940444088049</p>\n\n<p>[[11/16/2019 11:40:45 AM]] max_threshold: 0.6499999999999999\n[[11/16/2019 11:40:45 AM]] threshold: 0.35\n[[11/16/2019 11:40:45 AM]] best dice: 0.5955454147587783</p>\n\n<p>[[11/16/2019 01:03:33 PM]] max_threshold: 0.7\n[[11/16/2019 01:03:33 PM]] threshold: 0.35\n[[11/16/2019 01:03:33 PM]] best dice: 0.605774435543988</p>",
          "rawMarkdown": "@hengck23  Thanks for your reply,\nI also agree that the threshold is lower than the default setting (0.7, 0.3)\nHere is the 5 fold with resnet 50 JPU ASPP, threshold found by hyperopt:\n-------- 0 -------------\n[[11/15/2019 11:55:35 PM]] max_threshold: 0.6499999999999999\n[[11/15/2019 11:55:35 PM]] threshold: 0.3\n[[11/15/2019 11:55:35 PM]] best dice: 0.6662194166950253\n\n[[11/16/2019 03:38:42 AM]] max_threshold: 0.5499999999999999\n[[11/16/2019 03:38:42 AM]] threshold: 0.25\n[[11/16/2019 03:38:42 AM]] best dice: 0.6269578569840203\n\n[[11/16/2019 09:26:30 AM]] max_threshold: 0.7\n[[11/16/2019 09:26:30 AM]] threshold: 0.35\n[[11/16/2019 09:26:30 AM]] best dice: 0.6324725536992442\n\n[[11/16/2019 11:38:55 AM]] max_threshold: 0.7499999999999998\n[[11/16/2019 11:38:55 AM]] threshold: 0.3\n[[11/16/2019 11:38:55 AM]] best dice: 0.6536089501721206\n\n[[11/16/2019 01:01:42 PM]] max_threshold: 0.6499999999999999\n[[11/16/2019 01:01:42 PM]] threshold: 0.25\n[[11/16/2019 01:01:42 PM]] best dice: 0.6198696383341208\n\n-------- 1 -------------\n[[11/15/2019 11:56:34 PM]] max_threshold: 0.7499999999999998\n[[11/15/2019 11:56:34 PM]] threshold: 0.35\n[[11/15/2019 11:56:34 PM]] best dice: 0.767333773896778\n\n[[11/16/2019 03:39:44 AM]] max_threshold: 0.7\n[[11/16/2019 03:39:44 AM]] threshold: 0.35\n[[11/16/2019 03:39:44 AM]] best dice: 0.7741650036825072\n\n[[11/16/2019 09:27:04 AM]] max_threshold: 0.6499999999999999\n[[11/16/2019 09:27:04 AM]] threshold: 0.35\n[[11/16/2019 09:27:04 AM]] best dice: 0.7717044965056499\n\n[[11/16/2019 11:39:28 AM]] max_threshold: 0.7\n[[11/16/2019 11:39:28 AM]] threshold: 0.25\n[[11/16/2019 11:39:28 AM]] best dice: 0.7702028346574968\n\n[[11/16/2019 01:02:16 PM]] max_threshold: 0.7\n[[11/16/2019 01:02:16 PM]] threshold: 0.44999999999999996\n[[11/16/2019 01:02:16 PM]] best dice: 0.7569014497993047\n\n-------- 2 -------------\n\n[[11/15/2019 11:57:37 PM]] max_threshold: 0.5999999999999999\n[[11/15/2019 11:57:37 PM]] threshold: 0.3\n[[11/15/2019 11:57:37 PM]] best dice: 0.6211370617538596\n\n[[11/16/2019 03:40:51 AM]] max_threshold: 0.49999999999999994\n[[11/16/2019 03:40:51 AM]] threshold: 0.3\n[[11/16/2019 03:40:51 AM]] best dice: 0.6190079927221759\n\n[[11/16/2019 09:27:40 AM]] max_threshold: 0.6499999999999999\n[[11/16/2019 09:27:40 AM]] threshold: 0.25\n[[11/16/2019 09:27:40 AM]] best dice: 0.6250109995914948\n\n[[11/16/2019 11:40:06 AM]] max_threshold: 0.5999999999999999\n[[11/16/2019 11:40:06 AM]] threshold: 0.39999999999999997\n[[11/16/2019 11:40:06 AM]] best dice: 0.6094671966791182\n\n[[11/16/2019 01:02:53 PM]] max_threshold: 0.49999999999999994\n[[11/16/2019 01:02:53 PM]] threshold: 0.25\n[[11/16/2019 01:02:53 PM]] best dice: 0.6099291254010601\n\n-------- 3 -------------\n\n[[11/15/2019 11:58:44 PM]] max_threshold: 0.5999999999999999\n[[11/15/2019 11:58:44 PM]] threshold: 0.3\n[[11/15/2019 11:58:44 PM]] best dice: 0.5958070768192975\n\n[[11/16/2019 03:42:02 AM]] max_threshold: 0.6499999999999999\n[[11/16/2019 03:42:02 AM]] threshold: 0.39999999999999997\n[[11/16/2019 03:42:02 AM]] best dice: 0.596168368005109\n\n[[11/16/2019 09:28:20 AM]] max_threshold: 0.5999999999999999\n[[11/16/2019 09:28:20 AM]] threshold: 0.3\n[[11/16/2019 09:28:20 AM]] best dice: 0.5984940444088049\n\n[[11/16/2019 11:40:45 AM]] max_threshold: 0.6499999999999999\n[[11/16/2019 11:40:45 AM]] threshold: 0.35\n[[11/16/2019 11:40:45 AM]] best dice: 0.5955454147587783\n\n[[11/16/2019 01:03:33 PM]] max_threshold: 0.7\n[[11/16/2019 01:03:33 PM]] threshold: 0.35\n[[11/16/2019 01:03:33 PM]] best dice: 0.605774435543988"
        }
      ]
    },
    {
      "id": 672859,
      "postDate": "2019-11-14T08:30:47.040Z",
      "content": "<p>Do you have a post process in validation?</p>",
      "rawMarkdown": "Do you have a post process in validation?\n",
      "replies": [
        {
          "id": 672861,
          "postDate": "2019-11-14T08:31:31.040Z",
          "content": "<p>no. except thresholding</p>",
          "rawMarkdown": "no. except thresholding",
          "votes": 1
        },
        {
          "id": 672864,
          "postDate": "2019-11-14T08:33:30.413Z",
          "content": "<p>Hi Heng, did you try different image size? wondering how much does it affect?</p>",
          "rawMarkdown": "Hi Heng, did you try different image size? wondering how much does it affect?"
        },
        {
          "id": 672953,
          "postDate": "2019-11-14T10:34:25.620Z",
          "content": "<p>\"did you try different image size? \"</p>\n\n<p>it will affect results</p>\n\n<p>the best results i have are \n576x384 (network scaling structure, see resnet34unet,resnet34-jpu-aspp )\n256x384 (network scaling structure, see resnet34unet)\n1050x700 (network scaling structure, see resnet34-fpn)</p>",
          "rawMarkdown": "\"did you try different image size? \"\n\nit will affect results\n\nthe best results i have are \n576x384 (network scaling structure, see resnet34unet,resnet34-jpu-aspp )\n256x384 (network scaling structure, see resnet34unet)\n1050x700 (network scaling structure, see resnet34-fpn)",
          "votes": 1
        }
      ]
    },
    {
      "id": 672189,
      "postDate": "2019-11-13T15:44:54.663Z",
      "content": "<p>Thanks, Heng. I ran your resnet34-fpn, got similar vaild performance but the model only gave 0.639 lb. Is this normal or I missed something?</p>",
      "rawMarkdown": "Thanks, Heng. I ran your resnet34-fpn, got similar vaild performance but the model only gave 0.639 lb. Is this normal or I missed something?",
      "replies": [
        {
          "id": 672197,
          "postDate": "2019-11-13T15:54:26.157Z",
          "content": "<p>check the output of summarise_submission_csv(df). if you don't mind, post the output here and i will take a look. (attach train log and submit log)</p>\n\n<p>i think it is the threshold problem</p>",
          "rawMarkdown": "check the output of summarise\\_submission\\_csv(df). if you don't mind, post the output here and i will take a look. (attach train log and submit log)\n\ni think it is the threshold problem"
        },
        {
          "id": 672235,
          "postDate": "2019-11-13T17:05:54.290Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F61b1c74a54b386657915ada8530d8599%2FSelection_052.png?generation=1573664718561797&amp;alt=media\" alt=\"\"></p>\n\n<p>here you can see that the same loss value 0.71 gives kaggle score from 0.64 to 0.65. tn is true negative rate and tp is true positive rate of image level label.</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F61b1c74a54b386657915ada8530d8599%2FSelection_052.png?generation=1573664718561797&amp;alt=media)\n\n\nhere you can see that the same loss value 0.71 gives kaggle score from 0.64 to 0.65. tn is true negative rate and tp is true positive rate of image level label.",
          "votes": 1
        },
        {
          "id": 672439,
          "postDate": "2019-11-13T21:57:20.717Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>  how do your kaggle column scores and losses change so fast? Are you using a larger batch size or learning rate?</p>",
          "rawMarkdown": "@hengck23  how do your kaggle column scores and losses change so fast? Are you using a larger batch size or learning rate?"
        },
        {
          "id": 672453,
          "postDate": "2019-11-13T23:01:34.703Z",
          "content": "<p>the initialization model is a one trained with different setting before.</p>\n\n<p>to speedup experiments, i do not train from scratch for different setting. rather i reuse previously trained model.</p>",
          "rawMarkdown": "the initialization model is a one trained with different setting before.\n\nto speedup experiments, i do not train from scratch for different setting. rather i reuse previously trained model.",
          "votes": 1
        },
        {
          "id": 672661,
          "postDate": "2019-11-14T03:52:43.470Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F847605%2F7cec2662cd0e9a11c48213d704bb0834%2FIMG20191114_120024.jpg?generation=1573703873692938&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F847605%2F791b66786ca7bbade4f7f62fb3f82b79%2FIMG20191114_115650.jpg?generation=1573703656470288&amp;alt=media\" alt=\"\"></p>\n\n<p>This result gives 0.64. Thanks Heng. Submission threshold is 0.7.</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F847605%2F7cec2662cd0e9a11c48213d704bb0834%2FIMG20191114_120024.jpg?generation=1573703873692938&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F847605%2F791b66786ca7bbade4f7f62fb3f82b79%2FIMG20191114_115650.jpg?generation=1573703656470288&amp;alt=media)\n\nThis result gives 0.64. Thanks Heng. Submission threshold is 0.7."
        },
        {
          "id": 672680,
          "postDate": "2019-11-14T04:13:00.043Z",
          "content": "<p>local cv 0.63875 is a bit low.  you can try threshold like 0.65,0.70 ... to see if it helps. you need to get local cv in the range of 0.65 to 0.66 to get good kaggle lb score</p>\n\n<p>in the training log you choose 0.652. but i note that it is close to 0.621.\nyou should  do a column copy and paste, and plot the graph in excel. we look are trend (rather then the individual values, since the values are fluctuating up and down). my code save models at very n-th iteration. try using other nearby models. </p>\n\n<p>i cannot see the whole of the log so i don't known what iteration and learning rate you are using.  i suggest you use high constant rate of 0.01. let it overfits and then choose the best model based on the best point of the \"smoothed graph\" (the overfitting curve i showed below)</p>\n\n<p>you haven't show the output of  summarise_submission_csv(df) part yet.  you just need to ensure that the predicted <strong>number</strong> of +ve, -ve and none mask is closed the reference values below.</p>",
          "rawMarkdown": "local cv 0.63875 is a bit low.  you can try threshold like 0.65,0.70 ... to see if it helps. you need to get local cv in the range of 0.65 to 0.66 to get good kaggle lb score\n\nin the training log you choose 0.652. but i note that it is close to 0.621.\nyou should  do a column copy and paste, and plot the graph in excel. we look are trend (rather then the individual values, since the values are fluctuating up and down). my code save models at very n-th iteration. try using other nearby models. \n\ni cannot see the whole of the log so i don't known what iteration and learning rate you are using.  i suggest you use high constant rate of 0.01. let it overfits and then choose the best model based on the best point of the \"smoothed graph\" (the overfitting curve i showed below)\n\nyou haven't show the output of  summarise\\_submission\\_csv(df) part yet.  you just need to ensure that the predicted **number** of +ve, -ve and none mask is closed the reference values below.\n ",
          "votes": 1
        },
        {
          "id": 672705,
          "postDate": "2019-11-14T04:33:01.643Z",
          "content": "<p>on a side note,\ntake a row of my log file and you have a bunch of numbers indicating validation loss and metric, e.g [ 0.2, 0.4 0.5 ,0.6, 0.6 ......] = x</p>\n\n<p>if you make a submission you have the kaggle score = s</p>\n\n<p>with enough submission samples, we can do this:\npredicted s = XgBoost(x).</p>\n\n<p>i.e. you can build a kaggle score predictor</p>",
          "rawMarkdown": "on a side note,\ntake a row of my log file and you have a bunch of numbers indicating validation loss and metric, e.g [ 0.2, 0.4 0.5 ,0.6, 0.6 ......] = x\n\nif you make a submission you have the kaggle score = s\n\nwith enough submission samples, we can do this:\npredicted s = XgBoost(x).\n\ni.e. you can build a kaggle score predictor",
          "votes": 3
        },
        {
          "id": 672719,
          "postDate": "2019-11-14T04:52:14.877Z",
          "content": "<p>Thanks Heng, I will retry model selection and threshold selection.\nAnd summarise_submission_csv here is similar to resnet34-unet result(lb 659)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F847605%2F6d31d3d7c4181e4a89e07646241a275c%2F.png?generation=1573707129225978&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thanks Heng, I will retry model selection and threshold selection.\nAnd summarise_submission_csv here is similar to resnet34-unet result(lb 659)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F847605%2F6d31d3d7c4181e4a89e07646241a275c%2F.png?generation=1573707129225978&amp;alt=media)\n"
        },
        {
          "id": 672737,
          "postDate": "2019-11-14T05:16:15.200Z",
          "content": "<p>i don't train your model so i cannot be sure if the below work. you may want to verify on validation set if threshold mask =0.4 is the best your you? e.g. 0.3</p>\n\n<p>this is not reflected in the  summarise_submission_csv becuase the function can only  output image level information and not pixel level information</p>",
          "rawMarkdown": "i don't train your model so i cannot be sure if the below work. you may want to verify on validation set if threshold mask =0.4 is the best your you? e.g. 0.3\n\nthis is not reflected in the  summarise\\_submission\\_csv becuase the function can only  output image level information and not pixel level information",
          "votes": 1
        },
        {
          "id": 672794,
          "postDate": "2019-11-14T06:55:11.570Z",
          "content": "<p>I really appreciate your help. I learnt so much from your code : )</p>",
          "rawMarkdown": "I really appreciate your help. I learnt so much from your code : )"
        }
      ]
    },
    {
      "id": 672085,
      "postDate": "2019-11-13T13:58:05.870Z",
      "content": "<p>amazing kenal 👍 </p>",
      "rawMarkdown": "amazing kenal 👍 "
    },
    {
      "id": 671716,
      "postDate": "2019-11-13T04:37:53.947Z",
      "content": "<p>hello, thanks for your sharing!\nI have a small question about the data:\nTake \"0a14f2b.png\" for example: how can I get the same mask as you did?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2154603%2F100193f640d51cad5c8e063bddd0c55b%2Fmyplot.png?generation=1573619814792392&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2154603%2F368a207e87078df3b936cc6796f5b1f6%2F0a14f2b.png?generation=1573619846740705&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "hello, thanks for your sharing!\nI have a small question about the data:\nTake \"0a14f2b.png\" for example: how can I get the same mask as you did?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2154603%2F100193f640d51cad5c8e063bddd0c55b%2Fmyplot.png?generation=1573619814792392&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2154603%2F368a207e87078df3b936cc6796f5b1f6%2F0a14f2b.png?generation=1573619846740705&amp;alt=media)\n",
      "replies": [
        {
          "id": 671757,
          "postDate": "2019-11-13T05:55:36.693Z",
          "content": "<p>use opencv to save as 4 channel png image.\ni think it is in kaggle.py, there is a function called dump_mask_to_png()</p>\n\n<p>the advantage is that you can use opencv to read back and use opencv flip and scale, rotate function later for augmentation</p>",
          "rawMarkdown": "use opencv to save as 4 channel png image.\ni think it is in kaggle.py, there is a function called dump\\_mask\\_to\\_png()\n\nthe advantage is that you can use opencv to read back and use opencv flip and scale, rotate function later for augmentation",
          "votes": 2
        }
      ]
    },
    {
      "id": 671705,
      "postDate": "2019-11-13T04:17:50.553Z",
      "content": "<p>hi Heng, I found you use constant lr 0.01 to train</p>\n\n<p>do you use any schduler, like ReduceLROnPlateau? or you just change lr by hand and restart from some checkpoint?</p>",
      "rawMarkdown": "hi Heng, I found you use constant lr 0.01 to train\n\ndo you use any schduler, like ReduceLROnPlateau? or you just change lr by hand and restart from some checkpoint?",
      "replies": [
        {
          "id": 671709,
          "postDate": "2019-11-13T04:24:34.800Z",
          "content": "<p>i usually use hand tuned learning rate. but for this competition, even initial learning rate overfits. so i only use single rate so far in my submission.</p>\n\n<p>in some experiments, initial rate rate =0.05 also works.</p>\n\n<p>i also try to reduce learning rate by hand, but most of them overfits very fast, some even before one epoach</p>",
          "rawMarkdown": "i usually use hand tuned learning rate. but for this competition, even initial learning rate overfits. so i only use single rate so far in my submission.\n\nin some experiments, initial rate rate =0.05 also works.\n\ni also try to reduce learning rate by hand, but most of them overfits very fast, some even before one epoach",
          "votes": 1
        },
        {
          "id": 671711,
          "postDate": "2019-11-13T04:27:31.923Z",
          "content": "<p>on a side note, ensemble of large learning rate is less likely to overfit sometimes.</p>",
          "rawMarkdown": "on a side note, ensemble of large learning rate is less likely to overfit sometimes.",
          "votes": 2
        },
        {
          "id": 671743,
          "postDate": "2019-11-13T05:26:23.840Z",
          "content": "<p>thanks for your advice, good luck:)</p>",
          "rawMarkdown": "thanks for your advice, good luck:)"
        }
      ]
    },
    {
      "id": 671089,
      "postDate": "2019-11-12T08:44:21.787Z",
      "content": "<p>Resizing image is ok, but how to resize masks? 4 separate images corresponding to 4 classes?</p>",
      "rawMarkdown": "Resizing image is ok, but how to resize masks? 4 separate images corresponding to 4 classes?",
      "replies": [
        {
          "id": 671102,
          "postDate": "2019-11-12T08:54:31.227Z",
          "content": "<p>```\nprobability = net(resize_image)\nprobability = change_back_to_ground_truth_size(probability)</p>\n\n<p>loss = loss_function(probability,truth_mask)</p>\n\n<p>```</p>\n\n<p>for this challenge the truth mask are quite \"blocky\" and do have have fine detail.\nthe submission size is  525 x 350. it is ok if your output from the network is 1x ~ 0.25x that of 525 x 350.  (i.e.  change_back_to_ground_truth_size() will perform a 1 to 4x upsize) you should do experiment to confirm</p>\n\n<hr>\n\n<p>opencv can resize 4-channel image.\npytorch and tensorflow have their resize function as well in gpu/cpu.</p>",
          "rawMarkdown": "```\nprobability = net(resize_image)\nprobability = change_back_to_ground_truth_size(probability)\n\nloss = loss_function(probability,truth_mask)\n\n```\n\nfor this challenge the truth mask are quite \"blocky\" and do have have fine detail.\nthe submission size is  525 x 350. it is ok if your output from the network is 1x ~ 0.25x that of 525 x 350.  (i.e.  change\\_back\\_to\\_ground\\_truth_size() will perform a 1 to 4x upsize) you should do experiment to confirm\n\n\n---\n\nopencv can resize 4-channel image.\npytorch and tensorflow have their resize function as well in gpu/cpu."
        },
        {
          "id": 671112,
          "postDate": "2019-11-12T09:07:54.140Z",
          "content": "<p>Thanks, Heng</p>",
          "rawMarkdown": "Thanks, Heng"
        }
      ]
    },
    {
      "id": 670403,
      "postDate": "2019-11-11T12:34:00.813Z",
      "content": "<p>Thank you very very much. I will do my best from now on. </p>",
      "rawMarkdown": "Thank you very very much. I will do my best from now on. "
    },
    {
      "id": 670135,
      "postDate": "2019-11-11T04:04:31.597Z",
      "content": "<p>best single fold .658, ensemble .667 ??</p>",
      "rawMarkdown": "best single fold .658, ensemble .667 ??",
      "replies": [
        {
          "id": 670172,
          "postDate": "2019-11-11T05:53:44.490Z",
          "content": "<p>yes, this is what i get. but different optimal thresholds were found for single fold and ensemble.</p>\n\n<p>my solution does not differs much from the top notebook solution. i suspect good blending of those notebook solution might achieve similar results?</p>",
          "rawMarkdown": "yes, this is what i get. but different optimal thresholds were found for single fold and ensemble.\n\nmy solution does not differs much from the top notebook solution. i suspect good blending of those notebook solution might achieve similar results?"
        },
        {
          "id": 670806,
          "postDate": "2019-11-11T22:08:28.893Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> how do you optimize your thresholds? on a hold out set or directly on LB? :)</p>",
          "rawMarkdown": "@hengck23 how do you optimize your thresholds? on a hold out set or directly on LB? :)"
        },
        {
          "id": 670919,
          "postDate": "2019-11-12T03:38:18.723Z",
          "content": "<p>i haven't find a good way yet. i would have to think about it in the last 7 days.</p>\n\n<p>unlike the steel competition, we get to see all public and private data. we should think of how to get stable private test results based on learning on the test+train and feedback on the public test.</p>\n\n<p>the test are not labelled but they still contain information, e.g. like input density distribution, nearest neighbors, clusters, etc. i would have to think of a way to mine this information.</p>\n\n<p>you can think of how to pull the private and public test features close together correctly</p>",
          "rawMarkdown": "i haven't find a good way yet. i would have to think about it in the last 7 days.\n\nunlike the steel competition, we get to see all public and private data. we should think of how to get stable private test results based on learning on the test+train and feedback on the public test.\n\nthe test are not labelled but they still contain information, e.g. like input density distribution, nearest neighbors, clusters, etc. i would have to think of a way to mine this information.\n\nyou can think of how to pull the private and public test features close together correctly",
          "votes": 2
        }
      ]
    },
    {
      "id": 669946,
      "postDate": "2019-11-10T18:26:49.187Z",
      "content": "<p>wow, nearly 0.01 ensembling boost? I got a single model/fold 0.666 but struggle to ensemble...</p>",
      "rawMarkdown": "wow, nearly 0.01 ensembling boost? I got a single model/fold 0.666 but struggle to ensemble...",
      "replies": [
        {
          "id": 670062,
          "postDate": "2019-11-11T00:13:59.827Z",
          "content": "<p>i think it is the threshold that works for me. another reason could be the use of different architecture and different input size</p>",
          "rawMarkdown": "i think it is the threshold that works for me. another reason could be the use of different architecture and different input size"
        },
        {
          "id": 670099,
          "postDate": "2019-11-11T02:25:06.720Z",
          "content": "<p>wow!  0.666 with single fold is  amazing</p>",
          "rawMarkdown": "wow!  0.666 with single fold is  amazing"
        }
      ]
    },
    {
      "id": 669222,
      "postDate": "2019-11-09T17:57:28.847Z",
      "content": "<p>the clouds are actually Shallow cumulus in the trade wind region? So maybe they have specific eddy pattern ...\nif the satellite image are taken at fixed time of the day, the sunlight and shadow are consistent. i suspect that by aligning the kaggle images back to the original orientation (the big black slant facing in the same direction), the data may be more consistent (e.g. test are more close to train?)</p>",
      "rawMarkdown": "the clouds are actually Shallow cumulus in the trade wind region? So maybe they have specific eddy pattern ...\nif the satellite image are taken at fixed time of the day, the sunlight and shadow are consistent. i suspect that by aligning the kaggle images back to the original orientation (the big black slant facing in the same direction), the data may be more consistent (e.g. test are more close to train?)",
      "replies": [
        {
          "id": 669230,
          "postDate": "2019-11-09T18:15:37.737Z",
          "content": "<p>I believe they're already in their original orientation. The lines travel from top-left to bottom-right for the Aqua satellite and from top-right to bottom-left for the Terra satellite.</p>\n\n<p>Here's where you can see the lines for a single satellite: <a href=\"https://worldview.earthdata.nasa.gov/\">https://worldview.earthdata.nasa.gov/</a></p>\n\n<p>Below are two snapshots of the same area, one from each satellite.</p>\n\n<p>Aqua:\n<a href=\"https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-24T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375\">https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-24T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375</a></p>\n\n<p>Terra:\n<a href=\"https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-24T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375\">https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-24T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375</a></p>",
          "rawMarkdown": "I believe they're already in their original orientation. The lines travel from top-left to bottom-right for the Aqua satellite and from top-right to bottom-left for the Terra satellite.\n\nHere's where you can see the lines for a single satellite: https://worldview.earthdata.nasa.gov/\n\nBelow are two snapshots of the same area, one from each satellite.\n\nAqua:\nhttps://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-24T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375\n\nTerra:\nhttps://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-24T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375\n\n"
        },
        {
          "id": 669238,
          "postDate": "2019-11-09T18:25:05.053Z",
          "content": "<p>This visualization also helps to understand why there are black regions:\n<a href=\"https://sos.noaa.gov/datasets/polar-orbiting-aqua-satellite-and-modis-swath/\">https://sos.noaa.gov/datasets/polar-orbiting-aqua-satellite-and-modis-swath/</a></p>",
          "rawMarkdown": "This visualization also helps to understand why there are black regions:\nhttps://sos.noaa.gov/datasets/polar-orbiting-aqua-satellite-and-modis-swath/"
        },
        {
          "id": 669242,
          "postDate": "2019-11-09T18:37:16.983Z",
          "content": "<p><img src=\"https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-24T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Terra_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375\" alt=\"\"></p>\n\n<p><img src=\"https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-24T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375\" alt=\"\"></p>",
          "rawMarkdown": " ![](https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-24T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Terra_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375)\n\n![](https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-24T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375)\n\n"
        },
        {
          "id": 669250,
          "postDate": "2019-11-09T18:50:16.790Z",
          "content": "<p>actually this means that we can download pairs of images from the the 2 different satellites and trained in unsupervised way. the \"consistency loss\" between the pair will be used for back progragtion</p>",
          "rawMarkdown": "actually this means that we can download pairs of images from the the 2 different satellites and trained in unsupervised way. the \"consistency loss\" between the pair will be used for back progragtion"
        },
        {
          "id": 669251,
          "postDate": "2019-11-09T18:53:41.747Z",
          "content": "<p>disposed old comment</p>",
          "rawMarkdown": "disposed old comment",
          "votes": 3
        },
        {
          "id": 669252,
          "postDate": "2019-11-09T18:53:54.793Z",
          "content": "<p>i further wonder if there are satellite image pair split into test and train (data leak)?</p>",
          "rawMarkdown": "i further wonder if there are satellite image pair split into test and train (data leak)?"
        },
        {
          "id": 669254,
          "postDate": "2019-11-09T18:57:47.700Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 669255,
          "postDate": "2019-11-09T18:59:02.077Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 669262,
          "postDate": "2019-11-09T19:18:58.847Z",
          "content": "<p>&gt; I further wonder if there are satellite image pair split into test and train (data leak)?</p>\n\n<p>I trained a network with contrastive loss to find pairs and was able to find</p>\n\n<p>1,434 pairs with two images in train (ie. 2,868 images total)\n1,813 pairs with one in train, one in test\n634 pairs with two images in test (ie. 1,276 images total)</p>\n\n<p>So I found pairs for 7,770 of 9, 244 total images (but all images should have a pair).</p>\n\n<p>I tried many ways of using this information but as <a href=\"/robga\">@robga</a> mentioned the noisy labels mean that even two images taken a few hours apart usually have drastically different labels. :(</p>\n\n<p>P.S. If you think you can make better use of these pairs than I have there are still two days left to create teams! ;)</p>",
          "rawMarkdown": "&gt; I further wonder if there are satellite image pair split into test and train (data leak)?\n\nI trained a network with contrastive loss to find pairs and was able to find\n\n1,434 pairs with two images in train (ie. 2,868 images total)\n1,813 pairs with one in train, one in test\n634 pairs with two images in test (ie. 1,276 images total)\n\nSo I found pairs for 7,770 of 9, 244 total images (but all images should have a pair).\n\nI tried many ways of using this information but as @robga mentioned the noisy labels mean that even two images taken a few hours apart usually have drastically different labels. :(\n\nP.S. If you think you can make better use of these pairs than I have there are still two days left to create teams! ;)",
          "votes": 1
        },
        {
          "id": 669264,
          "postDate": "2019-11-09T19:26:16.503Z",
          "content": "<p>disposed old comment</p>",
          "rawMarkdown": "disposed old comment",
          "votes": 1
        },
        {
          "id": 669928,
          "postDate": "2019-11-10T17:42:06.360Z",
          "content": "<p>“better use of these pairs ”...</p>\n\n<p>The pairs are less noisy for classification than segmentation . They are less noisy for certain  class</p>",
          "rawMarkdown": "“better use of these pairs ”...\n\nThe pairs are less noisy for classification than segmentation . They are less noisy for certain  class"
        },
        {
          "id": 670079,
          "postDate": "2019-11-11T01:15:42.287Z",
          "content": "<p>try mixup augmentation for satellite pairs. both input and target space will be smoothed</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fb55dee59e82d8cc6c4508749d47d9e2a%2Fmixup.png?generation=1573434923361204&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "try mixup augmentation for satellite pairs. both input and target space will be smoothed\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fb55dee59e82d8cc6c4508749d47d9e2a%2Fmixup.png?generation=1573434923361204&amp;alt=media)\n"
        },
        {
          "id": 670670,
          "postDate": "2019-11-11T17:58:53.557Z",
          "content": "<p>i find this interesting. although the images are already aligned,  using aligned image in TTA seems to improve results consistently in local validation . (i.e. you can use flip horizontal+vertical or rotate 180 degree)</p>",
          "rawMarkdown": "i find this interesting. although the images are already aligned,  using aligned image in TTA seems to improve results consistently in local validation . (i.e. you can use flip horizontal+vertical or rotate 180 degree)"
        }
      ]
    },
    {
      "id": 668579,
      "postDate": "2019-11-08T15:28:29.743Z",
      "content": "<p>mix augmentation</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa9129eca14eede16bc8b59a3c95b17bb%2FSelection_060.png?generation=1573226907253607&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "mix augmentation\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa9129eca14eede16bc8b59a3c95b17bb%2FSelection_060.png?generation=1573226907253607&amp;alt=media)\n"
    },
    {
      "id": 666390,
      "postDate": "2019-11-06T04:01:28.750Z",
      "content": "<p>Welcome <a href=\"/hengck23\">@hengck23</a>, I was missing you in this competition. Please be carefull this time and Best of luck!!!!</p>",
      "rawMarkdown": "Welcome @hengck23, I was missing you in this competition. Please be carefull this time and Best of luck!!!!"
    },
    {
      "id": 674438,
      "postDate": "2019-11-16T14:05:11.787Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true,
      "replies": [
        {
          "id": 674445,
          "postDate": "2019-11-16T14:19:13.650Z",
          "content": "<p>what works for one may not work for another. there are really a lot of factors involved, especially for this noisy set. Even when i do my experiments, same method can leads to very different results for different split.</p>\n\n<p>e.g. i bet people have different results when ensemble. another example is the use of temperature in another post. so don't expect magic if you just apply an idea in the discussion.</p>\n\n<p>so far i haven't disclose any thing very special. they are just a bunch of common methods that you will find in other kaggle competitions, and also some new ideas i have.</p>",
          "rawMarkdown": "what works for one may not work for another. there are really a lot of factors involved, especially for this noisy set. Even when i do my experiments, same method can leads to very different results for different split.\n\ne.g. i bet people have different results when ensemble. another example is the use of temperature in another post. so don't expect magic if you just apply an idea in the discussion.\n\nso far i haven't disclose any thing very special. they are just a bunch of common methods that you will find in other kaggle competitions, and also some new ideas i have.",
          "votes": 7
        }
      ]
    },
    {
      "id": 670673,
      "postDate": "2019-11-11T18:05:59.583Z",
      "rawMarkdown": "",
      "votes": -8,
      "isDeleted": true,
      "replies": [
        {
          "id": 670916,
          "postDate": "2019-11-12T03:33:25.833Z",
          "content": "<p>i don't think you have to transfer all my code in your private kernel.\nthe important thing are just the model and loss function.</p>\n\n<p>you can then just your original dataloader and training code on it.</p>",
          "rawMarkdown": "i don't think you have to transfer all my code in your private kernel.\nthe important thing are just the model and loss function.\n\nyou can then just your original dataloader and training code on it.",
          "votes": 4
        },
        {
          "id": 670985,
          "postDate": "2019-11-12T05:41:07.927Z",
          "content": "<p>What is the NullScheduler, couldn't find anything about it</p>",
          "rawMarkdown": "What is the NullScheduler, couldn't find anything about it"
        },
        {
          "id": 670987,
          "postDate": "2019-11-12T05:44:14.777Z",
          "content": "<p>you can ignore this. null scheduler is basically using constant learning rate</p>",
          "rawMarkdown": "you can ignore this. null scheduler is basically using constant learning rate"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 668036,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-07T22:26:00.563000",
      "content": "<p>here is one data leak!</p>\n\n<ul>\n<li>each image has at least one label .... you can use adaptive threshold to ensure each image has at least one type of cloud or you can design a loss function for that (e.g. ranking loss)</li>\n</ul>",
      "votes": 9,
      "replies": [
        {
          "id": 668901,
          "author_name": "Endi Niu",
          "author_url": "",
          "post_date": "2019-11-09T03:51:29.447000",
          "content": "<p>this could be helpful for postprocess\n.6609 model, found 100 images negative, force them to predict mask, improve .6614</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 668902,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-09T03:56:03.343000",
          "content": "<p>Training a softmax classification model for class selection may work</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 669537,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-10T06:19:34.493000",
      "content": "<p><a href=\"https://events.ecmwf.int/event/118/contributions/538/attachments/139/244/OCBWF-Bony.pdf\">https://events.ecmwf.int/event/118/contributions/538/attachments/139/244/OCBWF-Bony.pdf</a>\n<a href=\"https://rmets.onlinelibrary.wiley.com/doi/epdf/10.1002/qj.3662\">https://rmets.onlinelibrary.wiley.com/doi/epdf/10.1002/qj.3662</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe371d138ff40a6d50a9d09d41343d903%2FSelection_053.png?generation=1573366766042602&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd4d6206c17b3abb64bb2fa195e4ecf41%2FSelection_051.png?generation=1573366770599581&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F8527544c55b481b58e7a635cdffd1ed8%2FSelection_052.png?generation=1573366771716980&amp;alt=media\" alt=\"\"></p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 671393,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-12T16:25:41.047000",
      "content": "<p>0.6621 single model!</p>\n\n<p>i tried several segmentation head: unet decoder head, fpn head, ASPP head. I think ASPP is the best.</p>\n\n<p>```</p>\n\n<p>class Net(nn.Module):\n    def load_pretrain(self, skip=['logit.'], is_print=True):\n        load_pretrain(self, skip, pretrain_file=PRETRAIN_FILE, conversion=CONVERSION, is_print=is_print)</p>\n\n<pre><code>def __init__(self, num_class=4):\n    super(Net, self).__init__()\n\n    e = ResNet34()\n    self.block0 = e.block0\n    self.block1 = e.block1\n    self.block2 = e.block2\n    self.block3 = e.block3\n    self.block4 = e.block4\n    e = None  #dropped\n\n    self.jpu = JointPyramidUpsample([512,256,128],128)\n    #self.aspp = ASPP(512, 128, rate=[6,12,18], dropout_rate=0.1)\n    self.aspp = ASPP(512, 128, rate=[4,8,12], dropout_rate=0.1)\n    self.logit = nn.Conv2d(128,num_class,kernel_size=1)\n\n\ndef forward(self, x):\n    batch_size,C,H,W = x.shape\n\n    x0 = self.block0(x)\n    x1 = self.block1(x0)\n    x2 = self.block2(x1)\n    x3 = self.block3(x2)\n    x4 = self.block4(x3)\n\n    x = self.jpu([x4,x3,x2])\n    x = self.aspp(x)\n    logit = self.logit(x)\n\n    #---\n    probability_mask  = torch.sigmoid(logit)\n    probability_label = F.adaptive_max_pool2d(probability_mask,1).view(batch_size,-1)\n    return probability_label, probability_mask\n</code></pre>\n\n<p>```</p>",
      "votes": 7,
      "replies": [
        {
          "id": 671395,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-12T16:27:44.930000",
          "content": "<p>submission details:</p>\n\n<p>```</p>\n\n<p>submitting .... @ ['null', 'flip_lr', 'flip_ud', 'flip_both']\ninitial_checkpoint  = /root/share/project/kaggle/2019/cloud/result/run3/resnet34-jpu-aspp-576x384-fold_a2/checkpoint/00012000_model.pth\nthreshold_label = [0.7, 0.7, 0.7, 0.7]\nthreshold_mask  = [0.3, 0.3, 0.3, 0.3]\nthreshold_mask  = [1, 1, 1, 1]</p>\n\n<p>** net setting **\n 3698 / 3698   0 hr 02 min\n 0 min 02 sec\ntest submission .... @ ['null', 'flip_lr', 'flip_ud', 'flip_both']\ninitial_checkpoint=/root/share/project/kaggle/2019/cloud/result/run3/resnet34-jpu-aspp-576x384-fold_a2/checkpoint/00012000_model.pth</p>\n\n<p>compare with LB probing ... \n        num_image =  3698(3698) </p>\n\n<pre><code>    pos0 =  1303(1864)  0.699\n    pos1 =  1410(1508)  0.935\n    pos2 =  1084(1982)  0.547\n    pos3 =  2305(2382)  0.968\n\n    neg0 =  2395(1834)  1.306\n    neg1 =  2288(1940)  1.179\n    neg2 =  2614(2638)  0.991\n</code></pre>\n\n<h2>        neg3 =  1393(2017)  0.691</h2>\n\n<pre><code>    all_zero =   104 (?)\n</code></pre>\n\n<p>```</p>\n\n<p>validation</p>\n\n<p>```\ntest_dataset : \n    len = 300</p>\n\n<pre><code>mode    = train\nsplit   = ['by_random1/valid_fold_a2_300.npy']\ncsv     = ['train.csv']\nfolder  = {'image': '1050x700', 'mask': '525x350'}\nnum_image = 300\n            Fish   neg0, pos0 =   142  (0.473),    158  (0.527)\n          Flower   neg1, pos1 =   179  (0.597),    121  (0.403)\n          Gravel   neg2, pos2 =   144  (0.480),    156  (0.520)\n           Sugar   neg3, pos3 =   101  (0.337),    199  (0.663)\n</code></pre>\n\n<p>submitting .... @ ['null', 'flip_lr', 'flip_ud', 'flip_both']\ninitial_checkpoint  = /root/share/project/kaggle/2019/cloud/result/run3/resnet34-jpu-aspp-576x384-fold_a2/checkpoint/00012000_model.pth\nthreshold_label = [0.7, 0.7, 0.7, 0.7]\nthreshold_mask  = [0.3, 0.3, 0.3, 0.3]\nthreshold_mask  = [1, 1, 1, 1]</p>\n\n<p>** net setting **\n** all threshold **</p>\n\n<pre><code>             |   truth  |  predict |              |              |          \n</code></pre>\n\n<h2>                 | neg  pos | neg  pos | tn     tp    | dn     dp    | kaggle  </h2>\n\n<p>0      Fish     | 142  158 | 196  104 | 0.930  0.595 | 0.000  0.652 | 0.656 (0.761) <br>\n 1    Flower     | 179  121 | 188  112 | 0.911  0.793 | 0.000  0.776 | 0.790 (0.863) <br>\n 2    Gravel     | 144  156 | 213   87 | 0.931  0.494 | 0.000  0.703 | 0.618 (0.696) <br>\n 3     Sugar     | 101  199 | 113  187 | 0.743  0.809 | 0.000  0.686 | 0.622 (0.785)  </p>\n\n<p>kaggle (classification only) = 0.67159 (0.77636)</p>\n\n<p>** segmentation only **</p>\n\n<pre><code>             |   truth  |  predict |              |              |          \n</code></pre>\n\n<h2>                 | neg  pos | neg  pos | tn     tp    | dn     dp    | kaggle  </h2>\n\n<p>0      Fish     | 142  158 |   0  300 | 0.000  1.000 | 0.613  0.457 | 0.534 (0.504) <br>\n 1    Flower     | 179  121 |   0  300 | 0.000  1.000 | 0.754  0.665 | 0.718 (0.408) <br>\n 2    Gravel     | 144  156 |   0  300 | 0.000  1.000 | 0.611  0.477 | 0.539 (0.536) <br>\n 3     Sugar     | 101  199 |   0  300 | 0.000  1.000 | 0.297  0.616 | 0.503 (0.644)  </p>\n\n<p>kaggle (classification only) = 0.57352 (0.52299)</p>\n\n<p>```</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 671426,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2019-11-12T17:21:05.560000",
          "content": "<p>We got with UNet a single fold 0.6678. So it means that the decoder type has nothing to do with a high score?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 671652,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-13T01:56:51.070000",
          "content": "<p>\"We got with UNet a single fold 0.6678\"</p>\n\n<p>0.6678 is very good score and i am surprised. </p>\n\n<p>in my experiments, the score are quite sensitive. because of the negative images, the metric itself is non smooth. The BCE is smooth however. in my experiment, the ASPP head has a much BCE lower loss than others. so i conclude it is a better decoder type.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 671950,
          "author_name": "Thomas Yokota",
          "author_url": "",
          "post_date": "2019-11-13T11:10:05.607000",
          "content": "<p>Your findings look good <a href=\"/hengck23\">@hengck23</a>! What metric do you use for loss?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 671955,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-13T11:19:01.117000",
          "content": "<p>only BCE. you can check the code for more detail</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 671961,
          "author_name": "Thomas Yokota",
          "author_url": "",
          "post_date": "2019-11-13T11:26:46.627000",
          "content": "<p>Interesting. Thanks! Will do.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674442,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2019-11-16T14:15:01.717000",
          "content": "<p>JPU valid loss looks actually better than UNet. I accidently submitted a resnet34-JPU with a wrong label threshold 0.55 -&gt; LB 0.658. I think with a higher threshold it can achieve a good LB. The problem seems to be it predicts many masks in the test set -&gt; you have to search again some thresholds...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 670024,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2019-11-10T22:55:16.470000",
      "content": "<p>The competition is going to the final week. Please stop sharing. It's not nice to see solutions keeping posted like this.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 670914,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-12T03:30:59.567000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F07717d002f9a96888b2439d175227cbc%2Fguideline.png?generation=1573529402516418&amp;alt=media\" alt=\"\"></p>\n\n<p>kaggle guideline is stop publishing at the last week. it should be ok to put code publicly before this message comes out.</p>",
          "votes": 16,
          "replies": []
        }
      ]
    },
    {
      "id": 665655,
      "author_name": "Andrey Kiryasov",
      "author_url": "",
      "post_date": "2019-11-05T08:32:46.167000",
      "content": "<p>Please, be carefull )</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 670074,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-11T00:51:58.773000",
      "content": "<p>lesson from the steel competition is that  results is going to be sensitive to the image label threshold.</p>\n\n<p>one way to deal with this is to select proper high threshold, (like 0.70).</p>\n\n<p>another way is to modify the loss e.g loss = weight_pos*loss_pos +  weight_neg*loss_neg. By adjusting the weights, you can shift the threshold back to 0.50.</p>\n\n<p>to visualize the effects on classification, you can observe the roc curve of the weighted and un-weighted loss model. indicate tpr, fpr and threshold on your roc curve.</p>\n\n<p>for segmentation, just draw the probability heatmap in range (0,1). you can see the weighted version will be dilated (or eroded) version of the unweighted one.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 666095,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-05T18:14:38.337000",
      "content": "<p>daily tracking results:</p>\n\n<p>entry date : 2019-11-05</p>\n\n<p>```\nday-1, 2019-11-05: LB 0.644   (non-tuned resnet34-fpn) <br>\nday-3, 2019-11-08: LB 0.6586  (resnet34-fpn x3 fold) . for single fold : lb 0.656\nday-5, 2019-11-10: LB 0.6577  (resnet34-unet single). \n             ensemble of 1 x unet + 3 x fpn : lb 0.667  (finetune threshold, etc)</p>\n\n<p>day-6, 2019-11-11: LB 0.6709 (added a few misc models to ensemble, e.g. different input size, jpu+aspp (joint-pyramid-upsampling))</p>\n\n<p>day-8, 2019-11-13: LB 0.6713 (corrected a bug. wrongly resize image to widthxheight instead of heightxwidth). perform threshold search on LB to use up free submission slots, try 0.60 0.65, 0.67 as image label threshold. the lb score can ranges from  0.6664,0.6707, 0.6705. nevethelss, high threshold is preferred</p>\n\n<p>day-10 2019-11-14: LB 0.6721 (corrected the resize bug again! seems that this bug occurs in different part of my code). add 2 more models to my previous ensemble. i am still using the same models except for for more fold</p>\n\n<p>```</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 671545,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2019-11-12T21:24:11.957000",
      "content": "<p>You still are sharing top tips and even partial code for a near gold solutions. Can you please just stop and wait until it ends and share, I’m glad to discuss after that, not now. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 671558,
          "author_name": "Usmann Khan",
          "author_url": "",
          "post_date": "2019-11-12T21:42:09.380000",
          "content": "<p>This competition is really just \"ensemble hell\". I don't think <a href=\"/hengck23\">@hengck23</a> is spoiling much of anything by saying \"1xresnet34-unet + 3xfpn + TTA\" and publishing some scaffolding code.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 671561,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-11-12T21:47:03.923000",
          "content": "<p>This is the final week. Imagine if I share some tips to directly reach 0.671 on LB? If he is still sharing, who knows what he will share next? What’s the point of sharing here? Please just STOP.  Why not wait for 6 days and then he can share everything. But he rarely shared after the competition ends. That’s really strange. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 671646,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2019-11-13T01:42:28.397000",
          "content": "<p>Seems like the best way to get my first medal is to avoid join competition with Heng 😂\nBut still thankful for his share, I learned a lot.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 671649,
          "author_name": "Endi Niu",
          "author_url": "",
          "post_date": "2019-11-13T01:52:27.137000",
          "content": "<p>In the past segment competition, this can easily get you a medal, but now the community is getting familiar with segment task, it's harder, so it push me to learn more knowledge rather than copy my past code.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 666792,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-06T14:03:28.513000",
      "content": "<p>my plan  (to be updated as competition proceeds ...)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbd4150e743b7ac4b6eb9a1da4c03229d%2FSelection_053.png?generation=1573049004668282&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 676064,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-19T00:17:46.960000",
      "content": "<p>closing remarks:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F851c2335980981005516610ceb5348dc%2FSelection_061.png?generation=1574122453993841&amp;alt=media\" alt=\"\"></p>\n\n<p>best private score this code can get is LB 0.66441\nmethod: \n1. just train with all samples, use 1x  resnet34-unet (384x256), 1x resnet34-jpu-psp(576x384), 1x resnet34-\nfpn(1050x700)</p>\n\n<p>attached: resnet34-jpu-psp</p>\n\n<p>if you have suggestion on  how to improve starter kit and how to improve discussion process (e.g. how to select threshold, how to interpret train log)  to help kagglers, please leave your comments here. thanks!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 676099,
          "author_name": "Josh Varty",
          "author_url": "",
          "post_date": "2019-11-19T00:44:43.187000",
          "content": "<p>Thanks for your notes. I haven't had a chance to read all the material here yet. How did you end up merging threshold and min_size for multiple models? Did you average the values? Use the ones from your best model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 676102,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-19T00:50:23.470000",
          "content": "<p>use average results in ensemble. no classifier, no min size. use max pixel probability to threshold against 0.65 to get image level label</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 676109,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2019-11-19T00:57:42.483000",
          "content": "<p>You maybe underestimating your code...the best LB score that we got from your code is <code>0.66520</code>, although it was not the one we selected :(\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F93d269dd7c0012a7d850e24ae90a970c%2Fheng.png?generation=1574124988572045&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 676111,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-19T01:00:50.133000",
          "content": "<p>that is very good news!\nthis is the reason that i share my code (and ideas) before the competition ends.</p>\n\n<p>previous experiences always show that others can take my code (and ideas) to higher levels.</p>\n\n<p>good work and thanks!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 676112,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2019-11-19T01:03:07.743000",
          "content": "<p>Thank you for your amazing contribution. It wouldn't have been possible without your ideas/codes. We will write a post summarizing our experience and learning from your codes/ideas</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 675646,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-18T11:13:30.467000",
      "content": "<p>2 new arvix paper today:</p>\n\n<p><a href=\"https://arxiv.org/pdf/1911.06357.pdf\">https://arxiv.org/pdf/1911.06357.pdf</a>\n\"Give me (un)certainty - An exploration of parameters\nthat affect segmentation uncertainty\"</p>\n\n<p><a href=\"https://arxiv.org/pdf/1911.06667.pdf\">https://arxiv.org/pdf/1911.06667.pdf</a>\n\"CenterMask:Real-Time Anchor-Free Instance Segmentation\"</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 670923,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2019-11-12T03:55:55.943000",
      "content": "<p>Thanks, I'm getting better at reading your code😹 😹 </p>",
      "votes": 4,
      "replies": [
        {
          "id": 671239,
          "author_name": "哈尔的移动城堡",
          "author_url": "",
          "post_date": "2019-11-12T12:18:55.493000",
          "content": "<p>Hey , mate , I  also   read  your   code  before 😂 😂 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 671692,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2019-11-13T04:02:25.780000",
          "content": "<p>Oh, That's my pleasure:)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 672685,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-14T04:18:13.163000",
      "content": "<p><a href=\"/joshvarty\">@joshvarty</a> </p>\n\n<p>Kaggle RSNA challenge just ended today, and it gives me idea on how to use the satellite image  pair. for example instead of predicting single cloud image, predict pair of image together!!</p>\n\n<p>if you have timestamp or location stamp, you can predict them together. </p>\n\n<p>we assume there is some correlation between images taken are different time or location. e.g. if it is fish in latitude A, it is likely to be sugar in latitude B? e.g. fish today flower tomorrow?</p>\n\n<p>```</p>\n\n<p>input = concate[  satellite1 image, satellite2 image ] \n[  satellite1 mask, satellite2 mask] = net(input)</p>\n\n<p>other combinations:\ninput = concate[  first hour, 2nd hour, ....] \ninput = concate[  location 1, location 2, ....] \n```</p>\n\n<p><a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117235#latest-672655\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117235#latest-672655</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117228#latest-672613\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117228#latest-672613</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 672691,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-14T04:23:00.987000",
          "content": "<p>on a side note, TTA is what we usually do. but who knows if there is a better combination than just averaging or temperature sharpening?</p>\n\n<p>```\n mask = net(input)\n mask_tta1= net(input_tta1)\n mask_tta2= net(input_tta2)\n...\nmask_final = fuse_net([mask , mask_tta1, mask_tta2 ...])</p>\n\n<p>```</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 670198,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-11T06:55:30.203000",
      "content": "<p>overfitting</p>\n\n<p>```\nresnet34-unet</p>\n\n<p>constant learning rate=0.01\nbatch_size=16,  iter_accum=2 (100 iter is about 0.3 epoch)</p>\n\n<p>train_dataset : \n    len = 5246</p>\n\n<pre><code>mode    = train\nsplit   = ['by_random1/train_fold_a2_5246.npy']\ncsv     = ['train.csv']\nfolder  = {'image': '1050x700', 'mask': '525x350'}\nnum_image = 5246\n            Fish   neg0, pos0 =  2623  (0.500),   2623  (0.500)\n          Flower   neg1, pos1 =  3002  (0.572),   2244  (0.428)\n          Gravel   neg2, pos2 =  2463  (0.470),   2783  (0.530)\n           Sugar   neg3, pos3 =  1694  (0.323),   3552  (0.677)\n</code></pre>\n\n<p>valid_dataset : \n    len = 300</p>\n\n<pre><code>mode    = train\nsplit   = ['by_random1/valid_fold_a2_300.npy']\ncsv     = ['train.csv']\nfolder  = {'image': '1050x700', 'mask': '525x350'}\nnum_image = 300\n            Fish   neg0, pos0 =   142  (0.473),    158  (0.527)\n          Flower   neg1, pos1 =   179  (0.597),    121  (0.403)\n          Gravel   neg2, pos2 =   144  (0.480),    156  (0.520)\n           Sugar   neg3, pos3 =   101  (0.337),    199  (0.663)\n</code></pre>\n\n<p>```</p>\n\n<p>legend is wrong: </p>\n\n<p>orange=validation kaggle score\ngray=train loss\nblue=validation loss</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe194a1b8fad8eb11ae4aca98ef7ce392%2FSelection_060.png?generation=1573455219875682&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 670230,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-11T07:43:21.553000",
          "content": "<p>the reason why it can overfit at constant high learning rate is as follows:</p>\n\n<ul>\n<li>the cloud problem is a texture classification problem.</li>\n<li>we are looking at the texture at \"some large scale\", we ignore differences lower than \"this scale\".</li>\n<li>but model at good are picking differences at small scale. these are nuisance feature as they do not generalize and can cause overfitting</li>\n</ul>\n\n<p>hence to build a good model, scale is important here i think. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 671712,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-13T04:29:55.850000",
          "content": "<p>one thing to mention here is that resnet34 unet is trained with augmentation. </p>\n\n<p>unlike overfitting resnet18 unet where no augmentation is used. also, with augmentation, resnet18 is less likely to overfit. hence good augmentation also prevent overfitting</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 671723,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-13T04:43:00.513000",
          "content": "<p>to find augmentation that prevent overfits, you can try the below:\n(note this note necessary improve loss as it can lead to underfitting?)</p>\n\n<ol>\n<li><p>for a trained model, perform new test  augmentation. it can be: just add gaussian noise,  just scale, just zero out random part, etc ....</p></li>\n<li><p>measure the effect,change of loss, score, etc of these augmentation. you can sorted them in decreasing \"change of loss\".</p></li>\n<li><p>if the \"change of loss\" is within certain limits, they are probability safe to added for training.</p></li>\n</ol>\n\n<p>rather than \" just add gaussian noise,  just scale, just zero out random part, etc ....\", some adversarial loss should be better?</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 669929,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-10T17:47:07.353000",
      "content": "<p>I wonder if the jigsaw puzzle trick can be used here? The black slant is a good clue to stich back the big image </p>",
      "votes": 1,
      "replies": [
        {
          "id": 669931,
          "author_name": "robga",
          "author_url": "",
          "post_date": "2019-11-10T17:56:10.643000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 669987,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-11-10T20:53:26.310000",
          "content": "<p>My understanding is its not purely 3 locations that were imaged though, right? I thought it was several tiles from that specific region. So then you could jigsaw them together. You know one tile came above the other and you could take the top half of the bottom one and the bottom half of the top one and then you've got a new combination of the two spliced together images</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 670070,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-11T00:33:53.737000",
          "content": "<p>besides, you can fill in the missing black region from one satellite to another </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 671913,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2019-11-13T10:10:00.180000",
      "content": "<p>In your training log, there is a <code>kaggle</code> columns which prints out two values? what do they indicate? Higher values means better?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 671919,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-13T10:16:22.763000",
          "content": "<p>higher value is better.</p>\n\n<p>the first column is the kaggle score you would get on the leader board (classification + segmentation)</p>\n\n<p>the second column is only classification. (segmentation is assumed have perfect dice=1). it represents the upper limit if you try to freeze your classification results and improve segmentation only (e.g. via ensemble, post-processing)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 667136,
      "author_name": "Kenan Ajkunic",
      "author_url": "",
      "post_date": "2019-11-06T21:22:45.910000",
      "content": "<p>Good luck!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 670846,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-12T00:01:22.447000",
      "content": "<p>i wonder what do you get if you train only on images like these:</p>\n\n<p>train image contains only positive region only. maybe less noise?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fb8f393144067e469ee65ef984966951b%2F__results___28_0.png?generation=1573533605694538&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 671235,
          "author_name": "Cold",
          "author_url": "",
          "post_date": "2019-11-12T12:12:59.150000",
          "content": "<p>Well model might learn that the places that are not black are always positive</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 671659,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-13T02:24:03.083000",
          "content": "<p>how about using 2 dataset in stages, e.g. original  and masked-out version. use them only at pretrain, train, finetune ... there are many possible combinations and possibly  different results</p>\n\n<p>how about  training  with  original  + masked-out version, in this case the masked out is a form of augmentation</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 669247,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-09T18:47:27.600000",
      "content": "<p>how i overfit my model\n- use resnet18-unet\n- disable all augmentation\n- train for long epoch .... close 100 ...until loss = NAN</p>\n\n<p>i am surprised that even resnet18 can overfits. below shows how overfitting on train images looks like.\n(it also means we can clean up and create new \"correct label\" from these  train results just before  overfitting occurs.  however, clean label may not benefit the challenge as the test ground truth are also noisy)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F7e8a1370185a67ad02f937c6be034480%2F00031.png?generation=1573325055726255&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F24740b97fd2fcda9bbf457f633018ab9%2F00015.png?generation=1573325032499993&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ffd39a9c3b9181d36b230598620b065a1%2F00000.png?generation=1573324976914170&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 669256,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-09T19:06:27.423000",
          "content": "<p>overfitting reveals the inlier and outlier</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F2352e0b2c0b22c74b5efae2a2b83e77c%2F00023.png?generation=1573326264546928&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 669263,
          "author_name": "Miguel Pinto",
          "author_url": "",
          "post_date": "2019-11-09T19:20:11.547000",
          "content": "<p>A few examples of pairs of images for the same day/region:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2Fcf24962bc79e1b97507a7dc27b26504e%2Fim5.png?generation=1573327168294730&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2F9645162abf3ff12a974c17461b13a156%2Fim0.png?generation=1573327099727112&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2Fd9301cf3e1a0305837ded969cefd95c2%2Fim1.png?generation=1573327112040475&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2F7300c9859825ba029e0c9f2a04bc5bbb%2Fim2.png?generation=1573327123327558&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2Fcb7c5285aa0b26fdad39907d65473f37%2Fim3.png?generation=1573327136181854&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1532879%2F436733a096828758aaa10b0ac4cbbe11%2Fim4.png?generation=1573327149580874&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 669394,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-10T01:40:00.470000",
          "content": "<p>thanks. this explain</p>\n\n<ol>\n<li>the presence of label noise\n2.why label smoothing is not required\n3.why disable augmentation overfit .with almost zero loss... the network may have learned the black band</li>\n</ol>\n\n<p>i wonder if an annotator only label one satellite or both. if a annotator only deal with single satellite, training individual classifier may work?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 669990,
          "author_name": "Josh Varty",
          "author_url": "",
          "post_date": "2019-11-10T21:05:34.830000",
          "content": "<blockquote>\n  <p>i wonder if an annotator only label one satellite or both. if a annotator only deal with single satellite, training individual classifier may work?</p>\n</blockquote>\n\n<p>Annotators label images from both satellites. You can see how they annotated them here:</p>\n\n<p><a href=\"https://www.zooniverse.org/projects/raspstephan/sugar-flower-fish-or-gravel/classify\">https://www.zooniverse.org/projects/raspstephan/sugar-flower-fish-or-gravel/classify</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 670075,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-11T00:57:20.063000",
          "content": "<p>@Miguel Pinto</p>\n\n<p>how you get the date and domain of the images? is it provided?</p>\n\n<p>have you download the same images from the website for the website, you may have larger image and more information (like cloud precipitation) for that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 670382,
          "author_name": "Miguel Pinto",
          "author_url": "",
          "post_date": "2019-11-11T12:04:04.003000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> I downloaded all images from worldview (identical to the competition data) for the regions and period described in the paper about the data. Then I used perceptual hashing to match downloaded images to the identical ones provided in the competition data. I didn't download full-size images from worldview, 350 x 525 is faster to download and is enough to find the pairs. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 668561,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-08T15:05:25.203000",
      "content": "<p>some submission statistics</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F509471fed377077574518f8c124b140a%2FSelection_058.png?generation=1573225729280934&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 669918,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-10T17:16:37.710000",
          "content": "<p>how lb 0.6671 looks like:( there is a mistake in the spreadsheet. pos recall should be 0.64647, )</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ff437c2c01db64abc8dce18b34c7b758a%2FSelection_057.png?generation=1573406195419154&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 669937,
          "author_name": "robga",
          "author_url": "",
          "post_date": "2019-11-10T18:14:16.643000",
          "content": "<blockquote>\n  <p>how lb 0.6671 looks like:</p>\n</blockquote>\n\n<p>Publishing a silver level result with 8 days to go sounds more like a “finisher-kit” than a “starter-kit” ?</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 669958,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "2019-11-10T19:08:26.697000",
          "content": "<p>His single best model score is 0.657. There are public kernels which score same range. So I wouldn't say this is a \"finisher-kit\". What impressive is that he can get 0.6671 from such lower score models...</p>",
          "votes": -5,
          "replies": []
        },
        {
          "id": 669988,
          "author_name": "starsnew",
          "author_url": "",
          "post_date": "2019-11-10T20:57:57.140000",
          "content": "<p>it's not \"starter-kit\" ...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 670013,
          "author_name": "IgorMuniz",
          "author_url": "",
          "post_date": "2019-11-10T22:21:28.200000",
          "content": "<p>what's the point here? \ntips are cool\nstarter codes are cool\na complete silver medal solution is not cool</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 670066,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-11T00:23:12.210000",
          "content": "<p>\"Publishing a silver level result with 8 days to go sounds more like a “finisher-kit” than a “starter-kit” ?\"  </p>\n\n<p>you can't download my code and run as it is.  </p>\n\n<p>some files are intentionally left missing, but you can create them yourself, e.g. split file. in particular, the training iteration don't automatically stop. you will have to find the hyper-parameters yourself. </p>\n\n<p>What i would like to provide is a possible direction for new kagglers to follow and not a \"click and run solution\"</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 670134,
          "author_name": "哈尔的移动城堡",
          "author_url": "",
          "post_date": "2019-11-11T04:03:39.140000",
          "content": "<p>I   think  the  “finisher-kit”   here  means  that  you  will  get   higer  score .</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 670155,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-11T05:09:44.153000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa77f2aff2081e2851ab566e81f818f31%2FSelection_056.png?generation=1573549207608697&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 670234,
          "author_name": "Endi Niu",
          "author_url": "",
          "post_date": "2019-11-11T07:49:11.387000",
          "content": "<p>take a look at 'steel' competition private lb, lots of silver medal on public lb dropped far away, I image the same thing happen here</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 670241,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-11T07:52:42.470000",
          "content": "<p>i think the shakeup will be much smaller for cloud here.</p>\n\n<p>there is a lot of difference when you can see and cannot see both public+private test data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 670248,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-11-11T08:05:34.993000",
          "content": "<p>Am I missing something? What is the dice column in your image? Not sure where that is coming from. Seems very high. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 670251,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-11T08:11:00.723000",
          "content": "<p>dice of the positive classified image. it is value of pos_dice in the formula below:</p>\n\n<p>```\nkaggle_score = avergae(\n     tnr *num_true_negative _image +\\\n     tpr *num_true_positive_image*pos_dice +\\\n    (1-tnr)*num_true_negative_image*neg_dice +\\\n)</p>\n\n<p>```</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 670695,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-11T18:46:16.737000",
          "content": "<p>bestfitting score is a good guide. he is unusually unshaken and he don't overfits.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 670766,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2019-11-11T20:14:12.603000",
          "content": "<blockquote>\n  <p>bestfitting score is a good guide. he is unusually unshaken and he don't overfits.</p>\n</blockquote>\n\n<p>This sounds like top teams are overfitting and we might expect shakeup again?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 670963,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-12T05:00:39.923000",
          "content": "<p>@Bibek</p>\n\n<p>there will not be major  shakeup. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 672162,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-13T15:15:09.460000",
      "content": "<p><a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/70421\">https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/70421</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc800e5ea0c5d05dc4d44e4a9a9ad114f%2FSelection_051.png?generation=1573658106274069&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 672165,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-13T15:18:00.787000",
          "content": "<p>this is how they handle intersection of box (different from our cases)</p>\n\n<p>\"the intersection of boxes was used as opposed to the average. We tried to mimic this intersection process by taking the average of multiple intersections of boxes across models. This gave ~10-15% (!) improvement on stage 1 public LB. As an alternative to this, we simply resized the boxes by multiplying length/width by a fraction. We found 87.5% for each was a good reduction and worked a bit better than doing the intersection. It was also much easier to implement. We discovered this early on, and didn't submit anything without resizing after that\"</p>\n\n<p>... should we expand our box in ground truth ???</p>\n\n<p>it my code and experiment, it actually used 0.3 for threshold which works better. this is essentially a mask dilation </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 668913,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-09T04:43:07.590000",
      "content": "<p>when zero is not actually zero?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F02f6f6dc79fd8f87852ec5100df6d89d%2FSelection_061.png?generation=1573274564563490&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 669059,
          "author_name": "Miguel Pinto",
          "author_url": "",
          "post_date": "2019-11-09T12:09:17.997000",
          "content": "<p>According to the data section, the labels are actually the union of annotators and not the intersection. </p>\n\n<blockquote>\n  <p>\"Ground truth was determined by the union of the areas marked by all labelers for that image, after removing any black band area from the areas.\" </p>\n</blockquote>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 669067,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-09T12:31:14.610000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 669221,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-09T17:53:58.933000",
          "content": "<p>@Miguel Pinto</p>\n\n<p>you are right. hence it should be \"one is actually not one?\".</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 669299,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-11-09T20:53:54.807000",
          "content": "<p>Yes. Given any mask that is not shaped as a rectangle, you can split it into the separate annotators' masks and make more training data. Or as you say, you can relabel the 1's to be the correct values.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 665939,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-11-05T14:55:46.177000",
      "content": "<p>at first look, the images are randomly reflected or rotated (maybe to prevent the jigsaw puzzle assembling). you can realign all images so that all the \"black slant\" are in the same direction.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 665943,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2019-11-05T15:00:16.787000",
          "content": "<p>any helpful resources that can demonstrate how to design network with higher image sizes to predict small size masks? thanks in advance <a href=\"/hengck23\">@hengck23</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665986,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-05T16:00:32.727000",
          "content": "<p>simply:</p>\n\n<p>```\ninput --&gt; encode(scale by half)  --&gt; encode .....  --&gt;decode (scale by two)  ---&gt; decode</p>\n\n<p>if num of decoder  is less than num of encoder, mask size will be small \n```</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 666161,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-11-05T20:06:40.883000",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> You can do it with Qubvel's segmentation models. Each backbone has five 2x encoders. So the original image gets reduced by a factor of <code>2^5 = 32</code>. Now regarding decoders, the default is <code>decoder_filters = [256, 128, 64, 32, 16]</code> where <code>len(decoder_filters)</code> is the number of decoder filters and the numbers are how many convolutional maps each has. This says that there are five 2x decoders, so the reduced image becomes a mask similar to original size. If instead you use:</p>\n\n<pre><code>model = Unet('resnet18', decoder_filters=[256, 128, 64])\n</code></pre>\n\n<p>Then you only have three decoder filters. So if original is <code>640x960</code>, it gets encoded into <code>20x30</code>, and it only gets decoded to <code>160x240</code>. Hence the network outputs masks of size <code>160x240</code> and accepts images of size <code>640x960</code>.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 666358,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2019-11-06T03:06:32.447000",
          "content": "<p>thank you a lot for this nice explanation <a href=\"/cdeotte\">@cdeotte</a>  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 667957,
          "author_name": "Josh Varty",
          "author_url": "",
          "post_date": "2019-11-07T21:01:04.103000",
          "content": "<p>Edit: Sorry, I misunderstood.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1023955,
      "author_name": "Hari Mohan Rai",
      "author_url": "",
      "post_date": "2020-09-23T14:49:01.277000",
      "content": "<p>can any one will explain about this table please?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4444375%2F394d70364686cd678fbfa129736bd49f%2Fdicevsthreshold.png?generation=1600872408682015&amp;alt=media\" alt=\"\"></p>\n<p>any documents in support of this please?  i am trying very hard to find out but didn't find any documents. please someone help me in this regard.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 683084,
      "author_name": "byr_syx",
      "author_url": "",
      "post_date": "2019-11-28T04:22:51.257000",
      "content": "<p>my net </p>\n\n<p>class DinkNet101(nn.Module):\n    def <strong>init</strong>(self, num_classes=4):\n        super(DinkNet101, self).<strong>init</strong>()</p>\n\n<pre><code>    filters = [256, 512, 1024, 2048]\n    resnet = models.resnet101(pretrained=True)\n    self.firstconv = resnet.conv1\n    self.firstbn = resnet.bn1\n    self.firstrelu = resnet.relu\n    self.firstmaxpool = resnet.maxpool\n    self.encoder1 = resnet.layer1\n    self.encoder2 = resnet.layer2\n    self.encoder3 = resnet.layer3\n    self.encoder4 = resnet.layer4\n\n    self.dblock = Dblock_more_dilate(2048)\n\n    self.decoder4 = DecoderBlock(filters[3], filters[2])\n    self.decoder3 = DecoderBlock(filters[2], filters[1])\n    self.decoder2 = DecoderBlock(filters[1], filters[0])\n    self.decoder1 = DecoderBlock(filters[0], filters[0])\n\n    self.finaldeconv1 = nn.ConvTranspose2d(filters[0], 32, 4, 2, 1)\n    self.finalrelu1 = nonlinearity\n    self.finalconv2 = nn.Conv2d(32, 32, 3, padding=1)\n    self.finalrelu2 = nonlinearity\n    self.finalconv3 = nn.Conv2d(32, num_classes, 3, padding=1)\n\ndef forward(self, x):\n    # Encoder\n    x = self.firstconv(x)\n    x = self.firstbn(x)\n    x = self.firstrelu(x)\n    x = self.firstmaxpool(x)\n    e1 = self.encoder1(x)\n    e2 = self.encoder2(e1)\n    e3 = self.encoder3(e2)\n    e4 = self.encoder4(e3)\n\n    # Center\n    e4 = self.dblock(e4)\n\n    # Decoder\n    d4 = self.decoder4(e4) + e3\n    d3 = self.decoder3(d4) + e2\n    d2 = self.decoder2(d3) + e1\n    d1 = self.decoder1(d2)\n    out = self.finaldeconv1(d1)\n    out = self.finalrelu1(out)\n    out = self.finalconv2(out)\n    out = self.finalrelu2(out)\n    out = self.finalconv3(out)\n\n    return F.sigmoid(out)\n</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 675090,
      "author_name": "Qineng Cao",
      "author_url": "",
      "post_date": "2019-11-17T15:35:59.207000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 675096,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-17T15:42:19.850000",
          "content": "<p>don't think it is a bug.\n\"segmentation only\" uses threshold_label = -1</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 675107,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-17T15:55:32.600000",
          "content": "<p>i just check and i do have the same results as you if i use size threshold (i never use it before)\nlet me check if it is a bug\n<code>\nthreshold_size  = [ 30000, 30000, 30000, 30000]\n</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 675133,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-17T16:42:47.787000",
          "content": "<p>i check that the code should have no bug (but i am not 100% sure).\ni made a submission. using large area threshold like 30000 don't give good results</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 674993,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2019-11-17T12:06:27.133000",
      "content": "<p>I trained a model(different from yours) using your codes/pipeline and got a validation score of <code>0.659xx</code> but when I submit it, I get around <code>0.58xx</code>on LB. I have attached the screenshot...don't know what could be wrong?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F9cabc472aaa647c0c3d6e94e1f58ac4c%2Fval_score.bmp?generation=1573992298447004&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 674995,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2019-11-17T12:09:49.807000",
          "content": "<p>This is the training log...I used 0.658 model</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 675003,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-17T12:39:38.003000",
          "content": "<p>i suspect threshold problem.</p>\n\n<p>when you submit in 'test' mode, you should see the results of</p>\n\n<p><code>\n        print('initial_checkpoint=%s'%initial_checkpoint)\n        text = summarise_submission_csv(df)\n        log.write('\\n')\n        log.write('%s'%(text))\n</code></p>\n\n<p>you check the number of predicted masks, etc similar to what i have below</p>\n\n<p>```</p>\n\n<p>compare with LB probing ... \n        num_image =  3698(3698) </p>\n\n<pre><code>    pos0 =  1303(1864)  xxx\n    pos1 =  1410(1508)  xxx\n    pos2 =  1084(1982)  xxx\n    pos3 =  2305(2382)  xxx\n\n    neg0 =  xxx(1834)  1.306\n    neg1 =  xxx(1940)  1.179\n    neg2 =  xxx(2638)  0.991\n</code></pre>\n\n<h2>        neg3 =  xxx(2017)  0.691</h2>\n\n<pre><code>    all_zero =   xxx (?)\n</code></pre>\n\n<p>```\nadjust threshold until you get good number of predicted mask in the LB test data</p>\n\n<p>you can also see:\n<a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/117310#latest-674964\">https://www.kaggle.com/c/understanding_cloud_organization/discussion/117310#latest-674964</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 674509,
      "author_name": "makogarei",
      "author_url": "",
      "post_date": "2019-11-16T15:53:35.030000",
      "content": "<p>'compute_metric_label' is not defined ???</p>",
      "votes": 0,
      "replies": [
        {
          "id": 674511,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-16T15:55:29.467000",
          "content": "<p>kaggle.py?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 674515,
          "author_name": "makogarei",
          "author_url": "",
          "post_date": "2019-11-16T16:04:29.963000",
          "content": "<p>no\nsubmmit.py</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674574,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-16T18:14:44.900000",
          "content": "<p>The function should be found in kaggle.py. please the latest version of the code</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 674345,
      "author_name": "Qineng Cao",
      "author_url": "",
      "post_date": "2019-11-16T10:02:52.960000",
      "content": "<p>hi heng, do you feel strange about andrew's threshold_mask and your threshold_mask ? the threshold_mask &gt; 0.5 always have good result with andrew's model(i trained with FPN resnet34), but your model always need to set the threshold_mask &lt; 0.5. and you can see andrew's experimental result at <a href=\"https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools\">here</a>.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 674350,
          "author_name": "Qineng Cao",
          "author_url": "",
          "post_date": "2019-11-16T10:07:28.900000",
          "content": "<p>and you all use sigmod with model's output.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674353,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-16T10:16:24.223000",
          "content": "<p>i also use  threshold &gt;0.5 for image label. \nthreshold can be less than 0.5 for pixel label. this is to encourage dilation.</p>\n\n<p>did you train with loss_mask only or ( loss_label + loss_mask )?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 674415,
          "author_name": "Qineng Cao",
          "author_url": "",
          "post_date": "2019-11-16T13:18:12.767000",
          "content": "<p>“i also use threshold &gt;0.5 for image label.\nthreshold can be less than 0.5 for pixel label. this is to encourage dilation.”</p>\n\n<p>log of andrew's model(i trained with FPN resnet34) is lost, i will try it again. and if we do not consider the difference between Unet and FPN,  i think your submit mode of segmentation only is equal to Andrew's submit mode.  the threshold mask &gt; 0.5 with andrew's model will be better (see picture), and the threshold mask &gt; 0.5 with your model will be worse (see log).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F08a9a48a429996b20cd7872e839ebd7d%2Fs.png?generation=1573910110677859&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674417,
          "author_name": "Qineng Cao",
          "author_url": "",
          "post_date": "2019-11-16T13:19:48.740000",
          "content": "<p>\"did you train with loss_mask only or ( loss_label + loss_mask )?\"</p>\n\n<p>andrew' model is loss_mask only,  your model is ((loss_mask )/iter_accum).backward().</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674436,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-16T14:04:09.263000",
          "content": "<p>do note that my model do not use size post processing. It uses max pixel probability instead.</p>\n\n<p>\"i think your submit mode of segmentation only is equal to Andrew's submit mode\"\ni think that too. that is why i mentioned before that the notebook kernel are actually good enough for LB 0.67 if they are properly trained, post processed and ensembled.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 674444,
          "author_name": "Qineng Cao",
          "author_url": "",
          "post_date": "2019-11-16T14:18:54.780000",
          "content": "<p>thank u, heng, wait for me to traine andrew's model(fnp resnet) again, maybe after finished competion, because i am working on your code now.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674449,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-16T14:26:23.973000",
          "content": "<p>do also note that his fpn is not exactly the same as mine. the decoder head is differ slightly in number of convs, etc. but i think that does affect results. but do note that my fpn is tuned for size 1050x700 and not other size.</p>\n\n<p>i have tried many models, different sizes, etc in the past week. i carried about 80 experiments. most of my models will have 0.65 in validation, +/-0.05. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 674468,
          "author_name": "Qineng Cao",
          "author_url": "",
          "post_date": "2019-11-16T14:44:54.140000",
          "content": "<p>thank u.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674710,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-16T23:54:24.060000",
          "content": "<p>after i check andrew's code and results in detail, indeed our results are close but using very different parameters.</p>\n\n<p>you may want to check the correlation of the results of the two models. if there is diversity, there is a good chance for ensemble.</p>\n\n<p>but it is tricky for ensemble since both model use different parameter.</p>\n\n<p>```\neg \nmodel1 needs param1\nmodel2 needs param2</p>\n\n<p>but parm1+2 may/may not be good for model1+2?\nor param1 is better?\nor param2 is better?\n```</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 675093,
          "author_name": "Qineng Cao",
          "author_url": "",
          "post_date": "2019-11-17T15:39:27.223000",
          "content": "<p>thank you for your reply, i will do experiment about different between the two models after competion.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 674098,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-15T22:15:53.193000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 673776,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-15T13:26:22.197000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 673785,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T13:32:12.887000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 674140,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T23:41:22.063000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674197,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-16T01:52:34.757000",
          "content": "",
          "votes": 5,
          "replies": []
        },
        {
          "id": 674498,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-16T15:31:00.060000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 673752,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-15T12:45:48.167000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 673768,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T13:21:07.763000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 673792,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T13:47:15.400000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 673795,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T13:53:31.550000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 673799,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T14:00:01.357000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 673806,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T14:07:57.720000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 673835,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T14:40:20.730000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 673847,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T14:53:53.183000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 673853,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T15:01:46.120000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 673856,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T15:06:53.413000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 673857,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T15:06:54.590000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 673858,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T15:10:18.140000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 673867,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T15:21:59.620000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 673869,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T15:22:22.703000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 673870,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T15:24:49.950000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 673875,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T15:32:49.503000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 673230,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-14T17:35:29.107000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 673233,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-14T17:38:48.837000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 673372,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-14T22:38:18.980000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 673379,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-14T22:48:42.960000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 673388,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-14T23:07:32.947000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 673831,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-15T14:36:43.010000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 674502,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-16T15:34:39.037000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 672859,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-14T08:30:47.040000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 672861,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-14T08:31:31.040000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 672864,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-14T08:33:30.413000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 672953,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-14T10:34:25.620000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 672189,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-13T15:44:54.663000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 672197,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-13T15:54:26.157000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 672235,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-13T17:05:54.290000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 672439,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-13T21:57:20.717000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 672453,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-13T23:01:34.703000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 672661,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-14T03:52:43.470000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 672680,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-14T04:13:00.043000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 672705,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-14T04:33:01.643000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 672719,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-14T04:52:14.877000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 672737,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-14T05:16:15.200000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 672794,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-14T06:55:11.570000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 672085,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-13T13:58:05.870000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 671716,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-13T04:37:53.947000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 671757,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-13T05:55:36.693000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 671705,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-13T04:17:50.553000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 671709,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-13T04:24:34.800000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 671711,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-13T04:27:31.923000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 671743,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-13T05:26:23.840000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 671089,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-12T08:44:21.787000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 671102,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-12T08:54:31.227000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 671112,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-12T09:07:54.140000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 670403,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-11T12:34:00.813000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 670135,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-11T04:04:31.597000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 670172,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-11T05:53:44.490000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 670806,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-11T22:08:28.893000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 670919,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-12T03:38:18.723000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 669946,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-10T18:26:49.187000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 670062,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-11T00:13:59.827000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 670099,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-11T02:25:06.720000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 669222,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-09T17:57:28.847000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 669230,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-09T18:15:37.737000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 669238,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-09T18:25:05.053000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 669242,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-09T18:37:16.983000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 669250,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-09T18:50:16.790000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 669251,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-09T18:53:41.747000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 669252,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-09T18:53:54.793000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 669254,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-09T18:57:47.700000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 669255,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-09T18:59:02.077000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 669262,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-09T19:18:58.847000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 669264,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-09T19:26:16.503000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 669928,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-10T17:42:06.360000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 670079,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-11T01:15:42.287000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 670670,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-11T17:58:53.557000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 668579,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-08T15:28:29.743000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 666390,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-06T04:01:28.750000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 674438,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-16T14:05:11.787000",
      "content": "",
      "votes": -1,
      "replies": [
        {
          "id": 674445,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-16T14:19:13.650000",
          "content": "",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 670673,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-11T18:05:59.583000",
      "content": "",
      "votes": -8,
      "replies": [
        {
          "id": 670916,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-12T03:33:25.833000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 670985,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-12T05:41:07.927000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 670987,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-12T05:44:14.777000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "665639": "what this is not *NOT*\n- a click and run solution. do not expect just to run the scripts as it is and get the results\n\nwhat this is \n- the code base does indeed produce the results as mentioned, however, you have to find your hyper parameters and made some modification according to the discussion. you probably have to fill in some simple code yourself like making train splits, etc\n\nwhy another starter kit?\n- it is sometimes difficult to discuss without reference code.\n- training log is important to analyse model behavior. it is probably the most important reference for beginner to learn how to train a model. the provision of such log is to compare your version with the reference version to see how change of hyper-parameters affect results.\n- another tool is the visualization  of results.\n- this starter kit provides good logging and visualization. reference training log is also provided.\n\n\n**7 days is enough to make good results for this challenge due to the small dataset and noise in the annotation. have fun, work hard and good luck!**\n\nplease see\n\nhttps://drive.google.com/open?id=1FTAjX-rUDbvb2G8wWyC0YEvhPd5gN-B4\n\n\n----\n\npre-relase version  (2019-11-05):\n - dirty code, not all function are finished coded yet\n - but you can make  a submission of LB 0.644 using resnet34-fpn and trained with 3 to  4 hrs\n\n\n----\n\nversion.1: (2019-11-10)\n - LB 0.6586 (resnet34-fpn x3 fold, single fold lb 0.656)\n\nversion.1: (2019-11-10a)\n- added partial code for resnet34-unet(single fold lb 0.654 @ label threshold 0.6, 0.657 @ 0.7)\n- ensemble of 1xresnet34-unet + 3xfpn : lb 0.6671\n\nnote: \n  - you can improve the code by addition different input size in ensemble. Also add double flip = flip lr + up =rotate 180 as additional TTA. you should be about to get lb 0.670~",
    "668036": "here is one data leak!\n\n- each image has at least one label .... you can use adaptive threshold to ensure each image has at least one type of cloud or you can design a loss function for that (e.g. ranking loss)",
    "669537": "https://events.ecmwf.int/event/118/contributions/538/attachments/139/244/OCBWF-Bony.pdf\nhttps://rmets.onlinelibrary.wiley.com/doi/epdf/10.1002/qj.3662\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe371d138ff40a6d50a9d09d41343d903%2FSelection_053.png?generation=1573366766042602&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd4d6206c17b3abb64bb2fa195e4ecf41%2FSelection_051.png?generation=1573366770599581&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F8527544c55b481b58e7a635cdffd1ed8%2FSelection_052.png?generation=1573366771716980&amp;alt=media)\n",
    "671393": "0.6621 single model!\n\ni tried several segmentation head: unet decoder head, fpn head, ASPP head. I think ASPP is the best.\n\n```\n\nclass Net(nn.Module):\n    def load_pretrain(self, skip=['logit.'], is_print=True):\n        load_pretrain(self, skip, pretrain_file=PRETRAIN_FILE, conversion=CONVERSION, is_print=is_print)\n\n    def __init__(self, num_class=4):\n        super(Net, self).__init__()\n\n        e = ResNet34()\n        self.block0 = e.block0\n        self.block1 = e.block1\n        self.block2 = e.block2\n        self.block3 = e.block3\n        self.block4 = e.block4\n        e = None  #dropped\n\n        self.jpu = JointPyramidUpsample([512,256,128],128)\n        #self.aspp = ASPP(512, 128, rate=[6,12,18], dropout_rate=0.1)\n        self.aspp = ASPP(512, 128, rate=[4,8,12], dropout_rate=0.1)\n        self.logit = nn.Conv2d(128,num_class,kernel_size=1)\n\n\n    def forward(self, x):\n        batch_size,C,H,W = x.shape\n\n        x0 = self.block0(x)\n        x1 = self.block1(x0)\n        x2 = self.block2(x1)\n        x3 = self.block3(x2)\n        x4 = self.block4(x3)\n\n        x = self.jpu([x4,x3,x2])\n        x = self.aspp(x)\n        logit = self.logit(x)\n\n        #---\n        probability_mask  = torch.sigmoid(logit)\n        probability_label = F.adaptive_max_pool2d(probability_mask,1).view(batch_size,-1)\n        return probability_label, probability_mask\n\n\n```",
    "670024": "The competition is going to the final week. Please stop sharing. It's not nice to see solutions keeping posted like this.",
    "665655": "Please, be carefull )",
    "670074": "lesson from the steel competition is that  results is going to be sensitive to the image label threshold.\n\none way to deal with this is to select proper high threshold, (like 0.70).\n\nanother way is to modify the loss e.g loss = weight\\_pos*loss\\_pos +  weight\\_neg*loss\\_neg. By adjusting the weights, you can shift the threshold back to 0.50.\n\nto visualize the effects on classification, you can observe the roc curve of the weighted and un-weighted loss model. indicate tpr, fpr and threshold on your roc curve.\n\nfor segmentation, just draw the probability heatmap in range (0,1). you can see the weighted version will be dilated (or eroded) version of the unweighted one.\n",
    "666095": "daily tracking results:\n\nentry date : 2019-11-05\n\n```\nday-1, 2019-11-05: LB 0.644   (non-tuned resnet34-fpn)  \nday-3, 2019-11-08: LB 0.6586  (resnet34-fpn x3 fold) . for single fold : lb 0.656\nday-5, 2019-11-10: LB 0.6577  (resnet34-unet single). \n             ensemble of 1 x unet + 3 x fpn : lb 0.667  (finetune threshold, etc)\n\n\nday-6, 2019-11-11: LB 0.6709 (added a few misc models to ensemble, e.g. different input size, jpu+aspp (joint-pyramid-upsampling))\n\n\nday-8, 2019-11-13: LB 0.6713 (corrected a bug. wrongly resize image to widthxheight instead of heightxwidth). perform threshold search on LB to use up free submission slots, try 0.60 0.65, 0.67 as image label threshold. the lb score can ranges from  0.6664,0.6707, 0.6705. nevethelss, high threshold is preferred\n\nday-10 2019-11-14: LB 0.6721 (corrected the resize bug again! seems that this bug occurs in different part of my code). add 2 more models to my previous ensemble. i am still using the same models except for for more fold\n\n\n```",
    "671545": "You still are sharing top tips and even partial code for a near gold solutions. Can you please just stop and wait until it ends and share, I’m glad to discuss after that, not now. ",
    "666792": "my plan  (to be updated as competition proceeds ...)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbd4150e743b7ac4b6eb9a1da4c03229d%2FSelection_053.png?generation=1573049004668282&amp;alt=media)\n\n",
    "676064": "closing remarks:\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F851c2335980981005516610ceb5348dc%2FSelection_061.png?generation=1574122453993841&amp;alt=media)\n\n\nbest private score this code can get is LB 0.66441\nmethod: \n1. just train with all samples, use 1x  resnet34-unet (384x256), 1x resnet34-jpu-psp(576x384), 1x resnet34-\nfpn(1050x700)\n\nattached: resnet34-jpu-psp\n\n\nif you have suggestion on  how to improve starter kit and how to improve discussion process (e.g. how to select threshold, how to interpret train log)  to help kagglers, please leave your comments here. thanks!",
    "675646": "2 new arvix paper today:\n\nhttps://arxiv.org/pdf/1911.06357.pdf\n\"Give me (un)certainty - An exploration of parameters\nthat affect segmentation uncertainty\"\n\nhttps://arxiv.org/pdf/1911.06667.pdf\n\"CenterMask:Real-Time Anchor-Free Instance Segmentation\"",
    "670923": "Thanks, I'm getting better at reading your code😹 😹 ",
    "672685": "@joshvarty \n\nKaggle RSNA challenge just ended today, and it gives me idea on how to use the satellite image  pair. for example instead of predicting single cloud image, predict pair of image together!!\n\n\nif you have timestamp or location stamp, you can predict them together. \n\nwe assume there is some correlation between images taken are different time or location. e.g. if it is fish in latitude A, it is likely to be sugar in latitude B? e.g. fish today flower tomorrow?\n\n```\n\ninput = concate[  satellite1 image, satellite2 image ] \n[  satellite1 mask, satellite2 mask] = net(input)\n\nother combinations:\ninput = concate[  first hour, 2nd hour, ....] \ninput = concate[  location 1, location 2, ....] \n```\n\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117235#latest-672655\n\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117228#latest-672613",
    "670198": "overfitting\n\n```\nresnet34-unet\n\nconstant learning rate=0.01\nbatch_size=16,  iter_accum=2 (100 iter is about 0.3 epoch)\n\n\ntrain_dataset : \n\tlen = 5246\n\n\tmode    = train\n\tsplit   = ['by_random1/train_fold_a2_5246.npy']\n\tcsv     = ['train.csv']\n\tfolder  = {'image': '1050x700', 'mask': '525x350'}\n\tnum_image = 5246\n\t            Fish   neg0, pos0 =  2623  (0.500),   2623  (0.500)\n\t          Flower   neg1, pos1 =  3002  (0.572),   2244  (0.428)\n\t          Gravel   neg2, pos2 =  2463  (0.470),   2783  (0.530)\n\t           Sugar   neg3, pos3 =  1694  (0.323),   3552  (0.677)\n\nvalid_dataset : \n\tlen = 300\n\n\tmode    = train\n\tsplit   = ['by_random1/valid_fold_a2_300.npy']\n\tcsv     = ['train.csv']\n\tfolder  = {'image': '1050x700', 'mask': '525x350'}\n\tnum_image = 300\n\t            Fish   neg0, pos0 =   142  (0.473),    158  (0.527)\n\t          Flower   neg1, pos1 =   179  (0.597),    121  (0.403)\n\t          Gravel   neg2, pos2 =   144  (0.480),    156  (0.520)\n\t           Sugar   neg3, pos3 =   101  (0.337),    199  (0.663)\n```\n\nlegend is wrong: \n\norange=validation kaggle score\ngray=train loss\nblue=validation loss\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe194a1b8fad8eb11ae4aca98ef7ce392%2FSelection_060.png?generation=1573455219875682&amp;alt=media)\n",
    "669929": "I wonder if the jigsaw puzzle trick can be used here? The black slant is a good clue to stich back the big image ",
    "671913": "In your training log, there is a `kaggle` columns which prints out two values? what do they indicate? Higher values means better?",
    "667136": "Good luck!",
    "670846": "i wonder what do you get if you train only on images like these:\n\ntrain image contains only positive region only. maybe less noise?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fb8f393144067e469ee65ef984966951b%2F__results___28_0.png?generation=1573533605694538&amp;alt=media)\n",
    "669247": "how i overfit my model\n- use resnet18-unet\n- disable all augmentation\n- train for long epoch .... close 100 ...until loss = NAN\n\ni am surprised that even resnet18 can overfits. below shows how overfitting on train images looks like.\n(it also means we can clean up and create new \"correct label\" from these  train results just before  overfitting occurs.  however, clean label may not benefit the challenge as the test ground truth are also noisy)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F7e8a1370185a67ad02f937c6be034480%2F00031.png?generation=1573325055726255&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F24740b97fd2fcda9bbf457f633018ab9%2F00015.png?generation=1573325032499993&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Ffd39a9c3b9181d36b230598620b065a1%2F00000.png?generation=1573324976914170&amp;alt=media)\n",
    "668561": "some submission statistics\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F509471fed377077574518f8c124b140a%2FSelection_058.png?generation=1573225729280934&amp;alt=media)\n",
    "672162": "https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/70421\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc800e5ea0c5d05dc4d44e4a9a9ad114f%2FSelection_051.png?generation=1573658106274069&amp;alt=media)\n",
    "668913": "when zero is not actually zero?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F02f6f6dc79fd8f87852ec5100df6d89d%2FSelection_061.png?generation=1573274564563490&amp;alt=media)\n",
    "665939": "at first look, the images are randomly reflected or rotated (maybe to prevent the jigsaw puzzle assembling). you can realign all images so that all the \"black slant\" are in the same direction.",
    "1023955": "can any one will explain about this table please?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4444375%2F394d70364686cd678fbfa129736bd49f%2Fdicevsthreshold.png?generation=1600872408682015&alt=media)\n\nany documents in support of this please?  i am trying very hard to find out but didn't find any documents. please someone help me in this regard.",
    "683084": "my net \n\nclass DinkNet101(nn.Module):\n    def __init__(self, num_classes=4):\n        super(DinkNet101, self).__init__()\n\n        filters = [256, 512, 1024, 2048]\n        resnet = models.resnet101(pretrained=True)\n        self.firstconv = resnet.conv1\n        self.firstbn = resnet.bn1\n        self.firstrelu = resnet.relu\n        self.firstmaxpool = resnet.maxpool\n        self.encoder1 = resnet.layer1\n        self.encoder2 = resnet.layer2\n        self.encoder3 = resnet.layer3\n        self.encoder4 = resnet.layer4\n        \n        self.dblock = Dblock_more_dilate(2048)\n\n        self.decoder4 = DecoderBlock(filters[3], filters[2])\n        self.decoder3 = DecoderBlock(filters[2], filters[1])\n        self.decoder2 = DecoderBlock(filters[1], filters[0])\n        self.decoder1 = DecoderBlock(filters[0], filters[0])\n\n        self.finaldeconv1 = nn.ConvTranspose2d(filters[0], 32, 4, 2, 1)\n        self.finalrelu1 = nonlinearity\n        self.finalconv2 = nn.Conv2d(32, 32, 3, padding=1)\n        self.finalrelu2 = nonlinearity\n        self.finalconv3 = nn.Conv2d(32, num_classes, 3, padding=1)\n\n    def forward(self, x):\n        # Encoder\n        x = self.firstconv(x)\n        x = self.firstbn(x)\n        x = self.firstrelu(x)\n        x = self.firstmaxpool(x)\n        e1 = self.encoder1(x)\n        e2 = self.encoder2(e1)\n        e3 = self.encoder3(e2)\n        e4 = self.encoder4(e3)\n        \n        # Center\n        e4 = self.dblock(e4)\n\n        # Decoder\n        d4 = self.decoder4(e4) + e3\n        d3 = self.decoder3(d4) + e2\n        d2 = self.decoder2(d3) + e1\n        d1 = self.decoder1(d2)\n        out = self.finaldeconv1(d1)\n        out = self.finalrelu1(out)\n        out = self.finalconv2(out)\n        out = self.finalrelu2(out)\n        out = self.finalconv3(out)\n\n        return F.sigmoid(out)",
    "675090": "\n",
    "674993": "I trained a model(different from yours) using your codes/pipeline and got a validation score of `0.659xx` but when I submit it, I get around `0.58xx`on LB. I have attached the screenshot...don't know what could be wrong?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F9cabc472aaa647c0c3d6e94e1f58ac4c%2Fval_score.bmp?generation=1573992298447004&amp;alt=media)\n",
    "674509": "'compute_metric_label' is not defined ???\n",
    "674345": "hi heng, do you feel strange about andrew's threshold_mask and your threshold_mask ? the threshold_mask &gt; 0.5 always have good result with andrew's model(i trained with FPN resnet34), but your model always need to set the threshold_mask &lt; 0.5. and you can see andrew's experimental result at [here](https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools).",
    "674098": "@hengck23  this is great stuff for learning for sure! if you can follow this up with like additional details around \n\n&gt; \"training log is important to analyse model behavior. it is probably the most important reference for beginner to learn how to train a model. the provision of such log is to compare your version with the reference version to see how change of hyper-parameters affect results.\"\n\nwith examples for intuition building / reading you are platinum! ",
    "673776": "@hengck23 are you using for the ensembling your sharpening method?",
    "673752": "hi heng, I ran your code, [dn0,1,2,3] always zero, normal?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1672911%2F3ff73219a66bd3c20c7549fa44fa42d1%2FQQ20191115204236.png?generation=1573821978090365&amp;alt=media)\n\n",
    "673230": "@hengck23  thanks your sharing, my experiment also show that JPU + ASPP converge faster than other model\nMay I ask that what norm layer you use in JPU and ASPP? because with small batch size (4~6), BN won't behave well isn't it?\nThanks",
    "672859": "Do you have a post process in validation?\n",
    "672189": "Thanks, Heng. I ran your resnet34-fpn, got similar vaild performance but the model only gave 0.639 lb. Is this normal or I missed something?",
    "672085": "amazing kenal 👍 ",
    "671716": "hello, thanks for your sharing!\nI have a small question about the data:\nTake \"0a14f2b.png\" for example: how can I get the same mask as you did?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2154603%2F100193f640d51cad5c8e063bddd0c55b%2Fmyplot.png?generation=1573619814792392&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2154603%2F368a207e87078df3b936cc6796f5b1f6%2F0a14f2b.png?generation=1573619846740705&amp;alt=media)\n",
    "671705": "hi Heng, I found you use constant lr 0.01 to train\n\ndo you use any schduler, like ReduceLROnPlateau? or you just change lr by hand and restart from some checkpoint?",
    "671089": "Resizing image is ok, but how to resize masks? 4 separate images corresponding to 4 classes?",
    "670403": "Thank you very very much. I will do my best from now on. ",
    "670135": "best single fold .658, ensemble .667 ??",
    "669946": "wow, nearly 0.01 ensembling boost? I got a single model/fold 0.666 but struggle to ensemble...",
    "669222": "the clouds are actually Shallow cumulus in the trade wind region? So maybe they have specific eddy pattern ...\nif the satellite image are taken at fixed time of the day, the sunlight and shadow are consistent. i suspect that by aligning the kaggle images back to the original orientation (the big black slant facing in the same direction), the data may be more consistent (e.g. test are more close to train?)",
    "668579": "mix augmentation\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa9129eca14eede16bc8b59a3c95b17bb%2FSelection_060.png?generation=1573226907253607&amp;alt=media)\n",
    "666390": "Welcome @hengck23, I was missing you in this competition. Please be carefull this time and Best of luck!!!!",
    "674438": "",
    "670673": ""
  }
}