{
  "id": 115944,
  "title": "multi-label makes every prediction uncertain",
  "url": "/competitions/understanding_cloud_organization/discussion/115944",
  "author_name": "",
  "post_date": "2019-11-06T07:30:04.669381200Z",
  "votes": 10,
  "comment_count": 13,
  "views": 0,
  "content": "<p>this is what i get in my quick experiment:</p>\n\n<ul>\n<li>resnet18 unet</li>\n<li>train for 30 iterations , looping over same one batch = 12 images</li>\n<li>training errors are shown</li>\n</ul>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa5a4acc2bedd6e887e12135df4bd5045%2FSelection_052.png?generation=1573025401633090&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "666531",
      "postDate": "11/06/2019 07:30:04",
      "content": "<p>this is what i get in my quick experiment:</p>\n\n<ul>\n<li>resnet18 unet</li>\n<li>train for 30 iterations , looping over same one batch = 12 images</li>\n<li>training errors are shown</li>\n</ul>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa5a4acc2bedd6e887e12135df4bd5045%2FSelection_052.png?generation=1573025401633090&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "this is what i get in my quick experiment:\n\n- resnet18 unet\n- train for 30 iterations , looping over same one batch = 12 images\n- training errors are shown\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa5a4acc2bedd6e887e12135df4bd5045%2FSelection_052.png?generation=1573025401633090&amp;alt=media)",
      "votes": null
    },
    {
      "id": "666539",
      "postDate": "11/06/2019 07:45:47",
      "content": "<p>the data is really noisy, I checked my valid predictons, some predictions was right in real world,  but the annotation was wrong, and some predictions was wrong in real world, but the annotation shows that this prediction is correct. and the test data is the same distribution with this.</p>",
      "rawMarkdown": "the data is really noisy, I checked my valid predictons, some predictions was right in real world,  but the annotation was wrong, and some predictions was wrong in real world, but the annotation shows that this prediction is correct. and the test data is the same distribution with this.",
      "votes": null
    },
    {
      "id": "666546",
      "postDate": "11/06/2019 07:56:04",
      "content": "<p>I had tried to train 4 independent models for each mask , since I thought this might be helpful for overlap labels. But  the results seems worse than single model for 4 masks in my experiment. The model tended to overfit sooner than multi-label case and the loss was more unstable. </p>",
      "rawMarkdown": "I had tried to train 4 independent models for each mask , since I thought this might be helpful for overlap labels. But  the results seems worse than single model for 4 masks in my experiment. The model tended to overfit sooner than multi-label case and the loss was more unstable.",
      "votes": null
    },
    {
      "id": "666550",
      "postDate": "11/06/2019 08:03:00",
      "content": "<p>I wonder whether the remain private testset will have better annotation</p>",
      "rawMarkdown": "I wonder whether the remain private testset will have better annotation",
      "votes": null
    },
    {
      "id": "666580",
      "postDate": "11/06/2019 08:47:32",
      "content": "<p>\"the data is really noisy, ...\"</p>\n\n<p>this is the fun part of this competition :)</p>",
      "rawMarkdown": "\"the data is really noisy, ...\"\n\nthis is the fun part of this competition :)",
      "votes": null
    },
    {
      "id": "666589",
      "postDate": "11/06/2019 09:01:41",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> as same as steel....</p>",
      "rawMarkdown": "hengck23 as same as steel....",
      "votes": null
    },
    {
      "id": "666602",
      "postDate": "11/06/2019 09:23:34",
      "content": "<p>This is something beyond my understanding as well . I spent a lot of valuable time tuning individual class segmentation model . But it converges so soon and the result is quite worse . May be , there was something I was doing wrong . In theory , since the masks were overlapping in quite a few cases , I thought binary segmentations will be a good candidate here .</p>",
      "rawMarkdown": "This is something beyond my understanding as well . I spent a lot of valuable time tuning individual class segmentation model . But it converges so soon and the result is quite worse . May be , there was something I was doing wrong . In theory , since the masks were overlapping in quite a few cases , I thought binary segmentations will be a good candidate here .",
      "votes": null
    },
    {
      "id": "666635",
      "postDate": "11/06/2019 10:01:16",
      "content": "<p>I think maybe when we only train on single class, the portion of noisy label increased in a batch. Under same batch size, the labels of 4 classes case are 4 times of single class case. So the noisy label seems not that influential when calculating the loss. I'm not very sure, just my own understanding.</p>",
      "rawMarkdown": "I think maybe when we only train on single class, the portion of noisy label increased in a batch. Under same batch size, the labels of 4 classes case are 4 times of single class case. So the noisy label seems not that influential when calculating the loss. I'm not very sure, just my own understanding.",
      "votes": null
    },
    {
      "id": "666697",
      "postDate": "11/06/2019 11:47:07",
      "content": "<p>I  think  this  give  us  two   informations :\n1 :  Ensemble  is  important to this  competition\n2 :  How   about    delete   those    train_img    with   huge   loss ?</p>",
      "rawMarkdown": "I  think  this  give  us  two   informations :\n1 :  Ensemble  is  important to this  competition\n2 :  How   about    delete   those    train_img    with   huge   loss ?",
      "votes": null
    },
    {
      "id": "666717",
      "postDate": "11/06/2019 12:14:18",
      "content": "<p>I'm curious about other people's experiments with single class segmentation, I've tried as well, and also thought that would at least give similar results, but mine were worse. I have a feeling that I did something wrong.</p>",
      "rawMarkdown": "I'm curious about other people's experiments with single class segmentation, I've tried as well, and also thought that would at least give similar results, but mine were worse. I have a feeling that I did something wrong.",
      "votes": null
    },
    {
      "id": "666866",
      "postDate": "11/06/2019 15:13:43",
      "content": "<blockquote>\n  <p>How about delete those train_img with huge loss ?</p>\n</blockquote>\n\n<p>Nice suggestion. I plan to try this. In statistics when you build a linear regression model, it helps to remove outliers before fitting your line. Perhaps it will help here too.</p>",
      "rawMarkdown": "&gt; How about delete those train_img with huge loss ?\n\nNice suggestion. I plan to try this. In statistics when you build a linear regression model, it helps to remove outliers before fitting your line. Perhaps it will help here too.",
      "votes": null
    },
    {
      "id": "666876",
      "postDate": "11/06/2019 15:21:51",
      "content": "<p>I don't see why multi-label makes things complicated. You can build a classifier that labels clothing by its color and type. Then a shirt has two overlapping labels. It can be both a t-shirt and red. This doesn't confuse the model. It teaches the model to look very closely at all the clothing because we're going to ask it many questions about the clothing.</p>",
      "rawMarkdown": "I don't see why multi-label makes things complicated. You can build a classifier that labels clothing by its color and type. Then a shirt has two overlapping labels. It can be both a t-shirt and red. This doesn't confuse the model. It teaches the model to look very closely at all the clothing because we're going to ask it many questions about the clothing.",
      "votes": null
    },
    {
      "id": "666914",
      "postDate": "11/06/2019 16:10:50",
      "content": "<p>Great point</p>",
      "rawMarkdown": "Great point",
      "votes": null
    },
    {
      "id": "667034",
      "postDate": "11/06/2019 18:32:55",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> You could also say that the data is really bad and you are actually predicting bad stuff ;) </p>",
      "rawMarkdown": "hengck23 You could also say that the data is really bad and you are actually predicting bad stuff ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 666539,
      "author_name": "garybios",
      "author_url": "",
      "post_date": "11/06/2019 07:45:47",
      "content": "<p>the data is really noisy, I checked my valid predictons, some predictions was right in real world,  but the annotation was wrong, and some predictions was wrong in real world, but the annotation shows that this prediction is correct. and the test data is the same distribution with this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 666550,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/06/2019 08:03:00",
          "content": "<p>I wonder whether the remain private testset will have better annotation</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666580,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/06/2019 08:47:32",
          "content": "<p>\"the data is really noisy, ...\"</p>\n\n<p>this is the fun part of this competition :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666589,
          "author_name": "garybios",
          "author_url": "",
          "post_date": "11/06/2019 09:01:41",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> as same as steel....</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666697,
          "author_name": "xujingzhao",
          "author_url": "",
          "post_date": "11/06/2019 11:47:07",
          "content": "<p>I  think  this  give  us  two   informations :\n1 :  Ensemble  is  important to this  competition\n2 :  How   about    delete   those    train_img    with   huge   loss ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666866,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/06/2019 15:13:43",
          "content": "<blockquote>\n  <p>How about delete those train_img with huge loss ?</p>\n</blockquote>\n\n<p>Nice suggestion. I plan to try this. In statistics when you build a linear regression model, it helps to remove outliers before fitting your line. Perhaps it will help here too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 667034,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "11/06/2019 18:32:55",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> You could also say that the data is really bad and you are actually predicting bad stuff ;) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 666546,
      "author_name": "xiejialun",
      "author_url": "",
      "post_date": "11/06/2019 07:56:04",
      "content": "<p>I had tried to train 4 independent models for each mask , since I thought this might be helpful for overlap labels. But  the results seems worse than single model for 4 masks in my experiment. The model tended to overfit sooner than multi-label case and the loss was more unstable. </p>",
      "votes": null,
      "replies": [
        {
          "id": 666602,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "11/06/2019 09:23:34",
          "content": "<p>This is something beyond my understanding as well . I spent a lot of valuable time tuning individual class segmentation model . But it converges so soon and the result is quite worse . May be , there was something I was doing wrong . In theory , since the masks were overlapping in quite a few cases , I thought binary segmentations will be a good candidate here .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666635,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/06/2019 10:01:16",
          "content": "<p>I think maybe when we only train on single class, the portion of noisy label increased in a batch. Under same batch size, the labels of 4 classes case are 4 times of single class case. So the noisy label seems not that influential when calculating the loss. I'm not very sure, just my own understanding.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666717,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "11/06/2019 12:14:18",
          "content": "<p>I'm curious about other people's experiments with single class segmentation, I've tried as well, and also thought that would at least give similar results, but mine were worse. I have a feeling that I did something wrong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666876,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/06/2019 15:21:51",
          "content": "<p>I don't see why multi-label makes things complicated. You can build a classifier that labels clothing by its color and type. Then a shirt has two overlapping labels. It can be both a t-shirt and red. This doesn't confuse the model. It teaches the model to look very closely at all the clothing because we're going to ask it many questions about the clothing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666914,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "11/06/2019 16:10:50",
          "content": "<p>Great point</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "666531": "this is what i get in my quick experiment:\n\n- resnet18 unet\n- train for 30 iterations , looping over same one batch = 12 images\n- training errors are shown\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa5a4acc2bedd6e887e12135df4bd5045%2FSelection_052.png?generation=1573025401633090&amp;alt=media)",
    "666539": "the data is really noisy, I checked my valid predictons, some predictions was right in real world,  but the annotation was wrong, and some predictions was wrong in real world, but the annotation shows that this prediction is correct. and the test data is the same distribution with this.",
    "666546": "I had tried to train 4 independent models for each mask , since I thought this might be helpful for overlap labels. But  the results seems worse than single model for 4 masks in my experiment. The model tended to overfit sooner than multi-label case and the loss was more unstable.",
    "666550": "I wonder whether the remain private testset will have better annotation",
    "666580": "\"the data is really noisy, ...\"\n\nthis is the fun part of this competition :)",
    "666589": "hengck23 as same as steel....",
    "666602": "This is something beyond my understanding as well . I spent a lot of valuable time tuning individual class segmentation model . But it converges so soon and the result is quite worse . May be , there was something I was doing wrong . In theory , since the masks were overlapping in quite a few cases , I thought binary segmentations will be a good candidate here .",
    "666635": "I think maybe when we only train on single class, the portion of noisy label increased in a batch. Under same batch size, the labels of 4 classes case are 4 times of single class case. So the noisy label seems not that influential when calculating the loss. I'm not very sure, just my own understanding.",
    "666697": "I  think  this  give  us  two   informations :\n1 :  Ensemble  is  important to this  competition\n2 :  How   about    delete   those    train_img    with   huge   loss ?",
    "666717": "I'm curious about other people's experiments with single class segmentation, I've tried as well, and also thought that would at least give similar results, but mine were worse. I have a feeling that I did something wrong.",
    "666866": "&gt; How about delete those train_img with huge loss ?\n\nNice suggestion. I plan to try this. In statistics when you build a linear regression model, it helps to remove outliers before fitting your line. Perhaps it will help here too.",
    "666876": "I don't see why multi-label makes things complicated. You can build a classifier that labels clothing by its color and type. Then a shirt has two overlapping labels. It can be both a t-shirt and red. This doesn't confuse the model. It teaches the model to look very closely at all the clothing because we're going to ask it many questions about the clothing.",
    "666914": "Great point",
    "667034": "hengck23 You could also say that the data is really bad and you are actually predicting bad stuff ;)"
  },
  "source": "meta"
}