{
  "id": 117973,
  "title": "Does pseudo labelling work for you?",
  "url": "/competitions/understanding_cloud_organization/discussion/117973",
  "author_name": "Yirun Zhang",
  "post_date": "2019-11-19T02:34:35.269000",
  "votes": 4,
  "comment_count": 9,
  "views": 0,
  "content": "<p>In this competition, we tried pseudo label and it increased our public score from 0.670 to 0.673. The private score was also improved accordingly.</p>",
  "messages": [
    {
      "id": 676195,
      "postDate": "2019-11-19T02:34:35.270Z",
      "content": "<p>In this competition, we tried pseudo label and it increased our public score from 0.670 to 0.673. The private score was also improved accordingly.</p>",
      "rawMarkdown": "In this competition, we tried pseudo label and it increased our public score from 0.670 to 0.673. The private score was also improved accordingly.",
      "votes": 4
    },
    {
      "id": 677038,
      "postDate": "2019-11-19T18:15:33.937Z",
      "content": "<p>It worked well for Kha increasing his CV and LB by 0.003 described <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/118065\">here</a></p>",
      "rawMarkdown": "It worked well for Kha increasing his CV and LB by 0.003 described [here][1]\n\n[1]: https://www.kaggle.com/c/understanding_cloud_organization/discussion/118065",
      "votes": 1
    },
    {
      "id": 676440,
      "postDate": "2019-11-19T07:23:38.027Z",
      "content": "<p>Definitely, it's helped me get additional 0.01 on my CV (which one correlate with LB). </p>",
      "rawMarkdown": "Definitely, it's helped me get additional 0.01 on my CV (which one correlate with LB). ",
      "votes": 1
    },
    {
      "id": 676429,
      "postDate": "2019-11-19T07:08:11.537Z",
      "content": "<p>I failed. I scrapped 28730 images using modified code of <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/116892#671521\">https://www.kaggle.com/c/understanding_cloud_organization/discussion/116892#671521</a> , pseudo labeled them, concatenated them with test data, selected ~13000 most confident samples, added them to train set, and trained with more noise like in <a href=\"https://arxiv.org/abs/1911.04252\">https://arxiv.org/abs/1911.04252</a> with bce loss. It improved my classifier and Unet cv to about 0.008. However, cv after postprocessing was a bit lower than before I introduced pseudo label.\nThere might have been a bug in my code, or something with postprocessing that screwed up. Leaderboard scores were higher when I didn't use pseudo label.</p>",
      "rawMarkdown": "I failed. I scrapped 28730 images using modified code of https://www.kaggle.com/c/understanding_cloud_organization/discussion/116892#671521 , pseudo labeled them, concatenated them with test data, selected ~13000 most confident samples, added them to train set, and trained with more noise like in https://arxiv.org/abs/1911.04252 with bce loss. It improved my classifier and Unet cv to about 0.008. However, cv after postprocessing was a bit lower than before I introduced pseudo label.\nThere might have been a bug in my code, or something with postprocessing that screwed up. Leaderboard scores were higher when I didn't use pseudo label.",
      "votes": 1,
      "replies": [
        {
          "id": 676442,
          "postDate": "2019-11-19T07:29:46.813Z",
          "content": "<p><a href=\"/gogo827jz\">@gogo827jz</a>, <a href=\"/valyukov\">@valyukov</a>, did you use soft labels? If yes, what loss function did you use when training with soft labels? bce-dice loss didn't work for me.\n<em>Edit: soft labels for pseudo labeling</em></p>",
          "rawMarkdown": "@gogo827jz, @valyukov, did you use soft labels? If yes, what loss function did you use when training with soft labels? bce-dice loss didn't work for me.\n*Edit: soft labels for pseudo labeling*"
        },
        {
          "id": 676469,
          "postDate": "2019-11-19T07:59:52.580Z",
          "content": "<p>what I've done:\n1. select best min size for each threshold \n2. use different threshold for latest stage of training from 0.8 to 0.95\n3. build pseudo labels on single fold ensemble models that's get me best local CV score\n4. find best proportion of labeled/unlabelled data</p>",
          "rawMarkdown": "what I've done:\n1. select best min size for each threshold \n2. use different threshold for latest stage of training from 0.8 to 0.95\n3. build pseudo labels on single fold ensemble models that's get me best local CV score\n4. find best proportion of labeled/unlabelled data"
        },
        {
          "id": 676614,
          "postDate": "2019-11-19T10:55:16.413Z",
          "content": "<p>@YoonSoo</p>\n\n<p>I suggest you can create a new  kaggle dataset and upload your download cloud images.</p>",
          "rawMarkdown": "@YoonSoo\n\nI suggest you can create a new  kaggle dataset and upload your download cloud images.",
          "votes": 1
        },
        {
          "id": 676757,
          "postDate": "2019-11-19T13:36:51.297Z",
          "content": "<p>No, I didn't use soft labels for pseudo labelling. I pretrained the model on all pseudo labels (best submission) and then trained with KFold CV on training data.</p>",
          "rawMarkdown": "No, I didn't use soft labels for pseudo labelling. I pretrained the model on all pseudo labels (best submission) and then trained with KFold CV on training data.",
          "votes": 1
        }
      ]
    },
    {
      "id": 676409,
      "postDate": "2019-11-19T06:31:16.177Z",
      "content": "<p>I failed using pseudo label. It even doesn't work for public score.</p>",
      "rawMarkdown": "I failed using pseudo label. It even doesn't work for public score.",
      "votes": 1
    },
    {
      "id": 676479,
      "postDate": "2019-11-19T08:12:47.923Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 677038,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2019-11-19T18:15:33.937000",
      "content": "<p>It worked well for Kha increasing his CV and LB by 0.003 described <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/118065\">here</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 676440,
      "author_name": "Vlad A",
      "author_url": "",
      "post_date": "2019-11-19T07:23:38.027000",
      "content": "<p>Definitely, it's helped me get additional 0.01 on my CV (which one correlate with LB). </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 676429,
      "author_name": "YoonSoo",
      "author_url": "",
      "post_date": "2019-11-19T07:08:11.537000",
      "content": "<p>I failed. I scrapped 28730 images using modified code of <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/116892#671521\">https://www.kaggle.com/c/understanding_cloud_organization/discussion/116892#671521</a> , pseudo labeled them, concatenated them with test data, selected ~13000 most confident samples, added them to train set, and trained with more noise like in <a href=\"https://arxiv.org/abs/1911.04252\">https://arxiv.org/abs/1911.04252</a> with bce loss. It improved my classifier and Unet cv to about 0.008. However, cv after postprocessing was a bit lower than before I introduced pseudo label.\nThere might have been a bug in my code, or something with postprocessing that screwed up. Leaderboard scores were higher when I didn't use pseudo label.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 676442,
          "author_name": "YoonSoo",
          "author_url": "",
          "post_date": "2019-11-19T07:29:46.813000",
          "content": "<p><a href=\"/gogo827jz\">@gogo827jz</a>, <a href=\"/valyukov\">@valyukov</a>, did you use soft labels? If yes, what loss function did you use when training with soft labels? bce-dice loss didn't work for me.\n<em>Edit: soft labels for pseudo labeling</em></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 676469,
          "author_name": "Vlad A",
          "author_url": "",
          "post_date": "2019-11-19T07:59:52.580000",
          "content": "<p>what I've done:\n1. select best min size for each threshold \n2. use different threshold for latest stage of training from 0.8 to 0.95\n3. build pseudo labels on single fold ensemble models that's get me best local CV score\n4. find best proportion of labeled/unlabelled data</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 676614,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-11-19T10:55:16.413000",
          "content": "<p>@YoonSoo</p>\n\n<p>I suggest you can create a new  kaggle dataset and upload your download cloud images.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 676757,
          "author_name": "Yirun Zhang",
          "author_url": "",
          "post_date": "2019-11-19T13:36:51.297000",
          "content": "<p>No, I didn't use soft labels for pseudo labelling. I pretrained the model on all pseudo labels (best submission) and then trained with KFold CV on training data.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 676409,
      "author_name": "llh1818",
      "author_url": "",
      "post_date": "2019-11-19T06:31:16.177000",
      "content": "<p>I failed using pseudo label. It even doesn't work for public score.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 676479,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-19T08:12:47.923000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "676195": "In this competition, we tried pseudo label and it increased our public score from 0.670 to 0.673. The private score was also improved accordingly.",
    "677038": "It worked well for Kha increasing his CV and LB by 0.003 described [here][1]\n\n[1]: https://www.kaggle.com/c/understanding_cloud_organization/discussion/118065",
    "676440": "Definitely, it's helped me get additional 0.01 on my CV (which one correlate with LB). ",
    "676429": "I failed. I scrapped 28730 images using modified code of https://www.kaggle.com/c/understanding_cloud_organization/discussion/116892#671521 , pseudo labeled them, concatenated them with test data, selected ~13000 most confident samples, added them to train set, and trained with more noise like in https://arxiv.org/abs/1911.04252 with bce loss. It improved my classifier and Unet cv to about 0.008. However, cv after postprocessing was a bit lower than before I introduced pseudo label.\nThere might have been a bug in my code, or something with postprocessing that screwed up. Leaderboard scores were higher when I didn't use pseudo label.",
    "676409": "I failed using pseudo label. It even doesn't work for public score.",
    "676479": ""
  }
}