{
  "id": 38298,
  "title": "initial results on pseudo labeling",
  "url": "/competitions/carvana-image-masking-challenge/discussion/38298",
  "author_name": "",
  "post_date": "2017-08-18T17:04:55.818820400Z",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p>In goggle 2017 paper \"Revisiting Unreasonable Effectiveness of Data in Deep Learning Era\"- Chen Sun, <a href=\"https://research.googleblog.com/2017/07/revisiting-unreasonable-effectiveness.html\">https://research.googleblog.com/2017/07/revisiting-unreasonable-effectiveness.html</a>, it says that 300 million train images with 20% label noise beats 1 million train images without noise.</p>\n\n<p>This is the sheer power of data.</p>\n\n<p>In my experiment, i compare results of training with 100064 LB images using label predicted from my LB 0.997 CNN model. I use images from training set (completely disjoint set) for validation.</p>\n\n<p>Here are the results.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/214883/7169/1-Slide1.png\" alt=\"enter image description here\" title=\"\">\n  <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/214883/7170/1-Slide2.png\" alt=\"enter image description here\" title=\"\"></p>",
  "messages": [
    {
      "id": "214883",
      "postDate": "08/18/2017 17:04:55",
      "content": "<p>In goggle 2017 paper \"Revisiting Unreasonable Effectiveness of Data in Deep Learning Era\"- Chen Sun, <a href=\"https://research.googleblog.com/2017/07/revisiting-unreasonable-effectiveness.html\">https://research.googleblog.com/2017/07/revisiting-unreasonable-effectiveness.html</a>, it says that 300 million train images with 20% label noise beats 1 million train images without noise.</p>\n\n<p>This is the sheer power of data.</p>\n\n<p>In my experiment, i compare results of training with 100064 LB images using label predicted from my LB 0.997 CNN model. I use images from training set (completely disjoint set) for validation.</p>\n\n<p>Here are the results.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/214883/7169/1-Slide1.png\" alt=\"enter image description here\" title=\"\">\n  <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/214883/7170/1-Slide2.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "In goggle 2017 paper \"Revisiting Unreasonable Effectiveness of Data in Deep Learning Era\"- Chen Sun, https://research.googleblog.com/2017/07/revisiting-unreasonable-effectiveness.html, it says that 300 million train images with 20% label noise beats 1 million train images without noise.\n\nThis is the sheer power of data.\n\nIn my experiment, i compare results of training with 100064 LB images using label predicted from my LB 0.997 CNN model. I use images from training set (completely disjoint set) for validation.\n\nHere are the results.\n\n  ![enter image description here][1]\n  ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/214883/7169/1-Slide1.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/214883/7170/1-Slide2.png",
      "votes": null
    },
    {
      "id": "214908",
      "postDate": "08/18/2017 18:29:03",
      "content": "<p>I call this \"training on test data\". This is how people got 1.0 score in MNIST competition <a href=\"https://www.kaggle.com/c/digit-recognizer/discussion/35041\">https://www.kaggle.com/c/digit-recognizer/discussion/35041</a> .  Someone would call this \"cheating\". I don't think this is against the rules though. \nBtw it improved my LB score a little bit.</p>",
      "rawMarkdown": "I call this \"training on test data\". This is how people got 1.0 score in MNIST competition https://www.kaggle.com/c/digit-recognizer/discussion/35041 .  Someone would call this \"cheating\". I don't think this is against the rules though. \nBtw it improved my LB score a little bit.",
      "votes": null
    },
    {
      "id": "214917",
      "postDate": "08/18/2017 19:21:02",
      "content": "<p>I did the exact same thing on Planet competition. It improved my score quite a bit for a single model. The best thing about it is that after training a better model with it you can reproduce less noisy test data and repeat the same process.</p>",
      "rawMarkdown": "I did the exact same thing on Planet competition. It improved my score quite a bit for a single model. The best thing about it is that after training a better model with it you can reproduce less noisy test data and repeat the same process.",
      "votes": null
    },
    {
      "id": "214922",
      "postDate": "08/18/2017 19:28:39",
      "content": "<p>Isn't this effectively training to the test set? I thought that was taboo.</p>",
      "rawMarkdown": "Isn't this effectively training to the test set? I thought that was taboo.",
      "votes": null
    },
    {
      "id": "214924",
      "postDate": "08/18/2017 19:31:33",
      "content": "<p>I believe it shouldn't be a problem unless you're hand-labeling the test set.</p>",
      "rawMarkdown": "I believe it shouldn't be a problem unless you're hand-labeling the test set.",
      "votes": null
    },
    {
      "id": "214926",
      "postDate": "08/18/2017 19:36:52",
      "content": "<p>From my experience it works for near-perfect models only. I tried this trick a couple of times in other competitions years ago. It never worked</p>",
      "rawMarkdown": "From my experience it works for near-perfect models only. I tried this trick a couple of times in other competitions years ago. It never worked",
      "votes": null
    },
    {
      "id": "214940",
      "postDate": "08/18/2017 20:37:00",
      "content": "<p>it works fine if your \"noise\" hasn't bounded with image features.\nOtherwise you'll learn your own mistakes.</p>",
      "rawMarkdown": "it works fine if your \"noise\" hasn't bounded with image features.\nOtherwise you'll learn your own mistakes.",
      "votes": null
    },
    {
      "id": "215153",
      "postDate": "08/20/2017 05:33:03",
      "content": "<p>Such methods fall into the category of weak supervised or semi supervised learning. It is legitimate.  It is a natural trend. As data gets larger it will one day impossible to label all data. Such methods will become important one day.</p>\n\n<p>it is a common method used in kaggle and it is allowed.</p>\n\n<p>A better way is that kaggle can released additional unlabelled data for training, so that it is not mixed with the test data at all.</p>",
      "rawMarkdown": "Such methods fall into the category of weak supervised or semi supervised learning. It is legitimate.  It is a natural trend. As data gets larger it will one day impossible to label all data. Such methods will become important one day.\n\nit is a common method used in kaggle and it is allowed.\n\nA better way is that kaggle can released additional unlabelled data for training, so that it is not mixed with the test data at all.",
      "votes": null
    },
    {
      "id": "215154",
      "postDate": "08/20/2017 05:40:40",
      "content": "<p>the art of weak supervised or semi supervised learning is how not to learn noise. it is sometime not easy.</p>",
      "rawMarkdown": "the art of weak supervised or semi supervised learning is how not to learn noise. it is sometime not easy.",
      "votes": null
    },
    {
      "id": "215577",
      "postDate": "08/22/2017 07:27:02",
      "content": "<p>I think another approach can be to start from a lower resolution network to add pseudo-labeling to a larger network, i.e. unet512 test labels to train unet1024 (you can actually even start lower with smaller proportions of test labels and keep adding to higher networks). Have you tried that too?</p>",
      "rawMarkdown": "I think another approach can be to start from a lower resolution network to add pseudo-labeling to a larger network, i.e. unet512 test labels to train unet1024 (you can actually even start lower with smaller proportions of test labels and keep adding to higher networks). Have you tried that too?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 214908,
      "author_name": "pavelgonchar",
      "author_url": "",
      "post_date": "08/18/2017 18:29:03",
      "content": "<p>I call this \"training on test data\". This is how people got 1.0 score in MNIST competition <a href=\"https://www.kaggle.com/c/digit-recognizer/discussion/35041\">https://www.kaggle.com/c/digit-recognizer/discussion/35041</a> .  Someone would call this \"cheating\". I don't think this is against the rules though. \nBtw it improved my LB score a little bit.</p>",
      "votes": null,
      "replies": [
        {
          "id": 215153,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/20/2017 05:33:03",
          "content": "<p>Such methods fall into the category of weak supervised or semi supervised learning. It is legitimate.  It is a natural trend. As data gets larger it will one day impossible to label all data. Such methods will become important one day.</p>\n\n<p>it is a common method used in kaggle and it is allowed.</p>\n\n<p>A better way is that kaggle can released additional unlabelled data for training, so that it is not mixed with the test data at all.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 214917,
      "author_name": "harungunaydin",
      "author_url": "",
      "post_date": "08/18/2017 19:21:02",
      "content": "<p>I did the exact same thing on Planet competition. It improved my score quite a bit for a single model. The best thing about it is that after training a better model with it you can reproduce less noisy test data and repeat the same process.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 214922,
      "author_name": "thewonz",
      "author_url": "",
      "post_date": "08/18/2017 19:28:39",
      "content": "<p>Isn't this effectively training to the test set? I thought that was taboo.</p>",
      "votes": null,
      "replies": [
        {
          "id": 214924,
          "author_name": "harungunaydin",
          "author_url": "",
          "post_date": "08/18/2017 19:31:33",
          "content": "<p>I believe it shouldn't be a problem unless you're hand-labeling the test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 214926,
          "author_name": "pavelgonchar",
          "author_url": "",
          "post_date": "08/18/2017 19:36:52",
          "content": "<p>From my experience it works for near-perfect models only. I tried this trick a couple of times in other competitions years ago. It never worked</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 214940,
      "author_name": "gadgysaidoff",
      "author_url": "",
      "post_date": "08/18/2017 20:37:00",
      "content": "<p>it works fine if your \"noise\" hasn't bounded with image features.\nOtherwise you'll learn your own mistakes.</p>",
      "votes": null,
      "replies": [
        {
          "id": 215154,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/20/2017 05:40:40",
          "content": "<p>the art of weak supervised or semi supervised learning is how not to learn noise. it is sometime not easy.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 215577,
      "author_name": "timothyman",
      "author_url": "",
      "post_date": "08/22/2017 07:27:02",
      "content": "<p>I think another approach can be to start from a lower resolution network to add pseudo-labeling to a larger network, i.e. unet512 test labels to train unet1024 (you can actually even start lower with smaller proportions of test labels and keep adding to higher networks). Have you tried that too?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "214883": "In goggle 2017 paper \"Revisiting Unreasonable Effectiveness of Data in Deep Learning Era\"- Chen Sun, https://research.googleblog.com/2017/07/revisiting-unreasonable-effectiveness.html, it says that 300 million train images with 20% label noise beats 1 million train images without noise.\n\nThis is the sheer power of data.\n\nIn my experiment, i compare results of training with 100064 LB images using label predicted from my LB 0.997 CNN model. I use images from training set (completely disjoint set) for validation.\n\nHere are the results.\n\n  ![enter image description here][1]\n  ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/214883/7169/1-Slide1.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/214883/7170/1-Slide2.png",
    "214908": "I call this \"training on test data\". This is how people got 1.0 score in MNIST competition https://www.kaggle.com/c/digit-recognizer/discussion/35041 .  Someone would call this \"cheating\". I don't think this is against the rules though. \nBtw it improved my LB score a little bit.",
    "214917": "I did the exact same thing on Planet competition. It improved my score quite a bit for a single model. The best thing about it is that after training a better model with it you can reproduce less noisy test data and repeat the same process.",
    "214922": "Isn't this effectively training to the test set? I thought that was taboo.",
    "214924": "I believe it shouldn't be a problem unless you're hand-labeling the test set.",
    "214926": "From my experience it works for near-perfect models only. I tried this trick a couple of times in other competitions years ago. It never worked",
    "214940": "it works fine if your \"noise\" hasn't bounded with image features.\nOtherwise you'll learn your own mistakes.",
    "215153": "Such methods fall into the category of weak supervised or semi supervised learning. It is legitimate.  It is a natural trend. As data gets larger it will one day impossible to label all data. Such methods will become important one day.\n\nit is a common method used in kaggle and it is allowed.\n\nA better way is that kaggle can released additional unlabelled data for training, so that it is not mixed with the test data at all.",
    "215154": "the art of weak supervised or semi supervised learning is how not to learn noise. it is sometime not easy.",
    "215577": "I think another approach can be to start from a lower resolution network to add pseudo-labeling to a larger network, i.e. unet512 test labels to train unet1024 (you can actually even start lower with smaller proportions of test labels and keep adding to higher networks). Have you tried that too?"
  },
  "source": "meta"
}