{
  "id": 106177,
  "title": "Solve train-test distribution difference via PseudoLabeling",
  "url": "/competitions/aptos2019-blindness-detection/discussion/106177",
  "author_name": "Bibek",
  "post_date": "2019-08-28T16:15:53.417000",
  "votes": 11,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The difference in LB and CV has already been mentioned many times like <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/97860#latest-607462\">here</a> and <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815#latest-60\">here</a>.  This could be due to difference in train and test(public+private) distribution as shown <a href=\"https://www.kaggle.com/currypurin/image-shape-distribution-previous-and-present\">here</a> and discussed <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/104692#latest-603729\">here</a></p>\n\n<p>According to this <a href=\"https://www.youtube.com/watch?v=sfk5h0yC67o\">lecture</a> by Andrew Ng, one way to solve <code>train-test distribution difference</code> would be to add portion of <code>test set</code>(which come from different distribution than train) to the <code>train set</code> and  then training the model. </p>\n\n<p>In our case, we can do this by <code>PseudoLabeling</code>. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2Fb2c58b100fb88e3a981fcf3673b18e17%2Fpsuedo.png?generation=1567008187143187&amp;alt=media\" alt=\"\"></p>\n\n<p>Train ur model using train(+external data) and then make predictions on (public) test set; take the most confident predictions and add them to the original training set and train again. You can repeat this process for <code>K</code> rounds where <code>K</code> depends on u.</p>\n\n<h2>But how to know confident predictions?</h2>\n\n<p>My way of measuring confidence would be to take two(or more) individual models(regression or classification) and make predictions on the public test set. <code>Add the samples which have same predictions from both these models.</code></p>\n\n<p>I'm pretty sure there are other smart ways to do <code>pseudolabeling</code>. please share ur way :) </p>",
  "messages": [
    {
      "id": 610294,
      "postDate": "2019-08-28T16:15:53.417Z",
      "content": "<p>The difference in LB and CV has already been mentioned many times like <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/97860#latest-607462\">here</a> and <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815#latest-60\">here</a>.  This could be due to difference in train and test(public+private) distribution as shown <a href=\"https://www.kaggle.com/currypurin/image-shape-distribution-previous-and-present\">here</a> and discussed <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/104692#latest-603729\">here</a></p>\n\n<p>According to this <a href=\"https://www.youtube.com/watch?v=sfk5h0yC67o\">lecture</a> by Andrew Ng, one way to solve <code>train-test distribution difference</code> would be to add portion of <code>test set</code>(which come from different distribution than train) to the <code>train set</code> and  then training the model. </p>\n\n<p>In our case, we can do this by <code>PseudoLabeling</code>. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2Fb2c58b100fb88e3a981fcf3673b18e17%2Fpsuedo.png?generation=1567008187143187&amp;alt=media\" alt=\"\"></p>\n\n<p>Train ur model using train(+external data) and then make predictions on (public) test set; take the most confident predictions and add them to the original training set and train again. You can repeat this process for <code>K</code> rounds where <code>K</code> depends on u.</p>\n\n<h2>But how to know confident predictions?</h2>\n\n<p>My way of measuring confidence would be to take two(or more) individual models(regression or classification) and make predictions on the public test set. <code>Add the samples which have same predictions from both these models.</code></p>\n\n<p>I'm pretty sure there are other smart ways to do <code>pseudolabeling</code>. please share ur way :) </p>",
      "rawMarkdown": "The difference in LB and CV has already been mentioned many times like [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/97860#latest-607462) and [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815#latest-60).  This could be due to difference in train and test(public+private) distribution as shown [here](https://www.kaggle.com/currypurin/image-shape-distribution-previous-and-present) and discussed [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/104692#latest-603729)\n\nAccording to this [lecture](https://www.youtube.com/watch?v=sfk5h0yC67o) by Andrew Ng, one way to solve `train-test distribution difference` would be to add portion of `test set`(which come from different distribution than train) to the `train set` and  then training the model. \n\nIn our case, we can do this by `PseudoLabeling`. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2Fb2c58b100fb88e3a981fcf3673b18e17%2Fpsuedo.png?generation=1567008187143187&amp;alt=media)\n\n\nTrain ur model using train(+external data) and then make predictions on (public) test set; take the most confident predictions and add them to the original training set and train again. You can repeat this process for `K` rounds where `K` depends on u.\n\n## But how to know confident predictions? \nMy way of measuring confidence would be to take two(or more) individual models(regression or classification) and make predictions on the public test set. `Add the samples which have same predictions from both these models.`\n\nI'm pretty sure there are other smart ways to do `pseudolabeling`. please share ur way :) ",
      "votes": 11
    },
    {
      "id": 610295,
      "postDate": "2019-08-28T16:16:49.030Z",
      "content": "<p>[Reference] Pic for this post was taken from this <a href=\"https://arxiv.org/pdf/1904.04445.pdf\">paper</a></p>",
      "rawMarkdown": "[Reference] Pic for this post was taken from this [paper](https://arxiv.org/pdf/1904.04445.pdf)"
    },
    {
      "id": 621666,
      "postDate": "2019-09-08T20:16:14.917Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 610648,
      "postDate": "2019-08-28T23:42:16.230Z",
      "content": "<p>thanks <a href=\"/bibek777\">@bibek777</a> </p>",
      "rawMarkdown": "thanks @bibek777 \n"
    }
  ],
  "comments": [
    {
      "id": 610295,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2019-08-28T16:16:49.030000",
      "content": "<p>[Reference] Pic for this post was taken from this <a href=\"https://arxiv.org/pdf/1904.04445.pdf\">paper</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621666,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-08T20:16:14.917000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 610648,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2019-08-28T23:42:16.230000",
      "content": "<p>thanks <a href=\"/bibek777\">@bibek777</a> </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "610294": "The difference in LB and CV has already been mentioned many times like [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/97860#latest-607462) and [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815#latest-60).  This could be due to difference in train and test(public+private) distribution as shown [here](https://www.kaggle.com/currypurin/image-shape-distribution-previous-and-present) and discussed [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/104692#latest-603729)\n\nAccording to this [lecture](https://www.youtube.com/watch?v=sfk5h0yC67o) by Andrew Ng, one way to solve `train-test distribution difference` would be to add portion of `test set`(which come from different distribution than train) to the `train set` and  then training the model. \n\nIn our case, we can do this by `PseudoLabeling`. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2Fb2c58b100fb88e3a981fcf3673b18e17%2Fpsuedo.png?generation=1567008187143187&amp;alt=media)\n\n\nTrain ur model using train(+external data) and then make predictions on (public) test set; take the most confident predictions and add them to the original training set and train again. You can repeat this process for `K` rounds where `K` depends on u.\n\n## But how to know confident predictions? \nMy way of measuring confidence would be to take two(or more) individual models(regression or classification) and make predictions on the public test set. `Add the samples which have same predictions from both these models.`\n\nI'm pretty sure there are other smart ways to do `pseudolabeling`. please share ur way :) ",
    "610295": "[Reference] Pic for this post was taken from this [paper](https://arxiv.org/pdf/1904.04445.pdf)",
    "621666": "",
    "610648": "thanks @bibek777 \n"
  }
}