{
  "id": 76406,
  "title": "Pseudo labeling",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76406",
  "author_name": "",
  "post_date": "2019-01-02T14:20:17.533097900Z",
  "votes": 7,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Given that the test set will be changed in stage 2, does it make sense to do pseudo labeling on the test set provided at stage 1?</p>",
  "messages": [
    {
      "id": "449008",
      "postDate": "01/02/2019 14:20:17",
      "content": "<p>Given that the test set will be changed in stage 2, does it make sense to do pseudo labeling on the test set provided at stage 1?</p>",
      "rawMarkdown": "Given that the test set will be changed in stage 2, does it make sense to do pseudo labeling on the test set provided at stage 1?",
      "votes": null
    },
    {
      "id": "449009",
      "postDate": "01/02/2019 14:23:41",
      "content": "<p>I think stage 1 test dataset will be added to train dataset  </p>",
      "rawMarkdown": "I think stage 1 test dataset will be added to train dataset",
      "votes": null
    },
    {
      "id": "449014",
      "postDate": "01/02/2019 14:32:04",
      "content": "<p>In my opinion, I think if we can find a reasonable and reliable approach to label the test data within a short amount of time, it could be useful as information of the test data should be precious in the training process. Even in stage 2 where the data will be changed (or added, if that is the case), we can still use the same PL approach to relabel the new test data on-the-fly, and let that precious information goes to our learner.</p>\n\n<p>Having said that, unfortunately, I have tried 2 or 3 approaches, but I still could not be able to find any ‘reasonable’ labeling approach yet ;(. Perhaps someone else in the top already found out some useful ways to do this.</p>",
      "rawMarkdown": "In my opinion, I think if we can find a reasonable and reliable approach to label the test data within a short amount of time, it could be useful as information of the test data should be precious in the training process. Even in stage 2 where the data will be changed (or added, if that is the case), we can still use the same PL approach to relabel the new test data on-the-fly, and let that precious information goes to our learner.\n\nHaving said that, unfortunately, I have tried 2 or 3 approaches, but I still could not be able to find any ‘reasonable’ labeling approach yet ;(. Perhaps someone else in the top already found out some useful ways to do this.",
      "votes": null
    },
    {
      "id": "449055",
      "postDate": "01/02/2019 15:29:07",
      "content": "<p>Sounds good. So far what techniques have you followed for Pseudo labelling?</p>",
      "rawMarkdown": "Sounds good. So far what techniques have you followed for Pseudo labelling?",
      "votes": null
    },
    {
      "id": "449057",
      "postDate": "01/02/2019 15:30:56",
      "content": "<p>Is there any information on this or are you just guessing?</p>",
      "rawMarkdown": "Is there any information on this or are you just guessing?",
      "votes": null
    },
    {
      "id": "449090",
      "postDate": "01/02/2019 16:17:45",
      "content": "<p>Just guessing</p>",
      "rawMarkdown": "Just guessing",
      "votes": null
    },
    {
      "id": "449098",
      "postDate": "01/02/2019 16:32:44",
      "content": "<p>Definitely not, train dataset always stays the same. How should we ever judge runtime otherwise?</p>",
      "rawMarkdown": "Definitely not, train dataset always stays the same. How should we ever judge runtime otherwise?",
      "votes": null
    },
    {
      "id": "449118",
      "postDate": "01/02/2019 17:22:21",
      "content": "<p>Yup, there will only change in test data and not training. </p>\n\n<p>In the second stage of the competition, we will re-run your selected Kernels. The following files will be swapped with new data:</p>\n\n<p>test.csv - This will be swapped with the complete public and private test dataset. This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.\nsample_submission.csv - similar to test.csv, this will be changed from ~56k in stage 1 to ~376k rows in stage 2 . The file name will remain the same.</p>",
      "rawMarkdown": "Yup, there will only change in test data and not training. \n\nIn the second stage of the competition, we will re-run your selected Kernels. The following files will be swapped with new data:\n\ntest.csv - This will be swapped with the complete public and private test dataset. This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.\nsample_submission.csv - similar to test.csv, this will be changed from ~56k in stage 1 to ~376k rows in stage 2 . The file name will remain the same.",
      "votes": null
    },
    {
      "id": "449309",
      "postDate": "01/03/2019 00:08:29",
      "content": "<p>Hi Rajesh, I tried just simple and adhoc approaches:\n- Train one model first (try to choose the best model), and let it labels for the other models\n- Ensembling model 1, ..., N to label for model N+1\n- Use model 1 to label for model2, use model 3 for label for model4, etc. (half of the models are trained normally, and the other half then used PL &gt;&gt; the reason of this idea was to create diversity for pseudo labeling &gt;&gt; did not quite successful in the first try)\n- Add small noise to the above method</p>",
      "rawMarkdown": "Hi Rajesh, I tried just simple and adhoc approaches:\n- Train one model first (try to choose the best model), and let it labels for the other models\n- Ensembling model 1, ..., N to label for model N+1\n- Use model 1 to label for model2, use model 3 for label for model4, etc. (half of the models are trained normally, and the other half then used PL &gt;&gt; the reason of this idea was to create diversity for pseudo labeling &gt;&gt; did not quite successful in the first try)\n- Add small noise to the above method",
      "votes": null
    },
    {
      "id": "449311",
      "postDate": "01/03/2019 00:16:08",
      "content": "<p>For some approaches, I could see some performance improvements of the base models using PL (ie. increased from .68x to .69x). However, their prediction correlations increased (diversity decreased), and so the ensembling performance did not improve. </p>",
      "rawMarkdown": "For some approaches, I could see some performance improvements of the base models using PL (ie. increased from .68x to .69x). However, their prediction correlations increased (diversity decreased), and so the ensembling performance did not improve.",
      "votes": null
    },
    {
      "id": "449376",
      "postDate": "01/03/2019 03:45:05",
      "content": "<blockquote>\n  <p>does it make sense to do pseudo labeling on the test set provided at stage 1</p>\n</blockquote>\n\n<p>If you want your real final kernel to pseudo label stage 2 test set, then maybe?</p>",
      "rawMarkdown": "&gt; does it make sense to do pseudo labeling on the test set provided at stage 1\n\nIf you want your real final kernel to pseudo label stage 2 test set, then maybe?",
      "votes": null
    },
    {
      "id": "449991",
      "postDate": "01/04/2019 04:33:45",
      "content": "<p>I think time is the main constraint here, labelling test data and then use it for training, consumes quite some time I guess. I think best way to go for the problem is try to find best single model, in minimum time.</p>",
      "rawMarkdown": "I think time is the main constraint here, labelling test data and then use it for training, consumes quite some time I guess. I think best way to go for the problem is try to find best single model, in minimum time.",
      "votes": null
    },
    {
      "id": "450138",
      "postDate": "01/04/2019 11:33:04",
      "content": "<p>That might be true aintnosunshine. I may try my last shot on this method this week before going to try something else.</p>",
      "rawMarkdown": "That might be true aintnosunshine. I may try my last shot on this method this week before going to try something else.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 449009,
      "author_name": "",
      "author_url": "",
      "post_date": "01/02/2019 14:23:41",
      "content": "<p>I think stage 1 test dataset will be added to train dataset  </p>",
      "votes": null,
      "replies": [
        {
          "id": 449057,
          "author_name": "rajeshbhat",
          "author_url": "",
          "post_date": "01/02/2019 15:30:56",
          "content": "<p>Is there any information on this or are you just guessing?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 449090,
          "author_name": "",
          "author_url": "",
          "post_date": "01/02/2019 16:17:45",
          "content": "<p>Just guessing</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 449098,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "01/02/2019 16:32:44",
          "content": "<p>Definitely not, train dataset always stays the same. How should we ever judge runtime otherwise?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 449118,
          "author_name": "rajeshbhat",
          "author_url": "",
          "post_date": "01/02/2019 17:22:21",
          "content": "<p>Yup, there will only change in test data and not training. </p>\n\n<p>In the second stage of the competition, we will re-run your selected Kernels. The following files will be swapped with new data:</p>\n\n<p>test.csv - This will be swapped with the complete public and private test dataset. This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.\nsample_submission.csv - similar to test.csv, this will be changed from ~56k in stage 1 to ~376k rows in stage 2 . The file name will remain the same.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 449014,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "01/02/2019 14:32:04",
      "content": "<p>In my opinion, I think if we can find a reasonable and reliable approach to label the test data within a short amount of time, it could be useful as information of the test data should be precious in the training process. Even in stage 2 where the data will be changed (or added, if that is the case), we can still use the same PL approach to relabel the new test data on-the-fly, and let that precious information goes to our learner.</p>\n\n<p>Having said that, unfortunately, I have tried 2 or 3 approaches, but I still could not be able to find any ‘reasonable’ labeling approach yet ;(. Perhaps someone else in the top already found out some useful ways to do this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 449055,
          "author_name": "rajeshbhat",
          "author_url": "",
          "post_date": "01/02/2019 15:29:07",
          "content": "<p>Sounds good. So far what techniques have you followed for Pseudo labelling?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 449309,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "01/03/2019 00:08:29",
          "content": "<p>Hi Rajesh, I tried just simple and adhoc approaches:\n- Train one model first (try to choose the best model), and let it labels for the other models\n- Ensembling model 1, ..., N to label for model N+1\n- Use model 1 to label for model2, use model 3 for label for model4, etc. (half of the models are trained normally, and the other half then used PL &gt;&gt; the reason of this idea was to create diversity for pseudo labeling &gt;&gt; did not quite successful in the first try)\n- Add small noise to the above method</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 449311,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "01/03/2019 00:16:08",
          "content": "<p>For some approaches, I could see some performance improvements of the base models using PL (ie. increased from .68x to .69x). However, their prediction correlations increased (diversity decreased), and so the ensembling performance did not improve. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 449991,
          "author_name": "s4sarath",
          "author_url": "",
          "post_date": "01/04/2019 04:33:45",
          "content": "<p>I think time is the main constraint here, labelling test data and then use it for training, consumes quite some time I guess. I think best way to go for the problem is try to find best single model, in minimum time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 450138,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "01/04/2019 11:33:04",
          "content": "<p>That might be true aintnosunshine. I may try my last shot on this method this week before going to try something else.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 449376,
      "author_name": "peterhurford",
      "author_url": "",
      "post_date": "01/03/2019 03:45:05",
      "content": "<blockquote>\n  <p>does it make sense to do pseudo labeling on the test set provided at stage 1</p>\n</blockquote>\n\n<p>If you want your real final kernel to pseudo label stage 2 test set, then maybe?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "449008": "Given that the test set will be changed in stage 2, does it make sense to do pseudo labeling on the test set provided at stage 1?",
    "449009": "I think stage 1 test dataset will be added to train dataset",
    "449014": "In my opinion, I think if we can find a reasonable and reliable approach to label the test data within a short amount of time, it could be useful as information of the test data should be precious in the training process. Even in stage 2 where the data will be changed (or added, if that is the case), we can still use the same PL approach to relabel the new test data on-the-fly, and let that precious information goes to our learner.\n\nHaving said that, unfortunately, I have tried 2 or 3 approaches, but I still could not be able to find any ‘reasonable’ labeling approach yet ;(. Perhaps someone else in the top already found out some useful ways to do this.",
    "449055": "Sounds good. So far what techniques have you followed for Pseudo labelling?",
    "449057": "Is there any information on this or are you just guessing?",
    "449090": "Just guessing",
    "449098": "Definitely not, train dataset always stays the same. How should we ever judge runtime otherwise?",
    "449118": "Yup, there will only change in test data and not training. \n\nIn the second stage of the competition, we will re-run your selected Kernels. The following files will be swapped with new data:\n\ntest.csv - This will be swapped with the complete public and private test dataset. This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions. The file name will be the same (both test.csv) to ensure that your code will run.\nsample_submission.csv - similar to test.csv, this will be changed from ~56k in stage 1 to ~376k rows in stage 2 . The file name will remain the same.",
    "449309": "Hi Rajesh, I tried just simple and adhoc approaches:\n- Train one model first (try to choose the best model), and let it labels for the other models\n- Ensembling model 1, ..., N to label for model N+1\n- Use model 1 to label for model2, use model 3 for label for model4, etc. (half of the models are trained normally, and the other half then used PL &gt;&gt; the reason of this idea was to create diversity for pseudo labeling &gt;&gt; did not quite successful in the first try)\n- Add small noise to the above method",
    "449311": "For some approaches, I could see some performance improvements of the base models using PL (ie. increased from .68x to .69x). However, their prediction correlations increased (diversity decreased), and so the ensembling performance did not improve.",
    "449376": "&gt; does it make sense to do pseudo labeling on the test set provided at stage 1\n\nIf you want your real final kernel to pseudo label stage 2 test set, then maybe?",
    "449991": "I think time is the main constraint here, labelling test data and then use it for training, consumes quite some time I guess. I think best way to go for the problem is try to find best single model, in minimum time.",
    "450138": "That might be true aintnosunshine. I may try my last shot on this method this week before going to try something else."
  },
  "source": "meta"
}