{
  "id": 59893,
  "title": "Anyone found Pseudo-labeling useful for this competition?",
  "url": "/competitions/avito-demand-prediction/discussion/59893",
  "author_name": "W. Yifan",
  "post_date": "2018-06-28T03:13:51.439000",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>We trained one in the last several hours but didn't finish the process before the competition ends. So wondering anyone got the luck?</p>",
  "messages": [
    {
      "id": 349364,
      "postDate": "2018-06-28T03:13:51.440Z",
      "content": "<p>We trained one in the last several hours but didn't finish the process before the competition ends. So wondering anyone got the luck?</p>",
      "rawMarkdown": "We trained one in the last several hours but didn't finish the process before the competition ends. So wondering anyone got the luck?",
      "votes": 3
    },
    {
      "id": 349380,
      "postDate": "2018-06-28T03:40:28.823Z",
      "content": "<p>We spent some time with PL. For us, there are 2 major problems with PL in this comp</p>\n\n<ol>\n<li>PL is almost exclusively used in classification, not regression. We had no success applying PL for regression and found very little about PL in published literatures.</li>\n<li>We then turned the problem into a binary classification problem to classify if deal_probability = 0. We ran PL on active data but it still underperforms our regular model. This is because active data has no image_top_1 or image features, which are the most important features by far. </li>\n</ol>\n\n<p>We end up adding a few binary classification models trained with PL on active data as base models for diversity. Would love to hear other team's experience with PL. </p>",
      "rawMarkdown": "We spent some time with PL. For us, there are 2 major problems with PL in this comp\n\n 1. PL is almost exclusively used in classification, not regression. We had no success applying PL for regression and found very little about PL in published literatures.\n 2. We then turned the problem into a binary classification problem to classify if deal_probability = 0. We ran PL on active data but it still underperforms our regular model. This is because active data has no image_top_1 or image features, which are the most important features by far. \n\nWe end up adding a few binary classification models trained with PL on active data as base models for diversity. Would love to hear other team's experience with PL. ",
      "votes": 4
    },
    {
      "id": 349512,
      "postDate": "2018-06-28T08:07:16.890Z",
      "content": "<p>I tried it mid of this competition and it helped. Unfortunately we had no time to incorporate.\nWhat I did was to use the prediction on the test set directly as label for test, concatenate train and test and retrain on full dataset. Then predict again on test with newly trained model</p>",
      "rawMarkdown": "I tried it mid of this competition and it helped. Unfortunately we had no time to incorporate.\nWhat I did was to use the prediction on the test set directly as label for test, concatenate train and test and retrain on full dataset. Then predict again on test with newly trained model",
      "votes": 1
    },
    {
      "id": 349376,
      "postDate": "2018-06-28T03:36:40.073Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 349380,
      "author_name": "sijunhe",
      "author_url": "",
      "post_date": "2018-06-28T03:40:28.823000",
      "content": "<p>We spent some time with PL. For us, there are 2 major problems with PL in this comp</p>\n\n<ol>\n<li>PL is almost exclusively used in classification, not regression. We had no success applying PL for regression and found very little about PL in published literatures.</li>\n<li>We then turned the problem into a binary classification problem to classify if deal_probability = 0. We ran PL on active data but it still underperforms our regular model. This is because active data has no image_top_1 or image features, which are the most important features by far. </li>\n</ol>\n\n<p>We end up adding a few binary classification models trained with PL on active data as base models for diversity. Would love to hear other team's experience with PL. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 349512,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2018-06-28T08:07:16.890000",
      "content": "<p>I tried it mid of this competition and it helped. Unfortunately we had no time to incorporate.\nWhat I did was to use the prediction on the test set directly as label for test, concatenate train and test and retrain on full dataset. Then predict again on test with newly trained model</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349376,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T03:36:40.073000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "349364": "We trained one in the last several hours but didn't finish the process before the competition ends. So wondering anyone got the luck?",
    "349380": "We spent some time with PL. For us, there are 2 major problems with PL in this comp\n\n 1. PL is almost exclusively used in classification, not regression. We had no success applying PL for regression and found very little about PL in published literatures.\n 2. We then turned the problem into a binary classification problem to classify if deal_probability = 0. We ran PL on active data but it still underperforms our regular model. This is because active data has no image_top_1 or image features, which are the most important features by far. \n\nWe end up adding a few binary classification models trained with PL on active data as base models for diversity. Would love to hear other team's experience with PL. ",
    "349512": "I tried it mid of this competition and it helped. Unfortunately we had no time to incorporate.\nWhat I did was to use the prediction on the test set directly as label for test, concatenate train and test and retrain on full dataset. Then predict again on test with newly trained model",
    "349376": ""
  }
}