{
  "id": 69376,
  "title": "Clustering or Classification",
  "url": "/competitions/PLAsTiCC-2018/discussion/69376",
  "author_name": "",
  "post_date": "2018-10-23T10:09:19.817676Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>When the pdf states that </p>\n\n<blockquote>\n  <p>the training\n  data will be a small subset of the full data, and will also be a poor representation of the\n  test set</p>\n</blockquote>\n\n<p>Using the train set to do a straight classification will result in a poor score by definition no?</p>\n\n<p>On Olivier's marvellous script look at the weighted loss scores they are miles away from the LB score</p>",
  "messages": [
    {
      "id": "408699",
      "postDate": "10/23/2018 10:09:19",
      "content": "<p>When the pdf states that </p>\n\n<blockquote>\n  <p>the training\n  data will be a small subset of the full data, and will also be a poor representation of the\n  test set</p>\n</blockquote>\n\n<p>Using the train set to do a straight classification will result in a poor score by definition no?</p>\n\n<p>On Olivier's marvellous script look at the weighted loss scores they are miles away from the LB score</p>",
      "rawMarkdown": "When the pdf states that \n\n&gt; the training\ndata will be a small subset of the full data, and will also be a poor representation of the\ntest set\n\nUsing the train set to do a straight classification will result in a poor score by definition no?\n\nOn Olivier's marvellous script look at the weighted loss scores they are miles away from the LB score",
      "votes": null
    },
    {
      "id": "408709",
      "postDate": "10/23/2018 10:31:13",
      "content": "<p>I tried to do some pseudo-labeling (a form of semi-supervised learning). I took the observations in the test set with the highest predicted values (I used a threshold). I then added these observations to the training set and re-ran my model. It didn't help, in fact my score went down by a lot. I'm not sure what do to without more training data...</p>",
      "rawMarkdown": "I tried to do some pseudo-labeling (a form of semi-supervised learning). I took the observations in the test set with the highest predicted values (I used a threshold). I then added these observations to the training set and re-ran my model. It didn't help, in fact my score went down by a lot. I'm not sure what do to without more training data...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 408709,
      "author_name": "maxhalford",
      "author_url": "",
      "post_date": "10/23/2018 10:31:13",
      "content": "<p>I tried to do some pseudo-labeling (a form of semi-supervised learning). I took the observations in the test set with the highest predicted values (I used a threshold). I then added these observations to the training set and re-ran my model. It didn't help, in fact my score went down by a lot. I'm not sure what do to without more training data...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "408699": "When the pdf states that \n\n&gt; the training\ndata will be a small subset of the full data, and will also be a poor representation of the\ntest set\n\nUsing the train set to do a straight classification will result in a poor score by definition no?\n\nOn Olivier's marvellous script look at the weighted loss scores they are miles away from the LB score",
    "408709": "I tried to do some pseudo-labeling (a form of semi-supervised learning). I took the observations in the test set with the highest predicted values (I used a threshold). I then added these observations to the training set and re-ran my model. It didn't help, in fact my score went down by a lot. I'm not sure what do to without more training data..."
  },
  "source": "meta"
}