{
  "id": 74981,
  "title": "If it looks too good to be true, it probably is",
  "url": "/competitions/PLAsTiCC-2018/discussion/74981",
  "author_name": "",
  "post_date": "2018-12-17T19:33:19.274086400Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Latest CV score 0.14.  Guessing I have some serious overfitting going on.  Time to crank up the regularization and see if I can get CV up towards 0.4 before burning my last submission.</p>",
  "messages": [
    {
      "id": "440614",
      "postDate": "12/17/2018 19:33:19",
      "content": "<p>Latest CV score 0.14.  Guessing I have some serious overfitting going on.  Time to crank up the regularization and see if I can get CV up towards 0.4 before burning my last submission.</p>",
      "rawMarkdown": "Latest CV score 0.14.  Guessing I have some serious overfitting going on.  Time to crank up the regularization and see if I can get CV up towards 0.4 before burning my last submission.",
      "votes": null
    },
    {
      "id": "440622",
      "postDate": "12/17/2018 19:44:35",
      "content": "<p>My goodness, this is incredible. Double check everything, hope you don't have any target leakage in the training features. Good luck :) </p>",
      "rawMarkdown": "My goodness, this is incredible. Double check everything, hope you don't have any target leakage in the training features. Good luck :)",
      "votes": null
    },
    {
      "id": "440643",
      "postDate": "12/17/2018 20:26:41",
      "content": "<p>Data leakage I would think.  This is really low.</p>",
      "rawMarkdown": "Data leakage I would think.  This is really low.",
      "votes": null
    },
    {
      "id": "440694",
      "postDate": "12/17/2018 22:33:33",
      "content": "<p>I believe you are right.  I haven't yet figured out exactly why / how.  What happened was that I was down to my last submission and I wanted to experiment with different blends of my LGBM, NN, and SVM models.  I thought I could make an ML model find the optimal blend.  So I used the model's predictions as features (I had each model predict train and test).  So my training set had 44 columns ('object_id', 'target', 14 x 3 predicted probas) and my test set had 43 columns (ditto train but without the target).</p>\n\n<p>I guess since the target was used to train the models that predicted the probabilities that's where the leakage occurred.  Actual result was abysmal (1.685).  I kind of knew it would be (I couldn't raise the CV via regularization) but I was out of time anyways.</p>",
      "rawMarkdown": "I believe you are right.  I haven't yet figured out exactly why / how.  What happened was that I was down to my last submission and I wanted to experiment with different blends of my LGBM, NN, and SVM models.  I thought I could make an ML model find the optimal blend.  So I used the model's predictions as features (I had each model predict train and test).  So my training set had 44 columns ('object_id', 'target', 14 x 3 predicted probas) and my test set had 43 columns (ditto train but without the target).\n\nI guess since the target was used to train the models that predicted the probabilities that's where the leakage occurred.  Actual result was abysmal (1.685).  I kind of knew it would be (I couldn't raise the CV via regularization) but I was out of time anyways.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 440622,
      "author_name": "indranilbhattacharya",
      "author_url": "",
      "post_date": "12/17/2018 19:44:35",
      "content": "<p>My goodness, this is incredible. Double check everything, hope you don't have any target leakage in the training features. Good luck :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 440643,
      "author_name": "sdoria",
      "author_url": "",
      "post_date": "12/17/2018 20:26:41",
      "content": "<p>Data leakage I would think.  This is really low.</p>",
      "votes": null,
      "replies": [
        {
          "id": 440694,
          "author_name": "jimpsull",
          "author_url": "",
          "post_date": "12/17/2018 22:33:33",
          "content": "<p>I believe you are right.  I haven't yet figured out exactly why / how.  What happened was that I was down to my last submission and I wanted to experiment with different blends of my LGBM, NN, and SVM models.  I thought I could make an ML model find the optimal blend.  So I used the model's predictions as features (I had each model predict train and test).  So my training set had 44 columns ('object_id', 'target', 14 x 3 predicted probas) and my test set had 43 columns (ditto train but without the target).</p>\n\n<p>I guess since the target was used to train the models that predicted the probabilities that's where the leakage occurred.  Actual result was abysmal (1.685).  I kind of knew it would be (I couldn't raise the CV via regularization) but I was out of time anyways.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "440614": "Latest CV score 0.14.  Guessing I have some serious overfitting going on.  Time to crank up the regularization and see if I can get CV up towards 0.4 before burning my last submission.",
    "440622": "My goodness, this is incredible. Double check everything, hope you don't have any target leakage in the training features. Good luck :)",
    "440643": "Data leakage I would think.  This is really low.",
    "440694": "I believe you are right.  I haven't yet figured out exactly why / how.  What happened was that I was down to my last submission and I wanted to experiment with different blends of my LGBM, NN, and SVM models.  I thought I could make an ML model find the optimal blend.  So I used the model's predictions as features (I had each model predict train and test).  So my training set had 44 columns ('object_id', 'target', 14 x 3 predicted probas) and my test set had 43 columns (ditto train but without the target).\n\nI guess since the target was used to train the models that predicted the probabilities that's where the leakage occurred.  Actual result was abysmal (1.685).  I kind of knew it would be (I couldn't raise the CV via regularization) but I was out of time anyways."
  },
  "source": "meta"
}