{
  "id": 69538,
  "title": "LB/CV scores",
  "url": "/competitions/PLAsTiCC-2018/discussion/69538",
  "author_name": "olivier",
  "post_date": "2018-10-24T17:40:48.511000",
  "votes": 43,
  "comment_count": 193,
  "views": 0,
  "content": "<p>I'm opening this thread since I start wondering if I'm not completely off .</p>\n\n<p>My best local CV is around 0.981 (without any sort of class 99 modelization) and LB score at 1.455.</p>\n\n<p>This CV/LB difference has not really changed from my first submissions onward (at least from the day I started to compute a class 99 prediction). </p>\n\n<p>Are you, Dear Competitors, willing to share your own observations around this ? </p>\n\n<p>Thanks.</p>",
  "messages": [
    {
      "id": 409686,
      "postDate": "2018-10-24T17:40:48.513Z",
      "content": "<p>I'm opening this thread since I start wondering if I'm not completely off .</p>\n\n<p>My best local CV is around 0.981 (without any sort of class 99 modelization) and LB score at 1.455.</p>\n\n<p>This CV/LB difference has not really changed from my first submissions onward (at least from the day I started to compute a class 99 prediction). </p>\n\n<p>Are you, Dear Competitors, willing to share your own observations around this ? </p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "I'm opening this thread since I start wondering if I'm not completely off .\n\nMy best local CV is around 0.981 (without any sort of class 99 modelization) and LB score at 1.455.\n\nThis CV/LB difference has not really changed from my first submissions onward (at least from the day I started to compute a class 99 prediction). \n\nAre you, Dear Competitors, willing to share your own observations around this ? \n\nThanks.",
      "votes": 43
    },
    {
      "id": 409971,
      "postDate": "2018-10-25T06:14:38.850Z",
      "content": "<p>Galactic models CV: 0.21 +/- 0.01</p>\n\n<p>Extragalactic models CV: 0.88 +/- 0.02</p>\n\n<p>Local CV ~0.70</p>\n\n<p>LB: 1.155</p>\n\n<p>Overfitting region reached - many \"smart\" features decreasing local CV score by 0.02 increase LB score by 0.10 (the consequence of the fact, that the training data is a poor representation of the test set).\nClass_99 distribution set by intuition, further probing (using multilogloss probabilities instead of a coin or a dice) improved the score by 0.02 only.</p>\n\n<p>EDIT: To be clear, +/- 0.01 means here the range of CV of a few similar but different models, not an error of CV of one model</p>",
      "rawMarkdown": "Galactic models CV: 0.21 +/- 0.01\n\nExtragalactic models CV: 0.88 +/- 0.02\n\nLocal CV ~0.70\n\nLB: 1.155\n\nOverfitting region reached - many \"smart\" features decreasing local CV score by 0.02 increase LB score by 0.10 (the consequence of the fact, that the training data is a poor representation of the test set).\nClass_99 distribution set by intuition, further probing (using multilogloss probabilities instead of a coin or a dice) improved the score by 0.02 only.\n\nEDIT: To be clear, +/- 0.01 means here the range of CV of a few similar but different models, not an error of CV of one model",
      "votes": 11,
      "replies": [
        {
          "id": 409973,
          "postDate": "2018-10-25T06:24:06.143Z",
          "content": "<p>Thanks for sharing <a href=\"/sionek\">@sionek</a> ! very impressive results. Looking at the difference between LB/CV it seems you a have very good intuition of what class 99 should be :) </p>\n\n<p>I hope I'll be able to come anywhere close to your CV / LB score...</p>",
          "rawMarkdown": "Thanks for sharing @sionek ! very impressive results. Looking at the difference between LB/CV it seems you a have very good intuition of what class 99 should be :) \n\nI hope I'll be able to come anywhere close to your CV / LB score...",
          "votes": 1
        },
        {
          "id": 411818,
          "postDate": "2018-10-29T04:23:41.723Z",
          "content": "<blockquote>\n  <p>Overfitting region reached</p>\n</blockquote>\n\n<p>I think I found it too ;)</p>\n\n<p>You seem to have escaped yours!</p>\n\n<p>CV LB gap is decreasing a bit as my models are better.  Comparing those with same class_99 computation (probability that it is not another class):</p>\n\n<p>CV 0.622 LB 1.052  - GAP 0.43</p>\n\n<p>CV 0.733, LB 1.204 - GAP 0.47</p>\n\n<p>CV 0.902, LB 1.405 - GAP 0.50</p>",
          "rawMarkdown": "&gt; Overfitting region reached\n\nI think I found it too ;)\n\nYou seem to have escaped yours!\n\nCV LB gap is decreasing a bit as my models are better.  Comparing those with same class_99 computation (probability that it is not another class):\n\nCV 0.622 LB 1.052  - GAP 0.43\n\nCV 0.733, LB 1.204 - GAP 0.47\n\nCV 0.902, LB 1.405 - GAP 0.50",
          "votes": 2
        },
        {
          "id": 411914,
          "postDate": "2018-10-29T07:47:44.807Z",
          "content": "<blockquote>\n  <p>I think I found it too ;)</p>\n  \n  <p>You seem to have escaped yours!</p>\n</blockquote>\n\n<p>@CPMP,  If you are in a hopeless local minimum, it is time to trust yourself, not your local CV ;) Remember, that one of the best gains in this competition (~0.30) was obtained by removing hostgal_specz from the feature set in spite of CV score worse by ~0.07. It was simply logic. If you invent a new feature you think it should work, check it out in your submission, even if your local CV says \"forget it\". </p>\n\n<p>In many previous competitions, public LB scores vere only a small additional part of \"total\" validation, the main role was played by local CV, proportionally to the volume of the samples. In this competition, the roles are reversed - overproportionally, because of a significant difference between train and test files.</p>",
          "rawMarkdown": "&gt;I think I found it too ;)\n\n&gt;You seem to have escaped yours!\n\n@CPMP,  If you are in a hopeless local minimum, it is time to trust yourself, not your local CV ;) Remember, that one of the best gains in this competition (~0.30) was obtained by removing hostgal_specz from the feature set in spite of CV score worse by ~0.07. It was simply logic. If you invent a new feature you think it should work, check it out in your submission, even if your local CV says \"forget it\". \n\nIn many previous competitions, public LB scores vere only a small additional part of \"total\" validation, the main role was played by local CV, proportionally to the volume of the samples. In this competition, the roles are reversed - overproportionally, because of a significant difference between train and test files.\n",
          "votes": 7
        },
        {
          "id": 411936,
          "postDate": "2018-10-29T08:41:36.493Z",
          "content": "<p>Thanks for the encouragements.</p>\n\n<blockquote>\n  <p>hostgal_specz </p>\n</blockquote>\n\n<p>I did not include it as it was not much present in test data.  </p>\n\n<blockquote>\n  <p>f you invent a new feature you think it should work, check it out in your submission, even if your CV says \"forget it\".</p>\n</blockquote>\n\n<p>You are right in general, train is not representative of test, hence local CV is misleading.  But LB probing is against my nature, I'll resist a bit before giving into it.  And also because upload from my home machine is really slow unfortunately.</p>\n\n<p>I have an idea to try, but it will take a while to run it, and I'm traveling this week.  I don't expect progress before week end therefore.  I hope I won't be too far behind by then!</p>",
          "rawMarkdown": "Thanks for the encouragements.\n\n&gt; hostgal_specz \n\nI did not include it as it was not much present in test data.  \n\n&gt; f you invent a new feature you think it should work, check it out in your submission, even if your CV says \"forget it\".\n\nYou are right in general, train is not representative of test, hence local CV is misleading.  But LB probing is against my nature, I'll resist a bit before giving into it.  And also because upload from my home machine is really slow unfortunately.\n\nI have an idea to try, but it will take a while to run it, and I'm traveling this week.  I don't expect progress before week end therefore.  I hope I won't be too far behind by then!",
          "votes": 1
        },
        {
          "id": 412022,
          "postDate": "2018-10-29T12:08:48.323Z",
          "content": "<p>@CPMP, So, an approach without LB probing. I have found, that if you are in a hopeless local minimum (few unsuccesful submissions),  it makes no sense to submit another \"better\" results, if your local CV improvement is smaller than your CV error (in my case ~0.07).  I have escaped from my local minimum (LB=1.155) submitting results of local CV=0.62 (local CV improvement equal 0.08, LB gain equal 0.12). </p>",
          "rawMarkdown": "@CPMP, So, an approach without LB probing. I have found, that if you are in a hopeless local minimum (few unsuccesful submissions),  it makes no sense to submit another \"better\" results, if your local CV improvement is smaller than your CV error (in my case ~0.07).  I have escaped from my local minimum (LB=1.155) submitting results of local CV=0.62 (local CV improvement equal 0.08, LB gain equal 0.12). ",
          "votes": 3
        }
      ]
    },
    {
      "id": 423706,
      "postDate": "2018-11-18T22:36:25.497Z",
      "content": "<p>My current best submission has CV 0.49, LB 0.83 for a gap of 0.34. So there are ways of lowering the gap!</p>",
      "rawMarkdown": "My current best submission has CV 0.49, LB 0.83 for a gap of 0.34. So there are ways of lowering the gap!",
      "votes": 7,
      "replies": [
        {
          "id": 423721,
          "postDate": "2018-11-18T23:31:18.903Z",
          "content": "<p>Thanks, you are confirming there is room for more effective feature engineering.</p>",
          "rawMarkdown": "Thanks, you are confirming there is room for more effective feature engineering."
        },
        {
          "id": 425521,
          "postDate": "2018-11-21T18:16:09.417Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> </p>\n\n<p>We have a larger gap (0.4 -- 0.45). Do you think the gap is due\n1. over fitting (so parameters of the classifier) \n2. redundant features\n3. inappropriate features\n4. different distribution of classes in test (besides class 99, we still do not know if the other classes distribution is the same -- we maybe better ask)\n3. there is also class 99  (I doubt anyone has dealt with it yet), and it will add to the gap</p>",
          "rawMarkdown": "@cpmpml \n\nWe have a larger gap (0.4 -- 0.45). Do you think the gap is due\n1. over fitting (so parameters of the classifier) \n2. redundant features\n3. inappropriate features\n4. different distribution of classes in test (besides class 99, we still do not know if the other classes distribution is the same -- we maybe better ask)\n3. there is also class 99  (I doubt anyone has dealt with it yet), and it will add to the gap"
        },
        {
          "id": 425567,
          "postDate": "2018-11-21T19:36:25.907Z",
          "content": "<p>Blonde, it is hard to answer without knowing more about how you model the problem.  </p>\n\n<p>But I am not sure a larger gap is an issue.  As you noticed, when I decreased my gap I did not improve my LB.  </p>\n\n<p>I think it helps having a smaller gap; but it is hard to tell until we see private LB scores.  </p>\n\n<p>In the meantime, I would recommend you work on feature selection.  I improved my CV and LB score by removing features.</p>",
          "rawMarkdown": "Blonde, it is hard to answer without knowing more about how you model the problem.  \n\nBut I am not sure a larger gap is an issue.  As you noticed, when I decreased my gap I did not improve my LB.  \n\nI think it helps having a smaller gap; but it is hard to tell until we see private LB scores.  \n\nIn the meantime, I would recommend you work on feature selection.  I improved my CV and LB score by removing features.",
          "votes": 2
        },
        {
          "id": 425664,
          "postDate": "2018-11-21T23:33:47.207Z",
          "content": "<p>Thank you. The feature selection process is going ok, it just takes time to calculate submission. </p>\n\n<p>BTW, you noticed before that 400 features is too much for 7k, from your experience what it the range of optimal ratio between number of samples and number of features? (I have another project with signals, 400 samples, 20 features, but that signals I can easily augment and increase the number in a few times. So it would be nice to know the approximate merit which is ok, like the range of ratios between number of features and number of samples ) </p>",
          "rawMarkdown": "Thank you. The feature selection process is going ok, it just takes time to calculate submission. \n\nBTW, you noticed before that 400 features is too much for 7k, from your experience what it the range of optimal ratio between number of samples and number of features? (I have another project with signals, 400 samples, 20 features, but that signals I can easily augment and increase the number in a few times. So it would be nice to know the approximate merit which is ok, like the range of ratios between number of features and number of samples ) "
        },
        {
          "id": 426012,
          "postDate": "2018-11-22T12:39:31.847Z",
          "content": "<p>It depends on the model you are using, and how much regularization you use.  I reacted because 400 features for 7k examples means a ration of example per feature quite low, and prone to overfiting with lgb or xgb models.  General linear models with proper regularization are less sensitive to that.   And I know for sure that it is possible to get below 0.9 on the Lb with less than 199 features ;)</p>",
          "rawMarkdown": "It depends on the model you are using, and how much regularization you use.  I reacted because 400 features for 7k examples means a ration of example per feature quite low, and prone to overfiting with lgb or xgb models.  General linear models with proper regularization are less sensitive to that.   And I know for sure that it is possible to get below 0.9 on the Lb with less than 199 features ;)"
        },
        {
          "id": 426082,
          "postDate": "2018-11-22T15:20:25.447Z",
          "content": "<p>Thank you. I am free in choosing my model for that project, I used xgboost, but will look into linear models with regularization then. I have only 400 samples, which is not much... and selected 18 features. I'll try to minimize that number. I also tried extracting features with fully conv nets with large augmentation, but it has not work out so far. xgboost at least gave some results, but far from brilliant. 400 for 7k is about the same as 18 for 400, that's why I asked what is the practical merit for \"too many features\"...</p>",
          "rawMarkdown": "Thank you. I am free in choosing my model for that project, I used xgboost, but will look into linear models with regularization then. I have only 400 samples, which is not much... and selected 18 features. I'll try to minimize that number. I also tried extracting features with fully conv nets with large augmentation, but it has not work out so far. xgboost at least gave some results, but far from brilliant. 400 for 7k is about the same as 18 for 400, that's why I asked what is the practical merit for \"too many features\"..."
        },
        {
          "id": 426115,
          "postDate": "2018-11-22T16:35:42.510Z",
          "content": "<p>With 400 samples I would go with a glm.</p>",
          "rawMarkdown": "With 400 samples I would go with a glm.",
          "votes": 2
        },
        {
          "id": 428707,
          "postDate": "2018-11-27T18:33:43.050Z",
          "content": "<p>Update on this: my best single model is now at 0.43 CV / 0.776 LB.</p>",
          "rawMarkdown": "Update on this: my best single model is now at 0.43 CV / 0.776 LB.",
          "votes": 5
        },
        {
          "id": 428713,
          "postDate": "2018-11-27T18:43:55.560Z",
          "content": "<p>May I ask how many features you have for this one?</p>",
          "rawMarkdown": "May I ask how many features you have for this one?",
          "votes": 1
        },
        {
          "id": 428727,
          "postDate": "2018-11-27T19:05:24.947Z",
          "content": "<p><a href=\"/kyleboone\">@kyleboone</a> that is really amazing CV score and gap!! Just curious, is your current LB score a ensemble result?</p>",
          "rawMarkdown": "@kyleboone that is really amazing CV score and gap!! Just curious, is your current LB score a ensemble result?\n",
          "votes": 1
        },
        {
          "id": 428731,
          "postDate": "2018-11-27T19:19:56.927Z",
          "content": "<p>My latest model seems to have some issues with overfitting. I ensembled its predictions with an older one that performs worse but doesn't overfit as much to get the LB score.</p>",
          "rawMarkdown": "My latest model seems to have some issues with overfitting. I ensembled its predictions with an older one that performs worse but doesn't overfit as much to get the LB score.",
          "votes": 3
        },
        {
          "id": 429865,
          "postDate": "2018-11-29T13:28:47.593Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> </p>\n\n<p>LGB, around 100 features, single model, CV 0.52 -- LB 0.979, GAP 0.45 </p>\n\n<p>Any hints on lowering the gap? Which direction to look at: \nparameters of the classifier \nredundant features\ninappropriate features\nsomething else... \nit's my first GBM experience, any hints would be nice</p>",
          "rawMarkdown": "@cpmpml \n\nLGB, around 100 features, single model, CV 0.52 -- LB 0.979, GAP 0.45 \n\nAny hints on lowering the gap? Which direction to look at: \nparameters of the classifier \nredundant features\ninappropriate features\nsomething else... \nit's my first GBM experience, any hints would be nice"
        },
        {
          "id": 429923,
          "postDate": "2018-11-29T14:55:26.997Z",
          "content": "<blockquote>\n  <p>Any hints on lowering the gap?</p>\n</blockquote>\n\n<p>Priority is to lower the LB score ;)</p>\n\n<p>Your gap seems similar to what I had when my LB score was around 0.98, there is no real issue here I think.  Your number of feature does not seem too high either.  Just keep adding good features..</p>\n\n<p>My gap decreased drastically when my model became better.  Latest run is CV 0.420, LB 0.794, Gap 0.374.  I can't disclose what made it move from 0.9x LB to 0.7x before competition end ;)  And I think the two leader have even better single models than that.</p>\n\n<p>To answer some of your questions:</p>\n\n<ul>\n<li><p>Redundant features are not an issue for lgb or xgboost.  </p></li>\n<li><p>Parameter tuning makes sense.  I selected conservative settings given there is a risk of overfiting.  I may try less conservative settings at a point though.</p></li>\n<li><p>Inappropriate features.  Definitely an issue.  Watch for features that lower your CV score without lowering the LB score.</p></li>\n</ul>",
          "rawMarkdown": "&gt; Any hints on lowering the gap?\n\nPriority is to lower the LB score ;)\n\nYour gap seems similar to what I had when my LB score was around 0.98, there is no real issue here I think.  Your number of feature does not seem too high either.  Just keep adding good features..\n\nMy gap decreased drastically when my model became better.  Latest run is CV 0.420, LB 0.794, Gap 0.374.  I can't disclose what made it move from 0.9x LB to 0.7x before competition end ;)  And I think the two leader have even better single models than that.\n\nTo answer some of your questions:\n\n- Redundant features are not an issue for lgb or xgboost.  \n\n- Parameter tuning makes sense.  I selected conservative settings given there is a risk of overfiting.  I may try less conservative settings at a point though.\n\n- Inappropriate features.  Definitely an issue.  Watch for features that lower your CV score without lowering the LB score.",
          "votes": 5
        },
        {
          "id": 438829,
          "postDate": "2018-12-14T08:52:10.367Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a>, from 09x to 0.7x was due to one method/feature?</p>",
          "rawMarkdown": "@cpmpml, from 09x to 0.7x was due to one method/feature?"
        },
        {
          "id": 438842,
          "postDate": "2018-12-14T09:17:21.750Z",
          "content": "<p>You will know in few days.</p>",
          "rawMarkdown": "You will know in few days.",
          "votes": 1
        }
      ]
    },
    {
      "id": 426576,
      "postDate": "2018-11-23T13:23:44.750Z",
      "content": "<p>We've achieved a significant milestone by combining my best lgb model with ideas from my team mates.</p>\n\n<p>single run: CV 0.431, LB 0.801, Gap 0.370</p>\n\n<p>@yuval_r prediction that final best score can be below 0.7 may be true.</p>",
      "rawMarkdown": "We've achieved a significant milestone by combining my best lgb model with ideas from my team mates.\n\nsingle run: CV 0.431, LB 0.801, Gap 0.370\n\n@yuval_r prediction that final best score can be below 0.7 may be true.",
      "votes": 6,
      "replies": [
        {
          "id": 426578,
          "postDate": "2018-11-23T13:29:52.323Z",
          "content": "<p>Well done ! that's impressive.</p>",
          "rawMarkdown": "Well done ! that's impressive.",
          "votes": 2
        },
        {
          "id": 426825,
          "postDate": "2018-11-24T00:08:18.103Z",
          "content": "<p>wow that's great</p>",
          "rawMarkdown": "wow that's great"
        },
        {
          "id": 426971,
          "postDate": "2018-11-24T09:11:23.760Z",
          "content": "<p>Yes, I'm surprised NNs aren't better, but we are working on it ;)</p>",
          "rawMarkdown": "Yes, I'm surprised NNs aren't better, but we are working on it ;)"
        },
        {
          "id": 426985,
          "postDate": "2018-11-24T09:59:40.820Z",
          "content": "<p>that is true to me as well. when my LGBM was at 1.3 LB my NN was already at 1.1X with same features. But when my LGB reached .9X my NN stayed at 1.1X with same features</p>",
          "rawMarkdown": "that is true to me as well. when my LGBM was at 1.3 LB my NN was already at 1.1X with same features. But when my LGB reached .9X my NN stayed at 1.1X with same features"
        },
        {
          "id": 427281,
          "postDate": "2018-11-25T04:20:12.737Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> May I ask you to show us your team's confusion matrix?</p>",
          "rawMarkdown": "@cpmpml May I ask you to show us your team's confusion matrix?"
        },
        {
          "id": 427375,
          "postDate": "2018-11-25T11:49:52.293Z",
          "content": "<p>I'll create a post for it as I don't see how to upload an image to a comment, see <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72613\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72613</a></p>",
          "rawMarkdown": "I'll create a post for it as I don't see how to upload an image to a comment, see https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72613\n\n",
          "votes": 4
        }
      ]
    },
    {
      "id": 415424,
      "postDate": "2018-11-05T05:21:55.630Z",
      "content": "<p>My latest CV (single lightgbm) is 0.659 and LB 1.082. I think I finally found what I missed.</p>",
      "rawMarkdown": "My latest CV (single lightgbm) is 0.659 and LB 1.082. I think I finally found what I missed.",
      "votes": 5,
      "replies": [
        {
          "id": 415427,
          "postDate": "2018-11-05T05:25:58.273Z",
          "content": "<blockquote>\n  <p>I finally found what I missed</p>\n</blockquote>\n\n<p>Indeed, you now have a gap close to 0.43 ;)</p>\n\n<p>More seriously, this is good and steady progress!</p>",
          "rawMarkdown": "&gt; I finally found what I missed\n\nIndeed, you now have a gap close to 0.43 ;)\n\nMore seriously, this is good and steady progress!"
        },
        {
          "id": 415430,
          "postDate": "2018-11-05T05:34:46.917Z",
          "content": "<p>Thanks <a href=\"/cpmpml\">@cpmpml</a> ! I'm still far from top 10, all the more you're still with a single model...</p>",
          "rawMarkdown": "Thanks @cpmpml ! I'm still far from top 10, all the more you're still with a single model...",
          "votes": 1
        },
        {
          "id": 415537,
          "postDate": "2018-11-05T09:40:05.277Z",
          "content": "<blockquote>\n  <p>you're still with a single model</p>\n</blockquote>\n\n<p>Your best LB score comes from a single model too unless mistaken...</p>",
          "rawMarkdown": "&gt; you're still with a single model\n\nYour best LB score comes from a single model too unless mistaken..."
        },
        {
          "id": 415547,
          "postDate": "2018-11-05T09:52:33.157Z",
          "content": "<p>I should have been more explict... you're still 0.12 ahead which is a lot with a single model.</p>",
          "rawMarkdown": "I should have been more explict... you're still 0.12 ahead which is a lot with a single model.",
          "votes": 1
        },
        {
          "id": 415573,
          "postDate": "2018-11-05T10:41:21.340Z",
          "content": "<p>Leaders are way ahead of me as well, and they improve everyday too!  We will learn very interesting things after competition end.</p>",
          "rawMarkdown": "Leaders are way ahead of me as well, and they improve everyday too!  We will learn very interesting things after competition end."
        },
        {
          "id": 416188,
          "postDate": "2018-11-06T11:09:48.373Z",
          "content": "<p><a href=\"/ogrellier\">@ogrellier</a>, this is a nice progress!</p>\n\n<p>I've been stuck for the last few days with a local CV around 0.75 ~ 0.76. I have a feeling that I have too many features that are not that powerful (200+), but so far I couldn't come up with smarter features to improve on that. So I am still trying to find that thing you guys have already discovered :) </p>",
          "rawMarkdown": "@ogrellier, this is a nice progress!\n\nI've been stuck for the last few days with a local CV around 0.75 ~ 0.76. I have a feeling that I have too many features that are not that powerful (200+), but so far I couldn't come up with smarter features to improve on that. So I am still trying to find that thing you guys have already discovered :) "
        },
        {
          "id": 416198,
          "postDate": "2018-11-06T11:41:18.333Z",
          "content": "<p>Thanks <a href=\"/kozodoi\">@kozodoi</a>, depending on what objective/loss you use you may find my latest kernel update useful or not :) </p>\n\n<p><a href=\"https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data\">https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data</a></p>\n\n<p>If I'm right you may land in top 10, hopefully I'm right ;-)</p>",
          "rawMarkdown": "Thanks @kozodoi, depending on what objective/loss you use you may find my latest kernel update useful or not :) \n\nhttps://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data\n\nIf I'm right you may land in top 10, hopefully I'm right ;-)",
          "votes": 3
        },
        {
          "id": 416204,
          "postDate": "2018-11-06T11:54:36.530Z",
          "content": "<p>Thanks Olivier, you are really amazing!</p>",
          "rawMarkdown": "Thanks Olivier, you are really amazing!"
        },
        {
          "id": 416394,
          "postDate": "2018-11-06T16:02:30.757Z",
          "content": "<p>Thanks <a href=\"/ogrellier\">@ogrellier</a>, your kernels are always so helpful!</p>",
          "rawMarkdown": "Thanks @ogrellier, your kernels are always so helpful!"
        }
      ]
    },
    {
      "id": 414643,
      "postDate": "2018-11-03T09:01:51.827Z",
      "content": "<p>With same features</p>\n\n<ul>\n<li>lgb: CV 0.98 LB 1.454 gap 0.474  </li>\n<li>mlp: CV 0.82 LB 1.148 gap 0.328  </li>\n</ul>",
      "rawMarkdown": "With same features\n\n- lgb: CV 0.98 LB 1.454 gap 0.474  \n- mlp: CV 0.82 LB 1.148 gap 0.328  ",
      "votes": 5,
      "replies": [
        {
          "id": 414644,
          "postDate": "2018-11-03T09:12:35.180Z",
          "content": "<p>Seems I should move to mlp!</p>",
          "rawMarkdown": "Seems I should move to mlp!",
          "votes": 2
        },
        {
          "id": 414645,
          "postDate": "2018-11-03T09:29:15.267Z",
          "content": "<p>You certainly want ;-)</p>",
          "rawMarkdown": "You certainly want ;-)",
          "votes": 1
        },
        {
          "id": 414646,
          "postDate": "2018-11-03T09:30:39.427Z",
          "content": "<p>Definitely!</p>",
          "rawMarkdown": "Definitely!"
        },
        {
          "id": 414654,
          "postDate": "2018-11-03T09:50:10.243Z",
          "content": "<p>I will, don't worry !   I'm convinced deep learning is better on this competition as it can capture the overall shape of the light curve, something that is hard to achieve with feature engineering.  But I first want to squeeze out every bit I can from feature engineering as it makes me understand data better.  And I have not yet used XGBoost, which is stil my favorite algorithm.  I need to give it a chance!</p>",
          "rawMarkdown": "I will, don't worry !   I'm convinced deep learning is better on this competition as it can capture the overall shape of the light curve, something that is hard to achieve with feature engineering.  But I first want to squeeze out every bit I can from feature engineering as it makes me understand data better.  And I have not yet used XGBoost, which is stil my favorite algorithm.  I need to give it a chance!",
          "votes": 1
        },
        {
          "id": 414675,
          "postDate": "2018-11-03T11:12:26.520Z",
          "content": "<p>I feel it is strange that mlp is much better than lgbm.\nNN's merit is rnn for time siries features or cnn for image features.\nSo I don't know why mlp work well.\nanyway, I will try mlp.</p>",
          "rawMarkdown": "I feel it is strange that mlp is much better than lgbm.\nNN's merit is rnn for time siries features or cnn for image features.\nSo I don't know why mlp work well.\nanyway, I will try mlp."
        },
        {
          "id": 414690,
          "postDate": "2018-11-03T11:58:15.667Z",
          "content": "<p>Yes, I have the same feeling! <br>\nI guess RNN/CNN may be more proper than MLP for ts data here but found it troublesome to implement the preprocessing part (maybe reshape or padding or interpolating?) </p>",
          "rawMarkdown": "Yes, I have the same feeling!  \nI guess RNN/CNN may be more proper than MLP for ts data here but found it troublesome to implement the preprocessing part (maybe reshape or padding or interpolating?) "
        },
        {
          "id": 414697,
          "postDate": "2018-11-03T12:13:51.407Z",
          "content": "<p>This kernel may be useful when getting ts features.\n<a href=\"https://www.kaggle.com/scirpus/predict-by-row-then-average\">https://www.kaggle.com/scirpus/predict-by-row-then-average</a></p>",
          "rawMarkdown": "This kernel may be useful when getting ts features.\nhttps://www.kaggle.com/scirpus/predict-by-row-then-average",
          "votes": 3
        },
        {
          "id": 414702,
          "postDate": "2018-11-03T12:22:35.473Z",
          "content": "<p>Thanks for guiding! <br>\nThis brings me some new ideas:)</p>",
          "rawMarkdown": "Thanks for guiding!  \nThis brings me some new ideas:)",
          "votes": 1
        },
        {
          "id": 421041,
          "postDate": "2018-11-14T14:02:59.290Z",
          "content": "<p>It's weird. When I tried MLP structure in public kernel with same features of my lgbm, CV goes up to 0.82 from 0.58. Really wonder how did you manage to get lower CV score..</p>",
          "rawMarkdown": "It's weird. When I tried MLP structure in public kernel with same features of my lgbm, CV goes up to 0.82 from 0.58. Really wonder how did you manage to get lower CV score.."
        },
        {
          "id": 421049,
          "postDate": "2018-11-14T14:10:08.130Z",
          "content": "<p>Same here, for me many features that improve LB/CV in my NN model are useless when used in LGBM. It seems in this competition, Boosting and NN need considerably different set of features to work best. </p>",
          "rawMarkdown": "Same here, for me many features that improve LB/CV in my NN model are useless when used in LGBM. It seems in this competition, Boosting and NN need considerably different set of features to work best. ",
          "votes": 1
        },
        {
          "id": 421085,
          "postDate": "2018-11-14T15:06:30.383Z",
          "content": "<p>My best GBM gets around 0.65 CV / 1.085 LB and my best NN with the same features is around 0.75 CV / 1.200 LB.</p>",
          "rawMarkdown": "My best GBM gets around 0.65 CV / 1.085 LB and my best NN with the same features is around 0.75 CV / 1.200 LB."
        },
        {
          "id": 421092,
          "postDate": "2018-11-14T15:12:46.247Z",
          "content": "<p>It's good to hear that I'm not alone then.. May I ask is your current LB score blend of those two?</p>",
          "rawMarkdown": "It's good to hear that I'm not alone then.. May I ask is your current LB score blend of those two?"
        },
        {
          "id": 421108,
          "postDate": "2018-11-14T15:45:42.210Z",
          "content": "<p>I did not stack these 2 alone, but a 50/50 blend was really poor around 1.10</p>",
          "rawMarkdown": "I did not stack these 2 alone, but a 50/50 blend was really poor around 1.10"
        },
        {
          "id": 421139,
          "postDate": "2018-11-14T16:24:31.063Z",
          "content": "<p>I guess having a good blend depends on what you do with <code>class_99</code>. You can include the <code>class_99</code> columns into the blend, or you can blend the other columns and then recompute <code>class_99</code> with the new probabilities.</p>",
          "rawMarkdown": "I guess having a good blend depends on what you do with `class_99`. You can include the `class_99` columns into the blend, or you can blend the other columns and then recompute `class_99` with the new probabilities."
        },
        {
          "id": 421146,
          "postDate": "2018-11-14T16:38:18.090Z",
          "content": "<p>I confirm what others say, using a MLP on similar features than lgb leads to higher CV and LB score.  Moreover, many features lead to overfitting with the MLP.   The good news is that a weighted average of the MLP and lgb models is better than lgb alone still.</p>",
          "rawMarkdown": "I confirm what others say, using a MLP on similar features than lgb leads to higher CV and LB score.  Moreover, many features lead to overfitting with the MLP.   The good news is that a weighted average of the MLP and lgb models is better than lgb alone still.",
          "votes": 1
        },
        {
          "id": 421302,
          "postDate": "2018-11-14T21:49:37.170Z",
          "content": "<blockquote>\n  <p>I did not stack these 2 alone,</p>\n</blockquote>\n\n<p>You are stacking already?</p>\n\n<p>Beware, I'm  starting with XGBoost ;)</p>",
          "rawMarkdown": "&gt; I did not stack these 2 alone,\n\nYou are stacking already?\n\nBeware, I'm  starting with XGBoost ;)"
        },
        {
          "id": 421596,
          "postDate": "2018-11-15T07:04:09.110Z",
          "content": "<p>Yes, I'm on kernel only... and running a bit out of ideas !</p>\n\n<p>XGBoost is a very good option IMHO :)</p>",
          "rawMarkdown": "Yes, I'm on kernel only... and running a bit out of ideas !\n\nXGBoost is a very good option IMHO :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 410032,
      "postDate": "2018-10-25T08:46:49.367Z",
      "content": "<p>I love it when I see we are now dealing with galactic models and extra galactic models!  Machine Learning moving to entire new horizons! :)</p>",
      "rawMarkdown": "I love it when I see we are now dealing with galactic models and extra galactic models!  Machine Learning moving to entire new horizons! :)",
      "votes": 6
    },
    {
      "id": 409843,
      "postDate": "2018-10-24T23:37:38.203Z",
      "content": "<p>Galactic model CV: 0.23</p>\n\n<p>Extra-Galactic model CV: 0.99</p>\n\n<p>LB: 1.53</p>\n\n<p>I believe local CV is low because class_99 distribution is a mystery of the universe :-)</p>",
      "rawMarkdown": "Galactic model CV: 0.23\n\nExtra-Galactic model CV: 0.99\n\nLB: 1.53\n\nI believe local CV is low because class_99 distribution is a mystery of the universe :-)",
      "votes": 6,
      "replies": [
        {
          "id": 409931,
          "postDate": "2018-10-25T04:57:00.740Z",
          "content": "<p>Thanks Giba ! does that mean your LB score is a model blend ? or just Giba's magic ;)</p>",
          "rawMarkdown": "Thanks Giba ! does that mean your LB score is a model blend ? or just Giba's magic ;)",
          "votes": 1
        },
        {
          "id": 411643,
          "postDate": "2018-10-28T16:37:00.997Z",
          "content": "<p>It means I splited the data set in two parts and trained two completely separated models: Galactic and ExtraGalactic</p>",
          "rawMarkdown": "It means I splited the data set in two parts and trained two completely separated models: Galactic and ExtraGalactic"
        },
        {
          "id": 411662,
          "postDate": "2018-10-28T18:06:06.977Z",
          "content": "<p>@Giba I asked this because you talked about a LB score of 1.53 so I guess you meant 1.35 then ;-)</p>\n\n<p>I tried to split into galactic and extra galactic as well but did not get any sort of improvement with that setup. I must be missing something here and surely how CPMP went to 1.05 that fast ...</p>",
          "rawMarkdown": "@Giba I asked this because you talked about a LB score of 1.53 so I guess you meant 1.35 then ;-)\n\nI tried to split into galactic and extra galactic as well but did not get any sort of improvement with that setup. I must be missing something here and surely how CPMP went to 1.05 that fast ...",
          "votes": 1
        },
        {
          "id": 411683,
          "postDate": "2018-10-28T19:19:38.733Z",
          "content": "<blockquote>\n  <p>how CPMP went to 1.05 that fast ...</p>\n</blockquote>\n\n<p>I'm afraid I have reached a local optimum now.  I will need something really new to move under 1.0...</p>",
          "rawMarkdown": "&gt; how CPMP went to 1.05 that fast ...\n\nI'm afraid I have reached a local optimum now.  I will need something really new to move under 1.0..."
        },
        {
          "id": 411692,
          "postDate": "2018-10-28T19:59:44.710Z",
          "content": "<p>splitting the models gave me a bit of a boost (~0.01). tbh i never thought id be above <a href=\"/ogrellier\">@ogrellier</a> on the lb lol. usually my new ideas come from his kernels. i honestly feel like im flying by the 'skin of my teeth' here.</p>",
          "rawMarkdown": "splitting the models gave me a bit of a boost (~0.01). tbh i never thought id be above @ogrellier on the lb lol. usually my new ideas come from his kernels. i honestly feel like im flying by the 'skin of my teeth' here."
        },
        {
          "id": 412523,
          "postDate": "2018-10-30T10:34:52.307Z",
          "content": "<p>@CPMP, are you still using constant for class 99?</p>",
          "rawMarkdown": "@CPMP, are you still using constant for class 99?"
        },
        {
          "id": 421301,
          "postDate": "2018-11-14T21:48:21.190Z",
          "content": "<p>I don't use a constant for class_99, read what I shared ;)</p>",
          "rawMarkdown": "I don't use a constant for class_99, read what I shared ;)"
        }
      ]
    },
    {
      "id": 422993,
      "postDate": "2018-11-17T08:00:51.947Z",
      "content": "<p>My last features were badly overfitting, therefore I tried harder to understand why I had a much bigger gap than others.  And I found one reason, that leaders most probably dealt with weeks ago.  Using one of my last lightgbm runs, not the best on LB, but close to it, I got this.</p>\n\n<p>Before:</p>\n\n<p>CV 0.469 LB 0.903 GAP 0.434</p>\n\n<p>After:</p>\n\n<p>CV 0.527 LB 0.902 GAP 0.375</p>\n\n<p>I hope my feature evaluation will be more effective from now on.  </p>",
      "rawMarkdown": "My last features were badly overfitting, therefore I tried harder to understand why I had a much bigger gap than others.  And I found one reason, that leaders most probably dealt with weeks ago.  Using one of my last lightgbm runs, not the best on LB, but close to it, I got this.\n\nBefore:\n\nCV 0.469 LB 0.903 GAP 0.434\n\nAfter:\n\nCV 0.527 LB 0.902 GAP 0.375\n\nI hope my feature evaluation will be more effective from now on.  ",
      "votes": 3,
      "replies": [
        {
          "id": 423005,
          "postDate": "2018-11-17T08:39:45.710Z",
          "content": "<p>glad you found it. im still in the 0.4x gap.</p>",
          "rawMarkdown": "glad you found it. im still in the 0.4x gap.",
          "votes": 1
        },
        {
          "id": 423055,
          "postDate": "2018-11-17T11:07:30.547Z",
          "content": "<p>I could go below 0.9 LB with that kind of gap, there is hope therefore ;)</p>",
          "rawMarkdown": "I could go below 0.9 LB with that kind of gap, there is hope therefore ;)",
          "votes": 1
        },
        {
          "id": 423114,
          "postDate": "2018-11-17T14:17:19.407Z",
          "content": "<p>im rooting for you! most of my improvements are basically from your replies. I only started learning  ML this year and only learned about NNs in tutorials. I learned Gradient Boosting from <a href=\"/olivier\">@olivier</a> and a lot of feature engineering from you :)</p>",
          "rawMarkdown": "im rooting for you! most of my improvements are basically from your replies. I only started learning  ML this year and only learned about NNs in tutorials. I learned Gradient Boosting from @olivier and a lot of feature engineering from you :)",
          "votes": 1
        },
        {
          "id": 423120,
          "postDate": "2018-11-17T14:38:21.290Z",
          "content": "<p>Glad you find my posts useful!  </p>",
          "rawMarkdown": "Glad you find my posts useful!  ",
          "votes": 2
        },
        {
          "id": 423126,
          "postDate": "2018-11-17T15:00:16.400Z",
          "content": "<p>same for me... \n<a href=\"/cpmpml\">@cpmpml</a>\nI am learning ML on kaggle (coursera is not so much fun!)... what people do in theory with over fitting on boosted models (if adding more data is not an option) </p>\n\n<p>it's nice you found it, but on LB is still the same... will it add at the end, you think? </p>",
          "rawMarkdown": "same for me... \n@cpmpml\nI am learning ML on kaggle (coursera is not so much fun!)... what people do in theory with over fitting on boosted models (if adding more data is not an option) \n\nit's nice you found it, but on LB is still the same... will it add at the end, you think? \n",
          "votes": 1
        },
        {
          "id": 423174,
          "postDate": "2018-11-17T17:08:28.173Z",
          "content": "<blockquote>\n  <p>it's nice you found it, but on LB is still the same… will it add at the end, you think? </p>\n</blockquote>\n\n<p>Look at it the opposite way.  Suppose you have this:</p>\n\n<p>Before:</p>\n\n<p>CV 0.527 LB 0.902 GAP 0.375</p>\n\n<p>After:</p>\n\n<p>CV 0.469 LB 0.903 GAP 0.434</p>\n\n<p>You improve your CV by 0.058 but this does not show on the LB.  Would you keep the new code or the old one?</p>",
          "rawMarkdown": "&gt; it's nice you found it, but on LB is still the same… will it add at the end, you think? \n\nLook at it the opposite way.  Suppose you have this:\n\nBefore:\n\nCV 0.527 LB 0.902 GAP 0.375\n\nAfter:\n\nCV 0.469 LB 0.903 GAP 0.434\n\nYou improve your CV by 0.058 but this does not show on the LB.  Would you keep the new code or the old one?",
          "votes": 2
        },
        {
          "id": 423230,
          "postDate": "2018-11-17T18:26:07.750Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> and what would you say about : </p>\n\n<pre><code>CV 1.44 LB 1.57\n</code></pre>\n\n<p>and then</p>\n\n<pre><code>CV 0.99 LB 1.153\n</code></pre>\n\n<p>I think I'm on something interesting ;-)</p>",
          "rawMarkdown": "@cpmpml and what would you say about : \n\n    CV 1.44 LB 1.57\n\nand then\n\n    CV 0.99 LB 1.153\n\nI think I'm on something interesting ;-)",
          "votes": 4
        },
        {
          "id": 423244,
          "postDate": "2018-11-17T19:08:07.703Z",
          "content": "<p>I'd say its a small gap ;)  </p>\n\n<p>Keep pushing!</p>",
          "rawMarkdown": "I'd say its a small gap ;)  \n\nKeep pushing!",
          "votes": 2
        },
        {
          "id": 423262,
          "postDate": "2018-11-17T19:55:28.250Z",
          "content": "<p>it's good we have two submissions to choose from... but I got what you are talking about :-)</p>",
          "rawMarkdown": "it's good we have two submissions to choose from... but I got what you are talking about :-)",
          "votes": 1
        },
        {
          "id": 423268,
          "postDate": "2018-11-17T20:21:23.390Z",
          "content": "<p>My last submission has a CV of 0.528 and LB 0.983. GAP is 0.455. May I ask how did you eliminate those 'overfitting' features without submitting them? Looking distributions in train and test?</p>\n\n<p>Also this issue is not very clear to me. If you have two different models giving same Public LB but different CVs, why don't you choose the one having better CV? Wouldn't you think that, for the worst case, it will give you same private score?</p>",
          "rawMarkdown": "My last submission has a CV of 0.528 and LB 0.983. GAP is 0.455. May I ask how did you eliminate those 'overfitting' features without submitting them? Looking distributions in train and test?\n\nAlso this issue is not very clear to me. If you have two different models giving same Public LB but different CVs, why don't you choose the one having better CV? Wouldn't you think that, for the worst case, it will give you same private score?",
          "votes": 1
        },
        {
          "id": 423288,
          "postDate": "2018-11-17T21:24:35.313Z",
          "content": "<p>Yeah but it's a bit difficult to train. Hopefully I'll manage to lower the LB score ...</p>",
          "rawMarkdown": "Yeah but it's a bit difficult to train. Hopefully I'll manage to lower the LB score ...",
          "votes": 1
        },
        {
          "id": 423457,
          "postDate": "2018-11-18T09:48:04.510Z",
          "content": "<blockquote>\n  <p>If you have two different models giving same Public LB but different CVs, why don't you choose the one having better CV? </p>\n</blockquote>\n\n<p>The one that has best CV is overfiting to train data compared to the other one.  It generalizes less to new new data.  As we don't know if private test data is similar to public test data, keeping the one that generalizes best is safer IMHO.  </p>",
          "rawMarkdown": "&gt; If you have two different models giving same Public LB but different CVs, why don't you choose the one having better CV? \n\nThe one that has best CV is overfiting to train data compared to the other one.  It generalizes less to new new data.  As we don't know if private test data is similar to public test data, keeping the one that generalizes best is safer IMHO.  ",
          "votes": 2
        },
        {
          "id": 423615,
          "postDate": "2018-11-18T18:05:50.470Z",
          "content": "<p>I have 0.39 gap, but not so good score...</p>",
          "rawMarkdown": "I have 0.39 gap, but not so good score..."
        },
        {
          "id": 445280,
          "postDate": "2018-12-26T05:58:06.280Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> , can you share what you did here to reduce the CV - LB gap. Thanks in advance.</p>",
          "rawMarkdown": "@cpmpml , can you share what you did here to reduce the CV - LB gap. Thanks in advance."
        }
      ]
    },
    {
      "id": 413953,
      "postDate": "2018-11-01T21:22:00.190Z",
      "content": "<p>My latest model scores 0.92643 CV and 1.291 with a lot less overfitting. </p>",
      "rawMarkdown": "My latest model scores 0.92643 CV and 1.291 with a lot less overfitting. ",
      "votes": 3,
      "replies": [
        {
          "id": 414097,
          "postDate": "2018-11-02T06:08:37.710Z",
          "content": "<p>My CV LB gap is larger than yours.  This is with a single lgb model.  I am now looking at something different, will see if it helps...</p>\n\n<p>CV 0.570 LB 1.001 - GAP 0.43</p>\n\n<p>CV 0.622 LB 1.052 - GAP 0.43</p>\n\n<p>CV 0.733, LB 1.204 - GAP 0.47</p>\n\n<p>CV 0.902, LB 1.405 - GAP 0.50</p>",
          "rawMarkdown": "My CV LB gap is larger than yours.  This is with a single lgb model.  I am now looking at something different, will see if it helps...\n\nCV 0.570 LB 1.001 - GAP 0.43\n\nCV 0.622 LB 1.052 - GAP 0.43\n\nCV 0.733, LB 1.204 - GAP 0.47\n\nCV 0.902, LB 1.405 - GAP 0.50\n"
        }
      ]
    },
    {
      "id": 409833,
      "postDate": "2018-10-24T23:01:43.247Z",
      "content": "<p>1.07 CV / 1.53 LB here</p>\n\n<p>From the data note:</p>\n\n<blockquote>\n  <p>Crucially, the classifications will occur on a large test set, but the training data will be a small subset of the full data, and will also be a poor representation of the test set, to mimic the challenges we face observationally.</p>\n</blockquote>\n\n<p>Not much we can do I guess...?</p>",
      "rawMarkdown": "1.07 CV / 1.53 LB here\n\nFrom the data note:\n&gt; Crucially, the classifications will occur on a large test set, but the training data will be a small subset of the full data, and will also be a poor representation of the test set, to mimic the challenges we face observationally.\n\nNot much we can do I guess...?",
      "votes": 3
    },
    {
      "id": 427393,
      "postDate": "2018-11-25T12:33:10.820Z",
      "content": "<p>Two single models:\nCV: 0.46 LB: 0.931 (215 features)\nCV: 0.49 LB: 0.967 (57 features)</p>",
      "rawMarkdown": "Two single models:\nCV: 0.46 LB: 0.931 (215 features)\nCV: 0.49 LB: 0.967 (57 features)",
      "votes": 4
    },
    {
      "id": 425351,
      "postDate": "2018-11-21T13:32:57.920Z",
      "content": "<p>Inner-galactic model:  CV 0.091 with 141 features\nExtra-galactic model: CV 0.721 with 221 features\nOverall: CV 0.524, LB 0.889 - GAP 0.365</p>\n\n<p>There's still room for lowering the gap...</p>",
      "rawMarkdown": "Inner-galactic model:  CV 0.091 with 141 features\nExtra-galactic model: CV 0.721 with 221 features\nOverall: CV 0.524, LB 0.889 - GAP 0.365\n\nThere's still room for lowering the gap...",
      "votes": 4,
      "replies": [
        {
          "id": 425353,
          "postDate": "2018-11-21T13:35:08.273Z",
          "content": "<p>What kind of model?  lgb, NN, other? The gap depend son the type of model for me.</p>",
          "rawMarkdown": "What kind of model?  lgb, NN, other? The gap depend son the type of model for me."
        },
        {
          "id": 425365,
          "postDate": "2018-11-21T13:48:44.070Z",
          "content": "<p>lgb. Perhaps good NN model will get lower gap, but I love lgb :) What kind of your best model?</p>",
          "rawMarkdown": "lgb. Perhaps good NN model will get lower gap, but I love lgb :) What kind of your best model?"
        },
        {
          "id": 425381,
          "postDate": "2018-11-21T14:19:20.567Z",
          "content": "<p>My best is lgb, but for my team mates their best is NN.  I have similar cv/lb gap as you with lgb.</p>",
          "rawMarkdown": "My best is lgb, but for my team mates their best is NN.  I have similar cv/lb gap as you with lgb.",
          "votes": 5
        },
        {
          "id": 425531,
          "postDate": "2018-11-21T18:33:05.353Z",
          "content": "<p>Please pardon my noviceness, but 141 features for target 53?</p>",
          "rawMarkdown": "Please pardon my noviceness, but 141 features for target 53?"
        },
        {
          "id": 426127,
          "postDate": "2018-11-22T16:54:04.133Z",
          "content": "<p>Not only for target 53, but all galactic classes (6, 16, 53, 65 and 92).</p>",
          "rawMarkdown": "Not only for target 53, but all galactic classes (6, 16, 53, 65 and 92)."
        }
      ]
    },
    {
      "id": 413364,
      "postDate": "2018-10-31T19:44:54.810Z",
      "content": "<p>Current Lightgbm Model: CV: 0.598 , LB: 1.035.</p>",
      "rawMarkdown": "Current Lightgbm Model: CV: 0.598 , LB: 1.035.",
      "votes": 4,
      "replies": [
        {
          "id": 413366,
          "postDate": "2018-10-31T19:47:09.327Z",
          "content": "<p>That's a crazy local CV. I feel like I've missed something...</p>",
          "rawMarkdown": "That's a crazy local CV. I feel like I've missed something..."
        },
        {
          "id": 413374,
          "postDate": "2018-10-31T20:24:30.987Z",
          "content": "<p>Hi Faith, that's impressive cv score. Are you building two models on galactic and extragalactic level? Also, would you share how many features are you using now ?</p>",
          "rawMarkdown": "Hi Faith, that's impressive cv score. Are you building two models on galactic and extragalactic level? Also, would you share how many features are you using now ?"
        },
        {
          "id": 413380,
          "postDate": "2018-10-31T20:36:23.983Z",
          "content": "<p>I have very similar numbers with a single lgb model.  I also have models with lower CV, but higher LB.  One of the issue in this competition is to detect when fitting train data starts to be detrimental given test data is significantly different from it.</p>",
          "rawMarkdown": "I have very similar numbers with a single lgb model.  I also have models with lower CV, but higher LB.  One of the issue in this competition is to detect when fitting train data starts to be detrimental given test data is significantly different from it.",
          "votes": 3
        },
        {
          "id": 413392,
          "postDate": "2018-10-31T20:59:55.770Z",
          "content": "<p>Hi Indranil Bhattacharya. No, I have only one model. Also, I have 111 features in total.</p>\n\n<p>@CPMP I did not experience what you said yet. Still LB improves as CV improves.</p>",
          "rawMarkdown": "Hi Indranil Bhattacharya. No, I have only one model. Also, I have 111 features in total.\n\n@CPMP I did not experience what you said yet. Still LB improves as CV improves.",
          "votes": 5
        },
        {
          "id": 413397,
          "postDate": "2018-10-31T21:11:34.847Z",
          "content": "<p>Wow that's an amazing CV ! I'm really missing something here. You guys are really amazing.</p>",
          "rawMarkdown": "Wow that's an amazing CV ! I'm really missing something here. You guys are really amazing.",
          "votes": 1
        },
        {
          "id": 413485,
          "postDate": "2018-11-01T03:04:57.493Z",
          "content": "<p>If you don't mind sharing <a href=\"/fatihozturk\">@fatihozturk</a>, what kind of CV are you doing? I'm currently using straight class stratification, but was thinking of moving to something more versatile. My current best model scores 0.68777 CV (I haven't ran inference on test set for submission) but I fear I'm overfitting wildly. We'll soon see...</p>",
          "rawMarkdown": "If you don't mind sharing @fatihozturk, what kind of CV are you doing? I'm currently using straight class stratification, but was thinking of moving to something more versatile. My current best model scores 0.68777 CV (I haven't ran inference on test set for submission) but I fear I'm overfitting wildly. We'll soon see..."
        },
        {
          "id": 413591,
          "postDate": "2018-11-01T07:05:25.740Z",
          "content": "<p>I use simple stratified kfold. How do you fear without even submitting?</p>",
          "rawMarkdown": "I use simple stratified kfold. How do you fear without even submitting?",
          "votes": 1
        },
        {
          "id": 413646,
          "postDate": "2018-11-01T09:42:04.573Z",
          "content": "<p>@Olivier, I also think the same when I looked at top three's score :)</p>",
          "rawMarkdown": "@Olivier, I also think the same when I looked at top three's score :)",
          "votes": 2
        },
        {
          "id": 413678,
          "postDate": "2018-11-01T11:04:03.083Z",
          "content": "<p>@fatihöztürk, looks like you are a little bit over fitting, I get LB  ~0.96 with a similar or even worse CV.\nWhat do you do with the kfold models you get? throw them away as the theory suggest or ensemble the results? <br>\n<a href=\"/authman\">@authman</a>, Just submit. The LB doesn't bite. </p>",
          "rawMarkdown": "@fatihöztürk, looks like you are a little bit over fitting, I get LB  ~0.96 with a similar or even worse CV.\nWhat do you do with the kfold models you get? throw them away as the theory suggest or ensemble the results?  \n@authman, Just submit. The LB doesn't bite. "
        },
        {
          "id": 413681,
          "postDate": "2018-11-01T11:05:55.153Z",
          "content": "<p><a href=\"/fatihozturk\">@fatihozturk</a> I know I'm missing something on extra galactic objects. Galactic CV is at 0.203 but Extra Galactic is at 1.08...</p>\n\n<p>I'm not extraterrestrial enough I suppose !</p>",
          "rawMarkdown": "@fatihozturk I know I'm missing something on extra galactic objects. Galactic CV is at 0.203 but Extra Galactic is at 1.08...\n\nI'm not extraterrestrial enough I suppose !",
          "votes": 2
        },
        {
          "id": 413694,
          "postDate": "2018-11-01T11:33:24.633Z",
          "content": "<p><a href=\"/yuvalr\">@yuvalr</a> What do you mean by 'throw them away as the theory suggests'? I get test predictions for each fold and average them for submission. I don't know how to find the causes of such a sneaky overfitting if it exists.</p>\n\n<p><a href=\"/olivier\">@olivier</a> Since I did not work separately, I won't be able to help you :( But, I'm sure that you'll figure it out soon.</p>",
          "rawMarkdown": "@yuvalr What do you mean by 'throw them away as the theory suggests'? I get test predictions for each fold and average them for submission. I don't know how to find the causes of such a sneaky overfitting if it exists.\n\n@olivier Since I did not work separately, I won't be able to help you :( But, I'm sure that you'll figure it out soon.",
          "votes": 1
        },
        {
          "id": 413706,
          "postDate": "2018-11-01T11:50:15.797Z",
          "content": "<p><a href=\"/fatihozturk\">@fatihozturk</a> thanks. I'm doing the same. \nIn that case what is the CV you publish? The local score before averaging? </p>",
          "rawMarkdown": "@fatihozturk thanks. I'm doing the same. \nIn that case what is the CV you publish? The local score before averaging? "
        },
        {
          "id": 413717,
          "postDate": "2018-11-01T12:13:40.640Z",
          "content": "<p><a href=\"/yuvalr\">@yuvalr</a> Either the score of 'oof' predictions or mean of fold scores. You can think as the way done in Olivier's kernel.</p>",
          "rawMarkdown": "@yuvalr Either the score of 'oof' predictions or mean of fold scores. You can think as the way done in Olivier's kernel."
        }
      ]
    },
    {
      "id": 410895,
      "postDate": "2018-10-26T20:51:51.397Z",
      "content": "<p>For me also the difference is about 0.4 between Local CV and LB </p>",
      "rawMarkdown": "For me also the difference is about 0.4 between Local CV and LB ",
      "votes": 4
    },
    {
      "id": 410063,
      "postDate": "2018-10-25T09:59:53.530Z",
      "content": "<p>local 5f cv - 0.98\nlb - 1.505</p>\n\n<p>i can add more features that improve local cv but they all seem to hurt lb so far. probably to do with the many times mentioned here difference between train/test sets.</p>\n\n<p>as an aside. setting 0 for all class_99 preds seemed to yield a lb of 5.0 which kinda means (in my humble and probably uselss opinion) that this is probably the most critical part of the problem in terms of increasing lb score.</p>",
      "rawMarkdown": "local 5f cv - 0.98\nlb - 1.505\n\ni can add more features that improve local cv but they all seem to hurt lb so far. probably to do with the many times mentioned here difference between train/test sets.\n\nas an aside. setting 0 for all class_99 preds seemed to yield a lb of 5.0 which kinda means (in my humble and probably uselss opinion) that this is probably the most critical part of the problem in terms of increasing lb score.",
      "votes": 4
    },
    {
      "id": 409691,
      "postDate": "2018-10-24T17:50:21.107Z",
      "content": "<p>1.12 CV / 1.481 LB for me. IIRC, this is using different <code>class_99</code> predictions for galactic/extra-galactic objects.</p>",
      "rawMarkdown": "1.12 CV / 1.481 LB for me. IIRC, this is using different `class_99` predictions for galactic/extra-galactic objects.",
      "votes": 4,
      "replies": [
        {
          "id": 409700,
          "postDate": "2018-10-24T17:59:45.157Z",
          "content": "<p>Thanks Branden, very interesting.</p>\n\n<p>I use different predictions for galactic and extra-galactic objects as well. My model seems to really overfit then...</p>",
          "rawMarkdown": "Thanks Branden, very interesting.\n\nI use different predictions for galactic and extra-galactic objects as well. My model seems to really overfit then..."
        },
        {
          "id": 409859,
          "postDate": "2018-10-25T00:56:41.577Z",
          "content": "<p>To clarify, the function I use to score my CV includes a constant prediction of 1/9 for class_99 which artificially lowers my CV score some. I'd guess without that it'd probably be somewhere near the 0.98 that you're getting.</p>",
          "rawMarkdown": "To clarify, the function I use to score my CV includes a constant prediction of 1/9 for class_99 which artificially lowers my CV score some. I'd guess without that it'd probably be somewhere near the 0.98 that you're getting.",
          "votes": 2
        },
        {
          "id": 409933,
          "postDate": "2018-10-25T04:58:01.583Z",
          "content": "<p>Thanks for the clarification Branden ! I'm somewhat relieved.</p>",
          "rawMarkdown": "Thanks for the clarification Branden ! I'm somewhat relieved."
        }
      ]
    },
    {
      "id": 427890,
      "postDate": "2018-11-26T10:39:26.347Z",
      "content": "<p>Single NN Model:\nCV: 0.686  LB: 1.051\nGap: 0.365</p>",
      "rawMarkdown": "Single NN Model:\nCV: 0.686  LB: 1.051\nGap: 0.365",
      "votes": 1
    },
    {
      "id": 416093,
      "postDate": "2018-11-06T07:27:22.437Z",
      "content": "<p>When you say \"CV\", what does it stand for? multi log loss in lgb? weighted log loss used in this competition?</p>",
      "rawMarkdown": "When you say \"CV\", what does it stand for? multi log loss in lgb? weighted log loss used in this competition?",
      "votes": 1,
      "replies": [
        {
          "id": 416117,
          "postDate": "2018-11-06T08:05:20.057Z",
          "content": "<p><a href=\"/onodera\">@onodera</a>, as far as I'm concerned it is weighted log loss as defined for the competition.</p>",
          "rawMarkdown": "@onodera, as far as I'm concerned it is weighted log loss as defined for the competition.",
          "votes": 2
        }
      ]
    },
    {
      "id": 415903,
      "postDate": "2018-11-05T22:05:22.523Z",
      "content": "<p>For me it is CV 0.525, LB 0.989, and an enormous gap of 0.464. Previously it used to be around 0.43 - 0.45. I am using Olivier's method for assigning the probabilities of class 99</p>",
      "rawMarkdown": "For me it is CV 0.525, LB 0.989, and an enormous gap of 0.464. Previously it used to be around 0.43 - 0.45. I am using Olivier's method for assigning the probabilities of class 99\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 416027,
          "postDate": "2018-11-06T03:57:17.320Z",
          "content": "<p>My gap also increases as the model gets better, and I had to find additional ways of reducing overfitting.</p>",
          "rawMarkdown": "My gap also increases as the model gets better, and I had to find additional ways of reducing overfitting.",
          "votes": 1
        },
        {
          "id": 416824,
          "postDate": "2018-11-07T10:25:03.720Z",
          "content": "<p>My gap is a constant 0.3 even when the model improves (it was 0.4 in the past and improved when the model improved). </p>",
          "rawMarkdown": "My gap is a constant 0.3 even when the model improves (it was 0.4 in the past and improved when the model improved). ",
          "votes": 2
        },
        {
          "id": 417029,
          "postDate": "2018-11-07T16:26:43.353Z",
          "content": "<p>0.45 for me, my model is still in the infancy period</p>",
          "rawMarkdown": "0.45 for me, my model is still in the infancy period"
        }
      ]
    },
    {
      "id": 431180,
      "postDate": "2018-12-01T18:32:44.477Z",
      "content": "<p>single lgb CV 0.431, LB 0.790 Gap 0.359</p>\n\n<p>Gap decreases as model improves.</p>",
      "rawMarkdown": "single lgb CV 0.431, LB 0.790 Gap 0.359\n\nGap decreases as model improves.",
      "votes": 2,
      "replies": [
        {
          "id": 436065,
          "postDate": "2018-12-09T13:48:55.893Z",
          "content": "<p>What loss do you used in your Local CV? Multi log loss in lgb?  Or weighted log loss by <a href=\"/olivier\">@olivier</a> ?</p>",
          "rawMarkdown": "What loss do you used in your Local CV? Multi log loss in lgb?  Or weighted log loss by @olivier ?",
          "votes": -1
        },
        {
          "id": 436082,
          "postDate": "2018-12-09T14:41:22.860Z",
          "content": "<p>Neither ones.  I use a loss that mimics the competition metric.</p>",
          "rawMarkdown": "Neither ones.  I use a loss that mimics the competition metric."
        },
        {
          "id": 436086,
          "postDate": "2018-12-09T14:52:07.377Z",
          "content": "<p>We've reached CV around 0.4 but the gap is still big. I feel like there is something behind.</p>",
          "rawMarkdown": "We've reached CV around 0.4 but the gap is still big. I feel like there is something behind."
        },
        {
          "id": 436259,
          "postDate": "2018-12-10T02:11:05.507Z",
          "content": "<p>Got it, thanks</p>",
          "rawMarkdown": "Got it, thanks"
        }
      ]
    },
    {
      "id": 430676,
      "postDate": "2018-11-30T20:42:31.707Z",
      "content": "<p>Single LGB:\nCV: 0.454 LB: 0.875 gap: 0.42 (my gap still big)</p>",
      "rawMarkdown": "Single LGB:\nCV: 0.454 LB: 0.875 gap: 0.42 (my gap still big)",
      "votes": 2
    },
    {
      "id": 427252,
      "postDate": "2018-11-25T00:40:27.933Z",
      "content": "<p>CV: 0.5\nLB: 1.2\nLooks like I'm gonna have to figure out some stuff here...</p>",
      "rawMarkdown": "CV: 0.5\nLB: 1.2\nLooks like I'm gonna have to figure out some stuff here...",
      "votes": 2
    },
    {
      "id": 423010,
      "postDate": "2018-11-17T08:50:53.677Z",
      "content": "<p>NN with 400 features: <br>\nCV 0.532 LB 1.017 - GAP 0.485  </p>\n\n<p>LGBM with same features: <br>\nCV 0.591 LB 1.127 - GAP 0.536  </p>\n\n<p>My CV-LB gap is bigger than others. I removed features which seems diffrent between train and test, but didn't improve LB score.</p>",
      "rawMarkdown": "NN with 400 features:  \nCV 0.532 LB 1.017 - GAP 0.485  \n\nLGBM with same features:   \nCV 0.591 LB 1.127 - GAP 0.536  \n\nMy CV-LB gap is bigger than others. I removed features which seems diffrent between train and test, but didn't improve LB score.",
      "votes": 2,
      "replies": [
        {
          "id": 423049,
          "postDate": "2018-11-17T10:53:00.867Z",
          "content": "<p>Try tuning lgb parameters to be more conservative.  Start with what is given at the bottom of this page: <a href=\"https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html\">https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html</a></p>",
          "rawMarkdown": "Try tuning lgb parameters to be more conservative.  Start with what is given at the bottom of this page: https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html",
          "votes": 2
        },
        {
          "id": 423110,
          "postDate": "2018-11-17T14:08:23.927Z",
          "content": "<p>Thanks CPMP, I will try parameters tuning!</p>",
          "rawMarkdown": "Thanks CPMP, I will try parameters tuning!"
        },
        {
          "id": 423261,
          "postDate": "2018-11-17T19:54:06.020Z",
          "content": "<p>thank you!</p>",
          "rawMarkdown": "thank you!"
        },
        {
          "id": 423872,
          "postDate": "2018-11-19T07:04:43.807Z",
          "content": "<p>400 features seems a lot for 7k examples...</p>",
          "rawMarkdown": "400 features seems a lot for 7k examples..."
        },
        {
          "id": 423977,
          "postDate": "2018-11-19T11:13:24.143Z",
          "content": "<p>How many features do you use for your best model? <br>\nI have created about 8k features, and selected 400 for my models based on feature importances. Indeed, I need to refine the number of features and feature selection method.</p>",
          "rawMarkdown": "How many features do you use for your best model?  \nI have created about 8k features, and selected 400 for my models based on feature importances. Indeed, I need to refine the number of features and feature selection method."
        },
        {
          "id": 423978,
          "postDate": "2018-11-19T11:14:18.150Z",
          "content": "<p>About 110, but I am trying to reduce this number.</p>",
          "rawMarkdown": "About 110, but I am trying to reduce this number.",
          "votes": 1
        }
      ]
    },
    {
      "id": 415389,
      "postDate": "2018-11-05T03:04:38.217Z",
      "content": "<p>Single lgb model, 100 features:\nCV 0.520 LB 0.957 - GAP 0.437</p>\n\n<p>The gap is stable as shown by previous results:</p>\n\n<p>CV 0.570 LB 1.001 - GAP 0.431</p>\n\n<p>CV 0.622 LB 1.052 - GAP 0.430</p>\n\n<p>Seems deep learning models have a smaller gap.</p>",
      "rawMarkdown": "Single lgb model, 100 features:\nCV 0.520 LB 0.957 - GAP 0.437\n\nThe gap is stable as shown by previous results:\n\nCV 0.570 LB 1.001 - GAP 0.431\n\nCV 0.622 LB 1.052 - GAP 0.430\n\nSeems deep learning models have a smaller gap.",
      "votes": 2,
      "replies": [
        {
          "id": 416754,
          "postDate": "2018-11-07T08:19:03.440Z",
          "content": "<p>Wow, 0.52 CV is awesome. It seems you have found another important feature like mjd_diff of detected ones and I'm extremely curious about it :) I've just got back from my vacation and it seems I have a lot to do...</p>",
          "rawMarkdown": "Wow, 0.52 CV is awesome. It seems you have found another important feature like mjd_diff of detected ones and I'm extremely curious about it :) I've just got back from my vacation and it seems I have a lot to do...",
          "votes": 1
        },
        {
          "id": 416764,
          "postDate": "2018-11-07T08:46:37.360Z",
          "content": "<p>My current best is 0.504 CV, 0.943 LB, single lgb model</p>\n\n<p>I did find some useful features indeed, but I am not the only one if I look at the LB!</p>\n\n<p>Anyway, I am now trying to build some NN models as I find it harder and harder to improve my lgb model.  And we know ensembling can help a lot, right?</p>",
          "rawMarkdown": "My current best is 0.504 CV, 0.943 LB, single lgb model\n\nI did find some useful features indeed, but I am not the only one if I look at the LB!\n\nAnyway, I am now trying to build some NN models as I find it harder and harder to improve my lgb model.  And we know ensembling can help a lot, right?",
          "votes": 1
        },
        {
          "id": 416766,
          "postDate": "2018-11-07T08:57:15.560Z",
          "content": "<p>I doubt the first three scores come from a single model. Especially the top two. Yeah, ensembling with some NN will help you a lot :)</p>",
          "rawMarkdown": "I doubt the first three scores come from a single model. Especially the top two. Yeah, ensembling with some NN will help you a lot :)",
          "votes": 1
        },
        {
          "id": 416773,
          "postDate": "2018-11-07T09:01:58.333Z",
          "content": "<p>@CPMP I wish I had your CV. my current lgb model has CV of 0.582 and the corresponding LB is 0.950. So maybe top3 guys have something like your CV and my LB-CV difference.</p>",
          "rawMarkdown": "@CPMP I wish I had your CV. my current lgb model has CV of 0.582 and the corresponding LB is 0.950. So maybe top3 guys have something like your CV and my LB-CV difference.",
          "votes": 3
        },
        {
          "id": 416780,
          "postDate": "2018-11-07T09:17:37.430Z",
          "content": "<p>Ahmet,  I wish I had you CV-LB gap indeed ;)</p>\n\n<p>And I am NOW trying to build NN models....</p>",
          "rawMarkdown": "Ahmet,  I wish I had you CV-LB gap indeed ;)\n\nAnd I am NOW trying to build NN models....",
          "votes": 3
        },
        {
          "id": 416791,
          "postDate": "2018-11-07T09:34:30.450Z",
          "content": "<p>if I see a significant jump in your score in a week, I will switch to NN too:)</p>",
          "rawMarkdown": "if I see a significant jump in your score in a week, I will switch to NN too:)",
          "votes": 1
        },
        {
          "id": 416794,
          "postDate": "2018-11-07T09:42:54.663Z",
          "content": "<p>My LB score is a lgb model still, I am just starting with NN.  First NN sub had a LB score of 2.26 ...</p>",
          "rawMarkdown": "My LB score is a lgb model still, I am just starting with NN.  First NN sub had a LB score of 2.26 ...",
          "votes": 1
        },
        {
          "id": 416800,
          "postDate": "2018-11-07T09:56:11.160Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a>, the good thing is you can only improve this score :)</p>",
          "rawMarkdown": "@cpmpml, the good thing is you can only improve this score :)",
          "votes": 2
        },
        {
          "id": 416808,
          "postDate": "2018-11-07T10:00:09.553Z",
          "content": "<blockquote>\n  <p>you can only improve this score :)</p>\n</blockquote>\n\n<p>Don't overestimate my DL skills ;)</p>",
          "rawMarkdown": "&gt; you can only improve this score :)\n\nDon't overestimate my DL skills ;)"
        },
        {
          "id": 416911,
          "postDate": "2018-11-07T12:31:55.477Z",
          "content": "<p>@AhmetErdem do you think that you have an intentional attempt to have such a low CV-LB gap compared to us, or it is just luck?</p>",
          "rawMarkdown": "@AhmetErdem do you think that you have an intentional attempt to have such a low CV-LB gap compared to us, or it is just luck?"
        },
        {
          "id": 416967,
          "postDate": "2018-11-07T14:41:59.557Z",
          "content": "<p><a href=\"/fatihozturk\">@fatihozturk</a> I worked on it. maybe there is also some luck factor.\n@CPMP just tried to use deep NN with only dense features. My LB-CV increased from 0.37 to 0.42. I don't know why some people observed the opposite.</p>",
          "rawMarkdown": "@fatihozturk I worked on it. maybe there is also some luck factor.\n@CPMP just tried to use deep NN with only dense features. My LB-CV increased from 0.37 to 0.42. I don't know why some people observed the opposite."
        },
        {
          "id": 416982,
          "postDate": "2018-11-07T15:09:50.383Z",
          "content": "<p>@AhmetErdem Is using NN the reason for your last improvement? Just looking for additional motivation to use NN ;)</p>",
          "rawMarkdown": "@AhmetErdem Is using NN the reason for your last improvement? Just looking for additional motivation to use NN ;)",
          "votes": -1
        },
        {
          "id": 417008,
          "postDate": "2018-11-07T15:59:16.297Z",
          "content": "<p>While my NN scored much worse than my LGB, it could still contribute a bit by blending.</p>",
          "rawMarkdown": "While my NN scored much worse than my LGB, it could still contribute a bit by blending.",
          "votes": 1
        }
      ]
    },
    {
      "id": 413589,
      "postDate": "2018-11-01T06:58:53.483Z",
      "content": "<p>My best lightgbm model scores 0.814CV (no gal/extra gal model) and LB 1.296. Gap is 0.482...</p>",
      "rawMarkdown": "My best lightgbm model scores 0.814CV (no gal/extra gal model) and LB 1.296. Gap is 0.482...",
      "votes": 2
    },
    {
      "id": 411621,
      "postDate": "2018-10-28T15:28:31.817Z",
      "content": "<p>0.85 CV / 1.365 LB in my single lightgbm model using Olivier's method for class 99.</p>",
      "rawMarkdown": "0.85 CV / 1.365 LB in my single lightgbm model using Olivier's method for class 99.",
      "votes": 2,
      "replies": [
        {
          "id": 412287,
          "postDate": "2018-10-29T23:13:47.310Z",
          "content": "<p>Now: 0.746 CV / 1.324 LB</p>\n\n<p>Seems the gap between my validation and my leaderboard is increasing.</p>",
          "rawMarkdown": "Now: 0.746 CV / 1.324 LB\n\nSeems the gap between my validation and my leaderboard is increasing."
        },
        {
          "id": 412525,
          "postDate": "2018-10-30T10:40:41.743Z",
          "content": "<p>Are you using \"hostgal_specz\"? if yes, this might be the issue.\nOtherwise, you might be over fitting or you treat class 99 wrongly (try to give it a constant value)</p>",
          "rawMarkdown": "Are you using \"hostgal_specz\"? if yes, this might be the issue.\nOtherwise, you might be over fitting or you treat class 99 wrongly (try to give it a constant value)",
          "votes": 1
        }
      ]
    },
    {
      "id": 410208,
      "postDate": "2018-10-25T15:43:47.237Z",
      "content": "<p>I am also using a single model:</p>\n\n<ul>\n<li>Local CV = 0.845</li>\n<li>LB = 1.335</li>\n</ul>\n\n<p>Seems like the CV/LB gap is pretty stable across the participants at 0.4 - 0.5 :) I guess dealing with class 99 and different data distribution in train/test would be the only way to significantly reduce it. </p>",
      "rawMarkdown": "I am also using a single model:\n\n - Local CV = 0.845\n - LB = 1.335\n\nSeems like the CV/LB gap is pretty stable across the participants at 0.4 - 0.5 :) I guess dealing with class 99 and different data distribution in train/test would be the only way to significantly reduce it. ",
      "votes": 2
    },
    {
      "id": 410194,
      "postDate": "2018-10-25T15:23:20.027Z",
      "content": "<p>Single model,  CV 0.796 LB 1.246   class_99 as in Olivier kernel.</p>",
      "rawMarkdown": "Single model,  CV 0.796 LB 1.246   class_99 as in Olivier kernel.",
      "votes": 2
    },
    {
      "id": 410166,
      "postDate": "2018-10-25T14:48:18.900Z",
      "content": "<p>Single (first) lightgbm model, CV 0.90, LB 1.405 constant value for class 99.</p>",
      "rawMarkdown": "Single (first) lightgbm model, CV 0.90, LB 1.405 constant value for class 99.",
      "votes": 2,
      "replies": [
        {
          "id": 410183,
          "postDate": "2018-10-25T15:13:22.283Z",
          "content": "<p>After just 2 submissions ? you must be a competition GM ;-)</p>",
          "rawMarkdown": "After just 2 submissions ? you must be a competition GM ;-)",
          "votes": 1
        },
        {
          "id": 410203,
          "postDate": "2018-10-25T15:37:28.950Z",
          "content": "<p>First sub was to check weights, see: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67194#409633\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67194#409633</a>  ;)</p>\n\n<p>And I find Jack's single sub to be way more impressive than mine!</p>",
          "rawMarkdown": "First sub was to check weights, see: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67194#409633  ;)\n\nAnd I find Jack's single sub to be way more impressive than mine!",
          "votes": 2
        },
        {
          "id": 410655,
          "postDate": "2018-10-26T12:14:56.500Z",
          "content": "<p>local CV 0.733, LB 1.204</p>\n\n<p>As pointed out by Nikita the gap seems to be rather constant.  It means we won't get LB score below 0.5 ;)</p>",
          "rawMarkdown": "local CV 0.733, LB 1.204\n\nAs pointed out by Nikita the gap seems to be rather constant.  It means we won't get LB score below 0.5 ;)",
          "votes": 2
        }
      ]
    },
    {
      "id": 410026,
      "postDate": "2018-10-25T08:32:49.963Z",
      "content": "<p>For a single LGBM model:</p>\n\n<ul>\n<li><p>5 fold CV:  0.90</p></li>\n<li><p>LB:  1.342 </p></li>\n</ul>\n\n<p>CV and and LB have followed the same path from the start but I think a) I need to work on the ghost class 99 and b) there are some bad overfitting pitfall are out there...I just have fallen into some of them along the way...;) <br>\nI also used different methods for predicting prob(class_99) but just about 1-2 of my ideas worked out of my 9 submissions.</p>",
      "rawMarkdown": "For a single LGBM model:\n\n+ 5 fold CV:  0.90\n\n+ LB:  1.342 \n \nCV and and LB have followed the same path from the start but I think a) I need to work on the ghost class 99 and b) there are some bad overfitting pitfall are out there...I just have fallen into some of them along the way...;)  \nI also used different methods for predicting prob(class_99) but just about 1-2 of my ideas worked out of my 9 submissions.",
      "votes": 2
    },
    {
      "id": 430937,
      "postDate": "2018-12-01T09:26:50.257Z",
      "content": "<p>I've tried my first MLP with same features of lgbm. Model structure is completly same with the one in public kernels.</p>\n\n<p>It's CV: 0.82 LB: 1.143. Gap is so good but CV :(. Also, tried simple weighted average with LGBM and it neither improves CV nor LB.</p>",
      "rawMarkdown": "I've tried my first MLP with same features of lgbm. Model structure is completly same with the one in public kernels.\n\nIt's CV: 0.82 LB: 1.143. Gap is so good but CV :(. Also, tried simple weighted average with LGBM and it neither improves CV nor LB.",
      "replies": [
        {
          "id": 430976,
          "postDate": "2018-12-01T11:08:28.053Z",
          "content": "<p>MLP with same features as lgb is also way worse for me.  </p>",
          "rawMarkdown": "MLP with same features as lgb is also way worse for me.  \n"
        },
        {
          "id": 430989,
          "postDate": "2018-12-01T11:25:02.527Z",
          "content": "<p>So you generated completely new features for MLP ? Also what about blending? I remember that you said blending is still better than a single lgbm. Do you blend all classes with same weight?</p>",
          "rawMarkdown": "So you generated completely new features for MLP ? Also what about blending? I remember that you said blending is still better than a single lgbm. Do you blend all classes with same weight?"
        },
        {
          "id": 431034,
          "postDate": "2018-12-01T13:41:03.730Z",
          "content": "<p>I tuned features for MLP, as I tuned features for lgb.  Can't say more before end of competition ;)  I blend all classes with same weight so far.  </p>",
          "rawMarkdown": "I tuned features for MLP, as I tuned features for lgb.  Can't say more before end of competition ;)  I blend all classes with same weight so far.  ",
          "votes": 1
        }
      ]
    },
    {
      "id": 428785,
      "postDate": "2018-11-27T22:19:42.203Z",
      "content": "<p>Maybe we could think about this differently?</p>\n\n<p>I have two models, one for Galactic and another for Extra-Galactic. Same features, both LGBM. They give:\nGalactic CV: 0.129\nExtra-Gal CV: 0.836</p>\n\n<p>The training data is split 2325/5523 Galactic/Extra-Galactic. Combining the CV scores of my two models using a weighting based on the training data gives:\nCombined CV: 0.627\nLB: 1.034\nGAP: 0.407</p>\n\n<p>That GAP is in the normal range according to this thread. However, the test data contains a higher proportion of Extra-Galactic sources than the training data. The mix is 390510/3102380. Using a weighting based on the test data gives:\nCombined CV: 0.757\nLB: 1.034\nGAP: 0.277</p>\n\n<p>Hopefully by removing the effect of different proportions of Galactic/Extra-Galactic sources between train and test, we can see how much of the GAP is down to other effects like class 99 and other train/test data differences.</p>",
      "rawMarkdown": "Maybe we could think about this differently?\n\nI have two models, one for Galactic and another for Extra-Galactic. Same features, both LGBM. They give:\nGalactic CV: 0.129\nExtra-Gal CV: 0.836\n\nThe training data is split 2325/5523 Galactic/Extra-Galactic. Combining the CV scores of my two models using a weighting based on the training data gives:\nCombined CV: 0.627\nLB: 1.034\nGAP: 0.407\n\nThat GAP is in the normal range according to this thread. However, the test data contains a higher proportion of Extra-Galactic sources than the training data. The mix is 390510/3102380. Using a weighting based on the test data gives:\nCombined CV: 0.757\nLB: 1.034\nGAP: 0.277\n\nHopefully by removing the effect of different proportions of Galactic/Extra-Galactic sources between train and test, we can see how much of the GAP is down to other effects like class 99 and other train/test data differences.\n",
      "replies": [
        {
          "id": 428800,
          "postDate": "2018-11-27T22:54:37.780Z",
          "content": "<p>Hi Andy, please check competition metric carefully. Difference in number of samples for each class is not really important since each class is divided by its total number of examples. So for both train and test, Galactic vs Extra-Galactic is 5/16 vs 11/16.</p>",
          "rawMarkdown": "Hi Andy, please check competition metric carefully. Difference in number of samples for each class is not really important since each class is divided by its total number of examples. So for both train and test, Galactic vs Extra-Galactic is 5/16 vs 11/16.",
          "votes": 2
        },
        {
          "id": 428877,
          "postDate": "2018-11-28T02:40:55.897Z",
          "content": "<p>The galactic and extragalactic objects are really completely separate in this problem. Your classifier should be assigning 0 probability to extragalactic targets if the photos is 0 and vice versa. I'm not sure how weighting differently would affect that. It seems like you just changed how you calculated your CV score without affecting the predictions.</p>",
          "rawMarkdown": "The galactic and extragalactic objects are really completely separate in this problem. Your classifier should be assigning 0 probability to extragalactic targets if the photos is 0 and vice versa. I'm not sure how weighting differently would affect that. It seems like you just changed how you calculated your CV score without affecting the predictions.",
          "votes": 2
        },
        {
          "id": 429433,
          "postDate": "2018-11-28T20:53:47.713Z",
          "content": "<p>Thanks for the replies and saving me from my own wrong-thinking!</p>\n\n<p>Ahmet, sorry, yes you're correct. It  seems I skimmed the evaluation instructions too quickly. Time for me to review more carefully.</p>\n\n<p>Kyle, yes, my galactic model gives 0 probability to extra-galactic classes and vice versa. I was trying to short-cut combining the CVs from my models... but I took a wrong turn. I'll just have to combine the OOF predictions properly next time to double check any short-cuts.</p>",
          "rawMarkdown": "Thanks for the replies and saving me from my own wrong-thinking!\n\nAhmet, sorry, yes you're correct. It  seems I skimmed the evaluation instructions too quickly. Time for me to review more carefully.\n\nKyle, yes, my galactic model gives 0 probability to extra-galactic classes and vice versa. I was trying to short-cut combining the CVs from my models... but I took a wrong turn. I'll just have to combine the OOF predictions properly next time to double check any short-cuts.",
          "votes": 1
        }
      ]
    },
    {
      "id": 428422,
      "postDate": "2018-11-27T08:54:00.167Z",
      "content": "<p>Update - Single NN Model:\nCV: 0.666 LB: 0.995\nGap: 0.329</p>",
      "rawMarkdown": "Update - Single NN Model:\nCV: 0.666 LB: 0.995\nGap: 0.329",
      "replies": [
        {
          "id": 428467,
          "postDate": "2018-11-27T10:30:24.357Z",
          "content": "<p>Nice score! Would you mind to tell us what kind of NN it is? a single MLP or something special?</p>",
          "rawMarkdown": "Nice score! Would you mind to tell us what kind of NN it is? a single MLP or something special?"
        },
        {
          "id": 428493,
          "postDate": "2018-11-27T11:35:09.003Z",
          "content": "<p>It is a single MLP based model with some modifications...</p>",
          "rawMarkdown": "It is a single MLP based model with some modifications...",
          "votes": 1
        },
        {
          "id": 428725,
          "postDate": "2018-11-27T18:58:46.200Z",
          "content": "<p>thanks!</p>",
          "rawMarkdown": "thanks!"
        }
      ]
    },
    {
      "id": 427417,
      "postDate": "2018-11-25T13:25:59.640Z",
      "content": "<p>Guys, could you help to understand this topic for new members... Thanks.</p>",
      "rawMarkdown": "Guys, could you help to understand this topic for new members... Thanks.",
      "replies": [
        {
          "id": 427433,
          "postDate": "2018-11-25T13:49:52.660Z",
          "content": "<p>This is about sharing the LB score of models as well as the cross validation score of the same model.  Olivier was probably surprised by the large gap between these two values, and asked if others had the same.  It turns out that yes, the gap is large for all.  The smallest gap that was reported is 0.34 I think.</p>",
          "rawMarkdown": "This is about sharing the LB score of models as well as the cross validation score of the same model.  Olivier was probably surprised by the large gap between these two values, and asked if others had the same.  It turns out that yes, the gap is large for all.  The smallest gap that was reported is 0.34 I think.",
          "votes": 2
        },
        {
          "id": 427434,
          "postDate": "2018-11-25T13:50:33.177Z",
          "content": "<p>Please have a look at this link: <a href=\"https://www.kaggle.com/questions-and-answers/61785\">https://www.kaggle.com/questions-and-answers/61785</a> </p>",
          "rawMarkdown": "Please have a look at this link: https://www.kaggle.com/questions-and-answers/61785 ",
          "votes": 3
        }
      ]
    },
    {
      "id": 423638,
      "postDate": "2018-11-18T19:13:03.747Z",
      "content": "<p>Since we're talking about overfitting in here a bit,  how do your fold training errors compare to your fold validation errors?\nI typically get around 0.15 for my training loss and 0.5+ for my validation losses. This is for a 5-fold LGBM, not separating galactic from extragalactic.\nI think I might benefit from some form of regularization given the gap between training and validation.</p>",
      "rawMarkdown": "Since we're talking about overfitting in here a bit,  how do your fold training errors compare to your fold validation errors?\nI typically get around 0.15 for my training loss and 0.5+ for my validation losses. This is for a 5-fold LGBM, not separating galactic from extragalactic.\nI think I might benefit from some form of regularization given the gap between training and validation.",
      "replies": [
        {
          "id": 423654,
          "postDate": "2018-11-18T19:47:17.020Z",
          "content": "<p>After fixing my overfitting, I get:</p>\n\n<p>Train 0.287 cv 0.525</p>\n\n<p>Before I had</p>\n\n<p>Train 0.215 cv 0.468</p>",
          "rawMarkdown": "After fixing my overfitting, I get:\n\nTrain 0.287 cv 0.525\n\nBefore I had\n\nTrain 0.215 cv 0.468",
          "votes": 1
        },
        {
          "id": 423749,
          "postDate": "2018-11-19T01:29:41.543Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> \nI though \"fixing over fitting\" is when train loss is tiny bit lower then CV</p>",
          "rawMarkdown": "@cpmpml \nI though \"fixing over fitting\" is when train loss is tiny bit lower then CV"
        },
        {
          "id": 423871,
          "postDate": "2018-11-19T07:03:54.023Z",
          "content": "<p>Sure, there is some overfititng left.</p>",
          "rawMarkdown": "Sure, there is some overfititng left."
        }
      ]
    },
    {
      "id": 423422,
      "postDate": "2018-11-18T08:00:20.010Z",
      "content": "<p>My single lgb: CV 0.494, LB 1.01</p>\n\n<p>overfitting?</p>",
      "rawMarkdown": "My single lgb: CV 0.494, LB 1.01\n\noverfitting?",
      "replies": [
        {
          "id": 423429,
          "postDate": "2018-11-18T08:11:14.140Z",
          "content": "<p>Who exactly knows ? </p>\n\n<p>Our local CVs miss class 99 and from that <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67194\">discussion</a> and in particular <a href=\"https://www.kaggle.com/titericz\">Giba's</a> probing feedback class_99 proba would be around 0.17 for extragalactic and 0.017 for galactic.</p>\n\n<p>This makes quite a difference from CV to LB :)</p>",
          "rawMarkdown": "Who exactly knows ? \n\nOur local CVs miss class 99 and from that [discussion](https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67194) and in particular [Giba's](https://www.kaggle.com/titericz) probing feedback class_99 proba would be around 0.17 for extragalactic and 0.017 for galactic.\n\nThis makes quite a difference from CV to LB :)\n",
          "votes": 1
        },
        {
          "id": 423733,
          "postDate": "2018-11-19T00:27:04.350Z",
          "content": "<p>Compared to other guys, my CV score is better than LB🤔</p>",
          "rawMarkdown": "Compared to other guys, my CV score is better than LB🤔"
        }
      ]
    },
    {
      "id": 421434,
      "postDate": "2018-11-15T02:15:41.957Z",
      "content": "<p>Has anyone been monitoring CV score on galactic and extragalactic objects separately? My cv score on the galactic part is 0.05 so I am guessing there is overfitting there. Anyone with a similar experience?</p>",
      "rawMarkdown": "Has anyone been monitoring CV score on galactic and extragalactic objects separately? My cv score on the galactic part is 0.05 so I am guessing there is overfitting there. Anyone with a similar experience?",
      "replies": [
        {
          "id": 421591,
          "postDate": "2018-11-15T07:00:07.343Z",
          "content": "<p>I have 2 nns: galactic cv = 0.12, extragalactic cv=0.91, lb score of them is 1.06</p>",
          "rawMarkdown": "I have 2 nns: galactic cv = 0.12, extragalactic cv=0.91, lb score of them is 1.06"
        },
        {
          "id": 421658,
          "postDate": "2018-11-15T08:35:35.677Z",
          "content": "<p>My best model so far is LGB wtih 0.075 galactic and 0.833 extragalactic. LB 1.028.</p>",
          "rawMarkdown": "My best model so far is LGB wtih 0.075 galactic and 0.833 extragalactic. LB 1.028."
        }
      ]
    },
    {
      "id": 415943,
      "postDate": "2018-11-06T00:07:08.617Z",
      "content": "<p>I'm trying to understand what magic all of you are doing here, these 0.5x in CV are very crazy to me, I need to figure out this magic very soon...</p>\n\n<p>Maybe the problem is my poorly knowledge in physics and astronomy...</p>",
      "rawMarkdown": "I'm trying to understand what magic all of you are doing here, these 0.5x in CV are very crazy to me, I need to figure out this magic very soon...\n\nMaybe the problem is my poorly knowledge in physics and astronomy...",
      "replies": [
        {
          "id": 416026,
          "postDate": "2018-11-06T03:56:11.020Z",
          "content": "<p>On my side, no magic, no physics trick either, just hand crafted feature engineering and good cv setting.  I am also using Olivier's way of computing class_99 probabilities, except I normalize its mean to 0.18 rather than 0.14.</p>",
          "rawMarkdown": "On my side, no magic, no physics trick either, just hand crafted feature engineering and good cv setting.  I am also using Olivier's way of computing class_99 probabilities, except I normalize its mean to 0.18 rather than 0.14."
        },
        {
          "id": 416915,
          "postDate": "2018-11-07T12:36:11.977Z",
          "content": "<p>@CPMP Is there a specific reason for switching to 0.18? I also use Olivier's method.</p>",
          "rawMarkdown": "@CPMP Is there a specific reason for switching to 0.18? I also use Olivier's method."
        },
        {
          "id": 416933,
          "postDate": "2018-11-07T13:07:27.273Z",
          "content": "<p>I did LB probing, see <a href=\"https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data/comments#410554\">https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data/comments#410554</a></p>",
          "rawMarkdown": "I did LB probing, see https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data/comments#410554"
        },
        {
          "id": 421611,
          "postDate": "2018-11-15T07:27:40.540Z",
          "content": "<p>@CPMP, do you have hand crafted features based on meta_data? or meta_data + time_series combo?</p>",
          "rawMarkdown": "@CPMP, do you have hand crafted features based on meta_data? or meta_data + time_series combo?"
        },
        {
          "id": 422988,
          "postDate": "2018-11-17T07:31:52.920Z",
          "content": "<p>I use all available data ;)</p>",
          "rawMarkdown": "I use all available data ;)"
        },
        {
          "id": 423066,
          "postDate": "2018-11-17T11:58:51.203Z",
          "content": "<p>@CPMP did you use astronomical equations for feature interactions? or mainly your skills in interpreting the relationship between data?</p>",
          "rawMarkdown": "@CPMP did you use astronomical equations for feature interactions? or mainly your skills in interpreting the relationship between data?"
        },
        {
          "id": 423107,
          "postDate": "2018-11-17T13:54:45.210Z",
          "content": "<p>I have not used any astronomical knowledge because I have none ;)</p>",
          "rawMarkdown": "I have not used any astronomical knowledge because I have none ;)"
        },
        {
          "id": 423370,
          "postDate": "2018-11-18T03:38:29.597Z",
          "content": "<p>I have tried reading papers and articles about it, and tried transformations but none improved my CV.</p>",
          "rawMarkdown": "I have tried reading papers and articles about it, and tried transformations but none improved my CV."
        }
      ]
    },
    {
      "id": 414140,
      "postDate": "2018-11-02T07:39:52.827Z",
      "content": "<p>I've dropped the \"hostgal_specz\" column, get a local CV of around 0.9 with LB 1.5, either my model overfits or my class 99 needs some extra work.</p>",
      "rawMarkdown": "I've dropped the \"hostgal_specz\" column, get a local CV of around 0.9 with LB 1.5, either my model overfits or my class 99 needs some extra work."
    },
    {
      "id": 413584,
      "postDate": "2018-11-01T06:55:00.057Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 409971,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2018-10-25T06:14:38.850000",
      "content": "<p>Galactic models CV: 0.21 +/- 0.01</p>\n\n<p>Extragalactic models CV: 0.88 +/- 0.02</p>\n\n<p>Local CV ~0.70</p>\n\n<p>LB: 1.155</p>\n\n<p>Overfitting region reached - many \"smart\" features decreasing local CV score by 0.02 increase LB score by 0.10 (the consequence of the fact, that the training data is a poor representation of the test set).\nClass_99 distribution set by intuition, further probing (using multilogloss probabilities instead of a coin or a dice) improved the score by 0.02 only.</p>\n\n<p>EDIT: To be clear, +/- 0.01 means here the range of CV of a few similar but different models, not an error of CV of one model</p>",
      "votes": 11,
      "replies": [
        {
          "id": 409973,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-25T06:24:06.143000",
          "content": "<p>Thanks for sharing <a href=\"/sionek\">@sionek</a> ! very impressive results. Looking at the difference between LB/CV it seems you a have very good intuition of what class 99 should be :) </p>\n\n<p>I hope I'll be able to come anywhere close to your CV / LB score...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 411818,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-10-29T04:23:41.723000",
          "content": "<blockquote>\n  <p>Overfitting region reached</p>\n</blockquote>\n\n<p>I think I found it too ;)</p>\n\n<p>You seem to have escaped yours!</p>\n\n<p>CV LB gap is decreasing a bit as my models are better.  Comparing those with same class_99 computation (probability that it is not another class):</p>\n\n<p>CV 0.622 LB 1.052  - GAP 0.43</p>\n\n<p>CV 0.733, LB 1.204 - GAP 0.47</p>\n\n<p>CV 0.902, LB 1.405 - GAP 0.50</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 411914,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-10-29T07:47:44.807000",
          "content": "<blockquote>\n  <p>I think I found it too ;)</p>\n  \n  <p>You seem to have escaped yours!</p>\n</blockquote>\n\n<p>@CPMP,  If you are in a hopeless local minimum, it is time to trust yourself, not your local CV ;) Remember, that one of the best gains in this competition (~0.30) was obtained by removing hostgal_specz from the feature set in spite of CV score worse by ~0.07. It was simply logic. If you invent a new feature you think it should work, check it out in your submission, even if your local CV says \"forget it\". </p>\n\n<p>In many previous competitions, public LB scores vere only a small additional part of \"total\" validation, the main role was played by local CV, proportionally to the volume of the samples. In this competition, the roles are reversed - overproportionally, because of a significant difference between train and test files.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 411936,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-10-29T08:41:36.493000",
          "content": "<p>Thanks for the encouragements.</p>\n\n<blockquote>\n  <p>hostgal_specz </p>\n</blockquote>\n\n<p>I did not include it as it was not much present in test data.  </p>\n\n<blockquote>\n  <p>f you invent a new feature you think it should work, check it out in your submission, even if your CV says \"forget it\".</p>\n</blockquote>\n\n<p>You are right in general, train is not representative of test, hence local CV is misleading.  But LB probing is against my nature, I'll resist a bit before giving into it.  And also because upload from my home machine is really slow unfortunately.</p>\n\n<p>I have an idea to try, but it will take a while to run it, and I'm traveling this week.  I don't expect progress before week end therefore.  I hope I won't be too far behind by then!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 412022,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-10-29T12:08:48.323000",
          "content": "<p>@CPMP, So, an approach without LB probing. I have found, that if you are in a hopeless local minimum (few unsuccesful submissions),  it makes no sense to submit another \"better\" results, if your local CV improvement is smaller than your CV error (in my case ~0.07).  I have escaped from my local minimum (LB=1.155) submitting results of local CV=0.62 (local CV improvement equal 0.08, LB gain equal 0.12). </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 423706,
      "author_name": "Kyle Boone",
      "author_url": "",
      "post_date": "2018-11-18T22:36:25.497000",
      "content": "<p>My current best submission has CV 0.49, LB 0.83 for a gap of 0.34. So there are ways of lowering the gap!</p>",
      "votes": 7,
      "replies": [
        {
          "id": 423721,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-18T23:31:18.903000",
          "content": "<p>Thanks, you are confirming there is room for more effective feature engineering.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 425521,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-11-21T18:16:09.417000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> </p>\n\n<p>We have a larger gap (0.4 -- 0.45). Do you think the gap is due\n1. over fitting (so parameters of the classifier) \n2. redundant features\n3. inappropriate features\n4. different distribution of classes in test (besides class 99, we still do not know if the other classes distribution is the same -- we maybe better ask)\n3. there is also class 99  (I doubt anyone has dealt with it yet), and it will add to the gap</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 425567,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-21T19:36:25.907000",
          "content": "<p>Blonde, it is hard to answer without knowing more about how you model the problem.  </p>\n\n<p>But I am not sure a larger gap is an issue.  As you noticed, when I decreased my gap I did not improve my LB.  </p>\n\n<p>I think it helps having a smaller gap; but it is hard to tell until we see private LB scores.  </p>\n\n<p>In the meantime, I would recommend you work on feature selection.  I improved my CV and LB score by removing features.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 425664,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-11-21T23:33:47.207000",
          "content": "<p>Thank you. The feature selection process is going ok, it just takes time to calculate submission. </p>\n\n<p>BTW, you noticed before that 400 features is too much for 7k, from your experience what it the range of optimal ratio between number of samples and number of features? (I have another project with signals, 400 samples, 20 features, but that signals I can easily augment and increase the number in a few times. So it would be nice to know the approximate merit which is ok, like the range of ratios between number of features and number of samples ) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426012,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-22T12:39:31.847000",
          "content": "<p>It depends on the model you are using, and how much regularization you use.  I reacted because 400 features for 7k examples means a ration of example per feature quite low, and prone to overfiting with lgb or xgb models.  General linear models with proper regularization are less sensitive to that.   And I know for sure that it is possible to get below 0.9 on the Lb with less than 199 features ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426082,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-11-22T15:20:25.447000",
          "content": "<p>Thank you. I am free in choosing my model for that project, I used xgboost, but will look into linear models with regularization then. I have only 400 samples, which is not much... and selected 18 features. I'll try to minimize that number. I also tried extracting features with fully conv nets with large augmentation, but it has not work out so far. xgboost at least gave some results, but far from brilliant. 400 for 7k is about the same as 18 for 400, that's why I asked what is the practical merit for \"too many features\"...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426115,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-22T16:35:42.510000",
          "content": "<p>With 400 samples I would go with a glm.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 428707,
          "author_name": "Kyle Boone",
          "author_url": "",
          "post_date": "2018-11-27T18:33:43.050000",
          "content": "<p>Update on this: my best single model is now at 0.43 CV / 0.776 LB.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 428713,
          "author_name": "KALE",
          "author_url": "",
          "post_date": "2018-11-27T18:43:55.560000",
          "content": "<p>May I ask how many features you have for this one?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 428727,
          "author_name": "Angus Chang",
          "author_url": "",
          "post_date": "2018-11-27T19:05:24.947000",
          "content": "<p><a href=\"/kyleboone\">@kyleboone</a> that is really amazing CV score and gap!! Just curious, is your current LB score a ensemble result?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 428731,
          "author_name": "Kyle Boone",
          "author_url": "",
          "post_date": "2018-11-27T19:19:56.927000",
          "content": "<p>My latest model seems to have some issues with overfitting. I ensembled its predictions with an older one that performs worse but doesn't overfit as much to get the LB score.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 429865,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-11-29T13:28:47.593000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> </p>\n\n<p>LGB, around 100 features, single model, CV 0.52 -- LB 0.979, GAP 0.45 </p>\n\n<p>Any hints on lowering the gap? Which direction to look at: \nparameters of the classifier \nredundant features\ninappropriate features\nsomething else... \nit's my first GBM experience, any hints would be nice</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 429923,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-29T14:55:26.997000",
          "content": "<blockquote>\n  <p>Any hints on lowering the gap?</p>\n</blockquote>\n\n<p>Priority is to lower the LB score ;)</p>\n\n<p>Your gap seems similar to what I had when my LB score was around 0.98, there is no real issue here I think.  Your number of feature does not seem too high either.  Just keep adding good features..</p>\n\n<p>My gap decreased drastically when my model became better.  Latest run is CV 0.420, LB 0.794, Gap 0.374.  I can't disclose what made it move from 0.9x LB to 0.7x before competition end ;)  And I think the two leader have even better single models than that.</p>\n\n<p>To answer some of your questions:</p>\n\n<ul>\n<li><p>Redundant features are not an issue for lgb or xgboost.  </p></li>\n<li><p>Parameter tuning makes sense.  I selected conservative settings given there is a risk of overfiting.  I may try less conservative settings at a point though.</p></li>\n<li><p>Inappropriate features.  Definitely an issue.  Watch for features that lower your CV score without lowering the LB score.</p></li>\n</ul>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 438829,
          "author_name": "dylonLL",
          "author_url": "",
          "post_date": "2018-12-14T08:52:10.367000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a>, from 09x to 0.7x was due to one method/feature?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438842,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-14T09:17:21.750000",
          "content": "<p>You will know in few days.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 426576,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-11-23T13:23:44.750000",
      "content": "<p>We've achieved a significant milestone by combining my best lgb model with ideas from my team mates.</p>\n\n<p>single run: CV 0.431, LB 0.801, Gap 0.370</p>\n\n<p>@yuval_r prediction that final best score can be below 0.7 may be true.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 426578,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-11-23T13:29:52.323000",
          "content": "<p>Well done ! that's impressive.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 426825,
          "author_name": "dylonLL",
          "author_url": "",
          "post_date": "2018-11-24T00:08:18.103000",
          "content": "<p>wow that's great</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426971,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-24T09:11:23.760000",
          "content": "<p>Yes, I'm surprised NNs aren't better, but we are working on it ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426985,
          "author_name": "dylonLL",
          "author_url": "",
          "post_date": "2018-11-24T09:59:40.820000",
          "content": "<p>that is true to me as well. when my LGBM was at 1.3 LB my NN was already at 1.1X with same features. But when my LGB reached .9X my NN stayed at 1.1X with same features</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427281,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2018-11-25T04:20:12.737000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> May I ask you to show us your team's confusion matrix?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427375,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-25T11:49:52.293000",
          "content": "<p>I'll create a post for it as I don't see how to upload an image to a comment, see <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72613\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72613</a></p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 415424,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2018-11-05T05:21:55.630000",
      "content": "<p>My latest CV (single lightgbm) is 0.659 and LB 1.082. I think I finally found what I missed.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 415427,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-05T05:25:58.273000",
          "content": "<blockquote>\n  <p>I finally found what I missed</p>\n</blockquote>\n\n<p>Indeed, you now have a gap close to 0.43 ;)</p>\n\n<p>More seriously, this is good and steady progress!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415430,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-11-05T05:34:46.917000",
          "content": "<p>Thanks <a href=\"/cpmpml\">@cpmpml</a> ! I'm still far from top 10, all the more you're still with a single model...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 415537,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-05T09:40:05.277000",
          "content": "<blockquote>\n  <p>you're still with a single model</p>\n</blockquote>\n\n<p>Your best LB score comes from a single model too unless mistaken...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415547,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-11-05T09:52:33.157000",
          "content": "<p>I should have been more explict... you're still 0.12 ahead which is a lot with a single model.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 415573,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-05T10:41:21.340000",
          "content": "<p>Leaders are way ahead of me as well, and they improve everyday too!  We will learn very interesting things after competition end.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416188,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2018-11-06T11:09:48.373000",
          "content": "<p><a href=\"/ogrellier\">@ogrellier</a>, this is a nice progress!</p>\n\n<p>I've been stuck for the last few days with a local CV around 0.75 ~ 0.76. I have a feeling that I have too many features that are not that powerful (200+), but so far I couldn't come up with smarter features to improve on that. So I am still trying to find that thing you guys have already discovered :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416198,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-11-06T11:41:18.333000",
          "content": "<p>Thanks <a href=\"/kozodoi\">@kozodoi</a>, depending on what objective/loss you use you may find my latest kernel update useful or not :) </p>\n\n<p><a href=\"https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data\">https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data</a></p>\n\n<p>If I'm right you may land in top 10, hopefully I'm right ;-)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 416204,
          "author_name": "João Pedro Peinado",
          "author_url": "",
          "post_date": "2018-11-06T11:54:36.530000",
          "content": "<p>Thanks Olivier, you are really amazing!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416394,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2018-11-06T16:02:30.757000",
          "content": "<p>Thanks <a href=\"/ogrellier\">@ogrellier</a>, your kernels are always so helpful!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 414643,
      "author_name": "Jiazhen Xi",
      "author_url": "",
      "post_date": "2018-11-03T09:01:51.827000",
      "content": "<p>With same features</p>\n\n<ul>\n<li>lgb: CV 0.98 LB 1.454 gap 0.474  </li>\n<li>mlp: CV 0.82 LB 1.148 gap 0.328  </li>\n</ul>",
      "votes": 5,
      "replies": [
        {
          "id": 414644,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-03T09:12:35.180000",
          "content": "<p>Seems I should move to mlp!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 414645,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-11-03T09:29:15.267000",
          "content": "<p>You certainly want ;-)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 414646,
          "author_name": "Jiazhen Xi",
          "author_url": "",
          "post_date": "2018-11-03T09:30:39.427000",
          "content": "<p>Definitely!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 414654,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-03T09:50:10.243000",
          "content": "<p>I will, don't worry !   I'm convinced deep learning is better on this competition as it can capture the overall shape of the light curve, something that is hard to achieve with feature engineering.  But I first want to squeeze out every bit I can from feature engineering as it makes me understand data better.  And I have not yet used XGBoost, which is stil my favorite algorithm.  I need to give it a chance!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 414675,
          "author_name": "takuoko",
          "author_url": "",
          "post_date": "2018-11-03T11:12:26.520000",
          "content": "<p>I feel it is strange that mlp is much better than lgbm.\nNN's merit is rnn for time siries features or cnn for image features.\nSo I don't know why mlp work well.\nanyway, I will try mlp.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 414690,
          "author_name": "Jiazhen Xi",
          "author_url": "",
          "post_date": "2018-11-03T11:58:15.667000",
          "content": "<p>Yes, I have the same feeling! <br>\nI guess RNN/CNN may be more proper than MLP for ts data here but found it troublesome to implement the preprocessing part (maybe reshape or padding or interpolating?) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 414697,
          "author_name": "takuoko",
          "author_url": "",
          "post_date": "2018-11-03T12:13:51.407000",
          "content": "<p>This kernel may be useful when getting ts features.\n<a href=\"https://www.kaggle.com/scirpus/predict-by-row-then-average\">https://www.kaggle.com/scirpus/predict-by-row-then-average</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 414702,
          "author_name": "Jiazhen Xi",
          "author_url": "",
          "post_date": "2018-11-03T12:22:35.473000",
          "content": "<p>Thanks for guiding! <br>\nThis brings me some new ideas:)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 421041,
          "author_name": "Fatih Öztürk",
          "author_url": "",
          "post_date": "2018-11-14T14:02:59.290000",
          "content": "<p>It's weird. When I tried MLP structure in public kernel with same features of my lgbm, CV goes up to 0.82 from 0.58. Really wonder how did you manage to get lower CV score..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421049,
          "author_name": "Kain",
          "author_url": "",
          "post_date": "2018-11-14T14:10:08.130000",
          "content": "<p>Same here, for me many features that improve LB/CV in my NN model are useless when used in LGBM. It seems in this competition, Boosting and NN need considerably different set of features to work best. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 421085,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-11-14T15:06:30.383000",
          "content": "<p>My best GBM gets around 0.65 CV / 1.085 LB and my best NN with the same features is around 0.75 CV / 1.200 LB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421092,
          "author_name": "Fatih Öztürk",
          "author_url": "",
          "post_date": "2018-11-14T15:12:46.247000",
          "content": "<p>It's good to hear that I'm not alone then.. May I ask is your current LB score blend of those two?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421108,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-11-14T15:45:42.210000",
          "content": "<p>I did not stack these 2 alone, but a 50/50 blend was really poor around 1.10</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421139,
          "author_name": "Max Halford",
          "author_url": "",
          "post_date": "2018-11-14T16:24:31.063000",
          "content": "<p>I guess having a good blend depends on what you do with <code>class_99</code>. You can include the <code>class_99</code> columns into the blend, or you can blend the other columns and then recompute <code>class_99</code> with the new probabilities.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421146,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-14T16:38:18.090000",
          "content": "<p>I confirm what others say, using a MLP on similar features than lgb leads to higher CV and LB score.  Moreover, many features lead to overfitting with the MLP.   The good news is that a weighted average of the MLP and lgb models is better than lgb alone still.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 421302,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-14T21:49:37.170000",
          "content": "<blockquote>\n  <p>I did not stack these 2 alone,</p>\n</blockquote>\n\n<p>You are stacking already?</p>\n\n<p>Beware, I'm  starting with XGBoost ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421596,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-11-15T07:04:09.110000",
          "content": "<p>Yes, I'm on kernel only... and running a bit out of ideas !</p>\n\n<p>XGBoost is a very good option IMHO :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 410032,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-10-25T08:46:49.367000",
      "content": "<p>I love it when I see we are now dealing with galactic models and extra galactic models!  Machine Learning moving to entire new horizons! :)</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 409843,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2018-10-24T23:37:38.203000",
      "content": "<p>Galactic model CV: 0.23</p>\n\n<p>Extra-Galactic model CV: 0.99</p>\n\n<p>LB: 1.53</p>\n\n<p>I believe local CV is low because class_99 distribution is a mystery of the universe :-)</p>",
      "votes": 6,
      "replies": [
        {
          "id": 409931,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-25T04:57:00.740000",
          "content": "<p>Thanks Giba ! does that mean your LB score is a model blend ? or just Giba's magic ;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 411643,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2018-10-28T16:37:00.997000",
          "content": "<p>It means I splited the data set in two parts and trained two completely separated models: Galactic and ExtraGalactic</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 411662,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-28T18:06:06.977000",
          "content": "<p>@Giba I asked this because you talked about a LB score of 1.53 so I guess you meant 1.35 then ;-)</p>\n\n<p>I tried to split into galactic and extra galactic as well but did not get any sort of improvement with that setup. I must be missing something here and surely how CPMP went to 1.05 that fast ...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 411683,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-10-28T19:19:38.733000",
          "content": "<blockquote>\n  <p>how CPMP went to 1.05 that fast ...</p>\n</blockquote>\n\n<p>I'm afraid I have reached a local optimum now.  I will need something really new to move under 1.0...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 411692,
          "author_name": "yvan",
          "author_url": "",
          "post_date": "2018-10-28T19:59:44.710000",
          "content": "<p>splitting the models gave me a bit of a boost (~0.01). tbh i never thought id be above <a href=\"/ogrellier\">@ogrellier</a> on the lb lol. usually my new ideas come from his kernels. i honestly feel like im flying by the 'skin of my teeth' here.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 412523,
          "author_name": "dylonLL",
          "author_url": "",
          "post_date": "2018-10-30T10:34:52.307000",
          "content": "<p>@CPMP, are you still using constant for class 99?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421301,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-14T21:48:21.190000",
          "content": "<p>I don't use a constant for class_99, read what I shared ;)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 422993,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-11-17T08:00:51.947000",
      "content": "<p>My last features were badly overfitting, therefore I tried harder to understand why I had a much bigger gap than others.  And I found one reason, that leaders most probably dealt with weeks ago.  Using one of my last lightgbm runs, not the best on LB, but close to it, I got this.</p>\n\n<p>Before:</p>\n\n<p>CV 0.469 LB 0.903 GAP 0.434</p>\n\n<p>After:</p>\n\n<p>CV 0.527 LB 0.902 GAP 0.375</p>\n\n<p>I hope my feature evaluation will be more effective from now on.  </p>",
      "votes": 3,
      "replies": [
        {
          "id": 423005,
          "author_name": "dylonLL",
          "author_url": "",
          "post_date": "2018-11-17T08:39:45.710000",
          "content": "<p>glad you found it. im still in the 0.4x gap.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 423055,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-17T11:07:30.547000",
          "content": "<p>I could go below 0.9 LB with that kind of gap, there is hope therefore ;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 423114,
          "author_name": "dylonLL",
          "author_url": "",
          "post_date": "2018-11-17T14:17:19.407000",
          "content": "<p>im rooting for you! most of my improvements are basically from your replies. I only started learning  ML this year and only learned about NNs in tutorials. I learned Gradient Boosting from <a href=\"/olivier\">@olivier</a> and a lot of feature engineering from you :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 423120,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-17T14:38:21.290000",
          "content": "<p>Glad you find my posts useful!  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 423126,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-11-17T15:00:16.400000",
          "content": "<p>same for me... \n<a href=\"/cpmpml\">@cpmpml</a>\nI am learning ML on kaggle (coursera is not so much fun!)... what people do in theory with over fitting on boosted models (if adding more data is not an option) </p>\n\n<p>it's nice you found it, but on LB is still the same... will it add at the end, you think? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 423174,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-17T17:08:28.173000",
          "content": "<blockquote>\n  <p>it's nice you found it, but on LB is still the same… will it add at the end, you think? </p>\n</blockquote>\n\n<p>Look at it the opposite way.  Suppose you have this:</p>\n\n<p>Before:</p>\n\n<p>CV 0.527 LB 0.902 GAP 0.375</p>\n\n<p>After:</p>\n\n<p>CV 0.469 LB 0.903 GAP 0.434</p>\n\n<p>You improve your CV by 0.058 but this does not show on the LB.  Would you keep the new code or the old one?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 423230,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-11-17T18:26:07.750000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> and what would you say about : </p>\n\n<pre><code>CV 1.44 LB 1.57\n</code></pre>\n\n<p>and then</p>\n\n<pre><code>CV 0.99 LB 1.153\n</code></pre>\n\n<p>I think I'm on something interesting ;-)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 423244,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-17T19:08:07.703000",
          "content": "<p>I'd say its a small gap ;)  </p>\n\n<p>Keep pushing!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 423262,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-11-17T19:55:28.250000",
          "content": "<p>it's good we have two submissions to choose from... but I got what you are talking about :-)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 423268,
          "author_name": "Fatih Öztürk",
          "author_url": "",
          "post_date": "2018-11-17T20:21:23.390000",
          "content": "<p>My last submission has a CV of 0.528 and LB 0.983. GAP is 0.455. May I ask how did you eliminate those 'overfitting' features without submitting them? Looking distributions in train and test?</p>\n\n<p>Also this issue is not very clear to me. If you have two different models giving same Public LB but different CVs, why don't you choose the one having better CV? Wouldn't you think that, for the worst case, it will give you same private score?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 423288,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-11-17T21:24:35.313000",
          "content": "<p>Yeah but it's a bit difficult to train. Hopefully I'll manage to lower the LB score ...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 423457,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-18T09:48:04.510000",
          "content": "<blockquote>\n  <p>If you have two different models giving same Public LB but different CVs, why don't you choose the one having better CV? </p>\n</blockquote>\n\n<p>The one that has best CV is overfiting to train data compared to the other one.  It generalizes less to new new data.  As we don't know if private test data is similar to public test data, keeping the one that generalizes best is safer IMHO.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 423615,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-11-18T18:05:50.470000",
          "content": "<p>I have 0.39 gap, but not so good score...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445280,
          "author_name": "Subrahmanyam V",
          "author_url": "",
          "post_date": "2018-12-26T05:58:06.280000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> , can you share what you did here to reduce the CV - LB gap. Thanks in advance.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 413953,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2018-11-01T21:22:00.190000",
      "content": "<p>My latest model scores 0.92643 CV and 1.291 with a lot less overfitting. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 414097,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-02T06:08:37.710000",
          "content": "<p>My CV LB gap is larger than yours.  This is with a single lgb model.  I am now looking at something different, will see if it helps...</p>\n\n<p>CV 0.570 LB 1.001 - GAP 0.43</p>\n\n<p>CV 0.622 LB 1.052 - GAP 0.43</p>\n\n<p>CV 0.733, LB 1.204 - GAP 0.47</p>\n\n<p>CV 0.902, LB 1.405 - GAP 0.50</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 409833,
      "author_name": "Ganfear",
      "author_url": "",
      "post_date": "2018-10-24T23:01:43.247000",
      "content": "<p>1.07 CV / 1.53 LB here</p>\n\n<p>From the data note:</p>\n\n<blockquote>\n  <p>Crucially, the classifications will occur on a large test set, but the training data will be a small subset of the full data, and will also be a poor representation of the test set, to mimic the challenges we face observationally.</p>\n</blockquote>\n\n<p>Not much we can do I guess...?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 427393,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2018-11-25T12:33:10.820000",
      "content": "<p>Two single models:\nCV: 0.46 LB: 0.931 (215 features)\nCV: 0.49 LB: 0.967 (57 features)</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 425351,
      "author_name": "nyanp",
      "author_url": "",
      "post_date": "2018-11-21T13:32:57.920000",
      "content": "<p>Inner-galactic model:  CV 0.091 with 141 features\nExtra-galactic model: CV 0.721 with 221 features\nOverall: CV 0.524, LB 0.889 - GAP 0.365</p>\n\n<p>There's still room for lowering the gap...</p>",
      "votes": 4,
      "replies": [
        {
          "id": 425353,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-21T13:35:08.273000",
          "content": "<p>What kind of model?  lgb, NN, other? The gap depend son the type of model for me.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 425365,
          "author_name": "nyanp",
          "author_url": "",
          "post_date": "2018-11-21T13:48:44.070000",
          "content": "<p>lgb. Perhaps good NN model will get lower gap, but I love lgb :) What kind of your best model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 425381,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-21T14:19:20.567000",
          "content": "<p>My best is lgb, but for my team mates their best is NN.  I have similar cv/lb gap as you with lgb.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 425531,
          "author_name": "Tord Malmgren",
          "author_url": "",
          "post_date": "2018-11-21T18:33:05.353000",
          "content": "<p>Please pardon my noviceness, but 141 features for target 53?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426127,
          "author_name": "nyanp",
          "author_url": "",
          "post_date": "2018-11-22T16:54:04.133000",
          "content": "<p>Not only for target 53, but all galactic classes (6, 16, 53, 65 and 92).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 413364,
      "author_name": "Fatih Öztürk",
      "author_url": "",
      "post_date": "2018-10-31T19:44:54.810000",
      "content": "<p>Current Lightgbm Model: CV: 0.598 , LB: 1.035.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 413366,
          "author_name": "Max Halford",
          "author_url": "",
          "post_date": "2018-10-31T19:47:09.327000",
          "content": "<p>That's a crazy local CV. I feel like I've missed something...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 413374,
          "author_name": "Indranil Bhattacharya",
          "author_url": "",
          "post_date": "2018-10-31T20:24:30.987000",
          "content": "<p>Hi Faith, that's impressive cv score. Are you building two models on galactic and extragalactic level? Also, would you share how many features are you using now ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 413380,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-10-31T20:36:23.983000",
          "content": "<p>I have very similar numbers with a single lgb model.  I also have models with lower CV, but higher LB.  One of the issue in this competition is to detect when fitting train data starts to be detrimental given test data is significantly different from it.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 413392,
          "author_name": "Fatih Öztürk",
          "author_url": "",
          "post_date": "2018-10-31T20:59:55.770000",
          "content": "<p>Hi Indranil Bhattacharya. No, I have only one model. Also, I have 111 features in total.</p>\n\n<p>@CPMP I did not experience what you said yet. Still LB improves as CV improves.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 413397,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-31T21:11:34.847000",
          "content": "<p>Wow that's an amazing CV ! I'm really missing something here. You guys are really amazing.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 413485,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-11-01T03:04:57.493000",
          "content": "<p>If you don't mind sharing <a href=\"/fatihozturk\">@fatihozturk</a>, what kind of CV are you doing? I'm currently using straight class stratification, but was thinking of moving to something more versatile. My current best model scores 0.68777 CV (I haven't ran inference on test set for submission) but I fear I'm overfitting wildly. We'll soon see...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 413591,
          "author_name": "Fatih Öztürk",
          "author_url": "",
          "post_date": "2018-11-01T07:05:25.740000",
          "content": "<p>I use simple stratified kfold. How do you fear without even submitting?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 413646,
          "author_name": "Fatih Öztürk",
          "author_url": "",
          "post_date": "2018-11-01T09:42:04.573000",
          "content": "<p>@Olivier, I also think the same when I looked at top three's score :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 413678,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-11-01T11:04:03.083000",
          "content": "<p>@fatihöztürk, looks like you are a little bit over fitting, I get LB  ~0.96 with a similar or even worse CV.\nWhat do you do with the kfold models you get? throw them away as the theory suggest or ensemble the results? <br>\n<a href=\"/authman\">@authman</a>, Just submit. The LB doesn't bite. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 413681,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-11-01T11:05:55.153000",
          "content": "<p><a href=\"/fatihozturk\">@fatihozturk</a> I know I'm missing something on extra galactic objects. Galactic CV is at 0.203 but Extra Galactic is at 1.08...</p>\n\n<p>I'm not extraterrestrial enough I suppose !</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 413694,
          "author_name": "Fatih Öztürk",
          "author_url": "",
          "post_date": "2018-11-01T11:33:24.633000",
          "content": "<p><a href=\"/yuvalr\">@yuvalr</a> What do you mean by 'throw them away as the theory suggests'? I get test predictions for each fold and average them for submission. I don't know how to find the causes of such a sneaky overfitting if it exists.</p>\n\n<p><a href=\"/olivier\">@olivier</a> Since I did not work separately, I won't be able to help you :( But, I'm sure that you'll figure it out soon.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 413706,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-11-01T11:50:15.797000",
          "content": "<p><a href=\"/fatihozturk\">@fatihozturk</a> thanks. I'm doing the same. \nIn that case what is the CV you publish? The local score before averaging? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 413717,
          "author_name": "Fatih Öztürk",
          "author_url": "",
          "post_date": "2018-11-01T12:13:40.640000",
          "content": "<p><a href=\"/yuvalr\">@yuvalr</a> Either the score of 'oof' predictions or mean of fold scores. You can think as the way done in Olivier's kernel.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 410895,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2018-10-26T20:51:51.397000",
      "content": "<p>For me also the difference is about 0.4 between Local CV and LB </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 410063,
      "author_name": "yvan",
      "author_url": "",
      "post_date": "2018-10-25T09:59:53.530000",
      "content": "<p>local 5f cv - 0.98\nlb - 1.505</p>\n\n<p>i can add more features that improve local cv but they all seem to hurt lb so far. probably to do with the many times mentioned here difference between train/test sets.</p>\n\n<p>as an aside. setting 0 for all class_99 preds seemed to yield a lb of 5.0 which kinda means (in my humble and probably uselss opinion) that this is probably the most critical part of the problem in terms of increasing lb score.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 409691,
      "author_name": "Branden Murray",
      "author_url": "",
      "post_date": "2018-10-24T17:50:21.107000",
      "content": "<p>1.12 CV / 1.481 LB for me. IIRC, this is using different <code>class_99</code> predictions for galactic/extra-galactic objects.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 409700,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-24T17:59:45.157000",
          "content": "<p>Thanks Branden, very interesting.</p>\n\n<p>I use different predictions for galactic and extra-galactic objects as well. My model seems to really overfit then...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 409859,
          "author_name": "Branden Murray",
          "author_url": "",
          "post_date": "2018-10-25T00:56:41.577000",
          "content": "<p>To clarify, the function I use to score my CV includes a constant prediction of 1/9 for class_99 which artificially lowers my CV score some. I'd guess without that it'd probably be somewhere near the 0.98 that you're getting.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 409933,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-25T04:58:01.583000",
          "content": "<p>Thanks for the clarification Branden ! I'm somewhat relieved.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 427890,
      "author_name": "mrxnew",
      "author_url": "",
      "post_date": "2018-11-26T10:39:26.347000",
      "content": "<p>Single NN Model:\nCV: 0.686  LB: 1.051\nGap: 0.365</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 416093,
      "author_name": "ONODERA",
      "author_url": "",
      "post_date": "2018-11-06T07:27:22.437000",
      "content": "<p>When you say \"CV\", what does it stand for? multi log loss in lgb? weighted log loss used in this competition?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 416117,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-11-06T08:05:20.057000",
          "content": "<p><a href=\"/onodera\">@onodera</a>, as far as I'm concerned it is weighted log loss as defined for the competition.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 415903,
      "author_name": "agarreta",
      "author_url": "",
      "post_date": "2018-11-05T22:05:22.523000",
      "content": "<p>For me it is CV 0.525, LB 0.989, and an enormous gap of 0.464. Previously it used to be around 0.43 - 0.45. I am using Olivier's method for assigning the probabilities of class 99</p>",
      "votes": 1,
      "replies": [
        {
          "id": 416027,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-06T03:57:17.320000",
          "content": "<p>My gap also increases as the model gets better, and I had to find additional ways of reducing overfitting.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 416824,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-11-07T10:25:03.720000",
          "content": "<p>My gap is a constant 0.3 even when the model improves (it was 0.4 in the past and improved when the model improved). </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 417029,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-11-07T16:26:43.353000",
          "content": "<p>0.45 for me, my model is still in the infancy period</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 431180,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-12-01T18:32:44.477000",
      "content": "<p>single lgb CV 0.431, LB 0.790 Gap 0.359</p>\n\n<p>Gap decreases as model improves.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 436065,
          "author_name": "LongYin/杰少",
          "author_url": "",
          "post_date": "2018-12-09T13:48:55.893000",
          "content": "<p>What loss do you used in your Local CV? Multi log loss in lgb?  Or weighted log loss by <a href=\"/olivier\">@olivier</a> ?</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 436082,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-09T14:41:22.860000",
          "content": "<p>Neither ones.  I use a loss that mimics the competition metric.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436086,
          "author_name": "Angus Chang",
          "author_url": "",
          "post_date": "2018-12-09T14:52:07.377000",
          "content": "<p>We've reached CV around 0.4 but the gap is still big. I feel like there is something behind.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436259,
          "author_name": "LongYin/杰少",
          "author_url": "",
          "post_date": "2018-12-10T02:11:05.507000",
          "content": "<p>Got it, thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 430676,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2018-11-30T20:42:31.707000",
      "content": "<p>Single LGB:\nCV: 0.454 LB: 0.875 gap: 0.42 (my gap still big)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 427252,
      "author_name": "S D",
      "author_url": "",
      "post_date": "2018-11-25T00:40:27.933000",
      "content": "<p>CV: 0.5\nLB: 1.2\nLooks like I'm gonna have to figure out some stuff here...</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 423010,
      "author_name": "ynktk",
      "author_url": "",
      "post_date": "2018-11-17T08:50:53.677000",
      "content": "<p>NN with 400 features: <br>\nCV 0.532 LB 1.017 - GAP 0.485  </p>\n\n<p>LGBM with same features: <br>\nCV 0.591 LB 1.127 - GAP 0.536  </p>\n\n<p>My CV-LB gap is bigger than others. I removed features which seems diffrent between train and test, but didn't improve LB score.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 423049,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-17T10:53:00.867000",
          "content": "<p>Try tuning lgb parameters to be more conservative.  Start with what is given at the bottom of this page: <a href=\"https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html\">https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 423110,
          "author_name": "ynktk",
          "author_url": "",
          "post_date": "2018-11-17T14:08:23.927000",
          "content": "<p>Thanks CPMP, I will try parameters tuning!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423261,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-11-17T19:54:06.020000",
          "content": "<p>thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423872,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-19T07:04:43.807000",
          "content": "<p>400 features seems a lot for 7k examples...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423977,
          "author_name": "ynktk",
          "author_url": "",
          "post_date": "2018-11-19T11:13:24.143000",
          "content": "<p>How many features do you use for your best model? <br>\nI have created about 8k features, and selected 400 for my models based on feature importances. Indeed, I need to refine the number of features and feature selection method.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423978,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-19T11:14:18.150000",
          "content": "<p>About 110, but I am trying to reduce this number.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 415389,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-11-05T03:04:38.217000",
      "content": "<p>Single lgb model, 100 features:\nCV 0.520 LB 0.957 - GAP 0.437</p>\n\n<p>The gap is stable as shown by previous results:</p>\n\n<p>CV 0.570 LB 1.001 - GAP 0.431</p>\n\n<p>CV 0.622 LB 1.052 - GAP 0.430</p>\n\n<p>Seems deep learning models have a smaller gap.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 416754,
          "author_name": "Fatih Öztürk",
          "author_url": "",
          "post_date": "2018-11-07T08:19:03.440000",
          "content": "<p>Wow, 0.52 CV is awesome. It seems you have found another important feature like mjd_diff of detected ones and I'm extremely curious about it :) I've just got back from my vacation and it seems I have a lot to do...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 416764,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-07T08:46:37.360000",
          "content": "<p>My current best is 0.504 CV, 0.943 LB, single lgb model</p>\n\n<p>I did find some useful features indeed, but I am not the only one if I look at the LB!</p>\n\n<p>Anyway, I am now trying to build some NN models as I find it harder and harder to improve my lgb model.  And we know ensembling can help a lot, right?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 416766,
          "author_name": "Fatih Öztürk",
          "author_url": "",
          "post_date": "2018-11-07T08:57:15.560000",
          "content": "<p>I doubt the first three scores come from a single model. Especially the top two. Yeah, ensembling with some NN will help you a lot :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 416773,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2018-11-07T09:01:58.333000",
          "content": "<p>@CPMP I wish I had your CV. my current lgb model has CV of 0.582 and the corresponding LB is 0.950. So maybe top3 guys have something like your CV and my LB-CV difference.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 416780,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-07T09:17:37.430000",
          "content": "<p>Ahmet,  I wish I had you CV-LB gap indeed ;)</p>\n\n<p>And I am NOW trying to build NN models....</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 416791,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2018-11-07T09:34:30.450000",
          "content": "<p>if I see a significant jump in your score in a week, I will switch to NN too:)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 416794,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-07T09:42:54.663000",
          "content": "<p>My LB score is a lgb model still, I am just starting with NN.  First NN sub had a LB score of 2.26 ...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 416800,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-11-07T09:56:11.160000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a>, the good thing is you can only improve this score :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 416808,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-07T10:00:09.553000",
          "content": "<blockquote>\n  <p>you can only improve this score :)</p>\n</blockquote>\n\n<p>Don't overestimate my DL skills ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416911,
          "author_name": "Fatih Öztürk",
          "author_url": "",
          "post_date": "2018-11-07T12:31:55.477000",
          "content": "<p>@AhmetErdem do you think that you have an intentional attempt to have such a low CV-LB gap compared to us, or it is just luck?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416967,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2018-11-07T14:41:59.557000",
          "content": "<p><a href=\"/fatihozturk\">@fatihozturk</a> I worked on it. maybe there is also some luck factor.\n@CPMP just tried to use deep NN with only dense features. My LB-CV increased from 0.37 to 0.42. I don't know why some people observed the opposite.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416982,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-07T15:09:50.383000",
          "content": "<p>@AhmetErdem Is using NN the reason for your last improvement? Just looking for additional motivation to use NN ;)</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 417008,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2018-11-07T15:59:16.297000",
          "content": "<p>While my NN scored much worse than my LGB, it could still contribute a bit by blending.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 413589,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2018-11-01T06:58:53.483000",
      "content": "<p>My best lightgbm model scores 0.814CV (no gal/extra gal model) and LB 1.296. Gap is 0.482...</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 411621,
      "author_name": "João Pedro Peinado",
      "author_url": "",
      "post_date": "2018-10-28T15:28:31.817000",
      "content": "<p>0.85 CV / 1.365 LB in my single lightgbm model using Olivier's method for class 99.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 412287,
          "author_name": "João Pedro Peinado",
          "author_url": "",
          "post_date": "2018-10-29T23:13:47.310000",
          "content": "<p>Now: 0.746 CV / 1.324 LB</p>\n\n<p>Seems the gap between my validation and my leaderboard is increasing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 412525,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-10-30T10:40:41.743000",
          "content": "<p>Are you using \"hostgal_specz\"? if yes, this might be the issue.\nOtherwise, you might be over fitting or you treat class 99 wrongly (try to give it a constant value)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 410208,
      "author_name": "Nikita Kozodoi",
      "author_url": "",
      "post_date": "2018-10-25T15:43:47.237000",
      "content": "<p>I am also using a single model:</p>\n\n<ul>\n<li>Local CV = 0.845</li>\n<li>LB = 1.335</li>\n</ul>\n\n<p>Seems like the CV/LB gap is pretty stable across the participants at 0.4 - 0.5 :) I guess dealing with class 99 and different data distribution in train/test would be the only way to significantly reduce it. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 410194,
      "author_name": "Vicens Gaitan",
      "author_url": "",
      "post_date": "2018-10-25T15:23:20.027000",
      "content": "<p>Single model,  CV 0.796 LB 1.246   class_99 as in Olivier kernel.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 410166,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-10-25T14:48:18.900000",
      "content": "<p>Single (first) lightgbm model, CV 0.90, LB 1.405 constant value for class 99.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 410183,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-25T15:13:22.283000",
          "content": "<p>After just 2 submissions ? you must be a competition GM ;-)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 410203,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-10-25T15:37:28.950000",
          "content": "<p>First sub was to check weights, see: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67194#409633\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67194#409633</a>  ;)</p>\n\n<p>And I find Jack's single sub to be way more impressive than mine!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 410655,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-10-26T12:14:56.500000",
          "content": "<p>local CV 0.733, LB 1.204</p>\n\n<p>As pointed out by Nikita the gap seems to be rather constant.  It means we won't get LB score below 0.5 ;)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 410026,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-10-25T08:32:49.963000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 430937,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-01T09:26:50.257000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 430976,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-01T11:08:28.053000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 430989,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-01T11:25:02.527000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 431034,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-01T13:41:03.730000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 428785,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-27T22:19:42.203000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 428800,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-27T22:54:37.780000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 428877,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-28T02:40:55.897000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 429433,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-28T20:53:47.713000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 428422,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-27T08:54:00.167000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 428467,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-27T10:30:24.357000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 428493,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-27T11:35:09.003000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 428725,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-27T18:58:46.200000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 427417,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-25T13:25:59.640000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 427433,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-25T13:49:52.660000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 427434,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-25T13:50:33.177000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 423638,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-18T19:13:03.747000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 423654,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-18T19:47:17.020000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 423749,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-19T01:29:41.543000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423871,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-19T07:03:54.023000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 423422,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-18T08:00:20.010000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 423429,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-18T08:11:14.140000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 423733,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-19T00:27:04.350000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 421434,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-15T02:15:41.957000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 421591,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-15T07:00:07.343000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421658,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-15T08:35:35.677000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 415943,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-06T00:07:08.617000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 416026,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-06T03:56:11.020000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416915,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-07T12:36:11.977000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416933,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-07T13:07:27.273000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421611,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-15T07:27:40.540000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422988,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-17T07:31:52.920000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423066,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-17T11:58:51.203000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423107,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-17T13:54:45.210000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423370,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-18T03:38:29.597000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 414140,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-02T07:39:52.827000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 413584,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-01T06:55:00.057000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "409686": "I'm opening this thread since I start wondering if I'm not completely off .\n\nMy best local CV is around 0.981 (without any sort of class 99 modelization) and LB score at 1.455.\n\nThis CV/LB difference has not really changed from my first submissions onward (at least from the day I started to compute a class 99 prediction). \n\nAre you, Dear Competitors, willing to share your own observations around this ? \n\nThanks.",
    "409971": "Galactic models CV: 0.21 +/- 0.01\n\nExtragalactic models CV: 0.88 +/- 0.02\n\nLocal CV ~0.70\n\nLB: 1.155\n\nOverfitting region reached - many \"smart\" features decreasing local CV score by 0.02 increase LB score by 0.10 (the consequence of the fact, that the training data is a poor representation of the test set).\nClass_99 distribution set by intuition, further probing (using multilogloss probabilities instead of a coin or a dice) improved the score by 0.02 only.\n\nEDIT: To be clear, +/- 0.01 means here the range of CV of a few similar but different models, not an error of CV of one model",
    "423706": "My current best submission has CV 0.49, LB 0.83 for a gap of 0.34. So there are ways of lowering the gap!",
    "426576": "We've achieved a significant milestone by combining my best lgb model with ideas from my team mates.\n\nsingle run: CV 0.431, LB 0.801, Gap 0.370\n\n@yuval_r prediction that final best score can be below 0.7 may be true.",
    "415424": "My latest CV (single lightgbm) is 0.659 and LB 1.082. I think I finally found what I missed.",
    "414643": "With same features\n\n- lgb: CV 0.98 LB 1.454 gap 0.474  \n- mlp: CV 0.82 LB 1.148 gap 0.328  ",
    "410032": "I love it when I see we are now dealing with galactic models and extra galactic models!  Machine Learning moving to entire new horizons! :)",
    "409843": "Galactic model CV: 0.23\n\nExtra-Galactic model CV: 0.99\n\nLB: 1.53\n\nI believe local CV is low because class_99 distribution is a mystery of the universe :-)",
    "422993": "My last features were badly overfitting, therefore I tried harder to understand why I had a much bigger gap than others.  And I found one reason, that leaders most probably dealt with weeks ago.  Using one of my last lightgbm runs, not the best on LB, but close to it, I got this.\n\nBefore:\n\nCV 0.469 LB 0.903 GAP 0.434\n\nAfter:\n\nCV 0.527 LB 0.902 GAP 0.375\n\nI hope my feature evaluation will be more effective from now on.  ",
    "413953": "My latest model scores 0.92643 CV and 1.291 with a lot less overfitting. ",
    "409833": "1.07 CV / 1.53 LB here\n\nFrom the data note:\n&gt; Crucially, the classifications will occur on a large test set, but the training data will be a small subset of the full data, and will also be a poor representation of the test set, to mimic the challenges we face observationally.\n\nNot much we can do I guess...?",
    "427393": "Two single models:\nCV: 0.46 LB: 0.931 (215 features)\nCV: 0.49 LB: 0.967 (57 features)",
    "425351": "Inner-galactic model:  CV 0.091 with 141 features\nExtra-galactic model: CV 0.721 with 221 features\nOverall: CV 0.524, LB 0.889 - GAP 0.365\n\nThere's still room for lowering the gap...",
    "413364": "Current Lightgbm Model: CV: 0.598 , LB: 1.035.",
    "410895": "For me also the difference is about 0.4 between Local CV and LB ",
    "410063": "local 5f cv - 0.98\nlb - 1.505\n\ni can add more features that improve local cv but they all seem to hurt lb so far. probably to do with the many times mentioned here difference between train/test sets.\n\nas an aside. setting 0 for all class_99 preds seemed to yield a lb of 5.0 which kinda means (in my humble and probably uselss opinion) that this is probably the most critical part of the problem in terms of increasing lb score.",
    "409691": "1.12 CV / 1.481 LB for me. IIRC, this is using different `class_99` predictions for galactic/extra-galactic objects.",
    "427890": "Single NN Model:\nCV: 0.686  LB: 1.051\nGap: 0.365",
    "416093": "When you say \"CV\", what does it stand for? multi log loss in lgb? weighted log loss used in this competition?",
    "415903": "For me it is CV 0.525, LB 0.989, and an enormous gap of 0.464. Previously it used to be around 0.43 - 0.45. I am using Olivier's method for assigning the probabilities of class 99\n\n",
    "431180": "single lgb CV 0.431, LB 0.790 Gap 0.359\n\nGap decreases as model improves.",
    "430676": "Single LGB:\nCV: 0.454 LB: 0.875 gap: 0.42 (my gap still big)",
    "427252": "CV: 0.5\nLB: 1.2\nLooks like I'm gonna have to figure out some stuff here...",
    "423010": "NN with 400 features:  \nCV 0.532 LB 1.017 - GAP 0.485  \n\nLGBM with same features:   \nCV 0.591 LB 1.127 - GAP 0.536  \n\nMy CV-LB gap is bigger than others. I removed features which seems diffrent between train and test, but didn't improve LB score.",
    "415389": "Single lgb model, 100 features:\nCV 0.520 LB 0.957 - GAP 0.437\n\nThe gap is stable as shown by previous results:\n\nCV 0.570 LB 1.001 - GAP 0.431\n\nCV 0.622 LB 1.052 - GAP 0.430\n\nSeems deep learning models have a smaller gap.",
    "413589": "My best lightgbm model scores 0.814CV (no gal/extra gal model) and LB 1.296. Gap is 0.482...",
    "411621": "0.85 CV / 1.365 LB in my single lightgbm model using Olivier's method for class 99.",
    "410208": "I am also using a single model:\n\n - Local CV = 0.845\n - LB = 1.335\n\nSeems like the CV/LB gap is pretty stable across the participants at 0.4 - 0.5 :) I guess dealing with class 99 and different data distribution in train/test would be the only way to significantly reduce it. ",
    "410194": "Single model,  CV 0.796 LB 1.246   class_99 as in Olivier kernel.",
    "410166": "Single (first) lightgbm model, CV 0.90, LB 1.405 constant value for class 99.",
    "410026": "For a single LGBM model:\n\n+ 5 fold CV:  0.90\n\n+ LB:  1.342 \n \nCV and and LB have followed the same path from the start but I think a) I need to work on the ghost class 99 and b) there are some bad overfitting pitfall are out there...I just have fallen into some of them along the way...;)  \nI also used different methods for predicting prob(class_99) but just about 1-2 of my ideas worked out of my 9 submissions.",
    "430937": "I've tried my first MLP with same features of lgbm. Model structure is completly same with the one in public kernels.\n\nIt's CV: 0.82 LB: 1.143. Gap is so good but CV :(. Also, tried simple weighted average with LGBM and it neither improves CV nor LB.",
    "428785": "Maybe we could think about this differently?\n\nI have two models, one for Galactic and another for Extra-Galactic. Same features, both LGBM. They give:\nGalactic CV: 0.129\nExtra-Gal CV: 0.836\n\nThe training data is split 2325/5523 Galactic/Extra-Galactic. Combining the CV scores of my two models using a weighting based on the training data gives:\nCombined CV: 0.627\nLB: 1.034\nGAP: 0.407\n\nThat GAP is in the normal range according to this thread. However, the test data contains a higher proportion of Extra-Galactic sources than the training data. The mix is 390510/3102380. Using a weighting based on the test data gives:\nCombined CV: 0.757\nLB: 1.034\nGAP: 0.277\n\nHopefully by removing the effect of different proportions of Galactic/Extra-Galactic sources between train and test, we can see how much of the GAP is down to other effects like class 99 and other train/test data differences.\n",
    "428422": "Update - Single NN Model:\nCV: 0.666 LB: 0.995\nGap: 0.329",
    "427417": "Guys, could you help to understand this topic for new members... Thanks.",
    "423638": "Since we're talking about overfitting in here a bit,  how do your fold training errors compare to your fold validation errors?\nI typically get around 0.15 for my training loss and 0.5+ for my validation losses. This is for a 5-fold LGBM, not separating galactic from extragalactic.\nI think I might benefit from some form of regularization given the gap between training and validation.",
    "423422": "My single lgb: CV 0.494, LB 1.01\n\noverfitting?",
    "421434": "Has anyone been monitoring CV score on galactic and extragalactic objects separately? My cv score on the galactic part is 0.05 so I am guessing there is overfitting there. Anyone with a similar experience?",
    "415943": "I'm trying to understand what magic all of you are doing here, these 0.5x in CV are very crazy to me, I need to figure out this magic very soon...\n\nMaybe the problem is my poorly knowledge in physics and astronomy...",
    "414140": "I've dropped the \"hostgal_specz\" column, get a local CV of around 0.9 with LB 1.5, either my model overfits or my class 99 needs some extra work.",
    "413584": ""
  }
}