{
  "id": 70908,
  "title": "Good practices",
  "url": "/competitions/PLAsTiCC-2018/discussion/70908",
  "author_name": "",
  "post_date": "2018-11-08T09:46:02.450575700Z",
  "votes": 130,
  "comment_count": 64,
  "views": 0,
  "content": "<p>Here are few things that worked well for me so far.  Sharing them in the hope that they will be useful to others.</p>\n\n<ol>\n<li><p>Be conservative when adding features if you use non parametric models like gbms (lightgbm, xgboosst, catboost), or NNs.  Indeed, we only have 7k examples or so, and too many features easily overfit.  For instance, my best LB score is obtained with lightgbm and less than 100 features.</p></li>\n<li><p>Make sure features convey information.  Features that do not involve what is measured (flux or flux_err) are detrimental.  Examples are <code>max(mjd)- min(mjd)</code>, or <code>mean(passband)</code>.</p></li>\n<li><p>Read the forum.  Some very interesting features ideas can be shared, see for instance Grzegorz Sionkovski' <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69696#410538\">great feature</a>, and his hint for another one in <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/70725#416740\">this discussion</a>.</p></li>\n<li><p>Overfitting is the enemy here.  Avoid features that give a huge CV boost and a very small LB boost.  This means you should care about the gap beteen LB and CV. Ahmet Erdem is making a way better job than me here for instance, and it will pay in the end I'm sure.</p></li>\n<li><p>Make sure you use an objective function (loss function) that mimics the competition metric. Default log loss or default multi class loss are not best here, by far.  See for instance the <a href=\"https://www.kaggle.com/mithrillion/know-your-objective#\">Know Your Objective kernel</a> by Mithrillion.  I think there are simpler way to do it, but the general idea is right.</p></li>\n<li><p>Experiment ideas.  Many will fail (no CV improvement, or no LB improvement), but some will work.  Don't add features just because they look good 'intuitively'.  Always validate via cross validation score, and, possibly, by LB score.</p></li>\n<li><p>Make only one change at a time, and validate it by CV, and possibly by LB, before moving to the next one.</p></li>\n<li><p>Don't do too much hyper parameter optimization.  Again, with a small number of examples, it is very easy to overfit to the training data, even if you use cross validation. I usually do it after a week or so in the competition, then keep same parameters until last week of competition.</p></li>\n<li><p>Write in the forum about your ideas or doubts.  You'll receive useful feedback in general.</p></li>\n<li><p>Share some of your work in kernels.  As in 9, you'll receive useful feedback in general.</p></li>\n<li><p>Apply all the above when you reuse something shared by others, esp items 1, 2, and 4.</p></li>\n<li><p>Most important.  Never be discouraged, the competition is not over till the last minute!  It happened to me in the past that I gained several ranks in the private LB with a last minute submission.  I really mean  submitted one minute before competition deadline.</p></li>\n</ol>\n\n<p>edit, two more topics:</p>\n\n<ol>\n<li><p>Make sure you have reproducible results.  Initialize all random seeds to the same value for each run.  If you don't then it is possible to think your last change is responsible for the improvement you see when it is just due to a different execution path in the algorithm.  For NN this is tricky, and keras is known to not be very good at this.  Pytorch seems better.</p></li>\n<li><p>Have fun!</p></li>\n</ol>",
  "messages": [
    {
      "id": "417449",
      "postDate": "11/08/2018 09:46:02",
      "content": "<p>Here are few things that worked well for me so far.  Sharing them in the hope that they will be useful to others.</p>\n\n<ol>\n<li><p>Be conservative when adding features if you use non parametric models like gbms (lightgbm, xgboosst, catboost), or NNs.  Indeed, we only have 7k examples or so, and too many features easily overfit.  For instance, my best LB score is obtained with lightgbm and less than 100 features.</p></li>\n<li><p>Make sure features convey information.  Features that do not involve what is measured (flux or flux_err) are detrimental.  Examples are <code>max(mjd)- min(mjd)</code>, or <code>mean(passband)</code>.</p></li>\n<li><p>Read the forum.  Some very interesting features ideas can be shared, see for instance Grzegorz Sionkovski' <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69696#410538\">great feature</a>, and his hint for another one in <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/70725#416740\">this discussion</a>.</p></li>\n<li><p>Overfitting is the enemy here.  Avoid features that give a huge CV boost and a very small LB boost.  This means you should care about the gap beteen LB and CV. Ahmet Erdem is making a way better job than me here for instance, and it will pay in the end I'm sure.</p></li>\n<li><p>Make sure you use an objective function (loss function) that mimics the competition metric. Default log loss or default multi class loss are not best here, by far.  See for instance the <a href=\"https://www.kaggle.com/mithrillion/know-your-objective#\">Know Your Objective kernel</a> by Mithrillion.  I think there are simpler way to do it, but the general idea is right.</p></li>\n<li><p>Experiment ideas.  Many will fail (no CV improvement, or no LB improvement), but some will work.  Don't add features just because they look good 'intuitively'.  Always validate via cross validation score, and, possibly, by LB score.</p></li>\n<li><p>Make only one change at a time, and validate it by CV, and possibly by LB, before moving to the next one.</p></li>\n<li><p>Don't do too much hyper parameter optimization.  Again, with a small number of examples, it is very easy to overfit to the training data, even if you use cross validation. I usually do it after a week or so in the competition, then keep same parameters until last week of competition.</p></li>\n<li><p>Write in the forum about your ideas or doubts.  You'll receive useful feedback in general.</p></li>\n<li><p>Share some of your work in kernels.  As in 9, you'll receive useful feedback in general.</p></li>\n<li><p>Apply all the above when you reuse something shared by others, esp items 1, 2, and 4.</p></li>\n<li><p>Most important.  Never be discouraged, the competition is not over till the last minute!  It happened to me in the past that I gained several ranks in the private LB with a last minute submission.  I really mean  submitted one minute before competition deadline.</p></li>\n</ol>\n\n<p>edit, two more topics:</p>\n\n<ol>\n<li><p>Make sure you have reproducible results.  Initialize all random seeds to the same value for each run.  If you don't then it is possible to think your last change is responsible for the improvement you see when it is just due to a different execution path in the algorithm.  For NN this is tricky, and keras is known to not be very good at this.  Pytorch seems better.</p></li>\n<li><p>Have fun!</p></li>\n</ol>",
      "rawMarkdown": "Here are few things that worked well for me so far.  Sharing them in the hope that they will be useful to others.\n\n1. Be conservative when adding features if you use non parametric models like gbms (lightgbm, xgboosst, catboost), or NNs.  Indeed, we only have 7k examples or so, and too many features easily overfit.  For instance, my best LB score is obtained with lightgbm and less than 100 features.\n\n2. Make sure features convey information.  Features that do not involve what is measured (flux or flux_err) are detrimental.  Examples are `max(mjd)- min(mjd)`, or `mean(passband)`.\n\n3. Read the forum.  Some very interesting features ideas can be shared, see for instance Grzegorz Sionkovski' [great feature][1], and his hint for another one in [this discussion][2].\n\n4.  Overfitting is the enemy here.  Avoid features that give a huge CV boost and a very small LB boost.  This means you should care about the gap beteen LB and CV. Ahmet Erdem is making a way better job than me here for instance, and it will pay in the end I'm sure.\n\n5. Make sure you use an objective function (loss function) that mimics the competition metric. Default log loss or default multi class loss are not best here, by far.  See for instance the [Know Your Objective kernel][3] by Mithrillion.  I think there are simpler way to do it, but the general idea is right.\n\n6. Experiment ideas.  Many will fail (no CV improvement, or no LB improvement), but some will work.  Don't add features just because they look good 'intuitively'.  Always validate via cross validation score, and, possibly, by LB score.\n\n7.  Make only one change at a time, and validate it by CV, and possibly by LB, before moving to the next one.\n\n8. Don't do too much hyper parameter optimization.  Again, with a small number of examples, it is very easy to overfit to the training data, even if you use cross validation. I usually do it after a week or so in the competition, then keep same parameters until last week of competition.\n\n9. Write in the forum about your ideas or doubts.  You'll receive useful feedback in general.\n\n10. Share some of your work in kernels.  As in 9, you'll receive useful feedback in general.\n\n11.  Apply all the above when you reuse something shared by others, esp items 1, 2, and 4.\n\n12. Most important.  Never be discouraged, the competition is not over till the last minute!  It happened to me in the past that I gained several ranks in the private LB with a last minute submission.  I really mean  submitted one minute before competition deadline.\n\nedit, two more topics:\n\n13.  Make sure you have reproducible results.  Initialize all random seeds to the same value for each run.  If you don't then it is possible to think your last change is responsible for the improvement you see when it is just due to a different execution path in the algorithm.  For NN this is tricky, and keras is known to not be very good at this.  Pytorch seems better.\n\n14. Have fun!\n\n  [1]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69696#410538\n  [2]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/70725#416740\n  [3]: https://www.kaggle.com/mithrillion/know-your-objective#",
      "votes": null
    },
    {
      "id": "417456",
      "postDate": "11/08/2018 10:06:21",
      "content": "<p>Thanks for sharing, fantastic masters like you make this community awesome... </p>",
      "rawMarkdown": "Thanks for sharing, fantastic masters like you make this community awesome...",
      "votes": null
    },
    {
      "id": "417457",
      "postDate": "11/08/2018 10:10:37",
      "content": "<p>Great work\nI am fortunate that I have 128Gb of memory and it significantly improves scores if I don't have to aggregate on chunks - if you have a small amount of memory competitors are at a disadvantage - I tried to produce a sorted dataset so that people with small memory would be able to compete but it died. </p>",
      "rawMarkdown": "Great work\nI am fortunate that I have 128Gb of memory and it significantly improves scores if I don't have to aggregate on chunks - if you have a small amount of memory competitors are at a disadvantage - I tried to produce a sorted dataset so that people with small memory would be able to compete but it died.",
      "votes": null
    },
    {
      "id": "417487",
      "postDate": "11/08/2018 11:04:33",
      "content": "<p>Thanks for your sharing:)\nI have a question.</p>\n\n<blockquote>\n  <p>too many features easily overfit. For instance, my best LB score is obtained with lightgbm and less than 100 features.\n  You mean your cv is better when you use more than 100 features? (more than 100 features cv &gt; less than 100 features cv, more than 100 features LB &lt; less than 100 features LB, right?)\n  In my case, about 150~200 features got my best cv.</p>\n</blockquote>",
      "rawMarkdown": "Thanks for your sharing:)\nI have a question.\n&gt;too many features easily overfit. For instance, my best LB score is obtained with lightgbm and less than 100 features.\nYou mean your cv is better when you use more than 100 features? (more than 100 features cv &gt; less than 100 features cv, more than 100 features LB &lt; less than 100 features LB, right?)\nIn my case, about 150~200 features got my best cv.",
      "votes": null
    },
    {
      "id": "417521",
      "postDate": "11/08/2018 12:12:52",
      "content": "<p>Thank you for sharing. My question: can you really trust those \"importance\" on the lightgbm? I get good gain for the kurtosis for example, but it does not reflect on LB. Some features i don't find meaningfull, but they have good \"importance\". I start to doubt it...</p>",
      "rawMarkdown": "Thank you for sharing. My question: can you really trust those \"importance\" on the lightgbm? I get good gain for the kurtosis for example, but it does not reflect on LB. Some features i don't find meaningfull, but they have good \"importance\". I start to doubt it...",
      "votes": null
    },
    {
      "id": "417529",
      "postDate": "11/08/2018 12:23:25",
      "content": "<p>I have 8G and compute in chunks. <code>128Gb of memory and it significantly improves scores</code> why do you think it is? I thought the score depends on features you choose and should not depend on how you compute them?</p>",
      "rawMarkdown": "I have 8G and compute in chunks. ```128Gb of memory and it significantly improves scores``` why do you think it is? I thought the score depends on features you choose and should not depend on how you compute them?",
      "votes": null
    },
    {
      "id": "417540",
      "postDate": "11/08/2018 12:36:52",
      "content": "<blockquote>\n  <p>can you really trust those \"importance\" on the lightgbm?</p>\n</blockquote>\n\n<p>No ;)</p>\n\n<p>If the importance is very low or 0, then you can try without the feature, and see how it goes.  But in one competition (Talking Data) I gained a lot by removing the feature that was said to be the most important by lightgbm (the <code>channel</code> feature).</p>",
      "rawMarkdown": "&gt; can you really trust those \"importance\" on the lightgbm?\n\nNo ;)\n\nIf the importance is very low or 0, then you can try without the feature, and see how it goes.  But in one competition (Talking Data) I gained a lot by removing the feature that was said to be the most important by lightgbm (the `channel` feature).",
      "votes": null
    },
    {
      "id": "417541",
      "postDate": "11/08/2018 12:38:09",
      "content": "<p>Both  CV and LB are better with less than 100 features.  My CV is well correlated with LB, when one improves the other improves.</p>",
      "rawMarkdown": "Both  CV and LB are better with less than 100 features.  My CV is well correlated with LB, when one improves the other improves.",
      "votes": null
    },
    {
      "id": "417543",
      "postDate": "11/08/2018 12:40:23",
      "content": "<p>My features only depend on each object id, hence having lots of memory would not help me.  I haven't measured precisely but my processes consume less than 6GB.</p>",
      "rawMarkdown": "My features only depend on each object id, hence having lots of memory would not help me.  I haven't measured precisely but my processes consume less than 6GB.",
      "votes": null
    },
    {
      "id": "417544",
      "postDate": "11/08/2018 12:40:43",
      "content": "<p>Thanks.</p>",
      "rawMarkdown": "Thanks.",
      "votes": null
    },
    {
      "id": "417550",
      "postDate": "11/08/2018 12:47:01",
      "content": "<p>Thanks!\nMy CV is well correlated too. The gap is 0.45.</p>",
      "rawMarkdown": "Thanks!\nMy CV is well correlated too. The gap is 0.45.",
      "votes": null
    },
    {
      "id": "417552",
      "postDate": "11/08/2018 12:48:19",
      "content": "<p>I get different results if I use chunks rather than the whole - I believe not all objectids are continuously ordered - might just be me</p>",
      "rawMarkdown": "I get different results if I use chunks rather than the whole - I believe not all objectids are continuously ordered - might just be me",
      "votes": null
    },
    {
      "id": "417553",
      "postDate": "11/08/2018 12:52:22",
      "content": "<p>What depends on chunk in your case?  </p>\n\n<p>I am basically reusing Olivier's way of processing test data by chunks, and there is only one object_id that seems to appear in two chunks.  </p>",
      "rawMarkdown": "What depends on chunk in your case?  \n\nI am basically reusing Olivier's way of processing test data by chunks, and there is only one object_id that seems to appear in two chunks.",
      "votes": null
    },
    {
      "id": "417565",
      "postDate": "11/08/2018 13:15:30",
      "content": "<p>Good read, thanks a lot for sharing!</p>",
      "rawMarkdown": "Good read, thanks a lot for sharing!",
      "votes": null
    },
    {
      "id": "417571",
      "postDate": "11/08/2018 13:25:56",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "417588",
      "postDate": "11/08/2018 13:47:33",
      "content": "<p>I added two topics ;)</p>",
      "rawMarkdown": "I added two topics ;)",
      "votes": null
    },
    {
      "id": "417589",
      "postDate": "11/08/2018 13:47:34",
      "content": "<p>I was using smaller chunks so that was probably it.  I check and get back to you</p>",
      "rawMarkdown": "I was using smaller chunks so that was probably it.  I check and get back to you",
      "votes": null
    },
    {
      "id": "417598",
      "postDate": "11/08/2018 13:54:45",
      "content": "<p>I have dedicated some time to split the test set into several chunks, and I have double checked that the same object_id is in consecutive rows (though the ids are not in increasing order).\nMaybe you have some feature that depends on different objects?</p>",
      "rawMarkdown": "I have dedicated some time to split the test set into several chunks, and I have double checked that the same object_id is in consecutive rows (though the ids are not in increasing order).\nMaybe you have some feature that depends on different objects?",
      "votes": null
    },
    {
      "id": "417616",
      "postDate": "11/08/2018 14:24:12",
      "content": "<p>Thank you for your post, good remainder. Totally agree with point 8... I have wasted a lot of time just to make a submission with LB score of 3.3 :'(</p>",
      "rawMarkdown": "Thank you for your post, good remainder. Totally agree with point 8... I have wasted a lot of time just to make a submission with LB score of 3.3 :'(",
      "votes": null
    },
    {
      "id": "417621",
      "postDate": "11/08/2018 14:34:07",
      "content": "<p>It must be to do with my pseudo clustering - once I understand the issue I'll produce a kernel</p>",
      "rawMarkdown": "It must be to do with my pseudo clustering - once I understand the issue I'll produce a kernel",
      "votes": null
    },
    {
      "id": "417633",
      "postDate": "11/08/2018 14:55:26",
      "content": "<p>Nice post! I'll follow your guideline from now on.</p>",
      "rawMarkdown": "Nice post! I'll follow your guideline from now on.",
      "votes": null
    },
    {
      "id": "417642",
      "postDate": "11/08/2018 15:07:26",
      "content": "<p>Thank you Grandmaster</p>",
      "rawMarkdown": "Thank you Grandmaster",
      "votes": null
    },
    {
      "id": "417651",
      "postDate": "11/08/2018 15:27:59",
      "content": "<p>I've been burned by HPO more than once ;)</p>",
      "rawMarkdown": "I've been burned by HPO more than once ;)",
      "votes": null
    },
    {
      "id": "417652",
      "postDate": "11/08/2018 15:28:25",
      "content": "<p>Are you sure you need help here? ;)</p>",
      "rawMarkdown": "Are you sure you need help here? ;)",
      "votes": null
    },
    {
      "id": "417660",
      "postDate": "11/08/2018 15:38:42",
      "content": "<p>thank you</p>",
      "rawMarkdown": "thank you",
      "votes": null
    },
    {
      "id": "417791",
      "postDate": "11/08/2018 19:35:23",
      "content": "<p>\"My features only depend on each object id\", same for me. Then I do not need chunks after I have features file to calculate predict, do I ? (it's &lt; 2G csv file at the moment)\n\"I am basically reusing Olivier's way of processing test data by chunks, and there is only one object_id that seems to appear in two chunks\" --I use the same kernel. object_id  appearing in two chunks -- have not spot it, does it hurt the performance, should I find it and delete the duplicate?</p>",
      "rawMarkdown": "\"My features only depend on each object id\", same for me. Then I do not need chunks after I have features file to calculate predict, do I ? (it's &lt; 2G csv file at the moment)\n\"I am basically reusing Olivier's way of processing test data by chunks, and there is only one object_id that seems to appear in two chunks\" --I use the same kernel. object_id  appearing in two chunks -- have not spot it, does it hurt the performance, should I find it and delete the duplicate?",
      "votes": null
    },
    {
      "id": "417825",
      "postDate": "11/08/2018 20:40:24",
      "content": "<p>Olivier's kernel code is fine, he groups by object_id and averages before submitting.</p>",
      "rawMarkdown": "Olivier's kernel code is fine, he groups by object_id and averages before submitting.",
      "votes": null
    },
    {
      "id": "417831",
      "postDate": "11/08/2018 20:56:51",
      "content": "<p>Thanks CPMP... great and timeless advice. I strive to have the patience and discipline for 7!</p>\n\n<p>Your two extra topics are also very important. Tonight I need to sort out the crazy CV non-determinism I'm getting from Keras! </p>\n\n<p>Might I also add: Look at the data. There are some great EDA kernels for this competition. I learnt a lot looking at <a href=\"https://www.kaggle.com/mithrillion/all-classes-light-curve-characteristics\">this one</a>.</p>",
      "rawMarkdown": "Thanks CPMP... great and timeless advice. I strive to have the patience and discipline for 7!\n\nYour two extra topics are also very important. Tonight I need to sort out the crazy CV non-determinism I'm getting from Keras! \n\nMight I also add: Look at the data. There are some great EDA kernels for this competition. I learnt a lot looking at <a href=\"https://www.kaggle.com/mithrillion/all-classes-light-curve-characteristics\">this one</a>.",
      "votes": null
    },
    {
      "id": "417855",
      "postDate": "11/08/2018 21:53:22",
      "content": "<p>Nice to get some good advices from a GrandMaster !</p>\n\n<p>Thanks a lot.</p>\n\n<p>I would just add : keep tracks of your scores on a piece of paper (or spreadsheet) with modification(s) made, and more than obvious : save your best submission source file and do not modify anything, just start playing with a new version based on it.</p>",
      "rawMarkdown": "Nice to get some good advices from a GrandMaster !\n\nThanks a lot.\n\nI would just add : keep tracks of your scores on a piece of paper (or spreadsheet) with modification(s) made, and more than obvious : save your best submission source file and do not modify anything, just start playing with a new version based on it.",
      "votes": null
    },
    {
      "id": "417865",
      "postDate": "11/08/2018 22:22:32",
      "content": "<p>Great advice, thanks for sharing as well!</p>",
      "rawMarkdown": "Great advice, thanks for sharing as well!",
      "votes": null
    },
    {
      "id": "417884",
      "postDate": "11/08/2018 23:46:53",
      "content": "<p>Right, looking at data is key.  Thanks for reminding us of it.</p>",
      "rawMarkdown": "Right, looking at data is key.  Thanks for reminding us of it.",
      "votes": null
    },
    {
      "id": "417901",
      "postDate": "11/09/2018 00:35:19",
      "content": "<p>Nice advices</p>",
      "rawMarkdown": "Nice advices",
      "votes": null
    },
    {
      "id": "417995",
      "postDate": "11/09/2018 04:53:23",
      "content": "<p>Thanks @CPMP\nAdding my two cents: NN's love to over fit, think of clever ways to augment the data.</p>",
      "rawMarkdown": "Thanks @CPMP\nAdding my two cents: NN's love to over fit, think of clever ways to augment the data.",
      "votes": null
    },
    {
      "id": "417999",
      "postDate": "11/09/2018 05:08:22",
      "content": "<p>Right, tta is always key for deep learning.  Thanks for pointing it out.</p>",
      "rawMarkdown": "Right, tta is always key for deep learning.  Thanks for pointing it out.",
      "votes": null
    },
    {
      "id": "418001",
      "postDate": "11/09/2018 05:08:44",
      "content": "<p>Thanks, not sure you need any actually ;)</p>",
      "rawMarkdown": "Thanks, not sure you need any actually ;)",
      "votes": null
    },
    {
      "id": "418062",
      "postDate": "11/09/2018 08:17:20",
      "content": "<p>Thanks for sharing such insights. I also want to hear about feature engineering. It seems you are doing really great at feature engineering in this competition. Do you have any suggestions/tips for the community to learn how to generate meaningful features specific to this competition? Maybe some papers, or just plotting curves and getting insights from it? I'm sure that this will encourage people more.</p>\n\n<p>For my single lgbm model having a CV of 0.58X:</p>\n\n<ul>\n<li>Found my most important features after reading starter kit kernel. <a href=\"https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit\">https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit</a></li>\n<li>Then saw and applied some of features mentioned in Cesium package here: <a href=\"http://cesium-ml.org/docs/feature_table.html\">http://cesium-ml.org/docs/feature_table.html</a></li>\n<li>Following discussions (especially Grzegorz's feature ) </li>\n<li>New features having minor changes on features mentioned above.</li>\n</ul>",
      "rawMarkdown": "Thanks for sharing such insights. I also want to hear about feature engineering. It seems you are doing really great at feature engineering in this competition. Do you have any suggestions/tips for the community to learn how to generate meaningful features specific to this competition? Maybe some papers, or just plotting curves and getting insights from it? I'm sure that this will encourage people more.\n\nFor my single lgbm model having a CV of 0.58X:\n\n- Found my most important features after reading starter kit kernel. https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit\n- Then saw and applied some of features mentioned in Cesium package here: http://cesium-ml.org/docs/feature_table.html\n- Following discussions (especially Grzegorz's feature ) \n- New features having minor changes on features mentioned above.",
      "votes": null
    },
    {
      "id": "418121",
      "postDate": "11/09/2018 10:20:44",
      "content": "<p>Some of my features are trying to capture what I 'see' when looking at light curves from different classes.  If two classes seem to be different in a way, can I compute a feature that captures that difference?</p>\n\n<p>The main difference is the general shape of the light curve, and this is hard to capture via hand crafted features. This is why I think deep learning should be better in the end.  </p>",
      "rawMarkdown": "Some of my features are trying to capture what I 'see' when looking at light curves from different classes.  If two classes seem to be different in a way, can I compute a feature that captures that difference?\n\nThe main difference is the general shape of the light curve, and this is hard to capture via hand crafted features. This is why I think deep learning should be better in the end.",
      "votes": null
    },
    {
      "id": "418491",
      "postDate": "11/10/2018 01:22:55",
      "content": "<blockquote>\n  <p>Read the forum. Some very interesting features ideas can be shared,\n  see for instance Grzegorz Sionkovski' great feature, and his hint for\n  another one in this discussion.</p>\n</blockquote>\n\n<p>one more hint by <a href=\"/cpmpml\">@cpmpml</a> for useful features <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/70346#415506\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/70346#415506</a> </p>",
      "rawMarkdown": "&gt; Read the forum. Some very interesting features ideas can be shared,\n&gt; see for instance Grzegorz Sionkovski' great feature, and his hint for\n&gt; another one in this discussion.\n\none more hint by @cpmpml for useful features https://www.kaggle.com/c/PLAsTiCC-2018/discussion/70346#415506",
      "votes": null
    },
    {
      "id": "418744",
      "postDate": "11/10/2018 14:32:06",
      "content": "<p>thanks @CPMP. this is really helpful</p>",
      "rawMarkdown": "thanks @CPMP. this is really helpful",
      "votes": null
    },
    {
      "id": "419064",
      "postDate": "11/11/2018 06:57:17",
      "content": "<p>Thanks sharing your thoughts <a href=\"/cpmpml\">@cpmpml</a></p>\n\n<p>Feature selection is particularly important here.</p>",
      "rawMarkdown": "Thanks sharing your thoughts @cpmpml\n\nFeature selection is particularly important here.",
      "votes": null
    },
    {
      "id": "419340",
      "postDate": "11/11/2018 18:42:23",
      "content": "<p>Another post that is going to my personal <strong>Grandmasters 101</strong> document. Thanks @CPMP</p>",
      "rawMarkdown": "Another post that is going to my personal **Grandmasters 101** document. Thanks @CPMP",
      "votes": null
    },
    {
      "id": "419686",
      "postDate": "11/12/2018 11:51:24",
      "content": "<p>Great, excelent post</p>",
      "rawMarkdown": "Great, excelent post",
      "votes": null
    },
    {
      "id": "419850",
      "postDate": "11/12/2018 16:22:04",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "419897",
      "postDate": "11/12/2018 18:26:37",
      "content": "<p>Thank you for sharing your thoughts.</p>",
      "rawMarkdown": "Thank you for sharing your thoughts.",
      "votes": null
    },
    {
      "id": "420278",
      "postDate": "11/13/2018 11:40:32",
      "content": "<p>@CPMP</p>\n\n<p>So why do some strange features get high importance? For example, 'flux_err_min'.\nIs it just a peculiarity of the train set? And this is overfitting.</p>",
      "rawMarkdown": "CPMP\n\nSo why do some strange features get high importance? For example, 'flux_err_min'.\nIs it just a peculiarity of the train set? And this is overfitting.",
      "votes": null
    },
    {
      "id": "420280",
      "postDate": "11/13/2018 11:45:20",
      "content": "<p>Sergey, feature importance is not reliable at all.  For instance, in a previous competition I got a significant boost by removing the feature that lgb said was most important.</p>",
      "rawMarkdown": "Sergey, feature importance is not reliable at all.  For instance, in a previous competition I got a significant boost by removing the feature that lgb said was most important.",
      "votes": null
    },
    {
      "id": "420298",
      "postDate": "11/13/2018 12:35:33",
      "content": "<p>Nice, Thank you!</p>",
      "rawMarkdown": "Nice, Thank you!",
      "votes": null
    },
    {
      "id": "420509",
      "postDate": "11/13/2018 18:42:30",
      "content": "<p>Thanks! This work is helpful</p>",
      "rawMarkdown": "Thanks! This work is helpful",
      "votes": null
    },
    {
      "id": "420659",
      "postDate": "11/14/2018 00:53:26",
      "content": "<p>Thanks again! I've incorporated some of your ideas (especially 3.) in a kernel: <a href=\"https://www.kaggle.com/iprapas/ideas-from-kernels-and-discussion-lb-1-135\">https://www.kaggle.com/iprapas/ideas-from-kernels-and-discussion-lb-1-135</a> </p>",
      "rawMarkdown": "Thanks again! I've incorporated some of your ideas (especially 3.) in a kernel: https://www.kaggle.com/iprapas/ideas-from-kernels-and-discussion-lb-1-135",
      "votes": null
    },
    {
      "id": "420704",
      "postDate": "11/14/2018 02:17:57",
      "content": "<p>@CPMP, are you using 'multi_logloss' as metric in your lgbm?</p>",
      "rawMarkdown": "CPMP, are you using 'multi_logloss' as metric in your lgbm?",
      "votes": null
    },
    {
      "id": "420894",
      "postDate": "11/14/2018 09:36:23",
      "content": "<p>I am using the metric defined in Olivier's public kernel.</p>",
      "rawMarkdown": "I am using the metric defined in Olivier's public kernel.",
      "votes": null
    },
    {
      "id": "420974",
      "postDate": "11/14/2018 12:19:41",
      "content": "<p>does checking for permutation Importance of the features is the good practice?\nwhat about your practice on this , how often you use this method?</p>",
      "rawMarkdown": "does checking for permutation Importance of the features is the good practice?\nwhat about your practice on this , how often you use this method?",
      "votes": null
    },
    {
      "id": "420999",
      "postDate": "11/14/2018 13:00:29",
      "content": "<p>You mean boruta?  I've never used it...</p>\n\n<p>It is valuable, but it is slow.  And you must make sure you use the same algorithm within it than the algorithm you use for modeling.  For instance using boruta with random forest to select features for lightgbm is silly IMHO.  </p>",
      "rawMarkdown": "You mean boruta?  I've never used it...\n\nIt is valuable, but it is slow.  And you must make sure you use the same algorithm within it than the algorithm you use for modeling.  For instance using boruta with random forest to select features for lightgbm is silly IMHO.",
      "votes": null
    },
    {
      "id": "421006",
      "postDate": "11/14/2018 13:12:56",
      "content": "<p>not that! \nwhat i have used in some dataset for checking importance of features by doing permutation Importance using ELI5.sklearn. have you ever used this ELI5 library?\nI want to know is this a good practice or there is some best out there?</p>",
      "rawMarkdown": "not that! \nwhat i have used in some dataset for checking importance of features by doing permutation Importance using ELI5.sklearn. have you ever used this ELI5 library?\nI want to know is this a good practice or there is some best out there?",
      "votes": null
    },
    {
      "id": "421018",
      "postDate": "11/14/2018 13:27:30",
      "content": "<p>I never used eli5, but reading its documentation makes me think its permutation algorithm is close to the boruta algorithm.  If not then I welcome an explanation about how they differ.  Boruta in python: <a href=\"https://github.com/scikit-learn-contrib/boruta_py\">https://github.com/scikit-learn-contrib/boruta_py</a></p>\n\n<p>Some kagglers have adapted boruta to lightgbm, a google search should give you their git repos.</p>",
      "rawMarkdown": "I never used eli5, but reading its documentation makes me think its permutation algorithm is close to the boruta algorithm.  If not then I welcome an explanation about how they differ.  Boruta in python: https://github.com/scikit-learn-contrib/boruta_py\n\nSome kagglers have adapted boruta to lightgbm, a google search should give you their git repos.",
      "votes": null
    },
    {
      "id": "421116",
      "postDate": "11/14/2018 15:51:02",
      "content": "<p>thanks!</p>",
      "rawMarkdown": "thanks!",
      "votes": null
    },
    {
      "id": "421714",
      "postDate": "11/15/2018 10:09:21",
      "content": "<p>Nice Insights, thanks.</p>",
      "rawMarkdown": "Nice Insights, thanks.",
      "votes": null
    },
    {
      "id": "422708",
      "postDate": "11/16/2018 16:53:55",
      "content": "<p>thanks will check all about boruta algo</p>",
      "rawMarkdown": "thanks will check all about boruta algo",
      "votes": null
    },
    {
      "id": "422751",
      "postDate": "11/16/2018 18:21:11",
      "content": "<p>Nicely put !</p>",
      "rawMarkdown": "Nicely put !",
      "votes": null
    },
    {
      "id": "422834",
      "postDate": "11/16/2018 21:05:54",
      "content": "<p>@CPMPml How does that make any sense? If you are looking at gain importances, then the top feature is the feature that reduces entropy/uncertainty of the TARGET variable the most. In the absence of this variable, you no longer have the feature that can best discriminate/regress to your TARGET. Do you have any idea why one in general could get a boost by removing the top feature?</p>",
      "rawMarkdown": "CPMPml How does that make any sense? If you are looking at gain importances, then the top feature is the feature that reduces entropy/uncertainty of the TARGET variable the most. In the absence of this variable, you no longer have the feature that can best discriminate/regress to your TARGET. Do you have any idea why one in general could get a boost by removing the top feature?",
      "votes": null
    },
    {
      "id": "422836",
      "postDate": "11/16/2018 21:13:58",
      "content": "<p>An easy way would be the train and test data are different.  Consider the hostgalspecz parameter </p>",
      "rawMarkdown": "An easy way would be the train and test data are different.  Consider the hostgalspecz parameter",
      "votes": null
    },
    {
      "id": "422871",
      "postDate": "11/16/2018 22:51:04",
      "content": "<blockquote>\n  <p>@CPMPml How does that make any sense?</p>\n</blockquote>\n\n<p>Corey, </p>\n\n<p>Not sure what you have an issue with.  What is it that does not make sense to you?</p>",
      "rawMarkdown": "&gt; @CPMPml How does that make any sense?\n\nCorey, \n\nNot sure what you have an issue with.  What is it that does not make sense to you?",
      "votes": null
    },
    {
      "id": "422885",
      "postDate": "11/17/2018 00:26:12",
      "content": "<p>@cmpmml Have you ever met boruta for xgboost ? (asking not for this competition, but for my internship project)</p>",
      "rawMarkdown": "cmpmml Have you ever met boruta for xgboost ? (asking not for this competition, but for my internship project)",
      "votes": null
    },
    {
      "id": "423051",
      "postDate": "11/17/2018 10:57:40",
      "content": "<p>Google search is my friend here:</p>\n\n<p><a href=\"https://github.com/chasedehan/BoostARoota\">https://github.com/chasedehan/BoostARoota</a></p>\n\n<p><a href=\"https://www.kaggle.com/ogrellier/noise-analysis-of-porto-seguro-s-features\">https://www.kaggle.com/ogrellier/noise-analysis-of-porto-seguro-s-features</a></p>\n\n<p><a href=\"https://www.kaggle.com/terabytes/using-xgboost-for-feature-selection/code\">https://www.kaggle.com/terabytes/using-xgboost-for-feature-selection/code</a></p>",
      "rawMarkdown": "Google search is my friend here:\n\nhttps://github.com/chasedehan/BoostARoota\n\nhttps://www.kaggle.com/ogrellier/noise-analysis-of-porto-seguro-s-features\n\nhttps://www.kaggle.com/terabytes/using-xgboost-for-feature-selection/code",
      "votes": null
    },
    {
      "id": "1647459",
      "postDate": "01/12/2022 16:15:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> ,</p>\n<p>\"Avoid features that give a huge CV boost and a very small LB boost. This means you should care about the gap between LB and CV.\"   Could you give any guidance as to what a good gap is?  Thanks again for sharing your knowledge</p>",
      "rawMarkdown": "Hi @cpmpml ,\n\n\"Avoid features that give a huge CV boost and a very small LB boost. This means you should care about the gap between LB and CV.\"   Could you give any guidance as to what a good gap is?  Thanks again for sharing your knowledge",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1647459,
      "author_name": "yvonnef",
      "author_url": "",
      "post_date": "01/12/2022 16:15:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> ,</p>\n<p>\"Avoid features that give a huge CV boost and a very small LB boost. This means you should care about the gap between LB and CV.\"   Could you give any guidance as to what a good gap is?  Thanks again for sharing your knowledge</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417456,
      "author_name": "alexpartisan",
      "author_url": "",
      "post_date": "11/08/2018 10:06:21",
      "content": "<p>Thanks for sharing, fantastic masters like you make this community awesome... </p>",
      "votes": null,
      "replies": [
        {
          "id": 417544,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/08/2018 12:40:43",
          "content": "<p>Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417457,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "11/08/2018 10:10:37",
      "content": "<p>Great work\nI am fortunate that I have 128Gb of memory and it significantly improves scores if I don't have to aggregate on chunks - if you have a small amount of memory competitors are at a disadvantage - I tried to produce a sorted dataset so that people with small memory would be able to compete but it died. </p>",
      "votes": null,
      "replies": [
        {
          "id": 417529,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/08/2018 12:23:25",
          "content": "<p>I have 8G and compute in chunks. <code>128Gb of memory and it significantly improves scores</code> why do you think it is? I thought the score depends on features you choose and should not depend on how you compute them?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417543,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/08/2018 12:40:23",
          "content": "<p>My features only depend on each object id, hence having lots of memory would not help me.  I haven't measured precisely but my processes consume less than 6GB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417552,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "11/08/2018 12:48:19",
          "content": "<p>I get different results if I use chunks rather than the whole - I believe not all objectids are continuously ordered - might just be me</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417553,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/08/2018 12:52:22",
          "content": "<p>What depends on chunk in your case?  </p>\n\n<p>I am basically reusing Olivier's way of processing test data by chunks, and there is only one object_id that seems to appear in two chunks.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417589,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "11/08/2018 13:47:34",
          "content": "<p>I was using smaller chunks so that was probably it.  I check and get back to you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417598,
          "author_name": "druidh",
          "author_url": "",
          "post_date": "11/08/2018 13:54:45",
          "content": "<p>I have dedicated some time to split the test set into several chunks, and I have double checked that the same object_id is in consecutive rows (though the ids are not in increasing order).\nMaybe you have some feature that depends on different objects?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417621,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "11/08/2018 14:34:07",
          "content": "<p>It must be to do with my pseudo clustering - once I understand the issue I'll produce a kernel</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417791,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/08/2018 19:35:23",
          "content": "<p>\"My features only depend on each object id\", same for me. Then I do not need chunks after I have features file to calculate predict, do I ? (it's &lt; 2G csv file at the moment)\n\"I am basically reusing Olivier's way of processing test data by chunks, and there is only one object_id that seems to appear in two chunks\" --I use the same kernel. object_id  appearing in two chunks -- have not spot it, does it hurt the performance, should I find it and delete the duplicate?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417825,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/08/2018 20:40:24",
          "content": "<p>Olivier's kernel code is fine, he groups by object_id and averages before submitting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420704,
          "author_name": "niclasdoce",
          "author_url": "",
          "post_date": "11/14/2018 02:17:57",
          "content": "<p>@CPMP, are you using 'multi_logloss' as metric in your lgbm?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420894,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/14/2018 09:36:23",
          "content": "<p>I am using the metric defined in Olivier's public kernel.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417487,
      "author_name": "takuok",
      "author_url": "",
      "post_date": "11/08/2018 11:04:33",
      "content": "<p>Thanks for your sharing:)\nI have a question.</p>\n\n<blockquote>\n  <p>too many features easily overfit. For instance, my best LB score is obtained with lightgbm and less than 100 features.\n  You mean your cv is better when you use more than 100 features? (more than 100 features cv &gt; less than 100 features cv, more than 100 features LB &lt; less than 100 features LB, right?)\n  In my case, about 150~200 features got my best cv.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 417541,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/08/2018 12:38:09",
          "content": "<p>Both  CV and LB are better with less than 100 features.  My CV is well correlated with LB, when one improves the other improves.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417550,
          "author_name": "takuok",
          "author_url": "",
          "post_date": "11/08/2018 12:47:01",
          "content": "<p>Thanks!\nMy CV is well correlated too. The gap is 0.45.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417521,
      "author_name": "blondinka",
      "author_url": "",
      "post_date": "11/08/2018 12:12:52",
      "content": "<p>Thank you for sharing. My question: can you really trust those \"importance\" on the lightgbm? I get good gain for the kurtosis for example, but it does not reflect on LB. Some features i don't find meaningfull, but they have good \"importance\". I start to doubt it...</p>",
      "votes": null,
      "replies": [
        {
          "id": 417540,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/08/2018 12:36:52",
          "content": "<blockquote>\n  <p>can you really trust those \"importance\" on the lightgbm?</p>\n</blockquote>\n\n<p>No ;)</p>\n\n<p>If the importance is very low or 0, then you can try without the feature, and see how it goes.  But in one competition (Talking Data) I gained a lot by removing the feature that was said to be the most important by lightgbm (the <code>channel</code> feature).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417660,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/08/2018 15:38:42",
          "content": "<p>thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420278,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "11/13/2018 11:40:32",
          "content": "<p>@CPMP</p>\n\n<p>So why do some strange features get high importance? For example, 'flux_err_min'.\nIs it just a peculiarity of the train set? And this is overfitting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420280,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/13/2018 11:45:20",
          "content": "<p>Sergey, feature importance is not reliable at all.  For instance, in a previous competition I got a significant boost by removing the feature that lgb said was most important.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 422834,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "11/16/2018 21:05:54",
          "content": "<p>@CPMPml How does that make any sense? If you are looking at gain importances, then the top feature is the feature that reduces entropy/uncertainty of the TARGET variable the most. In the absence of this variable, you no longer have the feature that can best discriminate/regress to your TARGET. Do you have any idea why one in general could get a boost by removing the top feature?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 422836,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "11/16/2018 21:13:58",
          "content": "<p>An easy way would be the train and test data are different.  Consider the hostgalspecz parameter </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 422871,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/16/2018 22:51:04",
          "content": "<blockquote>\n  <p>@CPMPml How does that make any sense?</p>\n</blockquote>\n\n<p>Corey, </p>\n\n<p>Not sure what you have an issue with.  What is it that does not make sense to you?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417565,
      "author_name": "andreusancho",
      "author_url": "",
      "post_date": "11/08/2018 13:15:30",
      "content": "<p>Good read, thanks a lot for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 417571,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/08/2018 13:25:56",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417588,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "11/08/2018 13:47:33",
      "content": "<p>I added two topics ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417616,
      "author_name": "druidh",
      "author_url": "",
      "post_date": "11/08/2018 14:24:12",
      "content": "<p>Thank you for your post, good remainder. Totally agree with point 8... I have wasted a lot of time just to make a submission with LB score of 3.3 :'(</p>",
      "votes": null,
      "replies": [
        {
          "id": 417651,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/08/2018 15:27:59",
          "content": "<p>I've been burned by HPO more than once ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417633,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "11/08/2018 14:55:26",
      "content": "<p>Nice post! I'll follow your guideline from now on.</p>",
      "votes": null,
      "replies": [
        {
          "id": 417652,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/08/2018 15:28:25",
          "content": "<p>Are you sure you need help here? ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417642,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "11/08/2018 15:07:26",
      "content": "<p>Thank you Grandmaster</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417831,
      "author_name": "andypenrose",
      "author_url": "",
      "post_date": "11/08/2018 20:56:51",
      "content": "<p>Thanks CPMP... great and timeless advice. I strive to have the patience and discipline for 7!</p>\n\n<p>Your two extra topics are also very important. Tonight I need to sort out the crazy CV non-determinism I'm getting from Keras! </p>\n\n<p>Might I also add: Look at the data. There are some great EDA kernels for this competition. I learnt a lot looking at <a href=\"https://www.kaggle.com/mithrillion/all-classes-light-curve-characteristics\">this one</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 417884,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/08/2018 23:46:53",
          "content": "<p>Right, looking at data is key.  Thanks for reminding us of it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417855,
      "author_name": "pfichou",
      "author_url": "",
      "post_date": "11/08/2018 21:53:22",
      "content": "<p>Nice to get some good advices from a GrandMaster !</p>\n\n<p>Thanks a lot.</p>\n\n<p>I would just add : keep tracks of your scores on a piece of paper (or spreadsheet) with modification(s) made, and more than obvious : save your best submission source file and do not modify anything, just start playing with a new version based on it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 417865,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/08/2018 22:22:32",
          "content": "<p>Great advice, thanks for sharing as well!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417901,
      "author_name": "marcuslin",
      "author_url": "",
      "post_date": "11/09/2018 00:35:19",
      "content": "<p>Nice advices</p>",
      "votes": null,
      "replies": [
        {
          "id": 418001,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/09/2018 05:08:44",
          "content": "<p>Thanks, not sure you need any actually ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417995,
      "author_name": "yuval6967",
      "author_url": "",
      "post_date": "11/09/2018 04:53:23",
      "content": "<p>Thanks @CPMP\nAdding my two cents: NN's love to over fit, think of clever ways to augment the data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 417999,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/09/2018 05:08:22",
          "content": "<p>Right, tta is always key for deep learning.  Thanks for pointing it out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 418062,
      "author_name": "fatihozturk",
      "author_url": "",
      "post_date": "11/09/2018 08:17:20",
      "content": "<p>Thanks for sharing such insights. I also want to hear about feature engineering. It seems you are doing really great at feature engineering in this competition. Do you have any suggestions/tips for the community to learn how to generate meaningful features specific to this competition? Maybe some papers, or just plotting curves and getting insights from it? I'm sure that this will encourage people more.</p>\n\n<p>For my single lgbm model having a CV of 0.58X:</p>\n\n<ul>\n<li>Found my most important features after reading starter kit kernel. <a href=\"https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit\">https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit</a></li>\n<li>Then saw and applied some of features mentioned in Cesium package here: <a href=\"http://cesium-ml.org/docs/feature_table.html\">http://cesium-ml.org/docs/feature_table.html</a></li>\n<li>Following discussions (especially Grzegorz's feature ) </li>\n<li>New features having minor changes on features mentioned above.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 418121,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/09/2018 10:20:44",
          "content": "<p>Some of my features are trying to capture what I 'see' when looking at light curves from different classes.  If two classes seem to be different in a way, can I compute a feature that captures that difference?</p>\n\n<p>The main difference is the general shape of the light curve, and this is hard to capture via hand crafted features. This is why I think deep learning should be better in the end.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 418491,
      "author_name": "iprapas",
      "author_url": "",
      "post_date": "11/10/2018 01:22:55",
      "content": "<blockquote>\n  <p>Read the forum. Some very interesting features ideas can be shared,\n  see for instance Grzegorz Sionkovski' great feature, and his hint for\n  another one in this discussion.</p>\n</blockquote>\n\n<p>one more hint by <a href=\"/cpmpml\">@cpmpml</a> for useful features <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/70346#415506\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/70346#415506</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 418744,
      "author_name": "niclasdoce",
      "author_url": "",
      "post_date": "11/10/2018 14:32:06",
      "content": "<p>thanks @CPMP. this is really helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 419064,
      "author_name": "ogrellier",
      "author_url": "",
      "post_date": "11/11/2018 06:57:17",
      "content": "<p>Thanks sharing your thoughts <a href=\"/cpmpml\">@cpmpml</a></p>\n\n<p>Feature selection is particularly important here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 419340,
      "author_name": "satian",
      "author_url": "",
      "post_date": "11/11/2018 18:42:23",
      "content": "<p>Another post that is going to my personal <strong>Grandmasters 101</strong> document. Thanks @CPMP</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 419686,
      "author_name": "luizep",
      "author_url": "",
      "post_date": "11/12/2018 11:51:24",
      "content": "<p>Great, excelent post</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 419850,
      "author_name": "darbin",
      "author_url": "",
      "post_date": "11/12/2018 16:22:04",
      "content": "<p>Thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 419897,
      "author_name": "abhishekmamidi",
      "author_url": "",
      "post_date": "11/12/2018 18:26:37",
      "content": "<p>Thank you for sharing your thoughts.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 420298,
      "author_name": "vinicioswentz",
      "author_url": "",
      "post_date": "11/13/2018 12:35:33",
      "content": "<p>Nice, Thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 420509,
      "author_name": "migeruj",
      "author_url": "",
      "post_date": "11/13/2018 18:42:30",
      "content": "<p>Thanks! This work is helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 420659,
      "author_name": "iprapas",
      "author_url": "",
      "post_date": "11/14/2018 00:53:26",
      "content": "<p>Thanks again! I've incorporated some of your ideas (especially 3.) in a kernel: <a href=\"https://www.kaggle.com/iprapas/ideas-from-kernels-and-discussion-lb-1-135\">https://www.kaggle.com/iprapas/ideas-from-kernels-and-discussion-lb-1-135</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 420974,
      "author_name": "mohittanwar2323",
      "author_url": "",
      "post_date": "11/14/2018 12:19:41",
      "content": "<p>does checking for permutation Importance of the features is the good practice?\nwhat about your practice on this , how often you use this method?</p>",
      "votes": null,
      "replies": [
        {
          "id": 420999,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/14/2018 13:00:29",
          "content": "<p>You mean boruta?  I've never used it...</p>\n\n<p>It is valuable, but it is slow.  And you must make sure you use the same algorithm within it than the algorithm you use for modeling.  For instance using boruta with random forest to select features for lightgbm is silly IMHO.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 421006,
          "author_name": "mohittanwar2323",
          "author_url": "",
          "post_date": "11/14/2018 13:12:56",
          "content": "<p>not that! \nwhat i have used in some dataset for checking importance of features by doing permutation Importance using ELI5.sklearn. have you ever used this ELI5 library?\nI want to know is this a good practice or there is some best out there?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 421018,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/14/2018 13:27:30",
          "content": "<p>I never used eli5, but reading its documentation makes me think its permutation algorithm is close to the boruta algorithm.  If not then I welcome an explanation about how they differ.  Boruta in python: <a href=\"https://github.com/scikit-learn-contrib/boruta_py\">https://github.com/scikit-learn-contrib/boruta_py</a></p>\n\n<p>Some kagglers have adapted boruta to lightgbm, a google search should give you their git repos.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 422708,
          "author_name": "mohittanwar2323",
          "author_url": "",
          "post_date": "11/16/2018 16:53:55",
          "content": "<p>thanks will check all about boruta algo</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 422885,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/17/2018 00:26:12",
          "content": "<p>@cmpmml Have you ever met boruta for xgboost ? (asking not for this competition, but for my internship project)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 423051,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/17/2018 10:57:40",
          "content": "<p>Google search is my friend here:</p>\n\n<p><a href=\"https://github.com/chasedehan/BoostARoota\">https://github.com/chasedehan/BoostARoota</a></p>\n\n<p><a href=\"https://www.kaggle.com/ogrellier/noise-analysis-of-porto-seguro-s-features\">https://www.kaggle.com/ogrellier/noise-analysis-of-porto-seguro-s-features</a></p>\n\n<p><a href=\"https://www.kaggle.com/terabytes/using-xgboost-for-feature-selection/code\">https://www.kaggle.com/terabytes/using-xgboost-for-feature-selection/code</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 421116,
      "author_name": "quynguyends",
      "author_url": "",
      "post_date": "11/14/2018 15:51:02",
      "content": "<p>thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 421714,
      "author_name": "varmajipericherla",
      "author_url": "",
      "post_date": "11/15/2018 10:09:21",
      "content": "<p>Nice Insights, thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 422751,
      "author_name": "tejaswi",
      "author_url": "",
      "post_date": "11/16/2018 18:21:11",
      "content": "<p>Nicely put !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "417449": "Here are few things that worked well for me so far.  Sharing them in the hope that they will be useful to others.\n\n1. Be conservative when adding features if you use non parametric models like gbms (lightgbm, xgboosst, catboost), or NNs.  Indeed, we only have 7k examples or so, and too many features easily overfit.  For instance, my best LB score is obtained with lightgbm and less than 100 features.\n\n2. Make sure features convey information.  Features that do not involve what is measured (flux or flux_err) are detrimental.  Examples are `max(mjd)- min(mjd)`, or `mean(passband)`.\n\n3. Read the forum.  Some very interesting features ideas can be shared, see for instance Grzegorz Sionkovski' [great feature][1], and his hint for another one in [this discussion][2].\n\n4.  Overfitting is the enemy here.  Avoid features that give a huge CV boost and a very small LB boost.  This means you should care about the gap beteen LB and CV. Ahmet Erdem is making a way better job than me here for instance, and it will pay in the end I'm sure.\n\n5. Make sure you use an objective function (loss function) that mimics the competition metric. Default log loss or default multi class loss are not best here, by far.  See for instance the [Know Your Objective kernel][3] by Mithrillion.  I think there are simpler way to do it, but the general idea is right.\n\n6. Experiment ideas.  Many will fail (no CV improvement, or no LB improvement), but some will work.  Don't add features just because they look good 'intuitively'.  Always validate via cross validation score, and, possibly, by LB score.\n\n7.  Make only one change at a time, and validate it by CV, and possibly by LB, before moving to the next one.\n\n8. Don't do too much hyper parameter optimization.  Again, with a small number of examples, it is very easy to overfit to the training data, even if you use cross validation. I usually do it after a week or so in the competition, then keep same parameters until last week of competition.\n\n9. Write in the forum about your ideas or doubts.  You'll receive useful feedback in general.\n\n10. Share some of your work in kernels.  As in 9, you'll receive useful feedback in general.\n\n11.  Apply all the above when you reuse something shared by others, esp items 1, 2, and 4.\n\n12. Most important.  Never be discouraged, the competition is not over till the last minute!  It happened to me in the past that I gained several ranks in the private LB with a last minute submission.  I really mean  submitted one minute before competition deadline.\n\nedit, two more topics:\n\n13.  Make sure you have reproducible results.  Initialize all random seeds to the same value for each run.  If you don't then it is possible to think your last change is responsible for the improvement you see when it is just due to a different execution path in the algorithm.  For NN this is tricky, and keras is known to not be very good at this.  Pytorch seems better.\n\n14. Have fun!\n\n  [1]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69696#410538\n  [2]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/70725#416740\n  [3]: https://www.kaggle.com/mithrillion/know-your-objective#",
    "417456": "Thanks for sharing, fantastic masters like you make this community awesome...",
    "417457": "Great work\nI am fortunate that I have 128Gb of memory and it significantly improves scores if I don't have to aggregate on chunks - if you have a small amount of memory competitors are at a disadvantage - I tried to produce a sorted dataset so that people with small memory would be able to compete but it died.",
    "417487": "Thanks for your sharing:)\nI have a question.\n&gt;too many features easily overfit. For instance, my best LB score is obtained with lightgbm and less than 100 features.\nYou mean your cv is better when you use more than 100 features? (more than 100 features cv &gt; less than 100 features cv, more than 100 features LB &lt; less than 100 features LB, right?)\nIn my case, about 150~200 features got my best cv.",
    "417521": "Thank you for sharing. My question: can you really trust those \"importance\" on the lightgbm? I get good gain for the kurtosis for example, but it does not reflect on LB. Some features i don't find meaningfull, but they have good \"importance\". I start to doubt it...",
    "417529": "I have 8G and compute in chunks. ```128Gb of memory and it significantly improves scores``` why do you think it is? I thought the score depends on features you choose and should not depend on how you compute them?",
    "417540": "&gt; can you really trust those \"importance\" on the lightgbm?\n\nNo ;)\n\nIf the importance is very low or 0, then you can try without the feature, and see how it goes.  But in one competition (Talking Data) I gained a lot by removing the feature that was said to be the most important by lightgbm (the `channel` feature).",
    "417541": "Both  CV and LB are better with less than 100 features.  My CV is well correlated with LB, when one improves the other improves.",
    "417543": "My features only depend on each object id, hence having lots of memory would not help me.  I haven't measured precisely but my processes consume less than 6GB.",
    "417544": "Thanks.",
    "417550": "Thanks!\nMy CV is well correlated too. The gap is 0.45.",
    "417552": "I get different results if I use chunks rather than the whole - I believe not all objectids are continuously ordered - might just be me",
    "417553": "What depends on chunk in your case?  \n\nI am basically reusing Olivier's way of processing test data by chunks, and there is only one object_id that seems to appear in two chunks.",
    "417565": "Good read, thanks a lot for sharing!",
    "417571": "Thanks!",
    "417588": "I added two topics ;)",
    "417589": "I was using smaller chunks so that was probably it.  I check and get back to you",
    "417598": "I have dedicated some time to split the test set into several chunks, and I have double checked that the same object_id is in consecutive rows (though the ids are not in increasing order).\nMaybe you have some feature that depends on different objects?",
    "417616": "Thank you for your post, good remainder. Totally agree with point 8... I have wasted a lot of time just to make a submission with LB score of 3.3 :'(",
    "417621": "It must be to do with my pseudo clustering - once I understand the issue I'll produce a kernel",
    "417633": "Nice post! I'll follow your guideline from now on.",
    "417642": "Thank you Grandmaster",
    "417651": "I've been burned by HPO more than once ;)",
    "417652": "Are you sure you need help here? ;)",
    "417660": "thank you",
    "417791": "\"My features only depend on each object id\", same for me. Then I do not need chunks after I have features file to calculate predict, do I ? (it's &lt; 2G csv file at the moment)\n\"I am basically reusing Olivier's way of processing test data by chunks, and there is only one object_id that seems to appear in two chunks\" --I use the same kernel. object_id  appearing in two chunks -- have not spot it, does it hurt the performance, should I find it and delete the duplicate?",
    "417825": "Olivier's kernel code is fine, he groups by object_id and averages before submitting.",
    "417831": "Thanks CPMP... great and timeless advice. I strive to have the patience and discipline for 7!\n\nYour two extra topics are also very important. Tonight I need to sort out the crazy CV non-determinism I'm getting from Keras! \n\nMight I also add: Look at the data. There are some great EDA kernels for this competition. I learnt a lot looking at <a href=\"https://www.kaggle.com/mithrillion/all-classes-light-curve-characteristics\">this one</a>.",
    "417855": "Nice to get some good advices from a GrandMaster !\n\nThanks a lot.\n\nI would just add : keep tracks of your scores on a piece of paper (or spreadsheet) with modification(s) made, and more than obvious : save your best submission source file and do not modify anything, just start playing with a new version based on it.",
    "417865": "Great advice, thanks for sharing as well!",
    "417884": "Right, looking at data is key.  Thanks for reminding us of it.",
    "417901": "Nice advices",
    "417995": "Thanks @CPMP\nAdding my two cents: NN's love to over fit, think of clever ways to augment the data.",
    "417999": "Right, tta is always key for deep learning.  Thanks for pointing it out.",
    "418001": "Thanks, not sure you need any actually ;)",
    "418062": "Thanks for sharing such insights. I also want to hear about feature engineering. It seems you are doing really great at feature engineering in this competition. Do you have any suggestions/tips for the community to learn how to generate meaningful features specific to this competition? Maybe some papers, or just plotting curves and getting insights from it? I'm sure that this will encourage people more.\n\nFor my single lgbm model having a CV of 0.58X:\n\n- Found my most important features after reading starter kit kernel. https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit\n- Then saw and applied some of features mentioned in Cesium package here: http://cesium-ml.org/docs/feature_table.html\n- Following discussions (especially Grzegorz's feature ) \n- New features having minor changes on features mentioned above.",
    "418121": "Some of my features are trying to capture what I 'see' when looking at light curves from different classes.  If two classes seem to be different in a way, can I compute a feature that captures that difference?\n\nThe main difference is the general shape of the light curve, and this is hard to capture via hand crafted features. This is why I think deep learning should be better in the end.",
    "418491": "&gt; Read the forum. Some very interesting features ideas can be shared,\n&gt; see for instance Grzegorz Sionkovski' great feature, and his hint for\n&gt; another one in this discussion.\n\none more hint by @cpmpml for useful features https://www.kaggle.com/c/PLAsTiCC-2018/discussion/70346#415506",
    "418744": "thanks @CPMP. this is really helpful",
    "419064": "Thanks sharing your thoughts @cpmpml\n\nFeature selection is particularly important here.",
    "419340": "Another post that is going to my personal **Grandmasters 101** document. Thanks @CPMP",
    "419686": "Great, excelent post",
    "419850": "Thank you!",
    "419897": "Thank you for sharing your thoughts.",
    "420278": "CPMP\n\nSo why do some strange features get high importance? For example, 'flux_err_min'.\nIs it just a peculiarity of the train set? And this is overfitting.",
    "420280": "Sergey, feature importance is not reliable at all.  For instance, in a previous competition I got a significant boost by removing the feature that lgb said was most important.",
    "420298": "Nice, Thank you!",
    "420509": "Thanks! This work is helpful",
    "420659": "Thanks again! I've incorporated some of your ideas (especially 3.) in a kernel: https://www.kaggle.com/iprapas/ideas-from-kernels-and-discussion-lb-1-135",
    "420704": "CPMP, are you using 'multi_logloss' as metric in your lgbm?",
    "420894": "I am using the metric defined in Olivier's public kernel.",
    "420974": "does checking for permutation Importance of the features is the good practice?\nwhat about your practice on this , how often you use this method?",
    "420999": "You mean boruta?  I've never used it...\n\nIt is valuable, but it is slow.  And you must make sure you use the same algorithm within it than the algorithm you use for modeling.  For instance using boruta with random forest to select features for lightgbm is silly IMHO.",
    "421006": "not that! \nwhat i have used in some dataset for checking importance of features by doing permutation Importance using ELI5.sklearn. have you ever used this ELI5 library?\nI want to know is this a good practice or there is some best out there?",
    "421018": "I never used eli5, but reading its documentation makes me think its permutation algorithm is close to the boruta algorithm.  If not then I welcome an explanation about how they differ.  Boruta in python: https://github.com/scikit-learn-contrib/boruta_py\n\nSome kagglers have adapted boruta to lightgbm, a google search should give you their git repos.",
    "421116": "thanks!",
    "421714": "Nice Insights, thanks.",
    "422708": "thanks will check all about boruta algo",
    "422751": "Nicely put !",
    "422834": "CPMPml How does that make any sense? If you are looking at gain importances, then the top feature is the feature that reduces entropy/uncertainty of the TARGET variable the most. In the absence of this variable, you no longer have the feature that can best discriminate/regress to your TARGET. Do you have any idea why one in general could get a boost by removing the top feature?",
    "422836": "An easy way would be the train and test data are different.  Consider the hostgalspecz parameter",
    "422871": "&gt; @CPMPml How does that make any sense?\n\nCorey, \n\nNot sure what you have an issue with.  What is it that does not make sense to you?",
    "422885": "cmpmml Have you ever met boruta for xgboost ? (asking not for this competition, but for my internship project)",
    "423051": "Google search is my friend here:\n\nhttps://github.com/chasedehan/BoostARoota\n\nhttps://www.kaggle.com/ogrellier/noise-analysis-of-porto-seguro-s-features\n\nhttps://www.kaggle.com/terabytes/using-xgboost-for-feature-selection/code",
    "1647459": "Hi @cpmpml ,\n\n\"Avoid features that give a huge CV boost and a very small LB boost. This means you should care about the gap between LB and CV.\"   Could you give any guidance as to what a good gap is?  Thanks again for sharing your knowledge"
  },
  "source": "meta"
}