{
  "id": 75134,
  "title": "#13 Solution, true story: tries and fails",
  "url": "/competitions/PLAsTiCC-2018/writeups/day-meets-night-13-solution-true-story-tries-and-f",
  "author_name": "",
  "post_date": "2018-12-22T02:38:34.343Z",
  "votes": 60,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Thank you for this challenge, lots of fun and learning, data are good and clean, answers from organizers are timely. Very well organized. \nHere is our write up.</p>\n\n<p><strong>What worked:</strong> \nOur final solution was based on LGBM model inspired by Oliver kernel, feature design and selection, and augmentation.</p>\n\n<p><strong>Features:</strong> </p>\n\n<ol>\n<li><p>Magnitudes. Both aggregated and per band, but without k-correction till the last day. As we were not quite sure how to interpret detected flag exactly, we decided to try both: for detected == 1 and for all data. Surprisingly, magnitude features for detected == 1 worked better on LB (not on CV though). I attribute it to the outliers removal with detected == 1 and also class 99. Indeed, to detect unknown -- you have to be 100% sure, so the detected points may be important for 99 discovery. Known objects are a different story: we know them and the confidence level of error can be lower. </p></li>\n<li><p>Parametric curve fittings. We used models of Bazin (<a href=\"https://arxiv.org/pdf/0904.1066.pdf\">https://arxiv.org/pdf/0904.1066.pdf</a>) and Karpenka (<a href=\"https://arxiv.org/abs/1208.1264\">https://arxiv.org/abs/1208.1264</a>), Gaussian fits to define peak-width, and the decay slope fits of supernova light curves in the log scale. Fitting loss was one of the most important feature in all the parametric curves.</p></li>\n<li><p>Basic features from cesium package and statistics. Those based on ratios, std, and skew made it to the final. </p></li>\n<li><p>From the Bazin equation we derived the position of the max for the incomplete light curves and derived magnitudes from fits as well. They helped for some bands. We also tried a range of other parameters, m15 and m-10 parameters, peak width at different levels -- but they did not help. </p></li>\n<li><p>We checked all features from feets. CAR tau from feets was particularly helpful, but it takes ages to calculate. So, Mithrillion optimized a code to make it 10 times faster! If astronomers are interested, I think we can put it in a kernel.</p></li>\n<li><p>Autocorrelation lag 0, also estimated peak widths made it to the final. </p></li>\n<li><p>Finally, on the last 2 days I was pointed to the proper calculation of k-correction by Kyle (thank you). I used the calculator from here (<a href=\"http://kcor.sai.msu.ru/getthecode/\">http://kcor.sai.msu.ru/getthecode/</a>) and the colors as suggested here (<a href=\"https://arxiv.org/pdf/1410.8139.pdf\">https://arxiv.org/pdf/1410.8139.pdf</a>). It's only valid for relatively small red shifts, but the majority of data lie in this region. \nThe k-correction of magnitudes per band did not help, but along the way we calculated colors: both from detected magnitudes and Bazin fitted magnitudes --&gt; that's how we got from 0.836 to 0.815 on a single model </p></li>\n</ol>\n\n<p>Feature selection: we tried both RFE from sklearn and manual tuning while looking at eli5 importance and deleting highly correlated features. Manual approach and eli5 importance worked out much better than RFE.</p>\n\n<p><strong>Battling over fitting:</strong></p>\n\n<ol>\n<li><p>To overcome over-fitting we first augmented train 10x times using flux variation within normal distribution of the corresponding flux_err, similar for photoz values (except for galactic classes). We also introduced small time shift to bring translation symmetry. It was meant for NN mostly, that we did not use at the end.</p></li>\n<li><p>We twisted parameters of classifier, especially deep tree and max bin gave a change.</p></li>\n<li><p>We realized that we have ~30% of ddf samples in train and only 1% in test. So we are likely to over fit to the ddf samples. First, we down-sampled ddf in the current augmentation and re-selected features checking eli5 importance for wide field drilling samples only (ddf = 0). Secondly, we changed augmentation, adding 30x train from ddf = 0 samples and leaving ddf samples only in the original train fold. That helped. Unfortunately, this idea came to us in the last 48 hours, so we did what we could manage within that time, and feature selection -- optimization still could have been better. </p></li>\n</ol>\n\n<p>With 30x augmentation CV 0.5 LB 0.815, gap 0.315</p>\n\n<p>Not much ensembling, 50/50 blend of two LGBM models with slightly modified features and parameters was the final blend</p>\n\n<p><strong>What did not work:</strong></p>\n\n<ol>\n<li><p>Autoencoder. Although it has shown quite nice curves reconstruction, unfortunately, non of it's features made it to the final. We tried to include autoencoder loss as well -- did not help.</p></li>\n<li><p>More parametric fitting. Although we tried adding models with double-peak (<a href=\"https://arxiv.org/abs/1208.1264\">https://arxiv.org/abs/1208.1264</a> ), and add tau per band as well, nothing really improved LB compared to just our very first Bazin parametric fit.  </p></li>\n<li><p>Gaussian Processes augmentation. We tried to use GP to create more non-ddf samples for augmentation, but got a bit of an issues with detected and flux error, so far this did not work out, but the idea is good I think. Mithrillion gave more comments under Kyle's solution.</p></li>\n<li><p>PU classification to identify class 99. We tried to detect negative labels (99 labels that are not in train) for 99 class prediction using a PU classifier (<a href=\"https://arxiv.org/pdf/1605.06955.pdf\">https://arxiv.org/pdf/1605.06955.pdf</a> ). It did not work out...</p></li>\n<li><p>Class 99 probing. We tried to find class 99 in test using LB probing. We noticed that the percent of high probability 99 objects is in the far z region. Indeed, apart from themselves, people tend to know better what is closer than what is far. Also, our visible Universe keeps expanding as the light from far away objects keeps reaching the Earth, so we expect more unknown from far, I think. We tried to find 99 among photoz &gt; 2.5 and high probability of 99 class in Oliver and Scirpus methods. Guess what ? ... Bingo! It did not work out.</p></li>\n</ol>\n\n<p><strong>What we did right:</strong></p>\n\n<ol>\n<li><p>We teamed up! Teams have power, totally recommend, it's more fun and more learning and better results.</p></li>\n<li><p>We did not give up when we fell to place 19 just a three days before the end. The competition is not over till the last submit! It's in the last two-three days that we actually realized our mistakes, fixed what we could and managed to improve from 0.87 to 0.815 on a single model getting to place 12. Then it was just a bad luck on private...  </p></li>\n<li><p>We asked questions on forum. And people answered. Special thanks to Kyle, CPMP and organizers.</p></li>\n<li><p>We read, we learned, we tried, we discussed, we tried again and we learned more...</p></li>\n</ol>\n\n<p><strong>What we did wrong:</strong></p>\n\n<ol>\n<li><p>We did not utilize our submissions properly in the beginning of competition. When we had plenty of them -- that's when we should have tried more features on LB as well and more experimenting with parameters earlier. Last days every submission counts and you cannot just twist parameters while also checking on LB.</p></li>\n<li><p>We did not bring ddf percent close to the test data until the last 48 hours. We should have looked at data more and try in earlier. In general, I think 99% of answers are in the data, so if you feel stuck -- look at them again, look at them more.</p></li>\n<li><p>Strategy. We first did things that should work, like parametric fits from papers and magnitudes, and we left fun experiments to the end. At the end everything has speed up,  and we did not had time to properly utilize those experiments. Maybe we should have done it earlier when the time was not pressing. It's my third kaggle, so I still do not know what the best strategy is. More comments from Grandmasters here would be helpful.</p></li>\n<li><p>We did not plan submissions ahead. Trying to generate 5 submissions in evening to utilize all of them did not work out. I think it's nice to plan what we want to probe on LB and prepare ahead, so to have a full use of submissions. </p></li>\n<li><p>Trying to find a black cat in a dark room. We spent too much time trying to find \"the secret\", identify class 99 and build a classifier to find it. Instead, we could do more traditional things that bring incremental improvement, like better feature selection and parameters tuning. We left it to the last few days, but then its not enough time and submissions left.</p></li>\n</ol>\n\n<p>*The hardest thing of all is to find a black cat in a dark room. Especially if there is no cat. Confucius</p>\n\n<p><strong>What I still do not understand:</strong></p>\n\n<p>We noticed a considerable variations in folds loss with LGBM, it is stratified, it is still big for wdf only training. Anyone knows, why is that?</p>\n\n<p>NN classifier did not want to improve anymore after around 0.94 even with augmentation. Why is that? We tried to twist parameters.</p>\n\n<p>Feature selection: what are the best feature selection methods people use here? I've seen comment about boruta, but found it for random forest only, does it exist for lgbm and xgboost? What people use here?</p>\n\n<p>It looks from other solutions that we did things right, so what was the most important that we missed to cross 0.8 margin on a single model (we had 0.815)? And why there was such a shake on private? </p>\n\n<p><strong>Personal outcome:</strong> It was a hard work and lot's of learning. The last few days were especially crazy for me. I was literally holding a baby with one hand and programming with the other, trying to concentrate and compete for gold against people who can use both hands, lol. No much sleep, no much food, no time for anything else. </p>\n\n<p>Still, it was fun and learning experience! Thanks to kaggle, organizers and most of all -- to my best-ever-teammates!!!</p>\n\n<p>PS. Dear kaggle people, please consider: \n- adding emotions (happy, sad, thinking, and a little penguin dancing). \n- consider adding a sign: \nKAGGLE IS ADDICTIVE ! ENTER AT YOUR OWN RISK !!!</p>",
  "messages": [
    {
      "id": "441535",
      "postDate": "12/18/2018 18:59:21",
      "content": "<p>Thank you for this challenge, lots of fun and learning, data are good and clean, answers from organizers are timely. Very well organized. \nHere is our write up.</p>\n\n<p><strong>What worked:</strong> \nOur final solution was based on LGBM model inspired by Oliver kernel, feature design and selection, and augmentation.</p>\n\n<p><strong>Features:</strong> </p>\n\n<ol>\n<li><p>Magnitudes. Both aggregated and per band, but without k-correction till the last day. As we were not quite sure how to interpret detected flag exactly, we decided to try both: for detected == 1 and for all data. Surprisingly, magnitude features for detected == 1 worked better on LB (not on CV though). I attribute it to the outliers removal with detected == 1 and also class 99. Indeed, to detect unknown -- you have to be 100% sure, so the detected points may be important for 99 discovery. Known objects are a different story: we know them and the confidence level of error can be lower. </p></li>\n<li><p>Parametric curve fittings. We used models of Bazin (<a href=\"https://arxiv.org/pdf/0904.1066.pdf\">https://arxiv.org/pdf/0904.1066.pdf</a>) and Karpenka (<a href=\"https://arxiv.org/abs/1208.1264\">https://arxiv.org/abs/1208.1264</a>), Gaussian fits to define peak-width, and the decay slope fits of supernova light curves in the log scale. Fitting loss was one of the most important feature in all the parametric curves.</p></li>\n<li><p>Basic features from cesium package and statistics. Those based on ratios, std, and skew made it to the final. </p></li>\n<li><p>From the Bazin equation we derived the position of the max for the incomplete light curves and derived magnitudes from fits as well. They helped for some bands. We also tried a range of other parameters, m15 and m-10 parameters, peak width at different levels -- but they did not help. </p></li>\n<li><p>We checked all features from feets. CAR tau from feets was particularly helpful, but it takes ages to calculate. So, Mithrillion optimized a code to make it 10 times faster! If astronomers are interested, I think we can put it in a kernel.</p></li>\n<li><p>Autocorrelation lag 0, also estimated peak widths made it to the final. </p></li>\n<li><p>Finally, on the last 2 days I was pointed to the proper calculation of k-correction by Kyle (thank you). I used the calculator from here (<a href=\"http://kcor.sai.msu.ru/getthecode/\">http://kcor.sai.msu.ru/getthecode/</a>) and the colors as suggested here (<a href=\"https://arxiv.org/pdf/1410.8139.pdf\">https://arxiv.org/pdf/1410.8139.pdf</a>). It's only valid for relatively small red shifts, but the majority of data lie in this region. \nThe k-correction of magnitudes per band did not help, but along the way we calculated colors: both from detected magnitudes and Bazin fitted magnitudes --&gt; that's how we got from 0.836 to 0.815 on a single model </p></li>\n</ol>\n\n<p>Feature selection: we tried both RFE from sklearn and manual tuning while looking at eli5 importance and deleting highly correlated features. Manual approach and eli5 importance worked out much better than RFE.</p>\n\n<p><strong>Battling over fitting:</strong></p>\n\n<ol>\n<li><p>To overcome over-fitting we first augmented train 10x times using flux variation within normal distribution of the corresponding flux_err, similar for photoz values (except for galactic classes). We also introduced small time shift to bring translation symmetry. It was meant for NN mostly, that we did not use at the end.</p></li>\n<li><p>We twisted parameters of classifier, especially deep tree and max bin gave a change.</p></li>\n<li><p>We realized that we have ~30% of ddf samples in train and only 1% in test. So we are likely to over fit to the ddf samples. First, we down-sampled ddf in the current augmentation and re-selected features checking eli5 importance for wide field drilling samples only (ddf = 0). Secondly, we changed augmentation, adding 30x train from ddf = 0 samples and leaving ddf samples only in the original train fold. That helped. Unfortunately, this idea came to us in the last 48 hours, so we did what we could manage within that time, and feature selection -- optimization still could have been better. </p></li>\n</ol>\n\n<p>With 30x augmentation CV 0.5 LB 0.815, gap 0.315</p>\n\n<p>Not much ensembling, 50/50 blend of two LGBM models with slightly modified features and parameters was the final blend</p>\n\n<p><strong>What did not work:</strong></p>\n\n<ol>\n<li><p>Autoencoder. Although it has shown quite nice curves reconstruction, unfortunately, non of it's features made it to the final. We tried to include autoencoder loss as well -- did not help.</p></li>\n<li><p>More parametric fitting. Although we tried adding models with double-peak (<a href=\"https://arxiv.org/abs/1208.1264\">https://arxiv.org/abs/1208.1264</a> ), and add tau per band as well, nothing really improved LB compared to just our very first Bazin parametric fit.  </p></li>\n<li><p>Gaussian Processes augmentation. We tried to use GP to create more non-ddf samples for augmentation, but got a bit of an issues with detected and flux error, so far this did not work out, but the idea is good I think. Mithrillion gave more comments under Kyle's solution.</p></li>\n<li><p>PU classification to identify class 99. We tried to detect negative labels (99 labels that are not in train) for 99 class prediction using a PU classifier (<a href=\"https://arxiv.org/pdf/1605.06955.pdf\">https://arxiv.org/pdf/1605.06955.pdf</a> ). It did not work out...</p></li>\n<li><p>Class 99 probing. We tried to find class 99 in test using LB probing. We noticed that the percent of high probability 99 objects is in the far z region. Indeed, apart from themselves, people tend to know better what is closer than what is far. Also, our visible Universe keeps expanding as the light from far away objects keeps reaching the Earth, so we expect more unknown from far, I think. We tried to find 99 among photoz &gt; 2.5 and high probability of 99 class in Oliver and Scirpus methods. Guess what ? ... Bingo! It did not work out.</p></li>\n</ol>\n\n<p><strong>What we did right:</strong></p>\n\n<ol>\n<li><p>We teamed up! Teams have power, totally recommend, it's more fun and more learning and better results.</p></li>\n<li><p>We did not give up when we fell to place 19 just a three days before the end. The competition is not over till the last submit! It's in the last two-three days that we actually realized our mistakes, fixed what we could and managed to improve from 0.87 to 0.815 on a single model getting to place 12. Then it was just a bad luck on private...  </p></li>\n<li><p>We asked questions on forum. And people answered. Special thanks to Kyle, CPMP and organizers.</p></li>\n<li><p>We read, we learned, we tried, we discussed, we tried again and we learned more...</p></li>\n</ol>\n\n<p><strong>What we did wrong:</strong></p>\n\n<ol>\n<li><p>We did not utilize our submissions properly in the beginning of competition. When we had plenty of them -- that's when we should have tried more features on LB as well and more experimenting with parameters earlier. Last days every submission counts and you cannot just twist parameters while also checking on LB.</p></li>\n<li><p>We did not bring ddf percent close to the test data until the last 48 hours. We should have looked at data more and try in earlier. In general, I think 99% of answers are in the data, so if you feel stuck -- look at them again, look at them more.</p></li>\n<li><p>Strategy. We first did things that should work, like parametric fits from papers and magnitudes, and we left fun experiments to the end. At the end everything has speed up,  and we did not had time to properly utilize those experiments. Maybe we should have done it earlier when the time was not pressing. It's my third kaggle, so I still do not know what the best strategy is. More comments from Grandmasters here would be helpful.</p></li>\n<li><p>We did not plan submissions ahead. Trying to generate 5 submissions in evening to utilize all of them did not work out. I think it's nice to plan what we want to probe on LB and prepare ahead, so to have a full use of submissions. </p></li>\n<li><p>Trying to find a black cat in a dark room. We spent too much time trying to find \"the secret\", identify class 99 and build a classifier to find it. Instead, we could do more traditional things that bring incremental improvement, like better feature selection and parameters tuning. We left it to the last few days, but then its not enough time and submissions left.</p></li>\n</ol>\n\n<p>*The hardest thing of all is to find a black cat in a dark room. Especially if there is no cat. Confucius</p>\n\n<p><strong>What I still do not understand:</strong></p>\n\n<p>We noticed a considerable variations in folds loss with LGBM, it is stratified, it is still big for wdf only training. Anyone knows, why is that?</p>\n\n<p>NN classifier did not want to improve anymore after around 0.94 even with augmentation. Why is that? We tried to twist parameters.</p>\n\n<p>Feature selection: what are the best feature selection methods people use here? I've seen comment about boruta, but found it for random forest only, does it exist for lgbm and xgboost? What people use here?</p>\n\n<p>It looks from other solutions that we did things right, so what was the most important that we missed to cross 0.8 margin on a single model (we had 0.815)? And why there was such a shake on private? </p>\n\n<p><strong>Personal outcome:</strong> It was a hard work and lot's of learning. The last few days were especially crazy for me. I was literally holding a baby with one hand and programming with the other, trying to concentrate and compete for gold against people who can use both hands, lol. No much sleep, no much food, no time for anything else. </p>\n\n<p>Still, it was fun and learning experience! Thanks to kaggle, organizers and most of all -- to my best-ever-teammates!!!</p>\n\n<p>PS. Dear kaggle people, please consider: \n- adding emotions (happy, sad, thinking, and a little penguin dancing). \n- consider adding a sign: \nKAGGLE IS ADDICTIVE ! ENTER AT YOUR OWN RISK !!!</p>",
      "rawMarkdown": "Thank you for this challenge, lots of fun and learning, data are good and clean, answers from organizers are timely. Very well organized. \nHere is our write up.\n\n**What worked:** \nOur final solution was based on LGBM model inspired by Oliver kernel, feature design and selection, and augmentation.\n\n**Features:** \n\n1. Magnitudes. Both aggregated and per band, but without k-correction till the last day. As we were not quite sure how to interpret detected flag exactly, we decided to try both: for detected == 1 and for all data. Surprisingly, magnitude features for detected == 1 worked better on LB (not on CV though). I attribute it to the outliers removal with detected == 1 and also class 99. Indeed, to detect unknown -- you have to be 100% sure, so the detected points may be important for 99 discovery. Known objects are a different story: we know them and the confidence level of error can be lower. \n \n2. Parametric curve fittings. We used models of Bazin (https://arxiv.org/pdf/0904.1066.pdf) and Karpenka (https://arxiv.org/abs/1208.1264), Gaussian fits to define peak-width, and the decay slope fits of supernova light curves in the log scale. Fitting loss was one of the most important feature in all the parametric curves.\n\n3. Basic features from cesium package and statistics. Those based on ratios, std, and skew made it to the final. \n\n4. From the Bazin equation we derived the position of the max for the incomplete light curves and derived magnitudes from fits as well. They helped for some bands. We also tried a range of other parameters, m15 and m-10 parameters, peak width at different levels -- but they did not help. \n\n5. We checked all features from feets. CAR tau from feets was particularly helpful, but it takes ages to calculate. So, Mithrillion optimized a code to make it 10 times faster! If astronomers are interested, I think we can put it in a kernel.\n\n6. Autocorrelation lag 0, also estimated peak widths made it to the final. \n\n7. Finally, on the last 2 days I was pointed to the proper calculation of k-correction by Kyle (thank you). I used the calculator from here (http://kcor.sai.msu.ru/getthecode/) and the colors as suggested here (https://arxiv.org/pdf/1410.8139.pdf). It's only valid for relatively small red shifts, but the majority of data lie in this region. \nThe k-correction of magnitudes per band did not help, but along the way we calculated colors: both from detected magnitudes and Bazin fitted magnitudes --&gt; that's how we got from 0.836 to 0.815 on a single model \n\nFeature selection: we tried both RFE from sklearn and manual tuning while looking at eli5 importance and deleting highly correlated features. Manual approach and eli5 importance worked out much better than RFE.\n  \n\n**Battling over fitting:**\n\n1. To overcome over-fitting we first augmented train 10x times using flux variation within normal distribution of the corresponding flux_err, similar for photoz values (except for galactic classes). We also introduced small time shift to bring translation symmetry. It was meant for NN mostly, that we did not use at the end.\n\n2. We twisted parameters of classifier, especially deep tree and max bin gave a change.\n\n3. We realized that we have ~30% of ddf samples in train and only 1% in test. So we are likely to over fit to the ddf samples. First, we down-sampled ddf in the current augmentation and re-selected features checking eli5 importance for wide field drilling samples only (ddf = 0). Secondly, we changed augmentation, adding 30x train from ddf = 0 samples and leaving ddf samples only in the original train fold. That helped. Unfortunately, this idea came to us in the last 48 hours, so we did what we could manage within that time, and feature selection -- optimization still could have been better. \n\nWith 30x augmentation CV 0.5 LB 0.815, gap 0.315\n\nNot much ensembling, 50/50 blend of two LGBM models with slightly modified features and parameters was the final blend\n\n**What did not work:**\n\n1. Autoencoder. Although it has shown quite nice curves reconstruction, unfortunately, non of it's features made it to the final. We tried to include autoencoder loss as well -- did not help.\n\n2. More parametric fitting. Although we tried adding models with double-peak (https://arxiv.org/abs/1208.1264 ), and add tau per band as well, nothing really improved LB compared to just our very first Bazin parametric fit.  \n\n3. Gaussian Processes augmentation. We tried to use GP to create more non-ddf samples for augmentation, but got a bit of an issues with detected and flux error, so far this did not work out, but the idea is good I think. Mithrillion gave more comments under Kyle's solution.\n\n4. PU classification to identify class 99. We tried to detect negative labels (99 labels that are not in train) for 99 class prediction using a PU classifier (https://arxiv.org/pdf/1605.06955.pdf ). It did not work out...\n\n5. Class 99 probing. We tried to find class 99 in test using LB probing. We noticed that the percent of high probability 99 objects is in the far z region. Indeed, apart from themselves, people tend to know better what is closer than what is far. Also, our visible Universe keeps expanding as the light from far away objects keeps reaching the Earth, so we expect more unknown from far, I think. We tried to find 99 among photoz &gt; 2.5 and high probability of 99 class in Oliver and Scirpus methods. Guess what ? ... Bingo! It did not work out.\n\n**What we did right:**\n\n1. We teamed up! Teams have power, totally recommend, it's more fun and more learning and better results.\n\n2. We did not give up when we fell to place 19 just a three days before the end. The competition is not over till the last submit! It's in the last two-three days that we actually realized our mistakes, fixed what we could and managed to improve from 0.87 to 0.815 on a single model getting to place 12. Then it was just a bad luck on private...  \n\n3. We asked questions on forum. And people answered. Special thanks to Kyle, CPMP and organizers.\n\n4. We read, we learned, we tried, we discussed, we tried again and we learned more...\n\n**What we did wrong:**\n\n1. We did not utilize our submissions properly in the beginning of competition. When we had plenty of them -- that's when we should have tried more features on LB as well and more experimenting with parameters earlier. Last days every submission counts and you cannot just twist parameters while also checking on LB.\n\n2. We did not bring ddf percent close to the test data until the last 48 hours. We should have looked at data more and try in earlier. In general, I think 99% of answers are in the data, so if you feel stuck -- look at them again, look at them more.\n\n3. Strategy. We first did things that should work, like parametric fits from papers and magnitudes, and we left fun experiments to the end. At the end everything has speed up,  and we did not had time to properly utilize those experiments. Maybe we should have done it earlier when the time was not pressing. It's my third kaggle, so I still do not know what the best strategy is. More comments from Grandmasters here would be helpful.\n\n4. We did not plan submissions ahead. Trying to generate 5 submissions in evening to utilize all of them did not work out. I think it's nice to plan what we want to probe on LB and prepare ahead, so to have a full use of submissions. \n\n5. Trying to find a black cat in a dark room. We spent too much time trying to find \"the secret\", identify class 99 and build a classifier to find it. Instead, we could do more traditional things that bring incremental improvement, like better feature selection and parameters tuning. We left it to the last few days, but then its not enough time and submissions left.\n\n*The hardest thing of all is to find a black cat in a dark room. Especially if there is no cat. Confucius\n\n**What I still do not understand:**\n\nWe noticed a considerable variations in folds loss with LGBM, it is stratified, it is still big for wdf only training. Anyone knows, why is that?\n\nNN classifier did not want to improve anymore after around 0.94 even with augmentation. Why is that? We tried to twist parameters.\n\nFeature selection: what are the best feature selection methods people use here? I've seen comment about boruta, but found it for random forest only, does it exist for lgbm and xgboost? What people use here?\n\nIt looks from other solutions that we did things right, so what was the most important that we missed to cross 0.8 margin on a single model (we had 0.815)? And why there was such a shake on private? \n\n**Personal outcome:** It was a hard work and lot's of learning. The last few days were especially crazy for me. I was literally holding a baby with one hand and programming with the other, trying to concentrate and compete for gold against people who can use both hands, lol. No much sleep, no much food, no time for anything else. \n\nStill, it was fun and learning experience! Thanks to kaggle, organizers and most of all -- to my best-ever-teammates!!!\n\nPS. Dear kaggle people, please consider: \n- adding emotions (happy, sad, thinking, and a little penguin dancing). \n- consider adding a sign: \nKAGGLE IS ADDICTIVE ! ENTER AT YOUR OWN RISK !!!",
      "votes": null
    },
    {
      "id": "441561",
      "postDate": "12/18/2018 19:38:30",
      "content": "<p>Thanks for sharing and congrats on your result.  You share a lot fo what I have used as well, not sure why we fared better.  Probably blending with RNN and MLP was the difference.</p>",
      "rawMarkdown": "Thanks for sharing and congrats on your result.  You share a lot fo what I have used as well, not sure why we fared better.  Probably blending with RNN and MLP was the difference.",
      "votes": null
    },
    {
      "id": "441576",
      "postDate": "12/18/2018 19:55:07",
      "content": "<p>Thanks for sharing and congrats. \nThe term <strong>\"Black cat in a dark room\"</strong> fits class99 exactly! \nFrom our probing I can say for sure that some of class99 object were classified as class42 (or class52) with very high probability (close to 1) by all of our classifiers.   </p>",
      "rawMarkdown": "Thanks for sharing and congrats. \nThe term **\"Black cat in a dark room\"** fits class99 exactly! \nFrom our probing I can say for sure that some of class99 object were classified as class42 (or class52) with very high probability (close to 1) by all of our classifiers.",
      "votes": null
    },
    {
      "id": "441594",
      "postDate": "12/18/2018 20:26:16",
      "content": "<blockquote>\n  <p>Feature selection: what are the best feature selection methods people use here? I've seen comment about boruta, but found it for random forest only, does it exist for lgbm and xgboost? What people use here?</p>\n</blockquote>\n\n<p>Something I really want to know. Thanks for asking. I hope the masters/experts will share some of their strategy. </p>\n\n<p>P.S. Congrats and thanks for sharing your method. 😀</p>",
      "rawMarkdown": "&gt;Feature selection: what are the best feature selection methods people use here? I've seen comment about boruta, but found it for random forest only, does it exist for lgbm and xgboost? What people use here?\n\nSomething I really want to know. Thanks for asking. I hope the masters/experts will share some of their strategy. \n\nP.S. Congrats and thanks for sharing your method. 😀",
      "votes": null
    },
    {
      "id": "441730",
      "postDate": "12/19/2018 01:14:09",
      "content": "<p>This is so sad for us... I was the one with the strong belief that to tackle what was supposed to be 10-20% of the test data, we should not rely on a probed formula (kind of cheeky...) that does not make a lot of sense probabilisitically, so I figured we should utilise the one source of information we did have - class 99 is only in the test set. However, even after I constructed a subset of the test set with identical ddf/wdf ratio as the training set and mitigated hostgal_photoz differences as much as possible, it appears that there still exists residual differences unrelated to class 99 between the two sets, making utilising Positive-Unlabeled (PU) classification difficult. We were keen on discovering a class 99 approach that only relies on LB for validation, not model building, but we were not successful. It appears that even the very top performers in this competition did not manage to do it unfortunately. Who knows, science is hard...</p>",
      "rawMarkdown": "This is so sad for us... I was the one with the strong belief that to tackle what was supposed to be 10-20% of the test data, we should not rely on a probed formula (kind of cheeky...) that does not make a lot of sense probabilisitically, so I figured we should utilise the one source of information we did have - class 99 is only in the test set. However, even after I constructed a subset of the test set with identical ddf/wdf ratio as the training set and mitigated hostgal_photoz differences as much as possible, it appears that there still exists residual differences unrelated to class 99 between the two sets, making utilising Positive-Unlabeled (PU) classification difficult. We were keen on discovering a class 99 approach that only relies on LB for validation, not model building, but we were not successful. It appears that even the very top performers in this competition did not manage to do it unfortunately. Who knows, science is hard...",
      "votes": null
    },
    {
      "id": "441758",
      "postDate": "12/19/2018 02:49:45",
      "content": "<p>Congratulations and with you gold in the next competitions :)</p>",
      "rawMarkdown": "Congratulations and with you gold in the next competitions :)",
      "votes": null
    },
    {
      "id": "441871",
      "postDate": "12/19/2018 07:24:17",
      "content": "<p>Congrats and thank you:)\nI think it is too difficult question <code>what feature selection is best</code>. \nMy feature selection method was simple.\n- Training all features and calculate importance.\n- Select threshold with CV and cut features. ex. top250 features\n- Then drop one by one and check CV. If CV improves, it drops. (This part's improvement is little.)</p>\n\n<p>I think this feature selection method is not best. But it takes little time to select features so can focus on other things like feature engineering or post processing.</p>",
      "rawMarkdown": "Congrats and thank you:)\nI think it is too difficult question `what feature selection is best`. \nMy feature selection method was simple.\n- Training all features and calculate importance.\n- Select threshold with CV and cut features. ex. top250 features\n- Then drop one by one and check CV. If CV improves, it drops. (This part's improvement is little.)\n\nI think this feature selection method is not best. But it takes little time to select features so can focus on other things like feature engineering or post processing.",
      "votes": null
    },
    {
      "id": "441883",
      "postDate": "12/19/2018 07:43:02",
      "content": "<p>This is the process that we automated with RFECV. It works most of the time, especially when removing the bottom-ranking features, but towards the end of the process we found that it was pretty inadequate at determining which features are absolutely safe to remove. Sometimes features are correlated so when you remove one in a correlated group, you improve the CV a little bit by reducing feature redundancy, but you still lose a little bit information. Sometimes the redundant features are on the top of the importance list that you cannot eliminate from the bottom up. In some cases, I have to manually decorrelate the features to make them work, but it does not appear to do the trick for all correlated groups.</p>",
      "rawMarkdown": "This is the process that we automated with RFECV. It works most of the time, especially when removing the bottom-ranking features, but towards the end of the process we found that it was pretty inadequate at determining which features are absolutely safe to remove. Sometimes features are correlated so when you remove one in a correlated group, you improve the CV a little bit by reducing feature redundancy, but you still lose a little bit information. Sometimes the redundant features are on the top of the importance list that you cannot eliminate from the bottom up. In some cases, I have to manually decorrelate the features to make them work, but it does not appear to do the trick for all correlated groups.",
      "votes": null
    },
    {
      "id": "441894",
      "postDate": "12/19/2018 08:12:19",
      "content": "<p>LGBM is lobast for redundant features(ex.correlated features). It is true that removing some correlated features improve score, but it isn't much. I thought there are some features that was few importance but improved score well. So I added one by one on cuted features. But this improvement was also few.\nIt is the reason that I didn't focus on feature selection much. More important thing is to find good features.</p>",
      "rawMarkdown": "LGBM is lobast for redundant features(ex.correlated features). It is true that removing some correlated features improve score, but it isn't much. I thought there are some features that was few importance but improved score well. So I added one by one on cuted features. But this improvement was also few.\nIt is the reason that I didn't focus on feature selection much. More important thing is to find good features.",
      "votes": null
    },
    {
      "id": "441947",
      "postDate": "12/19/2018 09:21:04",
      "content": "<p>The default feature importance ranking I got from LGBM was slightly different from the SHAP ranking. I somehow thought the SHAP ranking made more sense. \n<a href=\"/mithrillion\">@mithrillion</a> Thanks. will check the RFECV next time.</p>",
      "rawMarkdown": "The default feature importance ranking I got from LGBM was slightly different from the SHAP ranking. I somehow thought the SHAP ranking made more sense. \n@mithrillion Thanks. will check the RFECV next time.",
      "votes": null
    },
    {
      "id": "442096",
      "postDate": "12/19/2018 13:39:20",
      "content": "<p>So far i found eli5 importance to be the most reliable</p>",
      "rawMarkdown": "So far i found eli5 importance to be the most reliable",
      "votes": null
    },
    {
      "id": "442109",
      "postDate": "12/19/2018 13:53:12",
      "content": "<p>Thanks for sharing and big congratulations to you and your team. We missed the silver for one position and you missed gold. Probably you feel 100 times what we feel. :P </p>",
      "rawMarkdown": "Thanks for sharing and big congratulations to you and your team. We missed the silver for one position and you missed gold. Probably you feel 100 times what we feel. :P",
      "votes": null
    },
    {
      "id": "442116",
      "postDate": "12/19/2018 13:55:56",
      "content": "<p>Rereading, I wanted to use CAR but running times were way too long.  Mithrilion speedup is certainly something I would have used if I had access to it!</p>\n\n<p>What did you use for Gaussian process modeling?</p>",
      "rawMarkdown": "Rereading, I wanted to use CAR but running times were way too long.  Mithrilion speedup is certainly something I would have used if I had access to it!\n\nWhat did you use for Gaussian process modeling?",
      "votes": null
    },
    {
      "id": "442133",
      "postDate": "12/19/2018 14:15:23",
      "content": "<p>It's actually fairly simple. The template code we used was from the feets library. We found that the code repeated evaluates a likelihood function with nested loops, which are awfully inefficient in Python, so we used the numba.jit trick along with minor vectorisation to speed it up.\nThe original function looks like this:\n<code>\ndef _car_like(parameters, t, x, error_vars):\n    sigma, tau = parameters\n    t, x, error_vars = t.flatten(), x.flatten(), error_vars.flatten()\n    b = np.mean(x) / tau\n    num_datos = np.size(x)\n    Omega = [(tau * (sigma ** 2)) / 2.]\n    x_hat = [0.]\n    x_ast = [x[0] - b * tau]\n    loglik = 0.\n    for i in range(1, num_datos):\n        a_new = np.exp(-(t[i] - t[i - 1]) / tau)\n        x_ast.append(x[i] - b * tau)\n        x_hat.append(\n            a_new * x_hat[i - 1] +\n            (a_new * Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1])) *\n            (x_ast[i - 1] - x_hat[i - 1]))\n        Omega.append(\n            Omega[0] * (1 - (a_new ** 2)) + ((a_new ** 2)) * Omega[i - 1] *\n            (1 - (Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1]))))\n        loglik_inter = np.log(\n            ((2 * np.pi * (Omega[i] + error_vars[i])) ** -0.5) *\n            (np.exp(-0.5 * (((x_hat[i] - x_ast[i]) ** 2) /\n             (Omega[i] + error_vars[i]))) + EPSILON))\n        loglik = loglik + loglik_inter\n        if loglik &amp;lt;= CTE_NEG:\n            warnings.warn(\n                \"CAR log-likelihood to inf\", FeatureExtractionWarning)\n            return -np.infty\n    # the minus one is to perfor maximization using the minimize function\n    return -loglik\n</code>\nWe used numba and replaced python lists with numpy arrays:\n<code>\n@jit(nopython=True)\ndef _car_like(parameters, t, x, error_vars):\n    EPSILON = 1e-300\n    CTE_NEG = -np.infty\n    sigma, tau = parameters\n    # t, x, error_vars = t.flatten(), x.flatten(), error_vars.flatten()\n    m = np.mean(x)\n    N = x.shape[0]\n    Omega = np.zeros(N)\n    Omega[0] = (tau * (sigma ** 2)) / 2.\n    x_hat = np.zeros(N)\n    x_ast = x - m\n    A = np.exp(-(t[1:] - t[:-1]) / tau)\n    loglik = 0.\n    for i in range(1, N):\n        x_hat[i] = A[i - 1] * x_hat[i - 1] + (A[i - 1] * Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1])) * (\n                x_ast[i - 1] - x_hat[i - 1])\n        Omega[i] = Omega[0] * (1 - (A[i - 1] ** 2)) + (A[i - 1] ** 2) * Omega[i - 1] * (\n                1 - (Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1])))\n        loglik_inter = np.log(\n            ((2 * np.pi * (Omega[i] + error_vars[i])) ** -0.5) *\n            (np.exp(-0.5 * (((x_hat[i] - x_ast[i]) ** 2) /\n                            (Omega[i] + error_vars[i]))) + EPSILON))\n        loglik = loglik + loglik_inter\n        if loglik &amp;lt;= CTE_NEG:\n            break\n    return -loglik\n</code>\nWe tried to use the <code>parallel</code> mode in Numba but it leads to some nasty memory leak, so we instead chose to use Dask to automatically parallelise tasks. Speed increase by a factor of approx. 20. I tried writing the whole function in Cython but apparently my C knowledge was pretty rusty and it was less than ideal...</p>\n\n<p>As for GP, I used Pyro. It was pretty slow for us, but then now I realised that I did not have to do the fitting per band (as Kyle pointed out), and that I really should be using something better than gradient descent to optimise it. I wonder how you guys get more than 1 fit per second? I'm interested in the type of optimisation the fast algorithms use. Maybe my chosen tool is too general-purpose?</p>",
      "rawMarkdown": "It's actually fairly simple. The template code we used was from the feets library. We found that the code repeated evaluates a likelihood function with nested loops, which are awfully inefficient in Python, so we used the numba.jit trick along with minor vectorisation to speed it up.\nThe original function looks like this:\n```\ndef _car_like(parameters, t, x, error_vars):\n    sigma, tau = parameters\n    t, x, error_vars = t.flatten(), x.flatten(), error_vars.flatten()\n    b = np.mean(x) / tau\n    num_datos = np.size(x)\n    Omega = [(tau * (sigma ** 2)) / 2.]\n    x_hat = [0.]\n    x_ast = [x[0] - b * tau]\n    loglik = 0.\n    for i in range(1, num_datos):\n        a_new = np.exp(-(t[i] - t[i - 1]) / tau)\n        x_ast.append(x[i] - b * tau)\n        x_hat.append(\n            a_new * x_hat[i - 1] +\n            (a_new * Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1])) *\n            (x_ast[i - 1] - x_hat[i - 1]))\n        Omega.append(\n            Omega[0] * (1 - (a_new ** 2)) + ((a_new ** 2)) * Omega[i - 1] *\n            (1 - (Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1]))))\n        loglik_inter = np.log(\n            ((2 * np.pi * (Omega[i] + error_vars[i])) ** -0.5) *\n            (np.exp(-0.5 * (((x_hat[i] - x_ast[i]) ** 2) /\n             (Omega[i] + error_vars[i]))) + EPSILON))\n        loglik = loglik + loglik_inter\n        if loglik &lt;= CTE_NEG:\n            warnings.warn(\n                \"CAR log-likelihood to inf\", FeatureExtractionWarning)\n            return -np.infty\n    # the minus one is to perfor maximization using the minimize function\n    return -loglik\n```\nWe used numba and replaced python lists with numpy arrays:\n```\n@jit(nopython=True)\ndef _car_like(parameters, t, x, error_vars):\n    EPSILON = 1e-300\n    CTE_NEG = -np.infty\n    sigma, tau = parameters\n    # t, x, error_vars = t.flatten(), x.flatten(), error_vars.flatten()\n    m = np.mean(x)\n    N = x.shape[0]\n    Omega = np.zeros(N)\n    Omega[0] = (tau * (sigma ** 2)) / 2.\n    x_hat = np.zeros(N)\n    x_ast = x - m\n    A = np.exp(-(t[1:] - t[:-1]) / tau)\n    loglik = 0.\n    for i in range(1, N):\n        x_hat[i] = A[i - 1] * x_hat[i - 1] + (A[i - 1] * Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1])) * (\n                x_ast[i - 1] - x_hat[i - 1])\n        Omega[i] = Omega[0] * (1 - (A[i - 1] ** 2)) + (A[i - 1] ** 2) * Omega[i - 1] * (\n                1 - (Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1])))\n        loglik_inter = np.log(\n            ((2 * np.pi * (Omega[i] + error_vars[i])) ** -0.5) *\n            (np.exp(-0.5 * (((x_hat[i] - x_ast[i]) ** 2) /\n                            (Omega[i] + error_vars[i]))) + EPSILON))\n        loglik = loglik + loglik_inter\n        if loglik &lt;= CTE_NEG:\n            break\n    return -loglik\n```\nWe tried to use the `parallel` mode in Numba but it leads to some nasty memory leak, so we instead chose to use Dask to automatically parallelise tasks. Speed increase by a factor of approx. 20. I tried writing the whole function in Cython but apparently my C knowledge was pretty rusty and it was less than ideal...\n\nAs for GP, I used Pyro. It was pretty slow for us, but then now I realised that I did not have to do the fitting per band (as Kyle pointed out), and that I really should be using something better than gradient descent to optimise it. I wonder how you guys get more than 1 fit per second? I'm interested in the type of optimisation the fast algorithms use. Maybe my chosen tool is too general-purpose?",
      "votes": null
    },
    {
      "id": "442158",
      "postDate": "12/19/2018 14:54:26",
      "content": "<blockquote>\n  <p>I wonder how you guys get more than 1 fit per second?</p>\n</blockquote>\n\n<p>I used celerite as I explained quite a bit ;)  It is blazingly fast as it inverts the covariance matrix in O(n) where n is the length of the series.  Fitting all train takes a couple of minutes using scipy.optimize(). and full test takes 8 hours using 20 threads on a 2.3GHz Xeon machine.</p>",
      "rawMarkdown": "&gt; I wonder how you guys get more than 1 fit per second?\n\nI used celerite as I explained quite a bit ;)  It is blazingly fast as it inverts the covariance matrix in O(n) where n is the length of the series.  Fitting all train takes a couple of minutes using scipy.optimize(). and full test takes 8 hours using 20 threads on a 2.3GHz Xeon machine.",
      "votes": null
    },
    {
      "id": "442232",
      "postDate": "12/19/2018 16:40:04",
      "content": "<p>Great writeup, thanks for sharing!</p>",
      "rawMarkdown": "Great writeup, thanks for sharing!",
      "votes": null
    },
    {
      "id": "442668",
      "postDate": "12/20/2018 10:00:13",
      "content": "<p>Thank you for sharing solution!\nDoes eli5 importance you said mean permutation importance of eil5?\n<a href=\"https://eli5.readthedocs.io/en/latest/blackbox/permutation_importance.html\">https://eli5.readthedocs.io/en/latest/blackbox/permutation_importance.html</a></p>",
      "rawMarkdown": "Thank you for sharing solution!\nDoes eli5 importance you said mean permutation importance of eil5?\nhttps://eli5.readthedocs.io/en/latest/blackbox/permutation_importance.html",
      "votes": null
    },
    {
      "id": "442682",
      "postDate": "12/20/2018 10:24:11",
      "content": "<blockquote>\n  <p>Does eli5 importance you said mean permutation importance of eil5?</p>\n</blockquote>\n\n<p>Yes, that's it (PermutationImportance).</p>",
      "rawMarkdown": "&gt; Does eli5 importance you said mean permutation importance of eil5?\n\nYes, that's it (PermutationImportance).",
      "votes": null
    },
    {
      "id": "442724",
      "postDate": "12/20/2018 11:39:26",
      "content": "<p>Thanks!!! \nAlthough I don't use eli5, I use permutation importance, too.\nI think it is more useful than LGBM feature importance.</p>",
      "rawMarkdown": "Thanks!!! \nAlthough I don't use eli5, I use permutation importance, too.\nI think it is more useful than LGBM feature importance.",
      "votes": null
    },
    {
      "id": "445370",
      "postDate": "12/26/2018 09:57:47",
      "content": "<p>Good starting point! thanks for sharing</p>",
      "rawMarkdown": "Good starting point! thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 441561,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/18/2018 19:38:30",
      "content": "<p>Thanks for sharing and congrats on your result.  You share a lot fo what I have used as well, not sure why we fared better.  Probably blending with RNN and MLP was the difference.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 441576,
      "author_name": "yuval6967",
      "author_url": "",
      "post_date": "12/18/2018 19:55:07",
      "content": "<p>Thanks for sharing and congrats. \nThe term <strong>\"Black cat in a dark room\"</strong> fits class99 exactly! \nFrom our probing I can say for sure that some of class99 object were classified as class42 (or class52) with very high probability (close to 1) by all of our classifiers.   </p>",
      "votes": null,
      "replies": [
        {
          "id": 441730,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "12/19/2018 01:14:09",
          "content": "<p>This is so sad for us... I was the one with the strong belief that to tackle what was supposed to be 10-20% of the test data, we should not rely on a probed formula (kind of cheeky...) that does not make a lot of sense probabilisitically, so I figured we should utilise the one source of information we did have - class 99 is only in the test set. However, even after I constructed a subset of the test set with identical ddf/wdf ratio as the training set and mitigated hostgal_photoz differences as much as possible, it appears that there still exists residual differences unrelated to class 99 between the two sets, making utilising Positive-Unlabeled (PU) classification difficult. We were keen on discovering a class 99 approach that only relies on LB for validation, not model building, but we were not successful. It appears that even the very top performers in this competition did not manage to do it unfortunately. Who knows, science is hard...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441594,
      "author_name": "vignam",
      "author_url": "",
      "post_date": "12/18/2018 20:26:16",
      "content": "<blockquote>\n  <p>Feature selection: what are the best feature selection methods people use here? I've seen comment about boruta, but found it for random forest only, does it exist for lgbm and xgboost? What people use here?</p>\n</blockquote>\n\n<p>Something I really want to know. Thanks for asking. I hope the masters/experts will share some of their strategy. </p>\n\n<p>P.S. Congrats and thanks for sharing your method. 😀</p>",
      "votes": null,
      "replies": [
        {
          "id": 441871,
          "author_name": "takuok",
          "author_url": "",
          "post_date": "12/19/2018 07:24:17",
          "content": "<p>Congrats and thank you:)\nI think it is too difficult question <code>what feature selection is best</code>. \nMy feature selection method was simple.\n- Training all features and calculate importance.\n- Select threshold with CV and cut features. ex. top250 features\n- Then drop one by one and check CV. If CV improves, it drops. (This part's improvement is little.)</p>\n\n<p>I think this feature selection method is not best. But it takes little time to select features so can focus on other things like feature engineering or post processing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441883,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "12/19/2018 07:43:02",
          "content": "<p>This is the process that we automated with RFECV. It works most of the time, especially when removing the bottom-ranking features, but towards the end of the process we found that it was pretty inadequate at determining which features are absolutely safe to remove. Sometimes features are correlated so when you remove one in a correlated group, you improve the CV a little bit by reducing feature redundancy, but you still lose a little bit information. Sometimes the redundant features are on the top of the importance list that you cannot eliminate from the bottom up. In some cases, I have to manually decorrelate the features to make them work, but it does not appear to do the trick for all correlated groups.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441894,
          "author_name": "takuok",
          "author_url": "",
          "post_date": "12/19/2018 08:12:19",
          "content": "<p>LGBM is lobast for redundant features(ex.correlated features). It is true that removing some correlated features improve score, but it isn't much. I thought there are some features that was few importance but improved score well. So I added one by one on cuted features. But this improvement was also few.\nIt is the reason that I didn't focus on feature selection much. More important thing is to find good features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441947,
          "author_name": "vignam",
          "author_url": "",
          "post_date": "12/19/2018 09:21:04",
          "content": "<p>The default feature importance ranking I got from LGBM was slightly different from the SHAP ranking. I somehow thought the SHAP ranking made more sense. \n<a href=\"/mithrillion\">@mithrillion</a> Thanks. will check the RFECV next time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442096,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "12/19/2018 13:39:20",
          "content": "<p>So far i found eli5 importance to be the most reliable</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442668,
          "author_name": "go5kuramubon",
          "author_url": "",
          "post_date": "12/20/2018 10:00:13",
          "content": "<p>Thank you for sharing solution!\nDoes eli5 importance you said mean permutation importance of eil5?\n<a href=\"https://eli5.readthedocs.io/en/latest/blackbox/permutation_importance.html\">https://eli5.readthedocs.io/en/latest/blackbox/permutation_importance.html</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442682,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "12/20/2018 10:24:11",
          "content": "<blockquote>\n  <p>Does eli5 importance you said mean permutation importance of eil5?</p>\n</blockquote>\n\n<p>Yes, that's it (PermutationImportance).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442724,
          "author_name": "go5kuramubon",
          "author_url": "",
          "post_date": "12/20/2018 11:39:26",
          "content": "<p>Thanks!!! \nAlthough I don't use eli5, I use permutation importance, too.\nI think it is more useful than LGBM feature importance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441758,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "12/19/2018 02:49:45",
      "content": "<p>Congratulations and with you gold in the next competitions :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 442109,
      "author_name": "indranilbhattacharya",
      "author_url": "",
      "post_date": "12/19/2018 13:53:12",
      "content": "<p>Thanks for sharing and big congratulations to you and your team. We missed the silver for one position and you missed gold. Probably you feel 100 times what we feel. :P </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 442116,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/19/2018 13:55:56",
      "content": "<p>Rereading, I wanted to use CAR but running times were way too long.  Mithrilion speedup is certainly something I would have used if I had access to it!</p>\n\n<p>What did you use for Gaussian process modeling?</p>",
      "votes": null,
      "replies": [
        {
          "id": 442133,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "12/19/2018 14:15:23",
          "content": "<p>It's actually fairly simple. The template code we used was from the feets library. We found that the code repeated evaluates a likelihood function with nested loops, which are awfully inefficient in Python, so we used the numba.jit trick along with minor vectorisation to speed it up.\nThe original function looks like this:\n<code>\ndef _car_like(parameters, t, x, error_vars):\n    sigma, tau = parameters\n    t, x, error_vars = t.flatten(), x.flatten(), error_vars.flatten()\n    b = np.mean(x) / tau\n    num_datos = np.size(x)\n    Omega = [(tau * (sigma ** 2)) / 2.]\n    x_hat = [0.]\n    x_ast = [x[0] - b * tau]\n    loglik = 0.\n    for i in range(1, num_datos):\n        a_new = np.exp(-(t[i] - t[i - 1]) / tau)\n        x_ast.append(x[i] - b * tau)\n        x_hat.append(\n            a_new * x_hat[i - 1] +\n            (a_new * Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1])) *\n            (x_ast[i - 1] - x_hat[i - 1]))\n        Omega.append(\n            Omega[0] * (1 - (a_new ** 2)) + ((a_new ** 2)) * Omega[i - 1] *\n            (1 - (Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1]))))\n        loglik_inter = np.log(\n            ((2 * np.pi * (Omega[i] + error_vars[i])) ** -0.5) *\n            (np.exp(-0.5 * (((x_hat[i] - x_ast[i]) ** 2) /\n             (Omega[i] + error_vars[i]))) + EPSILON))\n        loglik = loglik + loglik_inter\n        if loglik &amp;lt;= CTE_NEG:\n            warnings.warn(\n                \"CAR log-likelihood to inf\", FeatureExtractionWarning)\n            return -np.infty\n    # the minus one is to perfor maximization using the minimize function\n    return -loglik\n</code>\nWe used numba and replaced python lists with numpy arrays:\n<code>\n@jit(nopython=True)\ndef _car_like(parameters, t, x, error_vars):\n    EPSILON = 1e-300\n    CTE_NEG = -np.infty\n    sigma, tau = parameters\n    # t, x, error_vars = t.flatten(), x.flatten(), error_vars.flatten()\n    m = np.mean(x)\n    N = x.shape[0]\n    Omega = np.zeros(N)\n    Omega[0] = (tau * (sigma ** 2)) / 2.\n    x_hat = np.zeros(N)\n    x_ast = x - m\n    A = np.exp(-(t[1:] - t[:-1]) / tau)\n    loglik = 0.\n    for i in range(1, N):\n        x_hat[i] = A[i - 1] * x_hat[i - 1] + (A[i - 1] * Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1])) * (\n                x_ast[i - 1] - x_hat[i - 1])\n        Omega[i] = Omega[0] * (1 - (A[i - 1] ** 2)) + (A[i - 1] ** 2) * Omega[i - 1] * (\n                1 - (Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1])))\n        loglik_inter = np.log(\n            ((2 * np.pi * (Omega[i] + error_vars[i])) ** -0.5) *\n            (np.exp(-0.5 * (((x_hat[i] - x_ast[i]) ** 2) /\n                            (Omega[i] + error_vars[i]))) + EPSILON))\n        loglik = loglik + loglik_inter\n        if loglik &amp;lt;= CTE_NEG:\n            break\n    return -loglik\n</code>\nWe tried to use the <code>parallel</code> mode in Numba but it leads to some nasty memory leak, so we instead chose to use Dask to automatically parallelise tasks. Speed increase by a factor of approx. 20. I tried writing the whole function in Cython but apparently my C knowledge was pretty rusty and it was less than ideal...</p>\n\n<p>As for GP, I used Pyro. It was pretty slow for us, but then now I realised that I did not have to do the fitting per band (as Kyle pointed out), and that I really should be using something better than gradient descent to optimise it. I wonder how you guys get more than 1 fit per second? I'm interested in the type of optimisation the fast algorithms use. Maybe my chosen tool is too general-purpose?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442158,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/19/2018 14:54:26",
          "content": "<blockquote>\n  <p>I wonder how you guys get more than 1 fit per second?</p>\n</blockquote>\n\n<p>I used celerite as I explained quite a bit ;)  It is blazingly fast as it inverts the covariance matrix in O(n) where n is the length of the series.  Fitting all train takes a couple of minutes using scipy.optimize(). and full test takes 8 hours using 20 threads on a 2.3GHz Xeon machine.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442232,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "12/19/2018 16:40:04",
      "content": "<p>Great writeup, thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 445370,
      "author_name": "sooperdooper",
      "author_url": "",
      "post_date": "12/26/2018 09:57:47",
      "content": "<p>Good starting point! thanks for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "441535": "Thank you for this challenge, lots of fun and learning, data are good and clean, answers from organizers are timely. Very well organized. \nHere is our write up.\n\n**What worked:** \nOur final solution was based on LGBM model inspired by Oliver kernel, feature design and selection, and augmentation.\n\n**Features:** \n\n1. Magnitudes. Both aggregated and per band, but without k-correction till the last day. As we were not quite sure how to interpret detected flag exactly, we decided to try both: for detected == 1 and for all data. Surprisingly, magnitude features for detected == 1 worked better on LB (not on CV though). I attribute it to the outliers removal with detected == 1 and also class 99. Indeed, to detect unknown -- you have to be 100% sure, so the detected points may be important for 99 discovery. Known objects are a different story: we know them and the confidence level of error can be lower. \n \n2. Parametric curve fittings. We used models of Bazin (https://arxiv.org/pdf/0904.1066.pdf) and Karpenka (https://arxiv.org/abs/1208.1264), Gaussian fits to define peak-width, and the decay slope fits of supernova light curves in the log scale. Fitting loss was one of the most important feature in all the parametric curves.\n\n3. Basic features from cesium package and statistics. Those based on ratios, std, and skew made it to the final. \n\n4. From the Bazin equation we derived the position of the max for the incomplete light curves and derived magnitudes from fits as well. They helped for some bands. We also tried a range of other parameters, m15 and m-10 parameters, peak width at different levels -- but they did not help. \n\n5. We checked all features from feets. CAR tau from feets was particularly helpful, but it takes ages to calculate. So, Mithrillion optimized a code to make it 10 times faster! If astronomers are interested, I think we can put it in a kernel.\n\n6. Autocorrelation lag 0, also estimated peak widths made it to the final. \n\n7. Finally, on the last 2 days I was pointed to the proper calculation of k-correction by Kyle (thank you). I used the calculator from here (http://kcor.sai.msu.ru/getthecode/) and the colors as suggested here (https://arxiv.org/pdf/1410.8139.pdf). It's only valid for relatively small red shifts, but the majority of data lie in this region. \nThe k-correction of magnitudes per band did not help, but along the way we calculated colors: both from detected magnitudes and Bazin fitted magnitudes --&gt; that's how we got from 0.836 to 0.815 on a single model \n\nFeature selection: we tried both RFE from sklearn and manual tuning while looking at eli5 importance and deleting highly correlated features. Manual approach and eli5 importance worked out much better than RFE.\n  \n\n**Battling over fitting:**\n\n1. To overcome over-fitting we first augmented train 10x times using flux variation within normal distribution of the corresponding flux_err, similar for photoz values (except for galactic classes). We also introduced small time shift to bring translation symmetry. It was meant for NN mostly, that we did not use at the end.\n\n2. We twisted parameters of classifier, especially deep tree and max bin gave a change.\n\n3. We realized that we have ~30% of ddf samples in train and only 1% in test. So we are likely to over fit to the ddf samples. First, we down-sampled ddf in the current augmentation and re-selected features checking eli5 importance for wide field drilling samples only (ddf = 0). Secondly, we changed augmentation, adding 30x train from ddf = 0 samples and leaving ddf samples only in the original train fold. That helped. Unfortunately, this idea came to us in the last 48 hours, so we did what we could manage within that time, and feature selection -- optimization still could have been better. \n\nWith 30x augmentation CV 0.5 LB 0.815, gap 0.315\n\nNot much ensembling, 50/50 blend of two LGBM models with slightly modified features and parameters was the final blend\n\n**What did not work:**\n\n1. Autoencoder. Although it has shown quite nice curves reconstruction, unfortunately, non of it's features made it to the final. We tried to include autoencoder loss as well -- did not help.\n\n2. More parametric fitting. Although we tried adding models with double-peak (https://arxiv.org/abs/1208.1264 ), and add tau per band as well, nothing really improved LB compared to just our very first Bazin parametric fit.  \n\n3. Gaussian Processes augmentation. We tried to use GP to create more non-ddf samples for augmentation, but got a bit of an issues with detected and flux error, so far this did not work out, but the idea is good I think. Mithrillion gave more comments under Kyle's solution.\n\n4. PU classification to identify class 99. We tried to detect negative labels (99 labels that are not in train) for 99 class prediction using a PU classifier (https://arxiv.org/pdf/1605.06955.pdf ). It did not work out...\n\n5. Class 99 probing. We tried to find class 99 in test using LB probing. We noticed that the percent of high probability 99 objects is in the far z region. Indeed, apart from themselves, people tend to know better what is closer than what is far. Also, our visible Universe keeps expanding as the light from far away objects keeps reaching the Earth, so we expect more unknown from far, I think. We tried to find 99 among photoz &gt; 2.5 and high probability of 99 class in Oliver and Scirpus methods. Guess what ? ... Bingo! It did not work out.\n\n**What we did right:**\n\n1. We teamed up! Teams have power, totally recommend, it's more fun and more learning and better results.\n\n2. We did not give up when we fell to place 19 just a three days before the end. The competition is not over till the last submit! It's in the last two-three days that we actually realized our mistakes, fixed what we could and managed to improve from 0.87 to 0.815 on a single model getting to place 12. Then it was just a bad luck on private...  \n\n3. We asked questions on forum. And people answered. Special thanks to Kyle, CPMP and organizers.\n\n4. We read, we learned, we tried, we discussed, we tried again and we learned more...\n\n**What we did wrong:**\n\n1. We did not utilize our submissions properly in the beginning of competition. When we had plenty of them -- that's when we should have tried more features on LB as well and more experimenting with parameters earlier. Last days every submission counts and you cannot just twist parameters while also checking on LB.\n\n2. We did not bring ddf percent close to the test data until the last 48 hours. We should have looked at data more and try in earlier. In general, I think 99% of answers are in the data, so if you feel stuck -- look at them again, look at them more.\n\n3. Strategy. We first did things that should work, like parametric fits from papers and magnitudes, and we left fun experiments to the end. At the end everything has speed up,  and we did not had time to properly utilize those experiments. Maybe we should have done it earlier when the time was not pressing. It's my third kaggle, so I still do not know what the best strategy is. More comments from Grandmasters here would be helpful.\n\n4. We did not plan submissions ahead. Trying to generate 5 submissions in evening to utilize all of them did not work out. I think it's nice to plan what we want to probe on LB and prepare ahead, so to have a full use of submissions. \n\n5. Trying to find a black cat in a dark room. We spent too much time trying to find \"the secret\", identify class 99 and build a classifier to find it. Instead, we could do more traditional things that bring incremental improvement, like better feature selection and parameters tuning. We left it to the last few days, but then its not enough time and submissions left.\n\n*The hardest thing of all is to find a black cat in a dark room. Especially if there is no cat. Confucius\n\n**What I still do not understand:**\n\nWe noticed a considerable variations in folds loss with LGBM, it is stratified, it is still big for wdf only training. Anyone knows, why is that?\n\nNN classifier did not want to improve anymore after around 0.94 even with augmentation. Why is that? We tried to twist parameters.\n\nFeature selection: what are the best feature selection methods people use here? I've seen comment about boruta, but found it for random forest only, does it exist for lgbm and xgboost? What people use here?\n\nIt looks from other solutions that we did things right, so what was the most important that we missed to cross 0.8 margin on a single model (we had 0.815)? And why there was such a shake on private? \n\n**Personal outcome:** It was a hard work and lot's of learning. The last few days were especially crazy for me. I was literally holding a baby with one hand and programming with the other, trying to concentrate and compete for gold against people who can use both hands, lol. No much sleep, no much food, no time for anything else. \n\nStill, it was fun and learning experience! Thanks to kaggle, organizers and most of all -- to my best-ever-teammates!!!\n\nPS. Dear kaggle people, please consider: \n- adding emotions (happy, sad, thinking, and a little penguin dancing). \n- consider adding a sign: \nKAGGLE IS ADDICTIVE ! ENTER AT YOUR OWN RISK !!!",
    "441561": "Thanks for sharing and congrats on your result.  You share a lot fo what I have used as well, not sure why we fared better.  Probably blending with RNN and MLP was the difference.",
    "441576": "Thanks for sharing and congrats. \nThe term **\"Black cat in a dark room\"** fits class99 exactly! \nFrom our probing I can say for sure that some of class99 object were classified as class42 (or class52) with very high probability (close to 1) by all of our classifiers.",
    "441594": "&gt;Feature selection: what are the best feature selection methods people use here? I've seen comment about boruta, but found it for random forest only, does it exist for lgbm and xgboost? What people use here?\n\nSomething I really want to know. Thanks for asking. I hope the masters/experts will share some of their strategy. \n\nP.S. Congrats and thanks for sharing your method. 😀",
    "441730": "This is so sad for us... I was the one with the strong belief that to tackle what was supposed to be 10-20% of the test data, we should not rely on a probed formula (kind of cheeky...) that does not make a lot of sense probabilisitically, so I figured we should utilise the one source of information we did have - class 99 is only in the test set. However, even after I constructed a subset of the test set with identical ddf/wdf ratio as the training set and mitigated hostgal_photoz differences as much as possible, it appears that there still exists residual differences unrelated to class 99 between the two sets, making utilising Positive-Unlabeled (PU) classification difficult. We were keen on discovering a class 99 approach that only relies on LB for validation, not model building, but we were not successful. It appears that even the very top performers in this competition did not manage to do it unfortunately. Who knows, science is hard...",
    "441758": "Congratulations and with you gold in the next competitions :)",
    "441871": "Congrats and thank you:)\nI think it is too difficult question `what feature selection is best`. \nMy feature selection method was simple.\n- Training all features and calculate importance.\n- Select threshold with CV and cut features. ex. top250 features\n- Then drop one by one and check CV. If CV improves, it drops. (This part's improvement is little.)\n\nI think this feature selection method is not best. But it takes little time to select features so can focus on other things like feature engineering or post processing.",
    "441883": "This is the process that we automated with RFECV. It works most of the time, especially when removing the bottom-ranking features, but towards the end of the process we found that it was pretty inadequate at determining which features are absolutely safe to remove. Sometimes features are correlated so when you remove one in a correlated group, you improve the CV a little bit by reducing feature redundancy, but you still lose a little bit information. Sometimes the redundant features are on the top of the importance list that you cannot eliminate from the bottom up. In some cases, I have to manually decorrelate the features to make them work, but it does not appear to do the trick for all correlated groups.",
    "441894": "LGBM is lobast for redundant features(ex.correlated features). It is true that removing some correlated features improve score, but it isn't much. I thought there are some features that was few importance but improved score well. So I added one by one on cuted features. But this improvement was also few.\nIt is the reason that I didn't focus on feature selection much. More important thing is to find good features.",
    "441947": "The default feature importance ranking I got from LGBM was slightly different from the SHAP ranking. I somehow thought the SHAP ranking made more sense. \n@mithrillion Thanks. will check the RFECV next time.",
    "442096": "So far i found eli5 importance to be the most reliable",
    "442109": "Thanks for sharing and big congratulations to you and your team. We missed the silver for one position and you missed gold. Probably you feel 100 times what we feel. :P",
    "442116": "Rereading, I wanted to use CAR but running times were way too long.  Mithrilion speedup is certainly something I would have used if I had access to it!\n\nWhat did you use for Gaussian process modeling?",
    "442133": "It's actually fairly simple. The template code we used was from the feets library. We found that the code repeated evaluates a likelihood function with nested loops, which are awfully inefficient in Python, so we used the numba.jit trick along with minor vectorisation to speed it up.\nThe original function looks like this:\n```\ndef _car_like(parameters, t, x, error_vars):\n    sigma, tau = parameters\n    t, x, error_vars = t.flatten(), x.flatten(), error_vars.flatten()\n    b = np.mean(x) / tau\n    num_datos = np.size(x)\n    Omega = [(tau * (sigma ** 2)) / 2.]\n    x_hat = [0.]\n    x_ast = [x[0] - b * tau]\n    loglik = 0.\n    for i in range(1, num_datos):\n        a_new = np.exp(-(t[i] - t[i - 1]) / tau)\n        x_ast.append(x[i] - b * tau)\n        x_hat.append(\n            a_new * x_hat[i - 1] +\n            (a_new * Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1])) *\n            (x_ast[i - 1] - x_hat[i - 1]))\n        Omega.append(\n            Omega[0] * (1 - (a_new ** 2)) + ((a_new ** 2)) * Omega[i - 1] *\n            (1 - (Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1]))))\n        loglik_inter = np.log(\n            ((2 * np.pi * (Omega[i] + error_vars[i])) ** -0.5) *\n            (np.exp(-0.5 * (((x_hat[i] - x_ast[i]) ** 2) /\n             (Omega[i] + error_vars[i]))) + EPSILON))\n        loglik = loglik + loglik_inter\n        if loglik &lt;= CTE_NEG:\n            warnings.warn(\n                \"CAR log-likelihood to inf\", FeatureExtractionWarning)\n            return -np.infty\n    # the minus one is to perfor maximization using the minimize function\n    return -loglik\n```\nWe used numba and replaced python lists with numpy arrays:\n```\n@jit(nopython=True)\ndef _car_like(parameters, t, x, error_vars):\n    EPSILON = 1e-300\n    CTE_NEG = -np.infty\n    sigma, tau = parameters\n    # t, x, error_vars = t.flatten(), x.flatten(), error_vars.flatten()\n    m = np.mean(x)\n    N = x.shape[0]\n    Omega = np.zeros(N)\n    Omega[0] = (tau * (sigma ** 2)) / 2.\n    x_hat = np.zeros(N)\n    x_ast = x - m\n    A = np.exp(-(t[1:] - t[:-1]) / tau)\n    loglik = 0.\n    for i in range(1, N):\n        x_hat[i] = A[i - 1] * x_hat[i - 1] + (A[i - 1] * Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1])) * (\n                x_ast[i - 1] - x_hat[i - 1])\n        Omega[i] = Omega[0] * (1 - (A[i - 1] ** 2)) + (A[i - 1] ** 2) * Omega[i - 1] * (\n                1 - (Omega[i - 1] / (Omega[i - 1] + error_vars[i - 1])))\n        loglik_inter = np.log(\n            ((2 * np.pi * (Omega[i] + error_vars[i])) ** -0.5) *\n            (np.exp(-0.5 * (((x_hat[i] - x_ast[i]) ** 2) /\n                            (Omega[i] + error_vars[i]))) + EPSILON))\n        loglik = loglik + loglik_inter\n        if loglik &lt;= CTE_NEG:\n            break\n    return -loglik\n```\nWe tried to use the `parallel` mode in Numba but it leads to some nasty memory leak, so we instead chose to use Dask to automatically parallelise tasks. Speed increase by a factor of approx. 20. I tried writing the whole function in Cython but apparently my C knowledge was pretty rusty and it was less than ideal...\n\nAs for GP, I used Pyro. It was pretty slow for us, but then now I realised that I did not have to do the fitting per band (as Kyle pointed out), and that I really should be using something better than gradient descent to optimise it. I wonder how you guys get more than 1 fit per second? I'm interested in the type of optimisation the fast algorithms use. Maybe my chosen tool is too general-purpose?",
    "442158": "&gt; I wonder how you guys get more than 1 fit per second?\n\nI used celerite as I explained quite a bit ;)  It is blazingly fast as it inverts the covariance matrix in O(n) where n is the length of the series.  Fitting all train takes a couple of minutes using scipy.optimize(). and full test takes 8 hours using 20 threads on a 2.3GHz Xeon machine.",
    "442232": "Great writeup, thanks for sharing!",
    "442668": "Thank you for sharing solution!\nDoes eli5 importance you said mean permutation importance of eil5?\nhttps://eli5.readthedocs.io/en/latest/blackbox/permutation_importance.html",
    "442682": "&gt; Does eli5 importance you said mean permutation importance of eil5?\n\nYes, that's it (PermutationImportance).",
    "442724": "Thanks!!! \nAlthough I don't use eli5, I use permutation importance, too.\nI think it is more useful than LGBM feature importance.",
    "445370": "Good starting point! thanks for sharing"
  },
  "source": "meta"
}