{
  "id": 75050,
  "title": "Solution #5 tidbits (revised with code)",
  "url": "/competitions/PLAsTiCC-2018/discussion/75050",
  "author_name": "",
  "post_date": "2018-12-18T06:18:03.018243Z",
  "votes": 75,
  "comment_count": 37,
  "views": 0,
  "content": "<p>First, I want to congratulate Kyle for his awesome victory, due to a data driven approach rather than sophisticated machine learning.  Congrats also to all who managed to get something out of this complex and very unusual data set.  I am still not sure about why there is such gap on the leaderboard, as it seems that crossing 0.8 in the private LB was hard.  Passing the 0.7 bar on the public LB is yet another story, and all those who did it deserve special kudos IMHO.  In particular,  Ahmet's solo performance is impressive.  I also want to thank two of my team mates, SomethingIsWrong and Kun Hao Yeh for their contributions, especially on NNs and on stacking.  The third team member hasn't been active once I joined the team.   I'd also like to thank those who shared useful info, Kyle, Olivier, Grzegorz, Giba, to name a few.  Please excuse me if I don't cite you.  Last, but not least, the organizers deserve special kudos for preparing such a challenging dataset.  I learned a lot during this competition, and have no regret about it, even if for some time we hoped we would win it.</p>\n\n<p>Here are some highlights of my part of our solution.  I almost only ran lightgbm given my team mates had decent NN models, RNN and MLP.  </p>\n\n<p>There were several difficulties in this competition compared to others:</p>\n\n<ol>\n<li>Unevenly sampled time series.</li>\n<li>Biased training dataset</li>\n<li>Small training dataset</li>\n<li>Open classification with a catch all class</li>\n</ol>\n\n<p><strong>Baseline</strong></p>\n\n<p>Before looking at these issue I set a baseline with feature engineering on the light curves. <br>\nI hand crafted all of them, as features from packages light gatspy, cesium, tsfresh where not as good, and were way slower to compute.  Features were computed with pandas, see  <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71398#latest-430428\">here</a> for how to do it efficiently.  For eahc object_id, features were computed passband per passband, but also on the overall flux.   Some feature groups that proved useful:\n - Std, skewness, kurtosis.  Goal is to differentiate curves with peaks (supernova) for the other ones.  It also captures a bit the shape of the peak\n - Ratios.  For instance take the max for each passband, then divide by the largest one.  This give a view of the spectra.\n - Features based on difference between successive measurements (delta): average of delta sign change from one measurement to the next one (very strong), ratio of delta std over flux std: this isolates rapidly varying sources from others.\n - Magnitude. First issue was to estimate the absolute flux given we only have a difference to an unknown background.  I used flux max minus flux min, given flux min corresponds to a positive absolute flux once you add the background.  Then I tried two ways: the exact way, using distmod (I gave a link to a <a href=\"https://en.wikipedia.org/wiki/Distance_modulus\">page containing the way to do it</a> in the forum).  But I found that using max minus min times the square of hostgal_photoz was best.  Not sure why. <br>\n - Mjd diff. This was shared by Grzegorz Sionkowski in the forum.  std of detected days was also useful.\n - Flux was normalized by dividing by the max value, for each object_id.  The normalization factor was kept as a feature, but it wasn't important, as magnitude was capturing the scale more appropriately.\n - Bazin and <a href=\"https://arxiv.org/pdf/1010.1005.pdf\">Newling</a> curve fitting.  I used fitted parameters as features, max values, rise and decay rates.  I used <code>scipy.optimize.curve_fit()</code> for fitting these curves.  These curves were only fit for extra galactic sources.  I extended the curves with a randomly generated measures before and after to capture cases where the peak was before or after the observation period.   </p>\n\n<p><strong>Objective function</strong></p>\n\n<p>Thanks to Giba probing we know the class weights in the competition evaluation function.  Compared to log loss, the evaluation function take the average class per class.  One  way to mimic this is to use sample weights.  For each sample, the weight is the class weight divided by the class frequency.  I found this right away which is why I reached top 10 scores within a couple of days after entering the competition.  There was absolutely no need to implement complex loss functions.  For NN, I saw some using frequencies computed batch per batch.  This is unstable IMHO.</p>\n\n<p><strong>Training time data augmentation</strong></p>\n\n<p>While a common practice in deep learning, training time augmentation is not widely used with other models.  Not sure why.  Here I generated variants of each training source using the flux err. I generated new flux by adding a Gaussian noise with flux err std.  Best I found was to add 5 randomized version of each object light curves. The only trick is to make sure that all derived light curves stay in the same fold than the original, otherwise we face a severe target leak issue.  This improved my LB by 0.003 if I remember well.</p>\n\n<p>This baseline led me to a 0.855 public LB obtained by blending a lgb model with a mlp trained on a subset of the features.   This is when I was invited to join my team mates, which I accepted easily given they were leading the competition LB at the time.</p>\n\n<p><strong>Stacking</strong></p>\n\n<p>An immediate benefit of teaming was stacking proposed by SomethingIsWrong.  Using out of fold prediction for train data, and test predictions from the Kun Hao Yeh RNN, I got a LB of 0.801.  This is the first confusion matrix I shared on the forum.  We used 2nd level stacking in a similar way.  This explains about half of the progress we made since teaming, yielding 0.775 LB score.</p>\n\n<p>Let's now revisit the remaining issues of the competitions</p>\n\n<p><strong>Unevenly sampled time series</strong></p>\n\n<p>The way to fix this is to interpolate the curves.  I tried Bazin, Newling, and Gaussian process.  My team mates tried autoencoders, but results were disappointing.  GP seemed the most promising, confirmed since by Kyle solution, but I started too late on it, finished the last day of the competition.  I used celerite with a Matern32 kernel.  Here are examples of fit using the three methods.  For Bazin and Newling I aligned the curves for each passband by refitting the curves using a fixed t0 obtained by averaging the t0 fitted on teach passband.  Here is an example of fit for object_id 4173 with the three techniques.</p>\n\n<p>celerite:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10892/celerite.png\" alt=\"celerite fit\"></p>\n\n<p>Bazin:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10894/bazin.png\" alt=\"bazin\"></p>\n\n<p>Newling:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10895/newling.png\" alt=\"newling\"></p>\n\n<p>I planned to use these to generate new, evenly samples time series and compute features on these.  Unfortunately, mastering celerite took me too much time and I could only make one submission the last day.  It was promising at 0.766, but a bit worse than my best sub.  Kyle solution makes me regret to not have started this before.  But I'm happy to have learned about this technique, hope I'll reuse in future projects or competitions.</p>\n\n<p><strong>Biased training dataset</strong></p>\n\n<p>Training data is biased compared to test data.  First, class 99 is missing.  Second, distribution of sources is different.  For instance ddf sources are way more frequent in train than in test.  Adjusting weights to cope with ddf frequency did not improve my LB, but it did not degrade it either.\nAnother bias comes from hostgal photoz. Here is the distribution of each class in train data by hostgal_photoz.  Blue background is the full train data.  Orange is each of the extra galactic classes.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10893/hostgal_photoz_all.png\" alt=\"photoz\"></p>\n\n<p>We can see that two classes have a very different distribution which does not make sense from a physics point of view.  It means that hostgal_photoz will isolate these two classes when it should not.  My fix was to remove hostgal photoz and only keep the binary version 0 or positive as it separates galactic sources from the rest.  I degraded CV but improved LB.   There is also a difference in distribution of source by hostgal photoz between train and test.  I wanted to address it by reweighting samples accordingly, but I only thought of it day before last, and we didn't have enough submissions to test it properly.</p>\n\n<p><strong>Small training dataset</strong></p>\n\n<p>I addressed it with TTA both for lgb (see above) and for NN, see above.  I also started using Gaussian process but could not finish in time really.  Issue is to augment data without introducing new bias. Using Gaussian noise with flux_err is easy, but interpolating curve along the time dimension is trickier.   I am not sure I would have done it properly anyway, and I am eager to look at Kyle's solution in detail to see how he did it. </p>\n\n<p><strong>Open classification</strong></p>\n\n<p>We were given an open classification problem, where train data does not contain all classes.  I did quite a lot of research and found relevant papers for deep learning approaches that I passed to my team mates (see Kun Hao Yeh write up for the links).  I also looked at some anomaly detection approaches but didn't found them useful.  In the end we used Olivier's approach with some scaling of class_99 probabilities.  We did not probe LB to find better ways, and maybe we should have.  For some reason, I am extremely reluctant to perform any LB probing. This time it may have cost us some ranks in the LB.</p>\n\n<p><strong>What did not work</strong></p>\n\n<p>Lots of things.  The foremost one is that cross validation was not reliable.  This is because train data is too small and biased compared to test..  I think that with more bias correction it can become reliable, but we did not invest enough time in it.  Then a number of techniques we tried unsuccessfully include:\n - Auto encoder, both latent vectors or generated curves\n - Anomaly detection\n - KNN\n - CNN\n - Adversarial validation\n - Gaussian Process (lack of time)\n - k_correction (undo redshift)\n - LombScargle</p>\n\n<p>Edit: I shared some of my code on <a href=\"https://github.com/jfpuget/Kaggle_PLAsTiCC\">github</a>.</p>",
  "messages": [
    {
      "id": "440975",
      "postDate": "12/18/2018 06:18:03",
      "content": "<p>First, I want to congratulate Kyle for his awesome victory, due to a data driven approach rather than sophisticated machine learning.  Congrats also to all who managed to get something out of this complex and very unusual data set.  I am still not sure about why there is such gap on the leaderboard, as it seems that crossing 0.8 in the private LB was hard.  Passing the 0.7 bar on the public LB is yet another story, and all those who did it deserve special kudos IMHO.  In particular,  Ahmet's solo performance is impressive.  I also want to thank two of my team mates, SomethingIsWrong and Kun Hao Yeh for their contributions, especially on NNs and on stacking.  The third team member hasn't been active once I joined the team.   I'd also like to thank those who shared useful info, Kyle, Olivier, Grzegorz, Giba, to name a few.  Please excuse me if I don't cite you.  Last, but not least, the organizers deserve special kudos for preparing such a challenging dataset.  I learned a lot during this competition, and have no regret about it, even if for some time we hoped we would win it.</p>\n\n<p>Here are some highlights of my part of our solution.  I almost only ran lightgbm given my team mates had decent NN models, RNN and MLP.  </p>\n\n<p>There were several difficulties in this competition compared to others:</p>\n\n<ol>\n<li>Unevenly sampled time series.</li>\n<li>Biased training dataset</li>\n<li>Small training dataset</li>\n<li>Open classification with a catch all class</li>\n</ol>\n\n<p><strong>Baseline</strong></p>\n\n<p>Before looking at these issue I set a baseline with feature engineering on the light curves. <br>\nI hand crafted all of them, as features from packages light gatspy, cesium, tsfresh where not as good, and were way slower to compute.  Features were computed with pandas, see  <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71398#latest-430428\">here</a> for how to do it efficiently.  For eahc object_id, features were computed passband per passband, but also on the overall flux.   Some feature groups that proved useful:\n - Std, skewness, kurtosis.  Goal is to differentiate curves with peaks (supernova) for the other ones.  It also captures a bit the shape of the peak\n - Ratios.  For instance take the max for each passband, then divide by the largest one.  This give a view of the spectra.\n - Features based on difference between successive measurements (delta): average of delta sign change from one measurement to the next one (very strong), ratio of delta std over flux std: this isolates rapidly varying sources from others.\n - Magnitude. First issue was to estimate the absolute flux given we only have a difference to an unknown background.  I used flux max minus flux min, given flux min corresponds to a positive absolute flux once you add the background.  Then I tried two ways: the exact way, using distmod (I gave a link to a <a href=\"https://en.wikipedia.org/wiki/Distance_modulus\">page containing the way to do it</a> in the forum).  But I found that using max minus min times the square of hostgal_photoz was best.  Not sure why. <br>\n - Mjd diff. This was shared by Grzegorz Sionkowski in the forum.  std of detected days was also useful.\n - Flux was normalized by dividing by the max value, for each object_id.  The normalization factor was kept as a feature, but it wasn't important, as magnitude was capturing the scale more appropriately.\n - Bazin and <a href=\"https://arxiv.org/pdf/1010.1005.pdf\">Newling</a> curve fitting.  I used fitted parameters as features, max values, rise and decay rates.  I used <code>scipy.optimize.curve_fit()</code> for fitting these curves.  These curves were only fit for extra galactic sources.  I extended the curves with a randomly generated measures before and after to capture cases where the peak was before or after the observation period.   </p>\n\n<p><strong>Objective function</strong></p>\n\n<p>Thanks to Giba probing we know the class weights in the competition evaluation function.  Compared to log loss, the evaluation function take the average class per class.  One  way to mimic this is to use sample weights.  For each sample, the weight is the class weight divided by the class frequency.  I found this right away which is why I reached top 10 scores within a couple of days after entering the competition.  There was absolutely no need to implement complex loss functions.  For NN, I saw some using frequencies computed batch per batch.  This is unstable IMHO.</p>\n\n<p><strong>Training time data augmentation</strong></p>\n\n<p>While a common practice in deep learning, training time augmentation is not widely used with other models.  Not sure why.  Here I generated variants of each training source using the flux err. I generated new flux by adding a Gaussian noise with flux err std.  Best I found was to add 5 randomized version of each object light curves. The only trick is to make sure that all derived light curves stay in the same fold than the original, otherwise we face a severe target leak issue.  This improved my LB by 0.003 if I remember well.</p>\n\n<p>This baseline led me to a 0.855 public LB obtained by blending a lgb model with a mlp trained on a subset of the features.   This is when I was invited to join my team mates, which I accepted easily given they were leading the competition LB at the time.</p>\n\n<p><strong>Stacking</strong></p>\n\n<p>An immediate benefit of teaming was stacking proposed by SomethingIsWrong.  Using out of fold prediction for train data, and test predictions from the Kun Hao Yeh RNN, I got a LB of 0.801.  This is the first confusion matrix I shared on the forum.  We used 2nd level stacking in a similar way.  This explains about half of the progress we made since teaming, yielding 0.775 LB score.</p>\n\n<p>Let's now revisit the remaining issues of the competitions</p>\n\n<p><strong>Unevenly sampled time series</strong></p>\n\n<p>The way to fix this is to interpolate the curves.  I tried Bazin, Newling, and Gaussian process.  My team mates tried autoencoders, but results were disappointing.  GP seemed the most promising, confirmed since by Kyle solution, but I started too late on it, finished the last day of the competition.  I used celerite with a Matern32 kernel.  Here are examples of fit using the three methods.  For Bazin and Newling I aligned the curves for each passband by refitting the curves using a fixed t0 obtained by averaging the t0 fitted on teach passband.  Here is an example of fit for object_id 4173 with the three techniques.</p>\n\n<p>celerite:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10892/celerite.png\" alt=\"celerite fit\"></p>\n\n<p>Bazin:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10894/bazin.png\" alt=\"bazin\"></p>\n\n<p>Newling:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10895/newling.png\" alt=\"newling\"></p>\n\n<p>I planned to use these to generate new, evenly samples time series and compute features on these.  Unfortunately, mastering celerite took me too much time and I could only make one submission the last day.  It was promising at 0.766, but a bit worse than my best sub.  Kyle solution makes me regret to not have started this before.  But I'm happy to have learned about this technique, hope I'll reuse in future projects or competitions.</p>\n\n<p><strong>Biased training dataset</strong></p>\n\n<p>Training data is biased compared to test data.  First, class 99 is missing.  Second, distribution of sources is different.  For instance ddf sources are way more frequent in train than in test.  Adjusting weights to cope with ddf frequency did not improve my LB, but it did not degrade it either.\nAnother bias comes from hostgal photoz. Here is the distribution of each class in train data by hostgal_photoz.  Blue background is the full train data.  Orange is each of the extra galactic classes.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10893/hostgal_photoz_all.png\" alt=\"photoz\"></p>\n\n<p>We can see that two classes have a very different distribution which does not make sense from a physics point of view.  It means that hostgal_photoz will isolate these two classes when it should not.  My fix was to remove hostgal photoz and only keep the binary version 0 or positive as it separates galactic sources from the rest.  I degraded CV but improved LB.   There is also a difference in distribution of source by hostgal photoz between train and test.  I wanted to address it by reweighting samples accordingly, but I only thought of it day before last, and we didn't have enough submissions to test it properly.</p>\n\n<p><strong>Small training dataset</strong></p>\n\n<p>I addressed it with TTA both for lgb (see above) and for NN, see above.  I also started using Gaussian process but could not finish in time really.  Issue is to augment data without introducing new bias. Using Gaussian noise with flux_err is easy, but interpolating curve along the time dimension is trickier.   I am not sure I would have done it properly anyway, and I am eager to look at Kyle's solution in detail to see how he did it. </p>\n\n<p><strong>Open classification</strong></p>\n\n<p>We were given an open classification problem, where train data does not contain all classes.  I did quite a lot of research and found relevant papers for deep learning approaches that I passed to my team mates (see Kun Hao Yeh write up for the links).  I also looked at some anomaly detection approaches but didn't found them useful.  In the end we used Olivier's approach with some scaling of class_99 probabilities.  We did not probe LB to find better ways, and maybe we should have.  For some reason, I am extremely reluctant to perform any LB probing. This time it may have cost us some ranks in the LB.</p>\n\n<p><strong>What did not work</strong></p>\n\n<p>Lots of things.  The foremost one is that cross validation was not reliable.  This is because train data is too small and biased compared to test..  I think that with more bias correction it can become reliable, but we did not invest enough time in it.  Then a number of techniques we tried unsuccessfully include:\n - Auto encoder, both latent vectors or generated curves\n - Anomaly detection\n - KNN\n - CNN\n - Adversarial validation\n - Gaussian Process (lack of time)\n - k_correction (undo redshift)\n - LombScargle</p>\n\n<p>Edit: I shared some of my code on <a href=\"https://github.com/jfpuget/Kaggle_PLAsTiCC\">github</a>.</p>",
      "rawMarkdown": "First, I want to congratulate Kyle for his awesome victory, due to a data driven approach rather than sophisticated machine learning.  Congrats also to all who managed to get something out of this complex and very unusual data set.  I am still not sure about why there is such gap on the leaderboard, as it seems that crossing 0.8 in the private LB was hard.  Passing the 0.7 bar on the public LB is yet another story, and all those who did it deserve special kudos IMHO.  In particular,  Ahmet's solo performance is impressive.  I also want to thank two of my team mates, SomethingIsWrong and Kun Hao Yeh for their contributions, especially on NNs and on stacking.  The third team member hasn't been active once I joined the team.   I'd also like to thank those who shared useful info, Kyle, Olivier, Grzegorz, Giba, to name a few.  Please excuse me if I don't cite you.  Last, but not least, the organizers deserve special kudos for preparing such a challenging dataset.  I learned a lot during this competition, and have no regret about it, even if for some time we hoped we would win it.\n\nHere are some highlights of my part of our solution.  I almost only ran lightgbm given my team mates had decent NN models, RNN and MLP.  \n\nThere were several difficulties in this competition compared to others:\n\n 1. Unevenly sampled time series.\n 2. Biased training dataset\n 3. Small training dataset\n 4. Open classification with a catch all class\n\n**Baseline**\n\nBefore looking at these issue I set a baseline with feature engineering on the light curves.  \nI hand crafted all of them, as features from packages light gatspy, cesium, tsfresh where not as good, and were way slower to compute.  Features were computed with pandas, see  [here][1] for how to do it efficiently.  For eahc object_id, features were computed passband per passband, but also on the overall flux.   Some feature groups that proved useful:\n - Std, skewness, kurtosis.  Goal is to differentiate curves with peaks (supernova) for the other ones.  It also captures a bit the shape of the peak\n - Ratios.  For instance take the max for each passband, then divide by the largest one.  This give a view of the spectra.\n - Features based on difference between successive measurements (delta): average of delta sign change from one measurement to the next one (very strong), ratio of delta std over flux std: this isolates rapidly varying sources from others.\n - Magnitude. First issue was to estimate the absolute flux given we only have a difference to an unknown background.  I used flux max minus flux min, given flux min corresponds to a positive absolute flux once you add the background.  Then I tried two ways: the exact way, using distmod (I gave a link to a [page containing the way to do it][2] in the forum).  But I found that using max minus min times the square of hostgal_photoz was best.  Not sure why.  \n - Mjd diff. This was shared by Grzegorz Sionkowski in the forum.  std of detected days was also useful.\n - Flux was normalized by dividing by the max value, for each object_id.  The normalization factor was kept as a feature, but it wasn't important, as magnitude was capturing the scale more appropriately.\n - Bazin and [Newling][3] curve fitting.  I used fitted parameters as features, max values, rise and decay rates.  I used `scipy.optimize.curve_fit()` for fitting these curves.  These curves were only fit for extra galactic sources.  I extended the curves with a randomly generated measures before and after to capture cases where the peak was before or after the observation period.   \n\n**Objective function**\n\nThanks to Giba probing we know the class weights in the competition evaluation function.  Compared to log loss, the evaluation function take the average class per class.  One  way to mimic this is to use sample weights.  For each sample, the weight is the class weight divided by the class frequency.  I found this right away which is why I reached top 10 scores within a couple of days after entering the competition.  There was absolutely no need to implement complex loss functions.  For NN, I saw some using frequencies computed batch per batch.  This is unstable IMHO.\n\n**Training time data augmentation**\n\nWhile a common practice in deep learning, training time augmentation is not widely used with other models.  Not sure why.  Here I generated variants of each training source using the flux err. I generated new flux by adding a Gaussian noise with flux err std.  Best I found was to add 5 randomized version of each object light curves. The only trick is to make sure that all derived light curves stay in the same fold than the original, otherwise we face a severe target leak issue.  This improved my LB by 0.003 if I remember well.\n\nThis baseline led me to a 0.855 public LB obtained by blending a lgb model with a mlp trained on a subset of the features.   This is when I was invited to join my team mates, which I accepted easily given they were leading the competition LB at the time.\n\n**Stacking**\n\nAn immediate benefit of teaming was stacking proposed by SomethingIsWrong.  Using out of fold prediction for train data, and test predictions from the Kun Hao Yeh RNN, I got a LB of 0.801.  This is the first confusion matrix I shared on the forum.  We used 2nd level stacking in a similar way.  This explains about half of the progress we made since teaming, yielding 0.775 LB score.\n\nLet's now revisit the remaining issues of the competitions\n\n**Unevenly sampled time series**\n\nThe way to fix this is to interpolate the curves.  I tried Bazin, Newling, and Gaussian process.  My team mates tried autoencoders, but results were disappointing.  GP seemed the most promising, confirmed since by Kyle solution, but I started too late on it, finished the last day of the competition.  I used celerite with a Matern32 kernel.  Here are examples of fit using the three methods.  For Bazin and Newling I aligned the curves for each passband by refitting the curves using a fixed t0 obtained by averaging the t0 fitted on teach passband.  Here is an example of fit for object_id 4173 with the three techniques.\n\ncelerite:\n\n![celerite fit][4]\n\nBazin:\n\n![bazin][5]\n\nNewling:\n\n![newling][6]\n\nI planned to use these to generate new, evenly samples time series and compute features on these.  Unfortunately, mastering celerite took me too much time and I could only make one submission the last day.  It was promising at 0.766, but a bit worse than my best sub.  Kyle solution makes me regret to not have started this before.  But I'm happy to have learned about this technique, hope I'll reuse in future projects or competitions.\n\n**Biased training dataset**\n\nTraining data is biased compared to test data.  First, class 99 is missing.  Second, distribution of sources is different.  For instance ddf sources are way more frequent in train than in test.  Adjusting weights to cope with ddf frequency did not improve my LB, but it did not degrade it either.\nAnother bias comes from hostgal photoz. Here is the distribution of each class in train data by hostgal_photoz.  Blue background is the full train data.  Orange is each of the extra galactic classes.\n\n![photoz][7]\n\nWe can see that two classes have a very different distribution which does not make sense from a physics point of view.  It means that hostgal_photoz will isolate these two classes when it should not.  My fix was to remove hostgal photoz and only keep the binary version 0 or positive as it separates galactic sources from the rest.  I degraded CV but improved LB.   There is also a difference in distribution of source by hostgal photoz between train and test.  I wanted to address it by reweighting samples accordingly, but I only thought of it day before last, and we didn't have enough submissions to test it properly.\n\n**Small training dataset**\n\nI addressed it with TTA both for lgb (see above) and for NN, see above.  I also started using Gaussian process but could not finish in time really.  Issue is to augment data without introducing new bias. Using Gaussian noise with flux_err is easy, but interpolating curve along the time dimension is trickier.   I am not sure I would have done it properly anyway, and I am eager to look at Kyle's solution in detail to see how he did it. \n\n**Open classification**\n\nWe were given an open classification problem, where train data does not contain all classes.  I did quite a lot of research and found relevant papers for deep learning approaches that I passed to my team mates (see Kun Hao Yeh write up for the links).  I also looked at some anomaly detection approaches but didn't found them useful.  In the end we used Olivier's approach with some scaling of class_99 probabilities.  We did not probe LB to find better ways, and maybe we should have.  For some reason, I am extremely reluctant to perform any LB probing. This time it may have cost us some ranks in the LB.\n\n**What did not work**\n\nLots of things.  The foremost one is that cross validation was not reliable.  This is because train data is too small and biased compared to test..  I think that with more bias correction it can become reliable, but we did not invest enough time in it.  Then a number of techniques we tried unsuccessfully include:\n - Auto encoder, both latent vectors or generated curves\n - Anomaly detection\n - KNN\n - CNN\n - Adversarial validation\n - Gaussian Process (lack of time)\n - k_correction (undo redshift)\n - LombScargle\n\nEdit: I shared some of my code on [github][8].\n\n\n  [1]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71398#latest-430428\n  [2]: https://en.wikipedia.org/wiki/Distance_modulus\n  [3]: https://arxiv.org/pdf/1010.1005.pdf\n  [4]: https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10892/celerite.png\n  [5]: https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10894/bazin.png\n  [6]: https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10895/newling.png\n  [7]: https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10893/hostgal_photoz_all.png\n  [8]: https://github.com/jfpuget/Kaggle_PLAsTiCC",
      "votes": null
    },
    {
      "id": "440981",
      "postDate": "12/18/2018 06:27:58",
      "content": "<p>Congrats. Thanks for sharing! </p>",
      "rawMarkdown": "Congrats. Thanks for sharing!",
      "votes": null
    },
    {
      "id": "441037",
      "postDate": "12/18/2018 07:39:47",
      "content": "<p>Thanks for your contributions too~ And thanks for detailed summary you shared here! Cheers!</p>",
      "rawMarkdown": "Thanks for your contributions too~ And thanks for detailed summary you shared here! Cheers!",
      "votes": null
    },
    {
      "id": "441086",
      "postDate": "12/18/2018 08:44:27",
      "content": "<p>Congrats and thanks for everything. Your comments were really helpful during the competition.</p>",
      "rawMarkdown": "Congrats and thanks for everything. Your comments were really helpful during the competition.",
      "votes": null
    },
    {
      "id": "441125",
      "postDate": "12/18/2018 09:36:43",
      "content": "<p>Congratulations and thanks for sharing! Also, thank you for your contributions in the discussions. You showed that you don't have to care to be nice to be the most valuable discussion contributor, just speak your mind, be present, insightful and respectful.</p>\n\n<p>One class of features that I am surprised I haven't seen in solutions, as it helped me a lot (around 0.04) is the\n<em>time width around maximum</em>:</p>\n\n<pre><code># extract mjd_diff from the following\ndf.groupby(['object_id'].apply(lambda x: x[x['flux'] &gt; x['flux'].max()/N]) # N=2,4,10\n</code></pre>",
      "rawMarkdown": "Congratulations and thanks for sharing! Also, thank you for your contributions in the discussions. You showed that you don't have to care to be nice to be the most valuable discussion contributor, just speak your mind, be present, insightful and respectful.\n\nOne class of features that I am surprised I haven't seen in solutions, as it helped me a lot (around 0.04) is the\n*time width around maximum*:\n\n    # extract mjd_diff from the following\n    df.groupby(['object_id'].apply(lambda x: x[x['flux'] &gt; x['flux'].max()/N]) # N=2,4,10",
      "votes": null
    },
    {
      "id": "441131",
      "postDate": "12/18/2018 09:46:24",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Congratulations.  Thanks for sharing your approach, very insightful.\nI got benefited by many of your comments during the competition.</p>",
      "rawMarkdown": "cpmpml Congratulations.  Thanks for sharing your approach, very insightful.\nI got benefited by many of your comments during the competition.",
      "votes": null
    },
    {
      "id": "441168",
      "postDate": "12/18/2018 11:09:02",
      "content": "<p>Merci, master!!</p>",
      "rawMarkdown": "Merci, master!!",
      "votes": null
    },
    {
      "id": "441188",
      "postDate": "12/18/2018 11:38:39",
      "content": "<p>Congrats CPMP, SomethingIsWrong and Kun Hao Yeh ! Very nice solution.</p>",
      "rawMarkdown": "Congrats CPMP, SomethingIsWrong and Kun Hao Yeh ! Very nice solution.",
      "votes": null
    },
    {
      "id": "441218",
      "postDate": "12/18/2018 12:29:36",
      "content": "<p>congrats <a href=\"/cpmpml\">@cpmpml</a>. i calculated the diff between passband flux and gave me a boost. i was always thinking about ratio but have always put it on least priority during experiments, too bad.</p>",
      "rawMarkdown": "congrats @cpmpml. i calculated the diff between passband flux and gave me a boost. i was always thinking about ratio but have always put it on least priority during experiments, too bad.",
      "votes": null
    },
    {
      "id": "441233",
      "postDate": "12/18/2018 12:49:57",
      "content": "<p>Thanks.  You are right, peak width is also very important. Bazin and Newling curve parameters capture the peak width quite precisely, which is why I didn't need another feature for it. They even capture the rate at which flux increase and decrease (half life).  </p>",
      "rawMarkdown": "Thanks.  You are right, peak width is also very important. Bazin and Newling curve parameters capture the peak width quite precisely, which is why I didn't need another feature for it. They even capture the rate at which flux increase and decrease (half life).",
      "votes": null
    },
    {
      "id": "441310",
      "postDate": "12/18/2018 14:33:33",
      "content": "<blockquote>\n  <p>you showed that you don't have to care to be nice to be the most valuable discussion contributor</p>\n</blockquote>\n\n<p>LOL about the 'nice' part.  I try not to be mean, but sometimes I fail dramatically on that...</p>",
      "rawMarkdown": "&gt; you showed that you don't have to care to be nice to be the most valuable discussion contributor\n\nLOL about the 'nice' part.  I try not to be mean, but sometimes I fail dramatically on that...",
      "votes": null
    },
    {
      "id": "441316",
      "postDate": "12/18/2018 14:38:47",
      "content": "<p>personally I I am a fan of your straightforward way, but I understand that others might get offended</p>\n\n<p>PS: the fact that you are ranked #1 discussion contributor shows that many people agree</p>",
      "rawMarkdown": "personally I I am a fan of your straightforward way, but I understand that others might get offended\n\nPS: the fact that you are ranked #1 discussion contributor shows that many people agree",
      "votes": null
    },
    {
      "id": "441325",
      "postDate": "12/18/2018 14:49:12",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a>. First of all congratulations. You have been a great teacher all throughout this competition.</p>\n\n<p>I have a question regarding your TTA.</p>\n\n<p>You wrote:</p>\n\n<blockquote>\n  <p>Here I generated variants of each training source using the flux err. I generated new flux by adding a Gaussian noise with flux err std. Best I found was to add 5 randomized version of each object light curves. </p>\n</blockquote>\n\n<p>How did you augment the meta data (i.e features not derived from the flux)? Did you also modify for instance the specz, photoz, distmod, etc. for the generated training sample.?</p>",
      "rawMarkdown": "cpmpml. First of all congratulations. You have been a great teacher all throughout this competition.\n\nI have a question regarding your TTA.\n\nYou wrote:\n&gt;Here I generated variants of each training source using the flux err. I generated new flux by adding a Gaussian noise with flux err std. Best I found was to add 5 randomized version of each object light curves. \n\nHow did you augment the meta data (i.e features not derived from the flux)? Did you also modify for instance the specz, photoz, distmod, etc. for the generated training sample.?",
      "votes": null
    },
    {
      "id": "441346",
      "postDate": "12/18/2018 15:13:29",
      "content": "<p>I computed features using the modified lightcurves the same way I used the original lightcurves.  I only modified flux.  I should have modified hostgal_photoz too indeed.  </p>",
      "rawMarkdown": "I computed features using the modified lightcurves the same way I used the original lightcurves.  I only modified flux.  I should have modified hostgal_photoz too indeed.",
      "votes": null
    },
    {
      "id": "441354",
      "postDate": "12/18/2018 15:22:24",
      "content": "<p>The reason why I asked was, I did try doing a similar TTA w.r.t flux and also modified the other features, but it didn't help me. Finally, I ended up using two different training sets (one set with the original flux and one with the modified flux) but with the same <code>hostgal_photoz</code> and trained them separately and did an ensemble. </p>",
      "rawMarkdown": "The reason why I asked was, I did try doing a similar TTA w.r.t flux and also modified the other features, but it didn't help me. Finally, I ended up using two different training sets (one set with the original flux and one with the modified flux) but with the same ```hostgal_photoz``` and trained them separately and did an ensemble.",
      "votes": null
    },
    {
      "id": "441356",
      "postDate": "12/18/2018 15:24:07",
      "content": "<p>Maybe you did not take care of the folds as I did.</p>",
      "rawMarkdown": "Maybe you did not take care of the folds as I did.",
      "votes": null
    },
    {
      "id": "441366",
      "postDate": "12/18/2018 15:34:10",
      "content": "<p>Yes, indeed. But, thank you very much for that insight. I'm quite new to the field and learning stuffs by picking titbits that you and others very generously sprinkle here. </p>",
      "rawMarkdown": "Yes, indeed. But, thank you very much for that insight. I'm quite new to the field and learning stuffs by picking titbits that you and others very generously sprinkle here.",
      "votes": null
    },
    {
      "id": "441498",
      "postDate": "12/18/2018 18:01:35",
      "content": "<p>That celerite chart belongs in a gallery of beautiful graphs.</p>",
      "rawMarkdown": "That celerite chart belongs in a gallery of beautiful graphs.",
      "votes": null
    },
    {
      "id": "441522",
      "postDate": "12/18/2018 18:28:39",
      "content": "<p>Thanks.  I'll share code to generate them.</p>",
      "rawMarkdown": "Thanks.  I'll share code to generate them.",
      "votes": null
    },
    {
      "id": "441525",
      "postDate": "12/18/2018 18:33:09",
      "content": "<p>Some more gp fit.  The last one shows a limit: curves should not go that far in the negative ranges.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/441525/10898/celerite2.png\" alt=\"gp\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/441525/10899/celerite3.png\" alt=\"gp\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/441525/10900/celerite4.png\" alt=\"gp\"></p>",
      "rawMarkdown": "Some more gp fit.  The last one shows a limit: curves should not go that far in the negative ranges.\n\n![gp][1]\n\n![gp][2]\n\n![gp][3]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/441525/10898/celerite2.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/441525/10899/celerite3.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/441525/10900/celerite4.png",
      "votes": null
    },
    {
      "id": "441666",
      "postDate": "12/18/2018 22:56:22",
      "content": "<p>Congratulations and thanks for sharing your solution - there is so much to learn from this competition! Also thanks for all your thoughts, help and comments in the discussion group - very much appreciated.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing your solution - there is so much to learn from this competition! Also thanks for all your thoughts, help and comments in the discussion group - very much appreciated.",
      "votes": null
    },
    {
      "id": "442041",
      "postDate": "12/19/2018 12:04:17",
      "content": "<p>Thank you so much for your detailed feedback, including what we shan't explore for future analysis. It looks like that handcrafting the features was the thing to do in this challenge, which is good to know.</p>\n\n<p>I'm a little surprised by your statement about LombScargle, as I saw it being able to provide interesting separation between intragalactic classes, especially the periodic type against the others (green = 92), but also regarding  microlensing (blue=6)  or transit type (red= 16) especially in the regime where transit occurs in a variable system. (see figure : x=frequency of max , y=power of max). Or does it mean that other features were already enough to test periodicity ?</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/442041/10903/Unknown.png\" alt=\"Power vs frequency of galactic types\"></p>",
      "rawMarkdown": "Thank you so much for your detailed feedback, including what we shan't explore for future analysis. It looks like that handcrafting the features was the thing to do in this challenge, which is good to know.\n\nI'm a little surprised by your statement about LombScargle, as I saw it being able to provide interesting separation between intragalactic classes, especially the periodic type against the others (green = 92), but also regarding  microlensing (blue=6)  or transit type (red= 16) especially in the regime where transit occurs in a variable system. (see figure : x=frequency of max , y=power of max). Or does it mean that other features were already enough to test periodicity ?\n\n![Power vs frequency of galactic types][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/442041/10903/Unknown.png",
      "votes": null
    },
    {
      "id": "442107",
      "postDate": "12/19/2018 13:49:55",
      "content": "<p>Thanks.</p>\n\n<p>I am sure LombScargle is useful for separating galactic sources, but features that are way faster to compute seem to be even better  When I added LombScargle periods I didn't see any improvement in the LB score.  Features that were great are the std of flux delta divided by std of flux, as well as the average of flux delta sign changes.  They isolate short period sources very effectively.  What i did not use thought was the quality of the LombScargle fit.  Maybe this is more useful than the period itself.</p>",
      "rawMarkdown": "Thanks.\n\nI am sure LombScargle is useful for separating galactic sources, but features that are way faster to compute seem to be even better  When I added LombScargle periods I didn't see any improvement in the LB score.  Features that were great are the std of flux delta divided by std of flux, as well as the average of flux delta sign changes.  They isolate short period sources very effectively.  What i did not use thought was the quality of the LombScargle fit.  Maybe this is more useful than the period itself.",
      "votes": null
    },
    {
      "id": "442181",
      "postDate": "12/19/2018 15:27:18",
      "content": "<p>Thank you for your answer ! Notice that the quality of the LombScargle can be poor for some periodic objects as they are clearly not sinusoids. What I would try is to fit an average shape (we had at least 2 variants of periodic objects in the same class...). But this will be slow to compute...</p>",
      "rawMarkdown": "Thank you for your answer ! Notice that the quality of the LombScargle can be poor for some periodic objects as they are clearly not sinusoids. What I would try is to fit an average shape (we had at least 2 variants of periodic objects in the same class...). But this will be slow to compute...",
      "votes": null
    },
    {
      "id": "442213",
      "postDate": "12/19/2018 16:10:49",
      "content": "<p>I wish we had an astronomer like you on the team!  It took us way too long to identify what might work  I wanted to do template fitting, we did it for Bazin and Newling, but we failed with Salt2 just because we didn't know what parameters to use.  And now I see your templates which may be killer features...</p>\n\n<p>Obviously you will probably continue working on these, and I wish you good luck.  Hope this competition solutions will help process LSST data more effectively!</p>",
      "rawMarkdown": "I wish we had an astronomer like you on the team!  It took us way too long to identify what might work  I wanted to do template fitting, we did it for Bazin and Newling, but we failed with Salt2 just because we didn't know what parameters to use.  And now I see your templates which may be killer features...\n\nObviously you will probably continue working on these, and I wish you good luck.  Hope this competition solutions will help process LSST data more effectively!",
      "votes": null
    },
    {
      "id": "442325",
      "postDate": "12/19/2018 19:53:43",
      "content": "<p>Congratulations and thanks for your solution. I suppose all competitiors including us learned a lot from you in this competition! I really enjoyed competing with you again, and would like to compete with you in my next competition too!</p>",
      "rawMarkdown": "Congratulations and thanks for your solution. I suppose all competitiors including us learned a lot from you in this competition! I really enjoyed competing with you again, and would like to compete with you in my next competition too!",
      "votes": null
    },
    {
      "id": "442353",
      "postDate": "12/19/2018 21:13:55",
      "content": "<p>Thanks for sharing your solution.</p>\n\n<p>I would also like to give you a big thanks for your discussion contributions, from which I learned a great deal. My experience in this competition would have been very different without them.</p>",
      "rawMarkdown": "Thanks for sharing your solution.\n\nI would also like to give you a big thanks for your discussion contributions, from which I learned a great deal. My experience in this competition would have been very different without them.",
      "votes": null
    },
    {
      "id": "442379",
      "postDate": "12/19/2018 22:00:54",
      "content": "<p>... and I wish I were a fast coder as yourself! This is why interdisciplinary work is so important.</p>\n\n<p>You're right: for us this is only a beginning as we have to understand all the whereabouts of what made your solutions great. We warmly thank all people who posted working kernels of top ranked solutions, there is a lot for us to be learnt there.</p>\n\n<p>I hope you'll consider submitting your work to attend one of the workshops, it would be a great opportunity for sharing ideas.</p>",
      "rawMarkdown": "... and I wish I were a fast coder as yourself! This is why interdisciplinary work is so important.\n\nYou're right: for us this is only a beginning as we have to understand all the whereabouts of what made your solutions great. We warmly thank all people who posted working kernels of top ranked solutions, there is a lot for us to be learnt there.\n\nI hope you'll consider submitting your work to attend one of the workshops, it would be a great opportunity for sharing ideas.",
      "votes": null
    },
    {
      "id": "442837",
      "postDate": "12/20/2018 15:35:16",
      "content": "<p>CPMP, could you upload your best sub here? \nI think it's possible to get to 0.5x with your best sub :)\n<a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75179\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75179</a></p>",
      "rawMarkdown": "CPMP, could you upload your best sub here? \nI think it's possible to get to 0.5x with your best sub :)\nhttps://www.kaggle.com/c/PLAsTiCC-2018/discussion/75179",
      "votes": null
    },
    {
      "id": "442845",
      "postDate": "12/20/2018 15:44:36",
      "content": "<p>Will do, but really busy at work today...  Maybe tonight, tomorrow for sure.</p>",
      "rawMarkdown": "Will do, but really busy at work today...  Maybe tonight, tomorrow for sure.",
      "votes": null
    },
    {
      "id": "442856",
      "postDate": "12/20/2018 15:59:50",
      "content": "<p>Hi, I'd be happy to team with you!  You beat me twice in a row, I can't stand a third time :)</p>\n\n<p>And thanks for the kind words.</p>",
      "rawMarkdown": "Hi, I'd be happy to team with you!  You beat me twice in a row, I can't stand a third time :)\n\nAnd thanks for the kind words.",
      "votes": null
    },
    {
      "id": "442857",
      "postDate": "12/20/2018 16:00:44",
      "content": "<p>Thanks, your performance was great too.  Glad you found some of my posts useful.</p>",
      "rawMarkdown": "Thanks, your performance was great too.  Glad you found some of my posts useful.",
      "votes": null
    },
    {
      "id": "442858",
      "postDate": "12/20/2018 16:01:35",
      "content": "<p>Uploading, hope it will succeed...  </p>",
      "rawMarkdown": "Uploading, hope it will succeed...",
      "votes": null
    },
    {
      "id": "444474",
      "postDate": "12/24/2018 04:53:50",
      "content": "<p>Hi CPMP, will you share the codes for your solution? Maybe some more on writing efficient code like \"Parallelism\".</p>",
      "rawMarkdown": "Hi CPMP, will you share the codes for your solution? Maybe some more on writing efficient code like \"Parallelism\".",
      "votes": null
    },
    {
      "id": "444610",
      "postDate": "12/24/2018 11:17:25",
      "content": "<p>I will share some of it. Working on it.</p>",
      "rawMarkdown": "I will share some of it. Working on it.",
      "votes": null
    },
    {
      "id": "444621",
      "postDate": "12/24/2018 12:11:13",
      "content": "<p>My Christmas gift to you all: ;)  Here is Some of the code I used for our solution: <a href=\"https://github.com/jfpuget/Kaggle_PLAsTiCC\">https://github.com/jfpuget/Kaggle_PLAsTiCC</a></p>\n\n<p>I made it self contained. Difference with our best model is that I removed the part that includes out of fold features from models produced by my team mates (RNN and MLP).  Other than that it is the code of a lgb model that scored 0.752 on the public LB (0.766 private).</p>",
      "rawMarkdown": "My Christmas gift to you all: ;)  Here is Some of the code I used for our solution: https://github.com/jfpuget/Kaggle_PLAsTiCC\n\nI made it self contained. Difference with our best model is that I removed the part that includes out of fold features from models produced by my team mates (RNN and MLP).  Other than that it is the code of a lgb model that scored 0.752 on the public LB (0.766 private).",
      "votes": null
    },
    {
      "id": "444647",
      "postDate": "12/24/2018 13:41:30",
      "content": "<p>Thank you CPMP! Merry Christmas and happy new year!</p>",
      "rawMarkdown": "Thank you CPMP! Merry Christmas and happy new year!",
      "votes": null
    },
    {
      "id": "447982",
      "postDate": "12/31/2018 01:40:03",
      "content": "<p>Happy New Year! Thank you for sharing!</p>",
      "rawMarkdown": "Happy New Year! Thank you for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 440981,
      "author_name": "infzero",
      "author_url": "",
      "post_date": "12/18/2018 06:27:58",
      "content": "<p>Congrats. Thanks for sharing! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 441037,
      "author_name": "khyeh0719",
      "author_url": "",
      "post_date": "12/18/2018 07:39:47",
      "content": "<p>Thanks for your contributions too~ And thanks for detailed summary you shared here! Cheers!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 441086,
      "author_name": "fatall",
      "author_url": "",
      "post_date": "12/18/2018 08:44:27",
      "content": "<p>Congrats and thanks for everything. Your comments were really helpful during the competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 441125,
      "author_name": "iprapas",
      "author_url": "",
      "post_date": "12/18/2018 09:36:43",
      "content": "<p>Congratulations and thanks for sharing! Also, thank you for your contributions in the discussions. You showed that you don't have to care to be nice to be the most valuable discussion contributor, just speak your mind, be present, insightful and respectful.</p>\n\n<p>One class of features that I am surprised I haven't seen in solutions, as it helped me a lot (around 0.04) is the\n<em>time width around maximum</em>:</p>\n\n<pre><code># extract mjd_diff from the following\ndf.groupby(['object_id'].apply(lambda x: x[x['flux'] &gt; x['flux'].max()/N]) # N=2,4,10\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 441233,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/18/2018 12:49:57",
          "content": "<p>Thanks.  You are right, peak width is also very important. Bazin and Newling curve parameters capture the peak width quite precisely, which is why I didn't need another feature for it. They even capture the rate at which flux increase and decrease (half life).  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441310,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/18/2018 14:33:33",
          "content": "<blockquote>\n  <p>you showed that you don't have to care to be nice to be the most valuable discussion contributor</p>\n</blockquote>\n\n<p>LOL about the 'nice' part.  I try not to be mean, but sometimes I fail dramatically on that...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441316,
          "author_name": "iprapas",
          "author_url": "",
          "post_date": "12/18/2018 14:38:47",
          "content": "<p>personally I I am a fan of your straightforward way, but I understand that others might get offended</p>\n\n<p>PS: the fact that you are ranked #1 discussion contributor shows that many people agree</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441131,
      "author_name": "subrahmanyamv",
      "author_url": "",
      "post_date": "12/18/2018 09:46:24",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Congratulations.  Thanks for sharing your approach, very insightful.\nI got benefited by many of your comments during the competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 441168,
      "author_name": "joxemi",
      "author_url": "",
      "post_date": "12/18/2018 11:09:02",
      "content": "<p>Merci, master!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 441188,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "12/18/2018 11:38:39",
      "content": "<p>Congrats CPMP, SomethingIsWrong and Kun Hao Yeh ! Very nice solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 441218,
      "author_name": "niclasdoce",
      "author_url": "",
      "post_date": "12/18/2018 12:29:36",
      "content": "<p>congrats <a href=\"/cpmpml\">@cpmpml</a>. i calculated the diff between passband flux and gave me a boost. i was always thinking about ratio but have always put it on least priority during experiments, too bad.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 441325,
      "author_name": "vignam",
      "author_url": "",
      "post_date": "12/18/2018 14:49:12",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a>. First of all congratulations. You have been a great teacher all throughout this competition.</p>\n\n<p>I have a question regarding your TTA.</p>\n\n<p>You wrote:</p>\n\n<blockquote>\n  <p>Here I generated variants of each training source using the flux err. I generated new flux by adding a Gaussian noise with flux err std. Best I found was to add 5 randomized version of each object light curves. </p>\n</blockquote>\n\n<p>How did you augment the meta data (i.e features not derived from the flux)? Did you also modify for instance the specz, photoz, distmod, etc. for the generated training sample.?</p>",
      "votes": null,
      "replies": [
        {
          "id": 441346,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/18/2018 15:13:29",
          "content": "<p>I computed features using the modified lightcurves the same way I used the original lightcurves.  I only modified flux.  I should have modified hostgal_photoz too indeed.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441354,
          "author_name": "vignam",
          "author_url": "",
          "post_date": "12/18/2018 15:22:24",
          "content": "<p>The reason why I asked was, I did try doing a similar TTA w.r.t flux and also modified the other features, but it didn't help me. Finally, I ended up using two different training sets (one set with the original flux and one with the modified flux) but with the same <code>hostgal_photoz</code> and trained them separately and did an ensemble. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441356,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/18/2018 15:24:07",
          "content": "<p>Maybe you did not take care of the folds as I did.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441366,
          "author_name": "vignam",
          "author_url": "",
          "post_date": "12/18/2018 15:34:10",
          "content": "<p>Yes, indeed. But, thank you very much for that insight. I'm quite new to the field and learning stuffs by picking titbits that you and others very generously sprinkle here. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441498,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "12/18/2018 18:01:35",
      "content": "<p>That celerite chart belongs in a gallery of beautiful graphs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 441522,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/18/2018 18:28:39",
          "content": "<p>Thanks.  I'll share code to generate them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441525,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/18/2018 18:33:09",
      "content": "<p>Some more gp fit.  The last one shows a limit: curves should not go that far in the negative ranges.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/441525/10898/celerite2.png\" alt=\"gp\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/441525/10899/celerite3.png\" alt=\"gp\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/441525/10900/celerite4.png\" alt=\"gp\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 441666,
      "author_name": "andypenrose",
      "author_url": "",
      "post_date": "12/18/2018 22:56:22",
      "content": "<p>Congratulations and thanks for sharing your solution - there is so much to learn from this competition! Also thanks for all your thoughts, help and comments in the discussion group - very much appreciated.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 442041,
      "author_name": "manugangler",
      "author_url": "",
      "post_date": "12/19/2018 12:04:17",
      "content": "<p>Thank you so much for your detailed feedback, including what we shan't explore for future analysis. It looks like that handcrafting the features was the thing to do in this challenge, which is good to know.</p>\n\n<p>I'm a little surprised by your statement about LombScargle, as I saw it being able to provide interesting separation between intragalactic classes, especially the periodic type against the others (green = 92), but also regarding  microlensing (blue=6)  or transit type (red= 16) especially in the regime where transit occurs in a variable system. (see figure : x=frequency of max , y=power of max). Or does it mean that other features were already enough to test periodicity ?</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/442041/10903/Unknown.png\" alt=\"Power vs frequency of galactic types\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 442107,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/19/2018 13:49:55",
          "content": "<p>Thanks.</p>\n\n<p>I am sure LombScargle is useful for separating galactic sources, but features that are way faster to compute seem to be even better  When I added LombScargle periods I didn't see any improvement in the LB score.  Features that were great are the std of flux delta divided by std of flux, as well as the average of flux delta sign changes.  They isolate short period sources very effectively.  What i did not use thought was the quality of the LombScargle fit.  Maybe this is more useful than the period itself.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442181,
          "author_name": "manugangler",
          "author_url": "",
          "post_date": "12/19/2018 15:27:18",
          "content": "<p>Thank you for your answer ! Notice that the quality of the LombScargle can be poor for some periodic objects as they are clearly not sinusoids. What I would try is to fit an average shape (we had at least 2 variants of periodic objects in the same class...). But this will be slow to compute...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442213,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/19/2018 16:10:49",
          "content": "<p>I wish we had an astronomer like you on the team!  It took us way too long to identify what might work  I wanted to do template fitting, we did it for Bazin and Newling, but we failed with Salt2 just because we didn't know what parameters to use.  And now I see your templates which may be killer features...</p>\n\n<p>Obviously you will probably continue working on these, and I wish you good luck.  Hope this competition solutions will help process LSST data more effectively!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442379,
          "author_name": "manugangler",
          "author_url": "",
          "post_date": "12/19/2018 22:00:54",
          "content": "<p>... and I wish I were a fast coder as yourself! This is why interdisciplinary work is so important.</p>\n\n<p>You're right: for us this is only a beginning as we have to understand all the whereabouts of what made your solutions great. We warmly thank all people who posted working kernels of top ranked solutions, there is a lot for us to be learnt there.</p>\n\n<p>I hope you'll consider submitting your work to attend one of the workshops, it would be a great opportunity for sharing ideas.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442325,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "12/19/2018 19:53:43",
      "content": "<p>Congratulations and thanks for your solution. I suppose all competitiors including us learned a lot from you in this competition! I really enjoyed competing with you again, and would like to compete with you in my next competition too!</p>",
      "votes": null,
      "replies": [
        {
          "id": 442856,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/20/2018 15:59:50",
          "content": "<p>Hi, I'd be happy to team with you!  You beat me twice in a row, I can't stand a third time :)</p>\n\n<p>And thanks for the kind words.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442353,
      "author_name": "agarreta",
      "author_url": "",
      "post_date": "12/19/2018 21:13:55",
      "content": "<p>Thanks for sharing your solution.</p>\n\n<p>I would also like to give you a big thanks for your discussion contributions, from which I learned a great deal. My experience in this competition would have been very different without them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 442857,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/20/2018 16:00:44",
          "content": "<p>Thanks, your performance was great too.  Glad you found some of my posts useful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442837,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "12/20/2018 15:35:16",
      "content": "<p>CPMP, could you upload your best sub here? \nI think it's possible to get to 0.5x with your best sub :)\n<a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75179\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75179</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 442845,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/20/2018 15:44:36",
          "content": "<p>Will do, but really busy at work today...  Maybe tonight, tomorrow for sure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442858,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/20/2018 16:01:35",
          "content": "<p>Uploading, hope it will succeed...  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 444474,
      "author_name": "cczaixian",
      "author_url": "",
      "post_date": "12/24/2018 04:53:50",
      "content": "<p>Hi CPMP, will you share the codes for your solution? Maybe some more on writing efficient code like \"Parallelism\".</p>",
      "votes": null,
      "replies": [
        {
          "id": 444610,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/24/2018 11:17:25",
          "content": "<p>I will share some of it. Working on it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 444621,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/24/2018 12:11:13",
      "content": "<p>My Christmas gift to you all: ;)  Here is Some of the code I used for our solution: <a href=\"https://github.com/jfpuget/Kaggle_PLAsTiCC\">https://github.com/jfpuget/Kaggle_PLAsTiCC</a></p>\n\n<p>I made it self contained. Difference with our best model is that I removed the part that includes out of fold features from models produced by my team mates (RNN and MLP).  Other than that it is the code of a lgb model that scored 0.752 on the public LB (0.766 private).</p>",
      "votes": null,
      "replies": [
        {
          "id": 444647,
          "author_name": "cczaixian",
          "author_url": "",
          "post_date": "12/24/2018 13:41:30",
          "content": "<p>Thank you CPMP! Merry Christmas and happy new year!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447982,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "12/31/2018 01:40:03",
          "content": "<p>Happy New Year! Thank you for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "440975": "First, I want to congratulate Kyle for his awesome victory, due to a data driven approach rather than sophisticated machine learning.  Congrats also to all who managed to get something out of this complex and very unusual data set.  I am still not sure about why there is such gap on the leaderboard, as it seems that crossing 0.8 in the private LB was hard.  Passing the 0.7 bar on the public LB is yet another story, and all those who did it deserve special kudos IMHO.  In particular,  Ahmet's solo performance is impressive.  I also want to thank two of my team mates, SomethingIsWrong and Kun Hao Yeh for their contributions, especially on NNs and on stacking.  The third team member hasn't been active once I joined the team.   I'd also like to thank those who shared useful info, Kyle, Olivier, Grzegorz, Giba, to name a few.  Please excuse me if I don't cite you.  Last, but not least, the organizers deserve special kudos for preparing such a challenging dataset.  I learned a lot during this competition, and have no regret about it, even if for some time we hoped we would win it.\n\nHere are some highlights of my part of our solution.  I almost only ran lightgbm given my team mates had decent NN models, RNN and MLP.  \n\nThere were several difficulties in this competition compared to others:\n\n 1. Unevenly sampled time series.\n 2. Biased training dataset\n 3. Small training dataset\n 4. Open classification with a catch all class\n\n**Baseline**\n\nBefore looking at these issue I set a baseline with feature engineering on the light curves.  \nI hand crafted all of them, as features from packages light gatspy, cesium, tsfresh where not as good, and were way slower to compute.  Features were computed with pandas, see  [here][1] for how to do it efficiently.  For eahc object_id, features were computed passband per passband, but also on the overall flux.   Some feature groups that proved useful:\n - Std, skewness, kurtosis.  Goal is to differentiate curves with peaks (supernova) for the other ones.  It also captures a bit the shape of the peak\n - Ratios.  For instance take the max for each passband, then divide by the largest one.  This give a view of the spectra.\n - Features based on difference between successive measurements (delta): average of delta sign change from one measurement to the next one (very strong), ratio of delta std over flux std: this isolates rapidly varying sources from others.\n - Magnitude. First issue was to estimate the absolute flux given we only have a difference to an unknown background.  I used flux max minus flux min, given flux min corresponds to a positive absolute flux once you add the background.  Then I tried two ways: the exact way, using distmod (I gave a link to a [page containing the way to do it][2] in the forum).  But I found that using max minus min times the square of hostgal_photoz was best.  Not sure why.  \n - Mjd diff. This was shared by Grzegorz Sionkowski in the forum.  std of detected days was also useful.\n - Flux was normalized by dividing by the max value, for each object_id.  The normalization factor was kept as a feature, but it wasn't important, as magnitude was capturing the scale more appropriately.\n - Bazin and [Newling][3] curve fitting.  I used fitted parameters as features, max values, rise and decay rates.  I used `scipy.optimize.curve_fit()` for fitting these curves.  These curves were only fit for extra galactic sources.  I extended the curves with a randomly generated measures before and after to capture cases where the peak was before or after the observation period.   \n\n**Objective function**\n\nThanks to Giba probing we know the class weights in the competition evaluation function.  Compared to log loss, the evaluation function take the average class per class.  One  way to mimic this is to use sample weights.  For each sample, the weight is the class weight divided by the class frequency.  I found this right away which is why I reached top 10 scores within a couple of days after entering the competition.  There was absolutely no need to implement complex loss functions.  For NN, I saw some using frequencies computed batch per batch.  This is unstable IMHO.\n\n**Training time data augmentation**\n\nWhile a common practice in deep learning, training time augmentation is not widely used with other models.  Not sure why.  Here I generated variants of each training source using the flux err. I generated new flux by adding a Gaussian noise with flux err std.  Best I found was to add 5 randomized version of each object light curves. The only trick is to make sure that all derived light curves stay in the same fold than the original, otherwise we face a severe target leak issue.  This improved my LB by 0.003 if I remember well.\n\nThis baseline led me to a 0.855 public LB obtained by blending a lgb model with a mlp trained on a subset of the features.   This is when I was invited to join my team mates, which I accepted easily given they were leading the competition LB at the time.\n\n**Stacking**\n\nAn immediate benefit of teaming was stacking proposed by SomethingIsWrong.  Using out of fold prediction for train data, and test predictions from the Kun Hao Yeh RNN, I got a LB of 0.801.  This is the first confusion matrix I shared on the forum.  We used 2nd level stacking in a similar way.  This explains about half of the progress we made since teaming, yielding 0.775 LB score.\n\nLet's now revisit the remaining issues of the competitions\n\n**Unevenly sampled time series**\n\nThe way to fix this is to interpolate the curves.  I tried Bazin, Newling, and Gaussian process.  My team mates tried autoencoders, but results were disappointing.  GP seemed the most promising, confirmed since by Kyle solution, but I started too late on it, finished the last day of the competition.  I used celerite with a Matern32 kernel.  Here are examples of fit using the three methods.  For Bazin and Newling I aligned the curves for each passband by refitting the curves using a fixed t0 obtained by averaging the t0 fitted on teach passband.  Here is an example of fit for object_id 4173 with the three techniques.\n\ncelerite:\n\n![celerite fit][4]\n\nBazin:\n\n![bazin][5]\n\nNewling:\n\n![newling][6]\n\nI planned to use these to generate new, evenly samples time series and compute features on these.  Unfortunately, mastering celerite took me too much time and I could only make one submission the last day.  It was promising at 0.766, but a bit worse than my best sub.  Kyle solution makes me regret to not have started this before.  But I'm happy to have learned about this technique, hope I'll reuse in future projects or competitions.\n\n**Biased training dataset**\n\nTraining data is biased compared to test data.  First, class 99 is missing.  Second, distribution of sources is different.  For instance ddf sources are way more frequent in train than in test.  Adjusting weights to cope with ddf frequency did not improve my LB, but it did not degrade it either.\nAnother bias comes from hostgal photoz. Here is the distribution of each class in train data by hostgal_photoz.  Blue background is the full train data.  Orange is each of the extra galactic classes.\n\n![photoz][7]\n\nWe can see that two classes have a very different distribution which does not make sense from a physics point of view.  It means that hostgal_photoz will isolate these two classes when it should not.  My fix was to remove hostgal photoz and only keep the binary version 0 or positive as it separates galactic sources from the rest.  I degraded CV but improved LB.   There is also a difference in distribution of source by hostgal photoz between train and test.  I wanted to address it by reweighting samples accordingly, but I only thought of it day before last, and we didn't have enough submissions to test it properly.\n\n**Small training dataset**\n\nI addressed it with TTA both for lgb (see above) and for NN, see above.  I also started using Gaussian process but could not finish in time really.  Issue is to augment data without introducing new bias. Using Gaussian noise with flux_err is easy, but interpolating curve along the time dimension is trickier.   I am not sure I would have done it properly anyway, and I am eager to look at Kyle's solution in detail to see how he did it. \n\n**Open classification**\n\nWe were given an open classification problem, where train data does not contain all classes.  I did quite a lot of research and found relevant papers for deep learning approaches that I passed to my team mates (see Kun Hao Yeh write up for the links).  I also looked at some anomaly detection approaches but didn't found them useful.  In the end we used Olivier's approach with some scaling of class_99 probabilities.  We did not probe LB to find better ways, and maybe we should have.  For some reason, I am extremely reluctant to perform any LB probing. This time it may have cost us some ranks in the LB.\n\n**What did not work**\n\nLots of things.  The foremost one is that cross validation was not reliable.  This is because train data is too small and biased compared to test..  I think that with more bias correction it can become reliable, but we did not invest enough time in it.  Then a number of techniques we tried unsuccessfully include:\n - Auto encoder, both latent vectors or generated curves\n - Anomaly detection\n - KNN\n - CNN\n - Adversarial validation\n - Gaussian Process (lack of time)\n - k_correction (undo redshift)\n - LombScargle\n\nEdit: I shared some of my code on [github][8].\n\n\n  [1]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71398#latest-430428\n  [2]: https://en.wikipedia.org/wiki/Distance_modulus\n  [3]: https://arxiv.org/pdf/1010.1005.pdf\n  [4]: https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10892/celerite.png\n  [5]: https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10894/bazin.png\n  [6]: https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10895/newling.png\n  [7]: https://storage.googleapis.com/kaggle-forum-message-attachments/440975/10893/hostgal_photoz_all.png\n  [8]: https://github.com/jfpuget/Kaggle_PLAsTiCC",
    "440981": "Congrats. Thanks for sharing!",
    "441037": "Thanks for your contributions too~ And thanks for detailed summary you shared here! Cheers!",
    "441086": "Congrats and thanks for everything. Your comments were really helpful during the competition.",
    "441125": "Congratulations and thanks for sharing! Also, thank you for your contributions in the discussions. You showed that you don't have to care to be nice to be the most valuable discussion contributor, just speak your mind, be present, insightful and respectful.\n\nOne class of features that I am surprised I haven't seen in solutions, as it helped me a lot (around 0.04) is the\n*time width around maximum*:\n\n    # extract mjd_diff from the following\n    df.groupby(['object_id'].apply(lambda x: x[x['flux'] &gt; x['flux'].max()/N]) # N=2,4,10",
    "441131": "cpmpml Congratulations.  Thanks for sharing your approach, very insightful.\nI got benefited by many of your comments during the competition.",
    "441168": "Merci, master!!",
    "441188": "Congrats CPMP, SomethingIsWrong and Kun Hao Yeh ! Very nice solution.",
    "441218": "congrats @cpmpml. i calculated the diff between passband flux and gave me a boost. i was always thinking about ratio but have always put it on least priority during experiments, too bad.",
    "441233": "Thanks.  You are right, peak width is also very important. Bazin and Newling curve parameters capture the peak width quite precisely, which is why I didn't need another feature for it. They even capture the rate at which flux increase and decrease (half life).",
    "441310": "&gt; you showed that you don't have to care to be nice to be the most valuable discussion contributor\n\nLOL about the 'nice' part.  I try not to be mean, but sometimes I fail dramatically on that...",
    "441316": "personally I I am a fan of your straightforward way, but I understand that others might get offended\n\nPS: the fact that you are ranked #1 discussion contributor shows that many people agree",
    "441325": "cpmpml. First of all congratulations. You have been a great teacher all throughout this competition.\n\nI have a question regarding your TTA.\n\nYou wrote:\n&gt;Here I generated variants of each training source using the flux err. I generated new flux by adding a Gaussian noise with flux err std. Best I found was to add 5 randomized version of each object light curves. \n\nHow did you augment the meta data (i.e features not derived from the flux)? Did you also modify for instance the specz, photoz, distmod, etc. for the generated training sample.?",
    "441346": "I computed features using the modified lightcurves the same way I used the original lightcurves.  I only modified flux.  I should have modified hostgal_photoz too indeed.",
    "441354": "The reason why I asked was, I did try doing a similar TTA w.r.t flux and also modified the other features, but it didn't help me. Finally, I ended up using two different training sets (one set with the original flux and one with the modified flux) but with the same ```hostgal_photoz``` and trained them separately and did an ensemble.",
    "441356": "Maybe you did not take care of the folds as I did.",
    "441366": "Yes, indeed. But, thank you very much for that insight. I'm quite new to the field and learning stuffs by picking titbits that you and others very generously sprinkle here.",
    "441498": "That celerite chart belongs in a gallery of beautiful graphs.",
    "441522": "Thanks.  I'll share code to generate them.",
    "441525": "Some more gp fit.  The last one shows a limit: curves should not go that far in the negative ranges.\n\n![gp][1]\n\n![gp][2]\n\n![gp][3]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/441525/10898/celerite2.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/441525/10899/celerite3.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/441525/10900/celerite4.png",
    "441666": "Congratulations and thanks for sharing your solution - there is so much to learn from this competition! Also thanks for all your thoughts, help and comments in the discussion group - very much appreciated.",
    "442041": "Thank you so much for your detailed feedback, including what we shan't explore for future analysis. It looks like that handcrafting the features was the thing to do in this challenge, which is good to know.\n\nI'm a little surprised by your statement about LombScargle, as I saw it being able to provide interesting separation between intragalactic classes, especially the periodic type against the others (green = 92), but also regarding  microlensing (blue=6)  or transit type (red= 16) especially in the regime where transit occurs in a variable system. (see figure : x=frequency of max , y=power of max). Or does it mean that other features were already enough to test periodicity ?\n\n![Power vs frequency of galactic types][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/442041/10903/Unknown.png",
    "442107": "Thanks.\n\nI am sure LombScargle is useful for separating galactic sources, but features that are way faster to compute seem to be even better  When I added LombScargle periods I didn't see any improvement in the LB score.  Features that were great are the std of flux delta divided by std of flux, as well as the average of flux delta sign changes.  They isolate short period sources very effectively.  What i did not use thought was the quality of the LombScargle fit.  Maybe this is more useful than the period itself.",
    "442181": "Thank you for your answer ! Notice that the quality of the LombScargle can be poor for some periodic objects as they are clearly not sinusoids. What I would try is to fit an average shape (we had at least 2 variants of periodic objects in the same class...). But this will be slow to compute...",
    "442213": "I wish we had an astronomer like you on the team!  It took us way too long to identify what might work  I wanted to do template fitting, we did it for Bazin and Newling, but we failed with Salt2 just because we didn't know what parameters to use.  And now I see your templates which may be killer features...\n\nObviously you will probably continue working on these, and I wish you good luck.  Hope this competition solutions will help process LSST data more effectively!",
    "442325": "Congratulations and thanks for your solution. I suppose all competitiors including us learned a lot from you in this competition! I really enjoyed competing with you again, and would like to compete with you in my next competition too!",
    "442353": "Thanks for sharing your solution.\n\nI would also like to give you a big thanks for your discussion contributions, from which I learned a great deal. My experience in this competition would have been very different without them.",
    "442379": "... and I wish I were a fast coder as yourself! This is why interdisciplinary work is so important.\n\nYou're right: for us this is only a beginning as we have to understand all the whereabouts of what made your solutions great. We warmly thank all people who posted working kernels of top ranked solutions, there is a lot for us to be learnt there.\n\nI hope you'll consider submitting your work to attend one of the workshops, it would be a great opportunity for sharing ideas.",
    "442837": "CPMP, could you upload your best sub here? \nI think it's possible to get to 0.5x with your best sub :)\nhttps://www.kaggle.com/c/PLAsTiCC-2018/discussion/75179",
    "442845": "Will do, but really busy at work today...  Maybe tonight, tomorrow for sure.",
    "442856": "Hi, I'd be happy to team with you!  You beat me twice in a row, I can't stand a third time :)\n\nAnd thanks for the kind words.",
    "442857": "Thanks, your performance was great too.  Glad you found some of my posts useful.",
    "442858": "Uploading, hope it will succeed...",
    "444474": "Hi CPMP, will you share the codes for your solution? Maybe some more on writing efficient code like \"Parallelism\".",
    "444610": "I will share some of it. Working on it.",
    "444621": "My Christmas gift to you all: ;)  Here is Some of the code I used for our solution: https://github.com/jfpuget/Kaggle_PLAsTiCC\n\nI made it self contained. Difference with our best model is that I removed the part that includes out of fold features from models produced by my team mates (RNN and MLP).  Other than that it is the code of a lgb model that scored 0.752 on the public LB (0.766 private).",
    "444647": "Thank you CPMP! Merry Christmas and happy new year!",
    "447982": "Happy New Year! Thank you for sharing!"
  },
  "source": "meta"
}