{
  "id": 75222,
  "title": "3rd Place Part III - Pesudo Astronomer's Feature Engineering",
  "url": "/competitions/PLAsTiCC-2018/discussion/75222",
  "author_name": "",
  "post_date": "2018-12-19T15:47:29.925558800Z",
  "votes": 27,
  "comment_count": 17,
  "views": 0,
  "content": "<p>First of all, thanks to the organizers and all participants in this competition! And thanks a lot to my great teammates, <a href=\"/mamasinkgs\">@mamasinkgs</a>, and <a href=\"/yuval6967\">@yuval6967</a>. I learn a lot from these 2 guys. And I'm very happy to get my first gold medal :)</p>\n\n<p>In this part, I want to share my findings and feature engineering. Please refer to the following discussions to see the overall description of our solution.</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75116\">3rd Place Part I - CNN</a></li>\n<li><a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75131\">3rd Place Part II - CatBoost, mamas feature, and class 99</a></li>\n</ul>\n\n<h2>What I learned from astronomer's work</h2>\n\n<p>At the beginning of this competition, I tried to install domain knowledge and tried to be \"pseudo astronomer\". In addition to the data note provided by the organizers, I carefully read the following resources.</p>\n\n<ul>\n<li><a href=\"https://www.lsst.org/scientists/scibook\">LSST science book</a></li>\n<li><a href=\"https://arxiv.org/abs/1008.1024\">Result from SNPCC challenge</a></li>\n<li><a href=\"https://arxiv.org/abs/1603.00882\">Photometric Supernova Classification With Machine Learning</a> (<a href=\"https://kicp-workshops.uchicago.edu/SNClassification_2016/depot/talk-lochner-michelle.pdf\">slide</a>)</li>\n</ul>\n\n<p>After reading these resources and exploring data a bit, I thought that this competition mainly consists of these 2 challenges:</p>\n\n<ul>\n<li>How to distinguish between various types of supernova classes?</li>\n<li>How to detect class99?</li>\n</ul>\n\n<p>I also tried to understand what each class really is. By 1) basic light curve characteristics 2) class frequency compared with expected LSST observation rate 3) redshift distribution, here is my assumption (order by confidence):</p>\n\n<ul>\n<li>class90: SN Ia</li>\n<li>class42: SN II</li>\n<li>class95: Superluminous Supernova </li>\n<li>class52/62: SN Ib/c</li>\n<li>class88: Strong Gravitational Lens (or AGN?)</li>\n<li>class67: TDE</li>\n</ul>\n\n<p>I hope the competition hosts will reveal the answer!\n<em>I don't think this assumption helped us directly</em>, but this gave us a very good interpretation of the result of our class99 probing (class99 is something similar to core-collapse supernova).</p>\n\n<h2>Feature Engineering</h2>\n\n<h3>template fitting by sncosmo package - the golden feature for us</h3>\n\n<p>sncosmo provides nice lc fitting API in python (<a href=\"https://sncosmo.readthedocs.io/en/v1.6.x/examples/plot_lc_fit.html#sphx-glr-examples-plot-lc-fit-py\">link</a>). I used a fitting parameter and it's chi-square value as a feature. I'm a bit surprised that no one except me uses this library because using template fitting is a major solution in the past competition.</p>\n\n<p>It took a bit long time to calculate (1~2 lines/sec), so I split light curves by object_id%30 and run 30 preemptible instances to finish calculating within 1 day. I repeated this process about 15 times to make various types of models to cover all variant of supernovae (salt-2, salt-2-extended, nugent-sn1bc, nugent-sn2, snana, sako,...), and use 8 of them in my final model. It was an exhausting process... really...</p>\n\n<p>But these bag of template features gave me a significant boost (~0.05 in intermediate, and 0.1+ in my final model). It helped Yuval's CNN as well (+0.04). 7 out of 10 my most important features (by LGBM gain) are these features. All template features were used only in an extragalactic model.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/442200/10905/object_id_161521.png\" alt=\"example of template fitting result\"></p>\n\n<p>Here is an example of a result of template fitting (SALT-2, object_id = 161521). y-band is ignored by its wavelength and estimated redshift. One may think that GP fitting gives better fitting, but it's robust to noise (in SALT-2, only 5 parameters used to generate this all 6-band curve). Probably combining our features with GP features of Kyle or CPMP's will give us another significant boost.</p>\n\n<em>UPDATED</em>\n\n<p>I've attached my template features. you can download, unzip and use:</p>\n\n<p><code>\ndf = pd.read_feather('sncosmo_template_features.f')\n</code></p>\n\n<em>UPDATED2</em>\n\n<p>Kernel is available to see how did we use sncosmo:\n<a href=\"https://www.kaggle.com/nyanpn/salt-2-feature-part-of-3rd-place-solution\">https://www.kaggle.com/nyanpn/salt-2-feature-part-of-3rd-place-solution</a></p>\n\n<h3>hostgal-specz model</h3>\n\n<p>Exactly same as 2nd and 4th place solution. This gave me a small boost (~0.004).</p>\n\n<h3>luminosity</h3>\n\n<p>As shared in the discussion, luminosity is an important intrinsic property of variable stars. I used the following fomla:</p>\n\n<p><code>\nluminosity = (max(flux) - min(flux)) * distance ** 2\n</code></p>\n\n<p>Distance (in MPc) is converted from estimated specz above. Astropy's luminosity_distance (<a href=\"http://docs.astropy.org/en/stable/cosmology/index.html?highlight=luminosity#using-astropy-cosmology\">link</a>) function gives distance from redshift.</p>\n\n<p>I added these luminosity and luminosity difference between channels.</p>\n\n<h3>time difference</h3>\n\n<p>I made some features related to time-to-time difference. for example:</p>\n\n<ul>\n<li>(max(mjd) where detected == 1) - (mjd on peak flux)</li>\n<li>(mjd on peak flux) - (min(mjd) where detected == 1)</li>\n<li>(min(mjd) where detected == 1) - (mjd on previous observation)\n<ul><li>tried to capture information where the important signal is lost by its observation schedule</li></ul></li>\n<li>(mjd on 50% tile after the peak flux) - (mjd no peak flux)</li>\n</ul>\n\n<h3>LombScargle</h3>\n\n<p>Adding power and frequency obtained from <code>astropy.LombScargle.autopower()</code> gave a small improvement.</p>\n\n<h2>Pseudo Labelling</h2>\n\n<p>I also tried to use pseudo-labeling in the early stage of the competition. I found that using pseudo-label only in class90 gave me a big boost (0.005 ~ 0.03, depends on the model), but using all classes didn't work. I think it's related class99 (if class99 is similar to class52/62, pseudo-label in these classes contains a lot of false signals). class90+class42 gave a good result too, but its difference from class90-only was very small. </p>",
  "messages": [
    {
      "id": "442200",
      "postDate": "12/19/2018 15:47:29",
      "content": "<p>First of all, thanks to the organizers and all participants in this competition! And thanks a lot to my great teammates, <a href=\"/mamasinkgs\">@mamasinkgs</a>, and <a href=\"/yuval6967\">@yuval6967</a>. I learn a lot from these 2 guys. And I'm very happy to get my first gold medal :)</p>\n\n<p>In this part, I want to share my findings and feature engineering. Please refer to the following discussions to see the overall description of our solution.</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75116\">3rd Place Part I - CNN</a></li>\n<li><a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75131\">3rd Place Part II - CatBoost, mamas feature, and class 99</a></li>\n</ul>\n\n<h2>What I learned from astronomer's work</h2>\n\n<p>At the beginning of this competition, I tried to install domain knowledge and tried to be \"pseudo astronomer\". In addition to the data note provided by the organizers, I carefully read the following resources.</p>\n\n<ul>\n<li><a href=\"https://www.lsst.org/scientists/scibook\">LSST science book</a></li>\n<li><a href=\"https://arxiv.org/abs/1008.1024\">Result from SNPCC challenge</a></li>\n<li><a href=\"https://arxiv.org/abs/1603.00882\">Photometric Supernova Classification With Machine Learning</a> (<a href=\"https://kicp-workshops.uchicago.edu/SNClassification_2016/depot/talk-lochner-michelle.pdf\">slide</a>)</li>\n</ul>\n\n<p>After reading these resources and exploring data a bit, I thought that this competition mainly consists of these 2 challenges:</p>\n\n<ul>\n<li>How to distinguish between various types of supernova classes?</li>\n<li>How to detect class99?</li>\n</ul>\n\n<p>I also tried to understand what each class really is. By 1) basic light curve characteristics 2) class frequency compared with expected LSST observation rate 3) redshift distribution, here is my assumption (order by confidence):</p>\n\n<ul>\n<li>class90: SN Ia</li>\n<li>class42: SN II</li>\n<li>class95: Superluminous Supernova </li>\n<li>class52/62: SN Ib/c</li>\n<li>class88: Strong Gravitational Lens (or AGN?)</li>\n<li>class67: TDE</li>\n</ul>\n\n<p>I hope the competition hosts will reveal the answer!\n<em>I don't think this assumption helped us directly</em>, but this gave us a very good interpretation of the result of our class99 probing (class99 is something similar to core-collapse supernova).</p>\n\n<h2>Feature Engineering</h2>\n\n<h3>template fitting by sncosmo package - the golden feature for us</h3>\n\n<p>sncosmo provides nice lc fitting API in python (<a href=\"https://sncosmo.readthedocs.io/en/v1.6.x/examples/plot_lc_fit.html#sphx-glr-examples-plot-lc-fit-py\">link</a>). I used a fitting parameter and it's chi-square value as a feature. I'm a bit surprised that no one except me uses this library because using template fitting is a major solution in the past competition.</p>\n\n<p>It took a bit long time to calculate (1~2 lines/sec), so I split light curves by object_id%30 and run 30 preemptible instances to finish calculating within 1 day. I repeated this process about 15 times to make various types of models to cover all variant of supernovae (salt-2, salt-2-extended, nugent-sn1bc, nugent-sn2, snana, sako,...), and use 8 of them in my final model. It was an exhausting process... really...</p>\n\n<p>But these bag of template features gave me a significant boost (~0.05 in intermediate, and 0.1+ in my final model). It helped Yuval's CNN as well (+0.04). 7 out of 10 my most important features (by LGBM gain) are these features. All template features were used only in an extragalactic model.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/442200/10905/object_id_161521.png\" alt=\"example of template fitting result\"></p>\n\n<p>Here is an example of a result of template fitting (SALT-2, object_id = 161521). y-band is ignored by its wavelength and estimated redshift. One may think that GP fitting gives better fitting, but it's robust to noise (in SALT-2, only 5 parameters used to generate this all 6-band curve). Probably combining our features with GP features of Kyle or CPMP's will give us another significant boost.</p>\n\n<em>UPDATED</em>\n\n<p>I've attached my template features. you can download, unzip and use:</p>\n\n<p><code>\ndf = pd.read_feather('sncosmo_template_features.f')\n</code></p>\n\n<em>UPDATED2</em>\n\n<p>Kernel is available to see how did we use sncosmo:\n<a href=\"https://www.kaggle.com/nyanpn/salt-2-feature-part-of-3rd-place-solution\">https://www.kaggle.com/nyanpn/salt-2-feature-part-of-3rd-place-solution</a></p>\n\n<h3>hostgal-specz model</h3>\n\n<p>Exactly same as 2nd and 4th place solution. This gave me a small boost (~0.004).</p>\n\n<h3>luminosity</h3>\n\n<p>As shared in the discussion, luminosity is an important intrinsic property of variable stars. I used the following fomla:</p>\n\n<p><code>\nluminosity = (max(flux) - min(flux)) * distance ** 2\n</code></p>\n\n<p>Distance (in MPc) is converted from estimated specz above. Astropy's luminosity_distance (<a href=\"http://docs.astropy.org/en/stable/cosmology/index.html?highlight=luminosity#using-astropy-cosmology\">link</a>) function gives distance from redshift.</p>\n\n<p>I added these luminosity and luminosity difference between channels.</p>\n\n<h3>time difference</h3>\n\n<p>I made some features related to time-to-time difference. for example:</p>\n\n<ul>\n<li>(max(mjd) where detected == 1) - (mjd on peak flux)</li>\n<li>(mjd on peak flux) - (min(mjd) where detected == 1)</li>\n<li>(min(mjd) where detected == 1) - (mjd on previous observation)\n<ul><li>tried to capture information where the important signal is lost by its observation schedule</li></ul></li>\n<li>(mjd on 50% tile after the peak flux) - (mjd no peak flux)</li>\n</ul>\n\n<h3>LombScargle</h3>\n\n<p>Adding power and frequency obtained from <code>astropy.LombScargle.autopower()</code> gave a small improvement.</p>\n\n<h2>Pseudo Labelling</h2>\n\n<p>I also tried to use pseudo-labeling in the early stage of the competition. I found that using pseudo-label only in class90 gave me a big boost (0.005 ~ 0.03, depends on the model), but using all classes didn't work. I think it's related class99 (if class99 is similar to class52/62, pseudo-label in these classes contains a lot of false signals). class90+class42 gave a good result too, but its difference from class90-only was very small. </p>",
      "rawMarkdown": "First of all, thanks to the organizers and all participants in this competition! And thanks a lot to my great teammates, @mamasinkgs, and @yuval6967. I learn a lot from these 2 guys. And I'm very happy to get my first gold medal :)\n\nIn this part, I want to share my findings and feature engineering. Please refer to the following discussions to see the overall description of our solution.\n\n- [3rd Place Part I - CNN][1]\n- [3rd Place Part II - CatBoost, mamas feature, and class 99][2]\n\n## What I learned from astronomer's work\nAt the beginning of this competition, I tried to install domain knowledge and tried to be \"pseudo astronomer\". In addition to the data note provided by the organizers, I carefully read the following resources.\n\n- [LSST science book][3]\n- [Result from SNPCC challenge][4]\n- [Photometric Supernova Classification With Machine Learning][5] ([slide][6])\n\nAfter reading these resources and exploring data a bit, I thought that this competition mainly consists of these 2 challenges:\n\n- How to distinguish between various types of supernova classes?\n- How to detect class99?\n\nI also tried to understand what each class really is. By 1) basic light curve characteristics 2) class frequency compared with expected LSST observation rate 3) redshift distribution, here is my assumption (order by confidence):\n\n- class90: SN Ia\n- class42: SN II\n- class95: Superluminous Supernova \n- class52/62: SN Ib/c\n- class88: Strong Gravitational Lens (or AGN?)\n- class67: TDE\n\nI hope the competition hosts will reveal the answer!\n*I don't think this assumption helped us directly*, but this gave us a very good interpretation of the result of our class99 probing (class99 is something similar to core-collapse supernova).\n\n## Feature Engineering\n### template fitting by sncosmo package - the golden feature for us\nsncosmo provides nice lc fitting API in python ([link](https://sncosmo.readthedocs.io/en/v1.6.x/examples/plot_lc_fit.html#sphx-glr-examples-plot-lc-fit-py)). I used a fitting parameter and it's chi-square value as a feature. I'm a bit surprised that no one except me uses this library because using template fitting is a major solution in the past competition.\n\nIt took a bit long time to calculate (1~2 lines/sec), so I split light curves by object_id%30 and run 30 preemptible instances to finish calculating within 1 day. I repeated this process about 15 times to make various types of models to cover all variant of supernovae (salt-2, salt-2-extended, nugent-sn1bc, nugent-sn2, snana, sako,...), and use 8 of them in my final model. It was an exhausting process... really...\n\nBut these bag of template features gave me a significant boost (~0.05 in intermediate, and 0.1+ in my final model). It helped Yuval's CNN as well (+0.04). 7 out of 10 my most important features (by LGBM gain) are these features. All template features were used only in an extragalactic model.\n\n![example of template fitting result](https://storage.googleapis.com/kaggle-forum-message-attachments/442200/10905/object_id_161521.png)\n\nHere is an example of a result of template fitting (SALT-2, object_id = 161521). y-band is ignored by its wavelength and estimated redshift. One may think that GP fitting gives better fitting, but it's robust to noise (in SALT-2, only 5 parameters used to generate this all 6-band curve). Probably combining our features with GP features of Kyle or CPMP's will give us another significant boost.\n\n#### *UPDATED*\nI've attached my template features. you can download, unzip and use:\n\n```\ndf = pd.read_feather('sncosmo_template_features.f')\n```\n\n#### *UPDATED2*\nKernel is available to see how did we use sncosmo:\nhttps://www.kaggle.com/nyanpn/salt-2-feature-part-of-3rd-place-solution\n\n\n### hostgal-specz model\nExactly same as 2nd and 4th place solution. This gave me a small boost (~0.004).\n\n\n### luminosity\nAs shared in the discussion, luminosity is an important intrinsic property of variable stars. I used the following fomla:\n\n```\nluminosity = (max(flux) - min(flux)) * distance ** 2\n```\n\nDistance (in MPc) is converted from estimated specz above. Astropy's luminosity_distance ([link](http://docs.astropy.org/en/stable/cosmology/index.html?highlight=luminosity#using-astropy-cosmology)) function gives distance from redshift.\n\nI added these luminosity and luminosity difference between channels.\n\n### time difference\nI made some features related to time-to-time difference. for example:\n\n- (max(mjd) where detected == 1) - (mjd on peak flux)\n- (mjd on peak flux) - (min(mjd) where detected == 1)\n- (min(mjd) where detected == 1) - (mjd on previous observation)\n    - tried to capture information where the important signal is lost by its observation schedule\n- (mjd on 50% tile after the peak flux) - (mjd no peak flux)\n\n### LombScargle\nAdding power and frequency obtained from `astropy.LombScargle.autopower()` gave a small improvement.\n\n## Pseudo Labelling\nI also tried to use pseudo-labeling in the early stage of the competition. I found that using pseudo-label only in class90 gave me a big boost (0.005 ~ 0.03, depends on the model), but using all classes didn't work. I think it's related class99 (if class99 is similar to class52/62, pseudo-label in these classes contains a lot of false signals). class90+class42 gave a good result too, but its difference from class90-only was very small. \n\n\n  [1]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75116\n  [2]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75131\n  [3]: https://www.lsst.org/scientists/scibook\n  [4]: https://arxiv.org/abs/1008.1024\n  [5]: https://arxiv.org/abs/1603.00882\n  [6]: https://kicp-workshops.uchicago.edu/SNClassification_2016/depot/talk-lochner-michelle.pdf",
      "votes": null
    },
    {
      "id": "442209",
      "postDate": "12/19/2018 16:05:15",
      "content": "<p>Ah I forgot to say that I majored in aerospace engineering and joined an astrometry satellite project as an engineer before. This is why I call myself as pseudo-astronomer. Anyway, learning about project background is really interesting for me, and motivated me a lot ;)</p>",
      "rawMarkdown": "Ah I forgot to say that I majored in aerospace engineering and joined an astrometry satellite project as an engineer before. This is why I call myself as pseudo-astronomer. Anyway, learning about project background is really interesting for me, and motivated me a lot ;)",
      "votes": null
    },
    {
      "id": "442230",
      "postDate": "12/19/2018 16:39:21",
      "content": "<p>Thanks for sharing, and congrats on the result.  You make me feel even worse than before for not having tried SALT2 harder.  And for combining, I think it is way better to combine with Kyle's GP than with mine ;) But yes, I'm pretty sure combining SALT2 with what we have would give a significant boost.  I can try submitting if you share these template feature values for test and train datasets.</p>",
      "rawMarkdown": "Thanks for sharing, and congrats on the result.  You make me feel even worse than before for not having tried SALT2 harder.  And for combining, I think it is way better to combine with Kyle's GP than with mine ;) But yes, I'm pretty sure combining SALT2 with what we have would give a significant boost.  I can try submitting if you share these template feature values for test and train datasets.",
      "votes": null
    },
    {
      "id": "442331",
      "postDate": "12/19/2018 20:03:18",
      "content": "<p>Thanks nyanp, your performance in this competition is really incredible! Without your template fitting features it's impossible to make suich a high-scored GBDT model, and without your ideas we wouldn't have noticed class99 handling :)</p>",
      "rawMarkdown": "Thanks nyanp, your performance in this competition is really incredible! Without your template fitting features it's impossible to make suich a high-scored GBDT model, and without your ideas we wouldn't have noticed class99 handling :)",
      "votes": null
    },
    {
      "id": "442434",
      "postDate": "12/20/2018 00:43:41",
      "content": "<p>Very interesting results! I'm very intrigued by the sncosmo fits. That package was developed by another Kyle who used to work in my group. I have typically found that the sncosmo fits require quite a bit of hand tuning to get them to converge properly. Did you have any issues with that? There is as nested sampler fitter that is supposed to be more robust. Did you use that?</p>\n\n<p>Another thing to think about is that the data used in this competition was probably generated using a lot of the same models that are in sncosmo. I'm not really sure what the impact of that is...</p>",
      "rawMarkdown": "Very interesting results! I'm very intrigued by the sncosmo fits. That package was developed by another Kyle who used to work in my group. I have typically found that the sncosmo fits require quite a bit of hand tuning to get them to converge properly. Did you have any issues with that? There is as nested sampler fitter that is supposed to be more robust. Did you use that?\n\nAnother thing to think about is that the data used in this competition was probably generated using a lot of the same models that are in sncosmo. I'm not really sure what the impact of that is...",
      "votes": null
    },
    {
      "id": "442627",
      "postDate": "12/20/2018 08:29:08",
      "content": "<p>Thanks @nyanp we learned a lot from you. You also neglected to say it was your intuition that led us to improve our class99 prediction</p>",
      "rawMarkdown": "Thanks @nyanp we learned a lot from you. You also neglected to say it was your intuition that led us to improve our class99 prediction",
      "votes": null
    },
    {
      "id": "442646",
      "postDate": "12/20/2018 09:08:52",
      "content": "<p>Congratulations on your amazing result! I tried sncosmo too but was not able to get useful features out of it. However I only tried salt2 and salt2-extended models. Were either of these 2 models in the final 8 that you used? Can you share what values you used for zp and zpsys? Maybe that was my problem...</p>",
      "rawMarkdown": "Congratulations on your amazing result! I tried sncosmo too but was not able to get useful features out of it. However I only tried salt2 and salt2-extended models. Were either of these 2 models in the final 8 that you used? Can you share what values you used for zp and zpsys? Maybe that was my problem...",
      "votes": null
    },
    {
      "id": "442795",
      "postDate": "12/20/2018 14:11:05",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Thanks for your comment! And I appreciate your sharing in the competition, I learned a lot from you. I've attached all template features (in feather-format), now you can try our features :)</p>",
      "rawMarkdown": "cpmpml Thanks for your comment! And I appreciate your sharing in the competition, I learned a lot from you. I've attached all template features (in feather-format), now you can try our features :)",
      "votes": null
    },
    {
      "id": "442797",
      "postDate": "12/20/2018 14:16:26",
      "content": "<p>Thanks a lot, too :) Your ambition, hard working and deep knowledge about GBDT drove me a lot!</p>",
      "rawMarkdown": "Thanks a lot, too :) Your ambition, hard working and deep knowledge about GBDT drove me a lot!",
      "votes": null
    },
    {
      "id": "442810",
      "postDate": "12/20/2018 14:45:06",
      "content": "<p>Thanks a lot!  Between your share and Kyle's share my week end is ruined ;) Just kidding of course.</p>\n\n<p>Seeing how diverse the top teams solutions are, there is room for killer result, with LB &lt; 0.5 if not lower.</p>",
      "rawMarkdown": "Thanks a lot!  Between your share and Kyle's share my week end is ruined ;) Just kidding of course.\n\nSeeing how diverse the top teams solutions are, there is room for killer result, with LB &lt; 0.5 if not lower.",
      "votes": null
    },
    {
      "id": "442838",
      "postDate": "12/20/2018 15:37:03",
      "content": "<p>First of all, congratulations on your 1st! </p>\n\n<pre><code>That package was developed by another Kyle who used to work in my group.\n</code></pre>\n\n<p>Wow, so why didn't you use SALT-2 feature? ;) Please say thank you to another Kyle, I love this library by it's clean and well-documented API. I tried <code>sncosmo.nest_lc</code> and multinest package but it failed because both were too slow to converge (over 30 seconds/object). I thought nested sampling isn't feasible to this competition, but maybe I made a mistake. </p>\n\n<blockquote>\n  <p>Did you have any issues with that?</p>\n</blockquote>\n\n<ul>\n<li>Sometimes I met SEGV while calculating salt-2 or salt-2-extended feature. I will try to reproduce it in this weekend.</li>\n<li>Its internal solver (iminuit2) does not scale with the number of CPUs. So I needed to run 30x 8-core preemptible instances to speedup whole process. It is nice if we can turn off its internal parallelism (data parallelism is a much better way in this competition).</li>\n</ul>\n\n<blockquote>\n  <p>Another thing to think about is that the data used in this competition was probably generated using a lot of the same models that are in sncosmo. I'm not really sure what the impact of that is…</p>\n</blockquote>\n\n<p>Yes, probably is. Before selecting 15 sources, I tried over 30 types of built-in sources in sncosmo, and choose it by local CV and my intuition. Some sources improved my CV, but some didn't. My final feature set consists of:</p>\n\n<ul>\n<li>salt-2</li>\n<li>salt-2-extended</li>\n<li>hsiao</li>\n<li>nugent-sn1bc</li>\n<li>nugent-sn2n</li>\n<li>snana-2004fe</li>\n<li>snana-2007Y</li>\n</ul>",
      "rawMarkdown": "First of all, congratulations on your 1st! \n\n\n    That package was developed by another Kyle who used to work in my group.\n\nWow, so why didn't you use SALT-2 feature? ;) Please say thank you to another Kyle, I love this library by it's clean and well-documented API. I tried `sncosmo.nest_lc` and multinest package but it failed because both were too slow to converge (over 30 seconds/object). I thought nested sampling isn't feasible to this competition, but maybe I made a mistake. \n\n&gt; Did you have any issues with that?\n\n- Sometimes I met SEGV while calculating salt-2 or salt-2-extended feature. I will try to reproduce it in this weekend.\n- Its internal solver (iminuit2) does not scale with the number of CPUs. So I needed to run 30x 8-core preemptible instances to speedup whole process. It is nice if we can turn off its internal parallelism (data parallelism is a much better way in this competition).\n\n\n&gt; Another thing to think about is that the data used in this competition was probably generated using a lot of the same models that are in sncosmo. I'm not really sure what the impact of that is…\n\n\nYes, probably is. Before selecting 15 sources, I tried over 30 types of built-in sources in sncosmo, and choose it by local CV and my intuition. Some sources improved my CV, but some didn't. My final feature set consists of:\n\n- salt-2\n- salt-2-extended\n- hsiao\n- nugent-sn1bc\n- nugent-sn2n\n- snana-2004fe\n- snana-2007Y",
      "votes": null
    },
    {
      "id": "442844",
      "postDate": "12/20/2018 15:43:55",
      "content": "<p>Thanks a lot, too :) I think class99 prediction is a result of our collaborative working! Combine my intuition with you and mamas's creative ideas on function engineering worked very well.</p>",
      "rawMarkdown": "Thanks a lot, too :) I think class99 prediction is a result of our collaborative working! Combine my intuition with you and mamas's creative ideas on function engineering worked very well.",
      "votes": null
    },
    {
      "id": "442846",
      "postDate": "12/20/2018 15:47:42",
      "content": "<p><a href=\"/taniaj\">@taniaj</a>\nThanks! :)</p>\n\n<blockquote>\n  <p>Were either of these 2 models in the final 8 that you used?</p>\n</blockquote>\n\n<p>I used both, and also added average (weigted by square inverse of error) parameters of these 2 sources. I've attached full features and I've also published simple kernel <a href=\"https://www.kaggle.com/nyanpn/salt-2-feature-part-of-3rd-place-solution\">here</a>, so you can check the difference between us.</p>\n\n<p>One question from me: did you met SEGV while calculating salt2 and salt2-extended models? If so, how to deal with it?</p>",
      "rawMarkdown": "taniaj\nThanks! :)\n\n&gt; Were either of these 2 models in the final 8 that you used?\n\nI used both, and also added average (weigted by square inverse of error) parameters of these 2 sources. I've attached full features and I've also published simple kernel [here](https://www.kaggle.com/nyanpn/salt-2-feature-part-of-3rd-place-solution), so you can check the difference between us.\n\nOne question from me: did you met SEGV while calculating salt2 and salt2-extended models? If so, how to deal with it?",
      "votes": null
    },
    {
      "id": "442849",
      "postDate": "12/20/2018 15:54:49",
      "content": "<blockquote>\n  <p>Seeing how diverse the top teams solutions are</p>\n</blockquote>\n\n<p>I feel there is a lot of insight we can learn from :) And I agree with you, there is a room for LB &lt; 0.5.</p>",
      "rawMarkdown": "&gt; Seeing how diverse the top teams solutions are\n\nI feel there is a lot of insight we can learn from :) And I agree with you, there is a room for LB &lt; 0.5.",
      "votes": null
    },
    {
      "id": "442891",
      "postDate": "12/20/2018 16:47:04",
      "content": "<p>From late submission, I noticed (again) that <strong>combining all template features were much better than single SALT-2 template</strong>. Here is a single LGBM result:</p>\n\n<ol>\n<li>Baseline : 0.911 (public), 0.943 (private)</li>\n<li>Baseline + SALT-2: <strong>0.860</strong> (public), <strong>0.886</strong> (private)</li>\n<li>Baseline + All templates: <strong>0.824</strong> (public), <strong>0.860</strong> (private)</li>\n</ol>",
      "rawMarkdown": "From late submission, I noticed (again) that **combining all template features were much better than single SALT-2 template**. Here is a single LGBM result:\n\n1. Baseline : 0.911 (public), 0.943 (private)\n1. Baseline + SALT-2: **0.860** (public), **0.886** (private)\n1. Baseline + All templates: **0.824** (public), **0.860** (private)",
      "votes": null
    },
    {
      "id": "442952",
      "postDate": "12/20/2018 19:27:05",
      "content": "<p>Thanks for sharing the kernel! Now I can see the differences - you have a different zp from me and your method of setting bounds on z is better. I used photoz directly. </p>\n\n<p>I also got some exceptions (can't remember if it was SEGV or something else) and I just caught them and set all values for those objects to null (there were not many and I assume they were the ones that were not supernovae anyway), plus I wanted to get some rough results fast to figure out whether I should spend more time on sncosmo or not. I guess I gave up on it too soon.</p>",
      "rawMarkdown": "Thanks for sharing the kernel! Now I can see the differences - you have a different zp from me and your method of setting bounds on z is better. I used photoz directly. \n\nI also got some exceptions (can't remember if it was SEGV or something else) and I just caught them and set all values for those objects to null (there were not many and I assume they were the ones that were not supernovae anyway), plus I wanted to get some rough results fast to figure out whether I should spend more time on sncosmo or not. I guess I gave up on it too soon.",
      "votes": null
    },
    {
      "id": "443015",
      "postDate": "12/20/2018 22:22:53",
      "content": "<p>Thank you and congrats ! I'll try to add salt-2 features to our model, it's 0.815 (public) single model at the moment, so we'll see what it will be with the salt-2, will give you an update. I wanted to try it, but it just was not enough time for everything</p>",
      "rawMarkdown": "Thank you and congrats ! I'll try to add salt-2 features to our model, it's 0.815 (public) single model at the moment, so we'll see what it will be with the salt-2, will give you an update. I wanted to try it, but it just was not enough time for everything",
      "votes": null
    },
    {
      "id": "443358",
      "postDate": "12/21/2018 13:47:42",
      "content": "<p>Thanks for the templates parameters.  Unfortunately, I can't read them with pandas.read_feather.  I am sure I could find a way using other packages, but the motivation is gone now..  Thanks anyway, and congrats again for having used sncosmo successfully.</p>",
      "rawMarkdown": "Thanks for the templates parameters.  Unfortunately, I can't read them with pandas.read_feather.  I am sure I could find a way using other packages, but the motivation is gone now..  Thanks anyway, and congrats again for having used sncosmo successfully.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 442209,
      "author_name": "nyanpn",
      "author_url": "",
      "post_date": "12/19/2018 16:05:15",
      "content": "<p>Ah I forgot to say that I majored in aerospace engineering and joined an astrometry satellite project as an engineer before. This is why I call myself as pseudo-astronomer. Anyway, learning about project background is really interesting for me, and motivated me a lot ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 442230,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/19/2018 16:39:21",
      "content": "<p>Thanks for sharing, and congrats on the result.  You make me feel even worse than before for not having tried SALT2 harder.  And for combining, I think it is way better to combine with Kyle's GP than with mine ;) But yes, I'm pretty sure combining SALT2 with what we have would give a significant boost.  I can try submitting if you share these template feature values for test and train datasets.</p>",
      "votes": null,
      "replies": [
        {
          "id": 442795,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "12/20/2018 14:11:05",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Thanks for your comment! And I appreciate your sharing in the competition, I learned a lot from you. I've attached all template features (in feather-format), now you can try our features :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442810,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/20/2018 14:45:06",
          "content": "<p>Thanks a lot!  Between your share and Kyle's share my week end is ruined ;) Just kidding of course.</p>\n\n<p>Seeing how diverse the top teams solutions are, there is room for killer result, with LB &lt; 0.5 if not lower.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442849,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "12/20/2018 15:54:49",
          "content": "<blockquote>\n  <p>Seeing how diverse the top teams solutions are</p>\n</blockquote>\n\n<p>I feel there is a lot of insight we can learn from :) And I agree with you, there is a room for LB &lt; 0.5.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443358,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/21/2018 13:47:42",
          "content": "<p>Thanks for the templates parameters.  Unfortunately, I can't read them with pandas.read_feather.  I am sure I could find a way using other packages, but the motivation is gone now..  Thanks anyway, and congrats again for having used sncosmo successfully.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442331,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "12/19/2018 20:03:18",
      "content": "<p>Thanks nyanp, your performance in this competition is really incredible! Without your template fitting features it's impossible to make suich a high-scored GBDT model, and without your ideas we wouldn't have noticed class99 handling :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 442797,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "12/20/2018 14:16:26",
          "content": "<p>Thanks a lot, too :) Your ambition, hard working and deep knowledge about GBDT drove me a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442434,
      "author_name": "kyleboone",
      "author_url": "",
      "post_date": "12/20/2018 00:43:41",
      "content": "<p>Very interesting results! I'm very intrigued by the sncosmo fits. That package was developed by another Kyle who used to work in my group. I have typically found that the sncosmo fits require quite a bit of hand tuning to get them to converge properly. Did you have any issues with that? There is as nested sampler fitter that is supposed to be more robust. Did you use that?</p>\n\n<p>Another thing to think about is that the data used in this competition was probably generated using a lot of the same models that are in sncosmo. I'm not really sure what the impact of that is...</p>",
      "votes": null,
      "replies": [
        {
          "id": 442838,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "12/20/2018 15:37:03",
          "content": "<p>First of all, congratulations on your 1st! </p>\n\n<pre><code>That package was developed by another Kyle who used to work in my group.\n</code></pre>\n\n<p>Wow, so why didn't you use SALT-2 feature? ;) Please say thank you to another Kyle, I love this library by it's clean and well-documented API. I tried <code>sncosmo.nest_lc</code> and multinest package but it failed because both were too slow to converge (over 30 seconds/object). I thought nested sampling isn't feasible to this competition, but maybe I made a mistake. </p>\n\n<blockquote>\n  <p>Did you have any issues with that?</p>\n</blockquote>\n\n<ul>\n<li>Sometimes I met SEGV while calculating salt-2 or salt-2-extended feature. I will try to reproduce it in this weekend.</li>\n<li>Its internal solver (iminuit2) does not scale with the number of CPUs. So I needed to run 30x 8-core preemptible instances to speedup whole process. It is nice if we can turn off its internal parallelism (data parallelism is a much better way in this competition).</li>\n</ul>\n\n<blockquote>\n  <p>Another thing to think about is that the data used in this competition was probably generated using a lot of the same models that are in sncosmo. I'm not really sure what the impact of that is…</p>\n</blockquote>\n\n<p>Yes, probably is. Before selecting 15 sources, I tried over 30 types of built-in sources in sncosmo, and choose it by local CV and my intuition. Some sources improved my CV, but some didn't. My final feature set consists of:</p>\n\n<ul>\n<li>salt-2</li>\n<li>salt-2-extended</li>\n<li>hsiao</li>\n<li>nugent-sn1bc</li>\n<li>nugent-sn2n</li>\n<li>snana-2004fe</li>\n<li>snana-2007Y</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442627,
      "author_name": "yuval6967",
      "author_url": "",
      "post_date": "12/20/2018 08:29:08",
      "content": "<p>Thanks @nyanp we learned a lot from you. You also neglected to say it was your intuition that led us to improve our class99 prediction</p>",
      "votes": null,
      "replies": [
        {
          "id": 442844,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "12/20/2018 15:43:55",
          "content": "<p>Thanks a lot, too :) I think class99 prediction is a result of our collaborative working! Combine my intuition with you and mamas's creative ideas on function engineering worked very well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442646,
      "author_name": "taniaj",
      "author_url": "",
      "post_date": "12/20/2018 09:08:52",
      "content": "<p>Congratulations on your amazing result! I tried sncosmo too but was not able to get useful features out of it. However I only tried salt2 and salt2-extended models. Were either of these 2 models in the final 8 that you used? Can you share what values you used for zp and zpsys? Maybe that was my problem...</p>",
      "votes": null,
      "replies": [
        {
          "id": 442846,
          "author_name": "nyanpn",
          "author_url": "",
          "post_date": "12/20/2018 15:47:42",
          "content": "<p><a href=\"/taniaj\">@taniaj</a>\nThanks! :)</p>\n\n<blockquote>\n  <p>Were either of these 2 models in the final 8 that you used?</p>\n</blockquote>\n\n<p>I used both, and also added average (weigted by square inverse of error) parameters of these 2 sources. I've attached full features and I've also published simple kernel <a href=\"https://www.kaggle.com/nyanpn/salt-2-feature-part-of-3rd-place-solution\">here</a>, so you can check the difference between us.</p>\n\n<p>One question from me: did you met SEGV while calculating salt2 and salt2-extended models? If so, how to deal with it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442952,
          "author_name": "taniaj",
          "author_url": "",
          "post_date": "12/20/2018 19:27:05",
          "content": "<p>Thanks for sharing the kernel! Now I can see the differences - you have a different zp from me and your method of setting bounds on z is better. I used photoz directly. </p>\n\n<p>I also got some exceptions (can't remember if it was SEGV or something else) and I just caught them and set all values for those objects to null (there were not many and I assume they were the ones that were not supernovae anyway), plus I wanted to get some rough results fast to figure out whether I should spend more time on sncosmo or not. I guess I gave up on it too soon.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442891,
      "author_name": "nyanpn",
      "author_url": "",
      "post_date": "12/20/2018 16:47:04",
      "content": "<p>From late submission, I noticed (again) that <strong>combining all template features were much better than single SALT-2 template</strong>. Here is a single LGBM result:</p>\n\n<ol>\n<li>Baseline : 0.911 (public), 0.943 (private)</li>\n<li>Baseline + SALT-2: <strong>0.860</strong> (public), <strong>0.886</strong> (private)</li>\n<li>Baseline + All templates: <strong>0.824</strong> (public), <strong>0.860</strong> (private)</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 443015,
      "author_name": "blondinka",
      "author_url": "",
      "post_date": "12/20/2018 22:22:53",
      "content": "<p>Thank you and congrats ! I'll try to add salt-2 features to our model, it's 0.815 (public) single model at the moment, so we'll see what it will be with the salt-2, will give you an update. I wanted to try it, but it just was not enough time for everything</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "442200": "First of all, thanks to the organizers and all participants in this competition! And thanks a lot to my great teammates, @mamasinkgs, and @yuval6967. I learn a lot from these 2 guys. And I'm very happy to get my first gold medal :)\n\nIn this part, I want to share my findings and feature engineering. Please refer to the following discussions to see the overall description of our solution.\n\n- [3rd Place Part I - CNN][1]\n- [3rd Place Part II - CatBoost, mamas feature, and class 99][2]\n\n## What I learned from astronomer's work\nAt the beginning of this competition, I tried to install domain knowledge and tried to be \"pseudo astronomer\". In addition to the data note provided by the organizers, I carefully read the following resources.\n\n- [LSST science book][3]\n- [Result from SNPCC challenge][4]\n- [Photometric Supernova Classification With Machine Learning][5] ([slide][6])\n\nAfter reading these resources and exploring data a bit, I thought that this competition mainly consists of these 2 challenges:\n\n- How to distinguish between various types of supernova classes?\n- How to detect class99?\n\nI also tried to understand what each class really is. By 1) basic light curve characteristics 2) class frequency compared with expected LSST observation rate 3) redshift distribution, here is my assumption (order by confidence):\n\n- class90: SN Ia\n- class42: SN II\n- class95: Superluminous Supernova \n- class52/62: SN Ib/c\n- class88: Strong Gravitational Lens (or AGN?)\n- class67: TDE\n\nI hope the competition hosts will reveal the answer!\n*I don't think this assumption helped us directly*, but this gave us a very good interpretation of the result of our class99 probing (class99 is something similar to core-collapse supernova).\n\n## Feature Engineering\n### template fitting by sncosmo package - the golden feature for us\nsncosmo provides nice lc fitting API in python ([link](https://sncosmo.readthedocs.io/en/v1.6.x/examples/plot_lc_fit.html#sphx-glr-examples-plot-lc-fit-py)). I used a fitting parameter and it's chi-square value as a feature. I'm a bit surprised that no one except me uses this library because using template fitting is a major solution in the past competition.\n\nIt took a bit long time to calculate (1~2 lines/sec), so I split light curves by object_id%30 and run 30 preemptible instances to finish calculating within 1 day. I repeated this process about 15 times to make various types of models to cover all variant of supernovae (salt-2, salt-2-extended, nugent-sn1bc, nugent-sn2, snana, sako,...), and use 8 of them in my final model. It was an exhausting process... really...\n\nBut these bag of template features gave me a significant boost (~0.05 in intermediate, and 0.1+ in my final model). It helped Yuval's CNN as well (+0.04). 7 out of 10 my most important features (by LGBM gain) are these features. All template features were used only in an extragalactic model.\n\n![example of template fitting result](https://storage.googleapis.com/kaggle-forum-message-attachments/442200/10905/object_id_161521.png)\n\nHere is an example of a result of template fitting (SALT-2, object_id = 161521). y-band is ignored by its wavelength and estimated redshift. One may think that GP fitting gives better fitting, but it's robust to noise (in SALT-2, only 5 parameters used to generate this all 6-band curve). Probably combining our features with GP features of Kyle or CPMP's will give us another significant boost.\n\n#### *UPDATED*\nI've attached my template features. you can download, unzip and use:\n\n```\ndf = pd.read_feather('sncosmo_template_features.f')\n```\n\n#### *UPDATED2*\nKernel is available to see how did we use sncosmo:\nhttps://www.kaggle.com/nyanpn/salt-2-feature-part-of-3rd-place-solution\n\n\n### hostgal-specz model\nExactly same as 2nd and 4th place solution. This gave me a small boost (~0.004).\n\n\n### luminosity\nAs shared in the discussion, luminosity is an important intrinsic property of variable stars. I used the following fomla:\n\n```\nluminosity = (max(flux) - min(flux)) * distance ** 2\n```\n\nDistance (in MPc) is converted from estimated specz above. Astropy's luminosity_distance ([link](http://docs.astropy.org/en/stable/cosmology/index.html?highlight=luminosity#using-astropy-cosmology)) function gives distance from redshift.\n\nI added these luminosity and luminosity difference between channels.\n\n### time difference\nI made some features related to time-to-time difference. for example:\n\n- (max(mjd) where detected == 1) - (mjd on peak flux)\n- (mjd on peak flux) - (min(mjd) where detected == 1)\n- (min(mjd) where detected == 1) - (mjd on previous observation)\n    - tried to capture information where the important signal is lost by its observation schedule\n- (mjd on 50% tile after the peak flux) - (mjd no peak flux)\n\n### LombScargle\nAdding power and frequency obtained from `astropy.LombScargle.autopower()` gave a small improvement.\n\n## Pseudo Labelling\nI also tried to use pseudo-labeling in the early stage of the competition. I found that using pseudo-label only in class90 gave me a big boost (0.005 ~ 0.03, depends on the model), but using all classes didn't work. I think it's related class99 (if class99 is similar to class52/62, pseudo-label in these classes contains a lot of false signals). class90+class42 gave a good result too, but its difference from class90-only was very small. \n\n\n  [1]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75116\n  [2]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75131\n  [3]: https://www.lsst.org/scientists/scibook\n  [4]: https://arxiv.org/abs/1008.1024\n  [5]: https://arxiv.org/abs/1603.00882\n  [6]: https://kicp-workshops.uchicago.edu/SNClassification_2016/depot/talk-lochner-michelle.pdf",
    "442209": "Ah I forgot to say that I majored in aerospace engineering and joined an astrometry satellite project as an engineer before. This is why I call myself as pseudo-astronomer. Anyway, learning about project background is really interesting for me, and motivated me a lot ;)",
    "442230": "Thanks for sharing, and congrats on the result.  You make me feel even worse than before for not having tried SALT2 harder.  And for combining, I think it is way better to combine with Kyle's GP than with mine ;) But yes, I'm pretty sure combining SALT2 with what we have would give a significant boost.  I can try submitting if you share these template feature values for test and train datasets.",
    "442331": "Thanks nyanp, your performance in this competition is really incredible! Without your template fitting features it's impossible to make suich a high-scored GBDT model, and without your ideas we wouldn't have noticed class99 handling :)",
    "442434": "Very interesting results! I'm very intrigued by the sncosmo fits. That package was developed by another Kyle who used to work in my group. I have typically found that the sncosmo fits require quite a bit of hand tuning to get them to converge properly. Did you have any issues with that? There is as nested sampler fitter that is supposed to be more robust. Did you use that?\n\nAnother thing to think about is that the data used in this competition was probably generated using a lot of the same models that are in sncosmo. I'm not really sure what the impact of that is...",
    "442627": "Thanks @nyanp we learned a lot from you. You also neglected to say it was your intuition that led us to improve our class99 prediction",
    "442646": "Congratulations on your amazing result! I tried sncosmo too but was not able to get useful features out of it. However I only tried salt2 and salt2-extended models. Were either of these 2 models in the final 8 that you used? Can you share what values you used for zp and zpsys? Maybe that was my problem...",
    "442795": "cpmpml Thanks for your comment! And I appreciate your sharing in the competition, I learned a lot from you. I've attached all template features (in feather-format), now you can try our features :)",
    "442797": "Thanks a lot, too :) Your ambition, hard working and deep knowledge about GBDT drove me a lot!",
    "442810": "Thanks a lot!  Between your share and Kyle's share my week end is ruined ;) Just kidding of course.\n\nSeeing how diverse the top teams solutions are, there is room for killer result, with LB &lt; 0.5 if not lower.",
    "442838": "First of all, congratulations on your 1st! \n\n\n    That package was developed by another Kyle who used to work in my group.\n\nWow, so why didn't you use SALT-2 feature? ;) Please say thank you to another Kyle, I love this library by it's clean and well-documented API. I tried `sncosmo.nest_lc` and multinest package but it failed because both were too slow to converge (over 30 seconds/object). I thought nested sampling isn't feasible to this competition, but maybe I made a mistake. \n\n&gt; Did you have any issues with that?\n\n- Sometimes I met SEGV while calculating salt-2 or salt-2-extended feature. I will try to reproduce it in this weekend.\n- Its internal solver (iminuit2) does not scale with the number of CPUs. So I needed to run 30x 8-core preemptible instances to speedup whole process. It is nice if we can turn off its internal parallelism (data parallelism is a much better way in this competition).\n\n\n&gt; Another thing to think about is that the data used in this competition was probably generated using a lot of the same models that are in sncosmo. I'm not really sure what the impact of that is…\n\n\nYes, probably is. Before selecting 15 sources, I tried over 30 types of built-in sources in sncosmo, and choose it by local CV and my intuition. Some sources improved my CV, but some didn't. My final feature set consists of:\n\n- salt-2\n- salt-2-extended\n- hsiao\n- nugent-sn1bc\n- nugent-sn2n\n- snana-2004fe\n- snana-2007Y",
    "442844": "Thanks a lot, too :) I think class99 prediction is a result of our collaborative working! Combine my intuition with you and mamas's creative ideas on function engineering worked very well.",
    "442846": "taniaj\nThanks! :)\n\n&gt; Were either of these 2 models in the final 8 that you used?\n\nI used both, and also added average (weigted by square inverse of error) parameters of these 2 sources. I've attached full features and I've also published simple kernel [here](https://www.kaggle.com/nyanpn/salt-2-feature-part-of-3rd-place-solution), so you can check the difference between us.\n\nOne question from me: did you met SEGV while calculating salt2 and salt2-extended models? If so, how to deal with it?",
    "442849": "&gt; Seeing how diverse the top teams solutions are\n\nI feel there is a lot of insight we can learn from :) And I agree with you, there is a room for LB &lt; 0.5.",
    "442891": "From late submission, I noticed (again) that **combining all template features were much better than single SALT-2 template**. Here is a single LGBM result:\n\n1. Baseline : 0.911 (public), 0.943 (private)\n1. Baseline + SALT-2: **0.860** (public), **0.886** (private)\n1. Baseline + All templates: **0.824** (public), **0.860** (private)",
    "442952": "Thanks for sharing the kernel! Now I can see the differences - you have a different zp from me and your method of setting bounds on z is better. I used photoz directly. \n\nI also got some exceptions (can't remember if it was SEGV or something else) and I just caught them and set all values for those objects to null (there were not many and I assume they were the ones that were not supernovae anyway), plus I wanted to get some rough results fast to figure out whether I should spend more time on sncosmo or not. I guess I gave up on it too soon.",
    "443015": "Thank you and congrats ! I'll try to add salt-2 features to our model, it's 0.815 (public) single model at the moment, so we'll see what it will be with the salt-2, will give you an update. I wanted to try it, but it just was not enough time for everything",
    "443358": "Thanks for the templates parameters.  Unfortunately, I can't read them with pandas.read_feather.  I am sure I could find a way using other packages, but the motivation is gone now..  Thanks anyway, and congrats again for having used sncosmo successfully."
  },
  "source": "meta"
}