{
  "id": 75237,
  "title": "12th Place Solution",
  "url": "/competitions/PLAsTiCC-2018/writeups/go-spartans-12th-place-solution",
  "author_name": "",
  "post_date": "2018-12-19T19:35:47.696144900Z",
  "votes": 18,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, thanks to Kaggle and LSST teams for holding this fantastic competition. It is also very interesting and made me know a lot more about what mysterious things our astronomers are doing. I am very grateful for the community and learned very much from kagglers' generous kernels and discussions. Congratulations to the incredible <a href=\"https://www.kaggle.com/kyleboone\">@Kyle</a>, all the medal winners, my teammates <a href=\"https://www.kaggle.com/lucaskg\"></a><a href=\"/lucaskg\">@lucaskg</a>, <a href=\"https://www.kaggle.com/zuoweijian\"></a><a href=\"/zuoweijian\">@zuoweijian</a>, <a href=\"https://www.kaggle.com/xietian6578\"></a><a href=\"/xietian6578\">@xietian6578</a>, and especially my teammate <a href=\"https://www.kaggle.com/strideradu\"></a><a href=\"/strideradu\">@strideradu</a> who became a Kaggle Competition Master.</p>\n\n<p>Our final model is an ensemble of LightGBM, XGBoost, and binary classifiers. Final CV=0.3252(with specz)/0.3842(without specz), Public LB=0.814, Private LB=0.826.</p>\n\n<p>Here is a summary of what we did during the last tough several weeks.</p>\n\n<h1><strong>1. Feature Engineering</strong></h1>\n\n<h2><strong>1.1 Frequency non-related features</strong></h2>\n\n<p>We performed feature extraction manually one by one, inspired by the functions listed on <a href=\"https://pypi.org/project/FATS/\">FATS</a>. I guess they are basically similar to those extracted from packages such as tsfresh and cesium. Around a total of 50-60 kinds of features are extracted to characterize each light curve. We did both passband-level and object-level feature extractions.</p>\n\n<h2><strong>1.2 Frequency related features</strong></h2>\n\n<p>We referred to <a href=\"https://www.kaggle.com/scirpus/lomb-scargle\">this kernel</a> published by <a href=\"https://www.kaggle.com/scirpus\">@Scirpus</a>. We evaluated the periodogram and grouped the result into ~20 frequency bins. The amplitude was summed in each frequency bin.</p>\n\n<h2><strong>1.3 Bazin</strong></h2>\n\n<p>My teammate <a href=\"https://www.kaggle.com/lucaskg\"></a><a href=\"/lucaskg\">@lucaskg</a> did the curve fitting, using <a href=\"https://arxiv.org/pdf/0904.1066.pdf\">this paper</a> for reference.</p>\n\n<h2><strong>1.4 Flux adjustment &amp; Features while detected == 1</strong></h2>\n\n<p>We performed feature extractions on different datasets, including:</p>\n\n<ul>\n<li><p>original flux features,</p></li>\n<li><p>original flux features while detected == 1,</p></li>\n<li><p>magnitude1 (flux * specz * specz) features,</p></li>\n<li><p>magnitude1 features while detected == 1,</p></li>\n<li><p>magnitude2 (flux * photoz * photoz) features,</p></li>\n<li><p>magnitude2 features while detected == 1</p></li>\n</ul>\n\n<h2><strong>1.5 Feature difference among different passbands or passband groups</strong></h2>\n\n<p>Besides aggregating passband-level features using 'std' and 'mean', we constructed some features that reflect the difference among passbands or passband groups, such as:</p>\n\n<ul>\n<li>'flux max passband i' - 'flux max passband j',</li>\n<li>('flux mean passband0' + 'flux mean passband1' + 'flux mean passband2') - ('flux mean passband3' + 'flux mean passband4' + 'flux mean passband5')</li>\n</ul>\n\n<h2><strong>1.6 Meta Features from passband-level prediction</strong></h2>\n\n<p>Please refer to Part 2.1 below.</p>\n\n<h1><strong>2. Modeling</strong></h1>\n\n<p>We built three models: galactic, extragalactic with specz, and extragalactic without specz.</p>\n\n<h2><strong>2.1 Passband-level modeling</strong></h2>\n\n<p>We ran a LightGBM on passband level. In this model, we only employed those features related to the shape of the light curve, such as flux skewness and frequency features. All features related to the magnitude of the flux, such as flux max and flux mean, are ignored. After we got the passband-level probability prediction, we flattened them as a group of meta-features.</p>\n\n<p>It may be noteworthy that <strong>in the galactic model, we removed all the data with passband == 0</strong>. We found the prediction can be more accurate without them. At least, CV improved with doing that.</p>\n\n<h2><strong>2.2 Object-level modeling</strong></h2>\n\n<p>We applied LightGBM at most of the time during the competition and tried XGBoost on the last day. To avoid overfitting, feature numbers are restricted to 100 for galactic model and 140 for extragalactic model, based on the feature importance obtained by LightGBM. The performance of single models is as follows.</p>\n\n<ul>\n<li>Galactic: LGB CV=0.03149, XGB (no time to run)</li>\n<li>Extragalactic with specz: LGB CV=0.48724, XGB CV=0.50902</li>\n<li>Extragalactic without specz: LGB CV=0.58838, XGB CV=0.61177</li>\n<li>Total: LGB CV=0.3448(with specz)/0.4143(without specz), Public LB=0.831, Private LB=0.847</li>\n</ul>\n\n<h2><strong>2.3 Binary Classification</strong></h2>\n\n<p>Besides the multi-class classification models, my teammate <a href=\"https://www.kaggle.com/lucaskg\"></a><a href=\"/lucaskg\">@lucaskg</a> conducted a binary classification on each of extragalactic classes.</p>\n\n<h2><strong>2.4 Ensemble</strong></h2>\n\n<p>I tried to stack LightGBM and XGBoost by allocating them different weights, but CV did not improve at all. I guess the reason could be the high correlation between them. So I turned to use LightGBM to attempt an ensemble on LightGBM, XGBoost and binary classification predictions. It turned to be useful with CV improved from 0.3448 to 0.3252(with specz), 0.4143 to 0.3842(without specz), public LB score improved from 0.831 to 0.814, and private LB score improved from 0.847 to 0.826.</p>\n\n<h1> <strong>3. Class_99 adjustment</strong></h1>\n\n<p>Actually, what we did on class_99 is not worth mentioning because we only got a 0.001 boost on the leaderboard. However, considering we are finally just 0.0004 ahead of 13th place (who is the first of silver medal winners) on private LB,  it was very crucial to help us achieve the gold medal. </p>\n\n<p>We applied <a href=\"https://www.kaggle.com/ogrellier\"></a><a href=\"/olivier\">@olivier</a>'s method and <a href=\"https://www.kaggle.com/scirpus\">@Scirpus</a>' <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72104\">method</a>, and found the latter is always a little better than the former on public LB. So I tried to interpolate class_99 by setting it to:</p>\n\n<p>class99Prob = class99scirpus + (class99scirpus - class99olivier)</p>\n\n<p>which boosted the LB score by 0.001. Here, we would like to express our sincere thanks and appreciation to <a href=\"https://www.kaggle.com/ogrellier\"></a><a href=\"/olivier\">@olivier</a> and <a href=\"https://www.kaggle.com/scirpus\">@Scirpus</a>, helping us make our gold dream come true.</p>\n\n<h2><strong>Thanks to the Kaggle community and You'll never walk alone!</strong></h2>",
  "messages": [
    {
      "id": "442316",
      "postDate": "12/19/2018 19:35:47",
      "content": "<p>First of all, thanks to Kaggle and LSST teams for holding this fantastic competition. It is also very interesting and made me know a lot more about what mysterious things our astronomers are doing. I am very grateful for the community and learned very much from kagglers' generous kernels and discussions. Congratulations to the incredible <a href=\"https://www.kaggle.com/kyleboone\">@Kyle</a>, all the medal winners, my teammates <a href=\"https://www.kaggle.com/lucaskg\"></a><a href=\"/lucaskg\">@lucaskg</a>, <a href=\"https://www.kaggle.com/zuoweijian\"></a><a href=\"/zuoweijian\">@zuoweijian</a>, <a href=\"https://www.kaggle.com/xietian6578\"></a><a href=\"/xietian6578\">@xietian6578</a>, and especially my teammate <a href=\"https://www.kaggle.com/strideradu\"></a><a href=\"/strideradu\">@strideradu</a> who became a Kaggle Competition Master.</p>\n\n<p>Our final model is an ensemble of LightGBM, XGBoost, and binary classifiers. Final CV=0.3252(with specz)/0.3842(without specz), Public LB=0.814, Private LB=0.826.</p>\n\n<p>Here is a summary of what we did during the last tough several weeks.</p>\n\n<h1><strong>1. Feature Engineering</strong></h1>\n\n<h2><strong>1.1 Frequency non-related features</strong></h2>\n\n<p>We performed feature extraction manually one by one, inspired by the functions listed on <a href=\"https://pypi.org/project/FATS/\">FATS</a>. I guess they are basically similar to those extracted from packages such as tsfresh and cesium. Around a total of 50-60 kinds of features are extracted to characterize each light curve. We did both passband-level and object-level feature extractions.</p>\n\n<h2><strong>1.2 Frequency related features</strong></h2>\n\n<p>We referred to <a href=\"https://www.kaggle.com/scirpus/lomb-scargle\">this kernel</a> published by <a href=\"https://www.kaggle.com/scirpus\">@Scirpus</a>. We evaluated the periodogram and grouped the result into ~20 frequency bins. The amplitude was summed in each frequency bin.</p>\n\n<h2><strong>1.3 Bazin</strong></h2>\n\n<p>My teammate <a href=\"https://www.kaggle.com/lucaskg\"></a><a href=\"/lucaskg\">@lucaskg</a> did the curve fitting, using <a href=\"https://arxiv.org/pdf/0904.1066.pdf\">this paper</a> for reference.</p>\n\n<h2><strong>1.4 Flux adjustment &amp; Features while detected == 1</strong></h2>\n\n<p>We performed feature extractions on different datasets, including:</p>\n\n<ul>\n<li><p>original flux features,</p></li>\n<li><p>original flux features while detected == 1,</p></li>\n<li><p>magnitude1 (flux * specz * specz) features,</p></li>\n<li><p>magnitude1 features while detected == 1,</p></li>\n<li><p>magnitude2 (flux * photoz * photoz) features,</p></li>\n<li><p>magnitude2 features while detected == 1</p></li>\n</ul>\n\n<h2><strong>1.5 Feature difference among different passbands or passband groups</strong></h2>\n\n<p>Besides aggregating passband-level features using 'std' and 'mean', we constructed some features that reflect the difference among passbands or passband groups, such as:</p>\n\n<ul>\n<li>'flux max passband i' - 'flux max passband j',</li>\n<li>('flux mean passband0' + 'flux mean passband1' + 'flux mean passband2') - ('flux mean passband3' + 'flux mean passband4' + 'flux mean passband5')</li>\n</ul>\n\n<h2><strong>1.6 Meta Features from passband-level prediction</strong></h2>\n\n<p>Please refer to Part 2.1 below.</p>\n\n<h1><strong>2. Modeling</strong></h1>\n\n<p>We built three models: galactic, extragalactic with specz, and extragalactic without specz.</p>\n\n<h2><strong>2.1 Passband-level modeling</strong></h2>\n\n<p>We ran a LightGBM on passband level. In this model, we only employed those features related to the shape of the light curve, such as flux skewness and frequency features. All features related to the magnitude of the flux, such as flux max and flux mean, are ignored. After we got the passband-level probability prediction, we flattened them as a group of meta-features.</p>\n\n<p>It may be noteworthy that <strong>in the galactic model, we removed all the data with passband == 0</strong>. We found the prediction can be more accurate without them. At least, CV improved with doing that.</p>\n\n<h2><strong>2.2 Object-level modeling</strong></h2>\n\n<p>We applied LightGBM at most of the time during the competition and tried XGBoost on the last day. To avoid overfitting, feature numbers are restricted to 100 for galactic model and 140 for extragalactic model, based on the feature importance obtained by LightGBM. The performance of single models is as follows.</p>\n\n<ul>\n<li>Galactic: LGB CV=0.03149, XGB (no time to run)</li>\n<li>Extragalactic with specz: LGB CV=0.48724, XGB CV=0.50902</li>\n<li>Extragalactic without specz: LGB CV=0.58838, XGB CV=0.61177</li>\n<li>Total: LGB CV=0.3448(with specz)/0.4143(without specz), Public LB=0.831, Private LB=0.847</li>\n</ul>\n\n<h2><strong>2.3 Binary Classification</strong></h2>\n\n<p>Besides the multi-class classification models, my teammate <a href=\"https://www.kaggle.com/lucaskg\"></a><a href=\"/lucaskg\">@lucaskg</a> conducted a binary classification on each of extragalactic classes.</p>\n\n<h2><strong>2.4 Ensemble</strong></h2>\n\n<p>I tried to stack LightGBM and XGBoost by allocating them different weights, but CV did not improve at all. I guess the reason could be the high correlation between them. So I turned to use LightGBM to attempt an ensemble on LightGBM, XGBoost and binary classification predictions. It turned to be useful with CV improved from 0.3448 to 0.3252(with specz), 0.4143 to 0.3842(without specz), public LB score improved from 0.831 to 0.814, and private LB score improved from 0.847 to 0.826.</p>\n\n<h1> <strong>3. Class_99 adjustment</strong></h1>\n\n<p>Actually, what we did on class_99 is not worth mentioning because we only got a 0.001 boost on the leaderboard. However, considering we are finally just 0.0004 ahead of 13th place (who is the first of silver medal winners) on private LB,  it was very crucial to help us achieve the gold medal. </p>\n\n<p>We applied <a href=\"https://www.kaggle.com/ogrellier\"></a><a href=\"/olivier\">@olivier</a>'s method and <a href=\"https://www.kaggle.com/scirpus\">@Scirpus</a>' <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72104\">method</a>, and found the latter is always a little better than the former on public LB. So I tried to interpolate class_99 by setting it to:</p>\n\n<p>class99Prob = class99scirpus + (class99scirpus - class99olivier)</p>\n\n<p>which boosted the LB score by 0.001. Here, we would like to express our sincere thanks and appreciation to <a href=\"https://www.kaggle.com/ogrellier\"></a><a href=\"/olivier\">@olivier</a> and <a href=\"https://www.kaggle.com/scirpus\">@Scirpus</a>, helping us make our gold dream come true.</p>\n\n<h2><strong>Thanks to the Kaggle community and You'll never walk alone!</strong></h2>",
      "rawMarkdown": "First of all, thanks to Kaggle and LSST teams for holding this fantastic competition. It is also very interesting and made me know a lot more about what mysterious things our astronomers are doing. I am very grateful for the community and learned very much from kagglers' generous kernels and discussions. Congratulations to the incredible [@Kyle][1], all the medal winners, my teammates [@lucaskg][2], [@zuoweijian][3], [@xietian6578][4], and especially my teammate [@strideradu][5] who became a Kaggle Competition Master.\n\nOur final model is an ensemble of LightGBM, XGBoost, and binary classifiers. Final CV=0.3252(with specz)/0.3842(without specz), Public LB=0.814, Private LB=0.826.\n\nHere is a summary of what we did during the last tough several weeks.\n\n**1. Feature Engineering**\n==========================\n\n**1.1 Frequency non-related features**\n----------------------------------------\n\nWe performed feature extraction manually one by one, inspired by the functions listed on [FATS][6]. I guess they are basically similar to those extracted from packages such as tsfresh and cesium. Around a total of 50-60 kinds of features are extracted to characterize each light curve. We did both passband-level and object-level feature extractions.\n\n**1.2 Frequency related features**\n------------------------------------\n\nWe referred to [this kernel][7] published by [@Scirpus][8]. We evaluated the periodogram and grouped the result into ~20 frequency bins. The amplitude was summed in each frequency bin.\n\n**1.3 Bazin**\n------------------------------------\n\nMy teammate [@lucaskg][9] did the curve fitting, using [this paper][10] for reference.\n\n**1.4 Flux adjustment &amp; Features while detected == 1**\n--------------------------------------------------------\n\nWe performed feature extractions on different datasets, including:\n \n - original flux features,\n\n - original flux features while detected == 1,\n\n - magnitude1 (flux * specz * specz) features,\n\n - magnitude1 features while detected == 1,\n\n - magnitude2 (flux * photoz * photoz) features,\n\n - magnitude2 features while detected == 1\n\n**1.5 Feature difference among different passbands or passband groups**\n---------------------------------------------------\n\nBesides aggregating passband-level features using 'std' and 'mean', we constructed some features that reflect the difference among passbands or passband groups, such as:\n\n - 'flux max passband i' - 'flux max passband j',\n - ('flux mean passband0' + 'flux mean passband1' + 'flux mean passband2') - ('flux mean passband3' + 'flux mean passband4' + 'flux mean passband5')\n\n**1.6 Meta Features from passband-level prediction**\n------------------------------------------------------\n\nPlease refer to Part 2.1 below.\n\n**2. Modeling**\n===============\n\nWe built three models: galactic, extragalactic with specz, and extragalactic without specz.\n\n**2.1 Passband-level modeling**\n-------------------------------\n\nWe ran a LightGBM on passband level. In this model, we only employed those features related to the shape of the light curve, such as flux skewness and frequency features. All features related to the magnitude of the flux, such as flux max and flux mean, are ignored. After we got the passband-level probability prediction, we flattened them as a group of meta-features.\n\nIt may be noteworthy that **in the galactic model, we removed all the data with passband == 0**. We found the prediction can be more accurate without them. At least, CV improved with doing that.\n\n**2.2 Object-level modeling**\n-----------------------------\n\nWe applied LightGBM at most of the time during the competition and tried XGBoost on the last day. To avoid overfitting, feature numbers are restricted to 100 for galactic model and 140 for extragalactic model, based on the feature importance obtained by LightGBM. The performance of single models is as follows.\n \n\n - Galactic: LGB CV=0.03149, XGB (no time to run)\n - Extragalactic with specz: LGB CV=0.48724, XGB CV=0.50902\n - Extragalactic without specz: LGB CV=0.58838, XGB CV=0.61177\n - Total: LGB CV=0.3448(with specz)/0.4143(without specz), Public LB=0.831, Private LB=0.847\n\n**2.3 Binary Classification**\n-----------------------------\n\nBesides the multi-class classification models, my teammate [@lucaskg][11] conducted a binary classification on each of extragalactic classes.\n\n**2.4 Ensemble**\n----------------\n\nI tried to stack LightGBM and XGBoost by allocating them different weights, but CV did not improve at all. I guess the reason could be the high correlation between them. So I turned to use LightGBM to attempt an ensemble on LightGBM, XGBoost and binary classification predictions. It turned to be useful with CV improved from 0.3448 to 0.3252(with specz), 0.4143 to 0.3842(without specz), public LB score improved from 0.831 to 0.814, and private LB score improved from 0.847 to 0.826.\n\n **3. Class_99 adjustment**\n===============\nActually, what we did on class_99 is not worth mentioning because we only got a 0.001 boost on the leaderboard. However, considering we are finally just 0.0004 ahead of 13th place (who is the first of silver medal winners) on private LB,  it was very crucial to help us achieve the gold medal. \n\nWe applied [@olivier][12]'s method and [@Scirpus][8]' [method][13], and found the latter is always a little better than the former on public LB. So I tried to interpolate class_99 by setting it to:\n\nclass99Prob = class99scirpus + (class99scirpus - class99olivier)\n\nwhich boosted the LB score by 0.001. Here, we would like to express our sincere thanks and appreciation to [@olivier][12] and [@Scirpus][8], helping us make our gold dream come true.\n\n\n**Thanks to the Kaggle community and You'll never walk alone!**\n---------------------------------------------------------------\n\n  [1]: https://www.kaggle.com/kyleboone\n  [2]: https://www.kaggle.com/lucaskg\n  [3]: https://www.kaggle.com/zuoweijian\n  [4]: https://www.kaggle.com/xietian6578\n  [5]: https://www.kaggle.com/strideradu\n  [6]: https://pypi.org/project/FATS/\n  [7]: https://www.kaggle.com/scirpus/lomb-scargle\n  [8]: https://www.kaggle.com/scirpus\n  [9]: https://www.kaggle.com/lucaskg\n  [10]: https://arxiv.org/pdf/0904.1066.pdf\n  [11]: https://www.kaggle.com/lucaskg\n  [12]: https://www.kaggle.com/ogrellier\n  [13]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72104",
      "votes": null
    },
    {
      "id": "445764",
      "postDate": "12/27/2018 03:17:42",
      "content": "<p>Seems feature engineering &amp; ensemble are very important in this task, thanks for your sharing.</p>",
      "rawMarkdown": "Seems feature engineering &amp; ensemble are very important in this task, thanks for your sharing.",
      "votes": null
    },
    {
      "id": "446029",
      "postDate": "12/27/2018 11:18:14",
      "content": "<p>Thank you for sharing!!</p>\n\n<blockquote>\n  <p>It may be noteworthy that in the galactic model, we removed all the data with passband == 0. \n  How did you find that interesting process?</p>\n</blockquote>",
      "rawMarkdown": "Thank you for sharing!!\n&gt; It may be noteworthy that in the galactic model, we removed all the data with passband == 0. \nHow did you find that interesting process?",
      "votes": null
    },
    {
      "id": "446116",
      "postDate": "12/27/2018 14:29:36",
      "content": "<p>We checked the confusion matrix of our passband-level model and found that the accuracy is always better if only using data of higher passbands. But we didn't have time to try others. </p>",
      "rawMarkdown": "We checked the confusion matrix of our passband-level model and found that the accuracy is always better if only using data of higher passbands. But we didn't have time to try others.",
      "votes": null
    },
    {
      "id": "446568",
      "postDate": "12/28/2018 09:42:53",
      "content": "<p>Checking passband-level confusion matrix is great idea!! <br>\nI think this idea is useful in general. <br>\nThank you!!!</p>",
      "rawMarkdown": "Checking passband-level confusion matrix is great idea!!  \nI think this idea is useful in general.  \nThank you!!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 445764,
      "author_name": "leekltw1",
      "author_url": "",
      "post_date": "12/27/2018 03:17:42",
      "content": "<p>Seems feature engineering &amp; ensemble are very important in this task, thanks for your sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 446029,
      "author_name": "go5kuramubon",
      "author_url": "",
      "post_date": "12/27/2018 11:18:14",
      "content": "<p>Thank you for sharing!!</p>\n\n<blockquote>\n  <p>It may be noteworthy that in the galactic model, we removed all the data with passband == 0. \n  How did you find that interesting process?</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 446116,
          "author_name": "lfcpeng17",
          "author_url": "",
          "post_date": "12/27/2018 14:29:36",
          "content": "<p>We checked the confusion matrix of our passband-level model and found that the accuracy is always better if only using data of higher passbands. But we didn't have time to try others. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 446568,
          "author_name": "go5kuramubon",
          "author_url": "",
          "post_date": "12/28/2018 09:42:53",
          "content": "<p>Checking passband-level confusion matrix is great idea!! <br>\nI think this idea is useful in general. <br>\nThank you!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "442316": "First of all, thanks to Kaggle and LSST teams for holding this fantastic competition. It is also very interesting and made me know a lot more about what mysterious things our astronomers are doing. I am very grateful for the community and learned very much from kagglers' generous kernels and discussions. Congratulations to the incredible [@Kyle][1], all the medal winners, my teammates [@lucaskg][2], [@zuoweijian][3], [@xietian6578][4], and especially my teammate [@strideradu][5] who became a Kaggle Competition Master.\n\nOur final model is an ensemble of LightGBM, XGBoost, and binary classifiers. Final CV=0.3252(with specz)/0.3842(without specz), Public LB=0.814, Private LB=0.826.\n\nHere is a summary of what we did during the last tough several weeks.\n\n**1. Feature Engineering**\n==========================\n\n**1.1 Frequency non-related features**\n----------------------------------------\n\nWe performed feature extraction manually one by one, inspired by the functions listed on [FATS][6]. I guess they are basically similar to those extracted from packages such as tsfresh and cesium. Around a total of 50-60 kinds of features are extracted to characterize each light curve. We did both passband-level and object-level feature extractions.\n\n**1.2 Frequency related features**\n------------------------------------\n\nWe referred to [this kernel][7] published by [@Scirpus][8]. We evaluated the periodogram and grouped the result into ~20 frequency bins. The amplitude was summed in each frequency bin.\n\n**1.3 Bazin**\n------------------------------------\n\nMy teammate [@lucaskg][9] did the curve fitting, using [this paper][10] for reference.\n\n**1.4 Flux adjustment &amp; Features while detected == 1**\n--------------------------------------------------------\n\nWe performed feature extractions on different datasets, including:\n \n - original flux features,\n\n - original flux features while detected == 1,\n\n - magnitude1 (flux * specz * specz) features,\n\n - magnitude1 features while detected == 1,\n\n - magnitude2 (flux * photoz * photoz) features,\n\n - magnitude2 features while detected == 1\n\n**1.5 Feature difference among different passbands or passband groups**\n---------------------------------------------------\n\nBesides aggregating passband-level features using 'std' and 'mean', we constructed some features that reflect the difference among passbands or passband groups, such as:\n\n - 'flux max passband i' - 'flux max passband j',\n - ('flux mean passband0' + 'flux mean passband1' + 'flux mean passband2') - ('flux mean passband3' + 'flux mean passband4' + 'flux mean passband5')\n\n**1.6 Meta Features from passband-level prediction**\n------------------------------------------------------\n\nPlease refer to Part 2.1 below.\n\n**2. Modeling**\n===============\n\nWe built three models: galactic, extragalactic with specz, and extragalactic without specz.\n\n**2.1 Passband-level modeling**\n-------------------------------\n\nWe ran a LightGBM on passband level. In this model, we only employed those features related to the shape of the light curve, such as flux skewness and frequency features. All features related to the magnitude of the flux, such as flux max and flux mean, are ignored. After we got the passband-level probability prediction, we flattened them as a group of meta-features.\n\nIt may be noteworthy that **in the galactic model, we removed all the data with passband == 0**. We found the prediction can be more accurate without them. At least, CV improved with doing that.\n\n**2.2 Object-level modeling**\n-----------------------------\n\nWe applied LightGBM at most of the time during the competition and tried XGBoost on the last day. To avoid overfitting, feature numbers are restricted to 100 for galactic model and 140 for extragalactic model, based on the feature importance obtained by LightGBM. The performance of single models is as follows.\n \n\n - Galactic: LGB CV=0.03149, XGB (no time to run)\n - Extragalactic with specz: LGB CV=0.48724, XGB CV=0.50902\n - Extragalactic without specz: LGB CV=0.58838, XGB CV=0.61177\n - Total: LGB CV=0.3448(with specz)/0.4143(without specz), Public LB=0.831, Private LB=0.847\n\n**2.3 Binary Classification**\n-----------------------------\n\nBesides the multi-class classification models, my teammate [@lucaskg][11] conducted a binary classification on each of extragalactic classes.\n\n**2.4 Ensemble**\n----------------\n\nI tried to stack LightGBM and XGBoost by allocating them different weights, but CV did not improve at all. I guess the reason could be the high correlation between them. So I turned to use LightGBM to attempt an ensemble on LightGBM, XGBoost and binary classification predictions. It turned to be useful with CV improved from 0.3448 to 0.3252(with specz), 0.4143 to 0.3842(without specz), public LB score improved from 0.831 to 0.814, and private LB score improved from 0.847 to 0.826.\n\n **3. Class_99 adjustment**\n===============\nActually, what we did on class_99 is not worth mentioning because we only got a 0.001 boost on the leaderboard. However, considering we are finally just 0.0004 ahead of 13th place (who is the first of silver medal winners) on private LB,  it was very crucial to help us achieve the gold medal. \n\nWe applied [@olivier][12]'s method and [@Scirpus][8]' [method][13], and found the latter is always a little better than the former on public LB. So I tried to interpolate class_99 by setting it to:\n\nclass99Prob = class99scirpus + (class99scirpus - class99olivier)\n\nwhich boosted the LB score by 0.001. Here, we would like to express our sincere thanks and appreciation to [@olivier][12] and [@Scirpus][8], helping us make our gold dream come true.\n\n\n**Thanks to the Kaggle community and You'll never walk alone!**\n---------------------------------------------------------------\n\n  [1]: https://www.kaggle.com/kyleboone\n  [2]: https://www.kaggle.com/lucaskg\n  [3]: https://www.kaggle.com/zuoweijian\n  [4]: https://www.kaggle.com/xietian6578\n  [5]: https://www.kaggle.com/strideradu\n  [6]: https://pypi.org/project/FATS/\n  [7]: https://www.kaggle.com/scirpus/lomb-scargle\n  [8]: https://www.kaggle.com/scirpus\n  [9]: https://www.kaggle.com/lucaskg\n  [10]: https://arxiv.org/pdf/0904.1066.pdf\n  [11]: https://www.kaggle.com/lucaskg\n  [12]: https://www.kaggle.com/ogrellier\n  [13]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72104",
    "445764": "Seems feature engineering &amp; ensemble are very important in this task, thanks for your sharing.",
    "446029": "Thank you for sharing!!\n&gt; It may be noteworthy that in the galactic model, we removed all the data with passband == 0. \nHow did you find that interesting process?",
    "446116": "We checked the confusion matrix of our passband-level model and found that the accuracy is always better if only using data of higher passbands. But we didn't have time to try others.",
    "446568": "Checking passband-level confusion matrix is great idea!!  \nI think this idea is useful in general.  \nThank you!!!"
  },
  "source": "meta"
}