{
  "id": 75316,
  "title": "9th place solution",
  "url": "/competitions/PLAsTiCC-2018/writeups/three-musketeers-9th-place-solution",
  "author_name": "",
  "post_date": "2018-12-20T13:42:01.449831700Z",
  "votes": 24,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thanks to everybody for this competition. Besides enjoying it greatly it has been extremely instructive for me. Thanks also to my teammates <a href=\"https://www.kaggle.com/meaninglesslives\">Siddhartha</a> and <a href=\"https://www.kaggle.com/naka2ka\">Ynktk</a> for their great spirit, hard work, and ideas. Finally, a big thanks to those who have shared their kernels and their thoughts in the discussion board throughout the competition. The open flow of ideas enrichened the competition invaluably.</p>\n\n<p>Our solution relied primarily on careful feature engineering and ensembling. Below I present an overview of it. </p>\n\n<p><strong>Stacking</strong>\nOur best submissions were obtained by stacking the predictions of different lgb, catboost, and nn models, and then taking a weighted average with our previous best submission. This last step improved the score by 0.02 with respect to submitting the meta-model predictions alone (for our last submissions).</p>\n\n<p>As meta-model we initially used a shallow lgbm with depth 1 and learning rate 0.01. At the end of the competition we increased the depth to 3. </p>\n\n<p><strong>First-level models</strong> \nMy teammate Siddhartha focused mostly on a densely connected CNN model. You can read more about his ideas <a href=\"https://www.kaggle.com/meaninglesslives/a-slightly-better-nn-arch-and-some-tricks\">here</a>. Ynktk created a variety of nn, xgb, lgb and catboost models. He will add some notes on these later in this thread. We also experimented a little with regularized greedy forests but they did not add much to the ensemble. </p>\n\n<p>Personally I focused on a single first-level lgb model. This was split into two: one was trained on galactic objects, and the other on extragalactic ones. Each training set used different features. This was the model with best lb score (0.863) of our ensemble as far as we know.</p>\n\n<p><strong>Features</strong> \nI engineered around 8000 features and used lgb importance to select the top 80 for the galactic set and the top 130 for the extragalactic set (approx). Some notes on this part:</p>\n\n<ul>\n<li><p>First of all I would like to point to <a href=\"/manugangler\">@manugangler</a>'s <a href=\"https://www.kaggle.com/manugangler/optimal-feature-extraction-for-class-6\">kernel</a>, where he provided a method for fitting light curves to microlensing events. The features obtained from there worked wonderfully for us, especially after taking passband-wise ratios and differences for the microamp and microbase features, respectively. Einstein time is the most important feature of my model.</p></li>\n<li><p>I browsed and read some texts on astrophysics which gave me some ideas for feature engineering. Below I comment the most useful ones for my model. </p>\n\n<ul><li><p><code>distmod/log10(hostgal_photoz)</code>. If I understood correctly, this roughly measures how far from being a so-called 'standard candle' the object is (setting standard candles to be 'standard' supernovae). The actual formula is more complicated but for a tree-based approach this suffices. This is the 2nd or 3rd most important feature of my model.</p></li>\n<li><p><code>log_10( luminosity_mean )</code> where <code>luminosity = 4pi * (distance in parsecs) * flux</code> (I may be very confused and this may have nothing to do with luminosity). Some similar feats explained by Kyle Boone like magnitude <code>-2.5log_10(max flux in passband i) - distmod</code> and their ratios. There are several other feats. If someone is interested I can explain further.</p></li></ul></li>\n<li><p>I used different sets of time observations for feature agreggation. One set used all observations, the other only the ones marked as detected, and the other only the ones marked as undetected. After aggregation I took ratios passband-wise and also I created some other ratio variables such as <code>passband_std_undetected/mjd_diff_detected</code> and many others.</p></li>\n</ul>\n\n<p><strong>Some notes</strong></p>\n\n<ul>\n<li><p>Siddhartha augmented the train set by randomly selecting from 30% to 70% time observations for each object id, and then creating new objects with the selected data. This improved his nn cv score by 0.01. It did not seem to improve my lgbm. Looking now at other solutions it may have been a good idea to spend some more time thinking on this idea.</p></li>\n<li><p>Class 99: we used Olivier's and Cpmp's method throughout the competition. The less than 5 submissions we used for probing this class were unsuccessful.</p></li>\n<li><p>We tried to model hostgal_specz but this did not work for us.</p></li>\n<li><p>Pseudolabeling also did not work.</p></li>\n<li><p>Our public lb top score was 0.792. However the best score obtained by one of our single first-level models was 0.863. Hence I believe one of the strong aspects of our solution is the diversity of models constructed, which may have been achieved thanks to the fact that each team member used mostly his own features and set up. These were quite different between the three of us. </p></li>\n</ul>\n\n<p>This is all. Thank you for reading! If you have any question I will be pleased to answer</p>",
  "messages": [
    {
      "id": "442778",
      "postDate": "12/20/2018 13:42:01",
      "content": "<p>Thanks to everybody for this competition. Besides enjoying it greatly it has been extremely instructive for me. Thanks also to my teammates <a href=\"https://www.kaggle.com/meaninglesslives\">Siddhartha</a> and <a href=\"https://www.kaggle.com/naka2ka\">Ynktk</a> for their great spirit, hard work, and ideas. Finally, a big thanks to those who have shared their kernels and their thoughts in the discussion board throughout the competition. The open flow of ideas enrichened the competition invaluably.</p>\n\n<p>Our solution relied primarily on careful feature engineering and ensembling. Below I present an overview of it. </p>\n\n<p><strong>Stacking</strong>\nOur best submissions were obtained by stacking the predictions of different lgb, catboost, and nn models, and then taking a weighted average with our previous best submission. This last step improved the score by 0.02 with respect to submitting the meta-model predictions alone (for our last submissions).</p>\n\n<p>As meta-model we initially used a shallow lgbm with depth 1 and learning rate 0.01. At the end of the competition we increased the depth to 3. </p>\n\n<p><strong>First-level models</strong> \nMy teammate Siddhartha focused mostly on a densely connected CNN model. You can read more about his ideas <a href=\"https://www.kaggle.com/meaninglesslives/a-slightly-better-nn-arch-and-some-tricks\">here</a>. Ynktk created a variety of nn, xgb, lgb and catboost models. He will add some notes on these later in this thread. We also experimented a little with regularized greedy forests but they did not add much to the ensemble. </p>\n\n<p>Personally I focused on a single first-level lgb model. This was split into two: one was trained on galactic objects, and the other on extragalactic ones. Each training set used different features. This was the model with best lb score (0.863) of our ensemble as far as we know.</p>\n\n<p><strong>Features</strong> \nI engineered around 8000 features and used lgb importance to select the top 80 for the galactic set and the top 130 for the extragalactic set (approx). Some notes on this part:</p>\n\n<ul>\n<li><p>First of all I would like to point to <a href=\"/manugangler\">@manugangler</a>'s <a href=\"https://www.kaggle.com/manugangler/optimal-feature-extraction-for-class-6\">kernel</a>, where he provided a method for fitting light curves to microlensing events. The features obtained from there worked wonderfully for us, especially after taking passband-wise ratios and differences for the microamp and microbase features, respectively. Einstein time is the most important feature of my model.</p></li>\n<li><p>I browsed and read some texts on astrophysics which gave me some ideas for feature engineering. Below I comment the most useful ones for my model. </p>\n\n<ul><li><p><code>distmod/log10(hostgal_photoz)</code>. If I understood correctly, this roughly measures how far from being a so-called 'standard candle' the object is (setting standard candles to be 'standard' supernovae). The actual formula is more complicated but for a tree-based approach this suffices. This is the 2nd or 3rd most important feature of my model.</p></li>\n<li><p><code>log_10( luminosity_mean )</code> where <code>luminosity = 4pi * (distance in parsecs) * flux</code> (I may be very confused and this may have nothing to do with luminosity). Some similar feats explained by Kyle Boone like magnitude <code>-2.5log_10(max flux in passband i) - distmod</code> and their ratios. There are several other feats. If someone is interested I can explain further.</p></li></ul></li>\n<li><p>I used different sets of time observations for feature agreggation. One set used all observations, the other only the ones marked as detected, and the other only the ones marked as undetected. After aggregation I took ratios passband-wise and also I created some other ratio variables such as <code>passband_std_undetected/mjd_diff_detected</code> and many others.</p></li>\n</ul>\n\n<p><strong>Some notes</strong></p>\n\n<ul>\n<li><p>Siddhartha augmented the train set by randomly selecting from 30% to 70% time observations for each object id, and then creating new objects with the selected data. This improved his nn cv score by 0.01. It did not seem to improve my lgbm. Looking now at other solutions it may have been a good idea to spend some more time thinking on this idea.</p></li>\n<li><p>Class 99: we used Olivier's and Cpmp's method throughout the competition. The less than 5 submissions we used for probing this class were unsuccessful.</p></li>\n<li><p>We tried to model hostgal_specz but this did not work for us.</p></li>\n<li><p>Pseudolabeling also did not work.</p></li>\n<li><p>Our public lb top score was 0.792. However the best score obtained by one of our single first-level models was 0.863. Hence I believe one of the strong aspects of our solution is the diversity of models constructed, which may have been achieved thanks to the fact that each team member used mostly his own features and set up. These were quite different between the three of us. </p></li>\n</ul>\n\n<p>This is all. Thank you for reading! If you have any question I will be pleased to answer</p>",
      "rawMarkdown": "Thanks to everybody for this competition. Besides enjoying it greatly it has been extremely instructive for me. Thanks also to my teammates [Siddhartha][1] and [Ynktk][2] for their great spirit, hard work, and ideas. Finally, a big thanks to those who have shared their kernels and their thoughts in the discussion board throughout the competition. The open flow of ideas enrichened the competition invaluably.\n\nOur solution relied primarily on careful feature engineering and ensembling. Below I present an overview of it. \n\n**Stacking**\nOur best submissions were obtained by stacking the predictions of different lgb, catboost, and nn models, and then taking a weighted average with our previous best submission. This last step improved the score by 0.02 with respect to submitting the meta-model predictions alone (for our last submissions).\n\nAs meta-model we initially used a shallow lgbm with depth 1 and learning rate 0.01. At the end of the competition we increased the depth to 3. \n\n**First-level models** \nMy teammate Siddhartha focused mostly on a densely connected CNN model. You can read more about his ideas [here][3]. Ynktk created a variety of nn, xgb, lgb and catboost models. He will add some notes on these later in this thread. We also experimented a little with regularized greedy forests but they did not add much to the ensemble. \n\nPersonally I focused on a single first-level lgb model. This was split into two: one was trained on galactic objects, and the other on extragalactic ones. Each training set used different features. This was the model with best lb score (0.863) of our ensemble as far as we know.\n\n**Features** \nI engineered around 8000 features and used lgb importance to select the top 80 for the galactic set and the top 130 for the extragalactic set (approx). Some notes on this part:\n\n - First of all I would like to point to @manugangler's [kernel][5], where he provided a method for fitting light curves to microlensing events. The features obtained from there worked wonderfully for us, especially after taking passband-wise ratios and differences for the microamp and microbase features, respectively. Einstein time is the most important feature of my model.\n\n - I browsed and read some texts on astrophysics which gave me some ideas for feature engineering. Below I comment the most useful ones for my model. \n\n  - `distmod/log10(hostgal_photoz)`. If I understood correctly, this roughly measures how far from being a so-called 'standard candle' the object is (setting standard candles to be 'standard' supernovae). The actual formula is more complicated but for a tree-based approach this suffices. This is the 2nd or 3rd most important feature of my model.\n  \n  - `log_10( luminosity_mean )` where `luminosity = 4pi * (distance in parsecs) * flux` (I may be very confused and this may have nothing to do with luminosity). Some similar feats explained by Kyle Boone like magnitude `-2.5log_10(max flux in passband i) - distmod` and their ratios. There are several other feats. If someone is interested I can explain further.\n\n - I used different sets of time observations for feature agreggation. One set used all observations, the other only the ones marked as detected, and the other only the ones marked as undetected. After aggregation I took ratios passband-wise and also I created some other ratio variables such as `passband_std_undetected/mjd_diff_detected` and many others.\n\n**Some notes**\n\n- Siddhartha augmented the train set by randomly selecting from 30% to 70% time observations for each object id, and then creating new objects with the selected data. This improved his nn cv score by 0.01. It did not seem to improve my lgbm. Looking now at other solutions it may have been a good idea to spend some more time thinking on this idea.\n\n- Class 99: we used Olivier's and Cpmp's method throughout the competition. The less than 5 submissions we used for probing this class were unsuccessful.\n\n- We tried to model hostgal_specz but this did not work for us.\n\n- Pseudolabeling also did not work.\n\n- Our public lb top score was 0.792. However the best score obtained by one of our single first-level models was 0.863. Hence I believe one of the strong aspects of our solution is the diversity of models constructed, which may have been achieved thanks to the fact that each team member used mostly his own features and set up. These were quite different between the three of us. \n\nThis is all. Thank you for reading! If you have any question I will be pleased to answer\n\n\n  [1]: https://www.kaggle.com/meaninglesslives\n  [2]: https://www.kaggle.com/naka2ka\n  [3]: https://www.kaggle.com/meaninglesslives/a-slightly-better-nn-arch-and-some-tricks\n  [5]: https://www.kaggle.com/manugangler/optimal-feature-extraction-for-class-6",
      "votes": null
    },
    {
      "id": "442811",
      "postDate": "12/20/2018 14:47:48",
      "content": "<p>Congrats on the result, and thanks for sharing.  </p>",
      "rawMarkdown": "Congrats on the result, and thanks for sharing.",
      "votes": null
    },
    {
      "id": "442918",
      "postDate": "12/20/2018 17:49:18",
      "content": "<p>Congrats ! Glad to see that my features were useful.</p>",
      "rawMarkdown": "Congrats ! Glad to see that my features were useful.",
      "votes": null
    },
    {
      "id": "444374",
      "postDate": "12/23/2018 22:29:22",
      "content": "<p>Zorionak!!! 8000 features?? Only Bilbao's people can get something like that ;-) </p>",
      "rawMarkdown": "Zorionak!!! 8000 features?? Only Bilbao's people can get something like that ;-)",
      "votes": null
    },
    {
      "id": "444555",
      "postDate": "12/24/2018 09:00:57",
      "content": "<p>Eskerrik asko! Good one :)</p>\n\n<p>In fact at some point I got to up to 20000 features but only found new overfitting feats, so I went back to the 8000 set (I only used around 200 of these for training though).</p>\n\n<p>It is mainly due to taking passband-wise ratios and having three big classes of variables (detected, undetected, all) that the number grew so much..</p>",
      "rawMarkdown": "Eskerrik asko! Good one :)\n\nIn fact at some point I got to up to 20000 features but only found new overfitting feats, so I went back to the 8000 set (I only used around 200 of these for training though).\n\nIt is mainly due to taking passband-wise ratios and having three big classes of variables (detected, undetected, all) that the number grew so much..",
      "votes": null
    },
    {
      "id": "452090",
      "postDate": "01/08/2019 07:13:06",
      "content": "<p>thanks for sharing, but I wonder how do you decide which 200 features to use ?</p>",
      "rawMarkdown": "thanks for sharing, but I wonder how do you decide which 200 features to use ?",
      "votes": null
    },
    {
      "id": "495584",
      "postDate": "03/21/2019 09:28:56",
      "content": "<p>Hi <a href=\"/yangddd\">@yangddd</a>, sorry for the late reply</p>\n\n<p>I iteratively used lgb feature importance assessment. More precisely, I started by running lgb with all features and afterwards selecting the top 50% best ranked features. Then I restarted the process with this selected features, and so on. I stopped approximately when my CV score stopped improving and when I suspected/verified that the improvement would not generalize to the LB</p>",
      "rawMarkdown": "Hi @yangddd, sorry for the late reply\n\nI iteratively used lgb feature importance assessment. More precisely, I started by running lgb with all features and afterwards selecting the top 50% best ranked features. Then I restarted the process with this selected features, and so on. I stopped approximately when my CV score stopped improving and when I suspected/verified that the improvement would not generalize to the LB",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 442811,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/20/2018 14:47:48",
      "content": "<p>Congrats on the result, and thanks for sharing.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 442918,
      "author_name": "manugangler",
      "author_url": "",
      "post_date": "12/20/2018 17:49:18",
      "content": "<p>Congrats ! Glad to see that my features were useful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 444374,
      "author_name": "joxemi",
      "author_url": "",
      "post_date": "12/23/2018 22:29:22",
      "content": "<p>Zorionak!!! 8000 features?? Only Bilbao's people can get something like that ;-) </p>",
      "votes": null,
      "replies": [
        {
          "id": 444555,
          "author_name": "agarreta",
          "author_url": "",
          "post_date": "12/24/2018 09:00:57",
          "content": "<p>Eskerrik asko! Good one :)</p>\n\n<p>In fact at some point I got to up to 20000 features but only found new overfitting feats, so I went back to the 8000 set (I only used around 200 of these for training though).</p>\n\n<p>It is mainly due to taking passband-wise ratios and having three big classes of variables (detected, undetected, all) that the number grew so much..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452090,
          "author_name": "yangddd",
          "author_url": "",
          "post_date": "01/08/2019 07:13:06",
          "content": "<p>thanks for sharing, but I wonder how do you decide which 200 features to use ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 495584,
          "author_name": "agarreta",
          "author_url": "",
          "post_date": "03/21/2019 09:28:56",
          "content": "<p>Hi <a href=\"/yangddd\">@yangddd</a>, sorry for the late reply</p>\n\n<p>I iteratively used lgb feature importance assessment. More precisely, I started by running lgb with all features and afterwards selecting the top 50% best ranked features. Then I restarted the process with this selected features, and so on. I stopped approximately when my CV score stopped improving and when I suspected/verified that the improvement would not generalize to the LB</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "442778": "Thanks to everybody for this competition. Besides enjoying it greatly it has been extremely instructive for me. Thanks also to my teammates [Siddhartha][1] and [Ynktk][2] for their great spirit, hard work, and ideas. Finally, a big thanks to those who have shared their kernels and their thoughts in the discussion board throughout the competition. The open flow of ideas enrichened the competition invaluably.\n\nOur solution relied primarily on careful feature engineering and ensembling. Below I present an overview of it. \n\n**Stacking**\nOur best submissions were obtained by stacking the predictions of different lgb, catboost, and nn models, and then taking a weighted average with our previous best submission. This last step improved the score by 0.02 with respect to submitting the meta-model predictions alone (for our last submissions).\n\nAs meta-model we initially used a shallow lgbm with depth 1 and learning rate 0.01. At the end of the competition we increased the depth to 3. \n\n**First-level models** \nMy teammate Siddhartha focused mostly on a densely connected CNN model. You can read more about his ideas [here][3]. Ynktk created a variety of nn, xgb, lgb and catboost models. He will add some notes on these later in this thread. We also experimented a little with regularized greedy forests but they did not add much to the ensemble. \n\nPersonally I focused on a single first-level lgb model. This was split into two: one was trained on galactic objects, and the other on extragalactic ones. Each training set used different features. This was the model with best lb score (0.863) of our ensemble as far as we know.\n\n**Features** \nI engineered around 8000 features and used lgb importance to select the top 80 for the galactic set and the top 130 for the extragalactic set (approx). Some notes on this part:\n\n - First of all I would like to point to @manugangler's [kernel][5], where he provided a method for fitting light curves to microlensing events. The features obtained from there worked wonderfully for us, especially after taking passband-wise ratios and differences for the microamp and microbase features, respectively. Einstein time is the most important feature of my model.\n\n - I browsed and read some texts on astrophysics which gave me some ideas for feature engineering. Below I comment the most useful ones for my model. \n\n  - `distmod/log10(hostgal_photoz)`. If I understood correctly, this roughly measures how far from being a so-called 'standard candle' the object is (setting standard candles to be 'standard' supernovae). The actual formula is more complicated but for a tree-based approach this suffices. This is the 2nd or 3rd most important feature of my model.\n  \n  - `log_10( luminosity_mean )` where `luminosity = 4pi * (distance in parsecs) * flux` (I may be very confused and this may have nothing to do with luminosity). Some similar feats explained by Kyle Boone like magnitude `-2.5log_10(max flux in passband i) - distmod` and their ratios. There are several other feats. If someone is interested I can explain further.\n\n - I used different sets of time observations for feature agreggation. One set used all observations, the other only the ones marked as detected, and the other only the ones marked as undetected. After aggregation I took ratios passband-wise and also I created some other ratio variables such as `passband_std_undetected/mjd_diff_detected` and many others.\n\n**Some notes**\n\n- Siddhartha augmented the train set by randomly selecting from 30% to 70% time observations for each object id, and then creating new objects with the selected data. This improved his nn cv score by 0.01. It did not seem to improve my lgbm. Looking now at other solutions it may have been a good idea to spend some more time thinking on this idea.\n\n- Class 99: we used Olivier's and Cpmp's method throughout the competition. The less than 5 submissions we used for probing this class were unsuccessful.\n\n- We tried to model hostgal_specz but this did not work for us.\n\n- Pseudolabeling also did not work.\n\n- Our public lb top score was 0.792. However the best score obtained by one of our single first-level models was 0.863. Hence I believe one of the strong aspects of our solution is the diversity of models constructed, which may have been achieved thanks to the fact that each team member used mostly his own features and set up. These were quite different between the three of us. \n\nThis is all. Thank you for reading! If you have any question I will be pleased to answer\n\n\n  [1]: https://www.kaggle.com/meaninglesslives\n  [2]: https://www.kaggle.com/naka2ka\n  [3]: https://www.kaggle.com/meaninglesslives/a-slightly-better-nn-arch-and-some-tricks\n  [5]: https://www.kaggle.com/manugangler/optimal-feature-extraction-for-class-6",
    "442811": "Congrats on the result, and thanks for sharing.",
    "442918": "Congrats ! Glad to see that my features were useful.",
    "444374": "Zorionak!!! 8000 features?? Only Bilbao's people can get something like that ;-)",
    "444555": "Eskerrik asko! Good one :)\n\nIn fact at some point I got to up to 20000 features but only found new overfitting feats, so I went back to the 8000 set (I only used around 200 of these for training though).\n\nIt is mainly due to taking passband-wise ratios and having three big classes of variables (detected, undetected, all) that the number grew so much..",
    "452090": "thanks for sharing, but I wonder how do you decide which 200 features to use ?",
    "495584": "Hi @yangddd, sorry for the late reply\n\nI iteratively used lgb feature importance assessment. More precisely, I started by running lgb with all features and afterwards selecting the top 50% best ranked features. Then I restarted the process with this selected features, and so on. I stopped approximately when my CV score stopped improving and when I suspected/verified that the improvement would not generalize to the LB"
  },
  "source": "meta"
}