{
  "id": 75054,
  "title": "14th place solution",
  "url": "/competitions/PLAsTiCC-2018/discussion/75054",
  "author_name": "Belinda Trotta",
  "post_date": "2018-12-18T07:18:14.113000",
  "votes": 72,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Code and full writeup (see the pdf in the github repo).  <a href=\"https://github.com/btrotta/kaggle-plasticc\">https://github.com/btrotta/kaggle-plasticc</a></p>\n\n<p><strong>Summary</strong></p>\n\n<p>My solution is implemented in Python and uses LightGBM gradient boosted classification tree models. It scores 0.84070 on the private leaderboard, and runs in around 6 hours on a 24 Gb laptop (including calculating features, training, and prediction). It uses only elementary operations to calculate the features: there's no curve fitting or optimisation, which helps keep the runtime down. Apart from the hints revealed in the forum discussions, my original insights that gave the most improvement in score are: </p>\n\n<ul>\n<li>Bayesian approach to removing noise from the flux measurements. Replace the flux with \n<code>\n(flux / flux_err**2 + flux_mean / flux_std ** 2) / (1 / flux_err**2 + 1 / flux_std **2)\n</code>\nwhere <code>flux_mean</code> and <code>flux_std</code> are the mean and standard deviation for the given object and passband. (Detailed justification for this calculation is in the pdf).</li>\n<li>Adding features based on scaled flux values. As well as a lot of features caclulated on the raw flux (after removing noise), I also calculated a scaled version of the flux by dividing by the maximum absolute value for each object and passband. I think this helps because it captures the shape of the curve, normalising for magnitude.</li>\n<li>Adding features to capture the behaviour around the peak.  Since these peaks can occur at any time during the observation period, we need a way to extract the data from these peaks. I did this by finding the \"most extreme\" minimum and maximum times for each object, defined as the time with detection=1 when one of the object's passbands differs the most from its median value. Then, for each passband, I retrieve the closest flux value before and after this peak time. I do this for both the raw (but de-noised) flux and the max-scaled flux. Also, I calculate the \"duration\" of the peak in each passband, defined as the period of continuous detections around the peak, and the time difference between the overall peak and the individual passband peaks.</li>\n<li>Understanding how to optimise the metric (including for class 99). (Basically you can just optimise normal log loss, then multiply the estimated probabilities by W/class frequency.)</li>\n</ul>",
  "messages": [
    {
      "id": 441011,
      "postDate": "2018-12-18T07:18:14.113Z",
      "content": "<p>Code and full writeup (see the pdf in the github repo).  <a href=\"https://github.com/btrotta/kaggle-plasticc\">https://github.com/btrotta/kaggle-plasticc</a></p>\n\n<p><strong>Summary</strong></p>\n\n<p>My solution is implemented in Python and uses LightGBM gradient boosted classification tree models. It scores 0.84070 on the private leaderboard, and runs in around 6 hours on a 24 Gb laptop (including calculating features, training, and prediction). It uses only elementary operations to calculate the features: there's no curve fitting or optimisation, which helps keep the runtime down. Apart from the hints revealed in the forum discussions, my original insights that gave the most improvement in score are: </p>\n\n<ul>\n<li>Bayesian approach to removing noise from the flux measurements. Replace the flux with \n<code>\n(flux / flux_err**2 + flux_mean / flux_std ** 2) / (1 / flux_err**2 + 1 / flux_std **2)\n</code>\nwhere <code>flux_mean</code> and <code>flux_std</code> are the mean and standard deviation for the given object and passband. (Detailed justification for this calculation is in the pdf).</li>\n<li>Adding features based on scaled flux values. As well as a lot of features caclulated on the raw flux (after removing noise), I also calculated a scaled version of the flux by dividing by the maximum absolute value for each object and passband. I think this helps because it captures the shape of the curve, normalising for magnitude.</li>\n<li>Adding features to capture the behaviour around the peak.  Since these peaks can occur at any time during the observation period, we need a way to extract the data from these peaks. I did this by finding the \"most extreme\" minimum and maximum times for each object, defined as the time with detection=1 when one of the object's passbands differs the most from its median value. Then, for each passband, I retrieve the closest flux value before and after this peak time. I do this for both the raw (but de-noised) flux and the max-scaled flux. Also, I calculate the \"duration\" of the peak in each passband, defined as the period of continuous detections around the peak, and the time difference between the overall peak and the individual passband peaks.</li>\n<li>Understanding how to optimise the metric (including for class 99). (Basically you can just optimise normal log loss, then multiply the estimated probabilities by W/class frequency.)</li>\n</ul>",
      "rawMarkdown": "Code and full writeup (see the pdf in the github repo).  https://github.com/btrotta/kaggle-plasticc\n\n**Summary**\n\nMy solution is implemented in Python and uses LightGBM gradient boosted classification tree models. It scores 0.84070 on the private leaderboard, and runs in around 6 hours on a 24 Gb laptop (including calculating features, training, and prediction). It uses only elementary operations to calculate the features: there's no curve fitting or optimisation, which helps keep the runtime down. Apart from the hints revealed in the forum discussions, my original insights that gave the most improvement in score are: \n\n- Bayesian approach to removing noise from the flux measurements. Replace the flux with \n```\n (flux / flux_err**2 + flux_mean / flux_std ** 2) / (1 / flux_err**2 + 1 / flux_std **2)\n```\nwhere `flux_mean` and `flux_std` are the mean and standard deviation for the given object and passband. (Detailed justification for this calculation is in the pdf).\n- Adding features based on scaled flux values. As well as a lot of features caclulated on the raw flux (after removing noise), I also calculated a scaled version of the flux by dividing by the maximum absolute value for each object and passband. I think this helps because it captures the shape of the curve, normalising for magnitude.\n- Adding features to capture the behaviour around the peak.  Since these peaks can occur at any time during the observation period, we need a way to extract the data from these peaks. I did this by finding the \"most extreme\" minimum and maximum times for each object, defined as the time with detection=1 when one of the object's passbands differs the most from its median value. Then, for each passband, I retrieve the closest flux value before and after this peak time. I do this for both the raw (but de-noised) flux and the max-scaled flux. Also, I calculate the \"duration\" of the peak in each passband, defined as the period of continuous detections around the peak, and the time difference between the overall peak and the individual passband peaks.\n- Understanding how to optimise the metric (including for class 99). (Basically you can just optimise normal log loss, then multiply the estimated probabilities by W/class frequency.)\n",
      "votes": 72
    },
    {
      "id": 441386,
      "postDate": "2018-12-18T15:55:00.627Z",
      "content": "<p>Congratulations and thanks for making it clear. \nI did try using your Bayesian approach (sorry for stalking your github, coz. I was curious to know how you were completely solo so long with the others at the top). But, for me it didn't seem to work at that time and the features I derived from the modified fluxes were not much different from the original ones. So, I didn't bother to train on them. \nIn any case, I was eagerly looking forward to your solution once the competition ends, and expected that your solution would involve Bayesian approach. Just wanted to know, how you would use it.</p>",
      "rawMarkdown": "Congratulations and thanks for making it clear. \nI did try using your Bayesian approach (sorry for stalking your github, coz. I was curious to know how you were completely solo so long with the others at the top). But, for me it didn't seem to work at that time and the features I derived from the modified fluxes were not much different from the original ones. So, I didn't bother to train on them. \nIn any case, I was eagerly looking forward to your solution once the competition ends, and expected that your solution would involve Bayesian approach. Just wanted to know, how you would use it.\n",
      "votes": 1
    },
    {
      "id": 441371,
      "postDate": "2018-12-18T15:37:22.893Z",
      "content": "<p>Congratulations! Clearly explained, simple and efficient solution!</p>",
      "rawMarkdown": "Congratulations! Clearly explained, simple and efficient solution!",
      "votes": 1
    },
    {
      "id": 441208,
      "postDate": "2018-12-18T12:08:29.133Z",
      "content": "<p>Congratulations Belinda for great solo finish,  Thanks for sharing your work, I was curious to learn as you have been in gold zone for a long time.</p>",
      "rawMarkdown": "Congratulations Belinda for great solo finish,  Thanks for sharing your work, I was curious to learn as you have been in gold zone for a long time.",
      "votes": 1
    },
    {
      "id": 441187,
      "postDate": "2018-12-18T11:37:12.903Z",
      "content": "<p>Thank you sharing clean code and detailed write-up. Congrats!</p>",
      "rawMarkdown": "Thank you sharing clean code and detailed write-up. Congrats!",
      "votes": 1
    },
    {
      "id": 441178,
      "postDate": "2018-12-18T11:28:01.660Z",
      "content": "<p>Congrats! Clean and fast solution</p>",
      "rawMarkdown": "Congrats! Clean and fast solution",
      "votes": 1
    },
    {
      "id": 441170,
      "postDate": "2018-12-18T11:10:53.837Z",
      "content": "<p>Good job, Belinda! Congrats!!!</p>",
      "rawMarkdown": "Good job, Belinda! Congrats!!!",
      "votes": 1
    },
    {
      "id": 441035,
      "postDate": "2018-12-18T07:38:44.397Z",
      "content": "<p>Congratulations Belinda ! Thanks for sharing your work and your nice write-up. That was fast !</p>",
      "rawMarkdown": "Congratulations Belinda ! Thanks for sharing your work and your nice write-up. That was fast !",
      "votes": 1
    },
    {
      "id": 441404,
      "postDate": "2018-12-18T16:13:53.840Z",
      "content": "<p>Congrats! very impressive and elegant approach. I really should have paid more attention to the stats class. </p>\n\n<p>Hope in future we could have a chance to team up. : )</p>",
      "rawMarkdown": "Congrats! very impressive and elegant approach. I really should have paid more attention to the stats class. \n\nHope in future we could have a chance to team up. : )",
      "votes": 2
    },
    {
      "id": 444532,
      "postDate": "2018-12-24T07:48:02.733Z",
      "content": "<p>Congratulations! Thanks for sharing your elegant and impressive work! </p>",
      "rawMarkdown": "Congratulations! Thanks for sharing your elegant and impressive work! "
    },
    {
      "id": 443288,
      "postDate": "2018-12-21T11:06:41.837Z",
      "content": "<p>Congratulations! The Bayesian approach is really clever. Thank you for sharing!</p>",
      "rawMarkdown": "Congratulations! The Bayesian approach is really clever. Thank you for sharing!"
    },
    {
      "id": 442080,
      "postDate": "2018-12-19T13:09:09.930Z",
      "content": "<p>Congratulations and thanks for sharing your code. Your Bayesian noise removal is brilliant!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing your code. Your Bayesian noise removal is brilliant!"
    },
    {
      "id": 441660,
      "postDate": "2018-12-18T22:41:22.147Z",
      "content": "<p>Congratulations and thanks for sharing. It's really interesting to see how well a solution without curve fitting can do and I'll have to read up on the Bayesian de-noising for future work!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing. It's really interesting to see how well a solution without curve fitting can do and I'll have to read up on the Bayesian de-noising for future work!"
    },
    {
      "id": 441012,
      "postDate": "2018-12-18T07:22:14.613Z",
      "content": "<p>Congrats for the performance, I was curious to leanr about what you did, but, if I may, it is a pity you make reading your writeup a bit of a burden.  Can't you paste it here?  Maybe you are so used to Latex you cannot write plain text anymore? ;)</p>",
      "rawMarkdown": "Congrats for the performance, I was curious to leanr about what you did, but, if I may, it is a pity you make reading your writeup a bit of a burden.  Can't you paste it here?  Maybe you are so used to Latex you cannot write plain text anymore? ;)",
      "replies": [
        {
          "id": 441038,
          "postDate": "2018-12-18T07:43:09.163Z",
          "content": "<p>Thanks, added a short summary :)</p>",
          "rawMarkdown": "Thanks, added a short summary :)",
          "votes": 1
        },
        {
          "id": 441060,
          "postDate": "2018-12-18T08:15:14.217Z",
          "content": "<p>The Bayesian way of removing noise is very interesting!</p>",
          "rawMarkdown": "The Bayesian way of removing noise is very interesting!"
        },
        {
          "id": 441077,
          "postDate": "2018-12-18T08:33:21.677Z",
          "content": "<p>It gave me around 0.03 gain on leaderboard.</p>",
          "rawMarkdown": "It gave me around 0.03 gain on leaderboard.",
          "votes": 1
        },
        {
          "id": 441237,
          "postDate": "2018-12-18T12:53:37.890Z",
          "content": "<p>It's great.  I've learned something here.  I usually make my own rank of other people's solution by what and how much I learn from them, and with that you rank high ;)  My use of noise was to filter out values that had an error larger than 6 error sigma, as there are some wild outliers out there.  But this is way better.</p>",
          "rawMarkdown": "It's great.  I've learned something here.  I usually make my own rank of other people's solution by what and how much I learn from them, and with that you rank high ;)  My use of noise was to filter out values that had an error larger than 6 error sigma, as there are some wild outliers out there.  But this is way better.",
          "votes": 2
        }
      ]
    },
    {
      "id": 442418,
      "postDate": "2018-12-19T23:53:41.410Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 448552,
      "postDate": "2019-01-01T13:27:32.803Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    },
    {
      "id": 441881,
      "postDate": "2018-12-19T07:39:40.147Z",
      "content": "<p>Thanks for sharing! It's amazing!</p>",
      "rawMarkdown": "Thanks for sharing! It's amazing!"
    }
  ],
  "comments": [
    {
      "id": 441386,
      "author_name": "Vig",
      "author_url": "",
      "post_date": "2018-12-18T15:55:00.627000",
      "content": "<p>Congratulations and thanks for making it clear. \nI did try using your Bayesian approach (sorry for stalking your github, coz. I was curious to know how you were completely solo so long with the others at the top). But, for me it didn't seem to work at that time and the features I derived from the modified fluxes were not much different from the original ones. So, I didn't bother to train on them. \nIn any case, I was eagerly looking forward to your solution once the competition ends, and expected that your solution would involve Bayesian approach. Just wanted to know, how you would use it.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 441371,
      "author_name": "Gabriel Preda",
      "author_url": "",
      "post_date": "2018-12-18T15:37:22.893000",
      "content": "<p>Congratulations! Clearly explained, simple and efficient solution!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 441208,
      "author_name": "Chinta",
      "author_url": "",
      "post_date": "2018-12-18T12:08:29.133000",
      "content": "<p>Congratulations Belinda for great solo finish,  Thanks for sharing your work, I was curious to learn as you have been in gold zone for a long time.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 441187,
      "author_name": "iprapas",
      "author_url": "",
      "post_date": "2018-12-18T11:37:12.903000",
      "content": "<p>Thank you sharing clean code and detailed write-up. Congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 441178,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2018-12-18T11:28:01.660000",
      "content": "<p>Congrats! Clean and fast solution</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 441170,
      "author_name": "joxemi",
      "author_url": "",
      "post_date": "2018-12-18T11:10:53.837000",
      "content": "<p>Good job, Belinda! Congrats!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 441035,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2018-12-18T07:38:44.397000",
      "content": "<p>Congratulations Belinda ! Thanks for sharing your work and your nice write-up. That was fast !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 441404,
      "author_name": "Jiwei Liu",
      "author_url": "",
      "post_date": "2018-12-18T16:13:53.840000",
      "content": "<p>Congrats! very impressive and elegant approach. I really should have paid more attention to the stats class. </p>\n\n<p>Hope in future we could have a chance to team up. : )</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 444532,
      "author_name": "Alexpartisan",
      "author_url": "",
      "post_date": "2018-12-24T07:48:02.733000",
      "content": "<p>Congratulations! Thanks for sharing your elegant and impressive work! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 443288,
      "author_name": "Chirag",
      "author_url": "",
      "post_date": "2018-12-21T11:06:41.837000",
      "content": "<p>Congratulations! The Bayesian approach is really clever. Thank you for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 442080,
      "author_name": "Tania J",
      "author_url": "",
      "post_date": "2018-12-19T13:09:09.930000",
      "content": "<p>Congratulations and thanks for sharing your code. Your Bayesian noise removal is brilliant!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 441660,
      "author_name": "Andy Penrose",
      "author_url": "",
      "post_date": "2018-12-18T22:41:22.147000",
      "content": "<p>Congratulations and thanks for sharing. It's really interesting to see how well a solution without curve fitting can do and I'll have to read up on the Bayesian de-noising for future work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 441012,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-12-18T07:22:14.613000",
      "content": "<p>Congrats for the performance, I was curious to leanr about what you did, but, if I may, it is a pity you make reading your writeup a bit of a burden.  Can't you paste it here?  Maybe you are so used to Latex you cannot write plain text anymore? ;)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 441038,
          "author_name": "Belinda Trotta",
          "author_url": "",
          "post_date": "2018-12-18T07:43:09.163000",
          "content": "<p>Thanks, added a short summary :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 441060,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-18T08:15:14.217000",
          "content": "<p>The Bayesian way of removing noise is very interesting!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441077,
          "author_name": "Belinda Trotta",
          "author_url": "",
          "post_date": "2018-12-18T08:33:21.677000",
          "content": "<p>It gave me around 0.03 gain on leaderboard.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 441237,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-18T12:53:37.890000",
          "content": "<p>It's great.  I've learned something here.  I usually make my own rank of other people's solution by what and how much I learn from them, and with that you rank high ;)  My use of noise was to filter out values that had an error larger than 6 error sigma, as there are some wild outliers out there.  But this is way better.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 442418,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-19T23:53:41.410000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 448552,
      "author_name": "LongYin/杰少",
      "author_url": "",
      "post_date": "2019-01-01T13:27:32.803000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 441881,
      "author_name": "zephoony",
      "author_url": "",
      "post_date": "2018-12-19T07:39:40.147000",
      "content": "<p>Thanks for sharing! It's amazing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "441011": "Code and full writeup (see the pdf in the github repo).  https://github.com/btrotta/kaggle-plasticc\n\n**Summary**\n\nMy solution is implemented in Python and uses LightGBM gradient boosted classification tree models. It scores 0.84070 on the private leaderboard, and runs in around 6 hours on a 24 Gb laptop (including calculating features, training, and prediction). It uses only elementary operations to calculate the features: there's no curve fitting or optimisation, which helps keep the runtime down. Apart from the hints revealed in the forum discussions, my original insights that gave the most improvement in score are: \n\n- Bayesian approach to removing noise from the flux measurements. Replace the flux with \n```\n (flux / flux_err**2 + flux_mean / flux_std ** 2) / (1 / flux_err**2 + 1 / flux_std **2)\n```\nwhere `flux_mean` and `flux_std` are the mean and standard deviation for the given object and passband. (Detailed justification for this calculation is in the pdf).\n- Adding features based on scaled flux values. As well as a lot of features caclulated on the raw flux (after removing noise), I also calculated a scaled version of the flux by dividing by the maximum absolute value for each object and passband. I think this helps because it captures the shape of the curve, normalising for magnitude.\n- Adding features to capture the behaviour around the peak.  Since these peaks can occur at any time during the observation period, we need a way to extract the data from these peaks. I did this by finding the \"most extreme\" minimum and maximum times for each object, defined as the time with detection=1 when one of the object's passbands differs the most from its median value. Then, for each passband, I retrieve the closest flux value before and after this peak time. I do this for both the raw (but de-noised) flux and the max-scaled flux. Also, I calculate the \"duration\" of the peak in each passband, defined as the period of continuous detections around the peak, and the time difference between the overall peak and the individual passband peaks.\n- Understanding how to optimise the metric (including for class 99). (Basically you can just optimise normal log loss, then multiply the estimated probabilities by W/class frequency.)\n",
    "441386": "Congratulations and thanks for making it clear. \nI did try using your Bayesian approach (sorry for stalking your github, coz. I was curious to know how you were completely solo so long with the others at the top). But, for me it didn't seem to work at that time and the features I derived from the modified fluxes were not much different from the original ones. So, I didn't bother to train on them. \nIn any case, I was eagerly looking forward to your solution once the competition ends, and expected that your solution would involve Bayesian approach. Just wanted to know, how you would use it.\n",
    "441371": "Congratulations! Clearly explained, simple and efficient solution!",
    "441208": "Congratulations Belinda for great solo finish,  Thanks for sharing your work, I was curious to learn as you have been in gold zone for a long time.",
    "441187": "Thank you sharing clean code and detailed write-up. Congrats!",
    "441178": "Congrats! Clean and fast solution",
    "441170": "Good job, Belinda! Congrats!!!",
    "441035": "Congratulations Belinda ! Thanks for sharing your work and your nice write-up. That was fast !",
    "441404": "Congrats! very impressive and elegant approach. I really should have paid more attention to the stats class. \n\nHope in future we could have a chance to team up. : )",
    "444532": "Congratulations! Thanks for sharing your elegant and impressive work! ",
    "443288": "Congratulations! The Bayesian approach is really clever. Thank you for sharing!",
    "442080": "Congratulations and thanks for sharing your code. Your Bayesian noise removal is brilliant!",
    "441660": "Congratulations and thanks for sharing. It's really interesting to see how well a solution without curve fitting can do and I'll have to read up on the Bayesian de-noising for future work!",
    "441012": "Congrats for the performance, I was curious to leanr about what you did, but, if I may, it is a pity you make reading your writeup a bit of a burden.  Can't you paste it here?  Maybe you are so used to Latex you cannot write plain text anymore? ;)",
    "442418": "",
    "448552": "Thanks for sharing.",
    "441881": "Thanks for sharing! It's amazing!"
  }
}