{
  "id": 75213,
  "title": "A solution and some learnings",
  "url": "/competitions/PLAsTiCC-2018/writeups/helgi-a-solution-and-some-learnings",
  "author_name": "",
  "post_date": "2018-12-19T13:26:08.442932800Z",
  "votes": 11,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First I want to congratulate the winners and thank the organizers for a fantastic competition. It is great to have the scientific background of the competition. I feel like I just started and hope there will be a continuation.\nI have read through the descriptions of the top solution. There were parts that I had included, others I had worked on but not managed to complete, as well as ideas I had not thought of. I would like to give a (hopefully) short description of the solution and what was missing.  </p>\n\n<ol>\n<li>Data normalization. \nFor data normalization I used features luminosity = flux*distmod**2 and magnitude = -2.5 log10(flux) – distmod (+ 27) for extra galactic objects. I did not use the normalization as in <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75054\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75054</a> I would like to try that out.  </li>\n<li>Initial Models.\nI split the models early into galactic and extra galactic. The initial models were NN, catboost, lgb and xbg models. \nFor NN I started with the model given by <a href=\"https://www.kaggle.com/meaninglesslives/simple-neural-net-for-time-series-classification\">https://www.kaggle.com/meaninglesslives/simple-neural-net-for-time-series-classification</a>\nbut changed the features to luminosity and added 1-3 layers with a tower of 3 CNNs(64/32/16), Dropout and MaxPooling/AveragePooling for processing a group of features for the 6 bands. \nThis was followed by a sequence of Dense layers. This was a simple NN that gave good initial results. I was not able to improve this essentially, but used it as input for stacking and averaging. \nI used catboost, lgb and xgb with depth 3, initially with 5 folds and later with 10 folds. The best results were initially with catboost followed by lgb, but xgb results were worse. Of course that required much feature engineering and took some time to pass the simple NN. When using NN as input for stacking the results of lgb became best and xgb were good enough to use in weighted averaging. \nThis was then the final Architecture with base models with NNs, catboost and lbg, and lgb and xgb models using stacking with NN and catboost results as input. This was then followed by weighted averaging including the base models. </li>\n<li>Feature Engineering and PCA\nThere are extremely many possibilities of creating features. The most difficult ones are for differentiating the hard classes that seem to be different types of supernovae. For this I created, as many others <br>\n•   Features restricted to some areas: detected area between first and last detected = 1; area where magnitude is below a limit; a growing/falling area prior/after maximal flux but within detected area, dividing by maximum to use ratios, and including counts to help the models <br>\n•   Taking differences between steps and divide by mjd difference to have derivatives and statistics based on the derivatives. This is to get the slope during the growth and falling of luminosity as well as the changes in the slope. <br>\n•   Taking average between steps and multiply by mjd difference for summing up to get integrals of the luminosity over the growing and falling area. This was followed by normalization to get a description of the shape. <br>\n•   Additional features such as A, eta described in <a href=\"https://arxiv.org/pdf/1512.01611.pdf\">https://arxiv.org/pdf/1512.01611.pdf</a> <br>\n•   autocorrelation__lag_1, fft_coefficients, entropy, skewness (but never kurtosis) (using tsfresh) and stetson_k <br>\nWhen adding such features to models this often lead to worse results. For handling that I included the use of PCA on the 6 bands to reduce the dimensionality of such features from 6 to a lower dimension like 1, 2 or 3.  This made it possible to use much more features and to use them together. I have not seen this use of PCA mentioned here before. </li>\n<li>What failed or could have been better\nThe part that I worked on and did not manage to complete successfully was curve fitting and use of Bayesian techniques. For curve fit I used Bazin. That gave extremely good results in CV, but the initial features gave bad results on LB. I was not able to complete this together with Baysian techniques in time. The main difficulty was the difference between train and test set. Here I did not manage to setup a reasonable process including augmentation. This was a great part of the solution by Kyle Boone. I think it is important to identify such essentials early enough to focus on the right parts.  That will be my main follow-ups including complete the missing parts of the solution (and get experience in using Gaussian Processes).</li>\n</ol>",
  "messages": [
    {
      "id": "442091",
      "postDate": "12/19/2018 13:26:08",
      "content": "<p>First I want to congratulate the winners and thank the organizers for a fantastic competition. It is great to have the scientific background of the competition. I feel like I just started and hope there will be a continuation.\nI have read through the descriptions of the top solution. There were parts that I had included, others I had worked on but not managed to complete, as well as ideas I had not thought of. I would like to give a (hopefully) short description of the solution and what was missing.  </p>\n\n<ol>\n<li>Data normalization. \nFor data normalization I used features luminosity = flux*distmod**2 and magnitude = -2.5 log10(flux) – distmod (+ 27) for extra galactic objects. I did not use the normalization as in <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75054\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75054</a> I would like to try that out.  </li>\n<li>Initial Models.\nI split the models early into galactic and extra galactic. The initial models were NN, catboost, lgb and xbg models. \nFor NN I started with the model given by <a href=\"https://www.kaggle.com/meaninglesslives/simple-neural-net-for-time-series-classification\">https://www.kaggle.com/meaninglesslives/simple-neural-net-for-time-series-classification</a>\nbut changed the features to luminosity and added 1-3 layers with a tower of 3 CNNs(64/32/16), Dropout and MaxPooling/AveragePooling for processing a group of features for the 6 bands. \nThis was followed by a sequence of Dense layers. This was a simple NN that gave good initial results. I was not able to improve this essentially, but used it as input for stacking and averaging. \nI used catboost, lgb and xgb with depth 3, initially with 5 folds and later with 10 folds. The best results were initially with catboost followed by lgb, but xgb results were worse. Of course that required much feature engineering and took some time to pass the simple NN. When using NN as input for stacking the results of lgb became best and xgb were good enough to use in weighted averaging. \nThis was then the final Architecture with base models with NNs, catboost and lbg, and lgb and xgb models using stacking with NN and catboost results as input. This was then followed by weighted averaging including the base models. </li>\n<li>Feature Engineering and PCA\nThere are extremely many possibilities of creating features. The most difficult ones are for differentiating the hard classes that seem to be different types of supernovae. For this I created, as many others <br>\n•   Features restricted to some areas: detected area between first and last detected = 1; area where magnitude is below a limit; a growing/falling area prior/after maximal flux but within detected area, dividing by maximum to use ratios, and including counts to help the models <br>\n•   Taking differences between steps and divide by mjd difference to have derivatives and statistics based on the derivatives. This is to get the slope during the growth and falling of luminosity as well as the changes in the slope. <br>\n•   Taking average between steps and multiply by mjd difference for summing up to get integrals of the luminosity over the growing and falling area. This was followed by normalization to get a description of the shape. <br>\n•   Additional features such as A, eta described in <a href=\"https://arxiv.org/pdf/1512.01611.pdf\">https://arxiv.org/pdf/1512.01611.pdf</a> <br>\n•   autocorrelation__lag_1, fft_coefficients, entropy, skewness (but never kurtosis) (using tsfresh) and stetson_k <br>\nWhen adding such features to models this often lead to worse results. For handling that I included the use of PCA on the 6 bands to reduce the dimensionality of such features from 6 to a lower dimension like 1, 2 or 3.  This made it possible to use much more features and to use them together. I have not seen this use of PCA mentioned here before. </li>\n<li>What failed or could have been better\nThe part that I worked on and did not manage to complete successfully was curve fitting and use of Bayesian techniques. For curve fit I used Bazin. That gave extremely good results in CV, but the initial features gave bad results on LB. I was not able to complete this together with Baysian techniques in time. The main difficulty was the difference between train and test set. Here I did not manage to setup a reasonable process including augmentation. This was a great part of the solution by Kyle Boone. I think it is important to identify such essentials early enough to focus on the right parts.  That will be my main follow-ups including complete the missing parts of the solution (and get experience in using Gaussian Processes).</li>\n</ol>",
      "rawMarkdown": "First I want to congratulate the winners and thank the organizers for a fantastic competition. It is great to have the scientific background of the competition. I feel like I just started and hope there will be a continuation.\nI have read through the descriptions of the top solution. There were parts that I had included, others I had worked on but not managed to complete, as well as ideas I had not thought of. I would like to give a (hopefully) short description of the solution and what was missing.  \n\n 1. Data normalization. \nFor data normalization I used features luminosity = flux*distmod**2 and magnitude = -2.5 log10(flux) – distmod (+ 27) for extra galactic objects. I did not use the normalization as in https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75054 I would like to try that out.  \n 2. Initial Models.\nI split the models early into galactic and extra galactic. The initial models were NN, catboost, lgb and xbg models. \nFor NN I started with the model given by https://www.kaggle.com/meaninglesslives/simple-neural-net-for-time-series-classification\nbut changed the features to luminosity and added 1-3 layers with a tower of 3 CNNs(64/32/16), Dropout and MaxPooling/AveragePooling for processing a group of features for the 6 bands. \nThis was followed by a sequence of Dense layers. This was a simple NN that gave good initial results. I was not able to improve this essentially, but used it as input for stacking and averaging. \nI used catboost, lgb and xgb with depth 3, initially with 5 folds and later with 10 folds. The best results were initially with catboost followed by lgb, but xgb results were worse. Of course that required much feature engineering and took some time to pass the simple NN. When using NN as input for stacking the results of lgb became best and xgb were good enough to use in weighted averaging. \nThis was then the final Architecture with base models with NNs, catboost and lbg, and lgb and xgb models using stacking with NN and catboost results as input. This was then followed by weighted averaging including the base models. \n 3. Feature Engineering and PCA\nThere are extremely many possibilities of creating features. The most difficult ones are for differentiating the hard classes that seem to be different types of supernovae. For this I created, as many others        \n•\tFeatures restricted to some areas: detected area between first and last detected = 1; area where magnitude is below a limit; a growing/falling area prior/after maximal flux but within detected area, dividing by maximum to use ratios, and including counts to help the models       \n•\tTaking differences between steps and divide by mjd difference to have derivatives and statistics based on the derivatives. This is to get the slope during the growth and falling of luminosity as well as the changes in the slope.     \n•\tTaking average between steps and multiply by mjd difference for summing up to get integrals of the luminosity over the growing and falling area. This was followed by normalization to get a description of the shape.      \n•\tAdditional features such as A, eta described in https://arxiv.org/pdf/1512.01611.pdf      \n•\tautocorrelation__lag_1, fft_coefficients, entropy, skewness (but never kurtosis) (using tsfresh) and stetson_k     \nWhen adding such features to models this often lead to worse results. For handling that I included the use of PCA on the 6 bands to reduce the dimensionality of such features from 6 to a lower dimension like 1, 2 or 3.  This made it possible to use much more features and to use them together. I have not seen this use of PCA mentioned here before. \n 5. What failed or could have been better\nThe part that I worked on and did not manage to complete successfully was curve fitting and use of Bayesian techniques. For curve fit I used Bazin. That gave extremely good results in CV, but the initial features gave bad results on LB. I was not able to complete this together with Baysian techniques in time. The main difficulty was the difference between train and test set. Here I did not manage to setup a reasonable process including augmentation. This was a great part of the solution by Kyle Boone. I think it is important to identify such essentials early enough to focus on the right parts.  That will be my main follow-ups including complete the missing parts of the solution (and get experience in using Gaussian Processes).",
      "votes": null
    },
    {
      "id": "442110",
      "postDate": "12/19/2018 13:53:38",
      "content": "<p>Congrats on the result, and thanks for sharing.  It is interesting to see how every good solution contains bits not used by others.</p>",
      "rawMarkdown": "Congrats on the result, and thanks for sharing.  It is interesting to see how every good solution contains bits not used by others.",
      "votes": null
    },
    {
      "id": "442544",
      "postDate": "12/20/2018 05:12:37",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "votes": null
    },
    {
      "id": "442677",
      "postDate": "12/20/2018 10:13:18",
      "content": "<p>Nice idea to use PCA to reduce bands before adding features, I will note it down for future use. Thanks for sharing!</p>",
      "rawMarkdown": "Nice idea to use PCA to reduce bands before adding features, I will note it down for future use. Thanks for sharing!",
      "votes": null
    },
    {
      "id": "442783",
      "postDate": "12/20/2018 13:53:43",
      "content": "<p>Congratulation  </p>",
      "rawMarkdown": "Congratulation",
      "votes": null
    },
    {
      "id": "757196",
      "postDate": "02/26/2020 14:21:33",
      "content": "<p><a href=\"/helgith\">@helgith</a> Hello, because I have not reached a level on kaggle, I cannot contact you or send a letter. I can only use this method to leave a message in your discussion. I am a student studying for a master's degree. My research is The big energy prediction of the ASHRAE competition, I think the competition is over now, so I would like to take this opportunity to ask you to use the lgbm model for prediction? How to adjust the parameters? Thank you!</p>",
      "rawMarkdown": "helgith Hello, because I have not reached a level on kaggle, I cannot contact you or send a letter. I can only use this method to leave a message in your discussion. I am a student studying for a master's degree. My research is The big energy prediction of the ASHRAE competition, I think the competition is over now, so I would like to take this opportunity to ask you to use the lgbm model for prediction? How to adjust the parameters? Thank you!",
      "votes": null
    },
    {
      "id": "757890",
      "postDate": "02/27/2020 08:05:43",
      "content": "<p>Hi James, I do not quite understand your request. (I have not prepared my models for releasing them on Kaggle for use by others)</p>",
      "rawMarkdown": "Hi James, I do not quite understand your request. (I have not prepared my models for releasing them on Kaggle for use by others)",
      "votes": null
    },
    {
      "id": "758773",
      "postDate": "02/28/2020 06:10:11",
      "content": "<p><a href=\"/helgith\">@helgith</a> Because I saw you in the ASHRAE competition in the top ranking, I would like to ask whether the model you are using is Lightgbm, and if so, may I ask you how to adjust the hyperparameters or the configuration of hyperparameters, because after all the competition ended. : )</p>",
      "rawMarkdown": "helgith Because I saw you in the ASHRAE competition in the top ranking, I would like to ask whether the model you are using is Lightgbm, and if so, may I ask you how to adjust the hyperparameters or the configuration of hyperparameters, because after all the competition ended. : )",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 442110,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/19/2018 13:53:38",
      "content": "<p>Congrats on the result, and thanks for sharing.  It is interesting to see how every good solution contains bits not used by others.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 442544,
      "author_name": "amar391",
      "author_url": "",
      "post_date": "12/20/2018 05:12:37",
      "content": "<p>Thank you for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 442677,
      "author_name": "taniaj",
      "author_url": "",
      "post_date": "12/20/2018 10:13:18",
      "content": "<p>Nice idea to use PCA to reduce bands before adding features, I will note it down for future use. Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 442783,
      "author_name": "avinashtayde",
      "author_url": "",
      "post_date": "12/20/2018 13:53:43",
      "content": "<p>Congratulation  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 757196,
      "author_name": "m0718993",
      "author_url": "",
      "post_date": "02/26/2020 14:21:33",
      "content": "<p><a href=\"/helgith\">@helgith</a> Hello, because I have not reached a level on kaggle, I cannot contact you or send a letter. I can only use this method to leave a message in your discussion. I am a student studying for a master's degree. My research is The big energy prediction of the ASHRAE competition, I think the competition is over now, so I would like to take this opportunity to ask you to use the lgbm model for prediction? How to adjust the parameters? Thank you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 757890,
          "author_name": "helgith",
          "author_url": "",
          "post_date": "02/27/2020 08:05:43",
          "content": "<p>Hi James, I do not quite understand your request. (I have not prepared my models for releasing them on Kaggle for use by others)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758773,
          "author_name": "m0718993",
          "author_url": "",
          "post_date": "02/28/2020 06:10:11",
          "content": "<p><a href=\"/helgith\">@helgith</a> Because I saw you in the ASHRAE competition in the top ranking, I would like to ask whether the model you are using is Lightgbm, and if so, may I ask you how to adjust the hyperparameters or the configuration of hyperparameters, because after all the competition ended. : )</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "442091": "First I want to congratulate the winners and thank the organizers for a fantastic competition. It is great to have the scientific background of the competition. I feel like I just started and hope there will be a continuation.\nI have read through the descriptions of the top solution. There were parts that I had included, others I had worked on but not managed to complete, as well as ideas I had not thought of. I would like to give a (hopefully) short description of the solution and what was missing.  \n\n 1. Data normalization. \nFor data normalization I used features luminosity = flux*distmod**2 and magnitude = -2.5 log10(flux) – distmod (+ 27) for extra galactic objects. I did not use the normalization as in https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75054 I would like to try that out.  \n 2. Initial Models.\nI split the models early into galactic and extra galactic. The initial models were NN, catboost, lgb and xbg models. \nFor NN I started with the model given by https://www.kaggle.com/meaninglesslives/simple-neural-net-for-time-series-classification\nbut changed the features to luminosity and added 1-3 layers with a tower of 3 CNNs(64/32/16), Dropout and MaxPooling/AveragePooling for processing a group of features for the 6 bands. \nThis was followed by a sequence of Dense layers. This was a simple NN that gave good initial results. I was not able to improve this essentially, but used it as input for stacking and averaging. \nI used catboost, lgb and xgb with depth 3, initially with 5 folds and later with 10 folds. The best results were initially with catboost followed by lgb, but xgb results were worse. Of course that required much feature engineering and took some time to pass the simple NN. When using NN as input for stacking the results of lgb became best and xgb were good enough to use in weighted averaging. \nThis was then the final Architecture with base models with NNs, catboost and lbg, and lgb and xgb models using stacking with NN and catboost results as input. This was then followed by weighted averaging including the base models. \n 3. Feature Engineering and PCA\nThere are extremely many possibilities of creating features. The most difficult ones are for differentiating the hard classes that seem to be different types of supernovae. For this I created, as many others        \n•\tFeatures restricted to some areas: detected area between first and last detected = 1; area where magnitude is below a limit; a growing/falling area prior/after maximal flux but within detected area, dividing by maximum to use ratios, and including counts to help the models       \n•\tTaking differences between steps and divide by mjd difference to have derivatives and statistics based on the derivatives. This is to get the slope during the growth and falling of luminosity as well as the changes in the slope.     \n•\tTaking average between steps and multiply by mjd difference for summing up to get integrals of the luminosity over the growing and falling area. This was followed by normalization to get a description of the shape.      \n•\tAdditional features such as A, eta described in https://arxiv.org/pdf/1512.01611.pdf      \n•\tautocorrelation__lag_1, fft_coefficients, entropy, skewness (but never kurtosis) (using tsfresh) and stetson_k     \nWhen adding such features to models this often lead to worse results. For handling that I included the use of PCA on the 6 bands to reduce the dimensionality of such features from 6 to a lower dimension like 1, 2 or 3.  This made it possible to use much more features and to use them together. I have not seen this use of PCA mentioned here before. \n 5. What failed or could have been better\nThe part that I worked on and did not manage to complete successfully was curve fitting and use of Bayesian techniques. For curve fit I used Bazin. That gave extremely good results in CV, but the initial features gave bad results on LB. I was not able to complete this together with Baysian techniques in time. The main difficulty was the difference between train and test set. Here I did not manage to setup a reasonable process including augmentation. This was a great part of the solution by Kyle Boone. I think it is important to identify such essentials early enough to focus on the right parts.  That will be my main follow-ups including complete the missing parts of the solution (and get experience in using Gaussian Processes).",
    "442110": "Congrats on the result, and thanks for sharing.  It is interesting to see how every good solution contains bits not used by others.",
    "442544": "Thank you for sharing",
    "442677": "Nice idea to use PCA to reduce bands before adding features, I will note it down for future use. Thanks for sharing!",
    "442783": "Congratulation",
    "757196": "helgith Hello, because I have not reached a level on kaggle, I cannot contact you or send a letter. I can only use this method to leave a message in your discussion. I am a student studying for a master's degree. My research is The big energy prediction of the ASHRAE competition, I think the competition is over now, so I would like to take this opportunity to ask you to use the lgbm model for prediction? How to adjust the parameters? Thank you!",
    "757890": "Hi James, I do not quite understand your request. (I have not prepared my models for releasing them on Kaggle for use by others)",
    "758773": "helgith Because I saw you in the ASHRAE competition in the top ranking, I would like to ask whether the model you are using is Lightgbm, and if so, may I ask you how to adjust the hyperparameters or the configuration of hyperparameters, because after all the competition ended. : )"
  },
  "source": "meta"
}