{
  "id": 75131,
  "title": "3rd Place Part II",
  "url": "/competitions/PLAsTiCC-2018/writeups/major-tom-3rd-place-part-ii",
  "author_name": "",
  "post_date": "2021-05-19T08:12:46.947Z",
  "votes": 46,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Thank you my briliant team, all teams who competed with us, and all people who participated in this competition. Congrats Kyle Boone, who won this competition with an incredible performance! <br><br>\nI'm happy because I and yuval became kaggle expert with 2 gold medals in this competition, and We finally became Prize Winner :) <br><br>\nIt's my second kaggle competition and it was fun to compete with AhmetErdem and CPMP again, who I competed with in my previous competition, TalkingData Adtracking Fraud Detection. Here, I won't explain all of my team's solution, please take a look at <a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> and <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> solutions, if you want to see what my teammate has done. <br><br>\n<code>yuval's solution</code> : <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75116\" target=\"_blank\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75116</a><br>\n<code>nyanp's solution</code> : <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75222\" target=\"_blank\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75222</a><br>\n<code>My github for the host</code>: <a href=\"https://github.com/takashioya/plasticc\" target=\"_blank\">https://github.com/takashioya/plasticc</a></p>\n<h1>My Model</h1>\n<p>my model is based on CatBoost, with lotta feature engineering. <br>\nI made 3 types of models.<br>\n(1) galactic model, the number of features are about 1000.<br>\n(2) extra-galactic model, the number of features are about 2000. <br>\n(3) extra-galactic model with hostgal_specz, the number of features are about 2000. It gave 0.004 improvement.<br>\nThen, I trained 7 * 3 types models with different feature set and get simple average of them, which scored 0.806 on Public LB, while my best single model scored 0.823. <br> <br>\nyou may think the number of features of my model is too big, but using small number of features degraded the score. tbh I prefer to use big features, because I think big features lead to big improvement in ensemble, from my experience. </p>\n<h1>My Features</h1>\n<p>actually, I used nyanp's features for my model and some of them have much bigger importance than mine (especially salt-2 features). However, some of my features still have high importance, so I'll explain a bit about them. </p>\n<ul>\n<li><code>features related to mjd</code>: <br>\nsomething like 'variance of mjd whose detected == 1' worked well.</li>\n<li><code>abs_curve_angle_features</code>: <br>\nby calculating np.arctan(flux_diff/mjd_diff), we can calculate 'angle' of light curve. <br>\nI calculated many stats from this.</li>\n<li><code>feets features</code>: <br>\nI used <a href=\"https://github.com/carpyncho/feets\" target=\"_blank\">https://github.com/carpyncho/feets</a>, which was effective.</li>\n<li><code>gaussian process features</code>: <br>\nIt worked better than the features like 'days_from_peak_n_percent_flux'. <br>\nfor gaussian process, I used Gpy, but it's too slow and I <br>\nhad to use 30 instances for feature extraction.</li>\n<li><code>shifted_flux * distance ** 2 features</code>: <br>\nI calculated shifted_flux = flux - flux.min() if flux.min() &gt; 0 per each object_id and passband, which worked better than just calculating <code>flux * distance ** 2</code></li>\n</ul>\n<h1>Ensemble</h1>\n<p>I used weighted average for ensemble. however, I gave different weights to each class, which gave us 0.006 improvement. I calculated the weights using oof-predictions by hyperopt. <br><br>\nBy getting weighted average of mamas's CatBoost(0.806) and nyanp's LightGBM(0.834), we got 0.780 on public LB.Then, in the same way, we calculated weighted average of 0.780 model and yuval's NN (0.791) and got 0.737.</p>\n<h1>probing class99</h1>\n<p>Actually, we did nothing special about class99 7 days before the end of the competition. <br>\nWe noticed class99 probing is very effective 6 days before the end and spent 5 days probing class99 distribution. <br>\nFinally, we got 0.057 improvement by probing class99, which took us to the Prize zone (0.680 on public LB). <br><br>\nFor extra-galactic, we used this special formula. <br>\n<code>c99 = 1.65 * (c42 + c52 + c62 + c95) ** 2.5 *  (1 - c62 - c95) ** 0.635</code> <br><br>\nFor galactic, we did nothing special, just a normal method like this. <br>\n<code>c99 = 0.2 * (1 - c42) * (1 - c52) * .... * (1 - c90)</code><br><br>\nHere, I'll explain how we found this special formula.<br><br>\nFirst, we started with the normal method like this.<br>\n<code>c99 = 0.9 * (1 - c42) * (1 - c52) * .... * (1 - c90)</code> <br><br>\nThen, we noticed this way works better, which means c90 is not similar to c99 at all. <br>\n<code>c99 = 0.9 * (1 - c42) * (1 - c52) * .... * (1 - c90) ** 2</code> <br><br>\nThen, we also tried changing power for other classes and got this formula. (We spent about 19 subs for this step) <br>\n<code>powers = {c42 : 0, c52 : 0, c62 : 0.5, c95 : 1, c90 : 2, c88 : 2, c67 : 2, c64 : 2,  c15 : 2}</code><br>\n<code>c99 = np.prod([(1 - pred[c]) ** powers[c] for c in classes])</code> <br> <br>\nThis formula means c99 is very similar to c42 and c52, and somewhat similar to c62 and c95. actually, this formula already improved our score by 0.049 and it was the last day of the competition. but, we didn't give up and finally found the problem of this formula. <br>                  <br>\nFor object_id whose <code>c90 = 0.5</code> and <code>c88 = 0.5</code>, we thought we should assign <code>c99 = 0</code>.</p>\n<p>However, this formula assigns <code>c99 = 0.5 ** 2 * 0.5 ** 2 = 0.0625</code>, which looks too high for us. so, we thought we should do something like <code>c99 *= (1 - c90 - c88)</code> and finally found <br>\n<code>c99 = 1.65 * (1 - c90 - c88 - c67 - c64 - c15) ** 2 * (1 - c62 - c95) ** 0.635</code> <br>\n        <code>= 1.65 * (c42 + c52 + c62 + c95) ** 2 * (1 - c62 - c95) ** 0.635</code> <br>\nis a very good approximation of the formula. <br><br>\nActually, it was our last submission in this competition and we couldn't tune coefficient at all. <br></p>\n<h1>what didn't work for me</h1>\n<p>some of the experiments that didn't work (Not all)<br></p>\n<ul>\n<li><code>Band-correction using hostgal_photoz</code><br></li>\n<li><code>target encoding using ra, decl</code><br></li>\n<li><code>using adversarial validation for feature selection</code><br></li>\n<li><code>stacking using GBDT as meta model</code><br></li>\n<li><code>UNet-like Autoencoder for feature extraction</code><br></li>\n<li><code>make binary classification model and use oof-predictions as features</code><br></li>\n</ul>\n<h1>what worked</h1>\n<ul>\n<li><code>class90 pseudo-labelling</code>: <br>\nIt gave 0.016 improvements for me.<br>\n<br></li>\n<li><code>adversarial pseudo-labelling</code>: <br>\nI selected the objects whose the prediction is very high and similar to train data and used them for training. It didn't give a significant improvement for me, but I suppose it somewhat contributed to diversity in ensemble. </li>\n</ul>\n<h1>Comments</h1>\n<p>Actually, I haven't figured out why we dropped from 2nd to 3rd on Private LB.  I'm wondering whether it's because of the luck or not.</p>\n<p>As for the competition design, I think making 'class99' is not so good, because it leads to LB probing and is not so practical. However, on the whole, I suppose this competition is really interesting and one of the most successful competition in kaggle. I'm especially really impressed with the yuval's beautiful and practical NN modelling :)</p>\n<p>Anyway, I really enjoyed my second kaggle competition. I'll continue kaggle and want to become the winner in the next competition :) </p>\n<h1>Thanks all, see you again!</h1>",
  "messages": [
    {
      "id": "441520",
      "postDate": "12/18/2018 18:23:32",
      "content": "<p>Thank you my briliant team, all teams who competed with us, and all people who participated in this competition. Congrats Kyle Boone, who won this competition with an incredible performance! <br><br>\nI'm happy because I and yuval became kaggle expert with 2 gold medals in this competition, and We finally became Prize Winner :) <br><br>\nIt's my second kaggle competition and it was fun to compete with AhmetErdem and CPMP again, who I competed with in my previous competition, TalkingData Adtracking Fraud Detection. Here, I won't explain all of my team's solution, please take a look at <a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> and <a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a> solutions, if you want to see what my teammate has done. <br><br>\n<code>yuval's solution</code> : <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75116\" target=\"_blank\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75116</a><br>\n<code>nyanp's solution</code> : <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75222\" target=\"_blank\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75222</a><br>\n<code>My github for the host</code>: <a href=\"https://github.com/takashioya/plasticc\" target=\"_blank\">https://github.com/takashioya/plasticc</a></p>\n<h1>My Model</h1>\n<p>my model is based on CatBoost, with lotta feature engineering. <br>\nI made 3 types of models.<br>\n(1) galactic model, the number of features are about 1000.<br>\n(2) extra-galactic model, the number of features are about 2000. <br>\n(3) extra-galactic model with hostgal_specz, the number of features are about 2000. It gave 0.004 improvement.<br>\nThen, I trained 7 * 3 types models with different feature set and get simple average of them, which scored 0.806 on Public LB, while my best single model scored 0.823. <br> <br>\nyou may think the number of features of my model is too big, but using small number of features degraded the score. tbh I prefer to use big features, because I think big features lead to big improvement in ensemble, from my experience. </p>\n<h1>My Features</h1>\n<p>actually, I used nyanp's features for my model and some of them have much bigger importance than mine (especially salt-2 features). However, some of my features still have high importance, so I'll explain a bit about them. </p>\n<ul>\n<li><code>features related to mjd</code>: <br>\nsomething like 'variance of mjd whose detected == 1' worked well.</li>\n<li><code>abs_curve_angle_features</code>: <br>\nby calculating np.arctan(flux_diff/mjd_diff), we can calculate 'angle' of light curve. <br>\nI calculated many stats from this.</li>\n<li><code>feets features</code>: <br>\nI used <a href=\"https://github.com/carpyncho/feets\" target=\"_blank\">https://github.com/carpyncho/feets</a>, which was effective.</li>\n<li><code>gaussian process features</code>: <br>\nIt worked better than the features like 'days_from_peak_n_percent_flux'. <br>\nfor gaussian process, I used Gpy, but it's too slow and I <br>\nhad to use 30 instances for feature extraction.</li>\n<li><code>shifted_flux * distance ** 2 features</code>: <br>\nI calculated shifted_flux = flux - flux.min() if flux.min() &gt; 0 per each object_id and passband, which worked better than just calculating <code>flux * distance ** 2</code></li>\n</ul>\n<h1>Ensemble</h1>\n<p>I used weighted average for ensemble. however, I gave different weights to each class, which gave us 0.006 improvement. I calculated the weights using oof-predictions by hyperopt. <br><br>\nBy getting weighted average of mamas's CatBoost(0.806) and nyanp's LightGBM(0.834), we got 0.780 on public LB.Then, in the same way, we calculated weighted average of 0.780 model and yuval's NN (0.791) and got 0.737.</p>\n<h1>probing class99</h1>\n<p>Actually, we did nothing special about class99 7 days before the end of the competition. <br>\nWe noticed class99 probing is very effective 6 days before the end and spent 5 days probing class99 distribution. <br>\nFinally, we got 0.057 improvement by probing class99, which took us to the Prize zone (0.680 on public LB). <br><br>\nFor extra-galactic, we used this special formula. <br>\n<code>c99 = 1.65 * (c42 + c52 + c62 + c95) ** 2.5 *  (1 - c62 - c95) ** 0.635</code> <br><br>\nFor galactic, we did nothing special, just a normal method like this. <br>\n<code>c99 = 0.2 * (1 - c42) * (1 - c52) * .... * (1 - c90)</code><br><br>\nHere, I'll explain how we found this special formula.<br><br>\nFirst, we started with the normal method like this.<br>\n<code>c99 = 0.9 * (1 - c42) * (1 - c52) * .... * (1 - c90)</code> <br><br>\nThen, we noticed this way works better, which means c90 is not similar to c99 at all. <br>\n<code>c99 = 0.9 * (1 - c42) * (1 - c52) * .... * (1 - c90) ** 2</code> <br><br>\nThen, we also tried changing power for other classes and got this formula. (We spent about 19 subs for this step) <br>\n<code>powers = {c42 : 0, c52 : 0, c62 : 0.5, c95 : 1, c90 : 2, c88 : 2, c67 : 2, c64 : 2,  c15 : 2}</code><br>\n<code>c99 = np.prod([(1 - pred[c]) ** powers[c] for c in classes])</code> <br> <br>\nThis formula means c99 is very similar to c42 and c52, and somewhat similar to c62 and c95. actually, this formula already improved our score by 0.049 and it was the last day of the competition. but, we didn't give up and finally found the problem of this formula. <br>                  <br>\nFor object_id whose <code>c90 = 0.5</code> and <code>c88 = 0.5</code>, we thought we should assign <code>c99 = 0</code>.</p>\n<p>However, this formula assigns <code>c99 = 0.5 ** 2 * 0.5 ** 2 = 0.0625</code>, which looks too high for us. so, we thought we should do something like <code>c99 *= (1 - c90 - c88)</code> and finally found <br>\n<code>c99 = 1.65 * (1 - c90 - c88 - c67 - c64 - c15) ** 2 * (1 - c62 - c95) ** 0.635</code> <br>\n        <code>= 1.65 * (c42 + c52 + c62 + c95) ** 2 * (1 - c62 - c95) ** 0.635</code> <br>\nis a very good approximation of the formula. <br><br>\nActually, it was our last submission in this competition and we couldn't tune coefficient at all. <br></p>\n<h1>what didn't work for me</h1>\n<p>some of the experiments that didn't work (Not all)<br></p>\n<ul>\n<li><code>Band-correction using hostgal_photoz</code><br></li>\n<li><code>target encoding using ra, decl</code><br></li>\n<li><code>using adversarial validation for feature selection</code><br></li>\n<li><code>stacking using GBDT as meta model</code><br></li>\n<li><code>UNet-like Autoencoder for feature extraction</code><br></li>\n<li><code>make binary classification model and use oof-predictions as features</code><br></li>\n</ul>\n<h1>what worked</h1>\n<ul>\n<li><code>class90 pseudo-labelling</code>: <br>\nIt gave 0.016 improvements for me.<br>\n<br></li>\n<li><code>adversarial pseudo-labelling</code>: <br>\nI selected the objects whose the prediction is very high and similar to train data and used them for training. It didn't give a significant improvement for me, but I suppose it somewhat contributed to diversity in ensemble. </li>\n</ul>\n<h1>Comments</h1>\n<p>Actually, I haven't figured out why we dropped from 2nd to 3rd on Private LB.  I'm wondering whether it's because of the luck or not.</p>\n<p>As for the competition design, I think making 'class99' is not so good, because it leads to LB probing and is not so practical. However, on the whole, I suppose this competition is really interesting and one of the most successful competition in kaggle. I'm especially really impressed with the yuval's beautiful and practical NN modelling :)</p>\n<p>Anyway, I really enjoyed my second kaggle competition. I'll continue kaggle and want to become the winner in the next competition :) </p>\n<h1>Thanks all, see you again!</h1>",
      "rawMarkdown": "Thank you my briliant team, all teams who competed with us, and all people who participated in this competition. Congrats Kyle Boone, who won this competition with an incredible performance! <br>\nI'm happy because I and yuval became kaggle expert with 2 gold medals in this competition, and We finally became Prize Winner :) <br>\nIt's my second kaggle competition and it was fun to compete with AhmetErdem and CPMP again, who I competed with in my previous competition, TalkingData Adtracking Fraud Detection. Here, I won't explain all of my team's solution, please take a look at @yuval6967 and @nyanpn solutions, if you want to see what my teammate has done. <br>\n`yuval's solution` : https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75116\n`nyanp's solution` : https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75222\n`My github for the host`: https://github.com/takashioya/plasticc\n\n# My Model\nmy model is based on CatBoost, with lotta feature engineering. \nI made 3 types of models.\n(1) galactic model, the number of features are about 1000.\n(2) extra-galactic model, the number of features are about 2000. \n(3) extra-galactic model with hostgal_specz, the number of features are about 2000. It gave 0.004 improvement.\nThen, I trained 7 * 3 types models with different feature set and get simple average of them, which scored 0.806 on Public LB, while my best single model scored 0.823. <br> \nyou may think the number of features of my model is too big, but using small number of features degraded the score. tbh I prefer to use big features, because I think big features lead to big improvement in ensemble, from my experience. \n\n# My Features\n\nactually, I used nyanp's features for my model and some of them have much bigger importance than mine (especially salt-2 features). However, some of my features still have high importance, so I'll explain a bit about them. \n\n- `features related to mjd`: \nsomething like 'variance of mjd whose detected == 1' worked well.\n- `abs_curve_angle_features`: \nby calculating np.arctan(flux\\_diff/mjd\\_diff), we can calculate 'angle' of light curve. \nI calculated many stats from this.\n- `feets features`: \nI used https://github.com/carpyncho/feets, which was effective.\n- `gaussian process features`: \nIt worked better than the features like 'days\\_from\\_peak\\_n\\_percent\\_flux'. \nfor gaussian process, I used Gpy, but it's too slow and I \nhad to use 30 instances for feature extraction.\n- `shifted_flux * distance ** 2 features`: \nI calculated shifted\\_flux = flux - flux.min() if flux.min() &gt; 0 per each object_id and passband, which worked better than just calculating `flux * distance ** 2`\n\n\n\n# Ensemble\nI used weighted average for ensemble. however, I gave different weights to each class, which gave us 0.006 improvement. I calculated the weights using oof-predictions by hyperopt. <br>\nBy getting weighted average of mamas's CatBoost(0.806) and nyanp's LightGBM(0.834), we got 0.780 on public LB.Then, in the same way, we calculated weighted average of 0.780 model and yuval's NN (0.791) and got 0.737.\n\n# probing class99\nActually, we did nothing special about class99 7 days before the end of the competition. \nWe noticed class99 probing is very effective 6 days before the end and spent 5 days probing class99 distribution. \nFinally, we got 0.057 improvement by probing class99, which took us to the Prize zone (0.680 on public LB). <br>\nFor extra-galactic, we used this special formula. \n`c99 = 1.65 * (c42 + c52 + c62 + c95) ** 2.5 *  (1 - c62 - c95) ** 0.635` <br>\nFor galactic, we did nothing special, just a normal method like this. \n`c99 = 0.2 * (1 - c42) * (1 - c52) * .... * (1 - c90)`<br>\nHere, I'll explain how we found this special formula.<br>\nFirst, we started with the normal method like this.\n`c99 = 0.9 * (1 - c42) * (1 - c52) * .... * (1 - c90)` <br>\nThen, we noticed this way works better, which means c90 is not similar to c99 at all. \n`c99 = 0.9 * (1 - c42) * (1 - c52) * .... * (1 - c90) ** 2` <br>\nThen, we also tried changing power for other classes and got this formula. (We spent about 19 subs for this step) \n`powers = {c42 : 0, c52 : 0, c62 : 0.5, c95 : 1, c90 : 2, c88 : 2, c67 : 2, c64 : 2,  c15 : 2}`\n`c99 = np.prod([(1 - pred[c]) ** powers[c] for c in classes])` <br> \nThis formula means c99 is very similar to c42 and c52, and somewhat similar to c62 and c95. actually, this formula already improved our score by 0.049 and it was the last day of the competition. but, we didn't give up and finally found the problem of this formula. <br>                  \nFor object_id whose `c90 = 0.5` and `c88 = 0.5`, we thought we should assign `c99 = 0`.\n\nHowever, this formula assigns `c99 = 0.5 ** 2 * 0.5 ** 2 = 0.0625`, which looks too high for us. so, we thought we should do something like `c99 *= (1 - c90 - c88)` and finally found \n`c99 = 1.65 * (1 - c90 - c88 - c67 - c64 - c15) ** 2 * (1 - c62 - c95) ** 0.635` \n        `= 1.65 * (c42 + c52 + c62 + c95) ** 2 * (1 - c62 - c95) ** 0.635` \nis a very good approximation of the formula. <br>\nActually, it was our last submission in this competition and we couldn't tune coefficient at all. <br>\n\n# what didn't work for me\nsome of the experiments that didn't work (Not all)<br>\n- `Band-correction using hostgal_photoz`<br>\n- `target encoding using ra, decl`<br>\n- `using adversarial validation for feature selection`<br>\n- `stacking using GBDT as meta model`<br>\n- `UNet-like Autoencoder for feature extraction`<br>\n- `make binary classification model and use oof-predictions as features`<br>\n\n# what worked\n- `class90 pseudo-labelling`: \nIt gave 0.016 improvements for me.\n<br>\n- `adversarial pseudo-labelling`: \nI selected the objects whose the prediction is very high and similar to train data and used them for training. It didn't give a significant improvement for me, but I suppose it somewhat contributed to diversity in ensemble. \n\n# Comments\nActually, I haven't figured out why we dropped from 2nd to 3rd on Private LB.  I'm wondering whether it's because of the luck or not.\n\nAs for the competition design, I think making 'class99' is not so good, because it leads to LB probing and is not so practical. However, on the whole, I suppose this competition is really interesting and one of the most successful competition in kaggle. I'm especially really impressed with the yuval's beautiful and practical NN modelling :)\n\nAnyway, I really enjoyed my second kaggle competition. I'll continue kaggle and want to become the winner in the next competition :) \n<h1>Thanks all, see you again!</h1>",
      "votes": null
    },
    {
      "id": "441527",
      "postDate": "12/18/2018 18:39:55",
      "content": "<p>Congrats on the final result!</p>\n\n<p>Thanks a lot for sharing.  I tried FEATS without much success, is feets really different?</p>\n\n<p>I see that once again class 99 probing is what we missed most.  Too bad for us.</p>",
      "rawMarkdown": "Congrats on the final result!\n\nThanks a lot for sharing.  I tried FEATS without much success, is feets really different?\n\nI see that once again class 99 probing is what we missed most.  Too bad for us.",
      "votes": null
    },
    {
      "id": "441528",
      "postDate": "12/18/2018 18:44:30",
      "content": "<p>Congrats and thanks for sharing! Did you calculate all your features on adjusted flux? </p>\n\n<p>I also tried <code>abs_curve_angle_features</code> in total curve, per bandpass and concentrated around maximum, but they only helped my model overfit to training set in all cases</p>",
      "rawMarkdown": "Congrats and thanks for sharing! Did you calculate all your features on adjusted flux? \n\nI also tried `abs_curve_angle_features` in total curve, per bandpass and concentrated around maximum, but they only helped my model overfit to training set in all cases",
      "votes": null
    },
    {
      "id": "441534",
      "postDate": "12/18/2018 18:59:17",
      "content": "<p>I remember feets gave a small improvement for me. (but maybe it's not significant)\nyes, actually I might have been 5th place if we had not succeeded in class99 probing, because our best private LB score without class99 probing is 0.758, which is worse than your team.</p>",
      "rawMarkdown": "I remember feets gave a small improvement for me. (but maybe it's not significant)\nyes, actually I might have been 5th place if we had not succeeded in class99 probing, because our best private LB score without class99 probing is 0.758, which is worse than your team.",
      "votes": null
    },
    {
      "id": "441536",
      "postDate": "12/18/2018 19:05:38",
      "content": "<p>no, some features are calculated on adjusted flux, others not. </p>\n\n<p><code>\nI also tried abs_curve_angle_features in total curve, per bandpass and concentrated around maximum, but they only helped my model overfit to training set in all cases\n</code>\nIt may be related to your model. actually, my LightGBM model doesn't work at all in this competition, it overfits to training set very fast.</p>",
      "rawMarkdown": "no, some features are calculated on adjusted flux, others not. \n\n`\nI also tried abs_curve_angle_features in total curve, per bandpass and concentrated around maximum, but they only helped my model overfit to training set in all cases\n`\nIt may be related to your model. actually, my LightGBM model doesn't work at all in this competition, it overfits to training set very fast.",
      "votes": null
    },
    {
      "id": "441547",
      "postDate": "12/18/2018 19:24:49",
      "content": "<p>Congrats! I'm amazed that you were able to use 2000 features on the training set without running into serious overfitting problems. For the galactic set, there are only 2325 objects in the training set! How did you regularize to deal with that?</p>",
      "rawMarkdown": "Congrats! I'm amazed that you were able to use 2000 features on the training set without running into serious overfitting problems. For the galactic set, there are only 2325 objects in the training set! How did you regularize to deal with that?",
      "votes": null
    },
    {
      "id": "441549",
      "postDate": "12/18/2018 19:27:36",
      "content": "<p>Congraturations! Thank you for sharing. I'll try your solutions.</p>",
      "rawMarkdown": "Congraturations! Thank you for sharing. I'll try your solutions.",
      "votes": null
    },
    {
      "id": "441550",
      "postDate": "12/18/2018 19:30:19",
      "content": "<p>Thanks, Kyle. Actually, I was able to use more than 5000 features without overfitting problem. The only reason I used 2000 features is inference time and loading time for the huge test set.</p>",
      "rawMarkdown": "Thanks, Kyle. Actually, I was able to use more than 5000 features without overfitting problem. The only reason I used 2000 features is inference time and loading time for the huge test set.",
      "votes": null
    },
    {
      "id": "441555",
      "postDate": "12/18/2018 19:35:44",
      "content": "<p>Another comment related @mamas adversarial pseudo-labeling, I used his pseudo-labeling to retrain the NN's, which gave nice improvement to 2 of the 3 NN (The only one which wasn't improved was the one with mamas's best features ;) )  </p>",
      "rawMarkdown": "Another comment related @mamas adversarial pseudo-labeling, I used his pseudo-labeling to retrain the NN's, which gave nice improvement to 2 of the 3 NN (The only one which wasn't improved was the one with mamas's best features ;) )",
      "votes": null
    },
    {
      "id": "441563",
      "postDate": "12/18/2018 19:40:16",
      "content": "<p>Very interesting! I thought about trying to do something like this but ran out of time. How much did the pseudo-labeling improve your final scores?</p>",
      "rawMarkdown": "Very interesting! I thought about trying to do something like this but ran out of time. How much did the pseudo-labeling improve your final scores?",
      "votes": null
    },
    {
      "id": "441567",
      "postDate": "12/18/2018 19:44:36",
      "content": "<p>Thanks yuval, I didn't know you succeeded in adversarial pseudo-labelling. </p>\n\n<p><code>The only one which wasn't improved was the one with mamas's best features</code></p>\n\n<p>It may be because these features are used in my adversarial model.</p>",
      "rawMarkdown": "Thanks yuval, I didn't know you succeeded in adversarial pseudo-labelling. \n\n`The only one which wasn't improved was the one with mamas's best features`\n\nIt may be because these features are used in my adversarial model.",
      "votes": null
    },
    {
      "id": "441580",
      "postDate": "12/18/2018 20:01:55",
      "content": "<p>The improvement for the NN average was about 0.01 which probably gave 0.005 in the final score (we can't know for sure because we didn't submit this improvement alone) </p>",
      "rawMarkdown": "The improvement for the NN average was about 0.01 which probably gave 0.005 in the final score (we can't know for sure because we didn't submit this improvement alone)",
      "votes": null
    },
    {
      "id": "441582",
      "postDate": "12/18/2018 20:03:29",
      "content": "<p>If you are right, we would've been 4th place if we hadn't tried adversarial pseudo-labelling. it's terrible!</p>",
      "rawMarkdown": "If you are right, we would've been 4th place if we hadn't tried adversarial pseudo-labelling. it's terrible!",
      "votes": null
    },
    {
      "id": "441598",
      "postDate": "12/18/2018 20:34:55",
      "content": "<p>Congrats and thanks for sharing a detailed explanation of Class 99 probing.</p>",
      "rawMarkdown": "Congrats and thanks for sharing a detailed explanation of Class 99 probing.",
      "votes": null
    },
    {
      "id": "441734",
      "postDate": "12/19/2018 01:30:23",
      "content": "<p>It is the first time I see catboost being better than lgb or xgb.  To make sure this is because of catboost and not something else I'm curious to know if you tried lgb with your features, or if you tried catboost with nyanpn features?</p>",
      "rawMarkdown": "It is the first time I see catboost being better than lgb or xgb.  To make sure this is because of catboost and not something else I'm curious to know if you tried lgb with your features, or if you tried catboost with nyanpn features?",
      "votes": null
    },
    {
      "id": "441846",
      "postDate": "12/19/2018 06:45:16",
      "content": "<p>I think it's not only because of catboost, it is very sensitive to the features we use. These are what we tried.</p>\n\n<p><code>lgb only with my features</code> : didn't work at all. I remember it's much worse than catboost by more than 0.05.</p>\n\n<p><code>lgb only with nyanp features</code>: it somewhat works, scored 0.834 on LB. (you can see detail in nyanp's solution)</p>\n\n<p><code>lgb with nyanp and mamas features</code>: didn't work at all.</p>\n\n<p><code>catboost only with my features</code>: actually I haven't checked it for a long time, but I remember it worked.</p>\n\n<p><code>catboost with my features and nyanp features</code> : it worked well, scored 0.806 on LB as I said. </p>\n\n<p><code>XGBoost only with my features</code>: It's not so bad as lgb, but works worse than catboost.</p>",
      "rawMarkdown": "I think it's not only because of catboost, it is very sensitive to the features we use. These are what we tried.\n\n`lgb only with my features` : didn't work at all. I remember it's much worse than catboost by more than 0.05.\n\n`lgb only with nyanp features`: it somewhat works, scored 0.834 on LB. (you can see detail in nyanp's solution)\n\n`lgb with nyanp and mamas features`: didn't work at all.\n\n`catboost only with my features`: actually I haven't checked it for a long time, but I remember it worked.\n\n`catboost with my features and nyanp features` : it worked well, scored 0.806 on LB as I said. \n\n`XGBoost only with my features`: It's not so bad as lgb, but works worse than catboost.",
      "votes": null
    },
    {
      "id": "441943",
      "postDate": "12/19/2018 09:18:29",
      "content": "<p>I'd like to try your class_99 computation with our best submission (so that it hurts even more ;) )  From your post I assume you normalize the sum of probabilities for known classes to 1 before applying your formula.  Is that right?</p>",
      "rawMarkdown": "I'd like to try your class_99 computation with our best submission (so that it hurts even more ;) )  From your post I assume you normalize the sum of probabilities for known classes to 1 before applying your formula.  Is that right?",
      "votes": null
    },
    {
      "id": "441965",
      "postDate": "12/19/2018 09:54:42",
      "content": "<p>yes, right.</p>",
      "rawMarkdown": "yes, right.",
      "votes": null
    },
    {
      "id": "442038",
      "postDate": "12/19/2018 11:56:57",
      "content": "<p>Yes, it will hurt a lot more. Our PB score would have been 0.656 with this Class_99 formula!</p>",
      "rawMarkdown": "Yes, it will hurt a lot more. Our PB score would have been 0.656 with this Class_99 formula!",
      "votes": null
    },
    {
      "id": "442100",
      "postDate": "12/19/2018 13:42:56",
      "content": "<p>I also tried Catboost the last day and I have been surprised by its performance. \nA bit less overfitting than LGBM with same features</p>",
      "rawMarkdown": "I also tried Catboost the last day and I have been surprised by its performance. \nA bit less overfitting than LGBM with same features",
      "votes": null
    },
    {
      "id": "442162",
      "postDate": "12/19/2018 15:00:17",
      "content": "<p>Did you try LGB with dart? Dart gave me a slightly less overfitted model than gbdt. Wondeing if Catboost would still outperform it. </p>",
      "rawMarkdown": "Did you try LGB with dart? Dart gave me a slightly less overfitted model than gbdt. Wondeing if Catboost would still outperform it.",
      "votes": null
    },
    {
      "id": "442166",
      "postDate": "12/19/2018 15:03:18",
      "content": "<p>Your class_99 calculation improved my best model by 0.039!</p>\n\n<p>I guess one of my biggest lessons out of this competition is: don't be afraid to LB probe!</p>",
      "rawMarkdown": "Your class_99 calculation improved my best model by 0.039!\n\nI guess one of my biggest lessons out of this competition is: don't be afraid to LB probe!",
      "votes": null
    },
    {
      "id": "442196",
      "postDate": "12/19/2018 15:45:54",
      "content": "<p>no, I didn't try dart</p>",
      "rawMarkdown": "no, I didn't try dart",
      "votes": null
    },
    {
      "id": "442267",
      "postDate": "12/19/2018 17:54:47",
      "content": "<p>Your class 99 probing strategy is great, congrats!</p>",
      "rawMarkdown": "Your class 99 probing strategy is great, congrats!",
      "votes": null
    },
    {
      "id": "443356",
      "postDate": "12/21/2018 13:40:04",
      "content": "<p>I yields to some class_99 probabilities being higher than 1 in my case.  I guess it is a sign that I need to move to something else ;)</p>",
      "rawMarkdown": "I yields to some class_99 probabilities being higher than 1 in my case.  I guess it is a sign that I need to move to something else ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 441527,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/18/2018 18:39:55",
      "content": "<p>Congrats on the final result!</p>\n\n<p>Thanks a lot for sharing.  I tried FEATS without much success, is feets really different?</p>\n\n<p>I see that once again class 99 probing is what we missed most.  Too bad for us.</p>",
      "votes": null,
      "replies": [
        {
          "id": 441534,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/18/2018 18:59:17",
          "content": "<p>I remember feets gave a small improvement for me. (but maybe it's not significant)\nyes, actually I might have been 5th place if we had not succeeded in class99 probing, because our best private LB score without class99 probing is 0.758, which is worse than your team.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441528,
      "author_name": "iprapas",
      "author_url": "",
      "post_date": "12/18/2018 18:44:30",
      "content": "<p>Congrats and thanks for sharing! Did you calculate all your features on adjusted flux? </p>\n\n<p>I also tried <code>abs_curve_angle_features</code> in total curve, per bandpass and concentrated around maximum, but they only helped my model overfit to training set in all cases</p>",
      "votes": null,
      "replies": [
        {
          "id": 441536,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/18/2018 19:05:38",
          "content": "<p>no, some features are calculated on adjusted flux, others not. </p>\n\n<p><code>\nI also tried abs_curve_angle_features in total curve, per bandpass and concentrated around maximum, but they only helped my model overfit to training set in all cases\n</code>\nIt may be related to your model. actually, my LightGBM model doesn't work at all in this competition, it overfits to training set very fast.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441547,
      "author_name": "kyleboone",
      "author_url": "",
      "post_date": "12/18/2018 19:24:49",
      "content": "<p>Congrats! I'm amazed that you were able to use 2000 features on the training set without running into serious overfitting problems. For the galactic set, there are only 2325 objects in the training set! How did you regularize to deal with that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 441550,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/18/2018 19:30:19",
          "content": "<p>Thanks, Kyle. Actually, I was able to use more than 5000 features without overfitting problem. The only reason I used 2000 features is inference time and loading time for the huge test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441549,
      "author_name": "takus4649",
      "author_url": "",
      "post_date": "12/18/2018 19:27:36",
      "content": "<p>Congraturations! Thank you for sharing. I'll try your solutions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 441555,
      "author_name": "yuval6967",
      "author_url": "",
      "post_date": "12/18/2018 19:35:44",
      "content": "<p>Another comment related @mamas adversarial pseudo-labeling, I used his pseudo-labeling to retrain the NN's, which gave nice improvement to 2 of the 3 NN (The only one which wasn't improved was the one with mamas's best features ;) )  </p>",
      "votes": null,
      "replies": [
        {
          "id": 441563,
          "author_name": "kyleboone",
          "author_url": "",
          "post_date": "12/18/2018 19:40:16",
          "content": "<p>Very interesting! I thought about trying to do something like this but ran out of time. How much did the pseudo-labeling improve your final scores?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441567,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/18/2018 19:44:36",
          "content": "<p>Thanks yuval, I didn't know you succeeded in adversarial pseudo-labelling. </p>\n\n<p><code>The only one which wasn't improved was the one with mamas's best features</code></p>\n\n<p>It may be because these features are used in my adversarial model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441580,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "12/18/2018 20:01:55",
          "content": "<p>The improvement for the NN average was about 0.01 which probably gave 0.005 in the final score (we can't know for sure because we didn't submit this improvement alone) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441582,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/18/2018 20:03:29",
          "content": "<p>If you are right, we would've been 4th place if we hadn't tried adversarial pseudo-labelling. it's terrible!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441598,
      "author_name": "vignam",
      "author_url": "",
      "post_date": "12/18/2018 20:34:55",
      "content": "<p>Congrats and thanks for sharing a detailed explanation of Class 99 probing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 441734,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/19/2018 01:30:23",
      "content": "<p>It is the first time I see catboost being better than lgb or xgb.  To make sure this is because of catboost and not something else I'm curious to know if you tried lgb with your features, or if you tried catboost with nyanpn features?</p>",
      "votes": null,
      "replies": [
        {
          "id": 441846,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/19/2018 06:45:16",
          "content": "<p>I think it's not only because of catboost, it is very sensitive to the features we use. These are what we tried.</p>\n\n<p><code>lgb only with my features</code> : didn't work at all. I remember it's much worse than catboost by more than 0.05.</p>\n\n<p><code>lgb only with nyanp features</code>: it somewhat works, scored 0.834 on LB. (you can see detail in nyanp's solution)</p>\n\n<p><code>lgb with nyanp and mamas features</code>: didn't work at all.</p>\n\n<p><code>catboost only with my features</code>: actually I haven't checked it for a long time, but I remember it worked.</p>\n\n<p><code>catboost with my features and nyanp features</code> : it worked well, scored 0.806 on LB as I said. </p>\n\n<p><code>XGBoost only with my features</code>: It's not so bad as lgb, but works worse than catboost.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442100,
          "author_name": "vlarmet",
          "author_url": "",
          "post_date": "12/19/2018 13:42:56",
          "content": "<p>I also tried Catboost the last day and I have been surprised by its performance. \nA bit less overfitting than LGBM with same features</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442162,
          "author_name": "taniaj",
          "author_url": "",
          "post_date": "12/19/2018 15:00:17",
          "content": "<p>Did you try LGB with dart? Dart gave me a slightly less overfitted model than gbdt. Wondeing if Catboost would still outperform it. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442196,
          "author_name": "vlarmet",
          "author_url": "",
          "post_date": "12/19/2018 15:45:54",
          "content": "<p>no, I didn't try dart</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441943,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/19/2018 09:18:29",
      "content": "<p>I'd like to try your class_99 computation with our best submission (so that it hurts even more ;) )  From your post I assume you normalize the sum of probabilities for known classes to 1 before applying your formula.  Is that right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 441965,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "12/19/2018 09:54:42",
          "content": "<p>yes, right.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442038,
          "author_name": "psilogram",
          "author_url": "",
          "post_date": "12/19/2018 11:56:57",
          "content": "<p>Yes, it will hurt a lot more. Our PB score would have been 0.656 with this Class_99 formula!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443356,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/21/2018 13:40:04",
          "content": "<p>I yields to some class_99 probabilities being higher than 1 in my case.  I guess it is a sign that I need to move to something else ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442166,
      "author_name": "taniaj",
      "author_url": "",
      "post_date": "12/19/2018 15:03:18",
      "content": "<p>Your class_99 calculation improved my best model by 0.039!</p>\n\n<p>I guess one of my biggest lessons out of this competition is: don't be afraid to LB probe!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 442267,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "12/19/2018 17:54:47",
      "content": "<p>Your class 99 probing strategy is great, congrats!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "441520": "Thank you my briliant team, all teams who competed with us, and all people who participated in this competition. Congrats Kyle Boone, who won this competition with an incredible performance! <br>\nI'm happy because I and yuval became kaggle expert with 2 gold medals in this competition, and We finally became Prize Winner :) <br>\nIt's my second kaggle competition and it was fun to compete with AhmetErdem and CPMP again, who I competed with in my previous competition, TalkingData Adtracking Fraud Detection. Here, I won't explain all of my team's solution, please take a look at @yuval6967 and @nyanpn solutions, if you want to see what my teammate has done. <br>\n`yuval's solution` : https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75116\n`nyanp's solution` : https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75222\n`My github for the host`: https://github.com/takashioya/plasticc\n\n# My Model\nmy model is based on CatBoost, with lotta feature engineering. \nI made 3 types of models.\n(1) galactic model, the number of features are about 1000.\n(2) extra-galactic model, the number of features are about 2000. \n(3) extra-galactic model with hostgal_specz, the number of features are about 2000. It gave 0.004 improvement.\nThen, I trained 7 * 3 types models with different feature set and get simple average of them, which scored 0.806 on Public LB, while my best single model scored 0.823. <br> \nyou may think the number of features of my model is too big, but using small number of features degraded the score. tbh I prefer to use big features, because I think big features lead to big improvement in ensemble, from my experience. \n\n# My Features\n\nactually, I used nyanp's features for my model and some of them have much bigger importance than mine (especially salt-2 features). However, some of my features still have high importance, so I'll explain a bit about them. \n\n- `features related to mjd`: \nsomething like 'variance of mjd whose detected == 1' worked well.\n- `abs_curve_angle_features`: \nby calculating np.arctan(flux\\_diff/mjd\\_diff), we can calculate 'angle' of light curve. \nI calculated many stats from this.\n- `feets features`: \nI used https://github.com/carpyncho/feets, which was effective.\n- `gaussian process features`: \nIt worked better than the features like 'days\\_from\\_peak\\_n\\_percent\\_flux'. \nfor gaussian process, I used Gpy, but it's too slow and I \nhad to use 30 instances for feature extraction.\n- `shifted_flux * distance ** 2 features`: \nI calculated shifted\\_flux = flux - flux.min() if flux.min() &gt; 0 per each object_id and passband, which worked better than just calculating `flux * distance ** 2`\n\n\n\n# Ensemble\nI used weighted average for ensemble. however, I gave different weights to each class, which gave us 0.006 improvement. I calculated the weights using oof-predictions by hyperopt. <br>\nBy getting weighted average of mamas's CatBoost(0.806) and nyanp's LightGBM(0.834), we got 0.780 on public LB.Then, in the same way, we calculated weighted average of 0.780 model and yuval's NN (0.791) and got 0.737.\n\n# probing class99\nActually, we did nothing special about class99 7 days before the end of the competition. \nWe noticed class99 probing is very effective 6 days before the end and spent 5 days probing class99 distribution. \nFinally, we got 0.057 improvement by probing class99, which took us to the Prize zone (0.680 on public LB). <br>\nFor extra-galactic, we used this special formula. \n`c99 = 1.65 * (c42 + c52 + c62 + c95) ** 2.5 *  (1 - c62 - c95) ** 0.635` <br>\nFor galactic, we did nothing special, just a normal method like this. \n`c99 = 0.2 * (1 - c42) * (1 - c52) * .... * (1 - c90)`<br>\nHere, I'll explain how we found this special formula.<br>\nFirst, we started with the normal method like this.\n`c99 = 0.9 * (1 - c42) * (1 - c52) * .... * (1 - c90)` <br>\nThen, we noticed this way works better, which means c90 is not similar to c99 at all. \n`c99 = 0.9 * (1 - c42) * (1 - c52) * .... * (1 - c90) ** 2` <br>\nThen, we also tried changing power for other classes and got this formula. (We spent about 19 subs for this step) \n`powers = {c42 : 0, c52 : 0, c62 : 0.5, c95 : 1, c90 : 2, c88 : 2, c67 : 2, c64 : 2,  c15 : 2}`\n`c99 = np.prod([(1 - pred[c]) ** powers[c] for c in classes])` <br> \nThis formula means c99 is very similar to c42 and c52, and somewhat similar to c62 and c95. actually, this formula already improved our score by 0.049 and it was the last day of the competition. but, we didn't give up and finally found the problem of this formula. <br>                  \nFor object_id whose `c90 = 0.5` and `c88 = 0.5`, we thought we should assign `c99 = 0`.\n\nHowever, this formula assigns `c99 = 0.5 ** 2 * 0.5 ** 2 = 0.0625`, which looks too high for us. so, we thought we should do something like `c99 *= (1 - c90 - c88)` and finally found \n`c99 = 1.65 * (1 - c90 - c88 - c67 - c64 - c15) ** 2 * (1 - c62 - c95) ** 0.635` \n        `= 1.65 * (c42 + c52 + c62 + c95) ** 2 * (1 - c62 - c95) ** 0.635` \nis a very good approximation of the formula. <br>\nActually, it was our last submission in this competition and we couldn't tune coefficient at all. <br>\n\n# what didn't work for me\nsome of the experiments that didn't work (Not all)<br>\n- `Band-correction using hostgal_photoz`<br>\n- `target encoding using ra, decl`<br>\n- `using adversarial validation for feature selection`<br>\n- `stacking using GBDT as meta model`<br>\n- `UNet-like Autoencoder for feature extraction`<br>\n- `make binary classification model and use oof-predictions as features`<br>\n\n# what worked\n- `class90 pseudo-labelling`: \nIt gave 0.016 improvements for me.\n<br>\n- `adversarial pseudo-labelling`: \nI selected the objects whose the prediction is very high and similar to train data and used them for training. It didn't give a significant improvement for me, but I suppose it somewhat contributed to diversity in ensemble. \n\n# Comments\nActually, I haven't figured out why we dropped from 2nd to 3rd on Private LB.  I'm wondering whether it's because of the luck or not.\n\nAs for the competition design, I think making 'class99' is not so good, because it leads to LB probing and is not so practical. However, on the whole, I suppose this competition is really interesting and one of the most successful competition in kaggle. I'm especially really impressed with the yuval's beautiful and practical NN modelling :)\n\nAnyway, I really enjoyed my second kaggle competition. I'll continue kaggle and want to become the winner in the next competition :) \n<h1>Thanks all, see you again!</h1>",
    "441527": "Congrats on the final result!\n\nThanks a lot for sharing.  I tried FEATS without much success, is feets really different?\n\nI see that once again class 99 probing is what we missed most.  Too bad for us.",
    "441528": "Congrats and thanks for sharing! Did you calculate all your features on adjusted flux? \n\nI also tried `abs_curve_angle_features` in total curve, per bandpass and concentrated around maximum, but they only helped my model overfit to training set in all cases",
    "441534": "I remember feets gave a small improvement for me. (but maybe it's not significant)\nyes, actually I might have been 5th place if we had not succeeded in class99 probing, because our best private LB score without class99 probing is 0.758, which is worse than your team.",
    "441536": "no, some features are calculated on adjusted flux, others not. \n\n`\nI also tried abs_curve_angle_features in total curve, per bandpass and concentrated around maximum, but they only helped my model overfit to training set in all cases\n`\nIt may be related to your model. actually, my LightGBM model doesn't work at all in this competition, it overfits to training set very fast.",
    "441547": "Congrats! I'm amazed that you were able to use 2000 features on the training set without running into serious overfitting problems. For the galactic set, there are only 2325 objects in the training set! How did you regularize to deal with that?",
    "441549": "Congraturations! Thank you for sharing. I'll try your solutions.",
    "441550": "Thanks, Kyle. Actually, I was able to use more than 5000 features without overfitting problem. The only reason I used 2000 features is inference time and loading time for the huge test set.",
    "441555": "Another comment related @mamas adversarial pseudo-labeling, I used his pseudo-labeling to retrain the NN's, which gave nice improvement to 2 of the 3 NN (The only one which wasn't improved was the one with mamas's best features ;) )",
    "441563": "Very interesting! I thought about trying to do something like this but ran out of time. How much did the pseudo-labeling improve your final scores?",
    "441567": "Thanks yuval, I didn't know you succeeded in adversarial pseudo-labelling. \n\n`The only one which wasn't improved was the one with mamas's best features`\n\nIt may be because these features are used in my adversarial model.",
    "441580": "The improvement for the NN average was about 0.01 which probably gave 0.005 in the final score (we can't know for sure because we didn't submit this improvement alone)",
    "441582": "If you are right, we would've been 4th place if we hadn't tried adversarial pseudo-labelling. it's terrible!",
    "441598": "Congrats and thanks for sharing a detailed explanation of Class 99 probing.",
    "441734": "It is the first time I see catboost being better than lgb or xgb.  To make sure this is because of catboost and not something else I'm curious to know if you tried lgb with your features, or if you tried catboost with nyanpn features?",
    "441846": "I think it's not only because of catboost, it is very sensitive to the features we use. These are what we tried.\n\n`lgb only with my features` : didn't work at all. I remember it's much worse than catboost by more than 0.05.\n\n`lgb only with nyanp features`: it somewhat works, scored 0.834 on LB. (you can see detail in nyanp's solution)\n\n`lgb with nyanp and mamas features`: didn't work at all.\n\n`catboost only with my features`: actually I haven't checked it for a long time, but I remember it worked.\n\n`catboost with my features and nyanp features` : it worked well, scored 0.806 on LB as I said. \n\n`XGBoost only with my features`: It's not so bad as lgb, but works worse than catboost.",
    "441943": "I'd like to try your class_99 computation with our best submission (so that it hurts even more ;) )  From your post I assume you normalize the sum of probabilities for known classes to 1 before applying your formula.  Is that right?",
    "441965": "yes, right.",
    "442038": "Yes, it will hurt a lot more. Our PB score would have been 0.656 with this Class_99 formula!",
    "442100": "I also tried Catboost the last day and I have been surprised by its performance. \nA bit less overfitting than LGBM with same features",
    "442162": "Did you try LGB with dart? Dart gave me a slightly less overfitted model than gbdt. Wondeing if Catboost would still outperform it.",
    "442166": "Your class_99 calculation improved my best model by 0.039!\n\nI guess one of my biggest lessons out of this competition is: don't be afraid to LB probe!",
    "442196": "no, I didn't try dart",
    "442267": "Your class 99 probing strategy is great, congrats!",
    "443356": "I yields to some class_99 probabilities being higher than 1 in my case.  I guess it is a sign that I need to move to something else ;)"
  },
  "source": "meta"
}