{
  "id": 75059,
  "title": "2nd-Place Solution Notes",
  "url": "/competitions/PLAsTiCC-2018/discussion/75059",
  "author_name": "Silogram",
  "post_date": "2018-12-18T08:00:55.709000",
  "votes": 69,
  "comment_count": 29,
  "views": 0,
  "content": "<p>First, thanks to Kaggle and the PLAsTiCC organizers for presenting such an interesting challenge, and congratulations to all the top teams, especially Kyle with his amazing performance.</p>\n\n<p>I haven't looked at all the other solution write-ups yet but there were clearly a lot of different ways to tackle this problem effectively. For us, NNs worked much better than LGB models. Our final ensemble included 9 models, 7 of which were NNs that scored as low as 0.75x individually. The two LGB models were much weaker, each scoring about 0.90x, but they did provide a bit of diversity to the ensemble.</p>\n\n<p>My teammate Mike built the NN models so I'll let him describe them separately. What follows is some more general notes about our overall solution:</p>\n\n<ol>\n<li><p>The 'Gap'. There was quite a bit of chatter about lowering the gap between CV and LB scores, most of which we ignored because our LB scores were tracking our CV scores very consistently, with a gap of 0.43-0.44. Our final solution, which scored 0.694 on the LB had a CV score of 0.264 (0.430 gap). In hindsight, it may have been a mistake not to look more closely at the gap. We really didn't pay much attention to the differences between the train and test sets until the final few days of the competition.</p></li>\n<li><p>'Class_99'. Especially in the final week, we did quite a bit of LB probing to come up with a better way to predict Class_99. In the end we settled on the following alorithm:\na) Set Class_99 to 1-top_prediction per row\nb) Move class_99 closer to 0.14 (for extra-galactic objects) and 0.014 (for galactic objects) with the following       formula: class_99 = (2*class_99 + 0.14) / 3\nThis improved our scores modestly (about 0.005).   </p></li>\n<li><p>Ensemble. The final ensemble was a very shallow (max_depth=2) LGB model that was very effective. </p></li>\n<li><p>'Detected' Flag. I still don't understand what this means or how it was calculated, but it's clearly very significant. As Kyle pointed out in a post, it's roughly set when the absolute flux value is 5 times greater than the flux error (the documentation says +- 3 sigmas, which doesn't make much sense to me), but there are many exceptions to this rule. The only thing I can think of is that the flag was set before some random jitter was applied to the flux values. If anyone knows more about this, please tell me because I devoted several days trying to reverse-engineer this feature with no success. That said, I did find that this ratio -- absolute flux/flux_error -- to be very significant in my models. One of the LGB models generates features for 4 different sets of observations -- all, detected only, flux error ratio &gt;3, flux error ratio &gt; 4. Because the NN model was able to continuously process the flux and flux error values in parallel, this relationship did not need to be manually specified, one reason I suspect that the NN model performed so well.</p></li>\n<li><p>Hostgal_specz pseudo-labeling. One thing that worked for both the NN and LGB models was to build a separate model to predict the hostgal_specz values (using both the training set and the test set objects that have this value), and then using oof predictions for these values in our models.</p></li>\n<li><p>Flux adjustments. For the LGB model, I found that adjusting the the flux values with the simple formula flux *= hostgal_photoz improved the results. All efforts to normalize flux values by object and/or passband did not help. In the NN models, on the other hand, Mike found a very effective normalization strategy.</p></li>\n<li><p>MJD adjustments. Adjustments to MJD based on redshift and wavelength did not help the LGB model.</p></li>\n<li><p>Augmentation. We tried a variety of different augmentation methods. Augmenting the training set with random variations helped the NN models but did not help the LGB models. We also tried augmentation with pseudo-labeled test objects but this did not help. As I mentioned above, we weren't really focused on the differences between the train and test sets. In the final days of the competition, I realized that we should be augmenting in such a way that the train distribution would more closely match the test distribution, but we didn't have time to try this idea.</p></li>\n</ol>",
  "messages": [
    {
      "id": 441048,
      "postDate": "2018-12-18T08:00:55.710Z",
      "content": "<p>First, thanks to Kaggle and the PLAsTiCC organizers for presenting such an interesting challenge, and congratulations to all the top teams, especially Kyle with his amazing performance.</p>\n\n<p>I haven't looked at all the other solution write-ups yet but there were clearly a lot of different ways to tackle this problem effectively. For us, NNs worked much better than LGB models. Our final ensemble included 9 models, 7 of which were NNs that scored as low as 0.75x individually. The two LGB models were much weaker, each scoring about 0.90x, but they did provide a bit of diversity to the ensemble.</p>\n\n<p>My teammate Mike built the NN models so I'll let him describe them separately. What follows is some more general notes about our overall solution:</p>\n\n<ol>\n<li><p>The 'Gap'. There was quite a bit of chatter about lowering the gap between CV and LB scores, most of which we ignored because our LB scores were tracking our CV scores very consistently, with a gap of 0.43-0.44. Our final solution, which scored 0.694 on the LB had a CV score of 0.264 (0.430 gap). In hindsight, it may have been a mistake not to look more closely at the gap. We really didn't pay much attention to the differences between the train and test sets until the final few days of the competition.</p></li>\n<li><p>'Class_99'. Especially in the final week, we did quite a bit of LB probing to come up with a better way to predict Class_99. In the end we settled on the following alorithm:\na) Set Class_99 to 1-top_prediction per row\nb) Move class_99 closer to 0.14 (for extra-galactic objects) and 0.014 (for galactic objects) with the following       formula: class_99 = (2*class_99 + 0.14) / 3\nThis improved our scores modestly (about 0.005).   </p></li>\n<li><p>Ensemble. The final ensemble was a very shallow (max_depth=2) LGB model that was very effective. </p></li>\n<li><p>'Detected' Flag. I still don't understand what this means or how it was calculated, but it's clearly very significant. As Kyle pointed out in a post, it's roughly set when the absolute flux value is 5 times greater than the flux error (the documentation says +- 3 sigmas, which doesn't make much sense to me), but there are many exceptions to this rule. The only thing I can think of is that the flag was set before some random jitter was applied to the flux values. If anyone knows more about this, please tell me because I devoted several days trying to reverse-engineer this feature with no success. That said, I did find that this ratio -- absolute flux/flux_error -- to be very significant in my models. One of the LGB models generates features for 4 different sets of observations -- all, detected only, flux error ratio &gt;3, flux error ratio &gt; 4. Because the NN model was able to continuously process the flux and flux error values in parallel, this relationship did not need to be manually specified, one reason I suspect that the NN model performed so well.</p></li>\n<li><p>Hostgal_specz pseudo-labeling. One thing that worked for both the NN and LGB models was to build a separate model to predict the hostgal_specz values (using both the training set and the test set objects that have this value), and then using oof predictions for these values in our models.</p></li>\n<li><p>Flux adjustments. For the LGB model, I found that adjusting the the flux values with the simple formula flux *= hostgal_photoz improved the results. All efforts to normalize flux values by object and/or passband did not help. In the NN models, on the other hand, Mike found a very effective normalization strategy.</p></li>\n<li><p>MJD adjustments. Adjustments to MJD based on redshift and wavelength did not help the LGB model.</p></li>\n<li><p>Augmentation. We tried a variety of different augmentation methods. Augmenting the training set with random variations helped the NN models but did not help the LGB models. We also tried augmentation with pseudo-labeled test objects but this did not help. As I mentioned above, we weren't really focused on the differences between the train and test sets. In the final days of the competition, I realized that we should be augmenting in such a way that the train distribution would more closely match the test distribution, but we didn't have time to try this idea.</p></li>\n</ol>",
      "rawMarkdown": "First, thanks to Kaggle and the PLAsTiCC organizers for presenting such an interesting challenge, and congratulations to all the top teams, especially Kyle with his amazing performance.\n\nI haven't looked at all the other solution write-ups yet but there were clearly a lot of different ways to tackle this problem effectively. For us, NNs worked much better than LGB models. Our final ensemble included 9 models, 7 of which were NNs that scored as low as 0.75x individually. The two LGB models were much weaker, each scoring about 0.90x, but they did provide a bit of diversity to the ensemble.\n\nMy teammate Mike built the NN models so I'll let him describe them separately. What follows is some more general notes about our overall solution:\n\n1. The 'Gap'. There was quite a bit of chatter about lowering the gap between CV and LB scores, most of which we ignored because our LB scores were tracking our CV scores very consistently, with a gap of 0.43-0.44. Our final solution, which scored 0.694 on the LB had a CV score of 0.264 (0.430 gap). In hindsight, it may have been a mistake not to look more closely at the gap. We really didn't pay much attention to the differences between the train and test sets until the final few days of the competition.\n\n2. 'Class_99'. Especially in the final week, we did quite a bit of LB probing to come up with a better way to predict Class_99. In the end we settled on the following alorithm:\n\ta) Set Class_99 to 1-top_prediction per row\n\tb) Move class_99 closer to 0.14 (for extra-galactic objects) and 0.014 (for galactic objects) with the following \t   formula: class_99 = (2*class_99 + 0.14) / 3\nThis improved our scores modestly (about 0.005).   \n\n3. Ensemble. The final ensemble was a very shallow (max_depth=2) LGB model that was very effective. \n\n4. 'Detected' Flag. I still don't understand what this means or how it was calculated, but it's clearly very significant. As Kyle pointed out in a post, it's roughly set when the absolute flux value is 5 times greater than the flux error (the documentation says +- 3 sigmas, which doesn't make much sense to me), but there are many exceptions to this rule. The only thing I can think of is that the flag was set before some random jitter was applied to the flux values. If anyone knows more about this, please tell me because I devoted several days trying to reverse-engineer this feature with no success. That said, I did find that this ratio -- absolute flux/flux_error -- to be very significant in my models. One of the LGB models generates features for 4 different sets of observations -- all, detected only, flux error ratio &gt;3, flux error ratio &gt; 4. Because the NN model was able to continuously process the flux and flux error values in parallel, this relationship did not need to be manually specified, one reason I suspect that the NN model performed so well.\n\n5. Hostgal_specz pseudo-labeling. One thing that worked for both the NN and LGB models was to build a separate model to predict the hostgal_specz values (using both the training set and the test set objects that have this value), and then using oof predictions for these values in our models.\n\n6. Flux adjustments. For the LGB model, I found that adjusting the the flux values with the simple formula flux *= hostgal_photoz improved the results. All efforts to normalize flux values by object and/or passband did not help. In the NN models, on the other hand, Mike found a very effective normalization strategy.\n\n7. MJD adjustments. Adjustments to MJD based on redshift and wavelength did not help the LGB model.\n\n8. Augmentation. We tried a variety of different augmentation methods. Augmenting the training set with random variations helped the NN models but did not help the LGB models. We also tried augmentation with pseudo-labeled test objects but this did not help. As I mentioned above, we weren't really focused on the differences between the train and test sets. In the final days of the competition, I realized that we should be augmenting in such a way that the train distribution would more closely match the test distribution, but we didn't have time to try this idea.",
      "votes": 69
    },
    {
      "id": 446148,
      "postDate": "2018-12-27T15:37:05.043Z",
      "content": "<p>I have added a kernel with our RNN model:\n<a href=\"https://www.kaggle.com/zerrxy/plasticc-rnn\">https://www.kaggle.com/zerrxy/plasticc-rnn</a></p>\n\n<p>Here is a short description:</p>\n\n<ol>\n<li>NN model is a simple one layer Bidirectional GRU (we used 80 - 160 units) with global max pooling. Then max pooling output is combined with meta data and followed by 3 additional dense layers and softmax activation.</li>\n<li>Input data for RNN consists of flux, flux_err, intervals between measurements, emitted wavelength and passband (with embedding layer). flux and fluxerr are divided by a value of (fluxmax - fluxmin) for each object.</li>\n<li>Meta data consists of hostgalphotoz, hostgalphotozerr, ddf, mwebv and log2 of flux range.</li>\n<li>When preparing data for RNN we converted all time and wavelength related features to not depend on red shift. Namely time and wavelength are divided by (hostgalphotoz + 1)</li>\n<li>Augmentation works well for this model. The following steps were used:\n<ul><li>Drop 30% of measurements</li>\n<li>Modify red shift using normal distribution with sigma = hostgalphotozerr * 2/3. When changing red shift, all time and wavelength related features are also modified accordingly.</li>\n<li>Modify flux using normal distribution with sigma = fluxerr * 2/3</li></ul></li>\n</ol>\n\n<p>After all these steps we got a model with about 0.3 CV and 0.75 public LB score (with Olivier' algorithm for class 99).</p>",
      "rawMarkdown": "I have added a kernel with our RNN model:\nhttps://www.kaggle.com/zerrxy/plasticc-rnn\n\nHere is a short description:\n\n1. NN model is a simple one layer Bidirectional GRU (we used 80 - 160 units) with global max pooling. Then max pooling output is combined with meta data and followed by 3 additional dense layers and softmax activation.\n2. Input data for RNN consists of flux, flux_err, intervals between measurements, emitted wavelength and passband (with embedding layer). flux and fluxerr are divided by a value of (fluxmax - fluxmin) for each object.\n3. Meta data consists of hostgalphotoz, hostgalphotozerr, ddf, mwebv and log2 of flux range.\n4. When preparing data for RNN we converted all time and wavelength related features to not depend on red shift. Namely time and wavelength are divided by (hostgalphotoz + 1)\n5. Augmentation works well for this model. The following steps were used:\n- Drop 30% of measurements\n- Modify red shift using normal distribution with sigma = hostgalphotozerr * 2/3. When changing red shift, all time and wavelength related features are also modified accordingly.\n- Modify flux using normal distribution with sigma = fluxerr * 2/3\n\nAfter all these steps we got a model with about 0.3 CV and 0.75 public LB score (with Olivier' algorithm for class 99).",
      "votes": 11,
      "replies": [
        {
          "id": 446319,
          "postDate": "2018-12-27T21:42:24.977Z",
          "content": "<p>Thanks a lot for shairng.  Your RNN seems simpler than the one we used, yet yields better results. One reason may be that you used far less meta data features than us.</p>",
          "rawMarkdown": "Thanks a lot for shairng.  Your RNN seems simpler than the one we used, yet yields better results. One reason may be that you used far less meta data features than us.",
          "votes": 1
        },
        {
          "id": 446342,
          "postDate": "2018-12-27T22:46:17.373Z",
          "content": "<p><code>emitted wavelength and passband (with embedding layer)</code> \nwow, using wavelength and passband as input is a really interesting idea! I think that's a big difference between your model and yuval's model. <br>\n<code>When preparing data for RNN we converted all time and wavelength related features to not depend on red shift. Namely time and wavelength are divided by (hostgalphotoz + 1)</code> \nI think maybe you are the only one who succeeded in wavelength correction with dividing by  (z + 1)!<br>\n<code>Drop 30% of measurements</code>\nWow, it's same as what yuval did! <br>\n<code>After all these steps we got a model with about 0.3 CV and 0.75 public LB score (with Olivier' algorithm for class 99).</code> \nIt's really terrible to get to 0.75 without special class99 handling. It's much better than yuval's score and Kun Hao Yeh's score, which means your NN model is the best of all of NN models in this competition.<br>\nOn the whole, your model is really clean (I know your team would have been 1st without class99 handling), simple and beautiful :)\nCongrats Mike, I really enjoyed competing with you!</p>",
          "rawMarkdown": "`emitted wavelength and passband (with embedding layer)` \nwow, using wavelength and passband as input is a really interesting idea! I think that's a big difference between your model and yuval's model. <br>\n`When preparing data for RNN we converted all time and wavelength related features to not depend on red shift. Namely time and wavelength are divided by (hostgalphotoz + 1)` \nI think maybe you are the only one who succeeded in wavelength correction with dividing by  (z + 1)!<br>\n`Drop 30% of measurements`\nWow, it's same as what yuval did! <br>\n`After all these steps we got a model with about 0.3 CV and 0.75 public LB score (with Olivier' algorithm for class 99).` \nIt's really terrible to get to 0.75 without special class99 handling. It's much better than yuval's score and Kun Hao Yeh's score, which means your NN model is the best of all of NN models in this competition.<br>\nOn the whole, your model is really clean (I know your team would have been 1st without class99 handling), simple and beautiful :)\nCongrats Mike, I really enjoyed competing with you!",
          "votes": 1
        },
        {
          "id": 446625,
          "postDate": "2018-12-28T11:59:04.377Z",
          "content": "<blockquote>\n  <p>When preparing data for RNN we converted all time and wavelength related features to not depend on red shift. </p>\n</blockquote>\n\n<p>This is probably extremely important.</p>",
          "rawMarkdown": "&gt; When preparing data for RNN we converted all time and wavelength related features to not depend on red shift. \n\nThis is probably extremely important."
        },
        {
          "id": 446690,
          "postDate": "2018-12-28T14:13:11.207Z",
          "content": "<p>This step itself doesn't help much but it allows to augmentate time series data simultaneously with red shift that significantly reduces overfitting.</p>",
          "rawMarkdown": "This step itself doesn't help much but it allows to augmentate time series data simultaneously with red shift that significantly reduces overfitting.",
          "votes": 1
        },
        {
          "id": 446742,
          "postDate": "2018-12-28T16:01:02.343Z",
          "content": "<p>I was experimenting with dividing wavelength by (1+z) but hadn't realized that time needed to be adjusted as well! Thanks for sharing.</p>",
          "rawMarkdown": "I was experimenting with dividing wavelength by (1+z) but hadn't realized that time needed to be adjusted as well! Thanks for sharing."
        }
      ]
    },
    {
      "id": 441052,
      "postDate": "2018-12-18T08:11:56.687Z",
      "content": "<p>Congratualtions for your 2nd place ! You've been consistently in the top teams during the competition.</p>\n\n<p>Thanks for sharing.</p>",
      "rawMarkdown": "Congratualtions for your 2nd place ! You've been consistently in the top teams during the competition.\n\nThanks for sharing.",
      "votes": 1
    },
    {
      "id": 445474,
      "postDate": "2018-12-26T14:37:37.377Z",
      "content": "<p>Congratulations! Thanks for your  kindly sharing</p>",
      "rawMarkdown": "Congratulations! Thanks for your  kindly sharing"
    },
    {
      "id": 442873,
      "postDate": "2018-12-20T16:17:23.913Z",
      "content": "<p>Congratulations and thanks for sharing. What I find interesting is that this solution does not contain any curve fitting or GP that were essential is some other top solutions. Or is something similar contained (like autoencoding in the NNs)? </p>",
      "rawMarkdown": "Congratulations and thanks for sharing. What I find interesting is that this solution does not contain any curve fitting or GP that were essential is some other top solutions. Or is something similar contained (like autoencoding in the NNs)? ",
      "replies": [
        {
          "id": 442881,
          "postDate": "2018-12-20T16:26:29.720Z",
          "content": "<p>I was about to answer this but realized I better wait for Mike to respond because he built the NN models and understands them much better than me. Unfortunately he's traveling now and won't be able to post a description of his amazing solution until next week. </p>",
          "rawMarkdown": "I was about to answer this but realized I better wait for Mike to respond because he built the NN models and understands them much better than me. Unfortunately he's traveling now and won't be able to post a description of his amazing solution until next week. "
        }
      ]
    },
    {
      "id": 442322,
      "postDate": "2018-12-19T19:49:16.367Z",
      "content": "<p>Congratulations! Using simple GBDT model(depth = 2) as meta model looks like a very good idea, I should have tried it! </p>",
      "rawMarkdown": "Congratulations! Using simple GBDT model(depth = 2) as meta model looks like a very good idea, I should have tried it! "
    },
    {
      "id": 441318,
      "postDate": "2018-12-18T14:41:27.907Z",
      "content": "<p>Congrats and thanks for sharing the solution...</p>",
      "rawMarkdown": "Congrats and thanks for sharing the solution..."
    },
    {
      "id": 441078,
      "postDate": "2018-12-18T08:33:45.070Z",
      "content": "<p>I also found NN was the strongest then (for me) genetic programming came second and lgb third.  I didn't do any feature engineering more than stealing oliviers stuff as I was using this competition to finally develop softmax for multi class classification for Genetic Programming - I am relieved that it worked so well.</p>\n\n<p>Congrats on your approach.</p>",
      "rawMarkdown": "I also found NN was the strongest then (for me) genetic programming came second and lgb third.  I didn't do any feature engineering more than stealing oliviers stuff as I was using this competition to finally develop softmax for multi class classification for Genetic Programming - I am relieved that it worked so well.\n\nCongrats on your approach."
    },
    {
      "id": 441050,
      "postDate": "2018-12-18T08:05:15.280Z",
      "content": "<p>Thanks for sharing and congrats for the result!  Re this:</p>\n\n<blockquote>\n  <p>the documentation says +- 3 sigmas</p>\n</blockquote>\n\n<p>It is 3 sigmas from background flux.  But we don't have that background flux.  But we could use it the other way round: compute the background flux using ddf cutoff value.  This can yield to better magnitude estimates.  I didn't had time to implement it however.</p>",
      "rawMarkdown": "Thanks for sharing and congrats for the result!  Re this:\n\n&gt; the documentation says +- 3 sigmas\n\nIt is 3 sigmas from background flux.  But we don't have that background flux.  But we could use it the other way round: compute the background flux using ddf cutoff value.  This can yield to better magnitude estimates.  I didn't had time to implement it however.",
      "replies": [
        {
          "id": 441071,
          "postDate": "2018-12-18T08:25:45.783Z",
          "content": "<p>And what does an 'undetected' observation mean? The object could not be detected from the background noise, and yet we still have flux data for that observation? Am I the only one confused by this?</p>",
          "rawMarkdown": "And what does an 'undetected' observation mean? The object could not be detected from the background noise, and yet we still have flux data for that observation? Am I the only one confused by this?",
          "votes": 2
        },
        {
          "id": 441083,
          "postDate": "2018-12-18T08:39:16.280Z",
          "content": "<p>My undertsanding is: </p>\n\n<ol>\n<li>They compute flux std</li>\n<li>They subtract background from all flux</li>\n<li>Detected is when resulting flux is above 3 std.</li>\n</ol>",
          "rawMarkdown": "My undertsanding is: \n\n 1. They compute flux std\n 2. They subtract background from all flux\n 3. Detected is when resulting flux is above 3 std."
        },
        {
          "id": 441088,
          "postDate": "2018-12-18T08:46:32.237Z",
          "content": "<p>How does flux_err work in this formula? The detected flag is very highly correlated with abs(flux)/flux_err&gt;5.</p>",
          "rawMarkdown": "How does flux_err work in this formula? The detected flag is very highly correlated with abs(flux)/flux_err&gt;5."
        },
        {
          "id": 441317,
          "postDate": "2018-12-18T14:39:12.090Z",
          "content": "<p>fux err is not use for computing detected IMHO.  From the data description:</p>\n\n<blockquote>\n  <p>If 1, the object's brightness is significantly different at the 3-sigma level relative to the reference template. Only objects with at least 2 detections are included in the dataset. Boolean</p>\n</blockquote>\n\n<p>There is no mention to flux error at all.  Just the reference template (what I called background).</p>",
          "rawMarkdown": "fux err is not use for computing detected IMHO.  From the data description:\n\n&gt; If 1, the object's brightness is significantly different at the 3-sigma level relative to the reference template. Only objects with at least 2 detections are included in the dataset. Boolean\n\nThere is no mention to flux error at all.  Just the reference template (what I called background)."
        },
        {
          "id": 441374,
          "postDate": "2018-12-18T15:42:53.900Z",
          "content": "<p>If this were true, then (per object/passband) the detected flag would depend solely on the flux  values since the reference template per object is static, yes? So how would you explain the following records?</p>\n\n<pre><code>\nobject_id   passband    flux    flux_err    detected\n713         0   -14.735178  2.326417    0\n713         0   -13.083604  2.663738    0\n713         0   -12.353376  2.357691    1\n713         0   -12.232555  1.708795    0\n713         0   -12.148479  2.243120    0\n713         0   -11.829331  2.358846    0\n713         0   -11.605895  1.778605    1\n713         0   -11.340659  1.930082    1\n713         0   -10.934606  2.143276    1\n713         0   -10.828177  1.470152    1\n713         0   -10.602926  1.838902    1\n713         0   -10.165054  1.726118    1\n713         0   -10.050170  3.275514    0\n713         0   -9.363182   3.042286    0\n713         0   -9.289350   1.992813    1\n</code></pre>",
          "rawMarkdown": "If this were true, then (per object/passband) the detected flag would depend solely on the flux  values since the reference template per object is static, yes? So how would you explain the following records?\n\n<pre><code>\nobject_id\tpassband\tflux\tflux_err\tdetected\n713\t\t\t0\t-14.735178\t2.326417\t0\n713\t\t\t0\t-13.083604\t2.663738\t0\n713\t\t\t0\t-12.353376\t2.357691\t1\n713\t\t\t0\t-12.232555\t1.708795\t0\n713\t\t\t0\t-12.148479\t2.243120\t0\n713\t\t\t0\t-11.829331\t2.358846\t0\n713\t\t\t0\t-11.605895\t1.778605\t1\n713\t\t\t0\t-11.340659\t1.930082\t1\n713\t\t\t0\t-10.934606\t2.143276\t1\n713\t\t\t0\t-10.828177\t1.470152\t1\n713\t\t\t0\t-10.602926\t1.838902\t1\n713\t\t\t0\t-10.165054\t1.726118\t1\n713\t\t\t0\t-10.050170\t3.275514\t0\n713\t\t\t0\t-9.363182\t3.042286\t0\n713\t\t\t0\t-9.289350\t1.992813\t1\n</code></pre>",
          "votes": 2
        },
        {
          "id": 441406,
          "postDate": "2018-12-18T16:17:05.180Z",
          "content": "<p>Indeed, this is at odd with what the data description says.  I guess we need an organizer answer here.</p>",
          "rawMarkdown": "Indeed, this is at odd with what the data description says.  I guess we need an organizer answer here."
        },
        {
          "id": 441501,
          "postDate": "2018-12-18T18:04:36.070Z",
          "content": "<p>In my understanding/guess, the flux and flux err calculation is based on a distribution. So flux is the mean and flux_err is the std of the distribution. This is because in the image of one measurement, we have an area of pixels which is identified as an object. So the flux becomes a distribution.</p>",
          "rawMarkdown": "In my understanding/guess, the flux and flux err calculation is based on a distribution. So flux is the mean and flux_err is the std of the distribution. This is because in the image of one measurement, we have an area of pixels which is identified as an object. So the flux becomes a distribution."
        },
        {
          "id": 441526,
          "postDate": "2018-12-18T18:36:15.957Z",
          "content": "<p>flux err is NOT the std of the flux distribution.  It is the std of flux measurement error.  I asked and got a response from organizers, somewhere in the forum.</p>",
          "rawMarkdown": "flux err is NOT the std of the flux distribution.  It is the std of flux measurement error.  I asked and got a response from organizers, somewhere in the forum."
        },
        {
          "id": 442175,
          "postDate": "2018-12-19T15:20:13.933Z",
          "content": "<p>From my knowledge of how LSST works, I can infer what detected means (the organizers would know better of course). \nOn individual exposures, you may not see a star (&lt;3sigmas), but you know it exists from other exposures. In which case you get back to the exposure where you didn't see it, but as you now know its position, you can extract the flux as if something was there. In fact there is information in this ~0 flux, as the average of undetected fluxes can amount to a significant detection.\nThen, the challenges presented the data with respect to a reference flux. So the 0 which is presented can be either a detected value then put to 0 as it is equal to the reference, or a non-existent object, where the reference is indeed 0.\nI think that not including the reference flux from the metadata was an interesting decision made by the challenge organizers: for some classes, it is not well known, and may have biassed the results of the challenge towards the assumptions about the environment of the objects, and not towards the sky as it is in reality. So it was more interesting to see what you could achieve without this information - and you did great indeed !</p>",
          "rawMarkdown": "From my knowledge of how LSST works, I can infer what detected means (the organizers would know better of course). \nOn individual exposures, you may not see a star (&lt;3sigmas), but you know it exists from other exposures. In which case you get back to the exposure where you didn't see it, but as you now know its position, you can extract the flux as if something was there. In fact there is information in this ~0 flux, as the average of undetected fluxes can amount to a significant detection.\nThen, the challenges presented the data with respect to a reference flux. So the 0 which is presented can be either a detected value then put to 0 as it is equal to the reference, or a non-existent object, where the reference is indeed 0.\nI think that not including the reference flux from the metadata was an interesting decision made by the challenge organizers: for some classes, it is not well known, and may have biassed the results of the challenge towards the assumptions about the environment of the objects, and not towards the sky as it is in reality. So it was more interesting to see what you could achieve without this information - and you did great indeed !\n",
          "votes": 2
        },
        {
          "id": 442215,
          "postDate": "2018-12-19T16:14:54.380Z",
          "content": "<p>Many, thanks, this is my understanding (roughly speaking) as well.  But how could detected fux not be the highest ones then?  Silogram shared data where detected fluxes are not the highest ones.  </p>",
          "rawMarkdown": "Many, thanks, this is my understanding (roughly speaking) as well.  But how could detected fux not be the highest ones then?  Silogram shared data where detected fluxes are not the highest ones.  "
        },
        {
          "id": 442376,
          "postDate": "2018-12-19T21:54:01.177Z",
          "content": "<p>My theory is that the simulator starts with idealized curves for each class and then adds noise to mimic the data they'll actually get. I think the detected flag was set before the addition of at least part of the noise, so it's really a 'leak'. But that's just a theory. I would love to get some confirmation from the organizers.</p>",
          "rawMarkdown": "My theory is that the simulator starts with idealized curves for each class and then adds noise to mimic the data they'll actually get. I think the detected flag was set before the addition of at least part of the noise, so it's really a 'leak'. But that's just a theory. I would love to get some confirmation from the organizers."
        },
        {
          "id": 443468,
          "postDate": "2018-12-21T17:42:20.567Z",
          "content": "<p>There were also detected == 0 points in nova huge flux outbreaks... something I completely do not understand... I thought it might also depend on weather conditions, so they may have introduced some detected jitter... </p>\n\n<p>Indeed, some more info from organizers would be helpful. Detected == 1 magnitude features worked better, and i still do not fully understand why </p>",
          "rawMarkdown": "There were also detected == 0 points in nova huge flux outbreaks... something I completely do not understand... I thought it might also depend on weather conditions, so they may have introduced some detected jitter... \n\nIndeed, some more info from organizers would be helpful. Detected == 1 magnitude features worked better, and i still do not fully understand why "
        },
        {
          "id": 462456,
          "postDate": "2019-01-28T10:30:33.500Z",
          "content": "<p>Hi CPMP, Silogram,</p>\n\n<p>Few things to reply to different parts of the conversation that I hope will clarify things:\n1. sigma refers to fluxerr.  The std of the flux distribution itself is not used anywhere. \n2. It is possible that the detected fluxes were not the highest ones in the light curve because the sigma is the not the same for all observations. The fluxerr depends on the brightness of the source itself (weakly) and the brightness of the sky in the band of the observation, at the time of the observation. The time of the observation is important: The most important dependence is if the time is closer to twilight rather than midnight, or a time when the moon is nearby (or brighter).\n3. abs(flux) / fluxerr &gt; threshold  is what should correspond to detections. So, while Silogram uses a threshold of 5, I would expect that to correlate pretty well.</p>\n\n<p>Thanks,\nRahul (PLASTICC Organizer)</p>",
          "rawMarkdown": "Hi CPMP, Silogram,\n\nFew things to reply to different parts of the conversation that I hope will clarify things:\n1. sigma refers to fluxerr.  The std of the flux distribution itself is not used anywhere. \n2. It is possible that the detected fluxes were not the highest ones in the light curve because the sigma is the not the same for all observations. The fluxerr depends on the brightness of the source itself (weakly) and the brightness of the sky in the band of the observation, at the time of the observation. The time of the observation is important: The most important dependence is if the time is closer to twilight rather than midnight, or a time when the moon is nearby (or brighter).\n3. abs(flux) / fluxerr &gt; threshold  is what should correspond to detections. So, while Silogram uses a threshold of 5, I would expect that to correlate pretty well.\n\nThanks,\nRahul (PLASTICC Organizer)",
          "votes": 2
        },
        {
          "id": 462457,
          "postDate": "2019-01-28T10:33:44.347Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 441169,
      "postDate": "2018-12-18T11:10:09.787Z",
      "content": "<p>thanks, master!!</p>",
      "rawMarkdown": "thanks, master!!"
    }
  ],
  "comments": [
    {
      "id": 446148,
      "author_name": "Mike",
      "author_url": "",
      "post_date": "2018-12-27T15:37:05.043000",
      "content": "<p>I have added a kernel with our RNN model:\n<a href=\"https://www.kaggle.com/zerrxy/plasticc-rnn\">https://www.kaggle.com/zerrxy/plasticc-rnn</a></p>\n\n<p>Here is a short description:</p>\n\n<ol>\n<li>NN model is a simple one layer Bidirectional GRU (we used 80 - 160 units) with global max pooling. Then max pooling output is combined with meta data and followed by 3 additional dense layers and softmax activation.</li>\n<li>Input data for RNN consists of flux, flux_err, intervals between measurements, emitted wavelength and passband (with embedding layer). flux and fluxerr are divided by a value of (fluxmax - fluxmin) for each object.</li>\n<li>Meta data consists of hostgalphotoz, hostgalphotozerr, ddf, mwebv and log2 of flux range.</li>\n<li>When preparing data for RNN we converted all time and wavelength related features to not depend on red shift. Namely time and wavelength are divided by (hostgalphotoz + 1)</li>\n<li>Augmentation works well for this model. The following steps were used:\n<ul><li>Drop 30% of measurements</li>\n<li>Modify red shift using normal distribution with sigma = hostgalphotozerr * 2/3. When changing red shift, all time and wavelength related features are also modified accordingly.</li>\n<li>Modify flux using normal distribution with sigma = fluxerr * 2/3</li></ul></li>\n</ol>\n\n<p>After all these steps we got a model with about 0.3 CV and 0.75 public LB score (with Olivier' algorithm for class 99).</p>",
      "votes": 11,
      "replies": [
        {
          "id": 446319,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-27T21:42:24.977000",
          "content": "<p>Thanks a lot for shairng.  Your RNN seems simpler than the one we used, yet yields better results. One reason may be that you used far less meta data features than us.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 446342,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2018-12-27T22:46:17.373000",
          "content": "<p><code>emitted wavelength and passband (with embedding layer)</code> \nwow, using wavelength and passband as input is a really interesting idea! I think that's a big difference between your model and yuval's model. <br>\n<code>When preparing data for RNN we converted all time and wavelength related features to not depend on red shift. Namely time and wavelength are divided by (hostgalphotoz + 1)</code> \nI think maybe you are the only one who succeeded in wavelength correction with dividing by  (z + 1)!<br>\n<code>Drop 30% of measurements</code>\nWow, it's same as what yuval did! <br>\n<code>After all these steps we got a model with about 0.3 CV and 0.75 public LB score (with Olivier' algorithm for class 99).</code> \nIt's really terrible to get to 0.75 without special class99 handling. It's much better than yuval's score and Kun Hao Yeh's score, which means your NN model is the best of all of NN models in this competition.<br>\nOn the whole, your model is really clean (I know your team would have been 1st without class99 handling), simple and beautiful :)\nCongrats Mike, I really enjoyed competing with you!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 446625,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-28T11:59:04.377000",
          "content": "<blockquote>\n  <p>When preparing data for RNN we converted all time and wavelength related features to not depend on red shift. </p>\n</blockquote>\n\n<p>This is probably extremely important.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 446690,
          "author_name": "Mike",
          "author_url": "",
          "post_date": "2018-12-28T14:13:11.207000",
          "content": "<p>This step itself doesn't help much but it allows to augmentate time series data simultaneously with red shift that significantly reduces overfitting.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 446742,
          "author_name": "S D",
          "author_url": "",
          "post_date": "2018-12-28T16:01:02.343000",
          "content": "<p>I was experimenting with dividing wavelength by (1+z) but hadn't realized that time needed to be adjusted as well! Thanks for sharing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 441052,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2018-12-18T08:11:56.687000",
      "content": "<p>Congratualtions for your 2nd place ! You've been consistently in the top teams during the competition.</p>\n\n<p>Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 445474,
      "author_name": "Kun-Lin Lee",
      "author_url": "",
      "post_date": "2018-12-26T14:37:37.377000",
      "content": "<p>Congratulations! Thanks for your  kindly sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 442873,
      "author_name": "Helgi",
      "author_url": "",
      "post_date": "2018-12-20T16:17:23.913000",
      "content": "<p>Congratulations and thanks for sharing. What I find interesting is that this solution does not contain any curve fitting or GP that were essential is some other top solutions. Or is something similar contained (like autoencoding in the NNs)? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 442881,
          "author_name": "Silogram",
          "author_url": "",
          "post_date": "2018-12-20T16:26:29.720000",
          "content": "<p>I was about to answer this but realized I better wait for Mike to respond because he built the NN models and understands them much better than me. Unfortunately he's traveling now and won't be able to post a description of his amazing solution until next week. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 442322,
      "author_name": "mamas",
      "author_url": "",
      "post_date": "2018-12-19T19:49:16.367000",
      "content": "<p>Congratulations! Using simple GBDT model(depth = 2) as meta model looks like a very good idea, I should have tried it! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 441318,
      "author_name": "mrxnew",
      "author_url": "",
      "post_date": "2018-12-18T14:41:27.907000",
      "content": "<p>Congrats and thanks for sharing the solution...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 441078,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2018-12-18T08:33:45.070000",
      "content": "<p>I also found NN was the strongest then (for me) genetic programming came second and lgb third.  I didn't do any feature engineering more than stealing oliviers stuff as I was using this competition to finally develop softmax for multi class classification for Genetic Programming - I am relieved that it worked so well.</p>\n\n<p>Congrats on your approach.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 441050,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-12-18T08:05:15.280000",
      "content": "<p>Thanks for sharing and congrats for the result!  Re this:</p>\n\n<blockquote>\n  <p>the documentation says +- 3 sigmas</p>\n</blockquote>\n\n<p>It is 3 sigmas from background flux.  But we don't have that background flux.  But we could use it the other way round: compute the background flux using ddf cutoff value.  This can yield to better magnitude estimates.  I didn't had time to implement it however.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 441071,
          "author_name": "Silogram",
          "author_url": "",
          "post_date": "2018-12-18T08:25:45.783000",
          "content": "<p>And what does an 'undetected' observation mean? The object could not be detected from the background noise, and yet we still have flux data for that observation? Am I the only one confused by this?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 441083,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-18T08:39:16.280000",
          "content": "<p>My undertsanding is: </p>\n\n<ol>\n<li>They compute flux std</li>\n<li>They subtract background from all flux</li>\n<li>Detected is when resulting flux is above 3 std.</li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441088,
          "author_name": "Silogram",
          "author_url": "",
          "post_date": "2018-12-18T08:46:32.237000",
          "content": "<p>How does flux_err work in this formula? The detected flag is very highly correlated with abs(flux)/flux_err&gt;5.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441317,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-18T14:39:12.090000",
          "content": "<p>fux err is not use for computing detected IMHO.  From the data description:</p>\n\n<blockquote>\n  <p>If 1, the object's brightness is significantly different at the 3-sigma level relative to the reference template. Only objects with at least 2 detections are included in the dataset. Boolean</p>\n</blockquote>\n\n<p>There is no mention to flux error at all.  Just the reference template (what I called background).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441374,
          "author_name": "Silogram",
          "author_url": "",
          "post_date": "2018-12-18T15:42:53.900000",
          "content": "<p>If this were true, then (per object/passband) the detected flag would depend solely on the flux  values since the reference template per object is static, yes? So how would you explain the following records?</p>\n\n<pre><code>\nobject_id   passband    flux    flux_err    detected\n713         0   -14.735178  2.326417    0\n713         0   -13.083604  2.663738    0\n713         0   -12.353376  2.357691    1\n713         0   -12.232555  1.708795    0\n713         0   -12.148479  2.243120    0\n713         0   -11.829331  2.358846    0\n713         0   -11.605895  1.778605    1\n713         0   -11.340659  1.930082    1\n713         0   -10.934606  2.143276    1\n713         0   -10.828177  1.470152    1\n713         0   -10.602926  1.838902    1\n713         0   -10.165054  1.726118    1\n713         0   -10.050170  3.275514    0\n713         0   -9.363182   3.042286    0\n713         0   -9.289350   1.992813    1\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 441406,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-18T16:17:05.180000",
          "content": "<p>Indeed, this is at odd with what the data description says.  I guess we need an organizer answer here.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441501,
          "author_name": "lucaskg",
          "author_url": "",
          "post_date": "2018-12-18T18:04:36.070000",
          "content": "<p>In my understanding/guess, the flux and flux err calculation is based on a distribution. So flux is the mean and flux_err is the std of the distribution. This is because in the image of one measurement, we have an area of pixels which is identified as an object. So the flux becomes a distribution.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441526,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-18T18:36:15.957000",
          "content": "<p>flux err is NOT the std of the flux distribution.  It is the std of flux measurement error.  I asked and got a response from organizers, somewhere in the forum.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442175,
          "author_name": "Manu Gangler",
          "author_url": "",
          "post_date": "2018-12-19T15:20:13.933000",
          "content": "<p>From my knowledge of how LSST works, I can infer what detected means (the organizers would know better of course). \nOn individual exposures, you may not see a star (&lt;3sigmas), but you know it exists from other exposures. In which case you get back to the exposure where you didn't see it, but as you now know its position, you can extract the flux as if something was there. In fact there is information in this ~0 flux, as the average of undetected fluxes can amount to a significant detection.\nThen, the challenges presented the data with respect to a reference flux. So the 0 which is presented can be either a detected value then put to 0 as it is equal to the reference, or a non-existent object, where the reference is indeed 0.\nI think that not including the reference flux from the metadata was an interesting decision made by the challenge organizers: for some classes, it is not well known, and may have biassed the results of the challenge towards the assumptions about the environment of the objects, and not towards the sky as it is in reality. So it was more interesting to see what you could achieve without this information - and you did great indeed !</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 442215,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-12-19T16:14:54.380000",
          "content": "<p>Many, thanks, this is my understanding (roughly speaking) as well.  But how could detected fux not be the highest ones then?  Silogram shared data where detected fluxes are not the highest ones.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442376,
          "author_name": "Silogram",
          "author_url": "",
          "post_date": "2018-12-19T21:54:01.177000",
          "content": "<p>My theory is that the simulator starts with idealized curves for each class and then adds noise to mimic the data they'll actually get. I think the detected flag was set before the addition of at least part of the noise, so it's really a 'leak'. But that's just a theory. I would love to get some confirmation from the organizers.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 443468,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-12-21T17:42:20.567000",
          "content": "<p>There were also detected == 0 points in nova huge flux outbreaks... something I completely do not understand... I thought it might also depend on weather conditions, so they may have introduced some detected jitter... </p>\n\n<p>Indeed, some more info from organizers would be helpful. Detected == 1 magnitude features worked better, and i still do not fully understand why </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 462456,
          "author_name": "rbiswas4",
          "author_url": "",
          "post_date": "2019-01-28T10:30:33.500000",
          "content": "<p>Hi CPMP, Silogram,</p>\n\n<p>Few things to reply to different parts of the conversation that I hope will clarify things:\n1. sigma refers to fluxerr.  The std of the flux distribution itself is not used anywhere. \n2. It is possible that the detected fluxes were not the highest ones in the light curve because the sigma is the not the same for all observations. The fluxerr depends on the brightness of the source itself (weakly) and the brightness of the sky in the band of the observation, at the time of the observation. The time of the observation is important: The most important dependence is if the time is closer to twilight rather than midnight, or a time when the moon is nearby (or brighter).\n3. abs(flux) / fluxerr &gt; threshold  is what should correspond to detections. So, while Silogram uses a threshold of 5, I would expect that to correlate pretty well.</p>\n\n<p>Thanks,\nRahul (PLASTICC Organizer)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 462457,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-28T10:33:44.347000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 441169,
      "author_name": "joxemi",
      "author_url": "",
      "post_date": "2018-12-18T11:10:09.787000",
      "content": "<p>thanks, master!!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "441048": "First, thanks to Kaggle and the PLAsTiCC organizers for presenting such an interesting challenge, and congratulations to all the top teams, especially Kyle with his amazing performance.\n\nI haven't looked at all the other solution write-ups yet but there were clearly a lot of different ways to tackle this problem effectively. For us, NNs worked much better than LGB models. Our final ensemble included 9 models, 7 of which were NNs that scored as low as 0.75x individually. The two LGB models were much weaker, each scoring about 0.90x, but they did provide a bit of diversity to the ensemble.\n\nMy teammate Mike built the NN models so I'll let him describe them separately. What follows is some more general notes about our overall solution:\n\n1. The 'Gap'. There was quite a bit of chatter about lowering the gap between CV and LB scores, most of which we ignored because our LB scores were tracking our CV scores very consistently, with a gap of 0.43-0.44. Our final solution, which scored 0.694 on the LB had a CV score of 0.264 (0.430 gap). In hindsight, it may have been a mistake not to look more closely at the gap. We really didn't pay much attention to the differences between the train and test sets until the final few days of the competition.\n\n2. 'Class_99'. Especially in the final week, we did quite a bit of LB probing to come up with a better way to predict Class_99. In the end we settled on the following alorithm:\n\ta) Set Class_99 to 1-top_prediction per row\n\tb) Move class_99 closer to 0.14 (for extra-galactic objects) and 0.014 (for galactic objects) with the following \t   formula: class_99 = (2*class_99 + 0.14) / 3\nThis improved our scores modestly (about 0.005).   \n\n3. Ensemble. The final ensemble was a very shallow (max_depth=2) LGB model that was very effective. \n\n4. 'Detected' Flag. I still don't understand what this means or how it was calculated, but it's clearly very significant. As Kyle pointed out in a post, it's roughly set when the absolute flux value is 5 times greater than the flux error (the documentation says +- 3 sigmas, which doesn't make much sense to me), but there are many exceptions to this rule. The only thing I can think of is that the flag was set before some random jitter was applied to the flux values. If anyone knows more about this, please tell me because I devoted several days trying to reverse-engineer this feature with no success. That said, I did find that this ratio -- absolute flux/flux_error -- to be very significant in my models. One of the LGB models generates features for 4 different sets of observations -- all, detected only, flux error ratio &gt;3, flux error ratio &gt; 4. Because the NN model was able to continuously process the flux and flux error values in parallel, this relationship did not need to be manually specified, one reason I suspect that the NN model performed so well.\n\n5. Hostgal_specz pseudo-labeling. One thing that worked for both the NN and LGB models was to build a separate model to predict the hostgal_specz values (using both the training set and the test set objects that have this value), and then using oof predictions for these values in our models.\n\n6. Flux adjustments. For the LGB model, I found that adjusting the the flux values with the simple formula flux *= hostgal_photoz improved the results. All efforts to normalize flux values by object and/or passband did not help. In the NN models, on the other hand, Mike found a very effective normalization strategy.\n\n7. MJD adjustments. Adjustments to MJD based on redshift and wavelength did not help the LGB model.\n\n8. Augmentation. We tried a variety of different augmentation methods. Augmenting the training set with random variations helped the NN models but did not help the LGB models. We also tried augmentation with pseudo-labeled test objects but this did not help. As I mentioned above, we weren't really focused on the differences between the train and test sets. In the final days of the competition, I realized that we should be augmenting in such a way that the train distribution would more closely match the test distribution, but we didn't have time to try this idea.",
    "446148": "I have added a kernel with our RNN model:\nhttps://www.kaggle.com/zerrxy/plasticc-rnn\n\nHere is a short description:\n\n1. NN model is a simple one layer Bidirectional GRU (we used 80 - 160 units) with global max pooling. Then max pooling output is combined with meta data and followed by 3 additional dense layers and softmax activation.\n2. Input data for RNN consists of flux, flux_err, intervals between measurements, emitted wavelength and passband (with embedding layer). flux and fluxerr are divided by a value of (fluxmax - fluxmin) for each object.\n3. Meta data consists of hostgalphotoz, hostgalphotozerr, ddf, mwebv and log2 of flux range.\n4. When preparing data for RNN we converted all time and wavelength related features to not depend on red shift. Namely time and wavelength are divided by (hostgalphotoz + 1)\n5. Augmentation works well for this model. The following steps were used:\n- Drop 30% of measurements\n- Modify red shift using normal distribution with sigma = hostgalphotozerr * 2/3. When changing red shift, all time and wavelength related features are also modified accordingly.\n- Modify flux using normal distribution with sigma = fluxerr * 2/3\n\nAfter all these steps we got a model with about 0.3 CV and 0.75 public LB score (with Olivier' algorithm for class 99).",
    "441052": "Congratualtions for your 2nd place ! You've been consistently in the top teams during the competition.\n\nThanks for sharing.",
    "445474": "Congratulations! Thanks for your  kindly sharing",
    "442873": "Congratulations and thanks for sharing. What I find interesting is that this solution does not contain any curve fitting or GP that were essential is some other top solutions. Or is something similar contained (like autoencoding in the NNs)? ",
    "442322": "Congratulations! Using simple GBDT model(depth = 2) as meta model looks like a very good idea, I should have tried it! ",
    "441318": "Congrats and thanks for sharing the solution...",
    "441078": "I also found NN was the strongest then (for me) genetic programming came second and lgb third.  I didn't do any feature engineering more than stealing oliviers stuff as I was using this competition to finally develop softmax for multi class classification for Genetic Programming - I am relieved that it worked so well.\n\nCongrats on your approach.",
    "441050": "Thanks for sharing and congrats for the result!  Re this:\n\n&gt; the documentation says +- 3 sigmas\n\nIt is 3 sigmas from background flux.  But we don't have that background flux.  But we could use it the other way round: compute the background flux using ddf cutoff value.  This can yield to better magnitude estimates.  I didn't had time to implement it however.",
    "441169": "thanks, master!!"
  }
}