{
  "id": 75040,
  "title": "5th Place Partial Solution (RNN)",
  "url": "/competitions/PLAsTiCC-2018/writeups/skz-lost-in-translation-5th-place-partial-solution",
  "author_name": "",
  "post_date": "2018-12-18T04:08:49.248891100Z",
  "votes": 48,
  "comment_count": 11,
  "views": 0,
  "content": "<p>First of all, thanks PlAsTiCC for holding such an interesting competition.\nThanks for the initial probing in the post by <a href=\"/titericz\">@titericz</a> and others, save us all some submissions :D\nThanks for teammates for the nice cooperation, we've learned a lot from each other!\nHere is solution of my parts, my teammates might share theirs too.\nI'll update a github repo later.</p>\n\n<p><strong>Basic Solution</strong></p>\n\n<p>Basically, my parts involoves in building RNN model, since I've already got teammates proficient in lgb :D</p>\n\n<ol>\n<li>For each timestamp, encode with vector composed of flux + flux_err + one-hot encoded passband + mjd-diff</li>\n<li>Manual RNN Architecture Search (gru\\lstm, several rnn layers, dropout, attention, conv1d with avg\\max pooling): \n<ul><li>Training set is small, this is a quick trial for me.</li>\n<li>Best RNN Architecture: Time-Series Branch + Meta Branch.</li>\n<li>Time-Series Branch: 64-bigru + Spatial Dropout (Prob=0.1) + 64-bigru + Attention.</li>\n<li>Meta Branch: Simple Fully Connected NN.</li></ul></li>\n<li>Add Gaussian Noise to have run-time augmentation, so that training won't fit to train set too quickly.</li>\n<li>OOF from RNN with time-series branch only help building better lgb models from <a href=\"/cpmpml\">@cpmpml</a> and <a href=\"/marcuslin\">@marcuslin</a></li>\n<li>Single Model Scores.\n<ul><li>Based on raw features only: Private 0.93560, Public: 0.91956, CV: 0.585</li>\n<li>With features from <a href=\"/marcuslin\">@marcuslin</a>:  Private 0.85371, Public: 0.84372, CV: 0.474</li></ul></li>\n<li>How the model looks like: see appendix.</li>\n</ol>\n\n<p><strong>What I've tried\\researched ?</strong></p>\n\n<ol>\n<li>RNN autoencoder: \n<ul><li>The latent vectors and reconstructed light curves (change different timing sampled) does not help in lgb modeling.</li></ul></li>\n<li>Research for open-classifcation\\unknown class detection problems:\n<ul><li><a href=\"https://arxiv.org/abs/1610.02136\"></a><a href=\"https://arxiv.org/abs/1610.02136\">https://arxiv.org/abs/1610.02136</a>: Use max softmax probability for class99. Not worked for me. Oliver's approach still make sense to me more and perform better.</li>\n<li><a href=\"https://openreview.net/forum?id=ryiAv2xAZ\"></a><a href=\"https://openreview.net/forum?id=ryiAv2xAZ\">https://openreview.net/forum?id=ryiAv2xAZ</a>: \" Use GAN to generate adversarial samples to calibrate the probability. For adversarial samples, make the confidence predicted for known classes the same =&gt; Maximize (1-p0)<em>...</em>(1-p13) =&gt; Maximize prob for class 99. Seems an interesting idea to me. But did not try it eventually, don't think I could code GAN with RNN in time.</li>\n<li><a href=\"https://arxiv.org/abs/1704.03976\"></a><a href=\"https://arxiv.org/abs/1704.03976\">https://arxiv.org/abs/1704.03976</a>: Virtual Adversarial Training. Use some tricks in loss to calibrate the confidence prediction. Still in half-way modifying the loss...</li></ul></li>\n</ol>\n\n<p><strong>What did not work for me ?</strong></p>\n\n<ol>\n<li>Adversarial Classification, predict the probabilities of how the samples are similar to test. Then use the probability as the weight for training set.</li>\n<li>CNN: Conv1d, Conv2d, dilated CNN...</li>\n<li>Using the RNN autoencoder to pretrain time-series branch, and use the pretrained weights as initial weights with warmup for classification model training.</li>\n</ol>",
  "messages": [
    {
      "id": "440885",
      "postDate": "12/18/2018 04:08:49",
      "content": "<p>First of all, thanks PlAsTiCC for holding such an interesting competition.\nThanks for the initial probing in the post by <a href=\"/titericz\">@titericz</a> and others, save us all some submissions :D\nThanks for teammates for the nice cooperation, we've learned a lot from each other!\nHere is solution of my parts, my teammates might share theirs too.\nI'll update a github repo later.</p>\n\n<p><strong>Basic Solution</strong></p>\n\n<p>Basically, my parts involoves in building RNN model, since I've already got teammates proficient in lgb :D</p>\n\n<ol>\n<li>For each timestamp, encode with vector composed of flux + flux_err + one-hot encoded passband + mjd-diff</li>\n<li>Manual RNN Architecture Search (gru\\lstm, several rnn layers, dropout, attention, conv1d with avg\\max pooling): \n<ul><li>Training set is small, this is a quick trial for me.</li>\n<li>Best RNN Architecture: Time-Series Branch + Meta Branch.</li>\n<li>Time-Series Branch: 64-bigru + Spatial Dropout (Prob=0.1) + 64-bigru + Attention.</li>\n<li>Meta Branch: Simple Fully Connected NN.</li></ul></li>\n<li>Add Gaussian Noise to have run-time augmentation, so that training won't fit to train set too quickly.</li>\n<li>OOF from RNN with time-series branch only help building better lgb models from <a href=\"/cpmpml\">@cpmpml</a> and <a href=\"/marcuslin\">@marcuslin</a></li>\n<li>Single Model Scores.\n<ul><li>Based on raw features only: Private 0.93560, Public: 0.91956, CV: 0.585</li>\n<li>With features from <a href=\"/marcuslin\">@marcuslin</a>:  Private 0.85371, Public: 0.84372, CV: 0.474</li></ul></li>\n<li>How the model looks like: see appendix.</li>\n</ol>\n\n<p><strong>What I've tried\\researched ?</strong></p>\n\n<ol>\n<li>RNN autoencoder: \n<ul><li>The latent vectors and reconstructed light curves (change different timing sampled) does not help in lgb modeling.</li></ul></li>\n<li>Research for open-classifcation\\unknown class detection problems:\n<ul><li><a href=\"https://arxiv.org/abs/1610.02136\"></a><a href=\"https://arxiv.org/abs/1610.02136\">https://arxiv.org/abs/1610.02136</a>: Use max softmax probability for class99. Not worked for me. Oliver's approach still make sense to me more and perform better.</li>\n<li><a href=\"https://openreview.net/forum?id=ryiAv2xAZ\"></a><a href=\"https://openreview.net/forum?id=ryiAv2xAZ\">https://openreview.net/forum?id=ryiAv2xAZ</a>: \" Use GAN to generate adversarial samples to calibrate the probability. For adversarial samples, make the confidence predicted for known classes the same =&gt; Maximize (1-p0)<em>...</em>(1-p13) =&gt; Maximize prob for class 99. Seems an interesting idea to me. But did not try it eventually, don't think I could code GAN with RNN in time.</li>\n<li><a href=\"https://arxiv.org/abs/1704.03976\"></a><a href=\"https://arxiv.org/abs/1704.03976\">https://arxiv.org/abs/1704.03976</a>: Virtual Adversarial Training. Use some tricks in loss to calibrate the confidence prediction. Still in half-way modifying the loss...</li></ul></li>\n</ol>\n\n<p><strong>What did not work for me ?</strong></p>\n\n<ol>\n<li>Adversarial Classification, predict the probabilities of how the samples are similar to test. Then use the probability as the weight for training set.</li>\n<li>CNN: Conv1d, Conv2d, dilated CNN...</li>\n<li>Using the RNN autoencoder to pretrain time-series branch, and use the pretrained weights as initial weights with warmup for classification model training.</li>\n</ol>",
      "rawMarkdown": "First of all, thanks PlAsTiCC for holding such an interesting competition.\nThanks for the initial probing in the post by @titericz and others, save us all some submissions :D\nThanks for teammates for the nice cooperation, we've learned a lot from each other!\nHere is solution of my parts, my teammates might share theirs too.\nI'll update a github repo later.\n\n**Basic Solution**\n\nBasically, my parts involoves in building RNN model, since I've already got teammates proficient in lgb :D\n\n1. For each timestamp, encode with vector composed of flux + flux_err + one-hot encoded passband + mjd-diff\n2. Manual RNN Architecture Search (gru\\lstm, several rnn layers, dropout, attention, conv1d with avg\\max pooling): \n   - Training set is small, this is a quick trial for me.\n   - Best RNN Architecture: Time-Series Branch + Meta Branch.\n   - Time-Series Branch: 64-bigru + Spatial Dropout (Prob=0.1) + 64-bigru + Attention.\n   - Meta Branch: Simple Fully Connected NN.\n3. Add Gaussian Noise to have run-time augmentation, so that training won't fit to train set too quickly.\n4. OOF from RNN with time-series branch only help building better lgb models from @cpmpml and @marcuslin\n5. Single Model Scores.\n   - Based on raw features only: Private 0.93560, Public: 0.91956, CV: 0.585\n   - With features from @marcuslin:  Private 0.85371, Public: 0.84372, CV: 0.474\n6. How the model looks like: see appendix.\n\n**What I've tried\\researched ?**\n\n 1. RNN autoencoder: \n   - The latent vectors and reconstructed light curves (change different timing sampled) does not help in lgb modeling.\n 2. Research for open-classifcation\\unknown class detection problems:\n   - https://arxiv.org/abs/1610.02136: Use max softmax probability for class99. Not worked for me. Oliver's approach still make sense to me more and perform better.\n   - https://openreview.net/forum?id=ryiAv2xAZ: \" Use GAN to generate adversarial samples to calibrate the probability. For adversarial samples, make the confidence predicted for known classes the same =&gt; Maximize (1-p0)*...*(1-p13) =&gt; Maximize prob for class 99. Seems an interesting idea to me. But did not try it eventually, don't think I could code GAN with RNN in time.\n   - https://arxiv.org/abs/1704.03976: Virtual Adversarial Training. Use some tricks in loss to calibrate the confidence prediction. Still in half-way modifying the loss...\n\n\n**What did not work for me ?**\n\n 1. Adversarial Classification, predict the probabilities of how the samples are similar to test. Then use the probability as the weight for training set.\n 2. CNN: Conv1d, Conv2d, dilated CNN...\n 3. Using the RNN autoencoder to pretrain time-series branch, and use the pretrained weights as initial weights with warmup for classification model training.",
      "votes": null
    },
    {
      "id": "440904",
      "postDate": "12/18/2018 04:36:35",
      "content": "<p>Great job, this is very interesting! I played around with RNNs at the beginning of the competition, specifically trying to use them as an autoencoder. I found that I couldn't really get the RNN to learn how to use the mjd-diff information properly and that it struggled with things like gaps in the data. Were you able to figure out how to address that problem?</p>",
      "rawMarkdown": "Great job, this is very interesting! I played around with RNNs at the beginning of the competition, specifically trying to use them as an autoencoder. I found that I couldn't really get the RNN to learn how to use the mjd-diff information properly and that it struggled with things like gaps in the data. Were you able to figure out how to address that problem?",
      "votes": null
    },
    {
      "id": "440914",
      "postDate": "12/18/2018 04:55:47",
      "content": "<p>Congrats and thx for sharing. Really amazing that you guys have success in RNN.</p>",
      "rawMarkdown": "Congrats and thx for sharing. Really amazing that you guys have success in RNN.",
      "votes": null
    },
    {
      "id": "440929",
      "postDate": "12/18/2018 05:31:13",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "440953",
      "postDate": "12/18/2018 05:47:43",
      "content": "<p>Congrats! could you please point me to a paper of the RNN you used? Or did you just come up with it? :D</p>",
      "rawMarkdown": "Congrats! could you please point me to a paper of the RNN you used? Or did you just come up with it? :D",
      "votes": null
    },
    {
      "id": "440960",
      "postDate": "12/18/2018 05:54:24",
      "content": "<p>Thanks~ I did notice the data is not sampled uniformly in time...\nSo what I did was \n 1.  Sorted each observations by mjd\n 2. Groupby object id, and calculate the difference of mjd (mjd-difference)\n 3. Add the mjd-difference as one of the dimension in the input vector per timestamp in the time-series branch.</p>\n\n<p>In my understandings, it should somehow provide gaps info to RNN.</p>\n\n<p>A very interesting I've tried is also autoencoder too, where I trained it with original mjd-diff, passbands... and test with different mjd-difference (uniform this time), passbands, and seemingly it could reconstruct how the data should look like according to the given information. Maybe it's showing RNN could digest the given mjd diff, passbands to construct the light curves. Unfortunately, the latent vectors learned\\reconstructed curves did not help us getting better model :( (We did not apply what you've done to get lots of training data!).\nMay I ask how you figure out RNN struggled with gaps in data ? From the reconstructed curves in autoencoder as well?</p>",
      "rawMarkdown": "Thanks~ I did notice the data is not sampled uniformly in time...\nSo what I did was \n 1.  Sorted each observations by mjd\n 2. Groupby object id, and calculate the difference of mjd (mjd-difference)\n 3. Add the mjd-difference as one of the dimension in the input vector per timestamp in the time-series branch.\n\nIn my understandings, it should somehow provide gaps info to RNN.\n\nA very interesting I've tried is also autoencoder too, where I trained it with original mjd-diff, passbands... and test with different mjd-difference (uniform this time), passbands, and seemingly it could reconstruct how the data should look like according to the given information. Maybe it's showing RNN could digest the given mjd diff, passbands to construct the light curves. Unfortunately, the latent vectors learned\\reconstructed curves did not help us getting better model :( (We did not apply what you've done to get lots of training data!).\nMay I ask how you figure out RNN struggled with gaps in data ? From the reconstructed curves in autoencoder as well?",
      "votes": null
    },
    {
      "id": "440966",
      "postDate": "12/18/2018 06:02:39",
      "content": "<p>Congrats to you too!\nFor classification, I came up with it from the experience in the past competition involving NLPs.. lol\nFor autoencoder, I referenced the following for the decoder part: \n 1. Paper: Effective Approaches to Attention-based Neural Machine Translation's Global Attention with Dot-based scoring function\n 2. Someone's code <a href=\"https://github.com/wanasit/katakana/blob/master/notebooks/Attention-based%20Sequence-to-Sequence%20in%20Keras.ipynb\">https://github.com/wanasit/katakana/blob/master/notebooks/Attention-based%20Sequence-to-Sequence%20in%20Keras.ipynb</a>\n 3. <a href=\"/cpmpml\">@cpmpml</a>'s post: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71949\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71949</a>, which contains a link to github + paper as well.  P.S. I noticed 6th place cloned this in gitlab 2 months ago...</p>",
      "rawMarkdown": "Congrats to you too!\nFor classification, I came up with it from the experience in the past competition involving NLPs.. lol\nFor autoencoder, I referenced the following for the decoder part: \n 1. Paper: Effective Approaches to Attention-based Neural Machine Translation's Global Attention with Dot-based scoring function\n 2. Someone's code https://github.com/wanasit/katakana/blob/master/notebooks/Attention-based%20Sequence-to-Sequence%20in%20Keras.ipynb\n 3. @cpmpml's post: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71949, which contains a link to github + paper as well.  P.S. I noticed 6th place cloned this in gitlab 2 months ago...",
      "votes": null
    },
    {
      "id": "440968",
      "postDate": "12/18/2018 06:03:58",
      "content": "<p>That's really helpful. I hope I could upvote more than once!</p>",
      "rawMarkdown": "That's really helpful. I hope I could upvote more than once!",
      "votes": null
    },
    {
      "id": "440985",
      "postDate": "12/18/2018 06:36:32",
      "content": "<p>Teaming with you was great, I would not have got as good RNN as you.  I'm a bit surprised still that RNN could not be more competitive compared to lgb, but they proved great when used for stacking.</p>",
      "rawMarkdown": "Teaming with you was great, I would not have got as good RNN as you.  I'm a bit surprised still that RNN could not be more competitive compared to lgb, but they proved great when used for stacking.",
      "votes": null
    },
    {
      "id": "441027",
      "postDate": "12/18/2018 07:35:11",
      "content": "<p>Thanks, you are an amazing teammate to us too! Really well-deserved as a discussion GM, not only for the information shared on the discussion channel; as teammates, we got lots of useful suggestion from you~</p>",
      "rawMarkdown": "Thanks, you are an amazing teammate to us too! Really well-deserved as a discussion GM, not only for the information shared on the discussion channel; as teammates, we got lots of useful suggestion from you~",
      "votes": null
    },
    {
      "id": "443026",
      "postDate": "12/20/2018 22:49:53",
      "content": "<p>congratulations, and thanks for sharing your solution</p>",
      "rawMarkdown": "congratulations, and thanks for sharing your solution",
      "votes": null
    },
    {
      "id": "443609",
      "postDate": "12/21/2018 23:35:57",
      "content": "<p>Congratulations on your achievement! Thank you for sharing your code with us, this is very helpful.</p>",
      "rawMarkdown": "Congratulations on your achievement! Thank you for sharing your code with us, this is very helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 440904,
      "author_name": "kyleboone",
      "author_url": "",
      "post_date": "12/18/2018 04:36:35",
      "content": "<p>Great job, this is very interesting! I played around with RNNs at the beginning of the competition, specifically trying to use them as an autoencoder. I found that I couldn't really get the RNN to learn how to use the mjd-diff information properly and that it struggled with things like gaps in the data. Were you able to figure out how to address that problem?</p>",
      "votes": null,
      "replies": [
        {
          "id": 440960,
          "author_name": "khyeh0719",
          "author_url": "",
          "post_date": "12/18/2018 05:54:24",
          "content": "<p>Thanks~ I did notice the data is not sampled uniformly in time...\nSo what I did was \n 1.  Sorted each observations by mjd\n 2. Groupby object id, and calculate the difference of mjd (mjd-difference)\n 3. Add the mjd-difference as one of the dimension in the input vector per timestamp in the time-series branch.</p>\n\n<p>In my understandings, it should somehow provide gaps info to RNN.</p>\n\n<p>A very interesting I've tried is also autoencoder too, where I trained it with original mjd-diff, passbands... and test with different mjd-difference (uniform this time), passbands, and seemingly it could reconstruct how the data should look like according to the given information. Maybe it's showing RNN could digest the given mjd diff, passbands to construct the light curves. Unfortunately, the latent vectors learned\\reconstructed curves did not help us getting better model :( (We did not apply what you've done to get lots of training data!).\nMay I ask how you figure out RNN struggled with gaps in data ? From the reconstructed curves in autoencoder as well?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 440914,
      "author_name": "andrew60909",
      "author_url": "",
      "post_date": "12/18/2018 04:55:47",
      "content": "<p>Congrats and thx for sharing. Really amazing that you guys have success in RNN.</p>",
      "votes": null,
      "replies": [
        {
          "id": 440929,
          "author_name": "khyeh0719",
          "author_url": "",
          "post_date": "12/18/2018 05:31:13",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 440953,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "12/18/2018 05:47:43",
      "content": "<p>Congrats! could you please point me to a paper of the RNN you used? Or did you just come up with it? :D</p>",
      "votes": null,
      "replies": [
        {
          "id": 440966,
          "author_name": "khyeh0719",
          "author_url": "",
          "post_date": "12/18/2018 06:02:39",
          "content": "<p>Congrats to you too!\nFor classification, I came up with it from the experience in the past competition involving NLPs.. lol\nFor autoencoder, I referenced the following for the decoder part: \n 1. Paper: Effective Approaches to Attention-based Neural Machine Translation's Global Attention with Dot-based scoring function\n 2. Someone's code <a href=\"https://github.com/wanasit/katakana/blob/master/notebooks/Attention-based%20Sequence-to-Sequence%20in%20Keras.ipynb\">https://github.com/wanasit/katakana/blob/master/notebooks/Attention-based%20Sequence-to-Sequence%20in%20Keras.ipynb</a>\n 3. <a href=\"/cpmpml\">@cpmpml</a>'s post: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71949\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71949</a>, which contains a link to github + paper as well.  P.S. I noticed 6th place cloned this in gitlab 2 months ago...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440968,
          "author_name": "jiweiliu",
          "author_url": "",
          "post_date": "12/18/2018 06:03:58",
          "content": "<p>That's really helpful. I hope I could upvote more than once!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 440985,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/18/2018 06:36:32",
      "content": "<p>Teaming with you was great, I would not have got as good RNN as you.  I'm a bit surprised still that RNN could not be more competitive compared to lgb, but they proved great when used for stacking.</p>",
      "votes": null,
      "replies": [
        {
          "id": 441027,
          "author_name": "khyeh0719",
          "author_url": "",
          "post_date": "12/18/2018 07:35:11",
          "content": "<p>Thanks, you are an amazing teammate to us too! Really well-deserved as a discussion GM, not only for the information shared on the discussion channel; as teammates, we got lots of useful suggestion from you~</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 443026,
      "author_name": "ahmedbaz",
      "author_url": "",
      "post_date": "12/20/2018 22:49:53",
      "content": "<p>congratulations, and thanks for sharing your solution</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 443609,
      "author_name": "aryan5401",
      "author_url": "",
      "post_date": "12/21/2018 23:35:57",
      "content": "<p>Congratulations on your achievement! Thank you for sharing your code with us, this is very helpful.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "440885": "First of all, thanks PlAsTiCC for holding such an interesting competition.\nThanks for the initial probing in the post by @titericz and others, save us all some submissions :D\nThanks for teammates for the nice cooperation, we've learned a lot from each other!\nHere is solution of my parts, my teammates might share theirs too.\nI'll update a github repo later.\n\n**Basic Solution**\n\nBasically, my parts involoves in building RNN model, since I've already got teammates proficient in lgb :D\n\n1. For each timestamp, encode with vector composed of flux + flux_err + one-hot encoded passband + mjd-diff\n2. Manual RNN Architecture Search (gru\\lstm, several rnn layers, dropout, attention, conv1d with avg\\max pooling): \n   - Training set is small, this is a quick trial for me.\n   - Best RNN Architecture: Time-Series Branch + Meta Branch.\n   - Time-Series Branch: 64-bigru + Spatial Dropout (Prob=0.1) + 64-bigru + Attention.\n   - Meta Branch: Simple Fully Connected NN.\n3. Add Gaussian Noise to have run-time augmentation, so that training won't fit to train set too quickly.\n4. OOF from RNN with time-series branch only help building better lgb models from @cpmpml and @marcuslin\n5. Single Model Scores.\n   - Based on raw features only: Private 0.93560, Public: 0.91956, CV: 0.585\n   - With features from @marcuslin:  Private 0.85371, Public: 0.84372, CV: 0.474\n6. How the model looks like: see appendix.\n\n**What I've tried\\researched ?**\n\n 1. RNN autoencoder: \n   - The latent vectors and reconstructed light curves (change different timing sampled) does not help in lgb modeling.\n 2. Research for open-classifcation\\unknown class detection problems:\n   - https://arxiv.org/abs/1610.02136: Use max softmax probability for class99. Not worked for me. Oliver's approach still make sense to me more and perform better.\n   - https://openreview.net/forum?id=ryiAv2xAZ: \" Use GAN to generate adversarial samples to calibrate the probability. For adversarial samples, make the confidence predicted for known classes the same =&gt; Maximize (1-p0)*...*(1-p13) =&gt; Maximize prob for class 99. Seems an interesting idea to me. But did not try it eventually, don't think I could code GAN with RNN in time.\n   - https://arxiv.org/abs/1704.03976: Virtual Adversarial Training. Use some tricks in loss to calibrate the confidence prediction. Still in half-way modifying the loss...\n\n\n**What did not work for me ?**\n\n 1. Adversarial Classification, predict the probabilities of how the samples are similar to test. Then use the probability as the weight for training set.\n 2. CNN: Conv1d, Conv2d, dilated CNN...\n 3. Using the RNN autoencoder to pretrain time-series branch, and use the pretrained weights as initial weights with warmup for classification model training.",
    "440904": "Great job, this is very interesting! I played around with RNNs at the beginning of the competition, specifically trying to use them as an autoencoder. I found that I couldn't really get the RNN to learn how to use the mjd-diff information properly and that it struggled with things like gaps in the data. Were you able to figure out how to address that problem?",
    "440914": "Congrats and thx for sharing. Really amazing that you guys have success in RNN.",
    "440929": "Thanks!",
    "440953": "Congrats! could you please point me to a paper of the RNN you used? Or did you just come up with it? :D",
    "440960": "Thanks~ I did notice the data is not sampled uniformly in time...\nSo what I did was \n 1.  Sorted each observations by mjd\n 2. Groupby object id, and calculate the difference of mjd (mjd-difference)\n 3. Add the mjd-difference as one of the dimension in the input vector per timestamp in the time-series branch.\n\nIn my understandings, it should somehow provide gaps info to RNN.\n\nA very interesting I've tried is also autoencoder too, where I trained it with original mjd-diff, passbands... and test with different mjd-difference (uniform this time), passbands, and seemingly it could reconstruct how the data should look like according to the given information. Maybe it's showing RNN could digest the given mjd diff, passbands to construct the light curves. Unfortunately, the latent vectors learned\\reconstructed curves did not help us getting better model :( (We did not apply what you've done to get lots of training data!).\nMay I ask how you figure out RNN struggled with gaps in data ? From the reconstructed curves in autoencoder as well?",
    "440966": "Congrats to you too!\nFor classification, I came up with it from the experience in the past competition involving NLPs.. lol\nFor autoencoder, I referenced the following for the decoder part: \n 1. Paper: Effective Approaches to Attention-based Neural Machine Translation's Global Attention with Dot-based scoring function\n 2. Someone's code https://github.com/wanasit/katakana/blob/master/notebooks/Attention-based%20Sequence-to-Sequence%20in%20Keras.ipynb\n 3. @cpmpml's post: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71949, which contains a link to github + paper as well.  P.S. I noticed 6th place cloned this in gitlab 2 months ago...",
    "440968": "That's really helpful. I hope I could upvote more than once!",
    "440985": "Teaming with you was great, I would not have got as good RNN as you.  I'm a bit surprised still that RNN could not be more competitive compared to lgb, but they proved great when used for stacking.",
    "441027": "Thanks, you are an amazing teammate to us too! Really well-deserved as a discussion GM, not only for the information shared on the discussion channel; as teammates, we got lots of useful suggestion from you~",
    "443026": "congratulations, and thanks for sharing your solution",
    "443609": "Congratulations on your achievement! Thank you for sharing your code with us, this is very helpful."
  },
  "source": "meta"
}