{
  "id": 75031,
  "title": "Simple GRU Baseline (pub0.938/pri0.970)",
  "url": "/competitions/PLAsTiCC-2018/discussion/75031",
  "author_name": "",
  "post_date": "2018-12-18T02:42:07.966482600Z",
  "votes": 30,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Really great competition &amp; Congrats to everyone!!!  </p>\n\n<p>I just want to share a very simple rnn baseline (in pytorch), kernel is here(<em>results are not stable and maybe lr schedule should be set carefully</em>) -&gt; <a href=\"https://www.kaggle.com/johnfarrell/plasticc-2018-emb-gru\">https://www.kaggle.com/johnfarrell/plasticc-2018-emb-gru</a>   </p>\n\n<p>which only uses 5 raw series: 'mjd', 'flux', 'flux_err', 'detected', 'passband'(embedding dim is 16) &amp; meta features: 'ddf', 'hostgal_photoz', 'hostgal_photoz_err', 'distmod', 'mwebv'. The score result(<a href=\"https://drive.google.com/file/d/1GSkoyQ-kBU0ZpjQmSQ38X7roFl9gGi0m/view?usp=sharing\">training.ipynb source file</a>) is oof0.650/pub0.938/pri0.970.  And it has very low correlation with stats features so stacking works very well with other models using stats features.</p>\n\n<p>Finally, I'm really eager to know other better RNN/CNN/?NN solutions, if you have some please share with me! :)</p>",
  "messages": [
    {
      "id": "440838",
      "postDate": "12/18/2018 02:42:07",
      "content": "<p>Really great competition &amp; Congrats to everyone!!!  </p>\n\n<p>I just want to share a very simple rnn baseline (in pytorch), kernel is here(<em>results are not stable and maybe lr schedule should be set carefully</em>) -&gt; <a href=\"https://www.kaggle.com/johnfarrell/plasticc-2018-emb-gru\">https://www.kaggle.com/johnfarrell/plasticc-2018-emb-gru</a>   </p>\n\n<p>which only uses 5 raw series: 'mjd', 'flux', 'flux_err', 'detected', 'passband'(embedding dim is 16) &amp; meta features: 'ddf', 'hostgal_photoz', 'hostgal_photoz_err', 'distmod', 'mwebv'. The score result(<a href=\"https://drive.google.com/file/d/1GSkoyQ-kBU0ZpjQmSQ38X7roFl9gGi0m/view?usp=sharing\">training.ipynb source file</a>) is oof0.650/pub0.938/pri0.970.  And it has very low correlation with stats features so stacking works very well with other models using stats features.</p>\n\n<p>Finally, I'm really eager to know other better RNN/CNN/?NN solutions, if you have some please share with me! :)</p>",
      "rawMarkdown": "Really great competition &amp; Congrats to everyone!!!  \n\nI just want to share a very simple rnn baseline (in pytorch), kernel is here(*results are not stable and maybe lr schedule should be set carefully*) -&gt; https://www.kaggle.com/johnfarrell/plasticc-2018-emb-gru   \n\nwhich only uses 5 raw series: 'mjd', 'flux', 'flux\\_err', 'detected', 'passband'(embedding dim is 16) &amp; meta features: 'ddf', 'hostgal\\_photoz', 'hostgal\\_photoz\\_err', 'distmod', 'mwebv'. The score result([training.ipynb source file](https://drive.google.com/file/d/1GSkoyQ-kBU0ZpjQmSQ38X7roFl9gGi0m/view?usp=sharing)) is oof0.650/pub0.938/pri0.970.  And it has very low correlation with stats features so stacking works very well with other models using stats features.\n\nFinally, I'm really eager to know other better RNN/CNN/?NN solutions, if you have some please share with me! :)",
      "votes": null
    },
    {
      "id": "440867",
      "postDate": "12/18/2018 03:26:13",
      "content": "<p>Nice. My LSTM solution performs similarly (1.014 LB). It uses BiDirectional LSTM followed by separable depth-wise convolution layer.</p>\n\n<p>One thing I did different was I wrote code that takes the fluxes, and linearly interpolates them so long as the mjd gap&lt;80. This results in \"equispaced\" curves that look more like an 'image' or maybe more appropriately 'sound wave' since we're dealing with 1D convolutions here. The input is more dense though due to this and so it takes foreverrrrrr to run. I also augmented every iteration by adding a normal amount of flux_err to a copy of flux every epoch before performing the interpolation.</p>",
      "rawMarkdown": "Nice. My LSTM solution performs similarly (1.014 LB). It uses BiDirectional LSTM followed by separable depth-wise convolution layer.\n\nOne thing I did different was I wrote code that takes the fluxes, and linearly interpolates them so long as the mjd gap&lt;80. This results in \"equispaced\" curves that look more like an 'image' or maybe more appropriately 'sound wave' since we're dealing with 1D convolutions here. The input is more dense though due to this and so it takes foreverrrrrr to run. I also augmented every iteration by adding a normal amount of flux_err to a copy of flux every epoch before performing the interpolation.",
      "votes": null
    },
    {
      "id": "440872",
      "postDate": "12/18/2018 03:46:41",
      "content": "<p>I also thought about more preprocessing methods like interpolatings but got no time to run in only last 2~3 days so I only applied standardization on 'mjd' (finally, with such setup, 150 epochs X 5 folds training + predicting took ~6.5 Hr with one 1080TI). <br>\nFor raw (non-equispaced) series, I used 'pack_padded_sequence' in pytorch to process variable lengths in input tensor. <br>\nI also used adding 'flux_err' random noise to 'flux' in other run but it seemed valid score got worse, maybe some params need to be tuning, though...</p>",
      "rawMarkdown": "I also thought about more preprocessing methods like interpolatings but got no time to run in only last 2~3 days so I only applied standardization on 'mjd' (finally, with such setup, 150 epochs X 5 folds training + predicting took ~6.5 Hr with one 1080TI).   \nFor raw (non-equispaced) series, I used 'pack\\_padded\\_sequence' in pytorch to process variable lengths in input tensor.  \nI also used adding 'flux\\_err' random noise to 'flux' in other run but it seemed valid score got worse, maybe some params need to be tuning, though...",
      "votes": null
    },
    {
      "id": "440991",
      "postDate": "12/18/2018 06:44:55",
      "content": "<p>Thanks for sharing, and congrats for the result.  Maybe this one will be what makes me start using Pytorch ;)</p>",
      "rawMarkdown": "Thanks for sharing, and congrats for the result.  Maybe this one will be what makes me start using Pytorch ;)",
      "votes": null
    },
    {
      "id": "440999",
      "postDate": "12/18/2018 07:01:44",
      "content": "<p>Thanks for sharing. Very interesting kernel.i tried training an autoencoder to learn a better representation, but it didn't work out. I found MLP (inspired by densenet architecture) to give good results. You can find the kernel here <a href=\"https://www.kaggle.com/meaninglesslives/a-slightly-better-nn-arch-and-some-tricks\">https://www.kaggle.com/meaninglesslives/a-slightly-better-nn-arch-and-some-tricks</a></p>",
      "rawMarkdown": "Thanks for sharing. Very interesting kernel.i tried training an autoencoder to learn a better representation, but it didn't work out. I found MLP (inspired by densenet architecture) to give good results. You can find the kernel here https://www.kaggle.com/meaninglesslives/a-slightly-better-nn-arch-and-some-tricks",
      "votes": null
    },
    {
      "id": "441000",
      "postDate": "12/18/2018 07:02:08",
      "content": "<p>Congrats &amp; thanks for your sharings in discussion!! <br>\n<a href=\"https://twitter.com/karpathy/status/868178954032513024\">Everyone should try Pytorch ;p</a></p>",
      "rawMarkdown": "Congrats &amp; thanks for your sharings in discussion!!  \n[Everyone should try Pytorch ;p](https://twitter.com/karpathy/status/868178954032513024)",
      "votes": null
    },
    {
      "id": "441003",
      "postDate": "12/18/2018 07:09:47",
      "content": "<p>Thanks and congrats! <br>\nIt's exactly your very early MLP kernel which made me to consider some NN based methods :) <br>\nThis densenet-ish MLP is really amazing, both accuracy(loss) and speed performance! Thanks for sharing!!  </p>",
      "rawMarkdown": "Thanks and congrats!  \nIt's exactly your very early MLP kernel which made me to consider some NN based methods :)  \nThis densenet-ish MLP is really amazing, both accuracy(loss) and speed performance! Thanks for sharing!!",
      "votes": null
    },
    {
      "id": "441264",
      "postDate": "12/18/2018 13:40:52",
      "content": "<p>Love the pytorch api, but facebook '&gt;.&lt;</p>",
      "rawMarkdown": "Love the pytorch api, but facebook '&gt;.&lt;",
      "votes": null
    },
    {
      "id": "441531",
      "postDate": "12/18/2018 18:52:56",
      "content": "<p>Augmentation by randomly shifting along time axis improved my score by 0.1. I used cnn before gru so in my case it might be more useful than in the case of just rnn.</p>",
      "rawMarkdown": "Augmentation by randomly shifting along time axis improved my score by 0.1. I used cnn before gru so in my case it might be more useful than in the case of just rnn.",
      "votes": null
    },
    {
      "id": "441719",
      "postDate": "12/19/2018 00:39:22",
      "content": "<p>random shifting sounds interesting! with others' solutions, it seems augmentations in this comp are playing an important role to achieve better result.</p>",
      "rawMarkdown": "random shifting sounds interesting! with others' solutions, it seems augmentations in this comp are playing an important role to achieve better result.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 440867,
      "author_name": "authman",
      "author_url": "",
      "post_date": "12/18/2018 03:26:13",
      "content": "<p>Nice. My LSTM solution performs similarly (1.014 LB). It uses BiDirectional LSTM followed by separable depth-wise convolution layer.</p>\n\n<p>One thing I did different was I wrote code that takes the fluxes, and linearly interpolates them so long as the mjd gap&lt;80. This results in \"equispaced\" curves that look more like an 'image' or maybe more appropriately 'sound wave' since we're dealing with 1D convolutions here. The input is more dense though due to this and so it takes foreverrrrrr to run. I also augmented every iteration by adding a normal amount of flux_err to a copy of flux every epoch before performing the interpolation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 440872,
          "author_name": "johnfarrell",
          "author_url": "",
          "post_date": "12/18/2018 03:46:41",
          "content": "<p>I also thought about more preprocessing methods like interpolatings but got no time to run in only last 2~3 days so I only applied standardization on 'mjd' (finally, with such setup, 150 epochs X 5 folds training + predicting took ~6.5 Hr with one 1080TI). <br>\nFor raw (non-equispaced) series, I used 'pack_padded_sequence' in pytorch to process variable lengths in input tensor. <br>\nI also used adding 'flux_err' random noise to 'flux' in other run but it seemed valid score got worse, maybe some params need to be tuning, though...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 440991,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/18/2018 06:44:55",
      "content": "<p>Thanks for sharing, and congrats for the result.  Maybe this one will be what makes me start using Pytorch ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 441000,
          "author_name": "johnfarrell",
          "author_url": "",
          "post_date": "12/18/2018 07:02:08",
          "content": "<p>Congrats &amp; thanks for your sharings in discussion!! <br>\n<a href=\"https://twitter.com/karpathy/status/868178954032513024\">Everyone should try Pytorch ;p</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441264,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/18/2018 13:40:52",
          "content": "<p>Love the pytorch api, but facebook '&gt;.&lt;</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 440999,
      "author_name": "meaninglesslives",
      "author_url": "",
      "post_date": "12/18/2018 07:01:44",
      "content": "<p>Thanks for sharing. Very interesting kernel.i tried training an autoencoder to learn a better representation, but it didn't work out. I found MLP (inspired by densenet architecture) to give good results. You can find the kernel here <a href=\"https://www.kaggle.com/meaninglesslives/a-slightly-better-nn-arch-and-some-tricks\">https://www.kaggle.com/meaninglesslives/a-slightly-better-nn-arch-and-some-tricks</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 441003,
          "author_name": "johnfarrell",
          "author_url": "",
          "post_date": "12/18/2018 07:09:47",
          "content": "<p>Thanks and congrats! <br>\nIt's exactly your very early MLP kernel which made me to consider some NN based methods :) <br>\nThis densenet-ish MLP is really amazing, both accuracy(loss) and speed performance! Thanks for sharing!!  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441531,
      "author_name": "lrunaways",
      "author_url": "",
      "post_date": "12/18/2018 18:52:56",
      "content": "<p>Augmentation by randomly shifting along time axis improved my score by 0.1. I used cnn before gru so in my case it might be more useful than in the case of just rnn.</p>",
      "votes": null,
      "replies": [
        {
          "id": 441719,
          "author_name": "johnfarrell",
          "author_url": "",
          "post_date": "12/19/2018 00:39:22",
          "content": "<p>random shifting sounds interesting! with others' solutions, it seems augmentations in this comp are playing an important role to achieve better result.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "440838": "Really great competition &amp; Congrats to everyone!!!  \n\nI just want to share a very simple rnn baseline (in pytorch), kernel is here(*results are not stable and maybe lr schedule should be set carefully*) -&gt; https://www.kaggle.com/johnfarrell/plasticc-2018-emb-gru   \n\nwhich only uses 5 raw series: 'mjd', 'flux', 'flux\\_err', 'detected', 'passband'(embedding dim is 16) &amp; meta features: 'ddf', 'hostgal\\_photoz', 'hostgal\\_photoz\\_err', 'distmod', 'mwebv'. The score result([training.ipynb source file](https://drive.google.com/file/d/1GSkoyQ-kBU0ZpjQmSQ38X7roFl9gGi0m/view?usp=sharing)) is oof0.650/pub0.938/pri0.970.  And it has very low correlation with stats features so stacking works very well with other models using stats features.\n\nFinally, I'm really eager to know other better RNN/CNN/?NN solutions, if you have some please share with me! :)",
    "440867": "Nice. My LSTM solution performs similarly (1.014 LB). It uses BiDirectional LSTM followed by separable depth-wise convolution layer.\n\nOne thing I did different was I wrote code that takes the fluxes, and linearly interpolates them so long as the mjd gap&lt;80. This results in \"equispaced\" curves that look more like an 'image' or maybe more appropriately 'sound wave' since we're dealing with 1D convolutions here. The input is more dense though due to this and so it takes foreverrrrrr to run. I also augmented every iteration by adding a normal amount of flux_err to a copy of flux every epoch before performing the interpolation.",
    "440872": "I also thought about more preprocessing methods like interpolatings but got no time to run in only last 2~3 days so I only applied standardization on 'mjd' (finally, with such setup, 150 epochs X 5 folds training + predicting took ~6.5 Hr with one 1080TI).   \nFor raw (non-equispaced) series, I used 'pack\\_padded\\_sequence' in pytorch to process variable lengths in input tensor.  \nI also used adding 'flux\\_err' random noise to 'flux' in other run but it seemed valid score got worse, maybe some params need to be tuning, though...",
    "440991": "Thanks for sharing, and congrats for the result.  Maybe this one will be what makes me start using Pytorch ;)",
    "440999": "Thanks for sharing. Very interesting kernel.i tried training an autoencoder to learn a better representation, but it didn't work out. I found MLP (inspired by densenet architecture) to give good results. You can find the kernel here https://www.kaggle.com/meaninglesslives/a-slightly-better-nn-arch-and-some-tricks",
    "441000": "Congrats &amp; thanks for your sharings in discussion!!  \n[Everyone should try Pytorch ;p](https://twitter.com/karpathy/status/868178954032513024)",
    "441003": "Thanks and congrats!  \nIt's exactly your very early MLP kernel which made me to consider some NN based methods :)  \nThis densenet-ish MLP is really amazing, both accuracy(loss) and speed performance! Thanks for sharing!!",
    "441264": "Love the pytorch api, but facebook '&gt;.&lt;",
    "441531": "Augmentation by randomly shifting along time axis improved my score by 0.1. I used cnn before gru so in my case it might be more useful than in the case of just rnn.",
    "441719": "random shifting sounds interesting! with others' solutions, it seems augmentations in this comp are playing an important role to achieve better result."
  },
  "source": "meta"
}