{
  "id": 75116,
  "title": "3rd Place Part I -CNN",
  "url": "/competitions/PLAsTiCC-2018/discussion/75116",
  "author_name": "",
  "post_date": "2018-12-18T16:52:57.074642200Z",
  "votes": 39,
  "comment_count": 23,
  "views": 0,
  "content": "<p>First I want to thank the organizers for a very challenging and exciting competition.\nI also want to thank all the kagglers who already shared a lot of very useful information during and after the competition.\nAnd mostly I want to thank my team mates @mamas and @nyanp who thought me a lot and were great company. </p>\n\n<p>Our solution was an ensemble of 3 different solutions: LGB and CatBoost  with different features by @nyanp and @mamas and a CNN. The 3 solution where very different which meant, averaging really worked well for us.</p>\n\n<p>In this post I will mostly describe the CNN.</p>\n\n<p>Our CNN is a Fully 1D Convolutional Neural Network with 256 * 8,5,3 convolution kernels followed by a GlobalMaxPulling , this FCN is close to the one described here <a href=\"https://arxiv.org/pdf/1611.06455.pdf\">\"Time Series Classiﬁcation from Scratch with Deep Neural Networks: A Strong Baseline\"</a></p>\n\n<p>To the output of the last layer is concatenated with the output of a small MLP whose inputs are some meta features which would be described later.</p>\n\n<p>The inputs to the convolution layer is an 18 channels 128 long vectors:</p>\n\n<ol>\n<li>6 channels are built by linear interpolation of the Flux time series</li>\n<li>6 channels are built by linear interpolation of Flux*detected time series</li>\n<li>6 channels describe the distance of the current sampling point to the nearest valid sampling point in the corresponding channel (i.e. this channel have a high value if the interpolation is done with very distant points)</li>\n</ol>\n\n<p>There were 3 different NNs with different meta features as inputs to the MLP:</p>\n\n<ol>\n<li>Minimal - only 4 meta features:  hostgal_photoz, distmod, delta mjd, Flux std (the two later ones are designed to keep the normalization data) </li>\n<li>Minimal + 16 best features from @mamas's model</li>\n<li>Minimal + 15 best features from @nyanp's model</li>\n</ol>\n\n<p>For augmentation during training we used:</p>\n\n<p><strong>Before interpolation</strong></p>\n\n<ol>\n<li>Deletion - up to 30% of the time samples in the training set where deleted (the deletion was done over all the set, not per object)</li>\n<li>Noise - random value proportional to  flux_err</li>\n</ol>\n\n<p><strong>After interpolation</strong></p>\n\n<ol>\n<li>Cyclic shift</li>\n<li>Skew - every flux channel was multiplied by 1+k where k is a small random number</li>\n</ol>\n\n<p>We didn't use TTA as it degraded the LB</p>\n\n<p>Training was done on 4 folds, with changing learning rate,  and  averaging the best weights (More details can be found in <a href=\"https://www.kaggle.com/yuval6967/3rd-place-cnn\">this kernel</a>) </p>\n\n<p>The best LBs scores of the individual NNs were 0.857, 0.857, 0.814  and the average scored 0.791</p>",
  "messages": [
    {
      "id": "441433",
      "postDate": "12/18/2018 16:52:57",
      "content": "<p>First I want to thank the organizers for a very challenging and exciting competition.\nI also want to thank all the kagglers who already shared a lot of very useful information during and after the competition.\nAnd mostly I want to thank my team mates @mamas and @nyanp who thought me a lot and were great company. </p>\n\n<p>Our solution was an ensemble of 3 different solutions: LGB and CatBoost  with different features by @nyanp and @mamas and a CNN. The 3 solution where very different which meant, averaging really worked well for us.</p>\n\n<p>In this post I will mostly describe the CNN.</p>\n\n<p>Our CNN is a Fully 1D Convolutional Neural Network with 256 * 8,5,3 convolution kernels followed by a GlobalMaxPulling , this FCN is close to the one described here <a href=\"https://arxiv.org/pdf/1611.06455.pdf\">\"Time Series Classiﬁcation from Scratch with Deep Neural Networks: A Strong Baseline\"</a></p>\n\n<p>To the output of the last layer is concatenated with the output of a small MLP whose inputs are some meta features which would be described later.</p>\n\n<p>The inputs to the convolution layer is an 18 channels 128 long vectors:</p>\n\n<ol>\n<li>6 channels are built by linear interpolation of the Flux time series</li>\n<li>6 channels are built by linear interpolation of Flux*detected time series</li>\n<li>6 channels describe the distance of the current sampling point to the nearest valid sampling point in the corresponding channel (i.e. this channel have a high value if the interpolation is done with very distant points)</li>\n</ol>\n\n<p>There were 3 different NNs with different meta features as inputs to the MLP:</p>\n\n<ol>\n<li>Minimal - only 4 meta features:  hostgal_photoz, distmod, delta mjd, Flux std (the two later ones are designed to keep the normalization data) </li>\n<li>Minimal + 16 best features from @mamas's model</li>\n<li>Minimal + 15 best features from @nyanp's model</li>\n</ol>\n\n<p>For augmentation during training we used:</p>\n\n<p><strong>Before interpolation</strong></p>\n\n<ol>\n<li>Deletion - up to 30% of the time samples in the training set where deleted (the deletion was done over all the set, not per object)</li>\n<li>Noise - random value proportional to  flux_err</li>\n</ol>\n\n<p><strong>After interpolation</strong></p>\n\n<ol>\n<li>Cyclic shift</li>\n<li>Skew - every flux channel was multiplied by 1+k where k is a small random number</li>\n</ol>\n\n<p>We didn't use TTA as it degraded the LB</p>\n\n<p>Training was done on 4 folds, with changing learning rate,  and  averaging the best weights (More details can be found in <a href=\"https://www.kaggle.com/yuval6967/3rd-place-cnn\">this kernel</a>) </p>\n\n<p>The best LBs scores of the individual NNs were 0.857, 0.857, 0.814  and the average scored 0.791</p>",
      "rawMarkdown": "First I want to thank the organizers for a very challenging and exciting competition.\nI also want to thank all the kagglers who already shared a lot of very useful information during and after the competition.\nAnd mostly I want to thank my team mates @mamas and @nyanp who thought me a lot and were great company. \n\nOur solution was an ensemble of 3 different solutions: LGB and CatBoost  with different features by @nyanp and @mamas and a CNN. The 3 solution where very different which meant, averaging really worked well for us.\n \nIn this post I will mostly describe the CNN.\n\nOur CNN is a Fully 1D Convolutional Neural Network with 256 * 8,5,3 convolution kernels followed by a GlobalMaxPulling , this FCN is close to the one described here [\"Time Series Classiﬁcation from Scratch with Deep Neural Networks: A Strong Baseline\"]( https://arxiv.org/pdf/1611.06455.pdf)\n\nTo the output of the last layer is concatenated with the output of a small MLP whose inputs are some meta features which would be described later.\n\nThe inputs to the convolution layer is an 18 channels 128 long vectors:\n\n1. 6 channels are built by linear interpolation of the Flux time series\n2. 6 channels are built by linear interpolation of Flux*detected time series\n3. 6 channels describe the distance of the current sampling point to the nearest valid sampling point in the corresponding channel (i.e. this channel have a high value if the interpolation is done with very distant points)\n\nThere were 3 different NNs with different meta features as inputs to the MLP:\n\n1. Minimal - only 4 meta features:  hostgal_photoz, distmod, delta mjd, Flux std (the two later ones are designed to keep the normalization data) \n2. Minimal + 16 best features from @mamas's model\n3.  Minimal + 15 best features from @nyanp's model\n\nFor augmentation during training we used:\n\n\n**Before interpolation**\n\n1. Deletion - up to 30% of the time samples in the training set where deleted (the deletion was done over all the set, not per object)\n2. Noise - random value proportional to  flux_err\n\n**After interpolation**\n\n1. Cyclic shift\n2.  Skew - every flux channel was multiplied by 1+k where k is a small random number\n\t\nWe didn't use TTA as it degraded the LB\n\nTraining was done on 4 folds, with changing learning rate,  and  averaging the best weights (More details can be found in [this kernel](https://www.kaggle.com/yuval6967/3rd-place-cnn)) \n\nThe best LBs scores of the individual NNs were 0.857, 0.857, 0.814  and the average scored 0.791",
      "votes": null
    },
    {
      "id": "441440",
      "postDate": "12/18/2018 17:01:34",
      "content": "<p>Congrats! did your NN models do better than LGB?</p>",
      "rawMarkdown": "Congrats! did your NN models do better than LGB?",
      "votes": null
    },
    {
      "id": "441441",
      "postDate": "12/18/2018 17:04:10",
      "content": "<p>Great solution, thanks for sharing, and congrats on the result!  Including distance to nearest interpolation point is what makes it work probably.  </p>\n\n<p>I am not sure I get 'We didn't use TTA as it degraded the LB' because you did use augmentation during training set. Can you clarify?</p>",
      "rawMarkdown": "Great solution, thanks for sharing, and congrats on the result!  Including distance to nearest interpolation point is what makes it work probably.  \n\nI am not sure I get 'We didn't use TTA as it degraded the LB' because you did use augmentation during training set. Can you clarify?",
      "votes": null
    },
    {
      "id": "441443",
      "postDate": "12/18/2018 17:06:12",
      "content": "<p>They were very different.  In the NN most of the features were actually extracted by the NN itself.\nThe average of the 3 NN had a weight of about 50% of our solution, the LBG+CatBoost were the other 50%. and we gained about 0.035 by class99 probing </p>",
      "rawMarkdown": "They were very different.  In the NN most of the features were actually extracted by the NN itself.\nThe average of the 3 NN had a weight of about 50% of our solution, the LBG+CatBoost were the other 50%. and we gained about 0.035 by class99 probing",
      "votes": null
    },
    {
      "id": "441447",
      "postDate": "12/18/2018 17:09:01",
      "content": "<p>I'm a little weak when in comes to all the professional terms, I actually meant  TTA - Test Time Augmentation </p>",
      "rawMarkdown": "I'm a little weak when in comes to all the professional terms, I actually meant  TTA - Test Time Augmentation",
      "votes": null
    },
    {
      "id": "441449",
      "postDate": "12/18/2018 17:10:18",
      "content": "<p>And, Yes, the distance to nearest interpolation point improved our LB a lot.</p>",
      "rawMarkdown": "And, Yes, the distance to nearest interpolation point improved our LB a lot.",
      "votes": null
    },
    {
      "id": "441458",
      "postDate": "12/18/2018 17:17:46",
      "content": "<p>Thanks for the answers.  Maybe I'm the one being wrong about what TTA means.  Let me update my write up to make it clearer if need be.</p>",
      "rawMarkdown": "Thanks for the answers.  Maybe I'm the one being wrong about what TTA means.  Let me update my write up to make it clearer if need be.",
      "votes": null
    },
    {
      "id": "441904",
      "postDate": "12/19/2018 08:34:58",
      "content": "<p>Congrats and thank you for sharing! mamas said your NN was super cool. I am looking forward that you post the kernel;)</p>",
      "rawMarkdown": "Congrats and thank you for sharing! mamas said your NN was super cool. I am looking forward that you post the kernel;)",
      "votes": null
    },
    {
      "id": "442024",
      "postDate": "12/19/2018 11:38:24",
      "content": "<p>Super interesting! A question - how do you get from the long time series to a 128 long vector? You linearly interpolate the entire 1000 days then subsample 128 days from that?</p>",
      "rawMarkdown": "Super interesting! A question - how do you get from the long time series to a 128 long vector? You linearly interpolate the entire 1000 days then subsample 128 days from that?",
      "votes": null
    },
    {
      "id": "442046",
      "postDate": "12/19/2018 12:09:40",
      "content": "<p>I just linear interpolate from the min(mjd) of that object to max(mjd) </p>",
      "rawMarkdown": "I just linear interpolate from the min(mjd) of that object to max(mjd)",
      "votes": null
    },
    {
      "id": "442048",
      "postDate": "12/19/2018 12:10:48",
      "content": "<p>Thanks, I will post it over the weekend</p>",
      "rawMarkdown": "Thanks, I will post it over the weekend",
      "votes": null
    },
    {
      "id": "442266",
      "postDate": "12/19/2018 17:54:06",
      "content": "<p>Great solution and congrats Yuval and team!</p>",
      "rawMarkdown": "Great solution and congrats Yuval and team!",
      "votes": null
    },
    {
      "id": "442332",
      "postDate": "12/19/2018 20:05:17",
      "content": "<p>Thanks yuval, I'm really impressed with your beautiful 1D-CNN model, and your many valuable ideas related to class99 handling. I'm really proud of working with such a great team :)</p>",
      "rawMarkdown": "Thanks yuval, I'm really impressed with your beautiful 1D-CNN model, and your many valuable ideas related to class99 handling. I'm really proud of working with such a great team :)",
      "votes": null
    },
    {
      "id": "442694",
      "postDate": "12/20/2018 10:52:34",
      "content": "<p>Thank you for great NN solution!!!\nI have two question.\n1. How did you normalize the flux value? (like std 1.0 and mean 0?)\n2. If there are multiple value in 128 vector, how did you process? (mean?)</p>",
      "rawMarkdown": "Thank you for great NN solution!!!\nI have two question.\n1. How did you normalize the flux value? (like std 1.0 and mean 0?)\n2. If there are multiple value in 128 vector, how did you process? (mean?)",
      "votes": null
    },
    {
      "id": "442741",
      "postDate": "12/20/2018 12:19:21",
      "content": "<ol>\n<li>I divide by the std of the values of the all 6 channels. I then use the std as an extra feature which is an input the MLP. I don't shift the mean, I don't really find it necessary if the std is close to 1 and the mean is low enough.</li>\n<li>There is no special treatment to multiple values which fall between two sampling points. I use np.interp for interpolation.</li>\n</ol>",
      "rawMarkdown": "1. I divide by the std of the values of the all 6 channels. I then use the std as an extra feature which is an input the MLP. I don't shift the mean, I don't really find it necessary if the std is close to 1 and the mean is low enough.\n2. There is no special treatment to multiple values which fall between two sampling points. I use np.interp for interpolation.",
      "votes": null
    },
    {
      "id": "443024",
      "postDate": "12/20/2018 22:45:06",
      "content": "<p>Congratulations, and thanks for sharing</p>",
      "rawMarkdown": "Congratulations, and thanks for sharing",
      "votes": null
    },
    {
      "id": "443031",
      "postDate": "12/20/2018 22:56:05",
      "content": "<p>Congratulations and thank you for the interesting solution.  It's inspiring how you manage to get exclusively gold medals :-) </p>\n\n<p><code>We didn't use TTA as it degraded the LB</code> -- we did not do TTA cause it's a pain! But it's nice to see that somebody tried :-)))</p>",
      "rawMarkdown": "Congratulations and thank you for the interesting solution.  It's inspiring how you manage to get exclusively gold medals :-) \n\n```We didn't use TTA as it degraded the LB``` -- we did not do TTA cause it's a pain! But it's nice to see that somebody tried :-)))",
      "votes": null
    },
    {
      "id": "443035",
      "postDate": "12/20/2018 23:02:31",
      "content": "<p>I think since the fastai course became popular people attribute TTA to a Test Time Augmentation (this is how Jeremy used it and now it's kind of reserved for the test augmentation :)) </p>",
      "rawMarkdown": "I think since the fastai course became popular people attribute TTA to a Test Time Augmentation (this is how Jeremy used it and now it's kind of reserved for the test augmentation :))",
      "votes": null
    },
    {
      "id": "443058",
      "postDate": "12/21/2018 00:33:22",
      "content": "<p>Thanks!!\nI will try making 1D-CNN model in reference to your solution.\nAnd I'm looking forward to your kernel!</p>",
      "rawMarkdown": "Thanks!!\nI will try making 1D-CNN model in reference to your solution.\nAnd I'm looking forward to your kernel!",
      "votes": null
    },
    {
      "id": "443110",
      "postDate": "12/21/2018 03:06:30",
      "content": "<p>It really is a pain ;) but I had to give it a try.</p>",
      "rawMarkdown": "It really is a pain ;) but I had to give it a try.",
      "votes": null
    },
    {
      "id": "443285",
      "postDate": "12/21/2018 11:00:37",
      "content": "<p><a href=\"/yuval6967\">@yuval6967</a> </p>\n\n<p>I actually have a few questions regarding your NN. I also tried to apply FCN with global max pooling as a feature extractor in my internship (medical signals), so far unsuccessfully. I used 1D FCN with 1 channel (around 10 layers with small kernels of 3-5)</p>\n\n<p>In your architecture:</p>\n\n<ol>\n<li><p>6 channels are built by linear interpolation of the Flux time series -- what are those 6 channels? How they differ? Was is critical to have 6 of them and not just 1, did it help ?</p></li>\n<li><p>256 * 8,5,3 convolution kernels  -- this is not quite clear, did you use 256 filters with kernel sizes 8, 5 and 3 for each 6-channel group or did you change the amount of filters, or size of filters? How deep was your net for 128 signals then ? Also, NN needs lot's of augmentation, so do you recall approx order of signals you generated with augmentation (1000 000? )</p></li>\n<li><p>To the output of the last layer is concatenated with the output of a small MLP whose inputs are some meta features -- so was it a three 500 fully connected layers MLP like in that paper? When you concatenate before the softmax, did you bring both outputs (MLP and NN to the same size, like 32-32 or something like that)? Was it important to have the MLP to make it work? I have around 18 features in med signals, so you take a vector of M features x N signals for MLP input then?</p></li>\n</ol>\n\n<p>Too many questions, sorry :), I want to try to make my FCN work for med signals, and see yours architecture is much more advanced, but do not fully understand how to implement those advances. Thank you for that paper link as well</p>",
      "rawMarkdown": "yuval6967 \n\nI actually have a few questions regarding your NN. I also tried to apply FCN with global max pooling as a feature extractor in my internship (medical signals), so far unsuccessfully. I used 1D FCN with 1 channel (around 10 layers with small kernels of 3-5)\n\nIn your architecture:\n\n1. 6 channels are built by linear interpolation of the Flux time series -- what are those 6 channels? How they differ? Was is critical to have 6 of them and not just 1, did it help ?\n\n2.  256 * 8,5,3 convolution kernels  -- this is not quite clear, did you use 256 filters with kernel sizes 8, 5 and 3 for each 6-channel group or did you change the amount of filters, or size of filters? How deep was your net for 128 signals then ? Also, NN needs lot's of augmentation, so do you recall approx order of signals you generated with augmentation (1000 000? )\n\n3. To the output of the last layer is concatenated with the output of a small MLP whose inputs are some meta features -- so was it a three 500 fully connected layers MLP like in that paper? When you concatenate before the softmax, did you bring both outputs (MLP and NN to the same size, like 32-32 or something like that)? Was it important to have the MLP to make it work? I have around 18 features in med signals, so you take a vector of M features x N signals for MLP input then?\n\nToo many questions, sorry :), I want to try to make my FCN work for med signals, and see yours architecture is much more advanced, but do not fully understand how to implement those advances. Thank you for that paper link as well",
      "votes": null
    },
    {
      "id": "443295",
      "postDate": "12/21/2018 11:30:39",
      "content": "<p>Below you can find the KERAS code with the model definition.</p>\n\n<p>The 6 channel inputs are the 6 passbands.</p>\n\n<p>The parameter num_samples in the definition is set in global scope to 128\nThis network is set to receive 6 extra features, in other networks I set it to 4, 19 (when added @nyanp features)  or 20 (with @mamas featurs).</p>\n\n<p>The layer MySwitch forces the galactic outputs to be 0 for extragalactic objects and vice versa.</p>\n\n<p>When training, I augmented and interpolated on the fly (thanks to @mamas great code), I used 4 * 58 epochs for every training and I also repeated training with more noise/deletion.</p>\n\n<p>Later today I will release a full kernel</p>\n\n<pre><code>def build_model():\n   input_timeseries = Input(shape=(num_samples, 6,),name='input_timeseries')\n   input_timeseries0 = Input(shape=(num_samples, 6,),name='input_timeseries0')\n   input_timeseriese = Input(shape=(num_samples, 6,),name='input_timeseriese')\n   input_meta = Input(shape=(6,),name='input_meta')\n   input_gal = Input(shape=(1,),name='input_gal')\n   _series=concatenate([input_timeseries,input_timeseries0,input_timeseriese])\n   x = Conv1D(256,8,padding='same',name='Conv1')(_series)\n   x = BatchNormalization()(x)\n   x = LeakyReLU(alpha=0.1)(x)\n   x = Dropout(0.2)(x)\n   x = Conv1D(256,5,padding='same',name='Conv2')(x)\n   x = BatchNormalization()(x)\n   x = LeakyReLU(alpha=0.1)(x)\n   x = Dropout(0.2)(x)\n   x = Conv1D(256,3,padding='same',name='Conv5')(x)\n   x = BatchNormalization()(x)\n   x = LeakyReLU(alpha=0.1)(x)\n   x = GlobalMaxPooling1D()(x)\n   x1 = Dense(16,activation='relu',name='dense0')(input_meta)\n   x1 = Dense(32,activation='relu',name='dense1')(x1)\n   xc = concatenate([x,x1],name='concat')\n   x = Dense(256,activation='relu',name='features')(xc)\n   x = Dense(real_targets.shape[0],name='bout')(x)\n   x = MySwitch(galactic_targets.shape[0])([input_gal,x])\n   out = Activation('softmax',name='out')(x)\n   model=Model([input_timeseries,input_timeseries0, \n              input_timeseriese,input_meta,input_gal],out)\n   return model\n</code></pre>",
      "rawMarkdown": "Below you can find the KERAS code with the model definition.\n\nThe 6 channel inputs are the 6 passbands.\n\n\nThe parameter num_samples in the definition is set in global scope to 128\nThis network is set to receive 6 extra features, in other networks I set it to 4, 19 (when added @nyanp features)  or 20 (with @mamas featurs).\n\nThe layer MySwitch forces the galactic outputs to be 0 for extragalactic objects and vice versa.\n\nWhen training, I augmented and interpolated on the fly (thanks to @mamas great code), I used 4 * 58 epochs for every training and I also repeated training with more noise/deletion.\n\nLater today I will release a full kernel\n\n\n    def build_model():\n       input_timeseries = Input(shape=(num_samples, 6,),name='input_timeseries')\n       input_timeseries0 = Input(shape=(num_samples, 6,),name='input_timeseries0')\n       input_timeseriese = Input(shape=(num_samples, 6,),name='input_timeseriese')\n       input_meta = Input(shape=(6,),name='input_meta')\n       input_gal = Input(shape=(1,),name='input_gal')\n       _series=concatenate([input_timeseries,input_timeseries0,input_timeseriese])\n       x = Conv1D(256,8,padding='same',name='Conv1')(_series)\n       x = BatchNormalization()(x)\n       x = LeakyReLU(alpha=0.1)(x)\n       x = Dropout(0.2)(x)\n       x = Conv1D(256,5,padding='same',name='Conv2')(x)\n       x = BatchNormalization()(x)\n       x = LeakyReLU(alpha=0.1)(x)\n       x = Dropout(0.2)(x)\n       x = Conv1D(256,3,padding='same',name='Conv5')(x)\n       x = BatchNormalization()(x)\n       x = LeakyReLU(alpha=0.1)(x)\n       x = GlobalMaxPooling1D()(x)\n       x1 = Dense(16,activation='relu',name='dense0')(input_meta)\n       x1 = Dense(32,activation='relu',name='dense1')(x1)\n       xc = concatenate([x,x1],name='concat')\n       x = Dense(256,activation='relu',name='features')(xc)\n       x = Dense(real_targets.shape[0],name='bout')(x)\n       x = MySwitch(galactic_targets.shape[0])([input_gal,x])\n       out = Activation('softmax',name='out')(x)\n       model=Model([input_timeseries,input_timeseries0, \n                  input_timeseriese,input_meta,input_gal],out)\n       return model",
      "votes": null
    },
    {
      "id": "443424",
      "postDate": "12/21/2018 15:58:05",
      "content": "<p>Thank you Yuval, now it's clear. I'll see if the trick with adding features in parallel works for medicine as well, looks like a good idea. :)</p>",
      "rawMarkdown": "Thank you Yuval, now it's clear. I'll see if the trick with adding features in parallel works for medicine as well, looks like a good idea. :)",
      "votes": null
    },
    {
      "id": "443479",
      "postDate": "12/21/2018 18:07:37",
      "content": "<p>A kernel with the CNN can be found <a href=\"https://www.kaggle.com/yuval6967/3rd-place-cnn\">here</a></p>",
      "rawMarkdown": "A kernel with the CNN can be found [here](https://www.kaggle.com/yuval6967/3rd-place-cnn)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 441440,
      "author_name": "godaibo",
      "author_url": "",
      "post_date": "12/18/2018 17:01:34",
      "content": "<p>Congrats! did your NN models do better than LGB?</p>",
      "votes": null,
      "replies": [
        {
          "id": 441443,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "12/18/2018 17:06:12",
          "content": "<p>They were very different.  In the NN most of the features were actually extracted by the NN itself.\nThe average of the 3 NN had a weight of about 50% of our solution, the LBG+CatBoost were the other 50%. and we gained about 0.035 by class99 probing </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441441,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/18/2018 17:04:10",
      "content": "<p>Great solution, thanks for sharing, and congrats on the result!  Including distance to nearest interpolation point is what makes it work probably.  </p>\n\n<p>I am not sure I get 'We didn't use TTA as it degraded the LB' because you did use augmentation during training set. Can you clarify?</p>",
      "votes": null,
      "replies": [
        {
          "id": 441447,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "12/18/2018 17:09:01",
          "content": "<p>I'm a little weak when in comes to all the professional terms, I actually meant  TTA - Test Time Augmentation </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441449,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "12/18/2018 17:10:18",
          "content": "<p>And, Yes, the distance to nearest interpolation point improved our LB a lot.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441458,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/18/2018 17:17:46",
          "content": "<p>Thanks for the answers.  Maybe I'm the one being wrong about what TTA means.  Let me update my write up to make it clearer if need be.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443035,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "12/20/2018 23:02:31",
          "content": "<p>I think since the fastai course became popular people attribute TTA to a Test Time Augmentation (this is how Jeremy used it and now it's kind of reserved for the test augmentation :)) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441904,
      "author_name": "takuok",
      "author_url": "",
      "post_date": "12/19/2018 08:34:58",
      "content": "<p>Congrats and thank you for sharing! mamas said your NN was super cool. I am looking forward that you post the kernel;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 442048,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "12/19/2018 12:10:48",
          "content": "<p>Thanks, I will post it over the weekend</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442024,
      "author_name": "henripal",
      "author_url": "",
      "post_date": "12/19/2018 11:38:24",
      "content": "<p>Super interesting! A question - how do you get from the long time series to a 128 long vector? You linearly interpolate the entire 1000 days then subsample 128 days from that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 442046,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "12/19/2018 12:09:40",
          "content": "<p>I just linear interpolate from the min(mjd) of that object to max(mjd) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442266,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "12/19/2018 17:54:06",
      "content": "<p>Great solution and congrats Yuval and team!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 442332,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "12/19/2018 20:05:17",
      "content": "<p>Thanks yuval, I'm really impressed with your beautiful 1D-CNN model, and your many valuable ideas related to class99 handling. I'm really proud of working with such a great team :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 442694,
      "author_name": "go5kuramubon",
      "author_url": "",
      "post_date": "12/20/2018 10:52:34",
      "content": "<p>Thank you for great NN solution!!!\nI have two question.\n1. How did you normalize the flux value? (like std 1.0 and mean 0?)\n2. If there are multiple value in 128 vector, how did you process? (mean?)</p>",
      "votes": null,
      "replies": [
        {
          "id": 442741,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "12/20/2018 12:19:21",
          "content": "<ol>\n<li>I divide by the std of the values of the all 6 channels. I then use the std as an extra feature which is an input the MLP. I don't shift the mean, I don't really find it necessary if the std is close to 1 and the mean is low enough.</li>\n<li>There is no special treatment to multiple values which fall between two sampling points. I use np.interp for interpolation.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443058,
          "author_name": "go5kuramubon",
          "author_url": "",
          "post_date": "12/21/2018 00:33:22",
          "content": "<p>Thanks!!\nI will try making 1D-CNN model in reference to your solution.\nAnd I'm looking forward to your kernel!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 443024,
      "author_name": "ahmedbaz",
      "author_url": "",
      "post_date": "12/20/2018 22:45:06",
      "content": "<p>Congratulations, and thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 443031,
      "author_name": "blondinka",
      "author_url": "",
      "post_date": "12/20/2018 22:56:05",
      "content": "<p>Congratulations and thank you for the interesting solution.  It's inspiring how you manage to get exclusively gold medals :-) </p>\n\n<p><code>We didn't use TTA as it degraded the LB</code> -- we did not do TTA cause it's a pain! But it's nice to see that somebody tried :-)))</p>",
      "votes": null,
      "replies": [
        {
          "id": 443110,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "12/21/2018 03:06:30",
          "content": "<p>It really is a pain ;) but I had to give it a try.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 443285,
      "author_name": "blondinka",
      "author_url": "",
      "post_date": "12/21/2018 11:00:37",
      "content": "<p><a href=\"/yuval6967\">@yuval6967</a> </p>\n\n<p>I actually have a few questions regarding your NN. I also tried to apply FCN with global max pooling as a feature extractor in my internship (medical signals), so far unsuccessfully. I used 1D FCN with 1 channel (around 10 layers with small kernels of 3-5)</p>\n\n<p>In your architecture:</p>\n\n<ol>\n<li><p>6 channels are built by linear interpolation of the Flux time series -- what are those 6 channels? How they differ? Was is critical to have 6 of them and not just 1, did it help ?</p></li>\n<li><p>256 * 8,5,3 convolution kernels  -- this is not quite clear, did you use 256 filters with kernel sizes 8, 5 and 3 for each 6-channel group or did you change the amount of filters, or size of filters? How deep was your net for 128 signals then ? Also, NN needs lot's of augmentation, so do you recall approx order of signals you generated with augmentation (1000 000? )</p></li>\n<li><p>To the output of the last layer is concatenated with the output of a small MLP whose inputs are some meta features -- so was it a three 500 fully connected layers MLP like in that paper? When you concatenate before the softmax, did you bring both outputs (MLP and NN to the same size, like 32-32 or something like that)? Was it important to have the MLP to make it work? I have around 18 features in med signals, so you take a vector of M features x N signals for MLP input then?</p></li>\n</ol>\n\n<p>Too many questions, sorry :), I want to try to make my FCN work for med signals, and see yours architecture is much more advanced, but do not fully understand how to implement those advances. Thank you for that paper link as well</p>",
      "votes": null,
      "replies": [
        {
          "id": 443295,
          "author_name": "yuval6967",
          "author_url": "",
          "post_date": "12/21/2018 11:30:39",
          "content": "<p>Below you can find the KERAS code with the model definition.</p>\n\n<p>The 6 channel inputs are the 6 passbands.</p>\n\n<p>The parameter num_samples in the definition is set in global scope to 128\nThis network is set to receive 6 extra features, in other networks I set it to 4, 19 (when added @nyanp features)  or 20 (with @mamas featurs).</p>\n\n<p>The layer MySwitch forces the galactic outputs to be 0 for extragalactic objects and vice versa.</p>\n\n<p>When training, I augmented and interpolated on the fly (thanks to @mamas great code), I used 4 * 58 epochs for every training and I also repeated training with more noise/deletion.</p>\n\n<p>Later today I will release a full kernel</p>\n\n<pre><code>def build_model():\n   input_timeseries = Input(shape=(num_samples, 6,),name='input_timeseries')\n   input_timeseries0 = Input(shape=(num_samples, 6,),name='input_timeseries0')\n   input_timeseriese = Input(shape=(num_samples, 6,),name='input_timeseriese')\n   input_meta = Input(shape=(6,),name='input_meta')\n   input_gal = Input(shape=(1,),name='input_gal')\n   _series=concatenate([input_timeseries,input_timeseries0,input_timeseriese])\n   x = Conv1D(256,8,padding='same',name='Conv1')(_series)\n   x = BatchNormalization()(x)\n   x = LeakyReLU(alpha=0.1)(x)\n   x = Dropout(0.2)(x)\n   x = Conv1D(256,5,padding='same',name='Conv2')(x)\n   x = BatchNormalization()(x)\n   x = LeakyReLU(alpha=0.1)(x)\n   x = Dropout(0.2)(x)\n   x = Conv1D(256,3,padding='same',name='Conv5')(x)\n   x = BatchNormalization()(x)\n   x = LeakyReLU(alpha=0.1)(x)\n   x = GlobalMaxPooling1D()(x)\n   x1 = Dense(16,activation='relu',name='dense0')(input_meta)\n   x1 = Dense(32,activation='relu',name='dense1')(x1)\n   xc = concatenate([x,x1],name='concat')\n   x = Dense(256,activation='relu',name='features')(xc)\n   x = Dense(real_targets.shape[0],name='bout')(x)\n   x = MySwitch(galactic_targets.shape[0])([input_gal,x])\n   out = Activation('softmax',name='out')(x)\n   model=Model([input_timeseries,input_timeseries0, \n              input_timeseriese,input_meta,input_gal],out)\n   return model\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443424,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "12/21/2018 15:58:05",
          "content": "<p>Thank you Yuval, now it's clear. I'll see if the trick with adding features in parallel works for medicine as well, looks like a good idea. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 443479,
      "author_name": "yuval6967",
      "author_url": "",
      "post_date": "12/21/2018 18:07:37",
      "content": "<p>A kernel with the CNN can be found <a href=\"https://www.kaggle.com/yuval6967/3rd-place-cnn\">here</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "441433": "First I want to thank the organizers for a very challenging and exciting competition.\nI also want to thank all the kagglers who already shared a lot of very useful information during and after the competition.\nAnd mostly I want to thank my team mates @mamas and @nyanp who thought me a lot and were great company. \n\nOur solution was an ensemble of 3 different solutions: LGB and CatBoost  with different features by @nyanp and @mamas and a CNN. The 3 solution where very different which meant, averaging really worked well for us.\n \nIn this post I will mostly describe the CNN.\n\nOur CNN is a Fully 1D Convolutional Neural Network with 256 * 8,5,3 convolution kernels followed by a GlobalMaxPulling , this FCN is close to the one described here [\"Time Series Classiﬁcation from Scratch with Deep Neural Networks: A Strong Baseline\"]( https://arxiv.org/pdf/1611.06455.pdf)\n\nTo the output of the last layer is concatenated with the output of a small MLP whose inputs are some meta features which would be described later.\n\nThe inputs to the convolution layer is an 18 channels 128 long vectors:\n\n1. 6 channels are built by linear interpolation of the Flux time series\n2. 6 channels are built by linear interpolation of Flux*detected time series\n3. 6 channels describe the distance of the current sampling point to the nearest valid sampling point in the corresponding channel (i.e. this channel have a high value if the interpolation is done with very distant points)\n\nThere were 3 different NNs with different meta features as inputs to the MLP:\n\n1. Minimal - only 4 meta features:  hostgal_photoz, distmod, delta mjd, Flux std (the two later ones are designed to keep the normalization data) \n2. Minimal + 16 best features from @mamas's model\n3.  Minimal + 15 best features from @nyanp's model\n\nFor augmentation during training we used:\n\n\n**Before interpolation**\n\n1. Deletion - up to 30% of the time samples in the training set where deleted (the deletion was done over all the set, not per object)\n2. Noise - random value proportional to  flux_err\n\n**After interpolation**\n\n1. Cyclic shift\n2.  Skew - every flux channel was multiplied by 1+k where k is a small random number\n\t\nWe didn't use TTA as it degraded the LB\n\nTraining was done on 4 folds, with changing learning rate,  and  averaging the best weights (More details can be found in [this kernel](https://www.kaggle.com/yuval6967/3rd-place-cnn)) \n\nThe best LBs scores of the individual NNs were 0.857, 0.857, 0.814  and the average scored 0.791",
    "441440": "Congrats! did your NN models do better than LGB?",
    "441441": "Great solution, thanks for sharing, and congrats on the result!  Including distance to nearest interpolation point is what makes it work probably.  \n\nI am not sure I get 'We didn't use TTA as it degraded the LB' because you did use augmentation during training set. Can you clarify?",
    "441443": "They were very different.  In the NN most of the features were actually extracted by the NN itself.\nThe average of the 3 NN had a weight of about 50% of our solution, the LBG+CatBoost were the other 50%. and we gained about 0.035 by class99 probing",
    "441447": "I'm a little weak when in comes to all the professional terms, I actually meant  TTA - Test Time Augmentation",
    "441449": "And, Yes, the distance to nearest interpolation point improved our LB a lot.",
    "441458": "Thanks for the answers.  Maybe I'm the one being wrong about what TTA means.  Let me update my write up to make it clearer if need be.",
    "441904": "Congrats and thank you for sharing! mamas said your NN was super cool. I am looking forward that you post the kernel;)",
    "442024": "Super interesting! A question - how do you get from the long time series to a 128 long vector? You linearly interpolate the entire 1000 days then subsample 128 days from that?",
    "442046": "I just linear interpolate from the min(mjd) of that object to max(mjd)",
    "442048": "Thanks, I will post it over the weekend",
    "442266": "Great solution and congrats Yuval and team!",
    "442332": "Thanks yuval, I'm really impressed with your beautiful 1D-CNN model, and your many valuable ideas related to class99 handling. I'm really proud of working with such a great team :)",
    "442694": "Thank you for great NN solution!!!\nI have two question.\n1. How did you normalize the flux value? (like std 1.0 and mean 0?)\n2. If there are multiple value in 128 vector, how did you process? (mean?)",
    "442741": "1. I divide by the std of the values of the all 6 channels. I then use the std as an extra feature which is an input the MLP. I don't shift the mean, I don't really find it necessary if the std is close to 1 and the mean is low enough.\n2. There is no special treatment to multiple values which fall between two sampling points. I use np.interp for interpolation.",
    "443024": "Congratulations, and thanks for sharing",
    "443031": "Congratulations and thank you for the interesting solution.  It's inspiring how you manage to get exclusively gold medals :-) \n\n```We didn't use TTA as it degraded the LB``` -- we did not do TTA cause it's a pain! But it's nice to see that somebody tried :-)))",
    "443035": "I think since the fastai course became popular people attribute TTA to a Test Time Augmentation (this is how Jeremy used it and now it's kind of reserved for the test augmentation :))",
    "443058": "Thanks!!\nI will try making 1D-CNN model in reference to your solution.\nAnd I'm looking forward to your kernel!",
    "443110": "It really is a pain ;) but I had to give it a try.",
    "443285": "yuval6967 \n\nI actually have a few questions regarding your NN. I also tried to apply FCN with global max pooling as a feature extractor in my internship (medical signals), so far unsuccessfully. I used 1D FCN with 1 channel (around 10 layers with small kernels of 3-5)\n\nIn your architecture:\n\n1. 6 channels are built by linear interpolation of the Flux time series -- what are those 6 channels? How they differ? Was is critical to have 6 of them and not just 1, did it help ?\n\n2.  256 * 8,5,3 convolution kernels  -- this is not quite clear, did you use 256 filters with kernel sizes 8, 5 and 3 for each 6-channel group or did you change the amount of filters, or size of filters? How deep was your net for 128 signals then ? Also, NN needs lot's of augmentation, so do you recall approx order of signals you generated with augmentation (1000 000? )\n\n3. To the output of the last layer is concatenated with the output of a small MLP whose inputs are some meta features -- so was it a three 500 fully connected layers MLP like in that paper? When you concatenate before the softmax, did you bring both outputs (MLP and NN to the same size, like 32-32 or something like that)? Was it important to have the MLP to make it work? I have around 18 features in med signals, so you take a vector of M features x N signals for MLP input then?\n\nToo many questions, sorry :), I want to try to make my FCN work for med signals, and see yours architecture is much more advanced, but do not fully understand how to implement those advances. Thank you for that paper link as well",
    "443295": "Below you can find the KERAS code with the model definition.\n\nThe 6 channel inputs are the 6 passbands.\n\n\nThe parameter num_samples in the definition is set in global scope to 128\nThis network is set to receive 6 extra features, in other networks I set it to 4, 19 (when added @nyanp features)  or 20 (with @mamas featurs).\n\nThe layer MySwitch forces the galactic outputs to be 0 for extragalactic objects and vice versa.\n\nWhen training, I augmented and interpolated on the fly (thanks to @mamas great code), I used 4 * 58 epochs for every training and I also repeated training with more noise/deletion.\n\nLater today I will release a full kernel\n\n\n    def build_model():\n       input_timeseries = Input(shape=(num_samples, 6,),name='input_timeseries')\n       input_timeseries0 = Input(shape=(num_samples, 6,),name='input_timeseries0')\n       input_timeseriese = Input(shape=(num_samples, 6,),name='input_timeseriese')\n       input_meta = Input(shape=(6,),name='input_meta')\n       input_gal = Input(shape=(1,),name='input_gal')\n       _series=concatenate([input_timeseries,input_timeseries0,input_timeseriese])\n       x = Conv1D(256,8,padding='same',name='Conv1')(_series)\n       x = BatchNormalization()(x)\n       x = LeakyReLU(alpha=0.1)(x)\n       x = Dropout(0.2)(x)\n       x = Conv1D(256,5,padding='same',name='Conv2')(x)\n       x = BatchNormalization()(x)\n       x = LeakyReLU(alpha=0.1)(x)\n       x = Dropout(0.2)(x)\n       x = Conv1D(256,3,padding='same',name='Conv5')(x)\n       x = BatchNormalization()(x)\n       x = LeakyReLU(alpha=0.1)(x)\n       x = GlobalMaxPooling1D()(x)\n       x1 = Dense(16,activation='relu',name='dense0')(input_meta)\n       x1 = Dense(32,activation='relu',name='dense1')(x1)\n       xc = concatenate([x,x1],name='concat')\n       x = Dense(256,activation='relu',name='features')(xc)\n       x = Dense(real_targets.shape[0],name='bout')(x)\n       x = MySwitch(galactic_targets.shape[0])([input_gal,x])\n       out = Activation('softmax',name='out')(x)\n       model=Model([input_timeseries,input_timeseries0, \n                  input_timeseriese,input_meta,input_gal],out)\n       return model",
    "443424": "Thank you Yuval, now it's clear. I'll see if the trick with adding features in parallel works for medicine as well, looks like a good idea. :)",
    "443479": "A kernel with the CNN can be found [here](https://www.kaggle.com/yuval6967/3rd-place-cnn)"
  },
  "source": "meta"
}