{
  "id": 75012,
  "title": "Congrats and 8th place Rapids solution updated!",
  "url": "/competitions/PLAsTiCC-2018/writeups/rapids-ai-congrats-and-8th-place-rapids-solution-u",
  "author_name": "",
  "post_date": "2019-01-09T12:15:05.953Z",
  "votes": 88,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Update: my solution will be posted <a href=\"https://github.com/daxiongshu/Rapids_PLAsTiCC_2018\">on github</a>. Now you can find a quick end2end demo with rapids. Forgive me that I'm always distracted to other stuff but I promise I'll get the whole solution done eventually!</p>\n\n<p>Big Congrats to the winners and new kaggle masters. I want to thank Olivier <a href=\"/ogrellier\">@ogrellier</a> and Kyle <a href=\"/kyleboone\">@kyleboone</a> specifically. Without your kernels, I can't even start this competition.</p>\n\n<p>Overall this competition is a fantastic learning process for me. If you have no astronomical background, this solution is for you. Considering I started climbing the LB from a model that is just 0.95 LB, what happened is a miracle :P</p>\n\n<p>Here is a preview of my solution and more to come! </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/440750/10889/image.png\" alt=\"Solution Overview\"></p>\n\n<p>Back to my laptop :P Ok, let me walk through this chart. </p>\n\n<h1>Baseline Models and feature engineering.</h1>\n\n<p>In case you haven't noticed, I'm using <a href=\"https://github.com/rapidsai/cudf\">cudf</a> to replace pandas for this competition whenever I can. In general, it is at least <strong>10x</strong> faster for csv reader and general groupby - aggregation <strong>with a single GPU</strong> than pandas on CPU. Sometimes I got <strong>100x</strong> speedup. I used cudf to generate most features in <a href=\"https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data/code\">Olivier's kernel</a>. However, as a merely two-month old library, you can expect some missing functionalities of cudf. I will start publish cudf tutorial in kaggle competitions and work with our developers to make your favorite pandas trick come true on GPU!  A demo notebook of using <strong>cudf</strong> will be published sometime tomorrow. Stay tuned!</p>\n\n<p>In addition, I also built some other features that gave me a small boost including decline rate from the peak, number of ups and downs of the flux curve, passband count percentage around the peak, <a href=\"https://www.kaggle.com/rejpalcz/feature-extraction-using-period-analysis\">period features</a> (optimized), and features from <a href=\"https://www.kaggle.com/cttsai/forked-lgbm-w-ideas-from-kernels-and-discuss\">Chia-Ta's kernels</a>. A small trick to add diversity of features is to only use detected==1 and redo the above feature engineering. So in the end I have two sets of features, one with whole times series and the other with detected==1 series.</p>\n\n<p>Unfortunately, my many other feature engineering attempts failed miserably. In the end, I have a mediocre lgb model with about 80 features of 0.58 cv and 1.05 LB. By the way, lgb always outperformed xgb by a small margin but the xgb gpu is just so much faster so in the end I used xgb gpu for experimental validation &amp; feature selection, and lgb for submission.</p>\n\n<h1>Stacking: non linear ensemble</h1>\n\n<p>As the great Kaggler Giba once said, stacking is all about diversity of base models. so I actually have three base models (level 1 modes):</p>\n\n<ol>\n<li>lgb as in <a href=\"https://www.kaggle.com/cttsai/forked-lgbm-w-ideas-from-kernels-and-discuss\">Chia-Ta's kernels</a> with whole series</li>\n<li>lgb with features from detected==1 and train for galaxy and ex-galaxy separately</li>\n<li>MLP as in <a href=\"https://www.kaggle.com/meaninglesslives/simple-neural-net-for-time-series-classification\">Siddhartha's kernel</a> with same features as 2. Mysteriously, this keras mlp is just better than my own TF implementation with the exact same network structure and hyperparameters.</li>\n</ol>\n\n<p>A key to successful non-linear stacking is to avoid overfitting. I have found two tricks for this:</p>\n\n<ol>\n<li><p>In level 1 base model, don't use the oof labels for early stopping. Instead, split the train data in that fold again to get its own validation set. This will degrade the level 1 model's score but it will help level 2 stacking.</p></li>\n<li><p>Make the 2nd level model simpler. For example, my 1st level lgb use depth of 7 and 3. In 2nd level lgb, the depth is always 1. The 1st level MLP has 4 layers but the 2nd level MLP only has 1 layer. And again, mysteriously, now it is my own TF MLP at 2nd level that is always better than the keras counterpart. In the end I just average the lgb/xgb and nn level 2 predictions with equal weights. </p></li>\n</ol>\n\n<p>At this stage, my stacking is 0.51 cv and 0.95 private LB which is 0.1 better than my best single model. So I have a good stacking framework setup and everything is great except for my non-impressive rank. Now <strong>what I need is just some magic</strong>.</p>\n\n<h1>RNN feature extraction with attention</h1>\n\n<p>The sudden jump of @hklee made me believe that deep learning could be it. I looked at the data again and the great posts by <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71949\">CPMP</a> and <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72646\">hklee</a>. I did the following preprocessing of flux:</p>\n\n<ol>\n<li>log transform the flux and calculate the diff(periods=1) and just use the diff</li>\n<li>calculate mjd the diff(periods=1) and just use the diff</li>\n<li>segment the time series and reset the mjd_diff and flux_diff at segment boundaries to 0. </li>\n<li>Insert random size of all zero data to the gap between the segments.</li>\n</ol>\n\n<p>And I use an RNN structure as the following:</p>\n\n<p>```</p>\n\n<p>with tf.variable_scope(name):</p>\n\n<pre><code>        net1,net2 = net[:,:,:-1],net[:,:,-1] # B, S, F\n\n        net2 = self._get_embedding(\"%s/passband\"%(name),net2,V,E)\n\n        net = tf.concat([net1,net2],axis=2)\n\n        state = None\n\n        net = self._bd_rnn_layer(net,\"%s/rnn3\"%name,cell_name,args,\n            state_fw=state,state_bw=state,output_size=-1,useproject=False)\n\n        args = {\"num_units\":H//4,\"activation\":'relu'}\n\n        self.next_flux_diff = self._bd_rnn_layer(net,\"%s/rnn5\"%name,cell_name,args,\n            state_fw=state,state_bw=state,output_size=1,useproject=True)[:,:-1,0]\n\n        args = {\"num_units\":H//4,\"activation\":'relu'}\n\n        net = self._bd_rnn_layer(net,\"%s/rnn4\"%name,cell_name,args,\n            state_fw=state,state_bw=state,output_size=-1,useproject=False)\n\n        w = self._get_variable(name, name='attn', shape=[1,args[\"num_units\"]])\n\n        w = tf.expand_dims(w,axis=0)\n\n        atten = tf.nn.softmax(w*net)\n\n        self.bottleneck = tf.reduce_sum(net*atten,axis=1)\n\n        net = self._fc(self.bottleneck, num_classes, layer_name='%s/out'%(name))\n</code></pre>\n\n<p>```</p>\n\n<p>I start to train this RNN using the 7K training data and as you can image it went nowhere.  <em>Deep learning demands big data unless it can be transferred.</em>  And we do have big test data with 3M samples. And then it struck me that I could use the pseudo label of test data for training purpose as my great kaggle buddy <a href=\"/xiaozhouwang\">@xiaozhouwang</a> did in <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47722\">tensorflow speech 3rd place</a>. Intuitively, I used my best model at that time (LB 0.95) to generate pseudo labels which is just np.argmax(test_predictions,axis=1) and <strong>Bazinga!</strong> it just worked! So as the chart shows I did this repeatedly for the last week of competion and my score is improved as [0.95 LB, 0.88, 0.85, 0.83, 0.82, 0.81, 0.80] and it just converged to where I am.</p>\n\n<p>Some tricks of this step:</p>\n\n<ol>\n<li><p>only provide time series data to the RNN without meta data. In this way, we force the RNN to focus on the timer series patterns only to fit a score that's obtained by all features and complex ensemble.</p></li>\n<li><p>use all the test data for training and use train data for validation and early stopping. The rational is that we expect the RNN to learn a middle ground to best fit the pseudo labels of test and true labels of train. </p></li>\n<li><p>multi objective function. in addition to the final classification, the model also predicts the next flux diff. It helps the model from cold start.</p></li>\n<li><p>Use different RNN cells in a interleaved way. For example, if GRU get the best score in this experiment, use its pseudo labels to train a LSTM  for the next experiment.</p></li>\n<li><p>Use bottleneck layer as features instead of the softmax output. Remember, we want the middle ground features that could be useful to train data, not the ones that fits the pseudo labels best.</p></li>\n</ol>\n\n<p>Yep, that's it. I'll also open source my solution on github. Thank you all. Happy kaggling! :D</p>",
  "messages": [
    {
      "id": "440750",
      "postDate": "12/18/2018 00:32:00",
      "content": "<p>Update: my solution will be posted <a href=\"https://github.com/daxiongshu/Rapids_PLAsTiCC_2018\">on github</a>. Now you can find a quick end2end demo with rapids. Forgive me that I'm always distracted to other stuff but I promise I'll get the whole solution done eventually!</p>\n\n<p>Big Congrats to the winners and new kaggle masters. I want to thank Olivier <a href=\"/ogrellier\">@ogrellier</a> and Kyle <a href=\"/kyleboone\">@kyleboone</a> specifically. Without your kernels, I can't even start this competition.</p>\n\n<p>Overall this competition is a fantastic learning process for me. If you have no astronomical background, this solution is for you. Considering I started climbing the LB from a model that is just 0.95 LB, what happened is a miracle :P</p>\n\n<p>Here is a preview of my solution and more to come! </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/440750/10889/image.png\" alt=\"Solution Overview\"></p>\n\n<p>Back to my laptop :P Ok, let me walk through this chart. </p>\n\n<h1>Baseline Models and feature engineering.</h1>\n\n<p>In case you haven't noticed, I'm using <a href=\"https://github.com/rapidsai/cudf\">cudf</a> to replace pandas for this competition whenever I can. In general, it is at least <strong>10x</strong> faster for csv reader and general groupby - aggregation <strong>with a single GPU</strong> than pandas on CPU. Sometimes I got <strong>100x</strong> speedup. I used cudf to generate most features in <a href=\"https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data/code\">Olivier's kernel</a>. However, as a merely two-month old library, you can expect some missing functionalities of cudf. I will start publish cudf tutorial in kaggle competitions and work with our developers to make your favorite pandas trick come true on GPU!  A demo notebook of using <strong>cudf</strong> will be published sometime tomorrow. Stay tuned!</p>\n\n<p>In addition, I also built some other features that gave me a small boost including decline rate from the peak, number of ups and downs of the flux curve, passband count percentage around the peak, <a href=\"https://www.kaggle.com/rejpalcz/feature-extraction-using-period-analysis\">period features</a> (optimized), and features from <a href=\"https://www.kaggle.com/cttsai/forked-lgbm-w-ideas-from-kernels-and-discuss\">Chia-Ta's kernels</a>. A small trick to add diversity of features is to only use detected==1 and redo the above feature engineering. So in the end I have two sets of features, one with whole times series and the other with detected==1 series.</p>\n\n<p>Unfortunately, my many other feature engineering attempts failed miserably. In the end, I have a mediocre lgb model with about 80 features of 0.58 cv and 1.05 LB. By the way, lgb always outperformed xgb by a small margin but the xgb gpu is just so much faster so in the end I used xgb gpu for experimental validation &amp; feature selection, and lgb for submission.</p>\n\n<h1>Stacking: non linear ensemble</h1>\n\n<p>As the great Kaggler Giba once said, stacking is all about diversity of base models. so I actually have three base models (level 1 modes):</p>\n\n<ol>\n<li>lgb as in <a href=\"https://www.kaggle.com/cttsai/forked-lgbm-w-ideas-from-kernels-and-discuss\">Chia-Ta's kernels</a> with whole series</li>\n<li>lgb with features from detected==1 and train for galaxy and ex-galaxy separately</li>\n<li>MLP as in <a href=\"https://www.kaggle.com/meaninglesslives/simple-neural-net-for-time-series-classification\">Siddhartha's kernel</a> with same features as 2. Mysteriously, this keras mlp is just better than my own TF implementation with the exact same network structure and hyperparameters.</li>\n</ol>\n\n<p>A key to successful non-linear stacking is to avoid overfitting. I have found two tricks for this:</p>\n\n<ol>\n<li><p>In level 1 base model, don't use the oof labels for early stopping. Instead, split the train data in that fold again to get its own validation set. This will degrade the level 1 model's score but it will help level 2 stacking.</p></li>\n<li><p>Make the 2nd level model simpler. For example, my 1st level lgb use depth of 7 and 3. In 2nd level lgb, the depth is always 1. The 1st level MLP has 4 layers but the 2nd level MLP only has 1 layer. And again, mysteriously, now it is my own TF MLP at 2nd level that is always better than the keras counterpart. In the end I just average the lgb/xgb and nn level 2 predictions with equal weights. </p></li>\n</ol>\n\n<p>At this stage, my stacking is 0.51 cv and 0.95 private LB which is 0.1 better than my best single model. So I have a good stacking framework setup and everything is great except for my non-impressive rank. Now <strong>what I need is just some magic</strong>.</p>\n\n<h1>RNN feature extraction with attention</h1>\n\n<p>The sudden jump of @hklee made me believe that deep learning could be it. I looked at the data again and the great posts by <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71949\">CPMP</a> and <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72646\">hklee</a>. I did the following preprocessing of flux:</p>\n\n<ol>\n<li>log transform the flux and calculate the diff(periods=1) and just use the diff</li>\n<li>calculate mjd the diff(periods=1) and just use the diff</li>\n<li>segment the time series and reset the mjd_diff and flux_diff at segment boundaries to 0. </li>\n<li>Insert random size of all zero data to the gap between the segments.</li>\n</ol>\n\n<p>And I use an RNN structure as the following:</p>\n\n<p>```</p>\n\n<p>with tf.variable_scope(name):</p>\n\n<pre><code>        net1,net2 = net[:,:,:-1],net[:,:,-1] # B, S, F\n\n        net2 = self._get_embedding(\"%s/passband\"%(name),net2,V,E)\n\n        net = tf.concat([net1,net2],axis=2)\n\n        state = None\n\n        net = self._bd_rnn_layer(net,\"%s/rnn3\"%name,cell_name,args,\n            state_fw=state,state_bw=state,output_size=-1,useproject=False)\n\n        args = {\"num_units\":H//4,\"activation\":'relu'}\n\n        self.next_flux_diff = self._bd_rnn_layer(net,\"%s/rnn5\"%name,cell_name,args,\n            state_fw=state,state_bw=state,output_size=1,useproject=True)[:,:-1,0]\n\n        args = {\"num_units\":H//4,\"activation\":'relu'}\n\n        net = self._bd_rnn_layer(net,\"%s/rnn4\"%name,cell_name,args,\n            state_fw=state,state_bw=state,output_size=-1,useproject=False)\n\n        w = self._get_variable(name, name='attn', shape=[1,args[\"num_units\"]])\n\n        w = tf.expand_dims(w,axis=0)\n\n        atten = tf.nn.softmax(w*net)\n\n        self.bottleneck = tf.reduce_sum(net*atten,axis=1)\n\n        net = self._fc(self.bottleneck, num_classes, layer_name='%s/out'%(name))\n</code></pre>\n\n<p>```</p>\n\n<p>I start to train this RNN using the 7K training data and as you can image it went nowhere.  <em>Deep learning demands big data unless it can be transferred.</em>  And we do have big test data with 3M samples. And then it struck me that I could use the pseudo label of test data for training purpose as my great kaggle buddy <a href=\"/xiaozhouwang\">@xiaozhouwang</a> did in <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47722\">tensorflow speech 3rd place</a>. Intuitively, I used my best model at that time (LB 0.95) to generate pseudo labels which is just np.argmax(test_predictions,axis=1) and <strong>Bazinga!</strong> it just worked! So as the chart shows I did this repeatedly for the last week of competion and my score is improved as [0.95 LB, 0.88, 0.85, 0.83, 0.82, 0.81, 0.80] and it just converged to where I am.</p>\n\n<p>Some tricks of this step:</p>\n\n<ol>\n<li><p>only provide time series data to the RNN without meta data. In this way, we force the RNN to focus on the timer series patterns only to fit a score that's obtained by all features and complex ensemble.</p></li>\n<li><p>use all the test data for training and use train data for validation and early stopping. The rational is that we expect the RNN to learn a middle ground to best fit the pseudo labels of test and true labels of train. </p></li>\n<li><p>multi objective function. in addition to the final classification, the model also predicts the next flux diff. It helps the model from cold start.</p></li>\n<li><p>Use different RNN cells in a interleaved way. For example, if GRU get the best score in this experiment, use its pseudo labels to train a LSTM  for the next experiment.</p></li>\n<li><p>Use bottleneck layer as features instead of the softmax output. Remember, we want the middle ground features that could be useful to train data, not the ones that fits the pseudo labels best.</p></li>\n</ol>\n\n<p>Yep, that's it. I'll also open source my solution on github. Thank you all. Happy kaggling! :D</p>",
      "rawMarkdown": "Update: my solution will be posted [on github][1]. Now you can find a quick end2end demo with rapids. Forgive me that I'm always distracted to other stuff but I promise I'll get the whole solution done eventually!\n\nBig Congrats to the winners and new kaggle masters. I want to thank Olivier @ogrellier and Kyle @kyleboone specifically. Without your kernels, I can't even start this competition.\n\nOverall this competition is a fantastic learning process for me. If you have no astronomical background, this solution is for you. Considering I started climbing the LB from a model that is just 0.95 LB, what happened is a miracle :P\n\nHere is a preview of my solution and more to come! \n\n![Solution Overview][2]\n\nBack to my laptop :P Ok, let me walk through this chart. \n\n# Baseline Models and feature engineering.\n\nIn case you haven't noticed, I'm using [cudf][3] to replace pandas for this competition whenever I can. In general, it is at least **10x** faster for csv reader and general groupby - aggregation **with a single GPU** than pandas on CPU. Sometimes I got **100x** speedup. I used cudf to generate most features in [Olivier's kernel][4]. However, as a merely two-month old library, you can expect some missing functionalities of cudf. I will start publish cudf tutorial in kaggle competitions and work with our developers to make your favorite pandas trick come true on GPU!  A demo notebook of using **cudf** will be published sometime tomorrow. Stay tuned!\n\nIn addition, I also built some other features that gave me a small boost including decline rate from the peak, number of ups and downs of the flux curve, passband count percentage around the peak, [period features][5] (optimized), and features from [Chia-Ta's kernels][6]. A small trick to add diversity of features is to only use detected==1 and redo the above feature engineering. So in the end I have two sets of features, one with whole times series and the other with detected==1 series.\n\nUnfortunately, my many other feature engineering attempts failed miserably. In the end, I have a mediocre lgb model with about 80 features of 0.58 cv and 1.05 LB. By the way, lgb always outperformed xgb by a small margin but the xgb gpu is just so much faster so in the end I used xgb gpu for experimental validation &amp; feature selection, and lgb for submission.\n\n#  Stacking: non linear ensemble\nAs the great Kaggler Giba once said, stacking is all about diversity of base models. so I actually have three base models (level 1 modes):\n\n1. lgb as in [Chia-Ta's kernels][7] with whole series\n2. lgb with features from detected==1 and train for galaxy and ex-galaxy separately\n3. MLP as in [Siddhartha's kernel][8] with same features as 2. Mysteriously, this keras mlp is just better than my own TF implementation with the exact same network structure and hyperparameters.\n\nA key to successful non-linear stacking is to avoid overfitting. I have found two tricks for this:\n\n1. In level 1 base model, don't use the oof labels for early stopping. Instead, split the train data in that fold again to get its own validation set. This will degrade the level 1 model's score but it will help level 2 stacking.\n\n2. Make the 2nd level model simpler. For example, my 1st level lgb use depth of 7 and 3. In 2nd level lgb, the depth is always 1. The 1st level MLP has 4 layers but the 2nd level MLP only has 1 layer. And again, mysteriously, now it is my own TF MLP at 2nd level that is always better than the keras counterpart. In the end I just average the lgb/xgb and nn level 2 predictions with equal weights. \n\nAt this stage, my stacking is 0.51 cv and 0.95 private LB which is 0.1 better than my best single model. So I have a good stacking framework setup and everything is great except for my non-impressive rank. Now **what I need is just some magic**.\n\n# RNN feature extraction with attention\n\nThe sudden jump of @hklee made me believe that deep learning could be it. I looked at the data again and the great posts by [CPMP][9] and [hklee][10]. I did the following preprocessing of flux:\n\n1. log transform the flux and calculate the diff(periods=1) and just use the diff\n2. calculate mjd the diff(periods=1) and just use the diff\n3. segment the time series and reset the mjd_diff and flux_diff at segment boundaries to 0. \n4. Insert random size of all zero data to the gap between the segments.\n\nAnd I use an RNN structure as the following:\n\n```\n\nwith tf.variable_scope(name):\n\n            net1,net2 = net[:,:,:-1],net[:,:,-1] # B, S, F\n\n            net2 = self._get_embedding(\"%s/passband\"%(name),net2,V,E)\n\n            net = tf.concat([net1,net2],axis=2)\n\n            state = None\n\n            net = self._bd_rnn_layer(net,\"%s/rnn3\"%name,cell_name,args,\n                state_fw=state,state_bw=state,output_size=-1,useproject=False)\n\n            args = {\"num_units\":H//4,\"activation\":'relu'}\n\n            self.next_flux_diff = self._bd_rnn_layer(net,\"%s/rnn5\"%name,cell_name,args,\n                state_fw=state,state_bw=state,output_size=1,useproject=True)[:,:-1,0]\n\n            args = {\"num_units\":H//4,\"activation\":'relu'}\n\n            net = self._bd_rnn_layer(net,\"%s/rnn4\"%name,cell_name,args,\n                state_fw=state,state_bw=state,output_size=-1,useproject=False)\n\n            w = self._get_variable(name, name='attn', shape=[1,args[\"num_units\"]])\n\n            w = tf.expand_dims(w,axis=0)\n\n            atten = tf.nn.softmax(w*net)\n\n            self.bottleneck = tf.reduce_sum(net*atten,axis=1)\n\n            net = self._fc(self.bottleneck, num_classes, layer_name='%s/out'%(name))\n```\n\nI start to train this RNN using the 7K training data and as you can image it went nowhere.  *Deep learning demands big data unless it can be transferred.*  And we do have big test data with 3M samples. And then it struck me that I could use the pseudo label of test data for training purpose as my great kaggle buddy @xiaozhouwang did in [tensorflow speech 3rd place][11]. Intuitively, I used my best model at that time (LB 0.95) to generate pseudo labels which is just np.argmax(test_predictions,axis=1) and **Bazinga!** it just worked! So as the chart shows I did this repeatedly for the last week of competion and my score is improved as [0.95 LB, 0.88, 0.85, 0.83, 0.82, 0.81, 0.80] and it just converged to where I am.\n\nSome tricks of this step:\n\n1. only provide time series data to the RNN without meta data. In this way, we force the RNN to focus on the timer series patterns only to fit a score that's obtained by all features and complex ensemble.\n\n2. use all the test data for training and use train data for validation and early stopping. The rational is that we expect the RNN to learn a middle ground to best fit the pseudo labels of test and true labels of train. \n\n3. multi objective function. in addition to the final classification, the model also predicts the next flux diff. It helps the model from cold start.\n\n4. Use different RNN cells in a interleaved way. For example, if GRU get the best score in this experiment, use its pseudo labels to train a LSTM  for the next experiment.\n\n5. Use bottleneck layer as features instead of the softmax output. Remember, we want the middle ground features that could be useful to train data, not the ones that fits the pseudo labels best.\n\nYep, that's it. I'll also open source my solution on github. Thank you all. Happy kaggling! :D\n\n\n  [1]: https://github.com/daxiongshu/Rapids_PLAsTiCC_2018\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/440750/10889/image.png\n  [3]: https://github.com/rapidsai/cudf\n  [4]: https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data/code\n  [5]: https://www.kaggle.com/rejpalcz/feature-extraction-using-period-analysis\n  [6]: https://www.kaggle.com/cttsai/forked-lgbm-w-ideas-from-kernels-and-discuss\n  [7]: https://www.kaggle.com/cttsai/forked-lgbm-w-ideas-from-kernels-and-discuss\n  [8]: https://www.kaggle.com/meaninglesslives/simple-neural-net-for-time-series-classification\n  [9]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71949\n  [10]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72646\n  [11]: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47722",
      "votes": null
    },
    {
      "id": "440757",
      "postDate": "12/18/2018 00:42:16",
      "content": "<p>这一招偷师舟神，向舟神致敬！\n顺便吐槽一下中间做不下去的时候，给5，6支队伍发了merge request都被拒绝了而且没有一个队给我发merge request，人缘也是差到爆了 ==</p>",
      "rawMarkdown": "这一招偷师舟神，向舟神致敬！\n顺便吐槽一下中间做不下去的时候，给5，6支队伍发了merge request都被拒绝了而且没有一个队给我发merge request，人缘也是差到爆了 ==",
      "votes": null
    },
    {
      "id": "440773",
      "postDate": "12/18/2018 01:11:47",
      "content": "<p>Congrats Jiwei for the solid #8 place! Great solution</p>",
      "rawMarkdown": "Congrats Jiwei for the solid #8 place! Great solution",
      "votes": null
    },
    {
      "id": "440775",
      "postDate": "12/18/2018 01:14:15",
      "content": "<p>Thank you so much! Looking forward to your solution sharing!</p>",
      "rawMarkdown": "Thank you so much! Looking forward to your solution sharing!",
      "votes": null
    },
    {
      "id": "440777",
      "postDate": "12/18/2018 01:14:55",
      "content": "<p>τι εννοεις;</p>",
      "rawMarkdown": "τι εννοεις;",
      "votes": null
    },
    {
      "id": "440784",
      "postDate": "12/18/2018 01:18:03",
      "content": "<p>祝贺大神！ PS：可否给个舟神的相关链接？入行不久，很多大师未曾耳闻</p>",
      "rawMarkdown": "祝贺大神！ PS：可否给个舟神的相关链接？入行不久，很多大师未曾耳闻",
      "votes": null
    },
    {
      "id": "440787",
      "postDate": "12/18/2018 01:20:17",
      "content": "<p>@ipraps Sorry about it. It is some personal feelings to share with my Chinese fella.</p>",
      "rawMarkdown": "ipraps Sorry about it. It is some personal feelings to share with my Chinese fella.",
      "votes": null
    },
    {
      "id": "440790",
      "postDate": "12/18/2018 01:22:29",
      "content": "<p>@SimonChen 恭喜恭喜你的🏅️，舟神就是  <a href=\"/xiaozhouwang\">@xiaozhouwang</a> 你已经关注啦！</p>",
      "rawMarkdown": "SimonChen 恭喜恭喜你的🏅️，舟神就是  @xiaozhouwang 你已经关注啦！",
      "votes": null
    },
    {
      "id": "440795",
      "postDate": "12/18/2018 01:24:55",
      "content": "<p>sorry for intervening, google translate returned some gibberish. Congrats by the way!</p>",
      "rawMarkdown": "sorry for intervening, google translate returned some gibberish. Congrats by the way!",
      "votes": null
    },
    {
      "id": "440799",
      "postDate": "12/18/2018 01:27:08",
      "content": "<p>感谢！！</p>",
      "rawMarkdown": "感谢！！",
      "votes": null
    },
    {
      "id": "440813",
      "postDate": "12/18/2018 01:50:34",
      "content": "<p>Gongrats! </p>",
      "rawMarkdown": "Gongrats!",
      "votes": null
    },
    {
      "id": "440824",
      "postDate": "12/18/2018 02:10:51",
      "content": "<p>Congratulations Jiwei!\nI have two questions about your pseudo-labeling.\nDid you use all the pseudo-label in the test set for the second stage of the training(RNN stage)?\nDid you mix the train set with the pseudo-labels for the second stage of the training(RNN stage)?</p>",
      "rawMarkdown": "Congratulations Jiwei!\nI have two questions about your pseudo-labeling.\nDid you use all the pseudo-label in the test set for the second stage of the training(RNN stage)?\nDid you mix the train set with the pseudo-labels for the second stage of the training(RNN stage)?",
      "votes": null
    },
    {
      "id": "440911",
      "postDate": "12/18/2018 04:47:34",
      "content": "<p>Very interesting approach. Starting with base models that returned 1.06 on LB and with RNN mix, you could stack and build such a powerful solution. Congratulations.\nIt will be a delight to see a mashup of your solution with Kyle Boone's. Guess it will break the 0.6  as well.</p>",
      "rawMarkdown": "Very interesting approach. Starting with base models that returned 1.06 on LB and with RNN mix, you could stack and build such a powerful solution. Congratulations.\nIt will be a delight to see a mashup of your solution with Kyle Boone's. Guess it will break the 0.6  as well.",
      "votes": null
    },
    {
      "id": "440925",
      "postDate": "12/18/2018 05:17:54",
      "content": "<p>I updated it. Did I answer your question?</p>",
      "rawMarkdown": "I updated it. Did I answer your question?",
      "votes": null
    },
    {
      "id": "440927",
      "postDate": "12/18/2018 05:29:27",
      "content": "<p>Congrats for gold medal. Nice solution..</p>",
      "rawMarkdown": "Congrats for gold medal. Nice solution..",
      "votes": null
    },
    {
      "id": "440936",
      "postDate": "12/18/2018 05:38:13",
      "content": "<p>Congrats &amp; Thanks for sharing! <br>\nRNN and pseudo-labeling strategy is really amazing!</p>",
      "rawMarkdown": "Congrats &amp; Thanks for sharing!  \nRNN and pseudo-labeling strategy is really amazing!",
      "votes": null
    },
    {
      "id": "440972",
      "postDate": "12/18/2018 06:06:18",
      "content": "<p>Yes it answered my question. <br>\nThanks for the detailed solution!</p>",
      "rawMarkdown": "Yes it answered my question.  \nThanks for the detailed solution!",
      "votes": null
    },
    {
      "id": "440983",
      "postDate": "12/18/2018 06:29:46",
      "content": "<p>Congrats Jiwei, Thanks for sharing the solution</p>",
      "rawMarkdown": "Congrats Jiwei, Thanks for sharing the solution",
      "votes": null
    },
    {
      "id": "440990",
      "postDate": "12/18/2018 06:43:00",
      "content": "<p>Congrats on your performance.  I am very surprised that pseudo labeling worked that well.  My team mates said they tried it and it didn't work for them. </p>",
      "rawMarkdown": "Congrats on your performance.  I am very surprised that pseudo labeling worked that well.  My team mates said they tried it and it didn't work for them.",
      "votes": null
    },
    {
      "id": "441029",
      "postDate": "12/18/2018 07:35:26",
      "content": "<p>Congratulations Jiwei Liu ! Amazing work.</p>\n\n<p>As I understand it, I would have been ranked 16 did I not publish my kernel  ;-) LOL</p>\n\n<p>I'm happy I did and I can read your solution write-up. I'll probably need another 3 months to fully understand what your work :)</p>",
      "rawMarkdown": "Congratulations Jiwei Liu ! Amazing work.\n\nAs I understand it, I would have been ranked 16 did I not publish my kernel  ;-) LOL\n\nI'm happy I did and I can read your solution write-up. I'll probably need another 3 months to fully understand what your work :)",
      "votes": null
    },
    {
      "id": "441403",
      "postDate": "12/18/2018 16:10:06",
      "content": "<p>Haha, you are the real hero! My sole standard to enter a kaggle competition is a consistent CV-LB. And I didn't have one until your kernel.</p>",
      "rawMarkdown": "Haha, you are the real hero! My sole standard to enter a kaggle competition is a consistent CV-LB. And I didn't have one until your kernel.",
      "votes": null
    },
    {
      "id": "441483",
      "postDate": "12/18/2018 17:36:48",
      "content": "<p><a href=\"/olivier\">@olivier</a> Thanks for sharing your kernel and it helped us developed more good performance kernels.</p>",
      "rawMarkdown": "olivier Thanks for sharing your kernel and it helped us developed more good performance kernels.",
      "votes": null
    },
    {
      "id": "441825",
      "postDate": "12/19/2018 06:16:38",
      "content": "<p><a href=\"/ogrellier\">@ogrellier</a> Thank you so much for sharing your kernel. If not for your Kernel I wouldn't have been able to start in this competition.</p>",
      "rawMarkdown": "ogrellier Thank you so much for sharing your kernel. If not for your Kernel I wouldn't have been able to start in this competition.",
      "votes": null
    },
    {
      "id": "443025",
      "postDate": "12/20/2018 22:48:54",
      "content": "<p>Thank you. It's a new approach for me, will try to understand how it works. (That's what a totally fail at the moment while looking it at, but hopefully your github will help.) Congrats, and thank you for sharing your code as well</p>",
      "rawMarkdown": "Thank you. It's a new approach for me, will try to understand how it works. (That's what a totally fail at the moment while looking it at, but hopefully your github will help.) Congrats, and thank you for sharing your code as well",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 440757,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "12/18/2018 00:42:16",
      "content": "<p>这一招偷师舟神，向舟神致敬！\n顺便吐槽一下中间做不下去的时候，给5，6支队伍发了merge request都被拒绝了而且没有一个队给我发merge request，人缘也是差到爆了 ==</p>",
      "votes": null,
      "replies": [
        {
          "id": 440777,
          "author_name": "iprapas",
          "author_url": "",
          "post_date": "12/18/2018 01:14:55",
          "content": "<p>τι εννοεις;</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440784,
          "author_name": "chenshaomeng",
          "author_url": "",
          "post_date": "12/18/2018 01:18:03",
          "content": "<p>祝贺大神！ PS：可否给个舟神的相关链接？入行不久，很多大师未曾耳闻</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440787,
          "author_name": "jiweiliu",
          "author_url": "",
          "post_date": "12/18/2018 01:20:17",
          "content": "<p>@ipraps Sorry about it. It is some personal feelings to share with my Chinese fella.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440790,
          "author_name": "jiweiliu",
          "author_url": "",
          "post_date": "12/18/2018 01:22:29",
          "content": "<p>@SimonChen 恭喜恭喜你的🏅️，舟神就是  <a href=\"/xiaozhouwang\">@xiaozhouwang</a> 你已经关注啦！</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440795,
          "author_name": "iprapas",
          "author_url": "",
          "post_date": "12/18/2018 01:24:55",
          "content": "<p>sorry for intervening, google translate returned some gibberish. Congrats by the way!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440799,
          "author_name": "chenshaomeng",
          "author_url": "",
          "post_date": "12/18/2018 01:27:08",
          "content": "<p>感谢！！</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 440773,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "12/18/2018 01:11:47",
      "content": "<p>Congrats Jiwei for the solid #8 place! Great solution</p>",
      "votes": null,
      "replies": [
        {
          "id": 440775,
          "author_name": "jiweiliu",
          "author_url": "",
          "post_date": "12/18/2018 01:14:15",
          "content": "<p>Thank you so much! Looking forward to your solution sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 440813,
      "author_name": "qqgeogor",
      "author_url": "",
      "post_date": "12/18/2018 01:50:34",
      "content": "<p>Gongrats! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 440824,
      "author_name": "pocketsuteado",
      "author_url": "",
      "post_date": "12/18/2018 02:10:51",
      "content": "<p>Congratulations Jiwei!\nI have two questions about your pseudo-labeling.\nDid you use all the pseudo-label in the test set for the second stage of the training(RNN stage)?\nDid you mix the train set with the pseudo-labels for the second stage of the training(RNN stage)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 440925,
          "author_name": "jiweiliu",
          "author_url": "",
          "post_date": "12/18/2018 05:17:54",
          "content": "<p>I updated it. Did I answer your question?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440972,
          "author_name": "pocketsuteado",
          "author_url": "",
          "post_date": "12/18/2018 06:06:18",
          "content": "<p>Yes it answered my question. <br>\nThanks for the detailed solution!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 440911,
      "author_name": "subrahmanyamv",
      "author_url": "",
      "post_date": "12/18/2018 04:47:34",
      "content": "<p>Very interesting approach. Starting with base models that returned 1.06 on LB and with RNN mix, you could stack and build such a powerful solution. Congratulations.\nIt will be a delight to see a mashup of your solution with Kyle Boone's. Guess it will break the 0.6  as well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 440927,
      "author_name": "adityakumarsinha",
      "author_url": "",
      "post_date": "12/18/2018 05:29:27",
      "content": "<p>Congrats for gold medal. Nice solution..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 440936,
      "author_name": "johnfarrell",
      "author_url": "",
      "post_date": "12/18/2018 05:38:13",
      "content": "<p>Congrats &amp; Thanks for sharing! <br>\nRNN and pseudo-labeling strategy is really amazing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 440983,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "12/18/2018 06:29:46",
      "content": "<p>Congrats Jiwei, Thanks for sharing the solution</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 440990,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/18/2018 06:43:00",
      "content": "<p>Congrats on your performance.  I am very surprised that pseudo labeling worked that well.  My team mates said they tried it and it didn't work for them. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 441029,
      "author_name": "ogrellier",
      "author_url": "",
      "post_date": "12/18/2018 07:35:26",
      "content": "<p>Congratulations Jiwei Liu ! Amazing work.</p>\n\n<p>As I understand it, I would have been ranked 16 did I not publish my kernel  ;-) LOL</p>\n\n<p>I'm happy I did and I can read your solution write-up. I'll probably need another 3 months to fully understand what your work :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 441403,
          "author_name": "jiweiliu",
          "author_url": "",
          "post_date": "12/18/2018 16:10:06",
          "content": "<p>Haha, you are the real hero! My sole standard to enter a kaggle competition is a consistent CV-LB. And I didn't have one until your kernel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441483,
          "author_name": "cttsai",
          "author_url": "",
          "post_date": "12/18/2018 17:36:48",
          "content": "<p><a href=\"/olivier\">@olivier</a> Thanks for sharing your kernel and it helped us developed more good performance kernels.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441825,
          "author_name": "subrahmanyamv",
          "author_url": "",
          "post_date": "12/19/2018 06:16:38",
          "content": "<p><a href=\"/ogrellier\">@ogrellier</a> Thank you so much for sharing your kernel. If not for your Kernel I wouldn't have been able to start in this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 443025,
      "author_name": "blondinka",
      "author_url": "",
      "post_date": "12/20/2018 22:48:54",
      "content": "<p>Thank you. It's a new approach for me, will try to understand how it works. (That's what a totally fail at the moment while looking it at, but hopefully your github will help.) Congrats, and thank you for sharing your code as well</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "440750": "Update: my solution will be posted [on github][1]. Now you can find a quick end2end demo with rapids. Forgive me that I'm always distracted to other stuff but I promise I'll get the whole solution done eventually!\n\nBig Congrats to the winners and new kaggle masters. I want to thank Olivier @ogrellier and Kyle @kyleboone specifically. Without your kernels, I can't even start this competition.\n\nOverall this competition is a fantastic learning process for me. If you have no astronomical background, this solution is for you. Considering I started climbing the LB from a model that is just 0.95 LB, what happened is a miracle :P\n\nHere is a preview of my solution and more to come! \n\n![Solution Overview][2]\n\nBack to my laptop :P Ok, let me walk through this chart. \n\n# Baseline Models and feature engineering.\n\nIn case you haven't noticed, I'm using [cudf][3] to replace pandas for this competition whenever I can. In general, it is at least **10x** faster for csv reader and general groupby - aggregation **with a single GPU** than pandas on CPU. Sometimes I got **100x** speedup. I used cudf to generate most features in [Olivier's kernel][4]. However, as a merely two-month old library, you can expect some missing functionalities of cudf. I will start publish cudf tutorial in kaggle competitions and work with our developers to make your favorite pandas trick come true on GPU!  A demo notebook of using **cudf** will be published sometime tomorrow. Stay tuned!\n\nIn addition, I also built some other features that gave me a small boost including decline rate from the peak, number of ups and downs of the flux curve, passband count percentage around the peak, [period features][5] (optimized), and features from [Chia-Ta's kernels][6]. A small trick to add diversity of features is to only use detected==1 and redo the above feature engineering. So in the end I have two sets of features, one with whole times series and the other with detected==1 series.\n\nUnfortunately, my many other feature engineering attempts failed miserably. In the end, I have a mediocre lgb model with about 80 features of 0.58 cv and 1.05 LB. By the way, lgb always outperformed xgb by a small margin but the xgb gpu is just so much faster so in the end I used xgb gpu for experimental validation &amp; feature selection, and lgb for submission.\n\n#  Stacking: non linear ensemble\nAs the great Kaggler Giba once said, stacking is all about diversity of base models. so I actually have three base models (level 1 modes):\n\n1. lgb as in [Chia-Ta's kernels][7] with whole series\n2. lgb with features from detected==1 and train for galaxy and ex-galaxy separately\n3. MLP as in [Siddhartha's kernel][8] with same features as 2. Mysteriously, this keras mlp is just better than my own TF implementation with the exact same network structure and hyperparameters.\n\nA key to successful non-linear stacking is to avoid overfitting. I have found two tricks for this:\n\n1. In level 1 base model, don't use the oof labels for early stopping. Instead, split the train data in that fold again to get its own validation set. This will degrade the level 1 model's score but it will help level 2 stacking.\n\n2. Make the 2nd level model simpler. For example, my 1st level lgb use depth of 7 and 3. In 2nd level lgb, the depth is always 1. The 1st level MLP has 4 layers but the 2nd level MLP only has 1 layer. And again, mysteriously, now it is my own TF MLP at 2nd level that is always better than the keras counterpart. In the end I just average the lgb/xgb and nn level 2 predictions with equal weights. \n\nAt this stage, my stacking is 0.51 cv and 0.95 private LB which is 0.1 better than my best single model. So I have a good stacking framework setup and everything is great except for my non-impressive rank. Now **what I need is just some magic**.\n\n# RNN feature extraction with attention\n\nThe sudden jump of @hklee made me believe that deep learning could be it. I looked at the data again and the great posts by [CPMP][9] and [hklee][10]. I did the following preprocessing of flux:\n\n1. log transform the flux and calculate the diff(periods=1) and just use the diff\n2. calculate mjd the diff(periods=1) and just use the diff\n3. segment the time series and reset the mjd_diff and flux_diff at segment boundaries to 0. \n4. Insert random size of all zero data to the gap between the segments.\n\nAnd I use an RNN structure as the following:\n\n```\n\nwith tf.variable_scope(name):\n\n            net1,net2 = net[:,:,:-1],net[:,:,-1] # B, S, F\n\n            net2 = self._get_embedding(\"%s/passband\"%(name),net2,V,E)\n\n            net = tf.concat([net1,net2],axis=2)\n\n            state = None\n\n            net = self._bd_rnn_layer(net,\"%s/rnn3\"%name,cell_name,args,\n                state_fw=state,state_bw=state,output_size=-1,useproject=False)\n\n            args = {\"num_units\":H//4,\"activation\":'relu'}\n\n            self.next_flux_diff = self._bd_rnn_layer(net,\"%s/rnn5\"%name,cell_name,args,\n                state_fw=state,state_bw=state,output_size=1,useproject=True)[:,:-1,0]\n\n            args = {\"num_units\":H//4,\"activation\":'relu'}\n\n            net = self._bd_rnn_layer(net,\"%s/rnn4\"%name,cell_name,args,\n                state_fw=state,state_bw=state,output_size=-1,useproject=False)\n\n            w = self._get_variable(name, name='attn', shape=[1,args[\"num_units\"]])\n\n            w = tf.expand_dims(w,axis=0)\n\n            atten = tf.nn.softmax(w*net)\n\n            self.bottleneck = tf.reduce_sum(net*atten,axis=1)\n\n            net = self._fc(self.bottleneck, num_classes, layer_name='%s/out'%(name))\n```\n\nI start to train this RNN using the 7K training data and as you can image it went nowhere.  *Deep learning demands big data unless it can be transferred.*  And we do have big test data with 3M samples. And then it struck me that I could use the pseudo label of test data for training purpose as my great kaggle buddy @xiaozhouwang did in [tensorflow speech 3rd place][11]. Intuitively, I used my best model at that time (LB 0.95) to generate pseudo labels which is just np.argmax(test_predictions,axis=1) and **Bazinga!** it just worked! So as the chart shows I did this repeatedly for the last week of competion and my score is improved as [0.95 LB, 0.88, 0.85, 0.83, 0.82, 0.81, 0.80] and it just converged to where I am.\n\nSome tricks of this step:\n\n1. only provide time series data to the RNN without meta data. In this way, we force the RNN to focus on the timer series patterns only to fit a score that's obtained by all features and complex ensemble.\n\n2. use all the test data for training and use train data for validation and early stopping. The rational is that we expect the RNN to learn a middle ground to best fit the pseudo labels of test and true labels of train. \n\n3. multi objective function. in addition to the final classification, the model also predicts the next flux diff. It helps the model from cold start.\n\n4. Use different RNN cells in a interleaved way. For example, if GRU get the best score in this experiment, use its pseudo labels to train a LSTM  for the next experiment.\n\n5. Use bottleneck layer as features instead of the softmax output. Remember, we want the middle ground features that could be useful to train data, not the ones that fits the pseudo labels best.\n\nYep, that's it. I'll also open source my solution on github. Thank you all. Happy kaggling! :D\n\n\n  [1]: https://github.com/daxiongshu/Rapids_PLAsTiCC_2018\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/440750/10889/image.png\n  [3]: https://github.com/rapidsai/cudf\n  [4]: https://www.kaggle.com/ogrellier/plasticc-in-a-kernel-meta-and-data/code\n  [5]: https://www.kaggle.com/rejpalcz/feature-extraction-using-period-analysis\n  [6]: https://www.kaggle.com/cttsai/forked-lgbm-w-ideas-from-kernels-and-discuss\n  [7]: https://www.kaggle.com/cttsai/forked-lgbm-w-ideas-from-kernels-and-discuss\n  [8]: https://www.kaggle.com/meaninglesslives/simple-neural-net-for-time-series-classification\n  [9]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71949\n  [10]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/72646\n  [11]: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47722",
    "440757": "这一招偷师舟神，向舟神致敬！\n顺便吐槽一下中间做不下去的时候，给5，6支队伍发了merge request都被拒绝了而且没有一个队给我发merge request，人缘也是差到爆了 ==",
    "440773": "Congrats Jiwei for the solid #8 place! Great solution",
    "440775": "Thank you so much! Looking forward to your solution sharing!",
    "440777": "τι εννοεις;",
    "440784": "祝贺大神！ PS：可否给个舟神的相关链接？入行不久，很多大师未曾耳闻",
    "440787": "ipraps Sorry about it. It is some personal feelings to share with my Chinese fella.",
    "440790": "SimonChen 恭喜恭喜你的🏅️，舟神就是  @xiaozhouwang 你已经关注啦！",
    "440795": "sorry for intervening, google translate returned some gibberish. Congrats by the way!",
    "440799": "感谢！！",
    "440813": "Gongrats!",
    "440824": "Congratulations Jiwei!\nI have two questions about your pseudo-labeling.\nDid you use all the pseudo-label in the test set for the second stage of the training(RNN stage)?\nDid you mix the train set with the pseudo-labels for the second stage of the training(RNN stage)?",
    "440911": "Very interesting approach. Starting with base models that returned 1.06 on LB and with RNN mix, you could stack and build such a powerful solution. Congratulations.\nIt will be a delight to see a mashup of your solution with Kyle Boone's. Guess it will break the 0.6  as well.",
    "440925": "I updated it. Did I answer your question?",
    "440927": "Congrats for gold medal. Nice solution..",
    "440936": "Congrats &amp; Thanks for sharing!  \nRNN and pseudo-labeling strategy is really amazing!",
    "440972": "Yes it answered my question.  \nThanks for the detailed solution!",
    "440983": "Congrats Jiwei, Thanks for sharing the solution",
    "440990": "Congrats on your performance.  I am very surprised that pseudo labeling worked that well.  My team mates said they tried it and it didn't work for them.",
    "441029": "Congratulations Jiwei Liu ! Amazing work.\n\nAs I understand it, I would have been ranked 16 did I not publish my kernel  ;-) LOL\n\nI'm happy I did and I can read your solution write-up. I'll probably need another 3 months to fully understand what your work :)",
    "441403": "Haha, you are the real hero! My sole standard to enter a kaggle competition is a consistent CV-LB. And I didn't have one until your kernel.",
    "441483": "olivier Thanks for sharing your kernel and it helped us developed more good performance kernels.",
    "441825": "ogrellier Thank you so much for sharing your kernel. If not for your Kernel I wouldn't have been able to start in this competition.",
    "443025": "Thank you. It's a new approach for me, will try to understand how it works. (That's what a totally fail at the moment while looking it at, but hopefully your github will help.) Congrats, and thank you for sharing your code as well"
  },
  "source": "meta"
}