{
  "id": 86839,
  "title": "Are there any good deep learning models that might fit for the prediction?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/86839",
  "author_name": "",
  "post_date": "2019-03-27T04:58:11.759994900Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": "501245",
      "postDate": "03/27/2019 04:58:11",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "503367",
      "postDate": "03/29/2019 21:37:24",
      "content": "<p>I see a few issues with deep learning.</p>\n\n<p>The first is that, if you follow most everyone's kernel in regards to feature generation, you actually only have 4,194 rows (due to the way they sample each 150,000 rows to mimic the test data set). </p>\n\n<p>Good luck trying to get any deep learning model to perform well with so little in terms of number of observations.</p>\n\n<p>Second is that deep learning, at least historically, generally is more preferred for data that is unstructured/has very little feature engineering. Again, if you follow everyone's kernel, feature engineering is through the wazoo, which is where things like xgboost/lightgbm tend to perform better.</p>\n\n<p>That's not to say that deep learning can't or won't work (I mean, for all I know, everyone at the top of leaderboard could be using some form of NNs), but it's going to require some out-of-the-box thinking if you want it to do well. </p>",
      "rawMarkdown": "I see a few issues with deep learning.\n\nThe first is that, if you follow most everyone's kernel in regards to feature generation, you actually only have 4,194 rows (due to the way they sample each 150,000 rows to mimic the test data set). \n\nGood luck trying to get any deep learning model to perform well with so little in terms of number of observations.\n\nSecond is that deep learning, at least historically, generally is more preferred for data that is unstructured/has very little feature engineering. Again, if you follow everyone's kernel, feature engineering is through the wazoo, which is where things like xgboost/lightgbm tend to perform better.\n\nThat's not to say that deep learning can't or won't work (I mean, for all I know, everyone at the top of leaderboard could be using some form of NNs), but it's going to require some out-of-the-box thinking if you want it to do well.",
      "votes": null
    },
    {
      "id": "511060",
      "postDate": "04/09/2019 18:07:35",
      "content": "<p>Interesting observations! However what if one is to produce training samples with overlapping? (for example training instance 1 covers row 200,000 to 350,000; training instance 2 covers row 300,000 to 450,000) - can we not then create more samples for training, or is it incorrect to treat data in such a way?</p>",
      "rawMarkdown": "Interesting observations! However what if one is to produce training samples with overlapping? (for example training instance 1 covers row 200,000 to 350,000; training instance 2 covers row 300,000 to 450,000) - can we not then create more samples for training, or is it incorrect to treat data in such a way?",
      "votes": null
    },
    {
      "id": "512047",
      "postDate": "04/10/2019 14:26:20",
      "content": "<p>Yeah that is okay I think as long as you are careful. Avoid leakage in the validation data.</p>",
      "rawMarkdown": "Yeah that is okay I think as long as you are careful. Avoid leakage in the validation data.",
      "votes": null
    },
    {
      "id": "514458",
      "postDate": "04/11/2019 16:42:37",
      "content": "<p>Indeed - i wonder how best to think about the \"degree of freedoms\" of the problem at hand. Ostensibly it is a 150k dimension problem but in reality the serial correlation probably means it DOF is much lower. Presumably the optimal model would have similar # of parameters as this right? (which can then be used to rule out overly deep NN)</p>",
      "rawMarkdown": "Indeed - i wonder how best to think about the \"degree of freedoms\" of the problem at hand. Ostensibly it is a 150k dimension problem but in reality the serial correlation probably means it DOF is much lower. Presumably the optimal model would have similar # of parameters as this right? (which can then be used to rule out overly deep NN)",
      "votes": null
    },
    {
      "id": "515706",
      "postDate": "04/13/2019 01:45:51",
      "content": "<p>I'm not sure whether this could be called success, but so far I've managed to get an LB score of 1.530 from an ensemble of 9 fairly simple GRU recurrent neural net models created with R-&gt;Keras-&gt;Tensorflow.  For each model I partitioned the training data into 20000 150000-point segments with randomly-selected starting points (meaning they do overlap).  I divided each segment into 150 consecutive bins, and for each bin computed 13 features composed of simple statistics, results of acoustic wave analysis algorithms, etc.  I trained my models using Keras's \"train_on_batch\" function.\nBy averaging 2 parts of a lightgbm submission that scored 1.494 on LB with 1 part of the GRU submission I got my current best LB score of 1.489.</p>",
      "rawMarkdown": "I'm not sure whether this could be called success, but so far I've managed to get an LB score of 1.530 from an ensemble of 9 fairly simple GRU recurrent neural net models created with R-&gt;Keras-&gt;Tensorflow.  For each model I partitioned the training data into 20000 150000-point segments with randomly-selected starting points (meaning they do overlap).  I divided each segment into 150 consecutive bins, and for each bin computed 13 features composed of simple statistics, results of acoustic wave analysis algorithms, etc.  I trained my models using Keras's \"train_on_batch\" function.\nBy averaging 2 parts of a lightgbm submission that scored 1.494 on LB with 1 part of the GRU submission I got my current best LB score of 1.489.",
      "votes": null
    },
    {
      "id": "518182",
      "postDate": "04/17/2019 00:19:00",
      "content": "<p>I used LSTM with 36 features and achieved 1.509 on LB. I'm newbie to the lightgbm. I'm wondering if it's possible to share some more insight about how to pair lightgbm with GRU?</p>",
      "rawMarkdown": "I used LSTM with 36 features and achieved 1.509 on LB. I'm newbie to the lightgbm. I'm wondering if it's possible to share some more insight about how to pair lightgbm with GRU?",
      "votes": null
    },
    {
      "id": "518227",
      "postDate": "04/17/2019 01:45:03",
      "content": "<p>If you look at the Santander Customer Transaction competition, there's a method called <code>Knowledge Distillation</code> that a few teams used to combine NN's and LightGBM. Besides knowledge distillation, there are a bunch of good examples of NN/ LightGbm hybrids that got Gold medals in the competiton</p>",
      "rawMarkdown": "If you look at the Santander Customer Transaction competition, there's a method called `Knowledge Distillation` that a few teams used to combine NN's and LightGBM. Besides knowledge distillation, there are a bunch of good examples of NN/ LightGbm hybrids that got Gold medals in the competiton",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 503367,
      "author_name": "conormcnamara",
      "author_url": "",
      "post_date": "03/29/2019 21:37:24",
      "content": "<p>I see a few issues with deep learning.</p>\n\n<p>The first is that, if you follow most everyone's kernel in regards to feature generation, you actually only have 4,194 rows (due to the way they sample each 150,000 rows to mimic the test data set). </p>\n\n<p>Good luck trying to get any deep learning model to perform well with so little in terms of number of observations.</p>\n\n<p>Second is that deep learning, at least historically, generally is more preferred for data that is unstructured/has very little feature engineering. Again, if you follow everyone's kernel, feature engineering is through the wazoo, which is where things like xgboost/lightgbm tend to perform better.</p>\n\n<p>That's not to say that deep learning can't or won't work (I mean, for all I know, everyone at the top of leaderboard could be using some form of NNs), but it's going to require some out-of-the-box thinking if you want it to do well. </p>",
      "votes": null,
      "replies": [
        {
          "id": 511060,
          "author_name": "heisenger",
          "author_url": "",
          "post_date": "04/09/2019 18:07:35",
          "content": "<p>Interesting observations! However what if one is to produce training samples with overlapping? (for example training instance 1 covers row 200,000 to 350,000; training instance 2 covers row 300,000 to 450,000) - can we not then create more samples for training, or is it incorrect to treat data in such a way?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 512047,
          "author_name": "aceplayer11",
          "author_url": "",
          "post_date": "04/10/2019 14:26:20",
          "content": "<p>Yeah that is okay I think as long as you are careful. Avoid leakage in the validation data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 514458,
          "author_name": "heisenger",
          "author_url": "",
          "post_date": "04/11/2019 16:42:37",
          "content": "<p>Indeed - i wonder how best to think about the \"degree of freedoms\" of the problem at hand. Ostensibly it is a 150k dimension problem but in reality the serial correlation probably means it DOF is much lower. Presumably the optimal model would have similar # of parameters as this right? (which can then be used to rule out overly deep NN)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 515706,
          "author_name": "dslate",
          "author_url": "",
          "post_date": "04/13/2019 01:45:51",
          "content": "<p>I'm not sure whether this could be called success, but so far I've managed to get an LB score of 1.530 from an ensemble of 9 fairly simple GRU recurrent neural net models created with R-&gt;Keras-&gt;Tensorflow.  For each model I partitioned the training data into 20000 150000-point segments with randomly-selected starting points (meaning they do overlap).  I divided each segment into 150 consecutive bins, and for each bin computed 13 features composed of simple statistics, results of acoustic wave analysis algorithms, etc.  I trained my models using Keras's \"train_on_batch\" function.\nBy averaging 2 parts of a lightgbm submission that scored 1.494 on LB with 1 part of the GRU submission I got my current best LB score of 1.489.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 518182,
          "author_name": "zhuchunlin1995",
          "author_url": "",
          "post_date": "04/17/2019 00:19:00",
          "content": "<p>I used LSTM with 36 features and achieved 1.509 on LB. I'm newbie to the lightgbm. I'm wondering if it's possible to share some more insight about how to pair lightgbm with GRU?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 518227,
          "author_name": "halldalton94",
          "author_url": "",
          "post_date": "04/17/2019 01:45:03",
          "content": "<p>If you look at the Santander Customer Transaction competition, there's a method called <code>Knowledge Distillation</code> that a few teams used to combine NN's and LightGBM. Besides knowledge distillation, there are a bunch of good examples of NN/ LightGbm hybrids that got Gold medals in the competiton</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "501245": "",
    "503367": "I see a few issues with deep learning.\n\nThe first is that, if you follow most everyone's kernel in regards to feature generation, you actually only have 4,194 rows (due to the way they sample each 150,000 rows to mimic the test data set). \n\nGood luck trying to get any deep learning model to perform well with so little in terms of number of observations.\n\nSecond is that deep learning, at least historically, generally is more preferred for data that is unstructured/has very little feature engineering. Again, if you follow everyone's kernel, feature engineering is through the wazoo, which is where things like xgboost/lightgbm tend to perform better.\n\nThat's not to say that deep learning can't or won't work (I mean, for all I know, everyone at the top of leaderboard could be using some form of NNs), but it's going to require some out-of-the-box thinking if you want it to do well.",
    "511060": "Interesting observations! However what if one is to produce training samples with overlapping? (for example training instance 1 covers row 200,000 to 350,000; training instance 2 covers row 300,000 to 450,000) - can we not then create more samples for training, or is it incorrect to treat data in such a way?",
    "512047": "Yeah that is okay I think as long as you are careful. Avoid leakage in the validation data.",
    "514458": "Indeed - i wonder how best to think about the \"degree of freedoms\" of the problem at hand. Ostensibly it is a 150k dimension problem but in reality the serial correlation probably means it DOF is much lower. Presumably the optimal model would have similar # of parameters as this right? (which can then be used to rule out overly deep NN)",
    "515706": "I'm not sure whether this could be called success, but so far I've managed to get an LB score of 1.530 from an ensemble of 9 fairly simple GRU recurrent neural net models created with R-&gt;Keras-&gt;Tensorflow.  For each model I partitioned the training data into 20000 150000-point segments with randomly-selected starting points (meaning they do overlap).  I divided each segment into 150 consecutive bins, and for each bin computed 13 features composed of simple statistics, results of acoustic wave analysis algorithms, etc.  I trained my models using Keras's \"train_on_batch\" function.\nBy averaging 2 parts of a lightgbm submission that scored 1.494 on LB with 1 part of the GRU submission I got my current best LB score of 1.489.",
    "518182": "I used LSTM with 36 features and achieved 1.509 on LB. I'm newbie to the lightgbm. I'm wondering if it's possible to share some more insight about how to pair lightgbm with GRU?",
    "518227": "If you look at the Santander Customer Transaction competition, there's a method called `Knowledge Distillation` that a few teams used to combine NN's and LightGBM. Besides knowledge distillation, there are a bunch of good examples of NN/ LightGbm hybrids that got Gold medals in the competiton"
  },
  "source": "meta"
}