{
  "id": 79247,
  "title": "How much overlap in creating training data?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/79247",
  "author_name": "",
  "post_date": "2019-02-01T17:01:51.207846700Z",
  "votes": 27,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Most public kernels work like this:</p>\n\n<ol>\n<li><p>Split time series into chunks of length 150'000 to match the size of the test pieces</p></li>\n<li><p>From each chunk, derive one line of training data</p></li>\n</ol>\n\n<p>One of the many choices to make is: Should the chunks in Step 1 overlap? If yes, by how much? </p>\n\n<p>The larger the overlap (resp. the smaller the stride), the more lines in the training set. But that does not necessarily mean a better performance because of increased statistical dependence of the lines... Where is the optimum given a set of features (and the chosen learning algorithms)?</p>\n\n<p>I will test at lest the following strides</p>\n\n<ul>\n<li><p>150'000 (no overlap), </p></li>\n<li><p>5'000 (97% overlap), </p></li>\n<li><p>10'000 (93% overlap) and </p></li>\n<li><p>15'000 (90% overlap between successive chunks) and post the results here over the weekend. </p></li>\n</ul>\n\n<p>To get the LB scores, I will plug 35 standard features into an untuned random forest using <a href=\"https://www.kaggle.com/mayer79/earthquake-forecast-with-random-forests\">https://www.kaggle.com/mayer79/earthquake-forecast-with-random-forests</a> .  CV scores are mean absolute errors of leave-one-out earthquake validation as suggested by Elliot. </p>\n\n<p>Any experience on this question you want to share? What is your impression?</p>\n\n<p>|Stride|Overlap|CV|LB|\n|-------|----------|---|----|\n|5'000|97%|2.333|1.589|\n|10'000|93%|2.297|1.560|\n|15'000|90%|2.294|1.548|\n|30'000|80%|2.294|1.543|\n|50'000|66%|2.284|1.537|\n|75'000|50%|2.283|<strong>1.522</strong>|\n|100'000|33%|2.282|1.531|\n|150'000|0%|<strong>2.267</strong>|1.548|</p>\n\n<p>Using the fixed feature set of above mentioned kernel, \"leave-one-earthquake-out\" cross-validation MAE worsens with higher overlap. LB scores might benefit from small overlaps between 33% and 67%. But honestly, I trust neither of the two scores in this competition...</p>\n\n<p>All random forests were run with 500 trees. Would more trees help with dependent rows? For the setting with 30'000 stride length, using 2000 trees resulted in the following, only slightly, better scores:</p>\n\n<ul>\n<li><p>CV: 2.292 </p></li>\n<li><p>LB 1.541 </p></li>\n</ul>",
  "messages": [
    {
      "id": "464859",
      "postDate": "02/01/2019 17:01:51",
      "content": "<p>Most public kernels work like this:</p>\n\n<ol>\n<li><p>Split time series into chunks of length 150'000 to match the size of the test pieces</p></li>\n<li><p>From each chunk, derive one line of training data</p></li>\n</ol>\n\n<p>One of the many choices to make is: Should the chunks in Step 1 overlap? If yes, by how much? </p>\n\n<p>The larger the overlap (resp. the smaller the stride), the more lines in the training set. But that does not necessarily mean a better performance because of increased statistical dependence of the lines... Where is the optimum given a set of features (and the chosen learning algorithms)?</p>\n\n<p>I will test at lest the following strides</p>\n\n<ul>\n<li><p>150'000 (no overlap), </p></li>\n<li><p>5'000 (97% overlap), </p></li>\n<li><p>10'000 (93% overlap) and </p></li>\n<li><p>15'000 (90% overlap between successive chunks) and post the results here over the weekend. </p></li>\n</ul>\n\n<p>To get the LB scores, I will plug 35 standard features into an untuned random forest using <a href=\"https://www.kaggle.com/mayer79/earthquake-forecast-with-random-forests\">https://www.kaggle.com/mayer79/earthquake-forecast-with-random-forests</a> .  CV scores are mean absolute errors of leave-one-out earthquake validation as suggested by Elliot. </p>\n\n<p>Any experience on this question you want to share? What is your impression?</p>\n\n<p>|Stride|Overlap|CV|LB|\n|-------|----------|---|----|\n|5'000|97%|2.333|1.589|\n|10'000|93%|2.297|1.560|\n|15'000|90%|2.294|1.548|\n|30'000|80%|2.294|1.543|\n|50'000|66%|2.284|1.537|\n|75'000|50%|2.283|<strong>1.522</strong>|\n|100'000|33%|2.282|1.531|\n|150'000|0%|<strong>2.267</strong>|1.548|</p>\n\n<p>Using the fixed feature set of above mentioned kernel, \"leave-one-earthquake-out\" cross-validation MAE worsens with higher overlap. LB scores might benefit from small overlaps between 33% and 67%. But honestly, I trust neither of the two scores in this competition...</p>\n\n<p>All random forests were run with 500 trees. Would more trees help with dependent rows? For the setting with 30'000 stride length, using 2000 trees resulted in the following, only slightly, better scores:</p>\n\n<ul>\n<li><p>CV: 2.292 </p></li>\n<li><p>LB 1.541 </p></li>\n</ul>",
      "rawMarkdown": "Most public kernels work like this:\n\n1. Split time series into chunks of length 150'000 to match the size of the test pieces\n\n2. From each chunk, derive one line of training data\n\nOne of the many choices to make is: Should the chunks in Step 1 overlap? If yes, by how much? \n\nThe larger the overlap (resp. the smaller the stride), the more lines in the training set. But that does not necessarily mean a better performance because of increased statistical dependence of the lines... Where is the optimum given a set of features (and the chosen learning algorithms)?\n\nI will test at lest the following strides\n\n- 150'000 (no overlap), \n\n- 5'000 (97% overlap), \n\n- 10'000 (93% overlap) and \n\n- 15'000 (90% overlap between successive chunks) and post the results here over the weekend. \n\nTo get the LB scores, I will plug 35 standard features into an untuned random forest using https://www.kaggle.com/mayer79/earthquake-forecast-with-random-forests .  CV scores are mean absolute errors of leave-one-out earthquake validation as suggested by Elliot. \n\nAny experience on this question you want to share? What is your impression?\n\n|Stride|Overlap|CV|LB|\n|-------|----------|---|----|\n|5'000|97%|2.333|1.589|\n|10'000|93%|2.297|1.560|\n|15'000|90%|2.294|1.548|\n|30'000|80%|2.294|1.543|\n|50'000|66%|2.284|1.537|\n|75'000|50%|2.283|**1.522**|\n|100'000|33%|2.282|1.531|\n|150'000|0%|**2.267**|1.548|\n\nUsing the fixed feature set of above mentioned kernel, \"leave-one-earthquake-out\" cross-validation MAE worsens with higher overlap. LB scores might benefit from small overlaps between 33% and 67%. But honestly, I trust neither of the two scores in this competition...\n\nAll random forests were run with 500 trees. Would more trees help with dependent rows? For the setting with 30'000 stride length, using 2000 trees resulted in the following, only slightly, better scores:\n\n- CV: 2.292 \n\n- LB 1.541",
      "votes": null
    },
    {
      "id": "464923",
      "postDate": "02/01/2019 20:20:28",
      "content": "<p>I have tested with stride of 75'000:\n<a href=\"https://www.kaggle.com/alinealmeida/basic-feature-benchmark-with-quantiles-augmenting\">https://www.kaggle.com/alinealmeida/basic-feature-benchmark-with-quantiles-augmenting</a></p>\n\n<p>Apparently this approach does not change the results significantly (at least not with the basic features used: ave, std, max, min, q95, q99, q05 and q01).</p>\n\n<p>I created a custom split iterator in order to avoid segments in training which share regions with segments in validation set.</p>",
      "rawMarkdown": "I have tested with stride of 75'000:\nhttps://www.kaggle.com/alinealmeida/basic-feature-benchmark-with-quantiles-augmenting\n\nApparently this approach does not change the results significantly (at least not with the basic features used: ave, std, max, min, q95, q99, q05 and q01).\n\nI created a custom split iterator in order to avoid segments in training which share regions with segments in validation set.",
      "votes": null
    },
    {
      "id": "464925",
      "postDate": "02/01/2019 20:26:21",
      "content": "<p>Thanks for your input and elegant code in your kernel.</p>",
      "rawMarkdown": "Thanks for your input and elegant code in your kernel.",
      "votes": null
    },
    {
      "id": "464990",
      "postDate": "02/02/2019 01:09:26",
      "content": "<p>I've experimented with stride size 150k,  75k, 50k and 30k using leave-one-earthquake-out cv.\nNo obvious differences were observed.</p>",
      "rawMarkdown": "I've experimented with stride size 150k,  75k, 50k and 30k using leave-one-earthquake-out cv.\nNo obvious differences were observed.",
      "votes": null
    },
    {
      "id": "465053",
      "postDate": "02/02/2019 07:51:10",
      "content": "<p>Ingenious, thx. I will add leave-one-earthquake-out cv scores as well. Did you evaluate it using the same overlap strategy (within fold) as for training or using contiguous chunks within fold?</p>",
      "rawMarkdown": "Ingenious, thx. I will add leave-one-earthquake-out cv scores as well. Did you evaluate it using the same overlap strategy (within fold) as for training or using contiguous chunks within fold?",
      "votes": null
    },
    {
      "id": "465355",
      "postDate": "02/02/2019 22:29:54",
      "content": "<p>Hi, in my setup, train and valid sets all share the same stride formats ~ And thank you for the summary work!! thank you!</p>",
      "rawMarkdown": "Hi, in my setup, train and valid sets all share the same stride formats ~ And thank you for the summary work!! thank you!",
      "votes": null
    },
    {
      "id": "465405",
      "postDate": "02/03/2019 02:51:32",
      "content": "<p>I tried to do random sampling with 150k rows in train - taking 150k from random starting points. It caused an overfitting.</p>",
      "rawMarkdown": "I tried to do random sampling with 150k rows in train - taking 150k from random starting points. It caused an overfitting.",
      "votes": null
    },
    {
      "id": "468514",
      "postDate": "02/09/2019 03:05:37",
      "content": "<p>Great work!</p>\n\n<p>Just want to add a bit. The 'trending' effect could be confounded with the model tunning itself.\nUsually for a larger dataset, one would need more training iterations, ceteris paribus, to achieve the similar level of convergence. \nAt least this is the 'feeling' in general.\nBeing translated into RF, this could mean 'number of trees'. \nThat is to have a larger number of trees in order to achieve the similar level of convergence when working with larger dataset. \nI could be wrong... Not really familiar with 'ranger' package. They seem to have some fancy internal operations, so...</p>",
      "rawMarkdown": "Great work!\n\nJust want to add a bit. The 'trending' effect could be confounded with the model tunning itself.\nUsually for a larger dataset, one would need more training iterations, ceteris paribus, to achieve the similar level of convergence. \nAt least this is the 'feeling' in general.\nBeing translated into RF, this could mean 'number of trees'. \nThat is to have a larger number of trees in order to achieve the similar level of convergence when working with larger dataset. \nI could be wrong... Not really familiar with 'ranger' package. They seem to have some fancy internal operations, so...",
      "votes": null
    },
    {
      "id": "468575",
      "postDate": "02/09/2019 07:17:54",
      "content": "<p>Interesting aspect. In my view, it might especially be the dependence across rows that would disturb convergence. After all, the algos that are mostly being used for this competition are not at all designed for dependent rows... random forests not being an exception here. </p>\n\n<p>I might add 1-2 results with larger number of trees for a high overlap situation to see its effect.</p>",
      "rawMarkdown": "Interesting aspect. In my view, it might especially be the dependence across rows that would disturb convergence. After all, the algos that are mostly being used for this competition are not at all designed for dependent rows... random forests not being an exception here. \n\nI might add 1-2 results with larger number of trees for a high overlap situation to see its effect.",
      "votes": null
    },
    {
      "id": "468845",
      "postDate": "02/09/2019 20:29:14",
      "content": "<p>great!</p>",
      "rawMarkdown": "great!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 464923,
      "author_name": "alinealmeida",
      "author_url": "",
      "post_date": "02/01/2019 20:20:28",
      "content": "<p>I have tested with stride of 75'000:\n<a href=\"https://www.kaggle.com/alinealmeida/basic-feature-benchmark-with-quantiles-augmenting\">https://www.kaggle.com/alinealmeida/basic-feature-benchmark-with-quantiles-augmenting</a></p>\n\n<p>Apparently this approach does not change the results significantly (at least not with the basic features used: ave, std, max, min, q95, q99, q05 and q01).</p>\n\n<p>I created a custom split iterator in order to avoid segments in training which share regions with segments in validation set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 464925,
          "author_name": "mayer79",
          "author_url": "",
          "post_date": "02/01/2019 20:26:21",
          "content": "<p>Thanks for your input and elegant code in your kernel.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 464990,
      "author_name": "tclf90",
      "author_url": "",
      "post_date": "02/02/2019 01:09:26",
      "content": "<p>I've experimented with stride size 150k,  75k, 50k and 30k using leave-one-earthquake-out cv.\nNo obvious differences were observed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 465053,
          "author_name": "mayer79",
          "author_url": "",
          "post_date": "02/02/2019 07:51:10",
          "content": "<p>Ingenious, thx. I will add leave-one-earthquake-out cv scores as well. Did you evaluate it using the same overlap strategy (within fold) as for training or using contiguous chunks within fold?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 465355,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "02/02/2019 22:29:54",
          "content": "<p>Hi, in my setup, train and valid sets all share the same stride formats ~ And thank you for the summary work!! thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 465405,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "02/03/2019 02:51:32",
      "content": "<p>I tried to do random sampling with 150k rows in train - taking 150k from random starting points. It caused an overfitting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 468514,
      "author_name": "tclf90",
      "author_url": "",
      "post_date": "02/09/2019 03:05:37",
      "content": "<p>Great work!</p>\n\n<p>Just want to add a bit. The 'trending' effect could be confounded with the model tunning itself.\nUsually for a larger dataset, one would need more training iterations, ceteris paribus, to achieve the similar level of convergence. \nAt least this is the 'feeling' in general.\nBeing translated into RF, this could mean 'number of trees'. \nThat is to have a larger number of trees in order to achieve the similar level of convergence when working with larger dataset. \nI could be wrong... Not really familiar with 'ranger' package. They seem to have some fancy internal operations, so...</p>",
      "votes": null,
      "replies": [
        {
          "id": 468575,
          "author_name": "mayer79",
          "author_url": "",
          "post_date": "02/09/2019 07:17:54",
          "content": "<p>Interesting aspect. In my view, it might especially be the dependence across rows that would disturb convergence. After all, the algos that are mostly being used for this competition are not at all designed for dependent rows... random forests not being an exception here. </p>\n\n<p>I might add 1-2 results with larger number of trees for a high overlap situation to see its effect.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 468845,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "02/09/2019 20:29:14",
          "content": "<p>great!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "464859": "Most public kernels work like this:\n\n1. Split time series into chunks of length 150'000 to match the size of the test pieces\n\n2. From each chunk, derive one line of training data\n\nOne of the many choices to make is: Should the chunks in Step 1 overlap? If yes, by how much? \n\nThe larger the overlap (resp. the smaller the stride), the more lines in the training set. But that does not necessarily mean a better performance because of increased statistical dependence of the lines... Where is the optimum given a set of features (and the chosen learning algorithms)?\n\nI will test at lest the following strides\n\n- 150'000 (no overlap), \n\n- 5'000 (97% overlap), \n\n- 10'000 (93% overlap) and \n\n- 15'000 (90% overlap between successive chunks) and post the results here over the weekend. \n\nTo get the LB scores, I will plug 35 standard features into an untuned random forest using https://www.kaggle.com/mayer79/earthquake-forecast-with-random-forests .  CV scores are mean absolute errors of leave-one-out earthquake validation as suggested by Elliot. \n\nAny experience on this question you want to share? What is your impression?\n\n|Stride|Overlap|CV|LB|\n|-------|----------|---|----|\n|5'000|97%|2.333|1.589|\n|10'000|93%|2.297|1.560|\n|15'000|90%|2.294|1.548|\n|30'000|80%|2.294|1.543|\n|50'000|66%|2.284|1.537|\n|75'000|50%|2.283|**1.522**|\n|100'000|33%|2.282|1.531|\n|150'000|0%|**2.267**|1.548|\n\nUsing the fixed feature set of above mentioned kernel, \"leave-one-earthquake-out\" cross-validation MAE worsens with higher overlap. LB scores might benefit from small overlaps between 33% and 67%. But honestly, I trust neither of the two scores in this competition...\n\nAll random forests were run with 500 trees. Would more trees help with dependent rows? For the setting with 30'000 stride length, using 2000 trees resulted in the following, only slightly, better scores:\n\n- CV: 2.292 \n\n- LB 1.541",
    "464923": "I have tested with stride of 75'000:\nhttps://www.kaggle.com/alinealmeida/basic-feature-benchmark-with-quantiles-augmenting\n\nApparently this approach does not change the results significantly (at least not with the basic features used: ave, std, max, min, q95, q99, q05 and q01).\n\nI created a custom split iterator in order to avoid segments in training which share regions with segments in validation set.",
    "464925": "Thanks for your input and elegant code in your kernel.",
    "464990": "I've experimented with stride size 150k,  75k, 50k and 30k using leave-one-earthquake-out cv.\nNo obvious differences were observed.",
    "465053": "Ingenious, thx. I will add leave-one-earthquake-out cv scores as well. Did you evaluate it using the same overlap strategy (within fold) as for training or using contiguous chunks within fold?",
    "465355": "Hi, in my setup, train and valid sets all share the same stride formats ~ And thank you for the summary work!! thank you!",
    "465405": "I tried to do random sampling with 150k rows in train - taking 150k from random starting points. It caused an overfitting.",
    "468514": "Great work!\n\nJust want to add a bit. The 'trending' effect could be confounded with the model tunning itself.\nUsually for a larger dataset, one would need more training iterations, ceteris paribus, to achieve the similar level of convergence. \nAt least this is the 'feeling' in general.\nBeing translated into RF, this could mean 'number of trees'. \nThat is to have a larger number of trees in order to achieve the similar level of convergence when working with larger dataset. \nI could be wrong... Not really familiar with 'ranger' package. They seem to have some fancy internal operations, so...",
    "468575": "Interesting aspect. In my view, it might especially be the dependence across rows that would disturb convergence. After all, the algos that are mostly being used for this competition are not at all designed for dependent rows... random forests not being an exception here. \n\nI might add 1-2 results with larger number of trees for a high overlap situation to see its effect.",
    "468845": "great!"
  },
  "source": "meta"
}