{
  "id": 91490,
  "title": "Oversampling and Data Augmentation",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91490",
  "author_name": "RNA",
  "post_date": "2019-05-05T15:46:20.292000",
  "votes": 11,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I've taken this post from the  @vettejeep <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91298\">Masters project thread</a> since it deserves its own topic.</p>\n\n<p>I was surprised that his data augmentation approach, which seemed prone to leakage, had such promising results. When I try naive 2x oversampling by running feature generation on 150,000 rows with a 75,000 row step, my LB score always gets worse. When I attributed this to shuffled K-fold leakage, I tried a quake-wise split with no (or at least limited) leakage. Again, my score got substantially worse with the oversampled data compared to the regular data.</p>\n\n<p>Am I missing something fundamental to the data augmentation process? My approach was designed to limit the number of times any range of the data could be sampled more than once. But every time I use it, my CV score decreases and the LB score increases. I had assumed that random sampling would suffer since:</p>\n\n<ul>\n<li>certain parts of train.csv could avoid sampling entirely</li>\n<li>certain parts of train.csv could be sampled more than twice</li>\n</ul>\n\n<p>I would really appreciate some input from any resident experts, as I feel there is a serious gap in my understanding here. Instead of oversampling, I have been considering generating rows with a mixture of genuine feature values, and values randomly sampled from the feature distribution in order to make my models less sensitive to random variance.  But I don't know if this approach is fundamentally flawed.</p>\n\n<p>What have other people's experiences been with oversampling/augmentation?</p>",
  "messages": [
    {
      "id": 527466,
      "postDate": "2019-05-05T15:46:20.293Z",
      "content": "<p>I've taken this post from the  @vettejeep <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91298\">Masters project thread</a> since it deserves its own topic.</p>\n\n<p>I was surprised that his data augmentation approach, which seemed prone to leakage, had such promising results. When I try naive 2x oversampling by running feature generation on 150,000 rows with a 75,000 row step, my LB score always gets worse. When I attributed this to shuffled K-fold leakage, I tried a quake-wise split with no (or at least limited) leakage. Again, my score got substantially worse with the oversampled data compared to the regular data.</p>\n\n<p>Am I missing something fundamental to the data augmentation process? My approach was designed to limit the number of times any range of the data could be sampled more than once. But every time I use it, my CV score decreases and the LB score increases. I had assumed that random sampling would suffer since:</p>\n\n<ul>\n<li>certain parts of train.csv could avoid sampling entirely</li>\n<li>certain parts of train.csv could be sampled more than twice</li>\n</ul>\n\n<p>I would really appreciate some input from any resident experts, as I feel there is a serious gap in my understanding here. Instead of oversampling, I have been considering generating rows with a mixture of genuine feature values, and values randomly sampled from the feature distribution in order to make my models less sensitive to random variance.  But I don't know if this approach is fundamentally flawed.</p>\n\n<p>What have other people's experiences been with oversampling/augmentation?</p>",
      "rawMarkdown": "I've taken this post from the  @vettejeep [Masters project thread](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91298) since it deserves its own topic.\n \nI was surprised that his data augmentation approach, which seemed prone to leakage, had such promising results. When I try naive 2x oversampling by running feature generation on 150,000 rows with a 75,000 row step, my LB score always gets worse. When I attributed this to shuffled K-fold leakage, I tried a quake-wise split with no (or at least limited) leakage. Again, my score got substantially worse with the oversampled data compared to the regular data.\n\nAm I missing something fundamental to the data augmentation process? My approach was designed to limit the number of times any range of the data could be sampled more than once. But every time I use it, my CV score decreases and the LB score increases. I had assumed that random sampling would suffer since:\n\n* certain parts of train.csv could avoid sampling entirely\n* certain parts of train.csv could be sampled more than twice\n\nI would really appreciate some input from any resident experts, as I feel there is a serious gap in my understanding here. Instead of oversampling, I have been considering generating rows with a mixture of genuine feature values, and values randomly sampled from the feature distribution in order to make my models less sensitive to random variance.  But I don't know if this approach is fundamentally flawed.\n\nWhat have other people's experiences been with oversampling/augmentation?",
      "votes": 9
    },
    {
      "id": 527528,
      "postDate": "2019-05-05T18:23:55.610Z",
      "content": "<p>Augmentation works well for me. I have not tried over-sampling yet but I definitely will</p>",
      "rawMarkdown": "Augmentation works well for me. I have not tried over-sampling yet but I definitely will",
      "votes": 1,
      "replies": [
        {
          "id": 527561,
          "postDate": "2019-05-05T20:25:29.700Z",
          "content": "<p>What is the difference between augmentation and over-sampling ?</p>",
          "rawMarkdown": "What is the difference between augmentation and over-sampling ?",
          "votes": 3
        },
        {
          "id": 527580,
          "postDate": "2019-05-05T21:21:26.503Z",
          "content": "<p>Good question :)\nI interpreted oversampling to be augmentation applied only to or mostly to a specific part of the spectrum. \nNote that a particular example of oversampling it's simply over-weighting\nI may be wrong, I hope the topic author can intervene and clarify</p>",
          "rawMarkdown": "Good question :)\nI interpreted oversampling to be augmentation applied only to or mostly to a specific part of the spectrum. \nNote that a particular example of oversampling it's simply over-weighting\nI may be wrong, I hope the topic author can intervene and clarify",
          "votes": 1
        },
        {
          "id": 527587,
          "postDate": "2019-05-05T22:08:24.347Z",
          "content": "<p><a href=\"/returnofsputnik\">@returnofsputnik</a> My understanding is that using the same data more than once is oversampling. Data augmentation is more broad, and involves engineering new training data that will increase the number of input samples that models can work with. I was considering spending tomorrow night trying to engineer some new rows, which would mostly include copies of existing data, but random features would be replaced with samples picked from their existing distribution. Hopefully it will make models that are less prone to identifying splits based on random variance instead of exploitable trends. </p>",
          "rawMarkdown": "@returnofsputnik My understanding is that using the same data more than once is oversampling. Data augmentation is more broad, and involves engineering new training data that will increase the number of input samples that models can work with. I was considering spending tomorrow night trying to engineer some new rows, which would mostly include copies of existing data, but random features would be replaced with samples picked from their existing distribution. Hopefully it will make models that are less prone to identifying splits based on random variance instead of exploitable trends. ",
          "votes": 2
        },
        {
          "id": 528453,
          "postDate": "2019-05-07T21:26:05.793Z",
          "content": "<p>I've run some basic data augmentation here: <a href=\"https://www.kaggle.com/bigironsphere/basic-data-augmentation-feature-reduction\">https://www.kaggle.com/bigironsphere/basic-data-augmentation-feature-reduction</a></p>\n\n<p>I'll update when I've sent off the submissions.</p>",
          "rawMarkdown": "I've run some basic data augmentation here: https://www.kaggle.com/bigironsphere/basic-data-augmentation-feature-reduction\n\nI'll update when I've sent off the submissions."
        }
      ]
    },
    {
      "id": 527507,
      "postDate": "2019-05-05T17:20:36.520Z",
      "content": "<p>You need to not have overlaps between segments in train and val, otherwise you have leakage. Anyways, oversampling has not been helpful for me as well.</p>",
      "rawMarkdown": "You need to not have overlaps between segments in train and val, otherwise you have leakage. Anyways, oversampling has not been helpful for me as well.",
      "votes": 1,
      "replies": [
        {
          "id": 527509,
          "postDate": "2019-05-05T17:23:55.587Z",
          "content": "<p>That's why I tried oversampled data using a quake-wise split. It led to worse scores just like K-fold. Surely it can't just be a matter of luck that some people have benefitted from oversampling and others haven't? </p>",
          "rawMarkdown": "That's why I tried oversampled data using a quake-wise split. It led to worse scores just like K-fold. Surely it can't just be a matter of luck that some people have benefitted from oversampling and others haven't? "
        },
        {
          "id": 527511,
          "postDate": "2019-05-05T17:31:22.137Z",
          "content": "<p>How do you know that some have benefited from it?</p>",
          "rawMarkdown": "How do you know that some have benefited from it?"
        },
        {
          "id": 527513,
          "postDate": "2019-05-05T17:43:51.803Z",
          "content": "<p>Vettejeep oversampled his data by selection of 4000 random starting indices for each 1/6th of the total data, for a total of 24000 rows. A few others have also commented on obtaining better results from oversampling.</p>",
          "rawMarkdown": "Vettejeep oversampled his data by selection of 4000 random starting indices for each 1/6th of the total data, for a total of 24000 rows. A few others have also commented on obtaining better results from oversampling."
        },
        {
          "id": 527551,
          "postDate": "2019-05-05T19:28:02.873Z",
          "content": "<p>But do you know whether his strategy is better than just using ~4k samples?</p>",
          "rawMarkdown": "But do you know whether his strategy is better than just using ~4k samples?",
          "votes": 3
        },
        {
          "id": 527609,
          "postDate": "2019-05-05T23:30:54.973Z",
          "content": "<p>I don't know, so you're right to ask. But if oversampling in general leads to bad results, it is unlikely he would've attained such a high score when creating 24000 rows of data. That's nearly a 6x oversampling!</p>",
          "rawMarkdown": "I don't know, so you're right to ask. But if oversampling in general leads to bad results, it is unlikely he would've attained such a high score when creating 24000 rows of data. That's nearly a 6x oversampling!",
          "votes": 1
        },
        {
          "id": 527729,
          "postDate": "2019-05-06T07:07:55.020Z",
          "content": "<p>If you do random sampling / oversampling results might be exactly the same.</p>",
          "rawMarkdown": "If you do random sampling / oversampling results might be exactly the same."
        }
      ]
    },
    {
      "id": 527494,
      "postDate": "2019-05-05T16:54:17.430Z",
      "content": "<p><a href=\"/bigironsphere\">@bigironsphere</a> I don't have much to add over my response in the original thread, but lets bump this up and see what the community says.  There may be a sweet spot to over sampling.  When I went to 40k samples my LB score went down.  It took too many computer resources to pursue further.</p>",
      "rawMarkdown": "@bigironsphere I don't have much to add over my response in the original thread, but lets bump this up and see what the community says.  There may be a sweet spot to over sampling.  When I went to 40k samples my LB score went down.  It took too many computer resources to pursue further.",
      "votes": 1,
      "replies": [
        {
          "id": 527588,
          "postDate": "2019-05-05T22:13:51.217Z",
          "content": "<p>I wouldn't think too much about the LB score in this competition. Find a leakage free CV and compare the results there. </p>",
          "rawMarkdown": "I wouldn't think too much about the LB score in this competition. Find a leakage free CV and compare the results there. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 527644,
      "postDate": "2019-05-06T02:14:14.730Z",
      "content": "<p>I think oversampling just reduces variance mostly. In some cases, it increases bias - that would lead to decrease is model performance. So if you're over-fitting, then adding more data, or what you call oversampling may help you.</p>",
      "rawMarkdown": "I think oversampling just reduces variance mostly. In some cases, it increases bias - that would lead to decrease is model performance. So if you're over-fitting, then adding more data, or what you call oversampling may help you."
    },
    {
      "id": 527603,
      "postDate": "2019-05-05T23:06:17.457Z",
      "content": "<p>For this data I have no idea augmentation(with overlap) works or not, maybe by chance it helps with limited LB data, but my experiments with p4581 data shows augmentation without leaking to validation data does not help.(based on testing with most of the public kernel  and my own models. .</p>",
      "rawMarkdown": "For this data I have no idea augmentation(with overlap) works or not, maybe by chance it helps with limited LB data, but my experiments with p4581 data shows augmentation without leaking to validation data does not help.(based on testing with most of the public kernel  and my own models. ."
    },
    {
      "id": 527851,
      "postDate": "2019-05-06T13:17:27.437Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 527859,
          "postDate": "2019-05-06T13:28:32.507Z",
          "content": "<p>I did the same for my oversampling. I'm starting to think a better approach might be training separate models on the two datasets and combining predictions rather than merging the data entirely.</p>",
          "rawMarkdown": "I did the same for my oversampling. I'm starting to think a better approach might be training separate models on the two datasets and combining predictions rather than merging the data entirely."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 527528,
      "author_name": "bluetrain",
      "author_url": "",
      "post_date": "2019-05-05T18:23:55.610000",
      "content": "<p>Augmentation works well for me. I have not tried over-sampling yet but I definitely will</p>",
      "votes": 1,
      "replies": [
        {
          "id": 527561,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2019-05-05T20:25:29.700000",
          "content": "<p>What is the difference between augmentation and over-sampling ?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 527580,
          "author_name": "bluetrain",
          "author_url": "",
          "post_date": "2019-05-05T21:21:26.503000",
          "content": "<p>Good question :)\nI interpreted oversampling to be augmentation applied only to or mostly to a specific part of the spectrum. \nNote that a particular example of oversampling it's simply over-weighting\nI may be wrong, I hope the topic author can intervene and clarify</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 527587,
          "author_name": "RNA",
          "author_url": "",
          "post_date": "2019-05-05T22:08:24.347000",
          "content": "<p><a href=\"/returnofsputnik\">@returnofsputnik</a> My understanding is that using the same data more than once is oversampling. Data augmentation is more broad, and involves engineering new training data that will increase the number of input samples that models can work with. I was considering spending tomorrow night trying to engineer some new rows, which would mostly include copies of existing data, but random features would be replaced with samples picked from their existing distribution. Hopefully it will make models that are less prone to identifying splits based on random variance instead of exploitable trends. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 528453,
          "author_name": "RNA",
          "author_url": "",
          "post_date": "2019-05-07T21:26:05.793000",
          "content": "<p>I've run some basic data augmentation here: <a href=\"https://www.kaggle.com/bigironsphere/basic-data-augmentation-feature-reduction\">https://www.kaggle.com/bigironsphere/basic-data-augmentation-feature-reduction</a></p>\n\n<p>I'll update when I've sent off the submissions.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 527507,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2019-05-05T17:20:36.520000",
      "content": "<p>You need to not have overlaps between segments in train and val, otherwise you have leakage. Anyways, oversampling has not been helpful for me as well.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 527509,
          "author_name": "RNA",
          "author_url": "",
          "post_date": "2019-05-05T17:23:55.587000",
          "content": "<p>That's why I tried oversampled data using a quake-wise split. It led to worse scores just like K-fold. Surely it can't just be a matter of luck that some people have benefitted from oversampling and others haven't? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527511,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-05-05T17:31:22.137000",
          "content": "<p>How do you know that some have benefited from it?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527513,
          "author_name": "RNA",
          "author_url": "",
          "post_date": "2019-05-05T17:43:51.803000",
          "content": "<p>Vettejeep oversampled his data by selection of 4000 random starting indices for each 1/6th of the total data, for a total of 24000 rows. A few others have also commented on obtaining better results from oversampling.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527551,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-05-05T19:28:02.873000",
          "content": "<p>But do you know whether his strategy is better than just using ~4k samples?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 527609,
          "author_name": "RNA",
          "author_url": "",
          "post_date": "2019-05-05T23:30:54.973000",
          "content": "<p>I don't know, so you're right to ask. But if oversampling in general leads to bad results, it is unlikely he would've attained such a high score when creating 24000 rows of data. That's nearly a 6x oversampling!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 527729,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-05-06T07:07:55.020000",
          "content": "<p>If you do random sampling / oversampling results might be exactly the same.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 527494,
      "author_name": "Vettejeep",
      "author_url": "",
      "post_date": "2019-05-05T16:54:17.430000",
      "content": "<p><a href=\"/bigironsphere\">@bigironsphere</a> I don't have much to add over my response in the original thread, but lets bump this up and see what the community says.  There may be a sweet spot to over sampling.  When I went to 40k samples my LB score went down.  It took too many computer resources to pursue further.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 527588,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-05-05T22:13:51.217000",
          "content": "<p>I wouldn't think too much about the LB score in this competition. Find a leakage free CV and compare the results there. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 527644,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2019-05-06T02:14:14.730000",
      "content": "<p>I think oversampling just reduces variance mostly. In some cases, it increases bias - that would lead to decrease is model performance. So if you're over-fitting, then adding more data, or what you call oversampling may help you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 527603,
      "author_name": "Sia",
      "author_url": "",
      "post_date": "2019-05-05T23:06:17.457000",
      "content": "<p>For this data I have no idea augmentation(with overlap) works or not, maybe by chance it helps with limited LB data, but my experiments with p4581 data shows augmentation without leaking to validation data does not help.(based on testing with most of the public kernel  and my own models. .</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 527851,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-06T13:17:27.437000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 527859,
          "author_name": "RNA",
          "author_url": "",
          "post_date": "2019-05-06T13:28:32.507000",
          "content": "<p>I did the same for my oversampling. I'm starting to think a better approach might be training separate models on the two datasets and combining predictions rather than merging the data entirely.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "527466": "I've taken this post from the  @vettejeep [Masters project thread](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91298) since it deserves its own topic.\n \nI was surprised that his data augmentation approach, which seemed prone to leakage, had such promising results. When I try naive 2x oversampling by running feature generation on 150,000 rows with a 75,000 row step, my LB score always gets worse. When I attributed this to shuffled K-fold leakage, I tried a quake-wise split with no (or at least limited) leakage. Again, my score got substantially worse with the oversampled data compared to the regular data.\n\nAm I missing something fundamental to the data augmentation process? My approach was designed to limit the number of times any range of the data could be sampled more than once. But every time I use it, my CV score decreases and the LB score increases. I had assumed that random sampling would suffer since:\n\n* certain parts of train.csv could avoid sampling entirely\n* certain parts of train.csv could be sampled more than twice\n\nI would really appreciate some input from any resident experts, as I feel there is a serious gap in my understanding here. Instead of oversampling, I have been considering generating rows with a mixture of genuine feature values, and values randomly sampled from the feature distribution in order to make my models less sensitive to random variance.  But I don't know if this approach is fundamentally flawed.\n\nWhat have other people's experiences been with oversampling/augmentation?",
    "527528": "Augmentation works well for me. I have not tried over-sampling yet but I definitely will",
    "527507": "You need to not have overlaps between segments in train and val, otherwise you have leakage. Anyways, oversampling has not been helpful for me as well.",
    "527494": "@bigironsphere I don't have much to add over my response in the original thread, but lets bump this up and see what the community says.  There may be a sweet spot to over sampling.  When I went to 40k samples my LB score went down.  It took too many computer resources to pursue further.",
    "527644": "I think oversampling just reduces variance mostly. In some cases, it increases bias - that would lead to decrease is model performance. So if you're over-fitting, then adding more data, or what you call oversampling may help you.",
    "527603": "For this data I have no idea augmentation(with overlap) works or not, maybe by chance it helps with limited LB data, but my experiments with p4581 data shows augmentation without leaking to validation data does not help.(based on testing with most of the public kernel  and my own models. .",
    "527851": ""
  }
}