{
  "id": 90111,
  "title": "data augmentation is helpful.",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90111",
  "author_name": "",
  "post_date": "2019-04-20T14:02:26.437824100Z",
  "votes": 8,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I use 150000 data points start from index 0 as a training sample (non overlap), but somehow I perform an augmentation with another 150000 data points extraction start from index 75000 and get 0.004 LB improvements. </p>\n\n<p>but doesn't seems to work when I triple it.</p>",
  "messages": [
    {
      "id": "520237",
      "postDate": "04/20/2019 14:02:26",
      "content": "<p>I use 150000 data points start from index 0 as a training sample (non overlap), but somehow I perform an augmentation with another 150000 data points extraction start from index 75000 and get 0.004 LB improvements. </p>\n\n<p>but doesn't seems to work when I triple it.</p>",
      "rawMarkdown": "I use 150000 data points start from index 0 as a training sample (non overlap), but somehow I perform an augmentation with another 150000 data points extraction start from index 75000 and get 0.004 LB improvements. \n\nbut doesn't seems to work when I triple it.",
      "votes": null
    },
    {
      "id": "520247",
      "postDate": "04/20/2019 14:40:07",
      "content": "<p>I tried a similar thing using bigger and smaller segment lengths (multiples and fractions of 150000) without any shifting (segments don't overlap). This did improve my local CV score since I had more shorter ttf targets to predict (and those are easier) but didn't improve the LB score. \n I might try this strategy more later with overlapping and focusing on segments with longer ttf. For now, I am focusing on more features engineering (using physical properties) and parallelization (since it takes time to process all the segments). </p>",
      "rawMarkdown": "I tried a similar thing using bigger and smaller segment lengths (multiples and fractions of 150000) without any shifting (segments don't overlap). This did improve my local CV score since I had more shorter ttf targets to predict (and those are easier) but didn't improve the LB score. \n I might try this strategy more later with overlapping and focusing on segments with longer ttf. For now, I am focusing on more features engineering (using physical properties) and parallelization (since it takes time to process all the segments).",
      "votes": null
    },
    {
      "id": "520248",
      "postDate": "04/20/2019 14:49:21",
      "content": "<p>I tried doing something similar, and although my CV improved, the LB did not. </p>\n\n<p>Can you disclose what model you are using?</p>",
      "rawMarkdown": "I tried doing something similar, and although my CV improved, the LB did not. \n\nCan you disclose what model you are using?",
      "votes": null
    },
    {
      "id": "520250",
      "postDate": "04/20/2019 14:59:09",
      "content": "<p>Yes, I too have some advanced features didnt work,  I'm using lgb.</p>",
      "rawMarkdown": "Yes, I too have some advanced features didnt work,  I'm using lgb.",
      "votes": null
    },
    {
      "id": "520275",
      "postDate": "04/20/2019 16:01:37",
      "content": "<blockquote>\n  <p>This did improve my local CV score since I had more shorter ttf targets to predict</p>\n</blockquote>\n\n<p>??</p>\n\n<p>You also had more longer ttf as well, didn't you?</p>",
      "rawMarkdown": "&gt; This did improve my local CV score since I had more shorter ttf targets to predict\n\n??\n\nYou also had more longer ttf as well, didn't you?",
      "votes": null
    },
    {
      "id": "520285",
      "postDate": "04/20/2019 16:19:09",
      "content": "<p>Indeed, I had more longer ttfs to predict as well but the ttf histogram looked similar to the one with only 150000 segments: more ttfs closer to 0 and the distribution decreases the longer it gets. My interpretation might be incorrect and I might be doing something wrong of course. :)</p>",
      "rawMarkdown": "Indeed, I had more longer ttfs to predict as well but the ttf histogram looked similar to the one with only 150000 segments: more ttfs closer to 0 and the distribution decreases the longer it gets. My interpretation might be incorrect and I might be doing something wrong of course. :)",
      "votes": null
    },
    {
      "id": "520321",
      "postDate": "04/20/2019 17:43:40",
      "content": "<p>The more segments overlap, the more your model overfits (with simple KFold validation).</p>",
      "rawMarkdown": "The more segments overlap, the more your model overfits (with simple KFold validation).",
      "votes": null
    },
    {
      "id": "520340",
      "postDate": "04/20/2019 18:39:20",
      "content": "<p>Why would it overfit more?</p>",
      "rawMarkdown": "Why would it overfit more?",
      "votes": null
    },
    {
      "id": "520352",
      "postDate": "04/20/2019 19:26:05",
      "content": "<p>Train data set with 12k observations (750 per earthquake) has 1.9 CV,  while a data set with 120k observations (7500 earthquake) has 0.7 CV. I think that in this case, the observations are strongly correlated, and this leads to a “time” leak.</p>",
      "rawMarkdown": "Train data set with 12k observations (750 per earthquake) has 1.9 CV,  while a data set with 120k observations (7500 earthquake) has 0.7 CV. I think that in this case, the observations are strongly correlated, and this leads to a “time” leak.",
      "votes": null
    },
    {
      "id": "520381",
      "postDate": "04/20/2019 21:37:29",
      "content": "<p>I don’t understand how you can get 0.7 cv.</p>",
      "rawMarkdown": "I don’t understand how you can get 0.7 cv.",
      "votes": null
    },
    {
      "id": "520476",
      "postDate": "04/21/2019 04:28:12",
      "content": "<p>When I did 10 times the observations after FE(~40,000) by using overlap I got my CV to .6-.7, but my LB was &gt;1.6. Massive overfitting</p>",
      "rawMarkdown": "When I did 10 times the observations after FE(~40,000) by using overlap I got my CV to .6-.7, but my LB was &gt;1.6. Massive overfitting",
      "votes": null
    },
    {
      "id": "520547",
      "postDate": "04/21/2019 08:40:10",
      "content": "<p>I still don't get how you can have cv so low. I've tried augmentation a bit, and my cv increased a bit.  You must be doing something wrong IMHO.</p>",
      "rawMarkdown": "I still don't get how you can have cv so low. I've tried augmentation a bit, and my cv increased a bit.  You must be doing something wrong IMHO.",
      "votes": null
    },
    {
      "id": "520622",
      "postDate": "04/21/2019 12:40:29",
      "content": "<p>It happens with simple KFold ( shuffling=True).</p>\n\n<p>I.E. Two consequence observations with 95% overlap. One is used in train, another is used in validation. They are very similar and model overfits. </p>\n\n<p>So, one should choose: another validation scheme or less augmentations. My example shows wrong way to handle the problem. </p>",
      "rawMarkdown": "It happens with simple KFold ( shuffling=True).\n\nI.E. Two consequence observations with 95% overlap. One is used in train, another is used in validation. They are very similar and model overfits. \n\nSo, one should choose: another validation scheme or less augmentations. My example shows wrong way to handle the problem.",
      "votes": null
    },
    {
      "id": "520625",
      "postDate": "04/21/2019 12:48:12",
      "content": "<p>I think you should \"augment\" inside the folds to avoid this. But like you said this kind of augmentation gives new values with quite similar distributions to the base data, so I don't think you can even call this an augmentation</p>",
      "rawMarkdown": "I think you should \"augment\" inside the folds to avoid this. But like you said this kind of augmentation gives new values with quite similar distributions to the base data, so I don't think you can even call this an augmentation",
      "votes": null
    },
    {
      "id": "525449",
      "postDate": "05/01/2019 03:28:47",
      "content": "<p>If I understand what Marcus did (I did the same thing and it's pretty easy to do) he created more than the 4194 segments.  I did 2X, 10X and 20X augments in what I think is the same method he used.</p>\n\n<p>I got big changes in my local CV.  Some have commented that this is \"over-fit\" when you get a very low local CV but than get a bad LB score with this type of augmentation.  I don't believe that over-fit is the right word to use, it should be something more like leak-fit or bad-fit.  My 10X has CV of 1.05 on model with only 4 engineered features.  Not going to bother wasting a LB submission because I know I have a leaking  model.</p>\n\n<p>When you do the augmentation before the folds than you can easily have a segment in the validation set that is almost the same as one in the train set.  CV is good because the train and validation are almost the same signal.  I would assume that augmentation in this fashion at 100X will give me a very good local CV and crap for a LB score.  I did not \"over-fit\" but rather I leaked between validation and train.  </p>\n\n<p>The correct approach is to do this type of augmentation after you have created the folds, doing it only on the four train folds and not the validation fold (for a 5 fold).  But that's a whole lot harder to do and my Python skill set is probably not up to that challenge.</p>",
      "rawMarkdown": "If I understand what Marcus did (I did the same thing and it's pretty easy to do) he created more than the 4194 segments.  I did 2X, 10X and 20X augments in what I think is the same method he used.\n\nI got big changes in my local CV.  Some have commented that this is \"over-fit\" when you get a very low local CV but than get a bad LB score with this type of augmentation.  I don't believe that over-fit is the right word to use, it should be something more like leak-fit or bad-fit.  My 10X has CV of 1.05 on model with only 4 engineered features.  Not going to bother wasting a LB submission because I know I have a leaking  model.\n\nWhen you do the augmentation before the folds than you can easily have a segment in the validation set that is almost the same as one in the train set.  CV is good because the train and validation are almost the same signal.  I would assume that augmentation in this fashion at 100X will give me a very good local CV and crap for a LB score.  I did not \"over-fit\" but rather I leaked between validation and train.  \n\nThe correct approach is to do this type of augmentation after you have created the folds, doing it only on the four train folds and not the validation fold (for a 5 fold).  But that's a whole lot harder to do and my Python skill set is probably not up to that challenge.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 520247,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "04/20/2019 14:40:07",
      "content": "<p>I tried a similar thing using bigger and smaller segment lengths (multiples and fractions of 150000) without any shifting (segments don't overlap). This did improve my local CV score since I had more shorter ttf targets to predict (and those are easier) but didn't improve the LB score. \n I might try this strategy more later with overlapping and focusing on segments with longer ttf. For now, I am focusing on more features engineering (using physical properties) and parallelization (since it takes time to process all the segments). </p>",
      "votes": null,
      "replies": [
        {
          "id": 520275,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/20/2019 16:01:37",
          "content": "<blockquote>\n  <p>This did improve my local CV score since I had more shorter ttf targets to predict</p>\n</blockquote>\n\n<p>??</p>\n\n<p>You also had more longer ttf as well, didn't you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520285,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "04/20/2019 16:19:09",
          "content": "<p>Indeed, I had more longer ttfs to predict as well but the ttf histogram looked similar to the one with only 150000 segments: more ttfs closer to 0 and the distribution decreases the longer it gets. My interpretation might be incorrect and I might be doing something wrong of course. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520321,
          "author_name": "simakov",
          "author_url": "",
          "post_date": "04/20/2019 17:43:40",
          "content": "<p>The more segments overlap, the more your model overfits (with simple KFold validation).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520340,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/20/2019 18:39:20",
          "content": "<p>Why would it overfit more?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520352,
          "author_name": "simakov",
          "author_url": "",
          "post_date": "04/20/2019 19:26:05",
          "content": "<p>Train data set with 12k observations (750 per earthquake) has 1.9 CV,  while a data set with 120k observations (7500 earthquake) has 0.7 CV. I think that in this case, the observations are strongly correlated, and this leads to a “time” leak.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520381,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/20/2019 21:37:29",
          "content": "<p>I don’t understand how you can get 0.7 cv.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520476,
          "author_name": "halldalton94",
          "author_url": "",
          "post_date": "04/21/2019 04:28:12",
          "content": "<p>When I did 10 times the observations after FE(~40,000) by using overlap I got my CV to .6-.7, but my LB was &gt;1.6. Massive overfitting</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520547,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/21/2019 08:40:10",
          "content": "<p>I still don't get how you can have cv so low. I've tried augmentation a bit, and my cv increased a bit.  You must be doing something wrong IMHO.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520622,
          "author_name": "simakov",
          "author_url": "",
          "post_date": "04/21/2019 12:40:29",
          "content": "<p>It happens with simple KFold ( shuffling=True).</p>\n\n<p>I.E. Two consequence observations with 95% overlap. One is used in train, another is used in validation. They are very similar and model overfits. </p>\n\n<p>So, one should choose: another validation scheme or less augmentations. My example shows wrong way to handle the problem. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520625,
          "author_name": "daijin12",
          "author_url": "",
          "post_date": "04/21/2019 12:48:12",
          "content": "<p>I think you should \"augment\" inside the folds to avoid this. But like you said this kind of augmentation gives new values with quite similar distributions to the base data, so I don't think you can even call this an augmentation</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 520248,
      "author_name": "robertoanzaldua",
      "author_url": "",
      "post_date": "04/20/2019 14:49:21",
      "content": "<p>I tried doing something similar, and although my CV improved, the LB did not. </p>\n\n<p>Can you disclose what model you are using?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 520250,
      "author_name": "marcuslin",
      "author_url": "",
      "post_date": "04/20/2019 14:59:09",
      "content": "<p>Yes, I too have some advanced features didnt work,  I'm using lgb.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 525449,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "05/01/2019 03:28:47",
      "content": "<p>If I understand what Marcus did (I did the same thing and it's pretty easy to do) he created more than the 4194 segments.  I did 2X, 10X and 20X augments in what I think is the same method he used.</p>\n\n<p>I got big changes in my local CV.  Some have commented that this is \"over-fit\" when you get a very low local CV but than get a bad LB score with this type of augmentation.  I don't believe that over-fit is the right word to use, it should be something more like leak-fit or bad-fit.  My 10X has CV of 1.05 on model with only 4 engineered features.  Not going to bother wasting a LB submission because I know I have a leaking  model.</p>\n\n<p>When you do the augmentation before the folds than you can easily have a segment in the validation set that is almost the same as one in the train set.  CV is good because the train and validation are almost the same signal.  I would assume that augmentation in this fashion at 100X will give me a very good local CV and crap for a LB score.  I did not \"over-fit\" but rather I leaked between validation and train.  </p>\n\n<p>The correct approach is to do this type of augmentation after you have created the folds, doing it only on the four train folds and not the validation fold (for a 5 fold).  But that's a whole lot harder to do and my Python skill set is probably not up to that challenge.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "520237": "I use 150000 data points start from index 0 as a training sample (non overlap), but somehow I perform an augmentation with another 150000 data points extraction start from index 75000 and get 0.004 LB improvements. \n\nbut doesn't seems to work when I triple it.",
    "520247": "I tried a similar thing using bigger and smaller segment lengths (multiples and fractions of 150000) without any shifting (segments don't overlap). This did improve my local CV score since I had more shorter ttf targets to predict (and those are easier) but didn't improve the LB score. \n I might try this strategy more later with overlapping and focusing on segments with longer ttf. For now, I am focusing on more features engineering (using physical properties) and parallelization (since it takes time to process all the segments).",
    "520248": "I tried doing something similar, and although my CV improved, the LB did not. \n\nCan you disclose what model you are using?",
    "520250": "Yes, I too have some advanced features didnt work,  I'm using lgb.",
    "520275": "&gt; This did improve my local CV score since I had more shorter ttf targets to predict\n\n??\n\nYou also had more longer ttf as well, didn't you?",
    "520285": "Indeed, I had more longer ttfs to predict as well but the ttf histogram looked similar to the one with only 150000 segments: more ttfs closer to 0 and the distribution decreases the longer it gets. My interpretation might be incorrect and I might be doing something wrong of course. :)",
    "520321": "The more segments overlap, the more your model overfits (with simple KFold validation).",
    "520340": "Why would it overfit more?",
    "520352": "Train data set with 12k observations (750 per earthquake) has 1.9 CV,  while a data set with 120k observations (7500 earthquake) has 0.7 CV. I think that in this case, the observations are strongly correlated, and this leads to a “time” leak.",
    "520381": "I don’t understand how you can get 0.7 cv.",
    "520476": "When I did 10 times the observations after FE(~40,000) by using overlap I got my CV to .6-.7, but my LB was &gt;1.6. Massive overfitting",
    "520547": "I still don't get how you can have cv so low. I've tried augmentation a bit, and my cv increased a bit.  You must be doing something wrong IMHO.",
    "520622": "It happens with simple KFold ( shuffling=True).\n\nI.E. Two consequence observations with 95% overlap. One is used in train, another is used in validation. They are very similar and model overfits. \n\nSo, one should choose: another validation scheme or less augmentations. My example shows wrong way to handle the problem.",
    "520625": "I think you should \"augment\" inside the folds to avoid this. But like you said this kind of augmentation gives new values with quite similar distributions to the base data, so I don't think you can even call this an augmentation",
    "525449": "If I understand what Marcus did (I did the same thing and it's pretty easy to do) he created more than the 4194 segments.  I did 2X, 10X and 20X augments in what I think is the same method he used.\n\nI got big changes in my local CV.  Some have commented that this is \"over-fit\" when you get a very low local CV but than get a bad LB score with this type of augmentation.  I don't believe that over-fit is the right word to use, it should be something more like leak-fit or bad-fit.  My 10X has CV of 1.05 on model with only 4 engineered features.  Not going to bother wasting a LB submission because I know I have a leaking  model.\n\nWhen you do the augmentation before the folds than you can easily have a segment in the validation set that is almost the same as one in the train set.  CV is good because the train and validation are almost the same signal.  I would assume that augmentation in this fashion at 100X will give me a very good local CV and crap for a LB score.  I did not \"over-fit\" but rather I leaked between validation and train.  \n\nThe correct approach is to do this type of augmentation after you have created the folds, doing it only on the four train folds and not the validation fold (for a 5 fold).  But that's a whole lot harder to do and my Python skill set is probably not up to that challenge."
  },
  "source": "meta"
}