{
  "id": 82256,
  "title": "Something I found in local CV loss and public loss about oversample",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/82256",
  "author_name": "",
  "post_date": "2019-02-28T08:50:33.197694500Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>Something I found in local CV loss and public loss</h1>\n\n<p>When I started this competition， I first refer to the solution of Gabriel Preda's kernels <a href=\"https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction\">LANL Earthquake EDA and Prediction</a>.</p>\n\n<p>This is a very good kernels and help me get started quickly.</p>\n\n<p>In the process, I came up with an idea : Why don't we generate more data?</p>\n\n<p>So , I used the oversampling method to generate more than 20,000 train data and used these data to train the xgboost and lgb model. Then I got better result than  when I used 4,000 pieces of data.</p>\n\n<p>But unfortunately, I get worse result in Public Leaderboard.</p>\n\n<p>It really confused me, So I did some visualization of validation data's results.</p>\n\n<p>And I found, more data means that the model has stronger fitting ability and can better fit those more extreme data.</p>\n\n<p>But I got the worse public loss.</p>\n\n<p>I thought about several possible reasons.</p>\n\n<ol>\n<li>Because the public leaderboard only use the 13% test data , my local\nresults are good, and I should trust my local CV loss </li>\n<li>More extreme  thinking, the data distribution of the test set may vary greatly\nfrom the training set.</li>\n</ol>\n\n<p>But if these  reasons are true , it means that I can't use the public loss as my model's feedback. In fact, I've had more than one such problem.\n<strong>But today I find this problem different from the past.</strong></p>\n\n<p>When I observed the training data, I found that in fact the target(\"time_to_failure\") between data fragments are very close.\nAnd when I oversample the data , a lot of overlapping data is generated.</p>\n\n<p>Does this mean that when I randomly divide the data into validation sets and training sets, many validation sets are actually very similar to training sets?\nSo actually my model was <strong>trained the model on the train and validation data</strong>, and that's why I got worse results on public leaderboard</p>\n\n<p>The above is just <strong>my personal guess</strong>. And I really hope people who meet the same problems can communicate with me.</p>",
  "messages": [
    {
      "id": "480500",
      "postDate": "02/28/2019 08:50:33",
      "content": "<h1>Something I found in local CV loss and public loss</h1>\n\n<p>When I started this competition， I first refer to the solution of Gabriel Preda's kernels <a href=\"https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction\">LANL Earthquake EDA and Prediction</a>.</p>\n\n<p>This is a very good kernels and help me get started quickly.</p>\n\n<p>In the process, I came up with an idea : Why don't we generate more data?</p>\n\n<p>So , I used the oversampling method to generate more than 20,000 train data and used these data to train the xgboost and lgb model. Then I got better result than  when I used 4,000 pieces of data.</p>\n\n<p>But unfortunately, I get worse result in Public Leaderboard.</p>\n\n<p>It really confused me, So I did some visualization of validation data's results.</p>\n\n<p>And I found, more data means that the model has stronger fitting ability and can better fit those more extreme data.</p>\n\n<p>But I got the worse public loss.</p>\n\n<p>I thought about several possible reasons.</p>\n\n<ol>\n<li>Because the public leaderboard only use the 13% test data , my local\nresults are good, and I should trust my local CV loss </li>\n<li>More extreme  thinking, the data distribution of the test set may vary greatly\nfrom the training set.</li>\n</ol>\n\n<p>But if these  reasons are true , it means that I can't use the public loss as my model's feedback. In fact, I've had more than one such problem.\n<strong>But today I find this problem different from the past.</strong></p>\n\n<p>When I observed the training data, I found that in fact the target(\"time_to_failure\") between data fragments are very close.\nAnd when I oversample the data , a lot of overlapping data is generated.</p>\n\n<p>Does this mean that when I randomly divide the data into validation sets and training sets, many validation sets are actually very similar to training sets?\nSo actually my model was <strong>trained the model on the train and validation data</strong>, and that's why I got worse results on public leaderboard</p>\n\n<p>The above is just <strong>my personal guess</strong>. And I really hope people who meet the same problems can communicate with me.</p>",
      "rawMarkdown": "# Something I found in local CV loss and public loss\nWhen I started this competition， I first refer to the solution of Gabriel Preda's kernels [LANL Earthquake EDA and Prediction](https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction).\n\nThis is a very good kernels and help me get started quickly.\n\nIn the process, I came up with an idea : Why don't we generate more data?\n\nSo , I used the oversampling method to generate more than 20,000 train data and used these data to train the xgboost and lgb model. Then I got better result than  when I used 4,000 pieces of data.\n\nBut unfortunately, I get worse result in Public Leaderboard.\n\nIt really confused me, So I did some visualization of validation data's results.\n\nAnd I found, more data means that the model has stronger fitting ability and can better fit those more extreme data.\n\nBut I got the worse public loss.\n\nI thought about several possible reasons.\n\n 1. Because the public leaderboard only use the 13% test data , my local\n    results are good, and I should trust my local CV loss \n 2. More extreme  thinking, the data distribution of the test set may vary greatly\n    from the training set.\n\nBut if these  reasons are true , it means that I can't use the public loss as my model's feedback. In fact, I've had more than one such problem.\n**But today I find this problem different from the past.**\n\nWhen I observed the training data, I found that in fact the target(\"time_to_failure\") between data fragments are very close.\nAnd when I oversample the data , a lot of overlapping data is generated.\n\nDoes this mean that when I randomly divide the data into validation sets and training sets, many validation sets are actually very similar to training sets?\nSo actually my model was **trained the model on the train and validation data**, and that's why I got worse results on public leaderboard\n\nThe above is just **my personal guess**. And I really hope people who meet the same problems can communicate with me.",
      "votes": null
    },
    {
      "id": "480747",
      "postDate": "02/28/2019 16:05:55",
      "content": "<p>Hi,\nIf you do data oversampling, as you say, your validation data must not overlap the train data, otherwise you will obviously overfit your CV.\nA good way not to overlap is to separate earthquakes: train on 13 earthquakes, test on 3 earthquakes. \nAnyway the public LB gives very low information on the quality of the model: I only trust my CV. That's quite frustrating not to have a valid public LB on this competition!</p>",
      "rawMarkdown": "Hi,\nIf you do data oversampling, as you say, your validation data must not overlap the train data, otherwise you will obviously overfit your CV.\nA good way not to overlap is to separate earthquakes: train on 13 earthquakes, test on 3 earthquakes. \nAnyway the public LB gives very low information on the quality of the model: I only trust my CV. That's quite frustrating not to have a valid public LB on this competition!",
      "votes": null
    },
    {
      "id": "480884",
      "postDate": "02/28/2019 20:18:15",
      "content": "<p>I am afraid you can trust your CV in this competition less than the private LB. The data has some temporal relationship at least with the features that I am using. I wonder whether I should trust the public LB more so than my CV because the public LB data at least is from the same or similar time-frame as the private LB. If I correctly understood, the test data is from a later time-frame than the training data. Would be helpful to know what exactly that time-frame is (for example train from t=0 to t=10, test from t=10 to t=20....). I think this will determine your best validation scheme.\nBut then, I might be completely wrong and helplessly over-fitting to the private LB.</p>",
      "rawMarkdown": "I am afraid you can trust your CV in this competition less than the private LB. The data has some temporal relationship at least with the features that I am using. I wonder whether I should trust the public LB more so than my CV because the public LB data at least is from the same or similar time-frame as the private LB. If I correctly understood, the test data is from a later time-frame than the training data. Would be helpful to know what exactly that time-frame is (for example train from t=0 to t=10, test from t=10 to t=20....). I think this will determine your best validation scheme.\nBut then, I might be completely wrong and helplessly over-fitting to the private LB.",
      "votes": null
    },
    {
      "id": "480901",
      "postDate": "02/28/2019 20:40:04",
      "content": "<p>I agree that the difference of features between train and test data is quite confusing.\nAccording to organisers, train and test data are contiguous. I understand it this way: \n- There is 250 seconds of data\n- Train data is 0-154s\n- Public LB is 154-166s\n- Private LB is 166-252s\nIf the difference in features appears at 154s, then the public LB gives information on the private LB. </p>",
      "rawMarkdown": "I agree that the difference of features between train and test data is quite confusing.\nAccording to organisers, train and test data are contiguous. I understand it this way: \n- There is 250 seconds of data\n- Train data is 0-154s\n- Public LB is 154-166s\n- Private LB is 166-252s\nIf the difference in features appears at 154s, then the public LB gives information on the private LB.",
      "votes": null
    },
    {
      "id": "480926",
      "postDate": "02/28/2019 21:55:01",
      "content": "<p>Looks like train and test data are abruptly different using my features. No smooth transition. So train and test data in and of itself are contiguous. But I am wondering whether train is contiguous with test. Doesn't look like it using my features. But I might be wrong.</p>",
      "rawMarkdown": "Looks like train and test data are abruptly different using my features. No smooth transition. So train and test data in and of itself are contiguous. But I am wondering whether train is contiguous with test. Doesn't look like it using my features. But I might be wrong.",
      "votes": null
    },
    {
      "id": "481060",
      "postDate": "03/01/2019 02:29:26",
      "content": "<p>Thank you very much for commenting on my post.I think separate earthquakes will be useful and I'm doing it right now. But as you discussed ，the rules for selecting test data affect our work. Actually，I'm a new kaggler , and I am wondering it's worth to spend a lot time to guess the distribution of test sets and tuning it？</p>",
      "rawMarkdown": "Thank you very much for commenting on my post.I think separate earthquakes will be useful and I'm doing it right now. But as you discussed ，the rules for selecting test data affect our work. Actually，I'm a new kaggler , and I am wondering it's worth to spend a lot time to guess the distribution of test sets and tuning it？",
      "votes": null
    },
    {
      "id": "522431",
      "postDate": "04/24/2019 12:27:28",
      "content": "<p><a href=\"/hartmut\">@hartmut</a> </p>\n\n<blockquote>\n  <p>If I correctly understood, the test data is from a later time-frame than the training data. </p>\n</blockquote>\n\n<p>How did you come to that conclusion?</p>",
      "rawMarkdown": "hartmut \n\n&gt;  If I correctly understood, the test data is from a later time-frame than the training data. \n\nHow did you come to that conclusion?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 480747,
      "author_name": "zidmie",
      "author_url": "",
      "post_date": "02/28/2019 16:05:55",
      "content": "<p>Hi,\nIf you do data oversampling, as you say, your validation data must not overlap the train data, otherwise you will obviously overfit your CV.\nA good way not to overlap is to separate earthquakes: train on 13 earthquakes, test on 3 earthquakes. \nAnyway the public LB gives very low information on the quality of the model: I only trust my CV. That's quite frustrating not to have a valid public LB on this competition!</p>",
      "votes": null,
      "replies": [
        {
          "id": 480884,
          "author_name": "hartmut",
          "author_url": "",
          "post_date": "02/28/2019 20:18:15",
          "content": "<p>I am afraid you can trust your CV in this competition less than the private LB. The data has some temporal relationship at least with the features that I am using. I wonder whether I should trust the public LB more so than my CV because the public LB data at least is from the same or similar time-frame as the private LB. If I correctly understood, the test data is from a later time-frame than the training data. Would be helpful to know what exactly that time-frame is (for example train from t=0 to t=10, test from t=10 to t=20....). I think this will determine your best validation scheme.\nBut then, I might be completely wrong and helplessly over-fitting to the private LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 480901,
          "author_name": "zidmie",
          "author_url": "",
          "post_date": "02/28/2019 20:40:04",
          "content": "<p>I agree that the difference of features between train and test data is quite confusing.\nAccording to organisers, train and test data are contiguous. I understand it this way: \n- There is 250 seconds of data\n- Train data is 0-154s\n- Public LB is 154-166s\n- Private LB is 166-252s\nIf the difference in features appears at 154s, then the public LB gives information on the private LB. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 480926,
          "author_name": "hartmut",
          "author_url": "",
          "post_date": "02/28/2019 21:55:01",
          "content": "<p>Looks like train and test data are abruptly different using my features. No smooth transition. So train and test data in and of itself are contiguous. But I am wondering whether train is contiguous with test. Doesn't look like it using my features. But I might be wrong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 481060,
          "author_name": "tudoni",
          "author_url": "",
          "post_date": "03/01/2019 02:29:26",
          "content": "<p>Thank you very much for commenting on my post.I think separate earthquakes will be useful and I'm doing it right now. But as you discussed ，the rules for selecting test data affect our work. Actually，I'm a new kaggler , and I am wondering it's worth to spend a lot time to guess the distribution of test sets and tuning it？</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 522431,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/24/2019 12:27:28",
          "content": "<p><a href=\"/hartmut\">@hartmut</a> </p>\n\n<blockquote>\n  <p>If I correctly understood, the test data is from a later time-frame than the training data. </p>\n</blockquote>\n\n<p>How did you come to that conclusion?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "480500": "# Something I found in local CV loss and public loss\nWhen I started this competition， I first refer to the solution of Gabriel Preda's kernels [LANL Earthquake EDA and Prediction](https://www.kaggle.com/gpreda/lanl-earthquake-eda-and-prediction).\n\nThis is a very good kernels and help me get started quickly.\n\nIn the process, I came up with an idea : Why don't we generate more data?\n\nSo , I used the oversampling method to generate more than 20,000 train data and used these data to train the xgboost and lgb model. Then I got better result than  when I used 4,000 pieces of data.\n\nBut unfortunately, I get worse result in Public Leaderboard.\n\nIt really confused me, So I did some visualization of validation data's results.\n\nAnd I found, more data means that the model has stronger fitting ability and can better fit those more extreme data.\n\nBut I got the worse public loss.\n\nI thought about several possible reasons.\n\n 1. Because the public leaderboard only use the 13% test data , my local\n    results are good, and I should trust my local CV loss \n 2. More extreme  thinking, the data distribution of the test set may vary greatly\n    from the training set.\n\nBut if these  reasons are true , it means that I can't use the public loss as my model's feedback. In fact, I've had more than one such problem.\n**But today I find this problem different from the past.**\n\nWhen I observed the training data, I found that in fact the target(\"time_to_failure\") between data fragments are very close.\nAnd when I oversample the data , a lot of overlapping data is generated.\n\nDoes this mean that when I randomly divide the data into validation sets and training sets, many validation sets are actually very similar to training sets?\nSo actually my model was **trained the model on the train and validation data**, and that's why I got worse results on public leaderboard\n\nThe above is just **my personal guess**. And I really hope people who meet the same problems can communicate with me.",
    "480747": "Hi,\nIf you do data oversampling, as you say, your validation data must not overlap the train data, otherwise you will obviously overfit your CV.\nA good way not to overlap is to separate earthquakes: train on 13 earthquakes, test on 3 earthquakes. \nAnyway the public LB gives very low information on the quality of the model: I only trust my CV. That's quite frustrating not to have a valid public LB on this competition!",
    "480884": "I am afraid you can trust your CV in this competition less than the private LB. The data has some temporal relationship at least with the features that I am using. I wonder whether I should trust the public LB more so than my CV because the public LB data at least is from the same or similar time-frame as the private LB. If I correctly understood, the test data is from a later time-frame than the training data. Would be helpful to know what exactly that time-frame is (for example train from t=0 to t=10, test from t=10 to t=20....). I think this will determine your best validation scheme.\nBut then, I might be completely wrong and helplessly over-fitting to the private LB.",
    "480901": "I agree that the difference of features between train and test data is quite confusing.\nAccording to organisers, train and test data are contiguous. I understand it this way: \n- There is 250 seconds of data\n- Train data is 0-154s\n- Public LB is 154-166s\n- Private LB is 166-252s\nIf the difference in features appears at 154s, then the public LB gives information on the private LB.",
    "480926": "Looks like train and test data are abruptly different using my features. No smooth transition. So train and test data in and of itself are contiguous. But I am wondering whether train is contiguous with test. Doesn't look like it using my features. But I might be wrong.",
    "481060": "Thank you very much for commenting on my post.I think separate earthquakes will be useful and I'm doing it right now. But as you discussed ，the rules for selecting test data affect our work. Actually，I'm a new kaggler , and I am wondering it's worth to spend a lot time to guess the distribution of test sets and tuning it？",
    "522431": "hartmut \n\n&gt;  If I correctly understood, the test data is from a later time-frame than the training data. \n\nHow did you come to that conclusion?"
  },
  "source": "meta"
}