{
  "id": 90592,
  "title": "Is most of time information of not useful ?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90592",
  "author_name": "Manoj",
  "post_date": "2019-04-25T03:55:02.918000",
  "votes": 3,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I am seeing in the kernals that people splitting the train acoustic series of 150k and using the last time data in the modeling.\nDoes that mean the rest of time information is useless ?. if you can shed some perspective on that. </p>\n\n<p>For eg:\n```</p>\n\n<h1>Split train data into Dataframe containing rows of 150000 data points.</h1>\n\n<p>rows = 150000\nsegments = int(np.floor(train.shape[0] / rows))\nX_train = pd.DataFrame(index=range(4194), dtype=np.float64, columns = range(rows))\ny_tr = pd.DataFrame(index=range(4194), dtype=np.float64, columns=range(1))</p>\n\n<p>for segment in tqdm_notebook(range(segments)):\n    seg = train.iloc[segment*rows:segment*rows+rows]\n    X_train.iloc[segment] = seg['acoustic_data'].values\n    y_tr.iloc[segment] = seg['time_to_failure'].values[-1]</p>\n\n<p>X_train.head()\n```\nin this kernal \n<a href=\"https://www.kaggle.com/rsaund/lanl-shap\">https://www.kaggle.com/rsaund/lanl-shap</a></p>",
  "messages": [
    {
      "id": 522872,
      "postDate": "2019-04-25T06:56:23.633Z",
      "content": "<p>The last time data is the target.</p>",
      "rawMarkdown": "The last time data is the target.",
      "votes": 1
    },
    {
      "id": 522884,
      "postDate": "2019-04-25T07:29:11.040Z",
      "content": "<p>Segments, in which the value of first time-to-failure is lower than the last one, should be removed from training, because such pattern doesn't exist in test. The organizers said, huge signals during the failure are removed, so the transition between the neighboring quakes in the train segment is not continuous, and test segments are continuous.\nEDIT: I cannot find this statement in the organizers comments or their papers. Maybe I read this in one of other papers, regarding other, similar experiment, so treat it as a false alarm and excuse me.\nEDIT II: See below.</p>",
      "rawMarkdown": "Segments, in which the value of first time-to-failure is lower than the last one, should be removed from training, because such pattern doesn't exist in test. The organizers said, huge signals during the failure are removed, so the transition between the neighboring quakes in the train segment is not continuous, and test segments are continuous.\nEDIT: I cannot find this statement in the organizers comments or their papers. Maybe I read this in one of other papers, regarding other, similar experiment, so treat it as a false alarm and excuse me.\nEDIT II: See below.",
      "votes": 2,
      "replies": [
        {
          "id": 523042,
          "postDate": "2019-04-25T12:49:13.250Z",
          "content": "<p>Hey <a href=\"/sionek\">@sionek</a> (Grzegorz), do you have any evidence for your claim? I just looked through the discussions and I did not see the Competition Host say anything like that.</p>",
          "rawMarkdown": "Hey @sionek (Grzegorz), do you have any evidence for your claim? I just looked through the discussions and I did not see the Competition Host say anything like that."
        },
        {
          "id": 523043,
          "postDate": "2019-04-25T12:52:16.700Z",
          "content": "<p>I am also interested in the answer as I remember reading something related but I cannot find where.</p>",
          "rawMarkdown": "I am also interested in the answer as I remember reading something related but I cannot find where."
        },
        {
          "id": 523044,
          "postDate": "2019-04-25T12:53:50.930Z",
          "content": "<p>Good point <a href=\"/sionek\">@sionek</a> . I was thinking about this just yesterday. The thing is that those segments include (at least, and likely only) two earthquakes. So in the first part you have a behaviour of a low ttf, in the second part you have the behaviour of a high ttf. Honestly, it doesn't make sense from a practical point of view to model this. Will try to follow up later with experiments</p>",
          "rawMarkdown": "Good point @sionek . I was thinking about this just yesterday. The thing is that those segments include (at least, and likely only) two earthquakes. So in the first part you have a behaviour of a low ttf, in the second part you have the behaviour of a high ttf. Honestly, it doesn't make sense from a practical point of view to model this. Will try to follow up later with experiments"
        },
        {
          "id": 523694,
          "postDate": "2019-04-26T18:57:47.507Z",
          "content": "<p>Hi <a href=\"/sionek\">@sionek</a> , how did you come to the conclusion that the first time-to-failure value in each segment of the test set is never lower than the last one since we do not have access to the time-to-failure in that set? </p>",
          "rawMarkdown": "Hi @sionek , how did you come to the conclusion that the first time-to-failure value in each segment of the test set is never lower than the last one since we do not have access to the time-to-failure in that set? ",
          "votes": 2
        },
        {
          "id": 524592,
          "postDate": "2019-04-29T06:26:01.537Z",
          "content": "<p>I think I read  something like this: <code>\"We   discard   clipped acoustic  events  with  amplitude equal  to 2^14 bits  during post-processing.\"</code>\nin <a href=\"http://www3.geosc.psu.edu/~cjm38/papers_talks/ShreedharanetalARMA2017.pdf\">http://www3.geosc.psu.edu/~cjm38/papers_talks/ShreedharanetalARMA2017.pdf</a></p>",
          "rawMarkdown": "I think I read  something like this: `\"We   discard   clipped acoustic  events  with  amplitude equal  to 2^14 bits  during post-processing.\"`\nin http://www3.geosc.psu.edu/~cjm38/papers_talks/ShreedharanetalARMA2017.pdf"
        }
      ]
    },
    {
      "id": 522804,
      "postDate": "2019-04-25T03:55:02.917Z",
      "content": "<p>I am seeing in the kernals that people splitting the train acoustic series of 150k and using the last time data in the modeling.\nDoes that mean the rest of time information is useless ?. if you can shed some perspective on that. </p>\n\n<p>For eg:\n```</p>\n\n<h1>Split train data into Dataframe containing rows of 150000 data points.</h1>\n\n<p>rows = 150000\nsegments = int(np.floor(train.shape[0] / rows))\nX_train = pd.DataFrame(index=range(4194), dtype=np.float64, columns = range(rows))\ny_tr = pd.DataFrame(index=range(4194), dtype=np.float64, columns=range(1))</p>\n\n<p>for segment in tqdm_notebook(range(segments)):\n    seg = train.iloc[segment*rows:segment*rows+rows]\n    X_train.iloc[segment] = seg['acoustic_data'].values\n    y_tr.iloc[segment] = seg['time_to_failure'].values[-1]</p>\n\n<p>X_train.head()\n```\nin this kernal \n<a href=\"https://www.kaggle.com/rsaund/lanl-shap\">https://www.kaggle.com/rsaund/lanl-shap</a></p>",
      "rawMarkdown": "I am seeing in the kernals that people splitting the train acoustic series of 150k and using the last time data in the modeling.\nDoes that mean the rest of time information is useless ?. if you can shed some perspective on that. \n\nFor eg:\n```\n# Split train data into Dataframe containing rows of 150000 data points.\n\nrows = 150000\nsegments = int(np.floor(train.shape[0] / rows))\nX_train = pd.DataFrame(index=range(4194), dtype=np.float64, columns = range(rows))\ny_tr = pd.DataFrame(index=range(4194), dtype=np.float64, columns=range(1))\n\nfor segment in tqdm_notebook(range(segments)):\n    seg = train.iloc[segment*rows:segment*rows+rows]\n    X_train.iloc[segment] = seg['acoustic_data'].values\n    y_tr.iloc[segment] = seg['time_to_failure'].values[-1]\n    \nX_train.head()\n```\nin this kernal \nhttps://www.kaggle.com/rsaund/lanl-shap",
      "votes": 2
    },
    {
      "id": 523708,
      "postDate": "2019-04-26T19:59:23.813Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 523753,
          "postDate": "2019-04-26T23:36:12.243Z",
          "content": "<p>You've got to be careful about not validating on training data when doing that.  My general strategy is to split the data into non-overlapping chunks that are bigger than a segment (300k - 1M), assign each chunk to either the training set or the validation set, and then within each chunk use the sliding window technique that you propose.  This allows me to get many more slightly different segments than you would get if you divided the 630M into roughly 4200 segments of exactly 150k, but also allows me to be sure that I'm not validating on data that I used for training. </p>",
          "rawMarkdown": "You've got to be careful about not validating on training data when doing that.  My general strategy is to split the data into non-overlapping chunks that are bigger than a segment (300k - 1M), assign each chunk to either the training set or the validation set, and then within each chunk use the sliding window technique that you propose.  This allows me to get many more slightly different segments than you would get if you divided the 630M into roughly 4200 segments of exactly 150k, but also allows me to be sure that I'm not validating on data that I used for training. "
        },
        {
          "id": 523755,
          "postDate": "2019-04-26T23:50:32.973Z",
          "content": "<p>I've tried oversampling by taking overlapping segments. Even taking data for a 150,000 row window in steps of 75,000 (only doubling the total data) led to really bad overfitting.</p>\n\n<p>Other competitors have stated that they got an improvement from oversampling up to this point, but if the steps were any smaller they got worse LB results. I however simply found that oversampling led to really bad results, albeit a good CV score.</p>",
          "rawMarkdown": "I've tried oversampling by taking overlapping segments. Even taking data for a 150,000 row window in steps of 75,000 (only doubling the total data) led to really bad overfitting.\n\nOther competitors have stated that they got an improvement from oversampling up to this point, but if the steps were any smaller they got worse LB results. I however simply found that oversampling led to really bad results, albeit a good CV score.",
          "votes": 1
        },
        {
          "id": 523907,
          "postDate": "2019-04-27T11:03:29.747Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 522872,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-04-25T06:56:23.633000",
      "content": "<p>The last time data is the target.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 522884,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2019-04-25T07:29:11.040000",
      "content": "<p>Segments, in which the value of first time-to-failure is lower than the last one, should be removed from training, because such pattern doesn't exist in test. The organizers said, huge signals during the failure are removed, so the transition between the neighboring quakes in the train segment is not continuous, and test segments are continuous.\nEDIT: I cannot find this statement in the organizers comments or their papers. Maybe I read this in one of other papers, regarding other, similar experiment, so treat it as a false alarm and excuse me.\nEDIT II: See below.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 523042,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2019-04-25T12:49:13.250000",
          "content": "<p>Hey <a href=\"/sionek\">@sionek</a> (Grzegorz), do you have any evidence for your claim? I just looked through the discussions and I did not see the Competition Host say anything like that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 523043,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-25T12:52:16.700000",
          "content": "<p>I am also interested in the answer as I remember reading something related but I cannot find where.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 523044,
          "author_name": "Fede",
          "author_url": "",
          "post_date": "2019-04-25T12:53:50.930000",
          "content": "<p>Good point <a href=\"/sionek\">@sionek</a> . I was thinking about this just yesterday. The thing is that those segments include (at least, and likely only) two earthquakes. So in the first part you have a behaviour of a low ttf, in the second part you have the behaviour of a high ttf. Honestly, it doesn't make sense from a practical point of view to model this. Will try to follow up later with experiments</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 523694,
          "author_name": "Philippe Lonjoux",
          "author_url": "",
          "post_date": "2019-04-26T18:57:47.507000",
          "content": "<p>Hi <a href=\"/sionek\">@sionek</a> , how did you come to the conclusion that the first time-to-failure value in each segment of the test set is never lower than the last one since we do not have access to the time-to-failure in that set? </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 524592,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2019-04-29T06:26:01.537000",
          "content": "<p>I think I read  something like this: <code>\"We   discard   clipped acoustic  events  with  amplitude equal  to 2^14 bits  during post-processing.\"</code>\nin <a href=\"http://www3.geosc.psu.edu/~cjm38/papers_talks/ShreedharanetalARMA2017.pdf\">http://www3.geosc.psu.edu/~cjm38/papers_talks/ShreedharanetalARMA2017.pdf</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 523708,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-26T19:59:23.813000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 523753,
          "author_name": "Aaron Koch",
          "author_url": "",
          "post_date": "2019-04-26T23:36:12.243000",
          "content": "<p>You've got to be careful about not validating on training data when doing that.  My general strategy is to split the data into non-overlapping chunks that are bigger than a segment (300k - 1M), assign each chunk to either the training set or the validation set, and then within each chunk use the sliding window technique that you propose.  This allows me to get many more slightly different segments than you would get if you divided the 630M into roughly 4200 segments of exactly 150k, but also allows me to be sure that I'm not validating on data that I used for training. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 523755,
          "author_name": "RNA",
          "author_url": "",
          "post_date": "2019-04-26T23:50:32.973000",
          "content": "<p>I've tried oversampling by taking overlapping segments. Even taking data for a 150,000 row window in steps of 75,000 (only doubling the total data) led to really bad overfitting.</p>\n\n<p>Other competitors have stated that they got an improvement from oversampling up to this point, but if the steps were any smaller they got worse LB results. I however simply found that oversampling led to really bad results, albeit a good CV score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 523907,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-27T11:03:29.747000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "522872": "The last time data is the target.",
    "522884": "Segments, in which the value of first time-to-failure is lower than the last one, should be removed from training, because such pattern doesn't exist in test. The organizers said, huge signals during the failure are removed, so the transition between the neighboring quakes in the train segment is not continuous, and test segments are continuous.\nEDIT: I cannot find this statement in the organizers comments or their papers. Maybe I read this in one of other papers, regarding other, similar experiment, so treat it as a false alarm and excuse me.\nEDIT II: See below.",
    "522804": "I am seeing in the kernals that people splitting the train acoustic series of 150k and using the last time data in the modeling.\nDoes that mean the rest of time information is useless ?. if you can shed some perspective on that. \n\nFor eg:\n```\n# Split train data into Dataframe containing rows of 150000 data points.\n\nrows = 150000\nsegments = int(np.floor(train.shape[0] / rows))\nX_train = pd.DataFrame(index=range(4194), dtype=np.float64, columns = range(rows))\ny_tr = pd.DataFrame(index=range(4194), dtype=np.float64, columns=range(1))\n\nfor segment in tqdm_notebook(range(segments)):\n    seg = train.iloc[segment*rows:segment*rows+rows]\n    X_train.iloc[segment] = seg['acoustic_data'].values\n    y_tr.iloc[segment] = seg['time_to_failure'].values[-1]\n    \nX_train.head()\n```\nin this kernal \nhttps://www.kaggle.com/rsaund/lanl-shap",
    "523708": ""
  }
}