{
  "id": 89824,
  "title": "Data property?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89824",
  "author_name": "",
  "post_date": "2019-04-17T23:48:00.511215500Z",
  "votes": 10,
  "comment_count": 8,
  "views": 0,
  "content": "<p>looking my best 1.988 MAE on a simple stack of models i get the following graph!\n<img src=\"https://ibb.co/5sC2MxN\" alt=\"MAE\">\ni see that the most cases that increase my MAE are those with close to 0 actual and predicted &gt;0. Looking deeper on these errors i see that this is due to the ttf reset within segments because a new earthquake begins. So, i shifted down my oof predictions by 5 values and i got a cv of ~1.92 if i play more with these lagged predictions i can make it to 1.91, which is a huge improvment. It seems that these high MAE errors are mostly due to the fact that the ttf is more predictable from previous segments. Here is a snapshot of my oofs </p>\n\n<p>segment ttf prediction\n30  0.26149884  0.312911034\n31  0.223196042 0.378040314\n32  0.183797749 4.653840065\n33  0.144499456 7.257835388\n34  0.106196657 8.431538582\n35  0.066798364 9.088373184\n36  0.028495566 9.173859596\n37  11.53009728 8.951961517\n38  11.49079898 9.734399796\n39  11.45249618 9.284646988\n40  11.41309789 9.626651764\nThe CV for segments 32 to 40 is 4.21 but if we shift the predictions  2 values down we get a CV of 3.27.</p>\n\n<p>the model begins to predict the upcoming higher ttf of segment 37 from the segment 32. This is more like a timeseries problem, but when it comes to the test set we only have to make predictions for these 150K rows batches which are not continous. Hence, something weird is going on.... this is like having data for a stock per day ( and intraday) and you have to predict the stock's return after 2 months or after 1 day with only the intraday data. I think that there is a data validation discrepancy there. What is your opinion?</p>",
  "messages": [
    {
      "id": "518851",
      "postDate": "04/17/2019 23:48:00",
      "content": "<p>looking my best 1.988 MAE on a simple stack of models i get the following graph!\n<img src=\"https://ibb.co/5sC2MxN\" alt=\"MAE\">\ni see that the most cases that increase my MAE are those with close to 0 actual and predicted &gt;0. Looking deeper on these errors i see that this is due to the ttf reset within segments because a new earthquake begins. So, i shifted down my oof predictions by 5 values and i got a cv of ~1.92 if i play more with these lagged predictions i can make it to 1.91, which is a huge improvment. It seems that these high MAE errors are mostly due to the fact that the ttf is more predictable from previous segments. Here is a snapshot of my oofs </p>\n\n<p>segment ttf prediction\n30  0.26149884  0.312911034\n31  0.223196042 0.378040314\n32  0.183797749 4.653840065\n33  0.144499456 7.257835388\n34  0.106196657 8.431538582\n35  0.066798364 9.088373184\n36  0.028495566 9.173859596\n37  11.53009728 8.951961517\n38  11.49079898 9.734399796\n39  11.45249618 9.284646988\n40  11.41309789 9.626651764\nThe CV for segments 32 to 40 is 4.21 but if we shift the predictions  2 values down we get a CV of 3.27.</p>\n\n<p>the model begins to predict the upcoming higher ttf of segment 37 from the segment 32. This is more like a timeseries problem, but when it comes to the test set we only have to make predictions for these 150K rows batches which are not continous. Hence, something weird is going on.... this is like having data for a stock per day ( and intraday) and you have to predict the stock's return after 2 months or after 1 day with only the intraday data. I think that there is a data validation discrepancy there. What is your opinion?</p>",
      "rawMarkdown": "looking my best 1.988 MAE on a simple stack of models i get the following graph!\n![MAE](https://ibb.co/5sC2MxN)\ni see that the most cases that increase my MAE are those with close to 0 actual and predicted &gt;0. Looking deeper on these errors i see that this is due to the ttf reset within segments because a new earthquake begins. So, i shifted down my oof predictions by 5 values and i got a cv of ~1.92 if i play more with these lagged predictions i can make it to 1.91, which is a huge improvment. It seems that these high MAE errors are mostly due to the fact that the ttf is more predictable from previous segments. Here is a snapshot of my oofs \n\nsegment\tttf\tprediction\n30\t0.26149884\t0.312911034\n31\t0.223196042\t0.378040314\n32\t0.183797749\t4.653840065\n33\t0.144499456\t7.257835388\n34\t0.106196657\t8.431538582\n35\t0.066798364\t9.088373184\n36\t0.028495566\t9.173859596\n37\t11.53009728\t8.951961517\n38\t11.49079898\t9.734399796\n39\t11.45249618\t9.284646988\n40\t11.41309789\t9.626651764\nThe CV for segments 32 to 40 is 4.21 but if we shift the predictions  2 values down we get a CV of 3.27.\n\n\nthe model begins to predict the upcoming higher ttf of segment 37 from the segment 32. This is more like a timeseries problem, but when it comes to the test set we only have to make predictions for these 150K rows batches which are not continous. Hence, something weird is going on.... this is like having data for a stock per day ( and intraday) and you have to predict the stock's return after 2 months or after 1 day with only the intraday data. I think that there is a data validation discrepancy there. What is your opinion?",
      "votes": null
    },
    {
      "id": "519240",
      "postDate": "04/18/2019 15:34:55",
      "content": "<p>Good point, I observe the same behavior. These 5 segments are indistinguishable from the following segments with small ttf. And given that the later are many more than the former, MAE loss goes with the majority. I guess this is something we need to accept as an unavoidable error in the predictions. </p>",
      "rawMarkdown": "Good point, I observe the same behavior. These 5 segments are indistinguishable from the following segments with small ttf. And given that the later are many more than the former, MAE loss goes with the majority. I guess this is something we need to accept as an unavoidable error in the predictions.",
      "votes": null
    },
    {
      "id": "519717",
      "postDate": "04/19/2019 14:25:13",
      "content": "<p><a href=\"/dkaraflos\">@dkaraflos</a> \n\"This is more like a timeseries problem, but when it comes to the test set we only have to make predictions for these 150K rows batches which are not continous.\" This.</p>\n\n<p>Thats why we cant model it with ARIMA etc, since AR and MA properties really fall off. \nMy bet is that the more we shuffle the more realistic CV will be because of this.</p>",
      "rawMarkdown": "dkaraflos \n\"This is more like a timeseries problem, but when it comes to the test set we only have to make predictions for these 150K rows batches which are not continous.\" This.\n\nThats why we cant model it with ARIMA etc, since AR and MA properties really fall off. \nMy bet is that the more we shuffle the more realistic CV will be because of this.",
      "votes": null
    },
    {
      "id": "519888",
      "postDate": "04/19/2019 19:38:25",
      "content": "<p>Indeed. I believe that shuffling is the key for the validation, although we deal with timeseries data</p>",
      "rawMarkdown": "Indeed. I believe that shuffling is the key for the validation, although we deal with timeseries data",
      "votes": null
    },
    {
      "id": "521520",
      "postDate": "04/23/2019 01:53:40",
      "content": "<p><a href=\"/dkaraflos\">@dkaraflos</a>, I can't see your picture anymore. </p>",
      "rawMarkdown": "dkaraflos, I can't see your picture anymore.",
      "votes": null
    },
    {
      "id": "529497",
      "postDate": "05/10/2019 03:34:34",
      "content": "<p><a href=\"/dkaraflos\">@dkaraflos</a>  Are you talking about creating new features by lagging the ttf values in the training set? Let's assume the test set is time sorted (I know it is not in this case) if that is true how will you create the ttf lagged values for test set to make predictions?\nThank you in advance for the help!</p>",
      "rawMarkdown": "dkaraflos  Are you talking about creating new features by lagging the ttf values in the training set? Let's assume the test set is time sorted (I know it is not in this case) if that is true how will you create the ttf lagged values for test set to make predictions?\nThank you in advance for the help!",
      "votes": null
    },
    {
      "id": "529500",
      "postDate": "05/10/2019 03:53:08",
      "content": "<p>Hi,what is oofs?\nAnd what is ttf?</p>",
      "rawMarkdown": "Hi,what is oofs?\nAnd what is ttf?",
      "votes": null
    },
    {
      "id": "529537",
      "postDate": "05/10/2019 06:38:39",
      "content": "<p>oof - out of fold prediction, i.e. prediction made of data, which was not used for training, general term.\nttf - time to failure, specific term in this competition, name of the target.</p>",
      "rawMarkdown": "oof - out of fold prediction, i.e. prediction made of data, which was not used for training, general term.\nttf - time to failure, specific term in this competition, name of the target.",
      "votes": null
    },
    {
      "id": "529546",
      "postDate": "05/10/2019 06:55:26",
      "content": "<p>I do not refer to creating new features, i just state the issue that exists on every model's predictions. </p>",
      "rawMarkdown": "I do not refer to creating new features, i just state the issue that exists on every model's predictions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 519240,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "04/18/2019 15:34:55",
      "content": "<p>Good point, I observe the same behavior. These 5 segments are indistinguishable from the following segments with small ttf. And given that the later are many more than the former, MAE loss goes with the majority. I guess this is something we need to accept as an unavoidable error in the predictions. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 519717,
      "author_name": "zikazika",
      "author_url": "",
      "post_date": "04/19/2019 14:25:13",
      "content": "<p><a href=\"/dkaraflos\">@dkaraflos</a> \n\"This is more like a timeseries problem, but when it comes to the test set we only have to make predictions for these 150K rows batches which are not continous.\" This.</p>\n\n<p>Thats why we cant model it with ARIMA etc, since AR and MA properties really fall off. \nMy bet is that the more we shuffle the more realistic CV will be because of this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 519888,
          "author_name": "dkaraflos",
          "author_url": "",
          "post_date": "04/19/2019 19:38:25",
          "content": "<p>Indeed. I believe that shuffling is the key for the validation, although we deal with timeseries data</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 521520,
      "author_name": "pukkinming",
      "author_url": "",
      "post_date": "04/23/2019 01:53:40",
      "content": "<p><a href=\"/dkaraflos\">@dkaraflos</a>, I can't see your picture anymore. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 529497,
      "author_name": "anwarsadique",
      "author_url": "",
      "post_date": "05/10/2019 03:34:34",
      "content": "<p><a href=\"/dkaraflos\">@dkaraflos</a>  Are you talking about creating new features by lagging the ttf values in the training set? Let's assume the test set is time sorted (I know it is not in this case) if that is true how will you create the ttf lagged values for test set to make predictions?\nThank you in advance for the help!</p>",
      "votes": null,
      "replies": [
        {
          "id": 529546,
          "author_name": "dkaraflos",
          "author_url": "",
          "post_date": "05/10/2019 06:55:26",
          "content": "<p>I do not refer to creating new features, i just state the issue that exists on every model's predictions. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 529500,
      "author_name": "zjp2origin",
      "author_url": "",
      "post_date": "05/10/2019 03:53:08",
      "content": "<p>Hi,what is oofs?\nAnd what is ttf?</p>",
      "votes": null,
      "replies": [
        {
          "id": 529537,
          "author_name": "alexfir",
          "author_url": "",
          "post_date": "05/10/2019 06:38:39",
          "content": "<p>oof - out of fold prediction, i.e. prediction made of data, which was not used for training, general term.\nttf - time to failure, specific term in this competition, name of the target.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "518851": "looking my best 1.988 MAE on a simple stack of models i get the following graph!\n![MAE](https://ibb.co/5sC2MxN)\ni see that the most cases that increase my MAE are those with close to 0 actual and predicted &gt;0. Looking deeper on these errors i see that this is due to the ttf reset within segments because a new earthquake begins. So, i shifted down my oof predictions by 5 values and i got a cv of ~1.92 if i play more with these lagged predictions i can make it to 1.91, which is a huge improvment. It seems that these high MAE errors are mostly due to the fact that the ttf is more predictable from previous segments. Here is a snapshot of my oofs \n\nsegment\tttf\tprediction\n30\t0.26149884\t0.312911034\n31\t0.223196042\t0.378040314\n32\t0.183797749\t4.653840065\n33\t0.144499456\t7.257835388\n34\t0.106196657\t8.431538582\n35\t0.066798364\t9.088373184\n36\t0.028495566\t9.173859596\n37\t11.53009728\t8.951961517\n38\t11.49079898\t9.734399796\n39\t11.45249618\t9.284646988\n40\t11.41309789\t9.626651764\nThe CV for segments 32 to 40 is 4.21 but if we shift the predictions  2 values down we get a CV of 3.27.\n\n\nthe model begins to predict the upcoming higher ttf of segment 37 from the segment 32. This is more like a timeseries problem, but when it comes to the test set we only have to make predictions for these 150K rows batches which are not continous. Hence, something weird is going on.... this is like having data for a stock per day ( and intraday) and you have to predict the stock's return after 2 months or after 1 day with only the intraday data. I think that there is a data validation discrepancy there. What is your opinion?",
    "519240": "Good point, I observe the same behavior. These 5 segments are indistinguishable from the following segments with small ttf. And given that the later are many more than the former, MAE loss goes with the majority. I guess this is something we need to accept as an unavoidable error in the predictions.",
    "519717": "dkaraflos \n\"This is more like a timeseries problem, but when it comes to the test set we only have to make predictions for these 150K rows batches which are not continous.\" This.\n\nThats why we cant model it with ARIMA etc, since AR and MA properties really fall off. \nMy bet is that the more we shuffle the more realistic CV will be because of this.",
    "519888": "Indeed. I believe that shuffling is the key for the validation, although we deal with timeseries data",
    "521520": "dkaraflos, I can't see your picture anymore.",
    "529497": "dkaraflos  Are you talking about creating new features by lagging the ttf values in the training set? Let's assume the test set is time sorted (I know it is not in this case) if that is true how will you create the ttf lagged values for test set to make predictions?\nThank you in advance for the help!",
    "529500": "Hi,what is oofs?\nAnd what is ttf?",
    "529537": "oof - out of fold prediction, i.e. prediction made of data, which was not used for training, general term.\nttf - time to failure, specific term in this competition, name of the target.",
    "529546": "I do not refer to creating new features, i just state the issue that exists on every model's predictions."
  },
  "source": "meta"
}