{
  "id": 86530,
  "title": "The distribution of the test data",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/86530",
  "author_name": "Eplistical",
  "post_date": "2019-03-24T23:00:23.465000",
  "votes": 53,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I submitted a simple model (~1.52 LB) with the following modifications:\n1) Force all predicted data points with time_to_failure &gt; 6.0 (1006 points in total) to be 10000.0\n2) Force all predicted data points with time_to_failure &lt; 6.0 (1618 points in total) to be 10000.0</p>\n\n<p>I got MAE ~2442 for (1) and ~7556 for (2).\nFrom these results we can approximate the time_to_failure distribution in the test data:</p>\n\n<p>For public data, there are about 83 points (24%) whose time_to_failure is large and 258 points (76%) whose time_to_faliure is small. \nFor private data, there are about 923 points (40%) whose time_to_failure is large and 1360 points (60%) whose time_to_faliure is small. </p>\n\n<p>Since most models have higher accuracy for those data points with small time_to_failure (as I observed), the above distribution partially explains why people get lower LB error than their CV: There are more small time_to_failure data points in the public test data. </p>\n\n<p>Hope this is helpful :)</p>",
  "messages": [
    {
      "id": 499538,
      "postDate": "2019-03-24T23:00:23.467Z",
      "content": "<p>I submitted a simple model (~1.52 LB) with the following modifications:\n1) Force all predicted data points with time_to_failure &gt; 6.0 (1006 points in total) to be 10000.0\n2) Force all predicted data points with time_to_failure &lt; 6.0 (1618 points in total) to be 10000.0</p>\n\n<p>I got MAE ~2442 for (1) and ~7556 for (2).\nFrom these results we can approximate the time_to_failure distribution in the test data:</p>\n\n<p>For public data, there are about 83 points (24%) whose time_to_failure is large and 258 points (76%) whose time_to_faliure is small. \nFor private data, there are about 923 points (40%) whose time_to_failure is large and 1360 points (60%) whose time_to_faliure is small. </p>\n\n<p>Since most models have higher accuracy for those data points with small time_to_failure (as I observed), the above distribution partially explains why people get lower LB error than their CV: There are more small time_to_failure data points in the public test data. </p>\n\n<p>Hope this is helpful :)</p>",
      "rawMarkdown": "I submitted a simple model (~1.52 LB) with the following modifications:\n1) Force all predicted data points with time\\_to\\_failure &gt; 6.0 (1006 points in total) to be 10000.0\n2) Force all predicted data points with time\\_to\\_failure &lt; 6.0 (1618 points in total) to be 10000.0\n\nI got MAE ~2442 for (1) and ~7556 for (2).\nFrom these results we can approximate the time\\_to\\_failure distribution in the test data:\n\nFor public data, there are about 83 points (24%) whose time\\_to\\_failure is large and 258 points (76%) whose time\\_to\\_faliure is small. \nFor private data, there are about 923 points (40%) whose time\\_to\\_failure is large and 1360 points (60%) whose time\\_to\\_faliure is small. \n\nSince most models have higher accuracy for those data points with small time\\_to\\_failure (as I observed), the above distribution partially explains why people get lower LB error than their CV: There are more small time\\_to\\_failure data points in the public test data. \n\nHope this is helpful :)",
      "votes": 52
    },
    {
      "id": 500056,
      "postDate": "2019-03-25T15:10:25.623Z",
      "content": "<p>This will certainly help to get a better LB score, but it won't create models that can generalise to new situations so there may be some unpleasant surprises with the private test data!</p>\n\n<p>LB being higher than CV makes sense - the closer to ttf==0, the easier it is to predict because there is more acoustic noise that implies an imminent earthquake. When ttf is large, it's harder to extract details that inform the model how far away from failure it is. Clearly the subsample of the test data used for the public LB is closer to quake events than in train.csv.</p>\n\n<p>To expand on this, I think a good way to improve LB score is to only train on data where ttf is below a certain threshold. But without information on the private test set, there's no way of knowing if this will help the final scores.</p>",
      "rawMarkdown": "This will certainly help to get a better LB score, but it won't create models that can generalise to new situations so there may be some unpleasant surprises with the private test data!\n\nLB being higher than CV makes sense - the closer to ttf==0, the easier it is to predict because there is more acoustic noise that implies an imminent earthquake. When ttf is large, it's harder to extract details that inform the model how far away from failure it is. Clearly the subsample of the test data used for the public LB is closer to quake events than in train.csv.\n\nTo expand on this, I think a good way to improve LB score is to only train on data where ttf is below a certain threshold. But without information on the private test set, there's no way of knowing if this will help the final scores.",
      "votes": 6,
      "replies": [
        {
          "id": 500328,
          "postDate": "2019-03-25T21:32:14.867Z",
          "content": "<p>Exactly. Waiting to see how the ranking will change in the end. \nAs far as I'm concerned, the most challenging task in this competition is to get data with moderate ttf correct. That being said, for a point that is extremely far away from the next failure (say ttf=15), probably no model can get a good prediction (the input signal and the output ttf are simply uncorrelated. e.g. no one can predict an earthquake that happens 2 years later with the signal collected today); for a point that is very close (say ttf=0.3), most model works pretty well. The crux is to extract features that work for a moderate ttf (e.g. ttf=8).</p>",
          "rawMarkdown": "Exactly. Waiting to see how the ranking will change in the end. \nAs far as I'm concerned, the most challenging task in this competition is to get data with moderate ttf correct. That being said, for a point that is extremely far away from the next failure (say ttf=15), probably no model can get a good prediction (the input signal and the output ttf are simply uncorrelated. e.g. no one can predict an earthquake that happens 2 years later with the signal collected today); for a point that is very close (say ttf=0.3), most model works pretty well. The crux is to extract features that work for a moderate ttf (e.g. ttf=8).",
          "votes": 5
        }
      ]
    },
    {
      "id": 518623,
      "postDate": "2019-04-17T14:31:24.303Z",
      "content": "<p>I don't understand how you can get private test distribution at all.  Are these numbers assuming your prediction is exact?  If yes then it is a very quesiotnable assumption given your MAE is 1.5 and not 0.  I may be missing the obvious here.</p>",
      "rawMarkdown": "I don't understand how you can get private test distribution at all.  Are these numbers assuming your prediction is exact?  If yes then it is a very quesiotnable assumption given your MAE is 1.5 and not 0.  I may be missing the obvious here.",
      "votes": 3
    },
    {
      "id": 520519,
      "postDate": "2019-04-21T07:24:21.420Z",
      "content": "<p>Thanks for your sharing, very interesting. But I am curious about how could you know the target distribution of the private data set? </p>",
      "rawMarkdown": "Thanks for your sharing, very interesting. But I am curious about how could you know the target distribution of the private data set? ",
      "votes": 1
    },
    {
      "id": 511532,
      "postDate": "2019-04-10T05:10:30.020Z",
      "content": "<p>interesting discussion topics,thanks for bringing it out</p>",
      "rawMarkdown": "interesting discussion topics,thanks for bringing it out",
      "votes": 1
    },
    {
      "id": 500493,
      "postDate": "2019-03-26T05:33:18.537Z",
      "content": "<p>Genius detective! Thanks for this useful information Eplistical!\nI also tried to submit something like: first (or last 87%) rows of submission have value 1000, to see if the public/private is simply split by order. And I got no result :-)</p>",
      "rawMarkdown": "Genius detective! Thanks for this useful information Eplistical!\nI also tried to submit something like: first (or last 87%) rows of submission have value 1000, to see if the public/private is simply split by order. And I got no result :-)",
      "votes": 1
    },
    {
      "id": 514728,
      "postDate": "2019-04-11T21:46:59.807Z",
      "content": "<p>This asumes that your predicted values are correct, while in reality a portion might be below the threshold of 6.0 instead of above or vice versa. There is also no guarantee that the private set has  a similar portion of samples incorrectly predicted as being below or above the threshold of 6.0, right? Or am I missing something here?</p>",
      "rawMarkdown": "This asumes that your predicted values are correct, while in reality a portion might be below the threshold of 6.0 instead of above or vice versa. There is also no guarantee that the private set has  a similar portion of samples incorrectly predicted as being below or above the threshold of 6.0, right? Or am I missing something here?",
      "votes": 2,
      "replies": [
        {
          "id": 514974,
          "postDate": "2019-04-12T04:41:55.977Z",
          "content": "<p>Yes. We assume that the model predict relatively well in terms of high or low ttf. This may not 100% trustworthy, but it may reflects the portion of ttf in public/private data.</p>",
          "rawMarkdown": "Yes. We assume that the model predict relatively well in terms of high or low ttf. This may not 100% trustworthy, but it may reflects the portion of ttf in public/private data.",
          "votes": 1
        },
        {
          "id": 515572,
          "postDate": "2019-04-12T19:38:22.710Z",
          "content": "<p>You are right. This is the assumption I made.\nIn addition to what Kha Vo said, the threshold (6.0) used here is somehow arbitrary, one should define a reasonable number according to the model.</p>",
          "rawMarkdown": "You are right. This is the assumption I made.\nIn addition to what Kha Vo said, the threshold (6.0) used here is somehow arbitrary, one should define a reasonable number according to the model."
        }
      ]
    },
    {
      "id": 518798,
      "postDate": "2019-04-17T20:23:52.487Z",
      "content": "<p>Interesting LB probing. I think your analysis would be fuller and more convincing (regarding the final conclusion) if you checked it on OOF using the same model, the same threshold and the proportions obtained for public LB and compared the result with the public LB result.</p>",
      "rawMarkdown": "Interesting LB probing. I think your analysis would be fuller and more convincing (regarding the final conclusion) if you checked it on OOF using the same model, the same threshold and the proportions obtained for public LB and compared the result with the public LB result."
    },
    {
      "id": 518809,
      "postDate": "2019-04-17T21:13:48.467Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 500056,
      "author_name": "RNA",
      "author_url": "",
      "post_date": "2019-03-25T15:10:25.623000",
      "content": "<p>This will certainly help to get a better LB score, but it won't create models that can generalise to new situations so there may be some unpleasant surprises with the private test data!</p>\n\n<p>LB being higher than CV makes sense - the closer to ttf==0, the easier it is to predict because there is more acoustic noise that implies an imminent earthquake. When ttf is large, it's harder to extract details that inform the model how far away from failure it is. Clearly the subsample of the test data used for the public LB is closer to quake events than in train.csv.</p>\n\n<p>To expand on this, I think a good way to improve LB score is to only train on data where ttf is below a certain threshold. But without information on the private test set, there's no way of knowing if this will help the final scores.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 500328,
          "author_name": "Eplistical",
          "author_url": "",
          "post_date": "2019-03-25T21:32:14.867000",
          "content": "<p>Exactly. Waiting to see how the ranking will change in the end. \nAs far as I'm concerned, the most challenging task in this competition is to get data with moderate ttf correct. That being said, for a point that is extremely far away from the next failure (say ttf=15), probably no model can get a good prediction (the input signal and the output ttf are simply uncorrelated. e.g. no one can predict an earthquake that happens 2 years later with the signal collected today); for a point that is very close (say ttf=0.3), most model works pretty well. The crux is to extract features that work for a moderate ttf (e.g. ttf=8).</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 518623,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-04-17T14:31:24.303000",
      "content": "<p>I don't understand how you can get private test distribution at all.  Are these numbers assuming your prediction is exact?  If yes then it is a very quesiotnable assumption given your MAE is 1.5 and not 0.  I may be missing the obvious here.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 520519,
      "author_name": "Morgon",
      "author_url": "",
      "post_date": "2019-04-21T07:24:21.420000",
      "content": "<p>Thanks for your sharing, very interesting. But I am curious about how could you know the target distribution of the private data set? </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 511532,
      "author_name": "ai1776",
      "author_url": "",
      "post_date": "2019-04-10T05:10:30.020000",
      "content": "<p>interesting discussion topics,thanks for bringing it out</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 500493,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2019-03-26T05:33:18.537000",
      "content": "<p>Genius detective! Thanks for this useful information Eplistical!\nI also tried to submit something like: first (or last 87%) rows of submission have value 1000, to see if the public/private is simply split by order. And I got no result :-)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 514728,
      "author_name": "Machinehead",
      "author_url": "",
      "post_date": "2019-04-11T21:46:59.807000",
      "content": "<p>This asumes that your predicted values are correct, while in reality a portion might be below the threshold of 6.0 instead of above or vice versa. There is also no guarantee that the private set has  a similar portion of samples incorrectly predicted as being below or above the threshold of 6.0, right? Or am I missing something here?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 514974,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-04-12T04:41:55.977000",
          "content": "<p>Yes. We assume that the model predict relatively well in terms of high or low ttf. This may not 100% trustworthy, but it may reflects the portion of ttf in public/private data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 515572,
          "author_name": "Eplistical",
          "author_url": "",
          "post_date": "2019-04-12T19:38:22.710000",
          "content": "<p>You are right. This is the assumption I made.\nIn addition to what Kha Vo said, the threshold (6.0) used here is somehow arbitrary, one should define a reasonable number according to the model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 518798,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2019-04-17T20:23:52.487000",
      "content": "<p>Interesting LB probing. I think your analysis would be fuller and more convincing (regarding the final conclusion) if you checked it on OOF using the same model, the same threshold and the proportions obtained for public LB and compared the result with the public LB result.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 518809,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-17T21:13:48.467000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "499538": "I submitted a simple model (~1.52 LB) with the following modifications:\n1) Force all predicted data points with time\\_to\\_failure &gt; 6.0 (1006 points in total) to be 10000.0\n2) Force all predicted data points with time\\_to\\_failure &lt; 6.0 (1618 points in total) to be 10000.0\n\nI got MAE ~2442 for (1) and ~7556 for (2).\nFrom these results we can approximate the time\\_to\\_failure distribution in the test data:\n\nFor public data, there are about 83 points (24%) whose time\\_to\\_failure is large and 258 points (76%) whose time\\_to\\_faliure is small. \nFor private data, there are about 923 points (40%) whose time\\_to\\_failure is large and 1360 points (60%) whose time\\_to\\_faliure is small. \n\nSince most models have higher accuracy for those data points with small time\\_to\\_failure (as I observed), the above distribution partially explains why people get lower LB error than their CV: There are more small time\\_to\\_failure data points in the public test data. \n\nHope this is helpful :)",
    "500056": "This will certainly help to get a better LB score, but it won't create models that can generalise to new situations so there may be some unpleasant surprises with the private test data!\n\nLB being higher than CV makes sense - the closer to ttf==0, the easier it is to predict because there is more acoustic noise that implies an imminent earthquake. When ttf is large, it's harder to extract details that inform the model how far away from failure it is. Clearly the subsample of the test data used for the public LB is closer to quake events than in train.csv.\n\nTo expand on this, I think a good way to improve LB score is to only train on data where ttf is below a certain threshold. But without information on the private test set, there's no way of knowing if this will help the final scores.",
    "518623": "I don't understand how you can get private test distribution at all.  Are these numbers assuming your prediction is exact?  If yes then it is a very quesiotnable assumption given your MAE is 1.5 and not 0.  I may be missing the obvious here.",
    "520519": "Thanks for your sharing, very interesting. But I am curious about how could you know the target distribution of the private data set? ",
    "511532": "interesting discussion topics,thanks for bringing it out",
    "500493": "Genius detective! Thanks for this useful information Eplistical!\nI also tried to submit something like: first (or last 87%) rows of submission have value 1000, to see if the public/private is simply split by order. And I got no result :-)",
    "514728": "This asumes that your predicted values are correct, while in reality a portion might be below the threshold of 6.0 instead of above or vice versa. There is also no guarantee that the private set has  a similar portion of samples incorrectly predicted as being below or above the threshold of 6.0, right? Or am I missing something here?",
    "518798": "Interesting LB probing. I think your analysis would be fuller and more convincing (regarding the final conclusion) if you checked it on OOF using the same model, the same threshold and the proportions obtained for public LB and compared the result with the public LB result.",
    "518809": ""
  }
}