{
  "id": 91137,
  "title": "Skewed distribution",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91137",
  "author_name": "",
  "post_date": "2019-05-01T09:40:52.369315900Z",
  "votes": 9,
  "comment_count": 12,
  "views": 0,
  "content": "<p>My best model prediction has a binomial flavor to it. I find it hard to predict those mid-sections in the 4 to 6 sec. range; turning knobs helps a bit but always at the expense of the LB score. How do you guys think I should handle this issue?</p>\n\n<p>(vertical is target, horizontal is oof prediction)</p>",
  "messages": [
    {
      "id": "525562",
      "postDate": "05/01/2019 09:40:52",
      "content": "<p>My best model prediction has a binomial flavor to it. I find it hard to predict those mid-sections in the 4 to 6 sec. range; turning knobs helps a bit but always at the expense of the LB score. How do you guys think I should handle this issue?</p>\n\n<p>(vertical is target, horizontal is oof prediction)</p>",
      "rawMarkdown": "My best model prediction has a binomial flavor to it. I find it hard to predict those mid-sections in the 4 to 6 sec. range; turning knobs helps a bit but always at the expense of the LB score. How do you guys think I should handle this issue?\n\n(vertical is target, horizontal is oof prediction)",
      "votes": null
    },
    {
      "id": "525809",
      "postDate": "05/01/2019 18:42:17",
      "content": "<p>My model has similar appearance.   Since the training data does not show the same shape, I conclude that the sampling of segments for the test was not random, but was intentional.  </p>\n\n<p>I would assume that the creators of the data did much the same work that we are doing.  My guess would therefore be that their models not much different than ours.  </p>\n\n<p>They probably would select samples in the areas where their model had its poorest performance.  On the plots of training vs actual my models don' look that good in the 8 and above.    In that area my predictions for training are generally lower than actual.    They also want samples closer to the actual event.</p>\n\n<p>If I wanted 10,000 folks to try to improve things I would intentionally select segments in the same two general piles that your showing.</p>\n\n<p>In reading some of the available science on this type of experiment it's clear to me that the 8 and above wave shape has completely different characteristics than the waves close to the quake.  So the features that predict time to failure for the 9 peak group need/are different from the 4-5 peak group.</p>\n\n<p>So this kind of reminds me of a common issue I faced over the years - in my case, it was two different machines producing the same product, but their output was binomial when piled into a single data set.  In my case I always needed to run DOE's and analysis on the machines as separate models.</p>\n\n<p>So my plan is to train two different models -model 1 will only look at data for 6 and above segments.  Model 2 will only look at data for less than 6.  Than blend or stack the results.   Might do some overlap, but that will be my general plan.</p>\n\n<p>Would rather find a single model that could separate the output from two similar but different flows - I could have used it for the past 50 years - of course, I need the way back machine to take a PC loaded with Python back to 1970, but maybe....</p>",
      "rawMarkdown": "My model has similar appearance.   Since the training data does not show the same shape, I conclude that the sampling of segments for the test was not random, but was intentional.  \n\nI would assume that the creators of the data did much the same work that we are doing.  My guess would therefore be that their models not much different than ours.  \n\nThey probably would select samples in the areas where their model had its poorest performance.  On the plots of training vs actual my models don' look that good in the 8 and above.    In that area my predictions for training are generally lower than actual.    They also want samples closer to the actual event.\n\nIf I wanted 10,000 folks to try to improve things I would intentionally select segments in the same two general piles that your showing.\n\nIn reading some of the available science on this type of experiment it's clear to me that the 8 and above wave shape has completely different characteristics than the waves close to the quake.  So the features that predict time to failure for the 9 peak group need/are different from the 4-5 peak group.\n\nSo this kind of reminds me of a common issue I faced over the years - in my case, it was two different machines producing the same product, but their output was binomial when piled into a single data set.  In my case I always needed to run DOE's and analysis on the machines as separate models.\n\nSo my plan is to train two different models -model 1 will only look at data for 6 and above segments.  Model 2 will only look at data for less than 6.  Than blend or stack the results.   Might do some overlap, but that will be my general plan.\n\nWould rather find a single model that could separate the output from two similar but different flows - I could have used it for the past 50 years - of course, I need the way back machine to take a PC loaded with Python back to 1970, but maybe....",
      "votes": null
    },
    {
      "id": "526000",
      "postDate": "05/02/2019 06:27:28",
      "content": "<p>??  What Phillipe sows is on train data.  Test data sampling isn’t relevant.</p>",
      "rawMarkdown": "??  What Phillipe sows is on train data.  Test data sampling isn’t relevant.",
      "votes": null
    },
    {
      "id": "526159",
      "postDate": "05/02/2019 13:30:26",
      "content": "<blockquote>\n  <p>Since the training data does not show the same shape</p>\n</blockquote>\n\n<p>Aha! Found the guy who has the  ground truth!</p>",
      "rawMarkdown": "&gt; Since the training data does not show the same shape\n\nAha! Found the guy who has the  ground truth!",
      "votes": null
    },
    {
      "id": "526182",
      "postDate": "05/02/2019 14:14:21",
      "content": "<p>Why not just to train a classification model that will output a probability of being above 6 and then use it as additional feature in regression model?</p>",
      "rawMarkdown": "Why not just to train a classification model that will output a probability of being above 6 and then use it as additional feature in regression model?",
      "votes": null
    },
    {
      "id": "526189",
      "postDate": "05/02/2019 14:28:58",
      "content": "<p>I'll try that for sure and report back soon!</p>",
      "rawMarkdown": "I'll try that for sure and report back soon!",
      "votes": null
    },
    {
      "id": "526203",
      "postDate": "05/02/2019 15:05:20",
      "content": "<p><a href=\"/redstr\">@redstr</a> I've actually found the same thing with the training data so don't be alarmed! In fact, I assumed OP was talking about the test data at first since my models typically predict the 4-6 range very well.</p>",
      "rawMarkdown": "redstr I've actually found the same thing with the training data so don't be alarmed! In fact, I assumed OP was talking about the test data at first since my models typically predict the 4-6 range very well.",
      "votes": null
    },
    {
      "id": "526212",
      "postDate": "05/02/2019 15:25:23",
      "content": "<p>It was just a joke. I suppose you are talking about a bimodal distribution of predictions.</p>",
      "rawMarkdown": "It was just a joke. I suppose you are talking about a bimodal distribution of predictions.",
      "votes": null
    },
    {
      "id": "526370",
      "postDate": "05/02/2019 22:20:22",
      "content": "<p>I was looking at test data from my best model prediction.  Pretty sure that's what OP was talking about.</p>\n\n<p>Guess it could always be possible that my model predicts bimodal rather than the test data having that distribution ? </p>",
      "rawMarkdown": "I was looking at test data from my best model prediction.  Pretty sure that's what OP was talking about.\n\nGuess it could always be possible that my model predicts bimodal rather than the test data having that distribution ?",
      "votes": null
    },
    {
      "id": "526388",
      "postDate": "05/02/2019 23:58:01",
      "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> Alternatively, we've generally found that the 4-6 range is one of the easiest parts of the data to predict. Maybe the organisers deliberately oversampled test segments with particularly high/low TTF in order to evaluate how well we can predict the more complex cases. </p>",
      "rawMarkdown": "pcjimmmy Alternatively, we've generally found that the 4-6 range is one of the easiest parts of the data to predict. Maybe the organisers deliberately oversampled test segments with particularly high/low TTF in order to evaluate how well we can predict the more complex cases.",
      "votes": null
    },
    {
      "id": "526475",
      "postDate": "05/03/2019 06:17:59",
      "content": "<p><a href=\"/bigironsphere\">@bigironsphere</a> how do you know you predict well the 4-6 range on test data?</p>",
      "rawMarkdown": "bigironsphere how do you know you predict well the 4-6 range on test data?",
      "votes": null
    },
    {
      "id": "526802",
      "postDate": "05/03/2019 20:05:38",
      "content": "<p>You are quite right of course - I have no idea. I can only only speculate given:</p>\n\n<ul>\n<li>my model's performance on that range with the training data</li>\n<li>the knowledge that the average TTF of the LB data is around 4</li>\n<li>the fact that our models generally score better on the LB than OOF</li>\n</ul>\n\n<p>The real test data, who knows? They may have reserved especially complicated cases for it. </p>\n\n<p>For the record, I don't obsess over the LB score or waste time trying to second-guess what the test set might look like. My priority is, and has always been, to develop a robust CV method and work to improve my score on that. My LB score is admittedly low, and I doubt I will do well in this competition. But I think I will do better than my score suggests, and the learning is more important anyway.  </p>",
      "rawMarkdown": "You are quite right of course - I have no idea. I can only only speculate given:\n\n* my model's performance on that range with the training data\n* the knowledge that the average TTF of the LB data is around 4\n* the fact that our models generally score better on the LB than OOF\n\nThe real test data, who knows? They may have reserved especially complicated cases for it. \n\nFor the record, I don't obsess over the LB score or waste time trying to second-guess what the test set might look like. My priority is, and has always been, to develop a robust CV method and work to improve my score on that. My LB score is admittedly low, and I doubt I will do well in this competition. But I think I will do better than my score suggests, and the learning is more important anyway.",
      "votes": null
    },
    {
      "id": "526824",
      "postDate": "05/03/2019 21:11:37",
      "content": "<blockquote>\n  <p>You are quite right of course </p>\n</blockquote>\n\n<p>I am not right or wrong, I just asked a question ;)  Thanks for answering it!</p>",
      "rawMarkdown": "&gt; You are quite right of course \n\nI am not right or wrong, I just asked a question ;)  Thanks for answering it!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 525809,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "05/01/2019 18:42:17",
      "content": "<p>My model has similar appearance.   Since the training data does not show the same shape, I conclude that the sampling of segments for the test was not random, but was intentional.  </p>\n\n<p>I would assume that the creators of the data did much the same work that we are doing.  My guess would therefore be that their models not much different than ours.  </p>\n\n<p>They probably would select samples in the areas where their model had its poorest performance.  On the plots of training vs actual my models don' look that good in the 8 and above.    In that area my predictions for training are generally lower than actual.    They also want samples closer to the actual event.</p>\n\n<p>If I wanted 10,000 folks to try to improve things I would intentionally select segments in the same two general piles that your showing.</p>\n\n<p>In reading some of the available science on this type of experiment it's clear to me that the 8 and above wave shape has completely different characteristics than the waves close to the quake.  So the features that predict time to failure for the 9 peak group need/are different from the 4-5 peak group.</p>\n\n<p>So this kind of reminds me of a common issue I faced over the years - in my case, it was two different machines producing the same product, but their output was binomial when piled into a single data set.  In my case I always needed to run DOE's and analysis on the machines as separate models.</p>\n\n<p>So my plan is to train two different models -model 1 will only look at data for 6 and above segments.  Model 2 will only look at data for less than 6.  Than blend or stack the results.   Might do some overlap, but that will be my general plan.</p>\n\n<p>Would rather find a single model that could separate the output from two similar but different flows - I could have used it for the past 50 years - of course, I need the way back machine to take a PC loaded with Python back to 1970, but maybe....</p>",
      "votes": null,
      "replies": [
        {
          "id": 526000,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/02/2019 06:27:28",
          "content": "<p>??  What Phillipe sows is on train data.  Test data sampling isn’t relevant.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526159,
          "author_name": "redstr",
          "author_url": "",
          "post_date": "05/02/2019 13:30:26",
          "content": "<blockquote>\n  <p>Since the training data does not show the same shape</p>\n</blockquote>\n\n<p>Aha! Found the guy who has the  ground truth!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526182,
          "author_name": "pavelvod",
          "author_url": "",
          "post_date": "05/02/2019 14:14:21",
          "content": "<p>Why not just to train a classification model that will output a probability of being above 6 and then use it as additional feature in regression model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526189,
          "author_name": "welcomeworld",
          "author_url": "",
          "post_date": "05/02/2019 14:28:58",
          "content": "<p>I'll try that for sure and report back soon!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526203,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "05/02/2019 15:05:20",
          "content": "<p><a href=\"/redstr\">@redstr</a> I've actually found the same thing with the training data so don't be alarmed! In fact, I assumed OP was talking about the test data at first since my models typically predict the 4-6 range very well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526212,
          "author_name": "redstr",
          "author_url": "",
          "post_date": "05/02/2019 15:25:23",
          "content": "<p>It was just a joke. I suppose you are talking about a bimodal distribution of predictions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526370,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "05/02/2019 22:20:22",
          "content": "<p>I was looking at test data from my best model prediction.  Pretty sure that's what OP was talking about.</p>\n\n<p>Guess it could always be possible that my model predicts bimodal rather than the test data having that distribution ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526388,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "05/02/2019 23:58:01",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> Alternatively, we've generally found that the 4-6 range is one of the easiest parts of the data to predict. Maybe the organisers deliberately oversampled test segments with particularly high/low TTF in order to evaluate how well we can predict the more complex cases. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526475,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/03/2019 06:17:59",
          "content": "<p><a href=\"/bigironsphere\">@bigironsphere</a> how do you know you predict well the 4-6 range on test data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526802,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "05/03/2019 20:05:38",
          "content": "<p>You are quite right of course - I have no idea. I can only only speculate given:</p>\n\n<ul>\n<li>my model's performance on that range with the training data</li>\n<li>the knowledge that the average TTF of the LB data is around 4</li>\n<li>the fact that our models generally score better on the LB than OOF</li>\n</ul>\n\n<p>The real test data, who knows? They may have reserved especially complicated cases for it. </p>\n\n<p>For the record, I don't obsess over the LB score or waste time trying to second-guess what the test set might look like. My priority is, and has always been, to develop a robust CV method and work to improve my score on that. My LB score is admittedly low, and I doubt I will do well in this competition. But I think I will do better than my score suggests, and the learning is more important anyway.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526824,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/03/2019 21:11:37",
          "content": "<blockquote>\n  <p>You are quite right of course </p>\n</blockquote>\n\n<p>I am not right or wrong, I just asked a question ;)  Thanks for answering it!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "525562": "My best model prediction has a binomial flavor to it. I find it hard to predict those mid-sections in the 4 to 6 sec. range; turning knobs helps a bit but always at the expense of the LB score. How do you guys think I should handle this issue?\n\n(vertical is target, horizontal is oof prediction)",
    "525809": "My model has similar appearance.   Since the training data does not show the same shape, I conclude that the sampling of segments for the test was not random, but was intentional.  \n\nI would assume that the creators of the data did much the same work that we are doing.  My guess would therefore be that their models not much different than ours.  \n\nThey probably would select samples in the areas where their model had its poorest performance.  On the plots of training vs actual my models don' look that good in the 8 and above.    In that area my predictions for training are generally lower than actual.    They also want samples closer to the actual event.\n\nIf I wanted 10,000 folks to try to improve things I would intentionally select segments in the same two general piles that your showing.\n\nIn reading some of the available science on this type of experiment it's clear to me that the 8 and above wave shape has completely different characteristics than the waves close to the quake.  So the features that predict time to failure for the 9 peak group need/are different from the 4-5 peak group.\n\nSo this kind of reminds me of a common issue I faced over the years - in my case, it was two different machines producing the same product, but their output was binomial when piled into a single data set.  In my case I always needed to run DOE's and analysis on the machines as separate models.\n\nSo my plan is to train two different models -model 1 will only look at data for 6 and above segments.  Model 2 will only look at data for less than 6.  Than blend or stack the results.   Might do some overlap, but that will be my general plan.\n\nWould rather find a single model that could separate the output from two similar but different flows - I could have used it for the past 50 years - of course, I need the way back machine to take a PC loaded with Python back to 1970, but maybe....",
    "526000": "??  What Phillipe sows is on train data.  Test data sampling isn’t relevant.",
    "526159": "&gt; Since the training data does not show the same shape\n\nAha! Found the guy who has the  ground truth!",
    "526182": "Why not just to train a classification model that will output a probability of being above 6 and then use it as additional feature in regression model?",
    "526189": "I'll try that for sure and report back soon!",
    "526203": "redstr I've actually found the same thing with the training data so don't be alarmed! In fact, I assumed OP was talking about the test data at first since my models typically predict the 4-6 range very well.",
    "526212": "It was just a joke. I suppose you are talking about a bimodal distribution of predictions.",
    "526370": "I was looking at test data from my best model prediction.  Pretty sure that's what OP was talking about.\n\nGuess it could always be possible that my model predicts bimodal rather than the test data having that distribution ?",
    "526388": "pcjimmmy Alternatively, we've generally found that the 4-6 range is one of the easiest parts of the data to predict. Maybe the organisers deliberately oversampled test segments with particularly high/low TTF in order to evaluate how well we can predict the more complex cases.",
    "526475": "bigironsphere how do you know you predict well the 4-6 range on test data?",
    "526802": "You are quite right of course - I have no idea. I can only only speculate given:\n\n* my model's performance on that range with the training data\n* the knowledge that the average TTF of the LB data is around 4\n* the fact that our models generally score better on the LB than OOF\n\nThe real test data, who knows? They may have reserved especially complicated cases for it. \n\nFor the record, I don't obsess over the LB score or waste time trying to second-guess what the test set might look like. My priority is, and has always been, to develop a robust CV method and work to improve my score on that. My LB score is admittedly low, and I doubt I will do well in this competition. But I think I will do better than my score suggests, and the learning is more important anyway.",
    "526824": "&gt; You are quite right of course \n\nI am not right or wrong, I just asked a question ;)  Thanks for answering it!"
  },
  "source": "meta"
}