{
  "id": 93511,
  "title": "Prediction histogram",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/93511",
  "author_name": "",
  "post_date": "2019-05-27T20:03:37.327913900Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I have been experimenting with an unusual model which performs poorly. I am not surprised and I am still improving but here is one diagnostic I have made. It is the histogram of the submission predictions for a few variations of the model.</p>\n\n<p>Note that I do suffer from the \"mini earthquake\" problem, and that the model variation is not huge. What surprised me was the preference of the submission for smaller ttf's (below 4). This seems odd if the test data are chosen randomly from a number of earthquake cycles.</p>\n\n<p>Are others seeing this? If not, any ideas on what I have done wrong (hard question given lack of info I provided!). Perhaps I have missed an important thread.</p>",
  "messages": [
    {
      "id": "537901",
      "postDate": "05/27/2019 20:03:37",
      "content": "<p>I have been experimenting with an unusual model which performs poorly. I am not surprised and I am still improving but here is one diagnostic I have made. It is the histogram of the submission predictions for a few variations of the model.</p>\n\n<p>Note that I do suffer from the \"mini earthquake\" problem, and that the model variation is not huge. What surprised me was the preference of the submission for smaller ttf's (below 4). This seems odd if the test data are chosen randomly from a number of earthquake cycles.</p>\n\n<p>Are others seeing this? If not, any ideas on what I have done wrong (hard question given lack of info I provided!). Perhaps I have missed an important thread.</p>",
      "rawMarkdown": "I have been experimenting with an unusual model which performs poorly. I am not surprised and I am still improving but here is one diagnostic I have made. It is the histogram of the submission predictions for a few variations of the model.\n\nNote that I do suffer from the \"mini earthquake\" problem, and that the model variation is not huge. What surprised me was the preference of the submission for smaller ttf's (below 4). This seems odd if the test data are chosen randomly from a number of earthquake cycles.\n\nAre others seeing this? If not, any ideas on what I have done wrong (hard question given lack of info I provided!). Perhaps I have missed an important thread.",
      "votes": null
    },
    {
      "id": "537907",
      "postDate": "05/27/2019 20:13:58",
      "content": "<p>Seeing. No comment until after competition is over ;)</p>",
      "rawMarkdown": "Seeing. No comment until after competition is over ;)",
      "votes": null
    },
    {
      "id": "538098",
      "postDate": "05/28/2019 06:06:30",
      "content": "<p><a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844\">This thread</a> is important, take a look. But besides that note that optimal (optimal in a sense that you extract and use all the information from each segment fully and correctly) histogram of predictions can be far away from histogram of actual test values. For example, test set contains some values &gt;16, but it is correct to never give such predictions, because for each segment you want give median of its estimated density function. Well, I skipped many details in this explanation, but hope it clarifies.</p>",
      "rawMarkdown": "[This thread](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844) is important, take a look. But besides that note that optimal (optimal in a sense that you extract and use all the information from each segment fully and correctly) histogram of predictions can be far away from histogram of actual test values. For example, test set contains some values &gt;16, but it is correct to never give such predictions, because for each segment you want give median of its estimated density function. Well, I skipped many details in this explanation, but hope it clarifies.",
      "votes": null
    },
    {
      "id": "540242",
      "postDate": "05/31/2019 07:59:41",
      "content": "<p>I guess the link is wrong?</p>",
      "rawMarkdown": "I guess the link is wrong?",
      "votes": null
    },
    {
      "id": "540265",
      "postDate": "05/31/2019 08:48:46",
      "content": "<p>Thank you, corrected. For some reason it doesn't open for me when I click it, but it works if you do \"open in new window\"</p>",
      "rawMarkdown": "Thank you, corrected. For some reason it doesn't open for me when I click it, but it works if you do \"open in new window\"",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 537907,
      "author_name": "teeyee314",
      "author_url": "",
      "post_date": "05/27/2019 20:13:58",
      "content": "<p>Seeing. No comment until after competition is over ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 538098,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "05/28/2019 06:06:30",
      "content": "<p><a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844\">This thread</a> is important, take a look. But besides that note that optimal (optimal in a sense that you extract and use all the information from each segment fully and correctly) histogram of predictions can be far away from histogram of actual test values. For example, test set contains some values &gt;16, but it is correct to never give such predictions, because for each segment you want give median of its estimated density function. Well, I skipped many details in this explanation, but hope it clarifies.</p>",
      "votes": null,
      "replies": [
        {
          "id": 540242,
          "author_name": "lucaskg",
          "author_url": "",
          "post_date": "05/31/2019 07:59:41",
          "content": "<p>I guess the link is wrong?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540265,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "05/31/2019 08:48:46",
          "content": "<p>Thank you, corrected. For some reason it doesn't open for me when I click it, but it works if you do \"open in new window\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "537901": "I have been experimenting with an unusual model which performs poorly. I am not surprised and I am still improving but here is one diagnostic I have made. It is the histogram of the submission predictions for a few variations of the model.\n\nNote that I do suffer from the \"mini earthquake\" problem, and that the model variation is not huge. What surprised me was the preference of the submission for smaller ttf's (below 4). This seems odd if the test data are chosen randomly from a number of earthquake cycles.\n\nAre others seeing this? If not, any ideas on what I have done wrong (hard question given lack of info I provided!). Perhaps I have missed an important thread.",
    "537907": "Seeing. No comment until after competition is over ;)",
    "538098": "[This thread](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844) is important, take a look. But besides that note that optimal (optimal in a sense that you extract and use all the information from each segment fully and correctly) histogram of predictions can be far away from histogram of actual test values. For example, test set contains some values &gt;16, but it is correct to never give such predictions, because for each segment you want give median of its estimated density function. Well, I skipped many details in this explanation, but hope it clarifies.",
    "540242": "I guess the link is wrong?",
    "540265": "Thank you, corrected. For some reason it doesn't open for me when I click it, but it works if you do \"open in new window\""
  },
  "source": "meta"
}