{
  "id": 85108,
  "title": "Quick and dirty booster",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/85108",
  "author_name": "",
  "post_date": "2019-03-21T19:58:56.431788800Z",
  "votes": 10,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I've been examining the distribution of the errors and it seems like my lgb models are consistently overestimating the target value. By adjusting the median of my prediction to match the target, I got a boost both in cv (2.0702 -&gt; 2.0538) and lb (1.512 -&gt; 1.489). I'm almost sure it's overfitting, but I thought I'd put in out there. Curious what you guys think.</p>",
  "messages": [
    {
      "id": "496015",
      "postDate": "03/21/2019 19:58:56",
      "content": "<p>I've been examining the distribution of the errors and it seems like my lgb models are consistently overestimating the target value. By adjusting the median of my prediction to match the target, I got a boost both in cv (2.0702 -&gt; 2.0538) and lb (1.512 -&gt; 1.489). I'm almost sure it's overfitting, but I thought I'd put in out there. Curious what you guys think.</p>",
      "rawMarkdown": "I've been examining the distribution of the errors and it seems like my lgb models are consistently overestimating the target value. By adjusting the median of my prediction to match the target, I got a boost both in cv (2.0702 -&gt; 2.0538) and lb (1.512 -&gt; 1.489). I'm almost sure it's overfitting, but I thought I'd put in out there. Curious what you guys think.",
      "votes": null
    },
    {
      "id": "496449",
      "postDate": "03/22/2019 08:19:21",
      "content": "<p>Thanks for sharing this trick! If I understood correctly, in this case your target is the testing set?</p>",
      "rawMarkdown": "Thanks for sharing this trick! If I understood correctly, in this case your target is the testing set?",
      "votes": null
    },
    {
      "id": "496454",
      "postDate": "03/22/2019 08:22:03",
      "content": "<p>Indeed.</p>",
      "rawMarkdown": "Indeed.",
      "votes": null
    },
    {
      "id": "496570",
      "postDate": "03/22/2019 10:52:26",
      "content": "<p>Why \"almost sure\"? ;)</p>",
      "rawMarkdown": "Why \"almost sure\"? ;)",
      "votes": null
    },
    {
      "id": "496873",
      "postDate": "03/22/2019 17:32:51",
      "content": "<p>The test dataset is small, and the necessary shift varies per folder quite a bit.</p>",
      "rawMarkdown": "The test dataset is small, and the necessary shift varies per folder quite a bit.",
      "votes": null
    },
    {
      "id": "499404",
      "postDate": "03/24/2019 18:36:11",
      "content": "<p>It's an interesting find, but seems like one of those things that wouldn't impact the private leaderboard. Like you said, the test set is pretty small. I have a theory as to why our LB scores are so much better than the CV scores, I'll post it when I'm more certain. Out of curiosity, how are you adjusting the median to match the target? Do you mean you found a function to transform the training predictions closer to the actual value, and then applied it to test? </p>",
      "rawMarkdown": "It's an interesting find, but seems like one of those things that wouldn't impact the private leaderboard. Like you said, the test set is pretty small. I have a theory as to why our LB scores are so much better than the CV scores, I'll post it when I'm more certain. Out of curiosity, how are you adjusting the median to match the target? Do you mean you found a function to transform the training predictions closer to the actual value, and then applied it to test?",
      "votes": null
    },
    {
      "id": "499458",
      "postDate": "03/24/2019 19:59:49",
      "content": "<p>Manually shifting the whole thing the match the median across the training set. I did call it dirty for a reason :-) </p>",
      "rawMarkdown": "Manually shifting the whole thing the match the median across the training set. I did call it dirty for a reason :-)",
      "votes": null
    },
    {
      "id": "499480",
      "postDate": "03/24/2019 20:45:52",
      "content": "<p>I really, really hope that doesn't increase my Private LB score!</p>",
      "rawMarkdown": "I really, really hope that doesn't increase my Private LB score!",
      "votes": null
    },
    {
      "id": "522445",
      "postDate": "04/24/2019 12:44:19",
      "content": "<p>Interesting.  Maybe because your loss function is mse when the metric is mae.  Former optimizes mean while the latter optimizes median.  I guess mean is higher than median here.  I'll give it a try.</p>",
      "rawMarkdown": "Interesting.  Maybe because your loss function is mse when the metric is mae.  Former optimizes mean while the latter optimizes median.  I guess mean is higher than median here.  I'll give it a try.",
      "votes": null
    },
    {
      "id": "522470",
      "postDate": "04/24/2019 13:28:17",
      "content": "<p>From a theoretical standpoint, yes - but given the sample size issues, reliability of either estimate (mean / median) is problematic. That's why I called it quick and dirty - \"caveat emptor\" for our time.</p>",
      "rawMarkdown": "From a theoretical standpoint, yes - but given the sample size issues, reliability of either estimate (mean / median) is problematic. That's why I called it quick and dirty - \"caveat emptor\" for our time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 496449,
      "author_name": "ricarddelgado",
      "author_url": "",
      "post_date": "03/22/2019 08:19:21",
      "content": "<p>Thanks for sharing this trick! If I understood correctly, in this case your target is the testing set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 496454,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/22/2019 08:22:03",
          "content": "<p>Indeed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 496570,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "03/22/2019 10:52:26",
      "content": "<p>Why \"almost sure\"? ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 496873,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/22/2019 17:32:51",
          "content": "<p>The test dataset is small, and the necessary shift varies per folder quite a bit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 499404,
      "author_name": "bigironsphere",
      "author_url": "",
      "post_date": "03/24/2019 18:36:11",
      "content": "<p>It's an interesting find, but seems like one of those things that wouldn't impact the private leaderboard. Like you said, the test set is pretty small. I have a theory as to why our LB scores are so much better than the CV scores, I'll post it when I'm more certain. Out of curiosity, how are you adjusting the median to match the target? Do you mean you found a function to transform the training predictions closer to the actual value, and then applied it to test? </p>",
      "votes": null,
      "replies": [
        {
          "id": 499458,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/24/2019 19:59:49",
          "content": "<p>Manually shifting the whole thing the match the median across the training set. I did call it dirty for a reason :-) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 499480,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "03/24/2019 20:45:52",
          "content": "<p>I really, really hope that doesn't increase my Private LB score!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 522445,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/24/2019 12:44:19",
      "content": "<p>Interesting.  Maybe because your loss function is mse when the metric is mae.  Former optimizes mean while the latter optimizes median.  I guess mean is higher than median here.  I'll give it a try.</p>",
      "votes": null,
      "replies": [
        {
          "id": 522470,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "04/24/2019 13:28:17",
          "content": "<p>From a theoretical standpoint, yes - but given the sample size issues, reliability of either estimate (mean / median) is problematic. That's why I called it quick and dirty - \"caveat emptor\" for our time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "496015": "I've been examining the distribution of the errors and it seems like my lgb models are consistently overestimating the target value. By adjusting the median of my prediction to match the target, I got a boost both in cv (2.0702 -&gt; 2.0538) and lb (1.512 -&gt; 1.489). I'm almost sure it's overfitting, but I thought I'd put in out there. Curious what you guys think.",
    "496449": "Thanks for sharing this trick! If I understood correctly, in this case your target is the testing set?",
    "496454": "Indeed.",
    "496570": "Why \"almost sure\"? ;)",
    "496873": "The test dataset is small, and the necessary shift varies per folder quite a bit.",
    "499404": "It's an interesting find, but seems like one of those things that wouldn't impact the private leaderboard. Like you said, the test set is pretty small. I have a theory as to why our LB scores are so much better than the CV scores, I'll post it when I'm more certain. Out of curiosity, how are you adjusting the median to match the target? Do you mean you found a function to transform the training predictions closer to the actual value, and then applied it to test?",
    "499458": "Manually shifting the whole thing the match the median across the training set. I did call it dirty for a reason :-)",
    "499480": "I really, really hope that doesn't increase my Private LB score!",
    "522445": "Interesting.  Maybe because your loss function is mse when the metric is mae.  Former optimizes mean while the latter optimizes median.  I guess mean is higher than median here.  I'll give it a try.",
    "522470": "From a theoretical standpoint, yes - but given the sample size issues, reliability of either estimate (mean / median) is problematic. That's why I called it quick and dirty - \"caveat emptor\" for our time."
  },
  "source": "meta"
}