{
  "id": 59756,
  "title": "Clipping deal probability. ",
  "url": "/competitions/avito-demand-prediction/discussion/59756",
  "author_name": "",
  "post_date": "2018-06-26T19:11:52.555906600Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>A quick question - is it good to clip those very high/low predicted probabilities? Specially near zero?\nI did such a thing in iceberg challenge and it ended up pretty bad.\nIs there any good rule of thumb to follow in such s situation?\nThanks. </p>",
  "messages": [
    {
      "id": "348459",
      "postDate": "06/26/2018 19:11:52",
      "content": "<p>A quick question - is it good to clip those very high/low predicted probabilities? Specially near zero?\nI did such a thing in iceberg challenge and it ended up pretty bad.\nIs there any good rule of thumb to follow in such s situation?\nThanks. </p>",
      "rawMarkdown": "A quick question - is it good to clip those very high/low predicted probabilities? Specially near zero?\nI did such a thing in iceberg challenge and it ended up pretty bad.\nIs there any good rule of thumb to follow in such s situation?\nThanks.",
      "votes": null
    },
    {
      "id": "348463",
      "postDate": "06/26/2018 19:15:50",
      "content": "<p>Tried - Did not work for me.</p>",
      "rawMarkdown": "Tried - Did not work for me.",
      "votes": null
    },
    {
      "id": "348481",
      "postDate": "06/26/2018 19:44:32",
      "content": "<p>I've never seen this work... but this could be the first time? Usually we trust our models to clip themselves correctly (except for the 0-1 hard bound).</p>",
      "rawMarkdown": "I've never seen this work... but this could be the first time? Usually we trust our models to clip themselves correctly (except for the 0-1 hard bound).",
      "votes": null
    },
    {
      "id": "348576",
      "postDate": "06/27/2018 00:12:51",
      "content": "<p>I think clipping usually works for competition measured by log loss, but not RMSE (Large penalty to falsely predicted extremely large or small value ) </p>",
      "rawMarkdown": "I think clipping usually works for competition measured by log loss, but not RMSE (Large penalty to falsely predicted extremely large or small value )",
      "votes": null
    },
    {
      "id": "348606",
      "postDate": "06/27/2018 02:02:29",
      "content": "<p>Don‘t work for me.</p>",
      "rawMarkdown": "Don‘t work for me.",
      "votes": null
    },
    {
      "id": "348658",
      "postDate": "06/27/2018 04:29:39",
      "content": "<p>Totally makes sense.</p>",
      "rawMarkdown": "Totally makes sense.",
      "votes": null
    },
    {
      "id": "348673",
      "postDate": "06/27/2018 05:26:42",
      "content": "<p>It is interesting to look at this approach. As dieter has said, I attempted this, but it did not work for me either. I originally set my model to set anything below a certain threshold to zero because I saw that there is a large portion that are set at exactly zero. I tried setting various thresholds to match this number and it actually very negatively affected the score. As others have said. RMSE is very unforgiving in this regard. </p>",
      "rawMarkdown": "It is interesting to look at this approach. As dieter has said, I attempted this, but it did not work for me either. I originally set my model to set anything below a certain threshold to zero because I saw that there is a large portion that are set at exactly zero. I tried setting various thresholds to match this number and it actually very negatively affected the score. As others have said. RMSE is very unforgiving in this regard.",
      "votes": null
    },
    {
      "id": "348765",
      "postDate": "06/27/2018 08:31:18",
      "content": "<p>I think it does not change the RMSE score itself, but in terms of 'better prediction' clipping have some significant meaning. \n -depending on the competition. Not to mention the cases when you are using boosting  and it gives invalid range of values - for example, we can clip range somewhere between 0 and 1.</p>\n\n<p>About near zero values, it didn't work for me this time, but I think in some cases it could work - for example when there exists extremely large percentage of '0' values (or other specific values) - I think most of the regressors, or boosting model would give the values somewhere near that value, but it would be rare to give exact value. So is we clip the values very carefully, I believe we can reduce RMSE by reducing losses from missed '0' values. </p>\n\n<p>However, this would only works\n1) if percentage of 0 (or specific value) is big enough, so reduced RMSE by clipping insufficient predictions to 0 &gt; increased RMSE by clipping non 0 values to 0. </p>\n\n<p>2) When we clipped so carefully, so clipping threshold does not include values which are not supposed to be clipped</p>",
      "rawMarkdown": "I think it does not change the RMSE score itself, but in terms of 'better prediction' clipping have some significant meaning. \n -depending on the competition. Not to mention the cases when you are using boosting  and it gives invalid range of values - for example, we can clip range somewhere between 0 and 1.\n\nAbout near zero values, it didn't work for me this time, but I think in some cases it could work - for example when there exists extremely large percentage of '0' values (or other specific values) - I think most of the regressors, or boosting model would give the values somewhere near that value, but it would be rare to give exact value. So is we clip the values very carefully, I believe we can reduce RMSE by reducing losses from missed '0' values. \n\nHowever, this would only works\n1) if percentage of 0 (or specific value) is big enough, so reduced RMSE by clipping insufficient predictions to 0 &gt; increased RMSE by clipping non 0 values to 0. \n\n2) When we clipped so carefully, so clipping threshold does not include values which are not supposed to be clipped",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 348463,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "06/26/2018 19:15:50",
      "content": "<p>Tried - Did not work for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 348481,
      "author_name": "peterhurford",
      "author_url": "",
      "post_date": "06/26/2018 19:44:32",
      "content": "<p>I've never seen this work... but this could be the first time? Usually we trust our models to clip themselves correctly (except for the 0-1 hard bound).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 348576,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "06/27/2018 00:12:51",
      "content": "<p>I think clipping usually works for competition measured by log loss, but not RMSE (Large penalty to falsely predicted extremely large or small value ) </p>",
      "votes": null,
      "replies": [
        {
          "id": 348658,
          "author_name": "nuhsikander",
          "author_url": "",
          "post_date": "06/27/2018 04:29:39",
          "content": "<p>Totally makes sense.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 348606,
      "author_name": "baomengjiao",
      "author_url": "",
      "post_date": "06/27/2018 02:02:29",
      "content": "<p>Don‘t work for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 348673,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "06/27/2018 05:26:42",
      "content": "<p>It is interesting to look at this approach. As dieter has said, I attempted this, but it did not work for me either. I originally set my model to set anything below a certain threshold to zero because I saw that there is a large portion that are set at exactly zero. I tried setting various thresholds to match this number and it actually very negatively affected the score. As others have said. RMSE is very unforgiving in this regard. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 348765,
      "author_name": "sukhyun9673",
      "author_url": "",
      "post_date": "06/27/2018 08:31:18",
      "content": "<p>I think it does not change the RMSE score itself, but in terms of 'better prediction' clipping have some significant meaning. \n -depending on the competition. Not to mention the cases when you are using boosting  and it gives invalid range of values - for example, we can clip range somewhere between 0 and 1.</p>\n\n<p>About near zero values, it didn't work for me this time, but I think in some cases it could work - for example when there exists extremely large percentage of '0' values (or other specific values) - I think most of the regressors, or boosting model would give the values somewhere near that value, but it would be rare to give exact value. So is we clip the values very carefully, I believe we can reduce RMSE by reducing losses from missed '0' values. </p>\n\n<p>However, this would only works\n1) if percentage of 0 (or specific value) is big enough, so reduced RMSE by clipping insufficient predictions to 0 &gt; increased RMSE by clipping non 0 values to 0. </p>\n\n<p>2) When we clipped so carefully, so clipping threshold does not include values which are not supposed to be clipped</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "348459": "A quick question - is it good to clip those very high/low predicted probabilities? Specially near zero?\nI did such a thing in iceberg challenge and it ended up pretty bad.\nIs there any good rule of thumb to follow in such s situation?\nThanks.",
    "348463": "Tried - Did not work for me.",
    "348481": "I've never seen this work... but this could be the first time? Usually we trust our models to clip themselves correctly (except for the 0-1 hard bound).",
    "348576": "I think clipping usually works for competition measured by log loss, but not RMSE (Large penalty to falsely predicted extremely large or small value )",
    "348606": "Don‘t work for me.",
    "348658": "Totally makes sense.",
    "348673": "It is interesting to look at this approach. As dieter has said, I attempted this, but it did not work for me either. I originally set my model to set anything below a certain threshold to zero because I saw that there is a large portion that are set at exactly zero. I tried setting various thresholds to match this number and it actually very negatively affected the score. As others have said. RMSE is very unforgiving in this regard.",
    "348765": "I think it does not change the RMSE score itself, but in terms of 'better prediction' clipping have some significant meaning. \n -depending on the competition. Not to mention the cases when you are using boosting  and it gives invalid range of values - for example, we can clip range somewhere between 0 and 1.\n\nAbout near zero values, it didn't work for me this time, but I think in some cases it could work - for example when there exists extremely large percentage of '0' values (or other specific values) - I think most of the regressors, or boosting model would give the values somewhere near that value, but it would be rare to give exact value. So is we clip the values very carefully, I believe we can reduce RMSE by reducing losses from missed '0' values. \n\nHowever, this would only works\n1) if percentage of 0 (or specific value) is big enough, so reduced RMSE by clipping insufficient predictions to 0 &gt; increased RMSE by clipping non 0 values to 0. \n\n2) When we clipped so carefully, so clipping threshold does not include values which are not supposed to be clipped"
  },
  "source": "meta"
}