{
  "id": 15550,
  "title": "Distribution of probabilities for test set",
  "url": "/competitions/avito-context-ad-clicks/discussion/15550",
  "author_name": "",
  "post_date": "2015-07-26T08:08:52.033Z",
  "votes": null,
  "comment_count": 6,
  "views": 1223,
  "content": "<p>I was curious about the prediction probabilities and did a histogram on it. </p>\n\n<p>For a reasonable submission with LB ~0.047XX (I know it is still way down, but reasonable regardless), I got the following values for a 10 bin histogram:</p>\n\n<p>Frequency:\n7760859   48281        4109        1843         637         322         158         132           0          20</p>\n\n<p>bin center (probability):\n0.0174    0.0503    0.0832    0.1161    0.1490    0.1819    0.2148    0.2477    0.2806    0.3135</p>\n\n<p>with max prob =  0.33 and min prob = 0.000947</p>\n\n<p>7760859 out of 7816361 total testing set got probability less than 0.05. </p>\n\n<p>Does this mean that the model is not calibrated? </p>\n\n<p>It seems that the model is almost useless as it won't predict even a single click with more confidence than what's possible by chance although it achieves an overall logloss ~0.047XX. </p>\n\n<p>What are the practical implications of this model?</p>",
  "messages": [
    {
      "id": "87007",
      "postDate": "07/26/2015 08:08:52",
      "content": "<p>I was curious about the prediction probabilities and did a histogram on it. </p>\n\n<p>For a reasonable submission with LB ~0.047XX (I know it is still way down, but reasonable regardless), I got the following values for a 10 bin histogram:</p>\n\n<p>Frequency:\n7760859   48281        4109        1843         637         322         158         132           0          20</p>\n\n<p>bin center (probability):\n0.0174    0.0503    0.0832    0.1161    0.1490    0.1819    0.2148    0.2477    0.2806    0.3135</p>\n\n<p>with max prob =  0.33 and min prob = 0.000947</p>\n\n<p>7760859 out of 7816361 total testing set got probability less than 0.05. </p>\n\n<p>Does this mean that the model is not calibrated? </p>\n\n<p>It seems that the model is almost useless as it won't predict even a single click with more confidence than what's possible by chance although it achieves an overall logloss ~0.047XX. </p>\n\n<p>What are the practical implications of this model?</p>",
      "rawMarkdown": "I was curious about the prediction probabilities and did a histogram on it. \r\n\r\nFor a reasonable submission with LB ~0.047XX (I know it is still way down, but reasonable regardless), I got the following values for a 10 bin histogram:\r\n\r\nFrequency:\r\n7760859   48281        4109        1843         637         322         158         132           0          20\r\n\r\nbin center (probability):\r\n0.0174    0.0503    0.0832    0.1161    0.1490    0.1819    0.2148    0.2477    0.2806    0.3135\r\n\r\nwith max prob =  0.33 and min prob = 0.000947\r\n\r\n7760859 out of 7816361 total testing set got probability less than 0.05. \r\n\r\nDoes this mean that the model is not calibrated? \r\n\r\nIt seems that the model is almost useless as it won't predict even a single click with more confidence than what's possible by chance although it achieves an overall logloss ~0.047XX. \r\n\r\nWhat are the practical implications of this model?",
      "votes": null
    },
    {
      "id": "87019",
      "postDate": "07/26/2015 11:43:54",
      "content": "<pre><code> What are the practical implications of this model?\n</code></pre>\n\n<p>In practice, you can show ads with higher probability and earn more money. </p>",
      "rawMarkdown": "What are the practical implications of this model?\r\n\r\nIn practice, you can show ads with higher probability and earn more money.",
      "votes": null
    },
    {
      "id": "87114",
      "postDate": "07/27/2015 01:49:49",
      "content": "<p>@Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. </p>\n\n<p>If the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!</p>",
      "rawMarkdown": "Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. \r\n\r\nIf the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!",
      "votes": null
    },
    {
      "id": "87119",
      "postDate": "07/27/2015 02:24:08",
      "content": "<p>[quote=Deep;87114]</p>\n\n<p>@Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. </p>\n\n<p>If the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!</p>\n\n<p>[/quote]</p>\n\n<p>Why ask if you dont bother to consider other people answers? \nIn this competition they wanted calibrated probabilities,  and on this kind of model one usually expect the avg predicted CTR to be very close to the actual CTRs. Did you check the actual CTR? So the predicted probabilities are within expected range. \nNot all models are build with a cut off of 0.5 in mind. Some apllications will use only probabilities, and other will use only the ranking. In this case a calibrated model will also give expected revenue (probability*price) and that is most desirable. </p>\n\n<p>So you really @Artem answered right because &quot;In practice, you can show ads with higher probability and earn more money.&quot;</p>",
      "rawMarkdown": "[quote=Deep;87114]\r\n\r\n@Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. \r\n\r\nIf the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!\r\n\r\n[/quote]\r\n\r\nWhy ask if you dont bother to consider other people answers? \r\nIn this competition they wanted calibrated probabilities,  and on this kind of model one usually expect the avg predicted CTR to be very close to the actual CTRs. Did you check the actual CTR? So the predicted probabilities are within expected range. \r\nNot all models are build with a cut off of 0.5 in mind. Some apllications will use only probabilities, and other will use only the ranking. In this case a calibrated model will also give expected revenue (probability*price) and that is most desirable. \r\n\r\nSo you really @Artem answered right because \"In practice, you can show ads with higher probability and earn more money.\"",
      "votes": null
    },
    {
      "id": "87122",
      "postDate": "07/27/2015 02:35:34",
      "content": "<p>@Leustagos, thanks for your answer. You answered my question. I was just wondering whether calibrated probabilities are what we are after. Sorry if I sounded like I didn't consider his answer. It was just too generic for my specific question.</p>",
      "rawMarkdown": "Leustagos, thanks for your answer. You answered my question. I was just wondering whether calibrated probabilities are what we are after. Sorry if I sounded like I didn't consider his answer. It was just too generic for my specific question.",
      "votes": null
    },
    {
      "id": "87134",
      "postDate": "07/27/2015 06:24:57",
      "content": "<p>[quote=Deep;87114]</p>\n\n<p>@Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. </p>\n\n<p>If the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!</p>\n\n<p>[/quote]\nBetter than nothing, isn't it?</p>\n\n<p>BTW, 0.33 is too high to be real.</p>",
      "rawMarkdown": "[quote=Deep;87114]\r\n\r\n@Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. \r\n\r\nIf the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!\r\n\r\n[/quote]\r\nBetter than nothing, isn't it?\r\n\r\nBTW, 0.33 is too high to be real.",
      "votes": null
    },
    {
      "id": "87224",
      "postDate": "07/27/2015 22:25:11",
      "content": "<p>[quote=Deep;87114]</p>\n\n<p>@Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. </p>\n\n<p>If the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!</p>\n\n<p>[/quote]</p>\n\n<p>Deep, </p>\n\n<p>If the average CTR rate is &lt; 1%, why would a .33 probability be less than a random guess? </p>\n\n<p>Remember, our models are attempting to predict HUMAN BEHAVIOR - which is about as stable as a toddler's temperament. For practicality, we really just want to know which ad is the user MOST LIKELY to click on.  </p>\n\n<p>Also, I agree with Artem - .33 likely too high.   </p>",
      "rawMarkdown": "[quote=Deep;87114]\r\n\r\n@Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. \r\n\r\nIf the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!\r\n\r\n[/quote]\r\n\r\nDeep, \r\n\r\nIf the average CTR rate is < 1%, why would a .33 probability be less than a random guess? \r\n\r\nRemember, our models are attempting to predict HUMAN BEHAVIOR - which is about as stable as a toddler's temperament. For practicality, we really just want to know which ad is the user MOST LIKELY to click on.  \r\n\r\nAlso, I agree with Artem - .33 likely too high.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 87019,
      "author_name": "rushter",
      "author_url": "",
      "post_date": "07/26/2015 11:43:54",
      "content": "<pre><code> What are the practical implications of this model?\n</code></pre>\n\n<p>In practice, you can show ads with higher probability and earn more money. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87114,
      "author_name": "deepcnn",
      "author_url": "",
      "post_date": "07/27/2015 01:49:49",
      "content": "<p>@Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. </p>\n\n<p>If the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87119,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "07/27/2015 02:24:08",
      "content": "<p>[quote=Deep;87114]</p>\n\n<p>@Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. </p>\n\n<p>If the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!</p>\n\n<p>[/quote]</p>\n\n<p>Why ask if you dont bother to consider other people answers? \nIn this competition they wanted calibrated probabilities,  and on this kind of model one usually expect the avg predicted CTR to be very close to the actual CTRs. Did you check the actual CTR? So the predicted probabilities are within expected range. \nNot all models are build with a cut off of 0.5 in mind. Some apllications will use only probabilities, and other will use only the ranking. In this case a calibrated model will also give expected revenue (probability*price) and that is most desirable. </p>\n\n<p>So you really @Artem answered right because &quot;In practice, you can show ads with higher probability and earn more money.&quot;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87122,
      "author_name": "deepcnn",
      "author_url": "",
      "post_date": "07/27/2015 02:35:34",
      "content": "<p>@Leustagos, thanks for your answer. You answered my question. I was just wondering whether calibrated probabilities are what we are after. Sorry if I sounded like I didn't consider his answer. It was just too generic for my specific question.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87134,
      "author_name": "rushter",
      "author_url": "",
      "post_date": "07/27/2015 06:24:57",
      "content": "<p>[quote=Deep;87114]</p>\n\n<p>@Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. </p>\n\n<p>If the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!</p>\n\n<p>[/quote]\nBetter than nothing, isn't it?</p>\n\n<p>BTW, 0.33 is too high to be real.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87224,
      "author_name": "rhinodaveb",
      "author_url": "",
      "post_date": "07/27/2015 22:25:11",
      "content": "<p>[quote=Deep;87114]</p>\n\n<p>@Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. </p>\n\n<p>If the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!</p>\n\n<p>[/quote]</p>\n\n<p>Deep, </p>\n\n<p>If the average CTR rate is &lt; 1%, why would a .33 probability be less than a random guess? </p>\n\n<p>Remember, our models are attempting to predict HUMAN BEHAVIOR - which is about as stable as a toddler's temperament. For practicality, we really just want to know which ad is the user MOST LIKELY to click on.  </p>\n\n<p>Also, I agree with Artem - .33 likely too high.   </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "87007": "I was curious about the prediction probabilities and did a histogram on it. \r\n\r\nFor a reasonable submission with LB ~0.047XX (I know it is still way down, but reasonable regardless), I got the following values for a 10 bin histogram:\r\n\r\nFrequency:\r\n7760859   48281        4109        1843         637         322         158         132           0          20\r\n\r\nbin center (probability):\r\n0.0174    0.0503    0.0832    0.1161    0.1490    0.1819    0.2148    0.2477    0.2806    0.3135\r\n\r\nwith max prob =  0.33 and min prob = 0.000947\r\n\r\n7760859 out of 7816361 total testing set got probability less than 0.05. \r\n\r\nDoes this mean that the model is not calibrated? \r\n\r\nIt seems that the model is almost useless as it won't predict even a single click with more confidence than what's possible by chance although it achieves an overall logloss ~0.047XX. \r\n\r\nWhat are the practical implications of this model?",
    "87019": "What are the practical implications of this model?\r\n\r\nIn practice, you can show ads with higher probability and earn more money.",
    "87114": "Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. \r\n\r\nIf the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!",
    "87119": "[quote=Deep;87114]\r\n\r\n@Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. \r\n\r\nIf the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!\r\n\r\n[/quote]\r\n\r\nWhy ask if you dont bother to consider other people answers? \r\nIn this competition they wanted calibrated probabilities,  and on this kind of model one usually expect the avg predicted CTR to be very close to the actual CTRs. Did you check the actual CTR? So the predicted probabilities are within expected range. \r\nNot all models are build with a cut off of 0.5 in mind. Some apllications will use only probabilities, and other will use only the ranking. In this case a calibrated model will also give expected revenue (probability*price) and that is most desirable. \r\n\r\nSo you really @Artem answered right because \"In practice, you can show ads with higher probability and earn more money.\"",
    "87122": "Leustagos, thanks for your answer. You answered my question. I was just wondering whether calibrated probabilities are what we are after. Sorry if I sounded like I didn't consider his answer. It was just too generic for my specific question.",
    "87134": "[quote=Deep;87114]\r\n\r\n@Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. \r\n\r\nIf the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!\r\n\r\n[/quote]\r\nBetter than nothing, isn't it?\r\n\r\nBTW, 0.33 is too high to be real.",
    "87224": "[quote=Deep;87114]\r\n\r\n@Artem, You don't think I knew that? I was asking about this specific model not the goal of the competition. \r\n\r\nIf the largest probability this model can give is 0.33 (which is less than just random guess), what is the practical usefulness of the model although LB score is reasonable? In other words, why not randomly guess?!\r\n\r\n[/quote]\r\n\r\nDeep, \r\n\r\nIf the average CTR rate is < 1%, why would a .33 probability be less than a random guess? \r\n\r\nRemember, our models are attempting to predict HUMAN BEHAVIOR - which is about as stable as a toddler's temperament. For practicality, we really just want to know which ad is the user MOST LIKELY to click on.  \r\n\r\nAlso, I agree with Artem - .33 likely too high."
  },
  "source": "meta"
}