{
  "id": 98678,
  "title": "Why score can go below zero?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/98678",
  "author_name": "",
  "post_date": "2019-07-05T12:33:07.085988300Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>In Evaluation page, it says: \" In the event that there is less agreement between the raters than expected by chance, this metric may go below 0\".</p>\n\n<p>I can not understand why quadratic weighted kappa score can go below 0.</p>\n\n<p>Example, make a submission includes all values is one, It scores 0.0 in LB.</p>\n\n<p>But, in normal quadratic weighted kappa, it scores over zero in any time.</p>\n\n<p>What means \"less agreement between the raters than expected by chance\" ?</p>\n\n<p>The ratio of five values hit by chance is 0.2, so should I multiply the Kappa score by 0.8? But then negative values do not appear. Or pull 0.2? No, LB has a score of 0.8 or more.</p>\n\n<p>I think the explanation of the Evaluation page is missing, or so hard to understand. It would be helpful to introduce a detailed formula.</p>",
  "messages": [
    {
      "id": "568772",
      "postDate": "07/05/2019 12:33:07",
      "content": "<p>In Evaluation page, it says: \" In the event that there is less agreement between the raters than expected by chance, this metric may go below 0\".</p>\n\n<p>I can not understand why quadratic weighted kappa score can go below 0.</p>\n\n<p>Example, make a submission includes all values is one, It scores 0.0 in LB.</p>\n\n<p>But, in normal quadratic weighted kappa, it scores over zero in any time.</p>\n\n<p>What means \"less agreement between the raters than expected by chance\" ?</p>\n\n<p>The ratio of five values hit by chance is 0.2, so should I multiply the Kappa score by 0.8? But then negative values do not appear. Or pull 0.2? No, LB has a score of 0.8 or more.</p>\n\n<p>I think the explanation of the Evaluation page is missing, or so hard to understand. It would be helpful to introduce a detailed formula.</p>",
      "rawMarkdown": "In Evaluation page, it says: \" In the event that there is less agreement between the raters than expected by chance, this metric may go below 0\".\n\nI can not understand why quadratic weighted kappa score can go below 0.\n\nExample, make a submission includes all values is one, It scores 0.0 in LB.\n\nBut, in normal quadratic weighted kappa, it scores over zero in any time.\n\nWhat means \"less agreement between the raters than expected by chance\" ?\n\nThe ratio of five values hit by chance is 0.2, so should I multiply the Kappa score by 0.8? But then negative values do not appear. Or pull 0.2? No, LB has a score of 0.8 or more.\n\nI think the explanation of the Evaluation page is missing, or so hard to understand. It would be helpful to introduce a detailed formula.",
      "votes": null
    },
    {
      "id": "568824",
      "postDate": "07/05/2019 13:47:15",
      "content": "<blockquote>\n  <p>What means \"less agreement between the raters than expected by chance\" ?</p>\n</blockquote>\n\n<p>\"Expected by chance\" means if you were to not use any machine learning model and just take a completely random guess of your classification. If your model performs worse than some generic person guessing randomly, then your model scores below zero because it would have been negatively effective. A zero evaluation score means it performs no better than randomly guessing.</p>",
      "rawMarkdown": "&gt; What means \"less agreement between the raters than expected by chance\" ?\n\n\"Expected by chance\" means if you were to not use any machine learning model and just take a completely random guess of your classification. If your model performs worse than some generic person guessing randomly, then your model scores below zero because it would have been negatively effective. A zero evaluation score means it performs no better than randomly guessing.",
      "votes": null
    },
    {
      "id": "568829",
      "postDate": "07/05/2019 13:55:13",
      "content": "<p>I understand you very well. I want the details.\nFor example, if all the data is 1, LB score is 0.0, but if half the value is randomly set to 0, the score will be negative. On the other hand, if you set a very small value to 2 randomly, the score will be slightly higher (around 0.05).\nHow is this calculated?</p>",
      "rawMarkdown": "I understand you very well. I want the details.\nFor example, if all the data is 1, LB score is 0.0, but if half the value is randomly set to 0, the score will be negative. On the other hand, if you set a very small value to 2 randomly, the score will be slightly higher (around 0.05).\nHow is this calculated?",
      "votes": null
    },
    {
      "id": "570277",
      "postDate": "07/08/2019 04:52:09",
      "content": "<p>Basically 2 histograms matrices are constructed. One for observations(predictions) and other for expectations. Probabilites  are calculated based on these histograms. It could be that probability for observation is less than probability for expected. In that case kappa can be negative.\nPlease refer to <a href=\"https://www.kaggle.com/ashwan1/understanding-kappa-using-dummy-classifier\">this kernel</a> to get complete details on histogram matrices, weights, and overall <em>kappa</em>. External links to fully worked examples provided.\nHope, it will clear your doubts regarding this metric.</p>",
      "rawMarkdown": "Basically 2 histograms matrices are constructed. One for observations(predictions) and other for expectations. Probabilites  are calculated based on these histograms. It could be that probability for observation is less than probability for expected. In that case kappa can be negative.\nPlease refer to [this kernel](https://www.kaggle.com/ashwan1/understanding-kappa-using-dummy-classifier) to get complete details on histogram matrices, weights, and overall *kappa*. External links to fully worked examples provided.\nHope, it will clear your doubts regarding this metric.",
      "votes": null
    },
    {
      "id": "570459",
      "postDate": "07/08/2019 10:16:37",
      "content": "<p>Thank you, Ashwani. I understood.</p>\n\n<p>I thought, but if you look closely at the LB, can you predict the histogram of the correct values in the test dataset?</p>",
      "rawMarkdown": "Thank you, Ashwani. I understood.\n\nI thought, but if you look closely at the LB, can you predict the histogram of the correct values in the test dataset?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 568824,
      "author_name": "jamalrahman",
      "author_url": "",
      "post_date": "07/05/2019 13:47:15",
      "content": "<blockquote>\n  <p>What means \"less agreement between the raters than expected by chance\" ?</p>\n</blockquote>\n\n<p>\"Expected by chance\" means if you were to not use any machine learning model and just take a completely random guess of your classification. If your model performs worse than some generic person guessing randomly, then your model scores below zero because it would have been negatively effective. A zero evaluation score means it performs no better than randomly guessing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 568829,
          "author_name": "tanreinama",
          "author_url": "",
          "post_date": "07/05/2019 13:55:13",
          "content": "<p>I understand you very well. I want the details.\nFor example, if all the data is 1, LB score is 0.0, but if half the value is randomly set to 0, the score will be negative. On the other hand, if you set a very small value to 2 randomly, the score will be slightly higher (around 0.05).\nHow is this calculated?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 570277,
      "author_name": "ashwan1",
      "author_url": "",
      "post_date": "07/08/2019 04:52:09",
      "content": "<p>Basically 2 histograms matrices are constructed. One for observations(predictions) and other for expectations. Probabilites  are calculated based on these histograms. It could be that probability for observation is less than probability for expected. In that case kappa can be negative.\nPlease refer to <a href=\"https://www.kaggle.com/ashwan1/understanding-kappa-using-dummy-classifier\">this kernel</a> to get complete details on histogram matrices, weights, and overall <em>kappa</em>. External links to fully worked examples provided.\nHope, it will clear your doubts regarding this metric.</p>",
      "votes": null,
      "replies": [
        {
          "id": 570459,
          "author_name": "tanreinama",
          "author_url": "",
          "post_date": "07/08/2019 10:16:37",
          "content": "<p>Thank you, Ashwani. I understood.</p>\n\n<p>I thought, but if you look closely at the LB, can you predict the histogram of the correct values in the test dataset?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "568772": "In Evaluation page, it says: \" In the event that there is less agreement between the raters than expected by chance, this metric may go below 0\".\n\nI can not understand why quadratic weighted kappa score can go below 0.\n\nExample, make a submission includes all values is one, It scores 0.0 in LB.\n\nBut, in normal quadratic weighted kappa, it scores over zero in any time.\n\nWhat means \"less agreement between the raters than expected by chance\" ?\n\nThe ratio of five values hit by chance is 0.2, so should I multiply the Kappa score by 0.8? But then negative values do not appear. Or pull 0.2? No, LB has a score of 0.8 or more.\n\nI think the explanation of the Evaluation page is missing, or so hard to understand. It would be helpful to introduce a detailed formula.",
    "568824": "&gt; What means \"less agreement between the raters than expected by chance\" ?\n\n\"Expected by chance\" means if you were to not use any machine learning model and just take a completely random guess of your classification. If your model performs worse than some generic person guessing randomly, then your model scores below zero because it would have been negatively effective. A zero evaluation score means it performs no better than randomly guessing.",
    "568829": "I understand you very well. I want the details.\nFor example, if all the data is 1, LB score is 0.0, but if half the value is randomly set to 0, the score will be negative. On the other hand, if you set a very small value to 2 randomly, the score will be slightly higher (around 0.05).\nHow is this calculated?",
    "570277": "Basically 2 histograms matrices are constructed. One for observations(predictions) and other for expectations. Probabilites  are calculated based on these histograms. It could be that probability for observation is less than probability for expected. In that case kappa can be negative.\nPlease refer to [this kernel](https://www.kaggle.com/ashwan1/understanding-kappa-using-dummy-classifier) to get complete details on histogram matrices, weights, and overall *kappa*. External links to fully worked examples provided.\nHope, it will clear your doubts regarding this metric.",
    "570459": "Thank you, Ashwani. I understood.\n\nI thought, but if you look closely at the LB, can you predict the histogram of the correct values in the test dataset?"
  },
  "source": "meta"
}