{
  "id": 100208,
  "title": "Is Kappa Score Reliable?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/100208",
  "author_name": "",
  "post_date": "2019-07-17T07:27:06.586548600Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I was wondering  how the kappa score correlates with the goodness of the model cause though my kappa score is increasing upto 0.909 but i see no improvement of the submission .The train_loss is coming down to 0.27 and valid loss to 0.25 ,any suggestions would be helpful.Thanks in advance.</p>",
  "messages": [
    {
      "id": "577909",
      "postDate": "07/17/2019 07:27:06",
      "content": "<p>I was wondering  how the kappa score correlates with the goodness of the model cause though my kappa score is increasing upto 0.909 but i see no improvement of the submission .The train_loss is coming down to 0.27 and valid loss to 0.25 ,any suggestions would be helpful.Thanks in advance.</p>",
      "rawMarkdown": "I was wondering  how the kappa score correlates with the goodness of the model cause though my kappa score is increasing upto 0.909 but i see no improvement of the submission .The train_loss is coming down to 0.27 and valid loss to 0.25 ,any suggestions would be helpful.Thanks in advance.",
      "votes": null
    },
    {
      "id": "577926",
      "postDate": "07/17/2019 07:46:39",
      "content": "<p>Are you using kappa coefficients/thresholds optimised on validation set?</p>",
      "rawMarkdown": "Are you using kappa coefficients/thresholds optimised on validation set?",
      "votes": null
    },
    {
      "id": "577933",
      "postDate": "07/17/2019 07:53:36",
      "content": "<p>Hello <a href=\"/amardeepganguly\">@amardeepganguly</a> ,</p>\n\n<p>First of all lets understand and agree that none of the metrics are 'Absolute' …</p>\n\n<p>What do I mean by that …. Ex : Will you agree with me that 'Temperature'  is NOT an absolute measure …..if you don't agree , then you have to explain why do we have something called 'Feels Like' &amp; 'Wind Chill' metrics as well to get a full sense of how is the weather outside.</p>\n\n<p>Similarly , there are instances where we don't look at Metrics in isolation but rather as a group of related metrics.</p>\n\n<p>Example - \n      For Regression , we have RMSE ( Root Mean Square Error , but we do have R-Square as well ) </p>\n\n<p>Now , lets come back to Kappa ..</p>\n\n<p>Cohen’s Kappa statistic is a very useful, but under-utilised, metric. Sometimes in machine learning we are faced with a multi-class classification problem. In those cases, measures such as the accuracy, or precision/recall do not provide the complete picture of the performance of our classifier.\nIn some other cases we might face a problem with imbalanced classes. E.g. we have two classes, say A and B, and A shows up on 5% of the time. Accuracy can be misleading, so we go for measures such as precision and recall. There are ways to combine the two, such as the F-measure, but the F-measure does not have a very good intuitive explanation, other than it being the harmonic mean of precision and recall.\nCohen’s kappa is always less than or equal to 1. Values of 0 or less, indicate that the classifier is useless. There is no standardized way to interpret its values. Landis and Koch (1977) provide a way to characterize values. According to their scheme a value &lt; 0 is indicating no agreement , 0–0.20 as slight, 0.21–0.40 as fair, 0.41–0.60 as moderate, 0.61–0.80 as substantial, and 0.81–1 as almost perfect agreement.</p>\n\n<p>Time for a Layman Example - </p>\n\n<p>Why can't we just see how often we agree?  </p>\n\n<p>Think of this. You and I are deciding if Kernels should be deleted from Kaggle due to plagiarization. Most Kernels shouldn't be deleted from Kaggle , but I'm not a very good judge. So I just randomly select 1% of kernels.</p>\n\n<p>You're a very good judge. You very carefully select 1% of Kernels too ( and these are the Kernels which should be actually deleted due to plagiarizing ).</p>\n\n<p>So we are in agreement in (approximately) 98% of cases ( as 1% identified by me different from 1% identified by you for deletion ) , and we disagree just under 2% of the time. 98% agreement sounds great!  What's the problem?</p>\n\n<p>The problem is that I was just picking kernels at random. Because most kernels are not to be deleted, we agree most of the time. But that doesn't mean that we are using similar ratings.  Kappa corrects for this; our kappa would be approximately zero.  </p>\n\n<p>Higher Kappa is better which indicates the Model is far better than random Chance.</p>",
      "rawMarkdown": "Hello @amardeepganguly ,\n\nFirst of all lets understand and agree that none of the metrics are 'Absolute' …\n\nWhat do I mean by that …. Ex : Will you agree with me that 'Temperature'  is NOT an absolute measure …..if you don't agree , then you have to explain why do we have something called 'Feels Like' &amp; 'Wind Chill' metrics as well to get a full sense of how is the weather outside.\n\nSimilarly , there are instances where we don't look at Metrics in isolation but rather as a group of related metrics.\n\nExample - \n      For Regression , we have RMSE ( Root Mean Square Error , but we do have R-Square as well ) \n\nNow , lets come back to Kappa ..\n\nCohen’s Kappa statistic is a very useful, but under-utilised, metric. Sometimes in machine learning we are faced with a multi-class classification problem. In those cases, measures such as the accuracy, or precision/recall do not provide the complete picture of the performance of our classifier.\nIn some other cases we might face a problem with imbalanced classes. E.g. we have two classes, say A and B, and A shows up on 5% of the time. Accuracy can be misleading, so we go for measures such as precision and recall. There are ways to combine the two, such as the F-measure, but the F-measure does not have a very good intuitive explanation, other than it being the harmonic mean of precision and recall.\nCohen’s kappa is always less than or equal to 1. Values of 0 or less, indicate that the classifier is useless. There is no standardized way to interpret its values. Landis and Koch (1977) provide a way to characterize values. According to their scheme a value &lt; 0 is indicating no agreement , 0–0.20 as slight, 0.21–0.40 as fair, 0.41–0.60 as moderate, 0.61–0.80 as substantial, and 0.81–1 as almost perfect agreement.\n\nTime for a Layman Example - \n\nWhy can't we just see how often we agree?  \n\nThink of this. You and I are deciding if Kernels should be deleted from Kaggle due to plagiarization. Most Kernels shouldn't be deleted from Kaggle , but I'm not a very good judge. So I just randomly select 1% of kernels.\n\nYou're a very good judge. You very carefully select 1% of Kernels too ( and these are the Kernels which should be actually deleted due to plagiarizing ).\n\nSo we are in agreement in (approximately) 98% of cases ( as 1% identified by me different from 1% identified by you for deletion ) , and we disagree just under 2% of the time. 98% agreement sounds great!  What's the problem?\n\nThe problem is that I was just picking kernels at random. Because most kernels are not to be deleted, we agree most of the time. But that doesn't mean that we are using similar ratings.  Kappa corrects for this; our kappa would be approximately zero.  \n\nHigher Kappa is better which indicates the Model is far better than random Chance.",
      "votes": null
    },
    {
      "id": "577940",
      "postDate": "07/17/2019 08:06:49",
      "content": "<p>Yes</p>",
      "rawMarkdown": "Yes",
      "votes": null
    },
    {
      "id": "577956",
      "postDate": "07/17/2019 08:27:21",
      "content": "<p>Okay. Here's my hypothesis: There are ~13k images in private test set, there's high chance that you may end up with a val set which doesn't represent test set closely enough and hence the difference in val kappa score and LB score. Generalization is big factor in this competition. </p>",
      "rawMarkdown": "Okay. Here's my hypothesis: There are ~13k images in private test set, there's high chance that you may end up with a val set which doesn't represent test set closely enough and hence the difference in val kappa score and LB score. Generalization is big factor in this competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 577926,
      "author_name": "rishabhiitbhu",
      "author_url": "",
      "post_date": "07/17/2019 07:46:39",
      "content": "<p>Are you using kappa coefficients/thresholds optimised on validation set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 577940,
          "author_name": "amardeepganguly",
          "author_url": "",
          "post_date": "07/17/2019 08:06:49",
          "content": "<p>Yes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 577956,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "07/17/2019 08:27:21",
          "content": "<p>Okay. Here's my hypothesis: There are ~13k images in private test set, there's high chance that you may end up with a val set which doesn't represent test set closely enough and hence the difference in val kappa score and LB score. Generalization is big factor in this competition. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 577933,
      "author_name": "prmohanty",
      "author_url": "",
      "post_date": "07/17/2019 07:53:36",
      "content": "<p>Hello <a href=\"/amardeepganguly\">@amardeepganguly</a> ,</p>\n\n<p>First of all lets understand and agree that none of the metrics are 'Absolute' …</p>\n\n<p>What do I mean by that …. Ex : Will you agree with me that 'Temperature'  is NOT an absolute measure …..if you don't agree , then you have to explain why do we have something called 'Feels Like' &amp; 'Wind Chill' metrics as well to get a full sense of how is the weather outside.</p>\n\n<p>Similarly , there are instances where we don't look at Metrics in isolation but rather as a group of related metrics.</p>\n\n<p>Example - \n      For Regression , we have RMSE ( Root Mean Square Error , but we do have R-Square as well ) </p>\n\n<p>Now , lets come back to Kappa ..</p>\n\n<p>Cohen’s Kappa statistic is a very useful, but under-utilised, metric. Sometimes in machine learning we are faced with a multi-class classification problem. In those cases, measures such as the accuracy, or precision/recall do not provide the complete picture of the performance of our classifier.\nIn some other cases we might face a problem with imbalanced classes. E.g. we have two classes, say A and B, and A shows up on 5% of the time. Accuracy can be misleading, so we go for measures such as precision and recall. There are ways to combine the two, such as the F-measure, but the F-measure does not have a very good intuitive explanation, other than it being the harmonic mean of precision and recall.\nCohen’s kappa is always less than or equal to 1. Values of 0 or less, indicate that the classifier is useless. There is no standardized way to interpret its values. Landis and Koch (1977) provide a way to characterize values. According to their scheme a value &lt; 0 is indicating no agreement , 0–0.20 as slight, 0.21–0.40 as fair, 0.41–0.60 as moderate, 0.61–0.80 as substantial, and 0.81–1 as almost perfect agreement.</p>\n\n<p>Time for a Layman Example - </p>\n\n<p>Why can't we just see how often we agree?  </p>\n\n<p>Think of this. You and I are deciding if Kernels should be deleted from Kaggle due to plagiarization. Most Kernels shouldn't be deleted from Kaggle , but I'm not a very good judge. So I just randomly select 1% of kernels.</p>\n\n<p>You're a very good judge. You very carefully select 1% of Kernels too ( and these are the Kernels which should be actually deleted due to plagiarizing ).</p>\n\n<p>So we are in agreement in (approximately) 98% of cases ( as 1% identified by me different from 1% identified by you for deletion ) , and we disagree just under 2% of the time. 98% agreement sounds great!  What's the problem?</p>\n\n<p>The problem is that I was just picking kernels at random. Because most kernels are not to be deleted, we agree most of the time. But that doesn't mean that we are using similar ratings.  Kappa corrects for this; our kappa would be approximately zero.  </p>\n\n<p>Higher Kappa is better which indicates the Model is far better than random Chance.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "577909": "I was wondering  how the kappa score correlates with the goodness of the model cause though my kappa score is increasing upto 0.909 but i see no improvement of the submission .The train_loss is coming down to 0.27 and valid loss to 0.25 ,any suggestions would be helpful.Thanks in advance.",
    "577926": "Are you using kappa coefficients/thresholds optimised on validation set?",
    "577933": "Hello @amardeepganguly ,\n\nFirst of all lets understand and agree that none of the metrics are 'Absolute' …\n\nWhat do I mean by that …. Ex : Will you agree with me that 'Temperature'  is NOT an absolute measure …..if you don't agree , then you have to explain why do we have something called 'Feels Like' &amp; 'Wind Chill' metrics as well to get a full sense of how is the weather outside.\n\nSimilarly , there are instances where we don't look at Metrics in isolation but rather as a group of related metrics.\n\nExample - \n      For Regression , we have RMSE ( Root Mean Square Error , but we do have R-Square as well ) \n\nNow , lets come back to Kappa ..\n\nCohen’s Kappa statistic is a very useful, but under-utilised, metric. Sometimes in machine learning we are faced with a multi-class classification problem. In those cases, measures such as the accuracy, or precision/recall do not provide the complete picture of the performance of our classifier.\nIn some other cases we might face a problem with imbalanced classes. E.g. we have two classes, say A and B, and A shows up on 5% of the time. Accuracy can be misleading, so we go for measures such as precision and recall. There are ways to combine the two, such as the F-measure, but the F-measure does not have a very good intuitive explanation, other than it being the harmonic mean of precision and recall.\nCohen’s kappa is always less than or equal to 1. Values of 0 or less, indicate that the classifier is useless. There is no standardized way to interpret its values. Landis and Koch (1977) provide a way to characterize values. According to their scheme a value &lt; 0 is indicating no agreement , 0–0.20 as slight, 0.21–0.40 as fair, 0.41–0.60 as moderate, 0.61–0.80 as substantial, and 0.81–1 as almost perfect agreement.\n\nTime for a Layman Example - \n\nWhy can't we just see how often we agree?  \n\nThink of this. You and I are deciding if Kernels should be deleted from Kaggle due to plagiarization. Most Kernels shouldn't be deleted from Kaggle , but I'm not a very good judge. So I just randomly select 1% of kernels.\n\nYou're a very good judge. You very carefully select 1% of Kernels too ( and these are the Kernels which should be actually deleted due to plagiarizing ).\n\nSo we are in agreement in (approximately) 98% of cases ( as 1% identified by me different from 1% identified by you for deletion ) , and we disagree just under 2% of the time. 98% agreement sounds great!  What's the problem?\n\nThe problem is that I was just picking kernels at random. Because most kernels are not to be deleted, we agree most of the time. But that doesn't mean that we are using similar ratings.  Kappa corrects for this; our kappa would be approximately zero.  \n\nHigher Kappa is better which indicates the Model is far better than random Chance.",
    "577940": "Yes",
    "577956": "Okay. Here's my hypothesis: There are ~13k images in private test set, there's high chance that you may end up with a val set which doesn't represent test set closely enough and hence the difference in val kappa score and LB score. Generalization is big factor in this competition."
  },
  "source": "meta"
}