{
  "id": 198135,
  "title": "Competition Metric",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/198135",
  "author_name": "Yaroslav Isaienkov",
  "post_date": "2020-11-19T23:01:20.613000",
  "votes": 57,
  "comment_count": 37,
  "views": 0,
  "content": "<p>Did you think about using another metric in this competition?</p>\n<p>For example categorical log loss.</p>\n<p>Because I think many players will be able to achieve a very high accuracy score, and for many, this indicator will be equal. And if your dataset has an imbalance, then the accuracy metric is not the best choice for evaluating models.</p>",
  "messages": [
    {
      "id": 1084305,
      "postDate": "2020-11-19T23:01:20.613Z",
      "content": "<p>Did you think about using another metric in this competition?</p>\n<p>For example categorical log loss.</p>\n<p>Because I think many players will be able to achieve a very high accuracy score, and for many, this indicator will be equal. And if your dataset has an imbalance, then the accuracy metric is not the best choice for evaluating models.</p>",
      "rawMarkdown": "Did you think about using another metric in this competition?\n\nFor example categorical log loss.\n\nBecause I think many players will be able to achieve a very high accuracy score, and for many, this indicator will be equal. And if your dataset has an imbalance, then the accuracy metric is not the best choice for evaluating models.",
      "votes": 56
    },
    {
      "id": 1084737,
      "postDate": "2020-11-20T10:42:51.350Z",
      "content": "<p>Agree, F1 or (even better) the Matthews Correlation Coefficient might be better alternatives</p>",
      "rawMarkdown": "Agree, F1 or (even better) the Matthews Correlation Coefficient might be better alternatives",
      "votes": 6
    },
    {
      "id": 1087385,
      "postDate": "2020-11-22T16:55:47.270Z",
      "content": "<p>whatever the metric is, who gets the most accurate predictions wins. so I don't think it matters much</p>",
      "rawMarkdown": "whatever the metric is, who gets the most accurate predictions wins. so I don't think it matters much",
      "votes": 5,
      "replies": [
        {
          "id": 1087413,
          "postDate": "2020-11-22T17:28:05.927Z",
          "content": "<p><a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a> That's correct. It doesn't matter to us but assume a scenario that the winner's model is working very well on imbalanced class (which have very high number of data points) but performs average on other classes. This will impact to organizers.</p>",
          "rawMarkdown": "@moewie94 That's correct. It doesn't matter to us but assume a scenario that the winner's model is working very well on imbalanced class (which have very high number of data points) but performs average on other classes. This will impact to organizers.",
          "votes": 3
        },
        {
          "id": 1087694,
          "postDate": "2020-11-23T02:20:10.700Z",
          "content": "<p>this can even be the way the organizer wants it. they surely aware of this matters when they split the test set (they can just exclude some class 3 samples in the test set to make to test set class balanced) but chose to not do it so it might not be a big problem. furthermore, this can even be the way this problem work in real-life scenario (for example the third class appears most in nature). so just split everything to be balanced is not a good way sometimes.</p>",
          "rawMarkdown": "this can even be the way the organizer wants it. they surely aware of this matters when they split the test set (they can just exclude some class 3 samples in the test set to make to test set class balanced) but chose to not do it so it might not be a big problem. furthermore, this can even be the way this problem work in real-life scenario (for example the third class appears most in nature). so just split everything to be balanced is not a good way sometimes.",
          "votes": 2
        },
        {
          "id": 1088086,
          "postDate": "2020-11-23T10:14:03.263Z",
          "content": "<p><a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a> It matters because finally we have to classify all the 4 diseases correctly. As there is a less data for remaining 3 diseases accuracy score will give less weightage to those 3 diseases. Our model wont get enough data to train properly for these 3 diseases compared to the class 3 disease(higher one)</p>",
          "rawMarkdown": "@moewie94 It matters because finally we have to classify all the 4 diseases correctly. As there is a less data for remaining 3 diseases accuracy score will give less weightage to those 3 diseases. Our model wont get enough data to train properly for these 3 diseases compared to the class 3 disease(higher one)",
          "votes": 2
        },
        {
          "id": 1088338,
          "postDate": "2020-11-23T14:30:54.187Z",
          "content": "<p>my point is the organizers are fully aware of this matter and they let it happen this way so this imbalance can be their very intention. every organizer has to pay tens of thousands of dollars for kaggle consultant services and every dataset must be analyzed by the kaggle ds team before being accepted for a competition. so i think it's unlikely they are naive enough to not know this class imbalance and how it affects the scoring system.</p>\n<p>sometimes the test set distribution should be imbalanced like the way they are in real life, not like the way they are in MNIST.</p>",
          "rawMarkdown": "my point is the organizers are fully aware of this matter and they let it happen this way so this imbalance can be their very intention. every organizer has to pay tens of thousands of dollars for kaggle consultant services and every dataset must be analyzed by the kaggle ds team before being accepted for a competition. so i think it's unlikely they are naive enough to not know this class imbalance and how it affects the scoring system.\n\nsometimes the test set distribution should be imbalanced like the way they are in real life, not like the way they are in MNIST.",
          "votes": 8
        },
        {
          "id": 1088344,
          "postDate": "2020-11-23T14:36:12.837Z",
          "content": "<p><a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a> I totally agree with you (I don't know why you got so many dislikes)<br>\nMy goal was to open this discussion to make sure that the metric will not be changed in the middle of the competition.</p>",
          "rawMarkdown": "@moewie94 I totally agree with you (I don't know why you got so many dislikes)\nMy goal was to open this discussion to make sure that the metric will not be changed in the middle of the competition.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1085513,
      "postDate": "2020-11-21T00:00:11.307Z",
      "content": "<p>Agreed! Accuracy is a very bad choice for this dataset. However I've never seen a change of metrics in the middle of a competition. </p>",
      "rawMarkdown": "Agreed! Accuracy is a very bad choice for this dataset. However I've never seen a change of metrics in the middle of a competition. ",
      "votes": 3,
      "replies": [
        {
          "id": 1085900,
          "postDate": "2020-11-21T09:11:17.543Z",
          "content": "<p>It happened many times (first Jigsaw competition , Ion Switching etc. ) </p>\n<p>For <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48639\" target=\"_blank\">Jigsaw</a> , the metric was changed about a month after the beginning.  But Kaggle extended the deadline thereafter.</p>",
          "rawMarkdown": "It happened many times (first Jigsaw competition , Ion Switching etc. ) \n\nFor [Jigsaw](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48639) , the metric was changed about a month after the beginning.  But Kaggle extended the deadline thereafter.",
          "votes": 5
        },
        {
          "id": 1085954,
          "postDate": "2020-11-21T09:50:34.197Z",
          "content": "<p>Thanks for the info. I hope this accuracy metrics can change! Otherwise I don't even know which loss function to choose.</p>",
          "rawMarkdown": "Thanks for the info. I hope this accuracy metrics can change! Otherwise I don't even know which loss function to choose.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1084365,
      "postDate": "2020-11-20T01:03:58.410Z",
      "content": "<p>Categorical log loss metric holds good for imbalance classification <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a> </p>",
      "rawMarkdown": "Categorical log loss metric holds good for imbalance classification @ihelon ",
      "votes": 3
    },
    {
      "id": 1085132,
      "postDate": "2020-11-20T17:35:28.750Z",
      "content": "<p>From my perspective, I am not sure how the competition would improve with a different metric.  Accuracy always increases as the number of correct predictions increase, and the competition is won by the submission with the highest number of correct predictions.</p>",
      "rawMarkdown": "From my perspective, I am not sure how the competition would improve with a different metric.  Accuracy always increases as the number of correct predictions increase, and the competition is won by the submission with the highest number of correct predictions.",
      "votes": 2,
      "replies": [
        {
          "id": 1085135,
          "postDate": "2020-11-20T17:38:26.397Z",
          "content": "<p><a href=\"https://www.kaggle.com/jeffalltogether\" target=\"_blank\">@jeffalltogether</a> Yes, your point of view may also be correct. In this case, we need to think over an approach for fitting the imbalance</p>",
          "rawMarkdown": "@jeffalltogether Yes, your point of view may also be correct. In this case, we need to think over an approach for fitting the imbalance",
          "votes": 3
        },
        {
          "id": 1088022,
          "postDate": "2020-11-23T08:47:33.197Z",
          "content": "<p><a href=\"https://www.kaggle.com/jeffalltogether\" target=\"_blank\">@jeffalltogether</a> Main aim of this competition is to classify these 4 diseases correctly because farmers has to use the fertilizers for that particular disease. <br>\n             As there are more class 3 diseases our model will learn perfectly for class 3 but for remaining class it will fail. As the data for remaining 3 diseases is less, in accuracy score they will have less weightage.  <br>\n              So along with accuracy score we also have to look about how accurately we are predicting for each class of disease</p>",
          "rawMarkdown": "@jeffalltogether Main aim of this competition is to classify these 4 diseases correctly because farmers has to use the fertilizers for that particular disease. \n             As there are more class 3 diseases our model will learn perfectly for class 3 but for remaining class it will fail. As the data for remaining 3 diseases is less, in accuracy score they will have less weightage.  \n              So along with accuracy score we also have to look about how accurately we are predicting for each class of disease",
          "votes": 2
        }
      ]
    },
    {
      "id": 1084459,
      "postDate": "2020-11-20T03:52:46.353Z",
      "content": "<p>I think ROC-AUC score is best fit for imbalance dataset. You can calculate ROC-AUC via sklearn.metrics.roc_auc_score on the test set.</p>",
      "rawMarkdown": "I think ROC-AUC score is best fit for imbalance dataset. You can calculate ROC-AUC via sklearn.metrics.roc_auc_score on the test set.",
      "votes": 2
    },
    {
      "id": 1090333,
      "postDate": "2020-11-25T09:03:38.980Z",
      "content": "<p><a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a> Thanks for creating this post on metric Yaroslav . I enjoyed going through the discussion . </p>",
      "rawMarkdown": "@ihelon Thanks for creating this post on metric Yaroslav . I enjoyed going through the discussion . ",
      "votes": 1,
      "replies": [
        {
          "id": 1095307,
          "postDate": "2020-11-29T13:17:25.253Z",
          "content": "<p><a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> u are welcome)</p>",
          "rawMarkdown": "@usharengaraju u are welcome)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1087967,
      "postDate": "2020-11-23T07:47:03.670Z",
      "content": "<p>Agree. Since at least the public dataset has imbalanced classes (<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198410\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198410</a> ), other evaluation metrics (e.g. F1) would be suitable for this competition.</p>",
      "rawMarkdown": "Agree. Since at least the public dataset has imbalanced classes (https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198410 ), other evaluation metrics (e.g. F1) would be suitable for this competition.",
      "votes": 1,
      "replies": [
        {
          "id": 1088041,
          "postDate": "2020-11-23T09:08:38.500Z",
          "content": "<p>Yes you are right<br>\nI showed this fact first in the notebook<br>\n<a href=\"https://www.kaggle.com/ihelon/cassava-leaf-disease-exploratory-data-analysis\" target=\"_blank\">https://www.kaggle.com/ihelon/cassava-leaf-disease-exploratory-data-analysis</a></p>",
          "rawMarkdown": "Yes you are right\nI showed this fact first in the notebook\nhttps://www.kaggle.com/ihelon/cassava-leaf-disease-exploratory-data-analysis",
          "votes": 1
        },
        {
          "id": 1088687,
          "postDate": "2020-11-23T21:18:47.843Z",
          "content": "<p>The imbalance in the dataset seems to reflect the proportions one would encounter in a practical scenario, based on domain research, as I discuss in the post linked below, so I think we can safely assume that the private test set is likely to be similar. <br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198584\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198584</a></p>",
          "rawMarkdown": "The imbalance in the dataset seems to reflect the proportions one would encounter in a practical scenario, based on domain research, as I discuss in the post linked below, so I think we can safely assume that the private test set is likely to be similar. \nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198584",
          "votes": 1
        }
      ]
    },
    {
      "id": 1089735,
      "postDate": "2020-11-24T18:21:06.080Z",
      "content": "<p>Good point! I would like to see an answer to this!</p>",
      "rawMarkdown": "Good point! I would like to see an answer to this!",
      "votes": 2,
      "replies": [
        {
          "id": 1095306,
          "postDate": "2020-11-29T13:17:00.043Z",
          "content": "<p><a href=\"https://www.kaggle.com/aldi1t\" target=\"_blank\">@aldi1t</a> I am waiting too)</p>",
          "rawMarkdown": "@aldi1t I am waiting too)"
        }
      ]
    },
    {
      "id": 1088438,
      "postDate": "2020-11-23T15:52:52.440Z",
      "content": "<p>I think this is a good point! Seems like competition hosts should reply here.</p>",
      "rawMarkdown": "I think this is a good point! Seems like competition hosts should reply here.",
      "votes": 2,
      "replies": [
        {
          "id": 1095305,
          "postDate": "2020-11-29T13:16:16.200Z",
          "content": "<p><a href=\"https://www.kaggle.com/bayartsogtya\" target=\"_blank\">@bayartsogtya</a> agree</p>",
          "rawMarkdown": "@bayartsogtya agree"
        }
      ]
    },
    {
      "id": 1087066,
      "postDate": "2020-11-22T10:03:21.363Z",
      "content": "<p>The organizer seems too silent on this until now even it's a serious matter. </p>",
      "rawMarkdown": "The organizer seems too silent on this until now even it's a serious matter. ",
      "votes": 2,
      "replies": [
        {
          "id": 1091939,
          "postDate": "2020-11-26T11:52:59.730Z",
          "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> Could you please respond to this ongoing discussion? Thanks!</p>",
          "rawMarkdown": "@juliaelliott Could you please respond to this ongoing discussion? Thanks!",
          "votes": 1
        },
        {
          "id": 1096875,
          "postDate": "2020-11-30T20:51:02.307Z",
          "content": "<p>The competition metric as defined will remain as-is.</p>",
          "rawMarkdown": "The competition metric as defined will remain as-is.",
          "votes": 4
        },
        {
          "id": 1096876,
          "postDate": "2020-11-30T20:54:12.097Z",
          "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> thanks for the clarification!</p>",
          "rawMarkdown": "@juliaelliott thanks for the clarification!"
        }
      ]
    },
    {
      "id": 1085927,
      "postDate": "2020-11-21T09:29:29.340Z",
      "content": "<p>Mathews Correlation Coefficient or F1-Score can handle the metrics for imbalanced class datasets.</p>",
      "rawMarkdown": "Mathews Correlation Coefficient or F1-Score can handle the metrics for imbalanced class datasets.",
      "votes": 2
    },
    {
      "id": 1085753,
      "postDate": "2020-11-21T07:15:11.210Z",
      "content": "<p>I would vote for categorical log loss or even micro-f1 score. Accuracy is not a stable indicator, especially when the validation\\testing set is small.</p>",
      "rawMarkdown": "I would vote for categorical log loss or even micro-f1 score. Accuracy is not a stable indicator, especially when the validation\\testing set is small.",
      "votes": 2
    },
    {
      "id": 1084807,
      "postDate": "2020-11-20T12:08:16.370Z",
      "content": "<p>Yes, as the data is highly imbalanced, only predicting class 3, in this case, would achieve decent accuracy!</p>\n<p>Edit: Here's the proof: <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198410\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198410</a></p>",
      "rawMarkdown": "Yes, as the data is highly imbalanced, only predicting class 3, in this case, would achieve decent accuracy!\n\nEdit: Here's the proof: https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198410",
      "votes": 2
    },
    {
      "id": 1084603,
      "postDate": "2020-11-20T07:48:44.140Z",
      "content": "<p>For this type of imbalanced data categorical loss is the best metric to evaluate the performance of the model. <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a> </p>",
      "rawMarkdown": "For this type of imbalanced data categorical loss is the best metric to evaluate the performance of the model. @ihelon ",
      "votes": 2,
      "replies": [
        {
          "id": 1084660,
          "postDate": "2020-11-20T09:00:29.887Z",
          "content": "<p><a href=\"https://www.kaggle.com/akshaychavan123\" target=\"_blank\">@akshaychavan123</a> But to use this metric, they need to change the prediction format - probabilities instead of classes</p>",
          "rawMarkdown": "@akshaychavan123 But to use this metric, they need to change the prediction format - probabilities instead of classes",
          "votes": 1
        }
      ]
    },
    {
      "id": 1084460,
      "postDate": "2020-11-20T03:57:16.933Z",
      "content": "<p>if the host want a category prediction rather than the probability, macro/micro fbeta maybe a good choice</p>",
      "rawMarkdown": "if the host want a category prediction rather than the probability, macro/micro fbeta maybe a good choice",
      "votes": 2,
      "replies": [
        {
          "id": 1084583,
          "postDate": "2020-11-20T07:27:22.080Z",
          "content": "<p><a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> yes, I think it's a good idea to leave category prediction but change the metric</p>",
          "rawMarkdown": "@steamedsheep yes, I think it's a good idea to leave category prediction but change the metric",
          "votes": 1
        }
      ]
    },
    {
      "id": 1150967,
      "postDate": "2021-01-13T02:44:14.533Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1086886,
      "postDate": "2020-11-22T06:37:32.433Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1084737,
      "author_name": "Jacopo Repossi",
      "author_url": "",
      "post_date": "2020-11-20T10:42:51.350000",
      "content": "<p>Agree, F1 or (even better) the Matthews Correlation Coefficient might be better alternatives</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1087385,
      "author_name": "DatNT",
      "author_url": "",
      "post_date": "2020-11-22T16:55:47.270000",
      "content": "<p>whatever the metric is, who gets the most accurate predictions wins. so I don't think it matters much</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1087413,
          "author_name": "Kaushal Shah",
          "author_url": "",
          "post_date": "2020-11-22T17:28:05.927000",
          "content": "<p><a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a> That's correct. It doesn't matter to us but assume a scenario that the winner's model is working very well on imbalanced class (which have very high number of data points) but performs average on other classes. This will impact to organizers.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1087694,
          "author_name": "DatNT",
          "author_url": "",
          "post_date": "2020-11-23T02:20:10.700000",
          "content": "<p>this can even be the way the organizer wants it. they surely aware of this matters when they split the test set (they can just exclude some class 3 samples in the test set to make to test set class balanced) but chose to not do it so it might not be a big problem. furthermore, this can even be the way this problem work in real-life scenario (for example the third class appears most in nature). so just split everything to be balanced is not a good way sometimes.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1088086,
          "author_name": "vinay",
          "author_url": "",
          "post_date": "2020-11-23T10:14:03.263000",
          "content": "<p><a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a> It matters because finally we have to classify all the 4 diseases correctly. As there is a less data for remaining 3 diseases accuracy score will give less weightage to those 3 diseases. Our model wont get enough data to train properly for these 3 diseases compared to the class 3 disease(higher one)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1088338,
          "author_name": "DatNT",
          "author_url": "",
          "post_date": "2020-11-23T14:30:54.187000",
          "content": "<p>my point is the organizers are fully aware of this matter and they let it happen this way so this imbalance can be their very intention. every organizer has to pay tens of thousands of dollars for kaggle consultant services and every dataset must be analyzed by the kaggle ds team before being accepted for a competition. so i think it's unlikely they are naive enough to not know this class imbalance and how it affects the scoring system.</p>\n<p>sometimes the test set distribution should be imbalanced like the way they are in real life, not like the way they are in MNIST.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1088344,
          "author_name": "Yaroslav Isaienkov",
          "author_url": "",
          "post_date": "2020-11-23T14:36:12.837000",
          "content": "<p><a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a> I totally agree with you (I don't know why you got so many dislikes)<br>\nMy goal was to open this discussion to make sure that the metric will not be changed in the middle of the competition.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1085513,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2020-11-21T00:00:11.307000",
      "content": "<p>Agreed! Accuracy is a very bad choice for this dataset. However I've never seen a change of metrics in the middle of a competition. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1085900,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-11-21T09:11:17.543000",
          "content": "<p>It happened many times (first Jigsaw competition , Ion Switching etc. ) </p>\n<p>For <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48639\" target=\"_blank\">Jigsaw</a> , the metric was changed about a month after the beginning.  But Kaggle extended the deadline thereafter.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1085954,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2020-11-21T09:50:34.197000",
          "content": "<p>Thanks for the info. I hope this accuracy metrics can change! Otherwise I don't even know which loss function to choose.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1084365,
      "author_name": "Praveen ",
      "author_url": "",
      "post_date": "2020-11-20T01:03:58.410000",
      "content": "<p>Categorical log loss metric holds good for imbalance classification <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a> </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1085132,
      "author_name": "jeffalltogether",
      "author_url": "",
      "post_date": "2020-11-20T17:35:28.750000",
      "content": "<p>From my perspective, I am not sure how the competition would improve with a different metric.  Accuracy always increases as the number of correct predictions increase, and the competition is won by the submission with the highest number of correct predictions.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1085135,
          "author_name": "Yaroslav Isaienkov",
          "author_url": "",
          "post_date": "2020-11-20T17:38:26.397000",
          "content": "<p><a href=\"https://www.kaggle.com/jeffalltogether\" target=\"_blank\">@jeffalltogether</a> Yes, your point of view may also be correct. In this case, we need to think over an approach for fitting the imbalance</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1088022,
          "author_name": "vinay",
          "author_url": "",
          "post_date": "2020-11-23T08:47:33.197000",
          "content": "<p><a href=\"https://www.kaggle.com/jeffalltogether\" target=\"_blank\">@jeffalltogether</a> Main aim of this competition is to classify these 4 diseases correctly because farmers has to use the fertilizers for that particular disease. <br>\n             As there are more class 3 diseases our model will learn perfectly for class 3 but for remaining class it will fail. As the data for remaining 3 diseases is less, in accuracy score they will have less weightage.  <br>\n              So along with accuracy score we also have to look about how accurately we are predicting for each class of disease</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1084459,
      "author_name": "Nguyen Duc Quoc",
      "author_url": "",
      "post_date": "2020-11-20T03:52:46.353000",
      "content": "<p>I think ROC-AUC score is best fit for imbalance dataset. You can calculate ROC-AUC via sklearn.metrics.roc_auc_score on the test set.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1090333,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2020-11-25T09:03:38.980000",
      "content": "<p><a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a> Thanks for creating this post on metric Yaroslav . I enjoyed going through the discussion . </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1095307,
          "author_name": "Yaroslav Isaienkov",
          "author_url": "",
          "post_date": "2020-11-29T13:17:25.253000",
          "content": "<p><a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> u are welcome)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1087967,
      "author_name": "tropicbird",
      "author_url": "",
      "post_date": "2020-11-23T07:47:03.670000",
      "content": "<p>Agree. Since at least the public dataset has imbalanced classes (<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198410\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198410</a> ), other evaluation metrics (e.g. F1) would be suitable for this competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1088041,
          "author_name": "Yaroslav Isaienkov",
          "author_url": "",
          "post_date": "2020-11-23T09:08:38.500000",
          "content": "<p>Yes you are right<br>\nI showed this fact first in the notebook<br>\n<a href=\"https://www.kaggle.com/ihelon/cassava-leaf-disease-exploratory-data-analysis\" target=\"_blank\">https://www.kaggle.com/ihelon/cassava-leaf-disease-exploratory-data-analysis</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1088687,
          "author_name": "Thomas Brekk Unnvik",
          "author_url": "",
          "post_date": "2020-11-23T21:18:47.843000",
          "content": "<p>The imbalance in the dataset seems to reflect the proportions one would encounter in a practical scenario, based on domain research, as I discuss in the post linked below, so I think we can safely assume that the private test set is likely to be similar. <br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198584\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198584</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1089735,
      "author_name": "Aldi Topalli",
      "author_url": "",
      "post_date": "2020-11-24T18:21:06.080000",
      "content": "<p>Good point! I would like to see an answer to this!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1095306,
          "author_name": "Yaroslav Isaienkov",
          "author_url": "",
          "post_date": "2020-11-29T13:17:00.043000",
          "content": "<p><a href=\"https://www.kaggle.com/aldi1t\" target=\"_blank\">@aldi1t</a> I am waiting too)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1088438,
      "author_name": "Bayartsogt Yadamsuren",
      "author_url": "",
      "post_date": "2020-11-23T15:52:52.440000",
      "content": "<p>I think this is a good point! Seems like competition hosts should reply here.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1095305,
          "author_name": "Yaroslav Isaienkov",
          "author_url": "",
          "post_date": "2020-11-29T13:16:16.200000",
          "content": "<p><a href=\"https://www.kaggle.com/bayartsogtya\" target=\"_blank\">@bayartsogtya</a> agree</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1087066,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-11-22T10:03:21.363000",
      "content": "<p>The organizer seems too silent on this until now even it's a serious matter. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1091939,
          "author_name": "tropicbird",
          "author_url": "",
          "post_date": "2020-11-26T11:52:59.730000",
          "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> Could you please respond to this ongoing discussion? Thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1096875,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-11-30T20:51:02.307000",
          "content": "<p>The competition metric as defined will remain as-is.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1096876,
          "author_name": "Yaroslav Isaienkov",
          "author_url": "",
          "post_date": "2020-11-30T20:54:12.097000",
          "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> thanks for the clarification!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1085927,
      "author_name": "Kevin Joseph Scaria",
      "author_url": "",
      "post_date": "2020-11-21T09:29:29.340000",
      "content": "<p>Mathews Correlation Coefficient or F1-Score can handle the metrics for imbalanced class datasets.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1085753,
      "author_name": "khyeh",
      "author_url": "",
      "post_date": "2020-11-21T07:15:11.210000",
      "content": "<p>I would vote for categorical log loss or even micro-f1 score. Accuracy is not a stable indicator, especially when the validation\\testing set is small.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1084807,
      "author_name": "Kaushal Shah",
      "author_url": "",
      "post_date": "2020-11-20T12:08:16.370000",
      "content": "<p>Yes, as the data is highly imbalanced, only predicting class 3, in this case, would achieve decent accuracy!</p>\n<p>Edit: Here's the proof: <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198410\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198410</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1084603,
      "author_name": "Akshay Chavan",
      "author_url": "",
      "post_date": "2020-11-20T07:48:44.140000",
      "content": "<p>For this type of imbalanced data categorical loss is the best metric to evaluate the performance of the model. <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a> </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1084660,
          "author_name": "Yaroslav Isaienkov",
          "author_url": "",
          "post_date": "2020-11-20T09:00:29.887000",
          "content": "<p><a href=\"https://www.kaggle.com/akshaychavan123\" target=\"_blank\">@akshaychavan123</a> But to use this metric, they need to change the prediction format - probabilities instead of classes</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1084460,
      "author_name": "sheep",
      "author_url": "",
      "post_date": "2020-11-20T03:57:16.933000",
      "content": "<p>if the host want a category prediction rather than the probability, macro/micro fbeta maybe a good choice</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1084583,
          "author_name": "Yaroslav Isaienkov",
          "author_url": "",
          "post_date": "2020-11-20T07:27:22.080000",
          "content": "<p><a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a> yes, I think it's a good idea to leave category prediction but change the metric</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1150967,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-13T02:44:14.533000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1086886,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-22T06:37:32.433000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1084305": "Did you think about using another metric in this competition?\n\nFor example categorical log loss.\n\nBecause I think many players will be able to achieve a very high accuracy score, and for many, this indicator will be equal. And if your dataset has an imbalance, then the accuracy metric is not the best choice for evaluating models.",
    "1084737": "Agree, F1 or (even better) the Matthews Correlation Coefficient might be better alternatives",
    "1087385": "whatever the metric is, who gets the most accurate predictions wins. so I don't think it matters much",
    "1085513": "Agreed! Accuracy is a very bad choice for this dataset. However I've never seen a change of metrics in the middle of a competition. ",
    "1084365": "Categorical log loss metric holds good for imbalance classification @ihelon ",
    "1085132": "From my perspective, I am not sure how the competition would improve with a different metric.  Accuracy always increases as the number of correct predictions increase, and the competition is won by the submission with the highest number of correct predictions.",
    "1084459": "I think ROC-AUC score is best fit for imbalance dataset. You can calculate ROC-AUC via sklearn.metrics.roc_auc_score on the test set.",
    "1090333": "@ihelon Thanks for creating this post on metric Yaroslav . I enjoyed going through the discussion . ",
    "1087967": "Agree. Since at least the public dataset has imbalanced classes (https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198410 ), other evaluation metrics (e.g. F1) would be suitable for this competition.",
    "1089735": "Good point! I would like to see an answer to this!",
    "1088438": "I think this is a good point! Seems like competition hosts should reply here.",
    "1087066": "The organizer seems too silent on this until now even it's a serious matter. ",
    "1085927": "Mathews Correlation Coefficient or F1-Score can handle the metrics for imbalanced class datasets.",
    "1085753": "I would vote for categorical log loss or even micro-f1 score. Accuracy is not a stable indicator, especially when the validation\\testing set is small.",
    "1084807": "Yes, as the data is highly imbalanced, only predicting class 3, in this case, would achieve decent accuracy!\n\nEdit: Here's the proof: https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198410",
    "1084603": "For this type of imbalanced data categorical loss is the best metric to evaluate the performance of the model. @ihelon ",
    "1084460": "if the host want a category prediction rather than the probability, macro/micro fbeta maybe a good choice",
    "1150967": "",
    "1086886": ""
  }
}