{
  "id": 227237,
  "title": "CV-LB inconsistency",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/227237",
  "author_name": "Nikita Kuzmenkov",
  "post_date": "2021-03-19T12:34:33.250000",
  "votes": 22,
  "comment_count": 13,
  "views": 0,
  "content": "<h3>Hello!</h3>\n<p>As the competition just started, there are only a few public works so far, but the one issue I see across all the publicly available solutions (and my own) is severe CV-LB inconsistency. E.g.:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/pegasos\" target=\"_blank\">@pegasos</a> <strong><a href=\"https://www.kaggle.com/pegasos/plant2021-pytorch-lightning-starter-training#Training\" target=\"_blank\">solution</a></strong>: 1 OOF - ~0.9, LB - 0.652</li>\n<li><a href=\"https://www.kaggle.com/demetrypascal\" target=\"_blank\">@demetrypascal</a> <strong><a href=\"https://www.kaggle.com/demetrypascal/better-train-csv-format-keras-starter\" target=\"_blank\">solution</a></strong>: 1 OOF - ~0.4455, LB - 0.066</li>\n<li>my <strong><a href=\"https://www.kaggle.com/nickuzmenkov/pp2021-tpu-tf-training\" target=\"_blank\">solution</a></strong>: CV - 0.8272, LB - 0.543</li>\n</ol>\n<p>This makes me think of this is due to:</p>\n<ol>\n<li>Issues in submission scoring, e.g. maybe diseases are supposed to be listed in a specific order which is not explicitly stated in the <strong>Evaluation</strong> part?</li>\n<li>There's no information of how the <strong>F1 score</strong> is calculated, i.e. row-wise (<code>f1_score(true, pred, average='micro')</code> or column-/label-wise (<code>f1_score(true, pred, average='macro')</code>, and those scores yield pretty different results for me when training: 0.9XX (the former) and 0.8XX (the latter). So I think that should be explicitly re-confirmed.</li>\n<li>Noise: from a small previous experience with <strong>Cassava</strong> competition I can assume that at least the train set doesn't contain a lot of noise because here I'm able to reach 0.92 CV accuracy with a plain, straightforward pipeline, while in Cassava (as an example of a competition with a lot of noise) the best solution was only 0.915.</li>\n</ol>\n<p>So what do you think about this?</p>",
  "messages": [
    {
      "id": 1245068,
      "postDate": "2021-03-19T12:34:33.250Z",
      "content": "<h3>Hello!</h3>\n<p>As the competition just started, there are only a few public works so far, but the one issue I see across all the publicly available solutions (and my own) is severe CV-LB inconsistency. E.g.:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/pegasos\" target=\"_blank\">@pegasos</a> <strong><a href=\"https://www.kaggle.com/pegasos/plant2021-pytorch-lightning-starter-training#Training\" target=\"_blank\">solution</a></strong>: 1 OOF - ~0.9, LB - 0.652</li>\n<li><a href=\"https://www.kaggle.com/demetrypascal\" target=\"_blank\">@demetrypascal</a> <strong><a href=\"https://www.kaggle.com/demetrypascal/better-train-csv-format-keras-starter\" target=\"_blank\">solution</a></strong>: 1 OOF - ~0.4455, LB - 0.066</li>\n<li>my <strong><a href=\"https://www.kaggle.com/nickuzmenkov/pp2021-tpu-tf-training\" target=\"_blank\">solution</a></strong>: CV - 0.8272, LB - 0.543</li>\n</ol>\n<p>This makes me think of this is due to:</p>\n<ol>\n<li>Issues in submission scoring, e.g. maybe diseases are supposed to be listed in a specific order which is not explicitly stated in the <strong>Evaluation</strong> part?</li>\n<li>There's no information of how the <strong>F1 score</strong> is calculated, i.e. row-wise (<code>f1_score(true, pred, average='micro')</code> or column-/label-wise (<code>f1_score(true, pred, average='macro')</code>, and those scores yield pretty different results for me when training: 0.9XX (the former) and 0.8XX (the latter). So I think that should be explicitly re-confirmed.</li>\n<li>Noise: from a small previous experience with <strong>Cassava</strong> competition I can assume that at least the train set doesn't contain a lot of noise because here I'm able to reach 0.92 CV accuracy with a plain, straightforward pipeline, while in Cassava (as an example of a competition with a lot of noise) the best solution was only 0.915.</li>\n</ol>\n<p>So what do you think about this?</p>",
      "rawMarkdown": "### Hello!\n\nAs the competition just started, there are only a few public works so far, but the one issue I see across all the publicly available solutions (and my own) is severe CV-LB inconsistency. E.g.:\n\n1. @pegasos **[solution](https://www.kaggle.com/pegasos/plant2021-pytorch-lightning-starter-training#Training)**: 1 OOF - ~0.9, LB - 0.652\n2. @demetrypascal **[solution](https://www.kaggle.com/demetrypascal/better-train-csv-format-keras-starter)**: 1 OOF - ~0.4455, LB - 0.066\n3. my **[solution](https://www.kaggle.com/nickuzmenkov/pp2021-tpu-tf-training)**: CV - 0.8272, LB - 0.543\n\nThis makes me think of this is due to:\n1. Issues in submission scoring, e.g. maybe diseases are supposed to be listed in a specific order which is not explicitly stated in the **Evaluation** part?\n2. There's no information of how the **F1 score** is calculated, i.e. row-wise (`f1_score(true, pred, average='micro')` or column-/label-wise (`f1_score(true, pred, average='macro')`, and those scores yield pretty different results for me when training: 0.9XX (the former) and 0.8XX (the latter). So I think that should be explicitly re-confirmed.\n3. Noise: from a small previous experience with **Cassava** competition I can assume that at least the train set doesn't contain a lot of noise because here I'm able to reach 0.92 CV accuracy with a plain, straightforward pipeline, while in Cassava (as an example of a competition with a lot of noise) the best solution was only 0.915.\n\nSo what do you think about this?",
      "votes": 21
    },
    {
      "id": 1246638,
      "postDate": "2021-03-21T00:18:59.263Z",
      "content": "<p><a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a> <a href=\"https://www.kaggle.com/maxneumann\" target=\"_blank\">@maxneumann</a> The order of classes in the prediction does not matter and it is beneficial to provide at least partially correct classification. Please find the explanation in <a href=\"https://www.kaggle.com/buinyi/understanding-the-evaluation-metric\" target=\"_blank\">this notebook</a>.</p>",
      "rawMarkdown": "@nickuzmenkov @maxneumann The order of classes in the prediction does not matter and it is beneficial to provide at least partially correct classification. Please find the explanation in [this notebook](https://www.kaggle.com/buinyi/understanding-the-evaluation-metric).",
      "votes": 5,
      "replies": [
        {
          "id": 1246950,
          "postDate": "2021-03-21T09:20:25.187Z",
          "content": "<p>Thanks for your work and clarification, <a href=\"https://www.kaggle.com/buinyi\" target=\"_blank\">@buinyi</a>. I've also quickly checked whether the order of classes matters by shuffling my submission and got the same score. I know that's obvious, but we have to check everything until the reason for inconsistency becomes clear 😀</p>",
          "rawMarkdown": "Thanks for your work and clarification, @buinyi. I've also quickly checked whether the order of classes matters by shuffling my submission and got the same score. I know that's obvious, but we have to check everything until the reason for inconsistency becomes clear 😀"
        }
      ]
    },
    {
      "id": 1245660,
      "postDate": "2021-03-20T02:20:03.293Z",
      "content": "<p>I agree, seeing a similar CV score but getting 0.460 on the LB? Quite a large discrepancy between the two scores.. </p>\n<ul>\n<li>NOTE: last year the labels to predict were <code>[ 'healthy' , 'multiple_diseases' , 'rust' , 'scab' ]</code></li>\n</ul>",
      "rawMarkdown": "I agree, seeing a similar CV score but getting 0.460 on the LB? Quite a large discrepancy between the two scores.. \n\n- NOTE: last year the labels to predict were `[ 'healthy' , 'multiple_diseases' , 'rust' , 'scab' ]`",
      "votes": 3
    },
    {
      "id": 1253137,
      "postDate": "2021-03-26T11:52:54.903Z",
      "content": "<p>I've experimented a bit with a metric for this competition. There are a lot of ways to calculate metric on a cross-validation dataset depending on the way you consider the problem - as multilabel or multiclass classification. With multiclass approach and macro F1 score as it was used in many previous competitions I was able to get CV &gt;0.92 and LB &lt;0.6. Multilabel approach looks more reasonable here and it was mentioned in <a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226327\" target=\"_blank\">https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226327</a> that it is multilabel problem. Multilabel model results in a little bit higher scores on LB comparing to multiclass, macro F1 is somewhere near 0.9 on CV and F1 with average='sample' is closer to LB and equals to ~0.72. Another strange thing I found that in my models loss drops much lower comparing to kernels even on the first epoch.</p>",
      "rawMarkdown": "I've experimented a bit with a metric for this competition. There are a lot of ways to calculate metric on a cross-validation dataset depending on the way you consider the problem - as multilabel or multiclass classification. With multiclass approach and macro F1 score as it was used in many previous competitions I was able to get CV >0.92 and LB <0.6. Multilabel approach looks more reasonable here and it was mentioned in [https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226327](https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226327) that it is multilabel problem. Multilabel model results in a little bit higher scores on LB comparing to multiclass, macro F1 is somewhere near 0.9 on CV and F1 with average='sample' is closer to LB and equals to ~0.72. Another strange thing I found that in my models loss drops much lower comparing to kernels even on the first epoch.",
      "votes": 1
    },
    {
      "id": 1245500,
      "postDate": "2021-03-19T20:58:45.860Z",
      "content": "<p>column-wise macro F1 would probably a better choice here, given the biased dataset. But definitely need clarification on the metric.</p>",
      "rawMarkdown": "column-wise macro F1 would probably a better choice here, given the biased dataset. But definitely need clarification on the metric.",
      "votes": 1,
      "replies": [
        {
          "id": 1247009,
          "postDate": "2021-03-21T10:11:59.810Z",
          "content": "<p>So the tensorflow implementation is this right ? Correct me if I am wrong.</p>\n<pre><code>tfa.metrics.F1Score(\n    num_classes = NUM_CLASSES,\n    average = \"macro\",\n    name = 'f1_score',\n)\n</code></pre>",
          "rawMarkdown": "So the tensorflow implementation is this right ? Correct me if I am wrong.\n```\ntfa.metrics.F1Score(\n    num_classes = NUM_CLASSES,\n    average = \"macro\",\n    name = 'f1_score',\n)\n```"
        },
        {
          "id": 1247089,
          "postDate": "2021-03-21T12:04:25.237Z",
          "content": "<p>Yes, it's correct</p>",
          "rawMarkdown": "Yes, it's correct",
          "votes": 1
        }
      ]
    },
    {
      "id": 1250570,
      "postDate": "2021-03-24T06:04:53.300Z",
      "content": "<p><a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a> I got upto 91 F1 score while using Efficientnet-b3 with img res of 512 but while submitting I just got 0.278. This was 5-fold CV score.</p>",
      "rawMarkdown": "@nickuzmenkov I got upto 91 F1 score while using Efficientnet-b3 with img res of 512 but while submitting I just got 0.278. This was 5-fold CV score.",
      "votes": 2
    },
    {
      "id": 1245593,
      "postDate": "2021-03-19T23:44:45.670Z",
      "content": "<p>I was also wondering about how the score is calculated. Clarification would be great!<br>\nI've looked through some images and there are still label errors but way less than in cassava </p>",
      "rawMarkdown": "I was also wondering about how the score is calculated. Clarification would be great!\nI've looked through some images and there are still label errors but way less than in cassava ",
      "votes": 2
    },
    {
      "id": 1253681,
      "postDate": "2021-03-26T23:56:50.903Z",
      "content": "<p>Although my LB score is greater than 0.7, I don't understand the mean F1 score. Maybe my Lb is just random result.</p>",
      "rawMarkdown": "Although my LB score is greater than 0.7, I don't understand the mean F1 score. Maybe my Lb is just random result."
    },
    {
      "id": 1246453,
      "postDate": "2021-03-20T18:46:45.587Z",
      "content": "<p>I'm pretty sure its row wise f1. So if a disease is missed or there is a disease to much it'll count as a wrong prediction and as you said maybe the order is important aswell. I guess column wise F1 would have been a lot \"better\" choice…<br>\nBy the way Nick I wrote you a DM over kaggle. Maybe you've seen it already but I just want to make sure it doesn't get lost in a spam filter ))</p>",
      "rawMarkdown": "I'm pretty sure its row wise f1. So if a disease is missed or there is a disease to much it'll count as a wrong prediction and as you said maybe the order is important aswell. I guess column wise F1 would have been a lot \"better\" choice...\nBy the way Nick I wrote you a DM over kaggle. Maybe you've seen it already but I just want to make sure it doesn't get lost in a spam filter ))",
      "replies": [
        {
          "id": 1247040,
          "postDate": "2021-03-21T10:45:56.280Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1258733,
      "postDate": "2021-03-31T21:00:09.803Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1246638,
      "author_name": "Igor Buinyi",
      "author_url": "",
      "post_date": "2021-03-21T00:18:59.263000",
      "content": "<p><a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a> <a href=\"https://www.kaggle.com/maxneumann\" target=\"_blank\">@maxneumann</a> The order of classes in the prediction does not matter and it is beneficial to provide at least partially correct classification. Please find the explanation in <a href=\"https://www.kaggle.com/buinyi/understanding-the-evaluation-metric\" target=\"_blank\">this notebook</a>.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1246950,
          "author_name": "Nikita Kuzmenkov",
          "author_url": "",
          "post_date": "2021-03-21T09:20:25.187000",
          "content": "<p>Thanks for your work and clarification, <a href=\"https://www.kaggle.com/buinyi\" target=\"_blank\">@buinyi</a>. I've also quickly checked whether the order of classes matters by shuffling my submission and got the same score. I know that's obvious, but we have to check everything until the reason for inconsistency becomes clear 😀</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1245660,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2021-03-20T02:20:03.293000",
      "content": "<p>I agree, seeing a similar CV score but getting 0.460 on the LB? Quite a large discrepancy between the two scores.. </p>\n<ul>\n<li>NOTE: last year the labels to predict were <code>[ 'healthy' , 'multiple_diseases' , 'rust' , 'scab' ]</code></li>\n</ul>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1253137,
      "author_name": "Oleg Panichev",
      "author_url": "",
      "post_date": "2021-03-26T11:52:54.903000",
      "content": "<p>I've experimented a bit with a metric for this competition. There are a lot of ways to calculate metric on a cross-validation dataset depending on the way you consider the problem - as multilabel or multiclass classification. With multiclass approach and macro F1 score as it was used in many previous competitions I was able to get CV &gt;0.92 and LB &lt;0.6. Multilabel approach looks more reasonable here and it was mentioned in <a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226327\" target=\"_blank\">https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226327</a> that it is multilabel problem. Multilabel model results in a little bit higher scores on LB comparing to multiclass, macro F1 is somewhere near 0.9 on CV and F1 with average='sample' is closer to LB and equals to ~0.72. Another strange thing I found that in my models loss drops much lower comparing to kernels even on the first epoch.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1245500,
      "author_name": "Jaideep Murkute",
      "author_url": "",
      "post_date": "2021-03-19T20:58:45.860000",
      "content": "<p>column-wise macro F1 would probably a better choice here, given the biased dataset. But definitely need clarification on the metric.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1247009,
          "author_name": "Shanmukh",
          "author_url": "",
          "post_date": "2021-03-21T10:11:59.810000",
          "content": "<p>So the tensorflow implementation is this right ? Correct me if I am wrong.</p>\n<pre><code>tfa.metrics.F1Score(\n    num_classes = NUM_CLASSES,\n    average = \"macro\",\n    name = 'f1_score',\n)\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1247089,
          "author_name": "Nikita Kuzmenkov",
          "author_url": "",
          "post_date": "2021-03-21T12:04:25.237000",
          "content": "<p>Yes, it's correct</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1250570,
      "author_name": "TEnsorSAge",
      "author_url": "",
      "post_date": "2021-03-24T06:04:53.300000",
      "content": "<p><a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a> I got upto 91 F1 score while using Efficientnet-b3 with img res of 512 but while submitting I just got 0.278. This was 5-fold CV score.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1245593,
      "author_name": "Max Neumann",
      "author_url": "",
      "post_date": "2021-03-19T23:44:45.670000",
      "content": "<p>I was also wondering about how the score is calculated. Clarification would be great!<br>\nI've looked through some images and there are still label errors but way less than in cassava </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1253681,
      "author_name": "hanjin",
      "author_url": "",
      "post_date": "2021-03-26T23:56:50.903000",
      "content": "<p>Although my LB score is greater than 0.7, I don't understand the mean F1 score. Maybe my Lb is just random result.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1246453,
      "author_name": "Max Neumann",
      "author_url": "",
      "post_date": "2021-03-20T18:46:45.587000",
      "content": "<p>I'm pretty sure its row wise f1. So if a disease is missed or there is a disease to much it'll count as a wrong prediction and as you said maybe the order is important aswell. I guess column wise F1 would have been a lot \"better\" choice…<br>\nBy the way Nick I wrote you a DM over kaggle. Maybe you've seen it already but I just want to make sure it doesn't get lost in a spam filter ))</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1247040,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-21T10:45:56.280000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1258733,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-31T21:00:09.803000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1245068": "### Hello!\n\nAs the competition just started, there are only a few public works so far, but the one issue I see across all the publicly available solutions (and my own) is severe CV-LB inconsistency. E.g.:\n\n1. @pegasos **[solution](https://www.kaggle.com/pegasos/plant2021-pytorch-lightning-starter-training#Training)**: 1 OOF - ~0.9, LB - 0.652\n2. @demetrypascal **[solution](https://www.kaggle.com/demetrypascal/better-train-csv-format-keras-starter)**: 1 OOF - ~0.4455, LB - 0.066\n3. my **[solution](https://www.kaggle.com/nickuzmenkov/pp2021-tpu-tf-training)**: CV - 0.8272, LB - 0.543\n\nThis makes me think of this is due to:\n1. Issues in submission scoring, e.g. maybe diseases are supposed to be listed in a specific order which is not explicitly stated in the **Evaluation** part?\n2. There's no information of how the **F1 score** is calculated, i.e. row-wise (`f1_score(true, pred, average='micro')` or column-/label-wise (`f1_score(true, pred, average='macro')`, and those scores yield pretty different results for me when training: 0.9XX (the former) and 0.8XX (the latter). So I think that should be explicitly re-confirmed.\n3. Noise: from a small previous experience with **Cassava** competition I can assume that at least the train set doesn't contain a lot of noise because here I'm able to reach 0.92 CV accuracy with a plain, straightforward pipeline, while in Cassava (as an example of a competition with a lot of noise) the best solution was only 0.915.\n\nSo what do you think about this?",
    "1246638": "@nickuzmenkov @maxneumann The order of classes in the prediction does not matter and it is beneficial to provide at least partially correct classification. Please find the explanation in [this notebook](https://www.kaggle.com/buinyi/understanding-the-evaluation-metric).",
    "1245660": "I agree, seeing a similar CV score but getting 0.460 on the LB? Quite a large discrepancy between the two scores.. \n\n- NOTE: last year the labels to predict were `[ 'healthy' , 'multiple_diseases' , 'rust' , 'scab' ]`",
    "1253137": "I've experimented a bit with a metric for this competition. There are a lot of ways to calculate metric on a cross-validation dataset depending on the way you consider the problem - as multilabel or multiclass classification. With multiclass approach and macro F1 score as it was used in many previous competitions I was able to get CV >0.92 and LB <0.6. Multilabel approach looks more reasonable here and it was mentioned in [https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226327](https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226327) that it is multilabel problem. Multilabel model results in a little bit higher scores on LB comparing to multiclass, macro F1 is somewhere near 0.9 on CV and F1 with average='sample' is closer to LB and equals to ~0.72. Another strange thing I found that in my models loss drops much lower comparing to kernels even on the first epoch.",
    "1245500": "column-wise macro F1 would probably a better choice here, given the biased dataset. But definitely need clarification on the metric.",
    "1250570": "@nickuzmenkov I got upto 91 F1 score while using Efficientnet-b3 with img res of 512 but while submitting I just got 0.278. This was 5-fold CV score.",
    "1245593": "I was also wondering about how the score is calculated. Clarification would be great!\nI've looked through some images and there are still label errors but way less than in cassava ",
    "1253681": "Although my LB score is greater than 0.7, I don't understand the mean F1 score. Maybe my Lb is just random result.",
    "1246453": "I'm pretty sure its row wise f1. So if a disease is missed or there is a disease to much it'll count as a wrong prediction and as you said maybe the order is important aswell. I guess column wise F1 would have been a lot \"better\" choice...\nBy the way Nick I wrote you a DM over kaggle. Maybe you've seen it already but I just want to make sure it doesn't get lost in a spam filter ))",
    "1258733": ""
  }
}