{
  "id": 234485,
  "title": "[Help Needed] Model Predicting Wrong Labels with High Confidence",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/234485",
  "author_name": "",
  "post_date": "2021-04-24T15:06:45.529512900Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>During training F score is improving with each epoch, and also the loss is decreasing. But when I am checking the model performance on a different notebook with validation dataset, model is predicting 4 labels for each image with more than 0.9 value.</p>\n<p>I haven't worked with multi-label classification problem before, so there is a chance I have made mistakes in one of the notebooks. I will really appreciate if someone can review my notebooks, and provide feedback.</p>\n<p><strong>Training Notebook:</strong> <a href=\"https://www.kaggle.com/jabertuhin/multi-label-classification-training-with-pl\" target=\"_blank\">Multi-label Classification Training with PL</a></p>\n<p><strong>Inference Notebook:</strong> <a href=\"https://www.kaggle.com/jabertuhin/debug-inference-notebook\" target=\"_blank\">[Debug] Inference Notebook</a></p>\n<p>Thanks in Advance.</p>",
  "messages": [
    {
      "id": "1283084",
      "postDate": "04/24/2021 15:06:45",
      "content": "<p>During training F score is improving with each epoch, and also the loss is decreasing. But when I am checking the model performance on a different notebook with validation dataset, model is predicting 4 labels for each image with more than 0.9 value.</p>\n<p>I haven't worked with multi-label classification problem before, so there is a chance I have made mistakes in one of the notebooks. I will really appreciate if someone can review my notebooks, and provide feedback.</p>\n<p><strong>Training Notebook:</strong> <a href=\"https://www.kaggle.com/jabertuhin/multi-label-classification-training-with-pl\" target=\"_blank\">Multi-label Classification Training with PL</a></p>\n<p><strong>Inference Notebook:</strong> <a href=\"https://www.kaggle.com/jabertuhin/debug-inference-notebook\" target=\"_blank\">[Debug] Inference Notebook</a></p>\n<p>Thanks in Advance.</p>",
      "rawMarkdown": "During training F score is improving with each epoch, and also the loss is decreasing. But when I am checking the model performance on a different notebook with validation dataset, model is predicting 4 labels for each image with more than 0.9 value.\n\nI haven't worked with multi-label classification problem before, so there is a chance I have made mistakes in one of the notebooks. I will really appreciate if someone can review my notebooks, and provide feedback.\n\n**Training Notebook:** [Multi-label Classification Training with PL](https://www.kaggle.com/jabertuhin/multi-label-classification-training-with-pl)\n\n**Inference Notebook:** [[Debug] Inference Notebook](https://www.kaggle.com/jabertuhin/debug-inference-notebook)\n\nThanks in Advance.",
      "votes": null
    },
    {
      "id": "1291990",
      "postDate": "05/03/2021 14:04:50",
      "content": "<p>Are those 4 labels present in highest amounts in the augmented dataset ?</p>",
      "rawMarkdown": "Are those 4 labels present in highest amounts in the augmented dataset ?",
      "votes": null
    },
    {
      "id": "1292233",
      "postDate": "05/03/2021 19:14:31",
      "content": "<p>I am not using any kind of augmentation,  just using image resizing, normalization and tensor conversion.</p>",
      "rawMarkdown": "I am not using any kind of augmentation,  just using image resizing, normalization and tensor conversion.",
      "votes": null
    },
    {
      "id": "1292555",
      "postDate": "05/04/2021 05:08:36",
      "content": "<p>Are they some particular labels if yes which are they? </p>",
      "rawMarkdown": "Are they some particular labels if yes which are they?",
      "votes": null
    },
    {
      "id": "1296859",
      "postDate": "05/07/2021 14:33:57",
      "content": "<p>It's predicting - ['scab', 'healthy',  'rust',  'complex'] labels for every images. You can find the outputs in the inference notebook.</p>",
      "rawMarkdown": "It's predicting - ['scab', 'healthy',  'rust',  'complex'] labels for every images. You can find the outputs in the inference notebook.",
      "votes": null
    },
    {
      "id": "1297502",
      "postDate": "05/08/2021 04:56:12",
      "content": "<p>Most likely there is an imbalance problem , let me explain :-</p>\n<p>You can see the distribution of the data in this notebook (<a href=\"https://www.kaggle.com/arnabs007/apple-leaf-diseases-with-inceptionresnetv2-keras\" target=\"_blank\">https://www.kaggle.com/arnabs007/apple-leaf-diseases-with-inceptionresnetv2-keras</a>)</p>\n<p>The labels which you are saying is present in quite high percentage though that may not be the case but there is a probability for this happening so i will say add accuracy as your new metric then write a code to sum up the counts of  ['scab', 'healthy', 'rust', 'complex'] then divide it by the sum of all labels </p>\n<p>Its just the percentage of these labels in dataset scaled between 0 and 1</p>\n<p>If you see that your acc in the last or final epoch is quite similar to the reffered value it's most likely an imbalance problem . I would recommend keeping an average steps_per_epoch paramter not too small so that the estimates are accurate . If you dont have that much computational power then you can train on a small steps_per_epoch paramter and run an exponentially weighted average and chose the final average of acc</p>",
      "rawMarkdown": "Most likely there is an imbalance problem , let me explain :-\n\nYou can see the distribution of the data in this notebook (https://www.kaggle.com/arnabs007/apple-leaf-diseases-with-inceptionresnetv2-keras)\n\nThe labels which you are saying is present in quite high percentage though that may not be the case but there is a probability for this happening so i will say add accuracy as your new metric then write a code to sum up the counts of  ['scab', 'healthy', 'rust', 'complex'] then divide it by the sum of all labels \n\nIts just the percentage of these labels in dataset scaled between 0 and 1\n\nIf you see that your acc in the last or final epoch is quite similar to the reffered value it's most likely an imbalance problem . I would recommend keeping an average steps_per_epoch paramter not too small so that the estimates are accurate . If you dont have that much computational power then you can train on a small steps_per_epoch paramter and run an exponentially weighted average and chose the final average of acc",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1291990,
      "author_name": "swaralipibose",
      "author_url": "",
      "post_date": "05/03/2021 14:04:50",
      "content": "<p>Are those 4 labels present in highest amounts in the augmented dataset ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1292233,
          "author_name": "jabertuhin",
          "author_url": "",
          "post_date": "05/03/2021 19:14:31",
          "content": "<p>I am not using any kind of augmentation,  just using image resizing, normalization and tensor conversion.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1292555,
          "author_name": "swaralipibose",
          "author_url": "",
          "post_date": "05/04/2021 05:08:36",
          "content": "<p>Are they some particular labels if yes which are they? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1296859,
          "author_name": "jabertuhin",
          "author_url": "",
          "post_date": "05/07/2021 14:33:57",
          "content": "<p>It's predicting - ['scab', 'healthy',  'rust',  'complex'] labels for every images. You can find the outputs in the inference notebook.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1297502,
          "author_name": "swaralipibose",
          "author_url": "",
          "post_date": "05/08/2021 04:56:12",
          "content": "<p>Most likely there is an imbalance problem , let me explain :-</p>\n<p>You can see the distribution of the data in this notebook (<a href=\"https://www.kaggle.com/arnabs007/apple-leaf-diseases-with-inceptionresnetv2-keras\" target=\"_blank\">https://www.kaggle.com/arnabs007/apple-leaf-diseases-with-inceptionresnetv2-keras</a>)</p>\n<p>The labels which you are saying is present in quite high percentage though that may not be the case but there is a probability for this happening so i will say add accuracy as your new metric then write a code to sum up the counts of  ['scab', 'healthy', 'rust', 'complex'] then divide it by the sum of all labels </p>\n<p>Its just the percentage of these labels in dataset scaled between 0 and 1</p>\n<p>If you see that your acc in the last or final epoch is quite similar to the reffered value it's most likely an imbalance problem . I would recommend keeping an average steps_per_epoch paramter not too small so that the estimates are accurate . If you dont have that much computational power then you can train on a small steps_per_epoch paramter and run an exponentially weighted average and chose the final average of acc</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1283084": "During training F score is improving with each epoch, and also the loss is decreasing. But when I am checking the model performance on a different notebook with validation dataset, model is predicting 4 labels for each image with more than 0.9 value.\n\nI haven't worked with multi-label classification problem before, so there is a chance I have made mistakes in one of the notebooks. I will really appreciate if someone can review my notebooks, and provide feedback.\n\n**Training Notebook:** [Multi-label Classification Training with PL](https://www.kaggle.com/jabertuhin/multi-label-classification-training-with-pl)\n\n**Inference Notebook:** [[Debug] Inference Notebook](https://www.kaggle.com/jabertuhin/debug-inference-notebook)\n\nThanks in Advance.",
    "1291990": "Are those 4 labels present in highest amounts in the augmented dataset ?",
    "1292233": "I am not using any kind of augmentation,  just using image resizing, normalization and tensor conversion.",
    "1292555": "Are they some particular labels if yes which are they?",
    "1296859": "It's predicting - ['scab', 'healthy',  'rust',  'complex'] labels for every images. You can find the outputs in the inference notebook.",
    "1297502": "Most likely there is an imbalance problem , let me explain :-\n\nYou can see the distribution of the data in this notebook (https://www.kaggle.com/arnabs007/apple-leaf-diseases-with-inceptionresnetv2-keras)\n\nThe labels which you are saying is present in quite high percentage though that may not be the case but there is a probability for this happening so i will say add accuracy as your new metric then write a code to sum up the counts of  ['scab', 'healthy', 'rust', 'complex'] then divide it by the sum of all labels \n\nIts just the percentage of these labels in dataset scaled between 0 and 1\n\nIf you see that your acc in the last or final epoch is quite similar to the reffered value it's most likely an imbalance problem . I would recommend keeping an average steps_per_epoch paramter not too small so that the estimates are accurate . If you dont have that much computational power then you can train on a small steps_per_epoch paramter and run an exponentially weighted average and chose the final average of acc"
  },
  "source": "meta"
}