{
  "id": 200722,
  "title": "Cassava Dataset V2",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/200722",
  "author_name": "",
  "post_date": "2020-12-01T15:32:31.111428700Z",
  "votes": 54,
  "comment_count": 31,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/ammarali32/cassava-datasetv2\" target=\"_blank\">https://www.kaggle.com/ammarali32/cassava-datasetv2</a><br>\nBy this link, you can find a dataset that includes both Cassava datasets. Also, the images were cropped using YOLO detector to detect one leaf in an image (the one with the highest confidence).<br>\nUnfortunately, the detector is not highly accurate because I trained yolo-tinyv3 with 1000 images (drawing boxes is really boring). And it was tiny related to GPU issues.<br>\nhope it will be useful. <br>\nps1. the data published by the competition is really noisy healthy class includes at least 1000 images that aren't. also mixing between class 1 and 0.</p>\n<ul>\n<li>a lot of duplicated images.<br>\nTo save your time Cleaning the data didn't increase the accuracy I already tried.    </li>\n</ul>",
  "messages": [
    {
      "id": "1098341",
      "postDate": "12/01/2020 15:32:31",
      "content": "<p><a href=\"https://www.kaggle.com/ammarali32/cassava-datasetv2\" target=\"_blank\">https://www.kaggle.com/ammarali32/cassava-datasetv2</a><br>\nBy this link, you can find a dataset that includes both Cassava datasets. Also, the images were cropped using YOLO detector to detect one leaf in an image (the one with the highest confidence).<br>\nUnfortunately, the detector is not highly accurate because I trained yolo-tinyv3 with 1000 images (drawing boxes is really boring). And it was tiny related to GPU issues.<br>\nhope it will be useful. <br>\nps1. the data published by the competition is really noisy healthy class includes at least 1000 images that aren't. also mixing between class 1 and 0.</p>\n<ul>\n<li>a lot of duplicated images.<br>\nTo save your time Cleaning the data didn't increase the accuracy I already tried.    </li>\n</ul>",
      "rawMarkdown": "https://www.kaggle.com/ammarali32/cassava-datasetv2\nBy this link, you can find a dataset that includes both Cassava datasets. Also, the images were cropped using YOLO detector to detect one leaf in an image (the one with the highest confidence).\nUnfortunately, the detector is not highly accurate because I trained yolo-tinyv3 with 1000 images (drawing boxes is really boring). And it was tiny related to GPU issues.\nhope it will be useful. \nps1. the data published by the competition is really noisy healthy class includes at least 1000 images that aren't. also mixing between class 1 and 0.\n+ a lot of duplicated images.\nTo save your time Cleaning the data didn't increase the accuracy I already tried.",
      "votes": null
    },
    {
      "id": "1098362",
      "postDate": "12/01/2020 15:50:51",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> for sharing … this is helpful for training!! </p>",
      "rawMarkdown": "Thanks @ammarali32 for sharing ... this is helpful for training!!",
      "votes": null
    },
    {
      "id": "1098392",
      "postDate": "12/01/2020 16:06:01",
      "content": "<p>a big Thanks ; <br>\n<code>ps1. the data published by the competition is really noisy healthy class includes at least 1000 images that aren't. also mixing between class 1 and 0.</code><br>\nSorry I couldn't get what are you trying to say</p>",
      "rawMarkdown": "a big Thanks ; \n`ps1. the data published by the competition is really noisy healthy class includes at least 1000 images that aren't. also mixing between class 1 and 0.`\nSorry I couldn't get what are you trying to say",
      "votes": null
    },
    {
      "id": "1098408",
      "postDate": "12/01/2020 16:12:49",
      "content": "<p>Welcome ))</p>",
      "rawMarkdown": "Welcome ))",
      "votes": null
    },
    {
      "id": "1098414",
      "postDate": "12/01/2020 16:15:03",
      "content": "<p>Put all the images with label 4 in a folder and take a look at the data. It should contain healthy images but actually, it didn't.</p>",
      "rawMarkdown": "Put all the images with label 4 in a folder and take a look at the data. It should contain healthy images but actually, it didn't.",
      "votes": null
    },
    {
      "id": "1098420",
      "postDate": "12/01/2020 16:23:55",
      "content": "<p>Thanks for the clarification; Does using the new data improve your public_LB</p>",
      "rawMarkdown": "Thanks for the clarification; Does using the new data improve your public_LB",
      "votes": null
    },
    {
      "id": "1098444",
      "postDate": "12/01/2020 16:43:04",
      "content": "<p>No, unfortunately, I got a CV = 90 and a 90.3 LB while 88.9 CV got 90.1 it is really strange. I think the testing data is noisy as well   </p>",
      "rawMarkdown": "No, unfortunately, I got a CV = 90 and a 90.3 LB while 88.9 CV got 90.1 it is really strange. I think the testing data is noisy as well",
      "votes": null
    },
    {
      "id": "1099180",
      "postDate": "12/02/2020 06:35:33",
      "content": "<p>This saves a lot of work. Thanks.</p>",
      "rawMarkdown": "This saves a lot of work. Thanks.",
      "votes": null
    },
    {
      "id": "1099723",
      "postDate": "12/02/2020 15:07:09",
      "content": "<p>Welcome ))</p>",
      "rawMarkdown": "Welcome ))",
      "votes": null
    },
    {
      "id": "1100362",
      "postDate": "12/03/2020 03:20:33",
      "content": "<p>Thanks !! ))</p>",
      "rawMarkdown": "Thanks !! ))",
      "votes": null
    },
    {
      "id": "1100546",
      "postDate": "12/03/2020 06:33:57",
      "content": "<p><a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> have you also uploaded the annotation files?</p>",
      "rawMarkdown": "ammarali32 have you also uploaded the annotation files?",
      "votes": null
    },
    {
      "id": "1100608",
      "postDate": "12/03/2020 07:29:42",
      "content": "<p>if you meant the annotation files for the detection then Not yet. but maybe I will </p>",
      "rawMarkdown": "if you meant the annotation files for the detection then Not yet. but maybe I will",
      "votes": null
    },
    {
      "id": "1100634",
      "postDate": "12/03/2020 07:55:36",
      "content": "<p>Thanks for the info.</p>",
      "rawMarkdown": "Thanks for the info.",
      "votes": null
    },
    {
      "id": "1100960",
      "postDate": "12/03/2020 14:09:48",
      "content": "<p>Thank you ammarali, I really appreciate it. I am new to this field and I have two questions, <br>\n1)How did you clean the image data (not asking for code but concepts)<br>\n2)I tried submitting my code 3 times, and all 3 times I got an error message even though it shows that submission was successful. How can I fix this?</p>",
      "rawMarkdown": "Thank you ammarali, I really appreciate it. I am new to this field and I have two questions, \n1)How did you clean the image data (not asking for code but concepts)\n2)I tried submitting my code 3 times, and all 3 times I got an error message even though it shows that submission was successful. How can I fix this?",
      "votes": null
    },
    {
      "id": "1100989",
      "postDate": "12/03/2020 14:34:42",
      "content": "<p>Hello, you are welcome, First of all, don't clean the data cleaning the data will give you worse results because it seems that the testing data is also noisy. For submission issues please give me info about the error message or you can use any baseline for submission. Many competitors shared a template for submissions just change the model architecture load your model and that is all.</p>",
      "rawMarkdown": "Hello, you are welcome, First of all, don't clean the data cleaning the data will give you worse results because it seems that the testing data is also noisy. For submission issues please give me info about the error message or you can use any baseline for submission. Many competitors shared a template for submissions just change the model architecture load your model and that is all.",
      "votes": null
    },
    {
      "id": "1101072",
      "postDate": "12/03/2020 16:00:08",
      "content": "<p>There is no error message. What is baseline submission? My submission is nothing special tbh, tbh it's one of the first competitions I'm taking seriously. I use model.fit_generator(training and validation data) plot accuracies, and then use a function that loads the test image from path, processes it (resizes it) and then predicts it. This prediction is then sent to my_submission.csv with proper image name and label. My submissions are successful, but it doesn't score my submission. \"Submission Scoring Error\"</p>",
      "rawMarkdown": "There is no error message. What is baseline submission? My submission is nothing special tbh, tbh it's one of the first competitions I'm taking seriously. I use model.fit_generator(training and validation data) plot accuracies, and then use a function that loads the test image from path, processes it (resizes it) and then predicts it. This prediction is then sent to my_submission.csv with proper image name and label. My submissions are successful, but it doesn't score my submission. \"Submission Scoring Error\"",
      "votes": null
    },
    {
      "id": "1101085",
      "postDate": "12/03/2020 16:14:32",
      "content": "<p>You can take this as an example <a href=\"https://www.kaggle.com/manojprabhaakr/leaf-classification-resnext-50-32-4d\" target=\"_blank\">https://www.kaggle.com/manojprabhaakr/leaf-classification-resnext-50-32-4d</a><br>\nplease make sure that you are reading the testing data from the test data folder. The data is hidden and will be accessible only after the submission (when you submit your solution they will run it on these hidden data and score it). if you are reading the data fine then check that the output of your model is an integer between 0 and 4, not a float number or out of these bounds.</p>",
      "rawMarkdown": "You can take this as an example https://www.kaggle.com/manojprabhaakr/leaf-classification-resnext-50-32-4d\nplease make sure that you are reading the testing data from the test data folder. The data is hidden and will be accessible only after the submission (when you submit your solution they will run it on these hidden data and score it). if you are reading the data fine then check that the output of your model is an integer between 0 and 4, not a float number or out of these bounds.",
      "votes": null
    },
    {
      "id": "1101775",
      "postDate": "12/04/2020 08:35:52",
      "content": "<p>Yes Ammarali, my output is an integer between 0 and 4, but, I am reading the specific csv file (from the folder) and predicting the output. Is there anything else too that needs to be done?</p>",
      "rawMarkdown": "Yes Ammarali, my output is an integer between 0 and 4, but, I am reading the specific csv file (from the folder) and predicting the output. Is there anything else too that needs to be done?",
      "votes": null
    },
    {
      "id": "1102880",
      "postDate": "12/05/2020 12:39:22",
      "content": "<p>Welcome ))</p>",
      "rawMarkdown": "Welcome ))",
      "votes": null
    },
    {
      "id": "1103133",
      "postDate": "12/05/2020 16:42:58",
      "content": "<p>Great work and thanks! </p>\n<p>In such a scenario where test set may also be mislabeled, what can we do?</p>",
      "rawMarkdown": "Great work and thanks! \n\nIn such a scenario where test set may also be mislabeled, what can we do?",
      "votes": null
    },
    {
      "id": "1103136",
      "postDate": "12/05/2020 16:45:11",
      "content": "<p>hi <a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> <br>\nI have removed the images I know to be wrong and not considered them in train</p>",
      "rawMarkdown": "hi @reighns \nI have removed the images I know to be wrong and not considered them in train",
      "votes": null
    },
    {
      "id": "1103146",
      "postDate": "12/05/2020 16:52:58",
      "content": "<p>Hi thanks))<br>\nAfter removing the mislabeled images I got a higher CV and lower LB. It is probably an indicator that the testing data also noisy. But after testing over all data that may differ. Anyway, I suggest keeping the noisy images. but it is up to you. Good Luck ))</p>",
      "rawMarkdown": "Hi thanks))\nAfter removing the mislabeled images I got a higher CV and lower LB. It is probably an indicator that the testing data also noisy. But after testing over all data that may differ. Anyway, I suggest keeping the noisy images. but it is up to you. Good Luck ))",
      "votes": null
    },
    {
      "id": "1103219",
      "postDate": "12/05/2020 18:23:07",
      "content": "<p><a href=\"https://www.kaggle.com/kmldas\" target=\"_blank\">@kmldas</a> have you removed it manually and how can you tell which ones are mislabeled?</p>",
      "rawMarkdown": "kmldas have you removed it manually and how can you tell which ones are mislabeled?",
      "votes": null
    },
    {
      "id": "1103223",
      "postDate": "12/05/2020 18:29:53",
      "content": "<p><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200201\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200201</a> this dataset have both years data with duplicate images removed; you can use your yolo detector here and get a dataset without duplicates. Good for everyone I guess.</p>",
      "rawMarkdown": "https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200201 this dataset have both years data with duplicate images removed; you can use your yolo detector here and get a dataset without duplicates. Good for everyone I guess.",
      "votes": null
    },
    {
      "id": "1103269",
      "postDate": "12/05/2020 19:18:20",
      "content": "<p>thanks ))</p>",
      "rawMarkdown": "thanks ))",
      "votes": null
    },
    {
      "id": "1103611",
      "postDate": "12/06/2020 04:34:02",
      "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> <br>\nYes, I have removed them manually… those that seemed mislabed from discussions or when I did random batch image show, and felt were wrong I removed the ids from the train.<br>\nI may be wrong but felt better to have a cleaner train set</p>",
      "rawMarkdown": "mrinath \nYes, I have removed them manually... those that seemed mislabed from discussions or when I did random batch image show, and felt were wrong I removed the ids from the train.\nI may be wrong but felt better to have a cleaner train set",
      "votes": null
    },
    {
      "id": "1103616",
      "postDate": "12/06/2020 04:39:21",
      "content": "<p><a href=\"https://www.kaggle.com/kmldas\" target=\"_blank\">@kmldas</a> any score improvements after cleaning?</p>",
      "rawMarkdown": "kmldas any score improvements after cleaning?",
      "votes": null
    },
    {
      "id": "1103627",
      "postDate": "12/06/2020 04:47:57",
      "content": "<p>My CV improved, the public LB/score not as much … but hoping it helps with pvt LB as CV is better! 🤘</p>",
      "rawMarkdown": "My CV improved, the public LB/score not as much ... but hoping it helps with pvt LB as CV is better! 🤘",
      "votes": null
    },
    {
      "id": "1115550",
      "postDate": "12/16/2020 11:07:38",
      "content": "<p>It did not improve my LB neither my CV but I noticed, that the test_images cannot be resized to <code>256x256</code> </p>\n<p>The submission fails.</p>",
      "rawMarkdown": "It did not improve my LB neither my CV but I noticed, that the test_images cannot be resized to `256x256` \n\nThe submission fails.",
      "votes": null
    },
    {
      "id": "1149694",
      "postDate": "01/12/2021 04:26:30",
      "content": "<p><a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> Have you shared your inference notebook using YOLO?</p>",
      "rawMarkdown": "ammarali32 Have you shared your inference notebook using YOLO?",
      "votes": null
    },
    {
      "id": "1179878",
      "postDate": "01/31/2021 23:13:28",
      "content": "<p><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215940\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215940</a></p>",
      "rawMarkdown": "https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215940",
      "votes": null
    },
    {
      "id": "1181756",
      "postDate": "02/02/2021 05:21:43",
      "content": "<p>I spent days to remove the mislabelled and duplicate data using DBSCAN to find out the accuracy(both cv and lb) actually decreased after doing it.  So it just does not work. Will try with just removing duplicate images to see which model works best with analysing the CV</p>",
      "rawMarkdown": "I spent days to remove the mislabelled and duplicate data using DBSCAN to find out the accuracy(both cv and lb) actually decreased after doing it.  So it just does not work. Will try with just removing duplicate images to see which model works best with analysing the CV",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1098362,
      "author_name": "kmldas",
      "author_url": "",
      "post_date": "12/01/2020 15:50:51",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> for sharing … this is helpful for training!! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1098408,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "12/01/2020 16:12:49",
          "content": "<p>Welcome ))</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1098392,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "12/01/2020 16:06:01",
      "content": "<p>a big Thanks ; <br>\n<code>ps1. the data published by the competition is really noisy healthy class includes at least 1000 images that aren't. also mixing between class 1 and 0.</code><br>\nSorry I couldn't get what are you trying to say</p>",
      "votes": null,
      "replies": [
        {
          "id": 1098414,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "12/01/2020 16:15:03",
          "content": "<p>Put all the images with label 4 in a folder and take a look at the data. It should contain healthy images but actually, it didn't.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1098420,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "12/01/2020 16:23:55",
          "content": "<p>Thanks for the clarification; Does using the new data improve your public_LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1098444,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "12/01/2020 16:43:04",
          "content": "<p>No, unfortunately, I got a CV = 90 and a 90.3 LB while 88.9 CV got 90.1 it is really strange. I think the testing data is noisy as well   </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1099180,
      "author_name": "fanbyprinciple",
      "author_url": "",
      "post_date": "12/02/2020 06:35:33",
      "content": "<p>This saves a lot of work. Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1099723,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "12/02/2020 15:07:09",
          "content": "<p>Welcome ))</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1100546,
      "author_name": "iamprateek",
      "author_url": "",
      "post_date": "12/03/2020 06:33:57",
      "content": "<p><a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> have you also uploaded the annotation files?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1100608,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "12/03/2020 07:29:42",
          "content": "<p>if you meant the annotation files for the detection then Not yet. but maybe I will </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149694,
          "author_name": "iamprateek",
          "author_url": "",
          "post_date": "01/12/2021 04:26:30",
          "content": "<p><a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> Have you shared your inference notebook using YOLO?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1179878,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "01/31/2021 23:13:28",
          "content": "<p><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215940\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215940</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1100634,
      "author_name": "walaaothman",
      "author_url": "",
      "post_date": "12/03/2020 07:55:36",
      "content": "<p>Thanks for the info.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1102880,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "12/05/2020 12:39:22",
          "content": "<p>Welcome ))</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1100960,
      "author_name": "eddwait",
      "author_url": "",
      "post_date": "12/03/2020 14:09:48",
      "content": "<p>Thank you ammarali, I really appreciate it. I am new to this field and I have two questions, <br>\n1)How did you clean the image data (not asking for code but concepts)<br>\n2)I tried submitting my code 3 times, and all 3 times I got an error message even though it shows that submission was successful. How can I fix this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1100989,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "12/03/2020 14:34:42",
          "content": "<p>Hello, you are welcome, First of all, don't clean the data cleaning the data will give you worse results because it seems that the testing data is also noisy. For submission issues please give me info about the error message or you can use any baseline for submission. Many competitors shared a template for submissions just change the model architecture load your model and that is all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1101072,
          "author_name": "eddwait",
          "author_url": "",
          "post_date": "12/03/2020 16:00:08",
          "content": "<p>There is no error message. What is baseline submission? My submission is nothing special tbh, tbh it's one of the first competitions I'm taking seriously. I use model.fit_generator(training and validation data) plot accuracies, and then use a function that loads the test image from path, processes it (resizes it) and then predicts it. This prediction is then sent to my_submission.csv with proper image name and label. My submissions are successful, but it doesn't score my submission. \"Submission Scoring Error\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1101085,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "12/03/2020 16:14:32",
          "content": "<p>You can take this as an example <a href=\"https://www.kaggle.com/manojprabhaakr/leaf-classification-resnext-50-32-4d\" target=\"_blank\">https://www.kaggle.com/manojprabhaakr/leaf-classification-resnext-50-32-4d</a><br>\nplease make sure that you are reading the testing data from the test data folder. The data is hidden and will be accessible only after the submission (when you submit your solution they will run it on these hidden data and score it). if you are reading the data fine then check that the output of your model is an integer between 0 and 4, not a float number or out of these bounds.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1101775,
          "author_name": "eddwait",
          "author_url": "",
          "post_date": "12/04/2020 08:35:52",
          "content": "<p>Yes Ammarali, my output is an integer between 0 and 4, but, I am reading the specific csv file (from the folder) and predicting the output. Is there anything else too that needs to be done?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1103133,
      "author_name": "reighns",
      "author_url": "",
      "post_date": "12/05/2020 16:42:58",
      "content": "<p>Great work and thanks! </p>\n<p>In such a scenario where test set may also be mislabeled, what can we do?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1103136,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "12/05/2020 16:45:11",
          "content": "<p>hi <a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> <br>\nI have removed the images I know to be wrong and not considered them in train</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103146,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "12/05/2020 16:52:58",
          "content": "<p>Hi thanks))<br>\nAfter removing the mislabeled images I got a higher CV and lower LB. It is probably an indicator that the testing data also noisy. But after testing over all data that may differ. Anyway, I suggest keeping the noisy images. but it is up to you. Good Luck ))</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103219,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "12/05/2020 18:23:07",
          "content": "<p><a href=\"https://www.kaggle.com/kmldas\" target=\"_blank\">@kmldas</a> have you removed it manually and how can you tell which ones are mislabeled?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103611,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "12/06/2020 04:34:02",
          "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> <br>\nYes, I have removed them manually… those that seemed mislabed from discussions or when I did random batch image show, and felt were wrong I removed the ids from the train.<br>\nI may be wrong but felt better to have a cleaner train set</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103616,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "12/06/2020 04:39:21",
          "content": "<p><a href=\"https://www.kaggle.com/kmldas\" target=\"_blank\">@kmldas</a> any score improvements after cleaning?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103627,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "12/06/2020 04:47:57",
          "content": "<p>My CV improved, the public LB/score not as much … but hoping it helps with pvt LB as CV is better! 🤘</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1103223,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "12/05/2020 18:29:53",
      "content": "<p><a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200201\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200201</a> this dataset have both years data with duplicate images removed; you can use your yolo detector here and get a dataset without duplicates. Good for everyone I guess.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1103269,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "12/05/2020 19:18:20",
          "content": "<p>thanks ))</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1115550,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "12/16/2020 11:07:38",
      "content": "<p>It did not improve my LB neither my CV but I noticed, that the test_images cannot be resized to <code>256x256</code> </p>\n<p>The submission fails.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1181756,
      "author_name": "vickygoyal",
      "author_url": "",
      "post_date": "02/02/2021 05:21:43",
      "content": "<p>I spent days to remove the mislabelled and duplicate data using DBSCAN to find out the accuracy(both cv and lb) actually decreased after doing it.  So it just does not work. Will try with just removing duplicate images to see which model works best with analysing the CV</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1100362,
      "author_name": "ammarali32",
      "author_url": "",
      "post_date": "12/03/2020 03:20:33",
      "content": "<p>Thanks !! ))</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1098341": "https://www.kaggle.com/ammarali32/cassava-datasetv2\nBy this link, you can find a dataset that includes both Cassava datasets. Also, the images were cropped using YOLO detector to detect one leaf in an image (the one with the highest confidence).\nUnfortunately, the detector is not highly accurate because I trained yolo-tinyv3 with 1000 images (drawing boxes is really boring). And it was tiny related to GPU issues.\nhope it will be useful. \nps1. the data published by the competition is really noisy healthy class includes at least 1000 images that aren't. also mixing between class 1 and 0.\n+ a lot of duplicated images.\nTo save your time Cleaning the data didn't increase the accuracy I already tried.",
    "1098362": "Thanks @ammarali32 for sharing ... this is helpful for training!!",
    "1098392": "a big Thanks ; \n`ps1. the data published by the competition is really noisy healthy class includes at least 1000 images that aren't. also mixing between class 1 and 0.`\nSorry I couldn't get what are you trying to say",
    "1098408": "Welcome ))",
    "1098414": "Put all the images with label 4 in a folder and take a look at the data. It should contain healthy images but actually, it didn't.",
    "1098420": "Thanks for the clarification; Does using the new data improve your public_LB",
    "1098444": "No, unfortunately, I got a CV = 90 and a 90.3 LB while 88.9 CV got 90.1 it is really strange. I think the testing data is noisy as well",
    "1099180": "This saves a lot of work. Thanks.",
    "1099723": "Welcome ))",
    "1100362": "Thanks !! ))",
    "1100546": "ammarali32 have you also uploaded the annotation files?",
    "1100608": "if you meant the annotation files for the detection then Not yet. but maybe I will",
    "1100634": "Thanks for the info.",
    "1100960": "Thank you ammarali, I really appreciate it. I am new to this field and I have two questions, \n1)How did you clean the image data (not asking for code but concepts)\n2)I tried submitting my code 3 times, and all 3 times I got an error message even though it shows that submission was successful. How can I fix this?",
    "1100989": "Hello, you are welcome, First of all, don't clean the data cleaning the data will give you worse results because it seems that the testing data is also noisy. For submission issues please give me info about the error message or you can use any baseline for submission. Many competitors shared a template for submissions just change the model architecture load your model and that is all.",
    "1101072": "There is no error message. What is baseline submission? My submission is nothing special tbh, tbh it's one of the first competitions I'm taking seriously. I use model.fit_generator(training and validation data) plot accuracies, and then use a function that loads the test image from path, processes it (resizes it) and then predicts it. This prediction is then sent to my_submission.csv with proper image name and label. My submissions are successful, but it doesn't score my submission. \"Submission Scoring Error\"",
    "1101085": "You can take this as an example https://www.kaggle.com/manojprabhaakr/leaf-classification-resnext-50-32-4d\nplease make sure that you are reading the testing data from the test data folder. The data is hidden and will be accessible only after the submission (when you submit your solution they will run it on these hidden data and score it). if you are reading the data fine then check that the output of your model is an integer between 0 and 4, not a float number or out of these bounds.",
    "1101775": "Yes Ammarali, my output is an integer between 0 and 4, but, I am reading the specific csv file (from the folder) and predicting the output. Is there anything else too that needs to be done?",
    "1102880": "Welcome ))",
    "1103133": "Great work and thanks! \n\nIn such a scenario where test set may also be mislabeled, what can we do?",
    "1103136": "hi @reighns \nI have removed the images I know to be wrong and not considered them in train",
    "1103146": "Hi thanks))\nAfter removing the mislabeled images I got a higher CV and lower LB. It is probably an indicator that the testing data also noisy. But after testing over all data that may differ. Anyway, I suggest keeping the noisy images. but it is up to you. Good Luck ))",
    "1103219": "kmldas have you removed it manually and how can you tell which ones are mislabeled?",
    "1103223": "https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/200201 this dataset have both years data with duplicate images removed; you can use your yolo detector here and get a dataset without duplicates. Good for everyone I guess.",
    "1103269": "thanks ))",
    "1103611": "mrinath \nYes, I have removed them manually... those that seemed mislabed from discussions or when I did random batch image show, and felt were wrong I removed the ids from the train.\nI may be wrong but felt better to have a cleaner train set",
    "1103616": "kmldas any score improvements after cleaning?",
    "1103627": "My CV improved, the public LB/score not as much ... but hoping it helps with pvt LB as CV is better! 🤘",
    "1115550": "It did not improve my LB neither my CV but I noticed, that the test_images cannot be resized to `256x256` \n\nThe submission fails.",
    "1149694": "ammarali32 Have you shared your inference notebook using YOLO?",
    "1179878": "https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/215940",
    "1181756": "I spent days to remove the mislabelled and duplicate data using DBSCAN to find out the accuracy(both cv and lb) actually decreased after doing it.  So it just does not work. Will try with just removing duplicate images to see which model works best with analysing the CV"
  },
  "source": "meta"
}