{
  "id": 36996,
  "title": "Some labels are not correct or missing",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/36996",
  "author_name": "",
  "post_date": "2017-07-25T13:07:52.978266100Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>We've noticed some labels are not correct. for example: fdb996a779e5d65d043eaa160ec2f09f has the threat under arm but the labels are on foot (zone 15 /16)\nSome labels are missing, for example 0a83698bce92a6824dcc37c1d7fc31f5\nAlso some headers for a3daps files are missing.</p>",
  "messages": [
    {
      "id": "206955",
      "postDate": "07/25/2017 13:07:52",
      "content": "<p>We've noticed some labels are not correct. for example: fdb996a779e5d65d043eaa160ec2f09f has the threat under arm but the labels are on foot (zone 15 /16)\nSome labels are missing, for example 0a83698bce92a6824dcc37c1d7fc31f5\nAlso some headers for a3daps files are missing.</p>",
      "rawMarkdown": "We've noticed some labels are not correct. for example: fdb996a779e5d65d043eaa160ec2f09f has the threat under arm but the labels are on foot (zone 15 /16)\nSome labels are missing, for example 0a83698bce92a6824dcc37c1d7fc31f5\nAlso some headers for a3daps files are missing.",
      "votes": null
    },
    {
      "id": "206960",
      "postDate": "07/25/2017 13:35:10",
      "content": "<p>For fdb996a779e5d65d043eaa160ec2f09f, there's a label for Zone 6 (right ribs). The labels are missing for 0a83698bce92a6824dcc37c1d7fc31f5 because it's part of the test set.</p>",
      "rawMarkdown": "For fdb996a779e5d65d043eaa160ec2f09f, there's a label for Zone 6 (right ribs). The labels are missing for 0a83698bce92a6824dcc37c1d7fc31f5 because it's part of the test set.",
      "votes": null
    },
    {
      "id": "206971",
      "postDate": "07/25/2017 13:49:37",
      "content": "<p>Thanks for the reply. but for fdb996a779e5d65d043eaa160ec2f09f, I didn't see any threats on zone 15 and 16 but there is a label for it.</p>",
      "rawMarkdown": "Thanks for the reply. but for fdb996a779e5d65d043eaa160ec2f09f, I didn't see any threats on zone 15 and 16 but there is a label for it.",
      "votes": null
    },
    {
      "id": "206985",
      "postDate": "07/25/2017 14:09:36",
      "content": "<p>Look at the side views. I can see them pretty clearly.</p>",
      "rawMarkdown": "Look at the side views. I can see them pretty clearly.",
      "votes": null
    },
    {
      "id": "207184",
      "postDate": "07/25/2017 23:06:43",
      "content": "<p>There are a small number of incorrectly classified images in the data, but they are rare. I recall finding a problem with an incorrect threat allocation between zones 11/13 or 12/14.  </p>\n\n<p>This shouldn't come as a surprise as I assume that a human has been involved in the labelling and there will have been some subjectivity or opportunity for error when allocating threat to zone.</p>\n\n<p>Afraid I can't provide details of the image as in the process of kaggling I've fried my pc's main drive and am going through the painful process of trying to rebuild it.</p>\n\n<p>What is more common are threats that are rather difficult to spot by eye, unless as Branden says you observe from all  angles.</p>",
      "rawMarkdown": "There are a small number of incorrectly classified images in the data, but they are rare. I recall finding a problem with an incorrect threat allocation between zones 11/13 or 12/14.  \n\nThis shouldn't come as a surprise as I assume that a human has been involved in the labelling and there will have been some subjectivity or opportunity for error when allocating threat to zone.\n\nAfraid I can't provide details of the image as in the process of kaggling I've fried my pc's main drive and am going through the painful process of trying to rebuild it.\n\nWhat is more common are threats that are rather difficult to spot by eye, unless as Branden says you observe from all  angles.",
      "votes": null
    },
    {
      "id": "207499",
      "postDate": "07/26/2017 18:46:46",
      "content": "<p>Wait. The training and test sets are mixed together?? Guess we have to cross reference the spreadsheet to separate; based on no labels &amp; at least one label. Thanks, I just looked at a few files, had not come across any without labels; but some with wrong ones.\n Kinda new to data science, but aren't training &amp; test sets usually separate?</p>",
      "rawMarkdown": "Wait. The training and test sets are mixed together?? Guess we have to cross reference the spreadsheet to separate; based on no labels &amp; at least one label. Thanks, I just looked at a few files, had not come across any without labels; but some with wrong ones.\n Kinda new to data science, but aren't training &amp; test sets usually separate?",
      "votes": null
    },
    {
      "id": "208853",
      "postDate": "07/31/2017 09:18:12",
      "content": "<blockquote>\n  <p>Afraid I can't provide details of the image as in the process of kaggling I've fried my pc's main drive and am going through the painful process of trying to rebuild it.</p>\n</blockquote>\n\n<p>I was wondering why you disappeared from WTF competition after a good start.  I am sorry to see it is because of a HW failure.   Hope you can come back in time as we need more competition there.</p>",
      "rawMarkdown": "&gt; Afraid I can't provide details of the image as in the process of kaggling I've fried my pc's main drive and am going through the painful process of trying to rebuild it.\n\nI was wondering why you disappeared from WTF competition after a good start.  I am sorry to see it is because of a HW failure.   Hope you can come back in time as we need more competition there.",
      "votes": null
    },
    {
      "id": "228814",
      "postDate": "10/07/2017 23:36:09",
      "content": "<p>It is in TSA's interest to revise stage1labels.csv for errata to elicit the best possible algorithms</p>",
      "rawMarkdown": "It is in TSA's interest to revise stage1labels.csv for errata to elicit the best possible algorithms",
      "votes": null
    },
    {
      "id": "1690913",
      "postDate": "02/15/2022 06:47:00",
      "content": "<p>Handling missing values or dealing with incorrect labels are some of the most common challenges data analysts face. Every other dataset you encounter will have issues that you must deal with before training and making predictions. We all know errors in data lead to inaccurate predictions. </p>\n<p>In fact, a recent MIT <a href=\"https://venturebeat.com/2021/03/28/mit-study-finds-systematic-labeling-errors-in-popular-ai-benchmark-datasets/\" target=\"_blank\">study</a> found labeling <a href=\"https://labelerrors.com/\" target=\"_blank\">errors</a> in popular AI datasets. Primarily, missing or incorrect values result from human errors, data corruption, mislabelled images, unanswered survey or interview questions, etc. </p>\n<p>Before handling any incorrect or missing data, you first need to identify it. Deepchecks provides several data checks to help you do that. For instance, I often use <a href=\"https://docs.deepchecks.com/en/stable/examples/checks/integrity/label_ambiguity.html\" target=\"_blank\">LabelAmbiguity</a> to help me identify samples with multiple labels using the following code. </p>\n<pre><code>from deepchecks.checks.integrity import LabelAmbiguity\nfrom deepchecks.base import Dataset\nimport pandas as pd\ndataset = Dataset(pd.DataFrame({\"col1\":[1,2,1,2,1,2,1,2,1,2],\n                                \"col2\":[1,2,1,2,5,2,5,2,3,2],\n                                \"my_label\":[2,3,4,4,4,3,4,5,6,4]}),\n                  label=\"my_label\",\n                  label_type=\"classification_label\")\nLabelAmbiguity().run(dataset)\n</code></pre>\n<p>After you have identified incorrect labels or missing data, its crucial to come up with some approaches you can use to handle them like some of the following: <br>\nDelete the rows and columns with null values. The approach leads to loss of info and is advised when you have large datasets. <br>\nReplace missing data with a measure of central tendency (mean, mode, median). It works better with linear data but can add variance in data. Better when data is small<br>\nAssign a unique category <br>\nWork with synthetic data<br>\nPredict missing values <br>\nWork with algorithms that accommodate missing values, e.g., KNN, Random forest.</p>\n<p>Each comes with its share of pros and cons of course. The important thing is to examine which one suits your scenario best, depending on the data. One of the best ways to go about it is to experiment with several methods to see which will give you optimal results. </p>",
      "rawMarkdown": "Handling missing values or dealing with incorrect labels are some of the most common challenges data analysts face. Every other dataset you encounter will have issues that you must deal with before training and making predictions. We all know errors in data lead to inaccurate predictions. \n\nIn fact, a recent MIT [study](https://venturebeat.com/2021/03/28/mit-study-finds-systematic-labeling-errors-in-popular-ai-benchmark-datasets/) found labeling [errors](https://labelerrors.com/) in popular AI datasets. Primarily, missing or incorrect values result from human errors, data corruption, mislabelled images, unanswered survey or interview questions, etc. \n\nBefore handling any incorrect or missing data, you first need to identify it. Deepchecks provides several data checks to help you do that. For instance, I often use [LabelAmbiguity](https://docs.deepchecks.com/en/stable/examples/checks/integrity/label_ambiguity.html) to help me identify samples with multiple labels using the following code. \n\n```\nfrom deepchecks.checks.integrity import LabelAmbiguity\nfrom deepchecks.base import Dataset\nimport pandas as pd\ndataset = Dataset(pd.DataFrame({\"col1\":[1,2,1,2,1,2,1,2,1,2],\n                                \"col2\":[1,2,1,2,5,2,5,2,3,2],\n                                \"my_label\":[2,3,4,4,4,3,4,5,6,4]}),\n                  label=\"my_label\",\n                  label_type=\"classification_label\")\nLabelAmbiguity().run(dataset)\n```\n\n\n\nAfter you have identified incorrect labels or missing data, its crucial to come up with some approaches you can use to handle them like some of the following: \nDelete the rows and columns with null values. The approach leads to loss of info and is advised when you have large datasets. \nReplace missing data with a measure of central tendency (mean, mode, median). It works better with linear data but can add variance in data. Better when data is small\nAssign a unique category \nWork with synthetic data\nPredict missing values \nWork with algorithms that accommodate missing values, e.g., KNN, Random forest.\n\nEach comes with its share of pros and cons of course. The important thing is to examine which one suits your scenario best, depending on the data. One of the best ways to go about it is to experiment with several methods to see which will give you optimal results.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1690913,
      "author_name": "fredbyrd",
      "author_url": "",
      "post_date": "02/15/2022 06:47:00",
      "content": "<p>Handling missing values or dealing with incorrect labels are some of the most common challenges data analysts face. Every other dataset you encounter will have issues that you must deal with before training and making predictions. We all know errors in data lead to inaccurate predictions. </p>\n<p>In fact, a recent MIT <a href=\"https://venturebeat.com/2021/03/28/mit-study-finds-systematic-labeling-errors-in-popular-ai-benchmark-datasets/\" target=\"_blank\">study</a> found labeling <a href=\"https://labelerrors.com/\" target=\"_blank\">errors</a> in popular AI datasets. Primarily, missing or incorrect values result from human errors, data corruption, mislabelled images, unanswered survey or interview questions, etc. </p>\n<p>Before handling any incorrect or missing data, you first need to identify it. Deepchecks provides several data checks to help you do that. For instance, I often use <a href=\"https://docs.deepchecks.com/en/stable/examples/checks/integrity/label_ambiguity.html\" target=\"_blank\">LabelAmbiguity</a> to help me identify samples with multiple labels using the following code. </p>\n<pre><code>from deepchecks.checks.integrity import LabelAmbiguity\nfrom deepchecks.base import Dataset\nimport pandas as pd\ndataset = Dataset(pd.DataFrame({\"col1\":[1,2,1,2,1,2,1,2,1,2],\n                                \"col2\":[1,2,1,2,5,2,5,2,3,2],\n                                \"my_label\":[2,3,4,4,4,3,4,5,6,4]}),\n                  label=\"my_label\",\n                  label_type=\"classification_label\")\nLabelAmbiguity().run(dataset)\n</code></pre>\n<p>After you have identified incorrect labels or missing data, its crucial to come up with some approaches you can use to handle them like some of the following: <br>\nDelete the rows and columns with null values. The approach leads to loss of info and is advised when you have large datasets. <br>\nReplace missing data with a measure of central tendency (mean, mode, median). It works better with linear data but can add variance in data. Better when data is small<br>\nAssign a unique category <br>\nWork with synthetic data<br>\nPredict missing values <br>\nWork with algorithms that accommodate missing values, e.g., KNN, Random forest.</p>\n<p>Each comes with its share of pros and cons of course. The important thing is to examine which one suits your scenario best, depending on the data. One of the best ways to go about it is to experiment with several methods to see which will give you optimal results. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 206960,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "07/25/2017 13:35:10",
      "content": "<p>For fdb996a779e5d65d043eaa160ec2f09f, there's a label for Zone 6 (right ribs). The labels are missing for 0a83698bce92a6824dcc37c1d7fc31f5 because it's part of the test set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 206971,
          "author_name": "",
          "author_url": "",
          "post_date": "07/25/2017 13:49:37",
          "content": "<p>Thanks for the reply. but for fdb996a779e5d65d043eaa160ec2f09f, I didn't see any threats on zone 15 and 16 but there is a label for it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 206985,
          "author_name": "brandenkmurray",
          "author_url": "",
          "post_date": "07/25/2017 14:09:36",
          "content": "<p>Look at the side views. I can see them pretty clearly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 207184,
      "author_name": "nigelcarpenter",
      "author_url": "",
      "post_date": "07/25/2017 23:06:43",
      "content": "<p>There are a small number of incorrectly classified images in the data, but they are rare. I recall finding a problem with an incorrect threat allocation between zones 11/13 or 12/14.  </p>\n\n<p>This shouldn't come as a surprise as I assume that a human has been involved in the labelling and there will have been some subjectivity or opportunity for error when allocating threat to zone.</p>\n\n<p>Afraid I can't provide details of the image as in the process of kaggling I've fried my pc's main drive and am going through the painful process of trying to rebuild it.</p>\n\n<p>What is more common are threats that are rather difficult to spot by eye, unless as Branden says you observe from all  angles.</p>",
      "votes": null,
      "replies": [
        {
          "id": 208853,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/31/2017 09:18:12",
          "content": "<blockquote>\n  <p>Afraid I can't provide details of the image as in the process of kaggling I've fried my pc's main drive and am going through the painful process of trying to rebuild it.</p>\n</blockquote>\n\n<p>I was wondering why you disappeared from WTF competition after a good start.  I am sorry to see it is because of a HW failure.   Hope you can come back in time as we need more competition there.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 207499,
      "author_name": "srdhaeayautnon",
      "author_url": "",
      "post_date": "07/26/2017 18:46:46",
      "content": "<p>Wait. The training and test sets are mixed together?? Guess we have to cross reference the spreadsheet to separate; based on no labels &amp; at least one label. Thanks, I just looked at a few files, had not come across any without labels; but some with wrong ones.\n Kinda new to data science, but aren't training &amp; test sets usually separate?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 228814,
      "author_name": "lwalzer",
      "author_url": "",
      "post_date": "10/07/2017 23:36:09",
      "content": "<p>It is in TSA's interest to revise stage1labels.csv for errata to elicit the best possible algorithms</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "206955": "We've noticed some labels are not correct. for example: fdb996a779e5d65d043eaa160ec2f09f has the threat under arm but the labels are on foot (zone 15 /16)\nSome labels are missing, for example 0a83698bce92a6824dcc37c1d7fc31f5\nAlso some headers for a3daps files are missing.",
    "206960": "For fdb996a779e5d65d043eaa160ec2f09f, there's a label for Zone 6 (right ribs). The labels are missing for 0a83698bce92a6824dcc37c1d7fc31f5 because it's part of the test set.",
    "206971": "Thanks for the reply. but for fdb996a779e5d65d043eaa160ec2f09f, I didn't see any threats on zone 15 and 16 but there is a label for it.",
    "206985": "Look at the side views. I can see them pretty clearly.",
    "207184": "There are a small number of incorrectly classified images in the data, but they are rare. I recall finding a problem with an incorrect threat allocation between zones 11/13 or 12/14.  \n\nThis shouldn't come as a surprise as I assume that a human has been involved in the labelling and there will have been some subjectivity or opportunity for error when allocating threat to zone.\n\nAfraid I can't provide details of the image as in the process of kaggling I've fried my pc's main drive and am going through the painful process of trying to rebuild it.\n\nWhat is more common are threats that are rather difficult to spot by eye, unless as Branden says you observe from all  angles.",
    "207499": "Wait. The training and test sets are mixed together?? Guess we have to cross reference the spreadsheet to separate; based on no labels &amp; at least one label. Thanks, I just looked at a few files, had not come across any without labels; but some with wrong ones.\n Kinda new to data science, but aren't training &amp; test sets usually separate?",
    "208853": "&gt; Afraid I can't provide details of the image as in the process of kaggling I've fried my pc's main drive and am going through the painful process of trying to rebuild it.\n\nI was wondering why you disappeared from WTF competition after a good start.  I am sorry to see it is because of a HW failure.   Hope you can come back in time as we need more competition there.",
    "228814": "It is in TSA's interest to revise stage1labels.csv for errata to elicit the best possible algorithms",
    "1690913": "Handling missing values or dealing with incorrect labels are some of the most common challenges data analysts face. Every other dataset you encounter will have issues that you must deal with before training and making predictions. We all know errors in data lead to inaccurate predictions. \n\nIn fact, a recent MIT [study](https://venturebeat.com/2021/03/28/mit-study-finds-systematic-labeling-errors-in-popular-ai-benchmark-datasets/) found labeling [errors](https://labelerrors.com/) in popular AI datasets. Primarily, missing or incorrect values result from human errors, data corruption, mislabelled images, unanswered survey or interview questions, etc. \n\nBefore handling any incorrect or missing data, you first need to identify it. Deepchecks provides several data checks to help you do that. For instance, I often use [LabelAmbiguity](https://docs.deepchecks.com/en/stable/examples/checks/integrity/label_ambiguity.html) to help me identify samples with multiple labels using the following code. \n\n```\nfrom deepchecks.checks.integrity import LabelAmbiguity\nfrom deepchecks.base import Dataset\nimport pandas as pd\ndataset = Dataset(pd.DataFrame({\"col1\":[1,2,1,2,1,2,1,2,1,2],\n                                \"col2\":[1,2,1,2,5,2,5,2,3,2],\n                                \"my_label\":[2,3,4,4,4,3,4,5,6,4]}),\n                  label=\"my_label\",\n                  label_type=\"classification_label\")\nLabelAmbiguity().run(dataset)\n```\n\n\n\nAfter you have identified incorrect labels or missing data, its crucial to come up with some approaches you can use to handle them like some of the following: \nDelete the rows and columns with null values. The approach leads to loss of info and is advised when you have large datasets. \nReplace missing data with a measure of central tendency (mean, mode, median). It works better with linear data but can add variance in data. Better when data is small\nAssign a unique category \nWork with synthetic data\nPredict missing values \nWork with algorithms that accommodate missing values, e.g., KNN, Random forest.\n\nEach comes with its share of pros and cons of course. The important thing is to examine which one suits your scenario best, depending on the data. One of the best ways to go about it is to experiment with several methods to see which will give you optimal results."
  },
  "source": "meta"
}