{
  "id": 127437,
  "title": "Bad training data",
  "url": "/competitions/pku-autonomous-driving/discussion/127437",
  "author_name": "Dmitriy L",
  "post_date": "2020-01-23T22:45:34.459000",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi, looking at the data more carefully, I noticed the following inconsistency in one image ID_8a6e65317:</p>\n\n<p>Image with mask applied, (x,y,z) and models marked:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2Ff73308b7821dacba05e38fb23718fc1f%2FID_8a6e65317_OUT.jpg?generation=1579818739390476&amp;alt=media\" alt=\"\"></p>\n\n<p>Data from train.csv:\nID_8a6e65317,\n16 0.254839 -2.57534 -3.10256 7.96539 3.20066 11.0225 \n56 0.181647 -1.46947 -3.12159 9.60332 4.66632 19.339 \n70 0.163072 -1.56865 -3.11754 10.39 11.2219 59.7825 \n70 0.141942 -3.1395 3.11969 -9.59236 5.13662 24.7337 \n46 0.163068 -2.08578 -3.11754 9.83335 13.2689 72.9323</p>\n\n<p>Model 16 has yaw of -2.57534, which is about -147 deg.  This is inconsistent with the picture.  Car of model 16 is parallel to the one of model 56.  Both cars are parked next to each other on the right at about -90 deg.  Car 56's yaw is -84 deg, which is consistent.</p>\n\n<p>Given previous training data inconsistencies, I am thinking that the training dataset is bad.\nThis is a shame, as this is my first kaggle competition.  </p>\n\n<p>Has anyone else seen similar issues besides yaw and roll being switched AND 5 broken images?\nAccording to this notebook, lots of translation points are simply far outside the images: <a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">https://www.kaggle.com/hocop1/centernet-baseline</a>\nSearch page for \"Many points are outside!\"</p>",
  "messages": [
    {
      "id": 730799,
      "postDate": "2020-01-27T23:53:38.793Z",
      "content": "<p>I have traced the lowest point on the left to image ID_0289b0bc8:\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2F63ebcba04a5966569107387609cafd12%2Fscatter_OUT.jpg?generation=1580169623154728&amp;alt=media\" alt=\"\"></p>\n\n<p>It indeed corresponds to a car mostly out of the picture:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2Fc7c7678b913875e8079eaaaa82c875cc%2FID_0289b0bc8.jpg?generation=1580169211248670&amp;alt=media\" alt=\"\"></p>\n\n<p>Note: the scatter plot of the center points came from Apolloscape dataset. It's the same as the comp dataset, except that the rotational angles were fixed.</p>",
      "rawMarkdown": "I have traced the lowest point on the left to image ID_0289b0bc8:\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2F63ebcba04a5966569107387609cafd12%2Fscatter_OUT.jpg?generation=1580169623154728&amp;alt=media)\n\n\nIt indeed corresponds to a car mostly out of the picture:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2Fc7c7678b913875e8079eaaaa82c875cc%2FID_0289b0bc8.jpg?generation=1580169211248670&amp;alt=media)\n\nNote: the scatter plot of the center points came from Apolloscape dataset. It's the same as the comp dataset, except that the rotational angles were fixed.",
      "votes": 1
    },
    {
      "id": 727622,
      "postDate": "2020-01-23T22:45:34.460Z",
      "content": "<p>Hi, looking at the data more carefully, I noticed the following inconsistency in one image ID_8a6e65317:</p>\n\n<p>Image with mask applied, (x,y,z) and models marked:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2Ff73308b7821dacba05e38fb23718fc1f%2FID_8a6e65317_OUT.jpg?generation=1579818739390476&amp;alt=media\" alt=\"\"></p>\n\n<p>Data from train.csv:\nID_8a6e65317,\n16 0.254839 -2.57534 -3.10256 7.96539 3.20066 11.0225 \n56 0.181647 -1.46947 -3.12159 9.60332 4.66632 19.339 \n70 0.163072 -1.56865 -3.11754 10.39 11.2219 59.7825 \n70 0.141942 -3.1395 3.11969 -9.59236 5.13662 24.7337 \n46 0.163068 -2.08578 -3.11754 9.83335 13.2689 72.9323</p>\n\n<p>Model 16 has yaw of -2.57534, which is about -147 deg.  This is inconsistent with the picture.  Car of model 16 is parallel to the one of model 56.  Both cars are parked next to each other on the right at about -90 deg.  Car 56's yaw is -84 deg, which is consistent.</p>\n\n<p>Given previous training data inconsistencies, I am thinking that the training dataset is bad.\nThis is a shame, as this is my first kaggle competition.  </p>\n\n<p>Has anyone else seen similar issues besides yaw and roll being switched AND 5 broken images?\nAccording to this notebook, lots of translation points are simply far outside the images: <a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">https://www.kaggle.com/hocop1/centernet-baseline</a>\nSearch page for \"Many points are outside!\"</p>",
      "rawMarkdown": "Hi, looking at the data more carefully, I noticed the following inconsistency in one image ID_8a6e65317:\n\nImage with mask applied, (x,y,z) and models marked:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2Ff73308b7821dacba05e38fb23718fc1f%2FID_8a6e65317_OUT.jpg?generation=1579818739390476&amp;alt=media)\n\nData from train.csv:\nID_8a6e65317,\n16 0.254839 -2.57534 -3.10256 7.96539 3.20066 11.0225 \n56 0.181647 -1.46947 -3.12159 9.60332 4.66632 19.339 \n70 0.163072 -1.56865 -3.11754 10.39 11.2219 59.7825 \n70 0.141942 -3.1395 3.11969 -9.59236 5.13662 24.7337 \n46 0.163068 -2.08578 -3.11754 9.83335 13.2689 72.9323\n\nModel 16 has yaw of -2.57534, which is about -147 deg.  This is inconsistent with the picture.  Car of model 16 is parallel to the one of model 56.  Both cars are parked next to each other on the right at about -90 deg.  Car 56's yaw is -84 deg, which is consistent.\n\nGiven previous training data inconsistencies, I am thinking that the training dataset is bad.\nThis is a shame, as this is my first kaggle competition.  \n\nHas anyone else seen similar issues besides yaw and roll being switched AND 5 broken images?\nAccording to this notebook, lots of translation points are simply far outside the images: https://www.kaggle.com/hocop1/centernet-baseline\nSearch page for \"Many points are outside!\"",
      "votes": 1
    },
    {
      "id": 729166,
      "postDate": "2020-01-25T20:51:52.790Z",
      "content": "<p>One question I have is how were the rotational data collected?  I can see the translational data being derived from lidar/radar.  The dataset creators might have used the 3d models to match the car masks best at lidar's x,y,z.  Masks were probably drawn manually(?). That could explain why the car that is partly outside of the image was not matched very well.  At any rate, the dataset should have been validated by people.</p>",
      "rawMarkdown": "One question I have is how were the rotational data collected?  I can see the translational data being derived from lidar/radar.  The dataset creators might have used the 3d models to match the car masks best at lidar's x,y,z.  Masks were probably drawn manually(?). That could explain why the car that is partly outside of the image was not matched very well.  At any rate, the dataset should have been validated by people.",
      "replies": [
        {
          "id": 729589,
          "postDate": "2020-01-26T12:11:38.510Z",
          "content": "<p>I assume labeling is done automatically with human in the loop. And there are always errors in these sets. Not a big deal as long as they are rare (or small) enough.\nThe much bigger issue in this comp was the unclear metric and how the private LB would deal with cropped pictures. </p>\n\n<p>Translation points outside the picture is OK, because they correspond to the center point, which is ofc outside the picture if less than 50% of the car is visible. That's why <a href=\"/hocop1\">@hocop1</a> padded quite some pixels to the sides.</p>",
          "rawMarkdown": "I assume labeling is done automatically with human in the loop. And there are always errors in these sets. Not a big deal as long as they are rare (or small) enough.\nThe much bigger issue in this comp was the unclear metric and how the private LB would deal with cropped pictures. \n\nTranslation points outside the picture is OK, because they correspond to the center point, which is ofc outside the picture if less than 50% of the car is visible. That's why @hocop1 padded quite some pixels to the sides."
        },
        {
          "id": 729832,
          "postDate": "2020-01-26T17:31:50.393Z",
          "content": "<p><a href=\"/ilu000\">@ilu000</a> What is private LB?  At what point are we dealing with the cropped pictures?  Is this cropping for data augmentation?  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2F668344573e3fb24c570d14cd781b99f2%2Fkaggle_6d_img.png?generation=1580059475724460&amp;alt=media\" alt=\"\"></p>\n\n<p>Some points on the left are almost behind the camera.  Others are way too far.  If you try to imagine the white car ahead at those center points, it would definitely be outside the image.</p>\n\n<p>I disagree that wrong data are ok.  Most likely, these translation data can be thrown out in the code.\nRotation data is a different matter.  And that was the 1st image in the set I checked.  That's why I asked if anyone has seen similar issues.</p>",
          "rawMarkdown": "@ilu000 What is private LB?  At what point are we dealing with the cropped pictures?  Is this cropping for data augmentation?  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2F668344573e3fb24c570d14cd781b99f2%2Fkaggle_6d_img.png?generation=1580059475724460&amp;alt=media)\n\nSome points on the left are almost behind the camera.  Others are way too far.  If you try to imagine the white car ahead at those center points, it would definitely be outside the image.\n\nI disagree that wrong data are ok.  Most likely, these translation data can be thrown out in the code.\nRotation data is a different matter.  And that was the 1st image in the set I checked.  That's why I asked if anyone has seen similar issues.\n"
        },
        {
          "id": 730928,
          "postDate": "2020-01-28T06:09:00.867Z",
          "content": "<p>private LB stands for the private leaderboard. Until the very end we had no idea how cropped and flipped images are dealed with, as the metric was unknown and they were not affecting the public leaderboard score.</p>\n\n<p>I see you checked the out-of-image centerpoints yourself. IMO they are fine, and few actual errors exist in the train set (like e.g. broken pictures, and i don't even know how much they affect the score if you train with or without them).</p>",
          "rawMarkdown": "private LB stands for the private leaderboard. Until the very end we had no idea how cropped and flipped images are dealed with, as the metric was unknown and they were not affecting the public leaderboard score.\n\nI see you checked the out-of-image centerpoints yourself. IMO they are fine, and few actual errors exist in the train set (like e.g. broken pictures, and i don't even know how much they affect the score if you train with or without them)."
        }
      ]
    },
    {
      "id": 728331,
      "postDate": "2020-01-24T16:08:51.177Z",
      "content": "<p>Lesson #1 learned from this comp: Check the dataset before entering the comp. </p>",
      "rawMarkdown": "\nLesson #1 learned from this comp: Check the dataset before entering the comp. "
    },
    {
      "id": 727940,
      "postDate": "2020-01-24T08:14:05.540Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 727941,
          "postDate": "2020-01-24T08:21:49.610Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 730799,
      "author_name": "Dmitriy L",
      "author_url": "",
      "post_date": "2020-01-27T23:53:38.793000",
      "content": "<p>I have traced the lowest point on the left to image ID_0289b0bc8:\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2F63ebcba04a5966569107387609cafd12%2Fscatter_OUT.jpg?generation=1580169623154728&amp;alt=media\" alt=\"\"></p>\n\n<p>It indeed corresponds to a car mostly out of the picture:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2Fc7c7678b913875e8079eaaaa82c875cc%2FID_0289b0bc8.jpg?generation=1580169211248670&amp;alt=media\" alt=\"\"></p>\n\n<p>Note: the scatter plot of the center points came from Apolloscape dataset. It's the same as the comp dataset, except that the rotational angles were fixed.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 729166,
      "author_name": "Dmitriy L",
      "author_url": "",
      "post_date": "2020-01-25T20:51:52.790000",
      "content": "<p>One question I have is how were the rotational data collected?  I can see the translational data being derived from lidar/radar.  The dataset creators might have used the 3d models to match the car masks best at lidar's x,y,z.  Masks were probably drawn manually(?). That could explain why the car that is partly outside of the image was not matched very well.  At any rate, the dataset should have been validated by people.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 729589,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-01-26T12:11:38.510000",
          "content": "<p>I assume labeling is done automatically with human in the loop. And there are always errors in these sets. Not a big deal as long as they are rare (or small) enough.\nThe much bigger issue in this comp was the unclear metric and how the private LB would deal with cropped pictures. </p>\n\n<p>Translation points outside the picture is OK, because they correspond to the center point, which is ofc outside the picture if less than 50% of the car is visible. That's why <a href=\"/hocop1\">@hocop1</a> padded quite some pixels to the sides.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 729832,
          "author_name": "Dmitriy L",
          "author_url": "",
          "post_date": "2020-01-26T17:31:50.393000",
          "content": "<p><a href=\"/ilu000\">@ilu000</a> What is private LB?  At what point are we dealing with the cropped pictures?  Is this cropping for data augmentation?  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2F668344573e3fb24c570d14cd781b99f2%2Fkaggle_6d_img.png?generation=1580059475724460&amp;alt=media\" alt=\"\"></p>\n\n<p>Some points on the left are almost behind the camera.  Others are way too far.  If you try to imagine the white car ahead at those center points, it would definitely be outside the image.</p>\n\n<p>I disagree that wrong data are ok.  Most likely, these translation data can be thrown out in the code.\nRotation data is a different matter.  And that was the 1st image in the set I checked.  That's why I asked if anyone has seen similar issues.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 730928,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-01-28T06:09:00.867000",
          "content": "<p>private LB stands for the private leaderboard. Until the very end we had no idea how cropped and flipped images are dealed with, as the metric was unknown and they were not affecting the public leaderboard score.</p>\n\n<p>I see you checked the out-of-image centerpoints yourself. IMO they are fine, and few actual errors exist in the train set (like e.g. broken pictures, and i don't even know how much they affect the score if you train with or without them).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 728331,
      "author_name": "Dmitriy L",
      "author_url": "",
      "post_date": "2020-01-24T16:08:51.177000",
      "content": "<p>Lesson #1 learned from this comp: Check the dataset before entering the comp. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727940,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-24T08:14:05.540000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 727941,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-24T08:21:49.610000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "730799": "I have traced the lowest point on the left to image ID_0289b0bc8:\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2F63ebcba04a5966569107387609cafd12%2Fscatter_OUT.jpg?generation=1580169623154728&amp;alt=media)\n\n\nIt indeed corresponds to a car mostly out of the picture:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2Fc7c7678b913875e8079eaaaa82c875cc%2FID_0289b0bc8.jpg?generation=1580169211248670&amp;alt=media)\n\nNote: the scatter plot of the center points came from Apolloscape dataset. It's the same as the comp dataset, except that the rotational angles were fixed.",
    "727622": "Hi, looking at the data more carefully, I noticed the following inconsistency in one image ID_8a6e65317:\n\nImage with mask applied, (x,y,z) and models marked:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3764314%2Ff73308b7821dacba05e38fb23718fc1f%2FID_8a6e65317_OUT.jpg?generation=1579818739390476&amp;alt=media)\n\nData from train.csv:\nID_8a6e65317,\n16 0.254839 -2.57534 -3.10256 7.96539 3.20066 11.0225 \n56 0.181647 -1.46947 -3.12159 9.60332 4.66632 19.339 \n70 0.163072 -1.56865 -3.11754 10.39 11.2219 59.7825 \n70 0.141942 -3.1395 3.11969 -9.59236 5.13662 24.7337 \n46 0.163068 -2.08578 -3.11754 9.83335 13.2689 72.9323\n\nModel 16 has yaw of -2.57534, which is about -147 deg.  This is inconsistent with the picture.  Car of model 16 is parallel to the one of model 56.  Both cars are parked next to each other on the right at about -90 deg.  Car 56's yaw is -84 deg, which is consistent.\n\nGiven previous training data inconsistencies, I am thinking that the training dataset is bad.\nThis is a shame, as this is my first kaggle competition.  \n\nHas anyone else seen similar issues besides yaw and roll being switched AND 5 broken images?\nAccording to this notebook, lots of translation points are simply far outside the images: https://www.kaggle.com/hocop1/centernet-baseline\nSearch page for \"Many points are outside!\"",
    "729166": "One question I have is how were the rotational data collected?  I can see the translational data being derived from lidar/radar.  The dataset creators might have used the 3d models to match the car masks best at lidar's x,y,z.  Masks were probably drawn manually(?). That could explain why the car that is partly outside of the image was not matched very well.  At any rate, the dataset should have been validated by people.",
    "728331": "\nLesson #1 learned from this comp: Check the dataset before entering the comp. ",
    "727940": ""
  }
}