{
  "id": 30895,
  "title": "Mismatched training images",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/30895",
  "author_name": "DataCanary",
  "post_date": "2017-03-30T20:55:11.833000",
  "votes": 26,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Over the last day or two, a few competitors have pointed out a series of discrepancies in the training data, where images in the Train folder don't match the corresponding ones in TrainDotted.</p>\n\n<p>After investigating, we've found 57 images (out of 948 total) that are not exact matches. Since this is only a small proportion of the entire dataset, we’ve posted a list of the affected train_ids in an extra file on the Data page, rather than rebuilding the entire archive. In the affected cases, the dotted images could still be useful for training your models, and the counts of sealions in the dotted images should match the values in train.csv. [Edit: reworded last sentence for clarity.]</p>\n\n<p>We've replaced the small subset of training images on the Data page. The file TrainSmall2.7z contains training images 41 to 50, which match their dotted equivalents.</p>\n\n<p>We’ve also rechecked the images in the Test folder, and confirmed that they have not been affected. Thanks for your patience while we worked through verifying the data.</p>\n\n<p>In general, please bear in mind that this data was generated by humans. There will be mistakes, transcription errors, and occasionally typos. We've done our best to detect those things and minimize their impact, but some will still sneak through. In the future, this output will be produced by an algorithm. Maybe your algorithm! We're excited to see what you come up with.</p>",
  "messages": [
    {
      "id": 171624,
      "postDate": "2017-03-30T20:55:11.833Z",
      "content": "<p>Over the last day or two, a few competitors have pointed out a series of discrepancies in the training data, where images in the Train folder don't match the corresponding ones in TrainDotted.</p>\n\n<p>After investigating, we've found 57 images (out of 948 total) that are not exact matches. Since this is only a small proportion of the entire dataset, we’ve posted a list of the affected train_ids in an extra file on the Data page, rather than rebuilding the entire archive. In the affected cases, the dotted images could still be useful for training your models, and the counts of sealions in the dotted images should match the values in train.csv. [Edit: reworded last sentence for clarity.]</p>\n\n<p>We've replaced the small subset of training images on the Data page. The file TrainSmall2.7z contains training images 41 to 50, which match their dotted equivalents.</p>\n\n<p>We’ve also rechecked the images in the Test folder, and confirmed that they have not been affected. Thanks for your patience while we worked through verifying the data.</p>\n\n<p>In general, please bear in mind that this data was generated by humans. There will be mistakes, transcription errors, and occasionally typos. We've done our best to detect those things and minimize their impact, but some will still sneak through. In the future, this output will be produced by an algorithm. Maybe your algorithm! We're excited to see what you come up with.</p>",
      "rawMarkdown": "Over the last day or two, a few competitors have pointed out a series of discrepancies in the training data, where images in the Train folder don't match the corresponding ones in TrainDotted.\n\nAfter investigating, we've found 57 images (out of 948 total) that are not exact matches. Since this is only a small proportion of the entire dataset, we’ve posted a list of the affected train_ids in an extra file on the Data page, rather than rebuilding the entire archive. In the affected cases, the dotted images could still be useful for training your models, and the counts of sealions in the dotted images should match the values in train.csv. [Edit: reworded last sentence for clarity.]\n\nWe've replaced the small subset of training images on the Data page. The file TrainSmall2.7z contains training images 41 to 50, which match their dotted equivalents.\n\nWe’ve also rechecked the images in the Test folder, and confirmed that they have not been affected. Thanks for your patience while we worked through verifying the data.\n\nIn general, please bear in mind that this data was generated by humans. There will be mistakes, transcription errors, and occasionally typos. We've done our best to detect those things and minimize their impact, but some will still sneak through. In the future, this output will be produced by an algorithm. Maybe your algorithm! We're excited to see what you come up with.\n",
      "votes": 26
    },
    {
      "id": 178809,
      "postDate": "2017-04-28T22:40:46.390Z",
      "content": "<p>@DataCanary can you please check image 901: it's in MismatchedTrainImages.txt, but as far as I can see the image in Train matches the image in TrainDotted. There may be other similar images in MismatchedTrainImages.txt, I didn't check all yet. Did it get to MismatchedTrainImages by mistake, or there is some reason I am missing?</p>\n\n<p>Edit: same for 840, 913 - so far about 1/4 of images I checked</p>",
      "rawMarkdown": "@DataCanary can you please check image 901: it's in MismatchedTrainImages.txt, but as far as I can see the image in Train matches the image in TrainDotted. There may be other similar images in MismatchedTrainImages.txt, I didn't check all yet. Did it get to MismatchedTrainImages by mistake, or there is some reason I am missing?\n\nEdit: same for 840, 913 - so far about 1/4 of images I checked",
      "votes": 1
    },
    {
      "id": 178272,
      "postDate": "2017-04-27T07:17:37.827Z",
      "content": "<p>Please have a look at the dotted version of  280.jpg \nOn the left there seem to be hundreds of pups that are unmarked\nIs this a mistake or are they just not sealions?</p>",
      "rawMarkdown": "Please have a look at the dotted version of  280.jpg \nOn the left there seem to be hundreds of pups that are unmarked\nIs this a mistake or are they just not sealions?",
      "votes": 1,
      "replies": [
        {
          "id": 178851,
          "postDate": "2017-04-29T02:18:33.847Z",
          "content": "<p>695 is very similar to 280. Could be same beach. I don't think they are sea lions. You don't see big groups of pups with no mothers in other images. Possibly seals. Possibly dead? Zooming in, some of the bodies look weird to me. </p>",
          "rawMarkdown": "695 is very similar to 280. Could be same beach. I don't think they are sea lions. You don't see big groups of pups with no mothers in other images. Possibly seals. Possibly dead? Zooming in, some of the bodies look weird to me. "
        },
        {
          "id": 179747,
          "postDate": "2017-05-02T18:27:33.067Z",
          "content": "<p>I think this may be images from Bogoslof which has both Steller sea lions and northern fur seals. It can be hard to tell Steller sea lion pups from northern fur seals, even when counting by eye so it will be a challenging task for this competition, I expect. A good challenge :)</p>",
          "rawMarkdown": "I think this may be images from Bogoslof which has both Steller sea lions and northern fur seals. It can be hard to tell Steller sea lion pups from northern fur seals, even when counting by eye so it will be a challenging task for this competition, I expect. A good challenge :)",
          "votes": 2
        }
      ]
    },
    {
      "id": 172046,
      "postDate": "2017-04-01T17:26:12.590Z",
      "content": "<p>@DataCanary</p>\n\n<p>I've created a tiny script that cuts dots of interest I'm using your 'black list' of images but still see a lot mismatches.</p>\n\n<p>I will post only images with small counts of missing/extra dots, so they can be checked manually\n2 - brown, blue\n13 - red\n18 - magenta\n324 - brown \n539 - red\netc(can post the whole list if you like)</p>\n\n<p>In general I can use my script as a source of truth, but what should we do with public/private leaderboard that potentially is built on noisy data?</p>\n\n<p>I can propose to update training data or generate training/test counts by some publicly available script(will share mine if you agree). </p>\n\n<p>We should predict real values, not a noise. I've seen some images with pretty huge difference like 20, so it will ruin even a vary good model.</p>\n\n<p>Would be nice to help you investigate and recover population at Aleutian Islands(Алеутские острова) </p>",
      "rawMarkdown": "@DataCanary\n\nI've created a tiny script that cuts dots of interest I'm using your 'black list' of images but still see a lot mismatches.\n\nI will post only images with small counts of missing/extra dots, so they can be checked manually\n2 - brown, blue\n13 - red\n18 - magenta\n324 - brown \n539 - red\netc(can post the whole list if you like)\n\nIn general I can use my script as a source of truth, but what should we do with public/private leaderboard that potentially is built on noisy data?\n\nI can propose to update training data or generate training/test counts by some publicly available script(will share mine if you agree). \n\nWe should predict real values, not a noise. I've seen some images with pretty huge difference like 20, so it will ruin even a vary good model.\n\n\nWould be nice to help you investigate and recover population at Aleutian Islands(Алеутские острова) ",
      "votes": 1
    },
    {
      "id": 171710,
      "postDate": "2017-03-31T05:19:04.820Z",
      "content": "<p>Hi,</p>\n\n<p>I just found out there is pictures like this in Train set which is not in the mismatch list</p>",
      "rawMarkdown": "Hi,\n\nI just found out there is pictures like this in Train set which is not in the mismatch list",
      "votes": 1
    },
    {
      "id": 173374,
      "postDate": "2017-04-06T16:59:50.417Z",
      "content": "<p>On image 50 from TrainSmall2.7z - only one red dot, but it seems there can be seen other lions.</p>",
      "rawMarkdown": "On image 50 from TrainSmall2.7z - only one red dot, but it seems there can be seen other lions.",
      "votes": 2,
      "replies": [
        {
          "id": 179744,
          "postDate": "2017-05-02T18:18:55.917Z",
          "content": "<p>The other animals are actually northern fur seals, not Steller sea lions. The images from this site are tough to even count by eye since it can be hard to tell Steller sea lions from northern fur seals.</p>",
          "rawMarkdown": "The other animals are actually northern fur seals, not Steller sea lions. The images from this site are tough to even count by eye since it can be hard to tell Steller sea lions from northern fur seals.",
          "votes": 2
        }
      ]
    },
    {
      "id": 172049,
      "postDate": "2017-04-01T17:43:28.697Z",
      "content": "<p>My favorite one is <strong>66</strong>, according to the csv file it contains none, my script extracts \nred 8\nmagenta 5\nbrown 23\nblue 17\ngreen 2</p>",
      "rawMarkdown": "My favorite one is **66**, according to the csv file it contains none, my script extracts \nred 8\nmagenta 5\nbrown 23\nblue 17\ngreen 2",
      "votes": 2
    },
    {
      "id": 180719,
      "postDate": "2017-05-06T19:52:08.570Z",
      "content": "<ol>\n<li>Can we train directly on TrainDotted? i.e. inpaint color dots and train on these images?</li>\n<li>Reported number of species is equal to number of dots in TrainDotted image (not the 'real' number of sea lions on Train image)?</li>\n<li>Can we manually edit number of dots and their location in train data?</li>\n<li>Can we manually remove images from train data? can we build patch based training data by hand? </li>\n</ol>",
      "rawMarkdown": "1. Can we train directly on TrainDotted? i.e. inpaint color dots and train on these images?\n2. Reported number of species is equal to number of dots in TrainDotted image (not the 'real' number of sea lions on Train image)?\n3. Can we manually edit number of dots and their location in train data?\n4. Can we manually remove images from train data? can we build patch based training data by hand? ",
      "replies": [
        {
          "id": 181021,
          "postDate": "2017-05-08T09:40:58.490Z",
          "content": "<p>I guess you can do whatever you like with the training data. The score is evaluated on a separated test set though. So keep in mind that you have to apply all  transformations to the test set as well. Also note that there are no colored dots in the test set.</p>",
          "rawMarkdown": "I guess you can do whatever you like with the training data. The score is evaluated on a separated test set though. So keep in mind that you have to apply all  transformations to the test set as well. Also note that there are no colored dots in the test set."
        }
      ]
    },
    {
      "id": 172930,
      "postDate": "2017-04-05T09:35:04.087Z",
      "content": "<p>Please check the counts in these images:</p>\n\n<p>66  235  292  426  529  593 643  816 857 </p>",
      "rawMarkdown": "Please check the counts in these images:\n\n66  235  292  426  529  593 643  816 857 "
    },
    {
      "id": 172877,
      "postDate": "2017-04-05T04:20:07.737Z",
      "content": "<p>Hi,</p>\n\n<p>I am new to Kaggle. I wanted to know if the kernels provided by kaggle contains \"ALL of the DATA\" or I will still need to download 95 GB for training data?</p>",
      "rawMarkdown": "Hi,\n\nI am new to Kaggle. I wanted to know if the kernels provided by kaggle contains \"ALL of the DATA\" or I will still need to download 95 GB for training data?",
      "replies": [
        {
          "id": 175340,
          "postDate": "2017-04-15T00:56:50.213Z",
          "content": "<p>Kaggle would provide the data if you're running a kernel. If you need to work locally, then you'll have to download the data.</p>",
          "rawMarkdown": "Kaggle would provide the data if you're running a kernel. If you need to work locally, then you'll have to download the data."
        }
      ]
    },
    {
      "id": 188137,
      "postDate": "2017-06-02T04:05:28.697Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 178809,
      "author_name": "Konstantin Lopukhin",
      "author_url": "",
      "post_date": "2017-04-28T22:40:46.390000",
      "content": "<p>@DataCanary can you please check image 901: it's in MismatchedTrainImages.txt, but as far as I can see the image in Train matches the image in TrainDotted. There may be other similar images in MismatchedTrainImages.txt, I didn't check all yet. Did it get to MismatchedTrainImages by mistake, or there is some reason I am missing?</p>\n\n<p>Edit: same for 840, 913 - so far about 1/4 of images I checked</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 178272,
      "author_name": "pmgpmg",
      "author_url": "",
      "post_date": "2017-04-27T07:17:37.827000",
      "content": "<p>Please have a look at the dotted version of  280.jpg \nOn the left there seem to be hundreds of pups that are unmarked\nIs this a mistake or are they just not sealions?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 178851,
          "author_name": "threeplusone",
          "author_url": "",
          "post_date": "2017-04-29T02:18:33.847000",
          "content": "<p>695 is very similar to 280. Could be same beach. I don't think they are sea lions. You don't see big groups of pups with no mothers in other images. Possibly seals. Possibly dead? Zooming in, some of the bodies look weird to me. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 179747,
          "author_name": "Katie",
          "author_url": "",
          "post_date": "2017-05-02T18:27:33.067000",
          "content": "<p>I think this may be images from Bogoslof which has both Steller sea lions and northern fur seals. It can be hard to tell Steller sea lion pups from northern fur seals, even when counting by eye so it will be a challenging task for this competition, I expect. A good challenge :)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 172046,
      "author_name": "sh1ng",
      "author_url": "",
      "post_date": "2017-04-01T17:26:12.590000",
      "content": "<p>@DataCanary</p>\n\n<p>I've created a tiny script that cuts dots of interest I'm using your 'black list' of images but still see a lot mismatches.</p>\n\n<p>I will post only images with small counts of missing/extra dots, so they can be checked manually\n2 - brown, blue\n13 - red\n18 - magenta\n324 - brown \n539 - red\netc(can post the whole list if you like)</p>\n\n<p>In general I can use my script as a source of truth, but what should we do with public/private leaderboard that potentially is built on noisy data?</p>\n\n<p>I can propose to update training data or generate training/test counts by some publicly available script(will share mine if you agree). </p>\n\n<p>We should predict real values, not a noise. I've seen some images with pretty huge difference like 20, so it will ruin even a vary good model.</p>\n\n<p>Would be nice to help you investigate and recover population at Aleutian Islands(Алеутские острова) </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 171710,
      "author_name": "qwer",
      "author_url": "",
      "post_date": "2017-03-31T05:19:04.820000",
      "content": "<p>Hi,</p>\n\n<p>I just found out there is pictures like this in Train set which is not in the mismatch list</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 173374,
      "author_name": "Vladimir Savinov",
      "author_url": "",
      "post_date": "2017-04-06T16:59:50.417000",
      "content": "<p>On image 50 from TrainSmall2.7z - only one red dot, but it seems there can be seen other lions.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 179744,
          "author_name": "Katie",
          "author_url": "",
          "post_date": "2017-05-02T18:18:55.917000",
          "content": "<p>The other animals are actually northern fur seals, not Steller sea lions. The images from this site are tough to even count by eye since it can be hard to tell Steller sea lions from northern fur seals.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 172049,
      "author_name": "sh1ng",
      "author_url": "",
      "post_date": "2017-04-01T17:43:28.697000",
      "content": "<p>My favorite one is <strong>66</strong>, according to the csv file it contains none, my script extracts \nred 8\nmagenta 5\nbrown 23\nblue 17\ngreen 2</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 180719,
      "author_name": "mrgloom",
      "author_url": "",
      "post_date": "2017-05-06T19:52:08.570000",
      "content": "<ol>\n<li>Can we train directly on TrainDotted? i.e. inpaint color dots and train on these images?</li>\n<li>Reported number of species is equal to number of dots in TrainDotted image (not the 'real' number of sea lions on Train image)?</li>\n<li>Can we manually edit number of dots and their location in train data?</li>\n<li>Can we manually remove images from train data? can we build patch based training data by hand? </li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 181021,
          "author_name": "Markus",
          "author_url": "",
          "post_date": "2017-05-08T09:40:58.490000",
          "content": "<p>I guess you can do whatever you like with the training data. The score is evaluated on a separated test set though. So keep in mind that you have to apply all  transformations to the test set as well. Also note that there are no colored dots in the test set.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 172930,
      "author_name": "guodongwu",
      "author_url": "",
      "post_date": "2017-04-05T09:35:04.087000",
      "content": "<p>Please check the counts in these images:</p>\n\n<p>66  235  292  426  529  593 643  816 857 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 172877,
      "author_name": "SubarnaRana",
      "author_url": "",
      "post_date": "2017-04-05T04:20:07.737000",
      "content": "<p>Hi,</p>\n\n<p>I am new to Kaggle. I wanted to know if the kernels provided by kaggle contains \"ALL of the DATA\" or I will still need to download 95 GB for training data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 175340,
          "author_name": "hirviö",
          "author_url": "",
          "post_date": "2017-04-15T00:56:50.213000",
          "content": "<p>Kaggle would provide the data if you're running a kernel. If you need to work locally, then you'll have to download the data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 188137,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-06-02T04:05:28.697000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "171624": "Over the last day or two, a few competitors have pointed out a series of discrepancies in the training data, where images in the Train folder don't match the corresponding ones in TrainDotted.\n\nAfter investigating, we've found 57 images (out of 948 total) that are not exact matches. Since this is only a small proportion of the entire dataset, we’ve posted a list of the affected train_ids in an extra file on the Data page, rather than rebuilding the entire archive. In the affected cases, the dotted images could still be useful for training your models, and the counts of sealions in the dotted images should match the values in train.csv. [Edit: reworded last sentence for clarity.]\n\nWe've replaced the small subset of training images on the Data page. The file TrainSmall2.7z contains training images 41 to 50, which match their dotted equivalents.\n\nWe’ve also rechecked the images in the Test folder, and confirmed that they have not been affected. Thanks for your patience while we worked through verifying the data.\n\nIn general, please bear in mind that this data was generated by humans. There will be mistakes, transcription errors, and occasionally typos. We've done our best to detect those things and minimize their impact, but some will still sneak through. In the future, this output will be produced by an algorithm. Maybe your algorithm! We're excited to see what you come up with.\n",
    "178809": "@DataCanary can you please check image 901: it's in MismatchedTrainImages.txt, but as far as I can see the image in Train matches the image in TrainDotted. There may be other similar images in MismatchedTrainImages.txt, I didn't check all yet. Did it get to MismatchedTrainImages by mistake, or there is some reason I am missing?\n\nEdit: same for 840, 913 - so far about 1/4 of images I checked",
    "178272": "Please have a look at the dotted version of  280.jpg \nOn the left there seem to be hundreds of pups that are unmarked\nIs this a mistake or are they just not sealions?",
    "172046": "@DataCanary\n\nI've created a tiny script that cuts dots of interest I'm using your 'black list' of images but still see a lot mismatches.\n\nI will post only images with small counts of missing/extra dots, so they can be checked manually\n2 - brown, blue\n13 - red\n18 - magenta\n324 - brown \n539 - red\netc(can post the whole list if you like)\n\nIn general I can use my script as a source of truth, but what should we do with public/private leaderboard that potentially is built on noisy data?\n\nI can propose to update training data or generate training/test counts by some publicly available script(will share mine if you agree). \n\nWe should predict real values, not a noise. I've seen some images with pretty huge difference like 20, so it will ruin even a vary good model.\n\n\nWould be nice to help you investigate and recover population at Aleutian Islands(Алеутские острова) ",
    "171710": "Hi,\n\nI just found out there is pictures like this in Train set which is not in the mismatch list",
    "173374": "On image 50 from TrainSmall2.7z - only one red dot, but it seems there can be seen other lions.",
    "172049": "My favorite one is **66**, according to the csv file it contains none, my script extracts \nred 8\nmagenta 5\nbrown 23\nblue 17\ngreen 2",
    "180719": "1. Can we train directly on TrainDotted? i.e. inpaint color dots and train on these images?\n2. Reported number of species is equal to number of dots in TrainDotted image (not the 'real' number of sea lions on Train image)?\n3. Can we manually edit number of dots and their location in train data?\n4. Can we manually remove images from train data? can we build patch based training data by hand? ",
    "172930": "Please check the counts in these images:\n\n66  235  292  426  529  593 643  816 857 ",
    "172877": "Hi,\n\nI am new to Kaggle. I wanted to know if the kernels provided by kaggle contains \"ALL of the DATA\" or I will still need to download 95 GB for training data?",
    "188137": ""
  }
}