{
  "id": 117621,
  "title": "5 broken picture files in train, list inside",
  "url": "/competitions/pku-autonomous-driving/discussion/117621",
  "author_name": "Pascal Pfeiffer",
  "post_date": "2019-11-16T19:53:53.929000",
  "votes": 55,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi everyone.\nA while ago I noticed that some pictures were broken like the one shown below. Downloading the files again didn't solve the problem. Apparently, all those pictures are from the same road.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2675447%2Fdbc409427d2bf9194ef3e3f7c8784110%2Fbroken_pictures.PNG?generation=1573933803767282&amp;alt=media\" alt=\"\"></p>\n\n<p>The broken train pictures are:\n<code>\nID_1a5a10365.jpg\nID_4d238ae90.jpg\nID_408f58e9f.jpg\nID_bb1d991f6.jpg\nID_c44983aeb.jpg\n</code></p>\n\n<p>I found those by visual inspection, could have missed some.\nFrom looking at the test set, all pictures there seem to be OK.</p>",
  "messages": [
    {
      "id": 674626,
      "postDate": "2019-11-16T19:53:53.930Z",
      "content": "<p>Hi everyone.\nA while ago I noticed that some pictures were broken like the one shown below. Downloading the files again didn't solve the problem. Apparently, all those pictures are from the same road.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2675447%2Fdbc409427d2bf9194ef3e3f7c8784110%2Fbroken_pictures.PNG?generation=1573933803767282&amp;alt=media\" alt=\"\"></p>\n\n<p>The broken train pictures are:\n<code>\nID_1a5a10365.jpg\nID_4d238ae90.jpg\nID_408f58e9f.jpg\nID_bb1d991f6.jpg\nID_c44983aeb.jpg\n</code></p>\n\n<p>I found those by visual inspection, could have missed some.\nFrom looking at the test set, all pictures there seem to be OK.</p>",
      "rawMarkdown": "Hi everyone.\nA while ago I noticed that some pictures were broken like the one shown below. Downloading the files again didn't solve the problem. Apparently, all those pictures are from the same road.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2675447%2Fdbc409427d2bf9194ef3e3f7c8784110%2Fbroken_pictures.PNG?generation=1573933803767282&amp;alt=media)\n\nThe broken train pictures are:\n```\nID_1a5a10365.jpg\nID_4d238ae90.jpg\nID_408f58e9f.jpg\nID_bb1d991f6.jpg\nID_c44983aeb.jpg\n```\n\nI found those by visual inspection, could have missed some.\nFrom looking at the test set, all pictures there seem to be OK.\n",
      "votes": 55
    },
    {
      "id": 720293,
      "postDate": "2020-01-16T09:54:11.547Z",
      "content": "<p><a href=\"/ilu000\">@ilu000</a> Thanks for this post! Question: Did you use resnext50 with pretrained weights for your current score?</p>",
      "rawMarkdown": "@ilu000 Thanks for this post! Question: Did you use resnext50 with pretrained weights for your current score?",
      "replies": [
        {
          "id": 720394,
          "postDate": "2020-01-16T11:38:49.113Z",
          "content": "<p>I am training from scratch and found the backbone to be of secondary importance. Replacing resnet/efficientnet or hourglass needs tuning but yields similar results. At least, that's how it turned out for me.</p>",
          "rawMarkdown": "I am training from scratch and found the backbone to be of secondary importance. Replacing resnet/efficientnet or hourglass needs tuning but yields similar results. At least, that's how it turned out for me.",
          "votes": 1
        },
        {
          "id": 720487,
          "postDate": "2020-01-16T13:22:40.453Z",
          "content": "<p>I see, so even with pretrained=False, I should expect after 10 epochs to get LB Score of Top 100?</p>",
          "rawMarkdown": "I see, so even with pretrained=False, I should expect after 10 epochs to get LB Score of Top 100?"
        },
        {
          "id": 720644,
          "postDate": "2020-01-16T16:02:19.903Z",
          "content": "<p>I haven't tried to exactly replicate the centernet paper (or adapt the github code for this competition), but I am guessing a 0.070+ score is easily possible without much additional tuning. So yes, top 100 is in reach. \n10 epochs might not be enough, though. I am usually training for a bit more until the nets overfit. Never quit before reaching that state ;)</p>",
          "rawMarkdown": "I haven't tried to exactly replicate the centernet paper (or adapt the github code for this competition), but I am guessing a 0.070+ score is easily possible without much additional tuning. So yes, top 100 is in reach. \n10 epochs might not be enough, though. I am usually training for a bit more until the nets overfit. Never quit before reaching that state ;)",
          "votes": 1
        },
        {
          "id": 720809,
          "postDate": "2020-01-16T19:00:24.987Z",
          "content": "<p>I see, thank you for the suggestion, what is the expected loss when it's overfitted / really good loss in your opinion? less than 8 should be good? </p>\n\n<p>Also, about removing the images, I removed them from the <code>train.csv</code> and also the images itselves, it did an error with shapes after that during training. What am I missing?</p>",
          "rawMarkdown": "I see, thank you for the suggestion, what is the expected loss when it's overfitted / really good loss in your opinion? less than 8 should be good? \n\nAlso, about removing the images, I removed them from the `train.csv` and also the images itselves, it did an error with shapes after that during training. What am I missing?",
          "votes": 1
        },
        {
          "id": 720816,
          "postDate": "2020-01-16T19:13:15.340Z",
          "content": "<p>i can't give you a number, as this totally depends on your model-size, your preprocessing, focal-loss or BCE, almost anything. My total loss is ~0.5 at the end of training, but as said above, it's just a number without meaning. \nKeep a validation set and test against that, as soon as the error goes up again, stop the training. You will see that having a good local validation is maybe the most important thing in ML.</p>",
          "rawMarkdown": "i can't give you a number, as this totally depends on your model-size, your preprocessing, focal-loss or BCE, almost anything. My total loss is ~0.5 at the end of training, but as said above, it's just a number without meaning. \nKeep a validation set and test against that, as soon as the error goes up again, stop the training. You will see that having a good local validation is maybe the most important thing in ML.",
          "votes": 1
        },
        {
          "id": 720818,
          "postDate": "2020-01-16T19:15:28.873Z",
          "content": "<p><code>\ntrain = pd.read_csv(PATH + 'train.csv')\nprint(\"Cleaning train set of corrupted image files\")\nprint(\"len train before: \", len(train))\nindexNames = train[(train['ImageId'] == \"ID_1a5a10365\") |\n                   (train['ImageId'] == \"ID_4d238ae90\") |\n                   (train['ImageId'] == \"ID_408f58e9f\") |\n                   (train['ImageId'] == \"ID_bb1d991f6\") |\n                   (train['ImageId'] == \"ID_c44983aeb\")].index\ntrain.drop(indexNames, inplace=True)\nprint(\"len cleaning list: \", len(indexNames))\nprint(\"len train after cleaning: \", len(train), \"\\n\")\n</code></p>\n\n<p>That's what I am doing at the very begining.</p>",
          "rawMarkdown": "```\ntrain = pd.read_csv(PATH + 'train.csv')\nprint(\"Cleaning train set of corrupted image files\")\nprint(\"len train before: \", len(train))\nindexNames = train[(train['ImageId'] == \"ID_1a5a10365\") |\n                   (train['ImageId'] == \"ID_4d238ae90\") |\n                   (train['ImageId'] == \"ID_408f58e9f\") |\n                   (train['ImageId'] == \"ID_bb1d991f6\") |\n                   (train['ImageId'] == \"ID_c44983aeb\")].index\ntrain.drop(indexNames, inplace=True)\nprint(\"len cleaning list: \", len(indexNames))\nprint(\"len train after cleaning: \", len(train), \"\\n\")\n```\n\nThat's what I am doing at the very begining.",
          "votes": 2
        },
        {
          "id": 720821,
          "postDate": "2020-01-16T19:28:21.290Z",
          "content": "<p>Thank you again for all the precious information,\nI'll try the code, looks like that it should work.</p>",
          "rawMarkdown": "Thank you again for all the precious information,\nI'll try the code, looks like that it should work."
        }
      ]
    },
    {
      "id": 710360,
      "postDate": "2020-01-04T16:32:37.727Z",
      "content": "<p>good job 👍 </p>",
      "rawMarkdown": "good job 👍 "
    },
    {
      "id": 674775,
      "postDate": "2019-11-17T03:49:03.483Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 679778,
      "postDate": "2019-11-23T10:55:20.060Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 683676,
      "postDate": "2019-11-28T16:51:47.947Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks"
    },
    {
      "id": 677894,
      "postDate": "2019-11-20T18:19:28.267Z",
      "content": "<p>Thanks for sharing it.</p>",
      "rawMarkdown": "Thanks for sharing it."
    }
  ],
  "comments": [
    {
      "id": 720293,
      "author_name": "Ilan Aizelman",
      "author_url": "",
      "post_date": "2020-01-16T09:54:11.547000",
      "content": "<p><a href=\"/ilu000\">@ilu000</a> Thanks for this post! Question: Did you use resnext50 with pretrained weights for your current score?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 720394,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-01-16T11:38:49.113000",
          "content": "<p>I am training from scratch and found the backbone to be of secondary importance. Replacing resnet/efficientnet or hourglass needs tuning but yields similar results. At least, that's how it turned out for me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 720487,
          "author_name": "Ilan Aizelman",
          "author_url": "",
          "post_date": "2020-01-16T13:22:40.453000",
          "content": "<p>I see, so even with pretrained=False, I should expect after 10 epochs to get LB Score of Top 100?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 720644,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-01-16T16:02:19.903000",
          "content": "<p>I haven't tried to exactly replicate the centernet paper (or adapt the github code for this competition), but I am guessing a 0.070+ score is easily possible without much additional tuning. So yes, top 100 is in reach. \n10 epochs might not be enough, though. I am usually training for a bit more until the nets overfit. Never quit before reaching that state ;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 720809,
          "author_name": "Ilan Aizelman",
          "author_url": "",
          "post_date": "2020-01-16T19:00:24.987000",
          "content": "<p>I see, thank you for the suggestion, what is the expected loss when it's overfitted / really good loss in your opinion? less than 8 should be good? </p>\n\n<p>Also, about removing the images, I removed them from the <code>train.csv</code> and also the images itselves, it did an error with shapes after that during training. What am I missing?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 720816,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-01-16T19:13:15.340000",
          "content": "<p>i can't give you a number, as this totally depends on your model-size, your preprocessing, focal-loss or BCE, almost anything. My total loss is ~0.5 at the end of training, but as said above, it's just a number without meaning. \nKeep a validation set and test against that, as soon as the error goes up again, stop the training. You will see that having a good local validation is maybe the most important thing in ML.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 720818,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-01-16T19:15:28.873000",
          "content": "<p><code>\ntrain = pd.read_csv(PATH + 'train.csv')\nprint(\"Cleaning train set of corrupted image files\")\nprint(\"len train before: \", len(train))\nindexNames = train[(train['ImageId'] == \"ID_1a5a10365\") |\n                   (train['ImageId'] == \"ID_4d238ae90\") |\n                   (train['ImageId'] == \"ID_408f58e9f\") |\n                   (train['ImageId'] == \"ID_bb1d991f6\") |\n                   (train['ImageId'] == \"ID_c44983aeb\")].index\ntrain.drop(indexNames, inplace=True)\nprint(\"len cleaning list: \", len(indexNames))\nprint(\"len train after cleaning: \", len(train), \"\\n\")\n</code></p>\n\n<p>That's what I am doing at the very begining.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 720821,
          "author_name": "Ilan Aizelman",
          "author_url": "",
          "post_date": "2020-01-16T19:28:21.290000",
          "content": "<p>Thank you again for all the precious information,\nI'll try the code, looks like that it should work.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 710360,
      "author_name": "Reza Sadoughi",
      "author_url": "",
      "post_date": "2020-01-04T16:32:37.727000",
      "content": "<p>good job 👍 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 674775,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-17T03:49:03.483000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 679778,
      "author_name": "BlackScreen",
      "author_url": "",
      "post_date": "2019-11-23T10:55:20.060000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 683676,
      "author_name": "Eric Xue",
      "author_url": "",
      "post_date": "2019-11-28T16:51:47.947000",
      "content": "<p>Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 677894,
      "author_name": "Nithesh K G",
      "author_url": "",
      "post_date": "2019-11-20T18:19:28.267000",
      "content": "<p>Thanks for sharing it.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "674626": "Hi everyone.\nA while ago I noticed that some pictures were broken like the one shown below. Downloading the files again didn't solve the problem. Apparently, all those pictures are from the same road.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2675447%2Fdbc409427d2bf9194ef3e3f7c8784110%2Fbroken_pictures.PNG?generation=1573933803767282&amp;alt=media)\n\nThe broken train pictures are:\n```\nID_1a5a10365.jpg\nID_4d238ae90.jpg\nID_408f58e9f.jpg\nID_bb1d991f6.jpg\nID_c44983aeb.jpg\n```\n\nI found those by visual inspection, could have missed some.\nFrom looking at the test set, all pictures there seem to be OK.\n",
    "720293": "@ilu000 Thanks for this post! Question: Did you use resnext50 with pretrained weights for your current score?",
    "710360": "good job 👍 ",
    "674775": "",
    "679778": "Thanks for sharing.",
    "683676": "Thanks",
    "677894": "Thanks for sharing it."
  }
}