{
  "id": 99297,
  "title": "Fruitful dataset, but don't take scores seriously",
  "url": "/competitions/imagenet-object-localization-challenge/discussion/99297",
  "author_name": "Le Yan",
  "post_date": "2019-07-10T07:12:25.906000",
  "votes": 37,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The dataset is very rich with 1000 classes and more than 544k images for training and 50k images for validation and 100k for the test. However, the evaluation metric seems to be very wrongly implemented. Right now we are on the top of the Leaderboard with just a single right prediction and none for all other images. </p>\n\n<p>We essentially followed the protocol of implementing YOLO v3 from <a href=\"https://www.kaggle.com/c/imagenet-object-localization-challenge/discussion/64498#latest-548894\">Ming</a> to initiate our version (the details of our implementation will be summarized in a blog with the link shared later). </p>\n\n<p>We've been training the YOLO v3 on four RTX 2080 for about two weeks with about more than 200 epochs. The validation error has been reduced to around 0.47. We followed the description of the evaluation metric made our own metric and tested on the validation dataset. There we find the score we hit should be about 0.26 with two threshold parameters in the YOLO prediction script set to 0.01, low value to allow many false positive predictions, as the metric described in the evaluation shall take the best prediction from the first five. However, we find the submission score is much worse, at 0.99. Instead, when we set a high threshold at 0.25, the submission score becomes much better with a lot of None predictions. </p>\n\n<p>We immediately noticed that the submission score appears to be strongly correlated with the fraction of None prediction in the submission, rather than how well we have trained the model or parameters we have tuned for better predictions for evaluations. So we finally tested the submission with just one prediction we believe to be true while all other images with None predictions and got the absurd result with score hitting 0.00000. We also just noticed that this problem of evaluation metric had been reported by <a href=\"https://www.kaggle.com/akbargumbira\">Akbar Gumbira</a>, <a href=\"https://www.kaggle.com/yjinnouchi\">yasuyuki</a>, and many other users. </p>\n\n<p>This dataset is very rich, big enough to make it a hard object detection problem. One can play with it with specific models to test. But we here remind you that please don't take the evaluation score seriously. The evaluation metric is not implemented properly as the description. Don't waste your time on fighting for a better score as we have done. Let's see if the organizers will fix this bug.</p>",
  "messages": [
    {
      "id": 571873,
      "postDate": "2019-07-10T07:12:25.907Z",
      "content": "<p>The dataset is very rich with 1000 classes and more than 544k images for training and 50k images for validation and 100k for the test. However, the evaluation metric seems to be very wrongly implemented. Right now we are on the top of the Leaderboard with just a single right prediction and none for all other images. </p>\n\n<p>We essentially followed the protocol of implementing YOLO v3 from <a href=\"https://www.kaggle.com/c/imagenet-object-localization-challenge/discussion/64498#latest-548894\">Ming</a> to initiate our version (the details of our implementation will be summarized in a blog with the link shared later). </p>\n\n<p>We've been training the YOLO v3 on four RTX 2080 for about two weeks with about more than 200 epochs. The validation error has been reduced to around 0.47. We followed the description of the evaluation metric made our own metric and tested on the validation dataset. There we find the score we hit should be about 0.26 with two threshold parameters in the YOLO prediction script set to 0.01, low value to allow many false positive predictions, as the metric described in the evaluation shall take the best prediction from the first five. However, we find the submission score is much worse, at 0.99. Instead, when we set a high threshold at 0.25, the submission score becomes much better with a lot of None predictions. </p>\n\n<p>We immediately noticed that the submission score appears to be strongly correlated with the fraction of None prediction in the submission, rather than how well we have trained the model or parameters we have tuned for better predictions for evaluations. So we finally tested the submission with just one prediction we believe to be true while all other images with None predictions and got the absurd result with score hitting 0.00000. We also just noticed that this problem of evaluation metric had been reported by <a href=\"https://www.kaggle.com/akbargumbira\">Akbar Gumbira</a>, <a href=\"https://www.kaggle.com/yjinnouchi\">yasuyuki</a>, and many other users. </p>\n\n<p>This dataset is very rich, big enough to make it a hard object detection problem. One can play with it with specific models to test. But we here remind you that please don't take the evaluation score seriously. The evaluation metric is not implemented properly as the description. Don't waste your time on fighting for a better score as we have done. Let's see if the organizers will fix this bug.</p>",
      "rawMarkdown": "The dataset is very rich with 1000 classes and more than 544k images for training and 50k images for validation and 100k for the test. However, the evaluation metric seems to be very wrongly implemented. Right now we are on the top of the Leaderboard with just a single right prediction and none for all other images. \n\nWe essentially followed the protocol of implementing YOLO v3 from [Ming](https://www.kaggle.com/c/imagenet-object-localization-challenge/discussion/64498#latest-548894) to initiate our version (the details of our implementation will be summarized in a blog with the link shared later). \n\nWe've been training the YOLO v3 on four RTX 2080 for about two weeks with about more than 200 epochs. The validation error has been reduced to around 0.47. We followed the description of the evaluation metric made our own metric and tested on the validation dataset. There we find the score we hit should be about 0.26 with two threshold parameters in the YOLO prediction script set to 0.01, low value to allow many false positive predictions, as the metric described in the evaluation shall take the best prediction from the first five. However, we find the submission score is much worse, at 0.99. Instead, when we set a high threshold at 0.25, the submission score becomes much better with a lot of None predictions. \n\nWe immediately noticed that the submission score appears to be strongly correlated with the fraction of None prediction in the submission, rather than how well we have trained the model or parameters we have tuned for better predictions for evaluations. So we finally tested the submission with just one prediction we believe to be true while all other images with None predictions and got the absurd result with score hitting 0.00000. We also just noticed that this problem of evaluation metric had been reported by [Akbar Gumbira](https://www.kaggle.com/akbargumbira), [yasuyuki](https://www.kaggle.com/yjinnouchi), and many other users. \n\nThis dataset is very rich, big enough to make it a hard object detection problem. One can play with it with specific models to test. But we here remind you that please don't take the evaluation score seriously. The evaluation metric is not implemented properly as the description. Don't waste your time on fighting for a better score as we have done. Let's see if the organizers will fix this bug.",
      "votes": 36
    },
    {
      "id": 688581,
      "postDate": "2019-12-05T19:08:25.017Z",
      "content": "<p>5 months later and still not fixed?</p>",
      "rawMarkdown": "5 months later and still not fixed?",
      "votes": 3
    },
    {
      "id": 690882,
      "postDate": "2019-12-09T09:43:50.833Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 637078,
      "postDate": "2019-09-30T16:21:50.257Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!"
    }
  ],
  "comments": [
    {
      "id": 688581,
      "author_name": "Chaskin Saroff",
      "author_url": "",
      "post_date": "2019-12-05T19:08:25.017000",
      "content": "<p>5 months later and still not fixed?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 690882,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-09T09:43:50.833000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 637078,
      "author_name": "EvanZamir",
      "author_url": "",
      "post_date": "2019-09-30T16:21:50.257000",
      "content": "<p>Thanks!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "571873": "The dataset is very rich with 1000 classes and more than 544k images for training and 50k images for validation and 100k for the test. However, the evaluation metric seems to be very wrongly implemented. Right now we are on the top of the Leaderboard with just a single right prediction and none for all other images. \n\nWe essentially followed the protocol of implementing YOLO v3 from [Ming](https://www.kaggle.com/c/imagenet-object-localization-challenge/discussion/64498#latest-548894) to initiate our version (the details of our implementation will be summarized in a blog with the link shared later). \n\nWe've been training the YOLO v3 on four RTX 2080 for about two weeks with about more than 200 epochs. The validation error has been reduced to around 0.47. We followed the description of the evaluation metric made our own metric and tested on the validation dataset. There we find the score we hit should be about 0.26 with two threshold parameters in the YOLO prediction script set to 0.01, low value to allow many false positive predictions, as the metric described in the evaluation shall take the best prediction from the first five. However, we find the submission score is much worse, at 0.99. Instead, when we set a high threshold at 0.25, the submission score becomes much better with a lot of None predictions. \n\nWe immediately noticed that the submission score appears to be strongly correlated with the fraction of None prediction in the submission, rather than how well we have trained the model or parameters we have tuned for better predictions for evaluations. So we finally tested the submission with just one prediction we believe to be true while all other images with None predictions and got the absurd result with score hitting 0.00000. We also just noticed that this problem of evaluation metric had been reported by [Akbar Gumbira](https://www.kaggle.com/akbargumbira), [yasuyuki](https://www.kaggle.com/yjinnouchi), and many other users. \n\nThis dataset is very rich, big enough to make it a hard object detection problem. One can play with it with specific models to test. But we here remind you that please don't take the evaluation score seriously. The evaluation metric is not implemented properly as the description. Don't waste your time on fighting for a better score as we have done. Let's see if the organizers will fix this bug.",
    "688581": "5 months later and still not fixed?",
    "690882": "",
    "637078": "Thanks!"
  }
}