{
  "id": 217288,
  "title": "Leaderboard, scores, thoughts",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/217288",
  "author_name": "Krishna Kishor Kammaje",
  "post_date": "2021-02-06T07:28:36.048000",
  "votes": 17,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Since accuracy is the evaluation metric here, it becomes easy to interpret it and compare the various scores.</p>\n<p>Out of 3500+ teams, Out of every 100 images in the public dataset,</p>\n<ul>\n<li>2600+ teams have got 86 images classified right.</li>\n<li>2500+ teams have got 87 images classified right.</li>\n<li>2300+ teams have got 88 images classified right.</li>\n<li>2000+ teams have got 89 images classified right.</li>\n<li>900+ teams have got 90 images classified right.</li>\n<li>5 teams have got 91 images right.</li>\n<li>0 teams have got 92 images right.</li>\n</ul>\n<p>These counts are with some reasonable approximations.  It is expected that as we get close to the 90s, the competition will get tougher. </p>\n<p>You can note that by classifying <strong>one</strong> extra image 'right', your ranking can improvise from 1000+ to single digits. But considering the noise in the dataset is it really worth going after getting one extra image 'right'?. Otherwise, we are definitely trying to model/overfit to the noise which is not going to be useful later in the real world.</p>\n<p>I always thought organizers should have talked about the data quality in the public dataset (compared to the train set) and by having a better (lesser label noise) public/private set they could have got a more real-world useful model at the end of the competition. </p>",
  "messages": [
    {
      "id": 1188371,
      "postDate": "2021-02-06T07:28:36.047Z",
      "content": "<p>Since accuracy is the evaluation metric here, it becomes easy to interpret it and compare the various scores.</p>\n<p>Out of 3500+ teams, Out of every 100 images in the public dataset,</p>\n<ul>\n<li>2600+ teams have got 86 images classified right.</li>\n<li>2500+ teams have got 87 images classified right.</li>\n<li>2300+ teams have got 88 images classified right.</li>\n<li>2000+ teams have got 89 images classified right.</li>\n<li>900+ teams have got 90 images classified right.</li>\n<li>5 teams have got 91 images right.</li>\n<li>0 teams have got 92 images right.</li>\n</ul>\n<p>These counts are with some reasonable approximations.  It is expected that as we get close to the 90s, the competition will get tougher. </p>\n<p>You can note that by classifying <strong>one</strong> extra image 'right', your ranking can improvise from 1000+ to single digits. But considering the noise in the dataset is it really worth going after getting one extra image 'right'?. Otherwise, we are definitely trying to model/overfit to the noise which is not going to be useful later in the real world.</p>\n<p>I always thought organizers should have talked about the data quality in the public dataset (compared to the train set) and by having a better (lesser label noise) public/private set they could have got a more real-world useful model at the end of the competition. </p>",
      "rawMarkdown": "Since accuracy is the evaluation metric here, it becomes easy to interpret it and compare the various scores.\n\nOut of 3500+ teams, Out of every 100 images in the public dataset,\n- 2600+ teams have got 86 images classified right.\n- 2500+ teams have got 87 images classified right.\n- 2300+ teams have got 88 images classified right.\n- 2000+ teams have got 89 images classified right.\n- 900+ teams have got 90 images classified right.\n- 5 teams have got 91 images right.\n- 0 teams have got 92 images right.\n\nThese counts are with some reasonable approximations.  It is expected that as we get close to the 90s, the competition will get tougher. \n\nYou can note that by classifying **one** extra image 'right', your ranking can improvise from 1000+ to single digits. But considering the noise in the dataset is it really worth going after getting one extra image 'right'?. Otherwise, we are definitely trying to model/overfit to the noise which is not going to be useful later in the real world.\n\nI always thought organizers should have talked about the data quality in the public dataset (compared to the train set) and by having a better (lesser label noise) public/private set they could have got a more real-world useful model at the end of the competition. ",
      "votes": 17
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1188371": "Since accuracy is the evaluation metric here, it becomes easy to interpret it and compare the various scores.\n\nOut of 3500+ teams, Out of every 100 images in the public dataset,\n- 2600+ teams have got 86 images classified right.\n- 2500+ teams have got 87 images classified right.\n- 2300+ teams have got 88 images classified right.\n- 2000+ teams have got 89 images classified right.\n- 900+ teams have got 90 images classified right.\n- 5 teams have got 91 images right.\n- 0 teams have got 92 images right.\n\nThese counts are with some reasonable approximations.  It is expected that as we get close to the 90s, the competition will get tougher. \n\nYou can note that by classifying **one** extra image 'right', your ranking can improvise from 1000+ to single digits. But considering the noise in the dataset is it really worth going after getting one extra image 'right'?. Otherwise, we are definitely trying to model/overfit to the noise which is not going to be useful later in the real world.\n\nI always thought organizers should have talked about the data quality in the public dataset (compared to the train set) and by having a better (lesser label noise) public/private set they could have got a more real-world useful model at the end of the competition. "
  }
}