{
  "id": 202943,
  "title": "Class distribution in public test set (public LB)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/202943",
  "author_name": "",
  "post_date": "2020-12-12T20:12:30.198862500Z",
  "votes": 32,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello all, below, the class distribution on the public test set (31% of the test data). We can suppose that it will roughly be the same for the private test set. Despite using this information can be seen as overfit, this can nevertheless be used as part of your strategy.</p>\n<table>\n<thead>\n<tr>\n<th>Class 0</th>\n<th>Class 1</th>\n<th>Class 2</th>\n<th>Class 3</th>\n<th>Class 4</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>4.8%</td>\n<td>10.6%</td>\n<td>10.3%</td>\n<td>60.2%</td>\n<td>14.1%</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3234224%2F70b724cd1d5f33ecfa66034e8274c44e%2FCapture%20dcran%20du%202020-12-13%2000-03-17.png?generation=1607803630128051&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1110505",
      "postDate": "12/12/2020 20:12:30",
      "content": "<p>Hello all, below, the class distribution on the public test set (31% of the test data). We can suppose that it will roughly be the same for the private test set. Despite using this information can be seen as overfit, this can nevertheless be used as part of your strategy.</p>\n<table>\n<thead>\n<tr>\n<th>Class 0</th>\n<th>Class 1</th>\n<th>Class 2</th>\n<th>Class 3</th>\n<th>Class 4</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>4.8%</td>\n<td>10.6%</td>\n<td>10.3%</td>\n<td>60.2%</td>\n<td>14.1%</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3234224%2F70b724cd1d5f33ecfa66034e8274c44e%2FCapture%20dcran%20du%202020-12-13%2000-03-17.png?generation=1607803630128051&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hello all, below, the class distribution on the public test set (31% of the test data). We can suppose that it will roughly be the same for the private test set. Despite using this information can be seen as overfit, this can nevertheless be used as part of your strategy.\n\n| Class 0 | Class 1 | Class 2 | Class 3 | Class 4 |\n| --- | --- | --- | --- | --- |\n| 4.8% | 10.6% | 10.3% | 60.2% | 14.1% |\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3234224%2F70b724cd1d5f33ecfa66034e8274c44e%2FCapture%20dcran%20du%202020-12-13%2000-03-17.png?generation=1607803630128051&alt=media)",
      "votes": null
    },
    {
      "id": "1110905",
      "postDate": "12/13/2020 07:32:24",
      "content": "<p>Thanks for this!!! We can  use it as a weight in our weighted loss(  reciprocal  it of course)  and see if our CV or LB increase</p>",
      "rawMarkdown": "Thanks for this!!! We can  use it as a weight in our weighted loss(  reciprocal  it of course)  and see if our CV or LB increase",
      "votes": null
    },
    {
      "id": "1150054",
      "postDate": "01/12/2021 10:23:30",
      "content": "<p><a href=\"https://www.kaggle.com/vincoux\" target=\"_blank\">@vincoux</a> thanks a lot for checking this. It is interesting that the class distribution in the public test is quite similar to the training data, whereas 2019 competition data has a somewhat different distribution.</p>",
      "rawMarkdown": "vincoux thanks a lot for checking this. It is interesting that the class distribution in the public test is quite similar to the training data, whereas 2019 competition data has a somewhat different distribution.",
      "votes": null
    },
    {
      "id": "1151126",
      "postDate": "01/13/2021 06:41:37",
      "content": "<p>how do you know this distribution?</p>",
      "rawMarkdown": "how do you know this distribution?",
      "votes": null
    },
    {
      "id": "1151939",
      "postDate": "01/13/2021 16:51:25",
      "content": "<p><a href=\"https://www.kaggle.com/qixinyan\" target=\"_blank\">@qixinyan</a> the competition metric is accuracy. Submitting constant predictions of the same class achieves the accuracy value that is equal to the ratio of this class in the public test set. The private test set distribution remains unclear though.</p>",
      "rawMarkdown": "qixinyan the competition metric is accuracy. Submitting constant predictions of the same class achieves the accuracy value that is equal to the ratio of this class in the public test set. The private test set distribution remains unclear though.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1110905,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "12/13/2020 07:32:24",
      "content": "<p>Thanks for this!!! We can  use it as a weight in our weighted loss(  reciprocal  it of course)  and see if our CV or LB increase</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1150054,
      "author_name": "kozodoi",
      "author_url": "",
      "post_date": "01/12/2021 10:23:30",
      "content": "<p><a href=\"https://www.kaggle.com/vincoux\" target=\"_blank\">@vincoux</a> thanks a lot for checking this. It is interesting that the class distribution in the public test is quite similar to the training data, whereas 2019 competition data has a somewhat different distribution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1151126,
      "author_name": "qixinyan",
      "author_url": "",
      "post_date": "01/13/2021 06:41:37",
      "content": "<p>how do you know this distribution?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1151939,
          "author_name": "kozodoi",
          "author_url": "",
          "post_date": "01/13/2021 16:51:25",
          "content": "<p><a href=\"https://www.kaggle.com/qixinyan\" target=\"_blank\">@qixinyan</a> the competition metric is accuracy. Submitting constant predictions of the same class achieves the accuracy value that is equal to the ratio of this class in the public test set. The private test set distribution remains unclear though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1110505": "Hello all, below, the class distribution on the public test set (31% of the test data). We can suppose that it will roughly be the same for the private test set. Despite using this information can be seen as overfit, this can nevertheless be used as part of your strategy.\n\n| Class 0 | Class 1 | Class 2 | Class 3 | Class 4 |\n| --- | --- | --- | --- | --- |\n| 4.8% | 10.6% | 10.3% | 60.2% | 14.1% |\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3234224%2F70b724cd1d5f33ecfa66034e8274c44e%2FCapture%20dcran%20du%202020-12-13%2000-03-17.png?generation=1607803630128051&alt=media)",
    "1110905": "Thanks for this!!! We can  use it as a weight in our weighted loss(  reciprocal  it of course)  and see if our CV or LB increase",
    "1150054": "vincoux thanks a lot for checking this. It is interesting that the class distribution in the public test is quite similar to the training data, whereas 2019 competition data has a somewhat different distribution.",
    "1151126": "how do you know this distribution?",
    "1151939": "qixinyan the competition metric is accuracy. Submitting constant predictions of the same class achieves the accuracy value that is equal to the ratio of this class in the public test set. The private test set distribution remains unclear though."
  },
  "source": "meta"
}