{
  "id": 61813,
  "title": "Category distribution of current LB test samples",
  "url": "/competitions/freesound-audio-tagging/discussion/61813",
  "author_name": "",
  "post_date": "2018-07-23T23:29:39.553486400Z",
  "votes": 3,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi organizers, thank you for supporting us.</p>\n\n<p>I have a question about how category distribution of the samples used to score current leaderboard is. LB page shows that:</p>\n\n<blockquote>\n  <p>This leaderboard is calculated with approximately 19% of the test data</p>\n</blockquote>\n\n<p>And data description says:</p>\n\n<blockquote>\n  <p>The test set is composed of ~1.6k samples with manually-verified annotations and with a similar category distribution than that of the train set. </p>\n</blockquote>\n\n<p>I'd really appreciate if we could have clarification for the 19%'s distribution design. Thank you for taking your time.</p>",
  "messages": [
    {
      "id": "361146",
      "postDate": "07/23/2018 23:29:39",
      "content": "<p>Hi organizers, thank you for supporting us.</p>\n\n<p>I have a question about how category distribution of the samples used to score current leaderboard is. LB page shows that:</p>\n\n<blockquote>\n  <p>This leaderboard is calculated with approximately 19% of the test data</p>\n</blockquote>\n\n<p>And data description says:</p>\n\n<blockquote>\n  <p>The test set is composed of ~1.6k samples with manually-verified annotations and with a similar category distribution than that of the train set. </p>\n</blockquote>\n\n<p>I'd really appreciate if we could have clarification for the 19%'s distribution design. Thank you for taking your time.</p>",
      "rawMarkdown": "Hi organizers, thank you for supporting us.\n\nI have a question about how category distribution of the samples used to score current leaderboard is. LB page shows that:\n\n&gt; This leaderboard is calculated with approximately 19% of the test data\n\nAnd data description says:\n\n&gt; The test set is composed of ~1.6k samples with manually-verified annotations and with a similar category distribution than that of the train set. \n\nI'd really appreciate if we could have clarification for the 19%'s distribution design. Thank you for taking your time.",
      "votes": null
    },
    {
      "id": "361255",
      "postDate": "07/24/2018 05:30:18",
      "content": "<p>I noticed this yesterday. The current LB considers only 19% of 1.6K samples. Rest are padded and not counted I believe. In my opinion, that's a very small number to rely on and evaluate :-/</p>",
      "rawMarkdown": "I noticed this yesterday. The current LB considers only 19% of 1.6K samples. Rest are padded and not counted I believe. In my opinion, that's a very small number to rely on and evaluate :-/",
      "votes": null
    },
    {
      "id": "361300",
      "postDate": "07/24/2018 07:42:14",
      "content": "<p>guess it's 19% of 9400, about 1.6k?</p>",
      "rawMarkdown": "guess it's 19% of 9400, about 1.6k?",
      "votes": null
    },
    {
      "id": "361332",
      "postDate": "07/24/2018 08:44:37",
      "content": "<p>I hope so ;) But the latter half of test set explanation has:</p>\n\n<blockquote>\n  <p>The test set is complemented with ~7.8k padding sounds which are not used for scoring the systems.</p>\n</blockquote>\n\n<p>So 1.6k * 19% = 304, only 300 samples is used to evaluate results if my understanding is correct.\nThen it might be very difficult to design/pick 19% so that it has qualitatively/quantitatively same distribution as the 1.6k true-test samples.</p>",
      "rawMarkdown": "I hope so ;) But the latter half of test set explanation has:\n\n&gt; The test set is complemented with ~7.8k padding sounds which are not used for scoring the systems.\n\nSo 1.6k * 19% = 304, only 300 samples is used to evaluate results if my understanding is correct.\nThen it might be very difficult to design/pick 19% so that it has qualitatively/quantitatively same distribution as the 1.6k true-test samples.",
      "votes": null
    },
    {
      "id": "361348",
      "postDate": "07/24/2018 09:30:18",
      "content": "<p>That is exactly what I am afraid of! Plus this number is very small :-/</p>",
      "rawMarkdown": "That is exactly what I am afraid of! Plus this number is very small :-/",
      "votes": null
    },
    {
      "id": "361380",
      "postDate": "07/24/2018 10:51:28",
      "content": "<p>I just noticed it too while I wrote the challenge paper.</p>\n\n<p>What I understand is :</p>\n\n<ul>\n<li><p>Among the test set (9.4k), 1.6k samples are manually-verified and only those are used for evaluating the system. </p></li>\n<li><p>The rest of the test set (7.8k)  don't have any meaning for evaluation. I think that's why organizers use the word of 'padding sounds'.</p></li>\n<li><p>Only about 300 samples (19% of real test samples) have been applied to the public leaderboard.</p></li>\n<li><p>The remaining 1.3k samples (81% of real test samples) will be applied to the private leaderboard which is a final result.</p></li>\n</ul>",
      "rawMarkdown": "I just noticed it too while I wrote the challenge paper.\n\nWhat I understand is :\n\n- Among the test set (9.4k), 1.6k samples are manually-verified and only those are used for evaluating the system. \n\n- The rest of the test set (7.8k)  don't have any meaning for evaluation. I think that's why organizers use the word of 'padding sounds'.\n\n- Only about 300 samples (19% of real test samples) have been applied to the public leaderboard.\n\n- The remaining 1.3k samples (81% of real test samples) will be applied to the private leaderboard which is a final result.",
      "votes": null
    },
    {
      "id": "361692",
      "postDate": "07/25/2018 00:27:07",
      "content": "<p>Hi all, </p>\n\n<p>the statements done by Hyungui Lim in the previous post are correct.</p>\n\n<p>I also want to anticipate that we are finalizing a paper for the DCASE Workshop describing this competition (the task, the dataset and the baseline). The goal of this paper is that you can cite it in your papers when referring to details of the task, dataset, baseline, etc. In this way you can also save space and focus on the description of your systems.</p>\n\n<p>We’ll provide you with the link to the paper in the next days.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi all, \n\nthe statements done by Hyungui Lim in the previous post are correct.\n\nI also want to anticipate that we are finalizing a paper for the DCASE Workshop describing this competition (the task, the dataset and the baseline). The goal of this paper is that you can cite it in your papers when referring to details of the task, dataset, baseline, etc. In this way you can also save space and focus on the description of your systems.\n\nWe’ll provide you with the link to the paper in the next days.\n\nThanks",
      "votes": null
    },
    {
      "id": "361703",
      "postDate": "07/25/2018 01:05:55",
      "content": "<p>Hi Eduardo,</p>\n\n<p>Thank you for your comment.\nMy original question is about category distribution design for LB scoring 19%.</p>\n\n<p>We appreciate if you could make any comment on this.</p>",
      "rawMarkdown": "Hi Eduardo,\n\nThank you for your comment.\nMy original question is about category distribution design for LB scoring 19%.\n\nWe appreciate if you could make any comment on this.",
      "votes": null
    },
    {
      "id": "361791",
      "postDate": "07/25/2018 05:11:01",
      "content": "<p>Hi Daisukelab,</p>\n\n<p>Kaggle and the host teams often make difficult decisions in determining the public/private LB split. This split is often designed to prevent overfitting, or to protect against underrepresented classes. We do not reveal the decision-making process that goes into this, or any distribution of the categories in the public/private split, as it could allow participants to take advantage of the evaluation metric as opposed to developing the most accurate model.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi Daisukelab,\n\nKaggle and the host teams often make difficult decisions in determining the public/private LB split. This split is often designed to prevent overfitting, or to protect against underrepresented classes. We do not reveal the decision-making process that goes into this, or any distribution of the categories in the public/private split, as it could allow participants to take advantage of the evaluation metric as opposed to developing the most accurate model.\n\nThanks!",
      "votes": null
    },
    {
      "id": "361828",
      "postDate": "07/25/2018 06:35:47",
      "content": "<p>Hi Addison,</p>\n\n<p>Thank you for clarification, I understood policy and that's reasonable. The question is coming from the situation that it is quite hard to use LB score as metrics for choosing final submission...</p>\n\n<p>Thanks again for spending time for explanation.</p>",
      "rawMarkdown": "Hi Addison,\n\nThank you for clarification, I understood policy and that's reasonable. The question is coming from the situation that it is quite hard to use LB score as metrics for choosing final submission...\n\nThanks again for spending time for explanation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 361255,
      "author_name": "gyat2017",
      "author_url": "",
      "post_date": "07/24/2018 05:30:18",
      "content": "<p>I noticed this yesterday. The current LB considers only 19% of 1.6K samples. Rest are padded and not counted I believe. In my opinion, that's a very small number to rely on and evaluate :-/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 361300,
      "author_name": "sailorwei",
      "author_url": "",
      "post_date": "07/24/2018 07:42:14",
      "content": "<p>guess it's 19% of 9400, about 1.6k?</p>",
      "votes": null,
      "replies": [
        {
          "id": 361332,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "07/24/2018 08:44:37",
          "content": "<p>I hope so ;) But the latter half of test set explanation has:</p>\n\n<blockquote>\n  <p>The test set is complemented with ~7.8k padding sounds which are not used for scoring the systems.</p>\n</blockquote>\n\n<p>So 1.6k * 19% = 304, only 300 samples is used to evaluate results if my understanding is correct.\nThen it might be very difficult to design/pick 19% so that it has qualitatively/quantitatively same distribution as the 1.6k true-test samples.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 361348,
          "author_name": "gyat2017",
          "author_url": "",
          "post_date": "07/24/2018 09:30:18",
          "content": "<p>That is exactly what I am afraid of! Plus this number is very small :-/</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 361380,
          "author_name": "hyunguilim",
          "author_url": "",
          "post_date": "07/24/2018 10:51:28",
          "content": "<p>I just noticed it too while I wrote the challenge paper.</p>\n\n<p>What I understand is :</p>\n\n<ul>\n<li><p>Among the test set (9.4k), 1.6k samples are manually-verified and only those are used for evaluating the system. </p></li>\n<li><p>The rest of the test set (7.8k)  don't have any meaning for evaluation. I think that's why organizers use the word of 'padding sounds'.</p></li>\n<li><p>Only about 300 samples (19% of real test samples) have been applied to the public leaderboard.</p></li>\n<li><p>The remaining 1.3k samples (81% of real test samples) will be applied to the private leaderboard which is a final result.</p></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 361692,
          "author_name": "eduardofonseca",
          "author_url": "",
          "post_date": "07/25/2018 00:27:07",
          "content": "<p>Hi all, </p>\n\n<p>the statements done by Hyungui Lim in the previous post are correct.</p>\n\n<p>I also want to anticipate that we are finalizing a paper for the DCASE Workshop describing this competition (the task, the dataset and the baseline). The goal of this paper is that you can cite it in your papers when referring to details of the task, dataset, baseline, etc. In this way you can also save space and focus on the description of your systems.</p>\n\n<p>We’ll provide you with the link to the paper in the next days.</p>\n\n<p>Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 361703,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "07/25/2018 01:05:55",
          "content": "<p>Hi Eduardo,</p>\n\n<p>Thank you for your comment.\nMy original question is about category distribution design for LB scoring 19%.</p>\n\n<p>We appreciate if you could make any comment on this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 361791,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "07/25/2018 05:11:01",
          "content": "<p>Hi Daisukelab,</p>\n\n<p>Kaggle and the host teams often make difficult decisions in determining the public/private LB split. This split is often designed to prevent overfitting, or to protect against underrepresented classes. We do not reveal the decision-making process that goes into this, or any distribution of the categories in the public/private split, as it could allow participants to take advantage of the evaluation metric as opposed to developing the most accurate model.</p>\n\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 361828,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "07/25/2018 06:35:47",
          "content": "<p>Hi Addison,</p>\n\n<p>Thank you for clarification, I understood policy and that's reasonable. The question is coming from the situation that it is quite hard to use LB score as metrics for choosing final submission...</p>\n\n<p>Thanks again for spending time for explanation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "361146": "Hi organizers, thank you for supporting us.\n\nI have a question about how category distribution of the samples used to score current leaderboard is. LB page shows that:\n\n&gt; This leaderboard is calculated with approximately 19% of the test data\n\nAnd data description says:\n\n&gt; The test set is composed of ~1.6k samples with manually-verified annotations and with a similar category distribution than that of the train set. \n\nI'd really appreciate if we could have clarification for the 19%'s distribution design. Thank you for taking your time.",
    "361255": "I noticed this yesterday. The current LB considers only 19% of 1.6K samples. Rest are padded and not counted I believe. In my opinion, that's a very small number to rely on and evaluate :-/",
    "361300": "guess it's 19% of 9400, about 1.6k?",
    "361332": "I hope so ;) But the latter half of test set explanation has:\n\n&gt; The test set is complemented with ~7.8k padding sounds which are not used for scoring the systems.\n\nSo 1.6k * 19% = 304, only 300 samples is used to evaluate results if my understanding is correct.\nThen it might be very difficult to design/pick 19% so that it has qualitatively/quantitatively same distribution as the 1.6k true-test samples.",
    "361348": "That is exactly what I am afraid of! Plus this number is very small :-/",
    "361380": "I just noticed it too while I wrote the challenge paper.\n\nWhat I understand is :\n\n- Among the test set (9.4k), 1.6k samples are manually-verified and only those are used for evaluating the system. \n\n- The rest of the test set (7.8k)  don't have any meaning for evaluation. I think that's why organizers use the word of 'padding sounds'.\n\n- Only about 300 samples (19% of real test samples) have been applied to the public leaderboard.\n\n- The remaining 1.3k samples (81% of real test samples) will be applied to the private leaderboard which is a final result.",
    "361692": "Hi all, \n\nthe statements done by Hyungui Lim in the previous post are correct.\n\nI also want to anticipate that we are finalizing a paper for the DCASE Workshop describing this competition (the task, the dataset and the baseline). The goal of this paper is that you can cite it in your papers when referring to details of the task, dataset, baseline, etc. In this way you can also save space and focus on the description of your systems.\n\nWe’ll provide you with the link to the paper in the next days.\n\nThanks",
    "361703": "Hi Eduardo,\n\nThank you for your comment.\nMy original question is about category distribution design for LB scoring 19%.\n\nWe appreciate if you could make any comment on this.",
    "361791": "Hi Daisukelab,\n\nKaggle and the host teams often make difficult decisions in determining the public/private LB split. This split is often designed to prevent overfitting, or to protect against underrepresented classes. We do not reveal the decision-making process that goes into this, or any distribution of the categories in the public/private split, as it could allow participants to take advantage of the evaluation metric as opposed to developing the most accurate model.\n\nThanks!",
    "361828": "Hi Addison,\n\nThank you for clarification, I understood policy and that's reasonable. The question is coming from the situation that it is quite hard to use LB score as metrics for choosing final submission...\n\nThanks again for spending time for explanation."
  },
  "source": "meta"
}