{
  "id": 58052,
  "title": "Self-verification of non-verified audio.",
  "url": "/competitions/freesound-audio-tagging/discussion/58052",
  "author_name": "",
  "post_date": "2018-06-01T13:37:21.442258800Z",
  "votes": 2,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hi, I have a question about 'non-verified' data.</p>\n\n<p>I believe self-hand-verification of non-verified data is prohibited for fair competition even if it is training data. However, I couldn't find the rule about it.</p>\n\n<p>Non-verified data accounts for more than 60% of training data, and it seems to be more effective to use the entire data than to use only verified data. Since the non-verified data is still quite noisy as shown in data description, I think sampling a clean subset of non-verified data can make further improvement. In this situation, I think the sampling process should be purely automatic without any person's decision. But I have no idea whether this process is available or not.</p>\n\n<p>So, I hope organizers clarify the rule about 'non-verified' data ASAP.\nThanks.</p>",
  "messages": [
    {
      "id": "336883",
      "postDate": "06/01/2018 13:37:21",
      "content": "<p>Hi, I have a question about 'non-verified' data.</p>\n\n<p>I believe self-hand-verification of non-verified data is prohibited for fair competition even if it is training data. However, I couldn't find the rule about it.</p>\n\n<p>Non-verified data accounts for more than 60% of training data, and it seems to be more effective to use the entire data than to use only verified data. Since the non-verified data is still quite noisy as shown in data description, I think sampling a clean subset of non-verified data can make further improvement. In this situation, I think the sampling process should be purely automatic without any person's decision. But I have no idea whether this process is available or not.</p>\n\n<p>So, I hope organizers clarify the rule about 'non-verified' data ASAP.\nThanks.</p>",
      "rawMarkdown": "Hi, I have a question about 'non-verified' data.\n\nI believe self-hand-verification of non-verified data is prohibited for fair competition even if it is training data. However, I couldn't find the rule about it.\n\nNon-verified data accounts for more than 60% of training data, and it seems to be more effective to use the entire data than to use only verified data. Since the non-verified data is still quite noisy as shown in data description, I think sampling a clean subset of non-verified data can make further improvement. In this situation, I think the sampling process should be purely automatic without any person's decision. But I have no idea whether this process is available or not.\n\nSo, I hope organizers clarify the rule about 'non-verified' data ASAP.\nThanks.",
      "votes": null
    },
    {
      "id": "337983",
      "postDate": "06/04/2018 07:02:33",
      "content": "<p>Hi, I also have related question about pseudo labeling.\nI (maybe we) 'd appreciate if organizers could clarify use of pseudo labeling.\nIt would not be manually hand labeled, but automatically labeled by first model.\nThen final purpose is to train better model by using pseudo-labeled test samples or non-verified training samples.\nThanks in advance.</p>",
      "rawMarkdown": "Hi, I also have related question about pseudo labeling.\nI (maybe we) 'd appreciate if organizers could clarify use of pseudo labeling.\nIt would not be manually hand labeled, but automatically labeled by first model.\nThen final purpose is to train better model by using pseudo-labeled test samples or non-verified training samples.\nThanks in advance.",
      "votes": null
    },
    {
      "id": "338428",
      "postDate": "06/05/2018 03:27:56",
      "content": "<p>I also get the same question. And I think if there's no rules, maybe someone may check and label the non-verified samples hand-by-hand.</p>",
      "rawMarkdown": "I also get the same question. And I think if there's no rules, maybe someone may check and label the non-verified samples hand-by-hand.",
      "votes": null
    },
    {
      "id": "338560",
      "postDate": "06/05/2018 10:11:40",
      "content": "<p>Hi, we are discussing among the organizers which is the best way to address this issue, and we'll provide clarification ASAP. Thanks!</p>",
      "rawMarkdown": "Hi, we are discussing among the organizers which is the best way to address this issue, and we'll provide clarification ASAP. Thanks!",
      "votes": null
    },
    {
      "id": "339916",
      "postDate": "06/08/2018 00:37:02",
      "content": "<p>Hi all, </p>\n\n<p>thanks for the patience. After discussing among the organizers, we've decided that:</p>\n\n<ul>\n<li>automatic re-labeling of the train set (e.g., of the non-verified portion) in order to refine the training data is allowed.</li>\n<li>manual re-labeling of the train set (e.g., of the non-verified portion) is also allowed, on the condition that it is clearly specified somewhere, e.g., in the submission.</li>\n</ul>\n\n<p>That being said, we have two additional comments:</p>\n\n<ul>\n<li>we encourage automatic (rather than manual) re-labeling, as it can imply certain methodological/algorithmic innovation</li>\n<li>if any participant decides to carry out manual re-labeling in the train set, and <strong>only after the challenge is concluded</strong>, we encourage her/him to release the new labels (or at least share them with the organizers). In this way, the labels could be considered for future dataset versions and thus benefit the community.</li>\n</ul>\n\n<p>Hope this clarifies! If you have any further questions, do not hesitate to ask</p>",
      "rawMarkdown": "Hi all, \n\nthanks for the patience. After discussing among the organizers, we've decided that:\n\n- automatic re-labeling of the train set (e.g., of the non-verified portion) in order to refine the training data is allowed.\n- manual re-labeling of the train set (e.g., of the non-verified portion) is also allowed, on the condition that it is clearly specified somewhere, e.g., in the submission.\n\nThat being said, we have two additional comments:\n\n- we encourage automatic (rather than manual) re-labeling, as it can imply certain methodological/algorithmic innovation\n- if any participant decides to carry out manual re-labeling in the train set, and **only after the challenge is concluded**, we encourage her/him to release the new labels (or at least share them with the organizers). In this way, the labels could be considered for future dataset versions and thus benefit the community.\n\nHope this clarifies! If you have any further questions, do not hesitate to ask",
      "votes": null
    },
    {
      "id": "339969",
      "postDate": "06/08/2018 03:31:59",
      "content": "<p>Thank you for clarification and encourage our innovation.\nLet me confirm about test set.\nI'd like also perform automatic pseudo-labeling test set, and use it to fine tune existing model.\nIs this allowed?</p>",
      "rawMarkdown": "Thank you for clarification and encourage our innovation.\nLet me confirm about test set.\nI'd like also perform automatic pseudo-labeling test set, and use it to fine tune existing model.\nIs this allowed?",
      "votes": null
    },
    {
      "id": "340031",
      "postDate": "06/08/2018 07:28:57",
      "content": "<p>I think test set should not be included in the model although the rule only prohibits hand validation of test set. This process breaks the general assumption that test set is completely invisible.</p>",
      "rawMarkdown": "I think test set should not be included in the model although the rule only prohibits hand validation of test set. This process breaks the general assumption that test set is completely invisible.",
      "votes": null
    },
    {
      "id": "340050",
      "postDate": "06/08/2018 08:32:19",
      "content": "<p>Thanks for your comment.\nActually recently I already have tried some, just because it's normal to do the pseudo labeling in kaggle competitions and not specifically restricted by this competition rules.\nBecause your question was a little similar to pseudo labeling, then I wanted to make sure that it is OK.\nSo can I wait for their clarification? :)</p>\n\n<p>NOTE: My attempts with pseudo labeling so far was ended with all degraded results... Current my best is NOT using it.</p>\n\n<p>BTW I was wondering if pseudo labeling is useful in real use-case or not. And I also have concluded that it is useful as same as what we can read in some pseudo labeling literatures like [1] [2].</p>\n\n<p>Considering real applications, many cases don't have much samples initially. After we get application up and running in (test) operation, we collect samples. We then need to label them to re-train model to match to the real sample distribution. We will repeat this cycle then. Pseudo-labeling will be useful for automatic cycle of this kind of system. So I'd like to learn how to make use of it, and this competition's datasets seem to be good learning materials - it seems not to be easy to use it for further improvements.</p>\n\n<p>[1] Semi-Supervised Deep Learning Using Pseudo Labels for Hyperspectral Image Classification, Hao Wu et al, 2017\n[2] Pseudo-Label : The Simple and Efficient Semi-Supervised Learning Method for Deep Neural Networks, Dong-Hyun Lee, 2013</p>",
      "rawMarkdown": "Thanks for your comment.\nActually recently I already have tried some, just because it's normal to do the pseudo labeling in kaggle competitions and not specifically restricted by this competition rules.\nBecause your question was a little similar to pseudo labeling, then I wanted to make sure that it is OK.\nSo can I wait for their clarification? :)\n\nNOTE: My attempts with pseudo labeling so far was ended with all degraded results... Current my best is NOT using it.\n\nBTW I was wondering if pseudo labeling is useful in real use-case or not. And I also have concluded that it is useful as same as what we can read in some pseudo labeling literatures like [1] [2].\n\nConsidering real applications, many cases don't have much samples initially. After we get application up and running in (test) operation, we collect samples. We then need to label them to re-train model to match to the real sample distribution. We will repeat this cycle then. Pseudo-labeling will be useful for automatic cycle of this kind of system. So I'd like to learn how to make use of it, and this competition's datasets seem to be good learning materials - it seems not to be easy to use it for further improvements.\n\n[1] Semi-Supervised Deep Learning Using Pseudo Labels for Hyperspectral Image Classification, Hao Wu et al, 2017\n[2] Pseudo-Label : The Simple and Efficient Semi-Supervised Learning Method for Deep Neural Networks, Dong-Hyun Lee, 2013",
      "votes": null
    },
    {
      "id": "340073",
      "postDate": "06/08/2018 09:27:17",
      "content": "<p>Thanks for the clarification, Eduardo.\nI think it's kind of overfit to the test set ,some kind of treat.<a href=\"/daisukelab\">@daisukelab</a></p>",
      "rawMarkdown": "Thanks for the clarification, Eduardo.\nI think it's kind of overfit to the test set ,some kind of treat.@daisukelab",
      "votes": null
    },
    {
      "id": "340098",
      "postDate": "06/08/2018 10:26:22",
      "content": "<p>@LEI.Jinggui Thanks for your comment. If it's kind of overfit to test set, I guess it would lead to better score.\nAnyway as I was observing samples selected by my models, I found something, more fundamental issue.\nI will write about it when I could confirm...</p>",
      "rawMarkdown": "LEI.Jinggui Thanks for your comment. If it's kind of overfit to test set, I guess it would lead to better score.\nAnyway as I was observing samples selected by my models, I found something, more fundamental issue.\nI will write about it when I could confirm...",
      "votes": null
    },
    {
      "id": "340611",
      "postDate": "06/09/2018 19:08:25",
      "content": "<p>Hi again,</p>\n\n<p>providing <strong>clarification about the test set usage</strong>:</p>\n\n<ul>\n<li>Automatic (and of course manual) labeling of the test set is <strong>not</strong> allowed. In the context of this competition, we do <strong>not</strong> allow the test set as input into the model in any form.</li>\n</ul>\n\n<p>As usual, please ask any questions that you may have. <a href=\"/daisukelab\">@daisukelab</a>, thanks for raising this issue, which was not properly specified in the competition rules.</p>",
      "rawMarkdown": "Hi again,\n\nproviding **clarification about the test set usage**:\n\n - Automatic (and of course manual) labeling of the test set is **not** allowed. In the context of this competition, we do **not** allow the test set as input into the model in any form.\n\nAs usual, please ask any questions that you may have. @daisukelab, thanks for raising this issue, which was not properly specified in the competition rules.",
      "votes": null
    },
    {
      "id": "340665",
      "postDate": "06/09/2018 23:31:48",
      "content": "<p>Hi @Eduardo, thank you for taking time for clarification!</p>",
      "rawMarkdown": "Hi @Eduardo, thank you for taking time for clarification!",
      "votes": null
    },
    {
      "id": "345980",
      "postDate": "06/20/2018 21:18:35",
      "content": "<p>Hi, <a href=\"/daisukelab\">@daisukelab</a>, thanks for sharing your experience. </p>\n\n<p>I also tried pseudo labeling and it also gave lower score on LB. To be precise, what I did was to train the models, predict labels on test set, take samples with predicted probability higher then, say, 0.95, add these samples to train, re-train, etc. Surprisingly no improvement, especially because probability is taken as geometric mean of 4 classifiers, two of which are ensembles of NN. So, given high score on LB I would assume that top-predictions should have correct labels. I am 100% sure that same is applicable to your work.</p>\n\n<p>I also had an idea of hand-labeling training samples with lowest predicted probability, but haven't got time to do it properly. However, it seems that it might NOT bring added value, because labels are assigned in the same way on train and test, so if they are partly assigned in the wrong way on train, also on test. So there should be integrity around labels which might not make it work. Some support to that assumption is given by the fact that adding train samples which are not verified increases LB score.</p>\n\n<p>I also checked some of these TEST samples with lowest predicted probability and most of them are so called padding sounds - music, voices, water, birds, animals, etc. So hand-labelling them also shouldn't increase score cause they are not part of the competition(some details here <a href=\"https://www.kaggle.com/c/freesound-audio-tagging/discussion/58862\">https://www.kaggle.com/c/freesound-audio-tagging/discussion/58862</a>) </p>",
      "rawMarkdown": "Hi, @daisukelab, thanks for sharing your experience. \n\nI also tried pseudo labeling and it also gave lower score on LB. To be precise, what I did was to train the models, predict labels on test set, take samples with predicted probability higher then, say, 0.95, add these samples to train, re-train, etc. Surprisingly no improvement, especially because probability is taken as geometric mean of 4 classifiers, two of which are ensembles of NN. So, given high score on LB I would assume that top-predictions should have correct labels. I am 100% sure that same is applicable to your work.\n\nI also had an idea of hand-labeling training samples with lowest predicted probability, but haven't got time to do it properly. However, it seems that it might NOT bring added value, because labels are assigned in the same way on train and test, so if they are partly assigned in the wrong way on train, also on test. So there should be integrity around labels which might not make it work. Some support to that assumption is given by the fact that adding train samples which are not verified increases LB score.\n\nI also checked some of these TEST samples with lowest predicted probability and most of them are so called padding sounds - music, voices, water, birds, animals, etc. So hand-labelling them also shouldn't increase score cause they are not part of the competition(some details here https://www.kaggle.com/c/freesound-audio-tagging/discussion/58862)",
      "votes": null
    },
    {
      "id": "345987",
      "postDate": "06/20/2018 21:32:01",
      "content": "<p>To summarize, I would say that we should have been given more daily submissions. It would help to study problem much better, cause as it is now, we are very restricted to checking ideas and verifying assumptions. Especially because the aim of the competition is research.</p>\n\n<p>There definitely are wrongly labelled train samples and as pseudo-labelling does not work, we are stuck in the dead-end. That's why I was claiming in other comments that LB scores between 0.92 and 0.96 are top-results and probably produced by very similar models. Which means that we can solve that task by achieving mapk of around 0.95, but to achieve better results, we need to pseudo- or hand- label BOTH train and test. </p>",
      "rawMarkdown": "To summarize, I would say that we should have been given more daily submissions. It would help to study problem much better, cause as it is now, we are very restricted to checking ideas and verifying assumptions. Especially because the aim of the competition is research.\n\nThere definitely are wrongly labelled train samples and as pseudo-labelling does not work, we are stuck in the dead-end. That's why I was claiming in other comments that LB scores between 0.92 and 0.96 are top-results and probably produced by very similar models. Which means that we can solve that task by achieving mapk of around 0.95, but to achieve better results, we need to pseudo- or hand- label BOTH train and test.",
      "votes": null
    },
    {
      "id": "346004",
      "postDate": "06/20/2018 22:03:54",
      "content": "<p>Hi, I just wanted to clarify the <strong>allowed usage of the test set</strong>. As already stated in this thread:</p>\n\n<ul>\n<li>Automatic (and of course manual) labeling of the test set is <strong>not</strong> allowed. In the context of this competition, we do <strong>not</strong> allow the test set as input into the model in any form.</li>\n</ul>\n\n<p><a href=\"https://www.kaggle.com/c/freesound-audio-tagging/discussion/58052#340611\">https://www.kaggle.com/c/freesound-audio-tagging/discussion/58052#340611</a></p>\n\n<p>In other words, the test set should ONLY be used for reporting model's performance, and nothing else. Therefore, the approaches you were describing are forbidden in this competition.</p>",
      "rawMarkdown": "Hi, I just wanted to clarify the **allowed usage of the test set**. As already stated in this thread:\n\n - Automatic (and of course manual) labeling of the test set is **not** allowed. In the context of this competition, we do **not** allow the test set as input into the model in any form.\n\nhttps://www.kaggle.com/c/freesound-audio-tagging/discussion/58052#340611\n\nIn other words, the test set should ONLY be used for reporting model's performance, and nothing else. Therefore, the approaches you were describing are forbidden in this competition.",
      "votes": null
    },
    {
      "id": "346044",
      "postDate": "06/21/2018 01:10:25",
      "content": "<p>Hi @Aleksandrs,</p>\n\n<p>'#metoo'</p>\n\n<p>Thank you for sharing the detail. I agree with you 100% regarding pseudo labeling attempts for both train and test samples, yes I did almost the same, and experienced the same things. So I’m glad to hear your detail, I really felt that I’m not alone on the planet. :)</p>\n\n<p>I have distilled a trustful set of training samples by 4-5 times cycle of picking up samples of higher probabilities like &gt;0.92, and yes this attempt failed in vain. It proved that some samples are correctly labeled though they don't sound as labeled nor have higher probabilities.</p>\n\n<p>Then I almost came to understand that final goal of this competition would be modeling the real distribution of what have been done to label the dataset.</p>\n\n<p>Hand checking for the detailed distribution could show something, though I'm almost losing motivation to do that, just because this action might not be helpful to gain knowledge...\nBut finding the real distribution would be a good training, trying myself to keep up. As long as other people have better score, there should be something else to do... :)</p>",
      "rawMarkdown": "Hi @Aleksandrs,\n\n'#metoo'\n\nThank you for sharing the detail. I agree with you 100% regarding pseudo labeling attempts for both train and test samples, yes I did almost the same, and experienced the same things. So I’m glad to hear your detail, I really felt that I’m not alone on the planet. :)\n\nI have distilled a trustful set of training samples by 4-5 times cycle of picking up samples of higher probabilities like &gt;0.92, and yes this attempt failed in vain. It proved that some samples are correctly labeled though they don't sound as labeled nor have higher probabilities.\n\nThen I almost came to understand that final goal of this competition would be modeling the real distribution of what have been done to label the dataset.\n\nHand checking for the detailed distribution could show something, though I'm almost losing motivation to do that, just because this action might not be helpful to gain knowledge...\nBut finding the real distribution would be a good training, trying myself to keep up. As long as other people have better score, there should be something else to do... :)",
      "votes": null
    },
    {
      "id": "346524",
      "postDate": "06/21/2018 22:01:47",
      "content": "<p>@Eduardo Fonseca yes, it is 100% clear that it is not allowed. I used approach which I described more than a month ago before it was clear that it is not allowed. I am not using it anymore as it is restrited and I will not use these old submissions. Moreover, as I described above, that approach gave me lower score on LB, so I wouldn't use these submissions even if I wanted to cheat.</p>",
      "rawMarkdown": "Eduardo Fonseca yes, it is 100% clear that it is not allowed. I used approach which I described more than a month ago before it was clear that it is not allowed. I am not using it anymore as it is restrited and I will not use these old submissions. Moreover, as I described above, that approach gave me lower score on LB, so I wouldn't use these submissions even if I wanted to cheat.",
      "votes": null
    },
    {
      "id": "346529",
      "postDate": "06/21/2018 22:14:20",
      "content": "<p><a href=\"/daisukelab\">@daisukelab</a> I agree - after achieving high scores aim of that competition becomes learning the way how datasets were labelled, not learning smt meaningful. However, organizers say that \"<strong>the main goal of this competition is to foster and advance research in sound recognition</strong>\".  I am just losing motivation to do smt extra, partly because I am not sure that improvement on LB will be a \"real\" improvement, partly because of only two submissions. And there is smt to try - LSTMs, RNNs and HMMs. </p>",
      "rawMarkdown": "daisukelab I agree - after achieving high scores aim of that competition becomes learning the way how datasets were labelled, not learning smt meaningful. However, organizers say that \"**the main goal of this competition is to foster and advance research in sound recognition**\".  I am just losing motivation to do smt extra, partly because I am not sure that improvement on LB will be a \"real\" improvement, partly because of only two submissions. And there is smt to try - LSTMs, RNNs and HMMs.",
      "votes": null
    },
    {
      "id": "346559",
      "postDate": "06/22/2018 00:03:12",
      "content": "<p>Hi @Aleksandrs, </p>\n\n<p>great, we wanted to clarify the issue again in case another participant read your post (but not the other post where we clarify the allowed usage of the test set ). Since, as you pointed out, this is not specified in the competition Rules, we wanted to avoid potential confusions.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi @Aleksandrs, \n\ngreat, we wanted to clarify the issue again in case another participant read your post (but not the other post where we clarify the allowed usage of the test set ). Since, as you pointed out, this is not specified in the competition Rules, we wanted to avoid potential confusions.\n\nThanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 337983,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "06/04/2018 07:02:33",
      "content": "<p>Hi, I also have related question about pseudo labeling.\nI (maybe we) 'd appreciate if organizers could clarify use of pseudo labeling.\nIt would not be manually hand labeled, but automatically labeled by first model.\nThen final purpose is to train better model by using pseudo-labeled test samples or non-verified training samples.\nThanks in advance.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 338428,
      "author_name": "lccever",
      "author_url": "",
      "post_date": "06/05/2018 03:27:56",
      "content": "<p>I also get the same question. And I think if there's no rules, maybe someone may check and label the non-verified samples hand-by-hand.</p>",
      "votes": null,
      "replies": [
        {
          "id": 338560,
          "author_name": "eduardofonseca",
          "author_url": "",
          "post_date": "06/05/2018 10:11:40",
          "content": "<p>Hi, we are discussing among the organizers which is the best way to address this issue, and we'll provide clarification ASAP. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 339916,
      "author_name": "eduardofonseca",
      "author_url": "",
      "post_date": "06/08/2018 00:37:02",
      "content": "<p>Hi all, </p>\n\n<p>thanks for the patience. After discussing among the organizers, we've decided that:</p>\n\n<ul>\n<li>automatic re-labeling of the train set (e.g., of the non-verified portion) in order to refine the training data is allowed.</li>\n<li>manual re-labeling of the train set (e.g., of the non-verified portion) is also allowed, on the condition that it is clearly specified somewhere, e.g., in the submission.</li>\n</ul>\n\n<p>That being said, we have two additional comments:</p>\n\n<ul>\n<li>we encourage automatic (rather than manual) re-labeling, as it can imply certain methodological/algorithmic innovation</li>\n<li>if any participant decides to carry out manual re-labeling in the train set, and <strong>only after the challenge is concluded</strong>, we encourage her/him to release the new labels (or at least share them with the organizers). In this way, the labels could be considered for future dataset versions and thus benefit the community.</li>\n</ul>\n\n<p>Hope this clarifies! If you have any further questions, do not hesitate to ask</p>",
      "votes": null,
      "replies": [
        {
          "id": 339969,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "06/08/2018 03:31:59",
          "content": "<p>Thank you for clarification and encourage our innovation.\nLet me confirm about test set.\nI'd like also perform automatic pseudo-labeling test set, and use it to fine tune existing model.\nIs this allowed?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 340031,
          "author_name": "hyunguilim",
          "author_url": "",
          "post_date": "06/08/2018 07:28:57",
          "content": "<p>I think test set should not be included in the model although the rule only prohibits hand validation of test set. This process breaks the general assumption that test set is completely invisible.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 340050,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "06/08/2018 08:32:19",
          "content": "<p>Thanks for your comment.\nActually recently I already have tried some, just because it's normal to do the pseudo labeling in kaggle competitions and not specifically restricted by this competition rules.\nBecause your question was a little similar to pseudo labeling, then I wanted to make sure that it is OK.\nSo can I wait for their clarification? :)</p>\n\n<p>NOTE: My attempts with pseudo labeling so far was ended with all degraded results... Current my best is NOT using it.</p>\n\n<p>BTW I was wondering if pseudo labeling is useful in real use-case or not. And I also have concluded that it is useful as same as what we can read in some pseudo labeling literatures like [1] [2].</p>\n\n<p>Considering real applications, many cases don't have much samples initially. After we get application up and running in (test) operation, we collect samples. We then need to label them to re-train model to match to the real sample distribution. We will repeat this cycle then. Pseudo-labeling will be useful for automatic cycle of this kind of system. So I'd like to learn how to make use of it, and this competition's datasets seem to be good learning materials - it seems not to be easy to use it for further improvements.</p>\n\n<p>[1] Semi-Supervised Deep Learning Using Pseudo Labels for Hyperspectral Image Classification, Hao Wu et al, 2017\n[2] Pseudo-Label : The Simple and Efficient Semi-Supervised Learning Method for Deep Neural Networks, Dong-Hyun Lee, 2013</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 340073,
          "author_name": "lccever",
          "author_url": "",
          "post_date": "06/08/2018 09:27:17",
          "content": "<p>Thanks for the clarification, Eduardo.\nI think it's kind of overfit to the test set ,some kind of treat.<a href=\"/daisukelab\">@daisukelab</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 340098,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "06/08/2018 10:26:22",
          "content": "<p>@LEI.Jinggui Thanks for your comment. If it's kind of overfit to test set, I guess it would lead to better score.\nAnyway as I was observing samples selected by my models, I found something, more fundamental issue.\nI will write about it when I could confirm...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 345980,
          "author_name": "agehsbarg",
          "author_url": "",
          "post_date": "06/20/2018 21:18:35",
          "content": "<p>Hi, <a href=\"/daisukelab\">@daisukelab</a>, thanks for sharing your experience. </p>\n\n<p>I also tried pseudo labeling and it also gave lower score on LB. To be precise, what I did was to train the models, predict labels on test set, take samples with predicted probability higher then, say, 0.95, add these samples to train, re-train, etc. Surprisingly no improvement, especially because probability is taken as geometric mean of 4 classifiers, two of which are ensembles of NN. So, given high score on LB I would assume that top-predictions should have correct labels. I am 100% sure that same is applicable to your work.</p>\n\n<p>I also had an idea of hand-labeling training samples with lowest predicted probability, but haven't got time to do it properly. However, it seems that it might NOT bring added value, because labels are assigned in the same way on train and test, so if they are partly assigned in the wrong way on train, also on test. So there should be integrity around labels which might not make it work. Some support to that assumption is given by the fact that adding train samples which are not verified increases LB score.</p>\n\n<p>I also checked some of these TEST samples with lowest predicted probability and most of them are so called padding sounds - music, voices, water, birds, animals, etc. So hand-labelling them also shouldn't increase score cause they are not part of the competition(some details here <a href=\"https://www.kaggle.com/c/freesound-audio-tagging/discussion/58862\">https://www.kaggle.com/c/freesound-audio-tagging/discussion/58862</a>) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 345987,
          "author_name": "agehsbarg",
          "author_url": "",
          "post_date": "06/20/2018 21:32:01",
          "content": "<p>To summarize, I would say that we should have been given more daily submissions. It would help to study problem much better, cause as it is now, we are very restricted to checking ideas and verifying assumptions. Especially because the aim of the competition is research.</p>\n\n<p>There definitely are wrongly labelled train samples and as pseudo-labelling does not work, we are stuck in the dead-end. That's why I was claiming in other comments that LB scores between 0.92 and 0.96 are top-results and probably produced by very similar models. Which means that we can solve that task by achieving mapk of around 0.95, but to achieve better results, we need to pseudo- or hand- label BOTH train and test. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 346004,
          "author_name": "eduardofonseca",
          "author_url": "",
          "post_date": "06/20/2018 22:03:54",
          "content": "<p>Hi, I just wanted to clarify the <strong>allowed usage of the test set</strong>. As already stated in this thread:</p>\n\n<ul>\n<li>Automatic (and of course manual) labeling of the test set is <strong>not</strong> allowed. In the context of this competition, we do <strong>not</strong> allow the test set as input into the model in any form.</li>\n</ul>\n\n<p><a href=\"https://www.kaggle.com/c/freesound-audio-tagging/discussion/58052#340611\">https://www.kaggle.com/c/freesound-audio-tagging/discussion/58052#340611</a></p>\n\n<p>In other words, the test set should ONLY be used for reporting model's performance, and nothing else. Therefore, the approaches you were describing are forbidden in this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 346044,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "06/21/2018 01:10:25",
          "content": "<p>Hi @Aleksandrs,</p>\n\n<p>'#metoo'</p>\n\n<p>Thank you for sharing the detail. I agree with you 100% regarding pseudo labeling attempts for both train and test samples, yes I did almost the same, and experienced the same things. So I’m glad to hear your detail, I really felt that I’m not alone on the planet. :)</p>\n\n<p>I have distilled a trustful set of training samples by 4-5 times cycle of picking up samples of higher probabilities like &gt;0.92, and yes this attempt failed in vain. It proved that some samples are correctly labeled though they don't sound as labeled nor have higher probabilities.</p>\n\n<p>Then I almost came to understand that final goal of this competition would be modeling the real distribution of what have been done to label the dataset.</p>\n\n<p>Hand checking for the detailed distribution could show something, though I'm almost losing motivation to do that, just because this action might not be helpful to gain knowledge...\nBut finding the real distribution would be a good training, trying myself to keep up. As long as other people have better score, there should be something else to do... :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 346524,
          "author_name": "agehsbarg",
          "author_url": "",
          "post_date": "06/21/2018 22:01:47",
          "content": "<p>@Eduardo Fonseca yes, it is 100% clear that it is not allowed. I used approach which I described more than a month ago before it was clear that it is not allowed. I am not using it anymore as it is restrited and I will not use these old submissions. Moreover, as I described above, that approach gave me lower score on LB, so I wouldn't use these submissions even if I wanted to cheat.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 346529,
          "author_name": "agehsbarg",
          "author_url": "",
          "post_date": "06/21/2018 22:14:20",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a> I agree - after achieving high scores aim of that competition becomes learning the way how datasets were labelled, not learning smt meaningful. However, organizers say that \"<strong>the main goal of this competition is to foster and advance research in sound recognition</strong>\".  I am just losing motivation to do smt extra, partly because I am not sure that improvement on LB will be a \"real\" improvement, partly because of only two submissions. And there is smt to try - LSTMs, RNNs and HMMs. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 346559,
          "author_name": "eduardofonseca",
          "author_url": "",
          "post_date": "06/22/2018 00:03:12",
          "content": "<p>Hi @Aleksandrs, </p>\n\n<p>great, we wanted to clarify the issue again in case another participant read your post (but not the other post where we clarify the allowed usage of the test set ). Since, as you pointed out, this is not specified in the competition Rules, we wanted to avoid potential confusions.</p>\n\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 340611,
      "author_name": "eduardofonseca",
      "author_url": "",
      "post_date": "06/09/2018 19:08:25",
      "content": "<p>Hi again,</p>\n\n<p>providing <strong>clarification about the test set usage</strong>:</p>\n\n<ul>\n<li>Automatic (and of course manual) labeling of the test set is <strong>not</strong> allowed. In the context of this competition, we do <strong>not</strong> allow the test set as input into the model in any form.</li>\n</ul>\n\n<p>As usual, please ask any questions that you may have. <a href=\"/daisukelab\">@daisukelab</a>, thanks for raising this issue, which was not properly specified in the competition rules.</p>",
      "votes": null,
      "replies": [
        {
          "id": 340665,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "06/09/2018 23:31:48",
          "content": "<p>Hi @Eduardo, thank you for taking time for clarification!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "336883": "Hi, I have a question about 'non-verified' data.\n\nI believe self-hand-verification of non-verified data is prohibited for fair competition even if it is training data. However, I couldn't find the rule about it.\n\nNon-verified data accounts for more than 60% of training data, and it seems to be more effective to use the entire data than to use only verified data. Since the non-verified data is still quite noisy as shown in data description, I think sampling a clean subset of non-verified data can make further improvement. In this situation, I think the sampling process should be purely automatic without any person's decision. But I have no idea whether this process is available or not.\n\nSo, I hope organizers clarify the rule about 'non-verified' data ASAP.\nThanks.",
    "337983": "Hi, I also have related question about pseudo labeling.\nI (maybe we) 'd appreciate if organizers could clarify use of pseudo labeling.\nIt would not be manually hand labeled, but automatically labeled by first model.\nThen final purpose is to train better model by using pseudo-labeled test samples or non-verified training samples.\nThanks in advance.",
    "338428": "I also get the same question. And I think if there's no rules, maybe someone may check and label the non-verified samples hand-by-hand.",
    "338560": "Hi, we are discussing among the organizers which is the best way to address this issue, and we'll provide clarification ASAP. Thanks!",
    "339916": "Hi all, \n\nthanks for the patience. After discussing among the organizers, we've decided that:\n\n- automatic re-labeling of the train set (e.g., of the non-verified portion) in order to refine the training data is allowed.\n- manual re-labeling of the train set (e.g., of the non-verified portion) is also allowed, on the condition that it is clearly specified somewhere, e.g., in the submission.\n\nThat being said, we have two additional comments:\n\n- we encourage automatic (rather than manual) re-labeling, as it can imply certain methodological/algorithmic innovation\n- if any participant decides to carry out manual re-labeling in the train set, and **only after the challenge is concluded**, we encourage her/him to release the new labels (or at least share them with the organizers). In this way, the labels could be considered for future dataset versions and thus benefit the community.\n\nHope this clarifies! If you have any further questions, do not hesitate to ask",
    "339969": "Thank you for clarification and encourage our innovation.\nLet me confirm about test set.\nI'd like also perform automatic pseudo-labeling test set, and use it to fine tune existing model.\nIs this allowed?",
    "340031": "I think test set should not be included in the model although the rule only prohibits hand validation of test set. This process breaks the general assumption that test set is completely invisible.",
    "340050": "Thanks for your comment.\nActually recently I already have tried some, just because it's normal to do the pseudo labeling in kaggle competitions and not specifically restricted by this competition rules.\nBecause your question was a little similar to pseudo labeling, then I wanted to make sure that it is OK.\nSo can I wait for their clarification? :)\n\nNOTE: My attempts with pseudo labeling so far was ended with all degraded results... Current my best is NOT using it.\n\nBTW I was wondering if pseudo labeling is useful in real use-case or not. And I also have concluded that it is useful as same as what we can read in some pseudo labeling literatures like [1] [2].\n\nConsidering real applications, many cases don't have much samples initially. After we get application up and running in (test) operation, we collect samples. We then need to label them to re-train model to match to the real sample distribution. We will repeat this cycle then. Pseudo-labeling will be useful for automatic cycle of this kind of system. So I'd like to learn how to make use of it, and this competition's datasets seem to be good learning materials - it seems not to be easy to use it for further improvements.\n\n[1] Semi-Supervised Deep Learning Using Pseudo Labels for Hyperspectral Image Classification, Hao Wu et al, 2017\n[2] Pseudo-Label : The Simple and Efficient Semi-Supervised Learning Method for Deep Neural Networks, Dong-Hyun Lee, 2013",
    "340073": "Thanks for the clarification, Eduardo.\nI think it's kind of overfit to the test set ,some kind of treat.@daisukelab",
    "340098": "LEI.Jinggui Thanks for your comment. If it's kind of overfit to test set, I guess it would lead to better score.\nAnyway as I was observing samples selected by my models, I found something, more fundamental issue.\nI will write about it when I could confirm...",
    "340611": "Hi again,\n\nproviding **clarification about the test set usage**:\n\n - Automatic (and of course manual) labeling of the test set is **not** allowed. In the context of this competition, we do **not** allow the test set as input into the model in any form.\n\nAs usual, please ask any questions that you may have. @daisukelab, thanks for raising this issue, which was not properly specified in the competition rules.",
    "340665": "Hi @Eduardo, thank you for taking time for clarification!",
    "345980": "Hi, @daisukelab, thanks for sharing your experience. \n\nI also tried pseudo labeling and it also gave lower score on LB. To be precise, what I did was to train the models, predict labels on test set, take samples with predicted probability higher then, say, 0.95, add these samples to train, re-train, etc. Surprisingly no improvement, especially because probability is taken as geometric mean of 4 classifiers, two of which are ensembles of NN. So, given high score on LB I would assume that top-predictions should have correct labels. I am 100% sure that same is applicable to your work.\n\nI also had an idea of hand-labeling training samples with lowest predicted probability, but haven't got time to do it properly. However, it seems that it might NOT bring added value, because labels are assigned in the same way on train and test, so if they are partly assigned in the wrong way on train, also on test. So there should be integrity around labels which might not make it work. Some support to that assumption is given by the fact that adding train samples which are not verified increases LB score.\n\nI also checked some of these TEST samples with lowest predicted probability and most of them are so called padding sounds - music, voices, water, birds, animals, etc. So hand-labelling them also shouldn't increase score cause they are not part of the competition(some details here https://www.kaggle.com/c/freesound-audio-tagging/discussion/58862)",
    "345987": "To summarize, I would say that we should have been given more daily submissions. It would help to study problem much better, cause as it is now, we are very restricted to checking ideas and verifying assumptions. Especially because the aim of the competition is research.\n\nThere definitely are wrongly labelled train samples and as pseudo-labelling does not work, we are stuck in the dead-end. That's why I was claiming in other comments that LB scores between 0.92 and 0.96 are top-results and probably produced by very similar models. Which means that we can solve that task by achieving mapk of around 0.95, but to achieve better results, we need to pseudo- or hand- label BOTH train and test.",
    "346004": "Hi, I just wanted to clarify the **allowed usage of the test set**. As already stated in this thread:\n\n - Automatic (and of course manual) labeling of the test set is **not** allowed. In the context of this competition, we do **not** allow the test set as input into the model in any form.\n\nhttps://www.kaggle.com/c/freesound-audio-tagging/discussion/58052#340611\n\nIn other words, the test set should ONLY be used for reporting model's performance, and nothing else. Therefore, the approaches you were describing are forbidden in this competition.",
    "346044": "Hi @Aleksandrs,\n\n'#metoo'\n\nThank you for sharing the detail. I agree with you 100% regarding pseudo labeling attempts for both train and test samples, yes I did almost the same, and experienced the same things. So I’m glad to hear your detail, I really felt that I’m not alone on the planet. :)\n\nI have distilled a trustful set of training samples by 4-5 times cycle of picking up samples of higher probabilities like &gt;0.92, and yes this attempt failed in vain. It proved that some samples are correctly labeled though they don't sound as labeled nor have higher probabilities.\n\nThen I almost came to understand that final goal of this competition would be modeling the real distribution of what have been done to label the dataset.\n\nHand checking for the detailed distribution could show something, though I'm almost losing motivation to do that, just because this action might not be helpful to gain knowledge...\nBut finding the real distribution would be a good training, trying myself to keep up. As long as other people have better score, there should be something else to do... :)",
    "346524": "Eduardo Fonseca yes, it is 100% clear that it is not allowed. I used approach which I described more than a month ago before it was clear that it is not allowed. I am not using it anymore as it is restrited and I will not use these old submissions. Moreover, as I described above, that approach gave me lower score on LB, so I wouldn't use these submissions even if I wanted to cheat.",
    "346529": "daisukelab I agree - after achieving high scores aim of that competition becomes learning the way how datasets were labelled, not learning smt meaningful. However, organizers say that \"**the main goal of this competition is to foster and advance research in sound recognition**\".  I am just losing motivation to do smt extra, partly because I am not sure that improvement on LB will be a \"real\" improvement, partly because of only two submissions. And there is smt to try - LSTMs, RNNs and HMMs.",
    "346559": "Hi @Aleksandrs, \n\ngreat, we wanted to clarify the issue again in case another participant read your post (but not the other post where we clarify the allowed usage of the test set ). Since, as you pointed out, this is not specified in the competition Rules, we wanted to avoid potential confusions.\n\nThanks!"
  },
  "source": "meta"
}