{
  "id": 91827,
  "title": "Paper: MixMatch: A Holistic Approach to Semi-Supervised Learning",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/91827",
  "author_name": "",
  "post_date": "2019-05-09T13:19:15.693485100Z",
  "votes": 17,
  "comment_count": 30,
  "views": 0,
  "content": "<p><a href=\"https://twitter.com/D_Berthelot_ML/status/1125996671664451584\">https://twitter.com/D_Berthelot_ML/status/1125996671664451584</a>\n<a href=\"http://arxiv.org/abs/1905.02249\">http://arxiv.org/abs/1905.02249</a></p>\n\n<p>Here's another new paper that perfect matches to this competition needs!</p>\n\n<ul>\n<li>'... MixMatch, that works by guessing low-entropy labels for data-augmented unlabeled examples and mixing labeled and unlabeled data using MixUp.' - Noisy set as unlabeled?</li>\n<li>'... For example, on CIFAR-10 with 250 labels, we reduce error rate by a factor of 4 (from 38% to 11%) and by a factor of 2 on STL-10.' - Big improvement.</li>\n<li>See algorithm 1, it's very interesting to create new batch from labeled and unlabeled samples with augmentation, pseudo labeling and mixup (<a href=\"https://arxiv.org/pdf/1710.09412.pdf\">https://arxiv.org/pdf/1710.09412.pdf</a>).</li>\n</ul>\n\n<p>Anyone is going to try? :) There seems no implementation found on github yet. The author said on twitter that they are going to opensource it in few days (<a href=\"https://twitter.com/D_Berthelot_ML/status/1126393903022727168\">https://twitter.com/D_Berthelot_ML/status/1126393903022727168</a>).</p>\n\n<h2>CAUTION: Don't use test samples in training</h2>\n\n<p>Be sure NOT to use test set as unlabeled input, it's banned as clarified here: <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#524949\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#524949</a></p>",
  "messages": [
    {
      "id": "529221",
      "postDate": "05/09/2019 13:19:15",
      "content": "<p><a href=\"https://twitter.com/D_Berthelot_ML/status/1125996671664451584\">https://twitter.com/D_Berthelot_ML/status/1125996671664451584</a>\n<a href=\"http://arxiv.org/abs/1905.02249\">http://arxiv.org/abs/1905.02249</a></p>\n\n<p>Here's another new paper that perfect matches to this competition needs!</p>\n\n<ul>\n<li>'... MixMatch, that works by guessing low-entropy labels for data-augmented unlabeled examples and mixing labeled and unlabeled data using MixUp.' - Noisy set as unlabeled?</li>\n<li>'... For example, on CIFAR-10 with 250 labels, we reduce error rate by a factor of 4 (from 38% to 11%) and by a factor of 2 on STL-10.' - Big improvement.</li>\n<li>See algorithm 1, it's very interesting to create new batch from labeled and unlabeled samples with augmentation, pseudo labeling and mixup (<a href=\"https://arxiv.org/pdf/1710.09412.pdf\">https://arxiv.org/pdf/1710.09412.pdf</a>).</li>\n</ul>\n\n<p>Anyone is going to try? :) There seems no implementation found on github yet. The author said on twitter that they are going to opensource it in few days (<a href=\"https://twitter.com/D_Berthelot_ML/status/1126393903022727168\">https://twitter.com/D_Berthelot_ML/status/1126393903022727168</a>).</p>\n\n<h2>CAUTION: Don't use test samples in training</h2>\n\n<p>Be sure NOT to use test set as unlabeled input, it's banned as clarified here: <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#524949\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#524949</a></p>",
      "rawMarkdown": "https://twitter.com/D_Berthelot_ML/status/1125996671664451584\nhttp://arxiv.org/abs/1905.02249\n\nHere's another new paper that perfect matches to this competition needs!\n\n- '... MixMatch, that works by guessing low-entropy labels for data-augmented unlabeled examples and mixing labeled and unlabeled data using MixUp.' - Noisy set as unlabeled?\n- '... For example, on CIFAR-10 with 250 labels, we reduce error rate by a factor of 4 (from 38% to 11%) and by a factor of 2 on STL-10.' - Big improvement.\n- See algorithm 1, it's very interesting to create new batch from labeled and unlabeled samples with augmentation, pseudo labeling and mixup (https://arxiv.org/pdf/1710.09412.pdf).\n\nAnyone is going to try? :) There seems no implementation found on github yet. The author said on twitter that they are going to opensource it in few days (https://twitter.com/D_Berthelot_ML/status/1126393903022727168).\n\n## CAUTION: Don't use test samples in training\nBe sure NOT to use test set as unlabeled input, it's banned as clarified here: https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#524949",
      "votes": null
    },
    {
      "id": "529360",
      "postDate": "05/09/2019 17:52:07",
      "content": "<p>I'm trying this :) The results described in the paper are really impressive!</p>\n\n<p>So far it's looking promising, better than using the given labels of the noisy set.  </p>",
      "rawMarkdown": "I'm trying this :) The results described in the paper are really impressive!\n\nSo far it's looking promising, better than using the given labels of the noisy set.",
      "votes": null
    },
    {
      "id": "529389",
      "postDate": "05/09/2019 19:18:53",
      "content": "<p>It's great, thanks for sharing! 👍 </p>",
      "rawMarkdown": "It's great, thanks for sharing! 👍",
      "votes": null
    },
    {
      "id": "529405",
      "postDate": "05/09/2019 20:04:46",
      "content": "<p>Got LB 0.543 on first try, and it is still underfitting. With a similar setup I got LB 0.698 only with curated set. But again, for a first try it seems promising as I still need to debug more carefully my quick and dirty implementation of MixMatch, tune the hyperparameters and train for longer to see how far it goes :)</p>",
      "rawMarkdown": "Got LB 0.543 on first try, and it is still underfitting. With a similar setup I got LB 0.698 only with curated set. But again, for a first try it seems promising as I still need to debug more carefully my quick and dirty implementation of MixMatch, tune the hyperparameters and train for longer to see how far it goes :)",
      "votes": null
    },
    {
      "id": "529419",
      "postDate": "05/09/2019 20:57:34",
      "content": "<p>Amazing it's so quick to implement! I'm happy to know that it would be working fine, new way to go.\nI'll keep checking how far you go on the LB! ;)</p>",
      "rawMarkdown": "Amazing it's so quick to implement! I'm happy to know that it would be working fine, new way to go.\nI'll keep checking how far you go on the LB! ;)",
      "votes": null
    },
    {
      "id": "531949",
      "postDate": "05/15/2019 22:37:42",
      "content": "<p>FYI: <a href=\"https://github.com/google-research/mixmatch\">https://github.com/google-research/mixmatch</a></p>",
      "rawMarkdown": "FYI: https://github.com/google-research/mixmatch",
      "votes": null
    },
    {
      "id": "531980",
      "postDate": "05/16/2019 00:39:45",
      "content": "<p>Thanks, you're always quick ;)</p>",
      "rawMarkdown": "Thanks, you're always quick ;)",
      "votes": null
    },
    {
      "id": "532155",
      "postDate": "05/16/2019 10:10:16",
      "content": "<p>And from the authors:\n<a href=\"https://github.com/google-research/mixmatch\">https://github.com/google-research/mixmatch</a> </p>",
      "rawMarkdown": "And from the authors:\nhttps://github.com/google-research/mixmatch",
      "votes": null
    },
    {
      "id": "534107",
      "postDate": "05/20/2019 17:37:34",
      "content": "<p>Hi <a href=\"/mnpinto\">@mnpinto</a> , may I ask you how many epochs it takes for MixMatch to converge? I am still training it but it seems like taking forever.</p>",
      "rawMarkdown": "Hi @mnpinto , may I ask you how many epochs it takes for MixMatch to converge? I am still training it but it seems like taking forever.",
      "votes": null
    },
    {
      "id": "534142",
      "postDate": "05/20/2019 18:46:16",
      "content": "<p>I'm not sure that I've got it working correctly. It's giving me a result very close but not better than with curated data only (even after training for very long), I may be missing something. I'm not using the authors code since I'm working in pytorch. </p>\n\n<p>And be aware that the sharpening function is not designed for multi-class classifications, that may also impact the results as it seems an important step according to the ablation study in the paper.</p>",
      "rawMarkdown": "I'm not sure that I've got it working correctly. It's giving me a result very close but not better than with curated data only (even after training for very long), I may be missing something. I'm not using the authors code since I'm working in pytorch. \n\nAnd be aware that the sharpening function is not designed for multi-class classifications, that may also impact the results as it seems an important step according to the ablation study in the paper.",
      "votes": null
    },
    {
      "id": "534153",
      "postDate": "05/20/2019 19:44:43",
      "content": "<p>I am using PyTorch as well. I think the sharpening function works for multi-label too since it just amplifies the magnitude of output in order to make it more \"binary\" (closer to either 0 or 1). </p>\n\n<p>While I am typing, I just realize a bug in my code.... I didn't take the sigmoid of my output before feeding it into the sharpening function.</p>",
      "rawMarkdown": "I am using PyTorch as well. I think the sharpening function works for multi-label too since it just amplifies the magnitude of output in order to make it more \"binary\" (closer to either 0 or 1). \n\nWhile I am typing, I just realize a bug in my code.... I didn't take the sigmoid of my output before feeding it into the sharpening function.",
      "votes": null
    },
    {
      "id": "534158",
      "postDate": "05/20/2019 19:52:23",
      "content": "<p>An easy reading blog of what mixmatch is <a href=\"https://medium.com/&lt;a href=\">@sanjeev</a>.vadiraj/eureka-mixmatch-a-holistic-approach-to-semi-supervised-learning-125b14e82d2f\"&gt;https://medium.com/<a href=\"/sanjeev\">@sanjeev</a>.vadiraj/eureka-mixmatch-a-holistic-approach-to-semi-supervised-learning-125b14e82d2f</p>",
      "rawMarkdown": "An easy reading blog of what mixmatch is https://medium.com/@sanjeev.vadiraj/eureka-mixmatch-a-holistic-approach-to-semi-supervised-learning-125b14e82d2f",
      "votes": null
    },
    {
      "id": "534166",
      "postDate": "05/20/2019 20:24:23",
      "content": "<p>But it makes the result sum to 1 over all classes, more like the output of a softmax. It should work but probably it's not optimal. Did you try to reproduce the CIFAR-10 results of the paper? I guess it's the only way to make sure the code is really working without any bugs.</p>",
      "rawMarkdown": "But it makes the result sum to 1 over all classes, more like the output of a softmax. It should work but probably it's not optimal. Did you try to reproduce the CIFAR-10 results of the paper? I guess it's the only way to make sure the code is really working without any bugs.",
      "votes": null
    },
    {
      "id": "534194",
      "postDate": "05/20/2019 22:25:19",
      "content": "<p>You are right. I didn't try to reproduce the results of the paper. I will try to write a multi-label version of sharpening function. </p>",
      "rawMarkdown": "You are right. I didn't try to reproduce the results of the paper. I will try to write a multi-label version of sharpening function.",
      "votes": null
    },
    {
      "id": "534214",
      "postDate": "05/20/2019 23:46:26",
      "content": "<p>I've tried to reproduce the results last week but there are many things to consider like the same model, and data processing. I had no time to look into all that yet but if you are trying this other paper may be useful <a href=\"https://arxiv.org/abs/1903.03825\">https://arxiv.org/abs/1903.03825</a>, it presents a similar method (mentioned in MixMatch paper), but the good news is that the code is in PyTorch so it can be a good starting point: <a href=\"https://github.com/vikasverma1077/ICT\">https://github.com/vikasverma1077/ICT</a> </p>",
      "rawMarkdown": "I've tried to reproduce the results last week but there are many things to consider like the same model, and data processing. I had no time to look into all that yet but if you are trying this other paper may be useful https://arxiv.org/abs/1903.03825, it presents a similar method (mentioned in MixMatch paper), but the good news is that the code is in PyTorch so it can be a good starting point: https://github.com/vikasverma1077/ICT",
      "votes": null
    },
    {
      "id": "534252",
      "postDate": "05/21/2019 01:40:36",
      "content": "<p><a href=\"/mnpinto\">@mnpinto</a> <a href=\"/jihangz\">@jihangz</a> Question for you guys tried mixmatch with pytorch, are you guys training with ema as the paper stated? </p>",
      "rawMarkdown": "mnpinto @jihangz Question for you guys tried mixmatch with pytorch, are you guys training with ema as the paper stated?",
      "votes": null
    },
    {
      "id": "534418",
      "postDate": "05/21/2019 08:04:17",
      "content": "<p>I haven't tried EMA, they say \"EMA of parameter values hurt MixMatch’s performance slightly\" on the Ablation study. </p>",
      "rawMarkdown": "I haven't tried EMA, they say \"EMA of parameter values hurt MixMatch’s performance slightly\" on the Ablation study.",
      "votes": null
    },
    {
      "id": "534490",
      "postDate": "05/21/2019 10:25:19",
      "content": "<p>I also tried with pytorch, it did not converge to a high lwlrap.\nIt seems we do not need Augment(x) in this audio tagging task, because crop or resize may affect the logmel spectrogram.</p>",
      "rawMarkdown": "I also tried with pytorch, it did not converge to a high lwlrap.\nIt seems we do not need Augment(x) in this audio tagging task, because crop or resize may affect the logmel spectrogram.",
      "votes": null
    },
    {
      "id": "534520",
      "postDate": "05/21/2019 12:03:35",
      "content": "<p>Oh cool, I have not got to that part yet. It was mentioned in the implementation detail section before the ablation.</p>",
      "rawMarkdown": "Oh cool, I have not got to that part yet. It was mentioned in the implementation detail section before the ablation.",
      "votes": null
    },
    {
      "id": "534526",
      "postDate": "05/21/2019 12:08:07",
      "content": "<p>I haven't tried common image augmentation methods yet.  For longer clips, we would need to randomly truncate along the time axis to get them into required shapes. I think it could consider to be an augment function. </p>\n\n<p>I have just kicked off my mixmatch training with just that as aug, will see how it goes after work.</p>",
      "rawMarkdown": "I haven't tried common image augmentation methods yet.  For longer clips, we would need to randomly truncate along the time axis to get them into required shapes. I think it could consider to be an augment function. \n\nI have just kicked off my mixmatch training with just that as aug, will see how it goes after work.",
      "votes": null
    },
    {
      "id": "534673",
      "postDate": "05/21/2019 16:50:21",
      "content": "<p>500 epochs of training, local lwlrap is 0.736 now, and it is still increasing. I am going to train another 500 epochs to see how it goes.</p>",
      "rawMarkdown": "500 epochs of training, local lwlrap is 0.736 now, and it is still increasing. I am going to train another 500 epochs to see how it goes.",
      "votes": null
    },
    {
      "id": "534676",
      "postDate": "05/21/2019 17:03:22",
      "content": "<p>Here's another unsupervised learning paper called \"Unsupervised Data Augmentation\" (UDA). It reaches lower error rate than MixMatch on CIFAR-10 given 4000 labels. <a href=\"https://arxiv.org/pdf/1904.12848.pdf\">https://arxiv.org/pdf/1904.12848.pdf</a></p>\n\n<p>Two papers use very similar loss functions (CE+L2 vs CE+KL). UDA is also trained on ImageNet and multiple text classification datasets.</p>",
      "rawMarkdown": "Here's another unsupervised learning paper called \"Unsupervised Data Augmentation\" (UDA). It reaches lower error rate than MixMatch on CIFAR-10 given 4000 labels. https://arxiv.org/pdf/1904.12848.pdf\n\nTwo papers use very similar loss functions (CE+L2 vs CE+KL). UDA is also trained on ImageNet and multiple text classification datasets.",
      "votes": null
    },
    {
      "id": "534682",
      "postDate": "05/21/2019 17:15:28",
      "content": "<p><a href=\"/sailorwei\">@sailorwei</a> I am using SpecAugment as Augment(x)</p>",
      "rawMarkdown": "sailorwei I am using SpecAugment as Augment(x)",
      "votes": null
    },
    {
      "id": "534868",
      "postDate": "05/22/2019 00:18:48",
      "content": "<p>I implemented MixMatch in TF eager. I used the same model for my best LB and was able to reach max 0.83 on one fold lwlrap (avg around 0.8 per fold for 5-fold). I used SpecAugment for augmentation component + random crops along the time_axis. However this only resulted in 0.65 LB so still not better than curated only.</p>",
      "rawMarkdown": "I implemented MixMatch in TF eager. I used the same model for my best LB and was able to reach max 0.83 on one fold lwlrap (avg around 0.8 per fold for 5-fold). I used SpecAugment for augmentation component + random crops along the time_axis. However this only resulted in 0.65 LB so still not better than curated only.",
      "votes": null
    },
    {
      "id": "538627",
      "postDate": "05/28/2019 22:03:11",
      "content": "<p><a href=\"/ryanzhang\">@ryanzhang</a> I almost fell off my chair! Is MixMatch the key of your amazing LB score?</p>",
      "rawMarkdown": "ryanzhang I almost fell off my chair! Is MixMatch the key of your amazing LB score?",
      "votes": null
    },
    {
      "id": "539221",
      "postDate": "05/29/2019 18:04:59",
      "content": "<p>No, mixmatch is not working for me.</p>",
      "rawMarkdown": "No, mixmatch is not working for me.",
      "votes": null
    },
    {
      "id": "539343",
      "postDate": "05/30/2019 00:04:22",
      "content": "<p>Got it...  Nice work, anyway!</p>",
      "rawMarkdown": "Got it...  Nice work, anyway!",
      "votes": null
    },
    {
      "id": "550292",
      "postDate": "06/11/2019 13:22:05",
      "content": "<p><a href=\"https://github.com/daisukelab/freesound-audio-tagging-2019\">https://github.com/daisukelab/freesound-audio-tagging-2019</a></p>\n\n<p>This repo has MixMatch implementation, though it would be incomplete. I was using this as super version of mixup...</p>\n\n<p>I'm hoping to have some more time to complete, to reproduce the original paper... fingers crossed..</p>",
      "rawMarkdown": "https://github.com/daisukelab/freesound-audio-tagging-2019\n\nThis repo has MixMatch implementation, though it would be incomplete. I was using this as super version of mixup...\n\nI'm hoping to have some more time to complete, to reproduce the original paper... fingers crossed..",
      "votes": null
    },
    {
      "id": "550316",
      "postDate": "06/11/2019 13:49:51",
      "content": "<p>Thanks for sharing! I haven't been able to improve my results using MixMatch but comparing with my experiments on Cifar-10 I think it's probably a matter of getting all the hyper-parameters and model choice in the right spot. </p>",
      "rawMarkdown": "Thanks for sharing! I haven't been able to improve my results using MixMatch but comparing with my experiments on Cifar-10 I think it's probably a matter of getting all the hyper-parameters and model choice in the right spot.",
      "votes": null
    },
    {
      "id": "550331",
      "postDate": "06/11/2019 14:04:58",
      "content": "<p>Thanks for sharing your experience, ummm I see. What I suffered was prediction results of label guessing gets closer and closer like this:</p>\n\n<p>I was turning MixMatch on after once model is trained basically well. Then in earlier epochs, model predicts (guesses) like this:</p>\n\n<pre><code>max of predicted tensor([0.7737, 0.4485, 0.3357,  ..., 0.5738, 0.4036, 0.3020])\nmin of tensor([1.4771e-09, 9.2836e-08, 1.6407e-07,  ..., 1.1334e-07, 1.1591e-08 ...\n</code></pre>\n\n<p>Then after some more epochs, it predicts like:</p>\n\n<pre><code>max of predicted tensor([0.0455, 0.0465, 0.0464,  ..., 0.0581, 0.0484, 0.0472])\nmin of tensor([0.0115, 0.0112, 0.0115,  ..., 0.0115, 0.0113, 0.0115])\n</code></pre>\n\n<p>It gets closer until all <code>sigmoid(logits)</code> converges to single value.\nMy understanding so far is, model could be motivated to do so by MSE loss which encourages minimizing difference between guessed predictions with non-augmented image and predictions with augmented one. It'd be quite easy to achieve if model predicts a single value for all predictions...</p>\n\n<p>I'll have some time to debug more....</p>",
      "rawMarkdown": "Thanks for sharing your experience, ummm I see. What I suffered was prediction results of label guessing gets closer and closer like this:\n\nI was turning MixMatch on after once model is trained basically well. Then in earlier epochs, model predicts (guesses) like this:\n\n    max of predicted tensor([0.7737, 0.4485, 0.3357,  ..., 0.5738, 0.4036, 0.3020])\n    min of tensor([1.4771e-09, 9.2836e-08, 1.6407e-07,  ..., 1.1334e-07, 1.1591e-08 ...\n\nThen after some more epochs, it predicts like:\n\n    max of predicted tensor([0.0455, 0.0465, 0.0464,  ..., 0.0581, 0.0484, 0.0472])\n    min of tensor([0.0115, 0.0112, 0.0115,  ..., 0.0115, 0.0113, 0.0115])\n\nIt gets closer until all `sigmoid(logits)` converges to single value.\nMy understanding so far is, model could be motivated to do so by MSE loss which encourages minimizing difference between guessed predictions with non-augmented image and predictions with augmented one. It'd be quite easy to achieve if model predicts a single value for all predictions...\n\nI'll have some time to debug more....",
      "votes": null
    },
    {
      "id": "550372",
      "postDate": "06/11/2019 14:56:21",
      "content": "<p>I was training with MixMatch from random initialization but the problem was similar. I noticed that lwlrap would increase to levels not far from mixup with curated data only, but Fbeta score does not increase unless I remove the MSE loss term as the predicted probabilities remained very small. This happened in my Cifar-10 experiments when the weight of MSE loss was increased to fast and also when using a larger model like resnet18 instead of the wide resnet used in the paper.  </p>\n\n<p>Perhaps for debugging a good idea is to start with just the curated data and only the single label  samples, using a few labeled samples and the remaining as unlabeled. If it works that way than it should be easier to move forward to add the multi-label samples and finally the noisy data.</p>\n\n<p>I will also try to find some time to organize and improve my implementation on fastai and to try on other datasets.</p>",
      "rawMarkdown": "I was training with MixMatch from random initialization but the problem was similar. I noticed that lwlrap would increase to levels not far from mixup with curated data only, but Fbeta score does not increase unless I remove the MSE loss term as the predicted probabilities remained very small. This happened in my Cifar-10 experiments when the weight of MSE loss was increased to fast and also when using a larger model like resnet18 instead of the wide resnet used in the paper.  \n\nPerhaps for debugging a good idea is to start with just the curated data and only the single label  samples, using a few labeled samples and the remaining as unlabeled. If it works that way than it should be easier to move forward to add the multi-label samples and finally the noisy data.\n\nI will also try to find some time to organize and improve my implementation on fastai and to try on other datasets.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 529360,
      "author_name": "mnpinto",
      "author_url": "",
      "post_date": "05/09/2019 17:52:07",
      "content": "<p>I'm trying this :) The results described in the paper are really impressive!</p>\n\n<p>So far it's looking promising, better than using the given labels of the noisy set.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 529389,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "05/09/2019 19:18:53",
          "content": "<p>It's great, thanks for sharing! 👍 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 529405,
          "author_name": "mnpinto",
          "author_url": "",
          "post_date": "05/09/2019 20:04:46",
          "content": "<p>Got LB 0.543 on first try, and it is still underfitting. With a similar setup I got LB 0.698 only with curated set. But again, for a first try it seems promising as I still need to debug more carefully my quick and dirty implementation of MixMatch, tune the hyperparameters and train for longer to see how far it goes :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 529419,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "05/09/2019 20:57:34",
          "content": "<p>Amazing it's so quick to implement! I'm happy to know that it would be working fine, new way to go.\nI'll keep checking how far you go on the LB! ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534107,
          "author_name": "jihangz",
          "author_url": "",
          "post_date": "05/20/2019 17:37:34",
          "content": "<p>Hi <a href=\"/mnpinto\">@mnpinto</a> , may I ask you how many epochs it takes for MixMatch to converge? I am still training it but it seems like taking forever.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534142,
          "author_name": "mnpinto",
          "author_url": "",
          "post_date": "05/20/2019 18:46:16",
          "content": "<p>I'm not sure that I've got it working correctly. It's giving me a result very close but not better than with curated data only (even after training for very long), I may be missing something. I'm not using the authors code since I'm working in pytorch. </p>\n\n<p>And be aware that the sharpening function is not designed for multi-class classifications, that may also impact the results as it seems an important step according to the ablation study in the paper.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534153,
          "author_name": "jihangz",
          "author_url": "",
          "post_date": "05/20/2019 19:44:43",
          "content": "<p>I am using PyTorch as well. I think the sharpening function works for multi-label too since it just amplifies the magnitude of output in order to make it more \"binary\" (closer to either 0 or 1). </p>\n\n<p>While I am typing, I just realize a bug in my code.... I didn't take the sigmoid of my output before feeding it into the sharpening function.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534166,
          "author_name": "mnpinto",
          "author_url": "",
          "post_date": "05/20/2019 20:24:23",
          "content": "<p>But it makes the result sum to 1 over all classes, more like the output of a softmax. It should work but probably it's not optimal. Did you try to reproduce the CIFAR-10 results of the paper? I guess it's the only way to make sure the code is really working without any bugs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534194,
          "author_name": "jihangz",
          "author_url": "",
          "post_date": "05/20/2019 22:25:19",
          "content": "<p>You are right. I didn't try to reproduce the results of the paper. I will try to write a multi-label version of sharpening function. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534214,
          "author_name": "mnpinto",
          "author_url": "",
          "post_date": "05/20/2019 23:46:26",
          "content": "<p>I've tried to reproduce the results last week but there are many things to consider like the same model, and data processing. I had no time to look into all that yet but if you are trying this other paper may be useful <a href=\"https://arxiv.org/abs/1903.03825\">https://arxiv.org/abs/1903.03825</a>, it presents a similar method (mentioned in MixMatch paper), but the good news is that the code is in PyTorch so it can be a good starting point: <a href=\"https://github.com/vikasverma1077/ICT\">https://github.com/vikasverma1077/ICT</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534252,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "05/21/2019 01:40:36",
          "content": "<p><a href=\"/mnpinto\">@mnpinto</a> <a href=\"/jihangz\">@jihangz</a> Question for you guys tried mixmatch with pytorch, are you guys training with ema as the paper stated? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534418,
          "author_name": "mnpinto",
          "author_url": "",
          "post_date": "05/21/2019 08:04:17",
          "content": "<p>I haven't tried EMA, they say \"EMA of parameter values hurt MixMatch’s performance slightly\" on the Ablation study. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534490,
          "author_name": "sailorwei",
          "author_url": "",
          "post_date": "05/21/2019 10:25:19",
          "content": "<p>I also tried with pytorch, it did not converge to a high lwlrap.\nIt seems we do not need Augment(x) in this audio tagging task, because crop or resize may affect the logmel spectrogram.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534520,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "05/21/2019 12:03:35",
          "content": "<p>Oh cool, I have not got to that part yet. It was mentioned in the implementation detail section before the ablation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534526,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "05/21/2019 12:08:07",
          "content": "<p>I haven't tried common image augmentation methods yet.  For longer clips, we would need to randomly truncate along the time axis to get them into required shapes. I think it could consider to be an augment function. </p>\n\n<p>I have just kicked off my mixmatch training with just that as aug, will see how it goes after work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534673,
          "author_name": "jihangz",
          "author_url": "",
          "post_date": "05/21/2019 16:50:21",
          "content": "<p>500 epochs of training, local lwlrap is 0.736 now, and it is still increasing. I am going to train another 500 epochs to see how it goes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534682,
          "author_name": "jihangz",
          "author_url": "",
          "post_date": "05/21/2019 17:15:28",
          "content": "<p><a href=\"/sailorwei\">@sailorwei</a> I am using SpecAugment as Augment(x)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534868,
          "author_name": "jamesrequa",
          "author_url": "",
          "post_date": "05/22/2019 00:18:48",
          "content": "<p>I implemented MixMatch in TF eager. I used the same model for my best LB and was able to reach max 0.83 on one fold lwlrap (avg around 0.8 per fold for 5-fold). I used SpecAugment for augmentation component + random crops along the time_axis. However this only resulted in 0.65 LB so still not better than curated only.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 538627,
          "author_name": "osciiart",
          "author_url": "",
          "post_date": "05/28/2019 22:03:11",
          "content": "<p><a href=\"/ryanzhang\">@ryanzhang</a> I almost fell off my chair! Is MixMatch the key of your amazing LB score?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 539221,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "05/29/2019 18:04:59",
          "content": "<p>No, mixmatch is not working for me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 539343,
          "author_name": "osciiart",
          "author_url": "",
          "post_date": "05/30/2019 00:04:22",
          "content": "<p>Got it...  Nice work, anyway!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 531949,
      "author_name": "mnpinto",
      "author_url": "",
      "post_date": "05/15/2019 22:37:42",
      "content": "<p>FYI: <a href=\"https://github.com/google-research/mixmatch\">https://github.com/google-research/mixmatch</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 531980,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "05/16/2019 00:39:45",
          "content": "<p>Thanks, you're always quick ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 532155,
      "author_name": "sailorwei",
      "author_url": "",
      "post_date": "05/16/2019 10:10:16",
      "content": "<p>And from the authors:\n<a href=\"https://github.com/google-research/mixmatch\">https://github.com/google-research/mixmatch</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 534158,
      "author_name": "robga",
      "author_url": "",
      "post_date": "05/20/2019 19:52:23",
      "content": "<p>An easy reading blog of what mixmatch is <a href=\"https://medium.com/&lt;a href=\">@sanjeev</a>.vadiraj/eureka-mixmatch-a-holistic-approach-to-semi-supervised-learning-125b14e82d2f\"&gt;https://medium.com/<a href=\"/sanjeev\">@sanjeev</a>.vadiraj/eureka-mixmatch-a-holistic-approach-to-semi-supervised-learning-125b14e82d2f</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 534676,
      "author_name": "jihangz",
      "author_url": "",
      "post_date": "05/21/2019 17:03:22",
      "content": "<p>Here's another unsupervised learning paper called \"Unsupervised Data Augmentation\" (UDA). It reaches lower error rate than MixMatch on CIFAR-10 given 4000 labels. <a href=\"https://arxiv.org/pdf/1904.12848.pdf\">https://arxiv.org/pdf/1904.12848.pdf</a></p>\n\n<p>Two papers use very similar loss functions (CE+L2 vs CE+KL). UDA is also trained on ImageNet and multiple text classification datasets.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 550292,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "06/11/2019 13:22:05",
      "content": "<p><a href=\"https://github.com/daisukelab/freesound-audio-tagging-2019\">https://github.com/daisukelab/freesound-audio-tagging-2019</a></p>\n\n<p>This repo has MixMatch implementation, though it would be incomplete. I was using this as super version of mixup...</p>\n\n<p>I'm hoping to have some more time to complete, to reproduce the original paper... fingers crossed..</p>",
      "votes": null,
      "replies": [
        {
          "id": 550316,
          "author_name": "mnpinto",
          "author_url": "",
          "post_date": "06/11/2019 13:49:51",
          "content": "<p>Thanks for sharing! I haven't been able to improve my results using MixMatch but comparing with my experiments on Cifar-10 I think it's probably a matter of getting all the hyper-parameters and model choice in the right spot. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 550331,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "06/11/2019 14:04:58",
          "content": "<p>Thanks for sharing your experience, ummm I see. What I suffered was prediction results of label guessing gets closer and closer like this:</p>\n\n<p>I was turning MixMatch on after once model is trained basically well. Then in earlier epochs, model predicts (guesses) like this:</p>\n\n<pre><code>max of predicted tensor([0.7737, 0.4485, 0.3357,  ..., 0.5738, 0.4036, 0.3020])\nmin of tensor([1.4771e-09, 9.2836e-08, 1.6407e-07,  ..., 1.1334e-07, 1.1591e-08 ...\n</code></pre>\n\n<p>Then after some more epochs, it predicts like:</p>\n\n<pre><code>max of predicted tensor([0.0455, 0.0465, 0.0464,  ..., 0.0581, 0.0484, 0.0472])\nmin of tensor([0.0115, 0.0112, 0.0115,  ..., 0.0115, 0.0113, 0.0115])\n</code></pre>\n\n<p>It gets closer until all <code>sigmoid(logits)</code> converges to single value.\nMy understanding so far is, model could be motivated to do so by MSE loss which encourages minimizing difference between guessed predictions with non-augmented image and predictions with augmented one. It'd be quite easy to achieve if model predicts a single value for all predictions...</p>\n\n<p>I'll have some time to debug more....</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 550372,
          "author_name": "mnpinto",
          "author_url": "",
          "post_date": "06/11/2019 14:56:21",
          "content": "<p>I was training with MixMatch from random initialization but the problem was similar. I noticed that lwlrap would increase to levels not far from mixup with curated data only, but Fbeta score does not increase unless I remove the MSE loss term as the predicted probabilities remained very small. This happened in my Cifar-10 experiments when the weight of MSE loss was increased to fast and also when using a larger model like resnet18 instead of the wide resnet used in the paper.  </p>\n\n<p>Perhaps for debugging a good idea is to start with just the curated data and only the single label  samples, using a few labeled samples and the remaining as unlabeled. If it works that way than it should be easier to move forward to add the multi-label samples and finally the noisy data.</p>\n\n<p>I will also try to find some time to organize and improve my implementation on fastai and to try on other datasets.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "529221": "https://twitter.com/D_Berthelot_ML/status/1125996671664451584\nhttp://arxiv.org/abs/1905.02249\n\nHere's another new paper that perfect matches to this competition needs!\n\n- '... MixMatch, that works by guessing low-entropy labels for data-augmented unlabeled examples and mixing labeled and unlabeled data using MixUp.' - Noisy set as unlabeled?\n- '... For example, on CIFAR-10 with 250 labels, we reduce error rate by a factor of 4 (from 38% to 11%) and by a factor of 2 on STL-10.' - Big improvement.\n- See algorithm 1, it's very interesting to create new batch from labeled and unlabeled samples with augmentation, pseudo labeling and mixup (https://arxiv.org/pdf/1710.09412.pdf).\n\nAnyone is going to try? :) There seems no implementation found on github yet. The author said on twitter that they are going to opensource it in few days (https://twitter.com/D_Berthelot_ML/status/1126393903022727168).\n\n## CAUTION: Don't use test samples in training\nBe sure NOT to use test set as unlabeled input, it's banned as clarified here: https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#524949",
    "529360": "I'm trying this :) The results described in the paper are really impressive!\n\nSo far it's looking promising, better than using the given labels of the noisy set.",
    "529389": "It's great, thanks for sharing! 👍",
    "529405": "Got LB 0.543 on first try, and it is still underfitting. With a similar setup I got LB 0.698 only with curated set. But again, for a first try it seems promising as I still need to debug more carefully my quick and dirty implementation of MixMatch, tune the hyperparameters and train for longer to see how far it goes :)",
    "529419": "Amazing it's so quick to implement! I'm happy to know that it would be working fine, new way to go.\nI'll keep checking how far you go on the LB! ;)",
    "531949": "FYI: https://github.com/google-research/mixmatch",
    "531980": "Thanks, you're always quick ;)",
    "532155": "And from the authors:\nhttps://github.com/google-research/mixmatch",
    "534107": "Hi @mnpinto , may I ask you how many epochs it takes for MixMatch to converge? I am still training it but it seems like taking forever.",
    "534142": "I'm not sure that I've got it working correctly. It's giving me a result very close but not better than with curated data only (even after training for very long), I may be missing something. I'm not using the authors code since I'm working in pytorch. \n\nAnd be aware that the sharpening function is not designed for multi-class classifications, that may also impact the results as it seems an important step according to the ablation study in the paper.",
    "534153": "I am using PyTorch as well. I think the sharpening function works for multi-label too since it just amplifies the magnitude of output in order to make it more \"binary\" (closer to either 0 or 1). \n\nWhile I am typing, I just realize a bug in my code.... I didn't take the sigmoid of my output before feeding it into the sharpening function.",
    "534158": "An easy reading blog of what mixmatch is https://medium.com/@sanjeev.vadiraj/eureka-mixmatch-a-holistic-approach-to-semi-supervised-learning-125b14e82d2f",
    "534166": "But it makes the result sum to 1 over all classes, more like the output of a softmax. It should work but probably it's not optimal. Did you try to reproduce the CIFAR-10 results of the paper? I guess it's the only way to make sure the code is really working without any bugs.",
    "534194": "You are right. I didn't try to reproduce the results of the paper. I will try to write a multi-label version of sharpening function.",
    "534214": "I've tried to reproduce the results last week but there are many things to consider like the same model, and data processing. I had no time to look into all that yet but if you are trying this other paper may be useful https://arxiv.org/abs/1903.03825, it presents a similar method (mentioned in MixMatch paper), but the good news is that the code is in PyTorch so it can be a good starting point: https://github.com/vikasverma1077/ICT",
    "534252": "mnpinto @jihangz Question for you guys tried mixmatch with pytorch, are you guys training with ema as the paper stated?",
    "534418": "I haven't tried EMA, they say \"EMA of parameter values hurt MixMatch’s performance slightly\" on the Ablation study.",
    "534490": "I also tried with pytorch, it did not converge to a high lwlrap.\nIt seems we do not need Augment(x) in this audio tagging task, because crop or resize may affect the logmel spectrogram.",
    "534520": "Oh cool, I have not got to that part yet. It was mentioned in the implementation detail section before the ablation.",
    "534526": "I haven't tried common image augmentation methods yet.  For longer clips, we would need to randomly truncate along the time axis to get them into required shapes. I think it could consider to be an augment function. \n\nI have just kicked off my mixmatch training with just that as aug, will see how it goes after work.",
    "534673": "500 epochs of training, local lwlrap is 0.736 now, and it is still increasing. I am going to train another 500 epochs to see how it goes.",
    "534676": "Here's another unsupervised learning paper called \"Unsupervised Data Augmentation\" (UDA). It reaches lower error rate than MixMatch on CIFAR-10 given 4000 labels. https://arxiv.org/pdf/1904.12848.pdf\n\nTwo papers use very similar loss functions (CE+L2 vs CE+KL). UDA is also trained on ImageNet and multiple text classification datasets.",
    "534682": "sailorwei I am using SpecAugment as Augment(x)",
    "534868": "I implemented MixMatch in TF eager. I used the same model for my best LB and was able to reach max 0.83 on one fold lwlrap (avg around 0.8 per fold for 5-fold). I used SpecAugment for augmentation component + random crops along the time_axis. However this only resulted in 0.65 LB so still not better than curated only.",
    "538627": "ryanzhang I almost fell off my chair! Is MixMatch the key of your amazing LB score?",
    "539221": "No, mixmatch is not working for me.",
    "539343": "Got it...  Nice work, anyway!",
    "550292": "https://github.com/daisukelab/freesound-audio-tagging-2019\n\nThis repo has MixMatch implementation, though it would be incomplete. I was using this as super version of mixup...\n\nI'm hoping to have some more time to complete, to reproduce the original paper... fingers crossed..",
    "550316": "Thanks for sharing! I haven't been able to improve my results using MixMatch but comparing with my experiments on Cifar-10 I think it's probably a matter of getting all the hyper-parameters and model choice in the right spot.",
    "550331": "Thanks for sharing your experience, ummm I see. What I suffered was prediction results of label guessing gets closer and closer like this:\n\nI was turning MixMatch on after once model is trained basically well. Then in earlier epochs, model predicts (guesses) like this:\n\n    max of predicted tensor([0.7737, 0.4485, 0.3357,  ..., 0.5738, 0.4036, 0.3020])\n    min of tensor([1.4771e-09, 9.2836e-08, 1.6407e-07,  ..., 1.1334e-07, 1.1591e-08 ...\n\nThen after some more epochs, it predicts like:\n\n    max of predicted tensor([0.0455, 0.0465, 0.0464,  ..., 0.0581, 0.0484, 0.0472])\n    min of tensor([0.0115, 0.0112, 0.0115,  ..., 0.0115, 0.0113, 0.0115])\n\nIt gets closer until all `sigmoid(logits)` converges to single value.\nMy understanding so far is, model could be motivated to do so by MSE loss which encourages minimizing difference between guessed predictions with non-augmented image and predictions with augmented one. It'd be quite easy to achieve if model predicts a single value for all predictions...\n\nI'll have some time to debug more....",
    "550372": "I was training with MixMatch from random initialization but the problem was similar. I noticed that lwlrap would increase to levels not far from mixup with curated data only, but Fbeta score does not increase unless I remove the MSE loss term as the predicted probabilities remained very small. This happened in my Cifar-10 experiments when the weight of MSE loss was increased to fast and also when using a larger model like resnet18 instead of the wide resnet used in the paper.  \n\nPerhaps for debugging a good idea is to start with just the curated data and only the single label  samples, using a few labeled samples and the remaining as unlabeled. If it works that way than it should be easier to move forward to add the multi-label samples and finally the noisy data.\n\nI will also try to find some time to organize and improve my implementation on fastai and to try on other datasets."
  },
  "source": "meta"
}