{
  "id": 126574,
  "title": "Implementation Details of Cutmix & Mixup?",
  "url": "/competitions/bengaliai-cv19/discussion/126574",
  "author_name": "",
  "post_date": "2020-01-18T12:46:33.328333500Z",
  "votes": 14,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Currently using cutmix + mixup (each .5 probability, no normal samples), I am able to achieve .970 CV. There is apparently room for improvement.</p>\n\n<p>Several interesting things I noticed when using cutmix + mixup\n- When using a single approach, cutmix is better\n- Training loss and validation loss fluctuate noticeably more, even when no cutmix/mixup is applied on validation samples (why???)\n- The validation loss is not necessarily lower while the validation metric is significantly higher (.5% boost for me)\n- I have to lengthen the training from 16 to 64 epochs (onecycle policy)</p>\n\n<p>According to some discussions, advantage of using cutmix+mixup should &gt;.5%. What are you able to achieve?\nAlso, I am using alpha=.4 (fast.ai) for mixup and beta=1 for cutmix (original paper). Is the performance sensitive to these parameters?</p>\n\n<p>Any comment appreciated, thx in advance</p>",
  "messages": [
    {
      "id": "722346",
      "postDate": "01/18/2020 12:46:33",
      "content": "<p>Currently using cutmix + mixup (each .5 probability, no normal samples), I am able to achieve .970 CV. There is apparently room for improvement.</p>\n\n<p>Several interesting things I noticed when using cutmix + mixup\n- When using a single approach, cutmix is better\n- Training loss and validation loss fluctuate noticeably more, even when no cutmix/mixup is applied on validation samples (why???)\n- The validation loss is not necessarily lower while the validation metric is significantly higher (.5% boost for me)\n- I have to lengthen the training from 16 to 64 epochs (onecycle policy)</p>\n\n<p>According to some discussions, advantage of using cutmix+mixup should &gt;.5%. What are you able to achieve?\nAlso, I am using alpha=.4 (fast.ai) for mixup and beta=1 for cutmix (original paper). Is the performance sensitive to these parameters?</p>\n\n<p>Any comment appreciated, thx in advance</p>",
      "rawMarkdown": "Currently using cutmix + mixup (each .5 probability, no normal samples), I am able to achieve .970 CV. There is apparently room for improvement.\n\nSeveral interesting things I noticed when using cutmix + mixup\n- When using a single approach, cutmix is better\n- Training loss and validation loss fluctuate noticeably more, even when no cutmix/mixup is applied on validation samples (why???)\n- The validation loss is not necessarily lower while the validation metric is significantly higher (.5% boost for me)\n- I have to lengthen the training from 16 to 64 epochs (onecycle policy)\n\nAccording to some discussions, advantage of using cutmix+mixup should &gt;.5%. What are you able to achieve?\nAlso, I am using alpha=.4 (fast.ai) for mixup and beta=1 for cutmix (original paper). Is the performance sensitive to these parameters?\n\nAny comment appreciated, thx in advance",
      "votes": null
    },
    {
      "id": "766088",
      "postDate": "03/07/2020 16:47:15",
      "content": "<p>I am sorry for late reply, let me put some thoughts on your points:\n- According to CutMix paper, it is better than Mixup which means CutMix is SOTA :)\n- The problem of fluctuation, as I understood, could be connected to the type of augmentation you implemented (it could be batchwise/samplewise and also mixing could be done between batch or within whole dataset). I tried all of these implementations, but had no time to compare because of lack of computational power :(\n- Sometimes loss and metric are not connected, so it's fine;\n- 64 epoch is not so much I think, what the best CV score did you get with it?\n- Alpha and beta should be tuned ;)</p>",
      "rawMarkdown": "I am sorry for late reply, let me put some thoughts on your points:\n- According to CutMix paper, it is better than Mixup which means CutMix is SOTA :)\n- The problem of fluctuation, as I understood, could be connected to the type of augmentation you implemented (it could be batchwise/samplewise and also mixing could be done between batch or within whole dataset). I tried all of these implementations, but had no time to compare because of lack of computational power :(\n- Sometimes loss and metric are not connected, so it's fine;\n- 64 epoch is not so much I think, what the best CV score did you get with it?\n- Alpha and beta should be tuned ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 766088,
      "author_name": "smivvla",
      "author_url": "",
      "post_date": "03/07/2020 16:47:15",
      "content": "<p>I am sorry for late reply, let me put some thoughts on your points:\n- According to CutMix paper, it is better than Mixup which means CutMix is SOTA :)\n- The problem of fluctuation, as I understood, could be connected to the type of augmentation you implemented (it could be batchwise/samplewise and also mixing could be done between batch or within whole dataset). I tried all of these implementations, but had no time to compare because of lack of computational power :(\n- Sometimes loss and metric are not connected, so it's fine;\n- 64 epoch is not so much I think, what the best CV score did you get with it?\n- Alpha and beta should be tuned ;)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "722346": "Currently using cutmix + mixup (each .5 probability, no normal samples), I am able to achieve .970 CV. There is apparently room for improvement.\n\nSeveral interesting things I noticed when using cutmix + mixup\n- When using a single approach, cutmix is better\n- Training loss and validation loss fluctuate noticeably more, even when no cutmix/mixup is applied on validation samples (why???)\n- The validation loss is not necessarily lower while the validation metric is significantly higher (.5% boost for me)\n- I have to lengthen the training from 16 to 64 epochs (onecycle policy)\n\nAccording to some discussions, advantage of using cutmix+mixup should &gt;.5%. What are you able to achieve?\nAlso, I am using alpha=.4 (fast.ai) for mixup and beta=1 for cutmix (original paper). Is the performance sensitive to these parameters?\n\nAny comment appreciated, thx in advance",
    "766088": "I am sorry for late reply, let me put some thoughts on your points:\n- According to CutMix paper, it is better than Mixup which means CutMix is SOTA :)\n- The problem of fluctuation, as I understood, could be connected to the type of augmentation you implemented (it could be batchwise/samplewise and also mixing could be done between batch or within whole dataset). I tried all of these implementations, but had no time to compare because of lack of computational power :(\n- Sometimes loss and metric are not connected, so it's fine;\n- 64 epoch is not so much I think, what the best CV score did you get with it?\n- Alpha and beta should be tuned ;)"
  },
  "source": "meta"
}