{
  "id": 243174,
  "title": "How to choose alpha of mixup?",
  "url": "/competitions/seti-breakthrough-listen/discussion/243174",
  "author_name": "Tawara",
  "post_date": "2021-06-01T12:29:05.912000",
  "votes": 19,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Mixup seems to work very well in this competition.</p>\n<p>But I don't know how to choose alpha properly because this is the first time I use mixup.</p>\n<p>Can I assume some reasonable values of alpha based on the type of task or data? or, do I have to repeat the experiment endlessly?</p>\n<p>FYI: I tried only 0.2, 0.5 and 1.0. 1.0 works well for me.</p>",
  "messages": [
    {
      "id": 1331370,
      "postDate": "2021-06-01T12:29:05.913Z",
      "content": "<p>Mixup seems to work very well in this competition.</p>\n<p>But I don't know how to choose alpha properly because this is the first time I use mixup.</p>\n<p>Can I assume some reasonable values of alpha based on the type of task or data? or, do I have to repeat the experiment endlessly?</p>\n<p>FYI: I tried only 0.2, 0.5 and 1.0. 1.0 works well for me.</p>",
      "rawMarkdown": "Mixup seems to work very well in this competition.\n\nBut I don't know how to choose alpha properly because this is the first time I use mixup.\n\nCan I assume some reasonable values of alpha based on the type of task or data? or, do I have to repeat the experiment endlessly?\n\nFYI: I tried only 0.2, 0.5 and 1.0. 1.0 works well for me.",
      "votes": 19
    },
    {
      "id": 1334039,
      "postDate": "2021-06-03T08:30:43.757Z",
      "content": "<p>chris use alpha = np.random.uniform(0.19,0.31) for cutmix <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136025\" target=\"_blank\">https://www.kaggle.com/c/bengaliai-cv19/discussion/136025</a>, I tried this method in mixup，The results will improve slightly in my experiment。</p>",
      "rawMarkdown": "chris use alpha = np.random.uniform(0.19,0.31) for cutmix https://www.kaggle.com/c/bengaliai-cv19/discussion/136025, I tried this method in mixup，The results will improve slightly in my experiment。",
      "votes": 7
    },
    {
      "id": 1332727,
      "postDate": "2021-06-02T09:06:17.040Z",
      "content": "<p>Try use alpha 5 to make both sample get around 0.5 percentage for mixup, to avoid single sample got too less weights for mixup.</p>",
      "rawMarkdown": "Try use alpha 5 to make both sample get around 0.5 percentage for mixup, to avoid single sample got too less weights for mixup.",
      "votes": 1
    },
    {
      "id": 1332626,
      "postDate": "2021-06-02T08:11:27.757Z",
      "content": "<p>You can think of mixup as a regularizer and higher alpha leads to higher regularization. (Since we ask the network to predict points harder samples i.e. far away points from the original examples being mixed.)</p>\n<p>You can choose values to experiment with by applying similar heuristics we apply for any other regularizer. i.e. amount of underfitting/overfitting for the base model without mixup, model capacity, training budget, etc.</p>\n<p>I am, for now, using smaller models and have been using alpha=1.0, with a 50% mixup probability.</p>",
      "rawMarkdown": "You can think of mixup as a regularizer and higher alpha leads to higher regularization. (Since we ask the network to predict points harder samples i.e. far away points from the original examples being mixed.)\n\nYou can choose values to experiment with by applying similar heuristics we apply for any other regularizer. i.e. amount of underfitting/overfitting for the base model without mixup, model capacity, training budget, etc.\n\nI am, for now, using smaller models and have been using alpha=1.0, with a 50% mixup probability.",
      "votes": 1,
      "replies": [
        {
          "id": 1347156,
          "postDate": "2021-06-13T03:13:09.817Z",
          "content": "<p>Could you tell me how to use code to achieve 50% mixup，I also thought about doing this, but I don’t know how to implement it in code.</p>",
          "rawMarkdown": "Could you tell me how to use code to achieve 50% mixup，I also thought about doing this, but I don’t know how to implement it in code."
        },
        {
          "id": 1348805,
          "postDate": "2021-06-14T09:26:46.553Z",
          "content": "<p>I am just calling the mixup function based on sampled random value… something like this:</p>\n<p>import numpy as np<br>\nmixup_prob = 0.5</p>\n<p>p = np.random.rand()    # uniform random sampling between 0 and 1<br>\nif p &lt;= mixup_prob:<br>\n    perform_mixup()</p>",
          "rawMarkdown": "I am just calling the mixup function based on sampled random value... something like this:\n\n\nimport numpy as np\nmixup_prob = 0.5\n\np = np.random.rand()    # uniform random sampling between 0 and 1\nif p <= mixup_prob:\n\tperform_mixup()",
          "votes": 1
        }
      ]
    },
    {
      "id": 1332048,
      "postDate": "2021-06-01T21:37:08.920Z",
      "content": "<p>Hello, I have the same question and for now, I find that according to the <a href=\"https://arxiv.org/pdf/1710.09412.pdf\" target=\"_blank\">paper, page 5</a>: </p>\n<blockquote>\n  <p>For mixup, we find that α ∈ [0.1, 0.4] leads to improved performance over ERM, whereas for large α, mixup leads to underfitting</p>\n</blockquote>\n<p>I haven't read the full article so maybe there are some tips depending on the nature of the data.</p>",
      "rawMarkdown": "Hello, I have the same question and for now, I find that according to the [paper, page 5](https://arxiv.org/pdf/1710.09412.pdf): \n> For mixup, we find that α ∈ [0.1, 0.4] leads to improved performance over ERM, whereas for large α, mixup leads to underfitting\n\nI haven't read the full article so maybe there are some tips depending on the nature of the data.",
      "votes": 1,
      "replies": [
        {
          "id": 1332128,
          "postDate": "2021-06-01T23:51:24.120Z",
          "content": "<p>Thanks.</p>\n<p>I read the paper roughly and started from alpha=0.2 because 0.2 was chosen relatively many times.</p>\n<p>As you say, the difference between natural images(ImageNet, CIFAR) and this competition data may have an impact.</p>",
          "rawMarkdown": "Thanks.\n\nI read the paper roughly and started from alpha=0.2 because 0.2 was chosen relatively many times.\n\nAs you say, the difference between natural images(ImageNet, CIFAR) and this competition data may have an impact.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1331562,
      "postDate": "2021-06-01T14:48:39.730Z",
      "content": "<p>Not related to this but how did you implement mixup ?</p>",
      "rawMarkdown": "Not related to this but how did you implement mixup ?",
      "votes": 1,
      "replies": [
        {
          "id": 1331587,
          "postDate": "2021-06-01T15:04:52.663Z",
          "content": "<p>I referenced the implementation in <a href=\"https://github.com/facebookresearch/mixup-cifar10\" target=\"_blank\">facebook research's  repository</a>:  <br>\n<a href=\"https://github.com/facebookresearch/mixup-cifar10/blob/master/train.py#L119\" target=\"_blank\">https://github.com/facebookresearch/mixup-cifar10/blob/master/train.py#L119</a></p>\n<p>I think this implementation is same as one in this notebook:  <br>\n<a href=\"https://www.kaggle.com/micheomaano/efficientnet-b4-mixup-cv-0-98-lb-0-97\" target=\"_blank\">https://www.kaggle.com/micheomaano/efficientnet-b4-mixup-cv-0-98-lb-0-97</a></p>",
          "rawMarkdown": "I referenced the implementation in [facebook research's  repository](https://github.com/facebookresearch/mixup-cifar10):  \nhttps://github.com/facebookresearch/mixup-cifar10/blob/master/train.py#L119\n\nI think this implementation is same as one in this notebook:  \nhttps://www.kaggle.com/micheomaano/efficientnet-b4-mixup-cv-0-98-lb-0-97\n\n\n",
          "votes": 3
        },
        {
          "id": 1331942,
          "postDate": "2021-06-01T20:16:20.273Z",
          "content": "<p>Here is for tensorflow and keras:<br>\n<a href=\"https://keras.io/examples/vision/mixup/\" target=\"_blank\">https://keras.io/examples/vision/mixup/</a></p>",
          "rawMarkdown": "Here is for tensorflow and keras:\nhttps://keras.io/examples/vision/mixup/",
          "votes": 1
        },
        {
          "id": 1360974,
          "postDate": "2021-06-22T13:30:06.057Z",
          "content": "<p>if you are using timm, then you can maybe checkout timm.data.Mixup</p>",
          "rawMarkdown": "if you are using timm, then you can maybe checkout timm.data.Mixup",
          "votes": 1
        }
      ]
    },
    {
      "id": 1342435,
      "postDate": "2021-06-09T12:51:22.010Z",
      "content": "<p>I tried 2,5,10 when I read <a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/153257\" target=\"_blank\">https://www.kaggle.com/c/global-wheat-detection/discussion/153257</a> , but alpha=1 is best for me.</p>",
      "rawMarkdown": "I tried 2,5,10 when I read https://www.kaggle.com/c/global-wheat-detection/discussion/153257 , but alpha=1 is best for me.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1334039,
      "author_name": "MeisterMorxrc",
      "author_url": "",
      "post_date": "2021-06-03T08:30:43.757000",
      "content": "<p>chris use alpha = np.random.uniform(0.19,0.31) for cutmix <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136025\" target=\"_blank\">https://www.kaggle.com/c/bengaliai-cv19/discussion/136025</a>, I tried this method in mixup，The results will improve slightly in my experiment。</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1332727,
      "author_name": "Hao",
      "author_url": "",
      "post_date": "2021-06-02T09:06:17.040000",
      "content": "<p>Try use alpha 5 to make both sample get around 0.5 percentage for mixup, to avoid single sample got too less weights for mixup.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1332626,
      "author_name": "jvm",
      "author_url": "",
      "post_date": "2021-06-02T08:11:27.757000",
      "content": "<p>You can think of mixup as a regularizer and higher alpha leads to higher regularization. (Since we ask the network to predict points harder samples i.e. far away points from the original examples being mixed.)</p>\n<p>You can choose values to experiment with by applying similar heuristics we apply for any other regularizer. i.e. amount of underfitting/overfitting for the base model without mixup, model capacity, training budget, etc.</p>\n<p>I am, for now, using smaller models and have been using alpha=1.0, with a 50% mixup probability.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1347156,
          "author_name": "Gainover",
          "author_url": "",
          "post_date": "2021-06-13T03:13:09.817000",
          "content": "<p>Could you tell me how to use code to achieve 50% mixup，I also thought about doing this, but I don’t know how to implement it in code.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1348805,
          "author_name": "jvm",
          "author_url": "",
          "post_date": "2021-06-14T09:26:46.553000",
          "content": "<p>I am just calling the mixup function based on sampled random value… something like this:</p>\n<p>import numpy as np<br>\nmixup_prob = 0.5</p>\n<p>p = np.random.rand()    # uniform random sampling between 0 and 1<br>\nif p &lt;= mixup_prob:<br>\n    perform_mixup()</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1332048,
      "author_name": "Sergei Nazarikov",
      "author_url": "",
      "post_date": "2021-06-01T21:37:08.920000",
      "content": "<p>Hello, I have the same question and for now, I find that according to the <a href=\"https://arxiv.org/pdf/1710.09412.pdf\" target=\"_blank\">paper, page 5</a>: </p>\n<blockquote>\n  <p>For mixup, we find that α ∈ [0.1, 0.4] leads to improved performance over ERM, whereas for large α, mixup leads to underfitting</p>\n</blockquote>\n<p>I haven't read the full article so maybe there are some tips depending on the nature of the data.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1332128,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-06-01T23:51:24.120000",
          "content": "<p>Thanks.</p>\n<p>I read the paper roughly and started from alpha=0.2 because 0.2 was chosen relatively many times.</p>\n<p>As you say, the difference between natural images(ImageNet, CIFAR) and this competition data may have an impact.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1331562,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-06-01T14:48:39.730000",
      "content": "<p>Not related to this but how did you implement mixup ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1331587,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-06-01T15:04:52.663000",
          "content": "<p>I referenced the implementation in <a href=\"https://github.com/facebookresearch/mixup-cifar10\" target=\"_blank\">facebook research's  repository</a>:  <br>\n<a href=\"https://github.com/facebookresearch/mixup-cifar10/blob/master/train.py#L119\" target=\"_blank\">https://github.com/facebookresearch/mixup-cifar10/blob/master/train.py#L119</a></p>\n<p>I think this implementation is same as one in this notebook:  <br>\n<a href=\"https://www.kaggle.com/micheomaano/efficientnet-b4-mixup-cv-0-98-lb-0-97\" target=\"_blank\">https://www.kaggle.com/micheomaano/efficientnet-b4-mixup-cv-0-98-lb-0-97</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1331942,
          "author_name": "Krishna kant singh",
          "author_url": "",
          "post_date": "2021-06-01T20:16:20.273000",
          "content": "<p>Here is for tensorflow and keras:<br>\n<a href=\"https://keras.io/examples/vision/mixup/\" target=\"_blank\">https://keras.io/examples/vision/mixup/</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1360974,
          "author_name": "datta",
          "author_url": "",
          "post_date": "2021-06-22T13:30:06.057000",
          "content": "<p>if you are using timm, then you can maybe checkout timm.data.Mixup</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1342435,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2021-06-09T12:51:22.010000",
      "content": "<p>I tried 2,5,10 when I read <a href=\"https://www.kaggle.com/c/global-wheat-detection/discussion/153257\" target=\"_blank\">https://www.kaggle.com/c/global-wheat-detection/discussion/153257</a> , but alpha=1 is best for me.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1331370": "Mixup seems to work very well in this competition.\n\nBut I don't know how to choose alpha properly because this is the first time I use mixup.\n\nCan I assume some reasonable values of alpha based on the type of task or data? or, do I have to repeat the experiment endlessly?\n\nFYI: I tried only 0.2, 0.5 and 1.0. 1.0 works well for me.",
    "1334039": "chris use alpha = np.random.uniform(0.19,0.31) for cutmix https://www.kaggle.com/c/bengaliai-cv19/discussion/136025, I tried this method in mixup，The results will improve slightly in my experiment。",
    "1332727": "Try use alpha 5 to make both sample get around 0.5 percentage for mixup, to avoid single sample got too less weights for mixup.",
    "1332626": "You can think of mixup as a regularizer and higher alpha leads to higher regularization. (Since we ask the network to predict points harder samples i.e. far away points from the original examples being mixed.)\n\nYou can choose values to experiment with by applying similar heuristics we apply for any other regularizer. i.e. amount of underfitting/overfitting for the base model without mixup, model capacity, training budget, etc.\n\nI am, for now, using smaller models and have been using alpha=1.0, with a 50% mixup probability.",
    "1332048": "Hello, I have the same question and for now, I find that according to the [paper, page 5](https://arxiv.org/pdf/1710.09412.pdf): \n> For mixup, we find that α ∈ [0.1, 0.4] leads to improved performance over ERM, whereas for large α, mixup leads to underfitting\n\nI haven't read the full article so maybe there are some tips depending on the nature of the data.",
    "1331562": "Not related to this but how did you implement mixup ?",
    "1342435": "I tried 2,5,10 when I read https://www.kaggle.com/c/global-wheat-detection/discussion/153257 , but alpha=1 is best for me."
  }
}