{
  "id": 245156,
  "title": "Should we assign probabilities to Mixup?",
  "url": "/competitions/seti-breakthrough-listen/discussion/245156",
  "author_name": "gao-hongnan",
  "post_date": "2021-06-10T01:39:42.581000",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Following <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> discussion <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/244125\" target=\"_blank\">here on Mixup's Alpha</a>, I have a few additional questions.</p>\n<ol>\n<li><p>I have noticed people apply probability during training. For example, <br>\n<code>if random(0,1) &gt; 0.5: do mixup else don't do mixup</code><br>\nIs this better than just do mixup on each batch? For normal augmentations I can understand that you should not have a probability of 1, just because that you want the model to see different variants in each batch. Should this logic be carried over in mixup?</p></li>\n<li><p>I have read somewhere that people do some warmup epochs with less augmentation, and then progressively increase the augmentations done, is this good as well?</p></li>\n</ol>",
  "messages": [
    {
      "id": 1343037,
      "postDate": "2021-06-10T01:39:42.580Z",
      "content": "<p>Following <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> discussion <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/244125\" target=\"_blank\">here on Mixup's Alpha</a>, I have a few additional questions.</p>\n<ol>\n<li><p>I have noticed people apply probability during training. For example, <br>\n<code>if random(0,1) &gt; 0.5: do mixup else don't do mixup</code><br>\nIs this better than just do mixup on each batch? For normal augmentations I can understand that you should not have a probability of 1, just because that you want the model to see different variants in each batch. Should this logic be carried over in mixup?</p></li>\n<li><p>I have read somewhere that people do some warmup epochs with less augmentation, and then progressively increase the augmentations done, is this good as well?</p></li>\n</ol>",
      "rawMarkdown": "Following @ttahara discussion [here on Mixup's Alpha](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/244125), I have a few additional questions.\n\n1. I have noticed people apply probability during training. For example, \n`if random(0,1) > 0.5: do mixup else don't do mixup`\nIs this better than just do mixup on each batch? For normal augmentations I can understand that you should not have a probability of 1, just because that you want the model to see different variants in each batch. Should this logic be carried over in mixup?\n\n2. I have read somewhere that people do some warmup epochs with less augmentation, and then progressively increase the augmentations done, is this good as well?\n\n",
      "votes": 11
    },
    {
      "id": 1343169,
      "postDate": "2021-06-10T04:56:48.230Z",
      "content": "<ol>\n<li><p>I usually do mixup with <code>p=1.</code>. I did try using <code>p=0.5</code> in the past, but did not find difference significant enough for me to want to choose <code>p=0.5</code>. The best is to implement both try and see the differences.</p></li>\n<li><p>It depends on the model, the loss, and the data. I trained SED models with warm-up epochs on log-melspec (in Birdclef) with minimal augmentations. The idea of warm-up is to let the model have a more stable training regime, else the model may not converge in some cases. This is discussed in some old-ish (2017) conference paper.</p></li>\n</ol>",
      "rawMarkdown": "1. I usually do mixup with `p=1.`. I did try using `p=0.5` in the past, but did not find difference significant enough for me to want to choose `p=0.5`. The best is to implement both try and see the differences.\n\n2. It depends on the model, the loss, and the data. I trained SED models with warm-up epochs on log-melspec (in Birdclef) with minimal augmentations. The idea of warm-up is to let the model have a more stable training regime, else the model may not converge in some cases. This is discussed in some old-ish (2017) conference paper.",
      "votes": 1
    },
    {
      "id": 1343157,
      "postDate": "2021-06-10T04:45:09.970Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1343169,
      "author_name": "JunYong Tong",
      "author_url": "",
      "post_date": "2021-06-10T04:56:48.230000",
      "content": "<ol>\n<li><p>I usually do mixup with <code>p=1.</code>. I did try using <code>p=0.5</code> in the past, but did not find difference significant enough for me to want to choose <code>p=0.5</code>. The best is to implement both try and see the differences.</p></li>\n<li><p>It depends on the model, the loss, and the data. I trained SED models with warm-up epochs on log-melspec (in Birdclef) with minimal augmentations. The idea of warm-up is to let the model have a more stable training regime, else the model may not converge in some cases. This is discussed in some old-ish (2017) conference paper.</p></li>\n</ol>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1343157,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-10T04:45:09.970000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1343037": "Following @ttahara discussion [here on Mixup's Alpha](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/244125), I have a few additional questions.\n\n1. I have noticed people apply probability during training. For example, \n`if random(0,1) > 0.5: do mixup else don't do mixup`\nIs this better than just do mixup on each batch? For normal augmentations I can understand that you should not have a probability of 1, just because that you want the model to see different variants in each batch. Should this logic be carried over in mixup?\n\n2. I have read somewhere that people do some warmup epochs with less augmentation, and then progressively increase the augmentations done, is this good as well?\n\n",
    "1343169": "1. I usually do mixup with `p=1.`. I did try using `p=0.5` in the past, but did not find difference significant enough for me to want to choose `p=0.5`. The best is to implement both try and see the differences.\n\n2. It depends on the model, the loss, and the data. I trained SED models with warm-up epochs on log-melspec (in Birdclef) with minimal augmentations. The idea of warm-up is to let the model have a more stable training regime, else the model may not converge in some cases. This is discussed in some old-ish (2017) conference paper.",
    "1343157": ""
  }
}