{
  "id": 246018,
  "title": "Doesn't mix-up training affect batch normalization?",
  "url": "/competitions/seti-breakthrough-listen/discussion/246018",
  "author_name": "WOOSUNG YOON",
  "post_date": "2021-06-13T14:23:26.036000",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The batch normalization learns the mean and variance of data in a mini-batch size. <br>\n(I am not sure in PyTorch and TensorFlow)<br>\nIn my experiment, I think the mix-up method is unsafe to the training batch normalization . The performance increases temporarily (4-5 epoch in transfer learning) and then decreases again.</p>\n<p>It seems the neural networks that depend on the convolution with batch normalization are reduced performance to predict the original data. (The mean of the mix-up is constant, but I am not sure if the standard deviation is also constant.)<br>\n(I think the ResNet was strong in the mix-up method or rotation augmentation.)</p>\n<p>I think the batch normalization function may not be trained after getting some training. </p>\n<p>(And since the needle image is a very small feature, I think that the learning rate should be started high and then sharply lowered. I think the intermediate range of learning rates limits the convergence of the model. Instead, it may be better to increase the batch size gradually.)</p>",
  "messages": [
    {
      "id": 1347865,
      "postDate": "2021-06-13T14:23:26.037Z",
      "content": "<p>The batch normalization learns the mean and variance of data in a mini-batch size. <br>\n(I am not sure in PyTorch and TensorFlow)<br>\nIn my experiment, I think the mix-up method is unsafe to the training batch normalization . The performance increases temporarily (4-5 epoch in transfer learning) and then decreases again.</p>\n<p>It seems the neural networks that depend on the convolution with batch normalization are reduced performance to predict the original data. (The mean of the mix-up is constant, but I am not sure if the standard deviation is also constant.)<br>\n(I think the ResNet was strong in the mix-up method or rotation augmentation.)</p>\n<p>I think the batch normalization function may not be trained after getting some training. </p>\n<p>(And since the needle image is a very small feature, I think that the learning rate should be started high and then sharply lowered. I think the intermediate range of learning rates limits the convergence of the model. Instead, it may be better to increase the batch size gradually.)</p>",
      "rawMarkdown": "The batch normalization learns the mean and variance of data in a mini-batch size. \n(I am not sure in PyTorch and TensorFlow)\nIn my experiment, I think the mix-up method is unsafe to the training batch normalization . The performance increases temporarily (4-5 epoch in transfer learning) and then decreases again.\n\nIt seems the neural networks that depend on the convolution with batch normalization are reduced performance to predict the original data. (The mean of the mix-up is constant, but I am not sure if the standard deviation is also constant.)\n(I think the ResNet was strong in the mix-up method or rotation augmentation.)\n\nI think the batch normalization function may not be trained after getting some training. \n\n(And since the needle image is a very small feature, I think that the learning rate should be started high and then sharply lowered. I think the intermediate range of learning rates limits the convergence of the model. Instead, it may be better to increase the batch size gradually.)\n",
      "votes": 2
    },
    {
      "id": 1348849,
      "postDate": "2021-06-14T10:23:33Z",
      "content": "<blockquote>\n  <p>(The mean of the mix-up is constant, but I am not sure if the standard deviation is also constant.) </p>\n</blockquote>\n<p>As mentioned <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/245251\" target=\"_blank\">here</a>, to regularize std for the input features you can simply square the lamba coeff (sum of normal distribution properties).</p>",
      "rawMarkdown": "> (The mean of the mix-up is constant, but I am not sure if the standard deviation is also constant.) \n\nAs mentioned [here](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/245251), to regularize std for the input features you can simply square the lamba coeff (sum of normal distribution properties)."
    },
    {
      "id": 1348283,
      "postDate": "2021-06-13T23:28:43.847Z",
      "content": "<p>in my case,<br>\nin batch mixup (order shuffle &amp; mixup) is better than two batch mixup (val auc &amp; time both)<br>\nbut I use two batch mixup</p>\n<p><code>The performance increases temporarily (4-5 epoch in transfer learning) and then decreases again.</code><br>\nI think it is related to <code>Deep Double Descent</code>, it might related to dataset noise</p>\n<p>try to read <a href=\"https://arxiv.org/pdf/1912.02292.pdf\" target=\"_blank\">paper</a> this is one of my favorite paper </p>",
      "rawMarkdown": "in my case,\nin batch mixup (order shuffle & mixup) is better than two batch mixup (val auc & time both)\nbut I use two batch mixup\n\n`The performance increases temporarily (4-5 epoch in transfer learning) and then decreases again.`\nI think it is related to `Deep Double Descent`, it might related to dataset noise\n\ntry to read [paper](https://arxiv.org/pdf/1912.02292.pdf) this is one of my favorite paper \n "
    },
    {
      "id": 1348217,
      "postDate": "2021-06-13T20:57:28.170Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1348849,
      "author_name": "Riccardo",
      "author_url": "",
      "post_date": "2021-06-14T10:23:33",
      "content": "<blockquote>\n  <p>(The mean of the mix-up is constant, but I am not sure if the standard deviation is also constant.) </p>\n</blockquote>\n<p>As mentioned <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/245251\" target=\"_blank\">here</a>, to regularize std for the input features you can simply square the lamba coeff (sum of normal distribution properties).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1348283,
      "author_name": "assign",
      "author_url": "",
      "post_date": "2021-06-13T23:28:43.847000",
      "content": "<p>in my case,<br>\nin batch mixup (order shuffle &amp; mixup) is better than two batch mixup (val auc &amp; time both)<br>\nbut I use two batch mixup</p>\n<p><code>The performance increases temporarily (4-5 epoch in transfer learning) and then decreases again.</code><br>\nI think it is related to <code>Deep Double Descent</code>, it might related to dataset noise</p>\n<p>try to read <a href=\"https://arxiv.org/pdf/1912.02292.pdf\" target=\"_blank\">paper</a> this is one of my favorite paper </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1348217,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-13T20:57:28.170000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1347865": "The batch normalization learns the mean and variance of data in a mini-batch size. \n(I am not sure in PyTorch and TensorFlow)\nIn my experiment, I think the mix-up method is unsafe to the training batch normalization . The performance increases temporarily (4-5 epoch in transfer learning) and then decreases again.\n\nIt seems the neural networks that depend on the convolution with batch normalization are reduced performance to predict the original data. (The mean of the mix-up is constant, but I am not sure if the standard deviation is also constant.)\n(I think the ResNet was strong in the mix-up method or rotation augmentation.)\n\nI think the batch normalization function may not be trained after getting some training. \n\n(And since the needle image is a very small feature, I think that the learning rate should be started high and then sharply lowered. I think the intermediate range of learning rates limits the convergence of the model. Instead, it may be better to increase the batch size gradually.)\n",
    "1348849": "> (The mean of the mix-up is constant, but I am not sure if the standard deviation is also constant.) \n\nAs mentioned [here](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/245251), to regularize std for the input features you can simply square the lamba coeff (sum of normal distribution properties).",
    "1348283": "in my case,\nin batch mixup (order shuffle & mixup) is better than two batch mixup (val auc & time both)\nbut I use two batch mixup\n\n`The performance increases temporarily (4-5 epoch in transfer learning) and then decreases again.`\nI think it is related to `Deep Double Descent`, it might related to dataset noise\n\ntry to read [paper](https://arxiv.org/pdf/1912.02292.pdf) this is one of my favorite paper \n ",
    "1348217": ""
  }
}