{
  "id": 76638,
  "title": "The batch_size magic and why this whole thing makes no sense whatsoever.",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76638",
  "author_name": "",
  "post_date": "2019-01-05T04:50:44.041158300Z",
  "votes": 19,
  "comment_count": 5,
  "views": 0,
  "content": "<p>So Kaggle just updated the kernel UI. One of the important changes is that you can now see the public scores of notebook kernels (which was then only available to script kernels, for whatever reason). </p>\n\n<p>Thanks to that new 'feature', I found this gem:</p>\n\n<p><a href=\"https://www.kaggle.com/hung96ad/pytorch-starter?scriptVersionId=8716013\">https://www.kaggle.com/hung96ad/pytorch-starter?scriptVersionId=8716013</a></p>\n\n<p><a href=\"https://www.kaggle.com/hung96ad/pytorch-starter?scriptVersionId=8734728\">https://www.kaggle.com/hung96ad/pytorch-starter?scriptVersionId=8734728</a></p>\n\n<p>The author literally just changed the batch size from 3072 to 1536, and the LB score went from 0.688 to 0.697.</p>\n\n<p>Not that I have anything against this kernel (I know the author btw, heh). This just shows how crazy the LB is and why you should be cautious when getting a sudden jump in score. And don't start some LB tuning by changing your batch size because of this thread, of course.</p>",
  "messages": [
    {
      "id": "450489",
      "postDate": "01/05/2019 04:50:44",
      "content": "<p>So Kaggle just updated the kernel UI. One of the important changes is that you can now see the public scores of notebook kernels (which was then only available to script kernels, for whatever reason). </p>\n\n<p>Thanks to that new 'feature', I found this gem:</p>\n\n<p><a href=\"https://www.kaggle.com/hung96ad/pytorch-starter?scriptVersionId=8716013\">https://www.kaggle.com/hung96ad/pytorch-starter?scriptVersionId=8716013</a></p>\n\n<p><a href=\"https://www.kaggle.com/hung96ad/pytorch-starter?scriptVersionId=8734728\">https://www.kaggle.com/hung96ad/pytorch-starter?scriptVersionId=8734728</a></p>\n\n<p>The author literally just changed the batch size from 3072 to 1536, and the LB score went from 0.688 to 0.697.</p>\n\n<p>Not that I have anything against this kernel (I know the author btw, heh). This just shows how crazy the LB is and why you should be cautious when getting a sudden jump in score. And don't start some LB tuning by changing your batch size because of this thread, of course.</p>",
      "rawMarkdown": "So Kaggle just updated the kernel UI. One of the important changes is that you can now see the public scores of notebook kernels (which was then only available to script kernels, for whatever reason). \n\nThanks to that new 'feature', I found this gem:\n\nhttps://www.kaggle.com/hung96ad/pytorch-starter?scriptVersionId=8716013\n\nhttps://www.kaggle.com/hung96ad/pytorch-starter?scriptVersionId=8734728\n\nThe author literally just changed the batch size from 3072 to 1536, and the LB score went from 0.688 to 0.697.\n\nNot that I have anything against this kernel (I know the author btw, heh). This just shows how crazy the LB is and why you should be cautious when getting a sudden jump in score. And don't start some LB tuning by changing your batch size because of this thread, of course.",
      "votes": null
    },
    {
      "id": "450493",
      "postDate": "01/05/2019 05:04:20",
      "content": "<p>Thanks for posting this issue. Actually, there are quite a number of literature saying that batch size does matter. In short, small size makes our gradient quite noisy and it can be thought of as regularization, and big size can smooth out the gradient (some papers compare this to a small learning rate). This is one example of one recent paper on ICLR 2018 (one of the best Deep Learning conference) which try to explore on the batch size power :</p>\n\n<p><strong>Don't Decay the Learning Rate, Increase the Batch Size</strong>\n<a href=\"https://openreview.net/forum?id=B1Yy1BxCZ\">https://openreview.net/forum?id=B1Yy1BxCZ</a></p>\n\n<p>And one example of discussion :\n<strong>Relation between learning rate, batch size and gradient noise in NN?</strong> (more interesting papers are pointed out here)\n<a href=\"https://www.reddit.com/r/MachineLearning/comments/84waz4/d_relation_between_learning_rate_batch_size_and/\">https://www.reddit.com/r/MachineLearning/comments/84waz4/d_relation_between_learning_rate_batch_size_and/</a></p>\n\n<p>So i think tuning batch size actually makes sense and is no magic in general (but still is a challenge how to find a best size or best schedule).</p>",
      "rawMarkdown": "Thanks for posting this issue. Actually, there are quite a number of literature saying that batch size does matter. In short, small size makes our gradient quite noisy and it can be thought of as regularization, and big size can smooth out the gradient (some papers compare this to a small learning rate). This is one example of one recent paper on ICLR 2018 (one of the best Deep Learning conference) which try to explore on the batch size power :\n\n**Don't Decay the Learning Rate, Increase the Batch Size**\nhttps://openreview.net/forum?id=B1Yy1BxCZ\n\nAnd one example of discussion :\n**Relation between learning rate, batch size and gradient noise in NN?** (more interesting papers are pointed out here)\nhttps://www.reddit.com/r/MachineLearning/comments/84waz4/d_relation_between_learning_rate_batch_size_and/\n\nSo i think tuning batch size actually makes sense and is no magic in general (but still is a challenge how to find a best size or best schedule).",
      "votes": null
    },
    {
      "id": "450494",
      "postDate": "01/05/2019 05:07:32",
      "content": "<p>Yup, batch_size is definitely one of the overlooked parameters in this competition, what I mean was tuning it with the LB which always returns an unexpected result it almost became expected. </p>\n\n<p>If you have a strong CV pipeline that you can trust then why not? </p>",
      "rawMarkdown": "Yup, batch_size is definitely one of the overlooked parameters in this competition, what I mean was tuning it with the LB which always returns an unexpected result it almost became expected. \n\nIf you have a strong CV pipeline that you can trust then why not?",
      "votes": null
    },
    {
      "id": "450791",
      "postDate": "01/05/2019 19:25:49",
      "content": "<p>Yes，that 0.697 kernel. I just change the random seed, and the LB become 0.69. It’s crazy for the LB only have such slightly distance less than 0.01 from gold medal to the public kernel.</p>",
      "rawMarkdown": "Yes，that 0.697 kernel. I just change the random seed, and the LB become 0.69. It’s crazy for the LB only have such slightly distance less than 0.01 from gold medal to the public kernel.",
      "votes": null
    },
    {
      "id": "451018",
      "postDate": "01/06/2019 09:31:37",
      "content": "<p>I totally agree @Neuron. Large batch_size can avoid the influence of data noisy, but it will weaken randomness. Before this thread, I try to set my batch_size as 1024 and find it cannot improve the score, so batch_size is not larger, the better.</p>",
      "rawMarkdown": "I totally agree @Neuron. Large batch_size can avoid the influence of data noisy, but it will weaken randomness. Before this thread, I try to set my batch_size as 1024 and find it cannot improve the score, so batch_size is not larger, the better.",
      "votes": null
    },
    {
      "id": "451289",
      "postDate": "01/06/2019 19:26:33",
      "content": "<p>I also noticed the same. Increasing Batch size decreased by local CV as well as LB</p>",
      "rawMarkdown": "I also noticed the same. Increasing Batch size decreased by local CV as well as LB",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 450493,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "01/05/2019 05:04:20",
      "content": "<p>Thanks for posting this issue. Actually, there are quite a number of literature saying that batch size does matter. In short, small size makes our gradient quite noisy and it can be thought of as regularization, and big size can smooth out the gradient (some papers compare this to a small learning rate). This is one example of one recent paper on ICLR 2018 (one of the best Deep Learning conference) which try to explore on the batch size power :</p>\n\n<p><strong>Don't Decay the Learning Rate, Increase the Batch Size</strong>\n<a href=\"https://openreview.net/forum?id=B1Yy1BxCZ\">https://openreview.net/forum?id=B1Yy1BxCZ</a></p>\n\n<p>And one example of discussion :\n<strong>Relation between learning rate, batch size and gradient noise in NN?</strong> (more interesting papers are pointed out here)\n<a href=\"https://www.reddit.com/r/MachineLearning/comments/84waz4/d_relation_between_learning_rate_batch_size_and/\">https://www.reddit.com/r/MachineLearning/comments/84waz4/d_relation_between_learning_rate_batch_size_and/</a></p>\n\n<p>So i think tuning batch size actually makes sense and is no magic in general (but still is a challenge how to find a best size or best schedule).</p>",
      "votes": null,
      "replies": [
        {
          "id": 450494,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "01/05/2019 05:07:32",
          "content": "<p>Yup, batch_size is definitely one of the overlooked parameters in this competition, what I mean was tuning it with the LB which always returns an unexpected result it almost became expected. </p>\n\n<p>If you have a strong CV pipeline that you can trust then why not? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 450791,
      "author_name": "peining",
      "author_url": "",
      "post_date": "01/05/2019 19:25:49",
      "content": "<p>Yes，that 0.697 kernel. I just change the random seed, and the LB become 0.69. It’s crazy for the LB only have such slightly distance less than 0.01 from gold medal to the public kernel.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 451018,
      "author_name": "salonsai",
      "author_url": "",
      "post_date": "01/06/2019 09:31:37",
      "content": "<p>I totally agree @Neuron. Large batch_size can avoid the influence of data noisy, but it will weaken randomness. Before this thread, I try to set my batch_size as 1024 and find it cannot improve the score, so batch_size is not larger, the better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 451289,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "01/06/2019 19:26:33",
          "content": "<p>I also noticed the same. Increasing Batch size decreased by local CV as well as LB</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "450489": "So Kaggle just updated the kernel UI. One of the important changes is that you can now see the public scores of notebook kernels (which was then only available to script kernels, for whatever reason). \n\nThanks to that new 'feature', I found this gem:\n\nhttps://www.kaggle.com/hung96ad/pytorch-starter?scriptVersionId=8716013\n\nhttps://www.kaggle.com/hung96ad/pytorch-starter?scriptVersionId=8734728\n\nThe author literally just changed the batch size from 3072 to 1536, and the LB score went from 0.688 to 0.697.\n\nNot that I have anything against this kernel (I know the author btw, heh). This just shows how crazy the LB is and why you should be cautious when getting a sudden jump in score. And don't start some LB tuning by changing your batch size because of this thread, of course.",
    "450493": "Thanks for posting this issue. Actually, there are quite a number of literature saying that batch size does matter. In short, small size makes our gradient quite noisy and it can be thought of as regularization, and big size can smooth out the gradient (some papers compare this to a small learning rate). This is one example of one recent paper on ICLR 2018 (one of the best Deep Learning conference) which try to explore on the batch size power :\n\n**Don't Decay the Learning Rate, Increase the Batch Size**\nhttps://openreview.net/forum?id=B1Yy1BxCZ\n\nAnd one example of discussion :\n**Relation between learning rate, batch size and gradient noise in NN?** (more interesting papers are pointed out here)\nhttps://www.reddit.com/r/MachineLearning/comments/84waz4/d_relation_between_learning_rate_batch_size_and/\n\nSo i think tuning batch size actually makes sense and is no magic in general (but still is a challenge how to find a best size or best schedule).",
    "450494": "Yup, batch_size is definitely one of the overlooked parameters in this competition, what I mean was tuning it with the LB which always returns an unexpected result it almost became expected. \n\nIf you have a strong CV pipeline that you can trust then why not?",
    "450791": "Yes，that 0.697 kernel. I just change the random seed, and the LB become 0.69. It’s crazy for the LB only have such slightly distance less than 0.01 from gold medal to the public kernel.",
    "451018": "I totally agree @Neuron. Large batch_size can avoid the influence of data noisy, but it will weaken randomness. Before this thread, I try to set my batch_size as 1024 and find it cannot improve the score, so batch_size is not larger, the better.",
    "451289": "I also noticed the same. Increasing Batch size decreased by local CV as well as LB"
  },
  "source": "meta"
}