{
  "id": 79205,
  "title": "Why so many kernel use batchnormalization with the dropout?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79205",
  "author_name": "",
  "post_date": "2019-02-01T07:02:33.730434500Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>according to this paper <a href=\"http://link.zhihu.com/?target=https%3A//arxiv.org/pdf/1801.05134.pdf\">http://link.zhihu.com/?target=https%3A//arxiv.org/pdf/1801.05134.pdf</a>\nit seems to have some disharmony between dropout and Batch Normalization, but in my experience, adding BN do improve the score in this competition, is that because the BN have a small momentum 0.5?</p>",
  "messages": [
    {
      "id": "464624",
      "postDate": "02/01/2019 07:02:33",
      "content": "<p>according to this paper <a href=\"http://link.zhihu.com/?target=https%3A//arxiv.org/pdf/1801.05134.pdf\">http://link.zhihu.com/?target=https%3A//arxiv.org/pdf/1801.05134.pdf</a>\nit seems to have some disharmony between dropout and Batch Normalization, but in my experience, adding BN do improve the score in this competition, is that because the BN have a small momentum 0.5?</p>",
      "rawMarkdown": "according to this paper http://link.zhihu.com/?target=https%3A//arxiv.org/pdf/1801.05134.pdf\nit seems to have some disharmony between dropout and Batch Normalization, but in my experience, adding BN do improve the score in this competition, is that because the BN have a small momentum 0.5?",
      "votes": null
    },
    {
      "id": "464848",
      "postDate": "02/01/2019 16:23:04",
      "content": "<p>Adding the BN layer should be the first in my kernel. \n<a href=\"https://www.kaggle.com/jetouxu/bilstm-attention-kfold-clr-extra-features-bn\">https://www.kaggle.com/jetouxu/bilstm-attention-kfold-clr-extra-features-bn</a> \n As for 0.5, I used it randomly.</p>",
      "rawMarkdown": "Adding the BN layer should be the first in my kernel. \nhttps://www.kaggle.com/jetouxu/bilstm-attention-kfold-clr-extra-features-bn \n As for 0.5, I used it randomly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 464848,
      "author_name": "jetouxu",
      "author_url": "",
      "post_date": "02/01/2019 16:23:04",
      "content": "<p>Adding the BN layer should be the first in my kernel. \n<a href=\"https://www.kaggle.com/jetouxu/bilstm-attention-kfold-clr-extra-features-bn\">https://www.kaggle.com/jetouxu/bilstm-attention-kfold-clr-extra-features-bn</a> \n As for 0.5, I used it randomly.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "464624": "according to this paper http://link.zhihu.com/?target=https%3A//arxiv.org/pdf/1801.05134.pdf\nit seems to have some disharmony between dropout and Batch Normalization, but in my experience, adding BN do improve the score in this competition, is that because the BN have a small momentum 0.5?",
    "464848": "Adding the BN layer should be the first in my kernel. \nhttps://www.kaggle.com/jetouxu/bilstm-attention-kfold-clr-extra-features-bn \n As for 0.5, I used it randomly."
  },
  "source": "meta"
}