{
  "id": 159689,
  "title": "DropHead transformer regularization (with code)",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/159689",
  "author_name": "",
  "post_date": "2020-06-18T11:07:16.092547800Z",
  "votes": 28,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi all! I recently saw a <a href=\"https://arxiv.org/pdf/2004.13342.pdf\">paper</a> about special dropout for transformers which is basically about dropping an entire self-attention head instead of single neurons. The main idea is that with usual dropout you sometimes end up with some self-attention heads dominating over another heads and this method helps to avoid such situations.</p>\n\n<p>I had time to try it in the tweet sentiment competition and it seems it improved cv and private scores by simply replacing default dropout with drophead in Roberta model (setting default dropout to 0 and drophead to 0.1). I don't participate in this competition, so I am not sure if it works here. But in case you want to give it a try - here is my simple <a href=\"https://github.com/Kirill-Kravtsov/drophead-pytorch\">implementation</a>.</p>",
  "messages": [
    {
      "id": "891653",
      "postDate": "06/18/2020 11:07:16",
      "content": "<p>Hi all! I recently saw a <a href=\"https://arxiv.org/pdf/2004.13342.pdf\">paper</a> about special dropout for transformers which is basically about dropping an entire self-attention head instead of single neurons. The main idea is that with usual dropout you sometimes end up with some self-attention heads dominating over another heads and this method helps to avoid such situations.</p>\n\n<p>I had time to try it in the tweet sentiment competition and it seems it improved cv and private scores by simply replacing default dropout with drophead in Roberta model (setting default dropout to 0 and drophead to 0.1). I don't participate in this competition, so I am not sure if it works here. But in case you want to give it a try - here is my simple <a href=\"https://github.com/Kirill-Kravtsov/drophead-pytorch\">implementation</a>.</p>",
      "rawMarkdown": "Hi all! I recently saw a [paper](https://arxiv.org/pdf/2004.13342.pdf) about special dropout for transformers which is basically about dropping an entire self-attention head instead of single neurons. The main idea is that with usual dropout you sometimes end up with some self-attention heads dominating over another heads and this method helps to avoid such situations.\n\nI had time to try it in the tweet sentiment competition and it seems it improved cv and private scores by simply replacing default dropout with drophead in Roberta model (setting default dropout to 0 and drophead to 0.1). I don't participate in this competition, so I am not sure if it works here. But in case you want to give it a try - here is my simple [implementation](https://github.com/Kirill-Kravtsov/drophead-pytorch).",
      "votes": null
    },
    {
      "id": "891723",
      "postDate": "06/18/2020 12:18:08",
      "content": "<p>I can't wait to have a try! Thank you! <a href=\"/altkirill\">@altkirill</a> </p>",
      "rawMarkdown": "I can't wait to have a try! Thank you! @altkirill",
      "votes": null
    },
    {
      "id": "891760",
      "postDate": "06/18/2020 12:46:59",
      "content": "<p>Thanks for sharing Kirill.</p>",
      "rawMarkdown": "Thanks for sharing Kirill.",
      "votes": null
    },
    {
      "id": "892230",
      "postDate": "06/18/2020 18:20:06",
      "content": "<p>lucky you finding time to read papers 😏 thanks for contributing 👍 </p>",
      "rawMarkdown": "lucky you finding time to read papers 😏 thanks for contributing 👍",
      "votes": null
    },
    {
      "id": "892622",
      "postDate": "06/19/2020 03:42:32",
      "content": "<p>hello, <a href=\"/altkirill\">@altkirill</a> , can I use drophead together with dropout. Setting dropout to 0.3 other than aggresive 0.5, with 0.05 or 0.1 drophead, I think it may bring more balanced regularization. I am not sure this is right or effective. So if i am wrong, please let me know. </p>",
      "rawMarkdown": "hello, @altkirill , can I use drophead together with dropout. Setting dropout to 0.3 other than aggresive 0.5, with 0.05 or 0.1 drophead, I think it may bring more balanced regularization. I am not sure this is right or effective. So if i am wrong, please let me know.",
      "votes": null
    },
    {
      "id": "892801",
      "postDate": "06/19/2020 07:02:48",
      "content": "<p>Hi, from conceptual point of view I don't see any problem of using them together. I wrote about replacing just as an example. But about effectiveness only your cv could tell you :)</p>",
      "rawMarkdown": "Hi, from conceptual point of view I don't see any problem of using them together. I wrote about replacing just as an example. But about effectiveness only your cv could tell you :)",
      "votes": null
    },
    {
      "id": "896650",
      "postDate": "06/22/2020 10:48:11",
      "content": "<p>Your implementation is a DropHead with a fixed p_drophead, but have you tried the Scheduled DropHead as well?\nIf you've tried it, I'd like to know if you've improved your accuracy in the tweet sentiment competition.</p>",
      "rawMarkdown": "Your implementation is a DropHead with a fixed p_drophead, but have you tried the Scheduled DropHead as well?\nIf you've tried it, I'd like to know if you've improved your accuracy in the tweet sentiment competition.",
      "votes": null
    },
    {
      "id": "896761",
      "postDate": "06/22/2020 12:22:21",
      "content": "<p>I would say my implementation is a basic implementation which supports setting and changing <code>p_drophead</code> parameter, so it allows you to use it also for scheduled case. Yes, you need to implement scheduling (function: batch number -&gt; drophead probability) by yourself, but it would be strange to implement it on a model level.\nI've tried scheduling only once and it wasn't better but I assume scheduled drophead requires <code>p_drophead</code> retuning (my max <code>p_drophead</code> was also 0.1 and with scheduling it results in less regularization on average).</p>",
      "rawMarkdown": "I would say my implementation is a basic implementation which supports setting and changing `p_drophead` parameter, so it allows you to use it also for scheduled case. Yes, you need to implement scheduling (function: batch number -&gt; drophead probability) by yourself, but it would be strange to implement it on a model level.\nI've tried scheduling only once and it wasn't better but I assume scheduled drophead requires `p_drophead` retuning (my max `p_drophead` was also 0.1 and with scheduling it results in less regularization on average).",
      "votes": null
    },
    {
      "id": "897014",
      "postDate": "06/22/2020 15:10:23",
      "content": "<p>Thanks for the very helpful information.\nFrom your description, it seems that Scheduled DropHead doesn't necessarily work best.\nI'll try to implement it myself if I get a chance in a future competition. Thanks.</p>",
      "rawMarkdown": "Thanks for the very helpful information.\nFrom your description, it seems that Scheduled DropHead doesn't necessarily work best.\nI'll try to implement it myself if I get a chance in a future competition. Thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 891723,
      "author_name": "shangweichen",
      "author_url": "",
      "post_date": "06/18/2020 12:18:08",
      "content": "<p>I can't wait to have a try! Thank you! <a href=\"/altkirill\">@altkirill</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 891760,
      "author_name": "drpatrickchan",
      "author_url": "",
      "post_date": "06/18/2020 12:46:59",
      "content": "<p>Thanks for sharing Kirill.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 892230,
      "author_name": "dronych",
      "author_url": "",
      "post_date": "06/18/2020 18:20:06",
      "content": "<p>lucky you finding time to read papers 😏 thanks for contributing 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 892622,
      "author_name": "shangweichen",
      "author_url": "",
      "post_date": "06/19/2020 03:42:32",
      "content": "<p>hello, <a href=\"/altkirill\">@altkirill</a> , can I use drophead together with dropout. Setting dropout to 0.3 other than aggresive 0.5, with 0.05 or 0.1 drophead, I think it may bring more balanced regularization. I am not sure this is right or effective. So if i am wrong, please let me know. </p>",
      "votes": null,
      "replies": [
        {
          "id": 892801,
          "author_name": "altkirill",
          "author_url": "",
          "post_date": "06/19/2020 07:02:48",
          "content": "<p>Hi, from conceptual point of view I don't see any problem of using them together. I wrote about replacing just as an example. But about effectiveness only your cv could tell you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 896650,
      "author_name": "kanbehmw",
      "author_url": "",
      "post_date": "06/22/2020 10:48:11",
      "content": "<p>Your implementation is a DropHead with a fixed p_drophead, but have you tried the Scheduled DropHead as well?\nIf you've tried it, I'd like to know if you've improved your accuracy in the tweet sentiment competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 896761,
          "author_name": "altkirill",
          "author_url": "",
          "post_date": "06/22/2020 12:22:21",
          "content": "<p>I would say my implementation is a basic implementation which supports setting and changing <code>p_drophead</code> parameter, so it allows you to use it also for scheduled case. Yes, you need to implement scheduling (function: batch number -&gt; drophead probability) by yourself, but it would be strange to implement it on a model level.\nI've tried scheduling only once and it wasn't better but I assume scheduled drophead requires <code>p_drophead</code> retuning (my max <code>p_drophead</code> was also 0.1 and with scheduling it results in less regularization on average).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 897014,
          "author_name": "kanbehmw",
          "author_url": "",
          "post_date": "06/22/2020 15:10:23",
          "content": "<p>Thanks for the very helpful information.\nFrom your description, it seems that Scheduled DropHead doesn't necessarily work best.\nI'll try to implement it myself if I get a chance in a future competition. Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "891653": "Hi all! I recently saw a [paper](https://arxiv.org/pdf/2004.13342.pdf) about special dropout for transformers which is basically about dropping an entire self-attention head instead of single neurons. The main idea is that with usual dropout you sometimes end up with some self-attention heads dominating over another heads and this method helps to avoid such situations.\n\nI had time to try it in the tweet sentiment competition and it seems it improved cv and private scores by simply replacing default dropout with drophead in Roberta model (setting default dropout to 0 and drophead to 0.1). I don't participate in this competition, so I am not sure if it works here. But in case you want to give it a try - here is my simple [implementation](https://github.com/Kirill-Kravtsov/drophead-pytorch).",
    "891723": "I can't wait to have a try! Thank you! @altkirill",
    "891760": "Thanks for sharing Kirill.",
    "892230": "lucky you finding time to read papers 😏 thanks for contributing 👍",
    "892622": "hello, @altkirill , can I use drophead together with dropout. Setting dropout to 0.3 other than aggresive 0.5, with 0.05 or 0.1 drophead, I think it may bring more balanced regularization. I am not sure this is right or effective. So if i am wrong, please let me know.",
    "892801": "Hi, from conceptual point of view I don't see any problem of using them together. I wrote about replacing just as an example. But about effectiveness only your cv could tell you :)",
    "896650": "Your implementation is a DropHead with a fixed p_drophead, but have you tried the Scheduled DropHead as well?\nIf you've tried it, I'd like to know if you've improved your accuracy in the tweet sentiment competition.",
    "896761": "I would say my implementation is a basic implementation which supports setting and changing `p_drophead` parameter, so it allows you to use it also for scheduled case. Yes, you need to implement scheduling (function: batch number -&gt; drophead probability) by yourself, but it would be strange to implement it on a model level.\nI've tried scheduling only once and it wasn't better but I assume scheduled drophead requires `p_drophead` retuning (my max `p_drophead` was also 0.1 and with scheduling it results in less regularization on average).",
    "897014": "Thanks for the very helpful information.\nFrom your description, it seems that Scheduled DropHead doesn't necessarily work best.\nI'll try to implement it myself if I get a chance in a future competition. Thanks."
  },
  "source": "meta"
}