{
  "id": 157296,
  "title": "Taming the randomness while fine-tuning - Stability of fine-tuning BERT",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/157296",
  "author_name": "",
  "post_date": "2020-06-10T05:15:30.040930200Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks to <a href=\"/grohith327\">@grohith327</a> who posted an interesting paper url in Tweet competition <a href=\"https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/157262\">here</a>.</p>\n\n<p>The paper notes: \"... fine-tuning a model multiple times on the same dataset, varying only the random seed, leads to a large standard deviation of the fine-tuning accuracy (Devlin et al., 2019;Dodge et al., 2020)\" . They propose a baseline to deal with the spurious pattern.</p>\n\n<p>From the paper:\"\n- Use small learning rates combined with bias correction to avoid vanishing gradients early intraining.\n- Increase the number of iterations considerably and train to (almost) zero training loss whilemaking use of early stopping.\"</p>\n\n<p>It would be interesting to know who does what to deal with the scores variability. I personally couldn't find any remedy to this.</p>",
  "messages": [
    {
      "id": "880204",
      "postDate": "06/10/2020 05:15:30",
      "content": "<p>Thanks to <a href=\"/grohith327\">@grohith327</a> who posted an interesting paper url in Tweet competition <a href=\"https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/157262\">here</a>.</p>\n\n<p>The paper notes: \"... fine-tuning a model multiple times on the same dataset, varying only the random seed, leads to a large standard deviation of the fine-tuning accuracy (Devlin et al., 2019;Dodge et al., 2020)\" . They propose a baseline to deal with the spurious pattern.</p>\n\n<p>From the paper:\"\n- Use small learning rates combined with bias correction to avoid vanishing gradients early intraining.\n- Increase the number of iterations considerably and train to (almost) zero training loss whilemaking use of early stopping.\"</p>\n\n<p>It would be interesting to know who does what to deal with the scores variability. I personally couldn't find any remedy to this.</p>",
      "rawMarkdown": "Thanks to @grohith327 who posted an interesting paper url in Tweet competition [here](https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/157262).\n\nThe paper notes: \"... fine-tuning a model multiple times on the same dataset, varying only the random seed, leads to a large standard deviation of the fine-tuning accuracy (Devlin et al., 2019;Dodge et al., 2020)\" . They propose a baseline to deal with the spurious pattern.\n\nFrom the paper:\"\n- Use small learning rates combined with bias correction to avoid vanishing gradients early intraining.\n- Increase the number of iterations considerably and train to (almost) zero training loss whilemaking use of early stopping.\"\n\nIt would be interesting to know who does what to deal with the scores variability. I personally couldn't find any remedy to this.",
      "votes": null
    },
    {
      "id": "880329",
      "postDate": "06/10/2020 07:20:00",
      "content": "<p>hello, <a href=\"/isakev\">@isakev</a> , This problem also bothers me. My keras based model cannot reproduce the results using same random seed. My pytorch based model can reproduce well on kaggle TPU, but <a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">Shonenkov's notebook</a> cannot reproduce on Colab TPU， I don' know why. My previous pytorch based model on kaggle TPU changes a lot(from 0.91 to 0.72 at public leaderboard) when I using different random seed.</p>",
      "rawMarkdown": "hello, @isakev , This problem also bothers me. My keras based model cannot reproduce the results using same random seed. My pytorch based model can reproduce well on kaggle TPU, but [Shonenkov's notebook](https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta) cannot reproduce on Colab TPU， I don' know why. My previous pytorch based model on kaggle TPU changes a lot(from 0.91 to 0.72 at public leaderboard) when I using different random seed.",
      "votes": null
    },
    {
      "id": "880535",
      "postDate": "06/10/2020 11:24:31",
      "content": "<p>With the same (for python, tensorflow, and numpy) random seed my scores changed once by 0.003 (0.9390-0.9420). I did not test it again, though. Explanation might be that the TPU brings randomness when splitting batches across its 8 units. I have no idea how to deal with this problem except for treating  the scores with a grain of doubt, or maybe manually managing distribution of batches at a low level programmatically.</p>",
      "rawMarkdown": "With the same (for python, tensorflow, and numpy) random seed my scores changed once by 0.003 (0.9390-0.9420). I did not test it again, though. Explanation might be that the TPU brings randomness when splitting batches across its 8 units. I have no idea how to deal with this problem except for treating  the scores with a grain of doubt, or maybe manually managing distribution of batches at a low level programmatically.",
      "votes": null
    },
    {
      "id": "881266",
      "postDate": "06/10/2020 21:44:42",
      "content": "<p>That is something thats been baffling me in this competition. And it makes experimenting really hard. i mostly worked with image data before this competition and I'm not used to this level of instability and scatter with CNNs. Honestly, I have no idea why does this happen.</p>\n\n<p>The best I could come up with was first training with MLM, and after that using a different learning rate for the head and the transformer layers. While the scatter is not completely gone this way, but it is somewhat smaller.\nI posted these  here: <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/156622\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/156622</a></p>",
      "rawMarkdown": "That is something thats been baffling me in this competition. And it makes experimenting really hard. i mostly worked with image data before this competition and I'm not used to this level of instability and scatter with CNNs. Honestly, I have no idea why does this happen.\n\nThe best I could come up with was first training with MLM, and after that using a different learning rate for the head and the transformer layers. While the scatter is not completely gone this way, but it is somewhat smaller.\nI posted these  here: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/156622",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 880329,
      "author_name": "shangweichen",
      "author_url": "",
      "post_date": "06/10/2020 07:20:00",
      "content": "<p>hello, <a href=\"/isakev\">@isakev</a> , This problem also bothers me. My keras based model cannot reproduce the results using same random seed. My pytorch based model can reproduce well on kaggle TPU, but <a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">Shonenkov's notebook</a> cannot reproduce on Colab TPU， I don' know why. My previous pytorch based model on kaggle TPU changes a lot(from 0.91 to 0.72 at public leaderboard) when I using different random seed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 880535,
          "author_name": "isakev",
          "author_url": "",
          "post_date": "06/10/2020 11:24:31",
          "content": "<p>With the same (for python, tensorflow, and numpy) random seed my scores changed once by 0.003 (0.9390-0.9420). I did not test it again, though. Explanation might be that the TPU brings randomness when splitting batches across its 8 units. I have no idea how to deal with this problem except for treating  the scores with a grain of doubt, or maybe manually managing distribution of batches at a low level programmatically.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 881266,
      "author_name": "riblidezso",
      "author_url": "",
      "post_date": "06/10/2020 21:44:42",
      "content": "<p>That is something thats been baffling me in this competition. And it makes experimenting really hard. i mostly worked with image data before this competition and I'm not used to this level of instability and scatter with CNNs. Honestly, I have no idea why does this happen.</p>\n\n<p>The best I could come up with was first training with MLM, and after that using a different learning rate for the head and the transformer layers. While the scatter is not completely gone this way, but it is somewhat smaller.\nI posted these  here: <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/156622\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/156622</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "880204": "Thanks to @grohith327 who posted an interesting paper url in Tweet competition [here](https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/157262).\n\nThe paper notes: \"... fine-tuning a model multiple times on the same dataset, varying only the random seed, leads to a large standard deviation of the fine-tuning accuracy (Devlin et al., 2019;Dodge et al., 2020)\" . They propose a baseline to deal with the spurious pattern.\n\nFrom the paper:\"\n- Use small learning rates combined with bias correction to avoid vanishing gradients early intraining.\n- Increase the number of iterations considerably and train to (almost) zero training loss whilemaking use of early stopping.\"\n\nIt would be interesting to know who does what to deal with the scores variability. I personally couldn't find any remedy to this.",
    "880329": "hello, @isakev , This problem also bothers me. My keras based model cannot reproduce the results using same random seed. My pytorch based model can reproduce well on kaggle TPU, but [Shonenkov's notebook](https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta) cannot reproduce on Colab TPU， I don' know why. My previous pytorch based model on kaggle TPU changes a lot(from 0.91 to 0.72 at public leaderboard) when I using different random seed.",
    "880535": "With the same (for python, tensorflow, and numpy) random seed my scores changed once by 0.003 (0.9390-0.9420). I did not test it again, though. Explanation might be that the TPU brings randomness when splitting batches across its 8 units. I have no idea how to deal with this problem except for treating  the scores with a grain of doubt, or maybe manually managing distribution of batches at a low level programmatically.",
    "881266": "That is something thats been baffling me in this competition. And it makes experimenting really hard. i mostly worked with image data before this competition and I'm not used to this level of instability and scatter with CNNs. Honestly, I have no idea why does this happen.\n\nThe best I could come up with was first training with MLM, and after that using a different learning rate for the head and the transformer layers. While the scatter is not completely gone this way, but it is somewhat smaller.\nI posted these  here: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/156622"
  },
  "source": "meta"
}