{
  "id": 143967,
  "title": "Where does the '192' maxlen number come from?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/143967",
  "author_name": "",
  "post_date": "2020-04-17T01:38:11.417160400Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi, I'm following the shared awesome notebooks.</p>\n\n<p>I just wonder almost notebooks use 192 as maximum sequence length.\nI guess this is a hyperparameter, but how does this number come?</p>",
  "messages": [
    {
      "id": "810457",
      "postDate": "04/17/2020 01:38:11",
      "content": "<p>Hi, I'm following the shared awesome notebooks.</p>\n\n<p>I just wonder almost notebooks use 192 as maximum sequence length.\nI guess this is a hyperparameter, but how does this number come?</p>",
      "rawMarkdown": "Hi, I'm following the shared awesome notebooks.\n\nI just wonder almost notebooks use 192 as maximum sequence length.\nI guess this is a hyperparameter, but how does this number come?",
      "votes": null
    },
    {
      "id": "811270",
      "postDate": "04/17/2020 18:10:38",
      "content": "<p>I think it came from here:\n<a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a></p>\n\n<p>originally probably just the author chose this number.</p>",
      "rawMarkdown": "I think it came from here:\nhttps://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\n\noriginally probably just the author chose this number.",
      "votes": null
    },
    {
      "id": "811754",
      "postDate": "04/18/2020 07:46:45",
      "content": "<p>And, besides, is is multiple of 64 --- 64*2=128, 64*3 = 192, and the maximum sequence length (number of words (ids) + special tokens) used in BERT is 64*8 = 512. Higher sequence length, they say, slows down training.</p>",
      "rawMarkdown": "And, besides, is is multiple of 64 --- 64\\*2=128, 64\\*3 = 192, and the maximum sequence length (number of words (ids) + special tokens) used in BERT is 64\\*8 = 512. Higher sequence length, they say, slows down training.",
      "votes": null
    },
    {
      "id": "812062",
      "postDate": "04/18/2020 12:56:28",
      "content": "<p>Wow, that’s a good point. Thx for the hint.</p>",
      "rawMarkdown": "Wow, that’s a good point. Thx for the hint.",
      "votes": null
    },
    {
      "id": "838968",
      "postDate": "05/09/2020 00:49:46",
      "content": "<p>If you look at the distribution of comment lengths, 192 also happens to be a good cut-off. There are relatively few comments longer than this. Using a length of 256 hasn't really been worth it for me, no significant performance improvement, and much slower training.</p>",
      "rawMarkdown": "If you look at the distribution of comment lengths, 192 also happens to be a good cut-off. There are relatively few comments longer than this. Using a length of 256 hasn't really been worth it for me, no significant performance improvement, and much slower training.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 811270,
      "author_name": "luohongchen1993",
      "author_url": "",
      "post_date": "04/17/2020 18:10:38",
      "content": "<p>I think it came from here:\n<a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a></p>\n\n<p>originally probably just the author chose this number.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 811754,
      "author_name": "isakev",
      "author_url": "",
      "post_date": "04/18/2020 07:46:45",
      "content": "<p>And, besides, is is multiple of 64 --- 64*2=128, 64*3 = 192, and the maximum sequence length (number of words (ids) + special tokens) used in BERT is 64*8 = 512. Higher sequence length, they say, slows down training.</p>",
      "votes": null,
      "replies": [
        {
          "id": 812062,
          "author_name": "luohongchen1993",
          "author_url": "",
          "post_date": "04/18/2020 12:56:28",
          "content": "<p>Wow, that’s a good point. Thx for the hint.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 838968,
      "author_name": "nikitabu",
      "author_url": "",
      "post_date": "05/09/2020 00:49:46",
      "content": "<p>If you look at the distribution of comment lengths, 192 also happens to be a good cut-off. There are relatively few comments longer than this. Using a length of 256 hasn't really been worth it for me, no significant performance improvement, and much slower training.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "810457": "Hi, I'm following the shared awesome notebooks.\n\nI just wonder almost notebooks use 192 as maximum sequence length.\nI guess this is a hyperparameter, but how does this number come?",
    "811270": "I think it came from here:\nhttps://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\n\noriginally probably just the author chose this number.",
    "811754": "And, besides, is is multiple of 64 --- 64\\*2=128, 64\\*3 = 192, and the maximum sequence length (number of words (ids) + special tokens) used in BERT is 64\\*8 = 512. Higher sequence length, they say, slows down training.",
    "812062": "Wow, that’s a good point. Thx for the hint.",
    "838968": "If you look at the distribution of comment lengths, 192 also happens to be a good cut-off. There are relatively few comments longer than this. Using a length of 256 hasn't really been worth it for me, no significant performance improvement, and much slower training."
  },
  "source": "meta"
}