{
  "id": 126371,
  "title": "Possible to work with Albert ?",
  "url": "/competitions/tensorflow2-question-answering/discussion/126371",
  "author_name": "Yih-Dar SHIEH",
  "post_date": "2020-01-17T04:56:04.937000",
  "votes": 8,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Very unlikely I am going to try Albert, but I wonder if anyone is able to work with it for this competition.</p>\n\n<p>Albert's vocabulary file looks like a binary file (at least for Hugging Face's model). For Google Albert repo., there is a .txt file and a binaray file, but in that .txt file, there is no <code>reserved</code> tokens, unlike in bert's vocabulary file, we have something <code>[unusedxxx]</code>.</p>\n\n<p>Even not for this competition, it would be good to know how to work with Albert with additional tokens and not to pre-train the LM model from scratch.</p>",
  "messages": [
    {
      "id": 721134,
      "postDate": "2020-01-17T04:56:04.937Z",
      "content": "<p>Very unlikely I am going to try Albert, but I wonder if anyone is able to work with it for this competition.</p>\n\n<p>Albert's vocabulary file looks like a binary file (at least for Hugging Face's model). For Google Albert repo., there is a .txt file and a binaray file, but in that .txt file, there is no <code>reserved</code> tokens, unlike in bert's vocabulary file, we have something <code>[unusedxxx]</code>.</p>\n\n<p>Even not for this competition, it would be good to know how to work with Albert with additional tokens and not to pre-train the LM model from scratch.</p>",
      "rawMarkdown": "Very unlikely I am going to try Albert, but I wonder if anyone is able to work with it for this competition.\n\nAlbert's vocabulary file looks like a binary file (at least for Hugging Face's model). For Google Albert repo., there is a .txt file and a binaray file, but in that .txt file, there is no `reserved` tokens, unlike in bert's vocabulary file, we have something `[unusedxxx]`.\n\nEven not for this competition, it would be good to know how to work with Albert with additional tokens and not to pre-train the LM model from scratch.",
      "votes": 8
    },
    {
      "id": 722814,
      "postDate": "2020-01-19T06:37:05.903Z",
      "content": "<p>A few people have already mentioned that size is an issue for ALBERT xxl (and ALBERT xl too really), but here is a little more educational info as to why. ALBERT stands for \"A Lite BERT\", in the sense that it is supposed to be a lighter version of BERT. As you may have noticed though, this is somewhat of a misnomer. It is faster to train and do backpropagation because all the layers share the same weights, but that unfortunately does not make inference faster. Overall each layer has ~200 million parameters, and with 12 layers you end up with almost 3 billion total parameters. That puts ALBERT xxl at around 10 times bigger than BERT Large, and running it in a Kaggle notebook over just the public dataset would take around 1.5-2 hours using the same sequence length and doc stride. Even ALBERT xl would probably take around an hour.</p>",
      "rawMarkdown": "A few people have already mentioned that size is an issue for ALBERT xxl (and ALBERT xl too really), but here is a little more educational info as to why. ALBERT stands for \"A Lite BERT\", in the sense that it is supposed to be a lighter version of BERT. As you may have noticed though, this is somewhat of a misnomer. It is faster to train and do backpropagation because all the layers share the same weights, but that unfortunately does not make inference faster. Overall each layer has ~200 million parameters, and with 12 layers you end up with almost 3 billion total parameters. That puts ALBERT xxl at around 10 times bigger than BERT Large, and running it in a Kaggle notebook over just the public dataset would take around 1.5-2 hours using the same sequence length and doc stride. Even ALBERT xl would probably take around an hour.",
      "votes": 5
    },
    {
      "id": 721819,
      "postDate": "2020-01-17T18:23:04.630Z",
      "content": "<p>My 0.67 score was with ALBERT xxlarge (predicted offline), straight out without any tuning or modification to the postprocessing. This gave an Timeout Error though.. I would conclude that it's not really possible to get it under the time limit, unless if you do xlarge with reduced seq_length and increased doc_stride, in which case the improvement over the BERT baseline is marginal. So personally I would advise against wasting time on ALBERT.</p>\n\n<p>I haven't worked with Hugging Face and I think the binary file might be the sentencepiece model file, which should provide you with some method to extract the vocabs. There are at least two ways to handle additional tokens: I just manually mapped additional tokens to indices nearer to the end of the vocab file (overriding some infrequent tokens) and it worked fine; in another thread I saw someone manually appended to the end of embedding matrix, which should also work.</p>",
      "rawMarkdown": "My 0.67 score was with ALBERT xxlarge (predicted offline), straight out without any tuning or modification to the postprocessing. This gave an Timeout Error though.. I would conclude that it's not really possible to get it under the time limit, unless if you do xlarge with reduced seq_length and increased doc_stride, in which case the improvement over the BERT baseline is marginal. So personally I would advise against wasting time on ALBERT.\n\nI haven't worked with Hugging Face and I think the binary file might be the sentencepiece model file, which should provide you with some method to extract the vocabs. There are at least two ways to handle additional tokens: I just manually mapped additional tokens to indices nearer to the end of the vocab file (overriding some infrequent tokens) and it worked fine; in another thread I saw someone manually appended to the end of embedding matrix, which should also work.",
      "votes": 4,
      "replies": [
        {
          "id": 722025,
          "postDate": "2020-01-18T02:30:43.930Z",
          "content": "<p><a href=\"/siriuself\">@siriuself</a> </p>\n\n<p>After manually mapping additional tokens to indices nearer to the end of the vocab file, do you need to update the binary file (sentencepiece model file)? I haven't worked with sentencepiece model file</p>",
          "rawMarkdown": "@siriuself \n\nAfter manually mapping additional tokens to indices nearer to the end of the vocab file, do you need to update the binary file (sentencepiece model file)? I haven't worked with sentencepiece model file"
        },
        {
          "id": 722035,
          "postDate": "2020-01-18T02:57:06.357Z",
          "content": "<p>No. This mapping process is separate from tokenization. As a result both \"[special]\" and some other token (e.g. \"_alabama\") will share the same index, which wouldn't affect the model much since the latter token is very rare anyway. By training, the special token will take the place of the original token (in the embedding matrix but not in the sentencepiece model).</p>",
          "rawMarkdown": "No. This mapping process is separate from tokenization. As a result both \"[special]\" and some other token (e.g. \"_alabama\") will share the same index, which wouldn't affect the model much since the latter token is very rare anyway. By training, the special token will take the place of the original token (in the embedding matrix but not in the sentencepiece model)."
        },
        {
          "id": 722151,
          "postDate": "2020-01-18T07:12:34.203Z",
          "content": "<p>Maybe it is possible to use model distillation. At first, we can fine-tune an ALBERT-xxlarge on our dataset, and then use a smaller model as the student model to learn the distribution of the teacher model. I have found a sample code on huggingface about distilling model fine-tuned in SQuaD.<a href=\"https://github.com/huggingface/transformers/blob/master/examples/distillation/run_squad_w_distillation.py\">run_squad_w_distillation.py</a></p>",
          "rawMarkdown": "Maybe it is possible to use model distillation. At first, we can fine-tune an ALBERT-xxlarge on our dataset, and then use a smaller model as the student model to learn the distribution of the teacher model. I have found a sample code on huggingface about distilling model fine-tuned in SQuaD.[run_squad_w_distillation.py](https://github.com/huggingface/transformers/blob/master/examples/distillation/run_squad_w_distillation.py)",
          "votes": 1
        },
        {
          "id": 722194,
          "postDate": "2020-01-18T08:55:41.453Z",
          "content": "<p>Good idea! I will try this for next nlp competition 🙂</p>",
          "rawMarkdown": "Good idea! I will try this for next nlp competition 🙂"
        },
        {
          "id": 722307,
          "postDate": "2020-01-18T11:30:25.940Z",
          "content": "<p><a href=\"/siriuself\">@siriuself</a> Hi, are you using V1 or V2 version?</p>",
          "rawMarkdown": "@siriuself Hi, are you using V1 or V2 version?"
        },
        {
          "id": 722412,
          "postDate": "2020-01-18T14:52:35.280Z",
          "content": "<p><a href=\"/mcggood\">@mcggood</a> i used the V2 version, it's on average 2 pt better at least on their github page</p>",
          "rawMarkdown": "@mcggood i used the V2 version, it's on average 2 pt better at least on their github page"
        },
        {
          "id": 722426,
          "postDate": "2020-01-18T14:58:15.647Z",
          "content": "<p><a href=\"/siriuself\">@siriuself</a> thanks</p>",
          "rawMarkdown": "@siriuself thanks"
        },
        {
          "id": 722428,
          "postDate": "2020-01-18T14:59:25.647Z",
          "content": "<p><a href=\"/guozhiyu0914\">@guozhiyu0914</a> Distillation would be interesting for research's sake. However for competition I personally think it's not worth it after playing with it myself. When I simply replace native bert-large with whole-word-masking bert-large, it's closer to what ALBERT-xxlarge can achieve (like from 4 pt to ~1-2 pt difference). ALBERT is not THAT good and any distillation would likely erase such little gain</p>",
          "rawMarkdown": "@guozhiyu0914 Distillation would be interesting for research's sake. However for competition I personally think it's not worth it after playing with it myself. When I simply replace native bert-large with whole-word-masking bert-large, it's closer to what ALBERT-xxlarge can achieve (like from 4 pt to ~1-2 pt difference). ALBERT is not THAT good and any distillation would likely erase such little gain"
        },
        {
          "id": 723322,
          "postDate": "2020-01-19T21:26:13.863Z",
          "content": "<p>You predicted offline. Could you explain how you did that? You know the testing set is not visible.</p>",
          "rawMarkdown": "You predicted offline. Could you explain how you did that? You know the testing set is not visible."
        },
        {
          "id": 723336,
          "postDate": "2020-01-19T22:12:32.957Z",
          "content": "<p>The test set is available. We don't have the annotations. So he predicted everything offline and submitted the file with the predictions. \nAdvantages:\n-&gt; No GPU time needed to run\nDisadvantages:\n-&gt; No results for the private dataset</p>",
          "rawMarkdown": "The test set is available. We don't have the annotations. So he predicted everything offline and submitted the file with the predictions. \nAdvantages:\n-&gt; No GPU time needed to run\nDisadvantages:\n-&gt; No results for the private dataset",
          "votes": 1
        },
        {
          "id": 723425,
          "postDate": "2020-01-20T02:42:33.737Z",
          "content": "<p><a href=\"/xiaokangwang\">@xiaokangwang</a> Pedro is right. I mostly just predicted offline with public test set and make sure I'm in the right direction. This will give quick results but it's useless for private set</p>",
          "rawMarkdown": "@xiaokangwang Pedro is right. I mostly just predicted offline with public test set and make sure I'm in the right direction. This will give quick results but it's useless for private set"
        },
        {
          "id": 723455,
          "postDate": "2020-01-20T04:11:27.723Z",
          "content": "<p>The problem is SentencePiece is very difficult for Squad like models. The LCS approach they are using is not at all feasible for Natural Questions, as the article is very long. Am i right <a href=\"/siriuself\">@siriuself</a> ? </p>",
          "rawMarkdown": "The problem is SentencePiece is very difficult for Squad like models. The LCS approach they are using is not at all feasible for Natural Questions, as the article is very long. Am i right @siriuself ? "
        },
        {
          "id": 723874,
          "postDate": "2020-01-20T15:00:02.047Z",
          "content": "<p><a href=\"/s4sarath\">@s4sarath</a> I'm not sure I've got you. I did the same doc_strid' ing as when using BERT (so like just converting bert-joint to albert-joint). SentencePiece was pre-trained and I can't see why it's not feasible for NQ</p>",
          "rawMarkdown": "@s4sarath I'm not sure I've got you. I did the same doc_strid' ing as when using BERT (so like just converting bert-joint to albert-joint). SentencePiece was pre-trained and I can't see why it's not feasible for NQ"
        },
        {
          "id": 723902,
          "postDate": "2020-01-20T15:30:59.100Z",
          "content": "<p>You can compare the code run_squad.py in BERT and ALBERT, the function <code>def convert_examples_to_features</code>  , you can find they are slightly different, it's better to add LCS  approach for ALBERT.</p>",
          "rawMarkdown": "You can compare the code run_squad.py in BERT and ALBERT, the function `def convert_examples_to_features`  , you can find they are slightly different, it's better to add LCS  approach for ALBERT."
        },
        {
          "id": 723931,
          "postDate": "2020-01-20T15:59:47.887Z",
          "content": "<p>What I mean by feasibility is, for NQ as entire Wikipedia article is so long, it approximately takes 3 to 4 seconds to perform LCS . So I itl thought 4 seconds is not feasible, as we try to create features online for test data.</p>",
          "rawMarkdown": "What I mean by feasibility is, for NQ as entire Wikipedia article is so long, it approximately takes 3 to 4 seconds to perform LCS . So I itl thought 4 seconds is not feasible, as we try to create features online for test data."
        },
        {
          "id": 724008,
          "postDate": "2020-01-20T17:47:12.113Z",
          "content": "<p><a href=\"/guozhiyu0914\">@guozhiyu0914</a> I see. Haven't looked at the LCS approach yet</p>",
          "rawMarkdown": "@guozhiyu0914 I see. Haven't looked at the LCS approach yet"
        }
      ]
    },
    {
      "id": 721361,
      "postDate": "2020-01-17T09:57:01.867Z",
      "content": "<p>When we use Hugging Face's transformers, we can add tokens by just calling AutoTokenizer.add_tokens, so the vocamulary file won't be problem.\nHowever, to obtain better score we should employ Albert-xxlarge, but it is too heavy. I think using Albert-xxlarge is almost impossible at least in the same way as bert-joint-baseline.</p>",
      "rawMarkdown": "When we use Hugging Face's transformers, we can add tokens by just calling AutoTokenizer.add_tokens, so the vocamulary file won't be problem.\nHowever, to obtain better score we should employ Albert-xxlarge, but it is too heavy. I think using Albert-xxlarge is almost impossible at least in the same way as bert-joint-baseline.",
      "votes": 1
    },
    {
      "id": 721483,
      "postDate": "2020-01-17T12:25:29.477Z",
      "content": "<p>Time limitation also should be considered for Albert</p>",
      "rawMarkdown": "Time limitation also should be considered for Albert"
    },
    {
      "id": 721180,
      "postDate": "2020-01-17T06:31:30.613Z",
      "content": "<p>Maintain a separate dictionary for extra tokens, with index of new vocab starting from , len(original_vocab) . </p>",
      "rawMarkdown": "Maintain a separate dictionary for extra tokens, with index of new vocab starting from , len(original_vocab) . ",
      "replies": [
        {
          "id": 721285,
          "postDate": "2020-01-17T08:57:25.820Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 721381,
          "postDate": "2020-01-17T10:27:01.407Z",
          "content": "<p>By default ALBERT tokenizer deals an unknowrd as UNK. So, shouldn't be a problem </p>",
          "rawMarkdown": "By default ALBERT tokenizer deals an unknowrd as UNK. So, shouldn't be a problem "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 722814,
      "author_name": "Ejmejm",
      "author_url": "",
      "post_date": "2020-01-19T06:37:05.903000",
      "content": "<p>A few people have already mentioned that size is an issue for ALBERT xxl (and ALBERT xl too really), but here is a little more educational info as to why. ALBERT stands for \"A Lite BERT\", in the sense that it is supposed to be a lighter version of BERT. As you may have noticed though, this is somewhat of a misnomer. It is faster to train and do backpropagation because all the layers share the same weights, but that unfortunately does not make inference faster. Overall each layer has ~200 million parameters, and with 12 layers you end up with almost 3 billion total parameters. That puts ALBERT xxl at around 10 times bigger than BERT Large, and running it in a Kaggle notebook over just the public dataset would take around 1.5-2 hours using the same sequence length and doc stride. Even ALBERT xl would probably take around an hour.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 721819,
      "author_name": "Zhengkai Tu",
      "author_url": "",
      "post_date": "2020-01-17T18:23:04.630000",
      "content": "<p>My 0.67 score was with ALBERT xxlarge (predicted offline), straight out without any tuning or modification to the postprocessing. This gave an Timeout Error though.. I would conclude that it's not really possible to get it under the time limit, unless if you do xlarge with reduced seq_length and increased doc_stride, in which case the improvement over the BERT baseline is marginal. So personally I would advise against wasting time on ALBERT.</p>\n\n<p>I haven't worked with Hugging Face and I think the binary file might be the sentencepiece model file, which should provide you with some method to extract the vocabs. There are at least two ways to handle additional tokens: I just manually mapped additional tokens to indices nearer to the end of the vocab file (overriding some infrequent tokens) and it worked fine; in another thread I saw someone manually appended to the end of embedding matrix, which should also work.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 722025,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-01-18T02:30:43.930000",
          "content": "<p><a href=\"/siriuself\">@siriuself</a> </p>\n\n<p>After manually mapping additional tokens to indices nearer to the end of the vocab file, do you need to update the binary file (sentencepiece model file)? I haven't worked with sentencepiece model file</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 722035,
          "author_name": "Zhengkai Tu",
          "author_url": "",
          "post_date": "2020-01-18T02:57:06.357000",
          "content": "<p>No. This mapping process is separate from tokenization. As a result both \"[special]\" and some other token (e.g. \"_alabama\") will share the same index, which wouldn't affect the model much since the latter token is very rare anyway. By training, the special token will take the place of the original token (in the embedding matrix but not in the sentencepiece model).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 722151,
          "author_name": "Zhiyu Guo",
          "author_url": "",
          "post_date": "2020-01-18T07:12:34.203000",
          "content": "<p>Maybe it is possible to use model distillation. At first, we can fine-tune an ALBERT-xxlarge on our dataset, and then use a smaller model as the student model to learn the distribution of the teacher model. I have found a sample code on huggingface about distilling model fine-tuned in SQuaD.<a href=\"https://github.com/huggingface/transformers/blob/master/examples/distillation/run_squad_w_distillation.py\">run_squad_w_distillation.py</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 722194,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-01-18T08:55:41.453000",
          "content": "<p>Good idea! I will try this for next nlp competition 🙂</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 722307,
          "author_name": "MaChaogong",
          "author_url": "",
          "post_date": "2020-01-18T11:30:25.940000",
          "content": "<p><a href=\"/siriuself\">@siriuself</a> Hi, are you using V1 or V2 version?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 722412,
          "author_name": "Zhengkai Tu",
          "author_url": "",
          "post_date": "2020-01-18T14:52:35.280000",
          "content": "<p><a href=\"/mcggood\">@mcggood</a> i used the V2 version, it's on average 2 pt better at least on their github page</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 722426,
          "author_name": "MaChaogong",
          "author_url": "",
          "post_date": "2020-01-18T14:58:15.647000",
          "content": "<p><a href=\"/siriuself\">@siriuself</a> thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 722428,
          "author_name": "Zhengkai Tu",
          "author_url": "",
          "post_date": "2020-01-18T14:59:25.647000",
          "content": "<p><a href=\"/guozhiyu0914\">@guozhiyu0914</a> Distillation would be interesting for research's sake. However for competition I personally think it's not worth it after playing with it myself. When I simply replace native bert-large with whole-word-masking bert-large, it's closer to what ALBERT-xxlarge can achieve (like from 4 pt to ~1-2 pt difference). ALBERT is not THAT good and any distillation would likely erase such little gain</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 723322,
          "author_name": "XiaokangWang",
          "author_url": "",
          "post_date": "2020-01-19T21:26:13.863000",
          "content": "<p>You predicted offline. Could you explain how you did that? You know the testing set is not visible.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 723336,
          "author_name": "Pedro Azevedo",
          "author_url": "",
          "post_date": "2020-01-19T22:12:32.957000",
          "content": "<p>The test set is available. We don't have the annotations. So he predicted everything offline and submitted the file with the predictions. \nAdvantages:\n-&gt; No GPU time needed to run\nDisadvantages:\n-&gt; No results for the private dataset</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 723425,
          "author_name": "Zhengkai Tu",
          "author_url": "",
          "post_date": "2020-01-20T02:42:33.737000",
          "content": "<p><a href=\"/xiaokangwang\">@xiaokangwang</a> Pedro is right. I mostly just predicted offline with public test set and make sure I'm in the right direction. This will give quick results but it's useless for private set</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 723455,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2020-01-20T04:11:27.723000",
          "content": "<p>The problem is SentencePiece is very difficult for Squad like models. The LCS approach they are using is not at all feasible for Natural Questions, as the article is very long. Am i right <a href=\"/siriuself\">@siriuself</a> ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 723874,
          "author_name": "Zhengkai Tu",
          "author_url": "",
          "post_date": "2020-01-20T15:00:02.047000",
          "content": "<p><a href=\"/s4sarath\">@s4sarath</a> I'm not sure I've got you. I did the same doc_strid' ing as when using BERT (so like just converting bert-joint to albert-joint). SentencePiece was pre-trained and I can't see why it's not feasible for NQ</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 723902,
          "author_name": "Zhiyu Guo",
          "author_url": "",
          "post_date": "2020-01-20T15:30:59.100000",
          "content": "<p>You can compare the code run_squad.py in BERT and ALBERT, the function <code>def convert_examples_to_features</code>  , you can find they are slightly different, it's better to add LCS  approach for ALBERT.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 723931,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2020-01-20T15:59:47.887000",
          "content": "<p>What I mean by feasibility is, for NQ as entire Wikipedia article is so long, it approximately takes 3 to 4 seconds to perform LCS . So I itl thought 4 seconds is not feasible, as we try to create features online for test data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 724008,
          "author_name": "Zhengkai Tu",
          "author_url": "",
          "post_date": "2020-01-20T17:47:12.113000",
          "content": "<p><a href=\"/guozhiyu0914\">@guozhiyu0914</a> I see. Haven't looked at the LCS approach yet</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 721361,
      "author_name": "yufuin",
      "author_url": "",
      "post_date": "2020-01-17T09:57:01.867000",
      "content": "<p>When we use Hugging Face's transformers, we can add tokens by just calling AutoTokenizer.add_tokens, so the vocamulary file won't be problem.\nHowever, to obtain better score we should employ Albert-xxlarge, but it is too heavy. I think using Albert-xxlarge is almost impossible at least in the same way as bert-joint-baseline.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 721483,
      "author_name": "Luke",
      "author_url": "",
      "post_date": "2020-01-17T12:25:29.477000",
      "content": "<p>Time limitation also should be considered for Albert</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 721180,
      "author_name": "aintnosunshine",
      "author_url": "",
      "post_date": "2020-01-17T06:31:30.613000",
      "content": "<p>Maintain a separate dictionary for extra tokens, with index of new vocab starting from , len(original_vocab) . </p>",
      "votes": 0,
      "replies": [
        {
          "id": 721285,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-17T08:57:25.820000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 721381,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2020-01-17T10:27:01.407000",
          "content": "<p>By default ALBERT tokenizer deals an unknowrd as UNK. So, shouldn't be a problem </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "721134": "Very unlikely I am going to try Albert, but I wonder if anyone is able to work with it for this competition.\n\nAlbert's vocabulary file looks like a binary file (at least for Hugging Face's model). For Google Albert repo., there is a .txt file and a binaray file, but in that .txt file, there is no `reserved` tokens, unlike in bert's vocabulary file, we have something `[unusedxxx]`.\n\nEven not for this competition, it would be good to know how to work with Albert with additional tokens and not to pre-train the LM model from scratch.",
    "722814": "A few people have already mentioned that size is an issue for ALBERT xxl (and ALBERT xl too really), but here is a little more educational info as to why. ALBERT stands for \"A Lite BERT\", in the sense that it is supposed to be a lighter version of BERT. As you may have noticed though, this is somewhat of a misnomer. It is faster to train and do backpropagation because all the layers share the same weights, but that unfortunately does not make inference faster. Overall each layer has ~200 million parameters, and with 12 layers you end up with almost 3 billion total parameters. That puts ALBERT xxl at around 10 times bigger than BERT Large, and running it in a Kaggle notebook over just the public dataset would take around 1.5-2 hours using the same sequence length and doc stride. Even ALBERT xl would probably take around an hour.",
    "721819": "My 0.67 score was with ALBERT xxlarge (predicted offline), straight out without any tuning or modification to the postprocessing. This gave an Timeout Error though.. I would conclude that it's not really possible to get it under the time limit, unless if you do xlarge with reduced seq_length and increased doc_stride, in which case the improvement over the BERT baseline is marginal. So personally I would advise against wasting time on ALBERT.\n\nI haven't worked with Hugging Face and I think the binary file might be the sentencepiece model file, which should provide you with some method to extract the vocabs. There are at least two ways to handle additional tokens: I just manually mapped additional tokens to indices nearer to the end of the vocab file (overriding some infrequent tokens) and it worked fine; in another thread I saw someone manually appended to the end of embedding matrix, which should also work.",
    "721361": "When we use Hugging Face's transformers, we can add tokens by just calling AutoTokenizer.add_tokens, so the vocamulary file won't be problem.\nHowever, to obtain better score we should employ Albert-xxlarge, but it is too heavy. I think using Albert-xxlarge is almost impossible at least in the same way as bert-joint-baseline.",
    "721483": "Time limitation also should be considered for Albert",
    "721180": "Maintain a separate dictionary for extra tokens, with index of new vocab starting from , len(original_vocab) . "
  }
}