{
  "id": 123977,
  "title": "Clarification on Kernel time limit",
  "url": "/competitions/tensorflow2-question-answering/discussion/123977",
  "author_name": "",
  "post_date": "2020-01-01T03:27:11.149619900Z",
  "votes": 2,
  "comment_count": 35,
  "views": 0,
  "content": "<p>Can someone confirm whether the time limit of 3hr is for <strong>committing</strong> or for <strong>submission</strong>? Since the public test set is only 12% of the total, I would assume a 1-to-9 ratio for the timing (which means any kernel should take &lt;20 mins to commit for safe submission later). I'm new to Kaggle and this is not completely explicit in the Code Requirement, which states the 3hr limit is \"In order for the 'Submit to Competition' button to be active after a commit\".</p>\n\n<p>For anyone else trying <strong>ALBERT</strong>, I would advise to go with <strong>xlarge</strong> rather than <strong>xxlarge</strong>. The latter took an hour on the public set and gave me a \"Notebook Timeout\" Error after I submitted the committed notebook.</p>",
  "messages": [
    {
      "id": "707545",
      "postDate": "01/01/2020 03:27:11",
      "content": "<p>Can someone confirm whether the time limit of 3hr is for <strong>committing</strong> or for <strong>submission</strong>? Since the public test set is only 12% of the total, I would assume a 1-to-9 ratio for the timing (which means any kernel should take &lt;20 mins to commit for safe submission later). I'm new to Kaggle and this is not completely explicit in the Code Requirement, which states the 3hr limit is \"In order for the 'Submit to Competition' button to be active after a commit\".</p>\n\n<p>For anyone else trying <strong>ALBERT</strong>, I would advise to go with <strong>xlarge</strong> rather than <strong>xxlarge</strong>. The latter took an hour on the public set and gave me a \"Notebook Timeout\" Error after I submitted the committed notebook.</p>",
      "rawMarkdown": "Can someone confirm whether the time limit of 3hr is for **committing** or for **submission**? Since the public test set is only 12% of the total, I would assume a 1-to-9 ratio for the timing (which means any kernel should take &lt;20 mins to commit for safe submission later). I'm new to Kaggle and this is not completely explicit in the Code Requirement, which states the 3hr limit is \"In order for the 'Submit to Competition' button to be active after a commit\".\n\nFor anyone else trying **ALBERT**, I would advise to go with **xlarge** rather than **xxlarge**. The latter took an hour on the public set and gave me a \"Notebook Timeout\" Error after I submitted the committed notebook.",
      "votes": null
    },
    {
      "id": "707576",
      "postDate": "01/01/2020 05:01:04",
      "content": "<p><a href=\"/siriuself\">@siriuself</a> my best solution is taking almost 3 hours for inference on private test set. It takes 19 mins to complete the run on public test set. Anything above 3 hours on private test set will give you a Notebook Timeout error.</p>",
      "rawMarkdown": "siriuself my best solution is taking almost 3 hours for inference on private test set. It takes 19 mins to complete the run on public test set. Anything above 3 hours on private test set will give you a Notebook Timeout error.",
      "votes": null
    },
    {
      "id": "707585",
      "postDate": "01/01/2020 05:29:50",
      "content": "<p><a href=\"/axel81\">@axel81</a> Thank you for the clarification. I would forgo ALBERT-xxlarge then..</p>",
      "rawMarkdown": "axel81 Thank you for the clarification. I would forgo ALBERT-xxlarge then..",
      "votes": null
    },
    {
      "id": "707589",
      "postDate": "01/01/2020 05:37:30",
      "content": "<p>My best solution is also taking almost 3 hours for private test set. 20 min for public test set.\nI wonder how we can do ensemble :)</p>",
      "rawMarkdown": "My best solution is also taking almost 3 hours for private test set. 20 min for public test set.\nI wonder how we can do ensemble :)",
      "votes": null
    },
    {
      "id": "707593",
      "postDate": "01/01/2020 05:51:43",
      "content": "<p>Yeah right xD, no scope for ensemble if we can't optimize.</p>",
      "rawMarkdown": "Yeah right xD, no scope for ensemble if we can't optimize.",
      "votes": null
    },
    {
      "id": "708818",
      "postDate": "01/02/2020 18:02:16",
      "content": "<p>The compute time constraint is applied to both the commit (on the public test set) and the submission (on the private test set). Since the private test set is ~7 times larger than the public, you should anticipate this when committing your model to estimate whether it will re-run successfully on the private set.</p>",
      "rawMarkdown": "The compute time constraint is applied to both the commit (on the public test set) and the submission (on the private test set). Since the private test set is ~7 times larger than the public, you should anticipate this when committing your model to estimate whether it will re-run successfully on the private set.",
      "votes": null
    },
    {
      "id": "709045",
      "postDate": "01/03/2020 01:11:25",
      "content": "<p>Thanks for the clarification</p>",
      "rawMarkdown": "Thanks for the clarification",
      "votes": null
    },
    {
      "id": "709071",
      "postDate": "01/03/2020 02:38:54",
      "content": "<p>hi Zhengkai Tu, for albert, do you process the special token the official add for this competiton?</p>",
      "rawMarkdown": "hi Zhengkai Tu, for albert, do you process the special token the official add for this competiton?",
      "votes": null
    },
    {
      "id": "711173",
      "postDate": "01/05/2020 19:17:05",
      "content": "<p>Even ALBERT has shared parameter via layers, the graph is still as same as BERT, so i guess ALBERT xlarge will also raised a Notebook Timeout error</p>",
      "rawMarkdown": "Even ALBERT has shared parameter via layers, the graph is still as same as BERT, so i guess ALBERT xlarge will also raised a Notebook Timeout error",
      "votes": null
    },
    {
      "id": "714110",
      "postDate": "01/09/2020 04:19:27",
      "content": "<p>You're right. I'd have to decrease the max_length and increase doc_stride in compensation even with xlarge to barely fit under the time limit. But then I can only get 0.64.. I'm now redoing the whole thing with a pipelined approach, with First Stage only selecting the candidates with BERT-LARGE, and Second Stage for short answers only with ALBERT xxlarge.</p>",
      "rawMarkdown": "You're right. I'd have to decrease the max_length and increase doc_stride in compensation even with xlarge to barely fit under the time limit. But then I can only get 0.64.. I'm now redoing the whole thing with a pipelined approach, with First Stage only selecting the candidates with BERT-LARGE, and Second Stage for short answers only with ALBERT xxlarge.",
      "votes": null
    },
    {
      "id": "714112",
      "postDate": "01/09/2020 04:21:57",
      "content": "<p>You meant those \"[Context=] [Table=]\" etc.? Yes I manually mapped those to override least frequent vocabs near the end of ALBERT vocab file. I decreases the max context so got around 200 of them. </p>",
      "rawMarkdown": "You meant those \"[Context=] [Table=]\" etc.? Yes I manually mapped those to override least frequent vocabs near the end of ALBERT vocab file. I decreases the max context so got around 200 of them.",
      "votes": null
    },
    {
      "id": "714465",
      "postDate": "01/09/2020 13:15:50",
      "content": "<p>ok so your best model is ALBERT right.. When I tried ALBERT and set others same as BERT did, my CV will decrease 8%...</p>",
      "rawMarkdown": "ok so your best model is ALBERT right.. When I tried ALBERT and set others same as BERT did, my CV will decrease 8%...",
      "votes": null
    },
    {
      "id": "714507",
      "postDate": "01/09/2020 14:03:19",
      "content": "<p>In another kaggle  NLP competition running now, I also found some people said that after using ALBERT, their performance got worse than BERT, but in the squad leaderboard, most top-ranked models are based on ALBERT.</p>",
      "rawMarkdown": "In another kaggle  NLP competition running now, I also found some people said that after using ALBERT, their performance got worse than BERT, but in the squad leaderboard, most top-ranked models are based on ALBERT.",
      "votes": null
    },
    {
      "id": "714567",
      "postDate": "01/09/2020 14:55:36",
      "content": "<p><a href=\"/httpwwwfszyc\">@httpwwwfszyc</a> If you've re-processed the train data the ALBERT way, then it's probably parameter setting and some issues with the use of sentencepiece instead of wordpiece (quite a lot of such when I did it). I had to manually map all the special vocabs into (and override) least frequent vocabs near the end of ALBERT vocab file by visual inspection... since it doen't have [unused] placeholders for custom vocabs like BERT. Also, A learning rate of 3e-5 seems to be the max, the loss doesn't decrease with 5e-5.</p>",
      "rawMarkdown": "httpwwwfszyc If you've re-processed the train data the ALBERT way, then it's probably parameter setting and some issues with the use of sentencepiece instead of wordpiece (quite a lot of such when I did it). I had to manually map all the special vocabs into (and override) least frequent vocabs near the end of ALBERT vocab file by visual inspection... since it doen't have [unused] placeholders for custom vocabs like BERT. Also, A learning rate of 3e-5 seems to be the max, the loss doesn't decrease with 5e-5.",
      "votes": null
    },
    {
      "id": "714570",
      "postDate": "01/09/2020 15:00:13",
      "content": "<p><a href=\"/guozhiyu0914\">@guozhiyu0914</a> It'd be interested to see whether it's really the case or just parameter setting issues. They call ALBERT the new SOTA not because of SQuaD performance but rather it beats BERT in ALL glue tasks if I'm not wrong. That's for ALBERT xxlarge though.</p>",
      "rawMarkdown": "guozhiyu0914 It'd be interested to see whether it's really the case or just parameter setting issues. They call ALBERT the new SOTA not because of SQuaD performance but rather it beats BERT in ALL glue tasks if I'm not wrong. That's for ALBERT xxlarge though.",
      "votes": null
    },
    {
      "id": "714869",
      "postDate": "01/09/2020 21:35:54",
      "content": "<p>my way is to enlarge the vocab size from 30000 to 30209 by adding np.random.normal(0,0.02) and change the embedding layer shape. not sure this will make a difference to yours preprocessing method.</p>",
      "rawMarkdown": "my way is to enlarge the vocab size from 30000 to 30209 by adding np.random.normal(0,0.02) and change the embedding layer shape. not sure this will make a difference to yours preprocessing method.",
      "votes": null
    },
    {
      "id": "714978",
      "postDate": "01/10/2020 01:58:22",
      "content": "<p>I'm new to NLP, and I think this might be a stupid question, but I found that ALBERT has much less parameters compared to BERT (especially xlarge), which I thought, means that it calculates much faster. I'm trying to implement ALBERT because of this. Am I wrong? </p>",
      "rawMarkdown": "I'm new to NLP, and I think this might be a stupid question, but I found that ALBERT has much less parameters compared to BERT (especially xlarge), which I thought, means that it calculates much faster. I'm trying to implement ALBERT because of this. Am I wrong?",
      "votes": null
    },
    {
      "id": "714984",
      "postDate": "01/10/2020 02:11:24",
      "content": "<p><a href=\"/kitkatbar0429\">@kitkatbar0429</a> ALBERT has mush fewer parameters, yes. It calculates much faster, no. ALBERT basically reduce the number of parameters by sharing them across different BERT layers. In this way it makes the performance of BERT more stable with increasing depth and size, and that's why ALBERT can go all the way to xlarge and xxlarge without drastic drop in performance as BERT does. However, there is still almost the same amount of computation (say for BERT-large and ALBERT-large). There's barely any reduction in training/inference time.</p>",
      "rawMarkdown": "kitkatbar0429 ALBERT has mush fewer parameters, yes. It calculates much faster, no. ALBERT basically reduce the number of parameters by sharing them across different BERT layers. In this way it makes the performance of BERT more stable with increasing depth and size, and that's why ALBERT can go all the way to xlarge and xxlarge without drastic drop in performance as BERT does. However, there is still almost the same amount of computation (say for BERT-large and ALBERT-large). There's barely any reduction in training/inference time.",
      "votes": null
    },
    {
      "id": "715017",
      "postDate": "01/10/2020 03:19:06",
      "content": "<p><a href=\"/siriuself\">@siriuself</a> I did not know that, thank you !!</p>",
      "rawMarkdown": "siriuself I did not know that, thank you !!",
      "votes": null
    },
    {
      "id": "715271",
      "postDate": "01/10/2020 10:38:00",
      "content": "<p>Do you mean you just use wordpiece rather than sentencepiece ？I had tried sentencepiece  but worse than bert</p>",
      "rawMarkdown": "Do you mean you just use wordpiece rather than sentencepiece ？I had tried sentencepiece  but worse than bert",
      "votes": null
    },
    {
      "id": "715521",
      "postDate": "01/10/2020 16:02:22",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "715524",
      "postDate": "01/10/2020 16:02:48",
      "content": "<p>After examine the tokenization of ALBERT carefully, I found only punctuations (for example, ',' and 'good,' ) will have difference between word encoding and sentence encoding.</p>",
      "rawMarkdown": "After examine the tokenization of ALBERT carefully, I found only punctuations (for example, ',' and 'good,' ) will have difference between word encoding and sentence encoding.",
      "votes": null
    },
    {
      "id": "715708",
      "postDate": "01/10/2020 18:34:51",
      "content": "<p><a href=\"/httpwwwfszyc\">@httpwwwfszyc</a> I don't get it. Are you using the SentencePiece model that came with ALBERT? If so the tokenization and encoding will be completely different from those of BERT. Or are you suggesting something else?</p>",
      "rawMarkdown": "httpwwwfszyc I don't get it. Are you using the SentencePiece model that came with ALBERT? If so the tokenization and encoding will be completely different from those of BERT. Or are you suggesting something else?",
      "votes": null
    },
    {
      "id": "715752",
      "postDate": "01/10/2020 19:26:22",
      "content": "<p>I just change function <em>def tokenize</em> in bert_utils.py and *def convert_tokens_to_ids* in tokenization.py and compare the result from tokenizing  sentencewise to tokenizing word by word . </p>",
      "rawMarkdown": "I just change function *def tokenize* in bert_utils.py and *def convert_tokens_to_ids* in tokenization.py and compare the result from tokenizing  sentencewise to tokenizing word by word .",
      "votes": null
    },
    {
      "id": "715835",
      "postDate": "01/10/2020 21:54:53",
      "content": "<p><a href=\"/httpwwwfszyc\">@httpwwwfszyc</a> That's expected. For SentencePiece the word boundary is maintained by default, in other words \"_\" will always be at the start (not in the middle) of a token. If so there's no difference in tokenize(sentence) vs [tokenize(word) for word in sentence.split()]</p>\n\n<p>Also, do make sure that whatever passed into the SP tokenizer has been lower()-ed, I think I encountered this disastrous issue along the way.</p>",
      "rawMarkdown": "httpwwwfszyc That's expected. For SentencePiece the word boundary is maintained by default, in other words \"_\" will always be at the start (not in the middle) of a token. If so there's no difference in tokenize(sentence) vs [tokenize(word) for word in sentence.split()]\n\nAlso, do make sure that whatever passed into the SP tokenizer has been lower()-ed, I think I encountered this disastrous issue along the way.",
      "votes": null
    },
    {
      "id": "715992",
      "postDate": "01/11/2020 05:44:43",
      "content": "<p>My current solution takes 10 minutes (625s) to run when I commit (public set). However, when I commit (pvt set), it times out.</p>",
      "rawMarkdown": "My current solution takes 10 minutes (625s) to run when I commit (public set). However, when I commit (pvt set), it times out.",
      "votes": null
    },
    {
      "id": "716006",
      "postDate": "01/11/2020 06:17:37",
      "content": "<p>That seems weird to me. I've got better with at least 2 pts F1 with ALBERT xlarge and xxlarge straight out with bert-joint parameters with nq_eval, without any tuning (batch_size=32, alpha=3e-5, epoch=1) and even when I decreases max_length and increase doc_stride.. This is with sentencepiece and corresponding modifications to the preprocessing code of course. I do suspect it might be due to not explicitly calling lower() before passing into sentencepiece, somewhere. I weakly remember having such issues in early experiments with ALBERT for some other tasks (even if the case check is passed), and it's gonna be disastrous since any tokens with upper case chars would be .</p>",
      "rawMarkdown": "That seems weird to me. I've got better with at least 2 pts F1 with ALBERT xlarge and xxlarge straight out with bert-joint parameters with nq_eval, without any tuning (batch_size=32, alpha=3e-5, epoch=1) and even when I decreases max_length and increase doc_stride.. This is with sentencepiece and corresponding modifications to the preprocessing code of course. I do suspect it might be due to not explicitly calling lower() before passing into sentencepiece, somewhere. I weakly remember having such issues in early experiments with ALBERT for some other tasks (even if the case check is passed), and it's gonna be disastrous since any tokens with upper case chars would be",
      "votes": null
    },
    {
      "id": "716027",
      "postDate": "01/11/2020 06:56:08",
      "content": "<p><a href=\"/kenkrige\">@kenkrige</a> that's weird my solutions takes 828s to commit but it doesn't timeout on private test set.</p>",
      "rawMarkdown": "kenkrige that's weird my solutions takes 828s to commit but it doesn't timeout on private test set.",
      "votes": null
    },
    {
      "id": "716053",
      "postDate": "01/11/2020 07:24:40",
      "content": "<p>Tghanks <a href=\"/axel81\">@axel81</a> thanks for that info, it is very helpful. I will search for another reason in my code. The error message is not very informative, it just says \"Notebook Exceeded Allowed Compute\". I am not sure if that means GPU time or too much RAM used.</p>",
      "rawMarkdown": "Tghanks @axel81 thanks for that info, it is very helpful. I will search for another reason in my code. The error message is not very informative, it just says \"Notebook Exceeded Allowed Compute\". I am not sure if that means GPU time or too much RAM used.",
      "votes": null
    },
    {
      "id": "716083",
      "postDate": "01/11/2020 08:03:46",
      "content": "<p><a href=\"/kenkrige\">@kenkrige</a> if it was a timeout issue you would have gotten this error <code>Notebook Timeout</code>. So in your case it must be related to GPU or RAM memory limit exceeded which caused private run to fail</p>",
      "rawMarkdown": "kenkrige if it was a timeout issue you would have gotten this error `Notebook Timeout`. So in your case it must be related to GPU or RAM memory limit exceeded which caused private run to fail",
      "votes": null
    },
    {
      "id": "716097",
      "postDate": "01/11/2020 08:29:47",
      "content": "<p>OK, thanks. That helps a lot.</p>",
      "rawMarkdown": "OK, thanks. That helps a lot.",
      "votes": null
    },
    {
      "id": "716732",
      "postDate": "01/12/2020 07:08:10",
      "content": "<p><a href=\"/siriuself\">@siriuself</a> I was experimenting with ALBERT x large v1 but it gives very worse results compared to Bert. I am using nq-vocab.txt but the accuracy is very worse. As far as I understand tokenization process for ALBERT and BERT is same right? Even with ALBERT vocab my model gives very bad results my LR is 3e-5</p>",
      "rawMarkdown": "siriuself I was experimenting with ALBERT x large v1 but it gives very worse results compared to Bert. I am using nq-vocab.txt but the accuracy is very worse. As far as I understand tokenization process for ALBERT and BERT is same right? Even with ALBERT vocab my model gives very bad results my LR is 3e-5",
      "votes": null
    },
    {
      "id": "716795",
      "postDate": "01/12/2020 08:57:45",
      "content": "<p>Thanks <a href=\"/axel81\">@axel81</a> You were absolutely right. It was a RAM problem, not timeout. By using module <code>gs</code> and being more careful about garbage collection, it ran perfectly. No garbage collection was necessary in the commit stage on the smaller public set.</p>",
      "rawMarkdown": "Thanks @axel81 You were absolutely right. It was a RAM problem, not timeout. By using module `gs` and being more careful about garbage collection, it ran perfectly. No garbage collection was necessary in the commit stage on the smaller public set.",
      "votes": null
    },
    {
      "id": "716800",
      "postDate": "01/12/2020 09:02:20",
      "content": "<p>Good to hear that :)</p>",
      "rawMarkdown": "Good to hear that :)",
      "votes": null
    },
    {
      "id": "716827",
      "postDate": "01/12/2020 10:23:11",
      "content": "<p><a href=\"/axel81\">@axel81</a> \nIf you use official pretrained model for BERT and ALBERT, the tokenization process is not same.\nBERT uses wordpieces and ALBERT uses sentencepiece.</p>",
      "rawMarkdown": "axel81 \nIf you use official pretrained model for BERT and ALBERT, the tokenization process is not same.\nBERT uses wordpieces and ALBERT uses sentencepiece.",
      "votes": null
    },
    {
      "id": "716829",
      "postDate": "01/12/2020 10:26:59",
      "content": "<p>Oh Okay makes sense now why my ALBERT version is not giving good results</p>",
      "rawMarkdown": "Oh Okay makes sense now why my ALBERT version is not giving good results",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 707576,
      "author_name": "axel81",
      "author_url": "",
      "post_date": "01/01/2020 05:01:04",
      "content": "<p><a href=\"/siriuself\">@siriuself</a> my best solution is taking almost 3 hours for inference on private test set. It takes 19 mins to complete the run on public test set. Anything above 3 hours on private test set will give you a Notebook Timeout error.</p>",
      "votes": null,
      "replies": [
        {
          "id": 707589,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "01/01/2020 05:37:30",
          "content": "<p>My best solution is also taking almost 3 hours for private test set. 20 min for public test set.\nI wonder how we can do ensemble :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707593,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/01/2020 05:51:43",
          "content": "<p>Yeah right xD, no scope for ensemble if we can't optimize.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 707585,
      "author_name": "siriuself",
      "author_url": "",
      "post_date": "01/01/2020 05:29:50",
      "content": "<p><a href=\"/axel81\">@axel81</a> Thank you for the clarification. I would forgo ALBERT-xxlarge then..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 708818,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "01/02/2020 18:02:16",
      "content": "<p>The compute time constraint is applied to both the commit (on the public test set) and the submission (on the private test set). Since the private test set is ~7 times larger than the public, you should anticipate this when committing your model to estimate whether it will re-run successfully on the private set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 709045,
          "author_name": "siriuself",
          "author_url": "",
          "post_date": "01/03/2020 01:11:25",
          "content": "<p>Thanks for the clarification</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715992,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "01/11/2020 05:44:43",
          "content": "<p>My current solution takes 10 minutes (625s) to run when I commit (public set). However, when I commit (pvt set), it times out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716027,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/11/2020 06:56:08",
          "content": "<p><a href=\"/kenkrige\">@kenkrige</a> that's weird my solutions takes 828s to commit but it doesn't timeout on private test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716053,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "01/11/2020 07:24:40",
          "content": "<p>Tghanks <a href=\"/axel81\">@axel81</a> thanks for that info, it is very helpful. I will search for another reason in my code. The error message is not very informative, it just says \"Notebook Exceeded Allowed Compute\". I am not sure if that means GPU time or too much RAM used.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716083,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/11/2020 08:03:46",
          "content": "<p><a href=\"/kenkrige\">@kenkrige</a> if it was a timeout issue you would have gotten this error <code>Notebook Timeout</code>. So in your case it must be related to GPU or RAM memory limit exceeded which caused private run to fail</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716097,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "01/11/2020 08:29:47",
          "content": "<p>OK, thanks. That helps a lot.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716795,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "01/12/2020 08:57:45",
          "content": "<p>Thanks <a href=\"/axel81\">@axel81</a> You were absolutely right. It was a RAM problem, not timeout. By using module <code>gs</code> and being more careful about garbage collection, it ran perfectly. No garbage collection was necessary in the commit stage on the smaller public set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716800,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/12/2020 09:02:20",
          "content": "<p>Good to hear that :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 709071,
      "author_name": "",
      "author_url": "",
      "post_date": "01/03/2020 02:38:54",
      "content": "<p>hi Zhengkai Tu, for albert, do you process the special token the official add for this competiton?</p>",
      "votes": null,
      "replies": [
        {
          "id": 714112,
          "author_name": "siriuself",
          "author_url": "",
          "post_date": "01/09/2020 04:21:57",
          "content": "<p>You meant those \"[Context=] [Table=]\" etc.? Yes I manually mapped those to override least frequent vocabs near the end of ALBERT vocab file. I decreases the max context so got around 200 of them. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 711173,
      "author_name": "httpwwwfszyc",
      "author_url": "",
      "post_date": "01/05/2020 19:17:05",
      "content": "<p>Even ALBERT has shared parameter via layers, the graph is still as same as BERT, so i guess ALBERT xlarge will also raised a Notebook Timeout error</p>",
      "votes": null,
      "replies": [
        {
          "id": 714110,
          "author_name": "siriuself",
          "author_url": "",
          "post_date": "01/09/2020 04:19:27",
          "content": "<p>You're right. I'd have to decrease the max_length and increase doc_stride in compensation even with xlarge to barely fit under the time limit. But then I can only get 0.64.. I'm now redoing the whole thing with a pipelined approach, with First Stage only selecting the candidates with BERT-LARGE, and Second Stage for short answers only with ALBERT xxlarge.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 714465,
          "author_name": "httpwwwfszyc",
          "author_url": "",
          "post_date": "01/09/2020 13:15:50",
          "content": "<p>ok so your best model is ALBERT right.. When I tried ALBERT and set others same as BERT did, my CV will decrease 8%...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 714507,
          "author_name": "guozhiyu0914",
          "author_url": "",
          "post_date": "01/09/2020 14:03:19",
          "content": "<p>In another kaggle  NLP competition running now, I also found some people said that after using ALBERT, their performance got worse than BERT, but in the squad leaderboard, most top-ranked models are based on ALBERT.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 714567,
          "author_name": "siriuself",
          "author_url": "",
          "post_date": "01/09/2020 14:55:36",
          "content": "<p><a href=\"/httpwwwfszyc\">@httpwwwfszyc</a> If you've re-processed the train data the ALBERT way, then it's probably parameter setting and some issues with the use of sentencepiece instead of wordpiece (quite a lot of such when I did it). I had to manually map all the special vocabs into (and override) least frequent vocabs near the end of ALBERT vocab file by visual inspection... since it doen't have [unused] placeholders for custom vocabs like BERT. Also, A learning rate of 3e-5 seems to be the max, the loss doesn't decrease with 5e-5.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 714570,
          "author_name": "siriuself",
          "author_url": "",
          "post_date": "01/09/2020 15:00:13",
          "content": "<p><a href=\"/guozhiyu0914\">@guozhiyu0914</a> It'd be interested to see whether it's really the case or just parameter setting issues. They call ALBERT the new SOTA not because of SQuaD performance but rather it beats BERT in ALL glue tasks if I'm not wrong. That's for ALBERT xxlarge though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 714869,
          "author_name": "httpwwwfszyc",
          "author_url": "",
          "post_date": "01/09/2020 21:35:54",
          "content": "<p>my way is to enlarge the vocab size from 30000 to 30209 by adding np.random.normal(0,0.02) and change the embedding layer shape. not sure this will make a difference to yours preprocessing method.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 714978,
          "author_name": "kitkatbar0429",
          "author_url": "",
          "post_date": "01/10/2020 01:58:22",
          "content": "<p>I'm new to NLP, and I think this might be a stupid question, but I found that ALBERT has much less parameters compared to BERT (especially xlarge), which I thought, means that it calculates much faster. I'm trying to implement ALBERT because of this. Am I wrong? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 714984,
          "author_name": "siriuself",
          "author_url": "",
          "post_date": "01/10/2020 02:11:24",
          "content": "<p><a href=\"/kitkatbar0429\">@kitkatbar0429</a> ALBERT has mush fewer parameters, yes. It calculates much faster, no. ALBERT basically reduce the number of parameters by sharing them across different BERT layers. In this way it makes the performance of BERT more stable with increasing depth and size, and that's why ALBERT can go all the way to xlarge and xxlarge without drastic drop in performance as BERT does. However, there is still almost the same amount of computation (say for BERT-large and ALBERT-large). There's barely any reduction in training/inference time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715017,
          "author_name": "kitkatbar0429",
          "author_url": "",
          "post_date": "01/10/2020 03:19:06",
          "content": "<p><a href=\"/siriuself\">@siriuself</a> I did not know that, thank you !!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715524,
          "author_name": "httpwwwfszyc",
          "author_url": "",
          "post_date": "01/10/2020 16:02:48",
          "content": "<p>After examine the tokenization of ALBERT carefully, I found only punctuations (for example, ',' and 'good,' ) will have difference between word encoding and sentence encoding.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715708,
          "author_name": "siriuself",
          "author_url": "",
          "post_date": "01/10/2020 18:34:51",
          "content": "<p><a href=\"/httpwwwfszyc\">@httpwwwfszyc</a> I don't get it. Are you using the SentencePiece model that came with ALBERT? If so the tokenization and encoding will be completely different from those of BERT. Or are you suggesting something else?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715752,
          "author_name": "httpwwwfszyc",
          "author_url": "",
          "post_date": "01/10/2020 19:26:22",
          "content": "<p>I just change function <em>def tokenize</em> in bert_utils.py and *def convert_tokens_to_ids* in tokenization.py and compare the result from tokenizing  sentencewise to tokenizing word by word . </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715835,
          "author_name": "siriuself",
          "author_url": "",
          "post_date": "01/10/2020 21:54:53",
          "content": "<p><a href=\"/httpwwwfszyc\">@httpwwwfszyc</a> That's expected. For SentencePiece the word boundary is maintained by default, in other words \"_\" will always be at the start (not in the middle) of a token. If so there's no difference in tokenize(sentence) vs [tokenize(word) for word in sentence.split()]</p>\n\n<p>Also, do make sure that whatever passed into the SP tokenizer has been lower()-ed, I think I encountered this disastrous issue along the way.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716732,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/12/2020 07:08:10",
          "content": "<p><a href=\"/siriuself\">@siriuself</a> I was experimenting with ALBERT x large v1 but it gives very worse results compared to Bert. I am using nq-vocab.txt but the accuracy is very worse. As far as I understand tokenization process for ALBERT and BERT is same right? Even with ALBERT vocab my model gives very bad results my LR is 3e-5</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716827,
          "author_name": "kentaronakanishi",
          "author_url": "",
          "post_date": "01/12/2020 10:23:11",
          "content": "<p><a href=\"/axel81\">@axel81</a> \nIf you use official pretrained model for BERT and ALBERT, the tokenization process is not same.\nBERT uses wordpieces and ALBERT uses sentencepiece.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716829,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/12/2020 10:26:59",
          "content": "<p>Oh Okay makes sense now why my ALBERT version is not giving good results</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 715271,
      "author_name": "zhaomeng1126",
      "author_url": "",
      "post_date": "01/10/2020 10:38:00",
      "content": "<p>Do you mean you just use wordpiece rather than sentencepiece ？I had tried sentencepiece  but worse than bert</p>",
      "votes": null,
      "replies": [
        {
          "id": 716006,
          "author_name": "siriuself",
          "author_url": "",
          "post_date": "01/11/2020 06:17:37",
          "content": "<p>That seems weird to me. I've got better with at least 2 pts F1 with ALBERT xlarge and xxlarge straight out with bert-joint parameters with nq_eval, without any tuning (batch_size=32, alpha=3e-5, epoch=1) and even when I decreases max_length and increase doc_stride.. This is with sentencepiece and corresponding modifications to the preprocessing code of course. I do suspect it might be due to not explicitly calling lower() before passing into sentencepiece, somewhere. I weakly remember having such issues in early experiments with ALBERT for some other tasks (even if the case check is passed), and it's gonna be disastrous since any tokens with upper case chars would be .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 715521,
      "author_name": "httpwwwfszyc",
      "author_url": "",
      "post_date": "01/10/2020 16:02:22",
      "content": "",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "707545": "Can someone confirm whether the time limit of 3hr is for **committing** or for **submission**? Since the public test set is only 12% of the total, I would assume a 1-to-9 ratio for the timing (which means any kernel should take &lt;20 mins to commit for safe submission later). I'm new to Kaggle and this is not completely explicit in the Code Requirement, which states the 3hr limit is \"In order for the 'Submit to Competition' button to be active after a commit\".\n\nFor anyone else trying **ALBERT**, I would advise to go with **xlarge** rather than **xxlarge**. The latter took an hour on the public set and gave me a \"Notebook Timeout\" Error after I submitted the committed notebook.",
    "707576": "siriuself my best solution is taking almost 3 hours for inference on private test set. It takes 19 mins to complete the run on public test set. Anything above 3 hours on private test set will give you a Notebook Timeout error.",
    "707585": "axel81 Thank you for the clarification. I would forgo ALBERT-xxlarge then..",
    "707589": "My best solution is also taking almost 3 hours for private test set. 20 min for public test set.\nI wonder how we can do ensemble :)",
    "707593": "Yeah right xD, no scope for ensemble if we can't optimize.",
    "708818": "The compute time constraint is applied to both the commit (on the public test set) and the submission (on the private test set). Since the private test set is ~7 times larger than the public, you should anticipate this when committing your model to estimate whether it will re-run successfully on the private set.",
    "709045": "Thanks for the clarification",
    "709071": "hi Zhengkai Tu, for albert, do you process the special token the official add for this competiton?",
    "711173": "Even ALBERT has shared parameter via layers, the graph is still as same as BERT, so i guess ALBERT xlarge will also raised a Notebook Timeout error",
    "714110": "You're right. I'd have to decrease the max_length and increase doc_stride in compensation even with xlarge to barely fit under the time limit. But then I can only get 0.64.. I'm now redoing the whole thing with a pipelined approach, with First Stage only selecting the candidates with BERT-LARGE, and Second Stage for short answers only with ALBERT xxlarge.",
    "714112": "You meant those \"[Context=] [Table=]\" etc.? Yes I manually mapped those to override least frequent vocabs near the end of ALBERT vocab file. I decreases the max context so got around 200 of them.",
    "714465": "ok so your best model is ALBERT right.. When I tried ALBERT and set others same as BERT did, my CV will decrease 8%...",
    "714507": "In another kaggle  NLP competition running now, I also found some people said that after using ALBERT, their performance got worse than BERT, but in the squad leaderboard, most top-ranked models are based on ALBERT.",
    "714567": "httpwwwfszyc If you've re-processed the train data the ALBERT way, then it's probably parameter setting and some issues with the use of sentencepiece instead of wordpiece (quite a lot of such when I did it). I had to manually map all the special vocabs into (and override) least frequent vocabs near the end of ALBERT vocab file by visual inspection... since it doen't have [unused] placeholders for custom vocabs like BERT. Also, A learning rate of 3e-5 seems to be the max, the loss doesn't decrease with 5e-5.",
    "714570": "guozhiyu0914 It'd be interested to see whether it's really the case or just parameter setting issues. They call ALBERT the new SOTA not because of SQuaD performance but rather it beats BERT in ALL glue tasks if I'm not wrong. That's for ALBERT xxlarge though.",
    "714869": "my way is to enlarge the vocab size from 30000 to 30209 by adding np.random.normal(0,0.02) and change the embedding layer shape. not sure this will make a difference to yours preprocessing method.",
    "714978": "I'm new to NLP, and I think this might be a stupid question, but I found that ALBERT has much less parameters compared to BERT (especially xlarge), which I thought, means that it calculates much faster. I'm trying to implement ALBERT because of this. Am I wrong?",
    "714984": "kitkatbar0429 ALBERT has mush fewer parameters, yes. It calculates much faster, no. ALBERT basically reduce the number of parameters by sharing them across different BERT layers. In this way it makes the performance of BERT more stable with increasing depth and size, and that's why ALBERT can go all the way to xlarge and xxlarge without drastic drop in performance as BERT does. However, there is still almost the same amount of computation (say for BERT-large and ALBERT-large). There's barely any reduction in training/inference time.",
    "715017": "siriuself I did not know that, thank you !!",
    "715271": "Do you mean you just use wordpiece rather than sentencepiece ？I had tried sentencepiece  but worse than bert",
    "715521": "",
    "715524": "After examine the tokenization of ALBERT carefully, I found only punctuations (for example, ',' and 'good,' ) will have difference between word encoding and sentence encoding.",
    "715708": "httpwwwfszyc I don't get it. Are you using the SentencePiece model that came with ALBERT? If so the tokenization and encoding will be completely different from those of BERT. Or are you suggesting something else?",
    "715752": "I just change function *def tokenize* in bert_utils.py and *def convert_tokens_to_ids* in tokenization.py and compare the result from tokenizing  sentencewise to tokenizing word by word .",
    "715835": "httpwwwfszyc That's expected. For SentencePiece the word boundary is maintained by default, in other words \"_\" will always be at the start (not in the middle) of a token. If so there's no difference in tokenize(sentence) vs [tokenize(word) for word in sentence.split()]\n\nAlso, do make sure that whatever passed into the SP tokenizer has been lower()-ed, I think I encountered this disastrous issue along the way.",
    "715992": "My current solution takes 10 minutes (625s) to run when I commit (public set). However, when I commit (pvt set), it times out.",
    "716006": "That seems weird to me. I've got better with at least 2 pts F1 with ALBERT xlarge and xxlarge straight out with bert-joint parameters with nq_eval, without any tuning (batch_size=32, alpha=3e-5, epoch=1) and even when I decreases max_length and increase doc_stride.. This is with sentencepiece and corresponding modifications to the preprocessing code of course. I do suspect it might be due to not explicitly calling lower() before passing into sentencepiece, somewhere. I weakly remember having such issues in early experiments with ALBERT for some other tasks (even if the case check is passed), and it's gonna be disastrous since any tokens with upper case chars would be",
    "716027": "kenkrige that's weird my solutions takes 828s to commit but it doesn't timeout on private test set.",
    "716053": "Tghanks @axel81 thanks for that info, it is very helpful. I will search for another reason in my code. The error message is not very informative, it just says \"Notebook Exceeded Allowed Compute\". I am not sure if that means GPU time or too much RAM used.",
    "716083": "kenkrige if it was a timeout issue you would have gotten this error `Notebook Timeout`. So in your case it must be related to GPU or RAM memory limit exceeded which caused private run to fail",
    "716097": "OK, thanks. That helps a lot.",
    "716732": "siriuself I was experimenting with ALBERT x large v1 but it gives very worse results compared to Bert. I am using nq-vocab.txt but the accuracy is very worse. As far as I understand tokenization process for ALBERT and BERT is same right? Even with ALBERT vocab my model gives very bad results my LR is 3e-5",
    "716795": "Thanks @axel81 You were absolutely right. It was a RAM problem, not timeout. By using module `gs` and being more careful about garbage collection, it ran perfectly. No garbage collection was necessary in the commit stage on the smaller public set.",
    "716800": "Good to hear that :)",
    "716827": "axel81 \nIf you use official pretrained model for BERT and ALBERT, the tokenization process is not same.\nBERT uses wordpieces and ALBERT uses sentencepiece.",
    "716829": "Oh Okay makes sense now why my ALBERT version is not giving good results"
  },
  "source": "meta"
}