{
  "id": 119000,
  "title": "PyTorch BERT training kernel",
  "url": "/competitions/tensorflow2-question-answering/discussion/119000",
  "author_name": "",
  "post_date": "2019-11-26T02:21:05.225994Z",
  "votes": 23,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I published a PyTorch BERT training kernel <a href=\"https://www.kaggle.com/sakami/tfqa-pytorch-baseline\">here</a>.</p>\n\n<p>Its inference part is not refined (can't evaluate more than 2000 examples...), so please give me some advice. :)</p>",
  "messages": [
    {
      "id": "681355",
      "postDate": "11/26/2019 02:21:05",
      "content": "<p>I published a PyTorch BERT training kernel <a href=\"https://www.kaggle.com/sakami/tfqa-pytorch-baseline\">here</a>.</p>\n\n<p>Its inference part is not refined (can't evaluate more than 2000 examples...), so please give me some advice. :)</p>",
      "rawMarkdown": "I published a PyTorch BERT training kernel [here](https://www.kaggle.com/sakami/tfqa-pytorch-baseline).\n\nIts inference part is not refined (can't evaluate more than 2000 examples...), so please give me some advice. :)",
      "votes": null
    },
    {
      "id": "681490",
      "postDate": "11/26/2019 06:57:36",
      "content": "<p>Thanks for posting! Btw I'm using a similar scheme of building an index between word tokenization and bert tokenization, and pytorch transformers are really slow here. Here is a small fix <a href=\"https://github.com/lopuhin/transformers/commit/d1c7ad100869865f580c441106fbeeecfeed8861\">https://github.com/lopuhin/transformers/commit/d1c7ad100869865f580c441106fbeeecfeed8861</a> which is relevant if you have many special tokens - I put top ~10 markup tags there. Likely it's possible to improve this further by calling sentencepiece tokenizer directly, instead of using transformers wrapper.</p>",
      "rawMarkdown": "Thanks for posting! Btw I'm using a similar scheme of building an index between word tokenization and bert tokenization, and pytorch transformers are really slow here. Here is a small fix https://github.com/lopuhin/transformers/commit/d1c7ad100869865f580c441106fbeeecfeed8861 which is relevant if you have many special tokens - I put top ~10 markup tags there. Likely it's possible to improve this further by calling sentencepiece tokenizer directly, instead of using transformers wrapper.",
      "votes": null
    },
    {
      "id": "681689",
      "postDate": "11/26/2019 12:25:53",
      "content": "<p>Thank you, it's quite smart!</p>",
      "rawMarkdown": "Thank you, it's quite smart!",
      "votes": null
    },
    {
      "id": "683862",
      "postDate": "11/28/2019 23:12:04",
      "content": "<p>I think you can make the batch size larger when doing inference, because you don't need to store Adam parameters anymore, which should make it faster.</p>",
      "rawMarkdown": "I think you can make the batch size larger when doing inference, because you don't need to store Adam parameters anymore, which should make it faster.",
      "votes": null
    },
    {
      "id": "684038",
      "postDate": "11/29/2019 06:44:32",
      "content": "<p>I think you can make the batch size larger</p>",
      "rawMarkdown": "I think you can make the batch size larger",
      "votes": null
    },
    {
      "id": "690102",
      "postDate": "12/08/2019 01:06:36",
      "content": "<p>New to this competition and just wondering, why am I seeing so few pytorch kernels in this competition? Is it because the competition is sponsored by Google?</p>",
      "rawMarkdown": "New to this competition and just wondering, why am I seeing so few pytorch kernels in this competition? Is it because the competition is sponsored by Google?",
      "votes": null
    },
    {
      "id": "690337",
      "postDate": "12/08/2019 12:06:21",
      "content": "<p>Sorry for late reply!\nThank you for your comment. I tried, but couldn't. :(</p>",
      "rawMarkdown": "Sorry for late reply!\nThank you for your comment. I tried, but couldn't. :(",
      "votes": null
    },
    {
      "id": "690339",
      "postDate": "12/08/2019 12:07:29",
      "content": "<p>Sorry for late reply!\nThank you for your advice. As I wrote above, I couldn't make batch size larger in kaggle kernel.</p>",
      "rawMarkdown": "Sorry for late reply!\nThank you for your advice. As I wrote above, I couldn't make batch size larger in kaggle kernel.",
      "votes": null
    },
    {
      "id": "690344",
      "postDate": "12/08/2019 12:17:34",
      "content": "<p>I think it is because there is few PyTorch code working on this problem and it is hard to understand official tensorflow kernel completely.\nSQuAD is similar problem to this competition, so you can use as reference. (<a href=\"https://rajpurkar.github.io/SQuAD-explorer/\">https://rajpurkar.github.io/SQuAD-explorer/</a>)\nEspecially, this code is very helpful. (<a href=\"https://github.com/huggingface/transformers/blob/master/examples/run_squad.py\">https://github.com/huggingface/transformers/blob/master/examples/run_squad.py</a>)</p>",
      "rawMarkdown": "I think it is because there is few PyTorch code working on this problem and it is hard to understand official tensorflow kernel completely.\nSQuAD is similar problem to this competition, so you can use as reference. (https://rajpurkar.github.io/SQuAD-explorer/)\nEspecially, this code is very helpful. (https://github.com/huggingface/transformers/blob/master/examples/run_squad.py)",
      "votes": null
    },
    {
      "id": "692531",
      "postDate": "12/11/2019 11:54:57",
      "content": "<p>Thanks for your nice Pytorch kernel! Would you mind also publish the evaluation part? Thanks.</p>",
      "rawMarkdown": "Thanks for your nice Pytorch kernel! Would you mind also publish the evaluation part? Thanks.",
      "votes": null
    },
    {
      "id": "697557",
      "postDate": "12/18/2019 04:41:59",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "697656",
      "postDate": "12/18/2019 08:30:00",
      "content": "<p>I also have inference kernel based on this sakami's awesome kernel.\nI hope <a href=\"https://www.kaggle.com/sakami/tfqa-pytorch-baseline/comments#696848\">https://www.kaggle.com/sakami/tfqa-pytorch-baseline/comments#696848</a> helps. The original json reader needs a lot of memory in initialize phase.</p>",
      "rawMarkdown": "I also have inference kernel based on this sakami's awesome kernel.\nI hope https://www.kaggle.com/sakami/tfqa-pytorch-baseline/comments#696848 helps. The original json reader needs a lot of memory in initialize phase.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 681490,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "11/26/2019 06:57:36",
      "content": "<p>Thanks for posting! Btw I'm using a similar scheme of building an index between word tokenization and bert tokenization, and pytorch transformers are really slow here. Here is a small fix <a href=\"https://github.com/lopuhin/transformers/commit/d1c7ad100869865f580c441106fbeeecfeed8861\">https://github.com/lopuhin/transformers/commit/d1c7ad100869865f580c441106fbeeecfeed8861</a> which is relevant if you have many special tokens - I put top ~10 markup tags there. Likely it's possible to improve this further by calling sentencepiece tokenizer directly, instead of using transformers wrapper.</p>",
      "votes": null,
      "replies": [
        {
          "id": 681689,
          "author_name": "sakami",
          "author_url": "",
          "post_date": "11/26/2019 12:25:53",
          "content": "<p>Thank you, it's quite smart!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 697557,
          "author_name": "",
          "author_url": "",
          "post_date": "12/18/2019 04:41:59",
          "content": "",
          "votes": null,
          "replies": []
        },
        {
          "id": 697656,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/18/2019 08:30:00",
          "content": "<p>I also have inference kernel based on this sakami's awesome kernel.\nI hope <a href=\"https://www.kaggle.com/sakami/tfqa-pytorch-baseline/comments#696848\">https://www.kaggle.com/sakami/tfqa-pytorch-baseline/comments#696848</a> helps. The original json reader needs a lot of memory in initialize phase.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 683862,
      "author_name": "mamamot",
      "author_url": "",
      "post_date": "11/28/2019 23:12:04",
      "content": "<p>I think you can make the batch size larger when doing inference, because you don't need to store Adam parameters anymore, which should make it faster.</p>",
      "votes": null,
      "replies": [
        {
          "id": 690339,
          "author_name": "sakami",
          "author_url": "",
          "post_date": "12/08/2019 12:07:29",
          "content": "<p>Sorry for late reply!\nThank you for your advice. As I wrote above, I couldn't make batch size larger in kaggle kernel.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 684038,
      "author_name": "jatin0123456",
      "author_url": "",
      "post_date": "11/29/2019 06:44:32",
      "content": "<p>I think you can make the batch size larger</p>",
      "votes": null,
      "replies": [
        {
          "id": 690337,
          "author_name": "sakami",
          "author_url": "",
          "post_date": "12/08/2019 12:06:21",
          "content": "<p>Sorry for late reply!\nThank you for your comment. I tried, but couldn't. :(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 690102,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "12/08/2019 01:06:36",
      "content": "<p>New to this competition and just wondering, why am I seeing so few pytorch kernels in this competition? Is it because the competition is sponsored by Google?</p>",
      "votes": null,
      "replies": [
        {
          "id": 690344,
          "author_name": "sakami",
          "author_url": "",
          "post_date": "12/08/2019 12:17:34",
          "content": "<p>I think it is because there is few PyTorch code working on this problem and it is hard to understand official tensorflow kernel completely.\nSQuAD is similar problem to this competition, so you can use as reference. (<a href=\"https://rajpurkar.github.io/SQuAD-explorer/\">https://rajpurkar.github.io/SQuAD-explorer/</a>)\nEspecially, this code is very helpful. (<a href=\"https://github.com/huggingface/transformers/blob/master/examples/run_squad.py\">https://github.com/huggingface/transformers/blob/master/examples/run_squad.py</a>)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 692531,
      "author_name": "niuddd",
      "author_url": "",
      "post_date": "12/11/2019 11:54:57",
      "content": "<p>Thanks for your nice Pytorch kernel! Would you mind also publish the evaluation part? Thanks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "681355": "I published a PyTorch BERT training kernel [here](https://www.kaggle.com/sakami/tfqa-pytorch-baseline).\n\nIts inference part is not refined (can't evaluate more than 2000 examples...), so please give me some advice. :)",
    "681490": "Thanks for posting! Btw I'm using a similar scheme of building an index between word tokenization and bert tokenization, and pytorch transformers are really slow here. Here is a small fix https://github.com/lopuhin/transformers/commit/d1c7ad100869865f580c441106fbeeecfeed8861 which is relevant if you have many special tokens - I put top ~10 markup tags there. Likely it's possible to improve this further by calling sentencepiece tokenizer directly, instead of using transformers wrapper.",
    "681689": "Thank you, it's quite smart!",
    "683862": "I think you can make the batch size larger when doing inference, because you don't need to store Adam parameters anymore, which should make it faster.",
    "684038": "I think you can make the batch size larger",
    "690102": "New to this competition and just wondering, why am I seeing so few pytorch kernels in this competition? Is it because the competition is sponsored by Google?",
    "690337": "Sorry for late reply!\nThank you for your comment. I tried, but couldn't. :(",
    "690339": "Sorry for late reply!\nThank you for your advice. As I wrote above, I couldn't make batch size larger in kaggle kernel.",
    "690344": "I think it is because there is few PyTorch code working on this problem and it is hard to understand official tensorflow kernel completely.\nSQuAD is similar problem to this competition, so you can use as reference. (https://rajpurkar.github.io/SQuAD-explorer/)\nEspecially, this code is very helpful. (https://github.com/huggingface/transformers/blob/master/examples/run_squad.py)",
    "692531": "Thanks for your nice Pytorch kernel! Would you mind also publish the evaluation part? Thanks.",
    "697557": "",
    "697656": "I also have inference kernel based on this sakami's awesome kernel.\nI hope https://www.kaggle.com/sakami/tfqa-pytorch-baseline/comments#696848 helps. The original json reader needs a lot of memory in initialize phase."
  },
  "source": "meta"
}