{
  "id": 114717,
  "title": "When next NLP match begins, one question",
  "url": "/competitions/tensorflow2-question-answering/discussion/114717",
  "author_name": "",
  "post_date": "2019-10-28T21:34:34.244569200Z",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Since Jigsaw II,  I have waited for so long. Hope everyone enjoys. \nOne question: Are BERT GPT2 XLNET ALBERT/ROBERTA T5 already pretrained with this training data (Wikipedia articles) ?</p>",
  "messages": [
    {
      "id": "660203",
      "postDate": "10/28/2019 21:34:34",
      "content": "<p>Since Jigsaw II,  I have waited for so long. Hope everyone enjoys. \nOne question: Are BERT GPT2 XLNET ALBERT/ROBERTA T5 already pretrained with this training data (Wikipedia articles) ?</p>",
      "rawMarkdown": "Since Jigsaw II,  I have waited for so long. Hope everyone enjoys. \nOne question: Are BERT GPT2 XLNET ALBERT/ROBERTA T5 already pretrained with this training data (Wikipedia articles) ?",
      "votes": null
    },
    {
      "id": "660209",
      "postDate": "10/28/2019 21:45:10",
      "content": "<p>Emm 16GB textdata... how to train via colab😲 ... </p>",
      "rawMarkdown": "Emm 16GB textdata... how to train via colab😲 ...",
      "votes": null
    },
    {
      "id": "660210",
      "postDate": "10/28/2019 21:48:34",
      "content": "<p>The competition of computational resource, hahahaha</p>",
      "rawMarkdown": "The competition of computational resource, hahahaha",
      "votes": null
    },
    {
      "id": "660302",
      "postDate": "10/29/2019 01:12:40",
      "content": "<p>usually they will provide GCP credit. One can hope :) </p>",
      "rawMarkdown": "usually they will provide GCP credit. One can hope :)",
      "votes": null
    },
    {
      "id": "661208",
      "postDate": "10/30/2019 02:42:22",
      "content": "<p>If you look at the data format, you will notice that there is a lot of redundancy. The actual amount of text data is more like 2-3GB.</p>",
      "rawMarkdown": "If you look at the data format, you will notice that there is a lot of redundancy. The actual amount of text data is more like 2-3GB.",
      "votes": null
    },
    {
      "id": "662121",
      "postDate": "10/31/2019 05:38:20",
      "content": "<p>...</p>",
      "rawMarkdown": "...",
      "votes": null
    },
    {
      "id": "662341",
      "postDate": "10/31/2019 12:55:59",
      "content": "<p>Did you solve the problem? I am also troubled by this problem....</p>",
      "rawMarkdown": "Did you solve the problem? I am also troubled by this problem....",
      "votes": null
    },
    {
      "id": "662353",
      "postDate": "10/31/2019 13:16:17",
      "content": "<p>no because i don't know how to slice the corpus of the document text with a efficient function...\nI think the process will be like:\n- get corpus of document texts\n- dealing y labels inside the sliced document texts\n- make a demo dataset\n- fit a BERT-based model with as much length as my GPU can afford (second problem: how to treat a thousand-token text)\n- train the whole data with checkpoint saving, stopping, and training the next block of data\n- tbc.</p>",
      "rawMarkdown": "no because i don't know how to slice the corpus of the document text with a efficient function...\nI think the process will be like:\n- get corpus of document texts\n- dealing y labels inside the sliced document texts\n- make a demo dataset\n- fit a BERT-based model with as much length as my GPU can afford (second problem: how to treat a thousand-token text)\n- train the whole data with checkpoint saving, stopping, and training the next block of data\n- tbc.",
      "votes": null
    },
    {
      "id": "662816",
      "postDate": "11/01/2019 03:11:43",
      "content": "<p>Looking forward...</p>",
      "rawMarkdown": "Looking forward...",
      "votes": null
    },
    {
      "id": "698259",
      "postDate": "12/19/2019 01:43:57",
      "content": "<p><a href=\"/liweicai\">@liweicai</a> What redundancy?</p>",
      "rawMarkdown": "liweicai What redundancy?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 660209,
      "author_name": "httpwwwfszyc",
      "author_url": "",
      "post_date": "10/28/2019 21:45:10",
      "content": "<p>Emm 16GB textdata... how to train via colab😲 ... </p>",
      "votes": null,
      "replies": [
        {
          "id": 660302,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "10/29/2019 01:12:40",
          "content": "<p>usually they will provide GCP credit. One can hope :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 661208,
          "author_name": "liweicai",
          "author_url": "",
          "post_date": "10/30/2019 02:42:22",
          "content": "<p>If you look at the data format, you will notice that there is a lot of redundancy. The actual amount of text data is more like 2-3GB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662121,
          "author_name": "lizyang29",
          "author_url": "",
          "post_date": "10/31/2019 05:38:20",
          "content": "<p>...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662341,
          "author_name": "retunrd1",
          "author_url": "",
          "post_date": "10/31/2019 12:55:59",
          "content": "<p>Did you solve the problem? I am also troubled by this problem....</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662353,
          "author_name": "httpwwwfszyc",
          "author_url": "",
          "post_date": "10/31/2019 13:16:17",
          "content": "<p>no because i don't know how to slice the corpus of the document text with a efficient function...\nI think the process will be like:\n- get corpus of document texts\n- dealing y labels inside the sliced document texts\n- make a demo dataset\n- fit a BERT-based model with as much length as my GPU can afford (second problem: how to treat a thousand-token text)\n- train the whole data with checkpoint saving, stopping, and training the next block of data\n- tbc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 698259,
          "author_name": "rohitagarwal",
          "author_url": "",
          "post_date": "12/19/2019 01:43:57",
          "content": "<p><a href=\"/liweicai\">@liweicai</a> What redundancy?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 660210,
      "author_name": "dafchen",
      "author_url": "",
      "post_date": "10/28/2019 21:48:34",
      "content": "<p>The competition of computational resource, hahahaha</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 662816,
      "author_name": "peterbon",
      "author_url": "",
      "post_date": "11/01/2019 03:11:43",
      "content": "<p>Looking forward...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "660203": "Since Jigsaw II,  I have waited for so long. Hope everyone enjoys. \nOne question: Are BERT GPT2 XLNET ALBERT/ROBERTA T5 already pretrained with this training data (Wikipedia articles) ?",
    "660209": "Emm 16GB textdata... how to train via colab😲 ...",
    "660210": "The competition of computational resource, hahahaha",
    "660302": "usually they will provide GCP credit. One can hope :)",
    "661208": "If you look at the data format, you will notice that there is a lot of redundancy. The actual amount of text data is more like 2-3GB.",
    "662121": "...",
    "662341": "Did you solve the problem? I am also troubled by this problem....",
    "662353": "no because i don't know how to slice the corpus of the document text with a efficient function...\nI think the process will be like:\n- get corpus of document texts\n- dealing y labels inside the sliced document texts\n- make a demo dataset\n- fit a BERT-based model with as much length as my GPU can afford (second problem: how to treat a thousand-token text)\n- train the whole data with checkpoint saving, stopping, and training the next block of data\n- tbc.",
    "662816": "Looking forward...",
    "698259": "liweicai What redundancy?"
  },
  "source": "meta"
}