{
  "id": 209633,
  "title": "Transformer Pretraining? Did anyone try it?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209633",
  "author_name": "",
  "post_date": "2021-01-08T04:30:21.733998600Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Pretraining of Transformers models is widely done on difficult tasks as a way to initialize the model weights before finetuning. Availability of large unsupervised data is what is helping BERT/GPT models to out perform regular models with same amount of labelled data even though they have more parameters. In this competition SAINT/SAKT are one of the top performing models. I am curious is anyone has tried any sort of pretraining these models?</p>\n<p>Typical ideas for pretraining are:</p>\n<ol>\n<li>Next Question Prediction (Difficult Task)</li>\n<li>Next Part Prediction (Simpler Task)</li>\n<li>Masked Question Prediction? (not sure)</li>\n</ol>\n<p>I want to try these out myself as this would give me a good understanding on improving my knowledge on Transformer models further, but I am limited by the compute capacity. So would like to know if it worked for anyone.</p>\n<p>Thanks. Cheers.</p>",
  "messages": [
    {
      "id": "1143804",
      "postDate": "01/08/2021 04:30:21",
      "content": "<p>Pretraining of Transformers models is widely done on difficult tasks as a way to initialize the model weights before finetuning. Availability of large unsupervised data is what is helping BERT/GPT models to out perform regular models with same amount of labelled data even though they have more parameters. In this competition SAINT/SAKT are one of the top performing models. I am curious is anyone has tried any sort of pretraining these models?</p>\n<p>Typical ideas for pretraining are:</p>\n<ol>\n<li>Next Question Prediction (Difficult Task)</li>\n<li>Next Part Prediction (Simpler Task)</li>\n<li>Masked Question Prediction? (not sure)</li>\n</ol>\n<p>I want to try these out myself as this would give me a good understanding on improving my knowledge on Transformer models further, but I am limited by the compute capacity. So would like to know if it worked for anyone.</p>\n<p>Thanks. Cheers.</p>",
      "rawMarkdown": "Pretraining of Transformers models is widely done on difficult tasks as a way to initialize the model weights before finetuning. Availability of large unsupervised data is what is helping BERT/GPT models to out perform regular models with same amount of labelled data even though they have more parameters. In this competition SAINT/SAKT are one of the top performing models. I am curious is anyone has tried any sort of pretraining these models?\n\nTypical ideas for pretraining are:\n1. Next Question Prediction (Difficult Task)\n2. Next Part Prediction (Simpler Task)\n3. Masked Question Prediction? (not sure)\n\nI want to try these out myself as this would give me a good understanding on improving my knowledge on Transformer models further, but I am limited by the compute capacity. So would like to know if it worked for anyone.\n\nThanks. Cheers.",
      "votes": null
    },
    {
      "id": "1143901",
      "postDate": "01/08/2021 05:58:29",
      "content": "<p>i do some other try that is not the same as what you have mentioned, but it is similar to the pretraining idea.<br>\ni mask some questions in  the sequence of history questions, and i predict the correctness of them.<br>\nbecause of deadline, i just got 788 at lb.<br>\nmay it s worse than sakt.</p>",
      "rawMarkdown": "i do some other try that is not the same as what you have mentioned, but it is similar to the pretraining idea.\ni mask some questions in  the sequence of history questions, and i predict the correctness of them.\nbecause of deadline, i just got 788 at lb.\nmay it s worse than sakt.",
      "votes": null
    },
    {
      "id": "1143988",
      "postDate": "01/08/2021 06:59:29",
      "content": "<p>Ohh. That is also a nice idea. No change in architecture is required. </p>",
      "rawMarkdown": "Ohh. That is also a nice idea. No change in architecture is required.",
      "votes": null
    },
    {
      "id": "1144159",
      "postDate": "01/08/2021 09:14:36",
      "content": "<p>I tried it with usual masked language model - it doesn't help even a bit worse. During pretraining, the accuracy of predicting question is only about 38% , quite low</p>",
      "rawMarkdown": "I tried it with usual masked language model - it doesn't help even a bit worse. During pretraining, the accuracy of predicting question is only about 38% , quite low",
      "votes": null
    },
    {
      "id": "1144165",
      "postDate": "01/08/2021 09:18:59",
      "content": "<p>38% accuracy seems pretty bad but it is better than random accuracy I think. </p>",
      "rawMarkdown": "38% accuracy seems pretty bad but it is better than random accuracy I think.",
      "votes": null
    },
    {
      "id": "1144177",
      "postDate": "01/08/2021 09:39:43",
      "content": "<p>Only if it helps, but the final results seems a bit worse …</p>",
      "rawMarkdown": "Only if it helps, but the final results seems a bit worse ...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1143901,
      "author_name": "luffy521",
      "author_url": "",
      "post_date": "01/08/2021 05:58:29",
      "content": "<p>i do some other try that is not the same as what you have mentioned, but it is similar to the pretraining idea.<br>\ni mask some questions in  the sequence of history questions, and i predict the correctness of them.<br>\nbecause of deadline, i just got 788 at lb.<br>\nmay it s worse than sakt.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143988,
          "author_name": "manikanthr5",
          "author_url": "",
          "post_date": "01/08/2021 06:59:29",
          "content": "<p>Ohh. That is also a nice idea. No change in architecture is required. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144159,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/08/2021 09:14:36",
          "content": "<p>I tried it with usual masked language model - it doesn't help even a bit worse. During pretraining, the accuracy of predicting question is only about 38% , quite low</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144165,
          "author_name": "manikanthr5",
          "author_url": "",
          "post_date": "01/08/2021 09:18:59",
          "content": "<p>38% accuracy seems pretty bad but it is better than random accuracy I think. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144177,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "01/08/2021 09:39:43",
          "content": "<p>Only if it helps, but the final results seems a bit worse …</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1143804": "Pretraining of Transformers models is widely done on difficult tasks as a way to initialize the model weights before finetuning. Availability of large unsupervised data is what is helping BERT/GPT models to out perform regular models with same amount of labelled data even though they have more parameters. In this competition SAINT/SAKT are one of the top performing models. I am curious is anyone has tried any sort of pretraining these models?\n\nTypical ideas for pretraining are:\n1. Next Question Prediction (Difficult Task)\n2. Next Part Prediction (Simpler Task)\n3. Masked Question Prediction? (not sure)\n\nI want to try these out myself as this would give me a good understanding on improving my knowledge on Transformer models further, but I am limited by the compute capacity. So would like to know if it worked for anyone.\n\nThanks. Cheers.",
    "1143901": "i do some other try that is not the same as what you have mentioned, but it is similar to the pretraining idea.\ni mask some questions in  the sequence of history questions, and i predict the correctness of them.\nbecause of deadline, i just got 788 at lb.\nmay it s worse than sakt.",
    "1143988": "Ohh. That is also a nice idea. No change in architecture is required.",
    "1144159": "I tried it with usual masked language model - it doesn't help even a bit worse. During pretraining, the accuracy of predicting question is only about 38% , quite low",
    "1144165": "38% accuracy seems pretty bad but it is better than random accuracy I think.",
    "1144177": "Only if it helps, but the final results seems a bit worse ..."
  },
  "source": "meta"
}