{
  "id": 406311,
  "title": "Has anybody tried pre-training the transformer?",
  "url": "/competitions/asl-signs/discussion/406311",
  "author_name": "",
  "post_date": "2023-05-02T00:43:51.577951300Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I thought that pre-training task could be beneficial for model initialization here, because it would provide the model the knowledge of the overall geometry of the human body, however I tried continious MLM (mask some keypoints -&gt; predict them as regression targets), but it didn't work.</p>\n<p>I found the <a href=\"https://arxiv.org/pdf/2110.05877.pdf\" target=\"_blank\">OpenHands</a> paper that claimed that random masking is too easy, since model just learns to interpolate the keypoints, so this representation is useless, and they showed that <a href=\"https://arxiv.org/abs/1909.04656\" target=\"_blank\">DPC (Dense Predictive Coding)</a> is a harder and better task for pretraining.</p>\n<p>Basically instead of randomly masking, model tries to predict it's own hidden state in autoregressive manner, however for me it didn't work either. Did pre-training work for anybody?</p>",
  "messages": [
    {
      "id": "2241921",
      "postDate": "05/02/2023 00:43:51",
      "content": "<p>I thought that pre-training task could be beneficial for model initialization here, because it would provide the model the knowledge of the overall geometry of the human body, however I tried continious MLM (mask some keypoints -&gt; predict them as regression targets), but it didn't work.</p>\n<p>I found the <a href=\"https://arxiv.org/pdf/2110.05877.pdf\" target=\"_blank\">OpenHands</a> paper that claimed that random masking is too easy, since model just learns to interpolate the keypoints, so this representation is useless, and they showed that <a href=\"https://arxiv.org/abs/1909.04656\" target=\"_blank\">DPC (Dense Predictive Coding)</a> is a harder and better task for pretraining.</p>\n<p>Basically instead of randomly masking, model tries to predict it's own hidden state in autoregressive manner, however for me it didn't work either. Did pre-training work for anybody?</p>",
      "rawMarkdown": "I thought that pre-training task could be beneficial for model initialization here, because it would provide the model the knowledge of the overall geometry of the human body, however I tried continious MLM (mask some keypoints -> predict them as regression targets), but it didn't work.\n\nI found the [OpenHands](https://arxiv.org/pdf/2110.05877.pdf) paper that claimed that random masking is too easy, since model just learns to interpolate the keypoints, so this representation is useless, and they showed that [DPC (Dense Predictive Coding)](https://arxiv.org/abs/1909.04656) is a harder and better task for pretraining.\n\nBasically instead of randomly masking, model tries to predict it's own hidden state in autoregressive manner, however for me it didn't work either. Did pre-training work for anybody?",
      "votes": null
    },
    {
      "id": "2241923",
      "postDate": "05/02/2023 00:46:55",
      "content": "<p>IMO, i think pretraining is most effective if we have more (external) data than the train data. For example, imagine that we have 400k external data videos with words that do not appear in the train dataset. Then we can pretrain on those (to learn about lips, hands, pose attention) and then afterward finetune on competition data</p>\n<p>I tried some external datasets but I could not get the distribution of Landmark x y coordinates to be similar enough to train data to be helpful</p>",
      "rawMarkdown": "IMO, i think pretraining is most effective if we have more (external) data than the train data. For example, imagine that we have 400k external data videos with words that do not appear in the train dataset. Then we can pretrain on those (to learn about lips, hands, pose attention) and then afterward finetune on competition data\n\nI tried some external datasets but I could not get the distribution of Landmark x y coordinates to be similar enough to train data to be helpful",
      "votes": null
    },
    {
      "id": "2241937",
      "postDate": "05/02/2023 00:57:55",
      "content": "<p>I tried too,  almost the same method as yours, cost me a lot of electricity fee but getting worse results than no pre-training.</p>",
      "rawMarkdown": "I tried too,  almost the same method as yours, cost me a lot of electricity fee but getting worse results than no pre-training.",
      "votes": null
    },
    {
      "id": "2246995",
      "postDate": "05/05/2023 15:51:43",
      "content": "<p>The problem of training a neural network is that when a mistake occurs in training, a person finds out too late, especially if the neural network has a large dataset. However, pre-training avoids most of these problems.</p>",
      "rawMarkdown": "The problem of training a neural network is that when a mistake occurs in training, a person finds out too late, especially if the neural network has a large dataset. However, pre-training avoids most of these problems.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2241923,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "05/02/2023 00:46:55",
      "content": "<p>IMO, i think pretraining is most effective if we have more (external) data than the train data. For example, imagine that we have 400k external data videos with words that do not appear in the train dataset. Then we can pretrain on those (to learn about lips, hands, pose attention) and then afterward finetune on competition data</p>\n<p>I tried some external datasets but I could not get the distribution of Landmark x y coordinates to be similar enough to train data to be helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2241937,
      "author_name": "xianghuang666",
      "author_url": "",
      "post_date": "05/02/2023 00:57:55",
      "content": "<p>I tried too,  almost the same method as yours, cost me a lot of electricity fee but getting worse results than no pre-training.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2246995,
      "author_name": "ivanisaev",
      "author_url": "",
      "post_date": "05/05/2023 15:51:43",
      "content": "<p>The problem of training a neural network is that when a mistake occurs in training, a person finds out too late, especially if the neural network has a large dataset. However, pre-training avoids most of these problems.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2241921": "I thought that pre-training task could be beneficial for model initialization here, because it would provide the model the knowledge of the overall geometry of the human body, however I tried continious MLM (mask some keypoints -> predict them as regression targets), but it didn't work.\n\nI found the [OpenHands](https://arxiv.org/pdf/2110.05877.pdf) paper that claimed that random masking is too easy, since model just learns to interpolate the keypoints, so this representation is useless, and they showed that [DPC (Dense Predictive Coding)](https://arxiv.org/abs/1909.04656) is a harder and better task for pretraining.\n\nBasically instead of randomly masking, model tries to predict it's own hidden state in autoregressive manner, however for me it didn't work either. Did pre-training work for anybody?",
    "2241923": "IMO, i think pretraining is most effective if we have more (external) data than the train data. For example, imagine that we have 400k external data videos with words that do not appear in the train dataset. Then we can pretrain on those (to learn about lips, hands, pose attention) and then afterward finetune on competition data\n\nI tried some external datasets but I could not get the distribution of Landmark x y coordinates to be similar enough to train data to be helpful",
    "2241937": "I tried too,  almost the same method as yours, cost me a lot of electricity fee but getting worse results than no pre-training.",
    "2246995": "The problem of training a neural network is that when a mistake occurs in training, a person finds out too late, especially if the neural network has a large dataset. However, pre-training avoids most of these problems."
  },
  "source": "meta"
}