{
  "id": 207030,
  "title": "Adaptive Softmax",
  "url": "/competitions/riiid-test-answer-prediction/discussion/207030",
  "author_name": "عثمان",
  "post_date": "2020-12-27T19:10:27.119000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>One of the things I've been holding out on with the intent of introducing it into my training pipeline once my cv hit 800 was pre-training. Transformers are known for being extremely data hungry and so far, all the public saint/sakt implementations are trained from scratch. Why is that?</p>\n<p>There are at least two discussion topics which have already in passive shared this idea, so I hope no one thinks this is a last minute bombshell.</p>\n<p>Theoretically, we should be able to pre-train our networks similar to how they did in the <a href=\"https://arxiv.org/abs/1910.13461\" target=\"_blank\">BART paper</a>. The ablated options are:</p>\n<ul>\n<li>w/ Token Masking</li>\n<li>w/ Token Deletion</li>\n<li>w/ Text Infilling</li>\n<li>w/ Document Rotation</li>\n<li>w/ Sentence Shuffling</li>\n<li>w/ Text Infilling + Sentence Shuffling</li>\n</ul>\n<p>Token Masking is relatively straightforward and has good NLP performance and should be dead easy to implement. I went ahead and added such training in front of my model's regular training and immediately discovered that the softmax across the questions embedding (I tied weights) caused me to run out of ram and also single steps in an epoch were taking like 5 minutes a pop! Then I discovered <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.AdaptiveLogSoftmaxWithLoss.html?highlight=adaptive%20softmax#torch.nn.AdaptiveLogSoftmaxWithLoss\" target=\"_blank\">this little thing</a> in the Pytorch documentation.</p>\n<p>It's usage seems simple enough, one would just need to re-enumerate their indexing s.t. the most frequent content_id's appear first. I have a <strong>very strong</strong> hunch top LB scorers are probably using this, insomuch as I haven't been able to find combinations of features (ts based, clustering based, target encoding based, interaction based, etc). that easily lead to &gt;800LB scores.</p>\n<p>Would any of the top 100 lb people be kind enough to give a head nod if this is a decent direction?</p>",
  "messages": [
    {
      "id": 1128820,
      "postDate": "2020-12-27T19:10:27.120Z",
      "content": "<p>One of the things I've been holding out on with the intent of introducing it into my training pipeline once my cv hit 800 was pre-training. Transformers are known for being extremely data hungry and so far, all the public saint/sakt implementations are trained from scratch. Why is that?</p>\n<p>There are at least two discussion topics which have already in passive shared this idea, so I hope no one thinks this is a last minute bombshell.</p>\n<p>Theoretically, we should be able to pre-train our networks similar to how they did in the <a href=\"https://arxiv.org/abs/1910.13461\" target=\"_blank\">BART paper</a>. The ablated options are:</p>\n<ul>\n<li>w/ Token Masking</li>\n<li>w/ Token Deletion</li>\n<li>w/ Text Infilling</li>\n<li>w/ Document Rotation</li>\n<li>w/ Sentence Shuffling</li>\n<li>w/ Text Infilling + Sentence Shuffling</li>\n</ul>\n<p>Token Masking is relatively straightforward and has good NLP performance and should be dead easy to implement. I went ahead and added such training in front of my model's regular training and immediately discovered that the softmax across the questions embedding (I tied weights) caused me to run out of ram and also single steps in an epoch were taking like 5 minutes a pop! Then I discovered <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.AdaptiveLogSoftmaxWithLoss.html?highlight=adaptive%20softmax#torch.nn.AdaptiveLogSoftmaxWithLoss\" target=\"_blank\">this little thing</a> in the Pytorch documentation.</p>\n<p>It's usage seems simple enough, one would just need to re-enumerate their indexing s.t. the most frequent content_id's appear first. I have a <strong>very strong</strong> hunch top LB scorers are probably using this, insomuch as I haven't been able to find combinations of features (ts based, clustering based, target encoding based, interaction based, etc). that easily lead to &gt;800LB scores.</p>\n<p>Would any of the top 100 lb people be kind enough to give a head nod if this is a decent direction?</p>",
      "rawMarkdown": "One of the things I've been holding out on with the intent of introducing it into my training pipeline once my cv hit 800 was pre-training. Transformers are known for being extremely data hungry and so far, all the public saint/sakt implementations are trained from scratch. Why is that?\n\nThere are at least two discussion topics which have already in passive shared this idea, so I hope no one thinks this is a last minute bombshell.\n\nTheoretically, we should be able to pre-train our networks similar to how they did in the [BART paper](https://arxiv.org/abs/1910.13461). The ablated options are:\n\n- w/ Token Masking\n- w/ Token Deletion\n- w/ Text Infilling\n- w/ Document Rotation\n- w/ Sentence Shuffling\n- w/ Text Infilling + Sentence Shuffling\n\nToken Masking is relatively straightforward and has good NLP performance and should be dead easy to implement. I went ahead and added such training in front of my model's regular training and immediately discovered that the softmax across the questions embedding (I tied weights) caused me to run out of ram and also single steps in an epoch were taking like 5 minutes a pop! Then I discovered [this little thing](https://pytorch.org/docs/stable/generated/torch.nn.AdaptiveLogSoftmaxWithLoss.html?highlight=adaptive%20softmax#torch.nn.AdaptiveLogSoftmaxWithLoss) in the Pytorch documentation.\n\nIt's usage seems simple enough, one would just need to re-enumerate their indexing s.t. the most frequent content_id's appear first. I have a **very strong** hunch top LB scorers are probably using this, insomuch as I haven't been able to find combinations of features (ts based, clustering based, target encoding based, interaction based, etc). that easily lead to >800LB scores.\n\nWould any of the top 100 lb people be kind enough to give a head nod if this is a decent direction?",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1128820": "One of the things I've been holding out on with the intent of introducing it into my training pipeline once my cv hit 800 was pre-training. Transformers are known for being extremely data hungry and so far, all the public saint/sakt implementations are trained from scratch. Why is that?\n\nThere are at least two discussion topics which have already in passive shared this idea, so I hope no one thinks this is a last minute bombshell.\n\nTheoretically, we should be able to pre-train our networks similar to how they did in the [BART paper](https://arxiv.org/abs/1910.13461). The ablated options are:\n\n- w/ Token Masking\n- w/ Token Deletion\n- w/ Text Infilling\n- w/ Document Rotation\n- w/ Sentence Shuffling\n- w/ Text Infilling + Sentence Shuffling\n\nToken Masking is relatively straightforward and has good NLP performance and should be dead easy to implement. I went ahead and added such training in front of my model's regular training and immediately discovered that the softmax across the questions embedding (I tied weights) caused me to run out of ram and also single steps in an epoch were taking like 5 minutes a pop! Then I discovered [this little thing](https://pytorch.org/docs/stable/generated/torch.nn.AdaptiveLogSoftmaxWithLoss.html?highlight=adaptive%20softmax#torch.nn.AdaptiveLogSoftmaxWithLoss) in the Pytorch documentation.\n\nIt's usage seems simple enough, one would just need to re-enumerate their indexing s.t. the most frequent content_id's appear first. I have a **very strong** hunch top LB scorers are probably using this, insomuch as I haven't been able to find combinations of features (ts based, clustering based, target encoding based, interaction based, etc). that easily lead to >800LB scores.\n\nWould any of the top 100 lb people be kind enough to give a head nod if this is a decent direction?"
  }
}