{
  "id": 393495,
  "title": "Transformer architecture help! [1.159 score]",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/393495",
  "author_name": "",
  "post_date": "2023-03-09T16:29:00.110130Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>I tried out the transformer architecture for fun and created a notebook, which can be found at this link: <a href=\"https://www.kaggle.com/code/viktorcikojevic/transfomer-model-with-pretrained-model\" target=\"_blank\">https://www.kaggle.com/code/viktorcikojevic/transfomer-model-with-pretrained-model</a></p>\n<p>Although the model scored 1.159 on the leaderboard, I believe it ended up in a bad part of weight space. Specifically, when examining the \"wei\" (key query product), it has a large variance, causing the softmax to essentially transform \"wei\" into one-hot vectors. This likely means that the values being aggregated by the self-attention mechanism are not good (I assume, is this right?).</p>\n<p>To address this issue, I tried setting very small values for the weights for all layers in the network, which reduced the variance of \"wei\" at t=0. However, during training, the network weights eventually led to large variance in \"wei\".</p>\n<p>Does anyone have any insight into what could be causing this problem? Thank you! 😊</p>",
  "messages": [
    {
      "id": "2175081",
      "postDate": "03/09/2023 16:29:00",
      "content": "<p>Hello everyone,</p>\n<p>I tried out the transformer architecture for fun and created a notebook, which can be found at this link: <a href=\"https://www.kaggle.com/code/viktorcikojevic/transfomer-model-with-pretrained-model\" target=\"_blank\">https://www.kaggle.com/code/viktorcikojevic/transfomer-model-with-pretrained-model</a></p>\n<p>Although the model scored 1.159 on the leaderboard, I believe it ended up in a bad part of weight space. Specifically, when examining the \"wei\" (key query product), it has a large variance, causing the softmax to essentially transform \"wei\" into one-hot vectors. This likely means that the values being aggregated by the self-attention mechanism are not good (I assume, is this right?).</p>\n<p>To address this issue, I tried setting very small values for the weights for all layers in the network, which reduced the variance of \"wei\" at t=0. However, during training, the network weights eventually led to large variance in \"wei\".</p>\n<p>Does anyone have any insight into what could be causing this problem? Thank you! 😊</p>",
      "rawMarkdown": "Hello everyone,\n\nI tried out the transformer architecture for fun and created a notebook, which can be found at this link: https://www.kaggle.com/code/viktorcikojevic/transfomer-model-with-pretrained-model\n\nAlthough the model scored 1.159 on the leaderboard, I believe it ended up in a bad part of weight space. Specifically, when examining the \"wei\" (key query product), it has a large variance, causing the softmax to essentially transform \"wei\" into one-hot vectors. This likely means that the values being aggregated by the self-attention mechanism are not good (I assume, is this right?).\n\nTo address this issue, I tried setting very small values for the weights for all layers in the network, which reduced the variance of \"wei\" at t=0. However, during training, the network weights eventually led to large variance in \"wei\".\n\nDoes anyone have any insight into what could be causing this problem? Thank you! 😊",
      "votes": null
    },
    {
      "id": "2175487",
      "postDate": "03/09/2023 21:50:33",
      "content": "<p>I'm not sure using MSE on xyz predictions will yield the best results. Have you tried vonMishesFisher loss the others have shared?</p>",
      "rawMarkdown": "I'm not sure using MSE on xyz predictions will yield the best results. Have you tried vonMishesFisher loss the others have shared?",
      "votes": null
    },
    {
      "id": "2175995",
      "postDate": "03/10/2023 10:00:38",
      "content": "<p>Thanks for the input! I'll try it out. I implemented it here, if someone finds it useful: <a href=\"https://www.kaggle.com/viktorcikojevic/von-mises-fisher-3d-loss\" target=\"_blank\">https://www.kaggle.com/viktorcikojevic/von-mises-fisher-3d-loss</a></p>",
      "rawMarkdown": "Thanks for the input! I'll try it out. I implemented it here, if someone finds it useful: https://www.kaggle.com/viktorcikojevic/von-mises-fisher-3d-loss",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2175487,
      "author_name": "dipamc77",
      "author_url": "",
      "post_date": "03/09/2023 21:50:33",
      "content": "<p>I'm not sure using MSE on xyz predictions will yield the best results. Have you tried vonMishesFisher loss the others have shared?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2175995,
          "author_name": "viktorcikojevic",
          "author_url": "",
          "post_date": "03/10/2023 10:00:38",
          "content": "<p>Thanks for the input! I'll try it out. I implemented it here, if someone finds it useful: <a href=\"https://www.kaggle.com/viktorcikojevic/von-mises-fisher-3d-loss\" target=\"_blank\">https://www.kaggle.com/viktorcikojevic/von-mises-fisher-3d-loss</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2175081": "Hello everyone,\n\nI tried out the transformer architecture for fun and created a notebook, which can be found at this link: https://www.kaggle.com/code/viktorcikojevic/transfomer-model-with-pretrained-model\n\nAlthough the model scored 1.159 on the leaderboard, I believe it ended up in a bad part of weight space. Specifically, when examining the \"wei\" (key query product), it has a large variance, causing the softmax to essentially transform \"wei\" into one-hot vectors. This likely means that the values being aggregated by the self-attention mechanism are not good (I assume, is this right?).\n\nTo address this issue, I tried setting very small values for the weights for all layers in the network, which reduced the variance of \"wei\" at t=0. However, during training, the network weights eventually led to large variance in \"wei\".\n\nDoes anyone have any insight into what could be causing this problem? Thank you! 😊",
    "2175487": "I'm not sure using MSE on xyz predictions will yield the best results. Have you tried vonMishesFisher loss the others have shared?",
    "2175995": "Thanks for the input! I'll try it out. I implemented it here, if someone finds it useful: https://www.kaggle.com/viktorcikojevic/von-mises-fisher-3d-loss"
  },
  "source": "meta"
}