{
  "id": 208769,
  "title": "SAINT's learning loss is not reduced.",
  "url": "/competitions/riiid-test-answer-prediction/discussion/208769",
  "author_name": "",
  "post_date": "2021-01-05T00:24:36.141444700Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I trained a SAINT implemented using Pytorch's Transformer module, but the training data loss stopped around 0.6.<br>\nTo make sure that the loss was not decreasing, I prepared data for evaluation to make sure that the AUC would be around 0.5, and it was still around 0.5.</p>\n<p>The implementation is in the following link.<br>\nI would like to hear the opinions of anyone who is familiar with SAINT and Transformer.</p>\n<p>notebook: <a href=\"https://www.kaggle.com/konumaru/saint-with-pytorch-transformer-module\" target=\"_blank\">https://www.kaggle.com/konumaru/saint-with-pytorch-transformer-module</a></p>",
  "messages": [
    {
      "id": "1138764",
      "postDate": "01/05/2021 00:24:36",
      "content": "<p>I trained a SAINT implemented using Pytorch's Transformer module, but the training data loss stopped around 0.6.<br>\nTo make sure that the loss was not decreasing, I prepared data for evaluation to make sure that the AUC would be around 0.5, and it was still around 0.5.</p>\n<p>The implementation is in the following link.<br>\nI would like to hear the opinions of anyone who is familiar with SAINT and Transformer.</p>\n<p>notebook: <a href=\"https://www.kaggle.com/konumaru/saint-with-pytorch-transformer-module\" target=\"_blank\">https://www.kaggle.com/konumaru/saint-with-pytorch-transformer-module</a></p>",
      "rawMarkdown": "I trained a SAINT implemented using Pytorch's Transformer module, but the training data loss stopped around 0.6.\nTo make sure that the loss was not decreasing, I prepared data for evaluation to make sure that the AUC would be around 0.5, and it was still around 0.5.\n\nThe implementation is in the following link.\nI would like to hear the opinions of anyone who is familiar with SAINT and Transformer.\n\nnotebook: https://www.kaggle.com/konumaru/saint-with-pytorch-transformer-module",
      "votes": null
    },
    {
      "id": "1138765",
      "postDate": "01/05/2021 00:24:48",
      "content": "<p>The following Issue may be the cause.<br>\nIn this case, does it mean that I need to implement Attention by myself?</p>\n<p><a href=\"https://github.com/pytorch/pytorch/issues/24816\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/24816</a></p>",
      "rawMarkdown": "The following Issue may be the cause.\nIn this case, does it mean that I need to implement Attention by myself?\n\nhttps://github.com/pytorch/pytorch/issues/24816",
      "votes": null
    },
    {
      "id": "1139405",
      "postDate": "01/05/2021 11:21:17",
      "content": "<p>not sure about your bugs with inference, but as I experimented, the Transformer-related methods in Pytorch have some bugs when you use padding masks and attention masks together when you pad your sequences to the left. Try to pad to the right and run it again, it will work. Could be some other issues with your case, but at least, I encountered it.</p>\n<p>Theoretically, you have to build up a transformer from scratch but, practically, I didn't do that. You can find everything you need here <a href=\"https://github.com/seewoo5/KT\" target=\"_blank\">https://github.com/seewoo5/KT</a>. I use the repo with some personal adjustment, it worked for me</p>",
      "rawMarkdown": "not sure about your bugs with inference, but as I experimented, the Transformer-related methods in Pytorch have some bugs when you use padding masks and attention masks together when you pad your sequences to the left. Try to pad to the right and run it again, it will work. Could be some other issues with your case, but at least, I encountered it.\n\nTheoretically, you have to build up a transformer from scratch but, practically, I didn't do that. You can find everything you need here https://github.com/seewoo5/KT. I use the repo with some personal adjustment, it worked for me",
      "votes": null
    },
    {
      "id": "1139496",
      "postDate": "01/05/2021 12:34:51",
      "content": "<p>I think you could try to reduce the size of your model or implement some warming-up at the beginning of training. Alternatively, you could use another learning optimizer: <a href=\"https://github.com/lancopku/AdaMod\" target=\"_blank\">AdaMod</a> works nicely for me.</p>",
      "rawMarkdown": "I think you could try to reduce the size of your model or implement some warming-up at the beginning of training. Alternatively, you could use another learning optimizer: [AdaMod](https://github.com/lancopku/AdaMod) works nicely for me.",
      "votes": null
    },
    {
      "id": "1139530",
      "postDate": "01/05/2021 13:10:45",
      "content": "<p>Thanks for your comments.</p>\n<p>I tried right padding, but the result did not change.<br>\nThe loss is now nan again.</p>\n<p>It looks like there may be another bug, so I will review it again with the repo you sent me.</p>",
      "rawMarkdown": "Thanks for your comments.\n\nI tried right padding, but the result did not change.\nThe loss is now nan again.\n\nIt looks like there may be another bug, so I will review it again with the repo you sent me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1138765,
      "author_name": "konumaru",
      "author_url": "",
      "post_date": "01/05/2021 00:24:48",
      "content": "<p>The following Issue may be the cause.<br>\nIn this case, does it mean that I need to implement Attention by myself?</p>\n<p><a href=\"https://github.com/pytorch/pytorch/issues/24816\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/24816</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1139405,
          "author_name": "shinomoriaoshi",
          "author_url": "",
          "post_date": "01/05/2021 11:21:17",
          "content": "<p>not sure about your bugs with inference, but as I experimented, the Transformer-related methods in Pytorch have some bugs when you use padding masks and attention masks together when you pad your sequences to the left. Try to pad to the right and run it again, it will work. Could be some other issues with your case, but at least, I encountered it.</p>\n<p>Theoretically, you have to build up a transformer from scratch but, practically, I didn't do that. You can find everything you need here <a href=\"https://github.com/seewoo5/KT\" target=\"_blank\">https://github.com/seewoo5/KT</a>. I use the repo with some personal adjustment, it worked for me</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1139530,
          "author_name": "konumaru",
          "author_url": "",
          "post_date": "01/05/2021 13:10:45",
          "content": "<p>Thanks for your comments.</p>\n<p>I tried right padding, but the result did not change.<br>\nThe loss is now nan again.</p>\n<p>It looks like there may be another bug, so I will review it again with the repo you sent me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1139496,
      "author_name": "hannes82",
      "author_url": "",
      "post_date": "01/05/2021 12:34:51",
      "content": "<p>I think you could try to reduce the size of your model or implement some warming-up at the beginning of training. Alternatively, you could use another learning optimizer: <a href=\"https://github.com/lancopku/AdaMod\" target=\"_blank\">AdaMod</a> works nicely for me.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1138764": "I trained a SAINT implemented using Pytorch's Transformer module, but the training data loss stopped around 0.6.\nTo make sure that the loss was not decreasing, I prepared data for evaluation to make sure that the AUC would be around 0.5, and it was still around 0.5.\n\nThe implementation is in the following link.\nI would like to hear the opinions of anyone who is familiar with SAINT and Transformer.\n\nnotebook: https://www.kaggle.com/konumaru/saint-with-pytorch-transformer-module",
    "1138765": "The following Issue may be the cause.\nIn this case, does it mean that I need to implement Attention by myself?\n\nhttps://github.com/pytorch/pytorch/issues/24816",
    "1139405": "not sure about your bugs with inference, but as I experimented, the Transformer-related methods in Pytorch have some bugs when you use padding masks and attention masks together when you pad your sequences to the left. Try to pad to the right and run it again, it will work. Could be some other issues with your case, but at least, I encountered it.\n\nTheoretically, you have to build up a transformer from scratch but, practically, I didn't do that. You can find everything you need here https://github.com/seewoo5/KT. I use the repo with some personal adjustment, it worked for me",
    "1139496": "I think you could try to reduce the size of your model or implement some warming-up at the beginning of training. Alternatively, you could use another learning optimizer: [AdaMod](https://github.com/lancopku/AdaMod) works nicely for me.",
    "1139530": "Thanks for your comments.\n\nI tried right padding, but the result did not change.\nThe loss is now nan again.\n\nIt looks like there may be another bug, so I will review it again with the repo you sent me."
  },
  "source": "meta"
}