{
  "id": 207093,
  "title": "Saint Transformer Overfitting help",
  "url": "/competitions/riiid-test-answer-prediction/discussion/207093",
  "author_name": "",
  "post_date": "2020-12-28T08:09:18.325927400Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi All,</p>\n<p>I have spent weeks learning about transformers and have now I think successfully adapted the code on <a href=\"https://www.tensorflow.org/tutorials/text/transformer\" target=\"_blank\">https://www.tensorflow.org/tutorials/text/transformer</a> to the original Saint paper. I.E. changed the masks to block out future info. I feed the question_id, part  into encoder and the answered_correctly into the decoder. Both have the standard positional encoding applied.</p>\n<p>Thanks for the knowledge sharing especially (<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632)\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632)</a>, I have learnt a lot just through reading through the comments</p>\n<p>I am struggling to get performance anywhere near what has been posted here. Has anyone experienced the same previously?</p>\n<p>The model I train has the below specs</p>\n<p>num_layers = 2<br>\nd_model = 128<br>\ndff = 1024<br>\nnum_heads = 8<br>\ndropout_rate = 0.1<br>\nmax_seq = 100<br>\nbatch_size = 128</p>\n<p>My validation strategy is a 90% train, 10% test split by user.<br>\nI then have a generator (subclass tf.keras.utils.Sequence) for the training set that implements claverru's sampling strategy (Sample N users with replacement, take a random sequence for every id etc)<br>\nFor the test set I evaluate on the last 100 interactions for the above 10% of users.</p>\n<p>My Results are as below, I overfit after only 6 epochs</p>\n<p>EPOCH 1<br>\n---Train---<br>\nEpoch 1 Batch 2650 Loss 0.5815 Accuracy 0.7006<br>\nEpoch 1 Batch 2700 Loss 0.5808 Accuracy 0.7010<br>\nEpoch 1 Batch 2750 Loss 0.5801 Accuracy 0.7013<br>\n---Test---<br>\nThe test loss is 0.5757718086242676<br>\nThe test accuracy is 0.6960946917533875<br>\nThe test AUC is 0.7530733346939087</p>\n<p>EPOCH 2<br>\n---Train---<br>\nEpoch 2 Batch 2650 Loss 0.5423 Accuracy 0.7237<br>\nEpoch 2 Batch 2700 Loss 0.5423 Accuracy 0.7237<br>\nEpoch 2 Batch 2750 Loss 0.5423 Accuracy 0.7237<br>\n---Test---<br>\nThe test loss is 0.5748566389083862<br>\nThe test accuracy is 0.6984637975692749<br>\nThe test AUC is 0.7554817199707031</p>\n<p>EPOCH 3<br>\n---Train---<br>\nEpoch 3 Batch 2650 Loss 0.5388 Accuracy 0.7260<br>\nEpoch 3 Batch 2700 Loss 0.5388 Accuracy 0.7260<br>\nEpoch 3 Batch 2750 Loss 0.5388 Accuracy 0.7260<br>\n---Test---<br>\nThe test loss is 0.5733250975608826<br>\nThe test accuracy is 0.6987239122390747<br>\nThe test AUC is 0.756464958190918</p>\n<p>EPOCH 4<br>\n---Train---<br>\nEpoch 4 Batch 2650 Loss 0.5374 Accuracy 0.7270<br>\nEpoch 4 Batch 2700 Loss 0.5374 Accuracy 0.7271<br>\nEpoch 4 Batch 2750 Loss 0.5373 Accuracy 0.7271</p>\n<p>---Test---<br>\nThe test loss is 0.5702834129333496<br>\nThe test accuracy is 0.7009978294372559<br>\nThe test AUC is 0.7590261697769165</p>\n<p>EPOCH 5<br>\n---Train---<br>\nEpoch 5 Batch 2650 Loss 0.5357 Accuracy 0.7283<br>\nEpoch 5 Batch 2700 Loss 0.5357 Accuracy 0.7283<br>\nEpoch 5 Batch 2750 Loss 0.5357 Accuracy 0.7283</p>\n<p>---Test---<br>\nThe test loss is 0.568941593170166<br>\nThe test accuracy is 0.701901912689209<br>\nThe test AUC is 0.7602261304855347</p>\n<p>EPOCH 5<br>\n---Train---<br>\nEpoch 6 Batch 2650 Loss 0.5352 Accuracy 0.7286<br>\nEpoch 6 Batch 2700 Loss 0.5352 Accuracy 0.7286<br>\nEpoch 6 Batch 2750 Loss 0.5352 Accuracy 0.7286</p>\n<p>---Test---<br>\nThe test loss is 0.5691953897476196<br>\nThe test accuracy is 0.7021515965461731<br>\nThe test AUC is 0.7605013251304626</p>\n<p>Thanks you</p>",
  "messages": [
    {
      "id": "1129290",
      "postDate": "12/28/2020 08:09:18",
      "content": "<p>Hi All,</p>\n<p>I have spent weeks learning about transformers and have now I think successfully adapted the code on <a href=\"https://www.tensorflow.org/tutorials/text/transformer\" target=\"_blank\">https://www.tensorflow.org/tutorials/text/transformer</a> to the original Saint paper. I.E. changed the masks to block out future info. I feed the question_id, part  into encoder and the answered_correctly into the decoder. Both have the standard positional encoding applied.</p>\n<p>Thanks for the knowledge sharing especially (<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632)\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632)</a>, I have learnt a lot just through reading through the comments</p>\n<p>I am struggling to get performance anywhere near what has been posted here. Has anyone experienced the same previously?</p>\n<p>The model I train has the below specs</p>\n<p>num_layers = 2<br>\nd_model = 128<br>\ndff = 1024<br>\nnum_heads = 8<br>\ndropout_rate = 0.1<br>\nmax_seq = 100<br>\nbatch_size = 128</p>\n<p>My validation strategy is a 90% train, 10% test split by user.<br>\nI then have a generator (subclass tf.keras.utils.Sequence) for the training set that implements claverru's sampling strategy (Sample N users with replacement, take a random sequence for every id etc)<br>\nFor the test set I evaluate on the last 100 interactions for the above 10% of users.</p>\n<p>My Results are as below, I overfit after only 6 epochs</p>\n<p>EPOCH 1<br>\n---Train---<br>\nEpoch 1 Batch 2650 Loss 0.5815 Accuracy 0.7006<br>\nEpoch 1 Batch 2700 Loss 0.5808 Accuracy 0.7010<br>\nEpoch 1 Batch 2750 Loss 0.5801 Accuracy 0.7013<br>\n---Test---<br>\nThe test loss is 0.5757718086242676<br>\nThe test accuracy is 0.6960946917533875<br>\nThe test AUC is 0.7530733346939087</p>\n<p>EPOCH 2<br>\n---Train---<br>\nEpoch 2 Batch 2650 Loss 0.5423 Accuracy 0.7237<br>\nEpoch 2 Batch 2700 Loss 0.5423 Accuracy 0.7237<br>\nEpoch 2 Batch 2750 Loss 0.5423 Accuracy 0.7237<br>\n---Test---<br>\nThe test loss is 0.5748566389083862<br>\nThe test accuracy is 0.6984637975692749<br>\nThe test AUC is 0.7554817199707031</p>\n<p>EPOCH 3<br>\n---Train---<br>\nEpoch 3 Batch 2650 Loss 0.5388 Accuracy 0.7260<br>\nEpoch 3 Batch 2700 Loss 0.5388 Accuracy 0.7260<br>\nEpoch 3 Batch 2750 Loss 0.5388 Accuracy 0.7260<br>\n---Test---<br>\nThe test loss is 0.5733250975608826<br>\nThe test accuracy is 0.6987239122390747<br>\nThe test AUC is 0.756464958190918</p>\n<p>EPOCH 4<br>\n---Train---<br>\nEpoch 4 Batch 2650 Loss 0.5374 Accuracy 0.7270<br>\nEpoch 4 Batch 2700 Loss 0.5374 Accuracy 0.7271<br>\nEpoch 4 Batch 2750 Loss 0.5373 Accuracy 0.7271</p>\n<p>---Test---<br>\nThe test loss is 0.5702834129333496<br>\nThe test accuracy is 0.7009978294372559<br>\nThe test AUC is 0.7590261697769165</p>\n<p>EPOCH 5<br>\n---Train---<br>\nEpoch 5 Batch 2650 Loss 0.5357 Accuracy 0.7283<br>\nEpoch 5 Batch 2700 Loss 0.5357 Accuracy 0.7283<br>\nEpoch 5 Batch 2750 Loss 0.5357 Accuracy 0.7283</p>\n<p>---Test---<br>\nThe test loss is 0.568941593170166<br>\nThe test accuracy is 0.701901912689209<br>\nThe test AUC is 0.7602261304855347</p>\n<p>EPOCH 5<br>\n---Train---<br>\nEpoch 6 Batch 2650 Loss 0.5352 Accuracy 0.7286<br>\nEpoch 6 Batch 2700 Loss 0.5352 Accuracy 0.7286<br>\nEpoch 6 Batch 2750 Loss 0.5352 Accuracy 0.7286</p>\n<p>---Test---<br>\nThe test loss is 0.5691953897476196<br>\nThe test accuracy is 0.7021515965461731<br>\nThe test AUC is 0.7605013251304626</p>\n<p>Thanks you</p>",
      "rawMarkdown": "Hi All,\n\nI have spent weeks learning about transformers and have now I think successfully adapted the code on https://www.tensorflow.org/tutorials/text/transformer to the original Saint paper. I.E. changed the masks to block out future info. I feed the question_id, part  into encoder and the answered_correctly into the decoder. Both have the standard positional encoding applied.\n\nThanks for the knowledge sharing especially (https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632), I have learnt a lot just through reading through the comments\n\nI am struggling to get performance anywhere near what has been posted here. Has anyone experienced the same previously?\n\nThe model I train has the below specs\n\nnum_layers = 2\nd_model = 128\ndff = 1024\nnum_heads = 8\ndropout_rate = 0.1\nmax_seq = 100\nbatch_size = 128\n\nMy validation strategy is a 90% train, 10% test split by user.\nI then have a generator (subclass tf.keras.utils.Sequence) for the training set that implements claverru's sampling strategy (Sample N users with replacement, take a random sequence for every id etc)\nFor the test set I evaluate on the last 100 interactions for the above 10% of users.\n\nMy Results are as below, I overfit after only 6 epochs\n\n\nEPOCH 1\n---Train---\nEpoch 1 Batch 2650 Loss 0.5815 Accuracy 0.7006\nEpoch 1 Batch 2700 Loss 0.5808 Accuracy 0.7010\nEpoch 1 Batch 2750 Loss 0.5801 Accuracy 0.7013\n---Test---\nThe test loss is 0.5757718086242676\nThe test accuracy is 0.6960946917533875\nThe test AUC is 0.7530733346939087\n\n\nEPOCH 2\n---Train---\nEpoch 2 Batch 2650 Loss 0.5423 Accuracy 0.7237\nEpoch 2 Batch 2700 Loss 0.5423 Accuracy 0.7237\nEpoch 2 Batch 2750 Loss 0.5423 Accuracy 0.7237\n---Test---\nThe test loss is 0.5748566389083862\nThe test accuracy is 0.6984637975692749\nThe test AUC is 0.7554817199707031\n\nEPOCH 3\n---Train---\nEpoch 3 Batch 2650 Loss 0.5388 Accuracy 0.7260\nEpoch 3 Batch 2700 Loss 0.5388 Accuracy 0.7260\nEpoch 3 Batch 2750 Loss 0.5388 Accuracy 0.7260\n---Test---\nThe test loss is 0.5733250975608826\nThe test accuracy is 0.6987239122390747\nThe test AUC is 0.756464958190918\n\nEPOCH 4\n---Train---\nEpoch 4 Batch 2650 Loss 0.5374 Accuracy 0.7270\nEpoch 4 Batch 2700 Loss 0.5374 Accuracy 0.7271\nEpoch 4 Batch 2750 Loss 0.5373 Accuracy 0.7271\n\n---Test---\nThe test loss is 0.5702834129333496\nThe test accuracy is 0.7009978294372559\nThe test AUC is 0.7590261697769165\n\nEPOCH 5\n---Train---\nEpoch 5 Batch 2650 Loss 0.5357 Accuracy 0.7283\nEpoch 5 Batch 2700 Loss 0.5357 Accuracy 0.7283\nEpoch 5 Batch 2750 Loss 0.5357 Accuracy 0.7283\n\n---Test---\nThe test loss is 0.568941593170166\nThe test accuracy is 0.701901912689209\nThe test AUC is 0.7602261304855347\n\nEPOCH 5\n---Train---\nEpoch 6 Batch 2650 Loss 0.5352 Accuracy 0.7286\nEpoch 6 Batch 2700 Loss 0.5352 Accuracy 0.7286\nEpoch 6 Batch 2750 Loss 0.5352 Accuracy 0.7286\n\n\n---Test---\nThe test loss is 0.5691953897476196\nThe test accuracy is 0.7021515965461731\nThe test AUC is 0.7605013251304626\n\nThanks you",
      "votes": null
    },
    {
      "id": "1129728",
      "postDate": "12/28/2020 14:11:22",
      "content": "<p>Your scores aren't half bad. Especially if you're implementing the original SAINT. Some suggestions:</p>\n<ul>\n<li>dff drop that down a bit</li>\n<li>when you say standard positional encoder, do you mean arange or sin?</li>\n<li>are you using any ts features? your post didn't mention any</li>\n<li>keep in mind that to reach 0.7811 val AUC, the SAINT authors used month, day and hour feature embeddings that you do not have access to</li>\n<li>as I mentioned elsewhere in a post, how you setup you CV can really change your evaluation scores. SAINT authors only tell us test=20% but we have no idea how they split that up. by user? time based? who knows! so take their results with a huge grain of salt.</li>\n</ul>",
      "rawMarkdown": "Your scores aren't half bad. Especially if you're implementing the original SAINT. Some suggestions:\n\n- dff drop that down a bit\n- when you say standard positional encoder, do you mean arange or sin?\n- are you using any ts features? your post didn't mention any\n- keep in mind that to reach 0.7811 val AUC, the SAINT authors used month, day and hour feature embeddings that you do not have access to\n- as I mentioned elsewhere in a post, how you setup you CV can really change your evaluation scores. SAINT authors only tell us test=20% but we have no idea how they split that up. by user? time based? who knows! so take their results with a huge grain of salt.",
      "votes": null
    },
    {
      "id": "1129797",
      "postDate": "12/28/2020 14:56:16",
      "content": "<p>Thanks for your response, much appreciated.</p>\n<p><strong>Update since my original post</strong></p>\n<p>I altered the code so that it doesn't stop training once the test loss stops decreasing. On the 10th epoch I got a test loss of 0.5653 and AUC of  0.764. I saved these weights and submitted and got 0.776 on the LB (really surprised me).</p>\n<p>I will try dropping dff and seeing how that goes.<br>\nFor the positional encoder I am using the sin one.<br>\nNo ts features are used at all, the only columns from the dataset I am using is 'question_id', 'part' and 'answered_correctly'. </p>",
      "rawMarkdown": "Thanks for your response, much appreciated.\n\n**Update since my original post**\n\nI altered the code so that it doesn't stop training once the test loss stops decreasing. On the 10th epoch I got a test loss of 0.5653 and AUC of  0.764. I saved these weights and submitted and got 0.776 on the LB (really surprised me).\n\nI will try dropping dff and seeing how that goes.\nFor the positional encoder I am using the sin one.\nNo ts features are used at all, the only columns from the dataset I am using is 'question_id', 'part' and 'answered_correctly'.",
      "votes": null
    },
    {
      "id": "1129834",
      "postDate": "12/28/2020 15:23:06",
      "content": "<p>Congrats and nice score. I'd imagine a +0.01 jump once you get <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1128932\" target=\"_blank\">ts in</a> as well.</p>",
      "rawMarkdown": "Congrats and nice score. I'd imagine a +0.01 jump once you get [ts in](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1128932) as well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1129728,
      "author_name": "authman",
      "author_url": "",
      "post_date": "12/28/2020 14:11:22",
      "content": "<p>Your scores aren't half bad. Especially if you're implementing the original SAINT. Some suggestions:</p>\n<ul>\n<li>dff drop that down a bit</li>\n<li>when you say standard positional encoder, do you mean arange or sin?</li>\n<li>are you using any ts features? your post didn't mention any</li>\n<li>keep in mind that to reach 0.7811 val AUC, the SAINT authors used month, day and hour feature embeddings that you do not have access to</li>\n<li>as I mentioned elsewhere in a post, how you setup you CV can really change your evaluation scores. SAINT authors only tell us test=20% but we have no idea how they split that up. by user? time based? who knows! so take their results with a huge grain of salt.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1129797,
          "author_name": "darren1515",
          "author_url": "",
          "post_date": "12/28/2020 14:56:16",
          "content": "<p>Thanks for your response, much appreciated.</p>\n<p><strong>Update since my original post</strong></p>\n<p>I altered the code so that it doesn't stop training once the test loss stops decreasing. On the 10th epoch I got a test loss of 0.5653 and AUC of  0.764. I saved these weights and submitted and got 0.776 on the LB (really surprised me).</p>\n<p>I will try dropping dff and seeing how that goes.<br>\nFor the positional encoder I am using the sin one.<br>\nNo ts features are used at all, the only columns from the dataset I am using is 'question_id', 'part' and 'answered_correctly'. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1129834,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/28/2020 15:23:06",
          "content": "<p>Congrats and nice score. I'd imagine a +0.01 jump once you get <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1128932\" target=\"_blank\">ts in</a> as well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1129290": "Hi All,\n\nI have spent weeks learning about transformers and have now I think successfully adapted the code on https://www.tensorflow.org/tutorials/text/transformer to the original Saint paper. I.E. changed the masks to block out future info. I feed the question_id, part  into encoder and the answered_correctly into the decoder. Both have the standard positional encoding applied.\n\nThanks for the knowledge sharing especially (https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632), I have learnt a lot just through reading through the comments\n\nI am struggling to get performance anywhere near what has been posted here. Has anyone experienced the same previously?\n\nThe model I train has the below specs\n\nnum_layers = 2\nd_model = 128\ndff = 1024\nnum_heads = 8\ndropout_rate = 0.1\nmax_seq = 100\nbatch_size = 128\n\nMy validation strategy is a 90% train, 10% test split by user.\nI then have a generator (subclass tf.keras.utils.Sequence) for the training set that implements claverru's sampling strategy (Sample N users with replacement, take a random sequence for every id etc)\nFor the test set I evaluate on the last 100 interactions for the above 10% of users.\n\nMy Results are as below, I overfit after only 6 epochs\n\n\nEPOCH 1\n---Train---\nEpoch 1 Batch 2650 Loss 0.5815 Accuracy 0.7006\nEpoch 1 Batch 2700 Loss 0.5808 Accuracy 0.7010\nEpoch 1 Batch 2750 Loss 0.5801 Accuracy 0.7013\n---Test---\nThe test loss is 0.5757718086242676\nThe test accuracy is 0.6960946917533875\nThe test AUC is 0.7530733346939087\n\n\nEPOCH 2\n---Train---\nEpoch 2 Batch 2650 Loss 0.5423 Accuracy 0.7237\nEpoch 2 Batch 2700 Loss 0.5423 Accuracy 0.7237\nEpoch 2 Batch 2750 Loss 0.5423 Accuracy 0.7237\n---Test---\nThe test loss is 0.5748566389083862\nThe test accuracy is 0.6984637975692749\nThe test AUC is 0.7554817199707031\n\nEPOCH 3\n---Train---\nEpoch 3 Batch 2650 Loss 0.5388 Accuracy 0.7260\nEpoch 3 Batch 2700 Loss 0.5388 Accuracy 0.7260\nEpoch 3 Batch 2750 Loss 0.5388 Accuracy 0.7260\n---Test---\nThe test loss is 0.5733250975608826\nThe test accuracy is 0.6987239122390747\nThe test AUC is 0.756464958190918\n\nEPOCH 4\n---Train---\nEpoch 4 Batch 2650 Loss 0.5374 Accuracy 0.7270\nEpoch 4 Batch 2700 Loss 0.5374 Accuracy 0.7271\nEpoch 4 Batch 2750 Loss 0.5373 Accuracy 0.7271\n\n---Test---\nThe test loss is 0.5702834129333496\nThe test accuracy is 0.7009978294372559\nThe test AUC is 0.7590261697769165\n\nEPOCH 5\n---Train---\nEpoch 5 Batch 2650 Loss 0.5357 Accuracy 0.7283\nEpoch 5 Batch 2700 Loss 0.5357 Accuracy 0.7283\nEpoch 5 Batch 2750 Loss 0.5357 Accuracy 0.7283\n\n---Test---\nThe test loss is 0.568941593170166\nThe test accuracy is 0.701901912689209\nThe test AUC is 0.7602261304855347\n\nEPOCH 5\n---Train---\nEpoch 6 Batch 2650 Loss 0.5352 Accuracy 0.7286\nEpoch 6 Batch 2700 Loss 0.5352 Accuracy 0.7286\nEpoch 6 Batch 2750 Loss 0.5352 Accuracy 0.7286\n\n\n---Test---\nThe test loss is 0.5691953897476196\nThe test accuracy is 0.7021515965461731\nThe test AUC is 0.7605013251304626\n\nThanks you",
    "1129728": "Your scores aren't half bad. Especially if you're implementing the original SAINT. Some suggestions:\n\n- dff drop that down a bit\n- when you say standard positional encoder, do you mean arange or sin?\n- are you using any ts features? your post didn't mention any\n- keep in mind that to reach 0.7811 val AUC, the SAINT authors used month, day and hour feature embeddings that you do not have access to\n- as I mentioned elsewhere in a post, how you setup you CV can really change your evaluation scores. SAINT authors only tell us test=20% but we have no idea how they split that up. by user? time based? who knows! so take their results with a huge grain of salt.",
    "1129797": "Thanks for your response, much appreciated.\n\n**Update since my original post**\n\nI altered the code so that it doesn't stop training once the test loss stops decreasing. On the 10th epoch I got a test loss of 0.5653 and AUC of  0.764. I saved these weights and submitted and got 0.776 on the LB (really surprised me).\n\nI will try dropping dff and seeing how that goes.\nFor the positional encoder I am using the sin one.\nNo ts features are used at all, the only columns from the dataset I am using is 'question_id', 'part' and 'answered_correctly'.",
    "1129834": "Congrats and nice score. I'd imagine a +0.01 jump once you get [ts in](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195632#1128932) as well."
  },
  "source": "meta"
}