{
  "id": 207135,
  "title": "About d_model and embedding's output dimensions in saint+",
  "url": "/competitions/riiid-test-answer-prediction/discussion/207135",
  "author_name": "",
  "post_date": "2020-12-28T11:59:18.117983600Z",
  "votes": 8,
  "comment_count": 33,
  "views": 0,
  "content": "<p>when i use big d_model ,like 512,the result is bad.because i set d_model=embedding's output dimensions and use addition to combine all the embedding's outputs.so I want to seek some advice.Is there anyone help me，thanks</p>",
  "messages": [
    {
      "id": "1129515",
      "postDate": "12/28/2020 11:59:18",
      "content": "<p>when i use big d_model ,like 512,the result is bad.because i set d_model=embedding's output dimensions and use addition to combine all the embedding's outputs.so I want to seek some advice.Is there anyone help me，thanks</p>",
      "rawMarkdown": "when i use big d_model ,like 512,the result is bad.because i set d_model=embedding's output dimensions and use addition to combine all the embedding's outputs.so I want to seek some advice.Is there anyone help me，thanks",
      "votes": null
    },
    {
      "id": "1130347",
      "postDate": "12/29/2020 00:37:42",
      "content": "<p>I can confirm this, model performance drops drastically for d_model 512</p>",
      "rawMarkdown": "I can confirm this, model performance drops drastically for d_model 512",
      "votes": null
    },
    {
      "id": "1130366",
      "postDate": "12/29/2020 01:05:17",
      "content": "<p>Well, for me, setting d_model 512 is much better than d_model 64, 128, but similar to d_model 256.</p>",
      "rawMarkdown": "Well, for me, setting d_model 512 is much better than d_model 64, 128, but similar to d_model 256.",
      "votes": null
    },
    {
      "id": "1130377",
      "postDate": "12/29/2020 01:25:46",
      "content": "<p>How do you manage to train such a monster, 512 is 40M parameters. </p>\n<p>For my case I can reach 0.788 with d_model 64,  and 0.795 with d_model 128 , seeing 512 underperform demotivated me from testing 256.</p>",
      "rawMarkdown": "How do you manage to train such a monster, 512 is 40M parameters. \n\nFor my case I can reach 0.788 with d_model 64,  and 0.795 with d_model 128 , seeing 512 underperform demotivated me from testing 256.",
      "votes": null
    },
    {
      "id": "1130384",
      "postDate": "12/29/2020 01:35:29",
      "content": "<p>yes,I can reach 0.790 with d_model 128，0.605 with d_model 512..😹😹😹</p>",
      "rawMarkdown": "yes,I can reach 0.790 with d_model 128，0.605 with d_model 512..😹😹😹",
      "votes": null
    },
    {
      "id": "1130385",
      "postDate": "12/29/2020 01:46:28",
      "content": "<p>The problem maybe is not about d_model, it's a mixture of data sampling, d_model, number of layers, dropout and so on… So it's hard to say…</p>",
      "rawMarkdown": "The problem maybe is not about d_model, it's a mixture of data sampling, d_model, number of layers, dropout and so on... So it's hard to say...",
      "votes": null
    },
    {
      "id": "1130386",
      "postDate": "12/29/2020 01:47:34",
      "content": "<p>oh,How big is your embedding's output dimensions.I only get 0.605 with d_model 512. My all embedding's output dimensions equal to d_model,I doubt it is caused by this.</p>",
      "rawMarkdown": "oh,How big is your embedding's output dimensions.I only get 0.605 with d_model 512. My all embedding's output dimensions equal to d_model,I doubt it is caused by this.",
      "votes": null
    },
    {
      "id": "1130388",
      "postDate": "12/29/2020 01:50:46",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> If you don't mind, can you disclose if your LB score is single model or an ensemble</p>",
      "rawMarkdown": "lihaorocky If you don't mind, can you disclose if your LB score is single model or an ensemble",
      "votes": null
    },
    {
      "id": "1130396",
      "postDate": "12/29/2020 02:04:51",
      "content": "<p>single model</p>",
      "rawMarkdown": "single model",
      "votes": null
    },
    {
      "id": "1130425",
      "postDate": "12/29/2020 03:25:59",
      "content": "<p>You need to change lr when d_model is 512.</p>",
      "rawMarkdown": "You need to change lr when d_model is 512.",
      "votes": null
    },
    {
      "id": "1130439",
      "postDate": "12/29/2020 03:58:13",
      "content": "<p>I used same LR specified in the paper though</p>",
      "rawMarkdown": "I used same LR specified in the paper though",
      "votes": null
    },
    {
      "id": "1130601",
      "postDate": "12/29/2020 06:52:59",
      "content": "<p>thx！！we solve it👍</p>",
      "rawMarkdown": "thx！！we solve it👍",
      "votes": null
    },
    {
      "id": "1131074",
      "postDate": "12/29/2020 14:29:37",
      "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> , can you let me know that for LB of 0.80, was the local CV also around 0.80?</p>",
      "rawMarkdown": "lihaorocky , can you let me know that for LB of 0.80, was the local CV also around 0.80?",
      "votes": null
    },
    {
      "id": "1131196",
      "postDate": "12/29/2020 15:32:14",
      "content": "<p>This discussion is loosely related to one part of your query: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584</a> if you are interested</p>",
      "rawMarkdown": "This discussion is loosely related to one part of your query: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584 if you are interested",
      "votes": null
    },
    {
      "id": "1136919",
      "postDate": "01/03/2021 14:33:57",
      "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a>  using flat LR with reduction works better or  Varrying LR  for part of training</p>",
      "rawMarkdown": "m10515009  using flat LR with reduction works better or  Varrying LR  for part of training",
      "votes": null
    },
    {
      "id": "1136945",
      "postDate": "01/03/2021 14:57:23",
      "content": "<p>Is the purpose of warmup to allow the larger transformer models to get into a good parameterization so that they can train at the desired, e.g. flatline lr with optional reduceonplateau or something else?</p>",
      "rawMarkdown": "Is the purpose of warmup to allow the larger transformer models to get into a good parameterization so that they can train at the desired, e.g. flatline lr with optional reduceonplateau or something else?",
      "votes": null
    },
    {
      "id": "1136947",
      "postDate": "01/03/2021 14:59:17",
      "content": "<p>I didn't use that.<br>\nonly use:<br>\n<code>optimizer = torch.optim.AdamW(model.parameters(), lr=5e-4)</code></p>",
      "rawMarkdown": "I didn't use that.\nonly use:\n`optimizer = torch.optim.AdamW(model.parameters(), lr=5e-4)`",
      "votes": null
    },
    {
      "id": "1137021",
      "postDate": "01/03/2021 15:52:55",
      "content": "<p>With such low lr, how many epochs would it take to converge? Mine takes around 60 - 70 <a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a> </p>",
      "rawMarkdown": "With such low lr, how many epochs would it take to converge? Mine takes around 60 - 70 @m10515009",
      "votes": null
    },
    {
      "id": "1137028",
      "postDate": "01/03/2021 15:58:15",
      "content": "<p>It's about 20 epochs with d_model 512, layer 2, batch size 256.</p>\n<pre><code>epoch - 0 train_loss - 0.4516 train_auc - 0.7477 val_loss - 0.3475 val_auc - 0.7776 time=833.81s\nepoch - 1 train_loss - 0.4281 train_auc - 0.7815 val_loss - 0.3441 val_auc - 0.7831 time=834.06s\nepoch - 2 train_loss - 0.4240 train_auc - 0.7871 val_loss - 0.3421 val_auc - 0.7878 time=840.03s\nepoch - 3 train_loss - 0.4216 train_auc - 0.7904 val_loss - 0.3406 val_auc - 0.7893 time=833.13s\nepoch - 4 train_loss - 0.4200 train_auc - 0.7928 val_loss - 0.3401 val_auc - 0.7917 time=832.33s\nepoch - 5 train_loss - 0.4187 train_auc - 0.7946 val_loss - 0.3398 val_auc - 0.7927 time=832.43s\nepoch - 6 train_loss - 0.4177 train_auc - 0.7962 val_loss - 0.3383 val_auc - 0.7942 time=847.83s\nepoch - 7 train_loss - 0.4167 train_auc - 0.7976 val_loss - 0.3378 val_auc - 0.7948 time=843.36s\nepoch - 8 train_loss - 0.4159 train_auc - 0.7989 val_loss - 0.3373 val_auc - 0.7952 time=842.63s\nepoch - 9 train_loss - 0.4151 train_auc - 0.8002 val_loss - 0.3368 val_auc - 0.7959 time=862.17s\nepoch - 10 train_loss - 0.4143 train_auc - 0.8013 val_loss - 0.3365 val_auc - 0.7967 time=857.08s\nepoch - 11 train_loss - 0.4135 train_auc - 0.8023 val_loss - 0.3359 val_auc - 0.7983 time=842.30s\nepoch - 12 train_loss - 0.4128 train_auc - 0.8033 val_loss - 0.3356 val_auc - 0.7982 time=853.73s\nepoch - 13 train_loss - 0.4121 train_auc - 0.8043 val_loss - 0.3358 val_auc - 0.7989 time=843.08s\nepoch - 14 train_loss - 0.4115 train_auc - 0.8051 val_loss - 0.3357 val_auc - 0.7985 time=843.85s\nepoch - 15 train_loss - 0.4109 train_auc - 0.8060 val_loss - 0.3357 val_auc - 0.7982 time=886.01s\nepoch - 16 train_loss - 0.4103 train_auc - 0.8067 val_loss - 0.3355 val_auc - 0.7992 time=889.84s\nepoch - 17 train_loss - 0.4098 train_auc - 0.8071 val_loss - 0.3356 val_auc - 0.7995 time=888.07s\nepoch - 18 train_loss - 0.4092 train_auc - 0.8078 val_loss - 0.3353 val_auc - 0.7995 time=890.12s\nepoch - 19 train_loss - 0.4087 train_auc - 0.8088 val_loss - 0.3353 val_auc - 0.8001 time=889.51s\nepoch - 20 train_loss - 0.4081 train_auc - 0.8095 val_loss - 0.3358 val_auc - 0.8000 time=887.66s\nepoch - 21 train_loss - 0.4076 train_auc - 0.8101 val_loss - 0.3359 val_auc - 0.7994 time=888.36s\nepoch - 22 train_loss - 0.4070 train_auc - 0.8106 val_loss - 0.3362 val_auc - 0.7999 time=887.26s\nepoch - 23 train_loss - 0.4064 train_auc - 0.8115 val_loss - 0.3363 val_auc - 0.7996 time=889.62s\nepoch - 24 train_loss - 0.4058 train_auc - 0.8119 val_loss - 0.3365 val_auc - 0.7989 time=889.73s\n</code></pre>",
      "rawMarkdown": "It's about 20 epochs with d_model 512, layer 2, batch size 256.\n\n```\nepoch - 0 train_loss - 0.4516 train_auc - 0.7477 val_loss - 0.3475 val_auc - 0.7776 time=833.81s\nepoch - 1 train_loss - 0.4281 train_auc - 0.7815 val_loss - 0.3441 val_auc - 0.7831 time=834.06s\nepoch - 2 train_loss - 0.4240 train_auc - 0.7871 val_loss - 0.3421 val_auc - 0.7878 time=840.03s\nepoch - 3 train_loss - 0.4216 train_auc - 0.7904 val_loss - 0.3406 val_auc - 0.7893 time=833.13s\nepoch - 4 train_loss - 0.4200 train_auc - 0.7928 val_loss - 0.3401 val_auc - 0.7917 time=832.33s\nepoch - 5 train_loss - 0.4187 train_auc - 0.7946 val_loss - 0.3398 val_auc - 0.7927 time=832.43s\nepoch - 6 train_loss - 0.4177 train_auc - 0.7962 val_loss - 0.3383 val_auc - 0.7942 time=847.83s\nepoch - 7 train_loss - 0.4167 train_auc - 0.7976 val_loss - 0.3378 val_auc - 0.7948 time=843.36s\nepoch - 8 train_loss - 0.4159 train_auc - 0.7989 val_loss - 0.3373 val_auc - 0.7952 time=842.63s\nepoch - 9 train_loss - 0.4151 train_auc - 0.8002 val_loss - 0.3368 val_auc - 0.7959 time=862.17s\nepoch - 10 train_loss - 0.4143 train_auc - 0.8013 val_loss - 0.3365 val_auc - 0.7967 time=857.08s\nepoch - 11 train_loss - 0.4135 train_auc - 0.8023 val_loss - 0.3359 val_auc - 0.7983 time=842.30s\nepoch - 12 train_loss - 0.4128 train_auc - 0.8033 val_loss - 0.3356 val_auc - 0.7982 time=853.73s\nepoch - 13 train_loss - 0.4121 train_auc - 0.8043 val_loss - 0.3358 val_auc - 0.7989 time=843.08s\nepoch - 14 train_loss - 0.4115 train_auc - 0.8051 val_loss - 0.3357 val_auc - 0.7985 time=843.85s\nepoch - 15 train_loss - 0.4109 train_auc - 0.8060 val_loss - 0.3357 val_auc - 0.7982 time=886.01s\nepoch - 16 train_loss - 0.4103 train_auc - 0.8067 val_loss - 0.3355 val_auc - 0.7992 time=889.84s\nepoch - 17 train_loss - 0.4098 train_auc - 0.8071 val_loss - 0.3356 val_auc - 0.7995 time=888.07s\nepoch - 18 train_loss - 0.4092 train_auc - 0.8078 val_loss - 0.3353 val_auc - 0.7995 time=890.12s\nepoch - 19 train_loss - 0.4087 train_auc - 0.8088 val_loss - 0.3353 val_auc - 0.8001 time=889.51s\nepoch - 20 train_loss - 0.4081 train_auc - 0.8095 val_loss - 0.3358 val_auc - 0.8000 time=887.66s\nepoch - 21 train_loss - 0.4076 train_auc - 0.8101 val_loss - 0.3359 val_auc - 0.7994 time=888.36s\nepoch - 22 train_loss - 0.4070 train_auc - 0.8106 val_loss - 0.3362 val_auc - 0.7999 time=887.26s\nepoch - 23 train_loss - 0.4064 train_auc - 0.8115 val_loss - 0.3363 val_auc - 0.7996 time=889.62s\nepoch - 24 train_loss - 0.4058 train_auc - 0.8119 val_loss - 0.3365 val_auc - 0.7989 time=889.73s\n```",
      "votes": null
    },
    {
      "id": "1137037",
      "postDate": "01/03/2021 16:05:54",
      "content": "<p>Interesting, thank you, nice model you got! <a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a> Best of luck</p>",
      "rawMarkdown": "Interesting, thank you, nice model you got! @m10515009 Best of luck",
      "votes": null
    },
    {
      "id": "1137042",
      "postDate": "01/03/2021 16:15:50",
      "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a>  are you using cv strategy of tito or default last 100 ?</p>",
      "rawMarkdown": "m10515009  are you using cv strategy of tito or default last 100 ?",
      "votes": null
    },
    {
      "id": "1137050",
      "postDate": "01/03/2021 16:24:02",
      "content": "<p>I used CV strategy from tito.</p>",
      "rawMarkdown": "I used CV strategy from tito.",
      "votes": null
    },
    {
      "id": "1137091",
      "postDate": "01/03/2021 17:08:55",
      "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a> what is ur dropout?<br>\nI got my access issues fixed now, but it is too late to move to Saint. So stuck with SAKT. Not sure if same params work for SAKT though..<br>\nbut great to see u with high score..hope u move into gold zone soon</p>",
      "rawMarkdown": "m10515009 what is ur dropout?\nI got my access issues fixed now, but it is too late to move to Saint. So stuck with SAKT. Not sure if same params work for SAKT though..\nbut great to see u with high score..hope u move into gold zone soon",
      "votes": null
    },
    {
      "id": "1137172",
      "postDate": "01/03/2021 18:26:00",
      "content": "<p>what's your length of each sample feed to saint plus plus? <a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a></p>",
      "rawMarkdown": "what's your length of each sample feed to saint plus plus? @m10515009",
      "votes": null
    },
    {
      "id": "1137483",
      "postDate": "01/04/2021 01:30:17",
      "content": "<p><a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> I using dropout rate 0.1 in both SAKT and SAINT.<br>\n<a href=\"https://www.kaggle.com/cswwp347724\" target=\"_blank\">@cswwp347724</a> 100, I have tried 160 but not improvement.</p>",
      "rawMarkdown": "allohvk I using dropout rate 0.1 in both SAKT and SAINT.\n@cswwp347724 100, I have tried 160 but not improvement.",
      "votes": null
    },
    {
      "id": "1137489",
      "postDate": "01/04/2021 01:36:49",
      "content": "<p>Guys, is there a SAINT ++ I am unaware of ?  I only know SAINT + . Can some one link me the paper?</p>",
      "rawMarkdown": "Guys, is there a SAINT ++ I am unaware of ?  I only know SAINT + . Can some one link me the paper?",
      "votes": null
    },
    {
      "id": "1137523",
      "postDate": "01/04/2021 02:45:49",
      "content": "<p>I only know SAINT+ too.😂</p>",
      "rawMarkdown": "I only know SAINT+ too.😂",
      "votes": null
    },
    {
      "id": "1137528",
      "postDate": "01/04/2021 02:55:09",
      "content": "<p>SAINT ++ -&gt; Kaggle folks are going to give one because as per paper they have AUC of .79x and top folks have AUC of .81x and SAKT itself can reach close to .79x it seems as others have reported here and there.</p>",
      "rawMarkdown": "SAINT ++ -> Kaggle folks are going to give one because as per paper they have AUC of .79x and top folks have AUC of .81x and SAKT itself can reach close to .79x it seems as others have reported here and there.",
      "votes": null
    },
    {
      "id": "1137583",
      "postDate": "01/04/2021 04:33:01",
      "content": "<p>:( <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> so u r saying if I use SAKT the highest I can hope to reach is 79x..<br>\nwould u recommend any good saint+ implementation which I can quickly fine-tune. Unfortunately I can only spend a couple of hours each night as my job keeps me quite occupied</p>",
      "rawMarkdown": ":( @adityaecdrid so u r saying if I use SAKT the highest I can hope to reach is 79x..\nwould u recommend any good saint+ implementation which I can quickly fine-tune. Unfortunately I can only spend a couple of hours each night as my job keeps me quite occupied",
      "votes": null
    },
    {
      "id": "1137590",
      "postDate": "01/04/2021 04:44:27",
      "content": "<p>I wish i knew one and had one; I couldn't get SAINT as strong as others, so the current score doesn't have a SAINT/SAINT+.</p>",
      "rawMarkdown": "I wish i knew one and had one; I couldn't get SAINT as strong as others, so the current score doesn't have a SAINT/SAINT+.",
      "votes": null
    },
    {
      "id": "1137598",
      "postDate": "01/04/2021 04:57:21",
      "content": "<p><a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> SAKT reaching 0.79 is difficult for me.<br>\nThe history of my model is as follows.</p>\n<ol>\n<li>SAKT LB 0.776</li>\n<li>SAINT LB 0.784</li>\n<li>SAINT+ LB 0.792</li>\n</ol>\n<p>In my opinion, no one will public successful SAINT/SAINT+  training notebook, because it will influence LB very much.</p>",
      "rawMarkdown": "allohvk SAKT reaching 0.79 is difficult for me.\nThe history of my model is as follows.\n\n1.  SAKT LB 0.776\n2. SAINT LB 0.784\n3. SAINT+ LB 0.792\n\nIn my opinion, no one will public successful SAINT/SAINT+  training notebook, because it will influence LB very much.",
      "votes": null
    },
    {
      "id": "1137625",
      "postDate": "01/04/2021 05:25:27",
      "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a> thanks for the info. So it is a solid 1.6% lead for Saint over sakt<br>\nI am trying a customised version of stack. I will try to reach 80 with customised SAKT and an ensemble</p>\n<p>if time permits, last day (Thursday) I will take a shot at saint. </p>",
      "rawMarkdown": "m10515009 thanks for the info. So it is a solid 1.6% lead for Saint over sakt\nI am trying a customised version of stack. I will try to reach 80 with customised SAKT and an ensemble\n\nif time permits, last day (Thursday) I will take a shot at saint.",
      "votes": null
    },
    {
      "id": "1137865",
      "postDate": "01/04/2021 08:29:02",
      "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> if it helps pl check the Saint paper. They have some alternate models more closer to SAKT. If like me you are already invested in SAKT, you may want to check out those models (UTMTI, LTMTI, SSAKT) etc. The other point to note is that unlike Saint, these models underperform at 512 and may or may not perform at 256. Likewise they are better at layer=2 or 3 at best.</p>",
      "rawMarkdown": "adityaecdrid if it helps pl check the Saint paper. They have some alternate models more closer to SAKT. If like me you are already invested in SAKT, you may want to check out those models (UTMTI, LTMTI, SSAKT) etc. The other point to note is that unlike Saint, these models underperform at 512 and may or may not perform at 256. Likewise they are better at layer=2 or 3 at best.",
      "votes": null
    },
    {
      "id": "1137869",
      "postDate": "01/04/2021 08:30:39",
      "content": "<p>most likely the clean segregation of exercise and response in the Saint architecture in some way prevents overfitting and model performs better with more dimensions and layers. </p>",
      "rawMarkdown": "most likely the clean segregation of exercise and response in the Saint architecture in some way prevents overfitting and model performs better with more dimensions and layers.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1130347,
      "author_name": "abdessalemboukil",
      "author_url": "",
      "post_date": "12/29/2020 00:37:42",
      "content": "<p>I can confirm this, model performance drops drastically for d_model 512</p>",
      "votes": null,
      "replies": [
        {
          "id": 1130384,
          "author_name": "hu5851447",
          "author_url": "",
          "post_date": "12/29/2020 01:35:29",
          "content": "<p>yes,I can reach 0.790 with d_model 128，0.605 with d_model 512..😹😹😹</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1130425,
          "author_name": "m10515009",
          "author_url": "",
          "post_date": "12/29/2020 03:25:59",
          "content": "<p>You need to change lr when d_model is 512.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1130439,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "12/29/2020 03:58:13",
          "content": "<p>I used same LR specified in the paper though</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1130601,
          "author_name": "lzcabc123456",
          "author_url": "",
          "post_date": "12/29/2020 06:52:59",
          "content": "<p>thx！！we solve it👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1136919,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/03/2021 14:33:57",
          "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a>  using flat LR with reduction works better or  Varrying LR  for part of training</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1136947,
          "author_name": "m10515009",
          "author_url": "",
          "post_date": "01/03/2021 14:59:17",
          "content": "<p>I didn't use that.<br>\nonly use:<br>\n<code>optimizer = torch.optim.AdamW(model.parameters(), lr=5e-4)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137021,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "01/03/2021 15:52:55",
          "content": "<p>With such low lr, how many epochs would it take to converge? Mine takes around 60 - 70 <a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137028,
          "author_name": "m10515009",
          "author_url": "",
          "post_date": "01/03/2021 15:58:15",
          "content": "<p>It's about 20 epochs with d_model 512, layer 2, batch size 256.</p>\n<pre><code>epoch - 0 train_loss - 0.4516 train_auc - 0.7477 val_loss - 0.3475 val_auc - 0.7776 time=833.81s\nepoch - 1 train_loss - 0.4281 train_auc - 0.7815 val_loss - 0.3441 val_auc - 0.7831 time=834.06s\nepoch - 2 train_loss - 0.4240 train_auc - 0.7871 val_loss - 0.3421 val_auc - 0.7878 time=840.03s\nepoch - 3 train_loss - 0.4216 train_auc - 0.7904 val_loss - 0.3406 val_auc - 0.7893 time=833.13s\nepoch - 4 train_loss - 0.4200 train_auc - 0.7928 val_loss - 0.3401 val_auc - 0.7917 time=832.33s\nepoch - 5 train_loss - 0.4187 train_auc - 0.7946 val_loss - 0.3398 val_auc - 0.7927 time=832.43s\nepoch - 6 train_loss - 0.4177 train_auc - 0.7962 val_loss - 0.3383 val_auc - 0.7942 time=847.83s\nepoch - 7 train_loss - 0.4167 train_auc - 0.7976 val_loss - 0.3378 val_auc - 0.7948 time=843.36s\nepoch - 8 train_loss - 0.4159 train_auc - 0.7989 val_loss - 0.3373 val_auc - 0.7952 time=842.63s\nepoch - 9 train_loss - 0.4151 train_auc - 0.8002 val_loss - 0.3368 val_auc - 0.7959 time=862.17s\nepoch - 10 train_loss - 0.4143 train_auc - 0.8013 val_loss - 0.3365 val_auc - 0.7967 time=857.08s\nepoch - 11 train_loss - 0.4135 train_auc - 0.8023 val_loss - 0.3359 val_auc - 0.7983 time=842.30s\nepoch - 12 train_loss - 0.4128 train_auc - 0.8033 val_loss - 0.3356 val_auc - 0.7982 time=853.73s\nepoch - 13 train_loss - 0.4121 train_auc - 0.8043 val_loss - 0.3358 val_auc - 0.7989 time=843.08s\nepoch - 14 train_loss - 0.4115 train_auc - 0.8051 val_loss - 0.3357 val_auc - 0.7985 time=843.85s\nepoch - 15 train_loss - 0.4109 train_auc - 0.8060 val_loss - 0.3357 val_auc - 0.7982 time=886.01s\nepoch - 16 train_loss - 0.4103 train_auc - 0.8067 val_loss - 0.3355 val_auc - 0.7992 time=889.84s\nepoch - 17 train_loss - 0.4098 train_auc - 0.8071 val_loss - 0.3356 val_auc - 0.7995 time=888.07s\nepoch - 18 train_loss - 0.4092 train_auc - 0.8078 val_loss - 0.3353 val_auc - 0.7995 time=890.12s\nepoch - 19 train_loss - 0.4087 train_auc - 0.8088 val_loss - 0.3353 val_auc - 0.8001 time=889.51s\nepoch - 20 train_loss - 0.4081 train_auc - 0.8095 val_loss - 0.3358 val_auc - 0.8000 time=887.66s\nepoch - 21 train_loss - 0.4076 train_auc - 0.8101 val_loss - 0.3359 val_auc - 0.7994 time=888.36s\nepoch - 22 train_loss - 0.4070 train_auc - 0.8106 val_loss - 0.3362 val_auc - 0.7999 time=887.26s\nepoch - 23 train_loss - 0.4064 train_auc - 0.8115 val_loss - 0.3363 val_auc - 0.7996 time=889.62s\nepoch - 24 train_loss - 0.4058 train_auc - 0.8119 val_loss - 0.3365 val_auc - 0.7989 time=889.73s\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137037,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "01/03/2021 16:05:54",
          "content": "<p>Interesting, thank you, nice model you got! <a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a> Best of luck</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137042,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/03/2021 16:15:50",
          "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a>  are you using cv strategy of tito or default last 100 ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137050,
          "author_name": "m10515009",
          "author_url": "",
          "post_date": "01/03/2021 16:24:02",
          "content": "<p>I used CV strategy from tito.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137091,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "01/03/2021 17:08:55",
          "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a> what is ur dropout?<br>\nI got my access issues fixed now, but it is too late to move to Saint. So stuck with SAKT. Not sure if same params work for SAKT though..<br>\nbut great to see u with high score..hope u move into gold zone soon</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137172,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "01/03/2021 18:26:00",
          "content": "<p>what's your length of each sample feed to saint plus plus? <a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137483,
          "author_name": "m10515009",
          "author_url": "",
          "post_date": "01/04/2021 01:30:17",
          "content": "<p><a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> I using dropout rate 0.1 in both SAKT and SAINT.<br>\n<a href=\"https://www.kaggle.com/cswwp347724\" target=\"_blank\">@cswwp347724</a> 100, I have tried 160 but not improvement.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137489,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "01/04/2021 01:36:49",
          "content": "<p>Guys, is there a SAINT ++ I am unaware of ?  I only know SAINT + . Can some one link me the paper?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137523,
          "author_name": "m10515009",
          "author_url": "",
          "post_date": "01/04/2021 02:45:49",
          "content": "<p>I only know SAINT+ too.😂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137528,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "01/04/2021 02:55:09",
          "content": "<p>SAINT ++ -&gt; Kaggle folks are going to give one because as per paper they have AUC of .79x and top folks have AUC of .81x and SAKT itself can reach close to .79x it seems as others have reported here and there.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137583,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "01/04/2021 04:33:01",
          "content": "<p>:( <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> so u r saying if I use SAKT the highest I can hope to reach is 79x..<br>\nwould u recommend any good saint+ implementation which I can quickly fine-tune. Unfortunately I can only spend a couple of hours each night as my job keeps me quite occupied</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137590,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "01/04/2021 04:44:27",
          "content": "<p>I wish i knew one and had one; I couldn't get SAINT as strong as others, so the current score doesn't have a SAINT/SAINT+.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137598,
          "author_name": "m10515009",
          "author_url": "",
          "post_date": "01/04/2021 04:57:21",
          "content": "<p><a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> SAKT reaching 0.79 is difficult for me.<br>\nThe history of my model is as follows.</p>\n<ol>\n<li>SAKT LB 0.776</li>\n<li>SAINT LB 0.784</li>\n<li>SAINT+ LB 0.792</li>\n</ol>\n<p>In my opinion, no one will public successful SAINT/SAINT+  training notebook, because it will influence LB very much.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137625,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "01/04/2021 05:25:27",
          "content": "<p><a href=\"https://www.kaggle.com/m10515009\" target=\"_blank\">@m10515009</a> thanks for the info. So it is a solid 1.6% lead for Saint over sakt<br>\nI am trying a customised version of stack. I will try to reach 80 with customised SAKT and an ensemble</p>\n<p>if time permits, last day (Thursday) I will take a shot at saint. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137865,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "01/04/2021 08:29:02",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> if it helps pl check the Saint paper. They have some alternate models more closer to SAKT. If like me you are already invested in SAKT, you may want to check out those models (UTMTI, LTMTI, SSAKT) etc. The other point to note is that unlike Saint, these models underperform at 512 and may or may not perform at 256. Likewise they are better at layer=2 or 3 at best.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137869,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "01/04/2021 08:30:39",
          "content": "<p>most likely the clean segregation of exercise and response in the Saint architecture in some way prevents overfitting and model performs better with more dimensions and layers. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1130366,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "12/29/2020 01:05:17",
      "content": "<p>Well, for me, setting d_model 512 is much better than d_model 64, 128, but similar to d_model 256.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1130377,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "12/29/2020 01:25:46",
          "content": "<p>How do you manage to train such a monster, 512 is 40M parameters. </p>\n<p>For my case I can reach 0.788 with d_model 64,  and 0.795 with d_model 128 , seeing 512 underperform demotivated me from testing 256.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1130385,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "12/29/2020 01:46:28",
          "content": "<p>The problem maybe is not about d_model, it's a mixture of data sampling, d_model, number of layers, dropout and so on… So it's hard to say…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1130386,
          "author_name": "hu5851447",
          "author_url": "",
          "post_date": "12/29/2020 01:47:34",
          "content": "<p>oh,How big is your embedding's output dimensions.I only get 0.605 with d_model 512. My all embedding's output dimensions equal to d_model,I doubt it is caused by this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1130388,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "12/29/2020 01:50:46",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> If you don't mind, can you disclose if your LB score is single model or an ensemble</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1130396,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "12/29/2020 02:04:51",
          "content": "<p>single model</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131074,
          "author_name": "anuragtr",
          "author_url": "",
          "post_date": "12/29/2020 14:29:37",
          "content": "<p><a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> , can you let me know that for LB of 0.80, was the local CV also around 0.80?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1131196,
      "author_name": "allohvk",
      "author_url": "",
      "post_date": "12/29/2020 15:32:14",
      "content": "<p>This discussion is loosely related to one part of your query: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584</a> if you are interested</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1136945,
      "author_name": "authman",
      "author_url": "",
      "post_date": "01/03/2021 14:57:23",
      "content": "<p>Is the purpose of warmup to allow the larger transformer models to get into a good parameterization so that they can train at the desired, e.g. flatline lr with optional reduceonplateau or something else?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1129515": "when i use big d_model ,like 512,the result is bad.because i set d_model=embedding's output dimensions and use addition to combine all the embedding's outputs.so I want to seek some advice.Is there anyone help me，thanks",
    "1130347": "I can confirm this, model performance drops drastically for d_model 512",
    "1130366": "Well, for me, setting d_model 512 is much better than d_model 64, 128, but similar to d_model 256.",
    "1130377": "How do you manage to train such a monster, 512 is 40M parameters. \n\nFor my case I can reach 0.788 with d_model 64,  and 0.795 with d_model 128 , seeing 512 underperform demotivated me from testing 256.",
    "1130384": "yes,I can reach 0.790 with d_model 128，0.605 with d_model 512..😹😹😹",
    "1130385": "The problem maybe is not about d_model, it's a mixture of data sampling, d_model, number of layers, dropout and so on... So it's hard to say...",
    "1130386": "oh,How big is your embedding's output dimensions.I only get 0.605 with d_model 512. My all embedding's output dimensions equal to d_model,I doubt it is caused by this.",
    "1130388": "lihaorocky If you don't mind, can you disclose if your LB score is single model or an ensemble",
    "1130396": "single model",
    "1130425": "You need to change lr when d_model is 512.",
    "1130439": "I used same LR specified in the paper though",
    "1130601": "thx！！we solve it👍",
    "1131074": "lihaorocky , can you let me know that for LB of 0.80, was the local CV also around 0.80?",
    "1131196": "This discussion is loosely related to one part of your query: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206584 if you are interested",
    "1136919": "m10515009  using flat LR with reduction works better or  Varrying LR  for part of training",
    "1136945": "Is the purpose of warmup to allow the larger transformer models to get into a good parameterization so that they can train at the desired, e.g. flatline lr with optional reduceonplateau or something else?",
    "1136947": "I didn't use that.\nonly use:\n`optimizer = torch.optim.AdamW(model.parameters(), lr=5e-4)`",
    "1137021": "With such low lr, how many epochs would it take to converge? Mine takes around 60 - 70 @m10515009",
    "1137028": "It's about 20 epochs with d_model 512, layer 2, batch size 256.\n\n```\nepoch - 0 train_loss - 0.4516 train_auc - 0.7477 val_loss - 0.3475 val_auc - 0.7776 time=833.81s\nepoch - 1 train_loss - 0.4281 train_auc - 0.7815 val_loss - 0.3441 val_auc - 0.7831 time=834.06s\nepoch - 2 train_loss - 0.4240 train_auc - 0.7871 val_loss - 0.3421 val_auc - 0.7878 time=840.03s\nepoch - 3 train_loss - 0.4216 train_auc - 0.7904 val_loss - 0.3406 val_auc - 0.7893 time=833.13s\nepoch - 4 train_loss - 0.4200 train_auc - 0.7928 val_loss - 0.3401 val_auc - 0.7917 time=832.33s\nepoch - 5 train_loss - 0.4187 train_auc - 0.7946 val_loss - 0.3398 val_auc - 0.7927 time=832.43s\nepoch - 6 train_loss - 0.4177 train_auc - 0.7962 val_loss - 0.3383 val_auc - 0.7942 time=847.83s\nepoch - 7 train_loss - 0.4167 train_auc - 0.7976 val_loss - 0.3378 val_auc - 0.7948 time=843.36s\nepoch - 8 train_loss - 0.4159 train_auc - 0.7989 val_loss - 0.3373 val_auc - 0.7952 time=842.63s\nepoch - 9 train_loss - 0.4151 train_auc - 0.8002 val_loss - 0.3368 val_auc - 0.7959 time=862.17s\nepoch - 10 train_loss - 0.4143 train_auc - 0.8013 val_loss - 0.3365 val_auc - 0.7967 time=857.08s\nepoch - 11 train_loss - 0.4135 train_auc - 0.8023 val_loss - 0.3359 val_auc - 0.7983 time=842.30s\nepoch - 12 train_loss - 0.4128 train_auc - 0.8033 val_loss - 0.3356 val_auc - 0.7982 time=853.73s\nepoch - 13 train_loss - 0.4121 train_auc - 0.8043 val_loss - 0.3358 val_auc - 0.7989 time=843.08s\nepoch - 14 train_loss - 0.4115 train_auc - 0.8051 val_loss - 0.3357 val_auc - 0.7985 time=843.85s\nepoch - 15 train_loss - 0.4109 train_auc - 0.8060 val_loss - 0.3357 val_auc - 0.7982 time=886.01s\nepoch - 16 train_loss - 0.4103 train_auc - 0.8067 val_loss - 0.3355 val_auc - 0.7992 time=889.84s\nepoch - 17 train_loss - 0.4098 train_auc - 0.8071 val_loss - 0.3356 val_auc - 0.7995 time=888.07s\nepoch - 18 train_loss - 0.4092 train_auc - 0.8078 val_loss - 0.3353 val_auc - 0.7995 time=890.12s\nepoch - 19 train_loss - 0.4087 train_auc - 0.8088 val_loss - 0.3353 val_auc - 0.8001 time=889.51s\nepoch - 20 train_loss - 0.4081 train_auc - 0.8095 val_loss - 0.3358 val_auc - 0.8000 time=887.66s\nepoch - 21 train_loss - 0.4076 train_auc - 0.8101 val_loss - 0.3359 val_auc - 0.7994 time=888.36s\nepoch - 22 train_loss - 0.4070 train_auc - 0.8106 val_loss - 0.3362 val_auc - 0.7999 time=887.26s\nepoch - 23 train_loss - 0.4064 train_auc - 0.8115 val_loss - 0.3363 val_auc - 0.7996 time=889.62s\nepoch - 24 train_loss - 0.4058 train_auc - 0.8119 val_loss - 0.3365 val_auc - 0.7989 time=889.73s\n```",
    "1137037": "Interesting, thank you, nice model you got! @m10515009 Best of luck",
    "1137042": "m10515009  are you using cv strategy of tito or default last 100 ?",
    "1137050": "I used CV strategy from tito.",
    "1137091": "m10515009 what is ur dropout?\nI got my access issues fixed now, but it is too late to move to Saint. So stuck with SAKT. Not sure if same params work for SAKT though..\nbut great to see u with high score..hope u move into gold zone soon",
    "1137172": "what's your length of each sample feed to saint plus plus? @m10515009",
    "1137483": "allohvk I using dropout rate 0.1 in both SAKT and SAINT.\n@cswwp347724 100, I have tried 160 but not improvement.",
    "1137489": "Guys, is there a SAINT ++ I am unaware of ?  I only know SAINT + . Can some one link me the paper?",
    "1137523": "I only know SAINT+ too.😂",
    "1137528": "SAINT ++ -> Kaggle folks are going to give one because as per paper they have AUC of .79x and top folks have AUC of .81x and SAKT itself can reach close to .79x it seems as others have reported here and there.",
    "1137583": ":( @adityaecdrid so u r saying if I use SAKT the highest I can hope to reach is 79x..\nwould u recommend any good saint+ implementation which I can quickly fine-tune. Unfortunately I can only spend a couple of hours each night as my job keeps me quite occupied",
    "1137590": "I wish i knew one and had one; I couldn't get SAINT as strong as others, so the current score doesn't have a SAINT/SAINT+.",
    "1137598": "allohvk SAKT reaching 0.79 is difficult for me.\nThe history of my model is as follows.\n\n1.  SAKT LB 0.776\n2. SAINT LB 0.784\n3. SAINT+ LB 0.792\n\nIn my opinion, no one will public successful SAINT/SAINT+  training notebook, because it will influence LB very much.",
    "1137625": "m10515009 thanks for the info. So it is a solid 1.6% lead for Saint over sakt\nI am trying a customised version of stack. I will try to reach 80 with customised SAKT and an ensemble\n\nif time permits, last day (Thursday) I will take a shot at saint.",
    "1137865": "adityaecdrid if it helps pl check the Saint paper. They have some alternate models more closer to SAKT. If like me you are already invested in SAKT, you may want to check out those models (UTMTI, LTMTI, SSAKT) etc. The other point to note is that unlike Saint, these models underperform at 512 and may or may not perform at 256. Likewise they are better at layer=2 or 3 at best.",
    "1137869": "most likely the clean segregation of exercise and response in the Saint architecture in some way prevents overfitting and model performs better with more dimensions and layers."
  },
  "source": "meta"
}