{
  "id": 201944,
  "title": "Query About Tags",
  "url": "/competitions/riiid-test-answer-prediction/discussion/201944",
  "author_name": "",
  "post_date": "2020-12-07T13:08:39.150883700Z",
  "votes": 1,
  "comment_count": 24,
  "views": 0,
  "content": "<p>in Saint Paper there is mention of Category Embeddings as one of the input.<br>\n does this  category  refers to Part or Tag  of the question?</p>",
  "messages": [
    {
      "id": "1105011",
      "postDate": "12/07/2020 13:08:39",
      "content": "<p>in Saint Paper there is mention of Category Embeddings as one of the input.<br>\n does this  category  refers to Part or Tag  of the question?</p>",
      "rawMarkdown": "in Saint Paper there is mention of Category Embeddings as one of the input.\n does this  category  refers to Part or Tag  of the question?",
      "votes": null
    },
    {
      "id": "1105277",
      "postDate": "12/07/2020 17:58:21",
      "content": "<p>I believe for SAINT/SAINT+ the embeddings used are</p>\n<ol>\n<li>Positional </li>\n<li>Content id</li>\n<li>Part id</li>\n</ol>",
      "rawMarkdown": "I believe for SAINT/SAINT+ the embeddings used are\n\n1. Positional \n1. Content id\n1. Part id",
      "votes": null
    },
    {
      "id": "1105677",
      "postDate": "12/08/2020 04:53:01",
      "content": "<p>In Saint plus they say part but in Saint its exercise category, while part is section of assessment comprising of various question </p>",
      "rawMarkdown": "In Saint plus they say part but in Saint its exercise category, while part is section of assessment comprising of various question",
      "votes": null
    },
    {
      "id": "1105688",
      "postDate": "12/08/2020 05:05:28",
      "content": "<p>Oh, I might have mixed them then. Thanks for clarifying. I don't see a reason for not using both at the same time. They both incorporate different information to the model.</p>",
      "rawMarkdown": "Oh, I might have mixed them then. Thanks for clarifying. I don't see a reason for not using both at the same time. They both incorporate different information to the model.",
      "votes": null
    },
    {
      "id": "1105696",
      "postDate": "12/08/2020 05:16:25",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>  wcm , I do see good improvement upon adding on top of Excercise and Output seq , further embeddings</p>\n<p>is your current LB score based on SAINT or LGB.. </p>\n<p>If yes how much CV you get with SAINT Base version. </p>\n<p>They also mention about Start token Embeddings i dint quite get that ..</p>",
      "rawMarkdown": "abdurrafae  wcm , I do see good improvement upon adding on top of Excercise and Output seq , further embeddings\n\n\n  is your current LB score based on SAINT or LGB.. \n\nIf yes how much CV you get with SAINT Base version. \n\nThey also mention about Start token Embeddings i dint quite get that ..",
      "votes": null
    },
    {
      "id": "1105821",
      "postDate": "12/08/2020 08:27:12",
      "content": "<p>I got upto 0.775 with SAINT+ on validation data. my validation is closely co-related to LB +/- 0.001.</p>\n<p>I only trained it for one session of Kaggle environment so not sure if it would have improved further or not, but the gain in each epoch was too low for me to invest in it further. </p>\n<p>The issue I was having was that the model diverged when I increase number of layers or model dimensions. So I built it from scratch again and made some modifications along the way. Now it doesn't match the exact SAINT/SAINT+ but it's quite similar.</p>\n<p>For start token, I was just padding the sequence with zeros on left (with length 1) and cropping (with length 1) from right</p>",
      "rawMarkdown": "I got upto 0.775 with SAINT+ on validation data. my validation is closely co-related to LB +/- 0.001.\n\nI only trained it for one session of Kaggle environment so not sure if it would have improved further or not, but the gain in each epoch was too low for me to invest in it further. \n\nThe issue I was having was that the model diverged when I increase number of layers or model dimensions. So I built it from scratch again and made some modifications along the way. Now it doesn't match the exact SAINT/SAINT+ but it's quite similar.\n\nFor start token, I was just padding the sequence with zeros on left (with length 1) and cropping (with length 1) from right",
      "votes": null
    },
    {
      "id": "1105889",
      "postDate": "12/08/2020 09:48:41",
      "content": "<p>ok  thanks , in Saint i get 75.75 as cv yet to test cv .</p>\n<ol>\n<li><p>Challenge m finding is how do I handle user ID not in test as i wont get the same in responses that iterations . </p></li>\n<li><p>In Saint +  did you add lag time also? </p></li>\n<li><p>While addong prior elapsed time were u able to get right chronology of question .there is some discussion on going about the caveat involved if rely   just on default sequence .</p></li>\n<li><p>What is purpose of start token ,I dint gave any. </p></li>\n</ol>",
      "rawMarkdown": "ok  thanks , in Saint i get 75.75 as cv yet to test cv .\n\n1. Challenge m finding is how do I handle user ID not in test as i wont get the same in responses that iterations . \n\n2. In Saint +  did you add lag time also? \n\n3. While addong prior elapsed time were u able to get right chronology of question .there is some discussion on going about the caveat involved if rely   just on default sequence .\n\n4. What is purpose of start token ,I dint gave any.",
      "votes": null
    },
    {
      "id": "1106001",
      "postDate": "12/08/2020 11:57:52",
      "content": "<p>1) I use dummy responses for new users till I get them in next iteration.<br>\n2) I used a feature based on lag time, so not lag time exactly.<br>\n3) I used lag time as absolute difference between bundles (i.e. disregarding the time elapsed field to calculate this)<br>\n4) Start token is to mis-align your questions and response streams since at test time you'll have access to question (1:t) and  responses (1:t-1).</p>",
      "rawMarkdown": "1) I use dummy responses for new users till I get them in next iteration.\n2) I used a feature based on lag time, so not lag time exactly.\n3) I used lag time as absolute difference between bundles (i.e. disregarding the time elapsed field to calculate this)\n4) Start token is to mis-align your questions and response streams since at test time you'll have access to question (1:t) and  responses (1:t-1).",
      "votes": null
    },
    {
      "id": "1107500",
      "postDate": "12/09/2020 18:31:06",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>  how did perform merge operation between train csv and any other df during inference</p>",
      "rawMarkdown": "abdurrafae  how did perform merge operation between train csv and any other df during inference",
      "votes": null
    },
    {
      "id": "1107510",
      "postDate": "12/09/2020 18:39:57",
      "content": "<p>I use pandas for that. I only have 1 dataframe to join to, so it doesn't take that long</p>",
      "rawMarkdown": "I use pandas for that. I only have 1 dataframe to join to, so it doesn't take that long",
      "votes": null
    },
    {
      "id": "1107512",
      "postDate": "12/09/2020 18:40:49",
      "content": "<p>I don't have the train_df at inference, I've a customized object that stores data for each user.</p>",
      "rawMarkdown": "I don't have the train_df at inference, I've a customized object that stores data for each user.",
      "votes": null
    },
    {
      "id": "1108158",
      "postDate": "12/10/2020 10:49:27",
      "content": "<p>Thanks</p>\n<p>1) <a href=\"https://www.kaggle.com/mpware/sakt-fork/\" target=\"_blank\">https://www.kaggle.com/mpware/sakt-fork/</a><br>\nin this kernel what is significance of starting question from 2:  rather than from 0 or 1: </p>\n<p><code>questions = np.append(q[2:], [target_id])</code></p>\n<p>2) what was your first sub score with SAINT .. i get good cv with base SAINT but LB not aligned with CV. 76.42/ 74.7  following the same input indexing as above and not accumulating the Response history ,to start with. <br>\n not sure if am  making any mistake in inference</p>",
      "rawMarkdown": "Thanks\n\n1) https://www.kaggle.com/mpware/sakt-fork/\nin this kernel what is significance of starting question from 2:  rather than from 0 or 1: \n       \n`questions = np.append(q[2:], [target_id])`\n\n2) what was your first sub score with SAINT .. i get good cv with base SAINT but LB not aligned with CV. 76.42/ 74.7  following the same input indexing as above and not accumulating the Response history ,to start with. \n not sure if am  making any mistake in inference",
      "votes": null
    },
    {
      "id": "1108249",
      "postDate": "12/10/2020 13:11:14",
      "content": "<p>No idea about the SAKT kernel.</p>\n<p>I had a validation score of 0.775 with SAINT, didn't get an LB on that.<br>\nMy Validation score is closely related to LB with +/- 0.001</p>",
      "rawMarkdown": "No idea about the SAKT kernel.\n\nI had a validation score of 0.775 with SAINT, didn't get an LB on that.\nMy Validation score is closely related to LB with +/- 0.001",
      "votes": null
    },
    {
      "id": "1108429",
      "postDate": "12/10/2020 16:28:26",
      "content": "<pre><code>from paper\nThe encoder takes the sequence\nEe = [Ee1,··· ,Eek] of exercise embeddings and feeds the processed output O = [O1,··· ,Ok\n] to the decoder. The decoder\ntakes O and another sequential input R\ne = [S,Re1,··· ,Rek−1] of\nresponse embeddings with the start token embedding S, and\nproduces the predicted responses ˆr = [rˆ1,··· ,rˆk\n].\n</code></pre>\n<p>I had one more doubt wrt to SAINT.. for any user id if i understand well  for any current Exercise E1 model has got previous response as correct or not  but in the case of test  , when we add in Test questions we dont have  Test users exercise history for test question ..<br>\nAt any given user iteration<br>\neg: User 1- Test question T1,T2,T3.<br>\nExercise seq could become : <br>\nE1(from Train or val) -T1,T2,T3<br>\nResponse sequence<br>\nE1,T1,T2,T3<br>\nR0 R1,???????</p>",
      "rawMarkdown": "```\nfrom paper\nThe encoder takes the sequence\nEe = [Ee1,··· ,Eek] of exercise embeddings and feeds the processed output O = [O1,··· ,Ok\n] to the decoder. The decoder\ntakes O and another sequential input R\ne = [S,Re1,··· ,Rek−1] of\nresponse embeddings with the start token embedding S, and\nproduces the predicted responses ˆr = [rˆ1,··· ,rˆk\n].\n```\nI had one more doubt wrt to SAINT.. for any user id if i understand well  for any current Exercise E1 model has got previous response as correct or not  but in the case of test  , when we add in Test questions we dont have  Test users exercise history for test question ..\nAt any given user iteration\neg: User 1- Test question T1,T2,T3.\nExercise seq could become : \nE1(from Train or val) -T1,T2,T3\nResponse sequence\nE1,T1,T2,T3\nR0 R1,???????",
      "votes": null
    },
    {
      "id": "1108559",
      "postDate": "12/10/2020 19:37:24",
      "content": "<p>We do get the correct answers in test as well, but it's lagging by one bundle. You'll have to figure out how to manage that according to your pipeline. I use dummy values to keep the seq_len maintained and update them when I get the actual values.</p>",
      "rawMarkdown": "We do get the correct answers in test as well, but it's lagging by one bundle. You'll have to figure out how to manage that according to your pipeline. I use dummy values to keep the seq_len maintained and update them when I get the actual values.",
      "votes": null
    },
    {
      "id": "1108829",
      "postDate": "12/11/2020 03:45:41",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>  thanks for responses. I am getting to depth of it now.<br>\nBetween would you like to team up . I  have access to multiple V16 gpus (2 to 4) which is helping me to  do lot of experiments. <br>\nIn addition to GPUs ,below are other values i can add for sure.<br>\n1) Hard work..<br>\nhere is related post that supports my words<br>\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189251\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189251</a></p>\n<p>2) Continuous ideas about features that can strengthen model. Once base gets going, additional temporal features</p>\n<p>3) Came across a New Paper released few months ago with potential of overshooting the SAINT/+ benchmark. In discussion with author for its implementation.</p>",
      "rawMarkdown": "abdurrafae  thanks for responses. I am getting to depth of it now.\nBetween would you like to team up . I  have access to multiple V16 gpus (2 to 4) which is helping me to  do lot of experiments. \nIn addition to GPUs ,below are other values i can add for sure.\n1) Hard work..\nhere is related post that supports my words\nhttps://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189251\n\n2) Continuous ideas about features that can strengthen model. Once base gets going, additional temporal features\n\n3) Came across a New Paper released few months ago with potential of overshooting the SAINT/+ benchmark. In discussion with author for its implementation.",
      "votes": null
    },
    {
      "id": "1108866",
      "postDate": "12/11/2020 05:01:23",
      "content": "<p>Thanks for the kind offer, your profile looks really good. However I'll be really busy and probably would not have time to be doing anything other than making minor changes and checking model results for next 2 weeks.</p>\n<p>Let's connect again around 23rd and see what we can do then.</p>",
      "rawMarkdown": "Thanks for the kind offer, your profile looks really good. However I'll be really busy and probably would not have time to be doing anything other than making minor changes and checking model results for next 2 weeks.\n\nLet's connect again around 23rd and see what we can do then.",
      "votes": null
    },
    {
      "id": "1108878",
      "postDate": "12/11/2020 05:25:43",
      "content": "<p>Thanks Sure I have good bandwidth to work on i think till end ,in osic just iafoss stood by  till end rest frustratingly gave up ,  <br>\nMy intent to team up was  to  try n do some thing on top of what some thing   like adding more temporal feature on top of some thing similar and successfully already done  .</p>",
      "rawMarkdown": "Thanks Sure I have good bandwidth to work on i think till end ,in osic just iafoss stood by  till end rest frustratingly gave up ,  \nMy intent to team up was  to  try n do some thing on top of what some thing   like adding more temporal feature on top of some thing similar and successfully already done  .",
      "votes": null
    },
    {
      "id": "1108894",
      "postDate": "12/11/2020 05:48:52",
      "content": "<p>I'll also be free from 23rd onwards till end of competition so that's at least one good thing going for me.</p>",
      "rawMarkdown": "I'll also be free from 23rd onwards till end of competition so that's at least one good thing going for me.",
      "votes": null
    },
    {
      "id": "1108951",
      "postDate": "12/11/2020 07:16:22",
      "content": "<p>good to hear…<br>\ntell me one thing<br>\n1) I see the usage of Masked BCE loss in saint topic discussion, is that useful <br>\n2 ) I dont do shifting of embeddings as mpware is showing through </p>\n<pre><code>response_size = 2\nself.response_embedding = nn.Embedding(response_size + 1, embedding_dim) # +1 to include start token\n...\nx_correctness = self.response_embedding(response_sequence)\n...\n# Add start token to correctness\nx_correctness = torch.roll(x_correctness, shifts=(0, 1, 0), dims=(0, 1, 0)) # Shift right the sequence\nx_correctness[:,0,:] = self.response_size # Start token\n</code></pre>\n<p>rather i just do shift of inputs and assume embeddings will be passed accordingly.</p>\n<pre><code>label = qa[1:]\nrt=qa[1:-1].copy()\n rt = np.append(np.zeros((1,)),rt)\n</code></pre>\n<p>will it make a difference ?</p>",
      "rawMarkdown": "good to hear...\ntell me one thing\n1) I see the usage of Masked BCE loss in saint topic discussion, is that useful \n2 ) I dont do shifting of embeddings as mpware is showing through \n```\nresponse_size = 2\nself.response_embedding = nn.Embedding(response_size + 1, embedding_dim) # +1 to include start token\n...\nx_correctness = self.response_embedding(response_sequence)\n...\n# Add start token to correctness\nx_correctness = torch.roll(x_correctness, shifts=(0, 1, 0), dims=(0, 1, 0)) # Shift right the sequence\nx_correctness[:,0,:] = self.response_size # Start token\n```\nrather i just do shift of inputs and assume embeddings will be passed accordingly.\n```\nlabel = qa[1:]\nrt=qa[1:-1].copy()\n rt = np.append(np.zeros((1,)),rt)\n```\nwill it make a difference ?",
      "votes": null
    },
    {
      "id": "1108974",
      "postDate": "12/11/2020 07:59:43",
      "content": "<p>For the shifting. I believe doing it either way is okay, I had implemented it by shifting the embeddings earlier but then did it by shifting the inputs in a later re-factoring of my code.</p>\n<p>For loss, I initially didn't use any mask (as no lectures were included). I have added a mask for loss for the lecture interactions now.</p>",
      "rawMarkdown": "For the shifting. I believe doing it either way is okay, I had implemented it by shifting the embeddings earlier but then did it by shifting the inputs in a later re-factoring of my code.\n\nFor loss, I initially didn't use any mask (as no lectures were included). I have added a mask for loss for the lecture interactions now.",
      "votes": null
    },
    {
      "id": "1117917",
      "postDate": "12/18/2020 15:28:33",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>  my saint baseline stands at 77.3 ,trying few more things .<br>\nSo just thought of checking are  you open for teaming  up next week .</p>",
      "rawMarkdown": "abdurrafae  my saint baseline stands at 77.3 ,trying few more things .\nSo just thought of checking are  you open for teaming  up next week .",
      "votes": null
    },
    {
      "id": "1133817",
      "postDate": "12/31/2020 15:01:32",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>  you are doing good alone . But still just last ask.<br>\nWould you like to team up. If solo gold/ silver medal standing  is of less  importance to you or you think reaching winning range/gold zone could be difficult for you.<br>\nour SAINT pipeline got us 78.8 <br>\nwe are fixing small bugs  bugs in saint plus (current is 79.3)  and we have good lgb 79.2.<br>\nwe may reach to winning range of 81+ after mixing your model with our lgb  </p>",
      "rawMarkdown": "abdurrafae  you are doing good alone . But still just last ask.\nWould you like to team up. If solo gold/ silver medal standing  is of less  importance to you or you think reaching winning range/gold zone could be difficult for you.\nour SAINT pipeline got us 78.8 \nwe are fixing small bugs  bugs in saint plus (current is 79.3)  and we have good lgb 79.2.\nwe may reach to winning range of 81+ after mixing your model with our lgb",
      "votes": null
    },
    {
      "id": "1133823",
      "postDate": "12/31/2020 15:06:43",
      "content": "<p>Thanks man, I'll be going solo for this one I think.<br>\nBest of luck mate</p>",
      "rawMarkdown": "Thanks man, I'll be going solo for this one I think.\nBest of luck mate",
      "votes": null
    },
    {
      "id": "1133825",
      "postDate": "12/31/2020 15:10:33",
      "content": "<blockquote>\n  <p>Thanks man, I'll be going solo for this one I think.<br>\n  Best of luck mate</p>\n</blockquote>\n<p>No issues  i appreciate your hard work till date..<br>\nmay u reach winning alone in last few days.</p>",
      "rawMarkdown": "> Thanks man, I'll be going solo for this one I think.\n> Best of luck mate\n\nNo issues  i appreciate your hard work till date..\nmay u reach winning alone in last few days.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1105277,
      "author_name": "abdurrafae",
      "author_url": "",
      "post_date": "12/07/2020 17:58:21",
      "content": "<p>I believe for SAINT/SAINT+ the embeddings used are</p>\n<ol>\n<li>Positional </li>\n<li>Content id</li>\n<li>Part id</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1105677,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/08/2020 04:53:01",
          "content": "<p>In Saint plus they say part but in Saint its exercise category, while part is section of assessment comprising of various question </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105688,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/08/2020 05:05:28",
          "content": "<p>Oh, I might have mixed them then. Thanks for clarifying. I don't see a reason for not using both at the same time. They both incorporate different information to the model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105696,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/08/2020 05:16:25",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>  wcm , I do see good improvement upon adding on top of Excercise and Output seq , further embeddings</p>\n<p>is your current LB score based on SAINT or LGB.. </p>\n<p>If yes how much CV you get with SAINT Base version. </p>\n<p>They also mention about Start token Embeddings i dint quite get that ..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105821,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/08/2020 08:27:12",
          "content": "<p>I got upto 0.775 with SAINT+ on validation data. my validation is closely co-related to LB +/- 0.001.</p>\n<p>I only trained it for one session of Kaggle environment so not sure if it would have improved further or not, but the gain in each epoch was too low for me to invest in it further. </p>\n<p>The issue I was having was that the model diverged when I increase number of layers or model dimensions. So I built it from scratch again and made some modifications along the way. Now it doesn't match the exact SAINT/SAINT+ but it's quite similar.</p>\n<p>For start token, I was just padding the sequence with zeros on left (with length 1) and cropping (with length 1) from right</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105889,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/08/2020 09:48:41",
          "content": "<p>ok  thanks , in Saint i get 75.75 as cv yet to test cv .</p>\n<ol>\n<li><p>Challenge m finding is how do I handle user ID not in test as i wont get the same in responses that iterations . </p></li>\n<li><p>In Saint +  did you add lag time also? </p></li>\n<li><p>While addong prior elapsed time were u able to get right chronology of question .there is some discussion on going about the caveat involved if rely   just on default sequence .</p></li>\n<li><p>What is purpose of start token ,I dint gave any. </p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1106001,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/08/2020 11:57:52",
          "content": "<p>1) I use dummy responses for new users till I get them in next iteration.<br>\n2) I used a feature based on lag time, so not lag time exactly.<br>\n3) I used lag time as absolute difference between bundles (i.e. disregarding the time elapsed field to calculate this)<br>\n4) Start token is to mis-align your questions and response streams since at test time you'll have access to question (1:t) and  responses (1:t-1).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107500,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/09/2020 18:31:06",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>  how did perform merge operation between train csv and any other df during inference</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107510,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/09/2020 18:39:57",
          "content": "<p>I use pandas for that. I only have 1 dataframe to join to, so it doesn't take that long</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107512,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/09/2020 18:40:49",
          "content": "<p>I don't have the train_df at inference, I've a customized object that stores data for each user.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108158,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/10/2020 10:49:27",
          "content": "<p>Thanks</p>\n<p>1) <a href=\"https://www.kaggle.com/mpware/sakt-fork/\" target=\"_blank\">https://www.kaggle.com/mpware/sakt-fork/</a><br>\nin this kernel what is significance of starting question from 2:  rather than from 0 or 1: </p>\n<p><code>questions = np.append(q[2:], [target_id])</code></p>\n<p>2) what was your first sub score with SAINT .. i get good cv with base SAINT but LB not aligned with CV. 76.42/ 74.7  following the same input indexing as above and not accumulating the Response history ,to start with. <br>\n not sure if am  making any mistake in inference</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108249,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/10/2020 13:11:14",
          "content": "<p>No idea about the SAKT kernel.</p>\n<p>I had a validation score of 0.775 with SAINT, didn't get an LB on that.<br>\nMy Validation score is closely related to LB with +/- 0.001</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108429,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/10/2020 16:28:26",
          "content": "<pre><code>from paper\nThe encoder takes the sequence\nEe = [Ee1,··· ,Eek] of exercise embeddings and feeds the processed output O = [O1,··· ,Ok\n] to the decoder. The decoder\ntakes O and another sequential input R\ne = [S,Re1,··· ,Rek−1] of\nresponse embeddings with the start token embedding S, and\nproduces the predicted responses ˆr = [rˆ1,··· ,rˆk\n].\n</code></pre>\n<p>I had one more doubt wrt to SAINT.. for any user id if i understand well  for any current Exercise E1 model has got previous response as correct or not  but in the case of test  , when we add in Test questions we dont have  Test users exercise history for test question ..<br>\nAt any given user iteration<br>\neg: User 1- Test question T1,T2,T3.<br>\nExercise seq could become : <br>\nE1(from Train or val) -T1,T2,T3<br>\nResponse sequence<br>\nE1,T1,T2,T3<br>\nR0 R1,???????</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108559,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/10/2020 19:37:24",
          "content": "<p>We do get the correct answers in test as well, but it's lagging by one bundle. You'll have to figure out how to manage that according to your pipeline. I use dummy values to keep the seq_len maintained and update them when I get the actual values.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108829,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/11/2020 03:45:41",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>  thanks for responses. I am getting to depth of it now.<br>\nBetween would you like to team up . I  have access to multiple V16 gpus (2 to 4) which is helping me to  do lot of experiments. <br>\nIn addition to GPUs ,below are other values i can add for sure.<br>\n1) Hard work..<br>\nhere is related post that supports my words<br>\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189251\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189251</a></p>\n<p>2) Continuous ideas about features that can strengthen model. Once base gets going, additional temporal features</p>\n<p>3) Came across a New Paper released few months ago with potential of overshooting the SAINT/+ benchmark. In discussion with author for its implementation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108866,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/11/2020 05:01:23",
          "content": "<p>Thanks for the kind offer, your profile looks really good. However I'll be really busy and probably would not have time to be doing anything other than making minor changes and checking model results for next 2 weeks.</p>\n<p>Let's connect again around 23rd and see what we can do then.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108878,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/11/2020 05:25:43",
          "content": "<p>Thanks Sure I have good bandwidth to work on i think till end ,in osic just iafoss stood by  till end rest frustratingly gave up ,  <br>\nMy intent to team up was  to  try n do some thing on top of what some thing   like adding more temporal feature on top of some thing similar and successfully already done  .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108894,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/11/2020 05:48:52",
          "content": "<p>I'll also be free from 23rd onwards till end of competition so that's at least one good thing going for me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108951,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/11/2020 07:16:22",
          "content": "<p>good to hear…<br>\ntell me one thing<br>\n1) I see the usage of Masked BCE loss in saint topic discussion, is that useful <br>\n2 ) I dont do shifting of embeddings as mpware is showing through </p>\n<pre><code>response_size = 2\nself.response_embedding = nn.Embedding(response_size + 1, embedding_dim) # +1 to include start token\n...\nx_correctness = self.response_embedding(response_sequence)\n...\n# Add start token to correctness\nx_correctness = torch.roll(x_correctness, shifts=(0, 1, 0), dims=(0, 1, 0)) # Shift right the sequence\nx_correctness[:,0,:] = self.response_size # Start token\n</code></pre>\n<p>rather i just do shift of inputs and assume embeddings will be passed accordingly.</p>\n<pre><code>label = qa[1:]\nrt=qa[1:-1].copy()\n rt = np.append(np.zeros((1,)),rt)\n</code></pre>\n<p>will it make a difference ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108974,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/11/2020 07:59:43",
          "content": "<p>For the shifting. I believe doing it either way is okay, I had implemented it by shifting the embeddings earlier but then did it by shifting the inputs in a later re-factoring of my code.</p>\n<p>For loss, I initially didn't use any mask (as no lectures were included). I have added a mask for loss for the lecture interactions now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1117917,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/18/2020 15:28:33",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>  my saint baseline stands at 77.3 ,trying few more things .<br>\nSo just thought of checking are  you open for teaming  up next week .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1133817,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/31/2020 15:01:32",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a>  you are doing good alone . But still just last ask.<br>\nWould you like to team up. If solo gold/ silver medal standing  is of less  importance to you or you think reaching winning range/gold zone could be difficult for you.<br>\nour SAINT pipeline got us 78.8 <br>\nwe are fixing small bugs  bugs in saint plus (current is 79.3)  and we have good lgb 79.2.<br>\nwe may reach to winning range of 81+ after mixing your model with our lgb  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1133823,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/31/2020 15:06:43",
          "content": "<p>Thanks man, I'll be going solo for this one I think.<br>\nBest of luck mate</p>",
          "votes": null,
          "replies": [
            {
              "id": 1133825,
              "author_name": "jaideepvalani",
              "author_url": "",
              "post_date": "12/31/2020 15:10:33",
              "content": "<blockquote>\n  <p>Thanks man, I'll be going solo for this one I think.<br>\n  Best of luck mate</p>\n</blockquote>\n<p>No issues  i appreciate your hard work till date..<br>\nmay u reach winning alone in last few days.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1105011": "in Saint Paper there is mention of Category Embeddings as one of the input.\n does this  category  refers to Part or Tag  of the question?",
    "1105277": "I believe for SAINT/SAINT+ the embeddings used are\n\n1. Positional \n1. Content id\n1. Part id",
    "1105677": "In Saint plus they say part but in Saint its exercise category, while part is section of assessment comprising of various question",
    "1105688": "Oh, I might have mixed them then. Thanks for clarifying. I don't see a reason for not using both at the same time. They both incorporate different information to the model.",
    "1105696": "abdurrafae  wcm , I do see good improvement upon adding on top of Excercise and Output seq , further embeddings\n\n\n  is your current LB score based on SAINT or LGB.. \n\nIf yes how much CV you get with SAINT Base version. \n\nThey also mention about Start token Embeddings i dint quite get that ..",
    "1105821": "I got upto 0.775 with SAINT+ on validation data. my validation is closely co-related to LB +/- 0.001.\n\nI only trained it for one session of Kaggle environment so not sure if it would have improved further or not, but the gain in each epoch was too low for me to invest in it further. \n\nThe issue I was having was that the model diverged when I increase number of layers or model dimensions. So I built it from scratch again and made some modifications along the way. Now it doesn't match the exact SAINT/SAINT+ but it's quite similar.\n\nFor start token, I was just padding the sequence with zeros on left (with length 1) and cropping (with length 1) from right",
    "1105889": "ok  thanks , in Saint i get 75.75 as cv yet to test cv .\n\n1. Challenge m finding is how do I handle user ID not in test as i wont get the same in responses that iterations . \n\n2. In Saint +  did you add lag time also? \n\n3. While addong prior elapsed time were u able to get right chronology of question .there is some discussion on going about the caveat involved if rely   just on default sequence .\n\n4. What is purpose of start token ,I dint gave any.",
    "1106001": "1) I use dummy responses for new users till I get them in next iteration.\n2) I used a feature based on lag time, so not lag time exactly.\n3) I used lag time as absolute difference between bundles (i.e. disregarding the time elapsed field to calculate this)\n4) Start token is to mis-align your questions and response streams since at test time you'll have access to question (1:t) and  responses (1:t-1).",
    "1107500": "abdurrafae  how did perform merge operation between train csv and any other df during inference",
    "1107510": "I use pandas for that. I only have 1 dataframe to join to, so it doesn't take that long",
    "1107512": "I don't have the train_df at inference, I've a customized object that stores data for each user.",
    "1108158": "Thanks\n\n1) https://www.kaggle.com/mpware/sakt-fork/\nin this kernel what is significance of starting question from 2:  rather than from 0 or 1: \n       \n`questions = np.append(q[2:], [target_id])`\n\n2) what was your first sub score with SAINT .. i get good cv with base SAINT but LB not aligned with CV. 76.42/ 74.7  following the same input indexing as above and not accumulating the Response history ,to start with. \n not sure if am  making any mistake in inference",
    "1108249": "No idea about the SAKT kernel.\n\nI had a validation score of 0.775 with SAINT, didn't get an LB on that.\nMy Validation score is closely related to LB with +/- 0.001",
    "1108429": "```\nfrom paper\nThe encoder takes the sequence\nEe = [Ee1,··· ,Eek] of exercise embeddings and feeds the processed output O = [O1,··· ,Ok\n] to the decoder. The decoder\ntakes O and another sequential input R\ne = [S,Re1,··· ,Rek−1] of\nresponse embeddings with the start token embedding S, and\nproduces the predicted responses ˆr = [rˆ1,··· ,rˆk\n].\n```\nI had one more doubt wrt to SAINT.. for any user id if i understand well  for any current Exercise E1 model has got previous response as correct or not  but in the case of test  , when we add in Test questions we dont have  Test users exercise history for test question ..\nAt any given user iteration\neg: User 1- Test question T1,T2,T3.\nExercise seq could become : \nE1(from Train or val) -T1,T2,T3\nResponse sequence\nE1,T1,T2,T3\nR0 R1,???????",
    "1108559": "We do get the correct answers in test as well, but it's lagging by one bundle. You'll have to figure out how to manage that according to your pipeline. I use dummy values to keep the seq_len maintained and update them when I get the actual values.",
    "1108829": "abdurrafae  thanks for responses. I am getting to depth of it now.\nBetween would you like to team up . I  have access to multiple V16 gpus (2 to 4) which is helping me to  do lot of experiments. \nIn addition to GPUs ,below are other values i can add for sure.\n1) Hard work..\nhere is related post that supports my words\nhttps://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189251\n\n2) Continuous ideas about features that can strengthen model. Once base gets going, additional temporal features\n\n3) Came across a New Paper released few months ago with potential of overshooting the SAINT/+ benchmark. In discussion with author for its implementation.",
    "1108866": "Thanks for the kind offer, your profile looks really good. However I'll be really busy and probably would not have time to be doing anything other than making minor changes and checking model results for next 2 weeks.\n\nLet's connect again around 23rd and see what we can do then.",
    "1108878": "Thanks Sure I have good bandwidth to work on i think till end ,in osic just iafoss stood by  till end rest frustratingly gave up ,  \nMy intent to team up was  to  try n do some thing on top of what some thing   like adding more temporal feature on top of some thing similar and successfully already done  .",
    "1108894": "I'll also be free from 23rd onwards till end of competition so that's at least one good thing going for me.",
    "1108951": "good to hear...\ntell me one thing\n1) I see the usage of Masked BCE loss in saint topic discussion, is that useful \n2 ) I dont do shifting of embeddings as mpware is showing through \n```\nresponse_size = 2\nself.response_embedding = nn.Embedding(response_size + 1, embedding_dim) # +1 to include start token\n...\nx_correctness = self.response_embedding(response_sequence)\n...\n# Add start token to correctness\nx_correctness = torch.roll(x_correctness, shifts=(0, 1, 0), dims=(0, 1, 0)) # Shift right the sequence\nx_correctness[:,0,:] = self.response_size # Start token\n```\nrather i just do shift of inputs and assume embeddings will be passed accordingly.\n```\nlabel = qa[1:]\nrt=qa[1:-1].copy()\n rt = np.append(np.zeros((1,)),rt)\n```\nwill it make a difference ?",
    "1108974": "For the shifting. I believe doing it either way is okay, I had implemented it by shifting the embeddings earlier but then did it by shifting the inputs in a later re-factoring of my code.\n\nFor loss, I initially didn't use any mask (as no lectures were included). I have added a mask for loss for the lecture interactions now.",
    "1117917": "abdurrafae  my saint baseline stands at 77.3 ,trying few more things .\nSo just thought of checking are  you open for teaming  up next week .",
    "1133817": "abdurrafae  you are doing good alone . But still just last ask.\nWould you like to team up. If solo gold/ silver medal standing  is of less  importance to you or you think reaching winning range/gold zone could be difficult for you.\nour SAINT pipeline got us 78.8 \nwe are fixing small bugs  bugs in saint plus (current is 79.3)  and we have good lgb 79.2.\nwe may reach to winning range of 81+ after mixing your model with our lgb",
    "1133823": "Thanks man, I'll be going solo for this one I think.\nBest of luck mate",
    "1133825": "> Thanks man, I'll be going solo for this one I think.\n> Best of luck mate\n\nNo issues  i appreciate your hard work till date..\nmay u reach winning alone in last few days."
  },
  "source": "meta"
}