{
  "id": 200232,
  "title": "SAINT Inference Tips (Valid for all DL models)",
  "url": "/competitions/riiid-test-answer-prediction/discussion/200232",
  "author_name": "Aditya Soni",
  "post_date": "2020-11-29T15:20:34.245000",
  "votes": 50,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hey Earthlings!</p>\n<p>I know writing inference pipeline is a pain and it's frustrating at times as well. But i hope the below will be useful. </p>\n<p>NB, Please spend a day writing your own. It's quite important actually to do it yourself. What matter's is your attempt, not the results.</p>\n<p>Without further delay, here it's</p>\n<pre><code># the below is pretty robust and generic and works for FIXED_QUESTION_SEQ_LENGTHS_WITH PADDING_ETC.\n\nmodel.eval();\nprevious_test_df = None;\nSTATE = {} # load your global dictionary which has let's say last 100 rows values you care for (which columns etc and \"key\" is the \"user_id\")\nTARGET = \"answered_correctly\"\nMAX_SEQ = 100\n\n\n# looping starts here.\nfor idx, (test_df, sample_prediction_df) in enumerate(iter_test):\n    # here you define your variables where you need to store things (optional)\n    after that there's the section where you have to update the sequences like content_id, answered_correctly, part etc.\n\n    if previous_test_df is not None:\n        # Here you just need to update the things you care for actually.\n        previous_test_df[TARGET] = eval(test_df[\"prior_group_answers_correct\"].iloc[0])\n        previous_test_df = previous_test_df[previous_test_df['content_type_id'] == 0].reset_index(drop=True)\n        # convert it to numpy\n        # update that user's state (here)\n\n    previous_test_df = test_df.copy()\n    test_df = test_df[test_df['content_type_id'] == 0].reset_index(drop=True) # drop videos, depends.\n    # fill-na section here.\n    # other stuffs you might think is important\n    # in the end, you convert your_test_df to numpy.\n\ncurr = test_df.to_numpy()\n    for (cols_name &lt;your's call what to call them&gt;) in curr:\n       # get the user_id\n       # now there are two possible scenarios here, you have already seen this user or else you haven't seen this user till now.\n      # And again, even if you have seen a returning user, there are two cases to that as well, the seq_len already has last 100 seen things and the other case where it doesn't have that.\n        user_id = user_id\n        if user_id in STATE:\n            # split into two cases here depending on the length of the seq which i just talked\n            if _len_of_the_key &gt;= MAX_SEQ:\n                # take the last 99 and stick the latest value at the end, you can get that from the loop itself.\n                .......\n            else:\n                 # handle the padding case here accordinly as you don't have MAX_SEQ elements in your keys.\n        else:\n               # now user_id is a new user\n               # add it's entry to the STATE. NB, add the keys you care about as well here\n               STATE[user_id] = {....}\n\n# looping ends here.\n# create the dataset/data-loader (if needed, i don't need it actually....)\n\n    with torch.no_grad():\n        # convert the dtypes accordingly, NB if use embeddings, ensure you have handled all the possible content_ids and padding index etc. (content_id is just an e.g.)\n\n        # with timer(\"inference here\"):\n        test_df['answered_correctly'] = model(.........).cpu() # convert them to CPU, match the shape here.\n        env.predict(test_df.loc[:, ['row_id', 'answered_correctly']])\n</code></pre>\n<p>The above runs in around 2.15-3 hours for more than 5 subs successfully. </p>\n<p>Thanks to others for inspiring me!</p>\n<p>Hope the above helps, Feel free to let me know if anything is unclear.</p>\n<p>Best,<br>\nAditya.</p>\n<hr>\n<p>Edits -:</p>\n<p>Do some whiteboard coding/ scribble on paper folks and think if you were a machine, what all you would expect to be taken care of by the USER operating you, the problem will reveal it's secrets… :)</p>\n<p>Also, keep a pointer (Len) which indicates where will be the next update index basically. (Ensure you increment this and clip this to a max of MAX_SEQ  100 let's say); And this will be same for all values. Plus, don't pad the target when you prepare the cache for the inference part as well. And then you can related whatever is written above directly!</p>\n<p>PS, I personally don't have a great score with SAINT (yet) on LB (0.74), but the above is exactly what works for me, hence just shared to save the efforts</p>",
  "messages": [
    {
      "id": 1095410,
      "postDate": "2020-11-29T15:20:34.247Z",
      "content": "<p>Hey Earthlings!</p>\n<p>I know writing inference pipeline is a pain and it's frustrating at times as well. But i hope the below will be useful. </p>\n<p>NB, Please spend a day writing your own. It's quite important actually to do it yourself. What matter's is your attempt, not the results.</p>\n<p>Without further delay, here it's</p>\n<pre><code># the below is pretty robust and generic and works for FIXED_QUESTION_SEQ_LENGTHS_WITH PADDING_ETC.\n\nmodel.eval();\nprevious_test_df = None;\nSTATE = {} # load your global dictionary which has let's say last 100 rows values you care for (which columns etc and \"key\" is the \"user_id\")\nTARGET = \"answered_correctly\"\nMAX_SEQ = 100\n\n\n# looping starts here.\nfor idx, (test_df, sample_prediction_df) in enumerate(iter_test):\n    # here you define your variables where you need to store things (optional)\n    after that there's the section where you have to update the sequences like content_id, answered_correctly, part etc.\n\n    if previous_test_df is not None:\n        # Here you just need to update the things you care for actually.\n        previous_test_df[TARGET] = eval(test_df[\"prior_group_answers_correct\"].iloc[0])\n        previous_test_df = previous_test_df[previous_test_df['content_type_id'] == 0].reset_index(drop=True)\n        # convert it to numpy\n        # update that user's state (here)\n\n    previous_test_df = test_df.copy()\n    test_df = test_df[test_df['content_type_id'] == 0].reset_index(drop=True) # drop videos, depends.\n    # fill-na section here.\n    # other stuffs you might think is important\n    # in the end, you convert your_test_df to numpy.\n\ncurr = test_df.to_numpy()\n    for (cols_name &lt;your's call what to call them&gt;) in curr:\n       # get the user_id\n       # now there are two possible scenarios here, you have already seen this user or else you haven't seen this user till now.\n      # And again, even if you have seen a returning user, there are two cases to that as well, the seq_len already has last 100 seen things and the other case where it doesn't have that.\n        user_id = user_id\n        if user_id in STATE:\n            # split into two cases here depending on the length of the seq which i just talked\n            if _len_of_the_key &gt;= MAX_SEQ:\n                # take the last 99 and stick the latest value at the end, you can get that from the loop itself.\n                .......\n            else:\n                 # handle the padding case here accordinly as you don't have MAX_SEQ elements in your keys.\n        else:\n               # now user_id is a new user\n               # add it's entry to the STATE. NB, add the keys you care about as well here\n               STATE[user_id] = {....}\n\n# looping ends here.\n# create the dataset/data-loader (if needed, i don't need it actually....)\n\n    with torch.no_grad():\n        # convert the dtypes accordingly, NB if use embeddings, ensure you have handled all the possible content_ids and padding index etc. (content_id is just an e.g.)\n\n        # with timer(\"inference here\"):\n        test_df['answered_correctly'] = model(.........).cpu() # convert them to CPU, match the shape here.\n        env.predict(test_df.loc[:, ['row_id', 'answered_correctly']])\n</code></pre>\n<p>The above runs in around 2.15-3 hours for more than 5 subs successfully. </p>\n<p>Thanks to others for inspiring me!</p>\n<p>Hope the above helps, Feel free to let me know if anything is unclear.</p>\n<p>Best,<br>\nAditya.</p>\n<hr>\n<p>Edits -:</p>\n<p>Do some whiteboard coding/ scribble on paper folks and think if you were a machine, what all you would expect to be taken care of by the USER operating you, the problem will reveal it's secrets… :)</p>\n<p>Also, keep a pointer (Len) which indicates where will be the next update index basically. (Ensure you increment this and clip this to a max of MAX_SEQ  100 let's say); And this will be same for all values. Plus, don't pad the target when you prepare the cache for the inference part as well. And then you can related whatever is written above directly!</p>\n<p>PS, I personally don't have a great score with SAINT (yet) on LB (0.74), but the above is exactly what works for me, hence just shared to save the efforts</p>",
      "rawMarkdown": "Hey Earthlings!\n\nI know writing inference pipeline is a pain and it's frustrating at times as well. But i hope the below will be useful. \n\nNB, Please spend a day writing your own. It's quite important actually to do it yourself. What matter's is your attempt, not the results.\n\nWithout further delay, here it's\n\n```\n# the below is pretty robust and generic and works for FIXED_QUESTION_SEQ_LENGTHS_WITH PADDING_ETC.\n\nmodel.eval();\nprevious_test_df = None;\nSTATE = {} # load your global dictionary which has let's say last 100 rows values you care for (which columns etc and \"key\" is the \"user_id\")\nTARGET = \"answered_correctly\"\nMAX_SEQ = 100\n\n\n# looping starts here.\nfor idx, (test_df, sample_prediction_df) in enumerate(iter_test):\n    # here you define your variables where you need to store things (optional)\n    after that there's the section where you have to update the sequences like content_id, answered_correctly, part etc.\n\n    if previous_test_df is not None:\n        # Here you just need to update the things you care for actually.\n        previous_test_df[TARGET] = eval(test_df[\"prior_group_answers_correct\"].iloc[0])\n        previous_test_df = previous_test_df[previous_test_df['content_type_id'] == 0].reset_index(drop=True)\n        # convert it to numpy\n        # update that user's state (here)\n\n    previous_test_df = test_df.copy()\n    test_df = test_df[test_df['content_type_id'] == 0].reset_index(drop=True) # drop videos, depends.\n    # fill-na section here.\n    # other stuffs you might think is important\n    # in the end, you convert your_test_df to numpy.\n    \ncurr = test_df.to_numpy()\n    for (cols_name <your's call what to call them>) in curr:\n       # get the user_id\n       # now there are two possible scenarios here, you have already seen this user or else you haven't seen this user till now.\n      # And again, even if you have seen a returning user, there are two cases to that as well, the seq_len already has last 100 seen things and the other case where it doesn't have that.\n        user_id = user_id\n        if user_id in STATE:\n            # split into two cases here depending on the length of the seq which i just talked\n            if _len_of_the_key >= MAX_SEQ:\n                # take the last 99 and stick the latest value at the end, you can get that from the loop itself.\n                .......\n            else:\n                 # handle the padding case here accordinly as you don't have MAX_SEQ elements in your keys.\n        else:\n               # now user_id is a new user\n               # add it's entry to the STATE. NB, add the keys you care about as well here\n               STATE[user_id] = {....}\n\n# looping ends here.\n# create the dataset/data-loader (if needed, i don't need it actually....)\n\n    with torch.no_grad():\n        # convert the dtypes accordingly, NB if use embeddings, ensure you have handled all the possible content_ids and padding index etc. (content_id is just an e.g.)\n\n        # with timer(\"inference here\"):\n        test_df['answered_correctly'] = model(.........).cpu() # convert them to CPU, match the shape here.\n        env.predict(test_df.loc[:, ['row_id', 'answered_correctly']])\n```\n\nThe above runs in around 2.15-3 hours for more than 5 subs successfully. \n\nThanks to others for inspiring me!\n\nHope the above helps, Feel free to let me know if anything is unclear.\n\nBest,\nAditya.\n\n---------------------------------------------------------------------\n\nEdits -:\n\nDo some whiteboard coding/ scribble on paper folks and think if you were a machine, what all you would expect to be taken care of by the USER operating you, the problem will reveal it's secrets... :)\n\nAlso, keep a pointer (Len) which indicates where will be the next update index basically. (Ensure you increment this and clip this to a max of MAX_SEQ  100 let's say); And this will be same for all values. Plus, don't pad the target when you prepare the cache for the inference part as well. And then you can related whatever is written above directly!\n\n\nPS, I personally don't have a great score with SAINT (yet) on LB (0.74), but the above is exactly what works for me, hence just shared to save the efforts",
      "votes": 50
    },
    {
      "id": 1107160,
      "postDate": "2020-12-09T12:59:49.993Z",
      "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> whats the best LB &amp; CV score you got using SAINT &amp; SAKT ?</p>",
      "rawMarkdown": "@adityaecdrid whats the best LB & CV score you got using SAINT & SAKT ?",
      "votes": 1,
      "replies": [
        {
          "id": 1107305,
          "postDate": "2020-12-09T15:23:25.327Z",
          "content": "<p>SAKT gave me close to .756 on LB. and SAINT gave me ~.74 on LB. (CV was around ~.76 and ~.766 respectively)</p>",
          "rawMarkdown": "SAKT gave me close to .756 on LB. and SAINT gave me ~.74 on LB. (CV was around ~.76 and ~.766 respectively)",
          "votes": 2
        },
        {
          "id": 1107326,
          "postDate": "2020-12-09T15:41:29.240Z",
          "content": "<p>Thanks, So your current LB score is not using transformers ? </p>",
          "rawMarkdown": "Thanks, So your current LB score is not using transformers ? ",
          "votes": 1
        },
        {
          "id": 1107438,
          "postDate": "2020-12-09T17:31:58.500Z",
          "content": "<p>It's with gbm's with hand-crafted FE's.</p>",
          "rawMarkdown": "It's with gbm's with hand-crafted FE's.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1098590,
      "postDate": "2020-12-01T18:04:00.757Z",
      "content": "<p>Nice <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>, keep improving, you will have your competitive score with SAINT soon!</p>",
      "rawMarkdown": "Nice @adityaecdrid, keep improving, you will have your competitive score with SAINT soon!",
      "votes": 1,
      "replies": [
        {
          "id": 1098628,
          "postDate": "2020-12-01T18:36:47.400Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> ! I hope so :)</p>",
          "rawMarkdown": "Thanks @claverru ! I hope so :)",
          "votes": 1
        },
        {
          "id": 1107435,
          "postDate": "2020-12-09T17:30:09.267Z",
          "content": "<p>So <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>, i have found some bugs with my target masking now; Add I feel that's the main problem as to why i had a less score till now!  I am feeling very sad for making such a blunder 🥺;  <br>\nVery excited to first fix and then try out again! </p>\n<p>Thanks for the wishes and inspirations (for yet another time!) :)</p>",
          "rawMarkdown": "So @claverru, i have found some bugs with my target masking now; Add I feel that's the main problem as to why i had a less score till now!  I am feeling very sad for making such a blunder 🥺;  \nVery excited to first fix and then try out again! \n\nThanks for the wishes and inspirations (for yet another time!) :)"
        }
      ]
    },
    {
      "id": 1138893,
      "postDate": "2021-01-05T03:40:27.940Z",
      "content": "<p>Have shared the inference snip <a href=\"https://www.kaggle.com/adityaecdrid/fork-of-saint-inference-ea970c\" target=\"_blank\">here</a>, it depicts the idea I shared above. It's buggy somewhere, so use it as a reference only.</p>\n<p>Thanks for all the help!</p>",
      "rawMarkdown": "Have shared the inference snip [here](https://www.kaggle.com/adityaecdrid/fork-of-saint-inference-ea970c), it depicts the idea I shared above. It's buggy somewhere, so use it as a reference only.\n\nThanks for all the help!",
      "votes": 2
    },
    {
      "id": 1103530,
      "postDate": "2020-12-06T02:23:55.527Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1104193,
          "postDate": "2020-12-06T17:37:23.557Z",
          "content": "<p>It's my pleasure!</p>",
          "rawMarkdown": "It's my pleasure!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1107160,
      "author_name": "RAHUL SINGH INDA",
      "author_url": "",
      "post_date": "2020-12-09T12:59:49.993000",
      "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> whats the best LB &amp; CV score you got using SAINT &amp; SAKT ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1107305,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-09T15:23:25.327000",
          "content": "<p>SAKT gave me close to .756 on LB. and SAINT gave me ~.74 on LB. (CV was around ~.76 and ~.766 respectively)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1107326,
          "author_name": "RAHUL SINGH INDA",
          "author_url": "",
          "post_date": "2020-12-09T15:41:29.240000",
          "content": "<p>Thanks, So your current LB score is not using transformers ? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107438,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-09T17:31:58.500000",
          "content": "<p>It's with gbm's with hand-crafted FE's.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1098590,
      "author_name": "Claudio Verdú Ruiz",
      "author_url": "",
      "post_date": "2020-12-01T18:04:00.757000",
      "content": "<p>Nice <a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a>, keep improving, you will have your competitive score with SAINT soon!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1098628,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-01T18:36:47.400000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a> ! I hope so :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1107435,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-09T17:30:09.267000",
          "content": "<p>So <a href=\"https://www.kaggle.com/claverru\" target=\"_blank\">@claverru</a>, i have found some bugs with my target masking now; Add I feel that's the main problem as to why i had a less score till now!  I am feeling very sad for making such a blunder 🥺;  <br>\nVery excited to first fix and then try out again! </p>\n<p>Thanks for the wishes and inspirations (for yet another time!) :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1138893,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2021-01-05T03:40:27.940000",
      "content": "<p>Have shared the inference snip <a href=\"https://www.kaggle.com/adityaecdrid/fork-of-saint-inference-ea970c\" target=\"_blank\">here</a>, it depicts the idea I shared above. It's buggy somewhere, so use it as a reference only.</p>\n<p>Thanks for all the help!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1103530,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-06T02:23:55.527000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1104193,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-06T17:37:23.557000",
          "content": "<p>It's my pleasure!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1095410": "Hey Earthlings!\n\nI know writing inference pipeline is a pain and it's frustrating at times as well. But i hope the below will be useful. \n\nNB, Please spend a day writing your own. It's quite important actually to do it yourself. What matter's is your attempt, not the results.\n\nWithout further delay, here it's\n\n```\n# the below is pretty robust and generic and works for FIXED_QUESTION_SEQ_LENGTHS_WITH PADDING_ETC.\n\nmodel.eval();\nprevious_test_df = None;\nSTATE = {} # load your global dictionary which has let's say last 100 rows values you care for (which columns etc and \"key\" is the \"user_id\")\nTARGET = \"answered_correctly\"\nMAX_SEQ = 100\n\n\n# looping starts here.\nfor idx, (test_df, sample_prediction_df) in enumerate(iter_test):\n    # here you define your variables where you need to store things (optional)\n    after that there's the section where you have to update the sequences like content_id, answered_correctly, part etc.\n\n    if previous_test_df is not None:\n        # Here you just need to update the things you care for actually.\n        previous_test_df[TARGET] = eval(test_df[\"prior_group_answers_correct\"].iloc[0])\n        previous_test_df = previous_test_df[previous_test_df['content_type_id'] == 0].reset_index(drop=True)\n        # convert it to numpy\n        # update that user's state (here)\n\n    previous_test_df = test_df.copy()\n    test_df = test_df[test_df['content_type_id'] == 0].reset_index(drop=True) # drop videos, depends.\n    # fill-na section here.\n    # other stuffs you might think is important\n    # in the end, you convert your_test_df to numpy.\n    \ncurr = test_df.to_numpy()\n    for (cols_name <your's call what to call them>) in curr:\n       # get the user_id\n       # now there are two possible scenarios here, you have already seen this user or else you haven't seen this user till now.\n      # And again, even if you have seen a returning user, there are two cases to that as well, the seq_len already has last 100 seen things and the other case where it doesn't have that.\n        user_id = user_id\n        if user_id in STATE:\n            # split into two cases here depending on the length of the seq which i just talked\n            if _len_of_the_key >= MAX_SEQ:\n                # take the last 99 and stick the latest value at the end, you can get that from the loop itself.\n                .......\n            else:\n                 # handle the padding case here accordinly as you don't have MAX_SEQ elements in your keys.\n        else:\n               # now user_id is a new user\n               # add it's entry to the STATE. NB, add the keys you care about as well here\n               STATE[user_id] = {....}\n\n# looping ends here.\n# create the dataset/data-loader (if needed, i don't need it actually....)\n\n    with torch.no_grad():\n        # convert the dtypes accordingly, NB if use embeddings, ensure you have handled all the possible content_ids and padding index etc. (content_id is just an e.g.)\n\n        # with timer(\"inference here\"):\n        test_df['answered_correctly'] = model(.........).cpu() # convert them to CPU, match the shape here.\n        env.predict(test_df.loc[:, ['row_id', 'answered_correctly']])\n```\n\nThe above runs in around 2.15-3 hours for more than 5 subs successfully. \n\nThanks to others for inspiring me!\n\nHope the above helps, Feel free to let me know if anything is unclear.\n\nBest,\nAditya.\n\n---------------------------------------------------------------------\n\nEdits -:\n\nDo some whiteboard coding/ scribble on paper folks and think if you were a machine, what all you would expect to be taken care of by the USER operating you, the problem will reveal it's secrets... :)\n\nAlso, keep a pointer (Len) which indicates where will be the next update index basically. (Ensure you increment this and clip this to a max of MAX_SEQ  100 let's say); And this will be same for all values. Plus, don't pad the target when you prepare the cache for the inference part as well. And then you can related whatever is written above directly!\n\n\nPS, I personally don't have a great score with SAINT (yet) on LB (0.74), but the above is exactly what works for me, hence just shared to save the efforts",
    "1107160": "@adityaecdrid whats the best LB & CV score you got using SAINT & SAKT ?",
    "1098590": "Nice @adityaecdrid, keep improving, you will have your competitive score with SAINT soon!",
    "1138893": "Have shared the inference snip [here](https://www.kaggle.com/adityaecdrid/fork-of-saint-inference-ea970c), it depicts the idea I shared above. It's buggy somewhere, so use it as a reference only.\n\nThanks for all the help!",
    "1103530": ""
  }
}