{
  "id": 205662,
  "title": "Submission Scoring Error Others Reasons Explained & fix",
  "url": "/competitions/riiid-test-answer-prediction/discussion/205662",
  "author_name": "",
  "post_date": "2020-12-21T08:42:28.726078400Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Recently i encountered back to back submission scoring ..here are possible reasons based on which i put a fix.</p>\n<p>1) If you are adding a new meta data related feature based question to your transformer like ,lectures,parts ..and if you get sub scoring error ,then it could quite be possible that your embedding layer throwing Out of index error because test data is giving an index higher than what you defined for embedding layer<br>\nFix : Clip metadata value  to be in limit of what model is trained</p>\n<p>2) OOM =Out of memory reason. If you are using merge for test data to join with the meta data  of question,lecture etc , use copy=False  which might save some crucial memory </p>\n<p>3)  Do left join and assign meta feature null  values to default values. </p>",
  "messages": [
    {
      "id": "1120959",
      "postDate": "12/21/2020 08:42:28",
      "content": "<p>Recently i encountered back to back submission scoring ..here are possible reasons based on which i put a fix.</p>\n<p>1) If you are adding a new meta data related feature based question to your transformer like ,lectures,parts ..and if you get sub scoring error ,then it could quite be possible that your embedding layer throwing Out of index error because test data is giving an index higher than what you defined for embedding layer<br>\nFix : Clip metadata value  to be in limit of what model is trained</p>\n<p>2) OOM =Out of memory reason. If you are using merge for test data to join with the meta data  of question,lecture etc , use copy=False  which might save some crucial memory </p>\n<p>3)  Do left join and assign meta feature null  values to default values. </p>",
      "rawMarkdown": "Recently i encountered back to back submission scoring ..here are possible reasons based on which i put a fix.\n\n1) If you are adding a new meta data related feature based question to your transformer like ,lectures,parts ..and if you get sub scoring error ,then it could quite be possible that your embedding layer throwing Out of index error because test data is giving an index higher than what you defined for embedding layer\nFix : Clip metadata value  to be in limit of what model is trained\n\n2) OOM =Out of memory reason. If you are using merge for test data to join with the meta data  of question,lecture etc , use copy=False  which might save some crucial memory \n\n3)  Do left join and assign meta feature null  values to default values.",
      "votes": null
    },
    {
      "id": "1121905",
      "postDate": "12/22/2020 02:34:42",
      "content": "<p>Do you have any tips for diagnosing the Submission Scoring Error?</p>\n<p>I'm using a really simple (pre-trained) model. My model data takes &lt; 100mb of RAM. And the sample <code>submission.csv</code> that I generate has exactly the same format as when I just use <code>sample_prediction_df</code>.</p>",
      "rawMarkdown": "Do you have any tips for diagnosing the Submission Scoring Error?\n\nI'm using a really simple (pre-trained) model. My model data takes < 100mb of RAM. And the sample `submission.csv` that I generate has exactly the same format as when I just use `sample_prediction_df`.",
      "votes": null
    },
    {
      "id": "1123531",
      "postDate": "12/23/2020 09:57:33",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> For third point, if the new feature is part, because the question in hidden test must include in train.csv, so is this necessary for add part feature in transform?  i just use <br>\n'<br>\n    # test df insert new part_id <br>\n    test_df['part_id'] = test_df['content_id'].map(part_ids_map)<br>\n' <br>\nDo you think this will result in the error? can you give more info plz? thank you</p>",
      "rawMarkdown": "jaideepvalani For third point, if the new feature is part, because the question in hidden test must include in train.csv, so is this necessary for add part feature in transform?  i just use \n'\n    # test df insert new part_id \n    test_df['part_id'] = test_df['content_id'].map(part_ids_map)\n' \nDo you think this will result in the error? can you give more info plz? thank you",
      "votes": null
    },
    {
      "id": "1123641",
      "postDate": "12/23/2020 11:46:03",
      "content": "<p>this one will not give error..<br>\nbut if we use merge it gives error.. unable to diagnose why..</p>",
      "rawMarkdown": "this one will not give error..\nbut if we use merge it gives error.. unable to diagnose why..",
      "votes": null
    },
    {
      "id": "1123663",
      "postDate": "12/23/2020 12:08:33",
      "content": "<p>I suspect one point, but i have no times today to test, maybe you can try, please set batch_size low, such as 512, not use too big batch size when you use real saint do submission</p>",
      "rawMarkdown": "I suspect one point, but i have no times today to test, maybe you can try, please set batch_size low, such as 512, not use too big batch size when you use real saint do submission",
      "votes": null
    },
    {
      "id": "1124638",
      "postDate": "12/24/2020 05:31:58",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>  finally i find the bug, previous group response include the sub item which is lecture, solve by delete for my issue, submitting</p>",
      "rawMarkdown": "jaideepvalani  finally i find the bug, previous group response include the sub item which is lecture, solve by delete for my issue, submitting",
      "votes": null
    },
    {
      "id": "1124673",
      "postDate": "12/24/2020 06:06:19",
      "content": "<p>may be you are right ,so far i have been getting this fine using dictionary method as you shown above.. i included now one more feature related to the question in same way.. this led to error again :(  not sure if it was because of batch size.. solving blind issues are cumbersome..<br>\n<a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <br>\n I dont know when will kaggle team can feel pity on us  and provide atleast some messages that can provide more info on what is underlying cause that results into Sub scoring error/kernel errors. <br>\nIn almost Every competition of this type there are so many discussion revolving around this same topic, lot of time waste, resources waste</p>\n<p>**Update: batch size was the issue this time. **</p>",
      "rawMarkdown": "may be you are right ,so far i have been getting this fine using dictionary method as you shown above.. i included now one more feature related to the question in same way.. this led to error again :(  not sure if it was because of batch size.. solving blind issues are cumbersome..\n@addisonhoward \n I dont know when will kaggle team can feel pity on us  and provide atleast some messages that can provide more info on what is underlying cause that results into Sub scoring error/kernel errors. \nIn almost Every competition of this type there are so many discussion revolving around this same topic, lot of time waste, resources waste\n\n**Update: batch size was the issue this time. **",
      "votes": null
    },
    {
      "id": "1124694",
      "postDate": "12/24/2020 06:24:35",
      "content": "<pre><code>here is my loop\n\nfor (test_df, sample_prediction_df) in iter_test:\n\n        #test_df=pd.merge(test_df,questions_df[['question_id','part']],\n        #                  left_on='content_id',right_on='question_id',copy=False)\n        test_df['part'] =test_df['content_id'].map(part_ids_map) \n        test_df['bundle_part'] =test_df['content_id'].map(bundle_question_dict) \n        #test_df['part'] =test_df['content_id'].map(lambda x:  part_ids_map.loc[x][0])\n        #test_df.loc[test_df.part.isnull(),'part']=5\n\n        if (prev_test_df is not None) &amp; (psutil.virtual_memory().percent&lt;90):\n            #print(psutil.virtual_memory().percent)\n            prev_test_df['answered_correctly'] = eval(test_df['prior_group_answers_correct'].iloc[0])\n            prev_test_df = prev_test_df[prev_test_df.content_type_id == False]\n            prev_group = prev_test_df[['user_id', 'content_id', 'answered_correctly','part','bundle_part']].groupby('user_id').apply(lambda r: (\n                r['content_id'].values,\n                r['answered_correctly'].values ,\n                r['part'].values ,r['bundle_part'].values                                \n            ))\n            for prev_user_id in prev_group.index:\n                prev_group_content = prev_group[prev_user_id][0] #get for each user added q,ans correctly\n                prev_group_ac = prev_group[prev_user_id][1]\n                prev_group_part=prev_group[prev_user_id][2]\n                prev_group_b_part=prev_group[prev_user_id][3]\n\n                if prev_user_id in valid_group.index: # if present in our train db ,then append the list\n                    valid_group[prev_user_id] = (np.append(valid_group[prev_user_id][0],\n                                                           prev_group_content), \n                                           np.append(valid_group[prev_user_id][1],prev_group_ac)\n                                           ,np.append(valid_group[prev_user_id][2],prev_group_part )\n                                                ,np.append(valid_group[prev_user_id][3],prev_group_b_part ))\n\n                else: # else add a new user to db this updating is useful for next time when user id comes in\n                    valid_group[prev_user_id] = (prev_group_content,prev_group_ac,prev_group_part,prev_group_b_part)\n\n                if len(valid_group[prev_user_id][0])&gt;MAX_SEQ:\n                    new_group_content = valid_group[prev_user_id][0][-MAX_SEQ:]\n                    new_group_ac = valid_group[prev_user_id][1][-MAX_SEQ:]\n                    new_group_prt = valid_group[prev_user_id][2][-MAX_SEQ:]\n                    new_group_prt_b = valid_group[prev_user_id][3][-MAX_SEQ:]\n\n                    valid_group[prev_user_id] = (new_group_content,new_group_ac,new_group_prt,\n                                                 new_group_prt_b)\n\n\n        prev_test_df = test_df.copy() \n        test_df = test_df[test_df.content_type_id == 0]\n\n\n\n        #get unique user ids\n        test_user_ids=test_df.user_id.unique().tolist()\n        test_dataset = TestDataset(valid_group, test_df, n_skill,\n                                   max_seq=conf.max_seq)\n        test_dataloader = DataLoader(test_dataset, batch_size=conf.BATCH_SIZE, \n                                     shuffle=False, drop_last=False)\n\n        outs = []\n\n\n        for item in test_dataloader:\n\n            target_id= item[0].to(device).long()\n\n            rt=item[1].to(device).long()\n\n            tag_emb=item[2].to(device).long()\n            b_part=item[3].to(device).long()\n\n            with torch.no_grad():\n                output  = model(target_id,rt, tag_emb,b_part)\n\n\n                output = torch.sigmoid(output)\n                output = output[:, -1]\n\n                outs.extend(output.view(-1).data.cpu().numpy())\n\n        test_df['answered_correctly'] = outs\n\n        env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n    #print('predicted')\n        del test_df\n        gc.collect()\n</code></pre>",
      "rawMarkdown": "```\nhere is my loop\n\nfor (test_df, sample_prediction_df) in iter_test:\n        \n        #test_df=pd.merge(test_df,questions_df[['question_id','part']],\n        #                  left_on='content_id',right_on='question_id',copy=False)\n        test_df['part'] =test_df['content_id'].map(part_ids_map) \n        test_df['bundle_part'] =test_df['content_id'].map(bundle_question_dict) \n        #test_df['part'] =test_df['content_id'].map(lambda x:  part_ids_map.loc[x][0])\n        #test_df.loc[test_df.part.isnull(),'part']=5\n        \n        if (prev_test_df is not None) & (psutil.virtual_memory().percent<90):\n            #print(psutil.virtual_memory().percent)\n            prev_test_df['answered_correctly'] = eval(test_df['prior_group_answers_correct'].iloc[0])\n            prev_test_df = prev_test_df[prev_test_df.content_type_id == False]\n            prev_group = prev_test_df[['user_id', 'content_id', 'answered_correctly','part','bundle_part']].groupby('user_id').apply(lambda r: (\n                r['content_id'].values,\n                r['answered_correctly'].values ,\n                r['part'].values ,r['bundle_part'].values                                \n            ))\n            for prev_user_id in prev_group.index:\n                prev_group_content = prev_group[prev_user_id][0] #get for each user added q,ans correctly\n                prev_group_ac = prev_group[prev_user_id][1]\n                prev_group_part=prev_group[prev_user_id][2]\n                prev_group_b_part=prev_group[prev_user_id][3]\n                \n                if prev_user_id in valid_group.index: # if present in our train db ,then append the list\n                    valid_group[prev_user_id] = (np.append(valid_group[prev_user_id][0],\n                                                           prev_group_content), \n                                           np.append(valid_group[prev_user_id][1],prev_group_ac)\n                                           ,np.append(valid_group[prev_user_id][2],prev_group_part )\n                                                ,np.append(valid_group[prev_user_id][3],prev_group_b_part ))\n\n                else: # else add a new user to db this updating is useful for next time when user id comes in\n                    valid_group[prev_user_id] = (prev_group_content,prev_group_ac,prev_group_part,prev_group_b_part)\n                \n                if len(valid_group[prev_user_id][0])>MAX_SEQ:\n                    new_group_content = valid_group[prev_user_id][0][-MAX_SEQ:]\n                    new_group_ac = valid_group[prev_user_id][1][-MAX_SEQ:]\n                    new_group_prt = valid_group[prev_user_id][2][-MAX_SEQ:]\n                    new_group_prt_b = valid_group[prev_user_id][3][-MAX_SEQ:]\n                                                 \n                    valid_group[prev_user_id] = (new_group_content,new_group_ac,new_group_prt,\n                                                 new_group_prt_b)\n\n       \n        prev_test_df = test_df.copy() \n        test_df = test_df[test_df.content_type_id == 0]\n\n       \n        \n        #get unique user ids\n        test_user_ids=test_df.user_id.unique().tolist()\n        test_dataset = TestDataset(valid_group, test_df, n_skill,\n                                   max_seq=conf.max_seq)\n        test_dataloader = DataLoader(test_dataset, batch_size=conf.BATCH_SIZE, \n                                     shuffle=False, drop_last=False)\n\n        outs = []\n         \n\n        for item in test_dataloader:\n            \n            target_id= item[0].to(device).long()\n            \n            rt=item[1].to(device).long()\n            \n            tag_emb=item[2].to(device).long()\n            b_part=item[3].to(device).long()\n            \n            with torch.no_grad():\n                output  = model(target_id,rt, tag_emb,b_part)\n                 \n\n                output = torch.sigmoid(output)\n                output = output[:, -1]\n\n                outs.extend(output.view(-1).data.cpu().numpy())\n\n        test_df['answered_correctly'] = outs\n\n        env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n    #print('predicted')\n        del test_df\n        gc.collect()\n```",
      "votes": null
    },
    {
      "id": "1124706",
      "postDate": "12/24/2020 06:41:36",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>  i suggest you firstly try emulator with <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">Emulator</a> i find my bug with this NB</p>",
      "rawMarkdown": "jaideepvalani  i suggest you firstly try emulator with [Emulator](https://www.kaggle.com/its7171/time-series-api-iter-test-emulator) i find my bug with this NB",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1121905,
      "author_name": "rickdz",
      "author_url": "",
      "post_date": "12/22/2020 02:34:42",
      "content": "<p>Do you have any tips for diagnosing the Submission Scoring Error?</p>\n<p>I'm using a really simple (pre-trained) model. My model data takes &lt; 100mb of RAM. And the sample <code>submission.csv</code> that I generate has exactly the same format as when I just use <code>sample_prediction_df</code>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1123531,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "12/23/2020 09:57:33",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> For third point, if the new feature is part, because the question in hidden test must include in train.csv, so is this necessary for add part feature in transform?  i just use <br>\n'<br>\n    # test df insert new part_id <br>\n    test_df['part_id'] = test_df['content_id'].map(part_ids_map)<br>\n' <br>\nDo you think this will result in the error? can you give more info plz? thank you</p>",
      "votes": null,
      "replies": [
        {
          "id": 1123641,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/23/2020 11:46:03",
          "content": "<p>this one will not give error..<br>\nbut if we use merge it gives error.. unable to diagnose why..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1123663,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "12/23/2020 12:08:33",
          "content": "<p>I suspect one point, but i have no times today to test, maybe you can try, please set batch_size low, such as 512, not use too big batch size when you use real saint do submission</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124638,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "12/24/2020 05:31:58",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>  finally i find the bug, previous group response include the sub item which is lecture, solve by delete for my issue, submitting</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124673,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/24/2020 06:06:19",
          "content": "<p>may be you are right ,so far i have been getting this fine using dictionary method as you shown above.. i included now one more feature related to the question in same way.. this led to error again :(  not sure if it was because of batch size.. solving blind issues are cumbersome..<br>\n<a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <br>\n I dont know when will kaggle team can feel pity on us  and provide atleast some messages that can provide more info on what is underlying cause that results into Sub scoring error/kernel errors. <br>\nIn almost Every competition of this type there are so many discussion revolving around this same topic, lot of time waste, resources waste</p>\n<p>**Update: batch size was the issue this time. **</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124694,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/24/2020 06:24:35",
          "content": "<pre><code>here is my loop\n\nfor (test_df, sample_prediction_df) in iter_test:\n\n        #test_df=pd.merge(test_df,questions_df[['question_id','part']],\n        #                  left_on='content_id',right_on='question_id',copy=False)\n        test_df['part'] =test_df['content_id'].map(part_ids_map) \n        test_df['bundle_part'] =test_df['content_id'].map(bundle_question_dict) \n        #test_df['part'] =test_df['content_id'].map(lambda x:  part_ids_map.loc[x][0])\n        #test_df.loc[test_df.part.isnull(),'part']=5\n\n        if (prev_test_df is not None) &amp; (psutil.virtual_memory().percent&lt;90):\n            #print(psutil.virtual_memory().percent)\n            prev_test_df['answered_correctly'] = eval(test_df['prior_group_answers_correct'].iloc[0])\n            prev_test_df = prev_test_df[prev_test_df.content_type_id == False]\n            prev_group = prev_test_df[['user_id', 'content_id', 'answered_correctly','part','bundle_part']].groupby('user_id').apply(lambda r: (\n                r['content_id'].values,\n                r['answered_correctly'].values ,\n                r['part'].values ,r['bundle_part'].values                                \n            ))\n            for prev_user_id in prev_group.index:\n                prev_group_content = prev_group[prev_user_id][0] #get for each user added q,ans correctly\n                prev_group_ac = prev_group[prev_user_id][1]\n                prev_group_part=prev_group[prev_user_id][2]\n                prev_group_b_part=prev_group[prev_user_id][3]\n\n                if prev_user_id in valid_group.index: # if present in our train db ,then append the list\n                    valid_group[prev_user_id] = (np.append(valid_group[prev_user_id][0],\n                                                           prev_group_content), \n                                           np.append(valid_group[prev_user_id][1],prev_group_ac)\n                                           ,np.append(valid_group[prev_user_id][2],prev_group_part )\n                                                ,np.append(valid_group[prev_user_id][3],prev_group_b_part ))\n\n                else: # else add a new user to db this updating is useful for next time when user id comes in\n                    valid_group[prev_user_id] = (prev_group_content,prev_group_ac,prev_group_part,prev_group_b_part)\n\n                if len(valid_group[prev_user_id][0])&gt;MAX_SEQ:\n                    new_group_content = valid_group[prev_user_id][0][-MAX_SEQ:]\n                    new_group_ac = valid_group[prev_user_id][1][-MAX_SEQ:]\n                    new_group_prt = valid_group[prev_user_id][2][-MAX_SEQ:]\n                    new_group_prt_b = valid_group[prev_user_id][3][-MAX_SEQ:]\n\n                    valid_group[prev_user_id] = (new_group_content,new_group_ac,new_group_prt,\n                                                 new_group_prt_b)\n\n\n        prev_test_df = test_df.copy() \n        test_df = test_df[test_df.content_type_id == 0]\n\n\n\n        #get unique user ids\n        test_user_ids=test_df.user_id.unique().tolist()\n        test_dataset = TestDataset(valid_group, test_df, n_skill,\n                                   max_seq=conf.max_seq)\n        test_dataloader = DataLoader(test_dataset, batch_size=conf.BATCH_SIZE, \n                                     shuffle=False, drop_last=False)\n\n        outs = []\n\n\n        for item in test_dataloader:\n\n            target_id= item[0].to(device).long()\n\n            rt=item[1].to(device).long()\n\n            tag_emb=item[2].to(device).long()\n            b_part=item[3].to(device).long()\n\n            with torch.no_grad():\n                output  = model(target_id,rt, tag_emb,b_part)\n\n\n                output = torch.sigmoid(output)\n                output = output[:, -1]\n\n                outs.extend(output.view(-1).data.cpu().numpy())\n\n        test_df['answered_correctly'] = outs\n\n        env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n    #print('predicted')\n        del test_df\n        gc.collect()\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124706,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "12/24/2020 06:41:36",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>  i suggest you firstly try emulator with <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">Emulator</a> i find my bug with this NB</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1120959": "Recently i encountered back to back submission scoring ..here are possible reasons based on which i put a fix.\n\n1) If you are adding a new meta data related feature based question to your transformer like ,lectures,parts ..and if you get sub scoring error ,then it could quite be possible that your embedding layer throwing Out of index error because test data is giving an index higher than what you defined for embedding layer\nFix : Clip metadata value  to be in limit of what model is trained\n\n2) OOM =Out of memory reason. If you are using merge for test data to join with the meta data  of question,lecture etc , use copy=False  which might save some crucial memory \n\n3)  Do left join and assign meta feature null  values to default values.",
    "1121905": "Do you have any tips for diagnosing the Submission Scoring Error?\n\nI'm using a really simple (pre-trained) model. My model data takes < 100mb of RAM. And the sample `submission.csv` that I generate has exactly the same format as when I just use `sample_prediction_df`.",
    "1123531": "jaideepvalani For third point, if the new feature is part, because the question in hidden test must include in train.csv, so is this necessary for add part feature in transform?  i just use \n'\n    # test df insert new part_id \n    test_df['part_id'] = test_df['content_id'].map(part_ids_map)\n' \nDo you think this will result in the error? can you give more info plz? thank you",
    "1123641": "this one will not give error..\nbut if we use merge it gives error.. unable to diagnose why..",
    "1123663": "I suspect one point, but i have no times today to test, maybe you can try, please set batch_size low, such as 512, not use too big batch size when you use real saint do submission",
    "1124638": "jaideepvalani  finally i find the bug, previous group response include the sub item which is lecture, solve by delete for my issue, submitting",
    "1124673": "may be you are right ,so far i have been getting this fine using dictionary method as you shown above.. i included now one more feature related to the question in same way.. this led to error again :(  not sure if it was because of batch size.. solving blind issues are cumbersome..\n@addisonhoward \n I dont know when will kaggle team can feel pity on us  and provide atleast some messages that can provide more info on what is underlying cause that results into Sub scoring error/kernel errors. \nIn almost Every competition of this type there are so many discussion revolving around this same topic, lot of time waste, resources waste\n\n**Update: batch size was the issue this time. **",
    "1124694": "```\nhere is my loop\n\nfor (test_df, sample_prediction_df) in iter_test:\n        \n        #test_df=pd.merge(test_df,questions_df[['question_id','part']],\n        #                  left_on='content_id',right_on='question_id',copy=False)\n        test_df['part'] =test_df['content_id'].map(part_ids_map) \n        test_df['bundle_part'] =test_df['content_id'].map(bundle_question_dict) \n        #test_df['part'] =test_df['content_id'].map(lambda x:  part_ids_map.loc[x][0])\n        #test_df.loc[test_df.part.isnull(),'part']=5\n        \n        if (prev_test_df is not None) & (psutil.virtual_memory().percent<90):\n            #print(psutil.virtual_memory().percent)\n            prev_test_df['answered_correctly'] = eval(test_df['prior_group_answers_correct'].iloc[0])\n            prev_test_df = prev_test_df[prev_test_df.content_type_id == False]\n            prev_group = prev_test_df[['user_id', 'content_id', 'answered_correctly','part','bundle_part']].groupby('user_id').apply(lambda r: (\n                r['content_id'].values,\n                r['answered_correctly'].values ,\n                r['part'].values ,r['bundle_part'].values                                \n            ))\n            for prev_user_id in prev_group.index:\n                prev_group_content = prev_group[prev_user_id][0] #get for each user added q,ans correctly\n                prev_group_ac = prev_group[prev_user_id][1]\n                prev_group_part=prev_group[prev_user_id][2]\n                prev_group_b_part=prev_group[prev_user_id][3]\n                \n                if prev_user_id in valid_group.index: # if present in our train db ,then append the list\n                    valid_group[prev_user_id] = (np.append(valid_group[prev_user_id][0],\n                                                           prev_group_content), \n                                           np.append(valid_group[prev_user_id][1],prev_group_ac)\n                                           ,np.append(valid_group[prev_user_id][2],prev_group_part )\n                                                ,np.append(valid_group[prev_user_id][3],prev_group_b_part ))\n\n                else: # else add a new user to db this updating is useful for next time when user id comes in\n                    valid_group[prev_user_id] = (prev_group_content,prev_group_ac,prev_group_part,prev_group_b_part)\n                \n                if len(valid_group[prev_user_id][0])>MAX_SEQ:\n                    new_group_content = valid_group[prev_user_id][0][-MAX_SEQ:]\n                    new_group_ac = valid_group[prev_user_id][1][-MAX_SEQ:]\n                    new_group_prt = valid_group[prev_user_id][2][-MAX_SEQ:]\n                    new_group_prt_b = valid_group[prev_user_id][3][-MAX_SEQ:]\n                                                 \n                    valid_group[prev_user_id] = (new_group_content,new_group_ac,new_group_prt,\n                                                 new_group_prt_b)\n\n       \n        prev_test_df = test_df.copy() \n        test_df = test_df[test_df.content_type_id == 0]\n\n       \n        \n        #get unique user ids\n        test_user_ids=test_df.user_id.unique().tolist()\n        test_dataset = TestDataset(valid_group, test_df, n_skill,\n                                   max_seq=conf.max_seq)\n        test_dataloader = DataLoader(test_dataset, batch_size=conf.BATCH_SIZE, \n                                     shuffle=False, drop_last=False)\n\n        outs = []\n         \n\n        for item in test_dataloader:\n            \n            target_id= item[0].to(device).long()\n            \n            rt=item[1].to(device).long()\n            \n            tag_emb=item[2].to(device).long()\n            b_part=item[3].to(device).long()\n            \n            with torch.no_grad():\n                output  = model(target_id,rt, tag_emb,b_part)\n                 \n\n                output = torch.sigmoid(output)\n                output = output[:, -1]\n\n                outs.extend(output.view(-1).data.cpu().numpy())\n\n        test_df['answered_correctly'] = outs\n\n        env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n    #print('predicted')\n        del test_df\n        gc.collect()\n```",
    "1124706": "jaideepvalani  i suggest you firstly try emulator with [Emulator](https://www.kaggle.com/its7171/time-series-api-iter-test-emulator) i find my bug with this NB"
  },
  "source": "meta"
}