{
  "id": 190748,
  "title": "Understanding prior_group_answers_correct and prior_group_responses columns in the test data",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190748",
  "author_name": "",
  "post_date": "2020-10-13T07:05:20.401845100Z",
  "votes": 24,
  "comment_count": 3,
  "views": 0,
  "content": "<p>NB I am showing it only for first few rows, i have verified the same follows for the rest as well. (at least for the sample test rows)</p>\n<pre><code>row_id,group_num,timestamp,user_id,content_id,content_type_id,task_container_id,prior_question_elapsed_time,prior_question_had_explanation,prior_group_answers_correct,prior_group_responses\n0,0,0,275030867,5729,0,0,,,[],[]\n1,0,13309898705,554169193,12010,0,4427,19000.0,True,,\n2,0,4213672059,1720860329,457,0,240,17000.0,True,,\n3,0,62798072960,288641214,13262,0,266,23000.0,True,,\n4,0,10585422061,1728340777,6119,0,162,72400.0,True,,\n5,0,18020362258,1364159702,12023,0,4424,18000.0,True,,\n6,0,2325432079,1521618396,574,0,1367,18000.0,True,,\n7,0,39456940781,1317245193,12043,0,5314,17000.0,True,,\n8,0,3460555189,1700555100,7910,0,532,21000.0,True,,\n9,0,2214770464,998511398,7908,0,393,21000.0,True,,\n10,0,516803182,1422853669,1143,0,85,15000.0,True,,\n11,0,2153839851,1096784725,11033,0,315,34250.0,True,,\n12,0,2153839851,1096784725,11032,0,315,34250.0,True,,\n13,0,2153839851,1096784725,11034,0,315,34250.0,True,,\n14,0,2153839851,1096784725,11031,0,315,34250.0,True,,\n15,0,1218852591,385471210,9538,0,378,11000.0,True,,\n16,0,32722340115,1202386221,1002,0,136,16000.0,True,,\n17,0,2059097926,2018567473,12148,0,589,17000.0,True,,\n18,1,23609,275030867,5502,0,1,34000.0,False,\"[0, 1, 1, 0, 0, 1, 0, 1, 0, 1, 1, 1, 1, 1, 0, 1, 0, 1]\",\"[0, 0, 1, 1, 0, 1, 3, 3, 1, 1, 0, 3, 1, 2, 2, 1, 0, 3]\"\n</code></pre>\n<p>Above is the sample data from the file example_test.csv. </p>\n<p>Focus on the group_num as 0, you will see 18 records there. [0-17]</p>\n<p>Now focus just on the very next row where group_num 1 starts, you will find the list contains 18 elements in it. <br>\nWhich coincides with the row counts of the previous group_num you have seen..</p>\n<p>And the same follows for the next group num.</p>\n<pre><code>18,1,23609,275030867,5502,0,1,34000.0,False,\"[0, 1, 1, 0, 0, 1, 0, 1, 0, 1, 1, 1, 1, 1, 0, 1, 0, 1]\",\"[0, 0, 1, 1, 0, 1, 3, 3, 1, 1, 0, 3, 1, 2, 2, 1, 0, 3]\"\n19,1,2035159380,1233875513,1512,0,1431,25000.0,True,,\n20,1,2035159380,1233875513,1511,0,1431,25000.0,True,,\n21,1,2035159380,1233875513,1513,0,1431,25000.0,True,,\n22,1,1217231,891955351,9145,0,20,23000.0,True,,\n23,1,11265012636,1981166446,3299,0,950,36333.0,True,,\n24,1,11265012636,1981166446,3297,0,950,36333.0,True,,\n25,1,11265012636,1981166446,3298,0,950,36333.0,True,,\n26,1,4693145319,1637273633,11373,0,3148,41000.0,True,,\n27,1,2294633294,2030979309,3348,0,170,30666.0,True,,\n28,1,2294633294,2030979309,3350,0,170,30666.0,True,,\n29,1,2294633294,2030979309,3349,0,170,30666.0,True,,\n30,1,13679981387,319060572,10951,0,621,30250.0,True,,\n31,1,13679981387,319060572,10953,0,621,30250.0,True,,\n32,1,13679981387,319060572,10954,0,621,30250.0,True,,\n33,1,13679981387,319060572,10952,0,621,30250.0,True,,\n34,1,62798100988,288641214,5418,0,267,24000.0,True,,\n35,1,251107302,98059812,5892,0,9,34000.0,True,,\n37,1,1254086842,674533997,5301,0,1044,29000.0,True,,\n38,1,44526338636,555691277,666,0,52,25000.0,True,,\n39,1,39456965438,1317245193,12195,0,5315,18000.0,True,,\n40,1,32722359498,1202386221,582,0,137,17000.0,True,,\n41,1,417208826,775113212,10078,0,128,28250.0,True,,\n42,1,417208826,775113212,10080,0,128,28250.0,True,,\n43,1,417208826,775113212,10081,0,128,28250.0,True,,\n44,1,417208826,775113212,10079,0,128,28250.0,True,,\n45,1,209087981,1219481379,3784,0,40,22000.0,True,,\n46,2,42512,275030867,6133,0,2,17000.0,False,\"[1, 0, 1, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 1, 1, 0, 1, 1, 1, 0, 1]\",\"[1, 2, 2, 2, 1, 0, 1, 2, 1, 1, 3, 2, 0, 0, 3, 1, 1, 3, 3, 3, 0, 0, 0, 1, 2, 0, 3]\"\n</code></pre>\n<p>At the row no 46th, the list has 27 elements in it which coincides with the row count of the group_num 1.</p>\n<p>So this is a great information in case people have missed till now.<br>\nPlease correct my understandings.</p>\n<pre><code>prior_group_responses (string) provides all of the user_answer entries for previous group in a string representation of a list in the first row of the group. All other rows in each group are null. If you are using Python, you will likely want to call eval on the non-null rows. Some rows may be null, or empty lists.\n\nprior_group_answers_correct (string) provides all the answered_correctly field for previous group, with the same format and caveats as prior_group_responses. Some rows may be null, or empty lists.\n</code></pre>\n<p>So we have the actual correct labels (as in the test set) in the column \"prior_group_responses\" and whether they were correct or not by the user's response in the column \"prior_group_answers_correct\".. i.e. the user thought option 0 is correct but actually it was option 1 which was the correct asnwer.</p>\n<p>So it seems we can test set to train our models? I mean we won't we doing update to our model for every batch, we will do it when we have k batches \"combined\"…</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190430#1046574\" target=\"_blank\">Credits to the comment by </a><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a></p>\n<p>Looking forward to hear your thoughts folks! <br>\nThanks!</p>\n<p>Edit - 1</p>\n<ul>\n<li>Updated the same with chunking df's and then updating the stats etc with a successful submission. <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856\" target=\"_blank\">link</a></li>\n</ul>",
  "messages": [
    {
      "id": "1048082",
      "postDate": "10/13/2020 07:05:20",
      "content": "<p>NB I am showing it only for first few rows, i have verified the same follows for the rest as well. (at least for the sample test rows)</p>\n<pre><code>row_id,group_num,timestamp,user_id,content_id,content_type_id,task_container_id,prior_question_elapsed_time,prior_question_had_explanation,prior_group_answers_correct,prior_group_responses\n0,0,0,275030867,5729,0,0,,,[],[]\n1,0,13309898705,554169193,12010,0,4427,19000.0,True,,\n2,0,4213672059,1720860329,457,0,240,17000.0,True,,\n3,0,62798072960,288641214,13262,0,266,23000.0,True,,\n4,0,10585422061,1728340777,6119,0,162,72400.0,True,,\n5,0,18020362258,1364159702,12023,0,4424,18000.0,True,,\n6,0,2325432079,1521618396,574,0,1367,18000.0,True,,\n7,0,39456940781,1317245193,12043,0,5314,17000.0,True,,\n8,0,3460555189,1700555100,7910,0,532,21000.0,True,,\n9,0,2214770464,998511398,7908,0,393,21000.0,True,,\n10,0,516803182,1422853669,1143,0,85,15000.0,True,,\n11,0,2153839851,1096784725,11033,0,315,34250.0,True,,\n12,0,2153839851,1096784725,11032,0,315,34250.0,True,,\n13,0,2153839851,1096784725,11034,0,315,34250.0,True,,\n14,0,2153839851,1096784725,11031,0,315,34250.0,True,,\n15,0,1218852591,385471210,9538,0,378,11000.0,True,,\n16,0,32722340115,1202386221,1002,0,136,16000.0,True,,\n17,0,2059097926,2018567473,12148,0,589,17000.0,True,,\n18,1,23609,275030867,5502,0,1,34000.0,False,\"[0, 1, 1, 0, 0, 1, 0, 1, 0, 1, 1, 1, 1, 1, 0, 1, 0, 1]\",\"[0, 0, 1, 1, 0, 1, 3, 3, 1, 1, 0, 3, 1, 2, 2, 1, 0, 3]\"\n</code></pre>\n<p>Above is the sample data from the file example_test.csv. </p>\n<p>Focus on the group_num as 0, you will see 18 records there. [0-17]</p>\n<p>Now focus just on the very next row where group_num 1 starts, you will find the list contains 18 elements in it. <br>\nWhich coincides with the row counts of the previous group_num you have seen..</p>\n<p>And the same follows for the next group num.</p>\n<pre><code>18,1,23609,275030867,5502,0,1,34000.0,False,\"[0, 1, 1, 0, 0, 1, 0, 1, 0, 1, 1, 1, 1, 1, 0, 1, 0, 1]\",\"[0, 0, 1, 1, 0, 1, 3, 3, 1, 1, 0, 3, 1, 2, 2, 1, 0, 3]\"\n19,1,2035159380,1233875513,1512,0,1431,25000.0,True,,\n20,1,2035159380,1233875513,1511,0,1431,25000.0,True,,\n21,1,2035159380,1233875513,1513,0,1431,25000.0,True,,\n22,1,1217231,891955351,9145,0,20,23000.0,True,,\n23,1,11265012636,1981166446,3299,0,950,36333.0,True,,\n24,1,11265012636,1981166446,3297,0,950,36333.0,True,,\n25,1,11265012636,1981166446,3298,0,950,36333.0,True,,\n26,1,4693145319,1637273633,11373,0,3148,41000.0,True,,\n27,1,2294633294,2030979309,3348,0,170,30666.0,True,,\n28,1,2294633294,2030979309,3350,0,170,30666.0,True,,\n29,1,2294633294,2030979309,3349,0,170,30666.0,True,,\n30,1,13679981387,319060572,10951,0,621,30250.0,True,,\n31,1,13679981387,319060572,10953,0,621,30250.0,True,,\n32,1,13679981387,319060572,10954,0,621,30250.0,True,,\n33,1,13679981387,319060572,10952,0,621,30250.0,True,,\n34,1,62798100988,288641214,5418,0,267,24000.0,True,,\n35,1,251107302,98059812,5892,0,9,34000.0,True,,\n37,1,1254086842,674533997,5301,0,1044,29000.0,True,,\n38,1,44526338636,555691277,666,0,52,25000.0,True,,\n39,1,39456965438,1317245193,12195,0,5315,18000.0,True,,\n40,1,32722359498,1202386221,582,0,137,17000.0,True,,\n41,1,417208826,775113212,10078,0,128,28250.0,True,,\n42,1,417208826,775113212,10080,0,128,28250.0,True,,\n43,1,417208826,775113212,10081,0,128,28250.0,True,,\n44,1,417208826,775113212,10079,0,128,28250.0,True,,\n45,1,209087981,1219481379,3784,0,40,22000.0,True,,\n46,2,42512,275030867,6133,0,2,17000.0,False,\"[1, 0, 1, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 1, 1, 0, 1, 1, 1, 0, 1]\",\"[1, 2, 2, 2, 1, 0, 1, 2, 1, 1, 3, 2, 0, 0, 3, 1, 1, 3, 3, 3, 0, 0, 0, 1, 2, 0, 3]\"\n</code></pre>\n<p>At the row no 46th, the list has 27 elements in it which coincides with the row count of the group_num 1.</p>\n<p>So this is a great information in case people have missed till now.<br>\nPlease correct my understandings.</p>\n<pre><code>prior_group_responses (string) provides all of the user_answer entries for previous group in a string representation of a list in the first row of the group. All other rows in each group are null. If you are using Python, you will likely want to call eval on the non-null rows. Some rows may be null, or empty lists.\n\nprior_group_answers_correct (string) provides all the answered_correctly field for previous group, with the same format and caveats as prior_group_responses. Some rows may be null, or empty lists.\n</code></pre>\n<p>So we have the actual correct labels (as in the test set) in the column \"prior_group_responses\" and whether they were correct or not by the user's response in the column \"prior_group_answers_correct\".. i.e. the user thought option 0 is correct but actually it was option 1 which was the correct asnwer.</p>\n<p>So it seems we can test set to train our models? I mean we won't we doing update to our model for every batch, we will do it when we have k batches \"combined\"…</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190430#1046574\" target=\"_blank\">Credits to the comment by </a><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a></p>\n<p>Looking forward to hear your thoughts folks! <br>\nThanks!</p>\n<p>Edit - 1</p>\n<ul>\n<li>Updated the same with chunking df's and then updating the stats etc with a successful submission. <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856\" target=\"_blank\">link</a></li>\n</ul>",
      "rawMarkdown": "NB I am showing it only for first few rows, i have verified the same follows for the rest as well. (at least for the sample test rows)\n\n```\nrow_id,group_num,timestamp,user_id,content_id,content_type_id,task_container_id,prior_question_elapsed_time,prior_question_had_explanation,prior_group_answers_correct,prior_group_responses\n0,0,0,275030867,5729,0,0,,,[],[]\n1,0,13309898705,554169193,12010,0,4427,19000.0,True,,\n2,0,4213672059,1720860329,457,0,240,17000.0,True,,\n3,0,62798072960,288641214,13262,0,266,23000.0,True,,\n4,0,10585422061,1728340777,6119,0,162,72400.0,True,,\n5,0,18020362258,1364159702,12023,0,4424,18000.0,True,,\n6,0,2325432079,1521618396,574,0,1367,18000.0,True,,\n7,0,39456940781,1317245193,12043,0,5314,17000.0,True,,\n8,0,3460555189,1700555100,7910,0,532,21000.0,True,,\n9,0,2214770464,998511398,7908,0,393,21000.0,True,,\n10,0,516803182,1422853669,1143,0,85,15000.0,True,,\n11,0,2153839851,1096784725,11033,0,315,34250.0,True,,\n12,0,2153839851,1096784725,11032,0,315,34250.0,True,,\n13,0,2153839851,1096784725,11034,0,315,34250.0,True,,\n14,0,2153839851,1096784725,11031,0,315,34250.0,True,,\n15,0,1218852591,385471210,9538,0,378,11000.0,True,,\n16,0,32722340115,1202386221,1002,0,136,16000.0,True,,\n17,0,2059097926,2018567473,12148,0,589,17000.0,True,,\n18,1,23609,275030867,5502,0,1,34000.0,False,\"[0, 1, 1, 0, 0, 1, 0, 1, 0, 1, 1, 1, 1, 1, 0, 1, 0, 1]\",\"[0, 0, 1, 1, 0, 1, 3, 3, 1, 1, 0, 3, 1, 2, 2, 1, 0, 3]\"\n```\n\nAbove is the sample data from the file example_test.csv. \n\nFocus on the group_num as 0, you will see 18 records there. [0-17]\n\nNow focus just on the very next row where group_num 1 starts, you will find the list contains 18 elements in it. \nWhich coincides with the row counts of the previous group_num you have seen..\n\nAnd the same follows for the next group num.\n\n```\n18,1,23609,275030867,5502,0,1,34000.0,False,\"[0, 1, 1, 0, 0, 1, 0, 1, 0, 1, 1, 1, 1, 1, 0, 1, 0, 1]\",\"[0, 0, 1, 1, 0, 1, 3, 3, 1, 1, 0, 3, 1, 2, 2, 1, 0, 3]\"\n19,1,2035159380,1233875513,1512,0,1431,25000.0,True,,\n20,1,2035159380,1233875513,1511,0,1431,25000.0,True,,\n21,1,2035159380,1233875513,1513,0,1431,25000.0,True,,\n22,1,1217231,891955351,9145,0,20,23000.0,True,,\n23,1,11265012636,1981166446,3299,0,950,36333.0,True,,\n24,1,11265012636,1981166446,3297,0,950,36333.0,True,,\n25,1,11265012636,1981166446,3298,0,950,36333.0,True,,\n26,1,4693145319,1637273633,11373,0,3148,41000.0,True,,\n27,1,2294633294,2030979309,3348,0,170,30666.0,True,,\n28,1,2294633294,2030979309,3350,0,170,30666.0,True,,\n29,1,2294633294,2030979309,3349,0,170,30666.0,True,,\n30,1,13679981387,319060572,10951,0,621,30250.0,True,,\n31,1,13679981387,319060572,10953,0,621,30250.0,True,,\n32,1,13679981387,319060572,10954,0,621,30250.0,True,,\n33,1,13679981387,319060572,10952,0,621,30250.0,True,,\n34,1,62798100988,288641214,5418,0,267,24000.0,True,,\n35,1,251107302,98059812,5892,0,9,34000.0,True,,\n37,1,1254086842,674533997,5301,0,1044,29000.0,True,,\n38,1,44526338636,555691277,666,0,52,25000.0,True,,\n39,1,39456965438,1317245193,12195,0,5315,18000.0,True,,\n40,1,32722359498,1202386221,582,0,137,17000.0,True,,\n41,1,417208826,775113212,10078,0,128,28250.0,True,,\n42,1,417208826,775113212,10080,0,128,28250.0,True,,\n43,1,417208826,775113212,10081,0,128,28250.0,True,,\n44,1,417208826,775113212,10079,0,128,28250.0,True,,\n45,1,209087981,1219481379,3784,0,40,22000.0,True,,\n46,2,42512,275030867,6133,0,2,17000.0,False,\"[1, 0, 1, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 1, 1, 0, 1, 1, 1, 0, 1]\",\"[1, 2, 2, 2, 1, 0, 1, 2, 1, 1, 3, 2, 0, 0, 3, 1, 1, 3, 3, 3, 0, 0, 0, 1, 2, 0, 3]\"\n```\n\nAt the row no 46th, the list has 27 elements in it which coincides with the row count of the group_num 1.\n\nSo this is a great information in case people have missed till now.\nPlease correct my understandings.\n\n```\nprior_group_responses (string) provides all of the user_answer entries for previous group in a string representation of a list in the first row of the group. All other rows in each group are null. If you are using Python, you will likely want to call eval on the non-null rows. Some rows may be null, or empty lists.\n\nprior_group_answers_correct (string) provides all the answered_correctly field for previous group, with the same format and caveats as prior_group_responses. Some rows may be null, or empty lists.\n```\n\nSo we have the actual correct labels (as in the test set) in the column \"prior_group_responses\" and whether they were correct or not by the user's response in the column \"prior_group_answers_correct\".. i.e. the user thought option 0 is correct but actually it was option 1 which was the correct asnwer.\n\nSo it seems we can test set to train our models? I mean we won't we doing update to our model for every batch, we will do it when we have k batches \"combined\"...\n\n[Credits to the comment by @spacelx](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190430#1046574)\n\nLooking forward to hear your thoughts folks! \nThanks!\n\nEdit - 1\n\n- Updated the same with chunking df's and then updating the stats etc with a successful submission. [link](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856)",
      "votes": null
    },
    {
      "id": "1048116",
      "postDate": "10/13/2020 07:59:32",
      "content": "<p>That sounds about right. Additionally important to note, <code>prior_group_answers_correct</code> and <code>prior_group_responses</code> always have the same length as the previous <code>test_df</code> - even if there were lectures included!<br>\nSo even for lecture rows, there will be a value in these two lists. We can't know which value that is though, because the public test set doesn't include any lectures unfortunately. Shouldn't make a difference though whether it's zero or NaN or whatnot.</p>",
      "rawMarkdown": "That sounds about right. Additionally important to note, `prior_group_answers_correct` and `prior_group_responses` always have the same length as the previous `test_df` - even if there were lectures included!\nSo even for lecture rows, there will be a value in these two lists. We can't know which value that is though, because the public test set doesn't include any lectures unfortunately. Shouldn't make a difference though whether it's zero or NaN or whatnot.",
      "votes": null
    },
    {
      "id": "1048118",
      "postDate": "10/13/2020 08:02:48",
      "content": "<p>We can use the content_id to verify that i think as it's -1 for it and the answers will be null for that, right? Well the hosts should have given us that scenario or we can check ourselves by wasting another sub (break the loop if we encounter such a row and make a dummy sub)</p>",
      "rawMarkdown": "We can use the content_id to verify that i think as it's -1 for it and the answers will be null for that, right? Well the hosts should have given us that scenario or we can check ourselves by wasting another sub (break the loop if we encounter such a row and make a dummy sub)",
      "votes": null
    },
    {
      "id": "1048124",
      "postDate": "10/13/2020 08:07:17",
      "content": "<p>Yeah you best just use <code>content_type_id</code> to drop all rows which are lectures - but <em>after</em> you matched the <code>prior_group_answers_correct</code> to the previous <code>test_df</code>, not before.</p>",
      "rawMarkdown": "Yeah you best just use `content_type_id` to drop all rows which are lectures - but *after* you matched the `prior_group_answers_correct` to the previous `test_df`, not before.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1048116,
      "author_name": "spacelx",
      "author_url": "",
      "post_date": "10/13/2020 07:59:32",
      "content": "<p>That sounds about right. Additionally important to note, <code>prior_group_answers_correct</code> and <code>prior_group_responses</code> always have the same length as the previous <code>test_df</code> - even if there were lectures included!<br>\nSo even for lecture rows, there will be a value in these two lists. We can't know which value that is though, because the public test set doesn't include any lectures unfortunately. Shouldn't make a difference though whether it's zero or NaN or whatnot.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1048118,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/13/2020 08:02:48",
          "content": "<p>We can use the content_id to verify that i think as it's -1 for it and the answers will be null for that, right? Well the hosts should have given us that scenario or we can check ourselves by wasting another sub (break the loop if we encounter such a row and make a dummy sub)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1048124,
          "author_name": "spacelx",
          "author_url": "",
          "post_date": "10/13/2020 08:07:17",
          "content": "<p>Yeah you best just use <code>content_type_id</code> to drop all rows which are lectures - but <em>after</em> you matched the <code>prior_group_answers_correct</code> to the previous <code>test_df</code>, not before.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1048082": "NB I am showing it only for first few rows, i have verified the same follows for the rest as well. (at least for the sample test rows)\n\n```\nrow_id,group_num,timestamp,user_id,content_id,content_type_id,task_container_id,prior_question_elapsed_time,prior_question_had_explanation,prior_group_answers_correct,prior_group_responses\n0,0,0,275030867,5729,0,0,,,[],[]\n1,0,13309898705,554169193,12010,0,4427,19000.0,True,,\n2,0,4213672059,1720860329,457,0,240,17000.0,True,,\n3,0,62798072960,288641214,13262,0,266,23000.0,True,,\n4,0,10585422061,1728340777,6119,0,162,72400.0,True,,\n5,0,18020362258,1364159702,12023,0,4424,18000.0,True,,\n6,0,2325432079,1521618396,574,0,1367,18000.0,True,,\n7,0,39456940781,1317245193,12043,0,5314,17000.0,True,,\n8,0,3460555189,1700555100,7910,0,532,21000.0,True,,\n9,0,2214770464,998511398,7908,0,393,21000.0,True,,\n10,0,516803182,1422853669,1143,0,85,15000.0,True,,\n11,0,2153839851,1096784725,11033,0,315,34250.0,True,,\n12,0,2153839851,1096784725,11032,0,315,34250.0,True,,\n13,0,2153839851,1096784725,11034,0,315,34250.0,True,,\n14,0,2153839851,1096784725,11031,0,315,34250.0,True,,\n15,0,1218852591,385471210,9538,0,378,11000.0,True,,\n16,0,32722340115,1202386221,1002,0,136,16000.0,True,,\n17,0,2059097926,2018567473,12148,0,589,17000.0,True,,\n18,1,23609,275030867,5502,0,1,34000.0,False,\"[0, 1, 1, 0, 0, 1, 0, 1, 0, 1, 1, 1, 1, 1, 0, 1, 0, 1]\",\"[0, 0, 1, 1, 0, 1, 3, 3, 1, 1, 0, 3, 1, 2, 2, 1, 0, 3]\"\n```\n\nAbove is the sample data from the file example_test.csv. \n\nFocus on the group_num as 0, you will see 18 records there. [0-17]\n\nNow focus just on the very next row where group_num 1 starts, you will find the list contains 18 elements in it. \nWhich coincides with the row counts of the previous group_num you have seen..\n\nAnd the same follows for the next group num.\n\n```\n18,1,23609,275030867,5502,0,1,34000.0,False,\"[0, 1, 1, 0, 0, 1, 0, 1, 0, 1, 1, 1, 1, 1, 0, 1, 0, 1]\",\"[0, 0, 1, 1, 0, 1, 3, 3, 1, 1, 0, 3, 1, 2, 2, 1, 0, 3]\"\n19,1,2035159380,1233875513,1512,0,1431,25000.0,True,,\n20,1,2035159380,1233875513,1511,0,1431,25000.0,True,,\n21,1,2035159380,1233875513,1513,0,1431,25000.0,True,,\n22,1,1217231,891955351,9145,0,20,23000.0,True,,\n23,1,11265012636,1981166446,3299,0,950,36333.0,True,,\n24,1,11265012636,1981166446,3297,0,950,36333.0,True,,\n25,1,11265012636,1981166446,3298,0,950,36333.0,True,,\n26,1,4693145319,1637273633,11373,0,3148,41000.0,True,,\n27,1,2294633294,2030979309,3348,0,170,30666.0,True,,\n28,1,2294633294,2030979309,3350,0,170,30666.0,True,,\n29,1,2294633294,2030979309,3349,0,170,30666.0,True,,\n30,1,13679981387,319060572,10951,0,621,30250.0,True,,\n31,1,13679981387,319060572,10953,0,621,30250.0,True,,\n32,1,13679981387,319060572,10954,0,621,30250.0,True,,\n33,1,13679981387,319060572,10952,0,621,30250.0,True,,\n34,1,62798100988,288641214,5418,0,267,24000.0,True,,\n35,1,251107302,98059812,5892,0,9,34000.0,True,,\n37,1,1254086842,674533997,5301,0,1044,29000.0,True,,\n38,1,44526338636,555691277,666,0,52,25000.0,True,,\n39,1,39456965438,1317245193,12195,0,5315,18000.0,True,,\n40,1,32722359498,1202386221,582,0,137,17000.0,True,,\n41,1,417208826,775113212,10078,0,128,28250.0,True,,\n42,1,417208826,775113212,10080,0,128,28250.0,True,,\n43,1,417208826,775113212,10081,0,128,28250.0,True,,\n44,1,417208826,775113212,10079,0,128,28250.0,True,,\n45,1,209087981,1219481379,3784,0,40,22000.0,True,,\n46,2,42512,275030867,6133,0,2,17000.0,False,\"[1, 0, 1, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 1, 1, 0, 1, 1, 1, 0, 1]\",\"[1, 2, 2, 2, 1, 0, 1, 2, 1, 1, 3, 2, 0, 0, 3, 1, 1, 3, 3, 3, 0, 0, 0, 1, 2, 0, 3]\"\n```\n\nAt the row no 46th, the list has 27 elements in it which coincides with the row count of the group_num 1.\n\nSo this is a great information in case people have missed till now.\nPlease correct my understandings.\n\n```\nprior_group_responses (string) provides all of the user_answer entries for previous group in a string representation of a list in the first row of the group. All other rows in each group are null. If you are using Python, you will likely want to call eval on the non-null rows. Some rows may be null, or empty lists.\n\nprior_group_answers_correct (string) provides all the answered_correctly field for previous group, with the same format and caveats as prior_group_responses. Some rows may be null, or empty lists.\n```\n\nSo we have the actual correct labels (as in the test set) in the column \"prior_group_responses\" and whether they were correct or not by the user's response in the column \"prior_group_answers_correct\".. i.e. the user thought option 0 is correct but actually it was option 1 which was the correct asnwer.\n\nSo it seems we can test set to train our models? I mean we won't we doing update to our model for every batch, we will do it when we have k batches \"combined\"...\n\n[Credits to the comment by @spacelx](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190430#1046574)\n\nLooking forward to hear your thoughts folks! \nThanks!\n\nEdit - 1\n\n- Updated the same with chunking df's and then updating the stats etc with a successful submission. [link](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856)",
    "1048116": "That sounds about right. Additionally important to note, `prior_group_answers_correct` and `prior_group_responses` always have the same length as the previous `test_df` - even if there were lectures included!\nSo even for lecture rows, there will be a value in these two lists. We can't know which value that is though, because the public test set doesn't include any lectures unfortunately. Shouldn't make a difference though whether it's zero or NaN or whatnot.",
    "1048118": "We can use the content_id to verify that i think as it's -1 for it and the answers will be null for that, right? Well the hosts should have given us that scenario or we can check ourselves by wasting another sub (break the loop if we encounter such a row and make a dummy sub)",
    "1048124": "Yeah you best just use `content_type_id` to drop all rows which are lectures - but *after* you matched the `prior_group_answers_correct` to the previous `test_df`, not before."
  },
  "source": "meta"
}