{
  "id": 188862,
  "title": "Welcome to the competition!",
  "url": "/competitions/riiid-test-answer-prediction/discussion/188862",
  "author_name": "Sohier Dane",
  "post_date": "2020-10-05T17:28:32.453000",
  "votes": 40,
  "comment_count": 50,
  "views": 0,
  "content": "<p>Welcome to the Riiid! Answer Correctness Prediction competition! Assessing student progress from their past work is a critical task for online education and this challenge provides a large amount of real world data for that task.<br>\nWith the competition API serving data, you'll also need to tackle the difficulties introduced by new students arriving on the site and new questions being added. We're excited to see how you tackle this challenge!</p>\n<p>Happy Kaggling!</p>",
  "messages": [
    {
      "id": 1038272,
      "postDate": "2020-10-05T17:28:32.453Z",
      "content": "<p>Welcome to the Riiid! Answer Correctness Prediction competition! Assessing student progress from their past work is a critical task for online education and this challenge provides a large amount of real world data for that task.<br>\nWith the competition API serving data, you'll also need to tackle the difficulties introduced by new students arriving on the site and new questions being added. We're excited to see how you tackle this challenge!</p>\n<p>Happy Kaggling!</p>",
      "rawMarkdown": "Welcome to the Riiid! Answer Correctness Prediction competition! Assessing student progress from their past work is a critical task for online education and this challenge provides a large amount of real world data for that task.\nWith the competition API serving data, you'll also need to tackle the difficulties introduced by new students arriving on the site and new questions being added. We're excited to see how you tackle this challenge!\n\nHappy Kaggling!\n",
      "votes": 38
    },
    {
      "id": 1038346,
      "postDate": "2020-10-05T18:38:36.697Z",
      "content": "<p>This competition is a code competition, so, the prediction phase must be done in a Kaggle kernel. But can the models training phase be done offline on a public cloud VM or a personal computer?</p>\n<p>Thanks.</p>",
      "rawMarkdown": "This competition is a code competition, so, the prediction phase must be done in a Kaggle kernel. But can the models training phase be done offline on a public cloud VM or a personal computer?\n\nThanks.",
      "votes": 6,
      "replies": [
        {
          "id": 1038441,
          "postDate": "2020-10-05T19:38:55.350Z",
          "content": "<p>Yes, you can train a model somewhere other than on Kaggle.</p>",
          "rawMarkdown": "Yes, you can train a model somewhere other than on Kaggle.",
          "votes": 6
        },
        {
          "id": 1049243,
          "postDate": "2020-10-14T08:48:51.320Z",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Can we generate the features for test data on on a public cloud VM or a personal computer, then upload them and the models into Kaggle to do the prediction?</p>",
          "rawMarkdown": "@sohier Can we generate the features for test data on on a public cloud VM or a personal computer, then upload them and the models into Kaggle to do the prediction?",
          "votes": 1
        },
        {
          "id": 1087584,
          "postDate": "2020-11-22T21:54:03.087Z",
          "content": "<p>Yes I think we can do that!<br>\nHere is a sample of submission notebook <a href=\"url\" target=\"_blank\">https://www.kaggle.com/sohier/quick-sample-submission</a></p>",
          "rawMarkdown": "Yes I think we can do that!\nHere is a sample of submission notebook [https://www.kaggle.com/sohier/quick-sample-submission](url)"
        }
      ]
    },
    {
      "id": 1044670,
      "postDate": "2020-10-10T03:04:03.830Z",
      "content": "<p>It would be nice to have at-least the runtime stats of our submitted kernel on the submission's page against the private test set?</p>",
      "rawMarkdown": "It would be nice to have at-least the runtime stats of our submitted kernel on the submission's page against the private test set?",
      "votes": 3
    },
    {
      "id": 1042185,
      "postDate": "2020-10-08T05:50:55.700Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>,</p>\n<p>To my understanding, there are some user_ids appear both in the train.csv and test set. When making predictions, is it allowed to use the data about the learning history (if there is any) of a user_id in the train.csv? I mean not just for training the model, but also feeding train.csv into the model when making predictions.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi @sohier,\n\nTo my understanding, there are some user_ids appear both in the train.csv and test set. When making predictions, is it allowed to use the data about the learning history (if there is any) of a user_id in the train.csv? I mean not just for training the model, but also feeding train.csv into the model when making predictions.\n\nThanks!",
      "votes": 3
    },
    {
      "id": 1042236,
      "postDate": "2020-10-08T06:20:28.393Z",
      "content": "<p>And one more question related  to the previous 2 by <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> and <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> : Can one assume that there are no gaps neither in the training data nor between training and test slices, nor between test slices? </p>\n<p>The <em>Data</em> tab states that the test batches follow historical order, but is says nothing about completeness:</p>\n<blockquote>\n  <p>The API provides user interactions groups in the order in which they occurred.</p>\n</blockquote>\n<p>If one wants to build a training curve for individual users it is important to have complete history per user. For example, if we want to count the number of lectures the given user had up to the specific question, we need to be sure that in the training and test data that have been all user interactions recorded</p>",
      "rawMarkdown": "And one more question related  to the previous 2 by @frankpanxj and @sirishks : Can one assume that there are no gaps neither in the training data nor between training and test slices, nor between test slices? \n\nThe _Data_ tab states that the test batches follow historical order, but is says nothing about completeness:\n> The API provides user interactions groups in the order in which they occurred.\n\nIf one wants to build a training curve for individual users it is important to have complete history per user. For example, if we want to count the number of lectures the given user had up to the specific question, we need to be sure that in the training and test data that have been all user interactions recorded",
      "votes": 4
    },
    {
      "id": 1040558,
      "postDate": "2020-10-07T08:12:47.783Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> In the Data page, section example_test.csv:<br>\n<code>There are two different rows that mirror what information the AI tutor actually has available at any given time,...</code></p>\n<p>I think it should be<br>\n<code>There are two different columns...</code></p>",
      "rawMarkdown": "@sohier In the Data page, section example_test.csv:\n`There are two different rows that mirror what information the AI tutor actually has available at any given time,...`\n\nI think it should be\n`There are two different columns...`",
      "votes": 4,
      "replies": [
        {
          "id": 1042998,
          "postDate": "2020-10-08T15:52:29.287Z",
          "content": "<p>Fixed, thanks!</p>",
          "rawMarkdown": "Fixed, thanks!",
          "votes": 2
        }
      ]
    },
    {
      "id": 1041155,
      "postDate": "2020-10-07T15:27:42.067Z",
      "content": "<p>Hi Sohier, when I downloaded using Kaggle API and try to import riiideducation, I get the following error:<br>\nModuleNotFoundError: No module named 'riiideducation.competition'<br>\nCould you help?<br>\nThanks!</p>",
      "rawMarkdown": "Hi Sohier, when I downloaded using Kaggle API and try to import riiideducation, I get the following error:\nModuleNotFoundError: No module named 'riiideducation.competition'\nCould you help?\nThanks!",
      "votes": 1,
      "replies": [
        {
          "id": 1041195,
          "postDate": "2020-10-07T15:50:29.467Z",
          "content": "<p>The most likely issue is that you haven't added the module's location to your python path, which you can resolve with <code>sys.path.append</code>.</p>",
          "rawMarkdown": "The most likely issue is that you haven't added the module's location to your python path, which you can resolve with `sys.path.append`.",
          "votes": 1
        },
        {
          "id": 1042178,
          "postDate": "2020-10-08T05:39:14.447Z",
          "content": "<p>Thanks Sohier. I tried this and also added library variable but it still shows the same error</p>",
          "rawMarkdown": "Thanks Sohier. I tried this and also added library variable but it still shows the same error",
          "votes": 1
        },
        {
          "id": 1049625,
          "postDate": "2020-10-14T15:37:10.403Z",
          "content": "<p>Sheema, did you ever find a solution?? I have the same problem, I downloaded and installed the kaggle api and it works, then I appended the riideducation directory using sys.path.append. I see the two files in that directory, a python file and an .so file that begins with 'competition', but python says no riiideducation.competition module found, so I'm glad I'm not the only one who had the problem.</p>",
          "rawMarkdown": "Sheema, did you ever find a solution?? I have the same problem, I downloaded and installed the kaggle api and it works, then I appended the riideducation directory using sys.path.append. I see the two files in that directory, a python file and an .so file that begins with 'competition', but python says no riiideducation.competition module found, so I'm glad I'm not the only one who had the problem."
        },
        {
          "id": 1051943,
          "postDate": "2020-10-17T05:05:43.867Z",
          "content": "<p>Hi Jeff, i saw someone's comment that this works with python 3.7. I haven't yet tried this out. Currently my workaround is :  I work on my local machine on the modelling and training and go to  Kaggle for testing only.<br>\nHope this helps.. </p>",
          "rawMarkdown": "Hi Jeff, i saw someone's comment that this works with python 3.7. I haven't yet tried this out. Currently my workaround is :  I work on my local machine on the modelling and training and go to  Kaggle for testing only.\nHope this helps.. "
        },
        {
          "id": 1052155,
          "postDate": "2020-10-17T11:52:20.433Z",
          "content": "<p>Thank you Sheema, it helps tremendously. I prefer to work on my desktop, so I'll adopt your workaround and use kaggle kernel for testing.</p>",
          "rawMarkdown": "Thank you Sheema, it helps tremendously. I prefer to work on my desktop, so I'll adopt your workaround and use kaggle kernel for testing."
        },
        {
          "id": 1307138,
          "postDate": "2021-05-14T09:23:27.133Z",
          "content": "<p>Very interesting</p>",
          "rawMarkdown": "Very interesting"
        }
      ]
    },
    {
      "id": 1038862,
      "postDate": "2020-10-06T05:55:46.610Z",
      "content": "<p>hi sir, i need to understand that do we have to use \"competition.cpython-37m-x86_64-linux-gnu.so\" this file from dataset. i am not aware anything about linux. can i proceed with only dataset files and not above mentioned file in my windows machine or kaggle kernel?</p>",
      "rawMarkdown": "hi sir, i need to understand that do we have to use \"competition.cpython-37m-x86_64-linux-gnu.so\" this file from dataset. i am not aware anything about linux. can i proceed with only dataset files and not above mentioned file in my windows machine or kaggle kernel?",
      "votes": 1,
      "replies": [
        {
          "id": 1042031,
          "postDate": "2020-10-08T03:01:43.590Z",
          "content": "<p>That file is used to generate batches during the test phase. You're not supposed to use it for training a model.</p>",
          "rawMarkdown": "That file is used to generate batches during the test phase. You're not supposed to use it for training a model.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1042115,
      "postDate": "2020-10-08T04:24:08.030Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Although you have already answered<br>\n<code>we are limited in what we can disclose</code><br>\nin discussion topic \"<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188899\" target=\"_blank\">Assumptions about the dataset?</a>\"</p>\n<p>I would like to highlight especially the first question one more time viz.,<br>\n<code>1) All questions in the test set come cronologically after the questions in the training dataset?</code><br>\ni.e. the <strong>row_id</strong> and <strong>timestamp</strong> of test data delivered by the time-series API are greater than those of the train rows.</p>\n<p>This question is to ensure that there is no <strong>leakage</strong> and as such deserves to be answered.</p>",
      "rawMarkdown": "@sohier Although you have already answered\n`we are limited in what we can disclose`\nin discussion topic \"[Assumptions about the dataset?](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188899)\"\n\nI would like to highlight especially the first question one more time viz.,\n`1) All questions in the test set come cronologically after the questions in the training dataset?`\ni.e. the **row_id** and **timestamp** of test data delivered by the time-series API are greater than those of the train rows.\n\nThis question is to ensure that there is no **leakage** and as such deserves to be answered.",
      "votes": 2
    },
    {
      "id": 1039696,
      "postDate": "2020-10-06T18:19:55.870Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> It seems that there is an issue with the \"questions\" table when the solution is evaluated for the Leaderboard.<br>\nAs you can see <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188986\" target=\"_blank\">here</a>, people cannot make submissions on the LB when they use the \"questions\" table.</p>\n<p>Can you help us ?</p>\n<p>Thomas</p>",
      "rawMarkdown": "@sohier It seems that there is an issue with the \"questions\" table when the solution is evaluated for the Leaderboard.\nAs you can see [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188986), people cannot make submissions on the LB when they use the \"questions\" table.\n\nCan you help us ?\n\nThomas",
      "votes": 2
    },
    {
      "id": 1038705,
      "postDate": "2020-10-06T01:27:09.513Z",
      "content": "<p>Hi Sohier, thank you for this competition. It is designed around Python. I already have software that predicts educational performance, but it is written in other computer languages. Can I extract the competition data (piece-wise if necessary) using the Python API to disk. Then analyze the data and make predictions using other software. Then upload the predictions from disk using the Python API? Thanks.</p>",
      "rawMarkdown": "Hi Sohier, thank you for this competition. It is designed around Python. I already have software that predicts educational performance, but it is written in other computer languages. Can I extract the competition data (piece-wise if necessary) using the Python API to disk. Then analyze the data and make predictions using other software. Then upload the predictions from disk using the Python API? Thanks.",
      "votes": 2,
      "replies": [
        {
          "id": 1039540,
          "postDate": "2020-10-06T16:29:14.650Z",
          "content": "<p>You're welcome to give it a shot as long as the software adheres to the competition rules, though I'd expect it will be quite slow. The API will force you to loop to and from the disk in a piece wise manner.</p>",
          "rawMarkdown": "You're welcome to give it a shot as long as the software adheres to the competition rules, though I'd expect it will be quite slow. The API will force you to loop to and from the disk in a piece wise manner.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1143991,
      "postDate": "2021-01-08T07:02:57.707Z",
      "content": "<p>It was a nice learning experience.</p>\n<p>Couldn't have 'competed' here, but tried to see what score one gets using minimum set of data and without AI / sophisticated statistics. (as in here. <a href=\"https://www.kaggle.com/avaniv/riiid-threefields-noai\" target=\"_blank\">https://www.kaggle.com/avaniv/riiid-threefields-noai</a>)</p>\n<p>May be AI competitions here can have non AI section that helps in better understanding of value addition of additional data and AI over minimalist / analytical approaches?</p>",
      "rawMarkdown": "It was a nice learning experience.\n\nCouldn't have 'competed' here, but tried to see what score one gets using minimum set of data and without AI / sophisticated statistics. (as in here. https://www.kaggle.com/avaniv/riiid-threefields-noai)\n\nMay be AI competitions here can have non AI section that helps in better understanding of value addition of additional data and AI over minimalist / analytical approaches?"
    },
    {
      "id": 1142520,
      "postDate": "2021-01-07T12:48:20.037Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1192725%2F5f39de855040973bdf4598e35dddd00b%2F1ef84cacfe6a3fad8cbd2578bed3e6cd.png?generation=1610023643061539&amp;alt=media\" alt=\"\"></p>\n<p>Since yesterday, there has been a bug in my kernel. Please help to have a look<br>\n<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209222\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209222</a></p>",
      "rawMarkdown": "@sohier  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1192725%2F5f39de855040973bdf4598e35dddd00b%2F1ef84cacfe6a3fad8cbd2578bed3e6cd.png?generation=1610023643061539&alt=media)\n\nSince yesterday, there has been a bug in my kernel. Please help to have a look\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209222"
    },
    {
      "id": 1126502,
      "postDate": "2020-12-25T16:56:29.160Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> - request you to fix the Kaggle bug for my userid. Sample submission (no changes to sample submission kernel of urs) is also not getting recognized and the status of the run is NOT updated under my submissions (no error..the status itself is not there). Request your help to fix it urgently as competition has reached its last stage</p>",
      "rawMarkdown": "@sohier - request you to fix the Kaggle bug for my userid. Sample submission (no changes to sample submission kernel of urs) is also not getting recognized and the status of the run is NOT updated under my submissions (no error..the status itself is not there). Request your help to fix it urgently as competition has reached its last stage",
      "replies": [
        {
          "id": 1129868,
          "postDate": "2020-12-28T15:49:04.313Z",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> request you to fix the Kaggle bug for my userId. It is the same issue. I believe couple of other users are also facing similar problems<br>\nI just gave a run a few mins back for the sample submission kernel. There is no status. </p>",
          "rawMarkdown": "@sohier request you to fix the Kaggle bug for my userId. It is the same issue. I believe couple of other users are also facing similar problems\nI just gave a run a few mins back for the sample submission kernel. There is no status. "
        }
      ]
    },
    {
      "id": 1117489,
      "postDate": "2020-12-18T05:57:48.233Z",
      "content": "<p>Thank you for this opportunity to learn.</p>",
      "rawMarkdown": "Thank you for this opportunity to learn."
    },
    {
      "id": 1098476,
      "postDate": "2020-12-01T16:59:29.067Z",
      "content": "<p>Quick Question:</p>\n<p>Kernel run for 1-2 hours but when we hit on the submisssion tab it takes more than 5 hours why so ?</p>\n<p>How many rows of test  it is computing?</p>",
      "rawMarkdown": "Quick Question:\n\nKernel run for 1-2 hours but when we hit on the submisssion tab it takes more than 5 hours why so ?\n\nHow many rows of test  it is computing?",
      "replies": [
        {
          "id": 1098479,
          "postDate": "2020-12-01T17:01:13.977Z",
          "content": "<p>2.5M rows. The faster your inference pipeline, the lesser the time it takes.</p>",
          "rawMarkdown": "2.5M rows. The faster your inference pipeline, the lesser the time it takes.",
          "votes": 1
        },
        {
          "id": 1098500,
          "postDate": "2020-12-01T17:15:01.153Z",
          "content": "<p>Also the notebook you use to submit to the competition can just be an inference notebook.</p>",
          "rawMarkdown": "Also the notebook you use to submit to the competition can just be an inference notebook.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1086670,
      "postDate": "2020-11-21T23:15:31.077Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>  </p>\n<p>I am a little perplexed by what some features mean. </p>\n<p>In the training data set we have:<br>\ncontent_id: (int16) ID code for the user interaction<br>\ntask_container_id: (int16) Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id.</p>\n<p>In the questions csv we have:<br>\nquestion_id: foreign key for the train/test content_id column, when the content type is question (0).<br>\nbundle_id: code for which questions are served together.</p>\n<p>Now my interpretation is (assuming content_type_id is 0 i.e. we are dealing with a question) the content_id in the training data tells us which row/id to use in the questions csv. </p>\n<p>Now if we check the task_container_id and check the value against the row in the questions csv matching the content_id I would expect them to match. In other words, I interpreted task_contain_id to tell us which questions are bundled together and bundle_id to do the same. This is not true.</p>\n<p>When I look at the data in the questions csv each row has its own unique bundle_id. In fact, the id column, the question_id column, and the bundle_id column all have the exact same entry in every row and match the row number. </p>\n<p>Am I misinterpreting something here?</p>\n<p>I would be thankful for any input.</p>",
      "rawMarkdown": "Hello @sohier  \n\nI am a little perplexed by what some features mean. \n\nIn the training data set we have:\ncontent_id: (int16) ID code for the user interaction\ntask_container_id: (int16) Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id.\n\nIn the questions csv we have:\nquestion_id: foreign key for the train/test content_id column, when the content type is question (0).\nbundle_id: code for which questions are served together.\n\nNow my interpretation is (assuming content_type_id is 0 i.e. we are dealing with a question) the content_id in the training data tells us which row/id to use in the questions csv. \n\nNow if we check the task_container_id and check the value against the row in the questions csv matching the content_id I would expect them to match. In other words, I interpreted task_contain_id to tell us which questions are bundled together and bundle_id to do the same. This is not true.\n\nWhen I look at the data in the questions csv each row has its own unique bundle_id. In fact, the id column, the question_id column, and the bundle_id column all have the exact same entry in every row and match the row number. \n\nAm I misinterpreting something here?\n\nI would be thankful for any input.\n"
    },
    {
      "id": 1079354,
      "postDate": "2020-11-16T00:44:42.030Z",
      "content": "<p>Hi Sohier, could you please tell me whether we can use an offline trained model for the final submission? Currently, I do data processing and model training offline. And only upload the trained model and related statistics to Kaggle kernel for batch inference.</p>",
      "rawMarkdown": "Hi Sohier, could you please tell me whether we can use an offline trained model for the final submission? Currently, I do data processing and model training offline. And only upload the trained model and related statistics to Kaggle kernel for batch inference.",
      "replies": [
        {
          "id": 1079491,
          "postDate": "2020-11-16T06:07:46.723Z",
          "content": "<p>The notebook you use to submit to the competition can just be an inference notebook. If you win the competition, you'd be required to share the full model to the hosts.</p>",
          "rawMarkdown": "The notebook you use to submit to the competition can just be an inference notebook. If you win the competition, you'd be required to share the full model to the hosts."
        }
      ]
    },
    {
      "id": 1048140,
      "postDate": "2020-10-13T08:27:32.213Z",
      "content": "<p>I doubt there is a wrong statement from <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> in the Data Page:</p>\n<p><code>Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual.</code></p>\n<p>When I check unique question_id in questions.csv and compare with unique content_id where content_type_id==0 (corresponding to questions) in train.csv, they are the same.</p>\n<p>Please clarify. Maybe during private test run, a different question.csv is swapped in?</p>",
      "rawMarkdown": "I doubt there is a wrong statement from @sohier in the Data Page:\n\n`Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual.`\n\nWhen I check unique question_id in questions.csv and compare with unique content_id where content_type_id==0 (corresponding to questions) in train.csv, they are the same.\n\nPlease clarify. Maybe during private test run, a different question.csv is swapped in?",
      "replies": [
        {
          "id": 1048542,
          "postDate": "2020-10-13T15:32:46.353Z",
          "content": "<blockquote>\n  <p>Maybe during private test run, a different question.csv is swapped in?</p>\n</blockquote>\n<p>Yes, that's the code competition. Same as we cannot see test which is evaluated.</p>",
          "rawMarkdown": "> Maybe during private test run, a different question.csv is swapped in?\n\nYes, that's the code competition. Same as we cannot see test which is evaluated.",
          "votes": 1
        },
        {
          "id": 1048797,
          "postDate": "2020-10-13T19:37:27.053Z",
          "content": "<p>I've raised the same question here - <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190415\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190415</a> and in this thread too.</p>\n<p>If a new questions.csv is swapped in, wouldn't it make using pre-trained models very tough?</p>",
          "rawMarkdown": "I've raised the same question here - https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190415 and in this thread too.\n\nIf a new questions.csv is swapped in, wouldn't it make using pre-trained models very tough?",
          "votes": 2
        },
        {
          "id": 1049431,
          "postDate": "2020-10-14T12:14:51.683Z",
          "content": "<p>Yeah, one possibility is that a different questions.csv is swapped in. I've been in Code Rerun competitions before and I fully understand how it works, but never seen any comp with another file swapped rather than test.csv. That's why I need a confirmation.</p>",
          "rawMarkdown": "Yeah, one possibility is that a different questions.csv is swapped in. I've been in Code Rerun competitions before and I fully understand how it works, but never seen any comp with another file swapped rather than test.csv. That's why I need a confirmation."
        }
      ]
    },
    {
      "id": 1047793,
      "postDate": "2020-10-13T00:20:23.880Z",
      "content": "<p>Could you tell us how many rows are there in each group_num?</p>",
      "rawMarkdown": "Could you tell us how many rows are there in each group_num?",
      "replies": [
        {
          "id": 1047822,
          "postDate": "2020-10-13T01:25:55.883Z",
          "content": "<p>From the data description:</p>\n<blockquote>\n  <p>The API provides user interactions groups in the order in which they occurred. Each group will contain interactions from many different users, but no more than one task_container_id of questions from any single user. Each group has between 1 and 1000 users.</p>\n</blockquote>",
          "rawMarkdown": "From the data description:\n> The API provides user interactions groups in the order in which they occurred. Each group will contain interactions from many different users, but no more than one task_container_id of questions from any single user. Each group has between 1 and 1000 users.",
          "votes": 2
        },
        {
          "id": 1047827,
          "postDate": "2020-10-13T01:28:05.133Z",
          "content": "<p>Thanks, I missed that</p>",
          "rawMarkdown": "Thanks, I missed that"
        }
      ]
    },
    {
      "id": 1046578,
      "postDate": "2020-10-11T19:43:39.403Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>Can you please confirm that the contents of the files <code>train.csv</code>, <code>questions.csv</code>, and <code>lectures.csv</code> won't change (new rows won't be added) in the test environment? </p>\n<p>I'm asking this because it came up in this discussion - <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190415\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190415</a></p>",
      "rawMarkdown": "Hi @sohier \n\nCan you please confirm that the contents of the files `train.csv`, `questions.csv`, and `lectures.csv` won't change (new rows won't be added) in the test environment? \n\nI'm asking this because it came up in this discussion - https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190415"
    },
    {
      "id": 1077795,
      "postDate": "2020-11-13T23:50:00.820Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1073638,
      "postDate": "2020-11-09T18:21:22.307Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1073552,
      "postDate": "2020-11-09T16:07:38.820Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1072634,
      "postDate": "2020-11-08T14:26:35.987Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1073398,
          "postDate": "2020-11-09T13:36:50.840Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1073440,
          "postDate": "2020-11-09T14:32:46.490Z",
          "content": "<p>This shouldn't happen, Check some kernel with successful sub.</p>",
          "rawMarkdown": "This shouldn't happen, Check some kernel with successful sub."
        },
        {
          "id": 1073491,
          "postDate": "2020-11-09T15:15:42.777Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1039761,
      "postDate": "2020-10-06T19:19:12.900Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true
    },
    {
      "id": 1506055,
      "postDate": "2021-09-07T19:27:08.543Z",
      "content": "<p>Thanks! The info is truly helpful!</p>",
      "rawMarkdown": "Thanks! The info is truly helpful!"
    }
  ],
  "comments": [
    {
      "id": 1038346,
      "author_name": "Thomas SELECK",
      "author_url": "",
      "post_date": "2020-10-05T18:38:36.697000",
      "content": "<p>This competition is a code competition, so, the prediction phase must be done in a Kaggle kernel. But can the models training phase be done offline on a public cloud VM or a personal computer?</p>\n<p>Thanks.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1038441,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2020-10-05T19:38:55.350000",
          "content": "<p>Yes, you can train a model somewhere other than on Kaggle.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1049243,
          "author_name": "Ethan",
          "author_url": "",
          "post_date": "2020-10-14T08:48:51.320000",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Can we generate the features for test data on on a public cloud VM or a personal computer, then upload them and the models into Kaggle to do the prediction?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1087584,
          "author_name": "Koubouratou IDJATON",
          "author_url": "",
          "post_date": "2020-11-22T21:54:03.087000",
          "content": "<p>Yes I think we can do that!<br>\nHere is a sample of submission notebook <a href=\"url\" target=\"_blank\">https://www.kaggle.com/sohier/quick-sample-submission</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1044670,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-10-10T03:04:03.830000",
      "content": "<p>It would be nice to have at-least the runtime stats of our submitted kernel on the submission's page against the private test set?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1042185,
      "author_name": "Frank Pan",
      "author_url": "",
      "post_date": "2020-10-08T05:50:55.700000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>,</p>\n<p>To my understanding, there are some user_ids appear both in the train.csv and test set. When making predictions, is it allowed to use the data about the learning history (if there is any) of a user_id in the train.csv? I mean not just for training the model, but also feeding train.csv into the model when making predictions.</p>\n<p>Thanks!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1042236,
      "author_name": "Misha Lisovyi",
      "author_url": "",
      "post_date": "2020-10-08T06:20:28.393000",
      "content": "<p>And one more question related  to the previous 2 by <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a> and <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> : Can one assume that there are no gaps neither in the training data nor between training and test slices, nor between test slices? </p>\n<p>The <em>Data</em> tab states that the test batches follow historical order, but is says nothing about completeness:</p>\n<blockquote>\n  <p>The API provides user interactions groups in the order in which they occurred.</p>\n</blockquote>\n<p>If one wants to build a training curve for individual users it is important to have complete history per user. For example, if we want to count the number of lectures the given user had up to the specific question, we need to be sure that in the training and test data that have been all user interactions recorded</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1040558,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2020-10-07T08:12:47.783000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> In the Data page, section example_test.csv:<br>\n<code>There are two different rows that mirror what information the AI tutor actually has available at any given time,...</code></p>\n<p>I think it should be<br>\n<code>There are two different columns...</code></p>",
      "votes": 4,
      "replies": [
        {
          "id": 1042998,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2020-10-08T15:52:29.287000",
          "content": "<p>Fixed, thanks!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1041155,
      "author_name": "Sheema Egbert",
      "author_url": "",
      "post_date": "2020-10-07T15:27:42.067000",
      "content": "<p>Hi Sohier, when I downloaded using Kaggle API and try to import riiideducation, I get the following error:<br>\nModuleNotFoundError: No module named 'riiideducation.competition'<br>\nCould you help?<br>\nThanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1041195,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2020-10-07T15:50:29.467000",
          "content": "<p>The most likely issue is that you haven't added the module's location to your python path, which you can resolve with <code>sys.path.append</code>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1042178,
          "author_name": "Sheema Egbert",
          "author_url": "",
          "post_date": "2020-10-08T05:39:14.447000",
          "content": "<p>Thanks Sohier. I tried this and also added library variable but it still shows the same error</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1049625,
          "author_name": "Jeff McLeod",
          "author_url": "",
          "post_date": "2020-10-14T15:37:10.403000",
          "content": "<p>Sheema, did you ever find a solution?? I have the same problem, I downloaded and installed the kaggle api and it works, then I appended the riideducation directory using sys.path.append. I see the two files in that directory, a python file and an .so file that begins with 'competition', but python says no riiideducation.competition module found, so I'm glad I'm not the only one who had the problem.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1051943,
          "author_name": "Sheema Egbert",
          "author_url": "",
          "post_date": "2020-10-17T05:05:43.867000",
          "content": "<p>Hi Jeff, i saw someone's comment that this works with python 3.7. I haven't yet tried this out. Currently my workaround is :  I work on my local machine on the modelling and training and go to  Kaggle for testing only.<br>\nHope this helps.. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1052155,
          "author_name": "Jeff McLeod",
          "author_url": "",
          "post_date": "2020-10-17T11:52:20.433000",
          "content": "<p>Thank you Sheema, it helps tremendously. I prefer to work on my desktop, so I'll adopt your workaround and use kaggle kernel for testing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1307138,
          "author_name": "Smolyuk Anastasia",
          "author_url": "",
          "post_date": "2021-05-14T09:23:27.133000",
          "content": "<p>Very interesting</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1038862,
      "author_name": "Aakash Barwad",
      "author_url": "",
      "post_date": "2020-10-06T05:55:46.610000",
      "content": "<p>hi sir, i need to understand that do we have to use \"competition.cpython-37m-x86_64-linux-gnu.so\" this file from dataset. i am not aware anything about linux. can i proceed with only dataset files and not above mentioned file in my windows machine or kaggle kernel?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1042031,
          "author_name": "Max Halford",
          "author_url": "",
          "post_date": "2020-10-08T03:01:43.590000",
          "content": "<p>That file is used to generate batches during the test phase. You're not supposed to use it for training a model.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1042115,
      "author_name": "Sirish Somanchi",
      "author_url": "",
      "post_date": "2020-10-08T04:24:08.030000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Although you have already answered<br>\n<code>we are limited in what we can disclose</code><br>\nin discussion topic \"<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188899\" target=\"_blank\">Assumptions about the dataset?</a>\"</p>\n<p>I would like to highlight especially the first question one more time viz.,<br>\n<code>1) All questions in the test set come cronologically after the questions in the training dataset?</code><br>\ni.e. the <strong>row_id</strong> and <strong>timestamp</strong> of test data delivered by the time-series API are greater than those of the train rows.</p>\n<p>This question is to ensure that there is no <strong>leakage</strong> and as such deserves to be answered.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1039696,
      "author_name": "Thomas SELECK",
      "author_url": "",
      "post_date": "2020-10-06T18:19:55.870000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> It seems that there is an issue with the \"questions\" table when the solution is evaluated for the Leaderboard.<br>\nAs you can see <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188986\" target=\"_blank\">here</a>, people cannot make submissions on the LB when they use the \"questions\" table.</p>\n<p>Can you help us ?</p>\n<p>Thomas</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1038705,
      "author_name": "Mike L.",
      "author_url": "",
      "post_date": "2020-10-06T01:27:09.513000",
      "content": "<p>Hi Sohier, thank you for this competition. It is designed around Python. I already have software that predicts educational performance, but it is written in other computer languages. Can I extract the competition data (piece-wise if necessary) using the Python API to disk. Then analyze the data and make predictions using other software. Then upload the predictions from disk using the Python API? Thanks.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1039540,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2020-10-06T16:29:14.650000",
          "content": "<p>You're welcome to give it a shot as long as the software adheres to the competition rules, though I'd expect it will be quite slow. The API will force you to loop to and from the disk in a piece wise manner.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1143991,
      "author_name": "Avani V",
      "author_url": "",
      "post_date": "2021-01-08T07:02:57.707000",
      "content": "<p>It was a nice learning experience.</p>\n<p>Couldn't have 'competed' here, but tried to see what score one gets using minimum set of data and without AI / sophisticated statistics. (as in here. <a href=\"https://www.kaggle.com/avaniv/riiid-threefields-noai\" target=\"_blank\">https://www.kaggle.com/avaniv/riiid-threefields-noai</a>)</p>\n<p>May be AI competitions here can have non AI section that helps in better understanding of value addition of additional data and AI over minimalist / analytical approaches?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1142520,
      "author_name": "林有夕",
      "author_url": "",
      "post_date": "2021-01-07T12:48:20.037000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1192725%2F5f39de855040973bdf4598e35dddd00b%2F1ef84cacfe6a3fad8cbd2578bed3e6cd.png?generation=1610023643061539&amp;alt=media\" alt=\"\"></p>\n<p>Since yesterday, there has been a bug in my kernel. Please help to have a look<br>\n<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209222\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209222</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1126502,
      "author_name": "Allohvk",
      "author_url": "",
      "post_date": "2020-12-25T16:56:29.160000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> - request you to fix the Kaggle bug for my userid. Sample submission (no changes to sample submission kernel of urs) is also not getting recognized and the status of the run is NOT updated under my submissions (no error..the status itself is not there). Request your help to fix it urgently as competition has reached its last stage</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1129868,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2020-12-28T15:49:04.313000",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> request you to fix the Kaggle bug for my userId. It is the same issue. I believe couple of other users are also facing similar problems<br>\nI just gave a run a few mins back for the sample submission kernel. There is no status. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1117489,
      "author_name": "kavin_hope",
      "author_url": "",
      "post_date": "2020-12-18T05:57:48.233000",
      "content": "<p>Thank you for this opportunity to learn.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1098476,
      "author_name": "AK",
      "author_url": "",
      "post_date": "2020-12-01T16:59:29.067000",
      "content": "<p>Quick Question:</p>\n<p>Kernel run for 1-2 hours but when we hit on the submisssion tab it takes more than 5 hours why so ?</p>\n<p>How many rows of test  it is computing?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1098479,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-01T17:01:13.977000",
          "content": "<p>2.5M rows. The faster your inference pipeline, the lesser the time it takes.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1098500,
          "author_name": "Koubouratou IDJATON",
          "author_url": "",
          "post_date": "2020-12-01T17:15:01.153000",
          "content": "<p>Also the notebook you use to submit to the competition can just be an inference notebook.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1086670,
      "author_name": "Patrick Brazil",
      "author_url": "",
      "post_date": "2020-11-21T23:15:31.077000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>  </p>\n<p>I am a little perplexed by what some features mean. </p>\n<p>In the training data set we have:<br>\ncontent_id: (int16) ID code for the user interaction<br>\ntask_container_id: (int16) Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id.</p>\n<p>In the questions csv we have:<br>\nquestion_id: foreign key for the train/test content_id column, when the content type is question (0).<br>\nbundle_id: code for which questions are served together.</p>\n<p>Now my interpretation is (assuming content_type_id is 0 i.e. we are dealing with a question) the content_id in the training data tells us which row/id to use in the questions csv. </p>\n<p>Now if we check the task_container_id and check the value against the row in the questions csv matching the content_id I would expect them to match. In other words, I interpreted task_contain_id to tell us which questions are bundled together and bundle_id to do the same. This is not true.</p>\n<p>When I look at the data in the questions csv each row has its own unique bundle_id. In fact, the id column, the question_id column, and the bundle_id column all have the exact same entry in every row and match the row number. </p>\n<p>Am I misinterpreting something here?</p>\n<p>I would be thankful for any input.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1079354,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2020-11-16T00:44:42.030000",
      "content": "<p>Hi Sohier, could you please tell me whether we can use an offline trained model for the final submission? Currently, I do data processing and model training offline. And only upload the trained model and related statistics to Kaggle kernel for batch inference.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1079491,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2020-11-16T06:07:46.723000",
          "content": "<p>The notebook you use to submit to the competition can just be an inference notebook. If you win the competition, you'd be required to share the full model to the hosts.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1048140,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2020-10-13T08:27:32.213000",
      "content": "<p>I doubt there is a wrong statement from <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> in the Data Page:</p>\n<p><code>Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual.</code></p>\n<p>When I check unique question_id in questions.csv and compare with unique content_id where content_type_id==0 (corresponding to questions) in train.csv, they are the same.</p>\n<p>Please clarify. Maybe during private test run, a different question.csv is swapped in?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1048542,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2020-10-13T15:32:46.353000",
          "content": "<blockquote>\n  <p>Maybe during private test run, a different question.csv is swapped in?</p>\n</blockquote>\n<p>Yes, that's the code competition. Same as we cannot see test which is evaluated.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1048797,
          "author_name": "GauthamKumaran",
          "author_url": "",
          "post_date": "2020-10-13T19:37:27.053000",
          "content": "<p>I've raised the same question here - <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190415\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190415</a> and in this thread too.</p>\n<p>If a new questions.csv is swapped in, wouldn't it make using pre-trained models very tough?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1049431,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2020-10-14T12:14:51.683000",
          "content": "<p>Yeah, one possibility is that a different questions.csv is swapped in. I've been in Code Rerun competitions before and I fully understand how it works, but never seen any comp with another file swapped rather than test.csv. That's why I need a confirmation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1047793,
      "author_name": "ONODERA",
      "author_url": "",
      "post_date": "2020-10-13T00:20:23.880000",
      "content": "<p>Could you tell us how many rows are there in each group_num?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1047822,
          "author_name": "Ming Pan",
          "author_url": "",
          "post_date": "2020-10-13T01:25:55.883000",
          "content": "<p>From the data description:</p>\n<blockquote>\n  <p>The API provides user interactions groups in the order in which they occurred. Each group will contain interactions from many different users, but no more than one task_container_id of questions from any single user. Each group has between 1 and 1000 users.</p>\n</blockquote>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1047827,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2020-10-13T01:28:05.133000",
          "content": "<p>Thanks, I missed that</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1046578,
      "author_name": "GauthamKumaran",
      "author_url": "",
      "post_date": "2020-10-11T19:43:39.403000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>Can you please confirm that the contents of the files <code>train.csv</code>, <code>questions.csv</code>, and <code>lectures.csv</code> won't change (new rows won't be added) in the test environment? </p>\n<p>I'm asking this because it came up in this discussion - <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190415\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190415</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1077795,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-13T23:50:00.820000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1073638,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-09T18:21:22.307000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1073552,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-09T16:07:38.820000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1072634,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-08T14:26:35.987000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1073398,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-09T13:36:50.840000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1073440,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-09T14:32:46.490000",
          "content": "<p>This shouldn't happen, Check some kernel with successful sub.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1073491,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-09T15:15:42.777000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1039761,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-06T19:19:12.900000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1506055,
      "author_name": "Tong Zhou",
      "author_url": "",
      "post_date": "2021-09-07T19:27:08.543000",
      "content": "<p>Thanks! The info is truly helpful!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1038272": "Welcome to the Riiid! Answer Correctness Prediction competition! Assessing student progress from their past work is a critical task for online education and this challenge provides a large amount of real world data for that task.\nWith the competition API serving data, you'll also need to tackle the difficulties introduced by new students arriving on the site and new questions being added. We're excited to see how you tackle this challenge!\n\nHappy Kaggling!\n",
    "1038346": "This competition is a code competition, so, the prediction phase must be done in a Kaggle kernel. But can the models training phase be done offline on a public cloud VM or a personal computer?\n\nThanks.",
    "1044670": "It would be nice to have at-least the runtime stats of our submitted kernel on the submission's page against the private test set?",
    "1042185": "Hi @sohier,\n\nTo my understanding, there are some user_ids appear both in the train.csv and test set. When making predictions, is it allowed to use the data about the learning history (if there is any) of a user_id in the train.csv? I mean not just for training the model, but also feeding train.csv into the model when making predictions.\n\nThanks!",
    "1042236": "And one more question related  to the previous 2 by @frankpanxj and @sirishks : Can one assume that there are no gaps neither in the training data nor between training and test slices, nor between test slices? \n\nThe _Data_ tab states that the test batches follow historical order, but is says nothing about completeness:\n> The API provides user interactions groups in the order in which they occurred.\n\nIf one wants to build a training curve for individual users it is important to have complete history per user. For example, if we want to count the number of lectures the given user had up to the specific question, we need to be sure that in the training and test data that have been all user interactions recorded",
    "1040558": "@sohier In the Data page, section example_test.csv:\n`There are two different rows that mirror what information the AI tutor actually has available at any given time,...`\n\nI think it should be\n`There are two different columns...`",
    "1041155": "Hi Sohier, when I downloaded using Kaggle API and try to import riiideducation, I get the following error:\nModuleNotFoundError: No module named 'riiideducation.competition'\nCould you help?\nThanks!",
    "1038862": "hi sir, i need to understand that do we have to use \"competition.cpython-37m-x86_64-linux-gnu.so\" this file from dataset. i am not aware anything about linux. can i proceed with only dataset files and not above mentioned file in my windows machine or kaggle kernel?",
    "1042115": "@sohier Although you have already answered\n`we are limited in what we can disclose`\nin discussion topic \"[Assumptions about the dataset?](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188899)\"\n\nI would like to highlight especially the first question one more time viz.,\n`1) All questions in the test set come cronologically after the questions in the training dataset?`\ni.e. the **row_id** and **timestamp** of test data delivered by the time-series API are greater than those of the train rows.\n\nThis question is to ensure that there is no **leakage** and as such deserves to be answered.",
    "1039696": "@sohier It seems that there is an issue with the \"questions\" table when the solution is evaluated for the Leaderboard.\nAs you can see [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188986), people cannot make submissions on the LB when they use the \"questions\" table.\n\nCan you help us ?\n\nThomas",
    "1038705": "Hi Sohier, thank you for this competition. It is designed around Python. I already have software that predicts educational performance, but it is written in other computer languages. Can I extract the competition data (piece-wise if necessary) using the Python API to disk. Then analyze the data and make predictions using other software. Then upload the predictions from disk using the Python API? Thanks.",
    "1143991": "It was a nice learning experience.\n\nCouldn't have 'competed' here, but tried to see what score one gets using minimum set of data and without AI / sophisticated statistics. (as in here. https://www.kaggle.com/avaniv/riiid-threefields-noai)\n\nMay be AI competitions here can have non AI section that helps in better understanding of value addition of additional data and AI over minimalist / analytical approaches?",
    "1142520": "@sohier  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1192725%2F5f39de855040973bdf4598e35dddd00b%2F1ef84cacfe6a3fad8cbd2578bed3e6cd.png?generation=1610023643061539&alt=media)\n\nSince yesterday, there has been a bug in my kernel. Please help to have a look\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209222",
    "1126502": "@sohier - request you to fix the Kaggle bug for my userid. Sample submission (no changes to sample submission kernel of urs) is also not getting recognized and the status of the run is NOT updated under my submissions (no error..the status itself is not there). Request your help to fix it urgently as competition has reached its last stage",
    "1117489": "Thank you for this opportunity to learn.",
    "1098476": "Quick Question:\n\nKernel run for 1-2 hours but when we hit on the submisssion tab it takes more than 5 hours why so ?\n\nHow many rows of test  it is computing?",
    "1086670": "Hello @sohier  \n\nI am a little perplexed by what some features mean. \n\nIn the training data set we have:\ncontent_id: (int16) ID code for the user interaction\ntask_container_id: (int16) Id code for the batch of questions or lectures. For example, a user might see three questions in a row before seeing the explanations for any of them. Those three would all share a task_container_id.\n\nIn the questions csv we have:\nquestion_id: foreign key for the train/test content_id column, when the content type is question (0).\nbundle_id: code for which questions are served together.\n\nNow my interpretation is (assuming content_type_id is 0 i.e. we are dealing with a question) the content_id in the training data tells us which row/id to use in the questions csv. \n\nNow if we check the task_container_id and check the value against the row in the questions csv matching the content_id I would expect them to match. In other words, I interpreted task_contain_id to tell us which questions are bundled together and bundle_id to do the same. This is not true.\n\nWhen I look at the data in the questions csv each row has its own unique bundle_id. In fact, the id column, the question_id column, and the bundle_id column all have the exact same entry in every row and match the row number. \n\nAm I misinterpreting something here?\n\nI would be thankful for any input.\n",
    "1079354": "Hi Sohier, could you please tell me whether we can use an offline trained model for the final submission? Currently, I do data processing and model training offline. And only upload the trained model and related statistics to Kaggle kernel for batch inference.",
    "1048140": "I doubt there is a wrong statement from @sohier in the Data Page:\n\n`Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual.`\n\nWhen I check unique question_id in questions.csv and compare with unique content_id where content_type_id==0 (corresponding to questions) in train.csv, they are the same.\n\nPlease clarify. Maybe during private test run, a different question.csv is swapped in?",
    "1047793": "Could you tell us how many rows are there in each group_num?",
    "1046578": "Hi @sohier \n\nCan you please confirm that the contents of the files `train.csv`, `questions.csv`, and `lectures.csv` won't change (new rows won't be added) in the test environment? \n\nI'm asking this because it came up in this discussion - https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190415",
    "1077795": "",
    "1073638": "",
    "1073552": "",
    "1072634": "",
    "1039761": "",
    "1506055": "Thanks! The info is truly helpful!"
  }
}