{
  "id": 191106,
  "title": "Updates, corrections, and clarifications",
  "url": "/competitions/riiid-test-answer-prediction/discussion/191106",
  "author_name": "Sohier Dane",
  "post_date": "2020-10-14T17:25:00.392000",
  "votes": 152,
  "comment_count": 60,
  "views": 0,
  "content": "<ul>\n<li><p>I'm updating the dataset now so lecture tags will match the tags in <strong>questions.csv</strong>. The updated <strong>lectures.csv</strong> should be live within the next hour. Apologies for the inconvenience.</p></li>\n<li><p>Correction- the hidden test set contains new <em>users</em> but not new <em>questions</em>.</p></li>\n<li><p>The train/test data is complete, in the sense that there are no missing interactions in the union of train and test data. It remains possible that some questions weren't logged due to other issues that all datasets of mobile users are susceptible to,<br>\nsuch as if a user lost their connection mid-question.</p></li>\n<li><p>The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.</p></li>\n</ul>",
  "messages": [
    {
      "id": 1049702,
      "postDate": "2020-10-14T17:25:00.393Z",
      "content": "<ul>\n<li><p>I'm updating the dataset now so lecture tags will match the tags in <strong>questions.csv</strong>. The updated <strong>lectures.csv</strong> should be live within the next hour. Apologies for the inconvenience.</p></li>\n<li><p>Correction- the hidden test set contains new <em>users</em> but not new <em>questions</em>.</p></li>\n<li><p>The train/test data is complete, in the sense that there are no missing interactions in the union of train and test data. It remains possible that some questions weren't logged due to other issues that all datasets of mobile users are susceptible to,<br>\nsuch as if a user lost their connection mid-question.</p></li>\n<li><p>The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.</p></li>\n</ul>",
      "rawMarkdown": "- I'm updating the dataset now so lecture tags will match the tags in **questions.csv**. The updated **lectures.csv** should be live within the next hour. Apologies for the inconvenience.\n\n- Correction- the hidden test set contains new _users_ but not new _questions_.\n\n- The train/test data is complete, in the sense that there are no missing interactions in the union of train and test data. It remains possible that some questions weren't logged due to other issues that all datasets of mobile users are susceptible to,\nsuch as if a user lost their connection mid-question.\n\n- The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.",
      "votes": 150
    },
    {
      "id": 1049825,
      "postDate": "2020-10-14T20:04:54.367Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Thanks for the clarification! Now we can start to build the sequential model. Would you mind to answer the following questions?</p>\n<p>Q1. When we submit a notebook for this competition, does it run on the whole private test dataset (and 20% of them are used to calculate the public LB score).</p>\n<p>Q2. Will our submissions be re-run after the deadline in order to compute the private LB score? Or it will just use the score for the remaining 80% private test dataset that is potentially already calculated when we submit?</p>\n<p>Q3. If the submission will be re-run after the deadline, would questions.csv, lectures.csv and the private test dataset remain the same as when we submit for the public LB score before the deadline?</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190791\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190791</a></p>",
      "rawMarkdown": "@sohier Thanks for the clarification! Now we can start to build the sequential model. Would you mind to answer the following questions?\n\n\nQ1. When we submit a notebook for this competition, does it run on the whole private test dataset (and 20% of them are used to calculate the public LB score).\n\nQ2. Will our submissions be re-run after the deadline in order to compute the private LB score? Or it will just use the score for the remaining 80% private test dataset that is potentially already calculated when we submit?\n\nQ3. If the submission will be re-run after the deadline, would questions.csv, lectures.csv and the private test dataset remain the same as when we submit for the public LB score before the deadline?\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190791",
      "votes": 5
    },
    {
      "id": 1055372,
      "postDate": "2020-10-20T17:50:47.367Z",
      "content": "<p>Roughly what percentage of users in the hidden test set will be new users?</p>",
      "rawMarkdown": "Roughly what percentage of users in the hidden test set will be new users?",
      "votes": 4
    },
    {
      "id": 1065394,
      "postDate": "2020-10-31T08:36:39.717Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> When I try to use TPUs to make a submission it errs with \"Your Notebook cannot use TPUs in this competition.\" But as per the code requirements we can use TPUs, so is this a glitch?</p>",
      "rawMarkdown": "@sohier When I try to use TPUs to make a submission it errs with \"Your Notebook cannot use TPUs in this competition.\" But as per the code requirements we can use TPUs, so is this a glitch?",
      "votes": 3,
      "replies": [
        {
          "id": 1065401,
          "postDate": "2020-10-31T08:51:48.433Z",
          "content": "<p>The Code Requirements have been updated / changed since the start of the competition.<br>\nInitially internet access was not allowed either but that condition isn't mentioned anymore.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nWould be good to get clarification / confirmation on this since there has been no official post about the  updates.</p>",
          "rawMarkdown": "The Code Requirements have been updated / changed since the start of the competition.\nInitially internet access was not allowed either but that condition isn't mentioned anymore.\n\n@sohier \nWould be good to get clarification / confirmation on this since there has been no official post about the  updates.",
          "votes": 4
        },
        {
          "id": 1068689,
          "postDate": "2020-11-03T16:43:26.677Z",
          "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> I just tried to submit a notebook with Internet access turned on, and I got </p>\n<blockquote>\n  <p>Your Notebook cannot use internet access in this competition. Please disable internet in the Notebook editor and save a new version.</p>\n</blockquote>",
          "rawMarkdown": "@rohanrao I just tried to submit a notebook with Internet access turned on, and I got \n> Your Notebook cannot use internet access in this competition. Please disable internet in the Notebook editor and save a new version.",
          "votes": 1
        },
        {
          "id": 1069885,
          "postDate": "2020-11-05T03:40:58.783Z",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> needs to confirm regarding the mismatch between Code Requirements and actual backend since something changed.</p>",
          "rawMarkdown": "@sohier needs to confirm regarding the mismatch between Code Requirements and actual backend since something changed.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1050278,
      "postDate": "2020-10-15T08:08:37.387Z",
      "content": "<p>Thanks for the clarifications. I have a request that could you please show more details of the \"Submission Scoring Error\" when we make submission. Each submission spends really long time and finnally \"Submission Scoring Error\"  without any addition details. It's better to add some logs of the error, so we will not waste our time in debug and forcus more on the feature engineering and modeling.</p>",
      "rawMarkdown": "Thanks for the clarifications. I have a request that could you please show more details of the \"Submission Scoring Error\" when we make submission. Each submission spends really long time and finnally \"Submission Scoring Error\"  without any addition details. It's better to add some logs of the error, so we will not waste our time in debug and forcus more on the feature engineering and modeling.",
      "votes": 3,
      "replies": [
        {
          "id": 1050342,
          "postDate": "2020-10-15T09:49:25.227Z",
          "content": "<p>They cannot add it else people will dump the whole data into logs, you can always raise an error at will.</p>",
          "rawMarkdown": "They cannot add it else people will dump the whole data into logs, you can always raise an error at will.",
          "votes": 5
        }
      ]
    },
    {
      "id": 1049944,
      "postDate": "2020-10-14T22:51:46.527Z",
      "content": "<p>Thanks for the clarifications. Will the hidden set not contain new lectures as well?</p>",
      "rawMarkdown": "Thanks for the clarifications. Will the hidden set not contain new lectures as well?\n ",
      "votes": 3,
      "replies": [
        {
          "id": 1060721,
          "postDate": "2020-10-26T14:00:22.507Z",
          "content": "<p>well, it's obvious that if there is no new question that implies there no new lectures.</p>",
          "rawMarkdown": "well, it's obvious that if there is no new question that implies there no new lectures.\n"
        },
        {
          "id": 1060731,
          "postDate": "2020-10-26T14:09:16.593Z",
          "content": "<p>questions and lectures are different datasets so it's not completely obvious. There can be new lectures.</p>",
          "rawMarkdown": "questions and lectures are different datasets so it's not completely obvious. There can be new lectures.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1060803,
      "postDate": "2020-10-26T14:50:11.390Z",
      "content": "<p>Roughly what percentage of users in the hidden test set will be new users?</p>",
      "rawMarkdown": "Roughly what percentage of users in the hidden test set will be new users?",
      "votes": 2,
      "replies": [
        {
          "id": 1060827,
          "postDate": "2020-10-26T14:57:50.530Z",
          "content": "<ol>\n<li><p>Just like in real-world scenarios, we look at historical (train) data, observe timestamps and frequency of users in order to estimate future (test) data distribution.</p></li>\n<li><p>Since the data is time-series, you could also take 1st 80% data as train and last 20% data as test, and then perform this calculation viz., how many users are \"new\" i.e. present only in the last 20% of the train data.</p></li>\n</ol>",
          "rawMarkdown": "1. Just like in real-world scenarios, we look at historical (train) data, observe timestamps and frequency of users in order to estimate future (test) data distribution.\n\n2. Since the data is time-series, you could also take 1st 80% data as train and last 20% data as test, and then perform this calculation viz., how many users are \"new\" i.e. present only in the last 20% of the train data.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1068603,
      "postDate": "2020-11-03T15:09:23.620Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> and community! I have one last question about the code competition set up. It has been a while since I have participated in code competition but it looks like we can train models offline these days and load them to the kernel environment.</p>\n<blockquote>\n  <p>Please note that for this competition training is not required in Notebooks.</p>\n</blockquote>\n<p>My question is if we have to make it available to the community as well? Does it count as a pre-trained model?</p>\n<blockquote>\n  <p>Freely &amp; publicly available external data is allowed, including pre-trained models</p>\n</blockquote>\n<p>I am sure this has been answered multiple times since the code competitions have been updated but I wasn't sure where to look.</p>",
      "rawMarkdown": "Hey @sohier and community! I have one last question about the code competition set up. It has been a while since I have participated in code competition but it looks like we can train models offline these days and load them to the kernel environment.\n\n> Please note that for this competition training is not required in Notebooks.\n\nMy question is if we have to make it available to the community as well? Does it count as a pre-trained model?\n\n> Freely & publicly available external data is allowed, including pre-trained models\n\nI am sure this has been answered multiple times since the code competitions have been updated but I wasn't sure where to look.",
      "votes": 1,
      "replies": [
        {
          "id": 1134308,
          "postDate": "2021-01-01T06:03:08.023Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rdizzl3\" target=\"_blank\">@rdizzl3</a>, did you find an answer for these questions? I am also interested. Thanks and good luck in this competition!</p>",
          "rawMarkdown": "Hi @rdizzl3, did you find an answer for these questions? I am also interested. Thanks and good luck in this competition!"
        }
      ]
    },
    {
      "id": 1066177,
      "postDate": "2020-11-01T12:22:47.757Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> for the clarification.</p>",
      "rawMarkdown": "Thank you @sohier for the clarification.",
      "votes": 1
    },
    {
      "id": 1061416,
      "postDate": "2020-10-27T02:16:20.263Z",
      "content": "<p>content_type_id: Please include a lecture in the test data (Example.csv) so that we can verify that lectures are handled correctly before submitting to the hidden data.</p>\n<p>Thanks.</p>",
      "rawMarkdown": "content_type_id: Please include a lecture in the test data (Example.csv) so that we can verify that lectures are handled correctly before submitting to the hidden data.\n\nThanks.",
      "votes": 1,
      "replies": [
        {
          "id": 1061521,
          "postDate": "2020-10-27T04:01:20.410Z",
          "content": "<p>Mike, you can take a look at <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856\" target=\"_blank\">this</a>!</p>",
          "rawMarkdown": "Mike, you can take a look at [this](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856)!"
        },
        {
          "id": 1061536,
          "postDate": "2020-10-27T04:28:47.707Z",
          "content": "<p>Thanks. Now have a solution to my problem. Here is how to insert lectures into the Example dataframe test_df. This works for me 😊</p>\n<p>Improvements welcome!</p>\n<p>Borrowing from: <a href=\"https://www.geeksforgeeks.org/insert-row-at-given-position-in-pandas-dataframe/\" target=\"_blank\">https://www.geeksforgeeks.org/insert-row-at-given-position-in-pandas-dataframe/</a></p>\n<h1>Function to insert row in the dataframe</h1>\n<p>def Insert_row(row_number, df, row_value): <br>\n    # Starting value of upper half <br>\n    start_upper = 0</p>\n<pre><code># End value of upper half \nend_upper = row_number \n\n# Start value of lower half \nstart_lower = row_number \n\n# End value of lower half \nend_lower = df.shape[0] \n\n# Create a list of upper_half index \nupper_half = [*range(start_upper, end_upper, 1)] \n\n# Create a list of lower_half index \nlower_half = [*range(start_lower, end_lower, 1)] \n\n# Increment the value of lower half by 1 \nlower_half = [x.__add__(1) for x in lower_half] \n\n# Combine the two lists \nindex_ = upper_half + lower_half \n\n# Update the index of the dataframe \ndf.index = index_ \n\n# Insert a row at the end \ndf.loc[row_number] = row_value \n\n# Sort the index labels \ndf = df.sort_index() \n\n# return the dataframe \nreturn df \n</code></pre>\n<p>iter_test = env.iter_test()<br>\nfor (test_df, sample_prediction_df) in iter_test:</p>\n<pre><code># Let's create a row which we want to insert \n row_number = 2\n row_value = [89,653762,2746,6808,1,14,-1,-1,0,False] \n\n if row_number &gt; test_df.index.max()+1: \n    print(\"Invalid row_number\", row_number, test_df.index.max()) \n else: \n\n# Let's call the function and insert the row \n    test_df = Insert_row(row_number, test_df, row_value) \n\n# Print the updated dataframe \n    print(test_df) \n</code></pre>\n<p>….</p>",
          "rawMarkdown": "Thanks. Now have a solution to my problem. Here is how to insert lectures into the Example dataframe test_df. This works for me 😊\n\nImprovements welcome!\n\nBorrowing from: https://www.geeksforgeeks.org/insert-row-at-given-position-in-pandas-dataframe/\n\n# Function to insert row in the dataframe \ndef Insert_row(row_number, df, row_value): \n    # Starting value of upper half \n    start_upper = 0\n   \n    # End value of upper half \n    end_upper = row_number \n   \n    # Start value of lower half \n    start_lower = row_number \n   \n    # End value of lower half \n    end_lower = df.shape[0] \n   \n    # Create a list of upper_half index \n    upper_half = [*range(start_upper, end_upper, 1)] \n   \n    # Create a list of lower_half index \n    lower_half = [*range(start_lower, end_lower, 1)] \n   \n    # Increment the value of lower half by 1 \n    lower_half = [x.__add__(1) for x in lower_half] \n   \n    # Combine the two lists \n    index_ = upper_half + lower_half \n   \n    # Update the index of the dataframe \n    df.index = index_ \n   \n    # Insert a row at the end \n    df.loc[row_number] = row_value \n    \n    # Sort the index labels \n    df = df.sort_index() \n   \n    # return the dataframe \n    return df \n\niter_test = env.iter_test()\nfor (test_df, sample_prediction_df) in iter_test:\n\n    # Let's create a row which we want to insert \n     row_number = 2\n     row_value = [89,653762,2746,6808,1,14,-1,-1,0,False] \n\n     if row_number > test_df.index.max()+1: \n        print(\"Invalid row_number\", row_number, test_df.index.max()) \n     else: \n      \n    # Let's call the function and insert the row \n        test_df = Insert_row(row_number, test_df, row_value) \n   \n    # Print the updated dataframe \n        print(test_df) \n\n....",
          "votes": 1
        },
        {
          "id": 1067188,
          "postDate": "2020-11-02T11:48:52.100Z",
          "content": "<p><a href=\"https://www.kaggle.com/mikel1\" target=\"_blank\">@mikel1</a> <br>\n1) could u help understand the need for  insert of new lectures ?</p>\n<p>2) Can test set contain only question ids that are present in questions.csv ?</p>",
          "rawMarkdown": "@mikel1 \n1) could u help understand the need for  insert of new lectures ?\n\n2) Can test set contain only question ids that are present in questions.csv ?"
        }
      ]
    },
    {
      "id": 1054875,
      "postDate": "2020-10-20T08:47:22.993Z",
      "content": "<p>Sohier, in your very instructive \"Competition API Detailed Introduction\", you stop a little too soon.</p>\n<p>Here are the next steps:</p>\n<p>1) Have a Notebook that successfully predicts the Example test data when \"Run All\" is clicked<br>\n    Be sure that Setting \"Internet\" is off.<br>\n2) Save to share.<br>\n3) Share with the public<br>\n4) On the Notebooks page, click on it <br>\n5) Execution info<br>\n6) Click on \"Submit\"<br>\nYou will be told that your Notebook is running.</p>",
      "rawMarkdown": "Sohier, in your very instructive \"Competition API Detailed Introduction\", you stop a little too soon.\n\nHere are the next steps:\n\n1) Have a Notebook that successfully predicts the Example test data when \"Run All\" is clicked\n    Be sure that Setting \"Internet\" is off.\n2) Save to share.\n3) Share with the public\n4) On the Notebooks page, click on it \n5) Execution info\n6) Click on \"Submit\"\nYou will be told that your Notebook is running.",
      "votes": 1
    },
    {
      "id": 1049921,
      "postDate": "2020-10-14T22:13:05.127Z",
      "content": "<p>Thanks for the clarification! The new <strong>question_id</strong> in the hidden test set and the lack of sequence of test data have bothered me for a long time.</p>",
      "rawMarkdown": "Thanks for the clarification! The new **question_id** in the hidden test set and the lack of sequence of test data have bothered me for a long time.",
      "votes": 1
    },
    {
      "id": 1123192,
      "postDate": "2020-12-23T02:58:12.217Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> apologies if this has been covered elsewhere - I read through the updates and searched the other discussions, but didn't see anything.</p>\n<p>I ran <a href=\"https://www.kaggle.com/calebeverett/test-tid-deltas\" target=\"_blank\">this book</a> to see whether there were many instances of gaps of more than one in the sequence of task_container_ids.  I set it to error out if the count got over 100,000 and it errored out soon after it started. I understand there may be gaps due to the timing of user interactions on multiple devices and as a result task_container_ids are not sequential when records are ordered for each user by timestamp. I calculated that there are 2.5 million of such instances, in the training set,  or approximately 2.5% of all of the observations. If the proportion was similar in the test set that would equate to approximately 62,500 instances, so erroring out after 100,000 would seem to indicate another source of them not being sequential.</p>",
      "rawMarkdown": "@sohier apologies if this has been covered elsewhere - I read through the updates and searched the other discussions, but didn't see anything.\n\nI ran [this book](https://www.kaggle.com/calebeverett/test-tid-deltas) to see whether there were many instances of gaps of more than one in the sequence of task_container_ids.  I set it to error out if the count got over 100,000 and it errored out soon after it started. I understand there may be gaps due to the timing of user interactions on multiple devices and as a result task_container_ids are not sequential when records are ordered for each user by timestamp. I calculated that there are 2.5 million of such instances, in the training set,  or approximately 2.5% of all of the observations. If the proportion was similar in the test set that would equate to approximately 62,500 instances, so erroring out after 100,000 would seem to indicate another source of them not being sequential.",
      "votes": 2
    },
    {
      "id": 1061101,
      "postDate": "2020-10-26T18:35:28.533Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> thank you for the corrections and clarifications. I am about to enter the competition but I feel like I don't quite understand how the test API works for the new test users. I am looking at the <code>example_test.csv</code> and there is a user present in the test data and not in the training data.</p>\n<p>Here is what that user looks like:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F291298%2Feb4ec9b30ff5ceae322d677dba13aff4%2FScreen%20Shot%202020-10-26%20at%2011.24.00%20AM.png?generation=1603736688404605&amp;alt=media\" alt=\"\"></p>\n<p>My first question is making sure I understand the groups correctly - </p>\n<ol>\n<li><p>The lists from group 1 are the correct answers, etc. from questions in group 0?</p></li>\n<li><p>If the list from group 1 is their data from group zero, shouldn't there be different content ids (possibly) for the different questions? There is only one content id in group 0.</p></li>\n</ol>\n<p>Sorry if this has been explained somewhere else. I read quite a few topics in the forum but couldn't figure out how to correctly map and use this info.</p>",
      "rawMarkdown": "Hey @sohier thank you for the corrections and clarifications. I am about to enter the competition but I feel like I don't quite understand how the test API works for the new test users. I am looking at the `example_test.csv` and there is a user present in the test data and not in the training data.\n\nHere is what that user looks like:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F291298%2Feb4ec9b30ff5ceae322d677dba13aff4%2FScreen%20Shot%202020-10-26%20at%2011.24.00%20AM.png?generation=1603736688404605&alt=media)\n\nMy first question is making sure I understand the groups correctly - \n1. The lists from group 1 are the correct answers, etc. from questions in group 0?\n\n2. If the list from group 1 is their data from group zero, shouldn't there be different content ids (possibly) for the different questions? There is only one content id in group 0.\n\nSorry if this has been explained somewhere else. I read quite a few topics in the forum but couldn't figure out how to correctly map and use this info.",
      "votes": 2,
      "replies": [
        {
          "id": 1061108,
          "postDate": "2020-10-26T18:45:51.650Z",
          "content": "<blockquote>\n  <p>The lists from group 1 are the correct answers, etc. from questions in group 0?</p>\n</blockquote>\n<p>Yes.</p>\n<blockquote>\n  <p>There is only one content id in group 0.</p>\n</blockquote>\n<p>Don't filter on a user (it is irrelevant / incorrect). The length of the answer list in <code>group n</code> will exactly match the number of rows in <code>group (n-1)</code> (considering all users together).</p>\n<p>Highly recommend to go through <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124\" target=\"_blank\">this</a>.</p>",
          "rawMarkdown": "> The lists from group 1 are the correct answers, etc. from questions in group 0?\n\nYes.\n\n> There is only one content id in group 0.\n\nDon't filter on a user (it is irrelevant / incorrect). The length of the answer list in `group n` will exactly match the number of rows in `group (n-1)` (considering all users together).\n\nHighly recommend to go through [this](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124).",
          "votes": 3
        },
        {
          "id": 1061115,
          "postDate": "2020-10-26T18:53:54.480Z",
          "content": "<p>Plus there's video's labels in that list as well! So be careful with that as we don't have to make preds on video_rows so to say (content_type_id == 1).</p>",
          "rawMarkdown": "Plus there's video's labels in that list as well! So be careful with that as we don't have to make preds on video_rows so to say (content_type_id == 1).",
          "votes": 2
        },
        {
          "id": 1061122,
          "postDate": "2020-10-26T19:01:19.230Z",
          "content": "<p>thank you <a href=\"https://www.kaggle.com/vopani\" target=\"_blank\">@vopani</a>! I did not realize that topic had all this info! But I was able to verify everything I was questioning. Appreciate the comment!</p>",
          "rawMarkdown": "thank you @vopani! I did not realize that topic had all this info! But I was able to verify everything I was questioning. Appreciate the comment!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1050025,
      "postDate": "2020-10-15T02:14:32.560Z",
      "content": "<p>Thanks for the clarifications. Will it be possible to simply return the data from questions and lectures pre-appended to the test-df as well? Because it will be joined by us either ways, so IMO it's a good thing to return it when we are iterating over the test_df as it's columns?</p>\n<p>And the only file updated is lectures.csv, right?</p>",
      "rawMarkdown": "Thanks for the clarifications. Will it be possible to simply return the data from questions and lectures pre-appended to the test-df as well? Because it will be joined by us either ways, so IMO it's a good thing to return it when we are iterating over the test_df as it's columns?\n\nAnd the only file updated is lectures.csv, right?",
      "votes": 2
    },
    {
      "id": 1050004,
      "postDate": "2020-10-15T01:17:38.407Z",
      "content": "<p>Thanks for the clarification about private set. Now it help and have better clarification about private set.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Do you mean.. <strong>train/test set is complete</strong>  having all user activities logged except <strong>few missing</strong> values are <strong>negligible ratio</strong> ?</p>",
      "rawMarkdown": "Thanks for the clarification about private set. Now it help and have better clarification about private set.\n\n@sohier Do you mean.. **train/test set is complete**  having all user activities logged except **few missing** values are **negligible ratio** ?",
      "votes": 2
    },
    {
      "id": 1134231,
      "postDate": "2021-01-01T03:14:54.877Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Thanks for the clarification about private set. Now it help and have better clarification about private set.</p>\n<blockquote>\n  <p>The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.</p>\n</blockquote>\n<p>Regarding this matter, does chronologically mean that in order of <strong>timestamp</strong>, rather than task_container_id?</p>",
      "rawMarkdown": "@sohier Thanks for the clarification about private set. Now it help and have better clarification about private set.\n\n> The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.\n\nRegarding this matter, does chronologically mean that in order of **timestamp**, rather than task_container_id?"
    },
    {
      "id": 1121782,
      "postDate": "2020-12-21T22:44:02.397Z",
      "content": "<p>I was curious if the task_container_ids continue in sequence from the training data to the test data. It looks like there are gaps between the training data and sample test data.</p>\n<p>Below are the first ten sample test records.</p>\n<table>\n  <thead>\n    <tr>\n      <th></th>\n      <th>user_id</th>\n      <th>row_id</th>\n      <th>task_container_id</th>\n      <th>max_train_tid</th>\n      <th>tid_delta</th>\n      <th>ts_delta_sec</th>\n    </tr>\n  </thead>\n  <tbody>\n    <tr>\n      <th>1</th>\n      <td>554169193</td>\n      <td>1</td>\n      <td>4427</td>\n      <td>4421.0</td>\n      <td>6.0</td>\n      <td>184.627</td>\n    </tr>\n    <tr>\n      <th>2</th>\n      <td>1720860329</td>\n      <td>2</td>\n      <td>240</td>\n      <td>235.0</td>\n      <td>5.0</td>\n      <td>299.297</td>\n    </tr>\n    <tr>\n      <th>3</th>\n      <td>288641214</td>\n      <td>3</td>\n      <td>266</td>\n      <td>262.0</td>\n      <td>4.0</td>\n      <td>72252.433</td>\n    </tr>\n    <tr>\n      <th>4</th>\n      <td>1728340777</td>\n      <td>4</td>\n      <td>162</td>\n      <td>161.0</td>\n      <td>1.0</td>\n      <td>67009.619</td>\n    </tr>\n    <tr>\n      <th>5</th>\n      <td>1364159702</td>\n      <td>5</td>\n      <td>4424</td>\n      <td>4421.0</td>\n      <td>3.0</td>\n      <td>173642.470</td>\n    </tr>\n    <tr>\n      <th>6</th>\n      <td>1521618396</td>\n      <td>6</td>\n      <td>1367</td>\n      <td>1361.0</td>\n      <td>6.0</td>\n      <td>202.134</td>\n    </tr>\n    <tr>\n      <th>7</th>\n      <td>1317245193</td>\n      <td>7</td>\n      <td>5314</td>\n      <td>5307.0</td>\n      <td>7.0</td>\n      <td>384674.701</td>\n    </tr>\n    <tr>\n      <th>8</th>\n      <td>1700555100</td>\n      <td>8</td>\n      <td>532</td>\n      <td>530.0</td>\n      <td>2.0</td>\n      <td>246.422</td>\n    </tr>\n    <tr>\n      <th>9</th>\n      <td>998511398</td>\n      <td>9</td>\n      <td>393</td>\n      <td>389.0</td>\n      <td>4.0</td>\n      <td>37118.438</td>\n    </tr>\n    <tr>\n      <th>10</th>\n      <td>1422853669</td>\n      <td>10</td>\n      <td>85</td>\n      <td>79.0</td>\n      <td>6.0</td>\n      <td>194.973</td>\n    </tr>\n  </tbody>\n</table>\n<p>Are those gaps related to the preparation of the sample test data or are there similar gaps in the actual test data as well?</p>\n<p>EDIT: updated to include timetamp delta as well, which look like they might span gaps in task_container_ids. I had a median gap of approx. 42 seconds on the last set of user training records.</p>",
      "rawMarkdown": "I was curious if the task_container_ids continue in sequence from the training data to the test data. It looks like there are gaps between the training data and sample test data.\n\nBelow are the first ten sample test records.\n\n<table border=\"1\" class=\"dataframe\">\n  <thead>\n    <tr style=\"text-align: right;\">\n      <th></th>\n      <th>user_id</th>\n      <th>row_id</th>\n      <th>task_container_id</th>\n      <th>max_train_tid</th>\n      <th>tid_delta</th>\n      <th>ts_delta_sec</th>\n    </tr>\n  </thead>\n  <tbody>\n    <tr>\n      <th>1</th>\n      <td>554169193</td>\n      <td>1</td>\n      <td>4427</td>\n      <td>4421.0</td>\n      <td>6.0</td>\n      <td>184.627</td>\n    </tr>\n    <tr>\n      <th>2</th>\n      <td>1720860329</td>\n      <td>2</td>\n      <td>240</td>\n      <td>235.0</td>\n      <td>5.0</td>\n      <td>299.297</td>\n    </tr>\n    <tr>\n      <th>3</th>\n      <td>288641214</td>\n      <td>3</td>\n      <td>266</td>\n      <td>262.0</td>\n      <td>4.0</td>\n      <td>72252.433</td>\n    </tr>\n    <tr>\n      <th>4</th>\n      <td>1728340777</td>\n      <td>4</td>\n      <td>162</td>\n      <td>161.0</td>\n      <td>1.0</td>\n      <td>67009.619</td>\n    </tr>\n    <tr>\n      <th>5</th>\n      <td>1364159702</td>\n      <td>5</td>\n      <td>4424</td>\n      <td>4421.0</td>\n      <td>3.0</td>\n      <td>173642.470</td>\n    </tr>\n    <tr>\n      <th>6</th>\n      <td>1521618396</td>\n      <td>6</td>\n      <td>1367</td>\n      <td>1361.0</td>\n      <td>6.0</td>\n      <td>202.134</td>\n    </tr>\n    <tr>\n      <th>7</th>\n      <td>1317245193</td>\n      <td>7</td>\n      <td>5314</td>\n      <td>5307.0</td>\n      <td>7.0</td>\n      <td>384674.701</td>\n    </tr>\n    <tr>\n      <th>8</th>\n      <td>1700555100</td>\n      <td>8</td>\n      <td>532</td>\n      <td>530.0</td>\n      <td>2.0</td>\n      <td>246.422</td>\n    </tr>\n    <tr>\n      <th>9</th>\n      <td>998511398</td>\n      <td>9</td>\n      <td>393</td>\n      <td>389.0</td>\n      <td>4.0</td>\n      <td>37118.438</td>\n    </tr>\n    <tr>\n      <th>10</th>\n      <td>1422853669</td>\n      <td>10</td>\n      <td>85</td>\n      <td>79.0</td>\n      <td>6.0</td>\n      <td>194.973</td>\n    </tr>\n  </tbody>\n</table>\n\nAre those gaps related to the preparation of the sample test data or are there similar gaps in the actual test data as well?\n\nEDIT: updated to include timetamp delta as well, which look like they might span gaps in task_container_ids. I had a median gap of approx. 42 seconds on the last set of user training records."
    },
    {
      "id": 1117061,
      "postDate": "2020-12-17T17:30:35.587Z",
      "content": "<p>Is the part in lectures same as part in questions?</p>",
      "rawMarkdown": "Is the part in lectures same as part in questions?\n"
    },
    {
      "id": 1095031,
      "postDate": "2020-11-29T07:45:25.093Z",
      "content": "<p>That makes it clearer</p>",
      "rawMarkdown": "That makes it clearer"
    },
    {
      "id": 1085024,
      "postDate": "2020-11-20T15:59:35.760Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> I had a query that, In the column \"prior_question_had_explanation\"  the student can explain the question at any time or just after the question is answered?<br>\nplease look at this query<br>\n thanks</p>",
      "rawMarkdown": "@sohier I had a query that, In the column \"prior_question_had_explanation\"  the student can explain the question at any time or just after the question is answered?\nplease look at this query\n thanks"
    },
    {
      "id": 1079151,
      "postDate": "2020-11-15T17:16:24.223Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> and community, I'm still confused about the tags/tag columns  in <strong>questions.csv</strong> and <strong>lectures.csv</strong>. In this post you write, that these would \"match\" now, but there are 37 tags that occur in questions.csv that have no matching entries in lectures.csv. Am I missing something obvious here?</p>",
      "rawMarkdown": "Hi @sohier and community, I'm still confused about the tags/tag columns  in **questions.csv** and **lectures.csv**. In this post you write, that these would \"match\" now, but there are 37 tags that occur in questions.csv that have no matching entries in lectures.csv. Am I missing something obvious here?",
      "replies": [
        {
          "id": 1079199,
          "postDate": "2020-11-15T18:32:48.267Z",
          "content": "<p>I think of the same but in the above author discussion it's written as follows, </p>\n<blockquote>\n  <p>a)I'm updating the dataset now so <strong>lecture tags will match the tags in questions.csv</strong>. <br>\n  b)Correction- the hidden test set contains <strong>new users but not new questions</strong></p>\n</blockquote>\n<p>Coming to point 'B' there are no new questions which means there is a chance of introducing new lectures which then the tag of questions.csv and lecture.csv may match completely. </p>",
          "rawMarkdown": "I think of the same but in the above author discussion it's written as follows, \n\n> \na)I'm updating the dataset now so **lecture tags will match the tags in questions.csv**. \nb)Correction- the hidden test set contains **new users but not new questions**\n\nComing to point 'B' there are no new questions which means there is a chance of introducing new lectures which then the tag of questions.csv and lecture.csv may match completely. "
        }
      ]
    },
    {
      "id": 1067278,
      "postDate": "2020-11-02T12:43:17.600Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a><br>\n1)   why have we got multiple lecture tags associated to a single question id ,how do we interpret those</p>\n<p>2) Can a test set have different   lecture id than what is there in   lecture,csv that we are provided with</p>",
      "rawMarkdown": "@sohier\n1)   why have we got multiple lecture tags associated to a single question id ,how do we interpret those\n\n2) Can a test set have different   lecture id than what is there in   lecture,csv that we are provided with\n",
      "replies": [
        {
          "id": 1079217,
          "postDate": "2020-11-15T18:52:41.153Z",
          "content": "<p>for the first point, <br>\nlecture tag and question id are different things actually question_id is a content_id(which is in train.csv when cotent_type_id is zero ). and for the understanding of tag, it just says particular lecture or question belongs to that topic.</p>",
          "rawMarkdown": "for the first point, \nlecture tag and question id are different things actually question_id is a content_id(which is in train.csv when cotent_type_id is zero ). and for the understanding of tag, it just says particular lecture or question belongs to that topic.\n  "
        }
      ]
    },
    {
      "id": 1065700,
      "postDate": "2020-10-31T15:49:35.450Z",
      "content": "<blockquote>\n  <p>You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save &amp; Run All\"</p>\n</blockquote>\n<p>This looks great!</p>",
      "rawMarkdown": "> You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\"\n\nThis looks great!",
      "replies": [
        {
          "id": 1068727,
          "postDate": "2020-11-03T17:21:07.413Z",
          "content": "<p>Please tell me where the quote came from? (not met)</p>",
          "rawMarkdown": "Please tell me where the quote came from? (not met)"
        },
        {
          "id": 1069053,
          "postDate": "2020-11-04T02:40:58.557Z",
          "content": "<p>It comes from the pre-typed section on the kernel. Here's the formal <a href=\"https://www.kaggle.com/product-feedback/195163\" target=\"_blank\">announcement</a> now though…</p>",
          "rawMarkdown": "It comes from the pre-typed section on the kernel. Here's the formal [announcement](https://www.kaggle.com/product-feedback/195163) now though...",
          "votes": 1
        }
      ]
    },
    {
      "id": 1063501,
      "postDate": "2020-10-29T02:10:52.560Z",
      "content": "<p>In the 'Time-series API Details' section of <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/data\" target=\"_blank\">Data Description</a>, there is a statement,  </p>\n<blockquote>\n  <p>Expect to see roughly 2.5 million questions in the hidden test set.  </p>\n</blockquote>\n<p>I think the term 'questions' is misleading. Would you change that to 'rows' or 'interactions'?</p>",
      "rawMarkdown": "In the 'Time-series API Details' section of [Data Description](https://www.kaggle.com/c/riiid-test-answer-prediction/data), there is a statement,  \n> Expect to see roughly 2.5 million questions in the hidden test set.  \n\nI think the term 'questions' is misleading. Would you change that to 'rows' or 'interactions'?",
      "replies": [
        {
          "id": 1065260,
          "postDate": "2020-10-31T05:03:15.990Z",
          "content": "<p>Has it been confirmed somewhere that the test set has 2.5M rows (and not questions)?</p>\n<p>I've been assuming there are 2.5M questions and (based on train data %) an additional few thousand lectures.</p>",
          "rawMarkdown": "Has it been confirmed somewhere that the test set has 2.5M rows (and not questions)?\n\nI've been assuming there are 2.5M questions and (based on train data %) an additional few thousand lectures.",
          "votes": 3
        },
        {
          "id": 1067968,
          "postDate": "2020-11-02T23:24:44.640Z",
          "content": "<p>Thank you for your reply, Vopani.<br>\nI just thought that '2.5 million rows' would be more appropriate to explain the volume of test set.<br>\nI don't know that someone verified the volume or not.</p>",
          "rawMarkdown": "Thank you for your reply, Vopani.\nI just thought that '2.5 million rows' would be more appropriate to explain the volume of test set.\nI don't know that someone verified the volume or not.",
          "votes": 1
        },
        {
          "id": 1068102,
          "postDate": "2020-11-03T04:04:20.383Z",
          "content": "<p>Yes I agree using rows or observations is more appropriate.</p>",
          "rawMarkdown": "Yes I agree using rows or observations is more appropriate.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1062003,
      "postDate": "2020-10-27T13:52:39.920Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> for the clarification. </p>",
      "rawMarkdown": "Thank you @sohier for the clarification. "
    },
    {
      "id": 1059534,
      "postDate": "2020-10-25T08:02:56.573Z",
      "content": "<p>The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.<br>\n-- Does it mean all samples in test data happen after train data?  And GROUP 0 of test data happened before GROUP 1 of test data?</p>",
      "rawMarkdown": "The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.\n-- Does it mean all samples in test data happen after train data?  And GROUP 0 of test data happened before GROUP 1 of test data?",
      "replies": [
        {
          "id": 1059585,
          "postDate": "2020-10-25T08:48:33.517Z",
          "content": "<p>Yes that is my understanding and I'm assuming the API iterates over the groups chronologically, so group <code>n</code> happens before group <code>n+1</code></p>",
          "rawMarkdown": "Yes that is my understanding and I'm assuming the API iterates over the groups chronologically, so group `n` happens before group `n+1`",
          "votes": 1
        },
        {
          "id": 1060025,
          "postDate": "2020-10-25T17:32:15.220Z",
          "content": "<p>Yes, the User_ID in train set and test set are same that means test set is continuing train set but one more point to notice that there are some new User_ID in test set.<br>\nAnd the group num proptionals to time series</p>",
          "rawMarkdown": "Yes, the User_ID in train set and test set are same that means test set is continuing train set but one more point to notice that there are some new User_ID in test set.\nAnd the group num proptionals to time series",
          "votes": 1
        }
      ]
    },
    {
      "id": 1059533,
      "postDate": "2020-10-25T08:02:30.847Z",
      "content": "<p>Maybe I missed something but the defination of timestamp is vague. Can the host clarify what does \"first event completion from that user.\" exactly mean?   What does event completion imply? What happens when the (mobile) user never completes that event? </p>",
      "rawMarkdown": "Maybe I missed something but the defination of timestamp is vague. Can the host clarify what does \"first event completion from that user.\" exactly mean?   What does event completion imply? What happens when the (mobile) user never completes that event? "
    },
    {
      "id": 1053755,
      "postDate": "2020-10-19T10:27:31.780Z",
      "content": "<p>yiiiiihaaa</p>",
      "rawMarkdown": "yiiiiihaaa"
    },
    {
      "id": 1054629,
      "postDate": "2020-10-20T04:09:04.433Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1053439,
      "postDate": "2020-10-19T02:25:56.423Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1051706,
      "postDate": "2020-10-16T18:46:22.690Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1051688,
      "postDate": "2020-10-16T18:21:16.407Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1050323,
      "postDate": "2020-10-15T09:22:00.957Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1052395,
      "postDate": "2020-10-17T17:08:15.867Z",
      "content": "<p>Thank you! I really appreciate it. </p>",
      "rawMarkdown": "Thank you! I really appreciate it. "
    },
    {
      "id": 1052090,
      "postDate": "2020-10-17T09:37:15.097Z",
      "content": "<p>Thanks for the clarifications.</p>",
      "rawMarkdown": "Thanks for the clarifications."
    },
    {
      "id": 1051983,
      "postDate": "2020-10-17T06:27:56.870Z",
      "content": "<p>Thank for clearance</p>",
      "rawMarkdown": "Thank for clearance"
    },
    {
      "id": 1050747,
      "postDate": "2020-10-15T17:23:01.023Z",
      "content": "<p>thanks for clearance</p>",
      "rawMarkdown": "thanks for clearance"
    }
  ],
  "comments": [
    {
      "id": 1049825,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2020-10-14T20:04:54.367000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Thanks for the clarification! Now we can start to build the sequential model. Would you mind to answer the following questions?</p>\n<p>Q1. When we submit a notebook for this competition, does it run on the whole private test dataset (and 20% of them are used to calculate the public LB score).</p>\n<p>Q2. Will our submissions be re-run after the deadline in order to compute the private LB score? Or it will just use the score for the remaining 80% private test dataset that is potentially already calculated when we submit?</p>\n<p>Q3. If the submission will be re-run after the deadline, would questions.csv, lectures.csv and the private test dataset remain the same as when we submit for the public LB score before the deadline?</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190791\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190791</a></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1055372,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-10-20T17:50:47.367000",
      "content": "<p>Roughly what percentage of users in the hidden test set will be new users?</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1065394,
      "author_name": "Trigram",
      "author_url": "",
      "post_date": "2020-10-31T08:36:39.717000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> When I try to use TPUs to make a submission it errs with \"Your Notebook cannot use TPUs in this competition.\" But as per the code requirements we can use TPUs, so is this a glitch?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1065401,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-10-31T08:51:48.433000",
          "content": "<p>The Code Requirements have been updated / changed since the start of the competition.<br>\nInitially internet access was not allowed either but that condition isn't mentioned anymore.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nWould be good to get clarification / confirmation on this since there has been no official post about the  updates.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1068689,
          "author_name": "Nya 🚀",
          "author_url": "",
          "post_date": "2020-11-03T16:43:26.677000",
          "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> I just tried to submit a notebook with Internet access turned on, and I got </p>\n<blockquote>\n  <p>Your Notebook cannot use internet access in this competition. Please disable internet in the Notebook editor and save a new version.</p>\n</blockquote>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1069885,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-11-05T03:40:58.783000",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> needs to confirm regarding the mismatch between Code Requirements and actual backend since something changed.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1050278,
      "author_name": "Ethan",
      "author_url": "",
      "post_date": "2020-10-15T08:08:37.387000",
      "content": "<p>Thanks for the clarifications. I have a request that could you please show more details of the \"Submission Scoring Error\" when we make submission. Each submission spends really long time and finnally \"Submission Scoring Error\"  without any addition details. It's better to add some logs of the error, so we will not waste our time in debug and forcus more on the feature engineering and modeling.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1050342,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-15T09:49:25.227000",
          "content": "<p>They cannot add it else people will dump the whole data into logs, you can always raise an error at will.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1049944,
      "author_name": "Elman Mansimov",
      "author_url": "",
      "post_date": "2020-10-14T22:51:46.527000",
      "content": "<p>Thanks for the clarifications. Will the hidden set not contain new lectures as well?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1060721,
          "author_name": "sahilabs",
          "author_url": "",
          "post_date": "2020-10-26T14:00:22.507000",
          "content": "<p>well, it's obvious that if there is no new question that implies there no new lectures.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1060731,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-10-26T14:09:16.593000",
          "content": "<p>questions and lectures are different datasets so it's not completely obvious. There can be new lectures.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1060803,
      "author_name": "ManjotSinghDhillon",
      "author_url": "",
      "post_date": "2020-10-26T14:50:11.390000",
      "content": "<p>Roughly what percentage of users in the hidden test set will be new users?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1060827,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2020-10-26T14:57:50.530000",
          "content": "<ol>\n<li><p>Just like in real-world scenarios, we look at historical (train) data, observe timestamps and frequency of users in order to estimate future (test) data distribution.</p></li>\n<li><p>Since the data is time-series, you could also take 1st 80% data as train and last 20% data as test, and then perform this calculation viz., how many users are \"new\" i.e. present only in the last 20% of the train data.</p></li>\n</ol>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1068603,
      "author_name": "RDizzl3",
      "author_url": "",
      "post_date": "2020-11-03T15:09:23.620000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> and community! I have one last question about the code competition set up. It has been a while since I have participated in code competition but it looks like we can train models offline these days and load them to the kernel environment.</p>\n<blockquote>\n  <p>Please note that for this competition training is not required in Notebooks.</p>\n</blockquote>\n<p>My question is if we have to make it available to the community as well? Does it count as a pre-trained model?</p>\n<blockquote>\n  <p>Freely &amp; publicly available external data is allowed, including pre-trained models</p>\n</blockquote>\n<p>I am sure this has been answered multiple times since the code competitions have been updated but I wasn't sure where to look.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1134308,
          "author_name": "Douglas K.G. Araujo",
          "author_url": "",
          "post_date": "2021-01-01T06:03:08.023000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rdizzl3\" target=\"_blank\">@rdizzl3</a>, did you find an answer for these questions? I am also interested. Thanks and good luck in this competition!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1066177,
      "author_name": "yash",
      "author_url": "",
      "post_date": "2020-11-01T12:22:47.757000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> for the clarification.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061416,
      "author_name": "Mike L.",
      "author_url": "",
      "post_date": "2020-10-27T02:16:20.263000",
      "content": "<p>content_type_id: Please include a lecture in the test data (Example.csv) so that we can verify that lectures are handled correctly before submitting to the hidden data.</p>\n<p>Thanks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1061521,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-27T04:01:20.410000",
          "content": "<p>Mike, you can take a look at <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856\" target=\"_blank\">this</a>!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1061536,
          "author_name": "Mike L.",
          "author_url": "",
          "post_date": "2020-10-27T04:28:47.707000",
          "content": "<p>Thanks. Now have a solution to my problem. Here is how to insert lectures into the Example dataframe test_df. This works for me 😊</p>\n<p>Improvements welcome!</p>\n<p>Borrowing from: <a href=\"https://www.geeksforgeeks.org/insert-row-at-given-position-in-pandas-dataframe/\" target=\"_blank\">https://www.geeksforgeeks.org/insert-row-at-given-position-in-pandas-dataframe/</a></p>\n<h1>Function to insert row in the dataframe</h1>\n<p>def Insert_row(row_number, df, row_value): <br>\n    # Starting value of upper half <br>\n    start_upper = 0</p>\n<pre><code># End value of upper half \nend_upper = row_number \n\n# Start value of lower half \nstart_lower = row_number \n\n# End value of lower half \nend_lower = df.shape[0] \n\n# Create a list of upper_half index \nupper_half = [*range(start_upper, end_upper, 1)] \n\n# Create a list of lower_half index \nlower_half = [*range(start_lower, end_lower, 1)] \n\n# Increment the value of lower half by 1 \nlower_half = [x.__add__(1) for x in lower_half] \n\n# Combine the two lists \nindex_ = upper_half + lower_half \n\n# Update the index of the dataframe \ndf.index = index_ \n\n# Insert a row at the end \ndf.loc[row_number] = row_value \n\n# Sort the index labels \ndf = df.sort_index() \n\n# return the dataframe \nreturn df \n</code></pre>\n<p>iter_test = env.iter_test()<br>\nfor (test_df, sample_prediction_df) in iter_test:</p>\n<pre><code># Let's create a row which we want to insert \n row_number = 2\n row_value = [89,653762,2746,6808,1,14,-1,-1,0,False] \n\n if row_number &gt; test_df.index.max()+1: \n    print(\"Invalid row_number\", row_number, test_df.index.max()) \n else: \n\n# Let's call the function and insert the row \n    test_df = Insert_row(row_number, test_df, row_value) \n\n# Print the updated dataframe \n    print(test_df) \n</code></pre>\n<p>….</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1067188,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-11-02T11:48:52.100000",
          "content": "<p><a href=\"https://www.kaggle.com/mikel1\" target=\"_blank\">@mikel1</a> <br>\n1) could u help understand the need for  insert of new lectures ?</p>\n<p>2) Can test set contain only question ids that are present in questions.csv ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1054875,
      "author_name": "Mike L.",
      "author_url": "",
      "post_date": "2020-10-20T08:47:22.993000",
      "content": "<p>Sohier, in your very instructive \"Competition API Detailed Introduction\", you stop a little too soon.</p>\n<p>Here are the next steps:</p>\n<p>1) Have a Notebook that successfully predicts the Example test data when \"Run All\" is clicked<br>\n    Be sure that Setting \"Internet\" is off.<br>\n2) Save to share.<br>\n3) Share with the public<br>\n4) On the Notebooks page, click on it <br>\n5) Execution info<br>\n6) Click on \"Submit\"<br>\nYou will be told that your Notebook is running.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1049921,
      "author_name": "lucius lu",
      "author_url": "",
      "post_date": "2020-10-14T22:13:05.127000",
      "content": "<p>Thanks for the clarification! The new <strong>question_id</strong> in the hidden test set and the lack of sequence of test data have bothered me for a long time.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1123192,
      "author_name": "Caleb",
      "author_url": "",
      "post_date": "2020-12-23T02:58:12.217000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> apologies if this has been covered elsewhere - I read through the updates and searched the other discussions, but didn't see anything.</p>\n<p>I ran <a href=\"https://www.kaggle.com/calebeverett/test-tid-deltas\" target=\"_blank\">this book</a> to see whether there were many instances of gaps of more than one in the sequence of task_container_ids.  I set it to error out if the count got over 100,000 and it errored out soon after it started. I understand there may be gaps due to the timing of user interactions on multiple devices and as a result task_container_ids are not sequential when records are ordered for each user by timestamp. I calculated that there are 2.5 million of such instances, in the training set,  or approximately 2.5% of all of the observations. If the proportion was similar in the test set that would equate to approximately 62,500 instances, so erroring out after 100,000 would seem to indicate another source of them not being sequential.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1061101,
      "author_name": "RDizzl3",
      "author_url": "",
      "post_date": "2020-10-26T18:35:28.533000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> thank you for the corrections and clarifications. I am about to enter the competition but I feel like I don't quite understand how the test API works for the new test users. I am looking at the <code>example_test.csv</code> and there is a user present in the test data and not in the training data.</p>\n<p>Here is what that user looks like:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F291298%2Feb4ec9b30ff5ceae322d677dba13aff4%2FScreen%20Shot%202020-10-26%20at%2011.24.00%20AM.png?generation=1603736688404605&amp;alt=media\" alt=\"\"></p>\n<p>My first question is making sure I understand the groups correctly - </p>\n<ol>\n<li><p>The lists from group 1 are the correct answers, etc. from questions in group 0?</p></li>\n<li><p>If the list from group 1 is their data from group zero, shouldn't there be different content ids (possibly) for the different questions? There is only one content id in group 0.</p></li>\n</ol>\n<p>Sorry if this has been explained somewhere else. I read quite a few topics in the forum but couldn't figure out how to correctly map and use this info.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1061108,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-10-26T18:45:51.650000",
          "content": "<blockquote>\n  <p>The lists from group 1 are the correct answers, etc. from questions in group 0?</p>\n</blockquote>\n<p>Yes.</p>\n<blockquote>\n  <p>There is only one content id in group 0.</p>\n</blockquote>\n<p>Don't filter on a user (it is irrelevant / incorrect). The length of the answer list in <code>group n</code> will exactly match the number of rows in <code>group (n-1)</code> (considering all users together).</p>\n<p>Highly recommend to go through <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124\" target=\"_blank\">this</a>.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1061115,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-26T18:53:54.480000",
          "content": "<p>Plus there's video's labels in that list as well! So be careful with that as we don't have to make preds on video_rows so to say (content_type_id == 1).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1061122,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2020-10-26T19:01:19.230000",
          "content": "<p>thank you <a href=\"https://www.kaggle.com/vopani\" target=\"_blank\">@vopani</a>! I did not realize that topic had all this info! But I was able to verify everything I was questioning. Appreciate the comment!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1050025,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-10-15T02:14:32.560000",
      "content": "<p>Thanks for the clarifications. Will it be possible to simply return the data from questions and lectures pre-appended to the test-df as well? Because it will be joined by us either ways, so IMO it's a good thing to return it when we are iterating over the test_df as it's columns?</p>\n<p>And the only file updated is lectures.csv, right?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1050004,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2020-10-15T01:17:38.407000",
      "content": "<p>Thanks for the clarification about private set. Now it help and have better clarification about private set.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Do you mean.. <strong>train/test set is complete</strong>  having all user activities logged except <strong>few missing</strong> values are <strong>negligible ratio</strong> ?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1134231,
      "author_name": "S.Kodai",
      "author_url": "",
      "post_date": "2021-01-01T03:14:54.877000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Thanks for the clarification about private set. Now it help and have better clarification about private set.</p>\n<blockquote>\n  <p>The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.</p>\n</blockquote>\n<p>Regarding this matter, does chronologically mean that in order of <strong>timestamp</strong>, rather than task_container_id?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1121782,
      "author_name": "Caleb",
      "author_url": "",
      "post_date": "2020-12-21T22:44:02.397000",
      "content": "<p>I was curious if the task_container_ids continue in sequence from the training data to the test data. It looks like there are gaps between the training data and sample test data.</p>\n<p>Below are the first ten sample test records.</p>\n<table>\n  <thead>\n    <tr>\n      <th></th>\n      <th>user_id</th>\n      <th>row_id</th>\n      <th>task_container_id</th>\n      <th>max_train_tid</th>\n      <th>tid_delta</th>\n      <th>ts_delta_sec</th>\n    </tr>\n  </thead>\n  <tbody>\n    <tr>\n      <th>1</th>\n      <td>554169193</td>\n      <td>1</td>\n      <td>4427</td>\n      <td>4421.0</td>\n      <td>6.0</td>\n      <td>184.627</td>\n    </tr>\n    <tr>\n      <th>2</th>\n      <td>1720860329</td>\n      <td>2</td>\n      <td>240</td>\n      <td>235.0</td>\n      <td>5.0</td>\n      <td>299.297</td>\n    </tr>\n    <tr>\n      <th>3</th>\n      <td>288641214</td>\n      <td>3</td>\n      <td>266</td>\n      <td>262.0</td>\n      <td>4.0</td>\n      <td>72252.433</td>\n    </tr>\n    <tr>\n      <th>4</th>\n      <td>1728340777</td>\n      <td>4</td>\n      <td>162</td>\n      <td>161.0</td>\n      <td>1.0</td>\n      <td>67009.619</td>\n    </tr>\n    <tr>\n      <th>5</th>\n      <td>1364159702</td>\n      <td>5</td>\n      <td>4424</td>\n      <td>4421.0</td>\n      <td>3.0</td>\n      <td>173642.470</td>\n    </tr>\n    <tr>\n      <th>6</th>\n      <td>1521618396</td>\n      <td>6</td>\n      <td>1367</td>\n      <td>1361.0</td>\n      <td>6.0</td>\n      <td>202.134</td>\n    </tr>\n    <tr>\n      <th>7</th>\n      <td>1317245193</td>\n      <td>7</td>\n      <td>5314</td>\n      <td>5307.0</td>\n      <td>7.0</td>\n      <td>384674.701</td>\n    </tr>\n    <tr>\n      <th>8</th>\n      <td>1700555100</td>\n      <td>8</td>\n      <td>532</td>\n      <td>530.0</td>\n      <td>2.0</td>\n      <td>246.422</td>\n    </tr>\n    <tr>\n      <th>9</th>\n      <td>998511398</td>\n      <td>9</td>\n      <td>393</td>\n      <td>389.0</td>\n      <td>4.0</td>\n      <td>37118.438</td>\n    </tr>\n    <tr>\n      <th>10</th>\n      <td>1422853669</td>\n      <td>10</td>\n      <td>85</td>\n      <td>79.0</td>\n      <td>6.0</td>\n      <td>194.973</td>\n    </tr>\n  </tbody>\n</table>\n<p>Are those gaps related to the preparation of the sample test data or are there similar gaps in the actual test data as well?</p>\n<p>EDIT: updated to include timetamp delta as well, which look like they might span gaps in task_container_ids. I had a median gap of approx. 42 seconds on the last set of user training records.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1117061,
      "author_name": "skw1990",
      "author_url": "",
      "post_date": "2020-12-17T17:30:35.587000",
      "content": "<p>Is the part in lectures same as part in questions?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1095031,
      "author_name": "Yueqi Wang 0001",
      "author_url": "",
      "post_date": "2020-11-29T07:45:25.093000",
      "content": "<p>That makes it clearer</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1085024,
      "author_name": "sahilabs",
      "author_url": "",
      "post_date": "2020-11-20T15:59:35.760000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> I had a query that, In the column \"prior_question_had_explanation\"  the student can explain the question at any time or just after the question is answered?<br>\nplease look at this query<br>\n thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1079151,
      "author_name": "Stefan Mandl",
      "author_url": "",
      "post_date": "2020-11-15T17:16:24.223000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> and community, I'm still confused about the tags/tag columns  in <strong>questions.csv</strong> and <strong>lectures.csv</strong>. In this post you write, that these would \"match\" now, but there are 37 tags that occur in questions.csv that have no matching entries in lectures.csv. Am I missing something obvious here?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1079199,
          "author_name": "sahilabs",
          "author_url": "",
          "post_date": "2020-11-15T18:32:48.267000",
          "content": "<p>I think of the same but in the above author discussion it's written as follows, </p>\n<blockquote>\n  <p>a)I'm updating the dataset now so <strong>lecture tags will match the tags in questions.csv</strong>. <br>\n  b)Correction- the hidden test set contains <strong>new users but not new questions</strong></p>\n</blockquote>\n<p>Coming to point 'B' there are no new questions which means there is a chance of introducing new lectures which then the tag of questions.csv and lecture.csv may match completely. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1067278,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-11-02T12:43:17.600000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a><br>\n1)   why have we got multiple lecture tags associated to a single question id ,how do we interpret those</p>\n<p>2) Can a test set have different   lecture id than what is there in   lecture,csv that we are provided with</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1079217,
          "author_name": "sahilabs",
          "author_url": "",
          "post_date": "2020-11-15T18:52:41.153000",
          "content": "<p>for the first point, <br>\nlecture tag and question id are different things actually question_id is a content_id(which is in train.csv when cotent_type_id is zero ). and for the understanding of tag, it just says particular lecture or question belongs to that topic.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1065700,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-10-31T15:49:35.450000",
      "content": "<blockquote>\n  <p>You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save &amp; Run All\"</p>\n</blockquote>\n<p>This looks great!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1068727,
          "author_name": "Pavel Orlov",
          "author_url": "",
          "post_date": "2020-11-03T17:21:07.413000",
          "content": "<p>Please tell me where the quote came from? (not met)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1069053,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-11-04T02:40:58.557000",
          "content": "<p>It comes from the pre-typed section on the kernel. Here's the formal <a href=\"https://www.kaggle.com/product-feedback/195163\" target=\"_blank\">announcement</a> now though…</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1063501,
      "author_name": "kuroyuli",
      "author_url": "",
      "post_date": "2020-10-29T02:10:52.560000",
      "content": "<p>In the 'Time-series API Details' section of <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/data\" target=\"_blank\">Data Description</a>, there is a statement,  </p>\n<blockquote>\n  <p>Expect to see roughly 2.5 million questions in the hidden test set.  </p>\n</blockquote>\n<p>I think the term 'questions' is misleading. Would you change that to 'rows' or 'interactions'?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1065260,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-10-31T05:03:15.990000",
          "content": "<p>Has it been confirmed somewhere that the test set has 2.5M rows (and not questions)?</p>\n<p>I've been assuming there are 2.5M questions and (based on train data %) an additional few thousand lectures.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1067968,
          "author_name": "kuroyuli",
          "author_url": "",
          "post_date": "2020-11-02T23:24:44.640000",
          "content": "<p>Thank you for your reply, Vopani.<br>\nI just thought that '2.5 million rows' would be more appropriate to explain the volume of test set.<br>\nI don't know that someone verified the volume or not.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1068102,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-11-03T04:04:20.383000",
          "content": "<p>Yes I agree using rows or observations is more appropriate.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1062003,
      "author_name": "Md. Abdullah Al Mamun",
      "author_url": "",
      "post_date": "2020-10-27T13:52:39.920000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> for the clarification. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1059534,
      "author_name": "Travis",
      "author_url": "",
      "post_date": "2020-10-25T08:02:56.573000",
      "content": "<p>The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.<br>\n-- Does it mean all samples in test data happen after train data?  And GROUP 0 of test data happened before GROUP 1 of test data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1059585,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-10-25T08:48:33.517000",
          "content": "<p>Yes that is my understanding and I'm assuming the API iterates over the groups chronologically, so group <code>n</code> happens before group <code>n+1</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1060025,
          "author_name": "sahilabs",
          "author_url": "",
          "post_date": "2020-10-25T17:32:15.220000",
          "content": "<p>Yes, the User_ID in train set and test set are same that means test set is continuing train set but one more point to notice that there are some new User_ID in test set.<br>\nAnd the group num proptionals to time series</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1059533,
      "author_name": "Samarpan",
      "author_url": "",
      "post_date": "2020-10-25T08:02:30.847000",
      "content": "<p>Maybe I missed something but the defination of timestamp is vague. Can the host clarify what does \"first event completion from that user.\" exactly mean?   What does event completion imply? What happens when the (mobile) user never completes that event? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1053755,
      "author_name": "EnricRovira",
      "author_url": "",
      "post_date": "2020-10-19T10:27:31.780000",
      "content": "<p>yiiiiihaaa</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1054629,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-20T04:09:04.433000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1053439,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-19T02:25:56.423000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1051706,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-16T18:46:22.690000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1051688,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-16T18:21:16.407000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1050323,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-15T09:22:00.957000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1052395,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-17T17:08:15.867000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1052090,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-17T09:37:15.097000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1051983,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-17T06:27:56.870000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1050747,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-15T17:23:01.023000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1049702": "- I'm updating the dataset now so lecture tags will match the tags in **questions.csv**. The updated **lectures.csv** should be live within the next hour. Apologies for the inconvenience.\n\n- Correction- the hidden test set contains new _users_ but not new _questions_.\n\n- The train/test data is complete, in the sense that there are no missing interactions in the union of train and test data. It remains possible that some questions weren't logged due to other issues that all datasets of mobile users are susceptible to,\nsuch as if a user lost their connection mid-question.\n\n- The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.",
    "1049825": "@sohier Thanks for the clarification! Now we can start to build the sequential model. Would you mind to answer the following questions?\n\n\nQ1. When we submit a notebook for this competition, does it run on the whole private test dataset (and 20% of them are used to calculate the public LB score).\n\nQ2. Will our submissions be re-run after the deadline in order to compute the private LB score? Or it will just use the score for the remaining 80% private test dataset that is potentially already calculated when we submit?\n\nQ3. If the submission will be re-run after the deadline, would questions.csv, lectures.csv and the private test dataset remain the same as when we submit for the public LB score before the deadline?\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190791",
    "1055372": "Roughly what percentage of users in the hidden test set will be new users?",
    "1065394": "@sohier When I try to use TPUs to make a submission it errs with \"Your Notebook cannot use TPUs in this competition.\" But as per the code requirements we can use TPUs, so is this a glitch?",
    "1050278": "Thanks for the clarifications. I have a request that could you please show more details of the \"Submission Scoring Error\" when we make submission. Each submission spends really long time and finnally \"Submission Scoring Error\"  without any addition details. It's better to add some logs of the error, so we will not waste our time in debug and forcus more on the feature engineering and modeling.",
    "1049944": "Thanks for the clarifications. Will the hidden set not contain new lectures as well?\n ",
    "1060803": "Roughly what percentage of users in the hidden test set will be new users?",
    "1068603": "Hey @sohier and community! I have one last question about the code competition set up. It has been a while since I have participated in code competition but it looks like we can train models offline these days and load them to the kernel environment.\n\n> Please note that for this competition training is not required in Notebooks.\n\nMy question is if we have to make it available to the community as well? Does it count as a pre-trained model?\n\n> Freely & publicly available external data is allowed, including pre-trained models\n\nI am sure this has been answered multiple times since the code competitions have been updated but I wasn't sure where to look.",
    "1066177": "Thank you @sohier for the clarification.",
    "1061416": "content_type_id: Please include a lecture in the test data (Example.csv) so that we can verify that lectures are handled correctly before submitting to the hidden data.\n\nThanks.",
    "1054875": "Sohier, in your very instructive \"Competition API Detailed Introduction\", you stop a little too soon.\n\nHere are the next steps:\n\n1) Have a Notebook that successfully predicts the Example test data when \"Run All\" is clicked\n    Be sure that Setting \"Internet\" is off.\n2) Save to share.\n3) Share with the public\n4) On the Notebooks page, click on it \n5) Execution info\n6) Click on \"Submit\"\nYou will be told that your Notebook is running.",
    "1049921": "Thanks for the clarification! The new **question_id** in the hidden test set and the lack of sequence of test data have bothered me for a long time.",
    "1123192": "@sohier apologies if this has been covered elsewhere - I read through the updates and searched the other discussions, but didn't see anything.\n\nI ran [this book](https://www.kaggle.com/calebeverett/test-tid-deltas) to see whether there were many instances of gaps of more than one in the sequence of task_container_ids.  I set it to error out if the count got over 100,000 and it errored out soon after it started. I understand there may be gaps due to the timing of user interactions on multiple devices and as a result task_container_ids are not sequential when records are ordered for each user by timestamp. I calculated that there are 2.5 million of such instances, in the training set,  or approximately 2.5% of all of the observations. If the proportion was similar in the test set that would equate to approximately 62,500 instances, so erroring out after 100,000 would seem to indicate another source of them not being sequential.",
    "1061101": "Hey @sohier thank you for the corrections and clarifications. I am about to enter the competition but I feel like I don't quite understand how the test API works for the new test users. I am looking at the `example_test.csv` and there is a user present in the test data and not in the training data.\n\nHere is what that user looks like:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F291298%2Feb4ec9b30ff5ceae322d677dba13aff4%2FScreen%20Shot%202020-10-26%20at%2011.24.00%20AM.png?generation=1603736688404605&alt=media)\n\nMy first question is making sure I understand the groups correctly - \n1. The lists from group 1 are the correct answers, etc. from questions in group 0?\n\n2. If the list from group 1 is their data from group zero, shouldn't there be different content ids (possibly) for the different questions? There is only one content id in group 0.\n\nSorry if this has been explained somewhere else. I read quite a few topics in the forum but couldn't figure out how to correctly map and use this info.",
    "1050025": "Thanks for the clarifications. Will it be possible to simply return the data from questions and lectures pre-appended to the test-df as well? Because it will be joined by us either ways, so IMO it's a good thing to return it when we are iterating over the test_df as it's columns?\n\nAnd the only file updated is lectures.csv, right?",
    "1050004": "Thanks for the clarification about private set. Now it help and have better clarification about private set.\n\n@sohier Do you mean.. **train/test set is complete**  having all user activities logged except **few missing** values are **negligible ratio** ?",
    "1134231": "@sohier Thanks for the clarification about private set. Now it help and have better clarification about private set.\n\n> The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.\n\nRegarding this matter, does chronologically mean that in order of **timestamp**, rather than task_container_id?",
    "1121782": "I was curious if the task_container_ids continue in sequence from the training data to the test data. It looks like there are gaps between the training data and sample test data.\n\nBelow are the first ten sample test records.\n\n<table border=\"1\" class=\"dataframe\">\n  <thead>\n    <tr style=\"text-align: right;\">\n      <th></th>\n      <th>user_id</th>\n      <th>row_id</th>\n      <th>task_container_id</th>\n      <th>max_train_tid</th>\n      <th>tid_delta</th>\n      <th>ts_delta_sec</th>\n    </tr>\n  </thead>\n  <tbody>\n    <tr>\n      <th>1</th>\n      <td>554169193</td>\n      <td>1</td>\n      <td>4427</td>\n      <td>4421.0</td>\n      <td>6.0</td>\n      <td>184.627</td>\n    </tr>\n    <tr>\n      <th>2</th>\n      <td>1720860329</td>\n      <td>2</td>\n      <td>240</td>\n      <td>235.0</td>\n      <td>5.0</td>\n      <td>299.297</td>\n    </tr>\n    <tr>\n      <th>3</th>\n      <td>288641214</td>\n      <td>3</td>\n      <td>266</td>\n      <td>262.0</td>\n      <td>4.0</td>\n      <td>72252.433</td>\n    </tr>\n    <tr>\n      <th>4</th>\n      <td>1728340777</td>\n      <td>4</td>\n      <td>162</td>\n      <td>161.0</td>\n      <td>1.0</td>\n      <td>67009.619</td>\n    </tr>\n    <tr>\n      <th>5</th>\n      <td>1364159702</td>\n      <td>5</td>\n      <td>4424</td>\n      <td>4421.0</td>\n      <td>3.0</td>\n      <td>173642.470</td>\n    </tr>\n    <tr>\n      <th>6</th>\n      <td>1521618396</td>\n      <td>6</td>\n      <td>1367</td>\n      <td>1361.0</td>\n      <td>6.0</td>\n      <td>202.134</td>\n    </tr>\n    <tr>\n      <th>7</th>\n      <td>1317245193</td>\n      <td>7</td>\n      <td>5314</td>\n      <td>5307.0</td>\n      <td>7.0</td>\n      <td>384674.701</td>\n    </tr>\n    <tr>\n      <th>8</th>\n      <td>1700555100</td>\n      <td>8</td>\n      <td>532</td>\n      <td>530.0</td>\n      <td>2.0</td>\n      <td>246.422</td>\n    </tr>\n    <tr>\n      <th>9</th>\n      <td>998511398</td>\n      <td>9</td>\n      <td>393</td>\n      <td>389.0</td>\n      <td>4.0</td>\n      <td>37118.438</td>\n    </tr>\n    <tr>\n      <th>10</th>\n      <td>1422853669</td>\n      <td>10</td>\n      <td>85</td>\n      <td>79.0</td>\n      <td>6.0</td>\n      <td>194.973</td>\n    </tr>\n  </tbody>\n</table>\n\nAre those gaps related to the preparation of the sample test data or are there similar gaps in the actual test data as well?\n\nEDIT: updated to include timetamp delta as well, which look like they might span gaps in task_container_ids. I had a median gap of approx. 42 seconds on the last set of user training records.",
    "1117061": "Is the part in lectures same as part in questions?\n",
    "1095031": "That makes it clearer",
    "1085024": "@sohier I had a query that, In the column \"prior_question_had_explanation\"  the student can explain the question at any time or just after the question is answered?\nplease look at this query\n thanks",
    "1079151": "Hi @sohier and community, I'm still confused about the tags/tag columns  in **questions.csv** and **lectures.csv**. In this post you write, that these would \"match\" now, but there are 37 tags that occur in questions.csv that have no matching entries in lectures.csv. Am I missing something obvious here?",
    "1067278": "@sohier\n1)   why have we got multiple lecture tags associated to a single question id ,how do we interpret those\n\n2) Can a test set have different   lecture id than what is there in   lecture,csv that we are provided with\n",
    "1065700": "> You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\"\n\nThis looks great!",
    "1063501": "In the 'Time-series API Details' section of [Data Description](https://www.kaggle.com/c/riiid-test-answer-prediction/data), there is a statement,  \n> Expect to see roughly 2.5 million questions in the hidden test set.  \n\nI think the term 'questions' is misleading. Would you change that to 'rows' or 'interactions'?",
    "1062003": "Thank you @sohier for the clarification. ",
    "1059534": "The test data follows chronologically after the train data. The test iterations give interactions of users chronologically.\n-- Does it mean all samples in test data happen after train data?  And GROUP 0 of test data happened before GROUP 1 of test data?",
    "1059533": "Maybe I missed something but the defination of timestamp is vague. Can the host clarify what does \"first event completion from that user.\" exactly mean?   What does event completion imply? What happens when the (mobile) user never completes that event? ",
    "1053755": "yiiiiihaaa",
    "1054629": "",
    "1053439": "",
    "1051706": "",
    "1051688": "",
    "1050323": "",
    "1052395": "Thank you! I really appreciate it. ",
    "1052090": "Thanks for the clarifications.",
    "1051983": "Thank for clearance",
    "1050747": "thanks for clearance"
  }
}