{
  "id": 199266,
  "title": "Doing a merge of lecture and train csv ? Finding Same content id for lectures as well as questions",
  "url": "/competitions/riiid-test-answer-prediction/discussion/199266",
  "author_name": "",
  "post_date": "2020-11-25T05:44:28.188065Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello all , <br>\nI am not able to understand the lectures data and train data <br>\nI filtered the train data by user = 2147482216 and merged it with lectures data. </p>\n<p>And i am getting content_id 641 and lecture_id 641 and in the content_type_id is False<br>\nI think it is not the right way to merge lectures and questions ? Any solution to this<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2F467d844b41d5273112bf4211708ee2ac%2FScreenshot_20201125_112202.png?generation=1606282776064072&amp;alt=media\" alt=\"\"> ..</p>\n<p>I am trying to utilize lectures by shifting them to the next timestamp and trying to calculate the time intervals between lectures . i tried first by creating different dataframe of lectures and train csv where content_type_id is true and later on merging this with original dataframe but while merging i get memory error ..</p>\n<p>So i am trying to merge first on train and do the calulations in place ?</p>\n<p>Thank you</p>",
  "messages": [
    {
      "id": "1090141",
      "postDate": "11/25/2020 05:44:28",
      "content": "<p>Hello all , <br>\nI am not able to understand the lectures data and train data <br>\nI filtered the train data by user = 2147482216 and merged it with lectures data. </p>\n<p>And i am getting content_id 641 and lecture_id 641 and in the content_type_id is False<br>\nI think it is not the right way to merge lectures and questions ? Any solution to this<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2F467d844b41d5273112bf4211708ee2ac%2FScreenshot_20201125_112202.png?generation=1606282776064072&amp;alt=media\" alt=\"\"> ..</p>\n<p>I am trying to utilize lectures by shifting them to the next timestamp and trying to calculate the time intervals between lectures . i tried first by creating different dataframe of lectures and train csv where content_type_id is true and later on merging this with original dataframe but while merging i get memory error ..</p>\n<p>So i am trying to merge first on train and do the calulations in place ?</p>\n<p>Thank you</p>",
      "rawMarkdown": "Hello all , \nI am not able to understand the lectures data and train data \nI filtered the train data by user = 2147482216 and merged it with lectures data. \n\nAnd i am getting content_id 641 and lecture_id 641 and in the content_type_id is False\nI think it is not the right way to merge lectures and questions ? Any solution to this![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2F467d844b41d5273112bf4211708ee2ac%2FScreenshot_20201125_112202.png?generation=1606282776064072&alt=media) ..\n\nI am trying to utilize lectures by shifting them to the next timestamp and trying to calculate the time intervals between lectures . i tried first by creating different dataframe of lectures and train csv where content_type_id is true and later on merging this with original dataframe but while merging i get memory error ..\n\nSo i am trying to merge first on train and do the calulations in place ?\n\nThank you",
      "votes": null
    },
    {
      "id": "1090157",
      "postDate": "11/25/2020 06:01:42",
      "content": "<p>After getting wrong merge do i do this <br>\n<code>df_1.loc[df_1['content_type_id'] == False, ['lecture_id', 'tag', 'part', 'type_of']] = np.nan\n</code><br>\nand then shift the ['tag', 'part', 'type_of', 'timestamp'] to flag the user who was watching the lecture  ?</p>\n<p>Does this method give correct data on lectures and train or is there something else that i am missing </p>",
      "rawMarkdown": "After getting wrong merge do i do this \n`df_1.loc[df_1['content_type_id'] == False, ['lecture_id', 'tag', 'part', 'type_of']] = np.nan\n`\nand then shift the ['tag', 'part', 'type_of', 'timestamp'] to flag the user who was watching the lecture  ?\n\nDoes this method give correct data on lectures and train or is there something else that i am missing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1090157,
      "author_name": "ptrikp",
      "author_url": "",
      "post_date": "11/25/2020 06:01:42",
      "content": "<p>After getting wrong merge do i do this <br>\n<code>df_1.loc[df_1['content_type_id'] == False, ['lecture_id', 'tag', 'part', 'type_of']] = np.nan\n</code><br>\nand then shift the ['tag', 'part', 'type_of', 'timestamp'] to flag the user who was watching the lecture  ?</p>\n<p>Does this method give correct data on lectures and train or is there something else that i am missing </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1090141": "Hello all , \nI am not able to understand the lectures data and train data \nI filtered the train data by user = 2147482216 and merged it with lectures data. \n\nAnd i am getting content_id 641 and lecture_id 641 and in the content_type_id is False\nI think it is not the right way to merge lectures and questions ? Any solution to this![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2F467d844b41d5273112bf4211708ee2ac%2FScreenshot_20201125_112202.png?generation=1606282776064072&alt=media) ..\n\nI am trying to utilize lectures by shifting them to the next timestamp and trying to calculate the time intervals between lectures . i tried first by creating different dataframe of lectures and train csv where content_type_id is true and later on merging this with original dataframe but while merging i get memory error ..\n\nSo i am trying to merge first on train and do the calulations in place ?\n\nThank you",
    "1090157": "After getting wrong merge do i do this \n`df_1.loc[df_1['content_type_id'] == False, ['lecture_id', 'tag', 'part', 'type_of']] = np.nan\n`\nand then shift the ['tag', 'part', 'type_of', 'timestamp'] to flag the user who was watching the lecture  ?\n\nDoes this method give correct data on lectures and train or is there something else that i am missing"
  },
  "source": "meta"
}