{
  "id": 206037,
  "title": "Saving all users info before starting predictions",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206037",
  "author_name": "",
  "post_date": "2020-12-22T22:03:21.132764400Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hey all !</p>\n<p>Since user-based feature are important, and since some of the ids from test set are already available in the train set, I was wondering if you were using the history of each user to calculate their parameters.</p>\n<p>I am currently struggling with it because of OOM, so I am trying to save each id parameters in a different file, but it takes time to load and I am not sure a single folder can handle the 300k files…</p>\n<p>I was wondering what you guys were doing about it ?</p>",
  "messages": [
    {
      "id": "1123074",
      "postDate": "12/22/2020 22:03:21",
      "content": "<p>Hey all !</p>\n<p>Since user-based feature are important, and since some of the ids from test set are already available in the train set, I was wondering if you were using the history of each user to calculate their parameters.</p>\n<p>I am currently struggling with it because of OOM, so I am trying to save each id parameters in a different file, but it takes time to load and I am not sure a single folder can handle the 300k files…</p>\n<p>I was wondering what you guys were doing about it ?</p>",
      "rawMarkdown": "Hey all !\n\nSince user-based feature are important, and since some of the ids from test set are already available in the train set, I was wondering if you were using the history of each user to calculate their parameters.\n\nI am currently struggling with it because of OOM, so I am trying to save each id parameters in a different file, but it takes time to load and I am not sure a single folder can handle the 300k files...\n\nI was wondering what you guys were doing about it ?",
      "votes": null
    },
    {
      "id": "1123116",
      "postDate": "12/22/2020 23:34:37",
      "content": "<p>Hi, I think you should use python dictionaries (key: user_id, value: user_based feature value)</p>",
      "rawMarkdown": "Hi, I think you should use python dictionaries (key: user_id, value: user_based feature value)",
      "votes": null
    },
    {
      "id": "1123121",
      "postDate": "12/22/2020 23:44:51",
      "content": "<p>Hi, thanks for the answer.</p>\n<p>This is already what I am doing, but my dictionnaries tend to be very big as I have about 20 parameters that I follow for each of the 300000 ids…</p>",
      "rawMarkdown": "Hi, thanks for the answer.\n\nThis is already what I am doing, but my dictionnaries tend to be very big as I have about 20 parameters that I follow for each of the 300000 ids...",
      "votes": null
    },
    {
      "id": "1123126",
      "postDate": "12/22/2020 23:58:45",
      "content": "<p>I have been suffering from <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/202965\" target=\"_blank\">similar problem</a>.<br>\n<a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> gave us useful comment to use sqlite3. <br>\nI have not tested it, but it must be promising.</p>",
      "rawMarkdown": "I have been suffering from [similar problem](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/202965).\n@higepon gave us useful comment to use sqlite3. \nI have not tested it, but it must be promising.",
      "votes": null
    },
    {
      "id": "1123128",
      "postDate": "12/23/2020 00:04:44",
      "content": "<p>You can only store the information of users that show up in test data. Since the train dataset is sorted. My method is:<br>\n<code>dict_uid_idxes = pd.concat([\n    train[['user_id']].drop_duplicates(keep='first'),\n    train[['user_id']].drop_duplicates(keep='last'),\n]).reset_index().sort_values(['user_id', 'index'])\\\n    .groupby(['user_id'])['index'].agg([list]).to_dict()['list']</code><br>\nThen every time you see a new user in test dataset, you query from the train dataset for that user using the above index dictionary and store the information(features, state etc.) for that user and you only need to do this once.</p>",
      "rawMarkdown": "You can only store the information of users that show up in test data. Since the train dataset is sorted. My method is:\n`dict_uid_idxes = pd.concat([\n    train[['user_id']].drop_duplicates(keep='first'),\n    train[['user_id']].drop_duplicates(keep='last'),\n]).reset_index().sort_values(['user_id', 'index'])\\\n    .groupby(['user_id'])['index'].agg([list]).to_dict()['list']`\nThen every time you see a new user in test dataset, you query from the train dataset for that user using the above index dictionary and store the information(features, state etc.) for that user and you only need to do this once.",
      "votes": null
    },
    {
      "id": "1123130",
      "postDate": "12/23/2020 00:08:32",
      "content": "<p>Thanks, i'll have a look! For now I managed to save all 300k files via google colab. I might dig sqlite also! </p>",
      "rawMarkdown": "Thanks, i'll have a look! For now I managed to save all 300k files via google colab. I might dig sqlite also!",
      "votes": null
    },
    {
      "id": "1123135",
      "postDate": "12/23/2020 00:16:24",
      "content": "<p>You manage to use this method on the full dataset within the ram limitation?</p>\n<p>Are you saving a lot of features? </p>",
      "rawMarkdown": "You manage to use this method on the full dataset within the ram limitation?\n\nAre you saving a lot of features?",
      "votes": null
    },
    {
      "id": "1123138",
      "postDate": "12/23/2020 00:20:58",
      "content": "<p>Yes, for my LGB model, I use 100+ features which means a lot of states for each user to be stored. Although we have 37k+ users in train dataset, but not all of them show up in test dataset, so we don't need to store the info for all 37k+ users.</p>",
      "rawMarkdown": "Yes, for my LGB model, I use 100+ features which means a lot of states for each user to be stored. Although we have 37k+ users in train dataset, but not all of them show up in test dataset, so we don't need to store the info for all 37k+ users.",
      "votes": null
    },
    {
      "id": "1123461",
      "postDate": "12/23/2020 08:36:23",
      "content": "<p>I can put all the features into the dictionary except for the number of attempts per question by the user,300000 user_id</p>",
      "rawMarkdown": "I can put all the features into the dictionary except for the number of attempts per question by the user,300000 user_id",
      "votes": null
    },
    {
      "id": "1123516",
      "postDate": "12/23/2020 09:39:37",
      "content": "<p>Thanks for the answer ! <br>\nI guess my issue is that I store all the history of questions and lectures seen for each user and it is taking a lot of place…</p>",
      "rawMarkdown": "Thanks for the answer ! \nI guess my issue is that I store all the history of questions and lectures seen for each user and it is taking a lot of place...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1123116,
      "author_name": "gaetanlopez",
      "author_url": "",
      "post_date": "12/22/2020 23:34:37",
      "content": "<p>Hi, I think you should use python dictionaries (key: user_id, value: user_based feature value)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1123121,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/22/2020 23:44:51",
          "content": "<p>Hi, thanks for the answer.</p>\n<p>This is already what I am doing, but my dictionnaries tend to be very big as I have about 20 parameters that I follow for each of the 300000 ids…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1123461,
          "author_name": "yangxiaoshuai",
          "author_url": "",
          "post_date": "12/23/2020 08:36:23",
          "content": "<p>I can put all the features into the dictionary except for the number of attempts per question by the user,300000 user_id</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1123126,
      "author_name": "tomooinubushi",
      "author_url": "",
      "post_date": "12/22/2020 23:58:45",
      "content": "<p>I have been suffering from <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/202965\" target=\"_blank\">similar problem</a>.<br>\n<a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a> gave us useful comment to use sqlite3. <br>\nI have not tested it, but it must be promising.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1123130,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/23/2020 00:08:32",
          "content": "<p>Thanks, i'll have a look! For now I managed to save all 300k files via google colab. I might dig sqlite also! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1123128,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "12/23/2020 00:04:44",
      "content": "<p>You can only store the information of users that show up in test data. Since the train dataset is sorted. My method is:<br>\n<code>dict_uid_idxes = pd.concat([\n    train[['user_id']].drop_duplicates(keep='first'),\n    train[['user_id']].drop_duplicates(keep='last'),\n]).reset_index().sort_values(['user_id', 'index'])\\\n    .groupby(['user_id'])['index'].agg([list]).to_dict()['list']</code><br>\nThen every time you see a new user in test dataset, you query from the train dataset for that user using the above index dictionary and store the information(features, state etc.) for that user and you only need to do this once.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1123135,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/23/2020 00:16:24",
          "content": "<p>You manage to use this method on the full dataset within the ram limitation?</p>\n<p>Are you saving a lot of features? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1123138,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "12/23/2020 00:20:58",
          "content": "<p>Yes, for my LGB model, I use 100+ features which means a lot of states for each user to be stored. Although we have 37k+ users in train dataset, but not all of them show up in test dataset, so we don't need to store the info for all 37k+ users.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1123516,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/23/2020 09:39:37",
          "content": "<p>Thanks for the answer ! <br>\nI guess my issue is that I store all the history of questions and lectures seen for each user and it is taking a lot of place…</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1123074": "Hey all !\n\nSince user-based feature are important, and since some of the ids from test set are already available in the train set, I was wondering if you were using the history of each user to calculate their parameters.\n\nI am currently struggling with it because of OOM, so I am trying to save each id parameters in a different file, but it takes time to load and I am not sure a single folder can handle the 300k files...\n\nI was wondering what you guys were doing about it ?",
    "1123116": "Hi, I think you should use python dictionaries (key: user_id, value: user_based feature value)",
    "1123121": "Hi, thanks for the answer.\n\nThis is already what I am doing, but my dictionnaries tend to be very big as I have about 20 parameters that I follow for each of the 300000 ids...",
    "1123126": "I have been suffering from [similar problem](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/202965).\n@higepon gave us useful comment to use sqlite3. \nI have not tested it, but it must be promising.",
    "1123128": "You can only store the information of users that show up in test data. Since the train dataset is sorted. My method is:\n`dict_uid_idxes = pd.concat([\n    train[['user_id']].drop_duplicates(keep='first'),\n    train[['user_id']].drop_duplicates(keep='last'),\n]).reset_index().sort_values(['user_id', 'index'])\\\n    .groupby(['user_id'])['index'].agg([list]).to_dict()['list']`\nThen every time you see a new user in test dataset, you query from the train dataset for that user using the above index dictionary and store the information(features, state etc.) for that user and you only need to do this once.",
    "1123130": "Thanks, i'll have a look! For now I managed to save all 300k files via google colab. I might dig sqlite also!",
    "1123135": "You manage to use this method on the full dataset within the ram limitation?\n\nAre you saving a lot of features?",
    "1123138": "Yes, for my LGB model, I use 100+ features which means a lot of states for each user to be stored. Although we have 37k+ users in train dataset, but not all of them show up in test dataset, so we don't need to store the info for all 37k+ users.",
    "1123461": "I can put all the features into the dictionary except for the number of attempts per question by the user,300000 user_id",
    "1123516": "Thanks for the answer ! \nI guess my issue is that I store all the history of questions and lectures seen for each user and it is taking a lot of place..."
  },
  "source": "meta"
}