{
  "id": 202376,
  "title": "Merge : Costly ",
  "url": "/competitions/riiid-test-answer-prediction/discussion/202376",
  "author_name": "",
  "post_date": "2020-12-09T17:58:51.719903100Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I need to perform merge between train and question to build features but it eats up too much memory that it dsnt leaves any thing for rest of stuff. </p>\n<p>Please suggest an alternative to this . </p>\n<p>I read the discussions which talks about only inference time replacement. </p>",
  "messages": [
    {
      "id": "1107467",
      "postDate": "12/09/2020 17:58:51",
      "content": "<p>I need to perform merge between train and question to build features but it eats up too much memory that it dsnt leaves any thing for rest of stuff. </p>\n<p>Please suggest an alternative to this . </p>\n<p>I read the discussions which talks about only inference time replacement. </p>",
      "rawMarkdown": "I need to perform merge between train and question to build features but it eats up too much memory that it dsnt leaves any thing for rest of stuff. \n\nPlease suggest an alternative to this . \n\nI read the discussions which talks about only inference time replacement.",
      "votes": null
    },
    {
      "id": "1107559",
      "postDate": "12/09/2020 19:18:09",
      "content": "<p>Do look at the datatypes created for the tags - that is usually the cause of the problem.</p>",
      "rawMarkdown": "Do look at the datatypes created for the tags - that is usually the cause of the problem.",
      "votes": null
    },
    {
      "id": "1108386",
      "postDate": "12/10/2020 15:38:34",
      "content": "<p>One way to solve the memory issue is using the following code.</p>\n<p>First convert the questions data into a dictionary</p>\n<pre><code>from collections import defaultdict\n\nquestions_df.set_index(\"question_id\", inplace=True)\nquestions_dict = {}\nfor question_id, row in questions_df.iterrows():\n    questions_dict[question_id] = row.values()\n</code></pre>\n<p>Now loop through the training data and add the question information into a NumPy array</p>\n<pre><code>question_features = np.zeros((len(train), questions_df.shape[1]))\nfor index, question_id in df['content_id'].iteritems():\n    question_features[index, :] = questions_dict[question_id].values()\n</code></pre>\n<p>Now add the features to the actual train data</p>\n<pre><code>for index, col in enumerate(questions_df.columns):\n    train[col] = question_features[:, index]\n</code></pre>\n<p>I hope this helps.</p>",
      "rawMarkdown": "One way to solve the memory issue is using the following code.\n\nFirst convert the questions data into a dictionary\n```python\nfrom collections import defaultdict\n\nquestions_df.set_index(\"question_id\", inplace=True)\nquestions_dict = {}\nfor question_id, row in questions_df.iterrows():\n    questions_dict[question_id] = row.values()\n```\n\nNow loop through the training data and add the question information into a NumPy array\n```python\nquestion_features = np.zeros((len(train), questions_df.shape[1]))\nfor index, question_id in df['content_id'].iteritems():\n    question_features[index, :] = questions_dict[question_id].values()\n```\n\nNow add the features to the actual train data\n```python\nfor index, col in enumerate(questions_df.columns):\n    train[col] = question_features[:, index]\n```\n\nI hope this helps.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1107559,
      "author_name": "watzisname",
      "author_url": "",
      "post_date": "12/09/2020 19:18:09",
      "content": "<p>Do look at the datatypes created for the tags - that is usually the cause of the problem.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1108386,
      "author_name": "manikanthr5",
      "author_url": "",
      "post_date": "12/10/2020 15:38:34",
      "content": "<p>One way to solve the memory issue is using the following code.</p>\n<p>First convert the questions data into a dictionary</p>\n<pre><code>from collections import defaultdict\n\nquestions_df.set_index(\"question_id\", inplace=True)\nquestions_dict = {}\nfor question_id, row in questions_df.iterrows():\n    questions_dict[question_id] = row.values()\n</code></pre>\n<p>Now loop through the training data and add the question information into a NumPy array</p>\n<pre><code>question_features = np.zeros((len(train), questions_df.shape[1]))\nfor index, question_id in df['content_id'].iteritems():\n    question_features[index, :] = questions_dict[question_id].values()\n</code></pre>\n<p>Now add the features to the actual train data</p>\n<pre><code>for index, col in enumerate(questions_df.columns):\n    train[col] = question_features[:, index]\n</code></pre>\n<p>I hope this helps.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1107467": "I need to perform merge between train and question to build features but it eats up too much memory that it dsnt leaves any thing for rest of stuff. \n\nPlease suggest an alternative to this . \n\nI read the discussions which talks about only inference time replacement.",
    "1107559": "Do look at the datatypes created for the tags - that is usually the cause of the problem.",
    "1108386": "One way to solve the memory issue is using the following code.\n\nFirst convert the questions data into a dictionary\n```python\nfrom collections import defaultdict\n\nquestions_df.set_index(\"question_id\", inplace=True)\nquestions_dict = {}\nfor question_id, row in questions_df.iterrows():\n    questions_dict[question_id] = row.values()\n```\n\nNow loop through the training data and add the question information into a NumPy array\n```python\nquestion_features = np.zeros((len(train), questions_df.shape[1]))\nfor index, question_id in df['content_id'].iteritems():\n    question_features[index, :] = questions_dict[question_id].values()\n```\n\nNow add the features to the actual train data\n```python\nfor index, col in enumerate(questions_df.columns):\n    train[col] = question_features[:, index]\n```\n\nI hope this helps."
  },
  "source": "meta"
}