{
  "id": 205043,
  "title": "Suggest alternative to Merge",
  "url": "/competitions/riiid-test-answer-prediction/discussion/205043",
  "author_name": "",
  "post_date": "2020-12-18T08:01:00.180025900Z",
  "votes": -2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Pandas Merge is very memory intensive. any better way to join the dfs for more features .</p>",
  "messages": [
    {
      "id": "1117570",
      "postDate": "12/18/2020 08:01:00",
      "content": "<p>Pandas Merge is very memory intensive. any better way to join the dfs for more features .</p>",
      "rawMarkdown": "Pandas Merge is very memory intensive. any better way to join the dfs for more features .",
      "votes": null
    },
    {
      "id": "1117584",
      "postDate": "12/18/2020 08:24:23",
      "content": "<p>If you only want to add a single feature then you could try pd.Series.map.</p>",
      "rawMarkdown": "If you only want to add a single feature then you could try pd.Series.map.",
      "votes": null
    },
    {
      "id": "1117591",
      "postDate": "12/18/2020 08:33:57",
      "content": "<p>The best way for this competition is to use a for loop and iterate through the full csv only once.</p>\n<p>I personnaly use 3 functions that i created:</p>\n<ol>\n<li>Create_cache that generate a new cache to store information of a new user when i see one</li>\n<li>Create feature that uses the cache to create my features for sample i</li>\n<li>Update cache that i use after creating the features to avoid leaks and that compute different metrics of a particular user and store them in memory</li>\n</ol>",
      "rawMarkdown": "The best way for this competition is to use a for loop and iterate through the full csv only once.\n\nI personnaly use 3 functions that i created:\n1. Create_cache that generate a new cache to store information of a new user when i see one\n2. Create feature that uses the cache to create my features for sample i\n3. Update cache that i use after creating the features to avoid leaks and that compute different metrics of a particular user and store them in memory",
      "votes": null
    },
    {
      "id": "1118973",
      "postDate": "12/19/2020 15:39:52",
      "content": "<p>you can convert the dfs into numpy arrays and can save memory. One thing more, coloumns takes int64 for the integer data so you can convert into int32 to reduce the size in dfs as well. </p>",
      "rawMarkdown": "you can convert the dfs into numpy arrays and can save memory. One thing more, coloumns takes int64 for the integer data so you can convert into int32 to reduce the size in dfs as well.",
      "votes": null
    },
    {
      "id": "1118974",
      "postDate": "12/19/2020 15:40:43",
      "content": "<p>good suggestion as well !</p>",
      "rawMarkdown": "good suggestion as well !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1117584,
      "author_name": "nicohrubec",
      "author_url": "",
      "post_date": "12/18/2020 08:24:23",
      "content": "<p>If you only want to add a single feature then you could try pd.Series.map.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1117591,
      "author_name": "bowaka",
      "author_url": "",
      "post_date": "12/18/2020 08:33:57",
      "content": "<p>The best way for this competition is to use a for loop and iterate through the full csv only once.</p>\n<p>I personnaly use 3 functions that i created:</p>\n<ol>\n<li>Create_cache that generate a new cache to store information of a new user when i see one</li>\n<li>Create feature that uses the cache to create my features for sample i</li>\n<li>Update cache that i use after creating the features to avoid leaks and that compute different metrics of a particular user and store them in memory</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1118974,
          "author_name": "",
          "author_url": "",
          "post_date": "12/19/2020 15:40:43",
          "content": "<p>good suggestion as well !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1118973,
      "author_name": "",
      "author_url": "",
      "post_date": "12/19/2020 15:39:52",
      "content": "<p>you can convert the dfs into numpy arrays and can save memory. One thing more, coloumns takes int64 for the integer data so you can convert into int32 to reduce the size in dfs as well. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1117570": "Pandas Merge is very memory intensive. any better way to join the dfs for more features .",
    "1117584": "If you only want to add a single feature then you could try pd.Series.map.",
    "1117591": "The best way for this competition is to use a for loop and iterate through the full csv only once.\n\nI personnaly use 3 functions that i created:\n1. Create_cache that generate a new cache to store information of a new user when i see one\n2. Create feature that uses the cache to create my features for sample i\n3. Update cache that i use after creating the features to avoid leaks and that compute different metrics of a particular user and store them in memory",
    "1118973": "you can convert the dfs into numpy arrays and can save memory. One thing more, coloumns takes int64 for the integer data so you can convert into int32 to reduce the size in dfs as well.",
    "1118974": "good suggestion as well !"
  },
  "source": "meta"
}