{
  "id": 202271,
  "title": "Converting data frame to numpy array ",
  "url": "/competitions/riiid-test-answer-prediction/discussion/202271",
  "author_name": "",
  "post_date": "2020-12-09T07:37:56.092393500Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello all, </p>\n<p>I have made some features for almost 40 million rows that i want to train using xgboost.<br>\nAnd i have 15 columns amounting to 2.7 gb pandas dataframe.</p>\n<p>When i convert it to numpy using to_numpy my memory crashes. By the way my memory limit is 24 gb in colab.<br>\nAny way to convert it efficiently into array ?</p>\n<p>Also memory doesnot crash when i do to_records but i cannot feed it into model as it complains the dtype cannot be casted to float rule unsafe</p>\n<p>Any way to remove the dtype from the records array and feed it into the model ?</p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": "1106879",
      "postDate": "12/09/2020 07:37:56",
      "content": "<p>Hello all, </p>\n<p>I have made some features for almost 40 million rows that i want to train using xgboost.<br>\nAnd i have 15 columns amounting to 2.7 gb pandas dataframe.</p>\n<p>When i convert it to numpy using to_numpy my memory crashes. By the way my memory limit is 24 gb in colab.<br>\nAny way to convert it efficiently into array ?</p>\n<p>Also memory doesnot crash when i do to_records but i cannot feed it into model as it complains the dtype cannot be casted to float rule unsafe</p>\n<p>Any way to remove the dtype from the records array and feed it into the model ?</p>\n<p>Thanks</p>",
      "rawMarkdown": "Hello all, \n\nI have made some features for almost 40 million rows that i want to train using xgboost.\nAnd i have 15 columns amounting to 2.7 gb pandas dataframe.\n\nWhen i convert it to numpy using to_numpy my memory crashes. By the way my memory limit is 24 gb in colab.\nAny way to convert it efficiently into array ?\n\nAlso memory doesnot crash when i do to_records but i cannot feed it into model as it complains the dtype cannot be casted to float rule unsafe\n\nAny way to remove the dtype from the records array and feed it into the model ?\n\nThanks",
      "votes": null
    },
    {
      "id": "1108408",
      "postDate": "12/10/2020 16:10:14",
      "content": "<p>One way to convert it is to first initialize the NumPy array and then loop through the dataframe to remove the column and save it to the NumPy array. </p>\n<pre><code>train_np = np.zeros((train.shape[0], train.shape[1]))\nfor index, col in enumerate(train.columns):\n    train_np[:, index] = train.pop(col)\n</code></pre>\n<p>I hope this helps</p>",
      "rawMarkdown": "One way to convert it is to first initialize the NumPy array and then loop through the dataframe to remove the column and save it to the NumPy array. \n\n```python\ntrain_np = np.zeros((train.shape[0], train.shape[1]))\nfor index, col in enumerate(train.columns):\n    train_np[:, index] = train.pop(col)\n```\n\nI hope this helps",
      "votes": null
    },
    {
      "id": "1108424",
      "postDate": "12/10/2020 16:22:09",
      "content": "<p>This takes lot of memory to process almost 40 million rows … <br>\npandas to_numpy, to_records, values is slow than this ?<br>\nI've not been able to do tests on this</p>\n<p>After two days of working on this i figured out a new solution with rapids<br>\nNow i can train almost 70 million rows</p>",
      "rawMarkdown": "This takes lot of memory to process almost 40 million rows ... \npandas to_numpy, to_records, values is slow than this ?\nI've not been able to do tests on this\n\nAfter two days of working on this i figured out a new solution with rapids\nNow i can train almost 70 million rows",
      "votes": null
    },
    {
      "id": "1108426",
      "postDate": "12/10/2020 16:23:22",
      "content": "<p>but kaggle will not support it , i am doing on colab 24gb</p>",
      "rawMarkdown": "but kaggle will not support it , i am doing on colab 24gb",
      "votes": null
    },
    {
      "id": "1108432",
      "postDate": "12/10/2020 16:34:36",
      "content": "<p>It's working fine for me. </p>",
      "rawMarkdown": "It's working fine for me.",
      "votes": null
    },
    {
      "id": "1108436",
      "postDate": "12/10/2020 16:41:24",
      "content": "<p>is to_numpy or values slow according to you ?</p>",
      "rawMarkdown": "is to_numpy or values slow according to you ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1108408,
      "author_name": "manikanthr5",
      "author_url": "",
      "post_date": "12/10/2020 16:10:14",
      "content": "<p>One way to convert it is to first initialize the NumPy array and then loop through the dataframe to remove the column and save it to the NumPy array. </p>\n<pre><code>train_np = np.zeros((train.shape[0], train.shape[1]))\nfor index, col in enumerate(train.columns):\n    train_np[:, index] = train.pop(col)\n</code></pre>\n<p>I hope this helps</p>",
      "votes": null,
      "replies": [
        {
          "id": 1108424,
          "author_name": "ptrikp",
          "author_url": "",
          "post_date": "12/10/2020 16:22:09",
          "content": "<p>This takes lot of memory to process almost 40 million rows … <br>\npandas to_numpy, to_records, values is slow than this ?<br>\nI've not been able to do tests on this</p>\n<p>After two days of working on this i figured out a new solution with rapids<br>\nNow i can train almost 70 million rows</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108426,
          "author_name": "ptrikp",
          "author_url": "",
          "post_date": "12/10/2020 16:23:22",
          "content": "<p>but kaggle will not support it , i am doing on colab 24gb</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108432,
          "author_name": "manikanthr5",
          "author_url": "",
          "post_date": "12/10/2020 16:34:36",
          "content": "<p>It's working fine for me. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108436,
          "author_name": "ptrikp",
          "author_url": "",
          "post_date": "12/10/2020 16:41:24",
          "content": "<p>is to_numpy or values slow according to you ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1106879": "Hello all, \n\nI have made some features for almost 40 million rows that i want to train using xgboost.\nAnd i have 15 columns amounting to 2.7 gb pandas dataframe.\n\nWhen i convert it to numpy using to_numpy my memory crashes. By the way my memory limit is 24 gb in colab.\nAny way to convert it efficiently into array ?\n\nAlso memory doesnot crash when i do to_records but i cannot feed it into model as it complains the dtype cannot be casted to float rule unsafe\n\nAny way to remove the dtype from the records array and feed it into the model ?\n\nThanks",
    "1108408": "One way to convert it is to first initialize the NumPy array and then loop through the dataframe to remove the column and save it to the NumPy array. \n\n```python\ntrain_np = np.zeros((train.shape[0], train.shape[1]))\nfor index, col in enumerate(train.columns):\n    train_np[:, index] = train.pop(col)\n```\n\nI hope this helps",
    "1108424": "This takes lot of memory to process almost 40 million rows ... \npandas to_numpy, to_records, values is slow than this ?\nI've not been able to do tests on this\n\nAfter two days of working on this i figured out a new solution with rapids\nNow i can train almost 70 million rows",
    "1108426": "but kaggle will not support it , i am doing on colab 24gb",
    "1108432": "It's working fine for me.",
    "1108436": "is to_numpy or values slow according to you ?"
  },
  "source": "meta"
}