{
  "id": 157481,
  "title": "Your notebook tried to allocate more memory than is available. It has restarted.",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/157481",
  "author_name": "satish chavan",
  "post_date": "2020-06-10T19:03:42.257000",
  "votes": 0,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I was getting started with TalkingData AdTracking, which is apparently my first entry to Kaggle cometitions. The first line was **pd.read_csv() **and I got this error. I thought my code was run in the cloud and that I don't have to worry about the memory requirements. Can someone help me with that?</p>",
  "messages": [
    {
      "id": 1112717,
      "postDate": "2020-12-14T20:40:39.337Z",
      "content": "<p>You may use Dask to load the data that is larger that your memory capacity.</p>",
      "rawMarkdown": "You may use Dask to load the data that is larger that your memory capacity.",
      "votes": 1
    },
    {
      "id": 886638,
      "postDate": "2020-06-15T07:10:00.957Z",
      "content": "<p>Hi. I am getting the same error now in spite of multiple runs. Any pointers on how to resolve this?</p>",
      "rawMarkdown": "Hi. I am getting the same error now in spite of multiple runs. Any pointers on how to resolve this?",
      "replies": [
        {
          "id": 886852,
          "postDate": "2020-06-15T10:05:27.633Z",
          "content": "<p>if you just try reading the entire file, you may get this error because there are close to 150 mil records on this file. Ideally, Kaggle cloud kernels should handle this load but sometimes you will get this error. Try running after sometime. </p>\n\n<p>You can only the first million records, with nrows attribute in the read_csv() function. This will reduce the load on the cloud. A million records should be sufficient to train your model. Adding more records above a million will have little to no effect on your accuracy. \nHope this helpls</p>",
          "rawMarkdown": "\nif you just try reading the entire file, you may get this error because there are close to 150 mil records on this file. Ideally, Kaggle cloud kernels should handle this load but sometimes you will get this error. Try running after sometime. \n\nYou can only the first million records, with nrows attribute in the read_csv() function. This will reduce the load on the cloud. A million records should be sufficient to train your model. Adding more records above a million will have little to no effect on your accuracy. \nHope this helpls",
          "votes": 1
        },
        {
          "id": 888028,
          "postDate": "2020-06-16T04:32:10.350Z",
          "content": "<p>Thank you. Wouldn't more training data imply better learning of the relationships and thus a better model over all?</p>",
          "rawMarkdown": "Thank you. Wouldn't more training data imply better learning of the relationships and thus a better model over all?"
        }
      ]
    },
    {
      "id": 881138,
      "postDate": "2020-06-10T19:03:42.257Z",
      "content": "<p>I was getting started with TalkingData AdTracking, which is apparently my first entry to Kaggle cometitions. The first line was **pd.read_csv() **and I got this error. I thought my code was run in the cloud and that I don't have to worry about the memory requirements. Can someone help me with that?</p>",
      "rawMarkdown": "I was getting started with TalkingData AdTracking, which is apparently my first entry to Kaggle cometitions. The first line was **pd.read_csv() **and I got this error. I thought my code was run in the cloud and that I don't have to worry about the memory requirements. Can someone help me with that?\n\n"
    },
    {
      "id": 881532,
      "postDate": "2020-06-11T06:29:30.467Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 881676,
          "postDate": "2020-06-11T08:50:09.273Z",
          "content": "<p>Aren't Kaggle Kernels run on Kaggle Clouds? I was getting the error on Kaggle and not on my local notebook. And, now for some odd reason I am not getting the error. </p>",
          "rawMarkdown": "Aren't Kaggle Kernels run on Kaggle Clouds? I was getting the error on Kaggle and not on my local notebook. And, now for some odd reason I am not getting the error. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1112717,
      "author_name": "emoticn",
      "author_url": "",
      "post_date": "2020-12-14T20:40:39.337000",
      "content": "<p>You may use Dask to load the data that is larger that your memory capacity.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 886638,
      "author_name": "Shrinidhi Narasimhan",
      "author_url": "",
      "post_date": "2020-06-15T07:10:00.957000",
      "content": "<p>Hi. I am getting the same error now in spite of multiple runs. Any pointers on how to resolve this?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 886852,
          "author_name": "satish chavan",
          "author_url": "",
          "post_date": "2020-06-15T10:05:27.633000",
          "content": "<p>if you just try reading the entire file, you may get this error because there are close to 150 mil records on this file. Ideally, Kaggle cloud kernels should handle this load but sometimes you will get this error. Try running after sometime. </p>\n\n<p>You can only the first million records, with nrows attribute in the read_csv() function. This will reduce the load on the cloud. A million records should be sufficient to train your model. Adding more records above a million will have little to no effect on your accuracy. \nHope this helpls</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 888028,
          "author_name": "Shrinidhi Narasimhan",
          "author_url": "",
          "post_date": "2020-06-16T04:32:10.350000",
          "content": "<p>Thank you. Wouldn't more training data imply better learning of the relationships and thus a better model over all?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 881532,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-11T06:29:30.467000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 881676,
          "author_name": "satish chavan",
          "author_url": "",
          "post_date": "2020-06-11T08:50:09.273000",
          "content": "<p>Aren't Kaggle Kernels run on Kaggle Clouds? I was getting the error on Kaggle and not on my local notebook. And, now for some odd reason I am not getting the error. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1112717": "You may use Dask to load the data that is larger that your memory capacity.",
    "886638": "Hi. I am getting the same error now in spite of multiple runs. Any pointers on how to resolve this?",
    "881138": "I was getting started with TalkingData AdTracking, which is apparently my first entry to Kaggle cometitions. The first line was **pd.read_csv() **and I got this error. I thought my code was run in the cloud and that I don't have to worry about the memory requirements. Can someone help me with that?\n\n",
    "881532": ""
  }
}