{
  "id": 92272,
  "title": "The kernel appears to have died. It will restart automatically? Please help me...",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/92272",
  "author_name": "",
  "post_date": "2019-05-15T01:44:35.101989400Z",
  "votes": 3,
  "comment_count": 17,
  "views": 0,
  "content": "<p>When I run the kernel with code:\n<code>train_data = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype={'acoustic_data': np.int16, 'time_to_file': np.float64})</code>，I get the message \"The kernel appears to have died. It will restart automatically\", and the kernel died immediately.\nEven though I change the code like this:\n<code>train_data = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype={'acoustic_data': np.int16, 'time_to_file': np.float32}),</code> the same message occurs too.\nI have tried a lot of times, anyone can tell me how to solve the problem?\nBy the way: I didn't get the trouble in the past days until yesterday.</p>",
  "messages": [
    {
      "id": "531478",
      "postDate": "05/15/2019 01:44:35",
      "content": "<p>When I run the kernel with code:\n<code>train_data = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype={'acoustic_data': np.int16, 'time_to_file': np.float64})</code>，I get the message \"The kernel appears to have died. It will restart automatically\", and the kernel died immediately.\nEven though I change the code like this:\n<code>train_data = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype={'acoustic_data': np.int16, 'time_to_file': np.float32}),</code> the same message occurs too.\nI have tried a lot of times, anyone can tell me how to solve the problem?\nBy the way: I didn't get the trouble in the past days until yesterday.</p>",
      "rawMarkdown": "When I run the kernel with code:\n`train_data = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype={'acoustic_data': np.int16, 'time_to_file': np.float64})`，I get the message \"The kernel appears to have died. It will restart automatically\", and the kernel died immediately.\nEven though I change the code like this:\n`train_data = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype={'acoustic_data': np.int16, 'time_to_file': np.float32}),` the same message occurs too.\nI have tried a lot of times, anyone can tell me how to solve the problem?\nBy the way: I didn't get the trouble in the past days until yesterday.",
      "votes": null
    },
    {
      "id": "531481",
      "postDate": "05/15/2019 01:57:43",
      "content": "<p>I ran it again, and this time I got the correct result. But I just want to know why the problem occurs and how I can sove it if I get trouble with it again.</p>",
      "rawMarkdown": "I ran it again, and this time I got the correct result. But I just want to know why the problem occurs and how I can sove it if I get trouble with it again.",
      "votes": null
    },
    {
      "id": "531556",
      "postDate": "05/15/2019 06:08:57",
      "content": "<p>Please!!! I get the trouble again...</p>",
      "rawMarkdown": "Please!!! I get the trouble again...",
      "votes": null
    },
    {
      "id": "531559",
      "postDate": "05/15/2019 06:10:02",
      "content": "<p>I find that it is because the RAM up to 16G. what's wrong with it?</p>",
      "rawMarkdown": "I find that it is because the RAM up to 16G. what's wrong with it?",
      "votes": null
    },
    {
      "id": "531614",
      "postDate": "05/15/2019 07:45:35",
      "content": "<p>I could not load the LANL dataset in my kernel. I think we may get the same problem.</p>",
      "rawMarkdown": "I could not load the LANL dataset in my kernel. I think we may get the same problem.",
      "votes": null
    },
    {
      "id": "531628",
      "postDate": "05/15/2019 08:27:58",
      "content": "<p>The other column is called <code>time_to_failure</code> not <code>time_to_file</code></p>",
      "rawMarkdown": "The other column is called `time_to_failure` not `time_to_file`",
      "votes": null
    },
    {
      "id": "531650",
      "postDate": "05/15/2019 09:20:53",
      "content": "<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you very much!",
      "votes": null
    },
    {
      "id": "531651",
      "postDate": "05/15/2019 09:22:01",
      "content": "<p>Maybe you can tell us about your problem.</p>",
      "rawMarkdown": "Maybe you can tell us about your problem.",
      "votes": null
    },
    {
      "id": "531710",
      "postDate": "05/15/2019 11:38:09",
      "content": "<p><code>16GB</code> of <code>RAM</code> is not enough to fruitfully process the entire <code>train.csv</code> at once, you should consider using the <code>chunksize</code> parameter of the <code>pandas.read_csv</code> method. If needed, use the <code>partial_fit</code> method of a suitable model. </p>",
      "rawMarkdown": "`16GB` of `RAM` is not enough to fruitfully process the entire `train.csv` at once, you should consider using the `chunksize` parameter of the `pandas.read_csv` method. If needed, use the `partial_fit` method of a suitable model.",
      "votes": null
    },
    {
      "id": "531741",
      "postDate": "05/15/2019 12:57:29",
      "content": "<p>Thank you. I thought anthoer method to read the original data with float64:\n```\n%%time</p>\n\n<p>train_data_ad = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype=np.int16, usecols=['acoustic_data'])\ntrain_data_ttf = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype=np.float64, usecols=['time_to_failure'])\ntrain_data = pd.concat([train_data_ad, train_data_ttf], axis=1)\ndel train_data_ad\ndel train_data_ttf\n```</p>",
      "rawMarkdown": "Thank you. I thought anthoer method to read the original data with float64:\n```\n%%time\n\ntrain_data_ad = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype=np.int16, usecols=['acoustic_data'])\ntrain_data_ttf = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype=np.float64, usecols=['time_to_failure'])\ntrain_data = pd.concat([train_data_ad, train_data_ttf], axis=1)\ndel train_data_ad\ndel train_data_ttf\n```",
      "votes": null
    },
    {
      "id": "531764",
      "postDate": "05/15/2019 13:58:00",
      "content": "<p>I got the error message: Failed to load input data, this could result in the Kernel failing to load.</p>",
      "rawMarkdown": "I got the error message: Failed to load input data, this could result in the Kernel failing to load.",
      "votes": null
    },
    {
      "id": "531765",
      "postDate": "05/15/2019 14:03:17",
      "content": "<p>I also got the error from time to time. But it didn't influence the runing of the kernel.</p>",
      "rawMarkdown": "I also got the error from time to time. But it didn't influence the runing of the kernel.",
      "votes": null
    },
    {
      "id": "531770",
      "postDate": "05/15/2019 14:15:50",
      "content": "<p>This doesn't make it easier, if anything it gets worse because you are reading in the whole amount of original data and after that you are making another copy of it just before deleting! The memory requirements increase here actually rather than decreasing. See, bifurcating the data into two dataframes does not lower the memory requirement, all the data is being loaded anyway. This is why reading in chunks is the way to go if memory is low.</p>",
      "rawMarkdown": "This doesn't make it easier, if anything it gets worse because you are reading in the whole amount of original data and after that you are making another copy of it just before deleting! The memory requirements increase here actually rather than decreasing. See, bifurcating the data into two dataframes does not lower the memory requirement, all the data is being loaded anyway. This is why reading in chunks is the way to go if memory is low.",
      "votes": null
    },
    {
      "id": "531784",
      "postDate": "05/15/2019 14:36:25",
      "content": "<p>Yeah, but I found a fact that the function <code>pd.read_csv()</code> occupy higher RAM when we running it. When the function finished, the RAM of occupation will be lower. So I bifurcating the data into two dataframes. In fact, it works.  Anyway, thank you very much!</p>",
      "rawMarkdown": "Yeah, but I found a fact that the function `pd.read_csv()` occupy higher RAM when we running it. When the function finished, the RAM of occupation will be lower. So I bifurcating the data into two dataframes. In fact, it works.  Anyway, thank you very much!",
      "votes": null
    },
    {
      "id": "531811",
      "postDate": "05/15/2019 15:26:03",
      "content": "<p>I checked it out and it does give some leg-room for sure, but hits the limits after that. Perhaps you could try reading in chunks if you are getting errors often. This is just the memory management part though, I wish you all the best for the competition!</p>",
      "rawMarkdown": "I checked it out and it does give some leg-room for sure, but hits the limits after that. Perhaps you could try reading in chunks if you are getting errors often. This is just the memory management part though, I wish you all the best for the competition!",
      "votes": null
    },
    {
      "id": "532055",
      "postDate": "05/16/2019 05:28:28",
      "content": "<p>It’s probably because of the huge size of the train data. Try to load a small portion just for checking.</p>",
      "rawMarkdown": "It’s probably because of the huge size of the train data. Try to load a small portion just for checking.",
      "votes": null
    },
    {
      "id": "532224",
      "postDate": "05/16/2019 12:55:49",
      "content": "<p>Thank you.</p>",
      "rawMarkdown": "Thank you.",
      "votes": null
    },
    {
      "id": "532225",
      "postDate": "05/16/2019 12:57:48",
      "content": "<p>Thank you.  I wish you all the best for the competition too!</p>",
      "rawMarkdown": "Thank you.  I wish you all the best for the competition too!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 531481,
      "author_name": "",
      "author_url": "",
      "post_date": "05/15/2019 01:57:43",
      "content": "<p>I ran it again, and this time I got the correct result. But I just want to know why the problem occurs and how I can sove it if I get trouble with it again.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 531556,
      "author_name": "",
      "author_url": "",
      "post_date": "05/15/2019 06:08:57",
      "content": "<p>Please!!! I get the trouble again...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 531559,
      "author_name": "",
      "author_url": "",
      "post_date": "05/15/2019 06:10:02",
      "content": "<p>I find that it is because the RAM up to 16G. what's wrong with it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 531710,
          "author_name": "kishupro",
          "author_url": "",
          "post_date": "05/15/2019 11:38:09",
          "content": "<p><code>16GB</code> of <code>RAM</code> is not enough to fruitfully process the entire <code>train.csv</code> at once, you should consider using the <code>chunksize</code> parameter of the <code>pandas.read_csv</code> method. If needed, use the <code>partial_fit</code> method of a suitable model. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 531741,
          "author_name": "",
          "author_url": "",
          "post_date": "05/15/2019 12:57:29",
          "content": "<p>Thank you. I thought anthoer method to read the original data with float64:\n```\n%%time</p>\n\n<p>train_data_ad = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype=np.int16, usecols=['acoustic_data'])\ntrain_data_ttf = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype=np.float64, usecols=['time_to_failure'])\ntrain_data = pd.concat([train_data_ad, train_data_ttf], axis=1)\ndel train_data_ad\ndel train_data_ttf\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 531770,
          "author_name": "kishupro",
          "author_url": "",
          "post_date": "05/15/2019 14:15:50",
          "content": "<p>This doesn't make it easier, if anything it gets worse because you are reading in the whole amount of original data and after that you are making another copy of it just before deleting! The memory requirements increase here actually rather than decreasing. See, bifurcating the data into two dataframes does not lower the memory requirement, all the data is being loaded anyway. This is why reading in chunks is the way to go if memory is low.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 531784,
          "author_name": "",
          "author_url": "",
          "post_date": "05/15/2019 14:36:25",
          "content": "<p>Yeah, but I found a fact that the function <code>pd.read_csv()</code> occupy higher RAM when we running it. When the function finished, the RAM of occupation will be lower. So I bifurcating the data into two dataframes. In fact, it works.  Anyway, thank you very much!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 531811,
          "author_name": "kishupro",
          "author_url": "",
          "post_date": "05/15/2019 15:26:03",
          "content": "<p>I checked it out and it does give some leg-room for sure, but hits the limits after that. Perhaps you could try reading in chunks if you are getting errors often. This is just the memory management part though, I wish you all the best for the competition!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 532225,
          "author_name": "",
          "author_url": "",
          "post_date": "05/16/2019 12:57:48",
          "content": "<p>Thank you.  I wish you all the best for the competition too!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 531614,
      "author_name": "dysonlin",
      "author_url": "",
      "post_date": "05/15/2019 07:45:35",
      "content": "<p>I could not load the LANL dataset in my kernel. I think we may get the same problem.</p>",
      "votes": null,
      "replies": [
        {
          "id": 531651,
          "author_name": "",
          "author_url": "",
          "post_date": "05/15/2019 09:22:01",
          "content": "<p>Maybe you can tell us about your problem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 531764,
          "author_name": "dysonlin",
          "author_url": "",
          "post_date": "05/15/2019 13:58:00",
          "content": "<p>I got the error message: Failed to load input data, this could result in the Kernel failing to load.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 531765,
          "author_name": "",
          "author_url": "",
          "post_date": "05/15/2019 14:03:17",
          "content": "<p>I also got the error from time to time. But it didn't influence the runing of the kernel.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 531628,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "05/15/2019 08:27:58",
      "content": "<p>The other column is called <code>time_to_failure</code> not <code>time_to_file</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 531650,
          "author_name": "",
          "author_url": "",
          "post_date": "05/15/2019 09:20:53",
          "content": "<p>Thank you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 532055,
      "author_name": "yunlongzhang",
      "author_url": "",
      "post_date": "05/16/2019 05:28:28",
      "content": "<p>It’s probably because of the huge size of the train data. Try to load a small portion just for checking.</p>",
      "votes": null,
      "replies": [
        {
          "id": 532224,
          "author_name": "",
          "author_url": "",
          "post_date": "05/16/2019 12:55:49",
          "content": "<p>Thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "531478": "When I run the kernel with code:\n`train_data = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype={'acoustic_data': np.int16, 'time_to_file': np.float64})`，I get the message \"The kernel appears to have died. It will restart automatically\", and the kernel died immediately.\nEven though I change the code like this:\n`train_data = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype={'acoustic_data': np.int16, 'time_to_file': np.float32}),` the same message occurs too.\nI have tried a lot of times, anyone can tell me how to solve the problem?\nBy the way: I didn't get the trouble in the past days until yesterday.",
    "531481": "I ran it again, and this time I got the correct result. But I just want to know why the problem occurs and how I can sove it if I get trouble with it again.",
    "531556": "Please!!! I get the trouble again...",
    "531559": "I find that it is because the RAM up to 16G. what's wrong with it?",
    "531614": "I could not load the LANL dataset in my kernel. I think we may get the same problem.",
    "531628": "The other column is called `time_to_failure` not `time_to_file`",
    "531650": "Thank you very much!",
    "531651": "Maybe you can tell us about your problem.",
    "531710": "`16GB` of `RAM` is not enough to fruitfully process the entire `train.csv` at once, you should consider using the `chunksize` parameter of the `pandas.read_csv` method. If needed, use the `partial_fit` method of a suitable model.",
    "531741": "Thank you. I thought anthoer method to read the original data with float64:\n```\n%%time\n\ntrain_data_ad = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype=np.int16, usecols=['acoustic_data'])\ntrain_data_ttf = pd.read_csv(os.path.join(PATH, 'train.csv'), dtype=np.float64, usecols=['time_to_failure'])\ntrain_data = pd.concat([train_data_ad, train_data_ttf], axis=1)\ndel train_data_ad\ndel train_data_ttf\n```",
    "531764": "I got the error message: Failed to load input data, this could result in the Kernel failing to load.",
    "531765": "I also got the error from time to time. But it didn't influence the runing of the kernel.",
    "531770": "This doesn't make it easier, if anything it gets worse because you are reading in the whole amount of original data and after that you are making another copy of it just before deleting! The memory requirements increase here actually rather than decreasing. See, bifurcating the data into two dataframes does not lower the memory requirement, all the data is being loaded anyway. This is why reading in chunks is the way to go if memory is low.",
    "531784": "Yeah, but I found a fact that the function `pd.read_csv()` occupy higher RAM when we running it. When the function finished, the RAM of occupation will be lower. So I bifurcating the data into two dataframes. In fact, it works.  Anyway, thank you very much!",
    "531811": "I checked it out and it does give some leg-room for sure, but hits the limits after that. Perhaps you could try reading in chunks if you are getting errors often. This is just the memory management part though, I wish you all the best for the competition!",
    "532055": "It’s probably because of the huge size of the train data. Try to load a small portion just for checking.",
    "532224": "Thank you.",
    "532225": "Thank you.  I wish you all the best for the competition too!"
  },
  "source": "meta"
}