{
  "id": 191463,
  "title": "RAM usage in a kaggle notebook",
  "url": "/competitions/riiid-test-answer-prediction/discussion/191463",
  "author_name": "",
  "post_date": "2020-10-16T16:14:09.670416500Z",
  "votes": 5,
  "comment_count": 13,
  "views": 0,
  "content": "<p>hi,<br>\nI am facing an issue related to memory usage of the kaggle notebook .<br>\nwhen I load the train file:</p>\n<pre><code> train = pd.read_csv('file_path')\n</code></pre>\n<p>it occupies around 9.5 GB of the RAM.<br>\nbut if I (for example) deleted the training DataFrame object directly after reading the csv file and use the garbage collector:</p>\n<pre><code>del train\ngc.collect()\n</code></pre>\n<p>the RAM is still 5.9 GB occupied although I cleared everything !</p>\n<p>why I still have things loaded into the RAM ? <br>\nand how can I free up the RAM ? <br>\nis there a way to know what is currently loaded into the kaggle RAM ?</p>",
  "messages": [
    {
      "id": "1051550",
      "postDate": "10/16/2020 16:14:09",
      "content": "<p>hi,<br>\nI am facing an issue related to memory usage of the kaggle notebook .<br>\nwhen I load the train file:</p>\n<pre><code> train = pd.read_csv('file_path')\n</code></pre>\n<p>it occupies around 9.5 GB of the RAM.<br>\nbut if I (for example) deleted the training DataFrame object directly after reading the csv file and use the garbage collector:</p>\n<pre><code>del train\ngc.collect()\n</code></pre>\n<p>the RAM is still 5.9 GB occupied although I cleared everything !</p>\n<p>why I still have things loaded into the RAM ? <br>\nand how can I free up the RAM ? <br>\nis there a way to know what is currently loaded into the kaggle RAM ?</p>",
      "rawMarkdown": "hi,\nI am facing an issue related to memory usage of the kaggle notebook .\nwhen I load the train file:\n```\n train = pd.read_csv('file_path')\n```\nit occupies around 9.5 GB of the RAM.\nbut if I (for example) deleted the training DataFrame object directly after reading the csv file and use the garbage collector:\n```\ndel train\ngc.collect()\n```\nthe RAM is still 5.9 GB occupied although I cleared everything !\n\nwhy I still have things loaded into the RAM ? \nand how can I free up the RAM ? \nis there a way to know what is currently loaded into the kaggle RAM ?",
      "votes": null
    },
    {
      "id": "1052265",
      "postDate": "10/17/2020 14:26:03",
      "content": "<p>I got the same problem. might be the system cache,<br>\nespecially when you do lots of pd.merge. Now I do the training in my pc and test in the kaggle notebook.</p>",
      "rawMarkdown": "I got the same problem. might be the system cache,\nespecially when you do lots of pd.merge. Now I do the training in my pc and test in the kaggle notebook.",
      "votes": null
    },
    {
      "id": "1053635",
      "postDate": "10/19/2020 07:26:31",
      "content": "<p>I had the same kind of problems, I removed some .head() functions and <em>magic</em> got no more RAM issues. I cannot say why unfortunately </p>",
      "rawMarkdown": "I had the same kind of problems, I removed some .head() functions and *magic* got no more RAM issues. I cannot say why unfortunately",
      "votes": null
    },
    {
      "id": "1053873",
      "postDate": "10/19/2020 12:43:37",
      "content": "<p><a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a> <br>\nbut I am not using any code, just reading csv file, and deleting the DataFrame directly after that.</p>",
      "rawMarkdown": "alexj21 \nbut I am not using any code, just reading csv file, and deleting the DataFrame directly after that.",
      "votes": null
    },
    {
      "id": "1101016",
      "postDate": "12/03/2020 14:57:21",
      "content": "<p>I have this same problem. Did anyone find a solution?</p>\n<p>I use del df and then gc.collect(), but RAM remains the same.</p>",
      "rawMarkdown": "I have this same problem. Did anyone find a solution?\n\nI use del df and then gc.collect(), but RAM remains the same.",
      "votes": null
    },
    {
      "id": "1101759",
      "postDate": "12/04/2020 08:01:58",
      "content": "<p>no, I didn't find solutions for that, it seems as a bug in the kaggle notebook. <br>\nbut I found some ways around the memory usage. for example, </p>\n<ul>\n<li>use datatable library to load the data (and then convert to pandas). </li>\n<li>reduce the datatypes of features, </li>\n<li>deal with each feature and then move to other feature instead of all features at once which exhausts the memory </li>\n<li>and some other ideas</li>\n</ul>",
      "rawMarkdown": "no, I didn't find solutions for that, it seems as a bug in the kaggle notebook. \nbut I found some ways around the memory usage. for example, \n- use datatable library to load the data (and then convert to pandas). \n- reduce the datatypes of features, \n- deal with each feature and then move to other feature instead of all features at once which exhausts the memory \n- and some other ideas",
      "votes": null
    },
    {
      "id": "1102040",
      "postDate": "12/04/2020 14:03:01",
      "content": "<p>Thank you!</p>\n<p>My experience:</p>\n<ul>\n<li>If I create a notebook from scratch then I have this issue, ie RAM does not get freed.</li>\n<li>However, when I copy someone else's notebook I do see RAM get freed.</li>\n</ul>\n<p>I'm a beginner and don't understand why there is a difference.</p>\n<p>Hopefully this helps someone else. I spent a ton of time on this.</p>",
      "rawMarkdown": "Thank you!\n\nMy experience:\n* If I create a notebook from scratch then I have this issue, ie RAM does not get freed.\n* However, when I copy someone else's notebook I do see RAM get freed.\n\nI'm a beginner and don't understand why there is a difference.\n\nHopefully this helps someone else. I spent a ton of time on this.",
      "votes": null
    },
    {
      "id": "1102218",
      "postDate": "12/04/2020 17:45:30",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mohamadnawfal\" target=\"_blank\">@mohamadnawfal</a>! I faced similar issues and after a dozen of trial and error I discovered that doing a .head() causes this issue for me. Similar to <a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a> mine disappeared after removing this one line of code. </p>\n<p>Another fix that I read on stackoverflow which might work for you is to do a <code>train = None</code> after <code>del train</code> which apparantly helps de reference the variable and help gc to do its job. </p>\n<p>Hope this helps :)</p>",
      "rawMarkdown": "Hi @mohamadnawfal! I faced similar issues and after a dozen of trial and error I discovered that doing a .head() causes this issue for me. Similar to @alexj21 mine disappeared after removing this one line of code. \n\nAnother fix that I read on stackoverflow which might work for you is to do a ```train = None``` after ```del train``` which apparantly helps de reference the variable and help gc to do its job. \n\nHope this helps :)",
      "votes": null
    },
    {
      "id": "1102236",
      "postDate": "12/04/2020 18:06:57",
      "content": "<p>Thanks for sharing your experience :)</p>",
      "rawMarkdown": "Thanks for sharing your experience :)",
      "votes": null
    },
    {
      "id": "1105214",
      "postDate": "12/07/2020 16:19:53",
      "content": "<p>Naresh, Alex,</p>\n<p>I also see that removing the .head() solves the RAM issue. Is there an alternative to head that allows me to see the data while freeing RAM?</p>\n<p>Thank you for any pointers!</p>",
      "rawMarkdown": "Naresh, Alex,\n\nI also see that removing the .head() solves the RAM issue. Is there an alternative to head that allows me to see the data while freeing RAM?\n\nThank you for any pointers!",
      "votes": null
    },
    {
      "id": "1105233",
      "postDate": "12/07/2020 16:46:34",
      "content": "<p>I am kindly seeking help to explain this behavior:</p>\n<p>My RAM is at 3.9G with full_train data loaded. I test 2 cases:</p>\n<p>(A) I run the following:<br>\n\"full_train = full_train[['row_id','user_id','content_id','content_type_id','answered_correctly']]\"<br>\nI see my RAM decrease from 3.9G to 1.7G as expected, since I have reduced columns.</p>\n<p>(B) I run \"full_train.head()\".<br>\nMy RAM is still at 3.9G as expected.<br>\nNow I run \"full_train = full_train[['row_id','user_id','content_id','content_type_id','answered_correctly']]\"<br>\nMy RAM increases from 3.9G to 5.4G.</p>\n<p>Why is this happening? <br>\nI theorize that my notebook is effectively keeping a copy of full_train associated with the head() command.<br>\nHow can I fix this? </p>\n<p>Thank you in advance!</p>",
      "rawMarkdown": "I am kindly seeking help to explain this behavior:\n\nMy RAM is at 3.9G with full_train data loaded. I test 2 cases:\n\n(A) I run the following:\n\"full_train = full_train[['row_id','user_id','content_id','content_type_id','answered_correctly']]\"\nI see my RAM decrease from 3.9G to 1.7G as expected, since I have reduced columns.\n\n\n(B) I run \"full_train.head()\".\nMy RAM is still at 3.9G as expected.\nNow I run \"full_train = full_train[['row_id','user_id','content_id','content_type_id','answered_correctly']]\"\nMy RAM increases from 3.9G to 5.4G.\n\n\nWhy is this happening? \nI theorize that my notebook is effectively keeping a copy of full_train associated with the head() command.\nHow can I fix this? \n\nThank you in advance!",
      "votes": null
    },
    {
      "id": "1105298",
      "postDate": "12/07/2020 18:31:48",
      "content": "<p>Work-around found:</p>\n<p>Replace \"full_train.head()\"<br>\nwith \"print(full_train.head()\"</p>\n<p>By adding the print() the RAM gets freed and everything behaves as expected, ie when I run \"del full_train; gc.collect()\" I can free the RAM associated with the deleted variable. Without the print, the RAM does not get freed with gc.collect()</p>\n<p>I do not understand why this happens, but I hope others can avoid the pain I've been through.</p>\n<p>If anyone understands this behavior, I'd love an explanation!</p>",
      "rawMarkdown": "Work-around found:\n\nReplace \"full_train.head()\"\nwith \"print(full_train.head()\"\n\nBy adding the print() the RAM gets freed and everything behaves as expected, ie when I run \"del full_train; gc.collect()\" I can free the RAM associated with the deleted variable. Without the print, the RAM does not get freed with gc.collect()\n\nI do not understand why this happens, but I hope others can avoid the pain I've been through.\n\nIf anyone understands this behavior, I'd love an explanation!",
      "votes": null
    },
    {
      "id": "1105300",
      "postDate": "12/07/2020 18:37:05",
      "content": "<p>I think you can try running same code on your personal computer (jupyter notebook) and check if same behavior persist, if not, then it means there might be a problem with kaggle notebooks. otherwise its an issue with pandas itself. if you find answers please share with me</p>\n<p>anyway, you can remove the .head() and avoid the issue as you mentioned in your other reply. </p>",
      "rawMarkdown": "I think you can try running same code on your personal computer (jupyter notebook) and check if same behavior persist, if not, then it means there might be a problem with kaggle notebooks. otherwise its an issue with pandas itself. if you find answers please share with me\n\nanyway, you can remove the .head() and avoid the issue as you mentioned in your other reply.",
      "votes": null
    },
    {
      "id": "1105790",
      "postDate": "12/08/2020 07:44:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/thomasliptay\" target=\"_blank\">@thomasliptay</a>, I haven't been able to find any alternatives to .head().  I observed that doing operations on the data by itself doesn't cause this issue, only <strong>displaying</strong> the output does (using .head, .sample, etc). All the transformations that need to be done, could be written as a function, perhaps chain it all together. Test this \"pipeline\" on a small subset of the actual data to ensure it works as intended. Once you have ensured that it works good, restart the notebook and run again, this time <em>not displaying</em> any intermediate results (on the entire dataset). </p>\n<p>A good rule of thumb would be to keep the explorations separate from the actual modeling. Modeling is where you would require dealing with large amounts of data.  </p>\n<p>Another <em>trick</em> that might assist you is the magic command: <code>%reset -f</code>. This would reset the memory and all the variables you had created, including any module imports you might have done. The files written on the hard disk for that session would remain unaffected though. Using this command with caution! </p>\n<p>Hope this helps! Have a nice day :)</p>",
      "rawMarkdown": "Hi @thomasliptay, I haven't been able to find any alternatives to .head().  I observed that doing operations on the data by itself doesn't cause this issue, only **displaying** the output does (using .head, .sample, etc). All the transformations that need to be done, could be written as a function, perhaps chain it all together. Test this \"pipeline\" on a small subset of the actual data to ensure it works as intended. Once you have ensured that it works good, restart the notebook and run again, this time *not displaying* any intermediate results (on the entire dataset). \n\nA good rule of thumb would be to keep the explorations separate from the actual modeling. Modeling is where you would require dealing with large amounts of data.  \n\nAnother *trick* that might assist you is the magic command: ```%reset -f```. This would reset the memory and all the variables you had created, including any module imports you might have done. The files written on the hard disk for that session would remain unaffected though. Using this command with caution! \n\nHope this helps! Have a nice day :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1052265,
      "author_name": "xmoremoe",
      "author_url": "",
      "post_date": "10/17/2020 14:26:03",
      "content": "<p>I got the same problem. might be the system cache,<br>\nespecially when you do lots of pd.merge. Now I do the training in my pc and test in the kaggle notebook.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1053635,
      "author_name": "alexj21",
      "author_url": "",
      "post_date": "10/19/2020 07:26:31",
      "content": "<p>I had the same kind of problems, I removed some .head() functions and <em>magic</em> got no more RAM issues. I cannot say why unfortunately </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1053873,
      "author_name": "mohamadnawfal",
      "author_url": "",
      "post_date": "10/19/2020 12:43:37",
      "content": "<p><a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a> <br>\nbut I am not using any code, just reading csv file, and deleting the DataFrame directly after that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1101016,
      "author_name": "thomasliptay",
      "author_url": "",
      "post_date": "12/03/2020 14:57:21",
      "content": "<p>I have this same problem. Did anyone find a solution?</p>\n<p>I use del df and then gc.collect(), but RAM remains the same.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1101759,
          "author_name": "mohamadnawfal",
          "author_url": "",
          "post_date": "12/04/2020 08:01:58",
          "content": "<p>no, I didn't find solutions for that, it seems as a bug in the kaggle notebook. <br>\nbut I found some ways around the memory usage. for example, </p>\n<ul>\n<li>use datatable library to load the data (and then convert to pandas). </li>\n<li>reduce the datatypes of features, </li>\n<li>deal with each feature and then move to other feature instead of all features at once which exhausts the memory </li>\n<li>and some other ideas</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1102040,
          "author_name": "thomasliptay",
          "author_url": "",
          "post_date": "12/04/2020 14:03:01",
          "content": "<p>Thank you!</p>\n<p>My experience:</p>\n<ul>\n<li>If I create a notebook from scratch then I have this issue, ie RAM does not get freed.</li>\n<li>However, when I copy someone else's notebook I do see RAM get freed.</li>\n</ul>\n<p>I'm a beginner and don't understand why there is a difference.</p>\n<p>Hopefully this helps someone else. I spent a ton of time on this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1102218,
      "author_name": "doctorkael",
      "author_url": "",
      "post_date": "12/04/2020 17:45:30",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mohamadnawfal\" target=\"_blank\">@mohamadnawfal</a>! I faced similar issues and after a dozen of trial and error I discovered that doing a .head() causes this issue for me. Similar to <a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a> mine disappeared after removing this one line of code. </p>\n<p>Another fix that I read on stackoverflow which might work for you is to do a <code>train = None</code> after <code>del train</code> which apparantly helps de reference the variable and help gc to do its job. </p>\n<p>Hope this helps :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1102236,
          "author_name": "mohamadnawfal",
          "author_url": "",
          "post_date": "12/04/2020 18:06:57",
          "content": "<p>Thanks for sharing your experience :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105214,
          "author_name": "thomasliptay",
          "author_url": "",
          "post_date": "12/07/2020 16:19:53",
          "content": "<p>Naresh, Alex,</p>\n<p>I also see that removing the .head() solves the RAM issue. Is there an alternative to head that allows me to see the data while freeing RAM?</p>\n<p>Thank you for any pointers!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105790,
          "author_name": "doctorkael",
          "author_url": "",
          "post_date": "12/08/2020 07:44:26",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/thomasliptay\" target=\"_blank\">@thomasliptay</a>, I haven't been able to find any alternatives to .head().  I observed that doing operations on the data by itself doesn't cause this issue, only <strong>displaying</strong> the output does (using .head, .sample, etc). All the transformations that need to be done, could be written as a function, perhaps chain it all together. Test this \"pipeline\" on a small subset of the actual data to ensure it works as intended. Once you have ensured that it works good, restart the notebook and run again, this time <em>not displaying</em> any intermediate results (on the entire dataset). </p>\n<p>A good rule of thumb would be to keep the explorations separate from the actual modeling. Modeling is where you would require dealing with large amounts of data.  </p>\n<p>Another <em>trick</em> that might assist you is the magic command: <code>%reset -f</code>. This would reset the memory and all the variables you had created, including any module imports you might have done. The files written on the hard disk for that session would remain unaffected though. Using this command with caution! </p>\n<p>Hope this helps! Have a nice day :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1105233,
      "author_name": "thomasliptay",
      "author_url": "",
      "post_date": "12/07/2020 16:46:34",
      "content": "<p>I am kindly seeking help to explain this behavior:</p>\n<p>My RAM is at 3.9G with full_train data loaded. I test 2 cases:</p>\n<p>(A) I run the following:<br>\n\"full_train = full_train[['row_id','user_id','content_id','content_type_id','answered_correctly']]\"<br>\nI see my RAM decrease from 3.9G to 1.7G as expected, since I have reduced columns.</p>\n<p>(B) I run \"full_train.head()\".<br>\nMy RAM is still at 3.9G as expected.<br>\nNow I run \"full_train = full_train[['row_id','user_id','content_id','content_type_id','answered_correctly']]\"<br>\nMy RAM increases from 3.9G to 5.4G.</p>\n<p>Why is this happening? <br>\nI theorize that my notebook is effectively keeping a copy of full_train associated with the head() command.<br>\nHow can I fix this? </p>\n<p>Thank you in advance!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1105298,
          "author_name": "thomasliptay",
          "author_url": "",
          "post_date": "12/07/2020 18:31:48",
          "content": "<p>Work-around found:</p>\n<p>Replace \"full_train.head()\"<br>\nwith \"print(full_train.head()\"</p>\n<p>By adding the print() the RAM gets freed and everything behaves as expected, ie when I run \"del full_train; gc.collect()\" I can free the RAM associated with the deleted variable. Without the print, the RAM does not get freed with gc.collect()</p>\n<p>I do not understand why this happens, but I hope others can avoid the pain I've been through.</p>\n<p>If anyone understands this behavior, I'd love an explanation!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105300,
          "author_name": "mohamadnawfal",
          "author_url": "",
          "post_date": "12/07/2020 18:37:05",
          "content": "<p>I think you can try running same code on your personal computer (jupyter notebook) and check if same behavior persist, if not, then it means there might be a problem with kaggle notebooks. otherwise its an issue with pandas itself. if you find answers please share with me</p>\n<p>anyway, you can remove the .head() and avoid the issue as you mentioned in your other reply. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1051550": "hi,\nI am facing an issue related to memory usage of the kaggle notebook .\nwhen I load the train file:\n```\n train = pd.read_csv('file_path')\n```\nit occupies around 9.5 GB of the RAM.\nbut if I (for example) deleted the training DataFrame object directly after reading the csv file and use the garbage collector:\n```\ndel train\ngc.collect()\n```\nthe RAM is still 5.9 GB occupied although I cleared everything !\n\nwhy I still have things loaded into the RAM ? \nand how can I free up the RAM ? \nis there a way to know what is currently loaded into the kaggle RAM ?",
    "1052265": "I got the same problem. might be the system cache,\nespecially when you do lots of pd.merge. Now I do the training in my pc and test in the kaggle notebook.",
    "1053635": "I had the same kind of problems, I removed some .head() functions and *magic* got no more RAM issues. I cannot say why unfortunately",
    "1053873": "alexj21 \nbut I am not using any code, just reading csv file, and deleting the DataFrame directly after that.",
    "1101016": "I have this same problem. Did anyone find a solution?\n\nI use del df and then gc.collect(), but RAM remains the same.",
    "1101759": "no, I didn't find solutions for that, it seems as a bug in the kaggle notebook. \nbut I found some ways around the memory usage. for example, \n- use datatable library to load the data (and then convert to pandas). \n- reduce the datatypes of features, \n- deal with each feature and then move to other feature instead of all features at once which exhausts the memory \n- and some other ideas",
    "1102040": "Thank you!\n\nMy experience:\n* If I create a notebook from scratch then I have this issue, ie RAM does not get freed.\n* However, when I copy someone else's notebook I do see RAM get freed.\n\nI'm a beginner and don't understand why there is a difference.\n\nHopefully this helps someone else. I spent a ton of time on this.",
    "1102218": "Hi @mohamadnawfal! I faced similar issues and after a dozen of trial and error I discovered that doing a .head() causes this issue for me. Similar to @alexj21 mine disappeared after removing this one line of code. \n\nAnother fix that I read on stackoverflow which might work for you is to do a ```train = None``` after ```del train``` which apparantly helps de reference the variable and help gc to do its job. \n\nHope this helps :)",
    "1102236": "Thanks for sharing your experience :)",
    "1105214": "Naresh, Alex,\n\nI also see that removing the .head() solves the RAM issue. Is there an alternative to head that allows me to see the data while freeing RAM?\n\nThank you for any pointers!",
    "1105233": "I am kindly seeking help to explain this behavior:\n\nMy RAM is at 3.9G with full_train data loaded. I test 2 cases:\n\n(A) I run the following:\n\"full_train = full_train[['row_id','user_id','content_id','content_type_id','answered_correctly']]\"\nI see my RAM decrease from 3.9G to 1.7G as expected, since I have reduced columns.\n\n\n(B) I run \"full_train.head()\".\nMy RAM is still at 3.9G as expected.\nNow I run \"full_train = full_train[['row_id','user_id','content_id','content_type_id','answered_correctly']]\"\nMy RAM increases from 3.9G to 5.4G.\n\n\nWhy is this happening? \nI theorize that my notebook is effectively keeping a copy of full_train associated with the head() command.\nHow can I fix this? \n\nThank you in advance!",
    "1105298": "Work-around found:\n\nReplace \"full_train.head()\"\nwith \"print(full_train.head()\"\n\nBy adding the print() the RAM gets freed and everything behaves as expected, ie when I run \"del full_train; gc.collect()\" I can free the RAM associated with the deleted variable. Without the print, the RAM does not get freed with gc.collect()\n\nI do not understand why this happens, but I hope others can avoid the pain I've been through.\n\nIf anyone understands this behavior, I'd love an explanation!",
    "1105300": "I think you can try running same code on your personal computer (jupyter notebook) and check if same behavior persist, if not, then it means there might be a problem with kaggle notebooks. otherwise its an issue with pandas itself. if you find answers please share with me\n\nanyway, you can remove the .head() and avoid the issue as you mentioned in your other reply.",
    "1105790": "Hi @thomasliptay, I haven't been able to find any alternatives to .head().  I observed that doing operations on the data by itself doesn't cause this issue, only **displaying** the output does (using .head, .sample, etc). All the transformations that need to be done, could be written as a function, perhaps chain it all together. Test this \"pipeline\" on a small subset of the actual data to ensure it works as intended. Once you have ensured that it works good, restart the notebook and run again, this time *not displaying* any intermediate results (on the entire dataset). \n\nA good rule of thumb would be to keep the explorations separate from the actual modeling. Modeling is where you would require dealing with large amounts of data.  \n\nAnother *trick* that might assist you is the magic command: ```%reset -f```. This would reset the memory and all the variables you had created, including any module imports you might have done. The files written on the hard disk for that session would remain unaffected though. Using this command with caution! \n\nHope this helps! Have a nice day :)"
  },
  "source": "meta"
}