{
  "id": 336926,
  "title": "My notebook memory is just exploding to max 16 GB and restarts",
  "url": "/competitions/amex-default-prediction/discussion/336926",
  "author_name": "",
  "post_date": "2022-07-13T16:23:14.100524Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Only thing i have done so far is </p>\n<pre><code>%%time\ntrain_df = pd.read_parquet('../input/amex-data-integer-dtypes-parquet-format/train.parquet')\ntarget = pd.read_csv('../input/amex-default-prediction/train_labels.csv')\ni8 = [f for f in train_df.columns if train_df[f].dtype=='int8']\ni16= [f for f in train_df.columns if train_df[f].dtype=='int16']\nf32= [f for f in train_df.columns if train_df[f].dtype=='float32']\nobj =[f for f in train_df.columns if train_df[f].dtype=='object']\ntrain_df['customer_ID'].nunique()\n</code></pre>\n<p>Any clue why is this happening ?</p>",
  "messages": [
    {
      "id": "1854357",
      "postDate": "07/13/2022 16:23:14",
      "content": "<p>Only thing i have done so far is </p>\n<pre><code>%%time\ntrain_df = pd.read_parquet('../input/amex-data-integer-dtypes-parquet-format/train.parquet')\ntarget = pd.read_csv('../input/amex-default-prediction/train_labels.csv')\ni8 = [f for f in train_df.columns if train_df[f].dtype=='int8']\ni16= [f for f in train_df.columns if train_df[f].dtype=='int16']\nf32= [f for f in train_df.columns if train_df[f].dtype=='float32']\nobj =[f for f in train_df.columns if train_df[f].dtype=='object']\ntrain_df['customer_ID'].nunique()\n</code></pre>\n<p>Any clue why is this happening ?</p>",
      "rawMarkdown": "Only thing i have done so far is \n\n```\n%%time\ntrain_df = pd.read_parquet('../input/amex-data-integer-dtypes-parquet-format/train.parquet')\ntarget = pd.read_csv('../input/amex-default-prediction/train_labels.csv')\ni8 = [f for f in train_df.columns if train_df[f].dtype=='int8']\ni16= [f for f in train_df.columns if train_df[f].dtype=='int16']\nf32= [f for f in train_df.columns if train_df[f].dtype=='float32']\nobj =[f for f in train_df.columns if train_df[f].dtype=='object']\ntrain_df['customer_ID'].nunique()\n\n```\nAny clue why is this happening ?",
      "votes": null
    },
    {
      "id": "1854384",
      "postDate": "07/13/2022 16:53:35",
      "content": "<p>are you using Linux? you can increase the size of swap to fix this problem.</p>",
      "rawMarkdown": "are you using Linux? you can increase the size of swap to fix this problem.",
      "votes": null
    },
    {
      "id": "1854399",
      "postDate": "07/13/2022 17:03:08",
      "content": "<p><a href=\"https://www.kaggle.com/ptrikp\" target=\"_blank\">@ptrikp</a> Below approach to iterate columns is a culprit. Find an alternative to this. I am facing memory errors for similar line of code in a project of mine. Iterate through columns in a normal loop instead of below approach. 🙉🙈🙊</p>\n<p><a href=\"url\" target=\"_blank\">f for f in train_df.columns if train_df[f].dtype=='int8'</a></p>",
      "rawMarkdown": "ptrikp Below approach to iterate columns is a culprit. Find an alternative to this. I am facing memory errors for similar line of code in a project of mine. Iterate through columns in a normal loop instead of below approach. 🙉🙈🙊\n\n[f for f in train_df.columns if train_df[f].dtype=='int8'](url)",
      "votes": null
    },
    {
      "id": "1854431",
      "postDate": "07/13/2022 17:30:04",
      "content": "<p>You are hinting it is do with my operating system not kaggle environment ?</p>",
      "rawMarkdown": "You are hinting it is do with my operating system not kaggle environment ?",
      "votes": null
    },
    {
      "id": "1854435",
      "postDate": "07/13/2022 17:31:39",
      "content": "<p>Is it not a normal loop? No fancy stuffs just making a list. Are you sure its because of this ?</p>",
      "rawMarkdown": "Is it not a normal loop? No fancy stuffs just making a list. Are you sure its because of this ?",
      "votes": null
    },
    {
      "id": "1854482",
      "postDate": "07/13/2022 18:18:52",
      "content": "<p>I am banging my head for the same situation. 😄</p>\n<p>For example, beow code works: -<br>\n<code>for col in df.get_column_names():\n    print(col)</code></p>\n<p>But this crashes:-<br>\n<code>col: vaex.agg.first(col) for col in df.get_column_names()</code></p>\n<p>If you find a solution or alternative then please share 💕. If I find something I will share as well. 😎</p>",
      "rawMarkdown": "I am banging my head for the same situation. 😄\n\nFor example, beow code works: -\n`for col in df.get_column_names():\n    print(col)`\n\nBut this crashes:-\n`col: vaex.agg.first(col) for col in df.get_column_names()`\n\nIf you find a solution or alternative then please share 💕. If I find something I will share as well. 😎",
      "votes": null
    },
    {
      "id": "1856713",
      "postDate": "07/15/2022 14:30:57",
      "content": "<p>That loop is devil for you. <br>\nTry to import files one by one and <br>\nDo approch using .astype convert using pandas.<br>\nTry skimming through this once :<br>\n<a href=\"https://www.kaggle.com/code/sarang210/american-express-correlations\" target=\"_blank\">https://www.kaggle.com/code/sarang210/american-express-correlations</a></p>",
      "rawMarkdown": "That loop is devil for you. \nTry to import files one by one and \nDo approch using .astype convert using pandas.\nTry skimming through this once :\nhttps://www.kaggle.com/code/sarang210/american-express-correlations",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1854384,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/13/2022 16:53:35",
      "content": "<p>are you using Linux? you can increase the size of swap to fix this problem.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1854431,
          "author_name": "ptrikp",
          "author_url": "",
          "post_date": "07/13/2022 17:30:04",
          "content": "<p>You are hinting it is do with my operating system not kaggle environment ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1854399,
      "author_name": "mirfanazam",
      "author_url": "",
      "post_date": "07/13/2022 17:03:08",
      "content": "<p><a href=\"https://www.kaggle.com/ptrikp\" target=\"_blank\">@ptrikp</a> Below approach to iterate columns is a culprit. Find an alternative to this. I am facing memory errors for similar line of code in a project of mine. Iterate through columns in a normal loop instead of below approach. 🙉🙈🙊</p>\n<p><a href=\"url\" target=\"_blank\">f for f in train_df.columns if train_df[f].dtype=='int8'</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1854435,
          "author_name": "ptrikp",
          "author_url": "",
          "post_date": "07/13/2022 17:31:39",
          "content": "<p>Is it not a normal loop? No fancy stuffs just making a list. Are you sure its because of this ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1854482,
          "author_name": "mirfanazam",
          "author_url": "",
          "post_date": "07/13/2022 18:18:52",
          "content": "<p>I am banging my head for the same situation. 😄</p>\n<p>For example, beow code works: -<br>\n<code>for col in df.get_column_names():\n    print(col)</code></p>\n<p>But this crashes:-<br>\n<code>col: vaex.agg.first(col) for col in df.get_column_names()</code></p>\n<p>If you find a solution or alternative then please share 💕. If I find something I will share as well. 😎</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1856713,
      "author_name": "sarang210",
      "author_url": "",
      "post_date": "07/15/2022 14:30:57",
      "content": "<p>That loop is devil for you. <br>\nTry to import files one by one and <br>\nDo approch using .astype convert using pandas.<br>\nTry skimming through this once :<br>\n<a href=\"https://www.kaggle.com/code/sarang210/american-express-correlations\" target=\"_blank\">https://www.kaggle.com/code/sarang210/american-express-correlations</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1854357": "Only thing i have done so far is \n\n```\n%%time\ntrain_df = pd.read_parquet('../input/amex-data-integer-dtypes-parquet-format/train.parquet')\ntarget = pd.read_csv('../input/amex-default-prediction/train_labels.csv')\ni8 = [f for f in train_df.columns if train_df[f].dtype=='int8']\ni16= [f for f in train_df.columns if train_df[f].dtype=='int16']\nf32= [f for f in train_df.columns if train_df[f].dtype=='float32']\nobj =[f for f in train_df.columns if train_df[f].dtype=='object']\ntrain_df['customer_ID'].nunique()\n\n```\nAny clue why is this happening ?",
    "1854384": "are you using Linux? you can increase the size of swap to fix this problem.",
    "1854399": "ptrikp Below approach to iterate columns is a culprit. Find an alternative to this. I am facing memory errors for similar line of code in a project of mine. Iterate through columns in a normal loop instead of below approach. 🙉🙈🙊\n\n[f for f in train_df.columns if train_df[f].dtype=='int8'](url)",
    "1854431": "You are hinting it is do with my operating system not kaggle environment ?",
    "1854435": "Is it not a normal loop? No fancy stuffs just making a list. Are you sure its because of this ?",
    "1854482": "I am banging my head for the same situation. 😄\n\nFor example, beow code works: -\n`for col in df.get_column_names():\n    print(col)`\n\nBut this crashes:-\n`col: vaex.agg.first(col) for col in df.get_column_names()`\n\nIf you find a solution or alternative then please share 💕. If I find something I will share as well. 😎",
    "1856713": "That loop is devil for you. \nTry to import files one by one and \nDo approch using .astype convert using pandas.\nTry skimming through this once :\nhttps://www.kaggle.com/code/sarang210/american-express-correlations"
  },
  "source": "meta"
}