{
  "id": 312649,
  "title": "Memory error while resizing image!!",
  "url": "/competitions/ultra-mnist/discussion/312649",
  "author_name": "",
  "post_date": "2022-03-13T09:02:25.381314900Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hey, folks, I am new to the image classification problem. I am getting a memory error while resizing the image into 255x255 pixels.</p>\n<p>Please provide me with an efficient method to do this.</p>",
  "messages": [
    {
      "id": "1720907",
      "postDate": "03/13/2022 09:02:25",
      "content": "<p>Hey, folks, I am new to the image classification problem. I am getting a memory error while resizing the image into 255x255 pixels.</p>\n<p>Please provide me with an efficient method to do this.</p>",
      "rawMarkdown": "Hey, folks, I am new to the image classification problem. I am getting a memory error while resizing the image into 255x255 pixels.\n\nPlease provide me with an efficient method to do this.",
      "votes": null
    },
    {
      "id": "1720917",
      "postDate": "03/13/2022 09:09:12",
      "content": "<p><a href=\"https://www.kaggle.com/mohammadkashifunique\" target=\"_blank\">@mohammadkashifunique</a> by any chance, are you trying to load several images in the memory? Since the images are very large, load one, convert it, and save it to the HD before working with the next one. I hope this helps. Else, could you elaborate more on the error?</p>",
      "rawMarkdown": "mohammadkashifunique by any chance, are you trying to load several images in the memory? Since the images are very large, load one, convert it, and save it to the HD before working with the next one. I hope this helps. Else, could you elaborate more on the error?",
      "votes": null
    },
    {
      "id": "1720973",
      "postDate": "03/13/2022 09:53:34",
      "content": "<p>`def create_dataset(training_df, image_dir):</p>\n<pre><code>images = []\ntargets = []\n\nfor index, row in tqdm(\n    training_df.iterrows(),\n    total=len(training_df),\n    desc=\"processing images\"\n):\n    image_id = row[\"id\"]\n    image_path = os.path.join(image_dir, image_id)\n    image = Image.open(image_path + \".jpeg\")\n    image = image.resize((256, 256), resample=Image.BILINEAR)\n    image = np.array(image)\n    image = image.ravel()\n    images.append(image)\n    targets.append(int(row['digit_sum']))\n\nimages = np.array(images)\nprint(images.shape)\nreturn images, targets`\n</code></pre>\n<p>`for fold_ in range(5):<br>\n    train_df = df[df.kfold != fold_].reset_index(drop=True)<br>\n    test_df  = df[df.kfold == fold_].reset_index(drop=True)</p>\n<pre><code>x_train, y_train = create_dataset(train_df, train_image_path)\nx_test, y_test   = create_dataset(test_df, train_image_path)\n\nclf = ensemble.RandomForestClassifier(n_jobs=-1)\nclf.fit(x_train, y_train)\n\npreds = clf.predict(x_test)\n\nprint(f\"FOLD: {fold_}\")\nprint(f\"ACCURACY = {metrics.accuracy_score(y_test, preds)}\")\nprint(\"\")`\n</code></pre>",
      "rawMarkdown": "`def create_dataset(training_df, image_dir):\n    \n    images = []\n    targets = []\n    \n    for index, row in tqdm(\n        training_df.iterrows(),\n        total=len(training_df),\n        desc=\"processing images\"\n    ):\n        image_id = row[\"id\"]\n        image_path = os.path.join(image_dir, image_id)\n        image = Image.open(image_path + \".jpeg\")\n        image = image.resize((256, 256), resample=Image.BILINEAR)\n        image = np.array(image)\n        image = image.ravel()\n        images.append(image)\n        targets.append(int(row['digit_sum']))\n        \n    images = np.array(images)\n    print(images.shape)\n    return images, targets`\n\n\n`for fold_ in range(5):\n    train_df = df[df.kfold != fold_].reset_index(drop=True)\n    test_df  = df[df.kfold == fold_].reset_index(drop=True)\n    \n    x_train, y_train = create_dataset(train_df, train_image_path)\n    x_test, y_test   = create_dataset(test_df, train_image_path)\n    \n    clf = ensemble.RandomForestClassifier(n_jobs=-1)\n    clf.fit(x_train, y_train)\n    \n    preds = clf.predict(x_test)\n    \n    print(f\"FOLD: {fold_}\")\n    print(f\"ACCURACY = {metrics.accuracy_score(y_test, preds)}\")\n    print(\"\")`",
      "votes": null
    },
    {
      "id": "1720977",
      "postDate": "03/13/2022 09:54:47",
      "content": "<p><a href=\"https://www.kaggle.com/dkgupta90\" target=\"_blank\">@dkgupta90</a> <br>\nThis is the code I am using for resizing images.</p>",
      "rawMarkdown": "dkgupta90 \nThis is the code I am using for resizing images.",
      "votes": null
    },
    {
      "id": "1720981",
      "postDate": "03/13/2022 10:01:08",
      "content": "<p>How many rows are there? And are these the 4000X4000 images that you are resizing?</p>",
      "rawMarkdown": "How many rows are there? And are these the 4000X4000 images that you are resizing?",
      "votes": null
    },
    {
      "id": "1720984",
      "postDate": "03/13/2022 10:08:45",
      "content": "<p>5600 for fold and 22400 for out of the fold. Yes, I am resizing 4000X400 images.</p>",
      "rawMarkdown": "5600 for fold and 22400 for out of the fold. Yes, I am resizing 4000X400 images.",
      "votes": null
    },
    {
      "id": "1720988",
      "postDate": "03/13/2022 10:17:38",
      "content": "<p>So if I understand correctly, you are loading all these images in the RAM at once which is not possible. Alternatively, I recommend to first resize all the images one by one. Load one image, resize it and then save it to the HD. Then move to the next.</p>",
      "rawMarkdown": "So if I understand correctly, you are loading all these images in the RAM at once which is not possible. Alternatively, I recommend to first resize all the images one by one. Load one image, resize it and then save it to the HD. Then move to the next.",
      "votes": null
    },
    {
      "id": "1721383",
      "postDate": "03/13/2022 16:28:37",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/dkgupta90\" target=\"_blank\">@dkgupta90</a> </p>",
      "rawMarkdown": "Thanks @dkgupta90",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1720917,
      "author_name": "dkgupta90",
      "author_url": "",
      "post_date": "03/13/2022 09:09:12",
      "content": "<p><a href=\"https://www.kaggle.com/mohammadkashifunique\" target=\"_blank\">@mohammadkashifunique</a> by any chance, are you trying to load several images in the memory? Since the images are very large, load one, convert it, and save it to the HD before working with the next one. I hope this helps. Else, could you elaborate more on the error?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1720973,
          "author_name": "mohammadkashifunique",
          "author_url": "",
          "post_date": "03/13/2022 09:53:34",
          "content": "<p>`def create_dataset(training_df, image_dir):</p>\n<pre><code>images = []\ntargets = []\n\nfor index, row in tqdm(\n    training_df.iterrows(),\n    total=len(training_df),\n    desc=\"processing images\"\n):\n    image_id = row[\"id\"]\n    image_path = os.path.join(image_dir, image_id)\n    image = Image.open(image_path + \".jpeg\")\n    image = image.resize((256, 256), resample=Image.BILINEAR)\n    image = np.array(image)\n    image = image.ravel()\n    images.append(image)\n    targets.append(int(row['digit_sum']))\n\nimages = np.array(images)\nprint(images.shape)\nreturn images, targets`\n</code></pre>\n<p>`for fold_ in range(5):<br>\n    train_df = df[df.kfold != fold_].reset_index(drop=True)<br>\n    test_df  = df[df.kfold == fold_].reset_index(drop=True)</p>\n<pre><code>x_train, y_train = create_dataset(train_df, train_image_path)\nx_test, y_test   = create_dataset(test_df, train_image_path)\n\nclf = ensemble.RandomForestClassifier(n_jobs=-1)\nclf.fit(x_train, y_train)\n\npreds = clf.predict(x_test)\n\nprint(f\"FOLD: {fold_}\")\nprint(f\"ACCURACY = {metrics.accuracy_score(y_test, preds)}\")\nprint(\"\")`\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1720977,
          "author_name": "mohammadkashifunique",
          "author_url": "",
          "post_date": "03/13/2022 09:54:47",
          "content": "<p><a href=\"https://www.kaggle.com/dkgupta90\" target=\"_blank\">@dkgupta90</a> <br>\nThis is the code I am using for resizing images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1720981,
          "author_name": "dkgupta90",
          "author_url": "",
          "post_date": "03/13/2022 10:01:08",
          "content": "<p>How many rows are there? And are these the 4000X4000 images that you are resizing?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1720984,
          "author_name": "mohammadkashifunique",
          "author_url": "",
          "post_date": "03/13/2022 10:08:45",
          "content": "<p>5600 for fold and 22400 for out of the fold. Yes, I am resizing 4000X400 images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1720988,
          "author_name": "dkgupta90",
          "author_url": "",
          "post_date": "03/13/2022 10:17:38",
          "content": "<p>So if I understand correctly, you are loading all these images in the RAM at once which is not possible. Alternatively, I recommend to first resize all the images one by one. Load one image, resize it and then save it to the HD. Then move to the next.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1721383,
          "author_name": "mohammadkashifunique",
          "author_url": "",
          "post_date": "03/13/2022 16:28:37",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/dkgupta90\" target=\"_blank\">@dkgupta90</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1720907": "Hey, folks, I am new to the image classification problem. I am getting a memory error while resizing the image into 255x255 pixels.\n\nPlease provide me with an efficient method to do this.",
    "1720917": "mohammadkashifunique by any chance, are you trying to load several images in the memory? Since the images are very large, load one, convert it, and save it to the HD before working with the next one. I hope this helps. Else, could you elaborate more on the error?",
    "1720973": "`def create_dataset(training_df, image_dir):\n    \n    images = []\n    targets = []\n    \n    for index, row in tqdm(\n        training_df.iterrows(),\n        total=len(training_df),\n        desc=\"processing images\"\n    ):\n        image_id = row[\"id\"]\n        image_path = os.path.join(image_dir, image_id)\n        image = Image.open(image_path + \".jpeg\")\n        image = image.resize((256, 256), resample=Image.BILINEAR)\n        image = np.array(image)\n        image = image.ravel()\n        images.append(image)\n        targets.append(int(row['digit_sum']))\n        \n    images = np.array(images)\n    print(images.shape)\n    return images, targets`\n\n\n`for fold_ in range(5):\n    train_df = df[df.kfold != fold_].reset_index(drop=True)\n    test_df  = df[df.kfold == fold_].reset_index(drop=True)\n    \n    x_train, y_train = create_dataset(train_df, train_image_path)\n    x_test, y_test   = create_dataset(test_df, train_image_path)\n    \n    clf = ensemble.RandomForestClassifier(n_jobs=-1)\n    clf.fit(x_train, y_train)\n    \n    preds = clf.predict(x_test)\n    \n    print(f\"FOLD: {fold_}\")\n    print(f\"ACCURACY = {metrics.accuracy_score(y_test, preds)}\")\n    print(\"\")`",
    "1720977": "dkgupta90 \nThis is the code I am using for resizing images.",
    "1720981": "How many rows are there? And are these the 4000X4000 images that you are resizing?",
    "1720984": "5600 for fold and 22400 for out of the fold. Yes, I am resizing 4000X400 images.",
    "1720988": "So if I understand correctly, you are loading all these images in the RAM at once which is not possible. Alternatively, I recommend to first resize all the images one by one. Load one image, resize it and then save it to the HD. Then move to the next.",
    "1721383": "Thanks @dkgupta90"
  },
  "source": "meta"
}