{
  "id": 224350,
  "title": "Preprocessing and padding",
  "url": "/competitions/bms-molecular-translation/discussion/224350",
  "author_name": "",
  "post_date": "2021-03-08T03:29:58.385241200Z",
  "votes": 4,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hello everyone - beginner here. I'm looking for some advice on preprocessing the image data so that the images are all the same size. I've tried padding with zeros based on maximum dimensions, but I'm running into a memory error from Kaggle. Does anyone have any suggestions or notebooks? Any help greatly appreciated and GLTA!</p>",
  "messages": [
    {
      "id": "1230342",
      "postDate": "03/08/2021 03:29:58",
      "content": "<p>Hello everyone - beginner here. I'm looking for some advice on preprocessing the image data so that the images are all the same size. I've tried padding with zeros based on maximum dimensions, but I'm running into a memory error from Kaggle. Does anyone have any suggestions or notebooks? Any help greatly appreciated and GLTA!</p>",
      "rawMarkdown": "Hello everyone - beginner here. I'm looking for some advice on preprocessing the image data so that the images are all the same size. I've tried padding with zeros based on maximum dimensions, but I'm running into a memory error from Kaggle. Does anyone have any suggestions or notebooks? Any help greatly appreciated and GLTA!",
      "votes": null
    },
    {
      "id": "1230344",
      "postDate": "03/08/2021 03:33:24",
      "content": "<p>I'm probably going to try 1024x1024 at most. Larger molecules will have to be scaled down.</p>",
      "rawMarkdown": "I'm probably going to try 1024x1024 at most. Larger molecules will have to be scaled down.",
      "votes": null
    },
    {
      "id": "1230360",
      "postDate": "03/08/2021 04:09:47",
      "content": "<p>Thanks for the comment. I suppose bigger is better especially for unknown test images. But do you have a systematic way of resetting size? </p>",
      "rawMarkdown": "Thanks for the comment. I suppose bigger is better especially for unknown test images. But do you have a systematic way of resetting size?",
      "votes": null
    },
    {
      "id": "1230364",
      "postDate": "03/08/2021 04:17:33",
      "content": "<p>I have not done this step yet. But it can probably be handled with albumentations.</p>",
      "rawMarkdown": "I have not done this step yet. But it can probably be handled with albumentations.",
      "votes": null
    },
    {
      "id": "1230365",
      "postDate": "03/08/2021 04:18:03",
      "content": "<p>So far I have just been using torchvision.transforms resize function. You can send your desired (H,W) and it will resize an image using some interpolation method. Documentation here: <a href=\"https://pytorch.org/vision/stable/transforms.html\" target=\"_blank\">https://pytorch.org/vision/stable/transforms.html</a>.</p>",
      "rawMarkdown": "So far I have just been using torchvision.transforms resize function. You can send your desired (H,W) and it will resize an image using some interpolation method. Documentation here: https://pytorch.org/vision/stable/transforms.html.",
      "votes": null
    },
    {
      "id": "1231145",
      "postDate": "03/08/2021 17:46:05",
      "content": "<p>You can use cv2.resize(image,(scale dim))<br>\nor PIL resize function  to just resize the image</p>",
      "rawMarkdown": "You can use cv2.resize(image,(scale dim))\nor PIL resize function  to just resize the image",
      "votes": null
    },
    {
      "id": "1231306",
      "postDate": "03/08/2021 21:20:43",
      "content": "<p>Thank you! I've tried this a bit ago and my notebook crashes with the same memory error. Do you have a small working example I could borrow?</p>",
      "rawMarkdown": "Thank you! I've tried this a bit ago and my notebook crashes with the same memory error. Do you have a small working example I could borrow?",
      "votes": null
    },
    {
      "id": "1231307",
      "postDate": "03/08/2021 21:20:55",
      "content": "<p>Thanks for the advice! I've tried this a bit ago and my notebook crashes with the same memory error. Do you have a small working example I could borrow?</p>",
      "rawMarkdown": "Thanks for the advice! I've tried this a bit ago and my notebook crashes with the same memory error. Do you have a small working example I could borrow?",
      "votes": null
    },
    {
      "id": "1231328",
      "postDate": "03/08/2021 22:02:14",
      "content": "<p>I have a notebook here that uses resizing: <a href=\"https://www.kaggle.com/pasewark/pytorch-resnet-lstm-with-attention\" target=\"_blank\">https://www.kaggle.com/pasewark/pytorch-resnet-lstm-with-attention</a>. The relevant lines of code are <br>\ntransform = Compose([<br>\n    Resize((256,256), PIL.Image.BICUBIC),<br>\n    ToTensor(),<br>\n    Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),<br>\n    ])</p>\n<p>This transform is then passed into a dataset object which uses the transform on images</p>",
      "rawMarkdown": "I have a notebook here that uses resizing: https://www.kaggle.com/pasewark/pytorch-resnet-lstm-with-attention. The relevant lines of code are \ntransform = Compose([\n    Resize((256,256), PIL.Image.BICUBIC),\n    ToTensor(),\n    Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n    ])\n\nThis transform is then passed into a dataset object which uses the transform on images",
      "votes": null
    },
    {
      "id": "1231508",
      "postDate": "03/09/2021 03:20:00",
      "content": "<p>You must be trying to put all that in an array which wud crash the memory you will need to make a generator for your model which will dynamically change a batch of images at a time instead of all the images</p>\n<p>If you just want a working example just try one image at a time <br>\nimport cv2<br>\nimport matplotlib.pyplot as plt<br>\nim=cv2.imread(img_path)<br>\nim=cv2.resize(im,(64,64))<br>\nplt.imshow(im)</p>\n<p>This is just a working example for one file</p>",
      "rawMarkdown": "You must be trying to put all that in an array which wud crash the memory you will need to make a generator for your model which will dynamically change a batch of images at a time instead of all the images\n\n\nIf you just want a working example just try one image at a time \nimport cv2\nimport matplotlib.pyplot as plt\nim=cv2.imread(img_path)\nim=cv2.resize(im,(64,64))\nplt.imshow(im)\n\nThis is just a working example for one file",
      "votes": null
    },
    {
      "id": "1231568",
      "postDate": "03/09/2021 04:39:38",
      "content": "<p>Thank you for the idea! I believe the generator concept is exactly what I'm looking for. But I'm still a bit confused on how then we would access the full training set when our model is ready?</p>",
      "rawMarkdown": "Thank you for the idea! I believe the generator concept is exactly what I'm looking for. But I'm still a bit confused on how then we would access the full training set when our model is ready?",
      "votes": null
    },
    {
      "id": "1231571",
      "postDate": "03/09/2021 04:43:49",
      "content": "<p>once the model is trained you can predict normal images .</p>",
      "rawMarkdown": "once the model is trained you can predict normal images .",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1230344,
      "author_name": "matthewmasters",
      "author_url": "",
      "post_date": "03/08/2021 03:33:24",
      "content": "<p>I'm probably going to try 1024x1024 at most. Larger molecules will have to be scaled down.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1230360,
          "author_name": "jacobadamczyk",
          "author_url": "",
          "post_date": "03/08/2021 04:09:47",
          "content": "<p>Thanks for the comment. I suppose bigger is better especially for unknown test images. But do you have a systematic way of resetting size? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1230364,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "03/08/2021 04:17:33",
          "content": "<p>I have not done this step yet. But it can probably be handled with albumentations.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231145,
          "author_name": "accountstatus",
          "author_url": "",
          "post_date": "03/08/2021 17:46:05",
          "content": "<p>You can use cv2.resize(image,(scale dim))<br>\nor PIL resize function  to just resize the image</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231307,
          "author_name": "jacobadamczyk",
          "author_url": "",
          "post_date": "03/08/2021 21:20:55",
          "content": "<p>Thanks for the advice! I've tried this a bit ago and my notebook crashes with the same memory error. Do you have a small working example I could borrow?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231508,
          "author_name": "accountstatus",
          "author_url": "",
          "post_date": "03/09/2021 03:20:00",
          "content": "<p>You must be trying to put all that in an array which wud crash the memory you will need to make a generator for your model which will dynamically change a batch of images at a time instead of all the images</p>\n<p>If you just want a working example just try one image at a time <br>\nimport cv2<br>\nimport matplotlib.pyplot as plt<br>\nim=cv2.imread(img_path)<br>\nim=cv2.resize(im,(64,64))<br>\nplt.imshow(im)</p>\n<p>This is just a working example for one file</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231568,
          "author_name": "jacobadamczyk",
          "author_url": "",
          "post_date": "03/09/2021 04:39:38",
          "content": "<p>Thank you for the idea! I believe the generator concept is exactly what I'm looking for. But I'm still a bit confused on how then we would access the full training set when our model is ready?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231571,
          "author_name": "accountstatus",
          "author_url": "",
          "post_date": "03/09/2021 04:43:49",
          "content": "<p>once the model is trained you can predict normal images .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1230365,
      "author_name": "pasewark",
      "author_url": "",
      "post_date": "03/08/2021 04:18:03",
      "content": "<p>So far I have just been using torchvision.transforms resize function. You can send your desired (H,W) and it will resize an image using some interpolation method. Documentation here: <a href=\"https://pytorch.org/vision/stable/transforms.html\" target=\"_blank\">https://pytorch.org/vision/stable/transforms.html</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1231306,
          "author_name": "jacobadamczyk",
          "author_url": "",
          "post_date": "03/08/2021 21:20:43",
          "content": "<p>Thank you! I've tried this a bit ago and my notebook crashes with the same memory error. Do you have a small working example I could borrow?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1231328,
          "author_name": "pasewark",
          "author_url": "",
          "post_date": "03/08/2021 22:02:14",
          "content": "<p>I have a notebook here that uses resizing: <a href=\"https://www.kaggle.com/pasewark/pytorch-resnet-lstm-with-attention\" target=\"_blank\">https://www.kaggle.com/pasewark/pytorch-resnet-lstm-with-attention</a>. The relevant lines of code are <br>\ntransform = Compose([<br>\n    Resize((256,256), PIL.Image.BICUBIC),<br>\n    ToTensor(),<br>\n    Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),<br>\n    ])</p>\n<p>This transform is then passed into a dataset object which uses the transform on images</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1230342": "Hello everyone - beginner here. I'm looking for some advice on preprocessing the image data so that the images are all the same size. I've tried padding with zeros based on maximum dimensions, but I'm running into a memory error from Kaggle. Does anyone have any suggestions or notebooks? Any help greatly appreciated and GLTA!",
    "1230344": "I'm probably going to try 1024x1024 at most. Larger molecules will have to be scaled down.",
    "1230360": "Thanks for the comment. I suppose bigger is better especially for unknown test images. But do you have a systematic way of resetting size?",
    "1230364": "I have not done this step yet. But it can probably be handled with albumentations.",
    "1230365": "So far I have just been using torchvision.transforms resize function. You can send your desired (H,W) and it will resize an image using some interpolation method. Documentation here: https://pytorch.org/vision/stable/transforms.html.",
    "1231145": "You can use cv2.resize(image,(scale dim))\nor PIL resize function  to just resize the image",
    "1231306": "Thank you! I've tried this a bit ago and my notebook crashes with the same memory error. Do you have a small working example I could borrow?",
    "1231307": "Thanks for the advice! I've tried this a bit ago and my notebook crashes with the same memory error. Do you have a small working example I could borrow?",
    "1231328": "I have a notebook here that uses resizing: https://www.kaggle.com/pasewark/pytorch-resnet-lstm-with-attention. The relevant lines of code are \ntransform = Compose([\n    Resize((256,256), PIL.Image.BICUBIC),\n    ToTensor(),\n    Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n    ])\n\nThis transform is then passed into a dataset object which uses the transform on images",
    "1231508": "You must be trying to put all that in an array which wud crash the memory you will need to make a generator for your model which will dynamically change a batch of images at a time instead of all the images\n\n\nIf you just want a working example just try one image at a time \nimport cv2\nimport matplotlib.pyplot as plt\nim=cv2.imread(img_path)\nim=cv2.resize(im,(64,64))\nplt.imshow(im)\n\nThis is just a working example for one file",
    "1231568": "Thank you for the idea! I believe the generator concept is exactly what I'm looking for. But I'm still a bit confused on how then we would access the full training set when our model is ready?",
    "1231571": "once the model is trained you can predict normal images ."
  },
  "source": "meta"
}