{
  "id": 154568,
  "title": "Resized Images",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/154568",
  "author_name": "Bojan Tunguz",
  "post_date": "2020-05-28T21:59:10.477000",
  "votes": 72,
  "comment_count": 34,
  "views": 0,
  "content": "<p>Since the images in this competition are very large (1024x1024), it will probably be very useful for most of us to first experiment with smaller sizes. I decided to create numpy arrays of the smaller-resized images and upload them to a separate dataset. It can be found here: </p>\n\n<p><a href=\"https://www.kaggle.com/tunguz/siimisic-melanoma-resized-images\">https://www.kaggle.com/tunguz/siimisic-melanoma-resized-images</a></p>\n\n<p>All resizings were done in kernels, and can be found below:</p>\n\n<p><a href=\"https://www.kaggle.com/tunguz/image-resizing-32x32-train\">https://www.kaggle.com/tunguz/image-resizing-32x32-train</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-32x32-test\">https://www.kaggle.com/tunguz/image-resizing-32x32-test</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-64x-train\">https://www.kaggle.com/tunguz/image-resizing-64x-train</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-64x64-test\">https://www.kaggle.com/tunguz/image-resizing-64x64-test</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-96x96-train\">https://www.kaggle.com/tunguz/image-resizing-96x96-train</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-96x96-test\">https://www.kaggle.com/tunguz/image-resizing-96x96-test</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-128x128-train\">https://www.kaggle.com/tunguz/image-resizing-128x128-train</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-128x128-test\">https://www.kaggle.com/tunguz/image-resizing-128x128-test</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-224x224-train\">https://www.kaggle.com/tunguz/image-resizing-224x224-train</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-224x224-test\">https://www.kaggle.com/tunguz/image-resizing-224x224-test</a></p>\n\n<p><strong>* Update 5-29-2020 *</strong></p>\n\n<p>I have added 96x96 and 224x224 images as well.</p>",
  "messages": [
    {
      "id": 865785,
      "postDate": "2020-05-28T21:59:10.477Z",
      "content": "<p>Since the images in this competition are very large (1024x1024), it will probably be very useful for most of us to first experiment with smaller sizes. I decided to create numpy arrays of the smaller-resized images and upload them to a separate dataset. It can be found here: </p>\n\n<p><a href=\"https://www.kaggle.com/tunguz/siimisic-melanoma-resized-images\">https://www.kaggle.com/tunguz/siimisic-melanoma-resized-images</a></p>\n\n<p>All resizings were done in kernels, and can be found below:</p>\n\n<p><a href=\"https://www.kaggle.com/tunguz/image-resizing-32x32-train\">https://www.kaggle.com/tunguz/image-resizing-32x32-train</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-32x32-test\">https://www.kaggle.com/tunguz/image-resizing-32x32-test</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-64x-train\">https://www.kaggle.com/tunguz/image-resizing-64x-train</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-64x64-test\">https://www.kaggle.com/tunguz/image-resizing-64x64-test</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-96x96-train\">https://www.kaggle.com/tunguz/image-resizing-96x96-train</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-96x96-test\">https://www.kaggle.com/tunguz/image-resizing-96x96-test</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-128x128-train\">https://www.kaggle.com/tunguz/image-resizing-128x128-train</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-128x128-test\">https://www.kaggle.com/tunguz/image-resizing-128x128-test</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-224x224-train\">https://www.kaggle.com/tunguz/image-resizing-224x224-train</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-224x224-test\">https://www.kaggle.com/tunguz/image-resizing-224x224-test</a></p>\n\n<p><strong>* Update 5-29-2020 *</strong></p>\n\n<p>I have added 96x96 and 224x224 images as well.</p>",
      "rawMarkdown": "Since the images in this competition are very large (1024x1024), it will probably be very useful for most of us to first experiment with smaller sizes. I decided to create numpy arrays of the smaller-resized images and upload them to a separate dataset. It can be found here: \n\nhttps://www.kaggle.com/tunguz/siimisic-melanoma-resized-images\n\nAll resizings were done in kernels, and can be found below:\n\nhttps://www.kaggle.com/tunguz/image-resizing-32x32-train\nhttps://www.kaggle.com/tunguz/image-resizing-32x32-test\nhttps://www.kaggle.com/tunguz/image-resizing-64x-train\nhttps://www.kaggle.com/tunguz/image-resizing-64x64-test\nhttps://www.kaggle.com/tunguz/image-resizing-96x96-train\nhttps://www.kaggle.com/tunguz/image-resizing-96x96-test\nhttps://www.kaggle.com/tunguz/image-resizing-128x128-train\nhttps://www.kaggle.com/tunguz/image-resizing-128x128-test\nhttps://www.kaggle.com/tunguz/image-resizing-224x224-train\nhttps://www.kaggle.com/tunguz/image-resizing-224x224-test\n\n*** Update 5-29-2020 ***\n\nI have added 96x96 and 224x224 images as well.",
      "votes": 70
    },
    {
      "id": 865821,
      "postDate": "2020-05-28T23:06:57.387Z",
      "content": "<p>Thanks. Note I've done 224x224 here: <a href=\"https://www.kaggle.com/arroqc/siic-isic-224x224-images\">https://www.kaggle.com/arroqc/siic-isic-224x224-images</a></p>",
      "rawMarkdown": "Thanks. Note I've done 224x224 here: https://www.kaggle.com/arroqc/siic-isic-224x224-images",
      "votes": 9,
      "replies": [
        {
          "id": 865822,
          "postDate": "2020-05-28T23:08:29.203Z",
          "content": "<p>Thanks, good to know - I will be adding my own version anyway soon, the script is almost finished running. </p>",
          "rawMarkdown": "Thanks, good to know - I will be adding my own version anyway soon, the script is almost finished running. ",
          "votes": 1
        },
        {
          "id": 866400,
          "postDate": "2020-05-29T11:14:09.420Z",
          "content": "<p>And here is my version of the rescaled 224x224 images: </p>\n\n<p><a href=\"https://www.kaggle.com/tunguz/image-resizing-224x224-train/\">https://www.kaggle.com/tunguz/image-resizing-224x224-train/</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-224x224-test\">https://www.kaggle.com/tunguz/image-resizing-224x224-test</a></p>",
          "rawMarkdown": "And here is my version of the rescaled 224x224 images: \n\nhttps://www.kaggle.com/tunguz/image-resizing-224x224-train/\nhttps://www.kaggle.com/tunguz/image-resizing-224x224-test"
        },
        {
          "id": 880433,
          "postDate": "2020-06-10T09:46:21.293Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 865847,
      "postDate": "2020-05-28T23:53:18.087Z",
      "content": "<p>Just a simple logistic regression model on flattened 32x32 images can give you LB score of 0.78 AUC:</p>\n\n<p><a href=\"https://www.kaggle.com/tunguz/melanoma-classification-eda-and-modeling\">https://www.kaggle.com/tunguz/melanoma-classification-eda-and-modeling</a></p>",
      "rawMarkdown": "Just a simple logistic regression model on flattened 32x32 images can give you LB score of 0.78 AUC:\n\nhttps://www.kaggle.com/tunguz/melanoma-classification-eda-and-modeling",
      "votes": 3,
      "replies": [
        {
          "id": 866772,
          "postDate": "2020-05-29T16:57:15.733Z",
          "content": "<p>Interesting. Is logistic regression only using the flattened images to score 0.78, or are you including the extra features such as age, gender, etc?</p>",
          "rawMarkdown": "Interesting. Is logistic regression only using the flattened images to score 0.78, or are you including the extra features such as age, gender, etc?",
          "votes": 2
        },
        {
          "id": 866813,
          "postDate": "2020-05-29T17:35:35.787Z",
          "content": "<p>You can reach 0.78 with only LR and flattened 32x32 images. I am now including those other features that you mentioned and assembling, and I am getting 0.793 in the latest edition of the kernel. I beleive I can reach 0.8 with just LR and these features. :) </p>",
          "rawMarkdown": "You can reach 0.78 with only LR and flattened 32x32 images. I am now including those other features that you mentioned and assembling, and I am getting 0.793 in the latest edition of the kernel. I beleive I can reach 0.8 with just LR and these features. :) ",
          "votes": 1
        },
        {
          "id": 872118,
          "postDate": "2020-06-02T23:40:21.207Z",
          "content": "<p>I can confirm what <a href=\"/tunguz\">@tunguz</a> said. I can achieve 0.8134 (10-folds CV) with just a simple LR model on flattened 32x32 images and the \"contextual\" features.  </p>",
          "rawMarkdown": "I can confirm what @tunguz said. I can achieve 0.8134 (10-folds CV) with just a simple LR model on flattened 32x32 images and the \"contextual\" features.  ",
          "votes": 1
        }
      ]
    },
    {
      "id": 865801,
      "postDate": "2020-05-28T22:34:09.427Z",
      "content": "<p>Thanks Bojan</p>",
      "rawMarkdown": "Thanks Bojan",
      "votes": 3,
      "replies": [
        {
          "id": 865803,
          "postDate": "2020-05-28T22:36:23.850Z",
          "content": "<p>Ah, don't mention it. </p>\n\n<p><img src=\"https://pbs.twimg.com/media/EZIy_xiWAAEAM-i?format=png&amp;name=900x900\" alt=\"\"></p>",
          "rawMarkdown": "Ah, don't mention it. \n\n![](https://pbs.twimg.com/media/EZIy_xiWAAEAM-i?format=png&amp;name=900x900)",
          "votes": 15
        }
      ]
    },
    {
      "id": 872158,
      "postDate": "2020-06-03T00:41:54.187Z",
      "content": "<p>Thank you for sharing your notebook <a href=\"/tunguz\">@tunguz</a> </p>\n\n<p>I've refactored and used multiprocessing to take advantage of all cores in my machine. I don't like much working in Kaggle notebooks so I'm attaching it here in case anyone is interested.</p>\n\n<p>```\nfrom functools import partial\nfrom multiprocessing import Pool</p>\n\n<p>import numpy as np\nimport pandas as pd\nfrom PIL import Image</p>\n\n<p>from siim_isic_melanoma_classification.constants import (\n    train_fpath,\n    test_fpath,\n    data_path,\n)</p>\n\n<p>IMAGE_SZ = 32</p>\n\n<p>def main():\n    train_df = pd.read_csv(train_fpath)\n    test_df = pd.read_csv(test_fpath)</p>\n\n<pre><code>for df, name in [(train_df, \"train\"), (test_df, \"test\")]:\n    print(f\"Converting {name} images to NumPy array...\")\n    ar = convert_to_array(df, name=name)\n    np.save(data_path / f\"x_{name}_32\", ar)\n</code></pre>\n\n<p>def convert_to_array(df, name: str):\n    pool = Pool()\n    routine = partial(image_to_flattened_array, name=name)\n    data = pool.map(routine, df[\"image_name\"])\n    return np.vstack(data)</p>\n\n<p>def image_to_flattened_array(image_id, name: str, desired_size=IMAGE_SZ):\n    image_fpath = data_path / f\"{name}/{image_id}.jpg\"\n    im = Image.open(image_fpath)\n    small_im = im.resize((desired_size,) * 2, resample=Image.LANCZOS)\n    return np.array(small_im).reshape((1, 32, 32, 3))</p>\n\n<p>if <strong>name</strong> == \"<strong>main</strong>\":\n    main()\n```</p>\n\n<p>It could be further refactored but it does the job :-)</p>",
      "rawMarkdown": "Thank you for sharing your notebook @tunguz \n\nI've refactored and used multiprocessing to take advantage of all cores in my machine. I don't like much working in Kaggle notebooks so I'm attaching it here in case anyone is interested.\n\n```\nfrom functools import partial\nfrom multiprocessing import Pool\n\nimport numpy as np\nimport pandas as pd\nfrom PIL import Image\n\nfrom siim_isic_melanoma_classification.constants import (\n    train_fpath,\n    test_fpath,\n    data_path,\n)\n\nIMAGE_SZ = 32\n\n\ndef main():\n    train_df = pd.read_csv(train_fpath)\n    test_df = pd.read_csv(test_fpath)\n\n    for df, name in [(train_df, \"train\"), (test_df, \"test\")]:\n        print(f\"Converting {name} images to NumPy array...\")\n        ar = convert_to_array(df, name=name)\n        np.save(data_path / f\"x_{name}_32\", ar)\n\n\ndef convert_to_array(df, name: str):\n    pool = Pool()\n    routine = partial(image_to_flattened_array, name=name)\n    data = pool.map(routine, df[\"image_name\"])\n    return np.vstack(data)\n\n\ndef image_to_flattened_array(image_id, name: str, desired_size=IMAGE_SZ):\n    image_fpath = data_path / f\"{name}/{image_id}.jpg\"\n    im = Image.open(image_fpath)\n    small_im = im.resize((desired_size,) * 2, resample=Image.LANCZOS)\n    return np.array(small_im).reshape((1, 32, 32, 3))\n\n\nif __name__ == \"__main__\":\n    main()\n```\n\nIt could be further refactored but it does the job :-)",
      "votes": 4
    },
    {
      "id": 865820,
      "postDate": "2020-05-28T23:06:03.220Z",
      "content": "<p>npy arrays for 512 x 512 would be also nice, not for the kernels use, but in general </p>",
      "rawMarkdown": "npy arrays for 512 x 512 would be also nice, not for the kernels use, but in general ",
      "votes": 4,
      "replies": [
        {
          "id": 865823,
          "postDate": "2020-05-28T23:08:59.960Z",
          "content": "<p>I will probably make those offline and upload them to the dataset mentioned above.</p>",
          "rawMarkdown": "I will probably make those offline and upload them to the dataset mentioned above.",
          "votes": 4
        },
        {
          "id": 886812,
          "postDate": "2020-06-15T09:44:18.050Z",
          "content": "<p>That would be great! lookin  forward for it</p>",
          "rawMarkdown": "That would be great! lookin  forward for it"
        }
      ]
    },
    {
      "id": 875616,
      "postDate": "2020-06-06T02:39:35.983Z",
      "content": "<p>Thanks <a href=\"/tunguz\">@tunguz</a> it is awesome you helping others to experiment and learn...Best</p>",
      "rawMarkdown": "Thanks @tunguz it is awesome you helping others to experiment and learn...Best",
      "votes": 1
    },
    {
      "id": 872432,
      "postDate": "2020-06-03T07:27:20.527Z",
      "content": "<p>Thanks <a href=\"/tunguz\">@tunguz</a>  for all these resized dataset. Hoping for 512x512 soon.</p>",
      "rawMarkdown": "Thanks @tunguz  for all these resized dataset. Hoping for 512x512 soon.",
      "votes": 1
    },
    {
      "id": 867534,
      "postDate": "2020-05-30T11:53:11.470Z",
      "content": "<p><a href=\"/tunguz\">@tunguz</a> hello, thanks for sharing. One thing, in <code>jpeg</code> , all images are not in uniform resolution. Do you think plain resize is OK or should it be scale down by a factor? </p>",
      "rawMarkdown": "@tunguz hello, thanks for sharing. One thing, in `jpeg` , all images are not in uniform resolution. Do you think plain resize is OK or should it be scale down by a factor? ",
      "votes": 1,
      "replies": [
        {
          "id": 868010,
          "postDate": "2020-05-30T20:14:50.950Z",
          "content": "<p>For the most part that should not be a major obstacle. It is a standard practice in most computer vision challenges to resize images ot the same fixed size.</p>",
          "rawMarkdown": "For the most part that should not be a major obstacle. It is a standard practice in most computer vision challenges to resize images ot the same fixed size.",
          "votes": 1
        },
        {
          "id": 868019,
          "postDate": "2020-05-30T20:23:23.777Z",
          "content": "<p>Yes, I know and I agree. Actually it reminds me of the bengali.ai competition. That time In my cases, using a non-square image (original shape) gave me promising scores.</p>",
          "rawMarkdown": "Yes, I know and I agree. Actually it reminds me of the bengali.ai competition. That time In my cases, using a non-square image (original shape) gave me promising scores."
        }
      ]
    },
    {
      "id": 870182,
      "postDate": "2020-06-01T14:43:45.707Z",
      "content": "<p>Thank a lot for this, people at Kaggle are best!</p>",
      "rawMarkdown": "Thank a lot for this, people at Kaggle are best!",
      "votes": 2
    },
    {
      "id": 867995,
      "postDate": "2020-05-30T20:04:21.620Z",
      "content": "<p><a href=\"/tunguz\">@tunguz</a> Thankyou! That will be useful.</p>",
      "rawMarkdown": "@tunguz Thankyou! That will be useful.",
      "votes": 2
    },
    {
      "id": 928032,
      "postDate": "2020-07-13T17:46:59.390Z",
      "content": "<p>This is so helpful! Thanks for doing this</p>",
      "rawMarkdown": "This is so helpful! Thanks for doing this"
    },
    {
      "id": 893938,
      "postDate": "2020-06-20T04:55:20.703Z",
      "content": "<p>good one chris</p>",
      "rawMarkdown": "good one chris"
    },
    {
      "id": 885628,
      "postDate": "2020-06-14T10:58:25.773Z",
      "content": "<p>It's my first competition and was wondering how I am supposed to work with 100GB+ data and found this. Thank you so much!</p>",
      "rawMarkdown": "It's my first competition and was wondering how I am supposed to work with 100GB+ data and found this. Thank you so much!"
    },
    {
      "id": 881234,
      "postDate": "2020-06-10T20:53:11.723Z",
      "content": "<p>I have problems using npy files from resized images\nImages loaded with np.memmap have different appearence that when loaded with np.load, the first appears someway shifted \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1828058%2F4a4015359f6eb4cadc2925246574342c%2FnpMemmload.png?generation=1591822315508337&amp;alt=media\" alt=\"\"></p>\n\n<p>left loaded with np.load right with np.memmap\nis anyone having this problem... or I may be missing something?\nThanks in advance</p>",
      "rawMarkdown": "I have problems using npy files from resized images\nImages loaded with np.memmap have different appearence that when loaded with np.load, the first appears someway shifted \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1828058%2F4a4015359f6eb4cadc2925246574342c%2FnpMemmload.png?generation=1591822315508337&amp;alt=media)\n\nleft loaded with np.load right with np.memmap\nis anyone having this problem... or I may be missing something?\nThanks in advance"
    },
    {
      "id": 870042,
      "postDate": "2020-06-01T13:01:02.967Z",
      "content": "<p>Just a stupid doubt but is there any reason why we should downsize to 224 and not 256</p>",
      "rawMarkdown": "Just a stupid doubt but is there any reason why we should downsize to 224 and not 256",
      "replies": [
        {
          "id": 870078,
          "postDate": "2020-06-01T13:30:53.147Z",
          "content": "<p>224x224 is the native resolution of many pretrained Imagenet models. The resized 224x224 is also the largest one that can fit in 5 GB of disk space, which is the limit that we have in Kaggle kernels. </p>\n\n<p>I am planning on adding the 256x256 resolution that I resize offline as well, but have not gotten to it yet.</p>",
          "rawMarkdown": "224x224 is the native resolution of many pretrained Imagenet models. The resized 224x224 is also the largest one that can fit in 5 GB of disk space, which is the limit that we have in Kaggle kernels. \n\nI am planning on adding the 256x256 resolution that I resize offline as well, but have not gotten to it yet.",
          "votes": 1
        }
      ]
    },
    {
      "id": 869772,
      "postDate": "2020-06-01T09:24:52.060Z",
      "content": "<p>Getting all those links must have taken some time! Thank you for taking the trouble.</p>",
      "rawMarkdown": "Getting all those links must have taken some time! Thank you for taking the trouble."
    },
    {
      "id": 868531,
      "postDate": "2020-05-31T10:10:06.703Z",
      "content": "<p><a href=\"/tunguz\">@tunguz</a> Thank you for sharing these images. I just wanted to ask that is there any way to directly download the dataset to our colab nnotebook, some kind of api or so</p>",
      "rawMarkdown": "@tunguz Thank you for sharing these images. I just wanted to ask that is there any way to directly download the dataset to our colab nnotebook, some kind of api or so",
      "replies": [
        {
          "id": 869228,
          "postDate": "2020-05-31T20:58:31.567Z",
          "content": "<p>Not that I know of. Colab purposefully makes it difficult to move data around. 😉  </p>",
          "rawMarkdown": "Not that I know of. Colab purposefully makes it difficult to move data around. 😉  "
        },
        {
          "id": 869766,
          "postDate": "2020-06-01T09:22:14.207Z",
          "content": "<p><a href=\"/chetan06\">@chetan06</a> wget can be used and the url of the images are available at the end of notebooks in the output section/ in the dataset for separate files. For example: (the url is long, so hidden)</p>\n\n<p>!wget <a href=\"https://www.kaggleusercontent.com/kf/35005193/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..YK7SDrkN2rnAKlulOthCGQ.JAuqrwGvctqSaAUcHXhkaLAI_bU1VsEkmmQS569fKzRehgiSbsSwLH7aF3aJRT3Vvo4z5yLj45rE7P73aAl2GTB3PrlD6iY88NDcBz6VDI5_y619ENAU8Ld76b9kpphgjKbyRr9cktsJCiFTe-o9kXmdctLOMo-hit0MZPcIgq1GNblNc9_H9cqVJ1vz13zr7cRZrSquf8vVoTULHbq5fXWk2oTuHM0hKYoDZLg2XHceVfirLEwSJ83m52_13XArk4I2UAYb86u82PISRsFxIFnYO6Xp-OlrIXNHut8xFWFqg81LTIOTroMJ3fc5-1dMxDEt-ULtXDjbZ9udFMrBwTD8w7HhNBhUB0Dh7G3vyi8nUvFgep5KFI54oAkp7UsJpIXjTZu6TO1uHkUZjkVdXFLBGxRlvnLnXL3IjyYYYv1PUTlk57_UYHbnrrc5cU4Ml4cP-2FVZUL2xopV6yQjYO4VfWaKJbittGU9O0WFNvOPhZJLKvqCxnHYhW6qjYU_PdFH4p7Hx1eIKrYhYyX1DuMNiFzR6o8U2NVzEmmzWsuRDbVxbAglXaoQbFyDXRjsZvbSQQnBJCgIxZHGOs0T8d1A1Lej4WPEwVW2mxOKW4mMbC0gdY1zllyx7JjO1BLxhRXRd4QbgtEeXD3g-gpCfdYb21ylOISBgIZl4XjGvRU.5wnbhbLkwGqVSOPjUfGnpg/x_train_224.npy\">224x224-train</a>\n!wget <a href=\"https://www.kaggleusercontent.com/kf/35005268/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..ERx_kfoJ9xWg9XNJZ7DJ6g.NkQZeFq6g-y0pgGUO-tXXv8cC2V3Hl_P4_Hyr1qa1O5gRwOOrfhRPdf4h5e1Atv4JotSz2c-ZcOyQFrUBM-5OM3xzkxkPR1-P91V5_wkVlqBRWwtukNKdIe60OT8Oes1DbtNV7xAQ66qnT7snS-1Guvr8HRCswCjbRLA_C5MNnjA3Ub_edbpVOjseWW0xiwv3EYeKG2-zn2f_XgsUNTDd-CSBbyLtV2BTgstosYOEjeo47dzD61gzXQnEeXYFKMHYRrUfAeCGXaVZpDkXuGG_BzC6Kh5MGBVcPR4z2PL_7lcjhBO_HgO2b0_g0ULsoImR0VfKO-9RiPVnr_Oz0yntI3CEGoe89ZRpTdAn8rO6LrXyFcqvOqkbTyovvaZFw4WJbGpkqP6QHLC93maPgPMECG3YoZdzRKdzn4ONwuzcMdrGsMGMiNx2x0NWvhGSQs9fWRbxhknffiaa60wm5w9OxTQljW0LKkAqwaJhA1aUjgp5fUF8tWEFZxG3HI-PzdknR7h3k-TuvpuXnUYeYLrI0r2s0qUcjI1nYGdwilVF9Redh3xBL-YvD2b5eeyNTo9fS96Gw2J7I6XlPVIK4DvCP4Zede5Z8xLpQRAlc73tHAAafH8850ZyBQhHn_YLkaDpv-51Xy1WVCJ51dyz_ijhmCr3RIu2JGD9TBjnWG7z7Q.MaABMGFvMFh4JL2wCwlvOA/x_test_224.npy\">224x224-test</a></p>",
          "rawMarkdown": "@chetan06 wget can be used and the url of the images are available at the end of notebooks in the output section/ in the dataset for separate files. For example: (the url is long, so hidden)\n\n!wget [224x224-train](https://www.kaggleusercontent.com/kf/35005193/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..YK7SDrkN2rnAKlulOthCGQ.JAuqrwGvctqSaAUcHXhkaLAI_bU1VsEkmmQS569fKzRehgiSbsSwLH7aF3aJRT3Vvo4z5yLj45rE7P73aAl2GTB3PrlD6iY88NDcBz6VDI5_y619ENAU8Ld76b9kpphgjKbyRr9cktsJCiFTe-o9kXmdctLOMo-hit0MZPcIgq1GNblNc9_H9cqVJ1vz13zr7cRZrSquf8vVoTULHbq5fXWk2oTuHM0hKYoDZLg2XHceVfirLEwSJ83m52_13XArk4I2UAYb86u82PISRsFxIFnYO6Xp-OlrIXNHut8xFWFqg81LTIOTroMJ3fc5-1dMxDEt-ULtXDjbZ9udFMrBwTD8w7HhNBhUB0Dh7G3vyi8nUvFgep5KFI54oAkp7UsJpIXjTZu6TO1uHkUZjkVdXFLBGxRlvnLnXL3IjyYYYv1PUTlk57_UYHbnrrc5cU4Ml4cP-2FVZUL2xopV6yQjYO4VfWaKJbittGU9O0WFNvOPhZJLKvqCxnHYhW6qjYU_PdFH4p7Hx1eIKrYhYyX1DuMNiFzR6o8U2NVzEmmzWsuRDbVxbAglXaoQbFyDXRjsZvbSQQnBJCgIxZHGOs0T8d1A1Lej4WPEwVW2mxOKW4mMbC0gdY1zllyx7JjO1BLxhRXRd4QbgtEeXD3g-gpCfdYb21ylOISBgIZl4XjGvRU.5wnbhbLkwGqVSOPjUfGnpg/x_train_224.npy)\n!wget [224x224-test](https://www.kaggleusercontent.com/kf/35005268/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..ERx_kfoJ9xWg9XNJZ7DJ6g.NkQZeFq6g-y0pgGUO-tXXv8cC2V3Hl_P4_Hyr1qa1O5gRwOOrfhRPdf4h5e1Atv4JotSz2c-ZcOyQFrUBM-5OM3xzkxkPR1-P91V5_wkVlqBRWwtukNKdIe60OT8Oes1DbtNV7xAQ66qnT7snS-1Guvr8HRCswCjbRLA_C5MNnjA3Ub_edbpVOjseWW0xiwv3EYeKG2-zn2f_XgsUNTDd-CSBbyLtV2BTgstosYOEjeo47dzD61gzXQnEeXYFKMHYRrUfAeCGXaVZpDkXuGG_BzC6Kh5MGBVcPR4z2PL_7lcjhBO_HgO2b0_g0ULsoImR0VfKO-9RiPVnr_Oz0yntI3CEGoe89ZRpTdAn8rO6LrXyFcqvOqkbTyovvaZFw4WJbGpkqP6QHLC93maPgPMECG3YoZdzRKdzn4ONwuzcMdrGsMGMiNx2x0NWvhGSQs9fWRbxhknffiaa60wm5w9OxTQljW0LKkAqwaJhA1aUjgp5fUF8tWEFZxG3HI-PzdknR7h3k-TuvpuXnUYeYLrI0r2s0qUcjI1nYGdwilVF9Redh3xBL-YvD2b5eeyNTo9fS96Gw2J7I6XlPVIK4DvCP4Zede5Z8xLpQRAlc73tHAAafH8850ZyBQhHn_YLkaDpv-51Xy1WVCJ51dyz_ijhmCr3RIu2JGD9TBjnWG7z7Q.MaABMGFvMFh4JL2wCwlvOA/x_test_224.npy)"
        },
        {
          "id": 893820,
          "postDate": "2020-06-20T00:54:58.357Z",
          "content": "<p>Kaggle API works in colab. For eg: <code>!kaggle kernels output tunguz/image-resizing-64x64-test</code> </p>",
          "rawMarkdown": "Kaggle API works in colab. For eg: `!kaggle kernels output tunguz/image-resizing-64x64-test` "
        }
      ]
    },
    {
      "id": 868194,
      "postDate": "2020-05-31T02:59:54.023Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 866109,
      "postDate": "2020-05-29T05:44:43.780Z",
      "content": "<p>Thanks <a href=\"/tunguz\">@tunguz</a> very handy indeed !</p>",
      "rawMarkdown": "Thanks @tunguz very handy indeed !",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 865821,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-05-28T23:06:57.387000",
      "content": "<p>Thanks. Note I've done 224x224 here: <a href=\"https://www.kaggle.com/arroqc/siic-isic-224x224-images\">https://www.kaggle.com/arroqc/siic-isic-224x224-images</a></p>",
      "votes": 9,
      "replies": [
        {
          "id": 865822,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2020-05-28T23:08:29.203000",
          "content": "<p>Thanks, good to know - I will be adding my own version anyway soon, the script is almost finished running. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 866400,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2020-05-29T11:14:09.420000",
          "content": "<p>And here is my version of the rescaled 224x224 images: </p>\n\n<p><a href=\"https://www.kaggle.com/tunguz/image-resizing-224x224-train/\">https://www.kaggle.com/tunguz/image-resizing-224x224-train/</a>\n<a href=\"https://www.kaggle.com/tunguz/image-resizing-224x224-test\">https://www.kaggle.com/tunguz/image-resizing-224x224-test</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 880433,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-10T09:46:21.293000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 865847,
      "author_name": "Bojan Tunguz",
      "author_url": "",
      "post_date": "2020-05-28T23:53:18.087000",
      "content": "<p>Just a simple logistic regression model on flattened 32x32 images can give you LB score of 0.78 AUC:</p>\n\n<p><a href=\"https://www.kaggle.com/tunguz/melanoma-classification-eda-and-modeling\">https://www.kaggle.com/tunguz/melanoma-classification-eda-and-modeling</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 866772,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-05-29T16:57:15.733000",
          "content": "<p>Interesting. Is logistic regression only using the flattened images to score 0.78, or are you including the extra features such as age, gender, etc?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 866813,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2020-05-29T17:35:35.787000",
          "content": "<p>You can reach 0.78 with only LR and flattened 32x32 images. I am now including those other features that you mentioned and assembling, and I am getting 0.793 in the latest edition of the kernel. I beleive I can reach 0.8 with just LR and these features. :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 872118,
          "author_name": "Gianluca Rossi",
          "author_url": "",
          "post_date": "2020-06-02T23:40:21.207000",
          "content": "<p>I can confirm what <a href=\"/tunguz\">@tunguz</a> said. I can achieve 0.8134 (10-folds CV) with just a simple LR model on flattened 32x32 images and the \"contextual\" features.  </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 865801,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-05-28T22:34:09.427000",
      "content": "<p>Thanks Bojan</p>",
      "votes": 3,
      "replies": [
        {
          "id": 865803,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2020-05-28T22:36:23.850000",
          "content": "<p>Ah, don't mention it. </p>\n\n<p><img src=\"https://pbs.twimg.com/media/EZIy_xiWAAEAM-i?format=png&amp;name=900x900\" alt=\"\"></p>",
          "votes": 15,
          "replies": []
        }
      ]
    },
    {
      "id": 872158,
      "author_name": "Gianluca Rossi",
      "author_url": "",
      "post_date": "2020-06-03T00:41:54.187000",
      "content": "<p>Thank you for sharing your notebook <a href=\"/tunguz\">@tunguz</a> </p>\n\n<p>I've refactored and used multiprocessing to take advantage of all cores in my machine. I don't like much working in Kaggle notebooks so I'm attaching it here in case anyone is interested.</p>\n\n<p>```\nfrom functools import partial\nfrom multiprocessing import Pool</p>\n\n<p>import numpy as np\nimport pandas as pd\nfrom PIL import Image</p>\n\n<p>from siim_isic_melanoma_classification.constants import (\n    train_fpath,\n    test_fpath,\n    data_path,\n)</p>\n\n<p>IMAGE_SZ = 32</p>\n\n<p>def main():\n    train_df = pd.read_csv(train_fpath)\n    test_df = pd.read_csv(test_fpath)</p>\n\n<pre><code>for df, name in [(train_df, \"train\"), (test_df, \"test\")]:\n    print(f\"Converting {name} images to NumPy array...\")\n    ar = convert_to_array(df, name=name)\n    np.save(data_path / f\"x_{name}_32\", ar)\n</code></pre>\n\n<p>def convert_to_array(df, name: str):\n    pool = Pool()\n    routine = partial(image_to_flattened_array, name=name)\n    data = pool.map(routine, df[\"image_name\"])\n    return np.vstack(data)</p>\n\n<p>def image_to_flattened_array(image_id, name: str, desired_size=IMAGE_SZ):\n    image_fpath = data_path / f\"{name}/{image_id}.jpg\"\n    im = Image.open(image_fpath)\n    small_im = im.resize((desired_size,) * 2, resample=Image.LANCZOS)\n    return np.array(small_im).reshape((1, 32, 32, 3))</p>\n\n<p>if <strong>name</strong> == \"<strong>main</strong>\":\n    main()\n```</p>\n\n<p>It could be further refactored but it does the job :-)</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 865820,
      "author_name": "Blonde",
      "author_url": "",
      "post_date": "2020-05-28T23:06:03.220000",
      "content": "<p>npy arrays for 512 x 512 would be also nice, not for the kernels use, but in general </p>",
      "votes": 4,
      "replies": [
        {
          "id": 865823,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2020-05-28T23:08:59.960000",
          "content": "<p>I will probably make those offline and upload them to the dataset mentioned above.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 886812,
          "author_name": "amir arqand",
          "author_url": "",
          "post_date": "2020-06-15T09:44:18.050000",
          "content": "<p>That would be great! lookin  forward for it</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 875616,
      "author_name": "Marcelo Kittlein",
      "author_url": "",
      "post_date": "2020-06-06T02:39:35.983000",
      "content": "<p>Thanks <a href=\"/tunguz\">@tunguz</a> it is awesome you helping others to experiment and learn...Best</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 872432,
      "author_name": "Karan",
      "author_url": "",
      "post_date": "2020-06-03T07:27:20.527000",
      "content": "<p>Thanks <a href=\"/tunguz\">@tunguz</a>  for all these resized dataset. Hoping for 512x512 soon.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 867534,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-05-30T11:53:11.470000",
      "content": "<p><a href=\"/tunguz\">@tunguz</a> hello, thanks for sharing. One thing, in <code>jpeg</code> , all images are not in uniform resolution. Do you think plain resize is OK or should it be scale down by a factor? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 868010,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2020-05-30T20:14:50.950000",
          "content": "<p>For the most part that should not be a major obstacle. It is a standard practice in most computer vision challenges to resize images ot the same fixed size.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 868019,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-05-30T20:23:23.777000",
          "content": "<p>Yes, I know and I agree. Actually it reminds me of the bengali.ai competition. That time In my cases, using a non-square image (original shape) gave me promising scores.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 870182,
      "author_name": "Animesh Nareda",
      "author_url": "",
      "post_date": "2020-06-01T14:43:45.707000",
      "content": "<p>Thank a lot for this, people at Kaggle are best!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 867995,
      "author_name": "Dhruv Aggarwal",
      "author_url": "",
      "post_date": "2020-05-30T20:04:21.620000",
      "content": "<p><a href=\"/tunguz\">@tunguz</a> Thankyou! That will be useful.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 928032,
      "author_name": "Jagdish Mirchandani",
      "author_url": "",
      "post_date": "2020-07-13T17:46:59.390000",
      "content": "<p>This is so helpful! Thanks for doing this</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 893938,
      "author_name": "Satya Muralidhar",
      "author_url": "",
      "post_date": "2020-06-20T04:55:20.703000",
      "content": "<p>good one chris</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 885628,
      "author_name": "M&M",
      "author_url": "",
      "post_date": "2020-06-14T10:58:25.773000",
      "content": "<p>It's my first competition and was wondering how I am supposed to work with 100GB+ data and found this. Thank you so much!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 881234,
      "author_name": "Marcelo Kittlein",
      "author_url": "",
      "post_date": "2020-06-10T20:53:11.723000",
      "content": "<p>I have problems using npy files from resized images\nImages loaded with np.memmap have different appearence that when loaded with np.load, the first appears someway shifted \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1828058%2F4a4015359f6eb4cadc2925246574342c%2FnpMemmload.png?generation=1591822315508337&amp;alt=media\" alt=\"\"></p>\n\n<p>left loaded with np.load right with np.memmap\nis anyone having this problem... or I may be missing something?\nThanks in advance</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 870042,
      "author_name": "darthgera123",
      "author_url": "",
      "post_date": "2020-06-01T13:01:02.967000",
      "content": "<p>Just a stupid doubt but is there any reason why we should downsize to 224 and not 256</p>",
      "votes": 0,
      "replies": [
        {
          "id": 870078,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2020-06-01T13:30:53.147000",
          "content": "<p>224x224 is the native resolution of many pretrained Imagenet models. The resized 224x224 is also the largest one that can fit in 5 GB of disk space, which is the limit that we have in Kaggle kernels. </p>\n\n<p>I am planning on adding the 256x256 resolution that I resize offline as well, but have not gotten to it yet.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 869772,
      "author_name": "G_R_S",
      "author_url": "",
      "post_date": "2020-06-01T09:24:52.060000",
      "content": "<p>Getting all those links must have taken some time! Thank you for taking the trouble.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 868531,
      "author_name": "Anonymous",
      "author_url": "",
      "post_date": "2020-05-31T10:10:06.703000",
      "content": "<p><a href=\"/tunguz\">@tunguz</a> Thank you for sharing these images. I just wanted to ask that is there any way to directly download the dataset to our colab nnotebook, some kind of api or so</p>",
      "votes": 0,
      "replies": [
        {
          "id": 869228,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2020-05-31T20:58:31.567000",
          "content": "<p>Not that I know of. Colab purposefully makes it difficult to move data around. 😉  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 869766,
          "author_name": "Shayekh Islam",
          "author_url": "",
          "post_date": "2020-06-01T09:22:14.207000",
          "content": "<p><a href=\"/chetan06\">@chetan06</a> wget can be used and the url of the images are available at the end of notebooks in the output section/ in the dataset for separate files. For example: (the url is long, so hidden)</p>\n\n<p>!wget <a href=\"https://www.kaggleusercontent.com/kf/35005193/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..YK7SDrkN2rnAKlulOthCGQ.JAuqrwGvctqSaAUcHXhkaLAI_bU1VsEkmmQS569fKzRehgiSbsSwLH7aF3aJRT3Vvo4z5yLj45rE7P73aAl2GTB3PrlD6iY88NDcBz6VDI5_y619ENAU8Ld76b9kpphgjKbyRr9cktsJCiFTe-o9kXmdctLOMo-hit0MZPcIgq1GNblNc9_H9cqVJ1vz13zr7cRZrSquf8vVoTULHbq5fXWk2oTuHM0hKYoDZLg2XHceVfirLEwSJ83m52_13XArk4I2UAYb86u82PISRsFxIFnYO6Xp-OlrIXNHut8xFWFqg81LTIOTroMJ3fc5-1dMxDEt-ULtXDjbZ9udFMrBwTD8w7HhNBhUB0Dh7G3vyi8nUvFgep5KFI54oAkp7UsJpIXjTZu6TO1uHkUZjkVdXFLBGxRlvnLnXL3IjyYYYv1PUTlk57_UYHbnrrc5cU4Ml4cP-2FVZUL2xopV6yQjYO4VfWaKJbittGU9O0WFNvOPhZJLKvqCxnHYhW6qjYU_PdFH4p7Hx1eIKrYhYyX1DuMNiFzR6o8U2NVzEmmzWsuRDbVxbAglXaoQbFyDXRjsZvbSQQnBJCgIxZHGOs0T8d1A1Lej4WPEwVW2mxOKW4mMbC0gdY1zllyx7JjO1BLxhRXRd4QbgtEeXD3g-gpCfdYb21ylOISBgIZl4XjGvRU.5wnbhbLkwGqVSOPjUfGnpg/x_train_224.npy\">224x224-train</a>\n!wget <a href=\"https://www.kaggleusercontent.com/kf/35005268/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..ERx_kfoJ9xWg9XNJZ7DJ6g.NkQZeFq6g-y0pgGUO-tXXv8cC2V3Hl_P4_Hyr1qa1O5gRwOOrfhRPdf4h5e1Atv4JotSz2c-ZcOyQFrUBM-5OM3xzkxkPR1-P91V5_wkVlqBRWwtukNKdIe60OT8Oes1DbtNV7xAQ66qnT7snS-1Guvr8HRCswCjbRLA_C5MNnjA3Ub_edbpVOjseWW0xiwv3EYeKG2-zn2f_XgsUNTDd-CSBbyLtV2BTgstosYOEjeo47dzD61gzXQnEeXYFKMHYRrUfAeCGXaVZpDkXuGG_BzC6Kh5MGBVcPR4z2PL_7lcjhBO_HgO2b0_g0ULsoImR0VfKO-9RiPVnr_Oz0yntI3CEGoe89ZRpTdAn8rO6LrXyFcqvOqkbTyovvaZFw4WJbGpkqP6QHLC93maPgPMECG3YoZdzRKdzn4ONwuzcMdrGsMGMiNx2x0NWvhGSQs9fWRbxhknffiaa60wm5w9OxTQljW0LKkAqwaJhA1aUjgp5fUF8tWEFZxG3HI-PzdknR7h3k-TuvpuXnUYeYLrI0r2s0qUcjI1nYGdwilVF9Redh3xBL-YvD2b5eeyNTo9fS96Gw2J7I6XlPVIK4DvCP4Zede5Z8xLpQRAlc73tHAAafH8850ZyBQhHn_YLkaDpv-51Xy1WVCJ51dyz_ijhmCr3RIu2JGD9TBjnWG7z7Q.MaABMGFvMFh4JL2wCwlvOA/x_test_224.npy\">224x224-test</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 893820,
          "author_name": "RB",
          "author_url": "",
          "post_date": "2020-06-20T00:54:58.357000",
          "content": "<p>Kaggle API works in colab. For eg: <code>!kaggle kernels output tunguz/image-resizing-64x64-test</code> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 868194,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-31T02:59:54.023000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 866109,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2020-05-29T05:44:43.780000",
      "content": "<p>Thanks <a href=\"/tunguz\">@tunguz</a> very handy indeed !</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "865785": "Since the images in this competition are very large (1024x1024), it will probably be very useful for most of us to first experiment with smaller sizes. I decided to create numpy arrays of the smaller-resized images and upload them to a separate dataset. It can be found here: \n\nhttps://www.kaggle.com/tunguz/siimisic-melanoma-resized-images\n\nAll resizings were done in kernels, and can be found below:\n\nhttps://www.kaggle.com/tunguz/image-resizing-32x32-train\nhttps://www.kaggle.com/tunguz/image-resizing-32x32-test\nhttps://www.kaggle.com/tunguz/image-resizing-64x-train\nhttps://www.kaggle.com/tunguz/image-resizing-64x64-test\nhttps://www.kaggle.com/tunguz/image-resizing-96x96-train\nhttps://www.kaggle.com/tunguz/image-resizing-96x96-test\nhttps://www.kaggle.com/tunguz/image-resizing-128x128-train\nhttps://www.kaggle.com/tunguz/image-resizing-128x128-test\nhttps://www.kaggle.com/tunguz/image-resizing-224x224-train\nhttps://www.kaggle.com/tunguz/image-resizing-224x224-test\n\n*** Update 5-29-2020 ***\n\nI have added 96x96 and 224x224 images as well.",
    "865821": "Thanks. Note I've done 224x224 here: https://www.kaggle.com/arroqc/siic-isic-224x224-images",
    "865847": "Just a simple logistic regression model on flattened 32x32 images can give you LB score of 0.78 AUC:\n\nhttps://www.kaggle.com/tunguz/melanoma-classification-eda-and-modeling",
    "865801": "Thanks Bojan",
    "872158": "Thank you for sharing your notebook @tunguz \n\nI've refactored and used multiprocessing to take advantage of all cores in my machine. I don't like much working in Kaggle notebooks so I'm attaching it here in case anyone is interested.\n\n```\nfrom functools import partial\nfrom multiprocessing import Pool\n\nimport numpy as np\nimport pandas as pd\nfrom PIL import Image\n\nfrom siim_isic_melanoma_classification.constants import (\n    train_fpath,\n    test_fpath,\n    data_path,\n)\n\nIMAGE_SZ = 32\n\n\ndef main():\n    train_df = pd.read_csv(train_fpath)\n    test_df = pd.read_csv(test_fpath)\n\n    for df, name in [(train_df, \"train\"), (test_df, \"test\")]:\n        print(f\"Converting {name} images to NumPy array...\")\n        ar = convert_to_array(df, name=name)\n        np.save(data_path / f\"x_{name}_32\", ar)\n\n\ndef convert_to_array(df, name: str):\n    pool = Pool()\n    routine = partial(image_to_flattened_array, name=name)\n    data = pool.map(routine, df[\"image_name\"])\n    return np.vstack(data)\n\n\ndef image_to_flattened_array(image_id, name: str, desired_size=IMAGE_SZ):\n    image_fpath = data_path / f\"{name}/{image_id}.jpg\"\n    im = Image.open(image_fpath)\n    small_im = im.resize((desired_size,) * 2, resample=Image.LANCZOS)\n    return np.array(small_im).reshape((1, 32, 32, 3))\n\n\nif __name__ == \"__main__\":\n    main()\n```\n\nIt could be further refactored but it does the job :-)",
    "865820": "npy arrays for 512 x 512 would be also nice, not for the kernels use, but in general ",
    "875616": "Thanks @tunguz it is awesome you helping others to experiment and learn...Best",
    "872432": "Thanks @tunguz  for all these resized dataset. Hoping for 512x512 soon.",
    "867534": "@tunguz hello, thanks for sharing. One thing, in `jpeg` , all images are not in uniform resolution. Do you think plain resize is OK or should it be scale down by a factor? ",
    "870182": "Thank a lot for this, people at Kaggle are best!",
    "867995": "@tunguz Thankyou! That will be useful.",
    "928032": "This is so helpful! Thanks for doing this",
    "893938": "good one chris",
    "885628": "It's my first competition and was wondering how I am supposed to work with 100GB+ data and found this. Thank you so much!",
    "881234": "I have problems using npy files from resized images\nImages loaded with np.memmap have different appearence that when loaded with np.load, the first appears someway shifted \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1828058%2F4a4015359f6eb4cadc2925246574342c%2FnpMemmload.png?generation=1591822315508337&amp;alt=media)\n\nleft loaded with np.load right with np.memmap\nis anyone having this problem... or I may be missing something?\nThanks in advance",
    "870042": "Just a stupid doubt but is there any reason why we should downsize to 224 and not 256",
    "869772": "Getting all those links must have taken some time! Thank you for taking the trouble.",
    "868531": "@tunguz Thank you for sharing these images. I just wanted to ask that is there any way to directly download the dataset to our colab nnotebook, some kind of api or so",
    "868194": "",
    "866109": "Thanks @tunguz very handy indeed !"
  }
}