{
  "id": 163501,
  "title": "Can we use TFrecords image format with ImageDataGenerator() ? ",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/163501",
  "author_name": "",
  "post_date": "2020-07-02T09:44:47.236188900Z",
  "votes": null,
  "comment_count": 9,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": "912197",
      "postDate": "07/02/2020 09:44:47",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "913373",
      "postDate": "07/03/2020 07:21:04",
      "content": "<p>I guess No. \nUsing ImageDataGenerators it won't be possible to extract data from tfrecords .\nTensorflow has particular methods to read and process tfrecords for this purpose. Checkout the a notebook from <a href=\"/cdeotte\">@cdeotte</a> in the notebooks to read and create tfrecords.</p>",
      "rawMarkdown": "I guess No. \nUsing ImageDataGenerators it won't be possible to extract data from tfrecords .\nTensorflow has particular methods to read and process tfrecords for this purpose. Checkout the a notebook from @cdeotte in the notebooks to read and create tfrecords.",
      "votes": null
    },
    {
      "id": "915357",
      "postDate": "07/04/2020 17:03:54",
      "content": "<p>No. If you want to use ImageDataGenerator, then use my JPEGs instead of TFRecords posted <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a></p>",
      "rawMarkdown": "No. If you want to use ImageDataGenerator, then use my JPEGs instead of TFRecords posted [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092",
      "votes": null
    },
    {
      "id": "916150",
      "postDate": "07/05/2020 12:18:21",
      "content": "<p>I used the JPEG images with ImageDataGenerator but I got the same prediction value\nfor each image in test data. Here is my kernel link.\n<a href=\"https://www.kaggle.com/anuragkushwah/kernel15cd6cce67\">https://www.kaggle.com/anuragkushwah/kernel15cd6cce67</a>\n<a href=\"/cdeotte\">@cdeotte</a> sir can you please help me in this ,why is this happening. Also thanks for the above given link 😊</p>",
      "rawMarkdown": "I used the JPEG images with ImageDataGenerator but I got the same prediction value\nfor each image in test data. Here is my kernel link.\nhttps://www.kaggle.com/anuragkushwah/kernel15cd6cce67\n@cdeotte sir can you please help me in this ,why is this happening. Also thanks for the above given link 😊",
      "votes": null
    },
    {
      "id": "916411",
      "postDate": "07/05/2020 15:56:22",
      "content": "<p>I have two ideas and i'm testing them now. First, you need to turn on internet so that your VGG can download its pretrained weights. Second, i suggest just trying </p>\n\n<pre><code>loss = tf.keras.losses.BinaryCrossentropy()\nmodel.compile(optimizer=opt, loss = loss, metrics=['accuracy'])\n</code></pre>\n\n<p>This will see if the problem is being caused by your custom focal loss. Another strange thing that i see is that your epochs are taking 45 seconds. The epochs should be more like 15 minutes, so that is probably the main cause of your problem. I'm not sure why your run was 45 seconds.</p>",
      "rawMarkdown": "I have two ideas and i'm testing them now. First, you need to turn on internet so that your VGG can download its pretrained weights. Second, i suggest just trying \n\n    loss = tf.keras.losses.BinaryCrossentropy()\n    model.compile(optimizer=opt, loss = loss, metrics=['accuracy'])\n\nThis will see if the problem is being caused by your custom focal loss. Another strange thing that i see is that your epochs are taking 45 seconds. The epochs should be more like 15 minutes, so that is probably the main cause of your problem. I'm not sure why your run was 45 seconds.",
      "votes": null
    },
    {
      "id": "916490",
      "postDate": "07/05/2020 17:44:21",
      "content": "<p>Okk sir, I will try with binarycrossentrory ().\nAnd the reason for epochs taking 45 sec. is may be because I haven't take all the images for the training but created a new data frame with a random sample of 2000 benign ( target=0) and all the malignant one (In [188]) because of unsymmetrical data. So there are only 2600 images approx. for training &amp; validation purpose. I think that's why it takes only 45 seconds.</p>",
      "rawMarkdown": "Okk sir, I will try with binarycrossentrory ().\nAnd the reason for epochs taking 45 sec. is may be because I haven't take all the images for the training but created a new data frame with a random sample of 2000 benign ( target=0) and all the malignant one (In [188]) because of unsymmetrical data. So there are only 2600 images approx. for training &amp; validation purpose. I think that's why it takes only 45 seconds.",
      "votes": null
    },
    {
      "id": "916579",
      "postDate": "07/05/2020 19:35:35",
      "content": "<p>Only 33 steps - think it should be more like 33000/64 for the number of steps ??    </p>",
      "rawMarkdown": "Only 33 steps - think it should be more like 33000/64 for the number of steps ??",
      "votes": null
    },
    {
      "id": "916773",
      "postDate": "07/06/2020 02:49:32",
      "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> as I said above I haven't take all the images for training purposes, there are only 2067 images for training purpose that's why 2067/64 is nearly 33 steps..</p>",
      "rawMarkdown": "pcjimmmy as I said above I haven't take all the images for training purposes, there are only 2067 images for training purpose that's why 2067/64 is nearly 33 steps..",
      "votes": null
    },
    {
      "id": "916817",
      "postDate": "07/06/2020 04:07:48",
      "content": "<p>That may be the reason for having the same prediction value. If you are using only 2067 training images, your CNN may memorize them quickly. Perhaps try training with more images. If using 224x224 takes too long, try using my JPEGs 128x128 <a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-128x128\">here</a> (or TFRecords 128x128 <a href=\"https://www.kaggle.com/cdeotte/melanoma-128x128\">here</a>).</p>",
      "rawMarkdown": "That may be the reason for having the same prediction value. If you are using only 2067 training images, your CNN may memorize them quickly. Perhaps try training with more images. If using 224x224 takes too long, try using my JPEGs 128x128 [here][1] (or TFRecords 128x128 [here][2]).\n\n[1]: https://www.kaggle.com/cdeotte/jpeg-melanoma-128x128\n[2]: https://www.kaggle.com/cdeotte/melanoma-128x128",
      "votes": null
    },
    {
      "id": "918056",
      "postDate": "07/06/2020 23:32:35",
      "content": "<p>One of my big lessons on many of the methods available in machine learning - they are powerful.  </p>\n\n<p>Running small sample sizes for debug purposes is a great tool.  But the powerful ML method needs large sample sizes to avoid CNN simply doing what Chris suspects - you just memorize and not generalize.</p>\n\n<p>There are a million paths that you can take to develop a model - sadly using small sample sizes will only inform when you have made HUGE robust changes to parameters.  But even then I have been led down the wrong path trying to use small sample sizes.</p>\n\n<p>In almost every model that I create I will have a debug that runs only a very limited number of steps per epoch.  When I make large code changes that have the potential for me being STUPID than I run the small steps to validate the code.   But I can't use the results to inform if the idea is good.</p>\n\n<p>Once the code is clean than I want each epoch to look at all the images.  For vision problems using small image sizes lets you set batch sizes for speed of processing the model.  As Chris points out in another discussion topic you model will be seeing different features but in general small images can be used to evaluate most parameter changes.</p>",
      "rawMarkdown": "One of my big lessons on many of the methods available in machine learning - they are powerful.  \n\nRunning small sample sizes for debug purposes is a great tool.  But the powerful ML method needs large sample sizes to avoid CNN simply doing what Chris suspects - you just memorize and not generalize.\n\nThere are a million paths that you can take to develop a model - sadly using small sample sizes will only inform when you have made HUGE robust changes to parameters.  But even then I have been led down the wrong path trying to use small sample sizes.\n\nIn almost every model that I create I will have a debug that runs only a very limited number of steps per epoch.  When I make large code changes that have the potential for me being STUPID than I run the small steps to validate the code.   But I can't use the results to inform if the idea is good.\n\nOnce the code is clean than I want each epoch to look at all the images.  For vision problems using small image sizes lets you set batch sizes for speed of processing the model.  As Chris points out in another discussion topic you model will be seeing different features but in general small images can be used to evaluate most parameter changes.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 913373,
      "author_name": "prashantarorat",
      "author_url": "",
      "post_date": "07/03/2020 07:21:04",
      "content": "<p>I guess No. \nUsing ImageDataGenerators it won't be possible to extract data from tfrecords .\nTensorflow has particular methods to read and process tfrecords for this purpose. Checkout the a notebook from <a href=\"/cdeotte\">@cdeotte</a> in the notebooks to read and create tfrecords.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 915357,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/04/2020 17:03:54",
      "content": "<p>No. If you want to use ImageDataGenerator, then use my JPEGs instead of TFRecords posted <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 916150,
          "author_name": "anuragkushwah",
          "author_url": "",
          "post_date": "07/05/2020 12:18:21",
          "content": "<p>I used the JPEG images with ImageDataGenerator but I got the same prediction value\nfor each image in test data. Here is my kernel link.\n<a href=\"https://www.kaggle.com/anuragkushwah/kernel15cd6cce67\">https://www.kaggle.com/anuragkushwah/kernel15cd6cce67</a>\n<a href=\"/cdeotte\">@cdeotte</a> sir can you please help me in this ,why is this happening. Also thanks for the above given link 😊</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916411,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/05/2020 15:56:22",
          "content": "<p>I have two ideas and i'm testing them now. First, you need to turn on internet so that your VGG can download its pretrained weights. Second, i suggest just trying </p>\n\n<pre><code>loss = tf.keras.losses.BinaryCrossentropy()\nmodel.compile(optimizer=opt, loss = loss, metrics=['accuracy'])\n</code></pre>\n\n<p>This will see if the problem is being caused by your custom focal loss. Another strange thing that i see is that your epochs are taking 45 seconds. The epochs should be more like 15 minutes, so that is probably the main cause of your problem. I'm not sure why your run was 45 seconds.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916490,
          "author_name": "anuragkushwah",
          "author_url": "",
          "post_date": "07/05/2020 17:44:21",
          "content": "<p>Okk sir, I will try with binarycrossentrory ().\nAnd the reason for epochs taking 45 sec. is may be because I haven't take all the images for the training but created a new data frame with a random sample of 2000 benign ( target=0) and all the malignant one (In [188]) because of unsymmetrical data. So there are only 2600 images approx. for training &amp; validation purpose. I think that's why it takes only 45 seconds.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916579,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/05/2020 19:35:35",
          "content": "<p>Only 33 steps - think it should be more like 33000/64 for the number of steps ??    </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916773,
          "author_name": "anuragkushwah",
          "author_url": "",
          "post_date": "07/06/2020 02:49:32",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> as I said above I haven't take all the images for training purposes, there are only 2067 images for training purpose that's why 2067/64 is nearly 33 steps..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916817,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/06/2020 04:07:48",
          "content": "<p>That may be the reason for having the same prediction value. If you are using only 2067 training images, your CNN may memorize them quickly. Perhaps try training with more images. If using 224x224 takes too long, try using my JPEGs 128x128 <a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-128x128\">here</a> (or TFRecords 128x128 <a href=\"https://www.kaggle.com/cdeotte/melanoma-128x128\">here</a>).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 918056,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/06/2020 23:32:35",
          "content": "<p>One of my big lessons on many of the methods available in machine learning - they are powerful.  </p>\n\n<p>Running small sample sizes for debug purposes is a great tool.  But the powerful ML method needs large sample sizes to avoid CNN simply doing what Chris suspects - you just memorize and not generalize.</p>\n\n<p>There are a million paths that you can take to develop a model - sadly using small sample sizes will only inform when you have made HUGE robust changes to parameters.  But even then I have been led down the wrong path trying to use small sample sizes.</p>\n\n<p>In almost every model that I create I will have a debug that runs only a very limited number of steps per epoch.  When I make large code changes that have the potential for me being STUPID than I run the small steps to validate the code.   But I can't use the results to inform if the idea is good.</p>\n\n<p>Once the code is clean than I want each epoch to look at all the images.  For vision problems using small image sizes lets you set batch sizes for speed of processing the model.  As Chris points out in another discussion topic you model will be seeing different features but in general small images can be used to evaluate most parameter changes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "912197": "",
    "913373": "I guess No. \nUsing ImageDataGenerators it won't be possible to extract data from tfrecords .\nTensorflow has particular methods to read and process tfrecords for this purpose. Checkout the a notebook from @cdeotte in the notebooks to read and create tfrecords.",
    "915357": "No. If you want to use ImageDataGenerator, then use my JPEGs instead of TFRecords posted [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092",
    "916150": "I used the JPEG images with ImageDataGenerator but I got the same prediction value\nfor each image in test data. Here is my kernel link.\nhttps://www.kaggle.com/anuragkushwah/kernel15cd6cce67\n@cdeotte sir can you please help me in this ,why is this happening. Also thanks for the above given link 😊",
    "916411": "I have two ideas and i'm testing them now. First, you need to turn on internet so that your VGG can download its pretrained weights. Second, i suggest just trying \n\n    loss = tf.keras.losses.BinaryCrossentropy()\n    model.compile(optimizer=opt, loss = loss, metrics=['accuracy'])\n\nThis will see if the problem is being caused by your custom focal loss. Another strange thing that i see is that your epochs are taking 45 seconds. The epochs should be more like 15 minutes, so that is probably the main cause of your problem. I'm not sure why your run was 45 seconds.",
    "916490": "Okk sir, I will try with binarycrossentrory ().\nAnd the reason for epochs taking 45 sec. is may be because I haven't take all the images for the training but created a new data frame with a random sample of 2000 benign ( target=0) and all the malignant one (In [188]) because of unsymmetrical data. So there are only 2600 images approx. for training &amp; validation purpose. I think that's why it takes only 45 seconds.",
    "916579": "Only 33 steps - think it should be more like 33000/64 for the number of steps ??",
    "916773": "pcjimmmy as I said above I haven't take all the images for training purposes, there are only 2067 images for training purpose that's why 2067/64 is nearly 33 steps..",
    "916817": "That may be the reason for having the same prediction value. If you are using only 2067 training images, your CNN may memorize them quickly. Perhaps try training with more images. If using 224x224 takes too long, try using my JPEGs 128x128 [here][1] (or TFRecords 128x128 [here][2]).\n\n[1]: https://www.kaggle.com/cdeotte/jpeg-melanoma-128x128\n[2]: https://www.kaggle.com/cdeotte/melanoma-128x128",
    "918056": "One of my big lessons on many of the methods available in machine learning - they are powerful.  \n\nRunning small sample sizes for debug purposes is a great tool.  But the powerful ML method needs large sample sizes to avoid CNN simply doing what Chris suspects - you just memorize and not generalize.\n\nThere are a million paths that you can take to develop a model - sadly using small sample sizes will only inform when you have made HUGE robust changes to parameters.  But even then I have been led down the wrong path trying to use small sample sizes.\n\nIn almost every model that I create I will have a debug that runs only a very limited number of steps per epoch.  When I make large code changes that have the potential for me being STUPID than I run the small steps to validate the code.   But I can't use the results to inform if the idea is good.\n\nOnce the code is clean than I want each epoch to look at all the images.  For vision problems using small image sizes lets you set batch sizes for speed of processing the model.  As Chris points out in another discussion topic you model will be seeing different features but in general small images can be used to evaluate most parameter changes."
  },
  "source": "meta"
}