{
  "id": 158395,
  "title": "TF/ TPU: from .tfrec to cross-validation. Step-by-step",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/158395",
  "author_name": "Alexey Pronin",
  "post_date": "2020-06-14T07:55:39.343000",
  "votes": 48,
  "comment_count": 16,
  "views": 0,
  "content": "<h2>Tensor Flow on TPU: from making .tfrec files to 5-fold cross-validation. Step-by-step guide.</h2>\n\n<p>UPDATE: I have found an annoying bug in version 5 of <a href=\"https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512\">my EfficientNet B0 public notebook</a> which resulted in a very low public LB score. The bug is fixed now.</p>\n\n<h3>Motivation</h3>\n\n<p>First of all, I must confess that I am a total newbie in computer vision and I really enjoy this competition because I view it as a great opportunity to learn new things, such as working with TPU in Tensor Flow. If you are coming from the Machine Learning background, a lot of things should be familiar to you. But the devil is always in details and certain things that you know how to implement with traditional sklearn library might get a bit tricky in Tensor Flow. In this discussion topic, I would like to share with you my experience so far. If you read on, ultimately, you will see an example of a working machine learning pipeline with Tensor Flow on TPU which will include \n* making Startified Group K-folds for 5-fold cross-validation, \n* making .tfrec files that include not only the image data but all other tabular data present in <code>train</code> and <code>test</code>,\n* example of a neural network that utilizes both the image and tabular data (with the cross-validation outlined above).</p>\n\n<p>Hopefully, this is going to save you some time and help you get started with the competition or improve your existing model. </p>\n\n<h3>Startified Group K-folds</h3>\n\n<p>The reasons why we need to use Stratified Group K-Folds in this competition are summarized in <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156002\">this discussion topic</a>. Some time ago, I created <a href=\"https://www.kaggle.com/graf10a/siim-stratified-groupkfold-5-folds\">this public kernel</a> showing how Stratified Group K-Folds can be implemented on the competition data.</p>\n\n<h3>Making .tfrec files</h3>\n\n<p>The most effective way of working with TPU is by feeding it .tfrec files instead of individual JPEG files. A single .tfrec file may contain hundreds or even thousands of JPEG images and feeding a single file like this to TPU can save you a lot of bandwidth, so you don't have to wait too long for your data to be transferred from the storage to the TPU unit as would be the case if you had to transfer the JPEG files individually, one-by-one. To help you to get started with .tfrec files I created the following public notebook giving an example of how these files can be made:</p>\n\n<p><a href=\"https://www.kaggle.com/graf10a/siim-example-of-making-tfrec-files-512x512\">SIIM: Example of Making .tfrec files 512x512 (kernel)</a></p>\n\n<p>In this notebook, I follow very closely the approach outlined in <a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0\">this lab</a>. The lab was created by Martin Görner as Part 1 of his Keras on TPU series. You may also want to take a look at Chris Deotte's public notebook <a href=\"https://www.kaggle.com/cdeotte/how-to-create-tfrecords\">here</a>.</p>\n\n<p>The output of this notebook, is a set of 20 .tfrec files which are tailor-made for Startified Group K-fold cross-validation with 5 folds (4 .tfrec files per fold). The .tfrec files contain not only information about images but also all other tabular data present in <code>train</code>. The categorical data such as sex and anatomic site are one-hot encoded for further processing. The age feature is scaled with sklearn <code>StandardScaler()</code>.</p>\n\n<p>When working with TPU, it is more convenient to work with Kaggle datasets rather than with the outputs of Kaggle kernels. For this reason, I saved the .tfrec files for the training data as the  following public dataset: </p>\n\n<p><a href=\"https://www.kaggle.com/graf10a/siim-512x512-tfrec-q95\">SIIM 512x512 .tfrec train (dataset)</a></p>\n\n<p>And here are the kernel and the corresponding dataset for the test data:</p>\n\n<p><a href=\"https://www.kaggle.com/graf10a/siim-tfrec-512-test-q95\">SIIM 512x512 .tfrec Test (kernel)</a>\n<a href=\"https://www.kaggle.com/graf10a/siim-512x512-tfrec-q95-test\">SIIM 512x512 .tfrec Test (dataset)</a></p>\n\n<h3>5-fold cross-validation with EfficientNet B0</h3>\n\n<p>Now we have everything that we need to build a successful model! The following kernel gives you an example of how such a model can be implemented using EfficientNet B0 as a backbone. </p>\n\n<p><a href=\"https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512\">EfficinetNet-B0 Tensor Flow CV5 Tabular Features</a> (B0 is used in Version 9 of the notebook)</p>\n\n<p>You may notice that to some extend this model is based on the <a href=\"https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\">following public kernel</a>. \nI also borrowed a lot of ideas from <a href=\"https://www.kaggle.com/ragnar123/4-kfold-densenet201\">ragnar's Flower Classification notebook</a>. It is obviously not the best model and it may be improved in many different ways. But it converges and it gives a reasonable cross-validation score. In my opinion, it is a good place to start -- you can easily add any extra features that you like.</p>\n\n<p>By the way, the notebook can be run on Colab as well. If you decide to do that you will have to make some minor adjustments such as setting <code>colab=1</code> and changing the path to your Google Drive folder (or something similar) where you would like to store the inputs and outputs (<code>train</code>, <code>test</code>, the submission file, and out of fold predictions). See the following parameters: <code>PATH</code>, <code>SAVE_FOLDER</code>, <code>OUT_FOLDER</code>. They will require some adjustments. But it might be worth it. On Kaggle, you can run a TPU kernel continuously for only 3 hours. The free version of Colab allows you to run your model for 12 hours on TPU; if you pay an extra $10/month you can upgrade to Colab Pro extending your time limit to 24 hours! With Colab Pro, you can get up to 36 Gb of memory (versus 16 Gb allocated for a Kaggle kernel).</p>\n\n<p>Also, I observed that for some reason, when you run this notebook on Kaggle, some of the training epochs get missing from the output listing. I do not know why it is happening but this does not occur on Colab. Even on Kaggle the notebook is executed successfully and the reported cross-validation score looks reasonable.</p>\n\n<p><em>Disclaimer:</em> I am still not sure if I am going to officially participate in this competition, so I have not submitted the output of the model yet (if you do, please let me know the result!). For this reason, I have no idea how well/poor it performs on the public leaderboard. The purpose of this kernel is not to provide an easy way to get to a high position on the public LB but to illustrate the workflow and to give you a solid starting point in this competition.</p>\n\n<p>Enjoy! And please do not forget to kindly upvote my public kernels/datasets if you find them helpful.</p>",
  "messages": [
    {
      "id": 885428,
      "postDate": "2020-06-14T07:55:39.343Z",
      "content": "<h2>Tensor Flow on TPU: from making .tfrec files to 5-fold cross-validation. Step-by-step guide.</h2>\n\n<p>UPDATE: I have found an annoying bug in version 5 of <a href=\"https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512\">my EfficientNet B0 public notebook</a> which resulted in a very low public LB score. The bug is fixed now.</p>\n\n<h3>Motivation</h3>\n\n<p>First of all, I must confess that I am a total newbie in computer vision and I really enjoy this competition because I view it as a great opportunity to learn new things, such as working with TPU in Tensor Flow. If you are coming from the Machine Learning background, a lot of things should be familiar to you. But the devil is always in details and certain things that you know how to implement with traditional sklearn library might get a bit tricky in Tensor Flow. In this discussion topic, I would like to share with you my experience so far. If you read on, ultimately, you will see an example of a working machine learning pipeline with Tensor Flow on TPU which will include \n* making Startified Group K-folds for 5-fold cross-validation, \n* making .tfrec files that include not only the image data but all other tabular data present in <code>train</code> and <code>test</code>,\n* example of a neural network that utilizes both the image and tabular data (with the cross-validation outlined above).</p>\n\n<p>Hopefully, this is going to save you some time and help you get started with the competition or improve your existing model. </p>\n\n<h3>Startified Group K-folds</h3>\n\n<p>The reasons why we need to use Stratified Group K-Folds in this competition are summarized in <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156002\">this discussion topic</a>. Some time ago, I created <a href=\"https://www.kaggle.com/graf10a/siim-stratified-groupkfold-5-folds\">this public kernel</a> showing how Stratified Group K-Folds can be implemented on the competition data.</p>\n\n<h3>Making .tfrec files</h3>\n\n<p>The most effective way of working with TPU is by feeding it .tfrec files instead of individual JPEG files. A single .tfrec file may contain hundreds or even thousands of JPEG images and feeding a single file like this to TPU can save you a lot of bandwidth, so you don't have to wait too long for your data to be transferred from the storage to the TPU unit as would be the case if you had to transfer the JPEG files individually, one-by-one. To help you to get started with .tfrec files I created the following public notebook giving an example of how these files can be made:</p>\n\n<p><a href=\"https://www.kaggle.com/graf10a/siim-example-of-making-tfrec-files-512x512\">SIIM: Example of Making .tfrec files 512x512 (kernel)</a></p>\n\n<p>In this notebook, I follow very closely the approach outlined in <a href=\"https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0\">this lab</a>. The lab was created by Martin Görner as Part 1 of his Keras on TPU series. You may also want to take a look at Chris Deotte's public notebook <a href=\"https://www.kaggle.com/cdeotte/how-to-create-tfrecords\">here</a>.</p>\n\n<p>The output of this notebook, is a set of 20 .tfrec files which are tailor-made for Startified Group K-fold cross-validation with 5 folds (4 .tfrec files per fold). The .tfrec files contain not only information about images but also all other tabular data present in <code>train</code>. The categorical data such as sex and anatomic site are one-hot encoded for further processing. The age feature is scaled with sklearn <code>StandardScaler()</code>.</p>\n\n<p>When working with TPU, it is more convenient to work with Kaggle datasets rather than with the outputs of Kaggle kernels. For this reason, I saved the .tfrec files for the training data as the  following public dataset: </p>\n\n<p><a href=\"https://www.kaggle.com/graf10a/siim-512x512-tfrec-q95\">SIIM 512x512 .tfrec train (dataset)</a></p>\n\n<p>And here are the kernel and the corresponding dataset for the test data:</p>\n\n<p><a href=\"https://www.kaggle.com/graf10a/siim-tfrec-512-test-q95\">SIIM 512x512 .tfrec Test (kernel)</a>\n<a href=\"https://www.kaggle.com/graf10a/siim-512x512-tfrec-q95-test\">SIIM 512x512 .tfrec Test (dataset)</a></p>\n\n<h3>5-fold cross-validation with EfficientNet B0</h3>\n\n<p>Now we have everything that we need to build a successful model! The following kernel gives you an example of how such a model can be implemented using EfficientNet B0 as a backbone. </p>\n\n<p><a href=\"https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512\">EfficinetNet-B0 Tensor Flow CV5 Tabular Features</a> (B0 is used in Version 9 of the notebook)</p>\n\n<p>You may notice that to some extend this model is based on the <a href=\"https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\">following public kernel</a>. \nI also borrowed a lot of ideas from <a href=\"https://www.kaggle.com/ragnar123/4-kfold-densenet201\">ragnar's Flower Classification notebook</a>. It is obviously not the best model and it may be improved in many different ways. But it converges and it gives a reasonable cross-validation score. In my opinion, it is a good place to start -- you can easily add any extra features that you like.</p>\n\n<p>By the way, the notebook can be run on Colab as well. If you decide to do that you will have to make some minor adjustments such as setting <code>colab=1</code> and changing the path to your Google Drive folder (or something similar) where you would like to store the inputs and outputs (<code>train</code>, <code>test</code>, the submission file, and out of fold predictions). See the following parameters: <code>PATH</code>, <code>SAVE_FOLDER</code>, <code>OUT_FOLDER</code>. They will require some adjustments. But it might be worth it. On Kaggle, you can run a TPU kernel continuously for only 3 hours. The free version of Colab allows you to run your model for 12 hours on TPU; if you pay an extra $10/month you can upgrade to Colab Pro extending your time limit to 24 hours! With Colab Pro, you can get up to 36 Gb of memory (versus 16 Gb allocated for a Kaggle kernel).</p>\n\n<p>Also, I observed that for some reason, when you run this notebook on Kaggle, some of the training epochs get missing from the output listing. I do not know why it is happening but this does not occur on Colab. Even on Kaggle the notebook is executed successfully and the reported cross-validation score looks reasonable.</p>\n\n<p><em>Disclaimer:</em> I am still not sure if I am going to officially participate in this competition, so I have not submitted the output of the model yet (if you do, please let me know the result!). For this reason, I have no idea how well/poor it performs on the public leaderboard. The purpose of this kernel is not to provide an easy way to get to a high position on the public LB but to illustrate the workflow and to give you a solid starting point in this competition.</p>\n\n<p>Enjoy! And please do not forget to kindly upvote my public kernels/datasets if you find them helpful.</p>",
      "rawMarkdown": "## Tensor Flow on TPU: from making .tfrec files to 5-fold cross-validation. Step-by-step guide.\n\nUPDATE: I have found an annoying bug in version 5 of [my EfficientNet B0 public notebook](https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512) which resulted in a very low public LB score. The bug is fixed now.\n\n### Motivation\n\nFirst of all, I must confess that I am a total newbie in computer vision and I really enjoy this competition because I view it as a great opportunity to learn new things, such as working with TPU in Tensor Flow. If you are coming from the Machine Learning background, a lot of things should be familiar to you. But the devil is always in details and certain things that you know how to implement with traditional sklearn library might get a bit tricky in Tensor Flow. In this discussion topic, I would like to share with you my experience so far. If you read on, ultimately, you will see an example of a working machine learning pipeline with Tensor Flow on TPU which will include \n* making Startified Group K-folds for 5-fold cross-validation, \n* making .tfrec files that include not only the image data but all other tabular data present in `train` and `test`,\n* example of a neural network that utilizes both the image and tabular data (with the cross-validation outlined above).\n\nHopefully, this is going to save you some time and help you get started with the competition or improve your existing model. \n\n### Startified Group K-folds\n\nThe reasons why we need to use Stratified Group K-Folds in this competition are summarized in [this discussion topic](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156002). Some time ago, I created [this public kernel](https://www.kaggle.com/graf10a/siim-stratified-groupkfold-5-folds) showing how Stratified Group K-Folds can be implemented on the competition data.\n\n### Making .tfrec files\n\nThe most effective way of working with TPU is by feeding it .tfrec files instead of individual JPEG files. A single .tfrec file may contain hundreds or even thousands of JPEG images and feeding a single file like this to TPU can save you a lot of bandwidth, so you don't have to wait too long for your data to be transferred from the storage to the TPU unit as would be the case if you had to transfer the JPEG files individually, one-by-one. To help you to get started with .tfrec files I created the following public notebook giving an example of how these files can be made:\n\n[SIIM: Example of Making .tfrec files 512x512 (kernel)](https://www.kaggle.com/graf10a/siim-example-of-making-tfrec-files-512x512)\n\nIn this notebook, I follow very closely the approach outlined in [this lab](https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0). The lab was created by Martin Görner as Part 1 of his Keras on TPU series. You may also want to take a look at Chris Deotte's public notebook [here](https://www.kaggle.com/cdeotte/how-to-create-tfrecords).\n\nThe output of this notebook, is a set of 20 .tfrec files which are tailor-made for Startified Group K-fold cross-validation with 5 folds (4 .tfrec files per fold). The .tfrec files contain not only information about images but also all other tabular data present in `train`. The categorical data such as sex and anatomic site are one-hot encoded for further processing. The age feature is scaled with sklearn `StandardScaler()`.\n\nWhen working with TPU, it is more convenient to work with Kaggle datasets rather than with the outputs of Kaggle kernels. For this reason, I saved the .tfrec files for the training data as the  following public dataset: \n\n[SIIM 512x512 .tfrec train (dataset)](https://www.kaggle.com/graf10a/siim-512x512-tfrec-q95)\n\nAnd here are the kernel and the corresponding dataset for the test data:\n\n[SIIM 512x512 .tfrec Test (kernel)](https://www.kaggle.com/graf10a/siim-tfrec-512-test-q95)\n[SIIM 512x512 .tfrec Test (dataset)](https://www.kaggle.com/graf10a/siim-512x512-tfrec-q95-test)\n\n### 5-fold cross-validation with EfficientNet B0\n\nNow we have everything that we need to build a successful model! The following kernel gives you an example of how such a model can be implemented using EfficientNet B0 as a backbone. \n\n[EfficinetNet-B0 Tensor Flow CV5 Tabular Features](https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512) (B0 is used in Version 9 of the notebook)\n\nYou may notice that to some extend this model is based on the [following public kernel](https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head). \nI also borrowed a lot of ideas from [ragnar's Flower Classification notebook](https://www.kaggle.com/ragnar123/4-kfold-densenet201). It is obviously not the best model and it may be improved in many different ways. But it converges and it gives a reasonable cross-validation score. In my opinion, it is a good place to start -- you can easily add any extra features that you like.\n\nBy the way, the notebook can be run on Colab as well. If you decide to do that you will have to make some minor adjustments such as setting `colab=1` and changing the path to your Google Drive folder (or something similar) where you would like to store the inputs and outputs (`train`, `test`, the submission file, and out of fold predictions). See the following parameters: `PATH`, `SAVE_FOLDER`, `OUT_FOLDER`. They will require some adjustments. But it might be worth it. On Kaggle, you can run a TPU kernel continuously for only 3 hours. The free version of Colab allows you to run your model for 12 hours on TPU; if you pay an extra $10/month you can upgrade to Colab Pro extending your time limit to 24 hours! With Colab Pro, you can get up to 36 Gb of memory (versus 16 Gb allocated for a Kaggle kernel).\n\nAlso, I observed that for some reason, when you run this notebook on Kaggle, some of the training epochs get missing from the output listing. I do not know why it is happening but this does not occur on Colab. Even on Kaggle the notebook is executed successfully and the reported cross-validation score looks reasonable.\n\n*Disclaimer:* I am still not sure if I am going to officially participate in this competition, so I have not submitted the output of the model yet (if you do, please let me know the result!). For this reason, I have no idea how well/poor it performs on the public leaderboard. The purpose of this kernel is not to provide an easy way to get to a high position on the public LB but to illustrate the workflow and to give you a solid starting point in this competition.\n\nEnjoy! And please do not forget to kindly upvote my public kernels/datasets if you find them helpful.\n",
      "votes": 47
    },
    {
      "id": 886368,
      "postDate": "2020-06-15T01:43:03.630Z",
      "content": "<p>Thanks Alexey, this is very helpful information on how to build a pipeline that uses both image and tabular data.</p>",
      "rawMarkdown": "Thanks Alexey, this is very helpful information on how to build a pipeline that uses both image and tabular data.",
      "votes": 5
    },
    {
      "id": 910522,
      "postDate": "2020-07-01T07:36:15.150Z",
      "content": "<p>The notebook was really helpful !</p>",
      "rawMarkdown": "The notebook was really helpful !",
      "votes": 1
    },
    {
      "id": 907393,
      "postDate": "2020-06-29T23:43:37.823Z",
      "content": "<p>Thanks <a href=\"/graf10a\">@graf10a</a>  this is a great starting point, I have one question, you mention that for Kaggle is more convenient to use KaggleDatasets, why would that be, is there any performance upgrade, or is just a matter of making things easier?</p>",
      "rawMarkdown": "Thanks @graf10a  this is a great starting point, I have one question, you mention that for Kaggle is more convenient to use KaggleDatasets, why would that be, is there any performance upgrade, or is just a matter of making things easier?",
      "votes": 1,
      "replies": [
        {
          "id": 907404,
          "postDate": "2020-06-30T00:06:39.963Z",
          "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> Great question! You see, the addresses of the GCS buckets are not set in stone -- they change after a few days. <code>KaggleDatasets()</code> utility does not care about this change because it can read the addresses using the data set name that you provide in your code. So, it makes sense to use this utility whenever it is possible. Otherwise, you will have to manually change the addresses every 5-7 days (this is what I have to do because I do all my training on Colab). </p>",
          "rawMarkdown": "@dimitreoliveira Great question! You see, the addresses of the GCS buckets are not set in stone -- they change after a few days. `KaggleDatasets()` utility does not care about this change because it can read the addresses using the data set name that you provide in your code. So, it makes sense to use this utility whenever it is possible. Otherwise, you will have to manually change the addresses every 5-7 days (this is what I have to do because I do all my training on Colab). ",
          "votes": 1
        },
        {
          "id": 907416,
          "postDate": "2020-06-30T00:25:18.613Z",
          "content": "<p>Thanks <a href=\"/graf10a\">@graf10a</a> , I did not know this, it makes a lot of sense, by what I meant is, if there is any gain in performance by using <code>KaggleDatasets</code> to load data or just add to your notebook the dataset with the <code>tfrecords</code> as an additional source. One improvement that I see here is in the case that you also use Colab, this way by using <code>KaggleDatasets</code> you have one source to load data here and on colab.</p>",
          "rawMarkdown": "Thanks @graf10a , I did not know this, it makes a lot of sense, by what I meant is, if there is any gain in performance by using `KaggleDatasets` to load data or just add to your notebook the dataset with the `tfrecords` as an additional source. One improvement that I see here is in the case that you also use Colab, this way by using `KaggleDatasets` you have one source to load data here and on colab.",
          "votes": 1
        },
        {
          "id": 907430,
          "postDate": "2020-06-30T00:45:16.643Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 907438,
          "postDate": "2020-06-30T00:54:38.227Z",
          "content": "<p><a href=\"/d2181664\">@d2181664</a> Sorry, I misunderstood your question. I believe that  in order to use <code>KaggleDataset()</code>  on Kaggle you need to add your tfrecords data to your kernel. So, it is not one or the other -- you have to do both. And the reason you need to use this utility is because you want to point your TPU to the Google Cloud Storage location where the data are stored. GCS is capable of a fetching large amounts of data very fast,  so if you do it this way, your TPU won't be spending most of its time waiting for a new batch of data to be read from your hard drive.  So, it is a performance rather than convenience issue.</p>",
          "rawMarkdown": "@d2181664 Sorry, I misunderstood your question. I believe that  in order to use `KaggleDataset()`  on Kaggle you need to add your tfrecords data to your kernel. So, it is not one or the other -- you have to do both. And the reason you need to use this utility is because you want to point your TPU to the Google Cloud Storage location where the data are stored. GCS is capable of a fetching large amounts of data very fast,  so if you do it this way, your TPU won't be spending most of its time waiting for a new batch of data to be read from your hard drive.  So, it is a performance rather than convenience issue.",
          "votes": 1
        },
        {
          "id": 907461,
          "postDate": "2020-06-30T01:27:44.733Z",
          "content": "<p>Oh now I see it, thanks <a href=\"/graf10a\">@graf10a</a> , it is clear now.</p>",
          "rawMarkdown": "Oh now I see it, thanks @graf10a , it is clear now.",
          "votes": 1
        }
      ]
    },
    {
      "id": 886679,
      "postDate": "2020-06-15T07:41:30.743Z",
      "content": "<p>Thanks for consolidating everything in one place. Really helpful work.</p>",
      "rawMarkdown": "Thanks for consolidating everything in one place. Really helpful work.",
      "votes": 1
    },
    {
      "id": 886112,
      "postDate": "2020-06-14T17:56:27.850Z",
      "content": "<p>Great man this is helpfull!. I believe that in the fit method you can use verbosity = 2 and you will resolve missing epochs.</p>",
      "rawMarkdown": "Great man this is helpfull!. I believe that in the fit method you can use verbosity = 2 and you will resolve missing epochs.",
      "votes": 1,
      "replies": [
        {
          "id": 886263,
          "postDate": "2020-06-14T20:57:12.990Z",
          "content": "<p><a href=\"/ragnar123\">@ragnar123</a> Thank you for the suggestion -- I will try it. BTW, thank you for your great Flower Competition notebook -- it was very helpful (the link is now in the post). </p>",
          "rawMarkdown": "@ragnar123 Thank you for the suggestion -- I will try it. BTW, thank you for your great Flower Competition notebook -- it was very helpful (the link is now in the post). ",
          "votes": 1
        },
        {
          "id": 886383,
          "postDate": "2020-06-15T02:22:06.313Z",
          "content": "<p>This competition is going to be fun! </p>",
          "rawMarkdown": "This competition is going to be fun! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 910688,
      "postDate": "2020-07-01T09:59:02.353Z",
      "content": "<p>Could anyone help me with the GPU I really don't know how to use it. Each time I turn it on,  instead of using the GPU it still uses a CPU.</p>",
      "rawMarkdown": "Could anyone help me with the GPU I really don't know how to use it. Each time I turn it on,  instead of using the GPU it still uses a CPU.",
      "votes": 2,
      "replies": [
        {
          "id": 911154,
          "postDate": "2020-07-01T15:40:48.143Z",
          "content": "<p>In principle, the approach used in <a href=\"https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512\">my kernel</a> should work for both TPU and GPU. The main trick is that you need to define your model inside <code>strategy.scope()</code>. See the definition of the <code>get_model()</code> function (<code>In [47]</code>) in my notebook.</p>\n\n<p>P.S. Of course, you need to define the <code>strategy</code> first (CPU, GPU, TPU) -- see <code>In [12]</code>.</p>",
          "rawMarkdown": "In principle, the approach used in [my kernel](https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512) should work for both TPU and GPU. The main trick is that you need to define your model inside `strategy.scope()`. See the definition of the `get_model()` function (`In [47]`) in my notebook.\n\nP.S. Of course, you need to define the `strategy` first (CPU, GPU, TPU) -- see `In [12]`.",
          "votes": 1
        }
      ]
    },
    {
      "id": 913275,
      "postDate": "2020-07-03T05:16:13.223Z",
      "content": "<p>This was very helpful. Thanks </p>",
      "rawMarkdown": "This was very helpful. Thanks ",
      "votes": 1
    },
    {
      "id": 912919,
      "postDate": "2020-07-02T20:13:59.190Z",
      "content": "<p>Thanks, that was really helpful</p>",
      "rawMarkdown": "Thanks, that was really helpful",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 886368,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-15T01:43:03.630000",
      "content": "<p>Thanks Alexey, this is very helpful information on how to build a pipeline that uses both image and tabular data.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 910522,
      "author_name": "Nitesh Chaudhry",
      "author_url": "",
      "post_date": "2020-07-01T07:36:15.150000",
      "content": "<p>The notebook was really helpful !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 907393,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-06-29T23:43:37.823000",
      "content": "<p>Thanks <a href=\"/graf10a\">@graf10a</a>  this is a great starting point, I have one question, you mention that for Kaggle is more convenient to use KaggleDatasets, why would that be, is there any performance upgrade, or is just a matter of making things easier?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 907404,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-06-30T00:06:39.963000",
          "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> Great question! You see, the addresses of the GCS buckets are not set in stone -- they change after a few days. <code>KaggleDatasets()</code> utility does not care about this change because it can read the addresses using the data set name that you provide in your code. So, it makes sense to use this utility whenever it is possible. Otherwise, you will have to manually change the addresses every 5-7 days (this is what I have to do because I do all my training on Colab). </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 907416,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2020-06-30T00:25:18.613000",
          "content": "<p>Thanks <a href=\"/graf10a\">@graf10a</a> , I did not know this, it makes a lot of sense, by what I meant is, if there is any gain in performance by using <code>KaggleDatasets</code> to load data or just add to your notebook the dataset with the <code>tfrecords</code> as an additional source. One improvement that I see here is in the case that you also use Colab, this way by using <code>KaggleDatasets</code> you have one source to load data here and on colab.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 907430,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-30T00:45:16.643000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 907438,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-06-30T00:54:38.227000",
          "content": "<p><a href=\"/d2181664\">@d2181664</a> Sorry, I misunderstood your question. I believe that  in order to use <code>KaggleDataset()</code>  on Kaggle you need to add your tfrecords data to your kernel. So, it is not one or the other -- you have to do both. And the reason you need to use this utility is because you want to point your TPU to the Google Cloud Storage location where the data are stored. GCS is capable of a fetching large amounts of data very fast,  so if you do it this way, your TPU won't be spending most of its time waiting for a new batch of data to be read from your hard drive.  So, it is a performance rather than convenience issue.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 907461,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2020-06-30T01:27:44.733000",
          "content": "<p>Oh now I see it, thanks <a href=\"/graf10a\">@graf10a</a> , it is clear now.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 886679,
      "author_name": "Karan",
      "author_url": "",
      "post_date": "2020-06-15T07:41:30.743000",
      "content": "<p>Thanks for consolidating everything in one place. Really helpful work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 886112,
      "author_name": "Martin Kovacevic Buvinic",
      "author_url": "",
      "post_date": "2020-06-14T17:56:27.850000",
      "content": "<p>Great man this is helpfull!. I believe that in the fit method you can use verbosity = 2 and you will resolve missing epochs.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 886263,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-06-14T20:57:12.990000",
          "content": "<p><a href=\"/ragnar123\">@ragnar123</a> Thank you for the suggestion -- I will try it. BTW, thank you for your great Flower Competition notebook -- it was very helpful (the link is now in the post). </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 886383,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2020-06-15T02:22:06.313000",
          "content": "<p>This competition is going to be fun! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 910688,
      "author_name": "Jonathan Mbatuegwu",
      "author_url": "",
      "post_date": "2020-07-01T09:59:02.353000",
      "content": "<p>Could anyone help me with the GPU I really don't know how to use it. Each time I turn it on,  instead of using the GPU it still uses a CPU.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 911154,
          "author_name": "Alexey Pronin",
          "author_url": "",
          "post_date": "2020-07-01T15:40:48.143000",
          "content": "<p>In principle, the approach used in <a href=\"https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512\">my kernel</a> should work for both TPU and GPU. The main trick is that you need to define your model inside <code>strategy.scope()</code>. See the definition of the <code>get_model()</code> function (<code>In [47]</code>) in my notebook.</p>\n\n<p>P.S. Of course, you need to define the <code>strategy</code> first (CPU, GPU, TPU) -- see <code>In [12]</code>.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 913275,
      "author_name": "Shreyansh Behani",
      "author_url": "",
      "post_date": "2020-07-03T05:16:13.223000",
      "content": "<p>This was very helpful. Thanks </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 912919,
      "author_name": "Jonathan Mbatuegwu",
      "author_url": "",
      "post_date": "2020-07-02T20:13:59.190000",
      "content": "<p>Thanks, that was really helpful</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "885428": "## Tensor Flow on TPU: from making .tfrec files to 5-fold cross-validation. Step-by-step guide.\n\nUPDATE: I have found an annoying bug in version 5 of [my EfficientNet B0 public notebook](https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512) which resulted in a very low public LB score. The bug is fixed now.\n\n### Motivation\n\nFirst of all, I must confess that I am a total newbie in computer vision and I really enjoy this competition because I view it as a great opportunity to learn new things, such as working with TPU in Tensor Flow. If you are coming from the Machine Learning background, a lot of things should be familiar to you. But the devil is always in details and certain things that you know how to implement with traditional sklearn library might get a bit tricky in Tensor Flow. In this discussion topic, I would like to share with you my experience so far. If you read on, ultimately, you will see an example of a working machine learning pipeline with Tensor Flow on TPU which will include \n* making Startified Group K-folds for 5-fold cross-validation, \n* making .tfrec files that include not only the image data but all other tabular data present in `train` and `test`,\n* example of a neural network that utilizes both the image and tabular data (with the cross-validation outlined above).\n\nHopefully, this is going to save you some time and help you get started with the competition or improve your existing model. \n\n### Startified Group K-folds\n\nThe reasons why we need to use Stratified Group K-Folds in this competition are summarized in [this discussion topic](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156002). Some time ago, I created [this public kernel](https://www.kaggle.com/graf10a/siim-stratified-groupkfold-5-folds) showing how Stratified Group K-Folds can be implemented on the competition data.\n\n### Making .tfrec files\n\nThe most effective way of working with TPU is by feeding it .tfrec files instead of individual JPEG files. A single .tfrec file may contain hundreds or even thousands of JPEG images and feeding a single file like this to TPU can save you a lot of bandwidth, so you don't have to wait too long for your data to be transferred from the storage to the TPU unit as would be the case if you had to transfer the JPEG files individually, one-by-one. To help you to get started with .tfrec files I created the following public notebook giving an example of how these files can be made:\n\n[SIIM: Example of Making .tfrec files 512x512 (kernel)](https://www.kaggle.com/graf10a/siim-example-of-making-tfrec-files-512x512)\n\nIn this notebook, I follow very closely the approach outlined in [this lab](https://codelabs.developers.google.com/codelabs/keras-flowers-data/#0). The lab was created by Martin Görner as Part 1 of his Keras on TPU series. You may also want to take a look at Chris Deotte's public notebook [here](https://www.kaggle.com/cdeotte/how-to-create-tfrecords).\n\nThe output of this notebook, is a set of 20 .tfrec files which are tailor-made for Startified Group K-fold cross-validation with 5 folds (4 .tfrec files per fold). The .tfrec files contain not only information about images but also all other tabular data present in `train`. The categorical data such as sex and anatomic site are one-hot encoded for further processing. The age feature is scaled with sklearn `StandardScaler()`.\n\nWhen working with TPU, it is more convenient to work with Kaggle datasets rather than with the outputs of Kaggle kernels. For this reason, I saved the .tfrec files for the training data as the  following public dataset: \n\n[SIIM 512x512 .tfrec train (dataset)](https://www.kaggle.com/graf10a/siim-512x512-tfrec-q95)\n\nAnd here are the kernel and the corresponding dataset for the test data:\n\n[SIIM 512x512 .tfrec Test (kernel)](https://www.kaggle.com/graf10a/siim-tfrec-512-test-q95)\n[SIIM 512x512 .tfrec Test (dataset)](https://www.kaggle.com/graf10a/siim-512x512-tfrec-q95-test)\n\n### 5-fold cross-validation with EfficientNet B0\n\nNow we have everything that we need to build a successful model! The following kernel gives you an example of how such a model can be implemented using EfficientNet B0 as a backbone. \n\n[EfficinetNet-B0 Tensor Flow CV5 Tabular Features](https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512) (B0 is used in Version 9 of the notebook)\n\nYou may notice that to some extend this model is based on the [following public kernel](https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head). \nI also borrowed a lot of ideas from [ragnar's Flower Classification notebook](https://www.kaggle.com/ragnar123/4-kfold-densenet201). It is obviously not the best model and it may be improved in many different ways. But it converges and it gives a reasonable cross-validation score. In my opinion, it is a good place to start -- you can easily add any extra features that you like.\n\nBy the way, the notebook can be run on Colab as well. If you decide to do that you will have to make some minor adjustments such as setting `colab=1` and changing the path to your Google Drive folder (or something similar) where you would like to store the inputs and outputs (`train`, `test`, the submission file, and out of fold predictions). See the following parameters: `PATH`, `SAVE_FOLDER`, `OUT_FOLDER`. They will require some adjustments. But it might be worth it. On Kaggle, you can run a TPU kernel continuously for only 3 hours. The free version of Colab allows you to run your model for 12 hours on TPU; if you pay an extra $10/month you can upgrade to Colab Pro extending your time limit to 24 hours! With Colab Pro, you can get up to 36 Gb of memory (versus 16 Gb allocated for a Kaggle kernel).\n\nAlso, I observed that for some reason, when you run this notebook on Kaggle, some of the training epochs get missing from the output listing. I do not know why it is happening but this does not occur on Colab. Even on Kaggle the notebook is executed successfully and the reported cross-validation score looks reasonable.\n\n*Disclaimer:* I am still not sure if I am going to officially participate in this competition, so I have not submitted the output of the model yet (if you do, please let me know the result!). For this reason, I have no idea how well/poor it performs on the public leaderboard. The purpose of this kernel is not to provide an easy way to get to a high position on the public LB but to illustrate the workflow and to give you a solid starting point in this competition.\n\nEnjoy! And please do not forget to kindly upvote my public kernels/datasets if you find them helpful.\n",
    "886368": "Thanks Alexey, this is very helpful information on how to build a pipeline that uses both image and tabular data.",
    "910522": "The notebook was really helpful !",
    "907393": "Thanks @graf10a  this is a great starting point, I have one question, you mention that for Kaggle is more convenient to use KaggleDatasets, why would that be, is there any performance upgrade, or is just a matter of making things easier?",
    "886679": "Thanks for consolidating everything in one place. Really helpful work.",
    "886112": "Great man this is helpfull!. I believe that in the fit method you can use verbosity = 2 and you will resolve missing epochs.",
    "910688": "Could anyone help me with the GPU I really don't know how to use it. Each time I turn it on,  instead of using the GPU it still uses a CPU.",
    "913275": "This was very helpful. Thanks ",
    "912919": "Thanks, that was really helpful"
  }
}