{
  "id": 155697,
  "title": "Confused about TFRecords",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/155697",
  "author_name": "",
  "post_date": "2020-06-02T16:34:04.494016900Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello kaggle,\nI'm a beginner at Deep Learning, Kaggle, Tensorflow and I'm trying out this competition for experience and hopefully learning something new. I've read a few notebooks on this competition and I could understand the ones where they used .jpg files for input, but I really want to use the TFRecord files, and I came across another question regarding that, and it solved most of my initial doubt but I am still confused about <strong>how are we supposed to split the train/test data and also make splits for cross validation? and what do the numbers on the name of the files signify?</strong></p>",
  "messages": [
    {
      "id": "871801",
      "postDate": "06/02/2020 16:34:04",
      "content": "<p>Hello kaggle,\nI'm a beginner at Deep Learning, Kaggle, Tensorflow and I'm trying out this competition for experience and hopefully learning something new. I've read a few notebooks on this competition and I could understand the ones where they used .jpg files for input, but I really want to use the TFRecord files, and I came across another question regarding that, and it solved most of my initial doubt but I am still confused about <strong>how are we supposed to split the train/test data and also make splits for cross validation? and what do the numbers on the name of the files signify?</strong></p>",
      "rawMarkdown": "Hello kaggle,\nI'm a beginner at Deep Learning, Kaggle, Tensorflow and I'm trying out this competition for experience and hopefully learning something new. I've read a few notebooks on this competition and I could understand the ones where they used .jpg files for input, but I really want to use the TFRecord files, and I came across another question regarding that, and it solved most of my initial doubt but I am still confused about **how are we supposed to split the train/test data and also make splits for cross validation? and what do the numbers on the name of the files signify?**",
      "votes": null
    },
    {
      "id": "871834",
      "postDate": "06/02/2020 17:09:02",
      "content": "<p>You split train and test by filename. For example <code>TRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')</code> and <code>TEST_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/test*.tfrec')</code>. You can also make validation sets by choosing specific numbers between 0 and 15. For example </p>\n\n<pre><code>VAL = [0,5,8,12]\nVALIDATION_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train%.2i-2071.tfrec'%VAL[0])\nfor k in VAL[1:]:\n    VALIDATION_FILENAMES += tf.io.gfile.glob(GCS_PATH + '/train%.2i-2071.tfrec'%k)\n</code></pre>\n\n<p>In the filename <code>train00-2071.tfrec</code>, the <code>00</code> says this is 0 out of 16th file. And 2071 says there are 2071 records (images i.e. samples) within.</p>\n\n<p>To learn more, review this Flower Comp notebook <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu\">here</a></p>",
      "rawMarkdown": "You split train and test by filename. For example `TRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')` and `TEST_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/test*.tfrec')`. You can also make validation sets by choosing specific numbers between 0 and 15. For example \n\n    VAL = [0,5,8,12]\n    VALIDATION_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train%.2i-2071.tfrec'%VAL[0])\n    for k in VAL[1:]:\n        VALIDATION_FILENAMES += tf.io.gfile.glob(GCS_PATH + '/train%.2i-2071.tfrec'%k)\n\nIn the filename `train00-2071.tfrec`, the `00` says this is 0 out of 16th file. And 2071 says there are 2071 records (images i.e. samples) within.\n\nTo learn more, review this Flower Comp notebook [here][1]\n\n[1]: https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu",
      "votes": null
    },
    {
      "id": "872352",
      "postDate": "06/03/2020 05:43:04",
      "content": "<p>Thank you so much for the detailed explanation! I think I finally understood what I have to do!</p>",
      "rawMarkdown": "Thank you so much for the detailed explanation! I think I finally understood what I have to do!",
      "votes": null
    },
    {
      "id": "872377",
      "postDate": "06/03/2020 06:17:37",
      "content": "<p>Also, I was wondering that does the given TFRecords in the competition only have two features? Which in case of training, would be <code>image</code> and <code>target</code> ? Because since the distribution of targets in heavily skewed, a <strong>Stratified KFold CV is recommended by Abhishek Thakur, in this Youtube video. So would that be possible with TFRecords ?</strong> I also saw that you uploaded multiple datasets with different sizes, and additional features, so <strong>can you help me on how I would conditionally slice the dataset or records, if that's possible?</strong></p>",
      "rawMarkdown": "Also, I was wondering that does the given TFRecords in the competition only have two features? Which in case of training, would be `image` and `target` ? Because since the distribution of targets in heavily skewed, a **Stratified KFold CV is recommended by Abhishek Thakur, in this Youtube video. So would that be possible with TFRecords ?** I also saw that you uploaded multiple datasets with different sizes, and additional features, so **can you help me on how I would conditionally slice the dataset or records, if that's possible?**",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 871834,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/02/2020 17:09:02",
      "content": "<p>You split train and test by filename. For example <code>TRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')</code> and <code>TEST_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/test*.tfrec')</code>. You can also make validation sets by choosing specific numbers between 0 and 15. For example </p>\n\n<pre><code>VAL = [0,5,8,12]\nVALIDATION_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train%.2i-2071.tfrec'%VAL[0])\nfor k in VAL[1:]:\n    VALIDATION_FILENAMES += tf.io.gfile.glob(GCS_PATH + '/train%.2i-2071.tfrec'%k)\n</code></pre>\n\n<p>In the filename <code>train00-2071.tfrec</code>, the <code>00</code> says this is 0 out of 16th file. And 2071 says there are 2071 records (images i.e. samples) within.</p>\n\n<p>To learn more, review this Flower Comp notebook <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 872352,
          "author_name": "sayantankarmakar",
          "author_url": "",
          "post_date": "06/03/2020 05:43:04",
          "content": "<p>Thank you so much for the detailed explanation! I think I finally understood what I have to do!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 872377,
          "author_name": "sayantankarmakar",
          "author_url": "",
          "post_date": "06/03/2020 06:17:37",
          "content": "<p>Also, I was wondering that does the given TFRecords in the competition only have two features? Which in case of training, would be <code>image</code> and <code>target</code> ? Because since the distribution of targets in heavily skewed, a <strong>Stratified KFold CV is recommended by Abhishek Thakur, in this Youtube video. So would that be possible with TFRecords ?</strong> I also saw that you uploaded multiple datasets with different sizes, and additional features, so <strong>can you help me on how I would conditionally slice the dataset or records, if that's possible?</strong></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "871801": "Hello kaggle,\nI'm a beginner at Deep Learning, Kaggle, Tensorflow and I'm trying out this competition for experience and hopefully learning something new. I've read a few notebooks on this competition and I could understand the ones where they used .jpg files for input, but I really want to use the TFRecord files, and I came across another question regarding that, and it solved most of my initial doubt but I am still confused about **how are we supposed to split the train/test data and also make splits for cross validation? and what do the numbers on the name of the files signify?**",
    "871834": "You split train and test by filename. For example `TRAINING_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')` and `TEST_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/test*.tfrec')`. You can also make validation sets by choosing specific numbers between 0 and 15. For example \n\n    VAL = [0,5,8,12]\n    VALIDATION_FILENAMES = tf.io.gfile.glob(GCS_PATH + '/train%.2i-2071.tfrec'%VAL[0])\n    for k in VAL[1:]:\n        VALIDATION_FILENAMES += tf.io.gfile.glob(GCS_PATH + '/train%.2i-2071.tfrec'%k)\n\nIn the filename `train00-2071.tfrec`, the `00` says this is 0 out of 16th file. And 2071 says there are 2071 records (images i.e. samples) within.\n\nTo learn more, review this Flower Comp notebook [here][1]\n\n[1]: https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu",
    "872352": "Thank you so much for the detailed explanation! I think I finally understood what I have to do!",
    "872377": "Also, I was wondering that does the given TFRecords in the competition only have two features? Which in case of training, would be `image` and `target` ? Because since the distribution of targets in heavily skewed, a **Stratified KFold CV is recommended by Abhishek Thakur, in this Youtube video. So would that be possible with TFRecords ?** I also saw that you uploaded multiple datasets with different sizes, and additional features, so **can you help me on how I would conditionally slice the dataset or records, if that's possible?**"
  },
  "source": "meta"
}