{
  "id": 222049,
  "title": "how can we join training and validation folders into one ",
  "url": "/competitions/tpu-getting-started/discussion/222049",
  "author_name": "Jithin Varghese",
  "post_date": "2021-02-25T05:42:49.416000",
  "votes": 3,
  "comment_count": 6,
  "views": null,
  "content": "<p>how can we join training and validation folders to one and later split according to our split ratio</p>",
  "messages": [
    {
      "id": 1217484,
      "postDate": "2021-02-25T05:42:49.417Z",
      "content": "<p>how can we join training and validation folders to one and later split according to our split ratio</p>",
      "rawMarkdown": "how can we join training and validation folders to one and later split according to our split ratio",
      "votes": 3
    },
    {
      "id": 1249240,
      "postDate": "2021-03-23T07:36:58.360Z",
      "content": "<p>The data are in the input folder.  I don't think you could move the files or combine the folders. If you really want to put them in the same folder, you could copy them into your /kaggle/working/ directory.  The following is from <a href=\"https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification/edit/run/57501789\" target=\"_blank\">my notebook</a> showing an example of copying the sample_submission.csv over to your working folder.  You could use this command: <code>!gsutil cp</code></p>\n<pre><code># You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved \n#   as output when you create a version using \"Save &amp; Run All\" \n!gsutil cp /kaggle/input/tpu-getting-started/sample_submission.csv /kaggle/working/test.csv\nfor dirpath, _, filenames in os.walk('/kaggle/working'):\n    for filename in filenames:\n        print(os.path.join(dirpath, filename))     \nprint('list of entries contained in /kaggle/working/:',tf.io.gfile.listdir('/kaggle/working'))   \n</code></pre>",
      "rawMarkdown": "The data are in the input folder.  I don't think you could move the files or combine the folders. If you really want to put them in the same folder, you could copy them into your /kaggle/working/ directory.  The following is from [my notebook](https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification/edit/run/57501789) showing an example of copying the sample_submission.csv over to your working folder.  You could use this command: `!gsutil cp`\n\n```\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved \n#   as output when you create a version using \"Save & Run All\" \n!gsutil cp /kaggle/input/tpu-getting-started/sample_submission.csv /kaggle/working/test.csv\nfor dirpath, _, filenames in os.walk('/kaggle/working'):\n    for filename in filenames:\n        print(os.path.join(dirpath, filename))     \nprint('list of entries contained in /kaggle/working/:',tf.io.gfile.listdir('/kaggle/working'))   \n```",
      "votes": 1,
      "replies": [
        {
          "id": 1249583,
          "postDate": "2021-03-23T11:42:20.957Z",
          "content": "<p>Thanks a lot.😃<br>\nI will check them out👍.</p>",
          "rawMarkdown": "Thanks a lot.😃\nI will check them out👍."
        },
        {
          "id": 1262182,
          "postDate": "2021-04-03T23:06:44.340Z",
          "content": "<p>Are you going to use the combined set to train with cross validation?  I think working with the indices are the better approach.    For each fold, I would split the indices of the training set into train_1 and val_1 and the indices of validation set into train_2 and val_2.  Then you could combine train_1 and train_2 indices for training and val_1 and val_2.  Good idea.  You get more images for training this way.  </p>\n<p>I found these external images that we could use for this challenge:<br>\n<a href=\"https://www.kaggle.com/kirillblinov/tf-flower-photo-tfrec\" target=\"_blank\">tf_flower_photo_tfrec (jpeg)</a> created from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> data. Most duplicates removed and images converted to tfrecords for convenience.  Thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for creating and to Kirill Blinov for converting and sharing! They definitely help boost the public score of <a href=\"https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification\" target=\"_blank\">my notebook</a>.  Please check it out! </p>",
          "rawMarkdown": "Are you going to use the combined set to train with cross validation?  I think working with the indices are the better approach.    For each fold, I would split the indices of the training set into train_1 and val_1 and the indices of validation set into train_2 and val_2.  Then you could combine train_1 and train_2 indices for training and val_1 and val_2.  Good idea.  You get more images for training this way.  \n\nI found these external images that we could use for this challenge:\n[tf_flower_photo_tfrec (jpeg)](https://www.kaggle.com/kirillblinov/tf-flower-photo-tfrec) created from @hengck23 data. Most duplicates removed and images converted to tfrecords for convenience.  Thanks to @hengck23 for creating and to Kirill Blinov for converting and sharing! They definitely help boost the public score of [my notebook](https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification).  Please check it out! ",
          "votes": 1
        },
        {
          "id": 1263112,
          "postDate": "2021-04-05T05:11:07.567Z",
          "content": "<p>Hi, I just came across these codes from <a href=\"https://www.tensorflow.org/tutorials/images/transfer_learning\" target=\"_blank\">Tensorflow</a>.  They could make partitioning your combined data a lot easier.  The codes show how to split the val data to create test data.  Good luck!</p>\n<p>As the original dataset doesn't contains a test set, you will create one. To do so, determine how many batches of data are available in the validation set using tf.data.experimental.cardinality, then move 20% of them to a test set.</p>\n<pre><code>val_batches = tf.data.experimental.cardinality(validation_dataset)\ntest_dataset = validation_dataset.take(val_batches // 5)\nvalidation_dataset = validation_dataset.skip(val_batches // 5)\n</code></pre>",
          "rawMarkdown": "Hi, I just came across these codes from [Tensorflow](https://www.tensorflow.org/tutorials/images/transfer_learning).  They could make partitioning your combined data a lot easier.  The codes show how to split the val data to create test data.  Good luck!\n\nAs the original dataset doesn't contains a test set, you will create one. To do so, determine how many batches of data are available in the validation set using tf.data.experimental.cardinality, then move 20% of them to a test set.\n```\nval_batches = tf.data.experimental.cardinality(validation_dataset)\ntest_dataset = validation_dataset.take(val_batches // 5)\nvalidation_dataset = validation_dataset.skip(val_batches // 5)\n```"
        },
        {
          "id": 1263759,
          "postDate": "2021-04-05T16:56:31.393Z",
          "content": "<p>Thank you😊,  Great idea to add  <a href=\"https://www.kaggle.com/kirillblinov/tf-flower-photo-tfrec\" target=\"_blank\">https://www.kaggle.com/kirillblinov/tf-flower-photo-tfrec</a> to the training set. It really improved the training and F1 score.</p>",
          "rawMarkdown": "Thank you😊,  Great idea to add  https://www.kaggle.com/kirillblinov/tf-flower-photo-tfrec to the training set. It really improved the training and F1 score.\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1249226,
      "postDate": "2021-03-23T07:23:50.147Z",
      "content": "<p>Hi, I have never thought of doing the challenge this way.  But, I have a proposal:</p>\n<pre><code>GCS_DS_PATH = KaggleDatasets().get_gcs_path(\"tpu-getting-started\")  # Google Cloud Storage\nprint(GCS_DS_PATH)\nprint('Entries in the bucket:')\n!gsutil ls $GCS_DS_PATH # list items in the bucket \n</code></pre>\n<p>I believe you could do the following to combine the list of file paths for training and validation datasets of the current competition:</p>\n<pre><code>TRAIN_FILENAMES = tf.io.gfile.glob(GCS_DS_PATH + '/tfrecords-jpeg-512x512/train/*.tfrec') \nVAL_FILENAMES   = tf.io.gfile.glob(GCS_DS_PATH + '/tfrecords-jpeg-512x512/val/*.tfrec') \nCOMBINED_FILENAMES   = TRAIN_FILENAMES + VAL_FILENAMES\n</code></pre>\n<p>The indexes of the COMBINED_FILENAMES list ranges from 0 to  len(COMBINED_FILENAMES) - 1.   You could split them into two sets of indexes for the training and the validation sets according to your desired ratio and take it from there.  Note that each file path is for the .tfrec file which may consist different number of flower images (as shown in the last three characters of the tfrec filename).  All but the last training tfrec files consist of 798 images while all validation tfrec files consist of 232 images.</p>\n<p>I hope this help.  </p>",
      "rawMarkdown": "Hi, I have never thought of doing the challenge this way.  But, I have a proposal:\n\n```\nGCS_DS_PATH = KaggleDatasets().get_gcs_path(\"tpu-getting-started\")  # Google Cloud Storage\nprint(GCS_DS_PATH)\nprint('Entries in the bucket:')\n!gsutil ls $GCS_DS_PATH # list items in the bucket \n```\nI believe you could do the following to combine the list of file paths for training and validation datasets of the current competition:\n\n```\nTRAIN_FILENAMES = tf.io.gfile.glob(GCS_DS_PATH + '/tfrecords-jpeg-512x512/train/*.tfrec') \nVAL_FILENAMES   = tf.io.gfile.glob(GCS_DS_PATH + '/tfrecords-jpeg-512x512/val/*.tfrec') \nCOMBINED_FILENAMES   = TRAIN_FILENAMES + VAL_FILENAMES\n\n```\nThe indexes of the COMBINED_FILENAMES list ranges from 0 to  len(COMBINED_FILENAMES) - 1.   You could split them into two sets of indexes for the training and the validation sets according to your desired ratio and take it from there.  Note that each file path is for the .tfrec file which may consist different number of flower images (as shown in the last three characters of the tfrec filename).  All but the last training tfrec files consist of 798 images while all validation tfrec files consist of 232 images.\n\nI hope this help.  \n",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1249240,
      "author_name": "Sau Kha",
      "author_url": "",
      "post_date": "2021-03-23T07:36:58.360000",
      "content": "<p>The data are in the input folder.  I don't think you could move the files or combine the folders. If you really want to put them in the same folder, you could copy them into your /kaggle/working/ directory.  The following is from <a href=\"https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification/edit/run/57501789\" target=\"_blank\">my notebook</a> showing an example of copying the sample_submission.csv over to your working folder.  You could use this command: <code>!gsutil cp</code></p>\n<pre><code># You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved \n#   as output when you create a version using \"Save &amp; Run All\" \n!gsutil cp /kaggle/input/tpu-getting-started/sample_submission.csv /kaggle/working/test.csv\nfor dirpath, _, filenames in os.walk('/kaggle/working'):\n    for filename in filenames:\n        print(os.path.join(dirpath, filename))     \nprint('list of entries contained in /kaggle/working/:',tf.io.gfile.listdir('/kaggle/working'))   \n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 1249583,
          "author_name": "Jithin Varghese",
          "author_url": "",
          "post_date": "2021-03-23T11:42:20.957000",
          "content": "<p>Thanks a lot.😃<br>\nI will check them out👍.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1262182,
          "author_name": "Sau Kha",
          "author_url": "",
          "post_date": "2021-04-03T23:06:44.340000",
          "content": "<p>Are you going to use the combined set to train with cross validation?  I think working with the indices are the better approach.    For each fold, I would split the indices of the training set into train_1 and val_1 and the indices of validation set into train_2 and val_2.  Then you could combine train_1 and train_2 indices for training and val_1 and val_2.  Good idea.  You get more images for training this way.  </p>\n<p>I found these external images that we could use for this challenge:<br>\n<a href=\"https://www.kaggle.com/kirillblinov/tf-flower-photo-tfrec\" target=\"_blank\">tf_flower_photo_tfrec (jpeg)</a> created from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> data. Most duplicates removed and images converted to tfrecords for convenience.  Thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for creating and to Kirill Blinov for converting and sharing! They definitely help boost the public score of <a href=\"https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification\" target=\"_blank\">my notebook</a>.  Please check it out! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1263112,
          "author_name": "Sau Kha",
          "author_url": "",
          "post_date": "2021-04-05T05:11:07.567000",
          "content": "<p>Hi, I just came across these codes from <a href=\"https://www.tensorflow.org/tutorials/images/transfer_learning\" target=\"_blank\">Tensorflow</a>.  They could make partitioning your combined data a lot easier.  The codes show how to split the val data to create test data.  Good luck!</p>\n<p>As the original dataset doesn't contains a test set, you will create one. To do so, determine how many batches of data are available in the validation set using tf.data.experimental.cardinality, then move 20% of them to a test set.</p>\n<pre><code>val_batches = tf.data.experimental.cardinality(validation_dataset)\ntest_dataset = validation_dataset.take(val_batches // 5)\nvalidation_dataset = validation_dataset.skip(val_batches // 5)\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1263759,
          "author_name": "Jithin Varghese",
          "author_url": "",
          "post_date": "2021-04-05T16:56:31.393000",
          "content": "<p>Thank you😊,  Great idea to add  <a href=\"https://www.kaggle.com/kirillblinov/tf-flower-photo-tfrec\" target=\"_blank\">https://www.kaggle.com/kirillblinov/tf-flower-photo-tfrec</a> to the training set. It really improved the training and F1 score.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1249226,
      "author_name": "Sau Kha",
      "author_url": "",
      "post_date": "2021-03-23T07:23:50.147000",
      "content": "<p>Hi, I have never thought of doing the challenge this way.  But, I have a proposal:</p>\n<pre><code>GCS_DS_PATH = KaggleDatasets().get_gcs_path(\"tpu-getting-started\")  # Google Cloud Storage\nprint(GCS_DS_PATH)\nprint('Entries in the bucket:')\n!gsutil ls $GCS_DS_PATH # list items in the bucket \n</code></pre>\n<p>I believe you could do the following to combine the list of file paths for training and validation datasets of the current competition:</p>\n<pre><code>TRAIN_FILENAMES = tf.io.gfile.glob(GCS_DS_PATH + '/tfrecords-jpeg-512x512/train/*.tfrec') \nVAL_FILENAMES   = tf.io.gfile.glob(GCS_DS_PATH + '/tfrecords-jpeg-512x512/val/*.tfrec') \nCOMBINED_FILENAMES   = TRAIN_FILENAMES + VAL_FILENAMES\n</code></pre>\n<p>The indexes of the COMBINED_FILENAMES list ranges from 0 to  len(COMBINED_FILENAMES) - 1.   You could split them into two sets of indexes for the training and the validation sets according to your desired ratio and take it from there.  Note that each file path is for the .tfrec file which may consist different number of flower images (as shown in the last three characters of the tfrec filename).  All but the last training tfrec files consist of 798 images while all validation tfrec files consist of 232 images.</p>\n<p>I hope this help.  </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1217484": "how can we join training and validation folders to one and later split according to our split ratio",
    "1249240": "The data are in the input folder.  I don't think you could move the files or combine the folders. If you really want to put them in the same folder, you could copy them into your /kaggle/working/ directory.  The following is from [my notebook](https://www.kaggle.com/saukha/petals-to-the-metals-flower-classification/edit/run/57501789) showing an example of copying the sample_submission.csv over to your working folder.  You could use this command: `!gsutil cp`\n\n```\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved \n#   as output when you create a version using \"Save & Run All\" \n!gsutil cp /kaggle/input/tpu-getting-started/sample_submission.csv /kaggle/working/test.csv\nfor dirpath, _, filenames in os.walk('/kaggle/working'):\n    for filename in filenames:\n        print(os.path.join(dirpath, filename))     \nprint('list of entries contained in /kaggle/working/:',tf.io.gfile.listdir('/kaggle/working'))   \n```",
    "1249226": "Hi, I have never thought of doing the challenge this way.  But, I have a proposal:\n\n```\nGCS_DS_PATH = KaggleDatasets().get_gcs_path(\"tpu-getting-started\")  # Google Cloud Storage\nprint(GCS_DS_PATH)\nprint('Entries in the bucket:')\n!gsutil ls $GCS_DS_PATH # list items in the bucket \n```\nI believe you could do the following to combine the list of file paths for training and validation datasets of the current competition:\n\n```\nTRAIN_FILENAMES = tf.io.gfile.glob(GCS_DS_PATH + '/tfrecords-jpeg-512x512/train/*.tfrec') \nVAL_FILENAMES   = tf.io.gfile.glob(GCS_DS_PATH + '/tfrecords-jpeg-512x512/val/*.tfrec') \nCOMBINED_FILENAMES   = TRAIN_FILENAMES + VAL_FILENAMES\n\n```\nThe indexes of the COMBINED_FILENAMES list ranges from 0 to  len(COMBINED_FILENAMES) - 1.   You could split them into two sets of indexes for the training and the validation sets according to your desired ratio and take it from there.  Note that each file path is for the .tfrec file which may consist different number of flower images (as shown in the last three characters of the tfrec filename).  All but the last training tfrec files consist of 798 images while all validation tfrec files consist of 232 images.\n\nI hope this help.  \n"
  }
}