{
  "id": 214004,
  "title": "ValueError: With n_samples=0, test_size=0.1 and train_size=None, the resulting train set will be empty. Adjust any of the aforementioned parameters.",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/214004",
  "author_name": "",
  "post_date": "2021-01-25T00:10:56.851220Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>After successfully executing my Notebook or this competition for several training runs, recently I started getting the error mentioned in the title of this post. This is the line of code that elicits the error message:</p>\n<pre><code>train_files, val_files = train_test_split(tf.io.gfile.glob(GCP_DATASET + '/train_tfrecords/ld_train*.tfrec'),\n                                                           test_size=0.10, \n                                                           random_state=RANDOM_SEED)\n\ntest_files = tf.io.gfile.glob(GCP_DATASET + '/test_tfrecords/ld_test*.tfrec')\n</code></pre>\n<p>It doesn't appear that I've altered any code blocks prior to this one and everything worked earlier. Based on Stack Overflow (and a rare easy-to-interpret error message), it appears that my train_test_split function is creating an empty training set. My guess is that something is not working on the Kaggle/GCP side. </p>\n<p>I should point out that the 'GCP_DATASET' variable is a direct reference to the GCS bucket storing the data sets as I'm using Colab to train my models. Currently, I have it as:</p>\n<pre><code>GCP_DATASET = 'gs://kds-8a8a0e757020ef17f93b37a540d540ccfa0003dbd9620ed6ef47ea9b'\n</code></pre>",
  "messages": [
    {
      "id": "1168442",
      "postDate": "01/25/2021 00:10:56",
      "content": "<p>After successfully executing my Notebook or this competition for several training runs, recently I started getting the error mentioned in the title of this post. This is the line of code that elicits the error message:</p>\n<pre><code>train_files, val_files = train_test_split(tf.io.gfile.glob(GCP_DATASET + '/train_tfrecords/ld_train*.tfrec'),\n                                                           test_size=0.10, \n                                                           random_state=RANDOM_SEED)\n\ntest_files = tf.io.gfile.glob(GCP_DATASET + '/test_tfrecords/ld_test*.tfrec')\n</code></pre>\n<p>It doesn't appear that I've altered any code blocks prior to this one and everything worked earlier. Based on Stack Overflow (and a rare easy-to-interpret error message), it appears that my train_test_split function is creating an empty training set. My guess is that something is not working on the Kaggle/GCP side. </p>\n<p>I should point out that the 'GCP_DATASET' variable is a direct reference to the GCS bucket storing the data sets as I'm using Colab to train my models. Currently, I have it as:</p>\n<pre><code>GCP_DATASET = 'gs://kds-8a8a0e757020ef17f93b37a540d540ccfa0003dbd9620ed6ef47ea9b'\n</code></pre>",
      "rawMarkdown": "After successfully executing my Notebook or this competition for several training runs, recently I started getting the error mentioned in the title of this post. This is the line of code that elicits the error message:\n\n```\ntrain_files, val_files = train_test_split(tf.io.gfile.glob(GCP_DATASET + '/train_tfrecords/ld_train*.tfrec'),\n                                                           test_size=0.10, \n                                                           random_state=RANDOM_SEED)\n\ntest_files = tf.io.gfile.glob(GCP_DATASET + '/test_tfrecords/ld_test*.tfrec')\n```\n\nIt doesn't appear that I've altered any code blocks prior to this one and everything worked earlier. Based on Stack Overflow (and a rare easy-to-interpret error message), it appears that my train_test_split function is creating an empty training set. My guess is that something is not working on the Kaggle/GCP side. \n\nI should point out that the 'GCP_DATASET' variable is a direct reference to the GCS bucket storing the data sets as I'm using Colab to train my models. Currently, I have it as:\n```\nGCP_DATASET = 'gs://kds-8a8a0e757020ef17f93b37a540d540ccfa0003dbd9620ed6ef47ea9b'\n```",
      "votes": null
    },
    {
      "id": "1169820",
      "postDate": "01/25/2021 18:49:41",
      "content": "<p>Some follow-up on this: It looks like the issue with this is that the GCP_DATASET bucket reference is wrong. Apparently, Kaggle data sets are subject to change buckets -- can anyone confirm this? I was hard-coding in the path to the bucket, but after running the same script on Kaggle Notebooks (not Colab), it ran fine. I checked the GCP_DATASET environmental variable and it was, indeed, different.</p>",
      "rawMarkdown": "Some follow-up on this: It looks like the issue with this is that the GCP_DATASET bucket reference is wrong. Apparently, Kaggle data sets are subject to change buckets -- can anyone confirm this? I was hard-coding in the path to the bucket, but after running the same script on Kaggle Notebooks (not Colab), it ran fine. I checked the GCP_DATASET environmental variable and it was, indeed, different.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1169820,
      "author_name": "jsukup",
      "author_url": "",
      "post_date": "01/25/2021 18:49:41",
      "content": "<p>Some follow-up on this: It looks like the issue with this is that the GCP_DATASET bucket reference is wrong. Apparently, Kaggle data sets are subject to change buckets -- can anyone confirm this? I was hard-coding in the path to the bucket, but after running the same script on Kaggle Notebooks (not Colab), it ran fine. I checked the GCP_DATASET environmental variable and it was, indeed, different.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1168442": "After successfully executing my Notebook or this competition for several training runs, recently I started getting the error mentioned in the title of this post. This is the line of code that elicits the error message:\n\n```\ntrain_files, val_files = train_test_split(tf.io.gfile.glob(GCP_DATASET + '/train_tfrecords/ld_train*.tfrec'),\n                                                           test_size=0.10, \n                                                           random_state=RANDOM_SEED)\n\ntest_files = tf.io.gfile.glob(GCP_DATASET + '/test_tfrecords/ld_test*.tfrec')\n```\n\nIt doesn't appear that I've altered any code blocks prior to this one and everything worked earlier. Based on Stack Overflow (and a rare easy-to-interpret error message), it appears that my train_test_split function is creating an empty training set. My guess is that something is not working on the Kaggle/GCP side. \n\nI should point out that the 'GCP_DATASET' variable is a direct reference to the GCS bucket storing the data sets as I'm using Colab to train my models. Currently, I have it as:\n```\nGCP_DATASET = 'gs://kds-8a8a0e757020ef17f93b37a540d540ccfa0003dbd9620ed6ef47ea9b'\n```",
    "1169820": "Some follow-up on this: It looks like the issue with this is that the GCP_DATASET bucket reference is wrong. Apparently, Kaggle data sets are subject to change buckets -- can anyone confirm this? I was hard-coding in the path to the bucket, but after running the same script on Kaggle Notebooks (not Colab), it ran fine. I checked the GCP_DATASET environmental variable and it was, indeed, different."
  },
  "source": "meta"
}