{
  "id": 137445,
  "title": "Use TPUs on private dataset",
  "url": "/competitions/flower-classification-with-tpus/discussion/137445",
  "author_name": "",
  "post_date": "2020-03-20T18:56:28.469378500Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I hope thsi is an appropriate place for this topic. It is somewhat related to <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130194\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130194</a>.</p>\n\n<p>From my understanding, TPUs benefit from having data stored on GCS (even more using the TFRecord file format), but will also work without that with a \"normal\" dataset stored on Kaggle? Currently GCS is only supported for public datasets. I can kind of understand that decision, however currently I am not in the position to make that dataset public, afaik.</p>\n\n<p>Currently I load the data into numpy arrays that take about 12 GB of RAM. When fitting a model with this setup, there is no real progress (ETA for one epoch is about 3 hours vs. about 2 minutes on GPU).</p>\n\n<p>Do you have suggestions on how to improve this?</p>\n\n<p>cc <a href=\"/martingorner\">@martingorner</a> </p>",
  "messages": [
    {
      "id": "780942",
      "postDate": "03/20/2020 18:56:28",
      "content": "<p>I hope thsi is an appropriate place for this topic. It is somewhat related to <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130194\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130194</a>.</p>\n\n<p>From my understanding, TPUs benefit from having data stored on GCS (even more using the TFRecord file format), but will also work without that with a \"normal\" dataset stored on Kaggle? Currently GCS is only supported for public datasets. I can kind of understand that decision, however currently I am not in the position to make that dataset public, afaik.</p>\n\n<p>Currently I load the data into numpy arrays that take about 12 GB of RAM. When fitting a model with this setup, there is no real progress (ETA for one epoch is about 3 hours vs. about 2 minutes on GPU).</p>\n\n<p>Do you have suggestions on how to improve this?</p>\n\n<p>cc <a href=\"/martingorner\">@martingorner</a> </p>",
      "rawMarkdown": "I hope thsi is an appropriate place for this topic. It is somewhat related to https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130194.\n\nFrom my understanding, TPUs benefit from having data stored on GCS (even more using the TFRecord file format), but will also work without that with a \"normal\" dataset stored on Kaggle? Currently GCS is only supported for public datasets. I can kind of understand that decision, however currently I am not in the position to make that dataset public, afaik.\n\nCurrently I load the data into numpy arrays that take about 12 GB of RAM. When fitting a model with this setup, there is no real progress (ETA for one epoch is about 3 hours vs. about 2 minutes on GPU).\n\nDo you have suggestions on how to improve this?\n\ncc @martingorner",
      "votes": null
    },
    {
      "id": "783993",
      "postDate": "03/23/2020 21:22:59",
      "content": "<p>Training on numpy arrays should work. Both <code>model.fit(array, ...)</code> and also <code>dataset = tf.data.Dataset.from_tensor_slices(...)</code> then <code>model.fit(dataset)</code>. </p>\n\n<p>To het help, please be more explicit about what you are doing and if possible, share your code.</p>",
      "rawMarkdown": "Training on numpy arrays should work. Both `model.fit(array, ...)` and also `dataset = tf.data.Dataset.from_tensor_slices(...)` then `model.fit(dataset)`. \n\nTo het help, please be more explicit about what you are doing and if possible, share your code.",
      "votes": null
    },
    {
      "id": "784009",
      "postDate": "03/23/2020 21:53:24",
      "content": "<p>Thanks for your answer. Currently I am doing it like you described, but it does not seem to work. Maybe there is some problem with setting up the TPU correctly or there is a huge bottleneck when transferring the array data to the TPU.</p>\n\n<p>I am running a classification task on Voxel based files that get loaded as numpy arrays with 128x128x128 resolution. There are about 3000 files.</p>\n\n<p>I shared the code with you now.</p>",
      "rawMarkdown": "Thanks for your answer. Currently I am doing it like you described, but it does not seem to work. Maybe there is some problem with setting up the TPU correctly or there is a huge bottleneck when transferring the array data to the TPU.\n\nI am running a classification task on Voxel based files that get loaded as numpy arrays with 128x128x128 resolution. There are about 3000 files.\n\nI shared the code with you now.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 783993,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/23/2020 21:22:59",
      "content": "<p>Training on numpy arrays should work. Both <code>model.fit(array, ...)</code> and also <code>dataset = tf.data.Dataset.from_tensor_slices(...)</code> then <code>model.fit(dataset)</code>. </p>\n\n<p>To het help, please be more explicit about what you are doing and if possible, share your code.</p>",
      "votes": null,
      "replies": [
        {
          "id": 784009,
          "author_name": "claell",
          "author_url": "",
          "post_date": "03/23/2020 21:53:24",
          "content": "<p>Thanks for your answer. Currently I am doing it like you described, but it does not seem to work. Maybe there is some problem with setting up the TPU correctly or there is a huge bottleneck when transferring the array data to the TPU.</p>\n\n<p>I am running a classification task on Voxel based files that get loaded as numpy arrays with 128x128x128 resolution. There are about 3000 files.</p>\n\n<p>I shared the code with you now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "780942": "I hope thsi is an appropriate place for this topic. It is somewhat related to https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130194.\n\nFrom my understanding, TPUs benefit from having data stored on GCS (even more using the TFRecord file format), but will also work without that with a \"normal\" dataset stored on Kaggle? Currently GCS is only supported for public datasets. I can kind of understand that decision, however currently I am not in the position to make that dataset public, afaik.\n\nCurrently I load the data into numpy arrays that take about 12 GB of RAM. When fitting a model with this setup, there is no real progress (ETA for one epoch is about 3 hours vs. about 2 minutes on GPU).\n\nDo you have suggestions on how to improve this?\n\ncc @martingorner",
    "783993": "Training on numpy arrays should work. Both `model.fit(array, ...)` and also `dataset = tf.data.Dataset.from_tensor_slices(...)` then `model.fit(dataset)`. \n\nTo het help, please be more explicit about what you are doing and if possible, share your code.",
    "784009": "Thanks for your answer. Currently I am doing it like you described, but it does not seem to work. Maybe there is some problem with setting up the TPU correctly or there is a huge bottleneck when transferring the array data to the TPU.\n\nI am running a classification task on Voxel based files that get loaded as numpy arrays with 128x128x128 resolution. There are about 3000 files.\n\nI shared the code with you now."
  },
  "source": "meta"
}