{
  "id": 191259,
  "title": "Making a TensorFlow dataset",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/191259",
  "author_name": "",
  "post_date": "2020-10-15T14:54:53.539234100Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>It would be very nice to have a code snippet creating a TensorFlow dataset from all the csv files, where each file would be an example. Isn't it?</p>",
  "messages": [
    {
      "id": "1050575",
      "postDate": "10/15/2020 14:54:53",
      "content": "<p>It would be very nice to have a code snippet creating a TensorFlow dataset from all the csv files, where each file would be an example. Isn't it?</p>",
      "rawMarkdown": "It would be very nice to have a code snippet creating a TensorFlow dataset from all the csv files, where each file would be an example. Isn't it?",
      "votes": null
    },
    {
      "id": "1050787",
      "postDate": "10/15/2020 18:16:36",
      "content": "<p>What do you mean? Can you explain it further?</p>",
      "rawMarkdown": "What do you mean? Can you explain it further?",
      "votes": null
    },
    {
      "id": "1051152",
      "postDate": "10/16/2020 08:05:18",
      "content": "<p>In this competition the data that is provided is preprocessed and does not require any preprocessing, only what you want to do with it for training yur model, so its not nice that the only thing you have to do you demand it …..</p>\n<p>If you want to learn to make Datasets this is the best competition for it </p>",
      "rawMarkdown": "In this competition the data that is provided is preprocessed and does not require any preprocessing, only what you want to do with it for training yur model, so its not nice that the only thing you have to do you demand it .....\n\nIf you want to learn to make Datasets this is the best competition for it",
      "votes": null
    },
    {
      "id": "1051182",
      "postDate": "10/16/2020 08:41:48",
      "content": "<p>I would rather spend my time on the model architecture and tuning.</p>",
      "rawMarkdown": "I would rather spend my time on the model architecture and tuning.",
      "votes": null
    },
    {
      "id": "1051716",
      "postDate": "10/16/2020 19:03:18",
      "content": "<p>Not ultra performant, but works on my laptop and creates a numpy array that is trivial to get a TS or keras input.  Do you mean something like this?</p>\n<p>I prefer working with xarrays though and save it as a compressed netcdf file once. Takes longer first, but is easier to work with later, </p>\n<p>`<br>\nimport numpy as np<br>\nimport pandas as pd<br>\nimport os<br>\nimport time</p>\n<p>start = time.time()</p>\n<p>your_path_to_data =  \"/predict-volcanic-eruptions-ingv-oe/train/\"</p>\n<p>directory = os.fsencode(your_path_to_data)<br>\nc = 0<br>\ndata = np.zeros((4431,60001,10))<br>\nfor file in os.listdir(directory):<br>\n       filename = os.fsdecode(os.path.join(directory, file))<br>\n       df = pd.read_csv(filename)<br>\n       data[c,:,:] = df.values<br>\n       c = c + 1</p>\n<p>end = time.time()<br>\nprint(\"start time: \" + str(start) )<br>\nprint(\"end time: \" + str(end) )<br>\nprint(\"elapsed time: \" + str(end - start))</p>\n<p>`</p>",
      "rawMarkdown": "Not ultra performant, but works on my laptop and creates a numpy array that is trivial to get a TS or keras input.  Do you mean something like this?\n\n\nI prefer working with xarrays though and save it as a compressed netcdf file once. Takes longer first, but is easier to work with later, \n\n`\nimport numpy as np\nimport pandas as pd\nimport os\nimport time\n\nstart = time.time()\n\nyour_path_to_data =  \"/predict-volcanic-eruptions-ingv-oe/train/\"\n\ndirectory = os.fsencode(your_path_to_data)\nc = 0\ndata = np.zeros((4431,60001,10))\nfor file in os.listdir(directory):\n       filename = os.fsdecode(os.path.join(directory, file))\n       df = pd.read_csv(filename)\n       data[c,:,:] = df.values\n       c = c + 1\n\nend = time.time()\nprint(\"start time: \" + str(start) )\nprint(\"end time: \" + str(end) )\nprint(\"elapsed time: \" + str(end - start))\n\n\n\n`",
      "votes": null
    },
    {
      "id": "1051757",
      "postDate": "10/16/2020 20:15:01",
      "content": "<p>I think he would like to try different fast experiment on TF with TPU. </p>",
      "rawMarkdown": "I think he would like to try different fast experiment on TF with TPU.",
      "votes": null
    },
    {
      "id": "1059768",
      "postDate": "10/25/2020 12:37:51",
      "content": "<p>If you're not comfortable in creating a TFRecords file, you won't be able to load it and use it with TPUs effectively anyway. It's not easy (or at least it was not for me) so think it is worth spending some time to use kaggle examples etc to learn about it and be able to modify for your own code. There are quite a lot of decent examples on kaggle.</p>\n<p>Otherwise I don't see how you can make progress just by requesting chunks of code because it will just 'break' the moment you need to change anything…</p>\n<p>This comp has examples which I think pretty much demo everything from creating records to using TPUs.<br>\n<a href=\"https://www.kaggle.com/c/tpu-getting-started\" target=\"_blank\">https://www.kaggle.com/c/tpu-getting-started</a></p>",
      "rawMarkdown": "If you're not comfortable in creating a TFRecords file, you won't be able to load it and use it with TPUs effectively anyway. It's not easy (or at least it was not for me) so think it is worth spending some time to use kaggle examples etc to learn about it and be able to modify for your own code. There are quite a lot of decent examples on kaggle.\n\nOtherwise I don't see how you can make progress just by requesting chunks of code because it will just 'break' the moment you need to change anything...\n\nThis comp has examples which I think pretty much demo everything from creating records to using TPUs.\nhttps://www.kaggle.com/c/tpu-getting-started",
      "votes": null
    },
    {
      "id": "1065115",
      "postDate": "10/30/2020 21:35:06",
      "content": "<p>I made two datasets containing the competition data converted to TFRecord files. I also included code to show how to convert it into a tf.data.Dataset.</p>\n<p><strong><a href=\"https://www.kaggle.com/c/predict-volcanic-eruptions-ingv-oe/discussion/194150\" target=\"_blank\">https://www.kaggle.com/c/predict-volcanic-eruptions-ingv-oe/discussion/194150</a></strong></p>\n<p>Hope that helps!</p>",
      "rawMarkdown": "I made two datasets containing the competition data converted to TFRecord files. I also included code to show how to convert it into a tf.data.Dataset.\n\n**https://www.kaggle.com/c/predict-volcanic-eruptions-ingv-oe/discussion/194150**\n\nHope that helps!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1050787,
      "author_name": "rude009",
      "author_url": "",
      "post_date": "10/15/2020 18:16:36",
      "content": "<p>What do you mean? Can you explain it further?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1051757,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "10/16/2020 20:15:01",
          "content": "<p>I think he would like to try different fast experiment on TF with TPU. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1051152,
      "author_name": "enric1296",
      "author_url": "",
      "post_date": "10/16/2020 08:05:18",
      "content": "<p>In this competition the data that is provided is preprocessed and does not require any preprocessing, only what you want to do with it for training yur model, so its not nice that the only thing you have to do you demand it …..</p>\n<p>If you want to learn to make Datasets this is the best competition for it </p>",
      "votes": null,
      "replies": [
        {
          "id": 1051182,
          "author_name": "adrienrenaud",
          "author_url": "",
          "post_date": "10/16/2020 08:41:48",
          "content": "<p>I would rather spend my time on the model architecture and tuning.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1051716,
      "author_name": "pldatacybernetics",
      "author_url": "",
      "post_date": "10/16/2020 19:03:18",
      "content": "<p>Not ultra performant, but works on my laptop and creates a numpy array that is trivial to get a TS or keras input.  Do you mean something like this?</p>\n<p>I prefer working with xarrays though and save it as a compressed netcdf file once. Takes longer first, but is easier to work with later, </p>\n<p>`<br>\nimport numpy as np<br>\nimport pandas as pd<br>\nimport os<br>\nimport time</p>\n<p>start = time.time()</p>\n<p>your_path_to_data =  \"/predict-volcanic-eruptions-ingv-oe/train/\"</p>\n<p>directory = os.fsencode(your_path_to_data)<br>\nc = 0<br>\ndata = np.zeros((4431,60001,10))<br>\nfor file in os.listdir(directory):<br>\n       filename = os.fsdecode(os.path.join(directory, file))<br>\n       df = pd.read_csv(filename)<br>\n       data[c,:,:] = df.values<br>\n       c = c + 1</p>\n<p>end = time.time()<br>\nprint(\"start time: \" + str(start) )<br>\nprint(\"end time: \" + str(end) )<br>\nprint(\"elapsed time: \" + str(end - start))</p>\n<p>`</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1059768,
      "author_name": "davidedwards1",
      "author_url": "",
      "post_date": "10/25/2020 12:37:51",
      "content": "<p>If you're not comfortable in creating a TFRecords file, you won't be able to load it and use it with TPUs effectively anyway. It's not easy (or at least it was not for me) so think it is worth spending some time to use kaggle examples etc to learn about it and be able to modify for your own code. There are quite a lot of decent examples on kaggle.</p>\n<p>Otherwise I don't see how you can make progress just by requesting chunks of code because it will just 'break' the moment you need to change anything…</p>\n<p>This comp has examples which I think pretty much demo everything from creating records to using TPUs.<br>\n<a href=\"https://www.kaggle.com/c/tpu-getting-started\" target=\"_blank\">https://www.kaggle.com/c/tpu-getting-started</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1065115,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "10/30/2020 21:35:06",
      "content": "<p>I made two datasets containing the competition data converted to TFRecord files. I also included code to show how to convert it into a tf.data.Dataset.</p>\n<p><strong><a href=\"https://www.kaggle.com/c/predict-volcanic-eruptions-ingv-oe/discussion/194150\" target=\"_blank\">https://www.kaggle.com/c/predict-volcanic-eruptions-ingv-oe/discussion/194150</a></strong></p>\n<p>Hope that helps!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1050575": "It would be very nice to have a code snippet creating a TensorFlow dataset from all the csv files, where each file would be an example. Isn't it?",
    "1050787": "What do you mean? Can you explain it further?",
    "1051152": "In this competition the data that is provided is preprocessed and does not require any preprocessing, only what you want to do with it for training yur model, so its not nice that the only thing you have to do you demand it .....\n\nIf you want to learn to make Datasets this is the best competition for it",
    "1051182": "I would rather spend my time on the model architecture and tuning.",
    "1051716": "Not ultra performant, but works on my laptop and creates a numpy array that is trivial to get a TS or keras input.  Do you mean something like this?\n\n\nI prefer working with xarrays though and save it as a compressed netcdf file once. Takes longer first, but is easier to work with later, \n\n`\nimport numpy as np\nimport pandas as pd\nimport os\nimport time\n\nstart = time.time()\n\nyour_path_to_data =  \"/predict-volcanic-eruptions-ingv-oe/train/\"\n\ndirectory = os.fsencode(your_path_to_data)\nc = 0\ndata = np.zeros((4431,60001,10))\nfor file in os.listdir(directory):\n       filename = os.fsdecode(os.path.join(directory, file))\n       df = pd.read_csv(filename)\n       data[c,:,:] = df.values\n       c = c + 1\n\nend = time.time()\nprint(\"start time: \" + str(start) )\nprint(\"end time: \" + str(end) )\nprint(\"elapsed time: \" + str(end - start))\n\n\n\n`",
    "1051757": "I think he would like to try different fast experiment on TF with TPU.",
    "1059768": "If you're not comfortable in creating a TFRecords file, you won't be able to load it and use it with TPUs effectively anyway. It's not easy (or at least it was not for me) so think it is worth spending some time to use kaggle examples etc to learn about it and be able to modify for your own code. There are quite a lot of decent examples on kaggle.\n\nOtherwise I don't see how you can make progress just by requesting chunks of code because it will just 'break' the moment you need to change anything...\n\nThis comp has examples which I think pretty much demo everything from creating records to using TPUs.\nhttps://www.kaggle.com/c/tpu-getting-started",
    "1065115": "I made two datasets containing the competition data converted to TFRecord files. I also included code to show how to convert it into a tf.data.Dataset.\n\n**https://www.kaggle.com/c/predict-volcanic-eruptions-ingv-oe/discussion/194150**\n\nHope that helps!"
  },
  "source": "meta"
}