{
  "id": 512438,
  "title": "Different dimension images",
  "url": "/competitions/tpu-getting-started/discussion/512438",
  "author_name": "Umer Tariq",
  "post_date": "2024-06-15T08:16:33.296000",
  "votes": 1,
  "comment_count": 3,
  "views": null,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17225803%2Fe86bb0bc53339600490543458229a4a2%2FScreenshot%202024-06-15%20131721.jpg?generation=1718439475774399&amp;alt=media\" alt=\"Data image\"><br>\nI was looking at the dataset given to us for training the classification model and I saw <strong>4 folders</strong>. Each folders has images of different dimensions which have been split into <strong>train, val and test</strong>.<br>\n<strong>Questions:</strong><br>\n1) Do all the 4 images folders have the <strong>same images</strong> with different dimension to train 4 different models(higher resolution images will give better result but more computationally expensive and time consuming)?<br>\n2) There is also a <strong>summary dataset</strong> below. What is that about. Is it the same thing?</p>\n<p>It would be very helpful if you answer the questions<br>\nRegards <br>\n<strong>Umer Tariq</strong></p>",
  "messages": [
    {
      "id": 3033634,
      "postDate": "2024-11-01T11:27:26.037Z",
      "content": "<p>hi, for the point 1) yes I confirm, you can use different image resolutions. You can also use this dataset and after downsample images to lower resolutions:</p>\n<p><a href=\"https://www.kaggle.com/datasets/alenic/flower-classification-512-png\" target=\"_blank\">https://www.kaggle.com/datasets/alenic/flower-classification-512-png</a></p>",
      "rawMarkdown": "hi, for the point 1) yes I confirm, you can use different image resolutions. You can also use this dataset and after downsample images to lower resolutions:\n\nhttps://www.kaggle.com/datasets/alenic/flower-classification-512-png\n\n",
      "votes": 1
    },
    {
      "id": 2981067,
      "postDate": "2024-09-06T13:16:57.387Z",
      "content": "<p>Hey dude, there are four different datasets for the given competition and it means that you can work on any of the given datasets and provide the results. They are just different in resolutions. </p>\n<p>Also, the sample_submission.csv isn't a summary dataset, its the dataset that you are required to submit your results on to the competition.</p>",
      "rawMarkdown": "Hey dude, there are four different datasets for the given competition and it means that you can work on any of the given datasets and provide the results. They are just different in resolutions. \n\nAlso, the sample_submission.csv isn't a summary dataset, its the dataset that you are required to submit your results on to the competition.",
      "votes": 1
    },
    {
      "id": 2872997,
      "postDate": "2024-06-15T08:16:33.297Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17225803%2Fe86bb0bc53339600490543458229a4a2%2FScreenshot%202024-06-15%20131721.jpg?generation=1718439475774399&amp;alt=media\" alt=\"Data image\"><br>\nI was looking at the dataset given to us for training the classification model and I saw <strong>4 folders</strong>. Each folders has images of different dimensions which have been split into <strong>train, val and test</strong>.<br>\n<strong>Questions:</strong><br>\n1) Do all the 4 images folders have the <strong>same images</strong> with different dimension to train 4 different models(higher resolution images will give better result but more computationally expensive and time consuming)?<br>\n2) There is also a <strong>summary dataset</strong> below. What is that about. Is it the same thing?</p>\n<p>It would be very helpful if you answer the questions<br>\nRegards <br>\n<strong>Umer Tariq</strong></p>",
      "rawMarkdown": "![Data image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17225803%2Fe86bb0bc53339600490543458229a4a2%2FScreenshot%202024-06-15%20131721.jpg?generation=1718439475774399&alt=media)\nI was looking at the dataset given to us for training the classification model and I saw **4 folders**. Each folders has images of different dimensions which have been split into **train, val and test**.\n**Questions:**\n1) Do all the 4 images folders have the **same images** with different dimension to train 4 different models(higher resolution images will give better result but more computationally expensive and time consuming)?\n2) There is also a **summary dataset** below. What is that about. Is it the same thing?\n\nIt would be very helpful if you answer the questions\nRegards \n**Umer Tariq**",
      "votes": 1
    },
    {
      "id": 3169770,
      "postDate": "2025-04-03T20:30:38.343Z",
      "content": "<p>The 4 folders contains the same pics with different resolutions !</p>\n<p>with this code</p>\n<blockquote>\n  <p>def count_examples_in_tfrecord(tfrecord_path):<br>\n           \"\"\"Compte le nombre d'images dans un fichier TFRecord.\"\"\"<br>\n           return sum(1 for _ in tf.data.TFRecordDataset(tfrecord_path))</p>\n</blockquote>\n<p>then</p>\n<blockquote>\n  <p>for resolution in ['192x192','224x224','331x331','512x512']:<br>\n      print(f\"\\nResolution {resolution}\")<br>\n      for doss in [\"train\", \"val\", \"test\"]:<br>\n          # Lister tous les fichiers TFRecord<br>\n          tfrecord_files = tf.io.gfile.glob(GCS_DS_PATH + '/tfrecords-jpeg-'+resolution+'/' + doss + '/*.tfrec')<br>\n          # Compter les images dans chaque fichier et faire la somme<br>\n          total_images = sum(count_examples_in_tfrecord(file) for file in tfrecord_files)<br>\n          print(f\"Nombre total d'images dans les TFRecords {doss}: {total_images}\")</p>\n</blockquote>\n<p>you get :</p>\n<blockquote>\n  <p>Resolution 192x192<br>\n  Nombre total d'images dans les TFRecords train: 12753<br>\n  Nombre total d'images dans les TFRecords val: 3712<br>\n  Nombre total d'images dans les TFRecords test: 7382<br>\n  _<br>\n  Resolution 224x224<br>\n  Nombre total d'images dans les TFRecords train: 12753<br>\n  Nombre total d'images dans les TFRecords val: 3712<br>\n  Nombre total d'images dans les TFRecords test: 7382<br>\n  _<br>\n  Resolution 331x331<br>\n  Nombre total d'images dans les TFRecords train: 12753<br>\n  Nombre total d'images dans les TFRecords val: 3712<br>\n  Nombre total d'images dans les TFRecords test: 7382<br>\n  _<br>\n  Resolution 512x512<br>\n  Nombre total d'images dans les TFRecords train: 12753<br>\n  Nombre total d'images dans les TFRecords val: 3712<br>\n  Nombre total d'images dans les TFRecords test: 7382</p>\n</blockquote>\n<p>Excuse my french 🤪</p>",
      "rawMarkdown": "The 4 folders contains the same pics with different resolutions !\n\nwith this code\n\n>def count_examples_in_tfrecord(tfrecord_path):\n         \"\"\"Compte le nombre d'images dans un fichier TFRecord.\"\"\"\n         return sum(1 for _ in tf.data.TFRecordDataset(tfrecord_path))\n\nthen\n\n>for resolution in ['192x192','224x224','331x331','512x512']:\n    print(f\"\\nResolution {resolution}\")\n    for doss in [\"train\", \"val\", \"test\"]:\n        # Lister tous les fichiers TFRecord\n        tfrecord_files = tf.io.gfile.glob(GCS_DS_PATH + '/tfrecords-jpeg-'+resolution+'/' + doss + '/*.tfrec')\n        # Compter les images dans chaque fichier et faire la somme\n        total_images = sum(count_examples_in_tfrecord(file) for file in tfrecord_files)\n        print(f\"Nombre total d'images dans les TFRecords {doss}: {total_images}\")\n\nyou get :\n\n>Resolution 192x192\nNombre total d'images dans les TFRecords train: 12753\nNombre total d'images dans les TFRecords val: 3712\nNombre total d'images dans les TFRecords test: 7382\n_\n>Resolution 224x224\nNombre total d'images dans les TFRecords train: 12753\nNombre total d'images dans les TFRecords val: 3712\nNombre total d'images dans les TFRecords test: 7382\n_\n>Resolution 331x331\nNombre total d'images dans les TFRecords train: 12753\nNombre total d'images dans les TFRecords val: 3712\nNombre total d'images dans les TFRecords test: 7382\n_\n>Resolution 512x512\nNombre total d'images dans les TFRecords train: 12753\nNombre total d'images dans les TFRecords val: 3712\nNombre total d'images dans les TFRecords test: 7382\n\nExcuse my french 🤪"
    }
  ],
  "comments": [
    {
      "id": 3033634,
      "author_name": "AleNic",
      "author_url": "",
      "post_date": "2024-11-01T11:27:26.037000",
      "content": "<p>hi, for the point 1) yes I confirm, you can use different image resolutions. You can also use this dataset and after downsample images to lower resolutions:</p>\n<p><a href=\"https://www.kaggle.com/datasets/alenic/flower-classification-512-png\" target=\"_blank\">https://www.kaggle.com/datasets/alenic/flower-classification-512-png</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2981067,
      "author_name": "darky",
      "author_url": "",
      "post_date": "2024-09-06T13:16:57.387000",
      "content": "<p>Hey dude, there are four different datasets for the given competition and it means that you can work on any of the given datasets and provide the results. They are just different in resolutions. </p>\n<p>Also, the sample_submission.csv isn't a summary dataset, its the dataset that you are required to submit your results on to the competition.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3169770,
      "author_name": "Didier Flamm",
      "author_url": "",
      "post_date": "2025-04-03T20:30:38.343000",
      "content": "<p>The 4 folders contains the same pics with different resolutions !</p>\n<p>with this code</p>\n<blockquote>\n  <p>def count_examples_in_tfrecord(tfrecord_path):<br>\n           \"\"\"Compte le nombre d'images dans un fichier TFRecord.\"\"\"<br>\n           return sum(1 for _ in tf.data.TFRecordDataset(tfrecord_path))</p>\n</blockquote>\n<p>then</p>\n<blockquote>\n  <p>for resolution in ['192x192','224x224','331x331','512x512']:<br>\n      print(f\"\\nResolution {resolution}\")<br>\n      for doss in [\"train\", \"val\", \"test\"]:<br>\n          # Lister tous les fichiers TFRecord<br>\n          tfrecord_files = tf.io.gfile.glob(GCS_DS_PATH + '/tfrecords-jpeg-'+resolution+'/' + doss + '/*.tfrec')<br>\n          # Compter les images dans chaque fichier et faire la somme<br>\n          total_images = sum(count_examples_in_tfrecord(file) for file in tfrecord_files)<br>\n          print(f\"Nombre total d'images dans les TFRecords {doss}: {total_images}\")</p>\n</blockquote>\n<p>you get :</p>\n<blockquote>\n  <p>Resolution 192x192<br>\n  Nombre total d'images dans les TFRecords train: 12753<br>\n  Nombre total d'images dans les TFRecords val: 3712<br>\n  Nombre total d'images dans les TFRecords test: 7382<br>\n  _<br>\n  Resolution 224x224<br>\n  Nombre total d'images dans les TFRecords train: 12753<br>\n  Nombre total d'images dans les TFRecords val: 3712<br>\n  Nombre total d'images dans les TFRecords test: 7382<br>\n  _<br>\n  Resolution 331x331<br>\n  Nombre total d'images dans les TFRecords train: 12753<br>\n  Nombre total d'images dans les TFRecords val: 3712<br>\n  Nombre total d'images dans les TFRecords test: 7382<br>\n  _<br>\n  Resolution 512x512<br>\n  Nombre total d'images dans les TFRecords train: 12753<br>\n  Nombre total d'images dans les TFRecords val: 3712<br>\n  Nombre total d'images dans les TFRecords test: 7382</p>\n</blockquote>\n<p>Excuse my french 🤪</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3033634": "hi, for the point 1) yes I confirm, you can use different image resolutions. You can also use this dataset and after downsample images to lower resolutions:\n\nhttps://www.kaggle.com/datasets/alenic/flower-classification-512-png\n\n",
    "2981067": "Hey dude, there are four different datasets for the given competition and it means that you can work on any of the given datasets and provide the results. They are just different in resolutions. \n\nAlso, the sample_submission.csv isn't a summary dataset, its the dataset that you are required to submit your results on to the competition.",
    "2872997": "![Data image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17225803%2Fe86bb0bc53339600490543458229a4a2%2FScreenshot%202024-06-15%20131721.jpg?generation=1718439475774399&alt=media)\nI was looking at the dataset given to us for training the classification model and I saw **4 folders**. Each folders has images of different dimensions which have been split into **train, val and test**.\n**Questions:**\n1) Do all the 4 images folders have the **same images** with different dimension to train 4 different models(higher resolution images will give better result but more computationally expensive and time consuming)?\n2) There is also a **summary dataset** below. What is that about. Is it the same thing?\n\nIt would be very helpful if you answer the questions\nRegards \n**Umer Tariq**",
    "3169770": "The 4 folders contains the same pics with different resolutions !\n\nwith this code\n\n>def count_examples_in_tfrecord(tfrecord_path):\n         \"\"\"Compte le nombre d'images dans un fichier TFRecord.\"\"\"\n         return sum(1 for _ in tf.data.TFRecordDataset(tfrecord_path))\n\nthen\n\n>for resolution in ['192x192','224x224','331x331','512x512']:\n    print(f\"\\nResolution {resolution}\")\n    for doss in [\"train\", \"val\", \"test\"]:\n        # Lister tous les fichiers TFRecord\n        tfrecord_files = tf.io.gfile.glob(GCS_DS_PATH + '/tfrecords-jpeg-'+resolution+'/' + doss + '/*.tfrec')\n        # Compter les images dans chaque fichier et faire la somme\n        total_images = sum(count_examples_in_tfrecord(file) for file in tfrecord_files)\n        print(f\"Nombre total d'images dans les TFRecords {doss}: {total_images}\")\n\nyou get :\n\n>Resolution 192x192\nNombre total d'images dans les TFRecords train: 12753\nNombre total d'images dans les TFRecords val: 3712\nNombre total d'images dans les TFRecords test: 7382\n_\n>Resolution 224x224\nNombre total d'images dans les TFRecords train: 12753\nNombre total d'images dans les TFRecords val: 3712\nNombre total d'images dans les TFRecords test: 7382\n_\n>Resolution 331x331\nNombre total d'images dans les TFRecords train: 12753\nNombre total d'images dans les TFRecords val: 3712\nNombre total d'images dans les TFRecords test: 7382\n_\n>Resolution 512x512\nNombre total d'images dans les TFRecords train: 12753\nNombre total d'images dans les TFRecords val: 3712\nNombre total d'images dans les TFRecords test: 7382\n\nExcuse my french 🤪"
  }
}