{
  "id": 186487,
  "title": "20GB GLD dataset/tfrecords, to train on colab",
  "url": "/competitions/landmark-recognition-2020/discussion/186487",
  "author_name": "",
  "post_date": "2020-09-24T16:57:29.948259400Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>Hope this dataset helps people in training their custom models running on colab/kaggle.<br>\nAll images are resized using tensorflow (224,320) minimum size, preserving aspect ratio and file tree structure.</p>\n<p>run: to check no images have missed. train.csv from original dataset.</p>\n<pre><code>print(len( [x for x in pathlib.Path(\"gld20gb/\").rglob('*.jpg')]))\ndf = pd.read_csv(\"images_per_class.csv\")\nprint(len(df.index))\ndf = pd.read_csv(\"train.csv\")\nprint(len(df.index))\n</code></pre>\n<p>K-fold training can also be done, by grouping all classes having same number of images. The <a href=\"https://www.kaggle.com/jkreddy/kaggle-github-integration-to-work-on-large-dataset#Preprocess-images-into-groups,-for-k-fold-training-later\" target=\"_blank\">code in notebook</a> prepares dictionary with key being number of images(of all landmarks/classes having same number of samples) and values being the paths of images.</p>\n<p>Image dataset available at the <a href=\"https://www.kaggle.com/jkreddy/gld20gb\" target=\"_blank\">link</a><br>\nTfrecords dataset available at the <a href=\"https://www.kaggle.com/jkreddy/tfrecords20gb\" target=\"_blank\">link</a><br>\n(train/validation set with 80-20 percent split)</p>\n<p>Use the below code in <strong>Colab</strong> to download and start using dataset</p>\n<pre><code>! pip install -q kaggle\nfrom google.colab import files\nfiles.upload()\n!mkdir ~/.kaggle\n!cp kaggle.json ~/.kaggle/\n!chmod 600 ~/.kaggle/kaggle.json\n!pip install --upgrade --force-reinstall --no-deps kaggle\n\n!kaggle datasets download -d jkreddy/gld20gb\nimport zipfile\nzipref = zipfile.ZipFile('gld20gb.zip', 'r') \nzipref.extractall(\"train\")\nzipref.close()\n</code></pre>",
  "messages": [
    {
      "id": "1025582",
      "postDate": "09/24/2020 16:57:29",
      "content": "<p>Hello all,</p>\n<p>Hope this dataset helps people in training their custom models running on colab/kaggle.<br>\nAll images are resized using tensorflow (224,320) minimum size, preserving aspect ratio and file tree structure.</p>\n<p>run: to check no images have missed. train.csv from original dataset.</p>\n<pre><code>print(len( [x for x in pathlib.Path(\"gld20gb/\").rglob('*.jpg')]))\ndf = pd.read_csv(\"images_per_class.csv\")\nprint(len(df.index))\ndf = pd.read_csv(\"train.csv\")\nprint(len(df.index))\n</code></pre>\n<p>K-fold training can also be done, by grouping all classes having same number of images. The <a href=\"https://www.kaggle.com/jkreddy/kaggle-github-integration-to-work-on-large-dataset#Preprocess-images-into-groups,-for-k-fold-training-later\" target=\"_blank\">code in notebook</a> prepares dictionary with key being number of images(of all landmarks/classes having same number of samples) and values being the paths of images.</p>\n<p>Image dataset available at the <a href=\"https://www.kaggle.com/jkreddy/gld20gb\" target=\"_blank\">link</a><br>\nTfrecords dataset available at the <a href=\"https://www.kaggle.com/jkreddy/tfrecords20gb\" target=\"_blank\">link</a><br>\n(train/validation set with 80-20 percent split)</p>\n<p>Use the below code in <strong>Colab</strong> to download and start using dataset</p>\n<pre><code>! pip install -q kaggle\nfrom google.colab import files\nfiles.upload()\n!mkdir ~/.kaggle\n!cp kaggle.json ~/.kaggle/\n!chmod 600 ~/.kaggle/kaggle.json\n!pip install --upgrade --force-reinstall --no-deps kaggle\n\n!kaggle datasets download -d jkreddy/gld20gb\nimport zipfile\nzipref = zipfile.ZipFile('gld20gb.zip', 'r') \nzipref.extractall(\"train\")\nzipref.close()\n</code></pre>",
      "rawMarkdown": "Hello all,\n\nHope this dataset helps people in training their custom models running on colab/kaggle.\nAll images are resized using tensorflow (224,320) minimum size, preserving aspect ratio and file tree structure.\n\n\nrun: to check no images have missed. train.csv from original dataset.\n```\nprint(len( [x for x in pathlib.Path(\"gld20gb/\").rglob('*.jpg')]))\ndf = pd.read_csv(\"images_per_class.csv\")\nprint(len(df.index))\ndf = pd.read_csv(\"train.csv\")\nprint(len(df.index))\n```\n\nK-fold training can also be done, by grouping all classes having same number of images. The [code in notebook](https://www.kaggle.com/jkreddy/kaggle-github-integration-to-work-on-large-dataset#Preprocess-images-into-groups,-for-k-fold-training-later) prepares dictionary with key being number of images(of all landmarks/classes having same number of samples) and values being the paths of images.\n\nImage dataset available at the [link](https://www.kaggle.com/jkreddy/gld20gb)\nTfrecords dataset available at the [link](https://www.kaggle.com/jkreddy/tfrecords20gb)\n(train/validation set with 80-20 percent split)\n\nUse the below code in **Colab** to download and start using dataset\n\n```\n! pip install -q kaggle\nfrom google.colab import files\nfiles.upload()\n!mkdir ~/.kaggle\n!cp kaggle.json ~/.kaggle/\n!chmod 600 ~/.kaggle/kaggle.json\n!pip install --upgrade --force-reinstall --no-deps kaggle\n\n!kaggle datasets download -d jkreddy/gld20gb\nimport zipfile\nzipref = zipfile.ZipFile('gld20gb.zip', 'r') \nzipref.extractall(\"train\")\nzipref.close()\n```",
      "votes": null
    },
    {
      "id": "1025919",
      "postDate": "09/24/2020 21:53:57",
      "content": "<p>Thank you for your great kernel and dataset, Good job 👍</p>",
      "rawMarkdown": "Thank you for your great kernel and dataset, Good job 👍",
      "votes": null
    },
    {
      "id": "1030030",
      "postDate": "09/28/2020 11:14:38",
      "content": "<p>Thanks for your kernel and dataset <a href=\"https://www.kaggle.com/jkreddy\" target=\"_blank\">@jkreddy</a> Awesome job</p>",
      "rawMarkdown": "Thanks for your kernel and dataset @jkreddy Awesome job",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1025919,
      "author_name": "",
      "author_url": "",
      "post_date": "09/24/2020 21:53:57",
      "content": "<p>Thank you for your great kernel and dataset, Good job 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1030030,
      "author_name": "vineeth1999",
      "author_url": "",
      "post_date": "09/28/2020 11:14:38",
      "content": "<p>Thanks for your kernel and dataset <a href=\"https://www.kaggle.com/jkreddy\" target=\"_blank\">@jkreddy</a> Awesome job</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1025582": "Hello all,\n\nHope this dataset helps people in training their custom models running on colab/kaggle.\nAll images are resized using tensorflow (224,320) minimum size, preserving aspect ratio and file tree structure.\n\n\nrun: to check no images have missed. train.csv from original dataset.\n```\nprint(len( [x for x in pathlib.Path(\"gld20gb/\").rglob('*.jpg')]))\ndf = pd.read_csv(\"images_per_class.csv\")\nprint(len(df.index))\ndf = pd.read_csv(\"train.csv\")\nprint(len(df.index))\n```\n\nK-fold training can also be done, by grouping all classes having same number of images. The [code in notebook](https://www.kaggle.com/jkreddy/kaggle-github-integration-to-work-on-large-dataset#Preprocess-images-into-groups,-for-k-fold-training-later) prepares dictionary with key being number of images(of all landmarks/classes having same number of samples) and values being the paths of images.\n\nImage dataset available at the [link](https://www.kaggle.com/jkreddy/gld20gb)\nTfrecords dataset available at the [link](https://www.kaggle.com/jkreddy/tfrecords20gb)\n(train/validation set with 80-20 percent split)\n\nUse the below code in **Colab** to download and start using dataset\n\n```\n! pip install -q kaggle\nfrom google.colab import files\nfiles.upload()\n!mkdir ~/.kaggle\n!cp kaggle.json ~/.kaggle/\n!chmod 600 ~/.kaggle/kaggle.json\n!pip install --upgrade --force-reinstall --no-deps kaggle\n\n!kaggle datasets download -d jkreddy/gld20gb\nimport zipfile\nzipref = zipfile.ZipFile('gld20gb.zip', 'r') \nzipref.extractall(\"train\")\nzipref.close()\n```",
    "1025919": "Thank you for your great kernel and dataset, Good job 👍",
    "1030030": "Thanks for your kernel and dataset @jkreddy Awesome job"
  },
  "source": "meta"
}