{
  "id": 273214,
  "title": "400K Extra Train Images and TFRecords of Minority Classes",
  "url": "/competitions/landmark-recognition-2021/discussion/273214",
  "author_name": "",
  "post_date": "2021-09-19T19:42:06.926415200Z",
  "votes": 17,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>The huge training dataset provided for this competition is a subset of the even larger <a href=\"https://github.com/cvdfoundation/google-landmark\" target=\"_blank\">Google Landmarks Dataset v2</a> containing over 4 million images belonging to ~200K landmarks. This competition is a subset of 1.5M images with 81313 different landmarks, however some images are left out. I wrote a notebook which crawls all 4M images to create a dataset with extra training data of minority classes, landmarks with very few samples, in this case less than 20. This way the training dataset becomes less imbalanced.</p>\n<p>The crawling notebook can be found <a href=\"https://www.kaggle.com/markwijkhuizen/goole-landmark-recognition-extra-train-data-pub\" target=\"_blank\">here</a>, with the resulting training images <a href=\"https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-train-data-pub\" target=\"_blank\">dataset</a>.</p>\n<p>An extra notebook where the images are converted to TFReocrds can be found <a href=\"https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-data-tfrec-pub\" target=\"_blank\">here</a>. The TFRecords dataset  can be found <a href=\"https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-train-tfrecs-pub\" target=\"_blank\">here</a>.</p>\n<p>Each class is filled up to 20 samples, meaning all extra images are for landmarks with less than 20 samples. Having more samples of minority classes should create more accurate models with less bias, the majority class is relatively smaller now.</p>\n<p>Hope this dataset helps with getting the final performance boost in these last 2 weeks of this competition.</p>\n<p>Happy Kaggling!</p>",
  "messages": [
    {
      "id": "1517600",
      "postDate": "09/19/2021 19:42:06",
      "content": "<p>Hi all,</p>\n<p>The huge training dataset provided for this competition is a subset of the even larger <a href=\"https://github.com/cvdfoundation/google-landmark\" target=\"_blank\">Google Landmarks Dataset v2</a> containing over 4 million images belonging to ~200K landmarks. This competition is a subset of 1.5M images with 81313 different landmarks, however some images are left out. I wrote a notebook which crawls all 4M images to create a dataset with extra training data of minority classes, landmarks with very few samples, in this case less than 20. This way the training dataset becomes less imbalanced.</p>\n<p>The crawling notebook can be found <a href=\"https://www.kaggle.com/markwijkhuizen/goole-landmark-recognition-extra-train-data-pub\" target=\"_blank\">here</a>, with the resulting training images <a href=\"https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-train-data-pub\" target=\"_blank\">dataset</a>.</p>\n<p>An extra notebook where the images are converted to TFReocrds can be found <a href=\"https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-data-tfrec-pub\" target=\"_blank\">here</a>. The TFRecords dataset  can be found <a href=\"https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-train-tfrecs-pub\" target=\"_blank\">here</a>.</p>\n<p>Each class is filled up to 20 samples, meaning all extra images are for landmarks with less than 20 samples. Having more samples of minority classes should create more accurate models with less bias, the majority class is relatively smaller now.</p>\n<p>Hope this dataset helps with getting the final performance boost in these last 2 weeks of this competition.</p>\n<p>Happy Kaggling!</p>",
      "rawMarkdown": "Hi all,\n\nThe huge training dataset provided for this competition is a subset of the even larger [Google Landmarks Dataset v2](https://github.com/cvdfoundation/google-landmark) containing over 4 million images belonging to ~200K landmarks. This competition is a subset of 1.5M images with 81313 different landmarks, however some images are left out. I wrote a notebook which crawls all 4M images to create a dataset with extra training data of minority classes, landmarks with very few samples, in this case less than 20. This way the training dataset becomes less imbalanced.\n\nThe crawling notebook can be found [here](https://www.kaggle.com/markwijkhuizen/goole-landmark-recognition-extra-train-data-pub), with the resulting training images [dataset](https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-train-data-pub).\n\nAn extra notebook where the images are converted to TFReocrds can be found [here](https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-data-tfrec-pub). The TFRecords dataset  can be found [here](https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-train-tfrecs-pub).\n\nEach class is filled up to 20 samples, meaning all extra images are for landmarks with less than 20 samples. Having more samples of minority classes should create more accurate models with less bias, the majority class is relatively smaller now.\n\nHope this dataset helps with getting the final performance boost in these last 2 weeks of this competition.\n\nHappy Kaggling!",
      "votes": null
    },
    {
      "id": "1519860",
      "postDate": "09/22/2021 01:52:31",
      "content": "<p>Thanks for sharing this <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> ..definitly will be helpful for many teams in the last 2 weeks (although my TPU quota for the week is already over 😭…will have to wait till it resets)</p>",
      "rawMarkdown": "Thanks for sharing this @markwijkhuizen ..definitly will be helpful for many teams in the last 2 weeks (although my TPU quota for the week is already over 😭...will have to wait till it resets)",
      "votes": null
    },
    {
      "id": "1520050",
      "postDate": "09/22/2021 05:45:27",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": null
    },
    {
      "id": "1520648",
      "postDate": "09/22/2021 13:03:58",
      "content": "<p>Computing Power is definitely scarce in this competition, we will need to spend the last 30 hours smart ;)</p>",
      "rawMarkdown": "Computing Power is definitely scarce in this competition, we will need to spend the last 30 hours smart ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1519860,
      "author_name": "sandy1112",
      "author_url": "",
      "post_date": "09/22/2021 01:52:31",
      "content": "<p>Thanks for sharing this <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> ..definitly will be helpful for many teams in the last 2 weeks (although my TPU quota for the week is already over 😭…will have to wait till it resets)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1520648,
          "author_name": "markwijkhuizen",
          "author_url": "",
          "post_date": "09/22/2021 13:03:58",
          "content": "<p>Computing Power is definitely scarce in this competition, we will need to spend the last 30 hours smart ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1520050,
      "author_name": "faaizhashmi",
      "author_url": "",
      "post_date": "09/22/2021 05:45:27",
      "content": "<p>Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1517600": "Hi all,\n\nThe huge training dataset provided for this competition is a subset of the even larger [Google Landmarks Dataset v2](https://github.com/cvdfoundation/google-landmark) containing over 4 million images belonging to ~200K landmarks. This competition is a subset of 1.5M images with 81313 different landmarks, however some images are left out. I wrote a notebook which crawls all 4M images to create a dataset with extra training data of minority classes, landmarks with very few samples, in this case less than 20. This way the training dataset becomes less imbalanced.\n\nThe crawling notebook can be found [here](https://www.kaggle.com/markwijkhuizen/goole-landmark-recognition-extra-train-data-pub), with the resulting training images [dataset](https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-train-data-pub).\n\nAn extra notebook where the images are converted to TFReocrds can be found [here](https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-data-tfrec-pub). The TFRecords dataset  can be found [here](https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-train-tfrecs-pub).\n\nEach class is filled up to 20 samples, meaning all extra images are for landmarks with less than 20 samples. Having more samples of minority classes should create more accurate models with less bias, the majority class is relatively smaller now.\n\nHope this dataset helps with getting the final performance boost in these last 2 weeks of this competition.\n\nHappy Kaggling!",
    "1519860": "Thanks for sharing this @markwijkhuizen ..definitly will be helpful for many teams in the last 2 weeks (although my TPU quota for the week is already over 😭...will have to wait till it resets)",
    "1520050": "Thank you for sharing!",
    "1520648": "Computing Power is definitely scarce in this competition, we will need to spend the last 30 hours smart ;)"
  },
  "source": "meta"
}