{
  "id": 273912,
  "title": "How to run 500G complete training set on kaggle platform?",
  "url": "/competitions/landmark-recognition-2021/discussion/273912",
  "author_name": "",
  "post_date": "2021-09-23T08:34:09.854626200Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I want to use the complete training set, but it is too large to upload to the kaggle dataset</p>",
  "messages": [
    {
      "id": "1521440",
      "postDate": "09/23/2021 08:34:09",
      "content": "<p>I want to use the complete training set, but it is too large to upload to the kaggle dataset</p>",
      "rawMarkdown": "I want to use the complete training set, but it is too large to upload to the kaggle dataset",
      "votes": null
    },
    {
      "id": "1524825",
      "postDate": "09/26/2021 22:13:31",
      "content": "<p>You can take my datasets:<br>\n<a href=\"https://www.kaggle.com/markbquant/landmark-recognition-2021-16-tfrecs\" target=\"_blank\">https://www.kaggle.com/markbquant/landmark-recognition-2021-16-tfrecs</a><br>\n<a href=\"https://www.kaggle.com/markbquant/landmark-recognition-2021-test-dataset\" target=\"_blank\">https://www.kaggle.com/markbquant/landmark-recognition-2021-test-dataset</a></p>\n<p>You can simplify the dataset by taking only landmarks with more than 20/30 classes:</p>\n<pre><code>landmark_ids = labels['landmark_id'].value_counts()\nlandmark_ids = landmark_ids[landmark_ids&gt;=20].index.values\ndev_mask = labels['landmark_id'].isin(landmark_ids).values\ndev_labels = labels[dev_mask]\ndev_labels = dev_labels.groupby('landmark_id').head(20)\nlabels['dev_label'] = dev_labels['landmark_id'].astype('category').cat.codes\n</code></pre>",
      "rawMarkdown": "You can take my datasets:\nhttps://www.kaggle.com/markbquant/landmark-recognition-2021-16-tfrecs\nhttps://www.kaggle.com/markbquant/landmark-recognition-2021-test-dataset\n\nYou can simplify the dataset by taking only landmarks with more than 20/30 classes:\n```\nlandmark_ids = labels['landmark_id'].value_counts()\nlandmark_ids = landmark_ids[landmark_ids>=20].index.values\ndev_mask = labels['landmark_id'].isin(landmark_ids).values\ndev_labels = labels[dev_mask]\ndev_labels = dev_labels.groupby('landmark_id').head(20)\nlabels['dev_label'] = dev_labels['landmark_id'].astype('category').cat.codes\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1524825,
      "author_name": "markbquant",
      "author_url": "",
      "post_date": "09/26/2021 22:13:31",
      "content": "<p>You can take my datasets:<br>\n<a href=\"https://www.kaggle.com/markbquant/landmark-recognition-2021-16-tfrecs\" target=\"_blank\">https://www.kaggle.com/markbquant/landmark-recognition-2021-16-tfrecs</a><br>\n<a href=\"https://www.kaggle.com/markbquant/landmark-recognition-2021-test-dataset\" target=\"_blank\">https://www.kaggle.com/markbquant/landmark-recognition-2021-test-dataset</a></p>\n<p>You can simplify the dataset by taking only landmarks with more than 20/30 classes:</p>\n<pre><code>landmark_ids = labels['landmark_id'].value_counts()\nlandmark_ids = landmark_ids[landmark_ids&gt;=20].index.values\ndev_mask = labels['landmark_id'].isin(landmark_ids).values\ndev_labels = labels[dev_mask]\ndev_labels = dev_labels.groupby('landmark_id').head(20)\nlabels['dev_label'] = dev_labels['landmark_id'].astype('category').cat.codes\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1521440": "I want to use the complete training set, but it is too large to upload to the kaggle dataset",
    "1524825": "You can take my datasets:\nhttps://www.kaggle.com/markbquant/landmark-recognition-2021-16-tfrecs\nhttps://www.kaggle.com/markbquant/landmark-recognition-2021-test-dataset\n\nYou can simplify the dataset by taking only landmarks with more than 20/30 classes:\n```\nlandmark_ids = labels['landmark_id'].value_counts()\nlandmark_ids = landmark_ids[landmark_ids>=20].index.values\ndev_mask = labels['landmark_id'].isin(landmark_ids).values\ndev_labels = labels[dev_mask]\ndev_labels = dev_labels.groupby('landmark_id').head(20)\nlabels['dev_label'] = dev_labels['landmark_id'].astype('category').cat.codes\n```"
  },
  "source": "meta"
}