{
  "id": 181555,
  "title": "How to use only a subset of training image?",
  "url": "/competitions/landmark-recognition-2020/discussion/181555",
  "author_name": "",
  "post_date": "2020-09-09T09:22:23.460036100Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I want to use only a subset of the training images dataset given. The absolute paths of image files that i want to use, I have saved in a file('files_to_use_single_first.txt') in my private dataset 'train-emb'. <br>\nContents of this file look like this:-<br>\n…….<br>\n/kaggle/input/landmark-recognition-2020/train/2/0/8/2080d735a316b867.jpg<br>\n/kaggle/input/landmark-recognition-2020/train/0/c/b/0cb4689bd73b47ea.jpg<br>\n/kaggle/input/landmark-recognition-2020/train/6/c/f/6cf9b3589e456170.jpg<br>\n…….</p>\n<p>So, only for training directory, i want to load that list of filepaths as below(to extract global features as per baseline model):-</p>\n<p>if image_root_dir == TRAIN_IMAGE_DIR:<br>\n-&gt;-&gt;-&gt;-&gt;list_paths = load_text_file_into_list('files_to_use_single_first.txt')<br>\n-&gt;-&gt;-&gt;-&gt;image_paths = [pathlib.PosixPath(x) for x in list_paths]<br>\nelse:<br>\n-&gt;-&gt;-&gt;-&gt;image_paths = [x for x in pathlib.Path(image_root_dir).rglob('*.jpg')]</p>\n<p>But I started getting error : 'Submission CSV Not Found'</p>\n<p>So, I want to know is this not allowed or I am doing it the wrong way?</p>\n<p>Any help is really appreciated</p>\n<p>Note : On running the notebook in a batch session it is generating submission.csv file properly after running for about 5 hours.</p>",
  "messages": [
    {
      "id": "1003819",
      "postDate": "09/09/2020 09:22:23",
      "content": "<p>I want to use only a subset of the training images dataset given. The absolute paths of image files that i want to use, I have saved in a file('files_to_use_single_first.txt') in my private dataset 'train-emb'. <br>\nContents of this file look like this:-<br>\n…….<br>\n/kaggle/input/landmark-recognition-2020/train/2/0/8/2080d735a316b867.jpg<br>\n/kaggle/input/landmark-recognition-2020/train/0/c/b/0cb4689bd73b47ea.jpg<br>\n/kaggle/input/landmark-recognition-2020/train/6/c/f/6cf9b3589e456170.jpg<br>\n…….</p>\n<p>So, only for training directory, i want to load that list of filepaths as below(to extract global features as per baseline model):-</p>\n<p>if image_root_dir == TRAIN_IMAGE_DIR:<br>\n-&gt;-&gt;-&gt;-&gt;list_paths = load_text_file_into_list('files_to_use_single_first.txt')<br>\n-&gt;-&gt;-&gt;-&gt;image_paths = [pathlib.PosixPath(x) for x in list_paths]<br>\nelse:<br>\n-&gt;-&gt;-&gt;-&gt;image_paths = [x for x in pathlib.Path(image_root_dir).rglob('*.jpg')]</p>\n<p>But I started getting error : 'Submission CSV Not Found'</p>\n<p>So, I want to know is this not allowed or I am doing it the wrong way?</p>\n<p>Any help is really appreciated</p>\n<p>Note : On running the notebook in a batch session it is generating submission.csv file properly after running for about 5 hours.</p>",
      "rawMarkdown": "I want to use only a subset of the training images dataset given. The absolute paths of image files that i want to use, I have saved in a file('files_to_use_single_first.txt') in my private dataset 'train-emb'. \nContents of this file look like this:-\n.......\n/kaggle/input/landmark-recognition-2020/train/2/0/8/2080d735a316b867.jpg\n/kaggle/input/landmark-recognition-2020/train/0/c/b/0cb4689bd73b47ea.jpg\n/kaggle/input/landmark-recognition-2020/train/6/c/f/6cf9b3589e456170.jpg\n.......\n\n\nSo, only for training directory, i want to load that list of filepaths as below(to extract global features as per baseline model):-\n\nif image_root_dir == TRAIN_IMAGE_DIR:\n->->->->list_paths = load_text_file_into_list('files_to_use_single_first.txt')\n->->->->image_paths = [pathlib.PosixPath(x) for x in list_paths]\nelse:\n->->->->image_paths = [x for x in pathlib.Path(image_root_dir).rglob('*.jpg')]\n\nBut I started getting error : 'Submission CSV Not Found'\n\nSo, I want to know is this not allowed or I am doing it the wrong way?\n\nAny help is really appreciated\n\nNote : On running the notebook in a batch session it is generating submission.csv file properly after running for about 5 hours.",
      "votes": null
    },
    {
      "id": "1003894",
      "postDate": "09/09/2020 10:48:22",
      "content": "<p>The problem is that during the submission you have the whole training dataset available. But when you submit your kernel for rerun on private data only a part of training dataset becomes available (100K images out of 1.5M), that's why your kernel fails to load specific images that are not present in the private training set and you get Submission Error as a result. Solution could be if you add the whole training dataset additionally to the one that's included. If you want to use only a part of the training dataset, feel free to use those 100K images, but don't load old image names saved in .txt</p>",
      "rawMarkdown": "The problem is that during the submission you have the whole training dataset available. But when you submit your kernel for rerun on private data only a part of training dataset becomes available (100K images out of 1.5M), that's why your kernel fails to load specific images that are not present in the private training set and you get Submission Error as a result. Solution could be if you add the whole training dataset additionally to the one that's included. If you want to use only a part of the training dataset, feel free to use those 100K images, but don't load old image names saved in .txt",
      "votes": null
    },
    {
      "id": "1004110",
      "postDate": "09/09/2020 13:17:51",
      "content": "<p><a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a> thanks a lot for your quick reply</p>\n<p>I will try to do that</p>",
      "rawMarkdown": "vostankovich thanks a lot for your quick reply\n\nI will try to do that",
      "votes": null
    },
    {
      "id": "1004118",
      "postDate": "09/09/2020 13:24:12",
      "content": "<p>you're welcome</p>",
      "rawMarkdown": "you're welcome",
      "votes": null
    },
    {
      "id": "1004125",
      "postDate": "09/09/2020 13:28:33",
      "content": "<p><a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a> One more query i have:-</p>\n<p>As you suggested, i can copy files to another dataset. For this I need to use that path for my training dataset.<br>\nSo, should i make a change to this parameter or create a new one?</p>\n<p>TRAIN_IMAGE_DIR = os.path.join(DATASET_DIR, 'train')</p>\n<p>I mean if this path is to be used by private training dataset re-run after submission, the i think I shouldn't change it?</p>",
      "rawMarkdown": "vostankovich One more query i have:-\n\nAs you suggested, i can copy files to another dataset. For this I need to use that path for my training dataset.\nSo, should i make a change to this parameter or create a new one?\n\nTRAIN_IMAGE_DIR = os.path.join(DATASET_DIR, 'train')\n\nI mean if this path is to be used by private training dataset re-run after submission, the i think I shouldn't change it?",
      "votes": null
    },
    {
      "id": "1004176",
      "postDate": "09/09/2020 14:11:18",
      "content": "<p>Not sure if I got you correctly.</p>\n<p>During re-run private train set: /kaggle/input/landmark-recognition-2020/train/ will contain 100K images (subset)<br>\nAnything you upload additionally will not change during rerun<br>\nOnly /kaggle/input/landmark-recognition-2020/ changes its content</p>",
      "rawMarkdown": "Not sure if I got you correctly.\n\nDuring re-run private train set: /kaggle/input/landmark-recognition-2020/train/ will contain 100K images (subset)\nAnything you upload additionally will not change during rerun\nOnly /kaggle/input/landmark-recognition-2020/ changes its content",
      "votes": null
    },
    {
      "id": "1004190",
      "postDate": "09/09/2020 14:20:43",
      "content": "<p><a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a> I am asking is it okay to change ONLY this parameter and submit the baseline code given as it is?</p>\n<p>TRAIN_IMAGE_DIR = os.path.join(, 'my_subset_train')</p>\n<p>Or will this also give an error in submission?</p>",
      "rawMarkdown": "vostankovich I am asking is it okay to change ONLY this parameter and submit the baseline code given as it is?\n\nTRAIN_IMAGE_DIR = os.path.join(<PATH_TO_MY_PRIVATE_DATASET_DIR>, 'my_subset_train')\n\nOr will this also give an error in submission?",
      "votes": null
    },
    {
      "id": "1004196",
      "postDate": "09/09/2020 14:25:55",
      "content": "<p>I don't know if it will give an error, try it. But with larger size of training dataset the baseline submission may crush due to lack of time to process all of the images (even 12 hours is not enough)</p>",
      "rawMarkdown": "I don't know if it will give an error, try it. But with larger size of training dataset the baseline submission may crush due to lack of time to process all of the images (even 12 hours is not enough)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1003894,
      "author_name": "vostankovich",
      "author_url": "",
      "post_date": "09/09/2020 10:48:22",
      "content": "<p>The problem is that during the submission you have the whole training dataset available. But when you submit your kernel for rerun on private data only a part of training dataset becomes available (100K images out of 1.5M), that's why your kernel fails to load specific images that are not present in the private training set and you get Submission Error as a result. Solution could be if you add the whole training dataset additionally to the one that's included. If you want to use only a part of the training dataset, feel free to use those 100K images, but don't load old image names saved in .txt</p>",
      "votes": null,
      "replies": [
        {
          "id": 1004110,
          "author_name": "rohitdeepu17",
          "author_url": "",
          "post_date": "09/09/2020 13:17:51",
          "content": "<p><a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a> thanks a lot for your quick reply</p>\n<p>I will try to do that</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1004118,
          "author_name": "vostankovich",
          "author_url": "",
          "post_date": "09/09/2020 13:24:12",
          "content": "<p>you're welcome</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1004125,
          "author_name": "rohitdeepu17",
          "author_url": "",
          "post_date": "09/09/2020 13:28:33",
          "content": "<p><a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a> One more query i have:-</p>\n<p>As you suggested, i can copy files to another dataset. For this I need to use that path for my training dataset.<br>\nSo, should i make a change to this parameter or create a new one?</p>\n<p>TRAIN_IMAGE_DIR = os.path.join(DATASET_DIR, 'train')</p>\n<p>I mean if this path is to be used by private training dataset re-run after submission, the i think I shouldn't change it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1004176,
          "author_name": "vostankovich",
          "author_url": "",
          "post_date": "09/09/2020 14:11:18",
          "content": "<p>Not sure if I got you correctly.</p>\n<p>During re-run private train set: /kaggle/input/landmark-recognition-2020/train/ will contain 100K images (subset)<br>\nAnything you upload additionally will not change during rerun<br>\nOnly /kaggle/input/landmark-recognition-2020/ changes its content</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1004190,
          "author_name": "rohitdeepu17",
          "author_url": "",
          "post_date": "09/09/2020 14:20:43",
          "content": "<p><a href=\"https://www.kaggle.com/vostankovich\" target=\"_blank\">@vostankovich</a> I am asking is it okay to change ONLY this parameter and submit the baseline code given as it is?</p>\n<p>TRAIN_IMAGE_DIR = os.path.join(, 'my_subset_train')</p>\n<p>Or will this also give an error in submission?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1004196,
          "author_name": "vostankovich",
          "author_url": "",
          "post_date": "09/09/2020 14:25:55",
          "content": "<p>I don't know if it will give an error, try it. But with larger size of training dataset the baseline submission may crush due to lack of time to process all of the images (even 12 hours is not enough)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1003819": "I want to use only a subset of the training images dataset given. The absolute paths of image files that i want to use, I have saved in a file('files_to_use_single_first.txt') in my private dataset 'train-emb'. \nContents of this file look like this:-\n.......\n/kaggle/input/landmark-recognition-2020/train/2/0/8/2080d735a316b867.jpg\n/kaggle/input/landmark-recognition-2020/train/0/c/b/0cb4689bd73b47ea.jpg\n/kaggle/input/landmark-recognition-2020/train/6/c/f/6cf9b3589e456170.jpg\n.......\n\n\nSo, only for training directory, i want to load that list of filepaths as below(to extract global features as per baseline model):-\n\nif image_root_dir == TRAIN_IMAGE_DIR:\n->->->->list_paths = load_text_file_into_list('files_to_use_single_first.txt')\n->->->->image_paths = [pathlib.PosixPath(x) for x in list_paths]\nelse:\n->->->->image_paths = [x for x in pathlib.Path(image_root_dir).rglob('*.jpg')]\n\nBut I started getting error : 'Submission CSV Not Found'\n\nSo, I want to know is this not allowed or I am doing it the wrong way?\n\nAny help is really appreciated\n\nNote : On running the notebook in a batch session it is generating submission.csv file properly after running for about 5 hours.",
    "1003894": "The problem is that during the submission you have the whole training dataset available. But when you submit your kernel for rerun on private data only a part of training dataset becomes available (100K images out of 1.5M), that's why your kernel fails to load specific images that are not present in the private training set and you get Submission Error as a result. Solution could be if you add the whole training dataset additionally to the one that's included. If you want to use only a part of the training dataset, feel free to use those 100K images, but don't load old image names saved in .txt",
    "1004110": "vostankovich thanks a lot for your quick reply\n\nI will try to do that",
    "1004118": "you're welcome",
    "1004125": "vostankovich One more query i have:-\n\nAs you suggested, i can copy files to another dataset. For this I need to use that path for my training dataset.\nSo, should i make a change to this parameter or create a new one?\n\nTRAIN_IMAGE_DIR = os.path.join(DATASET_DIR, 'train')\n\nI mean if this path is to be used by private training dataset re-run after submission, the i think I shouldn't change it?",
    "1004176": "Not sure if I got you correctly.\n\nDuring re-run private train set: /kaggle/input/landmark-recognition-2020/train/ will contain 100K images (subset)\nAnything you upload additionally will not change during rerun\nOnly /kaggle/input/landmark-recognition-2020/ changes its content",
    "1004190": "vostankovich I am asking is it okay to change ONLY this parameter and submit the baseline code given as it is?\n\nTRAIN_IMAGE_DIR = os.path.join(<PATH_TO_MY_PRIVATE_DATASET_DIR>, 'my_subset_train')\n\nOr will this also give an error in submission?",
    "1004196": "I don't know if it will give an error, try it. But with larger size of training dataset the baseline submission may crush due to lack of time to process all of the images (even 12 hours is not enough)"
  },
  "source": "meta"
}