{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Hello Fellow Kagglers,\n\nThis notebook demonstrates how to get extra train images of low occuring classes. The train data is highly unbalanced, with some classes having thousands of samples and others just a handful of sampples. All classes are filled up to a maximum of 20 samples, greatly increasing the training data for low occuring classes. This should result in a lower class inbalance, lower bias towards the majority class and better recognition for low occuring classes.\n\nAll data is crawled from [this](https://github.com/cvdfoundation/google-landmark) GitHub repository. Over 400,000 new images are added. The provided training set in this competition contains 1.5M images, whereas the complete dataset contains over 4M images!\n\nAll 4M images are downloaded, if the Kaggle training set does not contain the image and the image belongs to a class with less than 20 samples the images is kept, it's as simple as that. The complete dataset also contains over 200,000 classes, most of those are not present in the Kaggle dataset. Only classes present in the Kaggle dataset are added.\n\nThe dataset this notebook results in can be found [here](https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-train-data-pub) and [this](https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-data-tfrec-pub) notebook shows how to convert the images to TFRecords, resulting in [this](https://www.kaggle.com/markwijkhuizen/google-landmark-recognition-extra-train-tfrecs-pub) TFRecords dataset.","metadata":{}},{"cell_type":"code","source":"# Silence All Tensorflow Warnings\n!pip install -q silence_tensorflow","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:46:54.347153Z","iopub.execute_input":"2021-09-19T10:46:54.34758Z","iopub.status.idle":"2021-09-19T10:47:03.624138Z","shell.execute_reply.started":"2021-09-19T10:46:54.347491Z","shell.execute_reply":"2021-09-19T10:47:03.623152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Silence Tensorflow\nimport silence_tensorflow.auto","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:03.625672Z","iopub.execute_input":"2021-09-19T10:47:03.625942Z","iopub.status.idle":"2021-09-19T10:47:08.832511Z","shell.execute_reply.started":"2021-09-19T10:47:03.625913Z","shell.execute_reply":"2021-09-19T10:47:08.831653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport tensorflow as tf\n\nimport matplotlib.pyplot as plt\n\nfrom tqdm.notebook import tqdm\nfrom multiprocessing import cpu_count\n\nimport joblib\nimport imageio\nimport cv2\nimport os\nimport glob\nimport multiprocessing\n\ntqdm.pandas()","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:08.833943Z","iopub.execute_input":"2021-09-19T10:47:08.83425Z","iopub.status.idle":"2021-09-19T10:47:09.170781Z","shell.execute_reply.started":"2021-09-19T10:47:08.83422Z","shell.execute_reply":"2021-09-19T10:47:09.169913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set CV2 to run single threaded, speeds up multithreading\ncv2.setNumThreads(1)","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:09.172192Z","iopub.execute_input":"2021-09-19T10:47:09.17246Z","iopub.status.idle":"2021-09-19T10:47:09.182628Z","shell.execute_reply.started":"2021-09-19T10:47:09.172434Z","shell.execute_reply":"2021-09-19T10:47:09.18161Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train Original","metadata":{}},{"cell_type":"code","source":"# Original Train\ntrain_original = pd.read_csv('/kaggle/input/landmark-recognition-2021/train.csv')","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:09.183768Z","iopub.execute_input":"2021-09-19T10:47:09.184028Z","iopub.status.idle":"2021-09-19T10:47:10.557443Z","shell.execute_reply.started":"2021-09-19T10:47:09.184002Z","shell.execute_reply":"2021-09-19T10:47:10.556557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_original.head())","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:10.55852Z","iopub.execute_input":"2021-09-19T10:47:10.558775Z","iopub.status.idle":"2021-09-19T10:47:10.581184Z","shell.execute_reply.started":"2021-09-19T10:47:10.55875Z","shell.execute_reply":"2021-09-19T10:47:10.580454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_original.info())","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:10.582082Z","iopub.execute_input":"2021-09-19T10:47:10.582364Z","iopub.status.idle":"2021-09-19T10:47:10.66743Z","shell.execute_reply.started":"2021-09-19T10:47:10.582337Z","shell.execute_reply":"2021-09-19T10:47:10.666787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Print class value counts, many classes have just 2 samples!\ntrain_original['landmark_id'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:10.669043Z","iopub.execute_input":"2021-09-19T10:47:10.669457Z","iopub.status.idle":"2021-09-19T10:47:10.7114Z","shell.execute_reply.started":"2021-09-19T10:47:10.669429Z","shell.execute_reply":"2021-09-19T10:47:10.710367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_original_landmark_ids = set(train_original['landmark_id'].unique())\nprint(f'There are {len(train_original_landmark_ids)} unique landmarks')","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:10.712736Z","iopub.execute_input":"2021-09-19T10:47:10.71298Z","iopub.status.idle":"2021-09-19T10:47:10.747214Z","shell.execute_reply.started":"2021-09-19T10:47:10.712957Z","shell.execute_reply":"2021-09-19T10:47:10.74639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Landmark ID occurances, we fill them up to 20 images\noriginal_landmark_id2count = train_original.groupby('landmark_id').count().squeeze().to_dict()","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:10.748189Z","iopub.execute_input":"2021-09-19T10:47:10.748427Z","iopub.status.idle":"2021-09-19T10:47:10.918122Z","shell.execute_reply.started":"2021-09-19T10:47:10.748403Z","shell.execute_reply":"2021-09-19T10:47:10.917341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Original image ids to check for duplicates\noriginal_ids = set(train_original['id'])","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:10.919245Z","iopub.execute_input":"2021-09-19T10:47:10.919505Z","iopub.status.idle":"2021-09-19T10:47:11.373412Z","shell.execute_reply.started":"2021-09-19T10:47:10.919481Z","shell.execute_reply":"2021-09-19T10:47:11.372239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train GitHub","metadata":{}},{"cell_type":"code","source":"# GitHub Train of complete dataset\n!wget -cq \"https://s3.amazonaws.com/google-landmark/metadata/train.csv\"\ntrain_github = pd.read_csv('./train.csv')","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:11.374612Z","iopub.execute_input":"2021-09-19T10:47:11.374884Z","iopub.status.idle":"2021-09-19T10:47:38.014617Z","shell.execute_reply.started":"2021-09-19T10:47:11.374848Z","shell.execute_reply":"2021-09-19T10:47:38.013725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_github.head())","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:38.015878Z","iopub.execute_input":"2021-09-19T10:47:38.016316Z","iopub.status.idle":"2021-09-19T10:47:38.025885Z","shell.execute_reply.started":"2021-09-19T10:47:38.016271Z","shell.execute_reply":"2021-09-19T10:47:38.025287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The complete dataset contains 203094 classes, many more than the Kaggle dataset\ntrain_github['landmark_id'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:38.026838Z","iopub.execute_input":"2021-09-19T10:47:38.027244Z","iopub.status.idle":"2021-09-19T10:47:38.147409Z","shell.execute_reply.started":"2021-09-19T10:47:38.027205Z","shell.execute_reply":"2021-09-19T10:47:38.146676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_github.info())","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:38.148425Z","iopub.execute_input":"2021-09-19T10:47:38.14884Z","iopub.status.idle":"2021-09-19T10:47:38.159583Z","shell.execute_reply.started":"2021-09-19T10:47:38.148802Z","shell.execute_reply":"2021-09-19T10:47:38.158947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'There are {train_github[\"landmark_id\"].nunique()} unique landmarks')","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:38.160564Z","iopub.execute_input":"2021-09-19T10:47:38.160954Z","iopub.status.idle":"2021-09-19T10:47:38.259299Z","shell.execute_reply.started":"2021-09-19T10:47:38.160916Z","shell.execute_reply":"2021-09-19T10:47:38.258263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"github_id2landmark_id = train_github[['id', 'landmark_id']].set_index('id').squeeze().to_dict()","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:38.260475Z","iopub.execute_input":"2021-09-19T10:47:38.260731Z","iopub.status.idle":"2021-09-19T10:47:40.998579Z","shell.execute_reply.started":"2021-09-19T10:47:38.260706Z","shell.execute_reply":"2021-09-19T10:47:40.997728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Extra Image Potential","metadata":{}},{"cell_type":"markdown","source":"This function computes the maximum number of additional images for a given fill value. To compute this the assumption is made that each class is filled up to the fill value. Thus with a fill value of 20 the assumption is made the number of samples for each class will be filled up to 20. As can be seen, with a fill value of 100 the additional image potential is over 6 million!","metadata":{}},{"cell_type":"code","source":"res = []\nfor n in tqdm(range(101)):\n    potential = 0\n    for k, count in original_landmark_id2count.items():\n        potential += max(0, n - count)\n    res.append(potential)","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:40.99976Z","iopub.execute_input":"2021-09-19T10:47:41.000009Z","iopub.status.idle":"2021-09-19T10:47:43.842115Z","shell.execute_reply.started":"2021-09-19T10:47:40.999985Z","shell.execute_reply":"2021-09-19T10:47:43.841153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12, 6))\npd.Series(res).plot()\nplt.grid()\nplt.title(f'Potential Number of Extra Train Images per Threshold', size=18)\nplt.xlabel('Threshold', size=16)\nplt.ylabel('Potential Number of Extra Train Images', size=16)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:43.843574Z","iopub.execute_input":"2021-09-19T10:47:43.843956Z","iopub.status.idle":"2021-09-19T10:47:44.155818Z","shell.execute_reply.started":"2021-09-19T10:47:43.843917Z","shell.execute_reply":"2021-09-19T10:47:44.154876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.DataFrame({ 'Potential Number of Extra Train Images': res[:26] })","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:47:44.157224Z","iopub.execute_input":"2021-09-19T10:47:44.157621Z","iopub.status.idle":"2021-09-19T10:47:44.167841Z","shell.execute_reply.started":"2021-09-19T10:47:44.157572Z","shell.execute_reply":"2021-09-19T10:47:44.167012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Process Download","metadata":{}},{"cell_type":"code","source":"# Process Extraced Images, beating heart of this notebook\ndef process_download(idx):\n    # Get all paths to the newly downloaded images\n    file_paths = glob.glob('/kaggle/working/temp/*/*/*/*.jpg')\n    new_train_data = 0\n    for file_path in file_paths:\n        # Get Image ID to check for duplicates\n        image_id = file_path.split('/')[-1].split('.')[0]\n        landmark_id = github_id2landmark_id[image_id]\n        # Check for duplicates and check if class is under Kaggle dataset\n        if landmark_id in train_original_landmark_ids and image_id not in original_ids:\n            # Only add image if class count is below threshold\n            count = original_landmark_id2count[landmark_id]\n            if count < THRESHOLD:\n                # Increase class count\n                original_landmark_id2count[landmark_id] += 1\n                # Increase newly found images count\n                new_train_data += 1\n                # Continue, do not remove this image\n                continue\n        # Remove image\n        os.remove(file_path)\n\n    # Ratio of images kept\n    keep_ratio = new_train_data / len(file_paths) * 100\n    # Count total new training data\n    total_new_files = new_train_data + len(glob.glob('/kaggle/working/train/*/*/*/*.jpg'))\n    # Print info\n    if idx % 10 == 0:\n        print(\n            f'{idx:03d} | ' +\n            f'{str(new_train_data).rjust(4)}/{len(file_paths)} ' +\n            f'({keep_ratio:05.2f}%) images kept' +\n            f', total new files: {total_new_files}'\n        )","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:54:38.068345Z","iopub.execute_input":"2021-09-19T10:54:38.068681Z","iopub.status.idle":"2021-09-19T10:54:38.076282Z","shell.execute_reply.started":"2021-09-19T10:54:38.068652Z","shell.execute_reply":"2021-09-19T10:54:38.075211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Downsize Images","metadata":{}},{"cell_type":"markdown","source":"The notebook disk size limit is just 20GB, therefore the images are downsized to have a smaller side of 384 pixels. This allows for more new training data!","metadata":{}},{"cell_type":"code","source":"def downsize_single_image(fp):\n    img = imageio.imread(fp)\n    h, w, _ = img.shape\n\n    # Check whether image is bigger than IMG_SIZE\n    if min(h,w) > IMG_SIZE:\n        r = IMG_SIZE / min(w, h)\n        w_resize = int(w * r)\n        h_resize = int(h * r)\n        # Resize using high quality LANCZOS algorithm\n        img = cv2.resize(img, (w_resize, h_resize), interpolation=cv2.INTER_LANCZOS4)\n        # Save as JPEG with quality set to 70, just as original images\n        img_jpeg = tf.io.encode_jpeg(img, quality=70, optimize_size=True).numpy()\n        # Overwrite image with lower res version\n        with open(fp, 'wb') as f:\n            f.write(img_jpeg)\n\n# Downsize images in parallel, speeds up the whole process\ndef downsize_images_parallel():\n    jobs = [joblib.delayed(downsize_single_image)(fp) for fp in glob.glob('/kaggle/working/temp/*/*/*/*.jpg')]\n    joblib.Parallel(\n        n_jobs=cpu_count(),\n        verbose=0,\n        require='sharedmem'\n    )(jobs)","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:54:38.840607Z","iopub.execute_input":"2021-09-19T10:54:38.840933Z","iopub.status.idle":"2021-09-19T10:54:49.364657Z","shell.execute_reply.started":"2021-09-19T10:54:38.840907Z","shell.execute_reply":"2021-09-19T10:54:49.363613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Add Extra Training Data","metadata":{}},{"cell_type":"code","source":"!rm -rf *","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:54:49.366092Z","iopub.execute_input":"2021-09-19T10:54:49.366393Z","iopub.status.idle":"2021-09-19T10:54:50.638485Z","shell.execute_reply.started":"2021-09-19T10:54:49.366363Z","shell.execute_reply":"2021-09-19T10:54:50.637242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fill Value\nTHRESHOLD = 20\n# Downsize Image Resolution\nIMG_SIZE = 384\n# Number of cores\nN_CORES = cpu_count()","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:54:50.640448Z","iopub.execute_input":"2021-09-19T10:54:50.640715Z","iopub.status.idle":"2021-09-19T10:54:50.645448Z","shell.execute_reply.started":"2021-09-19T10:54:50.640687Z","shell.execute_reply":"2021-09-19T10:54:50.644385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!mkdir train temp","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:54:50.646817Z","iopub.execute_input":"2021-09-19T10:54:50.647141Z","iopub.status.idle":"2021-09-19T10:54:51.478473Z","shell.execute_reply.started":"2021-09-19T10:54:50.647115Z","shell.execute_reply":"2021-09-19T10:54:51.477275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Install AXEL for multithreading download\n!apt-get -qq install axel","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:54:51.480084Z","iopub.execute_input":"2021-09-19T10:54:51.480411Z","iopub.status.idle":"2021-09-19T10:54:53.926877Z","shell.execute_reply.started":"2021-09-19T10:54:51.48038Z","shell.execute_reply":"2021-09-19T10:54:53.925827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The files are split up in 500 TAR files, they will all be downloaded and processed. Yes, that's processing half a Terabyte, over 4 million images, in about 6 hours.","metadata":{}},{"cell_type":"code","source":"# Process all TAR files\nfor i in tqdm(range(0, 500)):\n    idx = str(i).rjust(3, '0')\n    file = f'images_{idx}.tar'\n\n    # Get tar file, downloaded in parallel for speedup\n    !axel -q -n \"$N_CORES\" \"https://s3.amazonaws.com/google-landmark/train/$file\" -o \"temp\"\n\n    # Extract tar file\n    !tar -xf \"/kaggle/working/temp/$file\" -C \"/kaggle/working/temp\"\n\n    # Process Download\n    process_download(i)\n    \n    # Downsize Images in parallel\n    downsize_images_parallel()\n    \n    # Remove tar file\n    !rm -rf \"/kaggle/working/temp/$file\"\n\n    # Move all accepted images\n    for source in glob.glob('/kaggle/working/temp/*'):\n        !cp -r \"$source\" \"/kaggle/working/train\"\n        !rm -rf \"$source\"","metadata":{"execution":{"iopub.status.busy":"2021-09-19T10:54:53.928482Z","iopub.execute_input":"2021-09-19T10:54:53.928768Z","iopub.status.idle":"2021-09-19T11:04:21.004953Z","shell.execute_reply.started":"2021-09-19T10:54:53.928739Z","shell.execute_reply":"2021-09-19T11:04:21.003246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Mean Image Size","metadata":{}},{"cell_type":"code","source":"# Computes the mean images size in bytes, used for debug purposes\nfile_paths = glob.glob('/kaggle/working/train/*/*/*/*.jpg')\nmean_img_size = 0\nfor fp in tqdm(file_paths):\n    with open(fp, 'rb') as f:\n        mean_img_size += len(f.read()) / len(file_paths)\n        \nprint(f'Mean image size: {mean_img_size / 2**10:.2f}KB')\nprint(f'Maximum amount of images in 20GB dataset: {20 * 2**30 / mean_img_size / 1000:.1f}K')","metadata":{"execution":{"iopub.status.busy":"2021-08-26T13:51:14.759677Z","iopub.execute_input":"2021-08-26T13:51:14.760376Z","iopub.status.idle":"2021-08-26T13:51:15.161246Z","shell.execute_reply.started":"2021-08-26T13:51:14.760335Z","shell.execute_reply":"2021-08-26T13:51:15.15977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create Train Extra DataFrame","metadata":{}},{"cell_type":"code","source":"train_extra_list = []\n\nfor file_path in glob.glob('/kaggle/working/train/*/*/*/*.jpg'):\n    image_id = file_path.split('/')[-1].split('.')[0]\n    landmark_id = github_id2landmark_id[image_id]\n    \n    train_extra_list.append({ 'id': image_id, 'landmark_id': landmark_id })","metadata":{"execution":{"iopub.status.busy":"2021-08-26T13:51:15.482748Z","iopub.execute_input":"2021-08-26T13:51:15.483666Z","iopub.status.idle":"2021-08-26T13:51:15.521081Z","shell.execute_reply.started":"2021-08-26T13:51:15.48362Z","shell.execute_reply":"2021-08-26T13:51:15.520178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_extra = pd.DataFrame.from_dict(train_extra_list)","metadata":{"execution":{"iopub.status.busy":"2021-08-26T13:51:16.973691Z","iopub.execute_input":"2021-08-26T13:51:16.974788Z","iopub.status.idle":"2021-08-26T13:51:16.996425Z","shell.execute_reply.started":"2021-08-26T13:51:16.97468Z","shell.execute_reply":"2021-08-26T13:51:16.995215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_extra.head())","metadata":{"execution":{"iopub.status.busy":"2021-08-26T13:51:17.453714Z","iopub.execute_input":"2021-08-26T13:51:17.45431Z","iopub.status.idle":"2021-08-26T13:51:17.474253Z","shell.execute_reply.started":"2021-08-26T13:51:17.454263Z","shell.execute_reply":"2021-08-26T13:51:17.472917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_extra.info())","metadata":{"execution":{"iopub.status.busy":"2021-08-26T13:51:17.615594Z","iopub.execute_input":"2021-08-26T13:51:17.615988Z","iopub.status.idle":"2021-08-26T13:51:17.644182Z","shell.execute_reply.started":"2021-08-26T13:51:17.615951Z","shell.execute_reply":"2021-08-26T13:51:17.64145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save Train Extra DataFrame with ID and Landmark ID\ntrain_extra.to_pickle('train_extra.pkl.xz')","metadata":{"execution":{"iopub.status.busy":"2021-08-26T13:51:17.730701Z","iopub.execute_input":"2021-08-26T13:51:17.731123Z","iopub.status.idle":"2021-08-26T13:51:17.787784Z","shell.execute_reply.started":"2021-08-26T13:51:17.731084Z","shell.execute_reply":"2021-08-26T13:51:17.78625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Sanity check, there should be no duplicate images present in both train and train_extra\nduplicate_landmark_ids = len(set(train_extra['id']).intersection(set(train_original['id'])))\nprint(f'Found {duplicate_landmark_ids} landmark-ids occuring both in the original and extra dataset')","metadata":{"execution":{"iopub.status.busy":"2021-08-26T13:51:19.632224Z","iopub.execute_input":"2021-08-26T13:51:19.632695Z","iopub.status.idle":"2021-08-26T13:51:20.490748Z","shell.execute_reply.started":"2021-08-26T13:51:19.632634Z","shell.execute_reply":"2021-08-26T13:51:20.489156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Zip Dataset","metadata":{}},{"cell_type":"markdown","source":"This step is extremely important, all images must be zipped. Otherwise the notebook crashes, it will add all images as output when the notebook result is converted to HTML. It will thus add over 400K images to a HTML file, the notebook will fail. When zipping the images this will not happen. When creating the dataset Kaggle will automatically unzip the files.","metadata":{}},{"cell_type":"code","source":"for source in tqdm(glob.glob('/kaggle/working/train/*')):\n    # Ignore files\n    if '.' not in source:\n        print(f'Zipping folder {source}')\n        folder = source.split('/')[-1]\n        target = f'{folder}.zip'\n        # Zip\n        !cd \"/kaggle/working/train\" ; zip -qr \"$target\" \"$folder\"\n        # Remove original folder\n        !rm -rf \"$source\"","metadata":{"execution":{"iopub.status.busy":"2021-08-26T13:51:20.493042Z","iopub.execute_input":"2021-08-26T13:51:20.49358Z","iopub.status.idle":"2021-08-26T13:51:26.563026Z","shell.execute_reply.started":"2021-08-26T13:51:20.493527Z","shell.execute_reply":"2021-08-26T13:51:26.56218Z"},"trusted":true},"execution_count":null,"outputs":[]}]}