{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Data Cleaning\n\nThis notebook is the continuation of my data science project after performing the [Exploratoty Analysis](http://www.kaggle.com/nikoladyulgerov/labelling-famous-landmarks-eda) (**EDA**) on **Google Landmark Recognition 2020** dataset. The aim of this step is to clean and prepare the image data with which I am going to feed later on the constructed models.","metadata":{}},{"cell_type":"markdown","source":"# Index of Contents\n\n1. **Importing libraries**\n2. **Approach overview**\n3. **Data preparation**\n4. **Image sizes**","metadata":{}},{"cell_type":"markdown","source":"# Importing libraries\n\nThe modules are the most basic and common one for this kind of work.","metadata":{}},{"cell_type":"code","source":"import os\nimport shutil\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt \n\nfrom PIL import Image\nfrom sklearn.model_selection import train_test_split\nfrom tqdm.auto import tqdm\n\n#Seed for making reproducible experiments\nseed = 612","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The next step is loading the dataset.","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv(\"../input/landmark-recognition-2020/train.csv\")\ntrain_data.sample(5, random_state=seed)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we have checked that is correctly loaded, we are ready for work.","metadata":{}},{"cell_type":"markdown","source":"# Approach Overview\n\n\nThe [Exploratoty Analysis](http://www.kaggle.com/nikoladyulgerov/labelling-famous-landmarks-eda) gave us clues about the nature of this particular dataset. Let's review some of its characteristics, just to sum up what we should do now to deal with it.\n\n* Huge unbalanced distribution of classes\n* High variance inside classes\n* Different image sizes\n\nDue to the computational limits I am experimenting with this project I am not going to use the entire dataset, it is impossible for me to handle this amount of data. That is why I will divide the process into two parts:\n\n* First, set the model construction with the top `200` classes on Kaggle and Google Colab platform.\n* Secondly, scale down the number of classes and see how models behaviour with the only the top `20` classes.\n\nAlthough these classes do not represent a high percent of the database, they have a great number of images what is the main reason why I have selected them. With that in mind, the balance of labels will be in some way guaranteed, although some variance stills remaining. However, it can be handle with simple techniques.","metadata":{}},{"cell_type":"code","source":"print(\"Number of total labels \",  train_data[\"landmark_id\"].nunique())\nprint(\"Number of total images \", train_data.shape[0])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I am getting rid of the most frequent landmark of the dataset because as it was seen in the exploratory analysis it has a huge difference of images against the rest. My intention with this is to minimize the unbalance on the sample data.","metadata":{}},{"cell_type":"code","source":"n = 201\ntop_200 = train_data['landmark_id'].value_counts()[1:n].index.tolist()\nimages_200 = train_data.loc[train_data[\"landmark_id\"].isin(top_200)]\nprint(\"Number of images of the top \",images_200[\"landmark_id\"].nunique(), \" landmarks: \", images_200.shape[0])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n = 21\ntop_20 = train_data['landmark_id'].value_counts()[1:n].index.tolist()\nimages_20 = train_data.loc[train_data[\"landmark_id\"].isin(top_20)]\nprint(\"Number of images of the top \",images_20[\"landmark_id\"].nunique(), \" landmarks: \", images_20.shape[0])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Before carrying on the data cleaning and preparation procees, let's take a look of the already sampled data.","metadata":{}},{"cell_type":"code","source":"# Calculating the n_images per landmark\ndb1 = images_200.groupby([\"landmark_id\"]).size().reset_index(name='n_images')\ndb2 = images_20.groupby([\"landmark_id\"]).size().reset_index(name='n_images')\n#Plotting a scatterplot of the distribution -> share y axis to compare\nfig, axes = plt.subplots(1, 2, sharey = True, figsize = (20,5))\n\nsns.scatterplot(ax=axes[0], x='landmark_id', y='n_images', data=db1, palette=\"mako\")\naxes[0].set_title(\"Top 200 most frequent landmarks\")\nsns.scatterplot(ax=axes[1], x='landmark_id', y='n_images', data=db2, palette=\"mako\")\naxes[1].set_title(\"Top 20 most frequent landmarks\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As it was mentioned, the images per labels are in some way more balanced than before. However, there are some outstanding classes with a huge difference of images against the others. Also, looking to the 200th classes graph it is very clear that as the range of classes increases, the images per class decreases because the points are highly concentrated on the bottom of the scatter plot.","metadata":{}},{"cell_type":"markdown","source":"To reduce even more the imbalance among classes, I am going to perform a downsample of the outstanding classes. My aim with that is stablishing a more uniform distribution among classes so that the performance of the models will not to be affected by this aspect.","metadata":{}},{"cell_type":"code","source":"def downsampling(data_df, counts_df, max_samples):\n    # Get the oustanding classes to downsample\n    outstand = counts_df[counts_df[\"n_images\"]>=max_samples]\n    rest = data_df[data_df[\"landmark_id\"].isin(outstand[\"landmark_id\"]) == False]\n    # Random downsample of these classes to max_samples\n    for row in outstand.itertuples():\n        # Get max_samples for specific landmark_id\n        temp_df = data_df[data_df[\"landmark_id\"] == row.landmark_id].sample(max_samples, random_state = seed)\n        \n        rest = pd.concat([rest, temp_df], axis = 0)\n        \n    return rest","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"images_200 = downsampling(images_200, db1, db1['n_images'].min()+200) #Get the minimum samples per class and add some more\nimages_20 = downsampling(images_20, db2, db2['n_images'].min()+200)\n\n# Calculating the n_images per landmark\ndb1 = images_200.groupby([\"landmark_id\"]).size().reset_index(name='n_images')\ndb2 = images_20.groupby([\"landmark_id\"]).size().reset_index(name='n_images')\n#Plotting a barplot of the distributions\nfig, axes = plt.subplots(1, 2,sharey = True, figsize = (20,5))\n\nsns.scatterplot(ax=axes[0], x='landmark_id', y='n_images', data=db1, palette=\"mako\")\naxes[0].set_title(\"Top 200 most frequent landmarks\")\n\nsns.scatterplot(ax=axes[1], x='landmark_id', y='n_images', data=db2, palette=\"flare\")\naxes[1].set_title(\"Top 20 most frequent landmarks\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With the downsampling function the graphs have changed their shapes. Now all the classes have very similar number of images due to the reduction of the oustanding ones. Moreover, it is more noticeable that as the number of classes increase the number of images per class decrease. The maximum samples per class are setted according to the minimum number of images plus a certain \"variation\" (`200`), just to avoid losing too many images. ","metadata":{}},{"cell_type":"markdown","source":"Making some calculus we can stablish the proportion of the original dataset that will be used in further project steps.","metadata":{}},{"cell_type":"code","source":"print(\"Top 200 landmarks represent {:.2f} % of the entire dataset\".format(images_200.shape[0]/train_data.shape[0] * 100))\nprint(\"Top 2000 landmarks represent {:.2f} % of the entire dataset\".format(images_20.shape[0]/train_data.shape[0] * 100))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Although I am not going to make use of a huge part of the dataset, these percentages represent a big number of images as we have seen previously. The problem is simplified due to time and cost constraints of the development of the project.","metadata":{}},{"cell_type":"markdown","source":"# Data Preparation\n\nThis stage consist on the preparation of the directory structure to be downloaded from the Kaggle platform. I will divide each of the two sample datasets obtained before in three main folders: **train**, **validation** and finally **test**. Here comes the other part of simplifying the problem. The test sets with which I am going to work are entirely composed by landmark images which is not the case for the given test folder by the Google competition.\n\nFollowing the standard norms and knowing that the sample datasets are big enough, the holdout of images that corresponds to each partition are going to be: `70%` train, `20%` validation and `10%` test.","metadata":{}},{"cell_type":"markdown","source":"The way of making the directory structure comes from how **Keras** works. The idea is to construct the optimal way to feed the models and do not waste time making unnecesary searches while training and evaluating them. It will look something like the following composition:\n\n`--data\n    |--train\n        |--landmark_1\n            |--image_1_landmark_1\n            |--image_2_landmark_1\n            ...\n        |--landmark_2\n        ...\n    |--validation\n    |--test\n  `","metadata":{}},{"cell_type":"markdown","source":"Before starting, there is something important to consider, the classes are not perfeclty balanced, so it is a good practice to make a stratified splits. For this task, I am going to use the well known module `sklearn` that has some built-in functions that make it easy.","metadata":{}},{"cell_type":"code","source":"X, y = images_200[\"id\"].to_list(), images_200[\"landmark_id\"].to_list()\nX_train, X_test_200, y_train, y_test_200 = train_test_split(X, y, test_size = 0.1, random_state = seed, stratify = y)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, the same process must be applied again in order to get the validation split. For that, the split is applied on the remaining train set.","metadata":{}},{"cell_type":"code","source":"X_train_200, X_val_200, y_train_200, y_val_200 = train_test_split(X_train, y_train, test_size = 0.22, random_state = seed, stratify = y_train) # 0.22 x 0.9 = 0.2","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For an ease of operations later, let's build the data frames for each set.","metadata":{}},{"cell_type":"code","source":"train = { 'id': X_train_200, 'landmark_id': y_train_200 }  \nval = { 'id': X_val_200, 'landmark_id': y_val_200 }  \ntest = { 'id': X_test_200, 'landmark_id': y_test_200 }  \n\ntrain_200 = pd.DataFrame(train)\nval_200 = pd.DataFrame(val)\ntest_200 = pd.DataFrame(test)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Making sure, the process is completed correctly","metadata":{}},{"cell_type":"code","source":"train_200.sample(5, random_state=seed)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_200.sample(5, random_state=seed)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_200.sample(5, random_state=seed)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Repeat the same with the sample of 20 classes.","metadata":{}},{"cell_type":"code","source":"X, y = images_20[\"id\"].to_list(), images_20[\"landmark_id\"].to_list()\nX_train, X_test_20, y_train, y_test_20 = train_test_split(X, y, test_size = 0.1, random_state = seed, stratify = y)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_20, X_val_20, y_train_20, y_val_20 = train_test_split(X_train, y_train, test_size = 0.22, random_state = seed, stratify = y_train) # 0.22 x 0.9 = 0.2","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = { 'id': X_train_20, 'landmark_id': y_train_20 }  \nval = { 'id': X_val_20, 'landmark_id': y_val_20 }  \ntest = { 'id': X_test_20, 'landmark_id': y_test_20 }  \n\ntrain_20 = pd.DataFrame(train)\nval_20 = pd.DataFrame(val)\ntest_20 = pd.DataFrame(test)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_20.sample(5, random_state=seed)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_20.sample(5, random_state=seed)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_20.sample(5, random_state=seed)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Before making the directories, let's save the `csv` files with the original paths. It is a good idea to keep track of the images we are going to use. By this way further experiments can be reproducible and easily shareable with the community.","metadata":{}},{"cell_type":"code","source":"def build_paths(df):\n    paths = [\"../input/landmark-recognition-2020/train/{}/{}/{}/{}.jpg\".format(row.id[0],row.id[1],row.id[2],row.id) for row in df.itertuples()]\n    df['path'] = paths\n    return df","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_200 = build_paths(train_200)\nval_200 = build_paths(val_200)\ntest_200 = build_paths(test_200)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Checking everything is correct.","metadata":{}},{"cell_type":"code","source":"train_200.sample(5, random_state=seed)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_200.sample(5, random_state=seed)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_200.sample(5, random_state=seed)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Saving the files.","metadata":{}},{"cell_type":"code","source":"train_200.to_csv('train_200.csv',index=False)\nval_200.to_csv('val_200.csv',index=False)\ntest_200.to_csv('test_200.csv',index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The same operation with the other sample.","metadata":{}},{"cell_type":"code","source":"train_20 = build_paths(train_20)\nval_20 = build_paths(val_20)\ntest_20 = build_paths(test_20)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Checking everything is correct.","metadata":{}},{"cell_type":"code","source":"train_20.sample(5, random_state=seed)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_20.sample(5, random_state=seed)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_20.sample(5, random_state=seed)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Saving the files.","metadata":{}},{"cell_type":"code","source":"train_20.to_csv('train_20.csv',index=False)\nval_20.to_csv('val_20.csv',index=False)\ntest_20.to_csv('test_20.csv',index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, we are ready for making the directory structure.","metadata":{}},{"cell_type":"code","source":"def making_folders(df, c_col, i_col, path):\n    shutil.os.mkdir(path)\n    classes = df[c_col].unique()\n    for c in tqdm(classes):\n        f_path = f\"{path}{c}\"\n        shutil.os.mkdir(f_path)\n        imgs = df.loc[df[c_col] == c][i_col].to_list()\n        for i in imgs:\n            image_path = \"../input/landmark-recognition-2020/train/{}/{}/{}/{}.jpg\".format(i[0],i[1],i[2],i)\n            shutil.copy(image_path, f_path)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# making_folders(train_200, \"landmark_id\", \"id\", \"..train_200/\")\n# making_folders(val_200, \"landmark_id\", \"id\", \"..val_200/\")\n# making_folders(test_200, \"landmark_id\", \"id\", \"..test_200/\")\nmaking_folders(train_20, \"landmark_id\", \"id\", \"..train_20/\")\n# making_folders(val_20, \"landmark_id\", \"id\", \"..val_20/\")\n# making_folders(test_20, \"landmark_id\", \"id\", \"..test_20/\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"! ls -a","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"At this point, some of the limitations of the Kaggle platform appeared. There is a constraint about the space disk one can use, so this task must be done one by one. That is the reason the code above is commented. Moreover, it is impossible to publish the notebook with all the created folders because the **output** or **working directory** only supports `5GB`. What I did is download one by one and just the first part, the dataset with the 200 labels because the other one exceeded all the limits. My intentions is once constructed the models, come back and download the rest to scale down the builded neural networks with less classes and images.","metadata":{}},{"cell_type":"markdown","source":"# Image sizes\n\nAfter thinking about the problem with the different image size of the dataset I considered to download the original ones. The main is reason is that I can always go back one step instead of resizing all without already knowing what type of neural network I am going to construct. So, I postpone this task for the next step to make it more suitable and with the aim of a bigger control on the factors of the efficiency of the future models.\n\n----Note **Latest Version**----\n\nOnce I modeled the first neural network I found that resizing the images with ImageDataGenerator class or special Keras layers, the time per epoch was very high due to the sizing task. For each execution, the images were resized what was a waste of time. To make the process faster what I did was resizing all images once and save them, ready to use without waiting too much. \n\nThe other problem that comes with resizing the images is deciding what size to give them. A very big size increase the time to process because there are more features to exctract from an image. On the other hand, a smaller size decrease the execution time, but some features can be overlooked by the neural network.","metadata":{}},{"cell_type":"code","source":"def resize(path, size):\n    subfolders = [f.path for f in os.scandir(path) if f.is_dir()]\n    for s in tqdm(subfolders):\n        for i in os.listdir(s):\n            full_path = os.path.join(s,i)\n            if os.path.isfile(full_path):\n                im = Image.open(full_path)\n                imResize = im.resize(size, Image.ANTIALIAS)\n                imResize.save(full_path)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Repeat the same for every folder\ntrain = \"..train_20/\"\nval = \"..val_20/\"\ntest = \"..test_20/\"\n\nresize(train, (128,128))\n# resize(val, (128,128) )\n# resize(test, (128,128)) ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Repeat the same for every folder \n! zip -qr train_20.zip \"..train_20/\"","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}