{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Google Landmark Retrieval 2020 - EDA\n<img src=\"https://github.com/seriousran/img_link/blob/master/kg/glr.jpg?raw=true\" alt=\"drawing\" style=\"width:520px;\"/>\n\n\"This year, we have worked to set this up as a code competition and we have completely refreshed the test and index image sets.\"\n\nLet's start to dig it! :)\n\n### Past Related Competitions\n\n1. [Google Landmark Retrieval 2019](https://www.kaggle.com/c/landmark-retrieval-2019)\n1. [Google Landmark Retrieval Challenge](https://www.kaggle.com/c/landmark-retrieval-challenge)\n\n### The winner in last competition\n<img src=\"https://github.com/seriousran/img_link/blob/master/kg/lb_glr_2019.PNG?raw=true\" alt=\"drawing\"/>\n\n\n\nReference: https://www.kaggle.com/huangxiaoquan/google-landmarks-v2-exploratory-data-analysis-eda","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## Outline\n1. [File Exploration](#1)\n    1. [Training data](#2)\n    1. [Index data](#3)\n    1. [Display examples](#4)","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nimport glob\nimport cv2\nimport numpy as np \nimport pandas as pd \nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\nfrom scipy import stats\n\n\n\n%matplotlib inline\n%config InlineBackend.figure_format = 'retina'","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 1.1 Train data\n\nIn this competition, you are asked to develop models that can efficiently retrieve landmark images from a large database. \nThe training set is available in the train/ folder, with corresponding landmark labels in train.csv. ","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('../input/landmark-retrieval-2020/train.csv')\ntrain_df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Landmark_id distribuition","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.title('landmark_id distribution')\nsns.distplot(train_df['landmark_id'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Training set: number of images per class(line plot)\n\n","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.set()\nplt.title('Training set: number of images per class(line plot)')\nlandmarks_fold = pd.DataFrame(train_df['landmark_id'].value_counts())\nlandmarks_fold.reset_index(inplace=True)\nlandmarks_fold.columns = ['landmark_id','count']\nax = landmarks_fold['count'].plot(logy=True, grid=True)\nlocs, labels = plt.xticks()\nplt.setp(labels, rotation=30)\nax.set(xlabel=\"Landmarks\", ylabel=\"Number of images\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Training set: number of images per class(scatter plot)","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.set()\nlandmarks_fold_sorted = pd.DataFrame(train_df['landmark_id'].value_counts())\nlandmarks_fold_sorted.reset_index(inplace=True)\nlandmarks_fold_sorted.columns = ['landmark_id','count']\nlandmarks_fold_sorted = landmarks_fold_sorted.sort_values('landmark_id')\nax = landmarks_fold_sorted.plot.scatter(\\\n     x='landmark_id',y='count',\n     title='Training set: number of images per class(statter plot)')\nlocs, labels = plt.xticks()\nplt.setp(labels, rotation=30)\nax.set(xlabel=\"Landmarks\", ylabel=\"Number of images\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 1.2 Test and Index data\n\nThe query images are listed in the test/ folder, while the \"index\" images from which you are retrieving are listed in index/. \n\nEach image has a unique id. Since there are a large number of images, each image is placed within three subfolders according to the first three characters of the image id (i.e. image abcdef.jpg is placed in a/b/c/abcdef.jpg).\n\n0-f in 0-f in 0-f","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"test_list = glob.glob('../input/landmark-retrieval-2020/test/*/*/*/*')\nindex_list = glob.glob('../input/landmark-retrieval-2020/index/*/*/*/*')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print( 'Query', len(test_list), ' test images in ', len(index_list), 'index images')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 1.3 Display examples","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.rcParams[\"axes.grid\"] = False\nf, axarr = plt.subplots(4, 3, figsize=(24, 22))\n\ncurr_row = 0\nfor i in range(12):\n    example = cv2.imread(test_list[i])\n    example = example[:,:,::-1]\n    \n    col = i%4\n    axarr[col, curr_row].imshow(example)\n    if col == 3:\n        curr_row += 1\n            \n#     plt.imshow(example)\n#     plt.show()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}