{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nWelcome to the fourth Landmark Retrieval competition! This year, we introduce a lot more diversity in the challenge’s test images in order to measure global landmark retrieval performance in a fairer manner. And following last year’s success, we set this up as a code competition.\n\nImage retrieval is a central problem in computer vision, relevant to many applications. The problem is usually posed as follows: given a query image, can you find similar images in a large database? This is especially important for query images containing landmarks, which accounts for a large portion of what people like to photograph.\n\n> In this competition, you are asked to develop models that can efficiently retrieve landmark images from a large database. The training set is available in the train/ folder, with corresponding landmark labels in train.csv. The query images are listed in the test/ folder, while the \"index\" images from which you are retrieving are listed in index/. Each image has a unique id. Since there are a large number of images, each image is placed within three subfolders according to the first three characters of the image id (i.e. image abcdef.jpg is placed in a/b/c/abcdef.jpg).","metadata":{}},{"cell_type":"markdown","source":"# Info\n\n\nSubmissions are given 12 hours to run, as compared to the site-wide session limit of 9 hours. While your commit must still finish in the 9 hour limit in order to be eligible to submit, the rerun may take the full 12 hours.\n\n* train.csv: This file contains, ids and targets\n - id: image id\n - landmark_id: target landmark id\n ","metadata":{}},{"cell_type":"code","source":"import os\n\n\nimport random\nimport seaborn as sns\nimport cv2\n\n# General packages\nimport pandas as pd\nimport numpy as np\nimport matplotlib\nimport matplotlib.pyplot as plt\nimport PIL\nimport IPython.display as ipd\nimport glob\nimport h5py\nimport plotly.graph_objs as go\nimport plotly.express as px\nfrom PIL import Image\nfrom tempfile import mktemp\n\nfrom bokeh.plotting import figure, output_notebook, show\nfrom math import pi\n\noutput_notebook()\n\n\nfrom IPython.display import Image, display\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:19.041048Z","iopub.execute_input":"2021-08-13T07:56:19.041735Z","iopub.status.idle":"2021-08-13T07:56:22.47041Z","shell.execute_reply.started":"2021-08-13T07:56:19.041614Z","shell.execute_reply":"2021-08-13T07:56:22.469306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.listdir('../input/landmark-retrieval-2021/')","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:22.472033Z","iopub.execute_input":"2021-08-13T07:56:22.472349Z","iopub.status.idle":"2021-08-13T07:56:22.483689Z","shell.execute_reply.started":"2021-08-13T07:56:22.472318Z","shell.execute_reply":"2021-08-13T07:56:22.482406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATASET_DIR = '../input/landmark-retrieval-2021/'\n\nTRAIN_IMAGE_DIR = f'{DATASET_DIR}/train'\nTEST_IMAGE_DIR = f'{DATASET_DIR}/test'\ntrain = pd.read_csv(f'{DATASET_DIR}/train.csv')\nSUB = pd.read_csv(f'{DATASET_DIR}/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:22.486359Z","iopub.execute_input":"2021-08-13T07:56:22.486817Z","iopub.status.idle":"2021-08-13T07:56:24.548007Z","shell.execute_reply.started":"2021-08-13T07:56:22.486769Z","shell.execute_reply":"2021-08-13T07:56:24.546812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train.head())\nprint(\"Shape of train_data :\", train.shape)","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:24.550086Z","iopub.execute_input":"2021-08-13T07:56:24.55054Z","iopub.status.idle":"2021-08-13T07:56:24.579665Z","shell.execute_reply.started":"2021-08-13T07:56:24.550489Z","shell.execute_reply":"2021-08-13T07:56:24.578642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark = train.landmark_id.value_counts()\nlandmark_df = pd.DataFrame({'landmark_id':landmark.index, 'frequency':landmark.values}).head(30)\n\nlandmark_df['landmark_id'] =   landmark_df.landmark_id.apply(lambda x: f'landmark_id_{x}')\n\nfig = px.bar(landmark_df, x=\"frequency\", y=\"landmark_id\",color='landmark_id',\n             hover_data=[\"landmark_id\", \"frequency\"],\n             height=1000,\n             title='Number of images per landmark_id (Top 30 landmark_ids)',\n             color_discrete_sequence=px.colors.sequential.RdBu)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:24.581009Z","iopub.execute_input":"2021-08-13T07:56:24.581305Z","iopub.status.idle":"2021-08-13T07:56:26.226665Z","shell.execute_reply.started":"2021-08-13T07:56:24.581274Z","shell.execute_reply":"2021-08-13T07:56:26.225606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark.hist()","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:26.2278Z","iopub.execute_input":"2021-08-13T07:56:26.228078Z","iopub.status.idle":"2021-08-13T07:56:26.482836Z","shell.execute_reply.started":"2021-08-13T07:56:26.228048Z","shell.execute_reply":"2021-08-13T07:56:26.481952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Landmark ID distribution\nplt.figure(figsize = (10, 8))\nplt.title('Landmark ID Distribuition')\nsns.distplot(train['landmark_id'])\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:26.484008Z","iopub.execute_input":"2021-08-13T07:56:26.484418Z","iopub.status.idle":"2021-08-13T07:56:35.229946Z","shell.execute_reply.started":"2021-08-13T07:56:26.484385Z","shell.execute_reply":"2021-08-13T07:56:35.228735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set()\nplt.title('Training set: number of images per class(line plot)')\nsns.set_color_codes(\"pastel\")\nlandmarks_fold = pd.DataFrame(train['landmark_id'].value_counts())\nlandmarks_fold.reset_index(inplace=True)\nlandmarks_fold.columns = ['landmark_id','count']\nax = landmarks_fold['count'].plot(logy=True, grid=True)\nlocs, labels = plt.xticks()\nplt.setp(labels, rotation=30)\nax.set(xlabel=\"Landmarks\", ylabel=\"Number of images\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:35.232575Z","iopub.execute_input":"2021-08-13T07:56:35.232931Z","iopub.status.idle":"2021-08-13T07:56:35.958419Z","shell.execute_reply.started":"2021-08-13T07:56:35.232897Z","shell.execute_reply":"2021-08-13T07:56:35.957255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualize outliers, min/max or quantiles of the landmarks count\nsns.set()\nax = landmarks_fold.boxplot(column='count')\nax.set_yscale('log')","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:35.960116Z","iopub.execute_input":"2021-08-13T07:56:35.96041Z","iopub.status.idle":"2021-08-13T07:56:36.490748Z","shell.execute_reply.started":"2021-08-13T07:56:35.960381Z","shell.execute_reply":"2021-08-13T07:56:36.489776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.landmark_id.nunique()","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:36.492161Z","iopub.execute_input":"2021-08-13T07:56:36.492493Z","iopub.status.idle":"2021-08-13T07:56:36.550355Z","shell.execute_reply.started":"2021-08-13T07:56:36.492459Z","shell.execute_reply":"2021-08-13T07:56:36.549255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- There are 81313 unique landmark_ids","metadata":{}},{"cell_type":"code","source":"landmark[:5]","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:36.551748Z","iopub.execute_input":"2021-08-13T07:56:36.552069Z","iopub.status.idle":"2021-08-13T07:56:36.560167Z","shell.execute_reply.started":"2021-08-13T07:56:36.552038Z","shell.execute_reply":"2021-08-13T07:56:36.55916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- There is only one landmark which has more than 2300 images (landmark_id: 138982)","metadata":{}},{"cell_type":"code","source":"landmark.describe()","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:36.561523Z","iopub.execute_input":"2021-08-13T07:56:36.561864Z","iopub.status.idle":"2021-08-13T07:56:36.583887Z","shell.execute_reply.started":"2021-08-13T07:56:36.561833Z","shell.execute_reply":"2021-08-13T07:56:36.58272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Number of images per landmark_id ranges from 2 to 6272.\n- median is 9, mean is 19\n","metadata":{}},{"cell_type":"code","source":"landmark[landmark < 100].shape","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:36.585321Z","iopub.execute_input":"2021-08-13T07:56:36.585756Z","iopub.status.idle":"2021-08-13T07:56:36.594965Z","shell.execute_reply.started":"2021-08-13T07:56:36.585716Z","shell.execute_reply":"2021-08-13T07:56:36.593788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark.shape","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:36.59655Z","iopub.execute_input":"2021-08-13T07:56:36.596937Z","iopub.status.idle":"2021-08-13T07:56:36.608954Z","shell.execute_reply.started":"2021-08-13T07:56:36.5969Z","shell.execute_reply":"2021-08-13T07:56:36.607799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n- Out of 81313, there are 79298 (97.5%) landmark_ids with less than 100 images.","metadata":{"execution":{"iopub.status.busy":"2021-08-11T15:17:47.150771Z","iopub.execute_input":"2021-08-11T15:17:47.15117Z","iopub.status.idle":"2021-08-11T15:17:47.158171Z","shell.execute_reply.started":"2021-08-11T15:17:47.151136Z","shell.execute_reply":"2021-08-11T15:17:47.156411Z"}}},{"cell_type":"code","source":"import PIL\nfrom PIL import Image, ImageDraw\n\n\ndef display_images(images, title=None): \n    \"\"\"\n    func for display images \n    Thank you @rohitsingh9990 for this fucntion\n    \"\"\"\n    f, ax = plt.subplots(5,5, figsize=(18,22))\n    if title:\n        f.suptitle(title, fontsize = 30)\n\n    for i, image_id in enumerate(images):\n        image_path = os.path.join(TRAIN_IMAGE_DIR, f'{image_id[0]}/{image_id[1]}/{image_id[2]}/{image_id}.jpg')\n        image = Image.open(image_path)\n        \n        ax[i//5, i%5].imshow(image) \n        image.close()       \n        ax[i//5, i%5].axis('off')\n\n        landmark_id = train[train.id==image_id.split('.')[0]].landmark_id.values[0]\n        ax[i//5, i%5].set_title(f\"ID: {image_id.split('.')[0]}\\nLandmark_id: {landmark_id}\", fontsize=\"12\")\n\n    plt.show() ","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-08-13T07:56:36.610569Z","iopub.execute_input":"2021-08-13T07:56:36.610907Z","iopub.status.idle":"2021-08-13T07:56:36.624408Z","shell.execute_reply.started":"2021-08-13T07:56:36.610875Z","shell.execute_reply":"2021-08-13T07:56:36.622977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualizing","metadata":{}},{"cell_type":"code","source":"samples = train.sample(25).id.values\ndisplay_images(samples, 'Random')","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:36.626356Z","iopub.execute_input":"2021-08-13T07:56:36.627267Z","iopub.status.idle":"2021-08-13T07:56:47.273979Z","shell.execute_reply.started":"2021-08-13T07:56:36.627208Z","shell.execute_reply":"2021-08-13T07:56:47.27284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samples = train[train.landmark_id == 138982].sample(25).id.values\ndisplay_images(samples, 'Top 1')","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:47.275331Z","iopub.execute_input":"2021-08-13T07:56:47.275851Z","iopub.status.idle":"2021-08-13T07:56:57.800791Z","shell.execute_reply.started":"2021-08-13T07:56:47.275809Z","shell.execute_reply":"2021-08-13T07:56:57.799313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samples = train[train.landmark_id == 126637].sample(25).id.values\ndisplay_images(samples, 'Top 2')","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:56:57.802314Z","iopub.execute_input":"2021-08-13T07:56:57.802629Z","iopub.status.idle":"2021-08-13T07:57:08.629067Z","shell.execute_reply.started":"2021-08-13T07:56:57.802598Z","shell.execute_reply":"2021-08-13T07:57:08.627964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samples = train[train.landmark_id == 20409].sample(25).id.values\ndisplay_images(samples, 'Top 3')","metadata":{"execution":{"iopub.status.busy":"2021-08-13T07:57:08.63039Z","iopub.execute_input":"2021-08-13T07:57:08.630705Z","iopub.status.idle":"2021-08-13T07:57:18.984911Z","shell.execute_reply.started":"2021-08-13T07:57:08.630673Z","shell.execute_reply":"2021-08-13T07:57:18.98281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}