{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nWelcome to the fourth Landmark Retrieval competition! This year, we introduce a lot more diversity in the challenge’s test images in order to measure global landmark retrieval performance in a fairer manner. And following last year’s success, we set this up as a code competition.\n\nImage retrieval is a central problem in computer vision, relevant to many applications. The problem is usually posed as follows: given a query image, can you find similar images in a large database? This is especially important for query images containing landmarks, which accounts for a large portion of what people like to photograph.\n\n> In this competition, you are asked to develop models that can efficiently retrieve landmark images from a large database. The training set is available in the train/ folder, with corresponding landmark labels in train.csv. The query images are listed in the test/ folder, while the \"index\" images from which you are retrieving are listed in index/. Each image has a unique id. Since there are a large number of images, each image is placed within three subfolders according to the first three characters of the image id (i.e. image abcdef.jpg is placed in a/b/c/abcdef.jpg).","metadata":{}},{"cell_type":"markdown","source":"# Info\n\n\nSubmissions are given 12 hours to run, as compared to the site-wide session limit of 9 hours. While your commit must still finish in the 9 hour limit in order to be eligible to submit, the rerun may take the full 12 hours.\n\n* train.csv: This file contains, ids and targets\n - id: image id\n - landmark_id: target landmark id\n ","metadata":{}},{"cell_type":"code","source":"import os\n\n\nimport random\nimport seaborn as sns\nimport cv2\n\n# General packages\nimport pandas as pd\nimport numpy as np\nimport matplotlib\nimport matplotlib.pyplot as plt\nimport PIL\nimport IPython.display as ipd\nimport glob\nimport h5py\nimport plotly.graph_objs as go\nimport plotly.express as px\nfrom PIL import Image\nfrom tempfile import mktemp\n\nfrom bokeh.plotting import figure, output_notebook, show\nfrom math import pi\n\noutput_notebook()\n\n\nfrom IPython.display import Image, display\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:42.711101Z","iopub.execute_input":"2022-01-17T11:35:42.711772Z","iopub.status.idle":"2022-01-17T11:35:46.009276Z","shell.execute_reply.started":"2022-01-17T11:35:42.711652Z","shell.execute_reply":"2022-01-17T11:35:46.007941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.listdir('../input/landmark-retrieval-2021/')","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:46.011031Z","iopub.execute_input":"2022-01-17T11:35:46.01135Z","iopub.status.idle":"2022-01-17T11:35:46.022762Z","shell.execute_reply.started":"2022-01-17T11:35:46.011319Z","shell.execute_reply":"2022-01-17T11:35:46.021296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATASET_DIR = '../input/landmark-retrieval-2021/'\n\nTRAIN_IMAGE_DIR = f'{DATASET_DIR}/train'\nTEST_IMAGE_DIR = f'{DATASET_DIR}/test'\ntrain = pd.read_csv(f'{DATASET_DIR}/train.csv')\nSUB = pd.read_csv(f'{DATASET_DIR}/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:46.025024Z","iopub.execute_input":"2022-01-17T11:35:46.025479Z","iopub.status.idle":"2022-01-17T11:35:47.945417Z","shell.execute_reply.started":"2022-01-17T11:35:46.025436Z","shell.execute_reply":"2022-01-17T11:35:47.944366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train.head())\nprint(\"Shape of train_data :\", train.shape)\ntrain.values()","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:47.947092Z","iopub.execute_input":"2022-01-17T11:35:47.947398Z","iopub.status.idle":"2022-01-17T11:35:48.378235Z","shell.execute_reply.started":"2022-01-17T11:35:47.94737Z","shell.execute_reply":"2022-01-17T11:35:48.376557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark = train.landmark_id.value_counts()\n#print(landmark)\n\nlandmark_df = pd.DataFrame({'landmark_id':landmark.index, 'frequency':landmark.values}).head(30)\nprint(landmark_df)\nlandmark_df['landmark_id'] =   landmark_df.landmark_id.apply(lambda x: f'landmark_id_{x}')\nprint(landmark_df)\n\nfig = px.bar(landmark_df, x=\"frequency\", y=\"landmark_id\",color='landmark_id',\n             hover_data=[\"landmark_id\", \"frequency\"],\n             height=1000,\n             title='Number of images per landmark_id (Top 30 landmark_ids)',\n             color_discrete_sequence=px.colors.sequential.RdBu)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:48.37914Z","iopub.status.idle":"2022-01-17T11:35:48.379666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark.hist()","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:48.380507Z","iopub.status.idle":"2022-01-17T11:35:48.381025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Landmark ID distribution\nplt.figure(figsize = (10, 8))\nplt.title('Landmark ID Distribuition')\nsns.distplot(train['landmark_id'])\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:48.381881Z","iopub.status.idle":"2022-01-17T11:35:48.382426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set()\nplt.title('Training set: number of images per class(line plot)')\nsns.set_color_codes(\"pastel\")\nlandmarks_fold = pd.DataFrame(train['landmark_id'].value_counts())\nlandmarks_fold.reset_index(inplace=True)\nlandmarks_fold.columns = ['landmark_id','count']\nax = landmarks_fold['count'].plot(logy=True, grid=True)\nlocs, labels = plt.xticks()\nplt.setp(labels, rotation=30)\nax.set(xlabel=\"Landmarks\", ylabel=\"Number of images\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:48.383211Z","iopub.status.idle":"2022-01-17T11:35:48.383705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualize outliers, min/max or quantiles of the landmarks count\nsns.set()\nax = landmarks_fold.boxplot(column='count')\nax.set_yscale('log')","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:48.384525Z","iopub.status.idle":"2022-01-17T11:35:48.385035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.landmark_id.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:48.385828Z","iopub.status.idle":"2022-01-17T11:35:48.386348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- There are 81313 unique landmark_ids","metadata":{}},{"cell_type":"code","source":"landmark[:5]","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:48.387106Z","iopub.status.idle":"2022-01-17T11:35:48.387602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- There is only one landmark which has more than 2300 images (landmark_id: 138982)","metadata":{}},{"cell_type":"code","source":"landmark.describe()","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:48.3884Z","iopub.status.idle":"2022-01-17T11:35:48.388901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Number of images per landmark_id ranges from 2 to 6272.\n- median is 9, mean is 19\n","metadata":{}},{"cell_type":"code","source":"landmark[landmark < 100].shape","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:48.389683Z","iopub.status.idle":"2022-01-17T11:35:48.390204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark.shape","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:48.390952Z","iopub.status.idle":"2022-01-17T11:35:48.391489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n- Out of 81313, there are 79298 (97.5%) landmark_ids with less than 100 images.","metadata":{"execution":{"iopub.status.busy":"2021-08-11T15:17:47.150771Z","iopub.execute_input":"2021-08-11T15:17:47.15117Z","iopub.status.idle":"2021-08-11T15:17:47.158171Z","shell.execute_reply.started":"2021-08-11T15:17:47.151136Z","shell.execute_reply":"2021-08-11T15:17:47.156411Z"}}},{"cell_type":"code","source":"import PIL\nfrom PIL import Image, ImageDraw\n\n\ndef display_images(images, title=None): \n    \"\"\"\n    func for display images \n    Thank you @rohitsingh9990 for this fucntion\n    \"\"\"\n    f, ax = plt.subplots(5,5, figsize=(18,22))\n    if title:\n        f.suptitle(title, fontsize = 30)\n\n    for i, image_id in enumerate(images):\n        image_path = os.path.join(TRAIN_IMAGE_DIR, f'{image_id[0]}/{image_id[1]}/{image_id[2]}/{image_id}.jpg')\n        image = Image.open(image_path)\n        \n        ax[i//5, i%5].imshow(image) \n        image.close()       \n        ax[i//5, i%5].axis('off')\n\n        landmark_id = train[train.id==image_id.split('.')[0]].landmark_id.values[0]\n        ax[i//5, i%5].set_title(f\"ID: {image_id.split('.')[0]}\\nLandmark_id: {landmark_id}\", fontsize=\"12\")\n\n    plt.show() ","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-01-17T11:35:48.392287Z","iopub.status.idle":"2022-01-17T11:35:48.392782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualizing","metadata":{}},{"cell_type":"code","source":"samples = train.sample(25).id.values\nprint(samples)\ndisplay_images(samples, 'Random')","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:48.393579Z","iopub.status.idle":"2022-01-17T11:35:48.394118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samples = train[train.landmark_id == 138982].sample(25).id.values\ndisplay_images(samples, 'Top 1')","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:48.394874Z","iopub.status.idle":"2022-01-17T11:35:48.395398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samples = train[train.landmark_id == 126637].sample(25).id.values\ndisplay_images(samples, 'Top 2')","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:48.396181Z","iopub.status.idle":"2022-01-17T11:35:48.396682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"samples = train[train.landmark_id == 20409].sample(25).id.values\ndisplay_images(samples, 'Top 3')","metadata":{"execution":{"iopub.status.busy":"2022-01-17T11:35:48.39744Z","iopub.status.idle":"2022-01-17T11:35:48.397948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}