{"cells":[{"metadata":{},"cell_type":"markdown","source":"# OverView"},{"metadata":{},"cell_type":"markdown","source":"Let's go over [the rules](https://www.kaggle.com/c/landmark-recognition-2020/overview/code-requirements) first.\n* TPUs will not be available for making submissions to this competition. You are still welcome to use them for training models.\n* No internet access enabled\n* Freely & publicly available external data is allowed, including pre-trained models"},{"metadata":{"trusted":true},"cell_type":"code","source":"# linear algebra\nimport numpy as np\n# data processing, CSV file I/O (e.g. pd.read_csv)\nimport pandas as pd\n#Unix commands\nimport os\n\n# import useful tools\nfrom glob import glob\nfrom PIL import Image\nimport cv2\n\n# import data visualization\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nimport seaborn as sns\n\nfrom bokeh.plotting import figure\nfrom bokeh.io import output_notebook, show, output_file\nfrom bokeh.models import ColumnDataSource, HoverTool, Panel\nfrom bokeh.models.widgets import Tabs\n# import data augmentation\nimport albumentations as albu\n\n# import math module\nimport math\n#Libraries\nimport pandas_profiling\nfrom xgboost import XGBClassifier\nfrom sklearn import preprocessing","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# linear algebra\nimport numpy as np\n# data processing, CSV file I/O (e.g. pd.read_csv)\nimport pandas as pd\n#Unix commands\nimport os\n\n# import useful tools\nfrom glob import glob\nfrom PIL import Image\nimport cv2\n\n# import data visualization\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nimport seaborn as sns\n\nfrom bokeh.plotting import figure\nfrom bokeh.io import output_notebook, show, output_file\nfrom bokeh.models import ColumnDataSource, HoverTool, Panel\nfrom bokeh.models.widgets import Tabs\n# import data augmentation\nimport albumentations as albu\n\n# import math module\nimport math\n#Libraries\nimport pandas_profiling\nfrom xgboost import XGBClassifier\nfrom sklearn import preprocessing","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Loading data"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"# Setup the paths to train and test images\nDATASET = '../input/delg-saved-models/'\nTEST_DIR = '../input/landmark-recognition-2020/test/'\nTRAIN_DIR = '../input/landmark-recognition-2020/'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Loading train Files for Submission\ntrain = pd.read_csv(TRAIN_DIR + \"train.csv\")\n#Loading Sample Files for Submission\nsample = pd.read_csv(TRAIN_DIR + \"sample_submission.csv\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Let us check the format of the file for submission"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display some of the training data\ntrain.head(10).style.applymap(lambda x: 'background-color:lightsteelblue')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There seems to be unique IDs and some landmark IDs"},{"metadata":{},"cell_type":"markdown","source":"Next, we will check for missing values"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Confirmation of the format of samples for submission\nsample.head(10).style.applymap(lambda x: 'background-color:lightsteelblue')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Check for missing values in the training data\ntrain.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Fortunately, the missing value is content!\n* Let us check how many types of Landmark IDs there are"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Find the unique number of landmark IDs. \nn = train['landmark_id'].nunique()\nprint('The unique number of landmark IDs is ' + str(n))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* We have made it possible to notice that there are so many different types of Landmark IDs"},{"metadata":{},"cell_type":"markdown","source":"* We will next show how the more than 80,000 landmark IDs are distributed in a scatterplot"},{"metadata":{"trusted":true},"cell_type":"code","source":"# First, I'll use Sturgess's formula to find the appropriate number of classes in the histogram \nk = 1 + math.log2(n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display a histogram of the FVC of the training data\nsns.distplot(train['landmark_id'], kde=True, rug=False, bins=int(k), color='c') \n# Graph Title\nplt.title('Distribuition of landmark_ids')\n# label\nplt.xlabel(\"landmark_ids\")\nplt.ylabel(\"Frequency\")\n# Show Histogram\nplt.show() ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* There doesn't seem to be much of a bias."},{"metadata":{},"cell_type":"markdown","source":"* We'll find the landmark IDs that appear most often"},{"metadata":{"trusted":true},"cell_type":"code","source":"# coding: utf-8\nfrom tqdm import tqdm\nimport time\n\n# Set the total value \nbar = tqdm(total = 1000)\n# Add description\nbar.set_description('Progress rate')\nfor i in range(100):\n    # Set the progress\n    bar.update(25)\n    time.sleep(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(train['landmark_id'].value_counts())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Check the number of landmark ID 138982"},{"metadata":{"trusted":true},"cell_type":"code","source":"s_bool = train['landmark_id'] == 138982\nm = s_bool.sum()\nprint('The number of landmark ID 138982' + str(m))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Landmark ID 138982 looks like there are 6272 of them.\n* This ID seems to be the most common, so let us focus our attention on this ID."},{"metadata":{"trusted":true},"cell_type":"code","source":"the_most_cmn_pics = train[train[\"landmark_id\"]==138982]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(the_most_cmn_pics)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Acknowledgements"},{"metadata":{},"cell_type":"markdown","source":"* [Visualizing Landmarks (+more EDA)](https://www.kaggle.com/jeffreybraun/visualizing-landmarks-more-eda)"},{"metadata":{},"cell_type":"markdown","source":"# Your upvote is the source of my motivation."},{"metadata":{},"cell_type":"markdown","source":"# To be continued, sir."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}