{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels = pd.read_csv(\"../input/hotel-id-2021-fgvc8/train.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels.dtypes","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#fix datatype for timestamp\nlabels.timestamp = pd.to_datetime(labels.timestamp)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#number of hotels per chain\ncounts = labels.groupby('chain').hotel_id.count()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\"The largest chain has {0} hotels and the smallest has {1}\".format(counts.max(), counts.min())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"counts.hist();\nplt.title(\"Number of Hotels per Chain\");","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The number of hotels per chain has a very noticable positive skew and a very high outlier."},{"metadata":{"trusted":true},"cell_type":"code","source":"counts_sorted = sorted(counts, reverse=True)\nprint(\"The largest outlier is {:.2f} times bigger than the next largest chain (where size is determined by the chain's hotel count).\".format(counts_sorted[0]/counts_sorted[1]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.hist(counts_sorted[1:])\nplt.title(\"Number of Hotels per Chain (no outlier)\");","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Even dropping the most problematic outlier, there is still a very noticable positive skew. As chains with exceptionally high numbers of hotels may have more variation between hotels, this could make classifying these chains harder. On the other hand, owning such a large number of hotels may force these chains to be more systematic and thus uniform with their hotel interior design. This could make classifying them easier."},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"The timestamp range of the labels is:\\nfrom {0}\\n\\nto {1}\".format(labels.timestamp.min(), labels.timestamp.max()))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels.timestamp.dt.year.hist(bins=4)\nplt.title(\"Timestamp Years of Hotel Photos\");","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#number of years between earliest photo taken and latest photo taken for each chain\nranges= labels.groupby(\"chain\").timestamp.apply(lambda x:x.dt.year.max() - x.dt.year.min())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ranges.hist(bins=3)\nplt.title(\"Number of Years Between Photos (for each Chain)\");","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The photos of hotels for each chain were all taken within a relatively short range of time (5 years). It is likely that the timestamp won't be a hugely determining feature. However, if the chain underwent renovations or changes within those 5 years, the timestamp could be an important factor."},{"metadata":{"trusted":true},"cell_type":"code","source":"import os","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#preview train_images directory (`train_dir`). The subdirectories appear to all be integers.\ntrain_dir = \"../input/hotel-id-2021-fgvc8/train_images\"\nos.listdir(train_dir)[0:5]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#The chain column in the `labels` dataframe links to the subdirectories of `train_dir`. The subdirectories of `train_dir` are thus hotel chain ids. \nset(labels.chain.apply(str)) == set(os.listdir(train_dir))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#preview train_images/{hotel chain id}/ directory. They appear to be all jpg files\nos.listdir(os.path.join(\"../input/hotel-id-2021-fgvc8/train_images\", \"7\"))[0:5]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"example_file = \"../input/hotel-id-2021-fgvc8/train_images/32/809e6fd11c8d555e.jpg\"\nexample_img = plt.imread(example_file)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"The example image file is {0}px by {1}px\".format(*example_img.shape))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for chain_id in range(4):\n    fig = plt.figure()\n    fig.suptitle(\"Chain Id {} Sample Images\".format(chain_id))\n    \n    chain_dir = os.path.join(train_dir, str(chain_id))\n    files = os.listdir(chain_dir)[:5]\n    for i in range(1,5):\n        plt.subplot(2,2,i)\n        img = plt.imread(os.path.join(chain_dir, files[i]))\n        plt.imshow(img)\n    ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"From taking a quick glance at the images, we can conclude that the images are all different dimensions and thus will need to be standardized (and likely reduced in pixel size for faster training). In addition, the images from within the same hotel chain appear to vary from each other, in some respects, more than they do from images beloning to different classes. For example, within the same hotel chain, there can be bathrooms, bedrooms, sitting areas, etc. Thus, looking for image features that identify objects will likely be unhelpful. Rather, the style of the image may be a more appropriate feature for classification."},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}