{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h2><center>Hotel-ID to Combat Human Trafficking 2021 - FGVC8</center></h2>\n<h3><center>Recognizing hotels to aid Human trafficking investigations</center></h3>\n<center><img src = \"https://polarisproject.org/wp-content/uploads/2019/01/800x640-marriott-blog.jpg\" width = \"800\" height = \"640\"/></center>  ","metadata":{}},{"cell_type":"markdown","source":"## Contents\n<ul>\n    <li><h3>Motive</h3></li>\n    <li><h3>Import Libraries</h3></li>\n    <li><h3>Chains Vs Num Hotels</h3></li>\n    <li><h3>Chains Vs Num Images</h3></li>\n    <li><h3>Hotels Vs Num Images</h3></li>\n    <li><h3>Conclusion</h3></li>\n    <li><h3>Random Images</h3></li>\n</ul>","metadata":{}},{"cell_type":"markdown","source":"<h2><center>Motive</center></h2>","metadata":{}},{"cell_type":"markdown","source":"Human trafficking, a form of modern dat slavery, is a global problem affecting people of all ages. It is estimated that approximately 1,000,000 people are trafficked each year globally and that between 20,000 and 50,000 are trafficked into the United States, which is one of the largest destinations for victims of the sex-trafficking trade.\n\nVictims of human trafficking are kept in hotel rooms and sometimes photographed which can be used in later part of this crime. Identifying the hotels  from these images will play a vital role in lessening the astrounding noted case of human trafficking. Besides, it will also help the law enforcers to catch the criminals.\n\nIn this competition, we are tasked to identify hotels from 88 different chains by their images which can be later used for the abovementioned cause.","metadata":{}},{"cell_type":"markdown","source":"# Import Libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport cv2\nimport matplotlib.pyplot as plt\nimport os\nimport random\nimport sys\nfrom tqdm.autonotebook import tqdm\nimport seaborn as sns\nimport glob","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.read_csv(\"../input/hotel-id-2021-fgvc8/train.csv\")\ndf.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.info()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><h2> Chains VS Num Hotels </h2></center>","metadata":{}},{"cell_type":"markdown","source":"There are 88 different chains. Each chains have some hotels under them. In this section, we will look at the number of hotels under different chains.","metadata":{}},{"cell_type":"markdown","source":"<h3>Let's look at the chains having minimum and maximum different hotels.</h3>","metadata":{}},{"cell_type":"code","source":"chain_ids = []\nchain_values = []\nfor chain_id in df.chain.unique():\n    chain_ids.append(str(chain_id))\n    chain_values.append(len(df[df.chain == chain_id].hotel_id.unique()))\n    \nchain_ids = [x for _, x in sorted(zip(chain_values, chain_ids))]\nchain_values = sorted(chain_values)\n\n\nfigure, axes = plt.subplots(nrows=1, ncols=2, figsize=(15,8), squeeze=False)\n\n\nnames = chain_ids[:20]\nvalues = chain_values[:20]\n\nsns.barplot(x=names, y=values, ax = axes[0][0])\nplt.xticks(rotation=45)\n\n\nnames = chain_ids[-20:]\nvalues = chain_values[-20:]\n\nsns.barplot(x=names, y=values, ax = axes[0][1])\nplt.xticks(rotation=45)\n\naxes[0, 0].set_title(\"Min 20\")\naxes[0, 1].set_title(\"Max 20\")\naxes[0,0].tick_params(labelrotation=45)\naxes[0,1].tick_params(labelrotation=45)\nplt.setp(axes[-1, :], xlabel='Chain Id')\nplt.setp(axes[:, 0], ylabel='Hotel Count')\nplt.tight_layout()    \nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3>Let's look at the distribution.</h3>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 8))\n_ = plt.hist(chain_values, bins=30)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><h2> Chains VS Num Images </h2></center>","metadata":{}},{"cell_type":"markdown","source":"<h3>Let's look at the chains having minimum and maximum images.</h3>","metadata":{}},{"cell_type":"code","source":"chain_ids = []\nchain_values = []\nfor chain_id in df.chain.unique():\n    chain_ids.append(str(chain_id))\n    chain_values.append(len(df[df.chain == chain_id]))\n    \nchain_ids = [x for _, x in sorted(zip(chain_values, chain_ids))]\nchain_values = sorted(chain_values)\n\n\n\nfigure, axes = plt.subplots(nrows=1, ncols=2, figsize=(15, 8), squeeze=False)\n\n\nnames = chain_ids[:20]\nvalues = chain_values[:20]\n\n#plt.figure(figsize=(20, 10))\nsns.barplot(x=names, y=values, ax = axes[0][0])\nplt.xticks(rotation=45)\n\n\n\n\nnames = chain_ids[-20:]\nvalues = chain_values[-20:]\n\n#plt.figure(figsize=(20, 10))\nsns.barplot(x=names, y=values, ax = axes[0][1])\nplt.xticks(rotation=45)\n\n\naxes[0, 0].set_title(\"Min 20\")\naxes[0, 1].set_title(\"Max 20\")\naxes[0,0].tick_params(labelrotation=45)\naxes[0,1].tick_params(labelrotation=45)\nplt.setp(axes[-1, :], xlabel='Chain Id')\nplt.setp(axes[:, 0], ylabel='Image Count')\nplt.tight_layout()    \nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2>The distribution is</h2>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 8))\n_ = plt.hist(chain_values, bins=30)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><h2> Hotels VS Num Images </h2></center>","metadata":{}},{"cell_type":"markdown","source":"<h2> Minimum and Maximum Images</h2>","metadata":{}},{"cell_type":"code","source":"hotel_ids = []\nimage_values = []\nfor hotel_id in df.hotel_id.unique():\n    hotel_ids.append(str(hotel_id))\n    image_values.append(len(df[df.hotel_id == hotel_id]))\n    \nhotel_ids = [x for _, x in sorted(zip(image_values, hotel_ids))]\nimage_values = sorted(image_values)\n\n\n\nfigure, axes = plt.subplots(nrows=1, ncols=2, figsize=(15, 8), squeeze=False)\n\nnames = hotel_ids[:20]\nvalues = image_values[:20]\n\nsns.barplot(x=names, y=values, ax = axes[0][0])\n\n\nnames = hotel_ids[-20:]\nvalues = image_values[-20:]\n\nsns.barplot(x=names, y=values, ax = axes[0][1])\n\n\naxes[0, 0].set_title(\"Min 20\")\naxes[0, 1].set_title(\"Max 20\")\naxes[0,0].tick_params(labelrotation=45)\naxes[0,1].tick_params(labelrotation=45)\nplt.setp(axes[-1, :], xlabel='Hotel Id')\nplt.setp(axes[:, 0], ylabel='Image Count')\nplt.tight_layout()    \nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Distribution","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15, 8))\n_ = plt.hist(image_values, bins=30)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><h2>Conclusion</h2></center>","metadata":{}},{"cell_type":"markdown","source":"From the above EDA we can come to these conclusions.\n\n<ul>\n    <li>There are 88 different chains. Each chain manages various number of hotels. The number of managed hotels starts from 1 and reaches a maximum of 1750. But the majority of the chains have less than 100 hotels.</li> \n    <li>Number of samples taken from a chain can range from 10 to 20,000 at max. 60% of the chains have less than 500 samples.</li>\n    <li>There are 7770 different hotels listed in the dataset. In worst case scenario, there is one hotel having only one sample. Whereas, the maximun number of images per hotel can be around 90. Majority of the hotels have around 20 samples</li>\n    \n</ul>","metadata":{}},{"cell_type":"markdown","source":"## Some Random Images","metadata":{}},{"cell_type":"code","source":"files = glob.glob(\"../input/hotel-id-2021-fgvc8/train_images/*/*\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"figure, axes = plt.subplots(nrows=5, ncols=3, figsize=(20,15))\nfor i in range(15):\n    path = np.random.choice(files)\n    image = cv2.imread(path)\n    image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n    axes[i//3, i%3].imshow(image)","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}