{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns\nimport warnings\nwarnings.filterwarnings('ignore')\nimport geopy.distance\n! pip install reverse_geocode\nimport reverse_geocode\nimport tensorflow as tf","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:56:58.668706Z","iopub.execute_input":"2022-08-14T08:56:58.669082Z","iopub.status.idle":"2022-08-14T08:57:17.236277Z","shell.execute_reply.started":"2022-08-14T08:56:58.668986Z","shell.execute_reply":"2022-08-14T08:57:17.235363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Introduction**\n\nWhen you look for nearby restaurants or plan an errand in an unknown area, you expect relevant, accurate information. To maintain quality data worldwide is a challenge, and one with implications beyond navigation.  Foursquare is the #1 independent provider of global POI data. The leading independent location technology and data cloud platform, Foursquare is dedicated to building meaningful bridges between digital spaces and physical places. By efficiently and successfully matching POIs, we will make it easier to identify where new stores or businesses would benefit people the most.","metadata":{}},{"cell_type":"markdown","source":"# Datasets\n**train.csv** - The training set, comprising eleven attribute fields for over one million place entries, together with:\nid - A unique identifier for each entry.\npoint_of_interest - An identifier for the POI the entry represents. There may be one or many entries describing the same POI. Two entries \"match\" when they describe a common POI.\n\n**pairs.csv** - A pregenerated set of pairs of place entries from train.csv designed to improve detection of matches. You may wish to generate additional pairs to improve your model's ability to discriminate POIs.\nmatch - Whether (True or False) the pair of entries describes a common POI.\n\n**test.csv** - A set of place entries with their recorded attribute fields, similar to the training set.","metadata":{}},{"cell_type":"markdown","source":"# **DATASET**","metadata":{}},{"cell_type":"code","source":"df_train = pd.read_csv(\"/kaggle/input/foursquare-location-matching/train.csv\")\ndf_test = pd.read_csv(\"/kaggle/input/foursquare-location-matching/test.csv\")\ndf_pairs = pd.read_csv(\"/kaggle/input/foursquare-location-matching/pairs.csv\")","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-14T08:57:17.238120Z","iopub.execute_input":"2022-08-14T08:57:17.238423Z","iopub.status.idle":"2022-08-14T08:57:33.830773Z","shell.execute_reply.started":"2022-08-14T08:57:17.238390Z","shell.execute_reply":"2022-08-14T08:57:33.829965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:57:33.831796Z","iopub.execute_input":"2022-08-14T08:57:33.832027Z","iopub.status.idle":"2022-08-14T08:57:33.859760Z","shell.execute_reply.started":"2022-08-14T08:57:33.832000Z","shell.execute_reply":"2022-08-14T08:57:33.858947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:57:33.861603Z","iopub.execute_input":"2022-08-14T08:57:33.861876Z","iopub.status.idle":"2022-08-14T08:57:33.867873Z","shell.execute_reply.started":"2022-08-14T08:57:33.861848Z","shell.execute_reply":"2022-08-14T08:57:33.867027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Observation of train dataset\ndf_train.iloc[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:57:33.869347Z","iopub.execute_input":"2022-08-14T08:57:33.869828Z","iopub.status.idle":"2022-08-14T08:57:33.883004Z","shell.execute_reply.started":"2022-08-14T08:57:33.869793Z","shell.execute_reply":"2022-08-14T08:57:33.882177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pairs.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:57:33.884444Z","iopub.execute_input":"2022-08-14T08:57:33.884695Z","iopub.status.idle":"2022-08-14T08:57:33.918159Z","shell.execute_reply.started":"2022-08-14T08:57:33.884668Z","shell.execute_reply":"2022-08-14T08:57:33.917277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pairs.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:57:33.919576Z","iopub.execute_input":"2022-08-14T08:57:33.920444Z","iopub.status.idle":"2022-08-14T08:57:33.926559Z","shell.execute_reply.started":"2022-08-14T08:57:33.920404Z","shell.execute_reply":"2022-08-14T08:57:33.925738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Observation of pairs dataset\ndf_pairs.iloc[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:57:33.927481Z","iopub.execute_input":"2022-08-14T08:57:33.928005Z","iopub.status.idle":"2022-08-14T08:57:33.939504Z","shell.execute_reply.started":"2022-08-14T08:57:33.927961Z","shell.execute_reply":"2022-08-14T08:57:33.938756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Observation of test dataset\ndf_test.iloc[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:57:33.940666Z","iopub.execute_input":"2022-08-14T08:57:33.941110Z","iopub.status.idle":"2022-08-14T08:57:33.950584Z","shell.execute_reply.started":"2022-08-14T08:57:33.941048Z","shell.execute_reply":"2022-08-14T08:57:33.950039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We first verified which country has the highest location\ncountry=df_train['country'].value_counts()*100/df_train['country'].value_counts().sum()\ncountry=country.head(10)\n\nplt.figure(figsize=(8,7))\ncolor=[\"yellow\"]*len(country.index)\ncolor[0]=\"darkblue\"\nsns.barplot(x=country.index, y=country.values,palette=color, saturation=.7)\nplt.title(\"% Country Data\")\nplt.xlabel('country')\n_=plt.ylabel('Percentage')","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:57:33.953100Z","iopub.execute_input":"2022-08-14T08:57:33.953422Z","iopub.status.idle":"2022-08-14T08:57:34.402629Z","shell.execute_reply.started":"2022-08-14T08:57:33.953397Z","shell.execute_reply":"2022-08-14T08:57:34.401729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We then verified which state in US has the maximum data\nstate=df_train[df_train['country']=='US']['state'].value_counts()*100/df_train[df_train['country']=='US']['state'].value_counts().sum()\nstate=state.head(10)\n\nplt.figure(figsize=(8,7))\ncolor=[\"yellow\"]*len(state.index)\ncolor[0]=\"darkblue\"\nsns.barplot(x=state.index, y=state.values,palette=color, saturation=.7)\nplt.title(\"% Data by State\")\nplt.xlabel('State')\n_=plt.ylabel('Percentage')","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:57:34.404011Z","iopub.execute_input":"2022-08-14T08:57:34.404306Z","iopub.status.idle":"2022-08-14T08:57:35.052737Z","shell.execute_reply.started":"2022-08-14T08:57:34.404271Z","shell.execute_reply":"2022-08-14T08:57:35.051999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"state=df_train[df_train['country']=='US']['state'].value_counts()*100/df_train[df_train['country']=='US']['state'].value_counts().sum()\nstate=state.reset_index().rename(columns={'index':'state','state':'% data'})\nstate.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:57:35.053996Z","iopub.execute_input":"2022-08-14T08:57:35.054301Z","iopub.status.idle":"2022-08-14T08:57:35.421340Z","shell.execute_reply.started":"2022-08-14T08:57:35.054258Z","shell.execute_reply":"2022-08-14T08:57:35.420285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"category=df_train['categories'].value_counts()*100/df_train['categories'].value_counts().sum()\ncategory=df_train['categories'].value_counts()*100/df_train['categories'].value_counts().sum()\ncategory=category.head(10)\n\nplt.figure(figsize=(8,7))\ncolor=[\"yellow\"]*len(category.index)\ncolor[0]=\"darkblue\"\nsns.barplot(x=category.index, y=category.values,palette=color, saturation=.7)\nplt.xticks(rotation=90)\nplt.title(\"% Data by Category\")\nplt.xlabel('category')\n_=plt.ylabel('Percentage')","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:57:35.422658Z","iopub.execute_input":"2022-08-14T08:57:35.422979Z","iopub.status.idle":"2022-08-14T08:57:36.390151Z","shell.execute_reply.started":"2022-08-14T08:57:35.422923Z","shell.execute_reply":"2022-08-14T08:57:36.389177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data-PreProcessing","metadata":{}},{"cell_type":"code","source":"#Converting all the strings values to lower case\ndef fun_lower(df):\n  for c in df.columns:\n    df[c] = df[c].fillna('').astype(str).apply(lambda x: x.lower())\n\nfun_lower(df_train)\nfun_lower(df_pairs)\nfun_lower(df_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:57:36.391659Z","iopub.execute_input":"2022-08-14T08:57:36.391881Z","iopub.status.idle":"2022-08-14T08:57:51.939412Z","shell.execute_reply.started":"2022-08-14T08:57:36.391856Z","shell.execute_reply":"2022-08-14T08:57:51.938541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def missing(df):\n  for c in df.columns:\n    df[c] = df[c].replace(\"\", np.nan)\n\nmissing(df_train)\nmissing(df_test)\nmissing(df_pairs)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:57:51.940719Z","iopub.execute_input":"2022-08-14T08:57:51.941281Z","iopub.status.idle":"2022-08-14T08:57:53.808592Z","shell.execute_reply.started":"2022-08-14T08:57:51.941240Z","shell.execute_reply":"2022-08-14T08:57:53.807870Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Converting all object columns to category\ndef obj_to_cat(df):\n  for i in df.columns:\n    if df[i].dtype == \"O\":\n      df[i] = df[i].astype('category')\nobj_to_cat(df_train)\nobj_to_cat(df_pairs)\nobj_to_cat(df_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:57:53.809815Z","iopub.execute_input":"2022-08-14T08:57:53.810147Z","iopub.status.idle":"2022-08-14T08:58:27.080621Z","shell.execute_reply.started":"2022-08-14T08:57:53.810111Z","shell.execute_reply":"2022-08-14T08:58:27.079744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test[\"zip\"] = df_test[\"zip\"].astype('category')\ndf_test[\"phone\"] = df_test[\"phone\"].astype('category')","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:27.082070Z","iopub.execute_input":"2022-08-14T08:58:27.082433Z","iopub.status.idle":"2022-08-14T08:58:27.088814Z","shell.execute_reply.started":"2022-08-14T08:58:27.082393Z","shell.execute_reply":"2022-08-14T08:58:27.087864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(df_test.isna().sum()/len(df_test))*100","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:27.090054Z","iopub.execute_input":"2022-08-14T08:58:27.090450Z","iopub.status.idle":"2022-08-14T08:58:27.108312Z","shell.execute_reply.started":"2022-08-14T08:58:27.090420Z","shell.execute_reply":"2022-08-14T08:58:27.107250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(df_train.isnull().sum()/len(df_train))*100","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:27.109689Z","iopub.execute_input":"2022-08-14T08:58:27.110470Z","iopub.status.idle":"2022-08-14T08:58:27.163319Z","shell.execute_reply.started":"2022-08-14T08:58:27.110420Z","shell.execute_reply":"2022-08-14T08:58:27.162443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(df_pairs.isnull().sum()/len(df_pairs))*100","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:27.164506Z","iopub.execute_input":"2022-08-14T08:58:27.164724Z","iopub.status.idle":"2022-08-14T08:58:27.200799Z","shell.execute_reply.started":"2022-08-14T08:58:27.164700Z","shell.execute_reply":"2022-08-14T08:58:27.200007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_country_codes(coords):\n    data = reverse_geocode.search(coords)\n    return [v['country_code'] for v in data]\n\ndf_train['country_code'] = get_country_codes(df_train[['latitude', 'longitude']])\n\ndf_pairs['country_code_1'] = get_country_codes(df_pairs[['latitude_1', 'longitude_1']])\ndf_pairs['country_code_2'] = get_country_codes(df_pairs[['latitude_2', 'longitude_2']])\n\nprint(f'Unique Country Code 1 in Pairs: {df_pairs[\"country_code_1\"].nunique()}\\n')\nprint('===== Top 10 Most Occuring Country Codes 1 =====')\ndisplay(df_pairs['country_code_1'].value_counts(dropna=False).head(10))","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:27.201884Z","iopub.execute_input":"2022-08-14T08:58:27.202207Z","iopub.status.idle":"2022-08-14T08:58:36.823041Z","shell.execute_reply.started":"2022-08-14T08:58:27.202180Z","shell.execute_reply":"2022-08-14T08:58:36.822171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_cities(coords):\n    data = reverse_geocode.search(coords)\n    return [v['city'] for v in data]\n\ndf_train['city_rg'] = get_cities(df_train[['latitude', 'longitude']])\n    \ndf_pairs['city_rg_1'] = get_cities(df_pairs[['latitude_1', 'longitude_1']])\ndf_pairs['city_rg_2'] = get_cities(df_pairs[['latitude_2', 'longitude_2']])\n\nprint(f'Unique City Reverse Geocode 1 in Pairs: {df_pairs[\"city_rg_1\"].nunique()}\\n')\nprint('===== Top 10 Most Occuring City Reverse Geocode 1 =====')\ndisplay(df_pairs['city_rg_1'].value_counts(dropna=False).head(10))","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:36.824190Z","iopub.execute_input":"2022-08-14T08:58:36.824485Z","iopub.status.idle":"2022-08-14T08:58:43.591810Z","shell.execute_reply.started":"2022-08-14T08:58:36.824448Z","shell.execute_reply":"2022-08-14T08:58:43.590927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_train[\"country_code\"].isna().sum())\nprint(df_train[\"city_rg\"].isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:43.593133Z","iopub.execute_input":"2022-08-14T08:58:43.593453Z","iopub.status.idle":"2022-08-14T08:58:43.747666Z","shell.execute_reply.started":"2022-08-14T08:58:43.593417Z","shell.execute_reply":"2022-08-14T08:58:43.746695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['country_code'] = df_train['country_code'].str.lower()\ndf_train['city_rg'] = df_train['city_rg'].str.lower()\ndf_pairs['country_code_1'] = df_pairs['country_code_1'].str.lower()\ndf_pairs['country_code_2'] = df_pairs['country_code_2'].str.lower()\ndf_pairs['city_rg_1'] = df_pairs['city_rg_1'].str.lower()\ndf_pairs['city_rg_2'] = df_pairs['city_rg_2'].str.lower()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:43.748860Z","iopub.execute_input":"2022-08-14T08:58:43.749122Z","iopub.status.idle":"2022-08-14T08:58:45.700994Z","shell.execute_reply.started":"2022-08-14T08:58:43.749094Z","shell.execute_reply":"2022-08-14T08:58:45.700340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"obj_to_cat(df_train)\nobj_to_cat(df_pairs)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:45.702097Z","iopub.execute_input":"2022-08-14T08:58:45.702457Z","iopub.status.idle":"2022-08-14T08:58:46.540655Z","shell.execute_reply.started":"2022-08-14T08:58:45.702431Z","shell.execute_reply":"2022-08-14T08:58:46.540057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Most common names are fast food restaurants\nprint(f'Unique Names in Train: {df_train[\"name\"].nunique()}\\n')\nprint('===== Top 10 Most Occuring Names =====')\ndisplay(df_train['name'].value_counts(dropna=False, normalize=True).head(10))","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:46.541775Z","iopub.execute_input":"2022-08-14T08:58:46.542149Z","iopub.status.idle":"2022-08-14T08:58:46.849017Z","shell.execute_reply.started":"2022-08-14T08:58:46.542108Z","shell.execute_reply":"2022-08-14T08:58:46.848151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# VALIDATION","metadata":{}},{"cell_type":"code","source":"#Splitting the pairs data\nfrom sklearn.metrics import accuracy_score, classification_report,confusion_matrix\nfrom sklearn.model_selection import train_test_split\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense\nimport tensorflow as tf","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:46.850536Z","iopub.execute_input":"2022-08-14T08:58:46.850935Z","iopub.status.idle":"2022-08-14T08:58:47.019788Z","shell.execute_reply.started":"2022-08-14T08:58:46.850891Z","shell.execute_reply":"2022-08-14T08:58:47.018672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pairs.columns","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:47.023696Z","iopub.execute_input":"2022-08-14T08:58:47.023981Z","iopub.status.idle":"2022-08-14T08:58:47.030518Z","shell.execute_reply.started":"2022-08-14T08:58:47.023923Z","shell.execute_reply":"2022-08-14T08:58:47.029554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pairs.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:47.031754Z","iopub.execute_input":"2022-08-14T08:58:47.032084Z","iopub.status.idle":"2022-08-14T08:58:47.063663Z","shell.execute_reply.started":"2022-08-14T08:58:47.032046Z","shell.execute_reply":"2022-08-14T08:58:47.062632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pairs.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:47.066267Z","iopub.execute_input":"2022-08-14T08:58:47.066861Z","iopub.status.idle":"2022-08-14T08:58:48.986093Z","shell.execute_reply.started":"2022-08-14T08:58:47.066821Z","shell.execute_reply":"2022-08-14T08:58:48.985321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pairs['match']= df_pairs['match'].cat.codes\ndf_pairs.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:48.987312Z","iopub.execute_input":"2022-08-14T08:58:48.987547Z","iopub.status.idle":"2022-08-14T08:58:49.017524Z","shell.execute_reply.started":"2022-08-14T08:58:48.987520Z","shell.execute_reply":"2022-08-14T08:58:49.016579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pairs['categories_1']= df_pairs['categories_1'].cat.codes\ndf_pairs['categories_2']= df_pairs['categories_2'].cat.codes","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:49.019010Z","iopub.execute_input":"2022-08-14T08:58:49.019295Z","iopub.status.idle":"2022-08-14T08:58:49.031313Z","shell.execute_reply.started":"2022-08-14T08:58:49.019233Z","shell.execute_reply":"2022-08-14T08:58:49.030396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = df_pairs[['latitude_1', 'longitude_1','latitude_2', 'longitude_2','categories_1','categories_2','match']]","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:49.032462Z","iopub.execute_input":"2022-08-14T08:58:49.033100Z","iopub.status.idle":"2022-08-14T08:58:49.049160Z","shell.execute_reply.started":"2022-08-14T08:58:49.033069Z","shell.execute_reply":"2022-08-14T08:58:49.048491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X['latitude_1']= X['latitude_1'].astype(np.float)\nX['longitude_1']= X['longitude_1'].astype(np.float)\nX['latitude_2']= X['latitude_2'].astype(np.float)\nX['longitude_2']= X['longitude_2'].astype(np.float)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:49.050212Z","iopub.execute_input":"2022-08-14T08:58:49.050673Z","iopub.status.idle":"2022-08-14T08:58:50.003650Z","shell.execute_reply.started":"2022-08-14T08:58:49.050630Z","shell.execute_reply":"2022-08-14T08:58:50.002722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train, test = train_test_split(X, test_size=0.2)\ntrain, val = train_test_split(train, test_size=0.2)\nprint(len(train), 'train examples')\nprint(len(val), 'validation examples')\nprint(len(test), 'test examples')","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:50.005182Z","iopub.execute_input":"2022-08-14T08:58:50.005807Z","iopub.status.idle":"2022-08-14T08:58:50.124523Z","shell.execute_reply.started":"2022-08-14T08:58:50.005769Z","shell.execute_reply":"2022-08-14T08:58:50.123559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# A utility method to create a tf.data dataset from a Pandas Dataframe\ndef df_to_dataset(dataframe, shuffle=True, batch_size=32):\n    dataframe = dataframe.copy()\n    labels = dataframe.pop('match')\n    ds = tf.data.Dataset.from_tensor_slices((dict(dataframe), labels))\n    if shuffle:\n        ds = ds.shuffle(buffer_size=len(dataframe))\n    ds = ds.batch(batch_size)\n    return ds","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:50.125770Z","iopub.execute_input":"2022-08-14T08:58:50.126082Z","iopub.status.idle":"2022-08-14T08:58:50.132407Z","shell.execute_reply.started":"2022-08-14T08:58:50.126044Z","shell.execute_reply":"2022-08-14T08:58:50.131760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"batch_size = 32 # A small batch sized is used for demonstration purposes\ntrain_ds = df_to_dataset(train, batch_size=batch_size)\nval_ds = df_to_dataset(val, shuffle=False, batch_size=batch_size)\ntest_ds = df_to_dataset(test, shuffle=False, batch_size=batch_size)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:50.133463Z","iopub.execute_input":"2022-08-14T08:58:50.133708Z","iopub.status.idle":"2022-08-14T08:58:50.216929Z","shell.execute_reply.started":"2022-08-14T08:58:50.133681Z","shell.execute_reply":"2022-08-14T08:58:50.216306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We will use this batch to demonstrate several types of feature columns\nexample_batch = next(iter(train_ds))[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:50.218071Z","iopub.execute_input":"2022-08-14T08:58:50.218438Z","iopub.status.idle":"2022-08-14T08:58:51.995396Z","shell.execute_reply.started":"2022-08-14T08:58:50.218411Z","shell.execute_reply":"2022-08-14T08:58:51.994453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for feature_batch, label_batch in train_ds.take(1):\n    print('Every feature:', list(feature_batch.keys()))\n    print('A batch of lat_1:', feature_batch['latitude_1'])\n    print('A batch of lat_2:', feature_batch['latitude_2'])\n    print('A batch of long_1:', feature_batch['longitude_1'])\n    print('A batch of long_2:', feature_batch['longitude_2'])\n    print('A batch of cat_1:', feature_batch['categories_1'])\n    print('A batch of cat_2:', feature_batch['categories_2'])\n    print('A batch of targets:', label_batch )","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:51.996728Z","iopub.execute_input":"2022-08-14T08:58:51.997050Z","iopub.status.idle":"2022-08-14T08:58:53.841963Z","shell.execute_reply.started":"2022-08-14T08:58:51.997014Z","shell.execute_reply":"2022-08-14T08:58:53.841017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# A utility method to create a feature column\n# and to transform a batch of data\nfrom tensorflow import feature_column\nfrom tensorflow.keras import layers\ndef demo(feature_column):\n    feature_layer = layers.DenseFeatures(feature_column)\n    print(feature_layer(example_batch).numpy())","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:53.843047Z","iopub.execute_input":"2022-08-14T08:58:53.843282Z","iopub.status.idle":"2022-08-14T08:58:53.848146Z","shell.execute_reply.started":"2022-08-14T08:58:53.843255Z","shell.execute_reply":"2022-08-14T08:58:53.847352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lat_1_count = feature_column.numeric_column('latitude_1')\nlat_2_count = feature_column.numeric_column('latitude_2')\nlong_1_count = feature_column.numeric_column('longitude_1')\nlong_2_count = feature_column.numeric_column('longitude_2')\ncat_1_count = feature_column.numeric_column('categories_1')\ncat_2_count = feature_column.numeric_column('categories_2')\ndemo(lat_1_count)\ndemo(lat_2_count)\ndemo(long_1_count)\ndemo(long_2_count)\ndemo(cat_1_count)\ndemo(cat_2_count)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:53.849283Z","iopub.execute_input":"2022-08-14T08:58:53.849481Z","iopub.status.idle":"2022-08-14T08:58:53.921771Z","shell.execute_reply.started":"2022-08-14T08:58:53.849458Z","shell.execute_reply":"2022-08-14T08:58:53.920708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feature_columns = []\n\n# numeric cols\nfor header in ['latitude_1', 'longitude_1','latitude_2', 'longitude_2','categories_1','categories_2']:\n    feature_columns.append(feature_column.numeric_column(header))","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:53.923291Z","iopub.execute_input":"2022-08-14T08:58:53.924034Z","iopub.status.idle":"2022-08-14T08:58:53.929685Z","shell.execute_reply.started":"2022-08-14T08:58:53.923993Z","shell.execute_reply":"2022-08-14T08:58:53.928675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feature_layer = tf.keras.layers.DenseFeatures(feature_columns)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:53.931109Z","iopub.execute_input":"2022-08-14T08:58:53.931465Z","iopub.status.idle":"2022-08-14T08:58:53.941780Z","shell.execute_reply.started":"2022-08-14T08:58:53.931427Z","shell.execute_reply":"2022-08-14T08:58:53.941037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = tf.keras.Sequential([\n  feature_layer,\n  layers.Dense(128, activation='relu'),\n  layers.Dense(100, activation='relu'),\n  layers.Dense(1, activation='sigmoid')\n])","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:53.944541Z","iopub.execute_input":"2022-08-14T08:58:53.945315Z","iopub.status.idle":"2022-08-14T08:58:53.966103Z","shell.execute_reply.started":"2022-08-14T08:58:53.945279Z","shell.execute_reply":"2022-08-14T08:58:53.965270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.compile(optimizer='adam',\n              loss=tf.keras.losses.BinaryCrossentropy(from_logits=True),\n              metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:53.967445Z","iopub.execute_input":"2022-08-14T08:58:53.967880Z","iopub.status.idle":"2022-08-14T08:58:53.987519Z","shell.execute_reply.started":"2022-08-14T08:58:53.967850Z","shell.execute_reply":"2022-08-14T08:58:53.986821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"early_stopping = tf.keras.callbacks.EarlyStopping(\n    monitor=\"val_loss\",\n    min_delta=0.1,\n    patience=10,\n    mode=\"auto\"\n)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:53.988454Z","iopub.execute_input":"2022-08-14T08:58:53.989134Z","iopub.status.idle":"2022-08-14T08:58:53.992898Z","shell.execute_reply.started":"2022-08-14T08:58:53.989102Z","shell.execute_reply":"2022-08-14T08:58:53.992357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nhistory = model.fit(train_ds,\n          validation_data=val_ds,\n          epochs=100,callbacks = early_stopping)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T08:58:53.993857Z","iopub.execute_input":"2022-08-14T08:58:53.994203Z","iopub.status.idle":"2022-08-14T09:06:23.852344Z","shell.execute_reply.started":"2022-08-14T08:58:53.994178Z","shell.execute_reply":"2022-08-14T09:06:23.851474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scores = model.evaluate(test_ds)\nprint(scores)  ","metadata":{"execution":{"iopub.status.busy":"2022-08-14T09:06:23.853429Z","iopub.execute_input":"2022-08-14T09:06:23.853880Z","iopub.status.idle":"2022-08-14T09:06:28.257594Z","shell.execute_reply.started":"2022-08-14T09:06:23.853850Z","shell.execute_reply":"2022-08-14T09:06:28.256665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"epochs = range(1, len(history.history['accuracy'])+1)\n\nplt.plot(epochs, history.history['accuracy'], '#21466C', label='Train accuracy')\nplt.plot(epochs, history.history['val_accuracy'], '#cc1123', label='Validation accuracy')\nplt.title('Training and validation accuracy')\nplt.legend()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T09:06:28.258577Z","iopub.execute_input":"2022-08-14T09:06:28.259120Z","iopub.status.idle":"2022-08-14T09:06:28.528735Z","shell.execute_reply.started":"2022-08-14T09:06:28.259091Z","shell.execute_reply":"2022-08-14T09:06:28.527926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **CONCLUSION**","metadata":{}},{"cell_type":"markdown","source":"Model was clearly the best performing  with a test accuracy of 70.49% on the 115782 test dataset using the epochs 100 where training dataset shows the accuracy of 70.26%.\n\n\n* Train Accuracy: 70.26%\n* Validation Accuracy: 70.55%\n* Test Accuracy: 70.49%\n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}