{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Problem Statement\n\n* I am working on the Kaggle Competition [FourSquare Location Matching](https://www.kaggle.com/competitions/foursquare-location-matching). This Challenge involved a dataset of more than one-and-a-half Million Places of interest, and  task is to build a algorithm to predict which point of interests represents the same point-of-interest.\n* Each point of interest includes attribute like the name, street information, coordinate information, and category related information.\n\n\n# Buisness Problem\n\n\n\n* Businesses make decisions on new sites for market expansion, analyze the competitive landscape, and show relevant ads informed by location data. \n* Finding similar and duplicated data from million of palaces of interest is very challenging because the dataset is collected from diverse sources, and may contain duplicate and incomplete information, raw data can contain noise, unstructured information, and incomplete or inaccurate attributes. \n* By efficiently and successfully matching POIs, business will make it easier to identify where new stores or businesses would benefit people the most.\n* A combination of machine-learning algorithms and rigorous human validation methods is optimal to find and collect accurate records.\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# Dataset Information\n\n\n* Foursquare Provided over 1 million place entries for thousands of commercial Point-of-Interest (POIs) around the globe. \n* Data entries may represent or resemble entried for real places, they are also added with extra noise and artificial information for obvious privacy reasons.\n\n# Training Data\n\n\n**1. train.csv** - The training set, comprising eleven attribute fields for over one million place entries, together with\n\n**id** - A unique identifier for each entry.\n\n**point_of_interest** - An identifier for the POI the entry represents. There may be one or many entries describing the same POI. Two entries \"match\" when they describe a common POI.\n\n**Name** -  Name of point of interest\n\n**Latitude** - Geo Latitude location\n\n**Longitude** - Geo longitude Location\n\n**Address** - Address string\n\n**State** - State  \n\n**Zip** - Zip code of address\n\n**City** - city of poi\n\n**Country** - \n\n**Url** - website url if any\n\n**Phone** - \n\n**Categories** - which categories this poi represents\n\n**Point of interest** **POI** : What POI this location represents\n\n**pairs.csv** - A pregenerated set of pairs of place entries from train.csv designed to improve detection of matches. You may wish to generate additional pairs to improve your model's ability to discriminate POIs.\n\n**match** - Whether (True or False) the pair of entries describes a common POI..\nOther columns are all the columns ( 12 for each POI so total 24 )\n","metadata":{}},{"cell_type":"markdown","source":"# Evaluation \n\n* Submissions are evaluated by the mean [Intersection over Union \\(IOU\\) ](https://en.wikipedia.org/wiki/Jaccard_index) (aka the Jaccard index) of the ground-truth entry matches and the predicted entry matches. The mean is taken sample-wise, meaning that an IoU score is calculated for each row in the submission file, and the final score is their average\n\n$$ Jaccard(U,V) = \\frac{|groundtruth\\_matches \\cap predicted\\_matches|}{|groundtruth\\_matches \\cup predicted\\_matches|}$$\n\n","metadata":{}},{"cell_type":"markdown","source":"# Load Library","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport seaborn as sns\nfrom collections import Counter\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom collections import Counter, defaultdict\nfrom wordcloud import WordCloud, STOPWORDS\nimport matplotlib.pyplot as plt\n!pip install jaro-winkler\nimport jaro","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:17:51.342290Z","iopub.execute_input":"2022-07-07T17:17:51.342668Z","iopub.status.idle":"2022-07-07T17:18:05.528873Z","shell.execute_reply.started":"2022-07-07T17:17:51.342639Z","shell.execute_reply":"2022-07-07T17:18:05.527207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv('../input/foursquare-location-matching/train.csv')\n\nprint(\"Total Number of Point Of Interest in Training Dataset\", train_df.shape[0])","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:05.531262Z","iopub.execute_input":"2022-07-07T17:18:05.531685Z","iopub.status.idle":"2022-07-07T17:18:14.593437Z","shell.execute_reply.started":"2022-07-07T17:18:05.531636Z","shell.execute_reply":"2022-07-07T17:18:14.592233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\npairs_df = pd.read_csv('../input/foursquare-location-matching/pairs.csv')\n\npairs_df.head(10)\n\nprint(\"Total Number of Pairs in Dataset {}\".format(len(pairs_df)))\n\nprint(\"Class Distribution\")\n\nmatched_values =pairs_df['match'].value_counts()\n\nprint(\"Total # of Matching pairs in dataset {} , Total Percent of MAtching PAirs {}\".format(matched_values[1],matched_values[1]/len(pairs_df)))\nprint(\"Total # of Not Matching pairs in dataset {} , Total Percent of Not MAtching PAirs {}\".format(matched_values[0],matched_values[0]/len(pairs_df)))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:14.595153Z","iopub.execute_input":"2022-07-07T17:18:14.596025Z","iopub.status.idle":"2022-07-07T17:18:22.844050Z","shell.execute_reply.started":"2022-07-07T17:18:14.595975Z","shell.execute_reply":"2022-07-07T17:18:22.842675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Point Of Interest Data Modeling\n","metadata":{}},{"cell_type":"code","source":"class Location(object):\n    def __init__(self, latitude=None, longitude=None):\n        self.latitude = latitude\n        self.longitude = longitude\n    \n    \n    def __eq__(self, other):\n        pass\n    \n    def __hash__(self):\n        pass\n    \n    def __str__(self):\n        return \"[Latitude : {},  Longitude: {}]\".format(self.latitude or \" \", self.longitude or \" \")\n    \n    def __repr__(self):\n        return \"[Latitude : {},  Longitude: {}]\".format(self.latitude or \" \", self.longitude or \" \")\n    ","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:22.846437Z","iopub.execute_input":"2022-07-07T17:18:22.846779Z","iopub.status.idle":"2022-07-07T17:18:22.854489Z","shell.execute_reply.started":"2022-07-07T17:18:22.846751Z","shell.execute_reply":"2022-07-07T17:18:22.853286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class PointOfInterest(object):\n    \n    def __init__(self,id,POI=None, name=None, latitude=None, longitude=None,\n                 address=None, city=None, state=None, zip=None, country=None, \n                url=None, phone=None, categories=None ):\n        \n        self.id = id\n        self.POI = POI\n        self.name = name\n        self.location  =  Location(latitude,longitude)\n        self.address = address\n        self.city = city\n        self.country =  country\n        self.state = state\n        self.zip = zip\n        self.url = url\n        self.phone = phone\n        self.categories = categories\n    \n    def __eq__(self, other):\n        pass\n    \n    def __repr__(self):\n        return self.__str__()\n    \n    def __hash__(self):\n        pass\n    \n    def __str__(self):\n        return \"[\\nPoint Of Interest : \\n Id : {}, Name : {}, address {}, location : {} , categories : {},\\\n        country {}, url {}, phone {}, state {}, POI {} \\n]\".format(\n            self.id,self.name or \"\", self.address or \" \", self.location.__str__(), self.categories or \" \",   self.country or \" \", self.url or \" \", self.phone or \" \", self.state or \" \", self.POI or \" \"\n        )","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:22.855888Z","iopub.execute_input":"2022-07-07T17:18:22.856235Z","iopub.status.idle":"2022-07-07T17:18:22.870201Z","shell.execute_reply.started":"2022-07-07T17:18:22.856205Z","shell.execute_reply":"2022-07-07T17:18:22.868917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Location Metrics\n* The most obvious approach to Points of Interest matching is to compare the (physical) distance between them. \n* We may assume that two similar (We can conclude using hypthesis)POIs are located a short distance apart, so we can calculate the geographical distance in degrees (distdeg) between them using the following formulas:\n\n$$dist_{deg}= \\sqrt{(POI_{1_{lat}} - POI_{2_{lat}})^2 + dist^{lon}_{deg}(POI_{1},POI_{2})^2} $$\n\nwhere\n$$\ndist^{lon}_{deg}(POI_{1},POI_{2})= \\cos(\\frac{POI_{1_{lat}} * \\pi}{180}) * (POI_{1_{lon}} - POI_{2_{lon}})\n$$\n\n* We have obtained the distance in degree, now we convert distance in degree to distance in meters using following formula\n\n$$\ndist^{m}_{deg}(POI_{1},POI_{2}) = dist_{deg}(POI_{1},POI_{2}) * \\frac{40,075.704}{360} * 1000 meters  \n$$ \n\n\n\n<!-- !(url) [https://latex.codecogs.com/svg.image?dist_%7Bdeg%7D=%20%5Csqrt%7B(POI_%7B1_%7Blat%7D%7D%20-%20POI_%7B2_%7Blat%7D%7D)%5E2%20&plus;%20dist%5E%7Blon%7D_%7Bdeg%7D(POI_%7B1%7D,POI_%7B2%7D)%5E2%7D]\n -->","metadata":{}},{"cell_type":"code","source":"class LocationMetrics(object):\n    \"\"\"\n    This class will be used to find distance or similarity between two\n    Point  of interest based on Location Latitude and Longitude Information\n    \"\"\"\n    \n    def get_geo_distance(self,loc1 , loc2):\n        \"\"\"\n        dist_{deg}= \\sqrt{(POI_{1_{lat}} - POI_{2_{lat}})^2 + dist^{lon}_{deg}(POI_{1},POI_{2})^2} \n\n        where\n        \n        dist^{lon}_{deg}(POI_{1},POI_{2})= \\cos(\\frac{POI_{1_{lat}} * \\pi}{180}) * (POI_{1_{lon}} - POI_{2_{lon}})\n    \n\n        We have obtained the distance in degree, now we convert distance in degree to distance in meters using following formula\n\n    \n        dist^{m}_{deg}(POI_{1},POI_{2}) = dist_{deg}(POI_{1},POI_{2}) * \\frac{40,075.704}{360} * 1000 meters  \n        \n        \n        \n        \"\"\"\n        \n        \n    \n        dist_long_degree = np.cos(loc1.latitude * np.pi/180) * (loc1.longitude-loc2.longitude)\n#         print(\"dist_long_degree\", type(dist_long_degree),dist_long_degree)\n        dist_degree = np.sqrt((loc1.latitude-loc2.latitude)**2  + dist_long_degree**2 )\n#         print(\"dist_degree\", type(dist_degree),dist_degree)\n        dist_meters = dist_degree *40075.704/360 * 1000\n        return dist_meters\n        \n        \n    def __init__(self, metric='geo_distance'):\n        \"\"\"\n        param \n            metric : Distance metric to use for distance calculation\n        \n        \"\"\"\n        self.metric= metric\n        \n    def get_distance(self, loc1, loc2):\n        \"\"\"\n        param\n            loc1 : POI1 location\n            loc2 : POI2 location\n            \n        returns:\n            get distance based on metric selected.\n        \"\"\"\n        if not isinstance(loc1, Location) or not isinstance(loc2, Location):\n            raise Exception(\"loc 1 and loc 2 should be type of Location\")\n            \n        if self.metric == 'geo_distance':\n            dist = self.get_geo_distance(loc1, loc2)\n#             print(\"dist\", dist)\n            return dist\n        \n        \n    \n    def get_similarity(self,location1:Location, location2:Location):\n        \"\"\"\n        param\n            loc1 : POI1 location\n            loc2 : POI2 location\n            \n        returns:\n            get similarity based on metric selected.\n        \"\"\"\n        if not isinstance(loc1, Location) or not isinstance(loc2, Location):\n            raise Exception(\"loc 1 and loc 2 should be type of Location\")\n    \n    \nlm = LocationMetrics()\n# lm.get_distance(Location(),Location())\n        ","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:22.872564Z","iopub.execute_input":"2022-07-07T17:18:22.873876Z","iopub.status.idle":"2022-07-07T17:18:22.888788Z","shell.execute_reply.started":"2022-07-07T17:18:22.873838Z","shell.execute_reply":"2022-07-07T17:18:22.887706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pairs_df['match'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:22.890056Z","iopub.execute_input":"2022-07-07T17:18:22.891303Z","iopub.status.idle":"2022-07-07T17:18:22.912463Z","shell.execute_reply.started":"2022-07-07T17:18:22.891258Z","shell.execute_reply":"2022-07-07T17:18:22.911462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nPOI_source1_features = ['id_1', 'name_1', 'latitude_1', 'longitude_1', 'address_1', 'city_1',\n       'state_1', 'zip_1', 'country_1', 'url_1', 'phone_1', 'categories_1']\n\nPOI_source2_features = ['id_2', 'name_2', 'latitude_2', 'longitude_2', 'address_2', 'city_2',\n       'state_2', 'zip_2', 'country_2', 'url_2', 'phone_2', 'categories_2']\n\n\nY = pairs_df['match']\n\nPOI1 = pairs_df[POI_source1_features]\nPOI1.columns = ['id', 'name', 'latitude', 'longitude', 'address', 'city', 'state',\n       'zip', 'country', 'url', 'phone', 'categories']\nPOI2 = pairs_df[POI_source2_features]\nPOI2.columns = ['id', 'name', 'latitude', 'longitude', 'address', 'city', 'state',\n       'zip', 'country', 'url', 'phone', 'categories']\n\n\nPOI_source1_list=[]\nfor kwargs in POI1.to_dict(orient='records'):\n    \n    POI_source1_list.append(PointOfInterest(**kwargs))\n    \nPOI_source2_list=[]\nfor kwargs in POI2.to_dict(orient='records'):\n    \n    POI_source2_list.append(PointOfInterest(**kwargs))\n    \n","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:22.914307Z","iopub.execute_input":"2022-07-07T17:18:22.914784Z","iopub.status.idle":"2022-07-07T17:18:46.659017Z","shell.execute_reply.started":"2022-07-07T17:18:22.914745Z","shell.execute_reply":"2022-07-07T17:18:46.658065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pre_generated_pairs = set()\n\nfor poi1, poi2 in zip(POI_source1_list,POI_source2_list):\n    pre_generated_pairs.add((poi1.id, poi2.id))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:46.660482Z","iopub.execute_input":"2022-07-07T17:18:46.660884Z","iopub.status.idle":"2022-07-07T17:18:48.094768Z","shell.execute_reply.started":"2022-07-07T17:18:46.660850Z","shell.execute_reply":"2022-07-07T17:18:48.093403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Distance Analysis on Matching And Non Matching Pairs","metadata":{}},{"cell_type":"code","source":"distances = []\nlm = LocationMetrics()\nfor index,data in enumerate(zip(POI_source1_list,POI_source2_list,Y)):\n    poi1,poi2, match = data\n    distances.append(lm.get_distance(poi1.location, poi2.location))\n    \npairs_df['geo_distance'] = distances\n\nsubset = pairs_df[['id_1','id_2','name_1','name_2','categories_1','categories_2','geo_distance','match']]\n\n# get distance between POIs for all matching pairs\nmatching_pairs = pairs_df[pairs_df['match']==True]\nnon_matching_pairs = pairs_df[pairs_df['match']==False]\nmatching_distance = matching_pairs['geo_distance']\nnon_matching_distance = non_matching_pairs['geo_distance']\n\n# _as=['id_1', 'id_2', 'name_1', 'name_2', 'categories_1', 'categories_2',\n#        'geo_distance', 'match']\n# _as.extend(['latitude_1','latitude_2','longitude_1','longitude_2'])\n# pairs_df[pairs_df['geo_distance']==0][pairs_df['match']==False]\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:48.100880Z","iopub.execute_input":"2022-07-07T17:18:48.101261Z","iopub.status.idle":"2022-07-07T17:18:53.390614Z","shell.execute_reply.started":"2022-07-07T17:18:48.101230Z","shell.execute_reply":"2022-07-07T17:18:53.389330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Percentile of Distance for Matching pair ","metadata":{}},{"cell_type":"code","source":"for i in range(0,101,10):\n    print(\"{}th percentile Distance For Matching Pair = {}\".format(i, np.percentile(matching_distance,i)))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:53.392497Z","iopub.execute_input":"2022-07-07T17:18:53.392844Z","iopub.status.idle":"2022-07-07T17:18:53.457785Z","shell.execute_reply.started":"2022-07-07T17:18:53.392813Z","shell.execute_reply":"2022-07-07T17:18:53.456527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Percentile of Distance for Non Matching pair ","metadata":{}},{"cell_type":"code","source":"for i in range(0,101,10):\n    print(\"{}th percentile Distance For Non-Matching Pair = {}\".format(i, np.percentile(non_matching_distance,i)))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:53.459320Z","iopub.execute_input":"2022-07-07T17:18:53.460597Z","iopub.status.idle":"2022-07-07T17:18:53.495329Z","shell.execute_reply.started":"2022-07-07T17:18:53.460548Z","shell.execute_reply":"2022-07-07T17:18:53.494123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observation:\n\n* **90% Matching Pairs** have distance between them **less than 3055 meters**\n* **90% Non-Matching Pairs** have distance between them **less than 1118 meters**\n* By only looking into distance it is not easy to conclude anything If they are matching or non matching, because we don't have direct pattern associated with distance and matching/ nonmatching.\n* We need to **combine** distance with other attribute like **Name, categories** to see if there any pattern jointly associated with categorie and/or name with location.\n","metadata":{}},{"cell_type":"markdown","source":"## Analyze  zero distance apart  pairs (pairs having exact same location)","metadata":{}},{"cell_type":"markdown","source":"### Zero Distance apart Matching pairs","metadata":{}},{"cell_type":"code","source":"# Matching Pairs having 0 distance\ncolumn_subset = ['id_1','id_2','geo_distance','name_1','name_2','categories_1','categories_2','latitude_1','latitude_2','longitude_1','longitude_2']\nmatching_pairs[matching_pairs['geo_distance']==0][column_subset].head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:53.496694Z","iopub.execute_input":"2022-07-07T17:18:53.497010Z","iopub.status.idle":"2022-07-07T17:18:53.552123Z","shell.execute_reply.started":"2022-07-07T17:18:53.496982Z","shell.execute_reply":"2022-07-07T17:18:53.550905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from IPython.display import Image\nImage(\"../input/poi-feature-engineering-analysis/Matching_Pairs_Distance_0.png\")","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:53.553720Z","iopub.execute_input":"2022-07-07T17:18:53.554834Z","iopub.status.idle":"2022-07-07T17:18:53.627032Z","shell.execute_reply.started":"2022-07-07T17:18:53.554798Z","shell.execute_reply":"2022-07-07T17:18:53.625682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Zero Distance apart Non Matching pairs","metadata":{}},{"cell_type":"code","source":"# Non-Matching Pairs having 0 distance\ncolumn_subset = ['id_1','id_2','geo_distance','name_1','name_2','categories_1','categories_2','latitude_1','latitude_2','longitude_1','longitude_2']\nnon_matching_pairs[non_matching_pairs['geo_distance']==0][column_subset].head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:53.628948Z","iopub.execute_input":"2022-07-07T17:18:53.630118Z","iopub.status.idle":"2022-07-07T17:18:53.662345Z","shell.execute_reply.started":"2022-07-07T17:18:53.630070Z","shell.execute_reply":"2022-07-07T17:18:53.660809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<!-- ![](http://)\n![Matching_pairs_zero_distance](../input/poi-feature-engineering-analysis/Matching_Pairs_Distance_0.png)\n\n -->","metadata":{}},{"cell_type":"code","source":"from IPython.display import Image\nImage(\"../input/poi-feature-engineering-analysis/Non_matching_pairs_with_distance_0.png\")","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:53.664372Z","iopub.execute_input":"2022-07-07T17:18:53.665214Z","iopub.status.idle":"2022-07-07T17:18:53.694707Z","shell.execute_reply.started":"2022-07-07T17:18:53.665169Z","shell.execute_reply":"2022-07-07T17:18:53.693659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observation:\n\n* **Name and Categories** very Important to decide If Two POIs matches or not\n* **red color** is used to **anotate** the difference for **Non-matching pairs** having zero distacen apart. i.e. There are couple of pair of POIS with college classroom categories, though they have same location information, they are not same. \n","metadata":{}},{"cell_type":"markdown","source":"## Analyze  Matching Pairs having larger distance ","metadata":{}},{"cell_type":"code","source":"matching_pairs[matching_pairs['geo_distance'] > 3000][column_subset]","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:53.696605Z","iopub.execute_input":"2022-07-07T17:18:53.697366Z","iopub.status.idle":"2022-07-07T17:18:53.764368Z","shell.execute_reply.started":"2022-07-07T17:18:53.697322Z","shell.execute_reply":"2022-07-07T17:18:53.763357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"matching_pairs[matching_pairs['geo_distance'] > 3000][column_subset]['categories_2'].value_counts()[:5]","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:53.765856Z","iopub.execute_input":"2022-07-07T17:18:53.766179Z","iopub.status.idle":"2022-07-07T17:18:53.831746Z","shell.execute_reply.started":"2022-07-07T17:18:53.766152Z","shell.execute_reply":"2022-07-07T17:18:53.830687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"non_matching_pairs[non_matching_pairs['geo_distance'] > 3000][column_subset]","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:53.833189Z","iopub.execute_input":"2022-07-07T17:18:53.833645Z","iopub.status.idle":"2022-07-07T17:18:53.874676Z","shell.execute_reply.started":"2022-07-07T17:18:53.833613Z","shell.execute_reply":"2022-07-07T17:18:53.873532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"non_matching_pairs[non_matching_pairs['geo_distance'] > 3000][column_subset]['categories_2'].value_counts()[:5]","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:53.876133Z","iopub.execute_input":"2022-07-07T17:18:53.877134Z","iopub.status.idle":"2022-07-07T17:18:53.908340Z","shell.execute_reply.started":"2022-07-07T17:18:53.877089Z","shell.execute_reply":"2022-07-07T17:18:53.907485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observation\n* distance between TWO POIS depends of categories as well, e.g **If TWO POIs belongs to Sports stadium, Airports, Hotels, Lakes, Beaches, Gas station or college, these POIs might be far from each other.**\n* We should analyze what's the expected distance for matching and non matching pairs based on categories they belong","metadata":{}},{"cell_type":"code","source":"train_df['categories'].value_counts().to_csv('value_counts.csv')\ntrain_df['categories'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:53.909372Z","iopub.execute_input":"2022-07-07T17:18:53.910385Z","iopub.status.idle":"2022-07-07T17:18:54.492675Z","shell.execute_reply.started":"2022-07-07T17:18:53.910348Z","shell.execute_reply":"2022-07-07T17:18:54.491602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pairs_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:54.533007Z","iopub.execute_input":"2022-07-07T17:18:54.533354Z","iopub.status.idle":"2022-07-07T17:18:54.565171Z","shell.execute_reply.started":"2022-07-07T17:18:54.533315Z","shell.execute_reply":"2022-07-07T17:18:54.563969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # synthetically generated pairs\n# matching_pairs= train_df.groupby('point_of_interest').id.agg('unique').to_dict()\n# print(\"len(matching_pairs),\",len(matching_pairs))\n# matching_paired_dict  = {K:V for K,V in matching_pairs.items() if len(V)>1}\n# print(\"len(matching_pairs),\",len(matching_paired_dict))\n\n    \n\n# synthetic_pairs = []\n# index=0\n# # print(max(matching_paired_dict.values()))\n# max_len = 0\n# for K,V in matching_paired_dict.items():\n#     for i_1 in range(len(V)-1):\n#         for i_2 in range(i_1+1, len(V)):\n#             poi_1 = V[i_1]\n#             poi_2 = V[i_2]\n# #             if (poi_1,poi_2) in pre_generated_pairs  (poi_2,poi_1) in   pre_generated_pairs:\n# #                 continue\n# #             else:\n# #                 synthetic_pairs.append((poi_1, poi_2))\n#             if ((poi_1, poi_2) not in pre_generated_pairs ) and ((poi_2, poi_1) not in pre_generated_pairs ):\n#                 synthetic_pairs.append((poi_1, poi_2))\n    \n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:54.596481Z","iopub.execute_input":"2022-07-07T17:18:54.597672Z","iopub.status.idle":"2022-07-07T17:18:54.604241Z","shell.execute_reply.started":"2022-07-07T17:18:54.597625Z","shell.execute_reply":"2022-07-07T17:18:54.603082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# String Similarity","metadata":{}},{"cell_type":"markdown","source":"* **String similarity metrics** are useful when we want to compare string that refers to the same thing but **written differently**, **have some spelling mistake, order of words are different, numbering format is different.**\n* Here are few string similarity metrics.\n1. levenshtein distance\n2. Jaro-Winkler distance\n3. Ratio Metric\n4. Partial Ratio Metric\n5. Token Sort Ratio Metric\n6. Token Set Ratio Metric\n\n* In the next section we will discuss and experiment on all of the Above String metrics\n\n","metadata":{}},{"cell_type":"markdown","source":"# 1. Levenshtein Distance\n\n* Levenshtein Distance is defined as minimum number of operations required one string to another string. All the possible operations allowed are below\n1. deleting character from string \n    $$ lev_{a,b}(i-1,j)+1 $$\n2. inserting new character into string\n    $$ lev_{a,b}(i,j-1)+1 $$\n3. replacing character inside string with another character.\n    $$ lev_{a,b}(i-1,j-1)+1_{a_{i} \\neq b_{j}} $$\n    \n\n\n","metadata":{}},{"cell_type":"code","source":"from IPython.display import Image\nImage(\"../input/poi-feature-engineering-analysis/levensht.png\")","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:54.610416Z","iopub.execute_input":"2022-07-07T17:18:54.611299Z","iopub.status.idle":"2022-07-07T17:18:54.632379Z","shell.execute_reply.started":"2022-07-07T17:18:54.611246Z","shell.execute_reply":"2022-07-07T17:18:54.631208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* We can define **levenshtein Similarity** based on Levenshtein Distance formula.\n\n$$\\frac{(|a| + |b|) - {lev}_{a,b}(i,j)}{|a| + |b|}$$\n\nwhere $|a|$ and $|b|$ are the lengths of sequence $a$ and sequence $b$ respectively.\n\n* **levenshtein distance** is sensitive to the lower or upper case letter, it treat them differently and **include the cost of changing to lower or upper into account**.\n\n* **string processing (convert string into lower case before calculating distance is good practice.**\n","metadata":{}},{"cell_type":"code","source":"# Basic Dynamic programming solution for levenshetein distance\n\ndef levenshtein_distance(str1, str2):\n    \"\"\"\n    Python code to find the minimum cost to convert str1 to str2\n    \n    cache[i,j] is the distance between str1 strating  ith character of str1 and str2 starting with jth character\n    \n    assume the cost of insertion 1, deletion 1, substitution 2 (insert + deletion)\n    \n    \"\"\"\n    \n    rows = len(str1)\n    cols = len(str2)\n    \n    cache = [[float(\"inf\")] * (cols+1) for row in range(rows+1)]\n    \n    \n    \n    # number of step required to convert str1 to empty str2\n    for j in range(cols+1):\n        cache[rows][j] = cols-j\n    # number of step required to convert empty str1 to  str2\n    for i in range(rows+1):\n        cache[i][cols]=rows-i\n        \n    for i in range(rows-1,-1,-1):\n        for j in range(cols-1,-1,-1):\n            if str1[i] == str2[j]:\n                cache[i][j] = cache[i+1][j+1]\n            else:\n                cache[i][j] = min(\n                                    1 +  cache[i+1][j], # deletion\n                                    1 + cache[i][j+1], # insertion\n                                    2 + cache[i+1][j+1] #replacement Cost of replacement is two ( deletion + insertion)\n                                 )\n                \n    distance = cache[0][0]\n    similarity = ((len(str1)+len(str2)+0.001) - distance )/(len(str1)+len(str2)+0.001)\n#     similarity = 1 - (distance/ max(len(str1), len(str2)))\n    return cache[0][0],similarity\n    \n    \nStr1 = \"Indian Ocean.\"\nStr2 = \"indian Ocean\"\nDistance,similarity = levenshtein_distance(Str1.lower(),Str2.lower())\nprint(\"The strings are {} edits away\".format(Distance))\n# Ratio = levenshtein_ratio_and_distance(Str1,Str2,ratio_calc = True)\nprint(\"similarity score {} \".format(similarity) )        \n","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:54.633883Z","iopub.execute_input":"2022-07-07T17:18:54.634226Z","iopub.status.idle":"2022-07-07T17:18:54.650385Z","shell.execute_reply.started":"2022-07-07T17:18:54.634197Z","shell.execute_reply":"2022-07-07T17:18:54.649177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Str1 = \"Indian Ocean.\"\nStr2 = \"indian Ocean\"\n\nDistance,similarity = levenshtein_distance(Str1.lower(),Str2.lower())\nprint(\"The strings are {} edits away\".format(Distance))\n# Ratio = levenshtein_ratio_and_distance(Str1,Str2,ratio_calc = True)\nprint(\"similarity score {} \".format(similarity) )  ","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:54.651959Z","iopub.execute_input":"2022-07-07T17:18:54.652314Z","iopub.status.idle":"2022-07-07T17:18:54.667127Z","shell.execute_reply.started":"2022-07-07T17:18:54.652282Z","shell.execute_reply":"2022-07-07T17:18:54.665547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Comparision with prebuilt package","metadata":{}},{"cell_type":"code","source":"import Levenshtein as lev\nprint(\"The strings are {} edits away\".format(lev.distance(Str1.lower(), Str2.lower())))\n\nprint(\"similarity score {} \".format(lev.ratio(Str1.lower(), Str2.lower())) )  ","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:54.669062Z","iopub.execute_input":"2022-07-07T17:18:54.669949Z","iopub.status.idle":"2022-07-07T17:18:54.686211Z","shell.execute_reply.started":"2022-07-07T17:18:54.669904Z","shell.execute_reply":"2022-07-07T17:18:54.684968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from fuzzywuzzy import fuzz\nprint(\"similarity score {} \".format(fuzz.ratio(Str1.lower(), Str2.lower())) ) \nfrom fuzzywuzzy import fuzz\nprint(\"similarity score {} \".format(fuzz.ratio(Str1.lower(), \"\".lower())) )","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:54.688105Z","iopub.execute_input":"2022-07-07T17:18:54.688961Z","iopub.status.idle":"2022-07-07T17:18:54.705876Z","shell.execute_reply.started":"2022-07-07T17:18:54.688916Z","shell.execute_reply":"2022-07-07T17:18:54.704484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **Before string processing similarity score was 0.88** and ***after string processing similarity score was 0.96.***\n* If we had used **threshold value of 0.9** for Similarity  threshold, **we would have missed this match.**\n* My custom implementation of levenshtein distance are equal","metadata":{}},{"cell_type":"markdown","source":"# 2. Jaro  Similarity\n","metadata":{}},{"cell_type":"markdown","source":"* **Jaro  Metric** determines the distance between two string with following equation\n\n$$ d_{JW}(s1,s2) =  \\frac{1}{3} \\left (  \\frac{m}{\\left| s1 \\right|} + \\frac{m}{\\left| s2 \\right|} + \\frac{m-t}{m} \\right ) $$\n\nwhere m is number of matching characters, $$ m \\neq 0 $$ and  **t is half the number of transposition** ( **matching chacters but at different position in the inscriptions being compared**)\n\nIn short Jaro  metric is average of \n\n* The **number of matching characters** in both string compare to **length of first string**\n* The **number of matching characters** in both string compare to **length of second string**\n* The **number of matching characters** that **don't require transposition**\n\n\n*  **jaro_winkler_metric(string1, string2)**\n        The Jaro metric adjusted with Winkler's modification, \n        which boosts the metric for strings whose  prefixes match\n","metadata":{}},{"cell_type":"code","source":"# https://github.com/richmilne/JaroWinkler/blob/43e7e186d791622ba6636402b1c2590ef765cd51/jaro/jaro.py\ndef count_transposition(str1, str2, bitmask_str1, bitmask_str2):\n    \"\"\"\n        Transposition is defined as characters that are matched in both string 1 and string 2 but are at \n        different position\n        \n        bitmask_str1, and bitmask_str2 is the bitmask returned by count_match function \n        which gives where in string1 and string2 characters match\n        \n      \n    \"\"\"\n    \n    transposition = 0\n    k = 0\n    \n    # iterate over every match char in string1\n    for index, bit in enumerate(bitmask_str1):\n        if not bit :\n            continue\n        # for every match character from string 1\n        # we will find match chacter in string 2 starting from where we left from last time k\n        while not bitmask_str1[k]:\n            k+=1\n            \n        # once we found matching chars in both string\n        # if they are not equal increment transposition\n        if str1[index] != str2[k]:\n            transposition+=1\n        k+=1\n    \n    return transposition\n        \n\ndef count_match(str1, str2, len1,len2):\n    \"\"\"\n    \n        For every char in str1, count the number fo chars in str2 that're matching \n        in given range.\n        \n        len1, len2 are length of str1 and str2 respectively.\n        \n        The function returns the count the number of matched chars, two bit mask array showing\n        which char in each string are matches, and location matched in str1 where\n    \n    \"\"\"\n    \n    matching_range = max(0, len2//2-1)\n    \n    num_matches = 0\n    \n    # bitmask array to find out where there is match\n    bitmask_str1=[0]*len1\n    bitmask_str2=[0]*len2\n    \n    matched_locations=[-1]*len1\n    \n    \n    \n    for index, ch in enumerate(str1):\n    # count within search distance, and get matching count\n        lower = max(0, index-matching_range)\n        higher = min(index+matching_range, len2-1)\n        for j in range(lower, higher+1):\n            if not bitmask_str2[j] and str2[j]==ch:\n                # which char in str2 matches\n                bitmask_str2[j]=1\n                # which char in str1 matches\n                bitmask_str1[index]=1\n                # characters in str1 matches with which index in string2\n                matched_locations[index]=j\n                num_matches+=1\n\n                break\n                \n    return num_matches,bitmask_str1, bitmask_str2, matched_locations\n        \n            \n        \n    \n        \n\ndef Jaro_similarity(str1, str2):\n    \"\"\"\n    The Jaro Winkler  distance between is the min no. of single-character transpositions\n    required to change one word into another.\n    \n    we will asuume that str1 length is less than or equals to str2 length\n    \n    jaro_winkler = 0 if m = 0 else 1/3 * (m/|s_1| + m/|s_2| + (m-t)/m)\n\n    where:\n        - |s_1| is the length of string s_1\n        - |s_2| is the length of string s_2\n        - m is the no. of matching characters\n        - t is the half no. of possible transpositions. ( matching characters that are at different positioin)\n    \n    \n    \"\"\"\n    if not str1:\n        if not str2: return 1.0\n        return 0.0\n    \n    len2=len(str2)\n    len1=len(str1)\n    \n    if len2 < len1:\n        str1,str2=str2, str1\n        len1, len2 = len2, len1\n        \n    assert len1<=len2\n    \n    num_of_matches, bitmap_str1, bitmap_str2, match_locations = count_match(str1, str2,len1, len2)\n    num_transposition  = count_transposition(str1, str2,bitmap_str1, bitmap_str2 )\n    \n    if not num_of_matches: return 0.0\n    \n    \n    \n    weight = (  num_of_matches / len1\n              + num_of_matches / len2\n              + (num_of_matches - num_transposition//2) / num_of_matches)\n\n    return weight / 3\n    \n    \n# http://www.alias-i.com/lingpipe/docs/api/com/aliasi/spell/JaroWinklerDistance.html\nimport jaro\njs=jaro.jaro_metric('Cinema City Poland'.lower(),'Cinema City'.lower())\nprint(\"JARO DIstance between \\\"{}\\\" and \\\"{}\\\"  using JARO package is {} \".format(\"Cinema City Poland\",\"Cinema City\",js))\n\ncustom_js = Jaro_similarity('Cinema City Poland','Cinema City')\nprint(\"JARO DIstance between \\\"{}\\\" and \\\"{}\\\"  using Custom implementation  is {} \".format(\"Cinema City Poland\",\"Cinema City\",custom_js))\n","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:54.708337Z","iopub.execute_input":"2022-07-07T17:18:54.709223Z","iopub.status.idle":"2022-07-07T17:18:54.731376Z","shell.execute_reply.started":"2022-07-07T17:18:54.709175Z","shell.execute_reply":"2022-07-07T17:18:54.730013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3.  Ratio Metric","metadata":{}},{"cell_type":"markdown","source":"* Ratio Metric is defined as Normalized similarity between string calculated as  levenshtein distance devided by length of longer string\n","metadata":{}},{"cell_type":"code","source":"Str1='Cinema City Poland'\nStr2='Cinema City'\nDistance,similarity = levenshtein_distance(Str1.lower(),Str2.lower())\nprint(\"Ratio Metric  between \\\"{}\\\" and \\\"{}\\\"  using Custom implementation  is {} \".format(\"Cinema City Poland\",\"Cinema City\",np.round(similarity,3)))\n\nfuzz_ratio = fuzz.ratio(Str1.lower(), Str2.lower())/100\nprint(\"Ratio Metric  between \\\"{}\\\" and \\\"{}\\\"  using fuzz package   is {} \".format(\"Cinema City Poland\",\"Cinema City\",np.round(fuzz_ratio,3)))\n","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:54.733191Z","iopub.execute_input":"2022-07-07T17:18:54.733982Z","iopub.status.idle":"2022-07-07T17:18:54.746868Z","shell.execute_reply.started":"2022-07-07T17:18:54.733936Z","shell.execute_reply":"2022-07-07T17:18:54.745631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Partial Ratio Metric\n\n* Partial ratio metric is based on substring matching, it takes the shortest string and compare with all the substring having same length from larger string. \n* e.g If we compare two string **\"Boston Celtic\"** and **\"Boston\"** then, **ratio metric** will return **0.63**, but **partial ratio** metric yields  **100% similarity**\n\n\n","metadata":{}},{"cell_type":"code","source":"print(\"Testing Strings: {} and {}\".format(\"Boston celtic\",\"Boston\"))\n\nprint(\"Raio Metric \",fuzz.ratio(\"Boston celtic\",\"Boston\"))\nprint(\"Partial Ratio Metric\",fuzz.partial_ratio(\"Boston celtic\",\"Boston\"))\n","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:54.748506Z","iopub.execute_input":"2022-07-07T17:18:54.748882Z","iopub.status.idle":"2022-07-07T17:18:54.758480Z","shell.execute_reply.started":"2022-07-07T17:18:54.748842Z","shell.execute_reply":"2022-07-07T17:18:54.757081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Token Sort Ratio Metric\n\n* Partial Ratio would not work if Both the string are same, but **tokens are in different order**.\n     \n     i.e **\"Board of Education vs Brown Case\" and \"Brown vs Board of Education case\"**\n\n\n* In **Token Sort ratio** we first **tokenize** the string and **preprocess** every token convert into **lower** case and **remove punctuation**, before joining string **tokens** are **sorted** alphabatically and then combined together. \n* After that simple** Ratio Metric** is applied to get similarity. If we apply Token Sort Ratio over above example we will **yield 100 % similarity**.\n","metadata":{}},{"cell_type":"code","source":"print(\"Testing Strings: {} and {}\".format(\"Board of Education vs Brown Case\",\"Brown vs Board of Education case\"))\nprint(\"Partial Ratio\",fuzz.partial_ratio(\"Board of Education vs Brown Case\",\"Brown vs Board of Education case\"))\nprint(\"Token Sort Ratio\",fuzz.token_sort_ratio(\"Board of Education vs Brown Case\",\"Brown vs Board of Education case\"))\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:54.761695Z","iopub.execute_input":"2022-07-07T17:18:54.763444Z","iopub.status.idle":"2022-07-07T17:18:54.769803Z","shell.execute_reply.started":"2022-07-07T17:18:54.763405Z","shell.execute_reply":"2022-07-07T17:18:54.768705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Token Set Ratio Metric\n\n* **Token Sort Ratio** would **not give good results**, If two strings are **dissimilar in length widely**, despite they share common things in different order.\n* **Token Set Ratio** perform all the operations done by token sort ratio, apart from that it **performs set operation that extracts the common tokens (set intersection)** and then use ratio metric between following strings.\n\n\n1. s1 = **Sorted_tokens_in_intersection**\n2. s2 = Sorted_tokens_in_intersection + **sorted_rest_of_str1_tokens**\n3. s3 = Sorted_tokens_in_intersection + **sorted_rest_of_str2_tokens**\n\n* So we already have intersection of tokens which is same, along with it If remaining tokens are closer to each other then it will improve similarity score.\n\n","metadata":{}},{"cell_type":"code","source":"str1=\"Board of Education vs Brown Case\"\nstr2=\"Supreme court case of Brown vs Board of Education case\"\nprint(\"Testing Strings: {} and {}\".format(str1, str2))\nprint(\"Ratio\",fuzz.ratio(str1, str2))\nprint(\"Partial Ratio\",fuzz.partial_ratio(str1, str2))\nprint(\"Token Sort Ratio\",fuzz.token_sort_ratio(str1, str2))\nprint(\"Token Set Ratio\",fuzz.token_set_ratio(str1, str2))\n","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:54.771727Z","iopub.execute_input":"2022-07-07T17:18:54.772725Z","iopub.status.idle":"2022-07-07T17:18:54.784374Z","shell.execute_reply.started":"2022-07-07T17:18:54.772645Z","shell.execute_reply":"2022-07-07T17:18:54.783229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testcases=[[\"Indian Ocean\",\"Indian ocean.\"],\n         [\"Cinema City Poland\",\"Cinema City\"],\n          [\"Boston Celtic\",\"Boston\"],\n          [\"Board of Education vs Brown Case\",\"Supreme court case of Brown vs Board of Education case\"],\n         [\"Board of Education vs Brown Case\",\"Brown vs Board of Education case\"],\n         [\"Henryka Kamienskiego 123 Krakow polska\",\"Generala Henryka Kamienskiego 30-644 Krakow\"],\n          [\"+1456321123\",\"456321123\"],\n          \n          [\"www.snap.com/region/us\",\"www.snap.com\"]\n         ]","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:18:54.786257Z","iopub.execute_input":"2022-07-07T17:18:54.787140Z","iopub.status.idle":"2022-07-07T17:18:54.797360Z","shell.execute_reply.started":"2022-07-07T17:18:54.787083Z","shell.execute_reply":"2022-07-07T17:18:54.796235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"myTable =[]\n\nfor test_case in testcases:\n    str1 = test_case[0]\n    str2 = test_case[1]\n    str1=str1.lower()\n    str2 = str2.lower()\n    \n    Distance,leven = np.round(levenshtein_distance(str1,str2),2)\n    jaro_distance = np.round(Jaro_similarity(str1, str2),2)\n    jaro_winkler = np.round(jaro.jaro_winkler_metric(str1, str2),2)\n    ratio = np.round(fuzz.ratio(str1, str2)/100,2)\n    partial_rat = np.round(fuzz.partial_ratio(str1, str2)/100,2)\n    token_sort_rat = np.round(fuzz.token_sort_ratio(str1, str2)/100,2)\n    token_set_rat = np.round(fuzz.partial_token_set_ratio(str1, str2)/100,2)\n    avg_rat = np.round((token_set_rat+partial_rat)/2,2)\n    myTable.append([str1, str2, leven, jaro_distance,jaro_winkler,\n                   ratio,partial_rat,token_sort_rat,token_set_rat,avg_rat])\n    \n# print(myTable)\nsimilarity_df = pd.DataFrame(myTable, columns = [\"string 1\", \"string 2\", \"Levenshtein\", \"Jaro\",\"JW\",\"Ratio\",\"Part Ratio\",\"Tkn_Sort_Rat\",\"Token Set Ratio\",\"average[Part,token set ratio]\"])\nsimilarity_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:22:37.590747Z","iopub.execute_input":"2022-07-07T17:22:37.591137Z","iopub.status.idle":"2022-07-07T17:22:37.629536Z","shell.execute_reply.started":"2022-07-07T17:22:37.591107Z","shell.execute_reply":"2022-07-07T17:22:37.628305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class StringMetrics(object):\n    \"\"\"\n    This class will be used to find distance or similarity between two\n    Point  of interest based on String attribute value\n    \"\"\"\n    \n    def __init__(self, metric='levenshtein'):\n        \"\"\"\n        param \n            metric : Distance metric to use for string similarity/ distance \n        \n        \"\"\"\n        self.metric= metric\n        \n    def get_distance(self, str1:str, str2:str):\n        \"\"\"\n        param\n            str1 : POI1 string feature1\n            str1 : POI2 string feature2\n            \n        returns:\n            get Distance based on metric selected default levenshtein\n        \"\"\"\n        if not isinstance(str1, str) or not isinstance(str2, str):\n            raise Exception(\"str1 and str2  should be type of String\")\n            \n        if self.metric == 'levenshtein':\n            distance, similarity = levenshtein_distance(str1, str2)\n            return distance\n        \n        \n    def get_similarity(self,str1:str, str2:str):\n        \"\"\"\n        param\n            str1 : POI1 string feature1\n            str1 : POI2 string feature2\n            \n        returns:\n            get similarity based on metric selected,default levenshtein\n            \n            Possible Metric Methods:\n            \n            \n                Levenshtein\n                Jaro Distance\n                Jaro Winkler Distance\n                Partial Ratio\n                token sort ratio\n                token set ratio\n                token partial sort ratio\n                token partial set ratio\n        \"\"\"\n        if not isinstance(str1, str) or not isinstance(str2, str):\n            raise Exception(\"str1 and str2  should be type of String\")\n        \n        str1=str1.lower()\n        str2=str2.lower()\n        if self.metric == 'levenshtein':\n            distance, similarity = levenshtein_distance(str1, str2)\n            return similarity\n        \n        if self.metric == 'jaro':\n            similarity = Jaro_similarity(str1, str2)\n            return similarity\n        \n        if self.metric == 'jaro_winkler':\n            return jaro.jaro_winkler_metric(str1, str2)\n        \n        if self.metric == 'partial_ratio':\n            return np.round(fuzz.partial_ratio(str1, str2)/100,2)\n        \n        if self.metric=='token_sort_ratio':\n            return np.round(fuzz.token_sort_ratio(str1, str2)/100,2)\n        \n        if self.metric=='token_set_ratio':\n            return np.round(fuzz.token_set_ratio(str1, str2)/100,2)\n        \n        if self.metric=='token_partial_sort_ratio':\n            return np.round(fuzz.partial_token_sort_ratio(str1, str2)/100,2)\n        \n        if self.metric=='token_partial_set_ratio':\n            return np.round(fuzz.partial_token_set_ratio(str1, str2)/100,2)\n        \n        raise Exception(\"Unknown Metric\")\n        \n        \n    \n    \n    \nlm = StringMetrics('token_partial_sort_ratio')\nlm.get_similarity(\"Boston celtic\",\"Boston\")\n        ","metadata":{"execution":{"iopub.status.busy":"2022-07-07T17:32:13.764326Z","iopub.execute_input":"2022-07-07T17:32:13.764816Z","iopub.status.idle":"2022-07-07T17:32:13.789753Z","shell.execute_reply.started":"2022-07-07T17:32:13.764781Z","shell.execute_reply":"2022-07-07T17:32:13.788521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}