{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Overview","metadata":{}},{"cell_type":"markdown","source":"In this notebook, I will try to clean/analyze/visualize data, create new, meaningfull features and choose a suitable algorithm for predictions. As competition rules states, my goal is to get predictions which gives  low Test RMSE (root mean squared error).\n\n**Result** : **RMSE** _ **3.12271**","metadata":{}},{"cell_type":"markdown","source":"<font size=\"4\">1) Imports</font>\n\n<font size=\"4\">2) Upload Data</font>\n\n<font size=\"4\">3) Data Cleaning</font>\n\n<font size=\"4\">4) Feature Engineering</font>\n\n<font size=\"4\">5) Data Analyisis/Visualisation</font>\n\n<font size=\"4\">6) Modeling</font>\n\n<font size=\"4\">7) Results</font>","metadata":{}},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom xgboost import XGBRegressor\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.metrics import mean_squared_error\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.model_selection import cross_val_score, RandomizedSearchCV, GridSearchCV\nfrom sklearn.cluster import KMeans\nfrom yellowbrick.cluster import KElbowVisualizer\n\n\nimport random\n\nimport warnings\n\nwarnings.filterwarnings('ignore')\nsns.set_style('darkgrid')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:33:31.523876Z","iopub.execute_input":"2022-08-09T18:33:31.524238Z","iopub.status.idle":"2022-08-09T18:33:31.531444Z","shell.execute_reply.started":"2022-08-09T18:33:31.524207Z","shell.execute_reply":"2022-08-09T18:33:31.530251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Upload Data","metadata":{}},{"cell_type":"markdown","source":"To have basic undestanding of data efficiently, I will use method offered in notebook: *https://www.kaggle.com/code/szelee/how-to-import-a-csv-file-of-55-million-rows*. Compared to pandas read_csv method, it is faster way to upload/reload large dataframes, detect length of dataframe, ect.","metadata":{}},{"cell_type":"code","source":"# TRAIN_PATH = '../input/new-york-city-taxi-fare-prediction/train.csv'","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:33:31.542379Z","iopub.execute_input":"2022-08-09T18:33:31.543503Z","iopub.status.idle":"2022-08-09T18:33:31.548289Z","shell.execute_reply.started":"2022-08-09T18:33:31.543467Z","shell.execute_reply":"2022-08-09T18:33:31.547370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n\n# with open(TRAIN_PATH) as file:\n#     n_rows = len(file.readlines())\n\n# print (f'Exact number of rows: {n_rows}')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:33:31.560106Z","iopub.execute_input":"2022-08-09T18:33:31.560621Z","iopub.status.idle":"2022-08-09T18:33:31.565557Z","shell.execute_reply.started":"2022-08-09T18:33:31.560584Z","shell.execute_reply":"2022-08-09T18:33:31.564350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Lets check first 3 rows\n# temp_df = pd.read_csv(TRAIN_PATH, nrows=3)\n# temp_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:33:31.577376Z","iopub.execute_input":"2022-08-09T18:33:31.578387Z","iopub.status.idle":"2022-08-09T18:33:31.582591Z","shell.execute_reply.started":"2022-08-09T18:33:31.578357Z","shell.execute_reply":"2022-08-09T18:33:31.581545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# temp_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:33:31.587919Z","iopub.execute_input":"2022-08-09T18:33:31.588907Z","iopub.status.idle":"2022-08-09T18:33:31.594544Z","shell.execute_reply.started":"2022-08-09T18:33:31.588813Z","shell.execute_reply":"2022-08-09T18:33:31.593622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This dataset contains 55,423,857 observations but training models with 55 million rows will be too time consuming. Firstly I will upload 5,000,000 observations from which I will randomly choose 1,000,000 observations to train models and see what it will give us; If needed, we can add more observations to our subsample.\n\nTo optimize memory usage, I want to change data types of variables.\n\nAlso, column ***key*** is described as *Unique string identifying each row*; we do not need it in train set, so I will remove this column.","metadata":{}},{"cell_type":"code","source":"%%time\n\ncolumn_types = {'fare_amount': 'float64',\n                'pickup_datetime': 'str',\n                'pickup_longitude': 'float32',\n                'pickup_latitude': 'float32',\n                'dropoff_longitude': 'float32',\n                'dropoff_latitude': 'float32',\n                'passenger_count': 'uint8'}\n\ncols = list(column_types.keys())\n\ntrain_df = pd.read_csv('../input/new-york-city-taxi-fare-prediction/train.csv', usecols=cols,\n                 dtype=column_types, nrows=5000000)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:33:31.614974Z","iopub.execute_input":"2022-08-09T18:33:31.615228Z","iopub.status.idle":"2022-08-09T18:33:37.388714Z","shell.execute_reply.started":"2022-08-09T18:33:31.615204Z","shell.execute_reply":"2022-08-09T18:33:37.387753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"random.seed(4)\n\nn = 5000000 # number of records in file\ns = 4000000 # number of observations not to include\nsubsample_indexes = sorted(random.sample(range(n),n-s))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:33:37.390509Z","iopub.execute_input":"2022-08-09T18:33:37.391079Z","iopub.status.idle":"2022-08-09T18:33:39.306429Z","shell.execute_reply.started":"2022-08-09T18:33:37.391041Z","shell.execute_reply":"2022-08-09T18:33:39.305459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = train_df.iloc[subsample_indexes]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:33:39.307975Z","iopub.execute_input":"2022-08-09T18:33:39.308344Z","iopub.status.idle":"2022-08-09T18:33:39.680574Z","shell.execute_reply.started":"2022-08-09T18:33:39.308307Z","shell.execute_reply":"2022-08-09T18:33:39.679534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del subsample_indexes","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:33:39.683578Z","iopub.execute_input":"2022-08-09T18:33:39.683983Z","iopub.status.idle":"2022-08-09T18:33:39.731962Z","shell.execute_reply.started":"2022-08-09T18:33:39.683945Z","shell.execute_reply":"2022-08-09T18:33:39.730574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now lets upload test dataset:","metadata":{}},{"cell_type":"code","source":"test_df = pd.read_csv('../input/new-york-city-taxi-fare-prediction/test.csv')\ndisplay(test_df.head())\nprint(f'Shape of test dataframe: {test_df.shape}')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:33:39.733792Z","iopub.execute_input":"2022-08-09T18:33:39.734844Z","iopub.status.idle":"2022-08-09T18:33:39.769635Z","shell.execute_reply.started":"2022-08-09T18:33:39.734798Z","shell.execute_reply":"2022-08-09T18:33:39.768711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# these columns are datetime so lets change datatypes\ntrain_df['pickup_datetime'] = train_df['pickup_datetime'].astype('datetime64[ns, UTC]')\ntest_df['pickup_datetime'] = test_df['pickup_datetime'].astype('datetime64[ns, UTC]')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:33:39.770972Z","iopub.execute_input":"2022-08-09T18:33:39.771604Z","iopub.status.idle":"2022-08-09T18:35:46.298404Z","shell.execute_reply.started":"2022-08-09T18:33:39.771566Z","shell.execute_reply":"2022-08-09T18:35:46.297430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Cleaning","metadata":{}},{"cell_type":"code","source":"# just save number of observations before data cleaning\nn_before = train_df.shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.299719Z","iopub.execute_input":"2022-08-09T18:35:46.300173Z","iopub.status.idle":"2022-08-09T18:35:46.305294Z","shell.execute_reply.started":"2022-08-09T18:35:46.300136Z","shell.execute_reply":"2022-08-09T18:35:46.304102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Handling Missing Values","metadata":{}},{"cell_type":"code","source":"# firstly, lets check whether we have missing values or not\ntrain_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.306830Z","iopub.execute_input":"2022-08-09T18:35:46.307453Z","iopub.status.idle":"2022-08-09T18:35:46.331555Z","shell.execute_reply.started":"2022-08-09T18:35:46.307417Z","shell.execute_reply":"2022-08-09T18:35:46.330617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our dataset is large and there are few missing values, so removing observations with missing values seems to be good solution.","metadata":{}},{"cell_type":"code","source":"train_df.dropna(axis=0, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.332752Z","iopub.execute_input":"2022-08-09T18:35:46.333153Z","iopub.status.idle":"2022-08-09T18:35:46.387863Z","shell.execute_reply.started":"2022-08-09T18:35:46.333118Z","shell.execute_reply":"2022-08-09T18:35:46.386912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Handling Outliers","metadata":{}},{"cell_type":"code","source":"train_df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.392113Z","iopub.execute_input":"2022-08-09T18:35:46.392837Z","iopub.status.idle":"2022-08-09T18:35:46.603587Z","shell.execute_reply.started":"2022-08-09T18:35:46.392787Z","shell.execute_reply":"2022-08-09T18:35:46.602535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"At first, we can see several issues:\n* ***fare amount*** has negative values (negative fare does not make sense);\n* **Latitudes** are not in range from -90 to 90;\n* **Longitudes** are not in range from -180 to 180;\n* ***passenger_count*** have values more than 7 (maximum number of passengers is 7 as google says).","metadata":{"execution":{"iopub.status.busy":"2022-08-02T12:22:48.150066Z","iopub.execute_input":"2022-08-02T12:22:48.151101Z","iopub.status.idle":"2022-08-02T12:22:48.158633Z","shell.execute_reply.started":"2022-08-02T12:22:48.151056Z","shell.execute_reply":"2022-08-02T12:22:48.157391Z"}}},{"cell_type":"markdown","source":"#### *fare amount*","metadata":{}},{"cell_type":"code","source":"# check number of observations with negative fare amount value\ntrain_df.loc[train_df['fare_amount'] < 0].shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.605382Z","iopub.execute_input":"2022-08-09T18:35:46.606068Z","iopub.status.idle":"2022-08-09T18:35:46.615257Z","shell.execute_reply.started":"2022-08-09T18:35:46.606027Z","shell.execute_reply":"2022-08-09T18:35:46.614346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# drop these outliers\ntrain_df = train_df.loc[train_df['fare_amount'] >= 0]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.616576Z","iopub.execute_input":"2022-08-09T18:35:46.617505Z","iopub.status.idle":"2022-08-09T18:35:46.656788Z","shell.execute_reply.started":"2022-08-09T18:35:46.617468Z","shell.execute_reply":"2022-08-09T18:35:46.655915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### *pickup_longitude*","metadata":{}},{"cell_type":"code","source":"# check number of observations with longitude not in [-180,180]\ntrain_df.loc[(train_df['pickup_longitude'] < -180) | (train_df['pickup_longitude'] > 180)].shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.658121Z","iopub.execute_input":"2022-08-09T18:35:46.658910Z","iopub.status.idle":"2022-08-09T18:35:46.670251Z","shell.execute_reply.started":"2022-08-09T18:35:46.658872Z","shell.execute_reply":"2022-08-09T18:35:46.669137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# drop these outliers\ntrain_df = train_df.loc[(train_df['pickup_longitude'] >= -180) & (train_df['pickup_longitude'] <= 180)]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.673436Z","iopub.execute_input":"2022-08-09T18:35:46.674168Z","iopub.status.idle":"2022-08-09T18:35:46.714464Z","shell.execute_reply.started":"2022-08-09T18:35:46.674114Z","shell.execute_reply":"2022-08-09T18:35:46.713523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### *pickup_latitude*","metadata":{}},{"cell_type":"code","source":"# check number of observations with latitude not in [-90,90]\ntrain_df.loc[(train_df['pickup_latitude'] < -90) | (train_df['pickup_latitude'] > 90)].shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.717712Z","iopub.execute_input":"2022-08-09T18:35:46.717989Z","iopub.status.idle":"2022-08-09T18:35:46.729983Z","shell.execute_reply.started":"2022-08-09T18:35:46.717963Z","shell.execute_reply":"2022-08-09T18:35:46.728833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# drop these outliers\ntrain_df = train_df.loc[(train_df['pickup_latitude'] >= -90) & (train_df['pickup_latitude'] <= 90)]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.731535Z","iopub.execute_input":"2022-08-09T18:35:46.732173Z","iopub.status.idle":"2022-08-09T18:35:46.771611Z","shell.execute_reply.started":"2022-08-09T18:35:46.732137Z","shell.execute_reply":"2022-08-09T18:35:46.770710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### *dropoff_longitude*","metadata":{}},{"cell_type":"code","source":"# check number of observations with longitude not in [-180,180]\ntrain_df.loc[(train_df['dropoff_longitude'] < -180) | (train_df['dropoff_longitude'] > 180)].shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.773571Z","iopub.execute_input":"2022-08-09T18:35:46.774374Z","iopub.status.idle":"2022-08-09T18:35:46.784970Z","shell.execute_reply.started":"2022-08-09T18:35:46.774329Z","shell.execute_reply":"2022-08-09T18:35:46.783732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# drop these outliers\ntrain_df = train_df.loc[(train_df['dropoff_longitude'] >= -180) & (train_df['dropoff_longitude'] <= 180)]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.786845Z","iopub.execute_input":"2022-08-09T18:35:46.787241Z","iopub.status.idle":"2022-08-09T18:35:46.827513Z","shell.execute_reply.started":"2022-08-09T18:35:46.787206Z","shell.execute_reply":"2022-08-09T18:35:46.826563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### *dropoff_latitude*","metadata":{}},{"cell_type":"code","source":"# check number of observations with latitude not in [-90,90]\ntrain_df.loc[(train_df['dropoff_latitude'] < -90) | (train_df['dropoff_latitude'] > 90)].shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.829197Z","iopub.execute_input":"2022-08-09T18:35:46.829724Z","iopub.status.idle":"2022-08-09T18:35:46.842852Z","shell.execute_reply.started":"2022-08-09T18:35:46.829665Z","shell.execute_reply":"2022-08-09T18:35:46.841897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# drop these outliers\ntrain_df = train_df.loc[(train_df['dropoff_latitude'] >= -90) & (train_df['dropoff_latitude'] <= 90)]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.844322Z","iopub.execute_input":"2022-08-09T18:35:46.844942Z","iopub.status.idle":"2022-08-09T18:35:46.885422Z","shell.execute_reply.started":"2022-08-09T18:35:46.844905Z","shell.execute_reply":"2022-08-09T18:35:46.884520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### *passenger_count*","metadata":{}},{"cell_type":"code","source":"# check number of observations with different number of passengers\ntrain_df['passenger_count'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.886908Z","iopub.execute_input":"2022-08-09T18:35:46.887553Z","iopub.status.idle":"2022-08-09T18:35:46.901837Z","shell.execute_reply.started":"2022-08-09T18:35:46.887515Z","shell.execute_reply":"2022-08-09T18:35:46.900660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# seems here we only have 1 outlier so lets drop them\ntrain_df = train_df.loc[train_df['passenger_count'] <= 7]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.903159Z","iopub.execute_input":"2022-08-09T18:35:46.904260Z","iopub.status.idle":"2022-08-09T18:35:46.945195Z","shell.execute_reply.started":"2022-08-09T18:35:46.904220Z","shell.execute_reply":"2022-08-09T18:35:46.944321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## More Cleaning","metadata":{}},{"cell_type":"markdown","source":"Now I want to check if there are any records with same pickup and dropoff coordinates.","metadata":{}},{"cell_type":"code","source":"train_df.loc[(train_df['pickup_longitude'] == train_df['dropoff_longitude']) &\n          (train_df['pickup_latitude'] == train_df['dropoff_latitude'])]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.946832Z","iopub.execute_input":"2022-08-09T18:35:46.947221Z","iopub.status.idle":"2022-08-09T18:35:46.973303Z","shell.execute_reply.started":"2022-08-09T18:35:46.947185Z","shell.execute_reply":"2022-08-09T18:35:46.972350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 28,723  observations which have exactly same pickup and dropoff coordinates; As there is no information about the path taken by the taxi, these can be round trips, so I don not think dropping these observations will be right choise.\n\nInstead, I want to check if there are observations with same coordinates and zero fare.","metadata":{}},{"cell_type":"code","source":"train_df.loc[(train_df['pickup_longitude'] == train_df['dropoff_longitude']) &\n          (train_df['pickup_latitude'] == train_df['dropoff_latitude']) &\n          (train_df['fare_amount'] == 0)]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.974913Z","iopub.execute_input":"2022-08-09T18:35:46.975537Z","iopub.status.idle":"2022-08-09T18:35:46.994048Z","shell.execute_reply.started":"2022-08-09T18:35:46.975499Z","shell.execute_reply":"2022-08-09T18:35:46.993041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have 4 observations which have same pickup and dropoff coordinates with zero fare amount (these observations does not make sense). Lets drop this records.","metadata":{}},{"cell_type":"code","source":"indexes = train_df.loc[(train_df['pickup_longitude'] == train_df['dropoff_longitude']) &\n          (train_df['pickup_latitude'] == train_df['dropoff_latitude']) &\n          (train_df['fare_amount'] == 0)].index\n\ntrain_df.drop(index=indexes, axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:46.995794Z","iopub.execute_input":"2022-08-09T18:35:46.996434Z","iopub.status.idle":"2022-08-09T18:35:47.103152Z","shell.execute_reply.started":"2022-08-09T18:35:46.996397Z","shell.execute_reply":"2022-08-09T18:35:47.102167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Also, I checked coordinates of New York City in Google map, and it shows that longitude and latitude of this city are:\n* **Latitude** _  40.730610\n* **Longitude** _ -73.935242\n\nlets check pickup/dropoff boundary coordinates in train/test sets:","metadata":{}},{"cell_type":"code","source":"train_df[['pickup_longitude', 'pickup_latitude', 'dropoff_longitude', 'dropoff_latitude']].describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:47.104870Z","iopub.execute_input":"2022-08-09T18:35:47.105458Z","iopub.status.idle":"2022-08-09T18:35:47.258388Z","shell.execute_reply.started":"2022-08-09T18:35:47.105417Z","shell.execute_reply":"2022-08-09T18:35:47.257503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df[['pickup_longitude', 'pickup_latitude', 'dropoff_longitude', 'dropoff_latitude']].describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:47.260747Z","iopub.execute_input":"2022-08-09T18:35:47.261508Z","iopub.status.idle":"2022-08-09T18:35:47.286293Z","shell.execute_reply.started":"2022-08-09T18:35:47.261421Z","shell.execute_reply":"2022-08-09T18:35:47.285359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is big difference between train and test data coordinates variation. Lets define boundraries to restrict our train dataset's pickup/dropoff coordinates.","metadata":{}},{"cell_type":"code","source":"long_bound = (\nmin(test_df['pickup_longitude'].min(), test_df['dropoff_longitude'].min()),\nmax(test_df['pickup_longitude'].max(), test_df['dropoff_longitude'].max())\n)\nprint('Longitude bound in test set:')\nlong_bound","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:47.294056Z","iopub.execute_input":"2022-08-09T18:35:47.294609Z","iopub.status.idle":"2022-08-09T18:35:47.305329Z","shell.execute_reply.started":"2022-08-09T18:35:47.294579Z","shell.execute_reply":"2022-08-09T18:35:47.303807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lat_bound = (\nmin(test_df['pickup_latitude'].min(), test_df['dropoff_latitude'].min()),\nmax(test_df['pickup_latitude'].max(), test_df['dropoff_latitude'].max())\n)\nprint('latitude bound in test set:')\nlat_bound","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:47.306928Z","iopub.execute_input":"2022-08-09T18:35:47.307748Z","iopub.status.idle":"2022-08-09T18:35:47.318781Z","shell.execute_reply.started":"2022-08-09T18:35:47.307667Z","shell.execute_reply":"2022-08-09T18:35:47.317273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.loc[~((train_df['pickup_longitude'] >= long_bound[0]) & (train_df['pickup_longitude']<=long_bound[1])) &\n             ((train_df['pickup_latitude'] >= lat_bound[0]) & (train_df['pickup_latitude']<=lat_bound[1])) &\n             ((train_df['dropoff_longitude'] >= long_bound[0]) & (train_df['dropoff_longitude']<=long_bound[1])) &\n             ((train_df['dropoff_latitude'] >= lat_bound[0]) & (train_df['dropoff_latitude']<=lat_bound[1]))].shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:47.320045Z","iopub.execute_input":"2022-08-09T18:35:47.321230Z","iopub.status.idle":"2022-08-09T18:35:47.341631Z","shell.execute_reply.started":"2022-08-09T18:35:47.321195Z","shell.execute_reply":"2022-08-09T18:35:47.340617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = train_df.loc[((train_df['pickup_longitude'] >= long_bound[0]) & (train_df['pickup_longitude']<=long_bound[1])) &\n             ((train_df['pickup_latitude'] >= lat_bound[0]) & (train_df['pickup_latitude']<=lat_bound[1])) &\n             ((train_df['dropoff_longitude'] >= long_bound[0]) & (train_df['dropoff_longitude']<=long_bound[1])) &\n             ((train_df['dropoff_latitude'] >= lat_bound[0]) & (train_df['dropoff_latitude']<=lat_bound[1]))]\nprint(train_df.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:47.343079Z","iopub.execute_input":"2022-08-09T18:35:47.343429Z","iopub.status.idle":"2022-08-09T18:35:47.393075Z","shell.execute_reply.started":"2022-08-09T18:35:47.343395Z","shell.execute_reply":"2022-08-09T18:35:47.392023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['pickup_longitude', 'pickup_latitude', 'dropoff_longitude', 'dropoff_latitude']].describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:47.396107Z","iopub.execute_input":"2022-08-09T18:35:47.396380Z","iopub.status.idle":"2022-08-09T18:35:47.547201Z","shell.execute_reply.started":"2022-08-09T18:35:47.396353Z","shell.execute_reply":"2022-08-09T18:35:47.546098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now our train dataset has little variation like test set.\n\nLast thing I want to check is minimum taxi fare amount in New York. According to google, it's about 2.5 dollar, so lets see how many observations we have with less than 2.5 dollar fare amount.","metadata":{}},{"cell_type":"code","source":"train_df.loc[train_df['fare_amount']<2.5].shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:47.549044Z","iopub.execute_input":"2022-08-09T18:35:47.549408Z","iopub.status.idle":"2022-08-09T18:35:47.559300Z","shell.execute_reply.started":"2022-08-09T18:35:47.549371Z","shell.execute_reply":"2022-08-09T18:35:47.558176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As number of observations are too small (31), probably it will be good choice to drop them.","metadata":{}},{"cell_type":"code","source":"train_df = train_df.loc[train_df['fare_amount']>=2.5]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:47.560928Z","iopub.execute_input":"2022-08-09T18:35:47.561489Z","iopub.status.idle":"2022-08-09T18:35:47.601511Z","shell.execute_reply.started":"2022-08-09T18:35:47.561453Z","shell.execute_reply":"2022-08-09T18:35:47.600542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_after = train_df.shape[0]\n\nprint(f'Number of observations before data cleaning: {n_before};')\nprint(f'Number of observations after data cleaning: {n_after};')\nprint(f'We have dropped {round((n_before-n_after)/n_after * 100, 3)} % of the observations.')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:47.602908Z","iopub.execute_input":"2022-08-09T18:35:47.603350Z","iopub.status.idle":"2022-08-09T18:35:47.610190Z","shell.execute_reply.started":"2022-08-09T18:35:47.603313Z","shell.execute_reply":"2022-08-09T18:35:47.609004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"To apply feature engineering both train and test datasets, I will concatenate them into one dataframe.","metadata":{}},{"cell_type":"code","source":"# lets create column which will help us to differenciate train and test datasets\ntrain_df['train'] = 'train'\ntest_df['train'] = 'test'\n\n# concatenate these 2 datasets\ndata = pd.concat([train_df,test_df],axis=0)\ndata","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:47.611786Z","iopub.execute_input":"2022-08-09T18:35:47.612523Z","iopub.status.idle":"2022-08-09T18:35:47.671712Z","shell.execute_reply.started":"2022-08-09T18:35:47.612485Z","shell.execute_reply":"2022-08-09T18:35:47.670742Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Firstly, I want to extract informationa about Year/Month/Day/Weekday/Hour from ***pickup_datetime*** column. I want to extract columns:\n* **Year** _ the year of the datetime.\n* **Month** _ the month number of datetime.\n* **Day** _ the day of the datetime.\n* **Dayofweek** _ the day of the week.\n* **Hour** _ the hours of the datetime","metadata":{}},{"cell_type":"code","source":"# extract year from pickup_datetime columns\ndata['year'] = pd.DatetimeIndex(data['pickup_datetime']).year\n\n# check if there is any issue\nprint('Year')\nprint(data['year'].value_counts().sort_index())\n\n# extract month from pickup_datetime columns\ndata['month'] = pd.DatetimeIndex(data['pickup_datetime']).month\n\nprint(' ')\nprint('Month')\nprint(data['month'].value_counts().sort_index())\n\n# extract day from pickup_datetime columns\ndata['day'] = pd.DatetimeIndex(data['pickup_datetime']).day\n\nprint(' ')\nprint('day')\nprint(data['day'].value_counts().sort_index())\n\n# extract weekday from pickup_datetime columns\ndata['dayofweek'] = pd.DatetimeIndex(data['pickup_datetime']).dayofweek\n\nprint(' ')\nprint('dayofweek')\nprint(data['dayofweek'].value_counts().sort_index())\n\n# extract weekday from pickup_datetime columns\ndata['hour'] = pd.DatetimeIndex(data['pickup_datetime']).hour\n\nprint(' ')\nprint('hour')\nprint(data['hour'].value_counts().sort_index())","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:47.673131Z","iopub.execute_input":"2022-08-09T18:35:47.673934Z","iopub.status.idle":"2022-08-09T18:35:48.242790Z","shell.execute_reply.started":"2022-08-09T18:35:47.673886Z","shell.execute_reply":"2022-08-09T18:35:48.241748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I would also like to add information about the season and period of day.","metadata":{}},{"cell_type":"code","source":"# function to create season column\ndef decodeseasons(df):\n    df['season'] = np.nan\n    \n    df.loc[df['month'].isin([1, 2, 12]), 'season'] = 1 # winter\n    df.loc[df['month'].isin([3, 4, 5]), 'season'] =  2 # spring\n    df.loc[df['month'].isin([6, 7, 8]), 'season'] =  3 # summer\n    df.loc[df['month'].isin([9, 10, 11]), 'season'] =4 # autumn\n    \n    return df\n\n# function to create periods column\ndef decodeperiods(df):\n    df['periods'] = np.nan\n    df.loc[df['hour'].isin([0, 1, 2, 3, 4, 5, 22, 23]), 'periods'] = 1 # night\n    df.loc[df['hour'].isin([6, 7, 8, 9, 10, 11]), 'periods']        = 2 # morning\n    df.loc[df['hour'].isin([12, 13, 14, 15, 16]), 'periods']        = 3 # afternoon\n    df.loc[df['hour'].isin([17, 18, 19, 20, 21]), 'periods']        = 4 # evening\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:48.244427Z","iopub.execute_input":"2022-08-09T18:35:48.245068Z","iopub.status.idle":"2022-08-09T18:35:48.255006Z","shell.execute_reply.started":"2022-08-09T18:35:48.245028Z","shell.execute_reply":"2022-08-09T18:35:48.253932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = decodeseasons(data)\ndata = decodeperiods(data)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:48.258373Z","iopub.execute_input":"2022-08-09T18:35:48.259001Z","iopub.status.idle":"2022-08-09T18:35:48.483924Z","shell.execute_reply.started":"2022-08-09T18:35:48.258973Z","shell.execute_reply":"2022-08-09T18:35:48.482594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now I want to calculate distance between pickup point and dropoff point because distance most probably will be key factor to determine fare. I want to calculate 2 types of distance:\n* **Euclidean distance**;\n* **Manhattan distance**.","metadata":{}},{"cell_type":"code","source":"# Euclidean distance\ndata['euclidean'] = np.sqrt((data['dropoff_longitude']-data['pickup_longitude'])**2 +\n                             (data['dropoff_latitude']-data['pickup_latitude'])**2)\n\n# Manhattan distance\ndata['manhattan'] = abs(data['dropoff_longitude']-data['pickup_longitude']) + abs(\n    data['dropoff_latitude']-data['pickup_latitude'])","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:48.485664Z","iopub.execute_input":"2022-08-09T18:35:48.486332Z","iopub.status.idle":"2022-08-09T18:35:48.533524Z","shell.execute_reply.started":"2022-08-09T18:35:48.486291Z","shell.execute_reply":"2022-08-09T18:35:48.530380Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After calculating distances, I want to create dummy variable to assign round trips (as I mentiond before, I did not drop observations which had same pickup/dropoff coordinates because these can be round trips records). I will assign round trips with 1 and other with 0.","metadata":{}},{"cell_type":"code","source":"data['round_tr'] = 0\ndata.loc[data['manhattan'] ==0, 'round_tr'] = 1\ndata['round_tr'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:48.534660Z","iopub.execute_input":"2022-08-09T18:35:48.535063Z","iopub.status.idle":"2022-08-09T18:35:48.571789Z","shell.execute_reply.started":"2022-08-09T18:35:48.535025Z","shell.execute_reply":"2022-08-09T18:35:48.569592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lastly, there is one thing I want to do: it is quite common and logical that Taxi Fare is especially high around airports. To capture this relationship, I want to create column which shows distance (manhattan distance) between airport and pickup location. According to google, there are 3 main airport in New York:\n* **John F. Kennedy International Airport (JFK)** _ (40.641766, -73.780968);\n* **LaGuardia Airport (LGA)** _ (40.7730, -73.8702);\n* **Newark Liberty International Airport (EWR)** _ (40.7026, -74.1878);","metadata":{}},{"cell_type":"code","source":"JFK_coor = (40.641766, -73.780968)\nLGA_coor = (40.7730, -73.8702)\nEWR_coor = (40.7026, -74.1878)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:48.574716Z","iopub.execute_input":"2022-08-09T18:35:48.576025Z","iopub.status.idle":"2022-08-09T18:35:48.585087Z","shell.execute_reply.started":"2022-08-09T18:35:48.575805Z","shell.execute_reply":"2022-08-09T18:35:48.583851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data['JFK'] = abs(data['dropoff_longitude']-JFK_coor[1]) + abs(\n              data['dropoff_latitude']-JFK_coor[0])\ndata['LGA'] = abs(data['dropoff_longitude']-LGA_coor[1]) + abs(\n              data['dropoff_latitude']-LGA_coor[0])\ndata['EWR'] = abs(data['dropoff_longitude']-EWR_coor[1]) + abs(\n              data['dropoff_latitude']-EWR_coor[0])","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:48.588519Z","iopub.execute_input":"2022-08-09T18:35:48.589629Z","iopub.status.idle":"2022-08-09T18:35:48.640752Z","shell.execute_reply.started":"2022-08-09T18:35:48.589593Z","shell.execute_reply":"2022-08-09T18:35:48.639758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data[['JFK', 'LGA', 'EWR']].describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:48.645341Z","iopub.execute_input":"2022-08-09T18:35:48.647034Z","iopub.status.idle":"2022-08-09T18:35:48.991663Z","shell.execute_reply.started":"2022-08-09T18:35:48.646997Z","shell.execute_reply":"2022-08-09T18:35:48.990751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First part of feature engineering is done, now I want to set optimal data types to the new variables and explore dataset by visualization and lets see if we find interesting relationships which we will later use to build new variables.","metadata":{}},{"cell_type":"code","source":"data.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:48.992886Z","iopub.execute_input":"2022-08-09T18:35:48.993434Z","iopub.status.idle":"2022-08-09T18:35:49.193270Z","shell.execute_reply.started":"2022-08-09T18:35:48.993397Z","shell.execute_reply":"2022-08-09T18:35:49.192323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data[['passenger_count', 'month', 'day', 'dayofweek', 'hour', 'season', 'periods', 'round_tr']] = data[\n    ['passenger_count', 'month', 'day', 'dayofweek', 'hour', 'season', 'periods', 'round_tr']].astype('uint8')\ndata[['year']] = data[['year']].astype('uint16')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:49.197825Z","iopub.execute_input":"2022-08-09T18:35:49.198508Z","iopub.status.idle":"2022-08-09T18:35:49.326178Z","shell.execute_reply.started":"2022-08-09T18:35:49.198471Z","shell.execute_reply":"2022-08-09T18:35:49.325224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:49.330544Z","iopub.execute_input":"2022-08-09T18:35:49.333085Z","iopub.status.idle":"2022-08-09T18:35:49.518544Z","shell.execute_reply.started":"2022-08-09T18:35:49.333047Z","shell.execute_reply":"2022-08-09T18:35:49.517495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Firstly lets separate train/test parts\ntrain_df = data.loc[data['train'] == 'train']\ntrain_df.drop(['key', 'train'], axis = 1, inplace = True)\nprint(train_df.shape)\nprint(train_df.columns)\n\ntest_df = data.loc[data['train'] == 'test']\ntest_df.drop(['fare_amount', 'train'], axis = 1, inplace = True)\nprint(test_df.shape)\nprint(test_df.columns)\n\ndel data","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:49.522503Z","iopub.execute_input":"2022-08-09T18:35:49.523919Z","iopub.status.idle":"2022-08-09T18:35:49.893866Z","shell.execute_reply.started":"2022-08-09T18:35:49.523883Z","shell.execute_reply":"2022-08-09T18:35:49.892727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:49.895461Z","iopub.execute_input":"2022-08-09T18:35:49.895824Z","iopub.status.idle":"2022-08-09T18:35:49.940168Z","shell.execute_reply.started":"2022-08-09T18:35:49.895787Z","shell.execute_reply":"2022-08-09T18:35:49.939072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Analyisis/Visualisation","metadata":{}},{"cell_type":"markdown","source":"In this section, I want to check and answer several relationships and questions.\nParticularly, I want to:\n* Check distribution of our target variable;\n* Compare average fare by time periods (year, month, day, hour, dayofweek);\n* Check categorical time groups (night/day, seasons, first/second halft of the month;\n* Check relationships between fare and distance;\n* Check how number of passengers affect the fare;\n* Draw longitude/latitude scatterplot;\n* Draw correlation matrix & heatmap.","metadata":{}},{"cell_type":"markdown","source":"#### Distribution of ***fare_amount***","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(16,7))\n\nsns.kdeplot(x=train_df['fare_amount'], shade=True, color='blue')\nplt.title('Distribution of fare_amount')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:49.941790Z","iopub.execute_input":"2022-08-09T18:35:49.942159Z","iopub.status.idle":"2022-08-09T18:35:53.656536Z","shell.execute_reply.started":"2022-08-09T18:35:49.942122Z","shell.execute_reply":"2022-08-09T18:35:53.655570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we see, fare_amount has skewed distribution, so probably we will have to use more non-linear models to get accurate predictions.","metadata":{}},{"cell_type":"markdown","source":"#### Time categories","metadata":{}},{"cell_type":"code","source":"ls = ['year', 'month', 'dayofweek', 'day', 'hour']\nfor per in ls:\n    table = pd.pivot_table(train_df, values='pickup_datetime', index=per, aggfunc='count')\n    \n    plt.figure(figsize=(18,5))\n    \n    plt.subplot(1,2,1)\n    sns.barplot(x=train_df[per], y=train_df['fare_amount'])\n    plt.title(f'Average taxi fare by {per}')\n    plt.ylabel('Avegare Fare Amount');\n    \n    \n    plt.subplot(1,2,2)\n    sns.barplot(x=table.index, y=table.values.reshape(-1));\n    plt.title(f'Number of trips by {per}')\n    plt.ylabel('Number of trips');","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:35:53.657992Z","iopub.execute_input":"2022-08-09T18:35:53.659039Z","iopub.status.idle":"2022-08-09T18:36:49.685220Z","shell.execute_reply.started":"2022-08-09T18:35:53.658998Z","shell.execute_reply":"2022-08-09T18:36:49.684326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Conclusions:\n* Taxi fare increases year by year;\n* Taxi fare does not vary much by month, even through fare was slightly lower in January, February and July (number of tripes were relatively lower in this time period too);\n* Taxi fare also does not vary much by dayofweek and day, but fare is slightly lower on saturdays. Number of trips are lower in 31th day of the month as expected: only 7 monthes have 31 days;\n* Taxi fare is relatively high between 3-6 and 13-16 hours. Number of trips are much lower during 3-6 hours and relatively low during 14-17 hours.","metadata":{}},{"cell_type":"markdown","source":"##### Season/Period\n\n","metadata":{}},{"cell_type":"code","source":"ls_2 = ['season', 'periods']\nls_3 = [['winter', 'spring', 'summer', 'autumn'], ['night', 'morning', 'afternoon', 'evening']]\nfor per in ls_2:\n    table = pd.pivot_table(train_df, values='pickup_datetime', index=per, aggfunc='count')\n    \n    plt.figure(figsize=(12,3))\n    \n    plt.subplot(1,2,1)\n    g = sns.barplot(x=train_df[per], y=train_df['fare_amount'])\n    g.set_xticklabels(ls_3[ls_2.index(per)])\n    g.set(xlabel=None)\n    plt.title(f'Average taxi fare by {per}')\n    plt.ylabel('Avegare Fare Amount');\n\n\n    plt.subplot(1,2,2)\n    p = sns.barplot(x=table.index, y=table.values.reshape(-1))\n    p.set_xticklabels(ls_3[ls_2.index(per)])\n    p.set(xlabel=None)\n    plt.title(f'Number of trips by {per}')\n    plt.ylabel('Number of trips');","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:36:49.686511Z","iopub.execute_input":"2022-08-09T18:36:49.687200Z","iopub.status.idle":"2022-08-09T18:37:12.281871Z","shell.execute_reply.started":"2022-08-09T18:36:49.687160Z","shell.execute_reply":"2022-08-09T18:37:12.280973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Counclusions:\n* Tax fare does not vary much by season;\n* Number of trips are slightly high during spring (Probably because it is the most rainy season in New York);\n* Tax fare is slightly low in the evening while number of trips are relatively high.","metadata":{}},{"cell_type":"markdown","source":"#### Distance Influence on Fare","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(16,7))\n\nplt.subplot(1,2,1)\nsns.scatterplot(x=train_df['euclidean'], y=train_df['fare_amount'])\nplt.title('Fare v euclidean distance');\n\n\nplt.subplot(1,2,2)\nsns.scatterplot(x=train_df['manhattan'], y=train_df['fare_amount'])\nplt.title('Fare v manhattan distance');","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:37:12.283725Z","iopub.execute_input":"2022-08-09T18:37:12.284396Z","iopub.status.idle":"2022-08-09T18:37:15.837123Z","shell.execute_reply.started":"2022-08-09T18:37:12.284354Z","shell.execute_reply":"2022-08-09T18:37:15.835731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Scatterplots show that relationship between fare_amount and distance (both type of distance) are positive but there are some groups that do not follow this relationship. Lets try to create some groups by distance:","metadata":{}},{"cell_type":"code","source":"print('Distance _ [0]')\nprint('Average fare: '+str(train_df.loc[(train_df['manhattan'] == 0),'fare_amount'].mean()))\nprint('Distance _ (0, 0.5]')\nprint('Average fare: '+str(train_df.loc[(train_df['manhattan'] > 0) & (train_df['manhattan'] <= 0.5), 'fare_amount'].mean()))\nprint('Distance _ (0.5, 1.15]')\nprint('Average fare: '+str(train_df.loc[(train_df['manhattan'] > 0.6 ) & (train_df['manhattan'] <= 1.15), 'fare_amount'].mean()))\nprint('Distance _ (0.5, 1.15]')\nprint('Average fare: '+str(train_df.loc[(train_df['manhattan'] >= 1.15 ), 'fare_amount'].mean()))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:37:15.838986Z","iopub.execute_input":"2022-08-09T18:37:15.839403Z","iopub.status.idle":"2022-08-09T18:37:15.874023Z","shell.execute_reply.started":"2022-08-09T18:37:15.839364Z","shell.execute_reply":"2022-08-09T18:37:15.872845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets compare train vs test distance distributions:","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(16,5))\n\nplt.subplot(1,2,1)\nsns.kdeplot(x=train_df['manhattan'])\nplt.title('train Data distance (manhattan) Distribution');\n\n\nplt.subplot(1,2,2)\nsns.kdeplot(x=test_df['manhattan'])\nplt.title('Test Data distance (manhattan) Distribution');","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:37:15.875397Z","iopub.execute_input":"2022-08-09T18:37:15.876277Z","iopub.status.idle":"2022-08-09T18:37:19.635535Z","shell.execute_reply.started":"2022-08-09T18:37:15.876246Z","shell.execute_reply":"2022-08-09T18:37:19.633195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Even through extreme values are different, distributions are quite similar. lets create these categories in our train/test data (we already have distance equals 0 group, round_tr, so we should create 3 groups).","metadata":{}},{"cell_type":"code","source":"train_df['distgr_2'] = 0\ntest_df['distgr_2'] = 0\ntrain_df.loc[train_df['manhattan'].between(0, 0.5, inclusive='right'), 'distgr_2'] = 1\ntest_df.loc[test_df['manhattan'].between(0, 0.5, inclusive='right'), 'distgr_2'] = 1\n\ntrain_df['distgr_3'] = 0\ntest_df['distgr_3'] = 0\ntrain_df.loc[train_df['manhattan'].between(0.5, 1.15, inclusive='right'), 'distgr_3'] = 1\ntest_df.loc[test_df['manhattan'].between(0.5, 1.15, inclusive='right'), 'distgr_3'] = 1\n\ntrain_df['distgr_4'] = 0\ntest_df['distgr_4'] = 0\ntrain_df.loc[train_df['manhattan'] > 1.15, 'distgr_4'] = 1\ntest_df.loc[test_df['manhattan'] > 1.15, 'distgr_4'] = 1\n\ntrain_df[['distgr_2', 'distgr_3', 'distgr_4']] = train_df[['distgr_2', 'distgr_3', 'distgr_4']].astype('uint8')\ntest_df[['distgr_2', 'distgr_3', 'distgr_4']] = test_df[['distgr_2', 'distgr_3', 'distgr_4']].astype('uint8')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:37:19.637539Z","iopub.execute_input":"2022-08-09T18:37:19.638321Z","iopub.status.idle":"2022-08-09T18:37:19.704009Z","shell.execute_reply.started":"2022-08-09T18:37:19.638280Z","shell.execute_reply":"2022-08-09T18:37:19.702937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Number of passengers influence on fare","metadata":{}},{"cell_type":"code","source":"train_df.columns","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:37:19.705430Z","iopub.execute_input":"2022-08-09T18:37:19.705929Z","iopub.status.idle":"2022-08-09T18:37:19.712934Z","shell.execute_reply.started":"2022-08-09T18:37:19.705885Z","shell.execute_reply":"2022-08-09T18:37:19.711957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(16,7))\n\nplt.subplot(1,2,1)\nsns.barplot(x=train_df['passenger_count'], y=train_df['fare_amount'])\nplt.title('Average taxi fare by number of passenger')\nplt.ylabel('Avegare Fare Amount');\n\n\nplt.subplot(1,2,2)\ntable = pd.pivot_table(train_df, values='pickup_datetime', index='passenger_count', aggfunc='count')\nsns.barplot(x=table.index, y=table.values.reshape(-1))\nplt.title('Number of trips by number of passenger');","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:37:19.714503Z","iopub.execute_input":"2022-08-09T18:37:19.715186Z","iopub.status.idle":"2022-08-09T18:37:31.973727Z","shell.execute_reply.started":"2022-08-09T18:37:19.715149Z","shell.execute_reply":"2022-08-09T18:37:31.972776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From there barplots we can see that single passengers are the most frequent travellers. To further explore relationship between fare amount and number of passengers, lets draw violinplot.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10,10))\nsns.violinplot(x=\"passenger_count\", y=\"fare_amount\", data=train_df)\nplt.title('Fare distribution by number of passengers')\nplt.xlabel('Number of passengers')\nplt.ylabel('Fare amount');","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:37:31.975399Z","iopub.execute_input":"2022-08-09T18:37:31.976101Z","iopub.status.idle":"2022-08-09T18:37:34.346505Z","shell.execute_reply.started":"2022-08-09T18:37:31.976060Z","shell.execute_reply":"2022-08-09T18:37:34.345616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Violinplot shows that extreme records are associated with trips with one passengers.","metadata":{}},{"cell_type":"markdown","source":"#### Location influence","metadata":{}},{"cell_type":"markdown","source":"Firstly lets look what we can see on simple scatterplot.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15,8))\nsns.scatterplot(x=train_df['pickup_longitude'], y=train_df['pickup_latitude']);","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:37:34.347973Z","iopub.execute_input":"2022-08-09T18:37:34.349950Z","iopub.status.idle":"2022-08-09T18:37:36.190545Z","shell.execute_reply.started":"2022-08-09T18:37:34.349908Z","shell.execute_reply":"2022-08-09T18:37:36.189479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Not much we can see here, lets try other ones:","metadata":{}},{"cell_type":"code","source":"# pickup coordinates\ntrain_df.plot(kind='scatter', x='pickup_longitude', y='pickup_latitude',\n                color='blue', \n                s=.05, alpha=0.8)\nplt.title('Pickups')\nplt.ylim([40.568973, 40.950000])\nplt.xlim([-74.200000, -73.600000]);","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:37:36.192223Z","iopub.execute_input":"2022-08-09T18:37:36.192570Z","iopub.status.idle":"2022-08-09T18:37:40.277505Z","shell.execute_reply.started":"2022-08-09T18:37:36.192536Z","shell.execute_reply":"2022-08-09T18:37:40.276332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# dropoff coordinates\ntrain_df.plot(kind='scatter', x='dropoff_longitude', y='dropoff_latitude',\n                color='blue', \n                s=.05, alpha=0.8)\nplt.title('Dropoffs')\nplt.ylim([40.568973, 40.950000])\nplt.xlim([-74.200000, -73.600000]);","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:37:40.282416Z","iopub.execute_input":"2022-08-09T18:37:40.287400Z","iopub.status.idle":"2022-08-09T18:37:44.388471Z","shell.execute_reply.started":"2022-08-09T18:37:40.287359Z","shell.execute_reply":"2022-08-09T18:37:44.387576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Seems there are several distincts by locations. It can be possible that fare amount varies by distinct. Lets try clustering using K-means.","metadata":{}},{"cell_type":"markdown","source":"##### K-Means","metadata":{}},{"cell_type":"code","source":"dat_1 = train_df[['pickup_latitude', 'pickup_longitude']].values[:100000,:]\ndat_2 = train_df[['dropoff_latitude', 'dropoff_longitude']].values[:100000,:]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:37:44.389932Z","iopub.execute_input":"2022-08-09T18:37:44.391050Z","iopub.status.idle":"2022-08-09T18:37:44.409711Z","shell.execute_reply.started":"2022-08-09T18:37:44.390965Z","shell.execute_reply":"2022-08-09T18:37:44.408882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Pickup Coordinates:","metadata":{}},{"cell_type":"code","source":"inertia = []\nfor n in range(1 , 11):\n    algorithm = (KMeans(n_clusters = n ,init='k-means++', n_init = 10, max_iter=300, \n                        tol=0.0001, random_state=4) )\n    algorithm.fit(dat_1)\n    inertia.append(algorithm.inertia_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:37:44.411169Z","iopub.execute_input":"2022-08-09T18:37:44.411539Z","iopub.status.idle":"2022-08-09T18:37:59.443328Z","shell.execute_reply.started":"2022-08-09T18:37:44.411504Z","shell.execute_reply":"2022-08-09T18:37:59.442494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(1 , figsize = (15 ,6))\nplt.plot(np.arange(1 , 11) , inertia , 'o')\nplt.plot(np.arange(1 , 11) , inertia , '-' , alpha = 0.5)\nplt.xlabel('Number of Clusters') , plt.ylabel('Inertia')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:37:59.447171Z","iopub.execute_input":"2022-08-09T18:37:59.448716Z","iopub.status.idle":"2022-08-09T18:37:59.737529Z","shell.execute_reply.started":"2022-08-09T18:37:59.448653Z","shell.execute_reply":"2022-08-09T18:37:59.736600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# by elbow method, 6 should be good number of clusters\nalgorithm = (KMeans(n_clusters = 6 ,init='k-means++', n_init = 10 ,max_iter=300, \n                        tol=0.0001,  random_state= 4) )\nalgorithm.fit(dat_1)\nlabels1 = algorithm.labels_\ncentroids1 = algorithm.cluster_centers_\ny_kmeans = algorithm.predict(dat_1)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:37:59.741788Z","iopub.execute_input":"2022-08-09T18:37:59.744043Z","iopub.status.idle":"2022-08-09T18:38:01.917818Z","shell.execute_reply.started":"2022-08-09T18:37:59.744001Z","shell.execute_reply":"2022-08-09T18:38:01.916944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10,7))\nplt.scatter(dat_1[:, 1], dat_1[:, 0], c=y_kmeans, s=0.01, cmap='viridis')\nplt.ylim([40.568973, 40.950000])\nplt.xlim([-74.200000, -73.600000]);\n\ncenters = algorithm.cluster_centers_\nplt.scatter(centers[:, 1], centers[:, 0], c='red', s=20, alpha=0.8);","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:01.921806Z","iopub.execute_input":"2022-08-09T18:38:01.923642Z","iopub.status.idle":"2022-08-09T18:38:03.326524Z","shell.execute_reply.started":"2022-08-09T18:38:01.923608Z","shell.execute_reply":"2022-08-09T18:38:03.325611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:03.328197Z","iopub.execute_input":"2022-08-09T18:38:03.329175Z","iopub.status.idle":"2022-08-09T18:38:03.354166Z","shell.execute_reply.started":"2022-08-09T18:38:03.329136Z","shell.execute_reply":"2022-08-09T18:38:03.353044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['pick_cls'] = algorithm.predict(train_df[['pickup_latitude', 'pickup_longitude']].values)\ntest_df['pick_cls'] = algorithm.predict(test_df[['pickup_latitude', 'pickup_longitude']].values)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:03.355568Z","iopub.execute_input":"2022-08-09T18:38:03.356014Z","iopub.status.idle":"2022-08-09T18:38:03.415155Z","shell.execute_reply.started":"2022-08-09T18:38:03.355978Z","shell.execute_reply":"2022-08-09T18:38:03.414397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Dropoff Coordinates","metadata":{}},{"cell_type":"code","source":"inertia = []\nfor n in range(1 , 11):\n    algorithm = (KMeans(n_clusters = n ,init='k-means++', n_init = 10, max_iter=300, \n                        tol=0.0001, random_state=4) )\n    algorithm.fit(dat_2)\n    inertia.append(algorithm.inertia_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:03.418735Z","iopub.execute_input":"2022-08-09T18:38:03.420790Z","iopub.status.idle":"2022-08-09T18:38:20.397293Z","shell.execute_reply.started":"2022-08-09T18:38:03.420755Z","shell.execute_reply":"2022-08-09T18:38:20.396459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(1 , figsize = (15 ,6))\nplt.plot(np.arange(1 , 11) , inertia , 'o')\nplt.plot(np.arange(1 , 11) , inertia , '-' , alpha = 0.5)\nplt.xlabel('Number of Clusters') , plt.ylabel('Inertia')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:20.401243Z","iopub.execute_input":"2022-08-09T18:38:20.402080Z","iopub.status.idle":"2022-08-09T18:38:20.636549Z","shell.execute_reply.started":"2022-08-09T18:38:20.402045Z","shell.execute_reply":"2022-08-09T18:38:20.635598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# by elbow method, 3 should be good number of clusters\nalgorithm = (KMeans(n_clusters = 3 ,init='k-means++', n_init = 10 ,max_iter=300, \n                        tol=0.0001,  random_state= 4) )\nalgorithm.fit(dat_2)\nlabels1 = algorithm.labels_\ncentroids1 = algorithm.cluster_centers_\ny_kmeans = algorithm.predict(dat_2)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:20.637978Z","iopub.execute_input":"2022-08-09T18:38:20.638332Z","iopub.status.idle":"2022-08-09T18:38:21.852905Z","shell.execute_reply.started":"2022-08-09T18:38:20.638304Z","shell.execute_reply":"2022-08-09T18:38:21.851490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10,7))\nplt.scatter(dat_2[:, 1], dat_2[:, 0], c=y_kmeans, s=0.01, cmap='viridis')\nplt.ylim([40.568973, 40.950000])\nplt.xlim([-74.200000, -73.600000]);\n\ncenters = algorithm.cluster_centers_\nplt.scatter(centers[:, 1], centers[:, 0], c='black', s=20, alpha=0.5);","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:21.854128Z","iopub.execute_input":"2022-08-09T18:38:21.854486Z","iopub.status.idle":"2022-08-09T18:38:23.008576Z","shell.execute_reply.started":"2022-08-09T18:38:21.854448Z","shell.execute_reply":"2022-08-09T18:38:23.007639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['dropoff_cls'] = algorithm.predict(train_df[['dropoff_latitude', 'dropoff_longitude']])\ntest_df['dropoff_cls'] = algorithm.predict(test_df[['dropoff_latitude', 'dropoff_longitude']])","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:23.010590Z","iopub.execute_input":"2022-08-09T18:38:23.011333Z","iopub.status.idle":"2022-08-09T18:38:23.062320Z","shell.execute_reply.started":"2022-08-09T18:38:23.011262Z","shell.execute_reply":"2022-08-09T18:38:23.061528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets create corresponding dummy variables:","metadata":{}},{"cell_type":"code","source":"train_df = pd.get_dummies(data=train_df, columns=['pick_cls', 'dropoff_cls'])\ntest_df = pd.get_dummies(data=test_df, columns=['pick_cls', 'dropoff_cls'])","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:23.063879Z","iopub.execute_input":"2022-08-09T18:38:23.064544Z","iopub.status.idle":"2022-08-09T18:38:23.199776Z","shell.execute_reply.started":"2022-08-09T18:38:23.064507Z","shell.execute_reply":"2022-08-09T18:38:23.198541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.columns","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:23.201169Z","iopub.execute_input":"2022-08-09T18:38:23.202562Z","iopub.status.idle":"2022-08-09T18:38:23.211373Z","shell.execute_reply.started":"2022-08-09T18:38:23.202523Z","shell.execute_reply":"2022-08-09T18:38:23.209879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Correlation Matrix & Heatmap","metadata":{}},{"cell_type":"code","source":"# now lets check correltion heatmap\nmask = np.triu(np.ones_like(train_df.corr(), dtype=np.bool))\n\nplt.figure(figsize=(16, 8))\nheatmap = sns.heatmap(train_df.corr(), vmin=-1, vmax=1, mask=mask, cmap='PuBu')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:23.213056Z","iopub.execute_input":"2022-08-09T18:38:23.214349Z","iopub.status.idle":"2022-08-09T18:38:28.857019Z","shell.execute_reply.started":"2022-08-09T18:38:23.214298Z","shell.execute_reply":"2022-08-09T18:38:28.854021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mask = np.triu(np.ones_like(train_df.iloc[:,:17].corr(), dtype=np.bool))\n\nplt.figure(figsize=(16, 8))\nheatmap = sns.heatmap(train_df.iloc[:,:17].corr(), vmin=-1, vmax=1, mask=mask, annot=True, cmap='PuBu')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:28.858575Z","iopub.execute_input":"2022-08-09T18:38:28.859146Z","iopub.status.idle":"2022-08-09T18:38:31.420473Z","shell.execute_reply.started":"2022-08-09T18:38:28.859107Z","shell.execute_reply":"2022-08-09T18:38:31.419564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mask = np.triu(np.ones_like(pd.concat([train_df.iloc[:,0],train_df.iloc[:,17:]],axis = 1).corr(), dtype=np.bool))\n\nplt.figure(figsize=(16, 8))\nheatmap = sns.heatmap(pd.concat([train_df.iloc[:,0],train_df.iloc[:,17:]],axis = 1).corr(), vmin=-1, vmax=1,\n                      mask=mask, annot=True, cmap='PuBu')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:31.421872Z","iopub.execute_input":"2022-08-09T18:38:31.422814Z","iopub.status.idle":"2022-08-09T18:38:34.004271Z","shell.execute_reply.started":"2022-08-09T18:38:31.422777Z","shell.execute_reply":"2022-08-09T18:38:34.003316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Counclusions:\n* Both **Euclidean** and **Manhattan** distances have almost same correlation coefficient to **fare_amount**;\n* After distances, most correlated variables to **fare_amount** are: **pickup/dropoff longitudes**, **pickup/dropoff latitudes**, **pick_cls_2/4** and **dropoff_cls_2**;\n* Least correlated variables are: **day**, **dayofweek**, **round_tr**, **distgr_2**, **distgr_4**.","metadata":{}},{"cell_type":"markdown","source":"# Modeling","metadata":{}},{"cell_type":"markdown","source":"In these section, I will try to use Machine Learning to get predictions about test data. To capture linear relationships, I want to use **Linear Regression** and  to capture non-linear relationship I will use two, one ensemble learning and one gradient boosting, tree based algorithms: **Random Forest** and **XGBoost**. Lets see which one will perform better on this data.","metadata":{}},{"cell_type":"markdown","source":"### Generate X & Y  Train & Validation Matrices","metadata":{}},{"cell_type":"markdown","source":"For predictors, I will remove variable pickup_datetime,  and also remove Euclidean distance and leave only Manhattan distance (I think this distance measure fits better to our case).","metadata":{}},{"cell_type":"code","source":"print('Train_df Columns:')\nprint(list(train_df.columns))\nprint(' ')\nprint('Test_df Columns:')\nprint(list(test_df.columns))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:34.007378Z","iopub.execute_input":"2022-08-09T18:38:34.011734Z","iopub.status.idle":"2022-08-09T18:38:34.019591Z","shell.execute_reply.started":"2022-08-09T18:38:34.011691Z","shell.execute_reply":"2022-08-09T18:38:34.018661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictors = ['pickup_longitude', 'pickup_latitude',\n              'dropoff_longitude', 'dropoff_latitude', 'passenger_count', 'year', 'month',\n              'day', 'dayofweek', 'hour', 'season', 'periods', 'manhattan', 'round_tr',\n              'JFK', 'LGA', 'EWR', 'distgr_2', 'distgr_3', 'distgr_4', 'pick_cls_0', 'pick_cls_1',\n              'pick_cls_2', 'pick_cls_3', 'pick_cls_4', 'pick_cls_5', 'dropoff_cls_0', 'dropoff_cls_1',\n              'dropoff_cls_2']\nresponse = 'fare_amount'","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:34.021008Z","iopub.execute_input":"2022-08-09T18:38:34.022342Z","iopub.status.idle":"2022-08-09T18:38:34.035809Z","shell.execute_reply.started":"2022-08-09T18:38:34.022304Z","shell.execute_reply":"2022-08-09T18:38:34.034846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = train_df[predictors].values\ny = train_df[response].values\n\nprint(X.shape)\nprint(y.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:34.050782Z","iopub.execute_input":"2022-08-09T18:38:34.051331Z","iopub.status.idle":"2022-08-09T18:38:34.208990Z","shell.execute_reply.started":"2022-08-09T18:38:34.051296Z","shell.execute_reply":"2022-08-09T18:38:34.207971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.1, random_state=4)\nprint(X_train.shape)\nprint(y_train.shape)\nprint(X_test.shape)\nprint(y_test.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:34.212134Z","iopub.execute_input":"2022-08-09T18:38:34.214846Z","iopub.status.idle":"2022-08-09T18:38:34.632625Z","shell.execute_reply.started":"2022-08-09T18:38:34.214807Z","shell.execute_reply":"2022-08-09T18:38:34.631385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For linear regression, I also want to create scaled X and y matrices; also, I want to drop some variables to avoid so called *dummy variable trap*.","metadata":{}},{"cell_type":"code","source":"# I will drop round_tr, pick_cls_0 and dropoff_cls_0.\npredictors_ln = ['pickup_longitude', 'pickup_latitude',\n              'dropoff_longitude', 'dropoff_latitude', 'passenger_count', 'year', 'month',\n              'day', 'dayofweek', 'hour', 'season', 'periods', 'manhattan',\n              'JFK', 'LGA', 'EWR', 'distgr_2', 'distgr_3', 'distgr_4', 'pick_cls_1',\n              'pick_cls_2', 'pick_cls_3', 'pick_cls_4', 'pick_cls_5', 'dropoff_cls_1', 'dropoff_cls_2']\n\n\nX_ln = train_df[predictors_ln].values\nX_train_ln, X_test_ln, y_train_ln, y_test_ln = train_test_split(X_ln, y, test_size=0.1, random_state=4)\n\nscaler = StandardScaler()\n\nX_train_sc = scaler.fit_transform(X_train_ln)\nX_test_sc = scaler.fit_transform(X_test_ln)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:34.634587Z","iopub.execute_input":"2022-08-09T18:38:34.635717Z","iopub.status.idle":"2022-08-09T18:38:35.640658Z","shell.execute_reply.started":"2022-08-09T18:38:34.635649Z","shell.execute_reply":"2022-08-09T18:38:35.639508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Base Results","metadata":{}},{"cell_type":"code","source":"# prepare dictionary to save base results\nbase_results = {}\ntuning_results = {}","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:35.642094Z","iopub.execute_input":"2022-08-09T18:38:35.642689Z","iopub.status.idle":"2022-08-09T18:38:35.649577Z","shell.execute_reply.started":"2022-08-09T18:38:35.642635Z","shell.execute_reply":"2022-08-09T18:38:35.648385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### LinearRegression","metadata":{}},{"cell_type":"code","source":"# y_train_ln and y_train are same so does not matter which one we will use.\nmodel_lin = LinearRegression()\nmodel_lin.fit(X_train_sc, y_train)\ny_hat = model_lin.predict(X_test_sc)\nbase_results['LinearRegression()'] = round(mean_squared_error(y_test, y_hat, squared=False),5)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:35.651052Z","iopub.execute_input":"2022-08-09T18:38:35.652330Z","iopub.status.idle":"2022-08-09T18:38:36.747520Z","shell.execute_reply.started":"2022-08-09T18:38:35.652293Z","shell.execute_reply":"2022-08-09T18:38:36.746067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Linear Regression can gve us nice interpretation of coefficients. Lets check ","metadata":{}},{"cell_type":"code","source":"# use set_printoptions to get non-scientific notation of coefficients\nnp.set_printoptions(suppress=True, precision=3)\n\nlin_df = pd.DataFrame(data=[model_lin.intercept_]+list(model_lin.coef_) , \n                      index=['intercept']+predictors_ln, columns=['coefficients'])","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:36.748988Z","iopub.execute_input":"2022-08-09T18:38:36.749335Z","iopub.status.idle":"2022-08-09T18:38:36.759864Z","shell.execute_reply.started":"2022-08-09T18:38:36.749299Z","shell.execute_reply":"2022-08-09T18:38:36.758494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lin_df.sort_values(by=['coefficients'], ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:36.761319Z","iopub.execute_input":"2022-08-09T18:38:36.761939Z","iopub.status.idle":"2022-08-09T18:38:36.784051Z","shell.execute_reply.started":"2022-08-09T18:38:36.761903Z","shell.execute_reply":"2022-08-09T18:38:36.782831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"According to this dataframe,                                 \n**most correlated continuous features** are:\n* **manhattan**\n* **JFK**\n* **dropoff_latitude**\n\n**most correlated categorical features** are:\n* **dropoff_latitude**\n* **dropoff_cls_2**\n* **dropoff_longitude**","metadata":{}},{"cell_type":"markdown","source":"#### RandomForestRegressor","metadata":{}},{"cell_type":"code","source":"model_rf = RandomForestRegressor(random_state = 4, n_estimators=5, n_jobs=-1)\nmodel_rf.fit(X_train, y_train)\ny_hat = model_rf.predict(X_test)\nbase_results['RandomForestRegressor()'] = round(mean_squared_error(y_test, y_hat, squared=False),5)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:38:36.785401Z","iopub.execute_input":"2022-08-09T18:38:36.786017Z","iopub.status.idle":"2022-08-09T18:39:37.544891Z","shell.execute_reply.started":"2022-08-09T18:38:36.785977Z","shell.execute_reply":"2022-08-09T18:39:37.543910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### XGBRegressor","metadata":{}},{"cell_type":"code","source":"model_xgb = XGBRegressor(random_state=4, verbosity=0, n_estimators=200, tree_method='gpu_hist')\nmodel_xgb.fit(X_train, y_train)\ny_hat = model_xgb.predict(X_test)\nbase_results['XGBRegressor()'] = round(mean_squared_error(y_test, y_hat, squared=False),5)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:39:37.546529Z","iopub.execute_input":"2022-08-09T18:39:37.546900Z","iopub.status.idle":"2022-08-09T18:39:40.153761Z","shell.execute_reply.started":"2022-08-09T18:39:37.546864Z","shell.execute_reply":"2022-08-09T18:39:40.152923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.barplot(y=predictors,x=list(model_xgb.feature_importances_))\nplt.title('XGBoost Feature Importance');","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:39:40.155196Z","iopub.execute_input":"2022-08-09T18:39:40.155957Z","iopub.status.idle":"2022-08-09T18:39:40.537492Z","shell.execute_reply.started":"2022-08-09T18:39:40.155910Z","shell.execute_reply":"2022-08-09T18:39:40.536575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Base Results","metadata":{}},{"cell_type":"code","source":"base_results_df = pd.DataFrame(data=base_results.values(), columns=['RMSE (TestSet)'], index=base_results.keys())\nbase_results_df","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:39:40.538934Z","iopub.execute_input":"2022-08-09T18:39:40.539800Z","iopub.status.idle":"2022-08-09T18:39:40.552522Z","shell.execute_reply.started":"2022-08-09T18:39:40.539751Z","shell.execute_reply":"2022-08-09T18:39:40.551555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Tuning RandomForestRegressor","metadata":{}},{"cell_type":"markdown","source":"> **note** As training models is too time consuming process, I will leave training codes as note and use models with tuned parameters.","metadata":{}},{"cell_type":"code","source":"# base hyperparameters\nmodel_rf.get_params()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:39:40.555771Z","iopub.execute_input":"2022-08-09T18:39:40.556054Z","iopub.status.idle":"2022-08-09T18:39:40.565851Z","shell.execute_reply.started":"2022-08-09T18:39:40.556027Z","shell.execute_reply":"2022-08-09T18:39:40.564591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n\n# model_rf = RandomForestRegressor(random_state = 4, n_jobs=-1)\n# model_params =  {'n_estimators': [10,20,33], \n#                'max_depth': [5, 9, 12],\n#                'max_features': ['auto', 'sqrt'],\n#                'min_samples_split': [3, 5, 10]}\n\n# model_rf = RandomizedSearchCV(model_rf, model_params, cv=3,verbose=True,\n#                               n_iter=5, n_jobs=-1, random_state=4,\n#                               scoring='neg_root_mean_squared_error')\n\n# model_rf = model_rf.fit(X_train, y_train)\n# print(model_rf.best_params_)\n# print(model_rf.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:39:40.567598Z","iopub.execute_input":"2022-08-09T18:39:40.568162Z","iopub.status.idle":"2022-08-09T18:39:40.573182Z","shell.execute_reply.started":"2022-08-09T18:39:40.568125Z","shell.execute_reply":"2022-08-09T18:39:40.572104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets check if we have improvement in test set:","metadata":{}},{"cell_type":"code","source":"# y_hat_1 = model_rf.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:39:40.574727Z","iopub.execute_input":"2022-08-09T18:39:40.575997Z","iopub.status.idle":"2022-08-09T18:39:40.582688Z","shell.execute_reply.started":"2022-08-09T18:39:40.575872Z","shell.execute_reply":"2022-08-09T18:39:40.581728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# base_res_rf = base_results['RandomForestRegressor()']\n# print(f'Base test RMSE: {base_res_rf}');\n# print(' ')\n# print(f'RandomizedGridSearch test RMSE: {round(mean_squared_error(y_test, y_hat_1, squared=False),5)}')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:39:40.586113Z","iopub.execute_input":"2022-08-09T18:39:40.587022Z","iopub.status.idle":"2022-08-09T18:39:40.593262Z","shell.execute_reply.started":"2022-08-09T18:39:40.586990Z","shell.execute_reply":"2022-08-09T18:39:40.592187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is improvement in test RMSE. Now lets use Gridsearch:","metadata":{}},{"cell_type":"code","source":"# %%time\n\n# model_rf = RandomForestRegressor(random_state = 4, n_jobs=-1)\n# model_params =  {'n_estimators': [33, 50], \n#                'max_depth': [10],\n#                'max_features': ['auto'],\n#                'min_samples_split': [5]}\n\n# model_rf = GridSearchCV(model_rf, model_params, cv=3,verbose=True,\n#                               n_jobs=-1,\n#                               scoring='neg_root_mean_squared_error')\n\n# model_rf = model_rf.fit(X_train, y_train)\n# print(model_rf.best_params_)\n# print(model_rf.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:39:40.594873Z","iopub.execute_input":"2022-08-09T18:39:40.595569Z","iopub.status.idle":"2022-08-09T18:39:40.603113Z","shell.execute_reply.started":"2022-08-09T18:39:40.595518Z","shell.execute_reply":"2022-08-09T18:39:40.602148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# y_hat_1 = model_rf.predict(X_test)\n# base_res_rf = base_results['RandomForestRegressor()']\n# print(f'Base test RMSE: {base_res_rf}');\n# print(' ')\n# print(f'GridSearch test RMSE: {round(mean_squared_error(y_test, y_hat_1, squared=False),5)}')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:39:40.606578Z","iopub.execute_input":"2022-08-09T18:39:40.606956Z","iopub.status.idle":"2022-08-09T18:39:40.613985Z","shell.execute_reply.started":"2022-08-09T18:39:40.606902Z","shell.execute_reply":"2022-08-09T18:39:40.612706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Test RMSE is again reduced. Now I want to increase n_estimators with tuned parameters:","metadata":{}},{"cell_type":"code","source":"# # lets try same model with more n_estimators\n# %%time\n\n# test_model_rf = RandomForestRegressor(random_state=4 , n_jobs=-1, max_depth=10, max_features='auto',\n#                                       min_samples_split=5, n_estimators=150)\n# test_model_rf.fit(X_train, y_train)\n# y_hat_test = test_model_rf.predict(X_test)\n# round(mean_squared_error(y_test, y_hat_test, squared=False),5)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:39:40.615638Z","iopub.execute_input":"2022-08-09T18:39:40.617447Z","iopub.status.idle":"2022-08-09T18:39:40.627270Z","shell.execute_reply.started":"2022-08-09T18:39:40.617421Z","shell.execute_reply":"2022-08-09T18:39:40.626367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Instead of decrease, by using large number of n_estimators , RMSE is increaed, so n_estimators=50 is fine for our model.\n\nLets define our final randomforest model with tuned parameters and save test RMSE:","metadata":{}},{"cell_type":"code","source":"model_rf = RandomForestRegressor(random_state=4 , n_estimators=50, max_depth=10, max_features='auto',\n                                 min_samples_split=5, n_jobs=-1)\nmodel_rf.fit(X_train, y_train)\n\ny_hat_rf = model_rf.predict(X_test)\ntuning_results['RandomForestRegressor()'] = round(mean_squared_error(y_test, y_hat_rf, squared=False),5)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:39:40.628456Z","iopub.execute_input":"2022-08-09T18:39:40.628753Z","iopub.status.idle":"2022-08-09T18:43:51.509155Z","shell.execute_reply.started":"2022-08-09T18:39:40.628718Z","shell.execute_reply":"2022-08-09T18:43:51.508168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###  Tuning XGBRegressor","metadata":{}},{"cell_type":"markdown","source":"> **note** As training models is too time consuming process, I will leave training codes as note and use models with tuned parameters.","metadata":{}},{"cell_type":"code","source":"# %%time\n\n# # tune max_depth, min_child_weight\n# model_xgb = XGBRegressor(learning_rate =0.1, n_estimators=250, \n#                           gamma=0, subsample=0.8, colsample_bytree=0.8, eval_metric='rmse',\n#                           scale_pos_weight=1, random_state=4, tree_method='gpu_hist')\n\n# model_params = {\n#     'max_depth': [3,5, 7, 10, 12],\n#     'min_child_weight': [1,3,6]\n# }\n\n# model_xgb = GridSearchCV(model_xgb, model_params, cv=3, verbose=True,\n#                          scoring='neg_root_mean_squared_error')\n\n# model_xgb = model_xgb.fit(X_train, y_train)\n# print(model_xgb.best_params_)\n# print(model_xgb.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:43:51.510564Z","iopub.execute_input":"2022-08-09T18:43:51.510942Z","iopub.status.idle":"2022-08-09T18:43:51.516791Z","shell.execute_reply.started":"2022-08-09T18:43:51.510901Z","shell.execute_reply":"2022-08-09T18:43:51.515735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n\n# # tune gamma\n# model_xgb = XGBRegressor(learning_rate=0.1, n_estimators=250, max_depth=10, min_child_weight=6,\n#                           gamma=0, subsample=0.8, colsample_bytree=0.8, eval_metric='rmse',\n#                           scale_pos_weight=1, random_state=4, tree_method='gpu_hist')\n\n# model_params = {\n#     'gamma': [0, 0.1, 0.2, 0.3, 0.4]\n# }\n\n# model_xgb = GridSearchCV(model_xgb, model_params, cv=3,verbose=True,\n#                          scoring='neg_root_mean_squared_error')\n\n# model_xgb = model_xgb.fit(X_train, y_train)\n# print(model_xgb.best_params_)\n# print(model_xgb.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:43:51.518543Z","iopub.execute_input":"2022-08-09T18:43:51.519422Z","iopub.status.idle":"2022-08-09T18:43:51.525824Z","shell.execute_reply.started":"2022-08-09T18:43:51.519384Z","shell.execute_reply":"2022-08-09T18:43:51.524889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n# model_xgb = XGBRegressor(learning_rate =0.1, n_estimators=250, max_depth=10, min_child_weight=6,\n#                           gamma=0.2, subsample=0.8, colsample_bytree=0.8, eval_metric='rmse',\n#                           scale_pos_weight=1, random_state=4, tree_method='gpu_hist')\n\n# model_params = {\n#     'subsample': [0.6, 0.7, 0.8, 0.9, 1],\n#     'colsample_bytree': [0.6, 0.7, 0.8, 0.9, 1]\n# }\n\n# model_xgb = GridSearchCV(model_xgb, model_params, cv=3,verbose=True,\n#                          scoring='neg_root_mean_squared_error')\n\n# model_xgb = model_xgb.fit(X_train, y_train)\n# print(model_xgb.best_params_)\n# print(model_xgb.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:43:51.526950Z","iopub.execute_input":"2022-08-09T18:43:51.529011Z","iopub.status.idle":"2022-08-09T18:43:51.538636Z","shell.execute_reply.started":"2022-08-09T18:43:51.528968Z","shell.execute_reply":"2022-08-09T18:43:51.537730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n# model_xgb = XGBRegressor(learning_rate =0.1, n_estimators=250, max_depth=10, min_child_weight=6,\n#                           gamma=0.2, subsample=0.8, colsample_bytree=0.8, eval_metric='rmse',\n#                           scale_pos_weight=1, random_state=4, tree_method='gpu_hist')\n\n# model_params = {\n#     'subsample': [0.95, 1],\n#     'colsample_bytree': [0.75, 0.8, 0.85]\n# }\n\n# model_xgb = GridSearchCV(model_xgb, model_params, cv=3,verbose=True,\n#                          scoring='neg_root_mean_squared_error')\n\n# model_xgb = model_xgb.fit(X_train, y_train)\n# print(model_xgb.best_params_)\n# print(model_xgb.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:43:51.540275Z","iopub.execute_input":"2022-08-09T18:43:51.540707Z","iopub.status.idle":"2022-08-09T18:43:51.548143Z","shell.execute_reply.started":"2022-08-09T18:43:51.540657Z","shell.execute_reply":"2022-08-09T18:43:51.547195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n# model_xgb = XGBRegressor(learning_rate =0.1, n_estimators=250, max_depth=10, min_child_weight=6,\n#                           gamma=0.2, subsample=1, colsample_bytree=0.8, eval_metric='rmse',\n#                           scale_pos_weight=1, random_state=4, tree_method='gpu_hist')\n\n# model_params = {\n#     'reg_alpha':[0.0001, 0.001, 0.1, 1, 5, 100]\n# }\n\n# model_xgb = GridSearchCV(model_xgb, model_params, cv=3,verbose=True,\n#                          scoring='neg_root_mean_squared_error')\n\n# model_xgb = model_xgb.fit(X_train, y_train)\n# print(model_xgb.best_params_)\n# print(model_xgb.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:43:51.549372Z","iopub.execute_input":"2022-08-09T18:43:51.549876Z","iopub.status.idle":"2022-08-09T18:43:51.558558Z","shell.execute_reply.started":"2022-08-09T18:43:51.549836Z","shell.execute_reply":"2022-08-09T18:43:51.557612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n# model_xgb = XGBRegressor(learning_rate =0.1, n_estimators=250, max_depth=10, min_child_weight=6,\n#                           gamma=0.2, subsample=1, colsample_bytree=0.8, eval_metric='rmse',\n#                           scale_pos_weight=1, random_state=4, tree_method='gpu_hist')\n\n# model_params = {\n#     'lambda':[30, 50, 70, 85]\n# }\n\n# model_xgb = GridSearchCV(model_xgb, model_params, cv=3,verbose=True,\n#                          scoring='neg_root_mean_squared_error')\n\n# model_xgb = model_xgb.fit(X_train, y_train)\n# print(model_xgb.best_params_)\n# print(model_xgb.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:43:51.560074Z","iopub.execute_input":"2022-08-09T18:43:51.560726Z","iopub.status.idle":"2022-08-09T18:43:51.571564Z","shell.execute_reply.started":"2022-08-09T18:43:51.560663Z","shell.execute_reply":"2022-08-09T18:43:51.570639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n# model_xgb = XGBRegressor(learning_rate =0.1, n_estimators=250, max_depth=10, min_child_weight=6,\n#                           gamma=0.2, subsample=1, colsample_bytree=0.8, eval_metric='rmse',\n#                           scale_pos_weight=1, random_state=4, tree_method='gpu_hist')\n\n# model_params = {\n#     'reg_alpha':[95, 105]\n# }\n\n# model_xgb = GridSearchCV(model_xgb, model_params, cv=3,verbose=True,\n#                          scoring='neg_root_mean_squared_error')\n\n# model_xgb = model_xgb.fit(X_train, y_train)\n# print(model_xgb.best_params_)\n# print(model_xgb.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:43:51.574961Z","iopub.execute_input":"2022-08-09T18:43:51.575742Z","iopub.status.idle":"2022-08-09T18:43:51.581017Z","shell.execute_reply.started":"2022-08-09T18:43:51.575707Z","shell.execute_reply":"2022-08-09T18:43:51.580110Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n# model_xgb = XGBRegressor(learning_rate =0.1, n_estimators=1000, max_depth=10, min_child_weight=6,\n#                           gamma=0.2, subsample=1, colsample_bytree=0.8, eval_metric='rmse',\n#                           scale_pos_weight=1, random_state=4, tree_method='gpu_hist',\n#                           reg_alpha=95)\n\n# model_params = {\n#     'learning_rate':[0.01, 0.05]\n# }\n\n# model_xgb = GridSearchCV(model_xgb, model_params, cv=3,verbose=True,\n#                          scoring='neg_root_mean_squared_error')\n\n# model_xgb = model_xgb.fit(X_train, y_train)\n# print(model_xgb.best_params_)\n# print(model_xgb.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:43:51.584328Z","iopub.execute_input":"2022-08-09T18:43:51.585051Z","iopub.status.idle":"2022-08-09T18:43:51.591004Z","shell.execute_reply.started":"2022-08-09T18:43:51.585017Z","shell.execute_reply":"2022-08-09T18:43:51.590121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n# model_xgb = XGBRegressor(learning_rate =0.1, n_estimators=1000, max_depth=10, min_child_weight=6,\n#                           gamma=0.2, subsample=1, colsample_bytree=0.8, eval_metric='rmse',\n#                           scale_pos_weight=1, random_state=4, tree_method='gpu_hist',\n#                           reg_alpha=95)\n\n# model_params = {\n#     'learning_rate':[0.03, 0.04]\n# }\n\n# model_xgb = GridSearchCV(model_xgb, model_params, cv=3,verbose=True,\n#                          scoring='neg_root_mean_squared_error')\n\n# model_xgb = model_xgb.fit(X_train, y_train)\n# print(model_xgb.best_params_)\n# print(model_xgb.best_score_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:43:51.592488Z","iopub.execute_input":"2022-08-09T18:43:51.593257Z","iopub.status.idle":"2022-08-09T18:43:51.602706Z","shell.execute_reply.started":"2022-08-09T18:43:51.593221Z","shell.execute_reply":"2022-08-09T18:43:51.601853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Tuning XGBoost hyperparameters is done, lets define our model and see tuned parameters:","metadata":{}},{"cell_type":"code","source":"model_xgb = XGBRegressor(learning_rate =0.04, n_estimators=1000, max_depth=10, min_child_weight=6,\n                          gamma=0.2, subsample=1, colsample_bytree=0.8, eval_metric='rmse',\n                          scale_pos_weight=1, random_state=4, tree_method='gpu_hist',\n                          reg_alpha=95)\n\nmodel_xgb.get_xgb_params()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:43:51.604452Z","iopub.execute_input":"2022-08-09T18:43:51.605118Z","iopub.status.idle":"2022-08-09T18:43:51.616666Z","shell.execute_reply.started":"2022-08-09T18:43:51.605083Z","shell.execute_reply":"2022-08-09T18:43:51.615489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# lets make predictions and save test RMSE\nmodel_xgb.fit(X_train, y_train)\ny_hat_xgb = model_xgb.predict(X_test)\ntuning_results['XGBRegressor()'] = round(mean_squared_error(y_test, y_hat_xgb, squared=False),5)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:43:51.618347Z","iopub.execute_input":"2022-08-09T18:43:51.618771Z","iopub.status.idle":"2022-08-09T18:44:31.399047Z","shell.execute_reply.started":"2022-08-09T18:43:51.618737Z","shell.execute_reply":"2022-08-09T18:44:31.398223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets look XGBoost built-in feature_importance plot:","metadata":{}},{"cell_type":"code","source":"sns.barplot(y=predictors,x=list(model_xgb.feature_importances_))\nplt.title('XGBoost Feature Importance');","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:44:31.400234Z","iopub.execute_input":"2022-08-09T18:44:31.400827Z","iopub.status.idle":"2022-08-09T18:44:31.801951Z","shell.execute_reply.started":"2022-08-09T18:44:31.400786Z","shell.execute_reply":"2022-08-09T18:44:31.800978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Compared to base feature_importance plot, here we have big differences: before tuning, the most \"important\" feature was manhattan distance, but after tuning, as we see, dummy variables (dropoff_cls_2, pick_cls_2, distgr_3) have highest values on the plot.","metadata":{}},{"cell_type":"markdown","source":"# Results","metadata":{}},{"cell_type":"code","source":"# Lets crate summery dataframe\nfinal_results = pd.DataFrame(data=base_results.values(), columns=['BASE RMSE (TestSet)'], index=base_results.keys())\nfinal_results['TUNED RMSE (TestSet)'] = np.nan\n\nfinal_results.loc[final_results.index == 'RandomForestRegressor()',\n                  'TUNED RMSE (TestSet)'] = round(mean_squared_error(y_test, y_hat_rf, squared=False),5)\nfinal_results.loc[final_results.index == 'XGBRegressor()',\n                  'TUNED RMSE (TestSet)'] = round(mean_squared_error(y_test, y_hat_xgb, squared=False),5)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:44:31.803762Z","iopub.execute_input":"2022-08-09T18:44:31.804362Z","iopub.status.idle":"2022-08-09T18:44:31.815549Z","shell.execute_reply.started":"2022-08-09T18:44:31.804321Z","shell.execute_reply":"2022-08-09T18:44:31.814479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_results","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:44:31.817413Z","iopub.execute_input":"2022-08-09T18:44:31.817780Z","iopub.status.idle":"2022-08-09T18:44:31.831512Z","shell.execute_reply.started":"2022-08-09T18:44:31.817743Z","shell.execute_reply":"2022-08-09T18:44:31.830461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Even through Test RMSE of RandomForest declined by more thatn 0.23, XGBoost still gives us lower Test RMSE. So for submission I will use XGBoostRegressor().","metadata":{}},{"cell_type":"code","source":"X_test_final = test_df[predictors].values\ntest_df['fare_amount'] = model_xgb.predict(X_test_final)\ntest_df[['key', 'fare_amount']].head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:44:31.860999Z","iopub.execute_input":"2022-08-09T18:44:31.861637Z","iopub.status.idle":"2022-08-09T18:44:32.252818Z","shell.execute_reply.started":"2022-08-09T18:44:31.861600Z","shell.execute_reply":"2022-08-09T18:44:32.252015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df[['key', 'fare_amount']].to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T18:44:32.256423Z","iopub.execute_input":"2022-08-09T18:44:32.256997Z","iopub.status.idle":"2022-08-09T18:44:32.285620Z","shell.execute_reply.started":"2022-08-09T18:44:32.256960Z","shell.execute_reply":"2022-08-09T18:44:32.284639Z"},"trusted":true},"execution_count":null,"outputs":[]}]}