{"cells":[{"metadata":{"_uuid":"f6207eb6ce63e4b5b04a3d40897b1a24089716b9"},"cell_type":"markdown","source":"# NYC taxi trip duration - EDA + Modeling"},{"metadata":{"_uuid":"ed99a91e439741bb05d6cae7a9035c223975bd95"},"cell_type":"markdown","source":"## Data Loading"},{"metadata":{"trusted":true,"_uuid":"51ac1d9bc17bab60c1b5d0523d82757069a2771b"},"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport seaborn as sns\nimport matplotlib.pyplot as plt \nfrom matplotlib.pyplot import plot\nfrom matplotlib.colors import LogNorm\n\n%matplotlib inline\nsns.set({'figure.figsize':(15,8)})\n\nimport os","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8c89f7cf41b61fd0a1abe215b4b545843eedea2f"},"cell_type":"code","source":"# Load the training and testing datasets\ndf_train = pd.read_csv('../input/train.csv')\ndf_test = pd.read_csv('../input/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8be59069b711910b1bf22319f0ced725fcc20db2"},"cell_type":"markdown","source":"## Data Exploration & Cleaning"},{"metadata":{"trusted":true,"_uuid":"2efc36e80282958ce62c3cac5bbaa6bcf1c1c926"},"cell_type":"code","source":"# Print the first 5 rows of the traing dataset\ndf_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"2c5cf6d72233a3eb84544204d5ca8ff585af8e8c"},"cell_type":"code","source":"# Print the first 5 rows of the testing dataset\ndf_test.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"d426efa606743631182fb8043b2fbaf20f07a982"},"cell_type":"code","source":"# List all the training dataset features, the number of values in each and their type\ndf_train.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"848abb698a40d74b1c6fd0b15c0d22f948dd504d"},"cell_type":"code","source":"# List all the training dataset features, the number of values in each and their type\ndf_test.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7959922029a7f49f200093227703f67ed775bada"},"cell_type":"markdown","source":"We can see that the training dataset counts 11 features versus 9 for the testing dataset. The two extra features in the training dataset are \"dropoff_datetime\" (numerical feature) and **the target feature: \"trip_duration\"**. The training dataset also counts more values per feature: 1.458.644 versus 625.134 for the testing dataset."},{"metadata":{"trusted":false,"_uuid":"48e32dce2b38e5ea55e001ab9e1518be0ff6914d"},"cell_type":"code","source":"# See if there are any missing values in the training dataset\nfor i, x in zip(list(df_train.isnull().sum().index),list(df_train.isnull().sum().values)):\n    print(f\"The feature {i} counts {x} missing value\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"44f683c876070853a477e4dca863bdc589b2317e"},"cell_type":"code","source":"# See if there are any missing value in the testing dataset as well\nfor i, x in zip(list(df_test.isnull().sum().index),list(df_test.isnull().sum().values)):\n    print(f\"The feature {i} counts {x} missing value\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"17481f75fd15b39f46fc33b0e3b8ad14bf0daa50"},"cell_type":"code","source":"# Now let's make sure there are no duplicated values\ntrain_mv = df_train.duplicated().sum()\ntest_mv = df_test.duplicated().sum()\nprint(f'The training dataset counts {train_mv} duplicated value and the testing dataset counts {test_mv} as well.')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ce6bba42bb07631c9b905b0d426040b0e0d91053"},"cell_type":"markdown","source":"It is not convenient to have both the date and the timing in the same cell, for the features \"pickup_datetime\" and \"dropoff_datetime\". Hence, we will split them accordingly."},{"metadata":{"trusted":false,"_uuid":"0bdf2aad2a9146e5c86138b4d4518ab78534c2b8"},"cell_type":"code","source":"# Split the cells content in two new features: pickup_day and pickup_time\ndf_train['pickup_day'] = pd.to_datetime(df_train['pickup_datetime']).dt.date\ndf_train['pickup_time'] = pd.to_datetime(df_train['pickup_datetime']).dt.time\n\n# We apply this same logic to the testing set\ndf_test['pickup_day'] = pd.to_datetime(df_test['pickup_datetime']).dt.date\ndf_test['pickup_time'] = pd.to_datetime(df_test['pickup_datetime']).dt.time\n\n# Then we do the same for the dropoff feature\ndf_train['dropoff_day'] = pd.to_datetime(df_train['dropoff_datetime']).dt.date\ndf_train['dropoff_time'] = pd.to_datetime(df_train['dropoff_datetime']).dt.time","execution_count":null,"outputs":[]},{"metadata":{"scrolled":false,"trusted":false,"_uuid":"b294a4ef6e2782ad912659a6b31c8d59aad025c1"},"cell_type":"code","source":"df_train.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"369100c4507ebc4948d83ae7e6e21c9ce97b1872"},"cell_type":"markdown","source":"Ok, now we can take a closer look. First, let's examine our most important feature: the target one, \"trip_duration\". "},{"metadata":{"trusted":false,"_uuid":"36ca7cc97d30eb598100d97ffa7575bc7b0ffc44"},"cell_type":"code","source":"df_train['trip_duration'].describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"9d8f0154881c6bf194083f705f182150d16aa4ae"},"cell_type":"code","source":"# Now that we now the mean and median, let's see if there are outliers\ndf_train.boxplot(['trip_duration']);","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d82191e09d83da5020423bceeee1849328f7d2f8"},"cell_type":"markdown","source":"As we can see, we need to handle the outliers considering their extent. Let's zoom on the box to get a better look."},{"metadata":{"scrolled":true,"trusted":false,"_uuid":"eebf1ed7df944627ca3cdab5eb129b1ae55389b0"},"cell_type":"code","source":"df_train.boxplot(['trip_duration'], showfliers=False, notch=True);","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"f54d88ed4ebc3d454ddd22139ae76cf98d19d7af"},"cell_type":"code","source":"len(df_train.trip_duration[df_train.trip_duration > 4000].values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"6c4e3e0f0e563f112b947e46384dbb3b56d35ee2"},"cell_type":"code","source":"len(df_train.trip_duration[df_train.trip_duration < 10].values)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fff066cd3e27368226ad6e75f7b3c8b1e231ecb8"},"cell_type":"markdown","source":"We can see that most of the trip durations last between 0 and roughly 2500 seconds (approximately 46 minutes), with the highest concentration between around 800 1100 seconds, and with less than 10.000 trips with a duration over 4000 seconds (approximately 66 minutes). As a result, **we will from now on consider the trip durations over 4000 seconds as outliers** and won't take them into account in our modeling process."},{"metadata":{"_uuid":"562331051bee7c0910a44b1400347cf3e2d48703"},"cell_type":"markdown","source":"## Features Engineering"},{"metadata":{"scrolled":true,"trusted":false,"_uuid":"86375affd7c818036f957aee329d6ee3b9285503"},"cell_type":"code","source":"# Let's create a new dataframe, without the outliers\ndf2_train = df_train[df_train.trip_duration < 4000]\ndf2_train.info()","execution_count":null,"outputs":[]},{"metadata":{"scrolled":true,"trusted":false,"_uuid":"523ee470dd0c1c59ca3ced2d702b6f7a9b1503d5"},"cell_type":"code","source":"# Let's see the values changes \ndf2_train['trip_duration'].describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"59e734f0732f7d24e82cd863597b0b433a5e11dc"},"cell_type":"markdown","source":"Now that we have normalized our target feature, let's plot those its points to see what their distribution looks like."},{"metadata":{"scrolled":true,"trusted":false,"_uuid":"586ca7b31063643507c3ef73ad630d4aa31ff7c1"},"cell_type":"code","source":"df2_train['trip_duration'].hist(bins=100, histtype='stepfilled')\nplt.title(\"Ditribution of the trip_duration feature points\");","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0807accfb633246ca57f3a13549238b069d28a71"},"cell_type":"markdown","source":"This is what appears to be a **normal distribution with a left skewness**."},{"metadata":{"_uuid":"802b26d26eab0d85df343c3ab491b07bc9cfcab6"},"cell_type":"markdown","source":"Now, we should handle the latitudes and longitudes. First, let's see if we can visualize some particularly high concentrations of points in specific places, in order to get a first insight."},{"metadata":{"trusted":false,"_uuid":"b086510f68c4fbf98e78ed3fc3bb4733f61e4181"},"cell_type":"code","source":"df2_train.boxplot(['pickup_longitude', 'dropoff_longitude', 'pickup_latitude', 'dropoff_latitude']);","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4e2b2d85b4977929ec16f7f1e3e21f5192b7197a"},"cell_type":"markdown","source":"Seems like we have outliers. Let's verify."},{"metadata":{"trusted":false,"_uuid":"076cdb01bf884c0c91e6bedd2606ba2391987b2b"},"cell_type":"code","source":"len(df2_train.pickup_longitude[df2_train.pickup_longitude < -80].values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"a7d6eb2f2b2e60625e98c4bfcbc5a9c21713c89b"},"cell_type":"code","source":"len(df2_train.pickup_longitude[df2_train.pickup_longitude > -50].values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"8fda320c14ad79fa5c1923720609044926020e6b"},"cell_type":"code","source":"len(df2_train.pickup_latitude[df2_train.pickup_latitude < 25].values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"6c97459f9c31744f9ec602db7f4cd65387e9b4e5"},"cell_type":"code","source":"len(df2_train.pickup_latitude[df2_train.pickup_latitude > 50].values)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e25955bc5c041b945212519d2d2cbcbdac423ee7"},"cell_type":"markdown","source":"Now we can set maximum and minimum ranges for our data visualization of the highest concentrations of pickup and dropoff locations."},{"metadata":{"trusted":false,"_uuid":"d685a1038fb797695b0c406af326f20002e4fa88"},"cell_type":"code","source":"min_long = -80\nmax_long = -50\nmin_lat = 25\nmax_lat = 50","execution_count":null,"outputs":[]},{"metadata":{"scrolled":true,"trusted":false,"_uuid":"c6024dfb744cab9735f9295f529267b725553f45"},"cell_type":"code","source":"fig = plt.figure(1, figsize=(10,5))\nhist = plt.hist2d(df2_train.pickup_longitude, df2_train.pickup_latitude, bins=50, range=[[min_long,max_long], [min_lat,max_lat]], norm=LogNorm())\nplt.xlabel('Longitude')\nplt.ylabel('Latitude')\nplt.colorbar(label='Value counts')\nplt.title('NYC pickup locations')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bd2f15e53c92fd677f6962fde149cb6c63c562ad"},"cell_type":"markdown","source":"One last thing, let's store the features according to their type. "},{"metadata":{"trusted":false,"_uuid":"d239ec7bc9323a28e8479c1f31b7784e3d1d6790"},"cell_type":"code","source":"num_col = [\n    col for col in df2_train.columns if \n    (df2_train[col].dtype=='int64' or df2_train[col].dtype=='float64') \n    and col != 'trip_duration']\n\nnum_col","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"ff72b401aac41c03c070f4494b80a545ad3c8469"},"cell_type":"code","source":"cat_col = [\n    col for col in df2_train.columns if \n    (df2_train[col].dtype=='object') \n    and col != 'trip_duration']\n\ncat_col","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"b6693758978a7f94dca53252e58e8454950469e7"},"cell_type":"code","source":"for col in cat_col:\n    df2_train[col] = df2_train[col].astype('category').cat.codes\n    \ndf2_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"2adfb8182d6ddd240f7a8320158bc2359bcc2757"},"cell_type":"code","source":"# Finally, we lock the target fature in a constant one\nTARGET = df2_train.trip_duration","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"7f1b663919df062e07e87dc133ee83f51fd1f87f"},"cell_type":"code","source":"df2_train.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"73d501aa264e0eba07c56dd00197a4d5a463743f"},"cell_type":"markdown","source":"# Model Validation & Training"},{"metadata":{"scrolled":true,"trusted":false,"_uuid":"169f9288ec3e80b01924110114971afb1a8b07dc"},"cell_type":"code","source":"X_train = df2_train[['vendor_id', 'passenger_count', 'pickup_longitude', 'pickup_latitude', 'dropoff_longitude', 'dropoff_latitude', 'pickup_day', 'pickup_time']]\nX_train.shape","execution_count":null,"outputs":[]},{"metadata":{"scrolled":true,"trusted":false,"_uuid":"8b9f91b49b7f40db6bd96cd8b250779ffbc227fb"},"cell_type":"code","source":"y_train = df2_train.trip_duration\ny_train.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a6801ee62ab1b57debe9abb9d7fc6de1c4646d36"},"cell_type":"markdown","source":"Let's split our model so we can validate the training before testing."},{"metadata":{"trusted":false,"_uuid":"a713435d6cf4f907125244faab40b77364f0a82b"},"cell_type":"code","source":"from sklearn.model_selection import train_test_split","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"6ef2f38ab2717176da64111091bf267091cfbd63"},"cell_type":"code","source":"X_train_new, X_valid, y_train_new, y_valid = train_test_split(X_train, y_train, \n                                                              test_size=.2, random_state=42, stratify=y_train)\nX_train.shape, y_train.shape, X_valid.shape, y_valid.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"9e1699a9acc9ec5f391ae94c30cbe9bc8284a5f7"},"cell_type":"code","source":"from sklearn.ensemble import RandomForestRegressor","execution_count":null,"outputs":[]},{"metadata":{"scrolled":true,"trusted":false,"_uuid":"2c9147fed039101483eb6e74e5ea03420980c303"},"cell_type":"code","source":"m1 = RandomForestRegressor(n_estimators=10, random_state=6)\nm1.fit(X_train, y_train)\nm1.score(X_valid, y_valid)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"556ac16654cba2a41172582ac1ad65cb322533dc"},"cell_type":"markdown","source":"Now that's a great score. Maybe too great, **our model is probably overfitting**. Let's try to reduce the number of trees."},{"metadata":{"_uuid":"bc45f791197f2bbc7c7f1795981b70efdc9f7841"},"cell_type":"markdown","source":"Let's try to reduce the number of trees."},{"metadata":{"trusted":false,"_uuid":"497e69b55cf671f5c4414a8d83ca2d25bb9e0dd5"},"cell_type":"code","source":"m2 = RandomForestRegressor(n_estimators=5, random_state=42)\nm2.fit(X_train, y_train)\nm2.score(X_valid, y_valid)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"500ae957a0fca74ba474a64942a7408cb71729e3"},"cell_type":"markdown","source":"Not working. Let's try with the number of leaves our trees can produce. "},{"metadata":{"trusted":false,"_uuid":"0fe7179cc36a404ac1106270a24c31f3f4258cf3"},"cell_type":"code","source":"m3 = RandomForestRegressor(n_estimators=8, random_state=42, max_leaf_nodes=100)\nm3.fit(X_train, y_train)\nm3.score(X_valid, y_valid)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7715c364b7ebd5cdba8eeee2fe9578cacacad1a4"},"cell_type":"markdown","source":"Looks like we can reduce the risk of overfitting this way. Now 0.61 is a bit of a low score, so we may have been too limiting on the number of leaves ; let's raise it back a little."},{"metadata":{"trusted":false,"_uuid":"6d97edee8470b9333fdbd2f8f850721fa652ecb3"},"cell_type":"code","source":"m4 = RandomForestRegressor(n_estimators=10, random_state=42, max_leaf_nodes=750)\nm4.fit(X_train, y_train)\nm4.score(X_valid, y_valid)","execution_count":null,"outputs":[]},{"metadata":{"scrolled":true,"trusted":false,"_uuid":"46222d4318714edcdd1b18b6f013cf88e509c198"},"cell_type":"code","source":"m5 = RandomForestRegressor(n_estimators=10, random_state=42, max_leaf_nodes=50000)\nm5.fit(X_train, y_train)\nm5.score(X_valid, y_valid)","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"b0daa4c7501809921cbeeee11ef58fd3cde4077d"},"cell_type":"code","source":"from sklearn.metrics import r2_score","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"cea62827891bdf37833b85cc52aedc4e6c40643e"},"cell_type":"code","source":"y_valid_pred = m5.predict(X_valid)\ny_valid_pred","execution_count":null,"outputs":[]},{"metadata":{"scrolled":false,"trusted":false,"_uuid":"59939152a8e60ab1431ae3405cc4a2d6b9d6864c"},"cell_type":"code","source":"r2_score(y_valid, y_valid_pred)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9943c8ae3ffd553a4f4f37f1741ff1f8d2f5786f"},"cell_type":"markdown","source":"Last thing, let's check the **mean squared error**, since the evaluation metric for this competition is Root Mean Squared Logarithmic Error. The lower the MSE, the higher the accuracy on the predictions."},{"metadata":{"trusted":false,"_uuid":"ad4fdef9adfb4b69f0d9f242d1337b7be01f2dd5"},"cell_type":"code","source":"from sklearn.metrics import mean_squared_error as MSE\nfrom sklearn.linear_model import SGDRegressor","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"b5a82daebe1715d29eeac99c08906dd71b57730e"},"cell_type":"code","source":"sgd = SGDRegressor()","execution_count":null,"outputs":[]},{"metadata":{"scrolled":true,"trusted":false,"_uuid":"b6f4a1f0761397a458a15ef93d97fccacf0de500"},"cell_type":"code","source":"sgd.fit(X_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"scrolled":true,"trusted":false,"_uuid":"ecc10df02babd1cbba2f4c9455c3e54450d76607"},"cell_type":"code","source":"MSE(y_valid, m5.predict(X_valid))","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"2c2582494a4d30c1198491a9a02fed0ecc2f5264"},"cell_type":"code","source":"loss = MSE(y_valid, sgd.predict(X_valid))\nloss","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"023b92fa6015e05466fc3c6380a9d5bf29cbfac9"},"cell_type":"code","source":"from sklearn.linear_model import LinearRegression","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"78217030d4dd297bcc556f52fac9ae63e62d120e"},"cell_type":"code","source":"lr = LinearRegression()\nlr.fit(X_train, y_train)\nMSE(y_valid, lr.predict(X_valid))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"17e29586ede7205dc60dfa23cf604d3d82f4c8ea"},"cell_type":"markdown","source":"Good ! Now we can cross validate our scores and start the predictions."},{"metadata":{"trusted":false,"_uuid":"30086ca517bd333915188823d696f427c849fd98"},"cell_type":"code","source":"from sklearn.model_selection import cross_val_score\n\ncv_scores = cross_val_score(m5, X_train, y_train, cv=5, scoring='neg_mean_squared_log_error')\ncv_scores","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"6d47d307e2b784831a770f4c1590b44f32a1b8be"},"cell_type":"code","source":"X_test = df_test[['vendor_id', 'passenger_count', 'pickup_longitude', 'pickup_latitude', 'dropoff_longitude', 'dropoff_latitude', 'pickup_day', 'pickup_time']]\ny_test_pred = m5.predict(X_test)\ny_test_pred[:10]","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"95cc8979f152bae08c90907ac77c6482e3ce21b2"},"cell_type":"code","source":"submission = pd.DataFrame(df_test.loc[:, 'id'])\nsubmission['trip_duration'] = y_test_pred\nprint(submission.shape)\nsubmission.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a3a843e1c6357195ca75c7a95e8f87daf73847c5"},"cell_type":"code","source":"submission.to_csv(\"submit_file.csv\", index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}