{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"# New York City Taxi Trip Duration"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"markdown","source":"The competition dataset is based on the 2016 NYC Yellow Cab trip record data made available in Big Query on Google Cloud Platform. The data was originally published by the NYC Taxi and Limousine Commission (TLC). The data was sampled and cleaned for the purposes of this playground competition. Based on individual trip attributes, participants should predict the duration of each trip in the test set.\n\n### Data fields\n\n- `id` - a unique identifier for each trip\n\n- `vendor_id` - a code indicating the provider associated with the trip record\n\n- `pickup_datetime` - date and time when the meter was engaged\n\n- `dropoff_datetime` - date and time when the meter was disengaged\n\n- `passenger_count` - the number of passengers in the vehicle (driver entered value)\n\n- `pickup_longitude` - the longitude where the meter was engaged\n\n- `pickup_latitude` - the latitude where the meter was engaged\n\n- `dropoff_longitude` - the longitude where the meter was disengaged\n\n- `dropoff_latitude` - the latitude where the meter was disengaged\n\n- `store_and_fwd_flag` - This flag indicates whether the trip record was held in vehicle memory before sending to the vendor because the vehicle did not have a connection to the server - Y=store and forward; N=not a store and forward trip\n\n- `trip_duration` - duration of the trip in seconds"},{"metadata":{"_uuid":"a909f91f5a79461a8f0c8b5677ab78ef0ab3f008"},"cell_type":"markdown","source":"### Module imports"},{"metadata":{"trusted":true,"_uuid":"9d49e9646568f270fe15cffd174a6162789eefe5"},"cell_type":"code","source":"import os\n\nimport math\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport warnings\n\nfrom datetime import datetime","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"aee48558b808a1910dbfd5ceca7c73b1b3bc6f75"},"cell_type":"code","source":"%matplotlib inline\nsns.set({'figure.figsize':(10,6), 'axes.titlesize':20, 'axes.labelsize':8})","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"55992aee7002aceb1fe04ad1b5b53a2520ced55c"},"cell_type":"markdown","source":"## 1 - Data loading"},{"metadata":{"trusted":true,"_uuid":"6bcb1bf6c040a58e4b88ff0db4a6201efd8d080e"},"cell_type":"code","source":"FILEPATH_TRAIN = os.path.join('../input/train.csv')\nFILEPATH_TEST = os.path.join('../input/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"943f1c442893ca75fe3f9409e3ae8288fe14d6f0"},"cell_type":"code","source":"df_train = pd.read_csv(FILEPATH_TRAIN)\ndf_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bb1a946f6b75805bd9dd3be0017782f4374fff52"},"cell_type":"code","source":"df_test = pd.read_csv(FILEPATH_TEST)\ndf_test.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8dd57fcc728353c7822b3ae3f25bf76736172bdb"},"cell_type":"markdown","source":"## 2 - Data exploration"},{"metadata":{"trusted":true,"_uuid":"0cbc120fac8cc6fd1121c3a17f5be14a1011f019"},"cell_type":"code","source":"df_train.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ece4e52ac07731b68eeb2c2435018d41f348a7b4"},"cell_type":"code","source":"df_train.describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8a0bbad721cd59d95a1104ad482f2f9a8f34f45f"},"cell_type":"markdown","source":"## 3 -  Data pre-processing"},{"metadata":{"_uuid":"6cd3dd59cbd5087045ec2789dccb43b156362338"},"cell_type":"markdown","source":"### 3.1 - Outliers"},{"metadata":{"trusted":true,"_uuid":"84d5efc1bd4596363ba8514470245e01ffdb1652"},"cell_type":"code","source":"plt.hist(df_train[df_train.trip_duration < 5000].trip_duration, bins = 100)\nplt.title('Trip duration distribution')\nplt.xlabel('Duration of a trip (in seconds)')\nplt.ylabel('Number of trips')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ffffd119d98779e71ba177ac70a46fd23101640b"},"cell_type":"markdown","source":"All `trip_duration` are less than 5000 seconds, a majority of them are between 0 and 3000 seconds, that is why we decide to focus on these data : only trips of less than 3000 seconds"},{"metadata":{"trusted":true,"_uuid":"ca8b38c09392ed771e6d19c66f0996c2ea0af068"},"cell_type":"code","source":"df_train = df_train[(df_train.trip_duration < 3000)]\ndf_train.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f6ced6521073f783ea780606ae7155912dca6d2"},"cell_type":"markdown","source":"We went from 1458644 entries against 1434783 after outliers management, a loss of less than 2% of the training data"},{"metadata":{"_uuid":"9a1e362602b1dbca0dd3a25883207cebf8470cc3"},"cell_type":"markdown","source":"### 3.2 - Missing & duplicate values"},{"metadata":{"trusted":true,"_uuid":"dc63c58ba684a80bfdd0e5036d6df4e7e78d3055"},"cell_type":"code","source":"df_train.isna().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f4c5af68ed0024f0721380fbfb3a82c1b3dc1bb0"},"cell_type":"code","source":"df_train.duplicated().sum()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"991bf9489975779646198ab4b5990672cb9556c9"},"cell_type":"markdown","source":"There is no missing value or duplicated rows in the dataset"},{"metadata":{"_uuid":"cebee1c70d6392ba3942c30aacca39852ac6b989"},"cell_type":"markdown","source":"### 3.3 - Categorical variables"},{"metadata":{"_uuid":"67991cec442fd40a96706beb545437e3bd11cec5"},"cell_type":"markdown","source":"The `store_and_fwd` column is a categorical variable, so we assign a numeric variable to each value in order to use it later\n\nThis change concerns training and test data"},{"metadata":{"trusted":true,"_uuid":"0b2490fede44650aa6a25ddc20a80f82ac1889da"},"cell_type":"code","source":"cat_vars = ['store_and_fwd_flag']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a8900dd2a4276373b5953399ff419a329409fa36"},"cell_type":"code","source":"for col in cat_vars:\n    df_train[col] = df_train[col].astype('category').cat.codes\ndf_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f86cdc8a461502c99dfce2aae5cfacb6382262dc"},"cell_type":"code","source":"for col in cat_vars:\n    df_test[col] = df_test[col].astype('category').cat.codes\ndf_test.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8bea734fe2d1cf2ebec68f53c809c0efdc7d9cff"},"cell_type":"markdown","source":"## 4 - Features engineering"},{"metadata":{"_uuid":"54c23ff270baadac62e897ca938c2a50debfb7bc"},"cell_type":"markdown","source":"### 4.1 - Features creation"},{"metadata":{"_uuid":"53d74e57f5640a30e2528ea923537ee07b1d252e"},"cell_type":"markdown","source":"During the analysis of the outliers, it turns out that the distribution of the data is mostly on the right.\n\nIt seems relevant to be interested in the logarithmic value of the trip duration, hence the creation of the feature `log_trip_duration`"},{"metadata":{"trusted":true,"_uuid":"e846139f02966f15fca3bdf7665594744798c644"},"cell_type":"code","source":"df_train['log_trip_duration'] = np.log(df_train.trip_duration)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f1a12f8b6c3de90bb6caad2ceb64edf047898ccb"},"cell_type":"markdown","source":"Moreover, it seems interesting to create a variable `distance` which represents the distance between the point of departure and arrival"},{"metadata":{"trusted":true,"_uuid":"e8aa74fb7fd8707e0d8f9da134673c70911aa72e"},"cell_type":"code","source":"df_train['distance'] = np.sqrt((df_train.pickup_latitude - df_train.dropoff_latitude)**2 + (df_train.pickup_longitude - df_train.dropoff_longitude)**2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5ab1298d6722d78d16ede4f2421d146f30ccd66a"},"cell_type":"code","source":"df_test['distance'] = np.sqrt((df_test.pickup_latitude - df_test.dropoff_latitude)**2 + (df_test.pickup_longitude - df_test.dropoff_longitude)**2)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1e8ca383c7ec7b3ec3626b6039b20d7418a4f2c0"},"cell_type":"markdown","source":"We are studying the logarithmic value of the duration trip, so we must do the same with the distance : `log_distance`"},{"metadata":{"_kg_hide-output":true,"trusted":true,"_uuid":"1c8c736c3e01db4226e9e43bc8c225c9db70b0f8"},"cell_type":"code","source":"df_train['log_distance'] = np.log(df_train.distance)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4eee9970a3de98c41143d2c9b478300288bae2aa"},"cell_type":"code","source":"df_test['log_distance'] = np.log(df_test.distance)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-output":true,"_uuid":"487660efd190c7713f3756bd3624746625cc21b9"},"cell_type":"markdown","source":"### 4.2 - Features selection"},{"metadata":{"_uuid":"6d5122339cfd623e0c532ab7e1bbcbcbc6f91327"},"cell_type":"markdown","source":"In order to predict travel time, we are interested in some features\n\nTherefor, we decide to ignore the following ones : `id`, `vendor_id`, and `store_and_fwd_flag`"},{"metadata":{"trusted":true,"_uuid":"b69ba6a602481ea4cc0def9c65bb0e6440726ee6"},"cell_type":"code","source":"df_train = df_train.drop(['vendor_id', 'store_and_fwd_flag'], axis=1)\ndf_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e2dd29082953eb707bc0e8a4c332e35a065fc12b"},"cell_type":"code","source":"df_test = df_test.drop(['vendor_id', 'store_and_fwd_flag'], axis=1)\ndf_test.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ccd35e2c47f778f4acfa87b8becdb1d25607f398"},"cell_type":"markdown","source":"As mentioned above, we will only select the following features"},{"metadata":{"trusted":true,"_uuid":"7dd91cdd4c1fccf5e3bd5a64cf3ec98a4529a0ef"},"cell_type":"code","source":"num_features = ['pickup_longitude', 'pickup_latitude', 'dropoff_longitude', 'dropoff_latitude']\ntarget = 'log_trip_duration'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d37edafe9c5838c6631f8cab3b46f36313e880aa"},"cell_type":"code","source":"X_train = df_train.loc[:, num_features]\ny_train = df_train[target]\nX_test = df_test.loc[:, num_features]\nX_train.shape, y_train.shape, X_test.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ab5752e9ca55c0c7bb53fa2df723aa5808ecc319"},"cell_type":"markdown","source":"## 5 - Model training"},{"metadata":{"trusted":true,"_uuid":"6f44d8706fb5018b14d8cce1bdefc0fbdcd45739"},"cell_type":"code","source":"from sklearn.ensemble import RandomForestRegressor","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ece0ae8fef3908d5e8be063147dabfcb56a70b16"},"cell_type":"code","source":"m = RandomForestRegressor(n_estimators=20)\nm.fit(X_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"60e5f2a06beccffd43b1e7f2bc2cf527e6802b1f"},"cell_type":"markdown","source":"## 6 - Validation"},{"metadata":{"trusted":true,"_uuid":"c4d5dfd639f1d51b01fa52eff052c04dced90e94"},"cell_type":"code","source":"from sklearn.model_selection import cross_val_score","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"333ccb258efbcb41b1049c4a8eba0aca3b07334e"},"cell_type":"code","source":"cv_scores = cross_val_score(m,X_train,y_train,cv=5,scoring='neg_mean_squared_log_error')\ncv_scores","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"569a178ae9001f3bf64bd06263d3efecba4501e0"},"cell_type":"code","source":"for i in range(len(cv_scores)):\n    cv_scores[i] = np.sqrt(abs(cv_scores[i]))\ncv_scores","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ca4d52c46a2399d6aecf4c01f38229d25697ae40"},"cell_type":"markdown","source":"## 7 - Predictions"},{"metadata":{"trusted":true,"_uuid":"b2dd5f2e2cf812307430e53ada4f50eee2751afd"},"cell_type":"code","source":"y_test_pred = m.predict(X_test)\ny_test_pred[:5]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3968c924f18fb2eba9d121b32b0caefcec3a71de"},"cell_type":"markdown","source":"## 8 - Submit predictions"},{"metadata":{"trusted":true,"_uuid":"5b8f536e9d1a31b395d260906b113c37ffb58cf0"},"cell_type":"code","source":"submission = pd.DataFrame({'id': df_test.id, 'trip_duration': np.exp(y_test_pred)})\nsubmission.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ab9c88153a78f227a001a3ffcde6fd0add33e57e"},"cell_type":"code","source":"submission.to_csv('Submission_file.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"612d82d17f5492fc9f1c1e8e47df9de9d49a3e3c"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"c4ff7353dc50919b1d64e0477b8cea1ea280682c"},"cell_type":"markdown","source":""}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}