{"cells":[{"metadata":{"_uuid":"6993d1a9c4ae776f1a98fbc036d7fc6fc000f2d1"},"cell_type":"markdown","source":"## Module imports"},{"metadata":{"trusted":true,"_uuid":"1a54d41a139fc58df2278e655b4b5abe764ac19c"},"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\n\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b19900054ae72d346373c25567cc13bae8b4f162"},"cell_type":"markdown","source":"## 1. Data loading"},{"metadata":{"trusted":true,"_uuid":"ecf475fb18216adda999b1d0308278433bf03a87"},"cell_type":"code","source":"TRAIN_PATH = os.path.join(\"..\", \"input\", \"train.csv\")\nTEST_PATH = os.path.join(\"..\", \"input\", \"test.csv\")\n\ndf_train = pd.read_csv(TRAIN_PATH,index_col=0)\ndf_test = pd.read_csv(TEST_PATH, index_col=0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d4737425d1a5977d7fe034b4190b81b3a411cdb8"},"cell_type":"code","source":"df_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a5251ee668187680f58de256ac871cc9d9c00a87"},"cell_type":"code","source":"df_test.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0d8ebb349e6c364a90d72509299e7da8d6b6b99d"},"cell_type":"markdown","source":"## 2. Data exploration"},{"metadata":{"trusted":true,"_uuid":"4923e2e349aa979be6ed46d9eafed73b554317fb"},"cell_type":"code","source":"df_train.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1f656daa9b95a5efbd74a9abf47c5297782429bd"},"cell_type":"code","source":"df_test.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"22a0ad269dca8a678d1e4c2714e6da990bc4d763"},"cell_type":"markdown","source":"**id** - a unique identifier for each trip\n\n**vendor_id** - a code indicating the provider associated with the trip record\n\n**pickup_datetime** - date and time when the meter was engaged\n\n**dropoff_datetime** - date and time when the meter was disengaged\n\n**passenger_count** - the number of passengers in the vehicle (driver entered value)\n\n**pickup_longitude** - the longitude where the meter was engaged\n\n**pickup_latitude** - the latitude where the meter was engaged\n\n**dropoff_longitude** - the longitude where the meter was disengaged\n\n**dropoff_latitude** - the latitude where the meter was disengaged\n\n**store_and_fwd_flag** - This flag indicates whether the trip record was held in vehicle memory before sending to the vendor because the vehicle did not have a connection to the server - Y=store and forward; N=not a store and forward trip\n\n**trip_duration** - duration of the trip in seconds\n\nDisclaimer: The decision was made to not remove dropoff coordinates from the dataset order to provide an expanded set of variables to use in Kernels."},{"metadata":{"trusted":true,"_uuid":"df30c1b34c119ffba54ad71978957c034d7cecf7"},"cell_type":"code","source":"df_train[\"pickup_datetime\"].head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0970b7cd006564935afd29314ccc7374d0fcef92"},"cell_type":"markdown","source":"## 3. Data preprocessing"},{"metadata":{"_uuid":"58760965ffd8cd5938280cbc982b2f03c63ec851"},"cell_type":"markdown","source":"### 3.a Outliers"},{"metadata":{"trusted":true,"_uuid":"d12a22101db9e0f68266aa531fcccbf4ed017017"},"cell_type":"code","source":"df_train[\"trip_duration\"].hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b9f4a9748d21332ebae90ec5a9bb927fc5000af4"},"cell_type":"markdown","source":"There is a large majority of the values near 0 - 10000.\n\nLet's have a look on this :"},{"metadata":{"trusted":true,"_uuid":"addabae6ae7ead3d8e0e743af06059a45ef471d1"},"cell_type":"code","source":"df_train[df_train.trip_duration < 10000].trip_duration.hist(bins=100)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7909f5ed3ff97f8b056c93f35880c5a9275fa60f"},"cell_type":"code","source":"# We check which partitions of the data to study\n# I think we would only take data which correspond to 95% of the data ?\n\nDURATION_MAX = 3500 # With this value we are studying 99% of the data\n\ndf_train[df_train.trip_duration < DURATION_MAX].trip_duration.hist(bins=100)\n\nprint(\"Count of values >\", DURATION_MAX, \": \", df_train[df_train.trip_duration > DURATION_MAX].trip_duration.count())\nprint(\"Count of values <=\", DURATION_MAX, \": \", df_train[df_train.trip_duration <= DURATION_MAX].trip_duration.count())\nprint(\"Lost data : \", df_train[df_train.trip_duration > DURATION_MAX].trip_duration.count() / df_train[df_train.trip_duration <= DURATION_MAX].trip_duration.count() * 100, \"%\")\n\ndf_train = df_train[df_train.trip_duration < DURATION_MAX]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4877463d171310af93b56a82d1cfd47107d65b9e"},"cell_type":"markdown","source":"I did this because I wanted to select only representative data, but it will be done in a further version."},{"metadata":{"_uuid":"a7e128fe6ff83d7aaf737bee396c05018864249f"},"cell_type":"markdown","source":"### 3.b Missing values"},{"metadata":{"trusted":true,"_uuid":"339199d85c19daba642972a567258e8489e97cd7"},"cell_type":"code","source":"df_train.isna().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"adbebba85e172efe4093565a1441ebc4108cdff0"},"cell_type":"code","source":"df_train.duplicated().sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"82ae5e10aa93e9253428c24f5c01a2ea13cd348f"},"cell_type":"code","source":"df_train[df_train.duplicated()]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ca1a493f6570331913d3e8dabc827430e1a5f539"},"cell_type":"markdown","source":"We can't consider these \"duplicated values\" as needed to be removed."},{"metadata":{"_uuid":"c65c5545d0759574c4326be664e7463b01dffb50"},"cell_type":"markdown","source":"## 4 Features engineering"},{"metadata":{"_uuid":"fa1eb9f7e10ebd735d1b4ef6a90f3b03fb861d58"},"cell_type":"markdown","source":"### 4.a Creation"},{"metadata":{"trusted":true,"_uuid":"645a36b91ad4653f4f54b5f09c5d2b70a2981a57"},"cell_type":"code","source":"df_train[\"pickup_hour\"] = pd.to_datetime(df_train.pickup_datetime).dt.hour\ndf_train[\"pickup_day_of_week\"] = pd.to_datetime(df_train.pickup_datetime).dt.dayofweek","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"0feb17d96f6567643464f73900488029df619b31"},"cell_type":"code","source":"df_train.pickup_hour.hist(bins=47)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b55c32be1dbe60b9896205f67ef7ed0e98379e34"},"cell_type":"code","source":"mean_trip_duration_by_hour = []\n\nfor i in df_train.pickup_hour.unique():\n    mean_trip_duration_by_hour.append(np.mean(df_train[df_train.pickup_hour == i].trip_duration))\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"efb8f6d954f9e23ad8423de829817b5cd4372af1"},"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(22,8))\nax.scatter(x=df_train[:1000].pickup_hour, y=df_train[:1000].trip_duration)\nax.bar(df_train.pickup_hour.unique(), mean_trip_duration_by_hour)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c5da442e901587dfa631e5dd2e75ba2d358d77fa"},"cell_type":"markdown","source":"About day of week :"},{"metadata":{"trusted":true,"_uuid":"8a0cd4834b6064b4c8aa265707698db75939fd78"},"cell_type":"code","source":"df_train.pickup_day_of_week.hist(bins=13)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a59290474ee3902e1a8d2a6618a96fd50e152b6d"},"cell_type":"code","source":"mean_trip_duration_by_day = []\n\nfor i in df_train.pickup_day_of_week.unique():\n    mean_trip_duration_by_day.append(np.mean(df_train[df_train.pickup_day_of_week == i].trip_duration))\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"414b63fab40832dce91a264ba3eb9c66df5b5ee9","scrolled":false},"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(22,8))\nax.scatter(x=df_train[:1000].pickup_day_of_week, y=df_train[:1000].trip_duration)\nax.bar(df_train.pickup_day_of_week.unique(), mean_trip_duration_by_day)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2ab52fc49f0afba38e5151b676454e43006a101d"},"cell_type":"markdown","source":"It doesn't seems to be a relevent study, I won't use it for training for the moment..."},{"metadata":{"trusted":true,"_uuid":"4251402b4b7d9dd288deb280db8b6a071d4f53fc"},"cell_type":"code","source":"df_test[\"pickup_hour\"] = pd.to_datetime(df_test.pickup_datetime).dt.hour\ndf_test[\"pickup_day_of_week\"] = pd.to_datetime(df_test.pickup_datetime).dt.dayofweek","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"374be2c761b7001352bf743a99833506fe749980"},"cell_type":"markdown","source":"### 4.b Features selection"},{"metadata":{"trusted":true,"_uuid":"6d2ead4d135333e8ebec9f28935e7b6b11d8f649"},"cell_type":"code","source":"# 1st test using only basic columns\nSELECTION = [\"pickup_longitude\", \"dropoff_longitude\", \"pickup_latitude\", \"dropoff_latitude\", \"pickup_hour\"]\nTARGET = \"trip_duration\"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e92e60652e21484a9b8a453497796af05614bdbe"},"cell_type":"markdown","source":""},{"metadata":{"trusted":true,"_uuid":"50257eb65510b4aff126aebfd5edee65a8342bfe"},"cell_type":"code","source":"X_train = df_train[SELECTION]\ny_train = df_train[TARGET]\nX_test = df_test[SELECTION]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7d4853c7c3ff36c4ecab665ccd4d1ed1aad7b03a"},"cell_type":"markdown","source":"## 5. Model selection"},{"metadata":{"trusted":true,"_uuid":"bad2ee07d146f83a52de78de7bf48e4604798b98"},"cell_type":"code","source":"from sklearn.ensemble import RandomForestRegressor\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.metrics import mean_squared_error","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d56cdc15f95a984f1866e4138497dc99116ef159"},"cell_type":"code","source":"# m1 = RandomForestRegressor(n_estimators=10)\n# m1.fit(X_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c02cf7ad26a6c877829e0f560680ce6c3cd9853d"},"cell_type":"code","source":"m2 = RandomForestRegressor(n_estimators=15)\nm2.fit(X_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"scrolled":true,"trusted":true,"_uuid":"175ac840f219e582b4c3b814a44301cb1f2b9104"},"cell_type":"code","source":"# m3 = RandomForestRegressor(n_estimators=20)\n# m3.fit(X_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1a96240f1cb05c5340638bc63a6f86adec5873a4"},"cell_type":"code","source":"m4 = RandomForestRegressor(n_estimators=15, min_samples_leaf=100, min_samples_split=150)\nm4.fit(X_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"86e60f14ed61d57fd4bcaf01b7414c1e1c7024f1"},"cell_type":"markdown","source":"## 6. Validation method"},{"metadata":{"_uuid":"cd071359005971b0ae7297e385c5e4a4046405c7"},"cell_type":"markdown","source":"Cross validation (NB : This might not be necessary because of the large number of rows we're working on)"},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"b1a81f1d3a37cbca06ae7b5461f15bdee4cfee76"},"cell_type":"code","source":"cv_scores_1 = cross_val_score(m2, X_train, y_train, cv=5, scoring='neg_mean_squared_log_error')\ncv_scores_1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"68cadf996b1fa99763bb5382142af91d2393795c","scrolled":true},"cell_type":"code","source":"# cv_scores_2 = cross_val_score(m4, X_train, y_train, cv=5, scoring='neg_mean_squared_log_error')\n# cv_scores_2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a8742d04f351e2a4cb7b8e1f92a9446d940495fd","scrolled":false},"cell_type":"code","source":"# def rmse(test, pred):\n#     return np.sqrt(mean_squared_error(test, pred))\n\ndef get_err(score):\n    err_test = []\n    for i in range(len(score)):\n        err_test.append(np.sqrt(abs(score[i])))\n    return err_test\n\nprint(np.mean(get_err(cv_scores_1)))\n# print(np.mean(get_err(cv_scores_2)))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"48cae632540b52303e37795e5703afa835545c07"},"cell_type":"markdown","source":"Results :\n  \n  cross val 1 :  0.4097527612968065\n  \n  cross val 2 : 0.43580030140648873"},{"metadata":{"_uuid":"f885c9a2a26bc2e9581b644cd774508b21dd4709"},"cell_type":"markdown","source":"## 7. Predictions"},{"metadata":{"trusted":true,"_uuid":"40fe1a949c2ae393392c3fbd18b9d2707cd51bf9"},"cell_type":"code","source":"y_test_pred = m2.predict(X_test)\nprint(y_test_pred[:10])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d3c54c4bafe34329fbb7ab9b13ac7ce6a6f6d4bb","scrolled":true},"cell_type":"code","source":"d = { \"id\": df_test.index, \"trip_duration\": y_test_pred}\nsubmission = pd.DataFrame(d)\nsubmission.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5bb0377e24bc07962dbf1f14b38b085c12aa4ea3"},"cell_type":"markdown","source":"## 8. Submission"},{"metadata":{"trusted":true,"_uuid":"7521a67ce85a51d9838ef65cea1b8627b405dec8"},"cell_type":"code","source":"submission.to_csv(\"submission.csv\", index=0)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}