{"cells":[{"metadata":{"_uuid":"7aabacca54b9b09098c22302fee0f1bacf03ff0a"},"cell_type":"markdown","source":"# Automated Machine Learning with H2O AutoML"},{"metadata":{"_uuid":"3e4917652bb7a689a2784ec1100c85d5be819efb"},"cell_type":"markdown","source":"H2O has been dedicated to make ML accessible by non data science expert. One of the major achievements is the AutoML package they developed (and the enterprise version: Driverless AI) It aims to eliminate some of the most time-consuming (yet repetitive) work out of the data science pipeline. (ex. feature engineering, crosss validation, proper targt encoding etc). . How it works is that you specify the feature, set how long you want it to run, and the max memory it can use, and then.........just let it run for as long as you want. At the end, it will evaluate all the models it created, stack ensemble them, and give you the final result.\n\nI am not sure whether this package is able to handle the complexity of industry-level data science project (where most of the time is spent on cleaning/find the right data...), but it seems like a \"Perfec\" tool for Kaggle. Knowing this actually makes me a little bit anxious. Indeed, we are here to learn about machine learning and practice our skills in data science, but at the end of the day, this is a competition site. Do you remember the frustration you had we you see an blend of blend of blend public kernel solution earns a silver? This is a similar feeling to me. As these packages become more and more sophisticated, a kaggle competition could turn into a pure competition of computation resources. Just randomly initiate some AutoML instance, let them run for as long you can, create ensemble, and then repeat the process. Indeed...someone with 0 machine learning experience will be able to get good scores this way, especially for competitions with anonymized data.\n\nThink of it another way, these packages might be able to raise the bar to a new level. I heard from a GM that, 5 years ago, if you know how to do stacking properly, you are already in the top 100 range. Nowadays, almost every kaggler knows how to do stacking (thanks to all generous contributors!) With packages like these, it is no longer enough to play around with sklearn and keras to get a descent score, you have to out-smart these auto-x packages (created by some of the most respectable Grand Masters) to be able to maintain a good stadnding. Challenging yet exiting! \n\nAnyways, here you go. All features come from public kernels. I will add references later (cannot remember which kernels I got them from...but I will figure it out)\n\n#### To replicate the 0.225 score, asumme you have 4 cores, set the run time to 28800s (on your own machine ofc), and give as much ram as you can.  Each run would have different iniialization so scores might be different from time to time. "},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import time\nnotebookstart= time.time()\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nimport gc\nprint(\"Data:\\n\",os.listdir(\"../input\"))\n\n# Models Packages\nfrom sklearn import metrics\nfrom sklearn.metrics import mean_squared_error\nfrom sklearn import feature_selection\nfrom catboost import CatBoostRegressor\nfrom sklearn.model_selection import train_test_split\nfrom sklearn import preprocessing\nfrom sklearn import metrics\nfrom sklearn.metrics import mean_squared_error\n\n# Viz\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d7de24a66cedfc8ee0a90c40ecf11c6c1c2740c1"},"cell_type":"markdown","source":"### Initializaing\nInitializing H2O AutoML instance, you can pass in your custom setting\n- See here for more details: http://docs.h2o.ai/h2o/latest-stable/h2o-docs/automl.html"},{"metadata":{"trusted":true,"_uuid":"7884b444eb791a0f576372b985367c3dddc834c2"},"cell_type":"code","source":"import h2o\nfrom h2o.automl import H2OAutoML\n# Set it according to kernel limits\nh2o.init(max_mem_size = \"10G\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c2dc377589e1260c4a6e42a5247eb640fba6e46f"},"cell_type":"code","source":"print(\"\\nData Load Stage\")\ntraining = pd.read_csv('../input/train.csv', index_col = \"item_id\", parse_dates = [\"activation_date\"])#.sample(1000)\ntraindex = training.index\ntesting = pd.read_csv('../input/test.csv', index_col = \"item_id\", parse_dates = [\"activation_date\"])#.sample(1000)\ntestdex = testing.index\ny = training.deal_probability.copy()\nprint('Train shape: {} Rows, {} Columns'.format(*training.shape))\nprint('Test shape: {} Rows, {} Columns'.format(*testing.shape))","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"# Combine Train and Test\ndf = pd.concat([training,testing],axis=0)\ndel training, testing\ngc.collect()\nprint('\\nAll Data shape: {} Rows, {} Columns'.format(*df.shape))\n\nprint(\"Feature Engineering\")\ndf[\"price\"] = np.log(df[\"price\"]+0.001)\ndf[\"price\"].fillna(-999,inplace=True)\ndf[\"image_top_1\"].fillna(-999,inplace=True)\n\nprint(\"\\nCreate Time Variables\")\ndf[\"Weekday\"] = df['activation_date'].dt.weekday\ndf[\"Weekd of Year\"] = df['activation_date'].dt.week\ndf[\"Day of Month\"] = df['activation_date'].dt.day\n\n# Remove Dead Variables\ndf.drop([\"activation_date\",\"image\"],axis=1,inplace=True)\n\nprint(\"\\nEncode Variables\")\ncategorical = [\"user_id\",\"region\",\"city\",\"parent_category_name\",\"category_name\",\"item_seq_number\",\"user_type\",\"image_top_1\"]\nmessy_categorical = [\"param_1\",\"param_2\",\"param_3\",\"title\",\"description\"] # Need to find better technique for these\nprint(\"Encoding :\",categorical + messy_categorical)\n\n# Encoder:\nlbl = preprocessing.LabelEncoder()\nfor col in categorical + messy_categorical:\n    df[col] = lbl.fit_transform(df[col].astype(str))\n    \nX = df.loc[traindex,:].copy()\nprint(\"Training Set shape\",X.shape)\ntest = df.loc[testdex,:].copy()\nprint(\"Submission Set Shape: {} Rows, {} Columns\".format(*test.shape))\ndel df\ngc.collect()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"c73c94e3f1c1ff89f312a959ab0e7b3266ab1559"},"cell_type":"code","source":"test.drop(\"deal_probability\",axis=1, inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"97d8f83e98196661b5406cd7a1f71a13894cf77e","collapsed":true},"cell_type":"code","source":"# Training and Validation Set\n#X_train, X_valid, y_train, y_valid = train_test_split(\n#    X, y, test_size=0.10, random_state=23)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1a17268c24b0bf009238ab7b310bb0d2b1e9ba8d","collapsed":true},"cell_type":"code","source":"#del X, y\n#gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a311175abbdea6ea8b369cf453a9c45ee38f249f","scrolled":true},"cell_type":"code","source":"htrain = h2o.H2OFrame(X)\n#hval = h2o.H2OFrame(X_valid)\nhtest = h2o.H2OFrame(test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c98b24450f90d1ccfa957361c93b48516523e599","collapsed":true},"cell_type":"code","source":"#del X_train,X_valid,test\n#gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3efcc8148fbccff94ab37bc7f84906ba0a7de067","scrolled":true},"cell_type":"code","source":"x =htrain.columns\ny ='deal_probability'\nx.remove(y)\n\n# Set maximum runtime according to Kaggle limits\naml = H2OAutoML(max_runtime_secs = 9989000)\naml.train(x=x, y =y, training_frame=htrain)\n\nprint('Generate predictions...')\nhtrain.drop(['deal_probability'])\n#preds = aml.leader.predict(hval)\n#preds = preds.as_data_frame()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"205c02d9e00d6cf8b05d81bc596303c1468eef20","collapsed":true},"cell_type":"code","source":"#print('RMSLE H2O automl leader: ', np.sqrt(metrics.mean_squared_error(y_valid, preds)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"73c9e283dd375a5af390e827c8332002098fd90d","collapsed":true},"cell_type":"code","source":"aml.leader","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8e53b5027f886a83d8aed8cc29f8df2b2b56cfbd","collapsed":true},"cell_type":"code","source":"preds = aml.leader.predict(htest)\npreds = preds.as_data_frame()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"27b1d41e7db6f5a7a0aa3171891c2ed3e259a397","collapsed":true},"cell_type":"code","source":"#preds","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8b8a3f96b1f1bb6390f2f5bc2eb1243fae596787","collapsed":true},"cell_type":"code","source":"lgsub = pd.DataFrame(preds.predict.values,columns=[\"deal_probability\"],index=testdex)\nlgsub['deal_probability'].clip(0.0, 1.0, inplace=True) # Between 0 and 1\nlgsub.to_csv(\"lgsub.csv\",index=True,header=True)\n#print(\"Model Runtime: %0.2f Minutes\"%((time.time() - modelstart)/60))\nprint(\"Notebook Runtime: %0.2f Minutes\"%((time.time() - notebookstart)/60))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4d4aa14c47c18459caa79ae0809ef03b1202fd5b","collapsed":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"b2fd466d8364e8f876070874d9a44bb705aa5e66"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.5","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":4}