{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Before going to main content\nnew feature was being created.\n\nV1 V319-V320\n\nV7 V109-V110\n\nV12 V329-V330\n\nV14 V316-V331\n\nV18 V4-V5\n \n\n**Sorry, I want to update quickly this kernel, but I can't because these days kernel's commiting is too busy! **\n\nThis kernel is using official kernel for judging whether my new feature is meaningful.\n\nThe official kernel is here(https://www.kaggle.com/inversion/ieee-simple-xgboost)\n\nAnd that the feature which I made is explained in detail here ( https://www.kaggle.com/yasagure/how-do-we-treat-with-similar-columns-v319-v321/edit/run/18988983)"},{"metadata":{},"cell_type":"markdown","source":"# Introduction - Do they match?\n**Some people throw away similar columns, but it is good thing?**\n\nEveryone found that there is a lot of similar columns in this dataset.\n\nV319-V320 and V109-V110 are good example of them.\n\nHow should we \"use\" it?\n\nMost people throw away the data, but I think that I can get useful information from it.\n\n**My idea is paying attention to whether they match each other.**"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport os\nprint(os.listdir(\"../input\"))\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn import preprocessing\nimport matplotlib.pylab as plt\n%matplotlib inline\nimport xgboost as xgb","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# load dataset"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_transaction = pd.read_csv('../input/ieee-fraud-detection/train_transaction.csv', index_col='TransactionID')\ntest_transaction = pd.read_csv('../input/ieee-fraud-detection/test_transaction.csv', index_col='TransactionID')\n\ntrain_identity = pd.read_csv('../input/ieee-fraud-detection/train_identity.csv', index_col='TransactionID')\ntest_identity = pd.read_csv('../input/ieee-fraud-detection/test_identity.csv', index_col='TransactionID')\n\nsample_submission = pd.read_csv('../input/ieee-fraud-detection/sample_submission.csv', index_col='TransactionID')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We want to both of data -\"identity\" and \"transaction\"-\n\nThe identity data should be added to transaction data. "},{"metadata":{"trusted":true},"cell_type":"code","source":"train = train_transaction.merge(train_identity, how='left', left_index=True, right_index=True)\ntest = test_transaction.merge(test_identity, how='left', left_index=True, right_index=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# How V-feature is?"},{"metadata":{"trusted":true},"cell_type":"code","source":"train.columns[54:393]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"V features have a lot of similar columns.\n\nLet's look at it."},{"metadata":{"trusted":true},"cell_type":"code","source":"train.iloc[:,54:393].corr()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Feature engineering\nI picked up these 3 pairs of columns. "},{"metadata":{"trusted":true},"cell_type":"code","source":"train.loc[:,[\"V319\",\"V320\"]].corr()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.loc[:,[\"V109\",\"V110\"]].corr()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.loc[:,[\"V329\",\"V330\"]].corr()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.loc[:,[\"V316\",\"V331\"]].corr()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.loc[:,[\"V4\",\"V5\"]].corr()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"They seem to be similar."},{"metadata":{"trusted":true},"cell_type":"code","source":"train[\"diff_V319_V320\"] = np.zeros(train.shape[0])\n\ntrain.loc[train[\"V319\"]!=train[\"V320\"],\"diff_V319_V320\"] = 1\n\ntest[\"diff_V319_V320\"] = np.zeros(test.shape[0])\n\ntest.loc[test[\"V319\"]!=test[\"V320\"],\"diff_V319_V320\"] = 1\n\ntrain[\"diff_V109_V110\"] = np.zeros(train.shape[0])\n\ntrain.loc[train[\"V109\"]!=train[\"V110\"],\"diff_V109_V110\"] = 1\n\ntest[\"diff_V109_V110\"] = np.zeros(test.shape[0])\n\ntest.loc[test[\"V109\"]!=test[\"V110\"],\"diff_V109_V110\"] = 1\n\ntrain[\"diff_V329_V330\"] = np.zeros(train.shape[0])\n\ntrain.loc[train[\"V329\"]!=train[\"V330\"],\"diff_V329_V330\"] = 1\n\ntest[\"diff_V329_V330\"] = np.zeros(test.shape[0])\n\ntest.loc[test[\"V329\"]!=test[\"V330\"],\"diff_V329_V330\"] = 1\n\n\ntrain[\"diff_V316_V331\"] = np.zeros(train.shape[0])\n\ntrain.loc[train[\"V331\"]!=train[\"V316\"],\"diff_V316_V331\"] = 1\n\ntest[\"diff_V316_V331\"] = np.zeros(test.shape[0])\n\ntest.loc[test[\"V316\"]!=test[\"V331\"],\"diff_V316_V331\"] = 1\n\n\ntrain[\"diff_V4_V5\"] = np.zeros(train.shape[0])\n\ntrain.loc[train[\"V4\"]!=train[\"V5\"],\"diff_V4_V5\"] = 1\n\ntest[\"diff_V4_V5\"] = np.zeros(test.shape[0])\n\ntest.loc[test[\"V4\"]!=test[\"V5\"],\"diff_V4_V5\"] = 1","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# look at appearance of this new feature "},{"metadata":{},"cell_type":"markdown","source":"## V319-V320"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.bar(train.groupby(\"diff_V319_V320\").mean().isFraud.index,train.groupby(\"diff_V319_V320\").mean().isFraud.values)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## V109-V110"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.bar(train.groupby(\"diff_V109_V110\").mean().isFraud.index,train.groupby(\"diff_V109_V110\").mean().isFraud.values)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## V329-V330"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.bar(train.groupby(\"diff_V329_V330\").mean().isFraud.index,train.groupby(\"diff_V329_V330\").mean().isFraud.values)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## V316-V331"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.bar(train.groupby(\"diff_V316_V331\").mean().isFraud.index,train.groupby(\"diff_V316_V331\").mean().isFraud.values)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## V4-V5"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.bar(train.groupby(\"diff_V4_V5\").mean().isFraud.index,train.groupby(\"diff_V4_V5\").mean().isFraud.values)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There seem to be difference, but the gap of \"diff_V109_V110\" is small.\n\nIn Version7, I found that \"diff_V109_V110\" is not meaningful. I deleted.\n\nIn Version13, I found that \"diff_V329_V330\" is not meaningful. I deleted.\n\nIn Version20, I found that \"diff_V4_V5\" is not meaningful. I deleted.\n\nI feel that when the gap is big, the column is meaningful."},{"metadata":{"trusted":true},"cell_type":"code","source":"train = train.drop(\"diff_V109_V110\",axis=1)\ntest = test.drop(\"diff_V109_V110\",axis=1)\n\ntrain = train.drop(\"diff_V329_V330\",axis=1)\ntest = test.drop(\"diff_V329_V330\",axis=1)\n\ntrain = train.drop(\"diff_V316_V331\",axis=1)\ntest = test.drop(\"diff_V316_V331\",axis=1)\n\n\ntrain = train.drop(\"diff_V4_V5\",axis=1)\ntest = test.drop(\"diff_V4_V5\",axis=1)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# data cleaning"},{"metadata":{"trusted":true},"cell_type":"code","source":"print(train.shape)\nprint(test.shape)\n\ny_train = train['isFraud'].copy()\n\n# Drop target, fill in NaNs\nX_train = train.drop('isFraud', axis=1)\nX_test = test.copy()\nX_train = X_train.fillna(-999)\nX_test = X_test.fillna(-999)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"del train, test, train_transaction, train_identity, test_transaction, test_identity\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Label Encoding\nWe cannot use literal features for XGB, so these features are changes.\n\nFor example, [H,G,W,A] →[0,1,2,3]\n\nThe number of the words is often related to the numeral([0,1,2,3])."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Label Encoding\nfor f in X_train.columns:\n    if X_train[f].dtype=='object' or X_test[f].dtype=='object': \n        lbl = preprocessing.LabelEncoder()\n        lbl.fit(list(X_train[f].values) + list(X_test[f].values))\n        X_train[f] = lbl.transform(list(X_train[f].values))\n        X_test[f] = lbl.transform(list(X_test[f].values))   ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Model"},{"metadata":{"trusted":true},"cell_type":"code","source":"clf = xgb.XGBClassifier(n_estimators=500,\n                        n_jobs=4,\n                        max_depth=9,\n                        learning_rate=0.05,\n                        subsample=0.9,\n                        colsample_bytree=0.9,\n                        missing=-999)\n\nclf.fit(X_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_submission['isFraud'] = clf.predict_proba(X_test)[:,1]\nsample_submission.to_csv('simple_xgboost.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Is it useful?\nThe score of V5 is 0.9367.\n\nThe official score was 0.9366.\n\nThe score seem to be  improved.\n\nThe score of V8(V109-V110 added) is not unknown. Soon, I tell you it. ←the diff V109-V110 does not seem to be useful.\n\nAlso, the diff V329-V330 does not seem to be useful."},{"metadata":{},"cell_type":"markdown","source":"# Conclusion\nSome people think this is useful, but others not.\n\nIf you interested in it, please use for your model and judge wheter these columns are useful for your model."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}