{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Only 3 features used. Limitting value enhances feature."},{"metadata":{},"cell_type":"markdown","source":"In this kernel, I will explain that limitting value make feature more useful in some case.  \nI will take \"std\" feature as an example.     \nWhen we ormit the 'acoustic_data' which is not in 0<=x<=10 ,  \nthe importance of the feature increased a lot.  \nIt is very simple code. So, maybe reading my code is easier to understand."},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport lightgbm as lgb\nimport eli5\nfrom eli5.sklearn import PermutationImportance\nfrom tsfresh.feature_extraction import feature_calculators\nfrom sklearn.model_selection import train_test_split\n\nparams={'bagging_fraction': 0.6364049179265991,\n        'bagging_freq': 17,\n        'feature_fraction': 0.8780002461376601,\n        'min_data_in_leaf': 100,\n        'num_leaves': 65,\n        'boost': 'gbdt',\n        'learning_rate': 0.01,\n        'max_depth': -1,\n        'metric': 'mae',\n        'num_threads': 4,\n        'tree_learner': 'serial',\n        'objective': 'huber',\n        'n_estimators': 100000}","execution_count":1,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Feature generation"},{"metadata":{"trusted":true},"cell_type":"code","source":"def cook_data(data):\n    output=pd.Series()\n    \n    if \"time_to_failure\" in data.columns:\n        output[\"target\"]=data[\"time_to_failure\"].iloc[-1]\n    \n    data=data[\"acoustic_data\"].values\n    output[\"std\"]=data.std()\n    \n    #Limit the range.\n    output[\"new_std\"]=data[np.logical_and(0<=data,data<=10)].std()\n    \n    #This feature is from public kernel.\n    output[\"numpeaks_10\"]=feature_calculators.number_peaks(data,10)\n    return output\n\ndef create_X_y():\n    reader = pd.read_csv(\"../input/train.csv\",chunksize=150000)\n    train=pd.DataFrame( [cook_data(r) for r in reader] )\n    y=train.pop(\"target\")\n    X=train\n    return X,y\n\n#80% for fit.\n#10% for early stopping.\n#10% for cv.\nX,y=create_X_y()\n(X_fit, _X,y_fit, _y) = train_test_split(X, y,train_size=0.8,test_size=0.2,random_state=0)\n(X_cv,X_es,y_cv,y_es) = train_test_split(_X, _y,train_size=0.5,test_size=0.5,random_state=0)\n\nmodel = lgb.LGBMRegressor(**params)\nmodel.fit(X_fit,y_fit,eval_set = [(X_es,y_es)],verbose = 5000,early_stopping_rounds=1000)\nperm = PermutationImportance(model, random_state=1).fit(X_cv,y_cv)\neli5.show_weights(perm, feature_names = X_cv.columns.tolist())","execution_count":2,"outputs":[{"output_type":"stream","text":"Training until validation scores don't improve for 1000 rounds.\nEarly stopping, best iteration is:\n[1155]\tvalid_0's l1: 2.09058\n","name":"stdout"},{"output_type":"execute_result","execution_count":2,"data":{"text/plain":"<IPython.core.display.HTML object>","text/html":"\n    <style>\n    table.eli5-weights tr:hover {\n        filter: brightness(85%);\n    }\n</style>\n\n\n\n    \n\n    \n\n    \n\n    \n\n    \n\n    \n\n\n    \n\n    \n\n    \n\n    \n\n    \n\n    \n\n\n    \n\n    \n\n    \n\n    \n\n    \n        <table class=\"eli5-weights eli5-feature-importances\" style=\"border-collapse: collapse; border: none; margin-top: 0em; table-layout: auto;\">\n    <thead>\n    <tr style=\"border: none;\">\n        <th style=\"padding: 0 1em 0 0.5em; text-align: right; border: none;\">Weight</th>\n        <th style=\"padding: 0 0.5em 0 0.5em; text-align: left; border: none;\">Feature</th>\n    </tr>\n    </thead>\n    <tbody>\n    \n        <tr style=\"background-color: hsl(120, 100.00%, 80.00%); border: none;\">\n            <td style=\"padding: 0 1em 0 0.5em; text-align: right; border: none;\">\n                0.5411\n                \n                    &plusmn; 0.0716\n                \n            </td>\n            <td style=\"padding: 0 0.5em 0 0.5em; text-align: left; border: none;\">\n                new_std\n            </td>\n        </tr>\n    \n        <tr style=\"background-color: hsl(120, 100.00%, 94.99%); border: none;\">\n            <td style=\"padding: 0 1em 0 0.5em; text-align: right; border: none;\">\n                0.0749\n                \n                    &plusmn; 0.0281\n                \n            </td>\n            <td style=\"padding: 0 0.5em 0 0.5em; text-align: left; border: none;\">\n                numpeaks_10\n            </td>\n        </tr>\n    \n        <tr style=\"background-color: hsl(120, 100.00%, 98.65%); border: none;\">\n            <td style=\"padding: 0 1em 0 0.5em; text-align: right; border: none;\">\n                0.0116\n                \n                    &plusmn; 0.0062\n                \n            </td>\n            <td style=\"padding: 0 0.5em 0 0.5em; text-align: left; border: none;\">\n                std\n            </td>\n        </tr>\n    \n    \n    </tbody>\n</table>\n    \n\n    \n\n\n    \n\n    \n\n    \n\n    \n\n    \n\n    \n\n\n\n"},"metadata":{}}]},{"metadata":{},"cell_type":"markdown","source":"## Conclusion\nComparing to \"std\", \"new_std\" is much more useful.   \nWe can also see that \"new_std\" is more useful than \"numpeaks_10\"  \nwhich is known as a strong feature in public kernels.  \nI think limitting value is worth doing.  　"},{"metadata":{},"cell_type":"markdown","source":"## Submission"},{"metadata":{"trusted":true},"cell_type":"code","source":"def create_prediction():\n    model = lgb.LGBMRegressor(**params)\n    model.fit(X_fit,y_fit,eval_set = [(X_es,y_es)],verbose = 5000,early_stopping_rounds=1000)\n    submission=pd.read_csv('../input/sample_submission.csv')\n    predictions=[cook_data(pd.read_csv('../input/test/'+s+'.csv')) for s in submission[\"seg_id\"]]\n    submission[\"time_to_failure\"]=model.predict(pd.DataFrame(predictions),num_iteration=model.best_iteration_)\n    submission.to_csv(\"submission.csv\",index=False)\n\ncreate_prediction()","execution_count":3,"outputs":[{"output_type":"stream","text":"Training until validation scores don't improve for 1000 rounds.\nEarly stopping, best iteration is:\n[1155]\tvalid_0's l1: 2.09058\n","name":"stdout"}]},{"metadata":{},"cell_type":"markdown","source":"Thank you for reading my kernel.  \nI hope this kernel helps your score better.\n"},{"metadata":{},"cell_type":"markdown","source":" \n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}