{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"**Thoughts:**\n\n* Distributed evenly in first 1-2 chars\n\n* First 3 or 4 chars of qid seems interesting (some id gives super high probability for prediction)... \n\n* Larger than 4 seems to be unique which is useless\n\n* If I print out some raw data, there is not nothing really interesting. The question is how qid was generated: related to topic or time or just randomly\n\n* Just to have some fun with qid. Be careful before using first 3 or 4 chars as feature ..."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"import pandas as pd\ndf_train = pd.read_csv(\"../input/train.csv\")\ndf_test = pd.read_csv(\"../input/test.csv\")\ndf_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e6eb901b131048d579b23fb8452c9fc000ac84b7"},"cell_type":"code","source":"df = df_train[['qid', 'question_text']].append(df_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f14f73f07c179190c698f1afe8bffa98584fa5fc"},"cell_type":"code","source":"df[df.duplicated(['qid'], keep=False)]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c82fccbb17e58c611d6dbc14f338d11d41a34b48"},"cell_type":"code","source":"for i in range(1, 21):\n    df_train['first_' + str(i)] = df_train['qid'].apply(lambda x: x[:i])\n    \ndf_train.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c3a6530fea02b15d5dfe61122a659c9b72c4f9a6"},"cell_type":"markdown","source":"**First char**"},{"metadata":{"trusted":true,"_uuid":"df2e4c6608148a45c27c46029a0d5c70716bcfb5"},"cell_type":"code","source":"df_train.groupby('first_1')['target'].agg(['count', 'mean']).reset_index().sort_values('mean', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"742749b70517bd3516d8d2a212839477f10ffaf0"},"cell_type":"markdown","source":"**Frist 2 chars**"},{"metadata":{"trusted":true,"_uuid":"1bf874abce42324f2a37f3d330fb893e1b111530"},"cell_type":"code","source":"df_train.groupby('first_2')['target'].agg(['count', 'mean']).reset_index().sort_values('mean', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"68cd3c8055fa129288d033a1fe9bab411d2c0597"},"cell_type":"markdown","source":"**Frist 3 chars**"},{"metadata":{"trusted":true,"_uuid":"6ab1f22b81f02d0c72e9b01087ce5a1ae9cc007d"},"cell_type":"code","source":"df_train.groupby('first_3')['target'].agg(['count', 'mean']).reset_index().sort_values('mean', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"81b61ff5bba1e1495d53f4af8bf8beaefc37e3eb"},"cell_type":"markdown","source":"**First 4 chars**"},{"metadata":{"trusted":true,"_uuid":"9232ac9ffa5bdeec2a8f2a0948c49426bf15e81c"},"cell_type":"code","source":"df_train.groupby('first_4')['target'].agg(['count', 'mean']).reset_index().sort_values('mean', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f3a6db566c4fa1fdd934ca72b5f687392cdd4c53"},"cell_type":"markdown","source":"**First 5 chars**"},{"metadata":{"trusted":true,"_uuid":"d93abf27a9780f94b713006171e7af897b96dca3"},"cell_type":"code","source":"df_train.groupby('first_5')['target'].agg(['count', 'mean']).reset_index().sort_values('mean', ascending=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dba9022a5d8952b49d374b9d770d511ede6814fd"},"cell_type":"markdown","source":"**Check the raw data**"},{"metadata":{"trusted":true,"_uuid":"79e7dbea708466566b8e811dcd6efe878a631f76"},"cell_type":"code","source":"df_2122 = df_train[df_train.first_4 == '2122']\nfor index, row in df_2122.iterrows():\n    print(row['target'], ': ',row['question_text'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ae20d2200602d9330691c08a6200c56ddac025c7"},"cell_type":"code","source":"df_623a = df_train[df_train.first_4 == '623a']\nfor index, row in df_623a.iterrows():\n    print(row['target'], ': ',row['question_text'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2d453b19cd0603615ed559a9845814dae8ea9406"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"108c868901a1ef206f99a4a25ade95231e34a109"},"cell_type":"markdown","source":"Do some prediction:"},{"metadata":{"trusted":true,"_uuid":"a81aa88decafa47104dd2458bf639c8b3eb1fd3c"},"cell_type":"code","source":"df_test['first_4'] = df_test['qid'].apply(lambda x: x[:4])\nstat = df_train.groupby('first_4')['target'].agg(['mean']).reset_index()\nresult = pd.merge(df_test, stat, how='left', on=['first_4'])\nresult.head(n=100)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"aae9adb5218b9e6ba37f1646928e5983454d400a"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"729d5f5afe37e9321ab2db18d9f8d33cd4fe83a1"},"cell_type":"markdown","source":"**Check the 20th row:**\n\nIs a decision tree better than logistic regres...  and 0.000000"},{"metadata":{"_uuid":"fe6820eedcc6b4a5108d74fa06c096a3ae656b70"},"cell_type":"markdown","source":"**I don't have courage to play with it further :)**"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}