{"cells":[{"metadata":{},"cell_type":"markdown","source":"I was curious to see if we can directly apply unsupervised learning to a two classes competition. \nIn this Kernel, the unsupervised learning KNN algorithm was applied to test data from TalkingDataAdTracking Fraud.\nThe KNN was applied directly to test data and it was tried to find two distinct classes.\nA submit file with all equal class 1 was used to find out the ration of classes in the test data. As it was expected the test data contained half class 1 and the other half class 0.\nAlthough the KNN divided the test data into two classes with almost the same numbers, the results were not promising (52%).\nI thought to share this Kernel with the fellow here at Kaggle (Kaggeleres). \nAny feedback is appreciated."},{"metadata":{},"cell_type":"markdown","source":"# libraries"},{"metadata":{"trusted":true},"cell_type":"code","source":"# import statements\nfrom sklearn.datasets import make_blobs\nimport numpy as np\nimport matplotlib.pyplot as plt\n# import KMeans\nfrom sklearn.cluster import KMeans\nimport pandas as pd\nfrom matplotlib import pyplot as plt\nfrom matplotlib import pyplot","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Functions"},{"metadata":{"trusted":true},"cell_type":"code","source":"def Parse_time(df):\n    df['day'] = df['click_time'].dt.day.astype('uint8')\n    df['hour'] = df['click_time'].dt.hour.astype('uint8')\n    df['minute'] = df['click_time'].dt.minute.astype('uint8')\n    df['second'] = df['click_time'].dt.second.astype('uint8')\n    \ndef Drop_cols(df, x):\n    Num_of_line = 100\n    print(Num_of_line*'=')\n    print('Before drop =\\n', df.head(3))\n    print(Num_of_line*'=')\n    df.drop(labels = x, axis = 1, inplace = True)\n    print('After drop =\\n', df.head(3))\n    return df","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Read data"},{"metadata":{"trusted":true},"cell_type":"code","source":"address_test = '../input/talkingdata-adtracking-fraud-detection/test.csv'\ndf_test = pd.read_csv(address_test, parse_dates=['click_time'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Parse data and drop columns"},{"metadata":{"trusted":true},"cell_type":"code","source":"Parse_time(df_test)\ncolmn_names = [\"click_time\", \"click_id\", \"ip\"]\ndf_test = Drop_cols(df_test, colmn_names); df_test.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Fit KNN clustering model"},{"metadata":{"trusted":true},"cell_type":"code","source":"# create kmeans object\nkmeans = KMeans(n_clusters=2)\n# fit kmeans object to data\nkmeans.fit(df_test)\n# print location of clusters learned by kmeans object\nprint(kmeans.cluster_centers_)\n# save new clusters for chart\ny_km = kmeans.fit_predict(df_test)\n\n\npredict = pd.DataFrame(y_km)\ndata_to_submit = pd.DataFrame()\ndata_to_submit['click_id'] = range(0, len(df_test))\ndata_to_submit['is_attributed'] = predict\nprint('data_to_submit = \\n', data_to_submit.head(5))\npyplot.hist(data_to_submit['is_attributed'], log = True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Save submit data"},{"metadata":{"trusted":true},"cell_type":"code","source":"data_to_submit.to_csv('Unsuper_csv_to_submit.csv', index = False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"See the submit data"},{"metadata":{"trusted":true},"cell_type":"code","source":"data_to_submit","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}