{"cells":[{"metadata":{},"cell_type":"markdown","source":"# UMAP\n\nI noticed @tunguz's notebook didn't include UMAP, which is something I use a lot at work for quickly visualising datasets so I thought I would quickly expand on his work\n\n\nAnother small thing I've added is to colour the train and test data differently (blue and orange). It's hardly scientific, but my first impression here is the train and test sets look fairly similar and if there is a difference it seems there isn't much in the test that isn't in the train set. Obviously that's assuming the fft doesn't remove weird things :)\n\nThere's nothing stopping you from also doing this to the labels ;)\n\n----------------"},{"metadata":{},"cell_type":"markdown","source":"In this notebook we'll do dimensionality reduction and visualization of the FFT features that were first used in this competition in [this Giba's notebook](https://www.kaggle.com/titericz/0-309-baseline-logisticregression-using-fft). I've created a stand-alone notebook that extracts those features, and it can be found [here](https://www.kaggle.com/tunguz/giba-s-fft-features-only).\n\nWe will make this visualization notebook with the Rapids library. [Rapids](https://rapids.ai) is an open-source GPU accelerated Data Sceince and Machine Learning library, developed and mainatained by [Nvidia](https://www.nvidia.com). It is designed to be compatible with many existing CPU tools, such as Pandas, scikit-learn, numpy, etc. It enables **massive** acceleration of many data-science and machine learning tasks, oftentimes by a factor fo 100X, or even more. \n\nRapids is still undergoing developemnt, and only recently has it become possible to use RAPIDS natively in the Kaggle Docker environment. If you are interested in installing and riunning Rapids locally on your own machine, then you should [refer to the followong instructions](https://rapids.ai/start.html)."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import cupy as cp\nimport cudf, cuml\nimport pandas as pd\nimport numpy as np\nfrom cuml.manifold import TSNE, UMAP\nimport matplotlib.pyplot as plt\nfrom sklearn.metrics import silhouette_score\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train = cp.load('../input/giba-s-fft-features-only/TRAIN.npy')\ntest = cp.load(\"../input/giba-s-fft-features-only/TEST.npy\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_test = cp.vstack([train, test])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# UMAP\n\nThe default number of neighbors (`n_neighbors`) used in UMAP is 15, playing with it will give you different results. I've typically found reducing it gives interesting results while increasing it gives you a blob with less structure. Perhaps that's just a reflection of the size of the datasets I'm playing with"},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\numap = UMAP(n_components=2, n_neighbors=15, random_state=42)\ntrain_test_2D = umap.fit_transform(train_test)\ntrain_test_2D = cp.asnumpy(train_test_2D)\n\ntrain_2D = train_test_2D[:train.shape[0], :]\ntest_2D = train_test_2D[train.shape[0]:, :]\n\nplt.figure(figsize=(10,10))\nplt.scatter(train_2D[:,0], train_2D[:,1], alpha=0.3)\nplt.scatter(test_2D[:,0], test_2D[:,1], alpha=0.3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## With n_neighbors=10"},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\numap = UMAP(n_components=2, n_neighbors=10, random_state=42)\ntrain_test_2D = umap.fit_transform(train_test)\ntrain_test_2D = cp.asnumpy(train_test_2D)\n\ntrain_2D = train_test_2D[:train.shape[0], :]\ntest_2D = train_test_2D[train.shape[0]:, :]\n\nplt.figure(figsize=(10,10))\nplt.scatter(train_2D[:,0], train_2D[:,1], alpha=0.3)\nplt.scatter(test_2D[:,0], test_2D[:,1], alpha=0.3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## with n_neighbors=7"},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\numap = UMAP(n_components=2, n_neighbors=7, random_state=42)\ntrain_test_2D = umap.fit_transform(train_test)\ntrain_test_2D = cp.asnumpy(train_test_2D)\n\ntrain_2D = train_test_2D[:train.shape[0], :]\ntest_2D = train_test_2D[train.shape[0]:, :]\n\nplt.figure(figsize=(10,10))\nplt.scatter(train_2D[:,0], train_2D[:,1], alpha=0.3)\nplt.scatter(test_2D[:,0], test_2D[:,1], alpha=0.3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## With n_neighbors=5\n\nBelow this you basically lose all structure again"},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\numap = UMAP(n_components=2, n_neighbors=5, random_state=42)\ntrain_test_2D = umap.fit_transform(train_test)\ntrain_test_2D = cp.asnumpy(train_test_2D)\n\ntrain_2D = train_test_2D[:train.shape[0], :]\ntest_2D = train_test_2D[train.shape[0]:, :]\n\nplt.figure(figsize=(10,10))\nplt.scatter(train_2D[:,0], train_2D[:,1], alpha=0.3)\nplt.scatter(test_2D[:,0], test_2D[:,1], alpha=0.3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}