{"cells":[{"metadata":{},"cell_type":"markdown","source":"In this notebook we'll do dimensionality reduction and visualization of the FFT features that were first used in this competition in [this Giba's notebook](https://www.kaggle.com/titericz/0-309-baseline-logisticregression-using-fft). I've created a stand-alone notebook that extracts those features, and it can be found [here](https://www.kaggle.com/tunguz/giba-s-fft-features-only).\n\nWe will make this visualization notebook with the Rapids library. [Rapids](https://rapids.ai) is an open-source GPU accelerated Data Sceince and Machine Learning library, developed and mainatained by [Nvidia](https://www.nvidia.com). It is designed to be compatible with many existing CPU tools, such as Pandas, scikit-learn, numpy, etc. It enables **massive** acceleration of many data-science and machine learning tasks, oftentimes by a factor fo 100X, or even more. \n\nRapids is still undergoing developemnt, and only recently has it become possible to use RAPIDS natively in the Kaggle Docker environment. If you are interested in installing and riunning Rapids locally on your own machine, then you should [refer to the followong instructions](https://rapids.ai/start.html)."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import cupy as cp\nimport cudf, cuml\nimport pandas as pd\nimport numpy as np\nfrom cuml.manifold import TSNE, UMAP\nimport matplotlib.pyplot as plt\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train = cp.load('../input/giba-s-fft-features-only/TRAIN.npy')\ntest = cp.load(\"../input/giba-s-fft-features-only/TEST.npy\")\nTRAIN_TAB = cudf.read_csv(\"../input/giba-s-fft-features-only/TRAIN_TAB.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"TRAIN_TAB.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"TARGET_VALUES = TRAIN_TAB.iloc[:,2:].values\nTARGET_VALUES = cp.asnumpy(TARGET_VALUES)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"TARGET_VALUES[:,0]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"First, we are going to combine train and test to try to visualize the overall shape of reduced data."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_test = cp.vstack([train, test])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\ntsne = TSNE(n_components=2)\ntrain_test_2D = tsne.fit_transform(train_test)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Well, that only took a few seconds!\n\nIn order to visualize the new reduced dataset, we'll need to convert it into a numpy array, as matplotlibe does not work on GPUs. "},{"metadata":{"trusted":true},"cell_type":"code","source":"train_test_2D = cp.asnumpy(train_test_2D)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now let's take a look at the data"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.scatter(train_test_2D[:,0], train_test_2D[:,1], s = 0.5)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are some hints of the structure, but nothing too dramatic. Maybe this is not surprizing; after all, the deataset is supposed to represent 24 different sound categories. \n\nNow let's look at what the dataset looks with UMAP dimensionality reduction."},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\numap = UMAP(n_components=2)\ntrain_test_2D = umap.fit_transform(train_test)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"That was even faster! \n\nLet's see what this dimensionality reduction looks like."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_test_2D = cp.asnumpy(train_test_2D)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.scatter(train_test_2D[:,0], train_test_2D[:,1], s = 0.5)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There seems to be more of a structure with this reduction, but still nothign dramatic, at least not at this scale. UMAP usually produces more outlyers, which tend to shrink make the bulk of the datapoints into small fraction of the visual representation."},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We'll now make dimensionality reductions for the train set only, and take a look how the distributions look with respect to the target."},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\ntsne = TSNE(n_components=2)\numap = UMAP(n_components=2)\ntrain_2D_tsne = tsne.fit_transform(train)\ntrain_2D_umap = umap.fit_transform(train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_2D_tsne = cp.asnumpy(train_2D_tsne)\ntrain_2D_umap = cp.asnumpy(train_2D_umap)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"y = TARGET_VALUES[:,0] #cp.asnumpy(TRAIN_TAB['s0'].values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.scatter(train_2D_tsne[:,0], train_2D_tsne[:,1], c = y, s = 0.5)\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It's very hard to see, but all of the yellow dots seem to be concentrated in the upper middle area."},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.scatter(train_2D_umap[:,0], train_2D_umap[:,1], c = y, s = 0.5)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, axs = plt.subplots(2, 2)\nfor i in range(2):\n    for j in range(2):\n        axs[i,j].scatter(train_2D_tsne[:,0], train_2D_tsne[:,1], c = TARGET_VALUES[:,i*2+j], s = 1.0)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, axs = plt.subplots(2, 2)\nfor i in range(2):\n    for j in range(2):\n        axs[i,j].scatter(train_2D_umap[:,0], train_2D_umap[:,1], c = TARGET_VALUES[:,i*2+j], s = 1.0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}