{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# `Understand Basic of PCA by Implementating it`","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:25.899550Z","iopub.execute_input":"2022-08-04T13:05:25.900051Z","iopub.status.idle":"2022-08-04T13:05:25.905599Z","shell.execute_reply.started":"2022-08-04T13:05:25.899995Z","shell.execute_reply":"2022-08-04T13:05:25.904781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load MNIST Data (train.csv)\nd0 = pd.read_csv('/kaggle/input/digit-recognizer/train.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:26.017531Z","iopub.execute_input":"2022-08-04T13:05:26.017869Z","iopub.status.idle":"2022-08-04T13:05:28.530364Z","shell.execute_reply.started":"2022-08-04T13:05:26.017823Z","shell.execute_reply":"2022-08-04T13:05:28.529647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print first five rows of d0\nd0.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:28.531780Z","iopub.execute_input":"2022-08-04T13:05:28.532153Z","iopub.status.idle":"2022-08-04T13:05:28.550318Z","shell.execute_reply.started":"2022-08-04T13:05:28.532122Z","shell.execute_reply":"2022-08-04T13:05:28.549653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# save the labels into a variable l.\n\nl = d0.label","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:28.552120Z","iopub.execute_input":"2022-08-04T13:05:28.552765Z","iopub.status.idle":"2022-08-04T13:05:28.557990Z","shell.execute_reply.started":"2022-08-04T13:05:28.552726Z","shell.execute_reply":"2022-08-04T13:05:28.557090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Drop the label feature from d0 and store the pixel data in d\nd = d0.drop('label',axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:28.559805Z","iopub.execute_input":"2022-08-04T13:05:28.560163Z","iopub.status.idle":"2022-08-04T13:05:28.677645Z","shell.execute_reply.started":"2022-08-04T13:05:28.560132Z","shell.execute_reply":"2022-08-04T13:05:28.676566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#print shape of pixel and label data\nprint(l.shape)\nprint(d.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:28.679023Z","iopub.execute_input":"2022-08-04T13:05:28.679376Z","iopub.status.idle":"2022-08-04T13:05:28.684892Z","shell.execute_reply.started":"2022-08-04T13:05:28.679317Z","shell.execute_reply":"2022-08-04T13:05:28.684129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"idx = 1\n# print label value for index 1\nprint(l[idx])","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:28.686135Z","iopub.execute_input":"2022-08-04T13:05:28.686583Z","iopub.status.idle":"2022-08-04T13:05:28.699763Z","shell.execute_reply.started":"2022-08-04T13:05:28.686523Z","shell.execute_reply":"2022-08-04T13:05:28.698752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Display or Plot above label","metadata":{}},{"cell_type":"code","source":"d.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:28.701289Z","iopub.execute_input":"2022-08-04T13:05:28.702229Z","iopub.status.idle":"2022-08-04T13:05:28.725684Z","shell.execute_reply.started":"2022-08-04T13:05:28.702179Z","shell.execute_reply":"2022-08-04T13:05:28.724929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(2,2))\n\n# reshape d from 1d to 2d pixel array for given idx ( prefer 28 X 28)\ngrid_data = d.loc[idx].values.reshape(28,28)\n\n#plot above grid image with cmap as gray and interpoltion as none\nplt.imshow(grid_data,interpolation='none',cmap='gray')\n\n#display plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:28.726788Z","iopub.execute_input":"2022-08-04T13:05:28.727413Z","iopub.status.idle":"2022-08-04T13:05:28.913793Z","shell.execute_reply.started":"2022-08-04T13:05:28.727376Z","shell.execute_reply":"2022-08-04T13:05:28.912914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2D Visualization using PCA","metadata":{}},{"cell_type":"code","source":"# Pick first 15K data-points to work on for time-effeciency\n#Excercise: Perform the same analysis on all of 42K data-points\n\nlabels = l.head(15000)#labels with 15k data points\ndata = d.head(15000)#data with 15k data points\n\nprint(\"the shape of sample data = \",data.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:28.915374Z","iopub.execute_input":"2022-08-04T13:05:28.915723Z","iopub.status.idle":"2022-08-04T13:05:28.922119Z","shell.execute_reply.started":"2022-08-04T13:05:28.915688Z","shell.execute_reply":"2022-08-04T13:05:28.921028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:28.925900Z","iopub.execute_input":"2022-08-04T13:05:28.926268Z","iopub.status.idle":"2022-08-04T13:05:28.948255Z","shell.execute_reply.started":"2022-08-04T13:05:28.926223Z","shell.execute_reply":"2022-08-04T13:05:28.947383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:28.949621Z","iopub.execute_input":"2022-08-04T13:05:28.950100Z","iopub.status.idle":"2022-08-04T13:05:28.963062Z","shell.execute_reply.started":"2022-08-04T13:05:28.950067Z","shell.execute_reply":"2022-08-04T13:05:28.962236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data-preprocessing\n### Standardizing data","metadata":{}},{"cell_type":"code","source":"# import standard scalar\nfrom sklearn.preprocessing import StandardScaler\n\n#fit transform data\nstandardized_data = StandardScaler().fit_transform(data)\n\n#print shape of standardized_data\ndata.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:28.965948Z","iopub.execute_input":"2022-08-04T13:05:28.966685Z","iopub.status.idle":"2022-08-04T13:05:29.302793Z","shell.execute_reply.started":"2022-08-04T13:05:28.966634Z","shell.execute_reply":"2022-08-04T13:05:29.302086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Find the co-variance matrix which is : \n\n<h1>$$A^{T}  * A^{'}$$ </h1>\n\n* Find covariance matrix of dataset by multiplying the matrix of features by its transpose \n* It is a measure of how much each of the dimensions vary from the mean with respect to each other\n\n\nCovariance is measured between 2 dimensions to see if there is a relationship between the 2 dimensions, e.g., relationship between height and weight of students\n\n* Positive value of covariance indicates that both dimensions are directly proportional to each other, where if one dimension increases the other dimension increases accordingly\n\n* Negative value of covariance indicates that both dimensions are indirectly proportional to each other, where if one dimension increases then other dimension decreases accordingly\n\n* If in case covariance is zero, then the two dimensions are independent of each other","metadata":{}},{"cell_type":"code","source":"sample_data = standardized_data\n\n#use matrix multiplication on sample_data using numpy to find covariance matrix\ncovar_matrix = np.matmul(sample_data.T,sample_data)\n\n\n#print shape of covar_matrix\nprint ( \"The shape of variance matrix \\n\",covar_matrix.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:29.304143Z","iopub.execute_input":"2022-08-04T13:05:29.304605Z","iopub.status.idle":"2022-08-04T13:05:29.508834Z","shell.execute_reply.started":"2022-08-04T13:05:29.304572Z","shell.execute_reply":"2022-08-04T13:05:29.507594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"covar_matrix","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:29.511185Z","iopub.execute_input":"2022-08-04T13:05:29.512089Z","iopub.status.idle":"2022-08-04T13:05:29.526696Z","shell.execute_reply.started":"2022-08-04T13:05:29.512025Z","shell.execute_reply":"2022-08-04T13:05:29.525420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Computing Eigenvectors and Eigenvalues\nEigenvectors and eigenvalues of a covariance (or correlation) matrix represent the “core” of a PCA: \n* Eigenvectors (principal components) determine the directions of new feature space\n* Eigenvalues determine their magnitude <br>\nIn other words,`eigenvalues explain variance of data along new feature axes`\n\nEigenvectors and Eigenvalues of covariance matrix will give the principal components and a vector that we can use to project high-dimensional inputs to lower-dimensional subspace","metadata":{}},{"cell_type":"code","source":"# finding the top two eigen-values and corresponding eigen-vectors \n# for projecting onto a 2-Dim space\n\nfrom scipy.linalg import eigh \n\n# the parameter 'eigvals' is defined (low value to heigh value) \n# eigh function will return the eigen values in asending order\n\n# this code generates only the top 2 (782 and 783) eigenvalues\nvalues, vectors = eigh(covar_matrix,eigvals=(782,783))\n\nprint(\"Shape of eigen vectors = \",vectors.shape)\nprint(vectors)\n\n# converting the eigen vectors into (2,d) shape for easyness of further computations\nvectors = vectors.T\nprint(\"Updated shape of eigen vectors = \",vectors.shape)\nprint(vectors)\n# here the vectors[1] represent the eigen vector corresponding 1st principal eigen vector\n# here the vectors[0] represent the eigen vector corresponding 2nd principal eigen vector","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:29.529221Z","iopub.execute_input":"2022-08-04T13:05:29.530351Z","iopub.status.idle":"2022-08-04T13:05:29.600914Z","shell.execute_reply.started":"2022-08-04T13:05:29.530278Z","shell.execute_reply":"2022-08-04T13:05:29.599714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**`projecting the original data sample on the plane formed by multiplication of two principal eigen vectors with transposed sample_data`**","metadata":{}},{"cell_type":"code","source":"# multiplication of two principal eigen vectors with transposed sample_data to get 2d projected data\nnew_coordinates = np.matmul(vectors,sample_data.T)\n\nprint (\" resultanat new data points' shape \",vectors.shape, \"X\",sample_data.T.shape,\" = \",new_coordinates.shape)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:29.603340Z","iopub.execute_input":"2022-08-04T13:05:29.604236Z","iopub.status.idle":"2022-08-04T13:05:29.633320Z","shell.execute_reply.started":"2022-08-04T13:05:29.604165Z","shell.execute_reply":"2022-08-04T13:05:29.631954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# appending label to the 2d projected data \nnew_coordinates = np.vstack((new_coordinates,labels)).T\n\n# creating a new data frame for ploting the labeled points.\ndataframe = pd.DataFrame(data=new_coordinates,columns=('1st_principal','2nd_principal','label'))\n\n#print dataframe head\ndataframe.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:29.636375Z","iopub.execute_input":"2022-08-04T13:05:29.637466Z","iopub.status.idle":"2022-08-04T13:05:29.664557Z","shell.execute_reply.started":"2022-08-04T13:05:29.637373Z","shell.execute_reply":"2022-08-04T13:05:29.663192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**`ploting the 2d data points with seaborn`**","metadata":{}},{"cell_type":"code","source":"sns.FacetGrid(dataframe,hue='label',height=8).map(plt.scatter,'1st_principal','2nd_principal','label').add_legend()\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:29.667290Z","iopub.execute_input":"2022-08-04T13:05:29.668321Z","iopub.status.idle":"2022-08-04T13:05:30.925991Z","shell.execute_reply.started":"2022-08-04T13:05:29.668248Z","shell.execute_reply":"2022-08-04T13:05:30.925032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# PCA using Scikit-Learn","metadata":{}},{"cell_type":"code","source":"# initializing the pca\nfrom sklearn import decomposition\n\n\npca = decomposition.PCA()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:30.927134Z","iopub.execute_input":"2022-08-04T13:05:30.927374Z","iopub.status.idle":"2022-08-04T13:05:30.932942Z","shell.execute_reply.started":"2022-08-04T13:05:30.927347Z","shell.execute_reply":"2022-08-04T13:05:30.931801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### configuring the parameteres","metadata":{"execution":{"iopub.status.busy":"2021-10-13T14:40:49.914461Z","iopub.execute_input":"2021-10-13T14:40:49.914723Z","iopub.status.idle":"2021-10-13T14:40:49.921235Z","shell.execute_reply.started":"2021-10-13T14:40:49.914698Z","shell.execute_reply":"2021-10-13T14:40:49.92046Z"}}},{"cell_type":"code","source":"# the number of components = 2\npca.n_components = 2\n\n# fit transform sample data using pca \npca_data = pca.fit_transform(sample_data)\n\n# pca_reduced will contain the 2-d projects of simple data\nprint(\"shape of pca_reduced.shape = \",pca_data.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:30.934230Z","iopub.execute_input":"2022-08-04T13:05:30.934473Z","iopub.status.idle":"2022-08-04T13:05:31.679414Z","shell.execute_reply.started":"2022-08-04T13:05:30.934445Z","shell.execute_reply":"2022-08-04T13:05:31.678187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" pd.DataFrame(data=np.vstack((pca_data.T,labels)).T,columns=('1st_principal','2nd_principal','label'))","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:31.681917Z","iopub.execute_input":"2022-08-04T13:05:31.682788Z","iopub.status.idle":"2022-08-04T13:05:31.708056Z","shell.execute_reply.started":"2022-08-04T13:05:31.682727Z","shell.execute_reply":"2022-08-04T13:05:31.707269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# attaching the label for each 2-d data point (Hint: Use np.vstack)\npca_data = np.vstack((pca_data.T,labels)).T\n\n# creating a new data fram which help us in ploting the result data\npca_df = pd.DataFrame(data=pca_data,columns=('1st_principal','2nd_principal','label'))\n\n#plotting the 2d data points\nsns.FacetGrid(pca_df,hue='label',height=8).map(plt.scatter,'1st_principal','2nd_principal','label').add_legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:31.709479Z","iopub.execute_input":"2022-08-04T13:05:31.710000Z","iopub.status.idle":"2022-08-04T13:05:32.931356Z","shell.execute_reply.started":"2022-08-04T13:05:31.709965Z","shell.execute_reply":"2022-08-04T13:05:32.930176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# PCA for dimensionality redcution (not for visualization)\nDistribution of explained variance for each principal component gives a sense of how much information will be represented and how much lost when the full, 64-dimensional input is reduced using a principal component model (i.e., a model that utilizes only the first N principal components).","metadata":{}},{"cell_type":"code","source":"# the number of components = 784\npca.n_components = 784\n\n\n# fit transform sample data using pca \npca_data = pca.fit_transform(sample_data)\n\n\n#calculating percentage of variance explained in the data\npercentage_var_explained = pca.explained_variance_ / np.sum(pca.explained_variance_)\n\n#cumulative sum of the percentage_var_explained\ncumulative_explained_variance = np.cumsum(percentage_var_explained)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:32.932666Z","iopub.execute_input":"2022-08-04T13:05:32.932925Z","iopub.status.idle":"2022-08-04T13:05:35.217484Z","shell.execute_reply.started":"2022-08-04T13:05:32.932894Z","shell.execute_reply":"2022-08-04T13:05:35.215873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot the PCA spectrum\nplt.figure(figsize=(6,4))\nplt.plot(cumulative_explained_variance,linewidth=4)\nplt.grid()\n\nplt.xlabel('n_components')\nplt.ylabel('Cumulative_explained_variance')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T13:05:35.220049Z","iopub.execute_input":"2022-08-04T13:05:35.220679Z","iopub.status.idle":"2022-08-04T13:05:35.455023Z","shell.execute_reply.started":"2022-08-04T13:05:35.220597Z","shell.execute_reply":"2022-08-04T13:05:35.454093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From above you can see that if we take 200-dimensions, approx. 90% of variance is expalined.\n\nOur intention with princpal component analysis is to reduce the high-dimensional input to a low-dimensional input \n\nUltimately that low-dimensional input is intended for use in a model, since adding more components increases the cost and the accuracy","metadata":{}},{"cell_type":"markdown","source":"# PCA is a method that brings together:\n* A measure of how each variable is associated with one another. (Covariance matrix.)\n\n* The directions in which our data are dispersed. (Eigenvectors.)\n\n* The relative importance of these different directions. (Eigenvalues.)\n\n* PCA combines our predictors and allows us to drop the Eigenvectors that are relatively unimportant.\n","metadata":{}},{"cell_type":"markdown","source":"----\n----\n\n# My Other Notebook:\n* [Simple Linear Regression](https://www.kaggle.com/mukeshmanral/linear-regression-basic)\n* [Multiple Linear Regression](https://www.kaggle.com/mukeshmanral/multiple-linear-regression-basic)\n* [Polynomial Regression](https://www.kaggle.com/mukeshmanral/polynomial-regression-basic)\n* [Advanced Linear Regression](https://www.kaggle.com/mukeshmanral/advance-linear-regression-basic-gridsearchcv-hpt)\n\n----\n\n* [Feature Engineering 1](https://www.kaggle.com/mukeshmanral/feature-engineering-diff-dataset-1)\n* [Feature Engineering 2](https://www.kaggle.com/mukeshmanral/feature-engineering-diff-dataset-2)\n* [Feature Engineering 3](https://www.kaggle.com/mukeshmanral/feature-engineering-diff-dataset-3)\n* [Feature Engineering 4](https://www.kaggle.com/mukeshmanral/feature-engineering-diff-dataset-4)\n\n----\n\n* [How KNN-Algorith(1) Works (Basic)](https://www.kaggle.com/mukeshmanral/k-nn-algorithm-1-basic)\n* [How KNN-Algorith(2) Works (Basic)](https://www.kaggle.com/mukeshmanral/k-nn-algorithm-2-basic)\n\n____\n\n* [Ensemble-Bagging-Random Forest-Extra Tree Basic](https://www.kaggle.com/mukeshmanral/bagging-ensemble-concept-rf-extree-basic) \n\n____\n____","metadata":{}},{"cell_type":"markdown","source":"<h1><center>Happy Learning </h1></center>","metadata":{}}]}