{"cells":[{"metadata":{"_uuid":"470f749fedf16f2c6b3a624ec2b058cc1db4cb2a"},"cell_type":"markdown","source":"I am trying to  go about with the MNIST data classification task using Python.\n\nAlthough I could have forked existing kernels for this task, I really wanted to build a kernel ground up and this seemed the easiest data to work off. Also I am using [this great kernel as my source](https://www.kaggle.com/archaeocharlie/a-beginner-s-approach-to-classification). \nI started with creating a new kernel and then adding the data under 'Draft Environment' and selecting the Digit Recognizer data. \n\n\n"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"\nimport os\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt,matplotlib.image as mpimg\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.svm import SVC\n#%matplotlib is a magic function in IPython\n#%matplotlib inline sets the backend of \n#matplotlib to the 'inline' backend: \n#With this backend, the output of plotting commands\n#is displayed inline within frontends like the Jupyter notebook, \n#directly below the code cell that produced it.\n%matplotlib inline\n\n#constants\nNUMBER_OF_TRAINING_IMGS=5000\nIMG_HEIGHT=28\nIMG_WIDTH=28\n#list all the contents of the ../input directory\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"11411336b0f523535477ac0392598420e1796b50"},"cell_type":"markdown","source":"Load the training data in and create a dataframe from the csv. \nThe first column in the ground truth label. The remaining are pixel values\nThe images are 28X28 pixels and these pixels are unrolled. See the data section to see more details.\n\nThe data is of the format: \nLabel, pixel0, pixel1, pixel2,..., pixel783\n\nLoading the data is done as follows:\n1. Read the csv using pandas read_csv to read into a dataframe\n2. separate into images ( design matrix X) and labels (Y)\n3. Split into training and test sets\n"},{"metadata":{"trusted":true,"_uuid":"e7da2a024480e620a6c2ed3bd6756745542659e1"},"cell_type":"code","source":"labeled_images=pd.read_csv('../input/train.csv')\nprint('shape of the dataframe ',labeled_images.shape)\nlabeled_images.head(n=3)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"98cadb567f2e81758b2a9be377ed51bdaa1bee6f"},"cell_type":"markdown","source":"*pandas loc is used to subset rows by labels while iloc is used to subset rows by row numbers.  \nBy default(?) the dataframe rows are assigned labels set to the row numbers.   \nLabels need not always be equal to the row numbers in all cases.   \n(row_number->label   0->A,1->B,2->C  or 0->Z,1->Y,2->X and so on).   \nloc => labels  \niloc => row numbers*"},{"metadata":{"trusted":true,"_uuid":"977aceb74e42eabda97b1a3707307e95f6c0b49b"},"cell_type":"code","source":"images=labeled_images.iloc[0:NUMBER_OF_TRAINING_IMGS,1:] # first NUMBER_OF_TRAINING_IMGS rows,column 2 onwards.\nlabels=labeled_images.iloc[0:NUMBER_OF_TRAINING_IMGS,:1] #first NUMBER_OF_TRAINING_IMGS rows, first column. \n                                                        #I could have used .iloc[0:NUMBER_OF_TRAINING_IMGS,0 ] \n                                                        #instead of        .iloc[0:NUMBER_OF_TRAINING_IMGS,:1] but the first case returns a Series and the second one a DataFrame.\n                                                        #prefer the latter.\ntrain_images,test_images,train_labels,test_labels=train_test_split(images,labels,test_size=0.2,random_state=13)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4ce5297ed361c5c601ed881667e998bb9085fb9c"},"cell_type":"markdown","source":"The output type from  [train_test_split](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html) is the same as the input type.  "},{"metadata":{"trusted":true,"_uuid":"48fbd2626770799d6386a52d2702efd97006ddac"},"cell_type":"code","source":"type(train_images)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5f17d7720b599cc1a9744d8e990c03ee53d4c640"},"cell_type":"markdown","source":"\n**Now to viewing an image**[](http://).  \nThe image is unrolled into a single row that represents an image. Convert the single data point (row) to a 28X28 matrix (to view the original image)\nAnd then use matplotlib to plot this matrix."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"ii=13\nimg=train_images.iloc[ii].values\nprint('shape of numpy array',img.shape,'reshaping to 28x28')\nimg=img.reshape(IMG_HEIGHT,IMG_WIDTH)\nplt.imshow(img,cmap='gray')\nplt.title(train_labels.iloc[ii,0])\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6ea6d5fa55e2c5c1f3f4540025da159b00aa5164"},"cell_type":"markdown","source":"**Statistics**  \nhow are the pixel values distributed?   \n"},{"metadata":{"trusted":true,"_uuid":"0eb23ac392f40fd0b3c3909dd14032afae684bfd"},"cell_type":"code","source":"plt.hist(train_images.iloc[ii].values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"20c7177cf35d6454dfc2f54143e546115210eecf"},"cell_type":"markdown","source":"More darker pixel values (close to 0 ) as compared to lighter( closer to 256).  \nEach of the pixels is a feature, so the data point is a 784-Dimensional vector.  \nWe need to standardize the features .   \n\n"},{"metadata":{"trusted":true,"_uuid":"8c04611d5e2e8185ad9e77c0615ac8df39d76179"},"cell_type":"code","source":"train_images.describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"51da91425e7f758b8fb77be1c71724f7deb820b8"},"cell_type":"markdown","source":"[How to standardize a dataframe while keeping the header information](https://stackoverflow.com/questions/26414913/normalize-columns-of-pandas-data-frame)  \nThe training set is assumed variable enough to be that all data (test and new) comes from a distribution similar to the training set distribution.  \nTherefore I am saving  the max  and the min values  from the training data distribution.  \nmax and min will then be used to perform the standardization of all the data."},{"metadata":{"trusted":true,"_uuid":"861826fc54116cdfc3728de30753a4b0a67a6237"},"cell_type":"code","source":"max=train_images.max()\nmin=train_images.min()\ntrain_images_std=(train_images-min)/(max-min)\ntest_images_std=(test_images-min)/(max-min)\ntrain_images_std.head(n=2)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c66f85ce3b1052da3c67b66f4effde7a1b7e632c"},"cell_type":"markdown","source":"> division by 0  causes  [NaNs](https://stackoverflow.com/questions/13295735/how-can-i-replace-all-the-nan-values-with-zeros-in-a-column-of-a-pandas-datafra).  and [infinities](https://stackoverflow.com/questions/17477979/dropping-infinite-values-from-dataframes-in-pandas) that need to be replaced\n"},{"metadata":{"trusted":true,"_uuid":"9d3ff729bb6bf1ea424cbd1ffd1c39dd93471e85"},"cell_type":"code","source":"train_images_std=train_images_std.replace([-np.inf,np.inf],np.nan)\ntest_images_std=test_images_std.replace([-np.inf,np.inf],np.nan)\ntrain_images_std=train_images_std.fillna(0)\ntest_images_std=test_images_std.fillna(0)\ntrain_images_std.head(n=2)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d99b887b10108211de7e9109d9c16908f28170c3"},"cell_type":"markdown","source":"what does the distribution look like now? "},{"metadata":{"trusted":true,"_uuid":"6944af1e7c238c0a26b52743d10f6ef3335d45ab"},"cell_type":"code","source":"plt.hist(train_images_std.iloc[13])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"70ff19f2be18fa677cc506eef3432ea3537606cc"},"cell_type":"markdown","source":"**Using [Support Vector Machines](https://scikit-learn.org/stable/modules/generated/sklearn.svm.SVC.html)**  \n\n"},{"metadata":{"trusted":true,"_uuid":"fb932c828ec46230649e3a513688aca5447dd3ce"},"cell_type":"code","source":"test_images_std.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f2556dd8eee1718e5742edecd7011f19d623f3b2","_kg_hide-input":false,"_kg_hide-output":false,"scrolled":true},"cell_type":"code","source":"clf=SVC(kernel='rbf',C=1.0,random_state=1,gamma=0.1)   # radial basis function K(a,b) = exp(-gamma * ||a-b||^2 ) where gamma=1/(2*stddev)^2\nclf.fit(train_images_std,train_labels.values.ravel())\nclf.score(test_images_std,\n          test_labels.values.ravel())\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"446375638fbb006e46f9cdeca7bff1931b9ca3c6"},"cell_type":"markdown","source":"**81%** accuracy with scaled data.   \nNow compare this with the case below where the data is used as is. (no scaling)"},{"metadata":{"trusted":true,"_uuid":"8deb8e22e084baf7e0422e4f7fa765ef3dc7f761"},"cell_type":"code","source":"## Will not be using this classifier. This  is to test the need for scaling in SVM\n\nclf_without_std=SVC(kernel='rbf',C=1.0,random_state=1,gamma=0.1)   # radial basis function K(a,b) = exp(-gamma * ||a-b||^2 ) where gamma=1/(2*stddev)^2\nclf_without_std.fit(train_images,train_labels.values.ravel())      # unscaled original  data.\nclf_without_std.score(test_images,\n          test_labels.values.ravel())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4266f1cedde2dd6f29b1203cb333ccd8fec07609"},"cell_type":"markdown","source":"**10%** accuracy  \nThis is because the radial basis function kernel aka the gaussian kernel is of the the form K(a,b) = exp(-gamma * ||a-b||^2 ) i.e the L2 norm is used between data points to determine similarity.   \nAnd when the L2 norm is been used, it's best to keep the scaling similar between features vectors ."},{"metadata":{"trusted":true,"_uuid":"225e27dcda2d2764b06f20e892c7c4459d69c1f8"},"cell_type":"markdown","source":"\n**Submission**  \n\nload the data in test.csv  \nscale it using the min and max values derived from the training set.  \nremove Nans and infinities as before.  \n"},{"metadata":{"trusted":true,"_uuid":"aafeaad9562e9fe9f9bb7a80eac12568a27d63c2"},"cell_type":"code","source":"#use the clf model. (trained on scaled data)\nraw_data=pd.read_csv('../input/test.csv')\nprint('shape of dataframe ',raw_data.shape)\nraw_data.head(n=3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"47bc3fecb1a4fe144c3ad1372514656e11cc711f"},"cell_type":"code","source":"raw_data_scaled=(raw_data-min)/(max-min)\nraw_data_scaled=raw_data_scaled.replace([-np.inf,np.inf],np.nan)\nraw_data_scaled=raw_data_scaled.fillna(0)\nraw_data_scaled.describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ec1a649d05ebcf1c2b36b14cc101e3903e3f2f56"},"cell_type":"markdown","source":"Spot check to see if the clf predicts correctly on new data.  \nView an image from the raw data and see if the prediction matches expected value"},{"metadata":{"trusted":true,"_uuid":"50eeb74ed4a651fdf78ebabe82612194f7d227d6"},"cell_type":"code","source":"from random import randint\n\njj=13\n\nfor i in range(1,3):\n    jx=randint(0,raw_data_scaled.shape[0])\n    smp=raw_data_scaled.iloc[[jx]]\n    img=smp.values        #get a random image\n    img=img.reshape(IMG_HEIGHT,IMG_WIDTH)\n    for j in range(1,3):\n        for k in range(1,2):\n            print('plt',i,j,k)\n            plt.subplot(i,j,k)\n            plt.imshow(img)\n            y=clf.predict(smp)\n            plt.title('predicted : '+str(y[0]))\n            \n\n'''\nimg=raw_data_scaled.iloc[[jj]].values\nimg=img.reshape(IMG_HEIGHT,IMG_WIDTH)\nimg1=raw_data_scaled.iloc[jj+1].values\nimg1=img1.reshape(IMG_HEIGHT,IMG_WIDTH)\nplt.subplot(121)\nplt.imshow(img)\nplt.title('predicted: ')\nplt.subplot(122)\nplt.imshow(img1)\nplt.title('predicted: ')\n'''","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7fdd2bf8a304d9e47a88c21350b3e8f52cf7645a"},"cell_type":"code","source":"jjthSample=raw_data_scaled.iloc[[jj]]     # iloc[j] returns a Series. We need a dataframe to pass to predict. which is returned by iloc[[jj]]  .i.e. a list of rows\ntype(jjthSample)\ny_pred_jj=clf.predict(jjthSample)\ny_pred_jj","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3849dae290ceccd127ec976ab7e906859564339f"},"cell_type":"code","source":"y_pred=clf.predict(raw_data_scaled)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4f6379fd31218e88b25338d9df023ae8eaa7e97c"},"cell_type":"code","source":"y_pred.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4151f0e0d2b971dc849b503e0fd5b81820f5132c"},"cell_type":"code","source":"submissions=pd.DataFrame({\"ImageId\":list(range(1,len(y_pred)+1)), \"Label\":y_pred})\nsubmissions.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7e16021f4975200e55e92ecaec4408f14e6827c0"},"cell_type":"code","source":"submissions.to_csv(\"mnist_svm_submit.csv\",index=False,header=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b3f2332cb8be54bd468064d35d1e97dd2fed7f6e"},"cell_type":"code","source":"!ls","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5a55d75cd9068e808d67b0d540cc3297f36f475a"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}