{"cells":[{"metadata":{"_uuid":"84379bcca723e122e02aa2086096d52ac228db8f"},"cell_type":"markdown","source":"In this notebook I  have  performed digit  classification using a  decision tree"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nfrom sklearn.tree import DecisionTreeClassifier\n\nfrom sklearn.model_selection import train_test_split\n#constants\n\nIMG_HEIGHT=28\nIMG_WIDTH=28\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2d64bd18345cc19235f45082130933cdb85afc93"},"cell_type":"markdown","source":"**Loading the data**  \nusing pandas.read_csv to load in the training data. "},{"metadata":{"trusted":true,"_uuid":"e93d16213d7daf5f5ccad92051e6d46ac76daa76"},"cell_type":"code","source":"loaded_images=pd.read_csv('../input/train.csv')\nloaded_images.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b24daa3e191e4f23b879924451b998b72a549187"},"cell_type":"markdown","source":"Split into images and labels.  "},{"metadata":{"trusted":true,"_uuid":"012b8d587c660ce638a8090327d6fa509a104f73"},"cell_type":"code","source":"images=loaded_images.iloc[:,1:]\nlabels=loaded_images.iloc[:,:1]   # for the labels to be a dataframe . iloc[:,0] returns a Series  iloc[:,:1] returns a Dataframe\nlabels.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fe7117e63c4099b165c9f867a3f2155c69740901"},"cell_type":"markdown","source":"further split into training and test sets"},{"metadata":{"trusted":true,"_uuid":"66075120f6e0bedf8d1dd4791460fe311f1016ec"},"cell_type":"code","source":"train_images,test_images,train_labels,test_labels=train_test_split(images,labels,test_size=0.2,random_state=13)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b9bb8e21517b258ef1f764a6d55a262930f86747"},"cell_type":"code","source":"train_images.describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"db70d9d1fc6eb75ac5b749d239da896deeb431a6"},"cell_type":"markdown","source":"Since out of all the features that are used to split a node, the one that maximizes the Information Gain is the one that the node is split  on.\n\n*IG is defined in terms of the impurity measure \"I\" which could be defined in terms of either entropy or the gini index. Each of these indices describes how pure a node is . So entropy of 1 implies a very impure node and an entropy of 0 implies that all members of the node belong to the same class i.e a pure node. similar with the gini index. Scaling is not an issue here since the impurity functions are defined in terms of probabilities which are values between 0 and 1*  \n\nTherefore features need not be scaled when dealing with DTs.  \n  \n**Classification using a Decision Tree**  \nusing the [DecisionTreeClassifier in sklearn](https://scikit-learn.org/stable/modules/generated/sklearn.tree.DecisionTreeClassifier.html)"},{"metadata":{"trusted":true,"_uuid":"4e445a09f1d0928067389aa6b4585613bcf44314"},"cell_type":"code","source":"tree=DecisionTreeClassifier(criterion='gini',random_state=1)\ntree.fit(train_images,train_labels)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"025c9694db946f9330fe447cb1fc31edb4c33e19"},"cell_type":"markdown","source":"get both the train and the test scores."},{"metadata":{"trusted":true,"_uuid":"97bdafbfd97b2fef82d6901b8e45f1872f2e3049"},"cell_type":"code","source":"tree.score(train_images,train_labels.values.ravel())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3c1edf2ab33f7e68557543e9813e05cf1d1e7088"},"cell_type":"markdown","source":"possible overfitting on the training data."},{"metadata":{"trusted":true,"_uuid":"3347596083e44fc80fcb051d03d02c847563bac9"},"cell_type":"code","source":"tree.score(test_images,test_labels.values.ravel())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f822dea5717b28faf18b15247abdc111991cad9a"},"cell_type":"markdown","source":"An 85% accuracy on the test data.\nSpot check to see if the predictions are correct . Plotting the predictions as labels."},{"metadata":{"trusted":true,"_uuid":"0d4fcc03d8a00fcd955e0b0e5aa1b192bd28dd02"},"cell_type":"code","source":"figr,axes=plt.subplots(figsize=(10,10),ncols=3,nrows=3)\naxes=axes.flatten()\nfor i in range(0,9):\n    jj=np.random.randint(0,test_images.shape[0])          #pick a random image\n    axes[i].imshow(test_images.iloc[[jj]].values.reshape(IMG_HEIGHT,IMG_WIDTH))\n    axes[i].set_title('predicted: '+str(tree.predict(test_images.iloc[[jj]])[0]))\n\n\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"25cb99687fe282585a6b00cd28a30b0ee16f81b0"},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"0be293c2ead85fdad9863676f2534f1398d2dfd7"},"cell_type":"markdown","source":"**Submission **  \n\nload the data in test.csv  \npredict using the model"},{"metadata":{"trusted":true,"_uuid":"6a5e6015143bcf89db2fe33a765c65c5f7c15e30"},"cell_type":"code","source":"new_data=pd.read_csv('../input/test.csv')\nnew_data.head(n=3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c28c01fcba3bf6b3bebdf28535fe0ce408c64910"},"cell_type":"code","source":"y_pred=tree.predict(new_data)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"22aa90c67f68dd4c215a54dbff992037726b2e9a"},"cell_type":"code","source":"y_pred.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ab9992f1d79c1d53478d0b8e2373e18bc58cc250"},"cell_type":"markdown","source":"create a dataframe which will then be exported as a csv file for submissions."},{"metadata":{"trusted":true,"_uuid":"90f8916f8a13b56dbcf79ba5cb946df9a79c8716"},"cell_type":"code","source":"submissions=pd.DataFrame({\"ImageId\":list(range(1,len(y_pred)+1)), \"Label\":y_pred})\nsubmissions.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4d0b86480cdeba5f30ac52f939ab3b6c32f42d6c"},"cell_type":"code","source":"submissions.to_csv(\"mnist_decision_tree_submit.csv\",index=False,header=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f5a5b4a4bf84b4ccc69fd24105f6d0296f79cdad"},"cell_type":"code","source":"!ls","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3e1083324b46e66b7068824c8f69716c6a6e69de"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}