{"cells":[{"metadata":{"_uuid":"aabe338b6e7cbe5fa8b33a34ed0f70094f252f1b"},"cell_type":"markdown","source":"Counting with Random Forests  \nThis is meant to be a baseline test of recognizing digits without extensive tuning or testing.\n\nImprovements Needed:\n\nGrid Search Cross Validation"},{"metadata":{"trusted":true,"_uuid":"d5638a6c94c4cfcbc6449a4e124a37cd55af8cf0"},"cell_type":"code","source":"#Import necessary packages\nimport pandas as pd\nimport numpy as np\nimport sklearn\nimport matplotlib.pyplot as plt\nimport math\nimport random as rand\n\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.model_selection import train_test_split,GridSearchCV","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ae8e3ec03262f021a0faafd9c625858254d75049"},"cell_type":"markdown","source":"Read in Training Data"},{"metadata":{"trusted":true,"_uuid":"0ae1a08e4bb56370372901b8ab8193b11faf6303"},"cell_type":"code","source":"filepath=\"../input/\"\ntrain_path = filepath+\"train.csv\"\ntraining = pd.read_csv(train_path)\ntraining.describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b69e69db9cc90aa14c941babd8c47e9cadbd943b"},"cell_type":"markdown","source":"Let's check out one of the numbers to see what it looks like.  \nFirst, we'll find out the shape of the training set to see how many pixels we are working with. The picture size will have a length and width of number of columns - 1 since there is a column for the labels."},{"metadata":{"trusted":true,"_uuid":"79c33f942d20d1d9f76925252436a2fcd3c3c429"},"cell_type":"code","source":"#Based on this, the picture size will be 28x28\nsize=int(math.sqrt(training.shape[1]-1))\nsize","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1cb55f099c6b34847e1930d9ab31724902452540"},"cell_type":"code","source":"#Pull a random integer between 0 and number of rows in dataaset to test\nrand_num =rand.randint(0,training.shape[0])\n\n#Create array based on pulling row corresponding to random int\nnumber=np.array(training.loc[rand_num][1:],dtype='uint8')\n\n#Create 2-D array based on image size\nnumber=number.reshape((size,size))\n\nprint(\"Labeled value is\",str(training.loc[rand_num][0]))\nplt.imshow(number,cmap='Greys')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"06bf8f5c398d67dba43557e05a2a58ff3a914635"},"cell_type":"markdown","source":"Separate labels from training data and split into training and testing datasets"},{"metadata":{"trusted":true,"_uuid":"0156abe3f9b46d3427bfd2296e1b90b6118908e4"},"cell_type":"code","source":"training_labels=training['label']\ntraining_without_labels=training.drop(labels='label',axis=1)\ntrain_data, test_data, train_labels, test_labels = train_test_split(training_without_labels, training_labels, test_size=0.7)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9ee3e5f6b04234ea298991a51762bc359cd3ca55"},"cell_type":"markdown","source":"Initialize Random Forest Classifier"},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"772f9b4f75177dedddcf57b7936125c41cb514a9"},"cell_type":"code","source":"#rf = RandomForestClassifier()\nparameters = {'n_estimators':[100,200],'max_features':['auto','sqrt',None]}\n\n#gs_cv_rf = GridSearchCV(rf,parameters,cv=3)\n#gs_cv_rf.fit(train_data,train_labels)\n#gs_cv_rf.best_params_\n#best params: max_features = 'sqrt', n_estimators = 200\n\nrf = RandomForestClassifier(n_estimators=200,max_features='sqrt')\nrf.fit(train_data,train_labels)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bfdf7c3caf5c64bbc6010c0f5c458235e1cb9046"},"cell_type":"markdown","source":"  Predict using Train/Test split data"},{"metadata":{"trusted":true,"_uuid":"efb2306b80572e4fffac9e8827eeafa3053d116e"},"cell_type":"code","source":"predict=gs_cv_rf.predict(test_data)\ncheck=pd.DataFrame(predict,columns=[\"Predict\"])\ncheck['true']=test_labels.reset_index()['label']\n\ncheck['Accuracy']=0\ncheck['Accuracy']=check['Accuracy'].where(check['Predict']!=check['true'],1)\nprint(\"Accuracy: \",(check.sum()[2]/check.shape[0]).round(3))\ncheck.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"db1305f486c9b1f1c63d54227f2e3f6b83f4130b"},"cell_type":"markdown","source":"What values aren't being predicted well?"},{"metadata":{"trusted":true,"_uuid":"d52bf808c6d6470c2c1ffd47b476fc46f1a3ce3e"},"cell_type":"code","source":"check['true'][check['Accuracy']==0].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d9a9227bf97d98705629b07d84b5dbe66c98ad37"},"cell_type":"code","source":"#The random forest seems to think that 4s and 3s are often 9s\nplt.hist(check['Predict'][(check['Accuracy']==0)&(check['true']==9)])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"88e7efbd8cea2841cb81d7a8aae7fc0fbdf97d72"},"cell_type":"markdown","source":"Print out accuracy based on accurate predictions"},{"metadata":{"_uuid":"ff2bd8fb343d9310446007bbc4b79f6ece55f634"},"cell_type":"markdown","source":"Read in testing dataset and predict using Random Forest Classifier\n\nGenerate submission file"},{"metadata":{"trusted":true,"_uuid":"d0212fa5ca06ee83be0885e5decea7c4f4ab59de"},"cell_type":"code","source":"test_path = filepath+\"test.csv\"\ntesting = pd.read_csv(test_path)\nsolutions=rf.predict(testing)\nsubmission=pd.DataFrame(solutions,columns=['Label'])\nsubmission.index.name = \"ImageID\"\nsubmission.index += 1\n\nsubmission.to_csv('submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"573b73077097b8d2c7e3ff744816d35e1452e66e"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c7102e8ea1fe0eb455779d62fb5c341feb97f4cc"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}