{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-17T04:24:30.464484Z","iopub.execute_input":"2022-07-17T04:24:30.465045Z","iopub.status.idle":"2022-07-17T04:24:30.500029Z","shell.execute_reply.started":"2022-07-17T04:24:30.464930Z","shell.execute_reply":"2022-07-17T04:24:30.498718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **About Kannada**\n\n**Kannada**, historically called Canarese, **is a classical Dravidian language spoken predominantly by the people of Karnataka in the southwestern region of India**. The language is also spoken by linguistic minorities in the states of Maharashtra, Andhra Pradesh, Tamil Nadu, Telangana, Kerala and Goa; and also by Kannadigas abroad. The language had roughly 43 million native speakers by 2011. Kannada is also spoken as a second and third language by over 12.9 million non-native speakers in Karnataka, which **adds up to 56.9 million speakers**.\n\nThe Kannada language is written using the Kannada script, which evolved from the 5th-century Kadamba script.\n\n![kannada-numbers](https://2.bp.blogspot.com/-e13ee8EcKxU/Wl7dQ32q44I/AAAAAAAAAB4/um6EcQ9gq0YL9un_WWQNpw_d_uTvrDpBgCLcBGAs/s1600/numbers-kannada1.jpg)","metadata":{}},{"cell_type":"markdown","source":"## **Reference**\n\nIf you wish to learn about CNNs alongwith Visualizations, you may refer:- https://www.kaggle.com/code/pythonkumar/a-z-cnn-tutorial-cats-vs-dogs\n\n**Alongwith Visualizations you will learn about:-**\n\n1. Colour Channels in RGB Images\n2. Simple CNNs & Maxpooling\n3. Dense Layers & Flattening Images\n4. **Padding & Strides**\n5. **Data Augmentation**\n6. **Regularization**","metadata":{}},{"cell_type":"markdown","source":"## **ANN vs CNN**\n\nThe fundamental difference between a densely connected layer and a convolution layer is this: Dense layers learn global patterns in their input feature space (for example, for a MNIST digit, patterns involving all pixels), whereas convolution layers learn local patterns : in the case of images, patterns found in small 2D windows of the inputs.\n\n### **IMPORTANT properties of CNNs:**\n\n1. **The patterns they learn are translation invariant.** After learning a certain pattern in the lower-right corner of a picture, a convnet can recognize it anywhere: for example, in the upper-left corner. A densely connected network would have to learn the pattern anew if it appeared at a new location.\n2. **They can learn spatial hierarchies of patterns.** A first convolution layer will learn small local patterns such as edges, a second convolution layer will learn larger patterns made of the features of the first layers, and so on. ","metadata":{}},{"cell_type":"markdown","source":"# **Imports**\n\n**more to be imported as needed**","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing\nimport matplotlib.pyplot as plt # data visualization\nimport seaborn as sns # data visualization\n\nimport tensorflow\nprint(tensorflow.__version__)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:24:30.587572Z","iopub.execute_input":"2022-07-17T04:24:30.588419Z","iopub.status.idle":"2022-07-17T04:24:41.166643Z","shell.execute_reply.started":"2022-07-17T04:24:30.588331Z","shell.execute_reply":"2022-07-17T04:24:41.165120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Data**\n\n1. We have  28 x 28  dimension handwritten pics.\n2. Dataset has been already flattened and has 784-pixel values for each pic.\n3. Total we have  60000  pics in training set.","metadata":{}},{"cell_type":"code","source":"train_1=pd.read_csv('../input/Kannada-MNIST/train.csv')\ntest_1=pd.read_csv('../input/Kannada-MNIST/test.csv')\nval_1=pd.read_csv('../input/Kannada-MNIST//Dig-MNIST.csv')\ntrain_1","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:24:41.168969Z","iopub.execute_input":"2022-07-17T04:24:41.169818Z","iopub.status.idle":"2022-07-17T04:24:47.803231Z","shell.execute_reply.started":"2022-07-17T04:24:41.169772Z","shell.execute_reply":"2022-07-17T04:24:47.801772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_1","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:24:47.805120Z","iopub.execute_input":"2022-07-17T04:24:47.806266Z","iopub.status.idle":"2022-07-17T04:24:47.832621Z","shell.execute_reply.started":"2022-07-17T04:24:47.806197Z","shell.execute_reply":"2022-07-17T04:24:47.830786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_1","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:24:47.835980Z","iopub.execute_input":"2022-07-17T04:24:47.836547Z","iopub.status.idle":"2022-07-17T04:24:47.862231Z","shell.execute_reply.started":"2022-07-17T04:24:47.836503Z","shell.execute_reply":"2022-07-17T04:24:47.860953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Merge the Train & Val Dataframes**\n\n**Merging the 60k Training Examples & 10k Validation Examples**","metadata":{}},{"cell_type":"code","source":"train=pd.concat([train_1,val_1],axis=0)\ntrain","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:24:47.863731Z","iopub.execute_input":"2022-07-17T04:24:47.864516Z","iopub.status.idle":"2022-07-17T04:24:48.199962Z","shell.execute_reply.started":"2022-07-17T04:24:47.864471Z","shell.execute_reply":"2022-07-17T04:24:48.198601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Explore**\n\n* It is important to know the distribution of data according to the labels they have.\n* This data set is homogeneously distributed as you see below.","metadata":{}},{"cell_type":"code","source":"num = train_1.label.value_counts()\nsns.barplot(num.index,num,palette='hot')","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:24:48.201459Z","iopub.execute_input":"2022-07-17T04:24:48.202235Z","iopub.status.idle":"2022-07-17T04:24:48.476937Z","shell.execute_reply.started":"2022-07-17T04:24:48.202184Z","shell.execute_reply":"2022-07-17T04:24:48.475586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can see that all of the classes has equal distribution.There are 6000 examples of each numbers in kannada in the the training dataset.Cool !\n\nIf the data **wasn't homogeneously distributed** what would we do?\n1. Then we could use data augmentation techniques to generate new data for low quantity labels,\n2. Or if we have enough data we can discard some high quantity labels","metadata":{}},{"cell_type":"code","source":"num = val_1.label.value_counts()\nsns.barplot(num.index,num,palette='viridis')","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:24:48.478698Z","iopub.execute_input":"2022-07-17T04:24:48.479236Z","iopub.status.idle":"2022-07-17T04:24:48.716223Z","shell.execute_reply.started":"2022-07-17T04:24:48.479179Z","shell.execute_reply":"2022-07-17T04:24:48.714951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:24:48.717845Z","iopub.execute_input":"2022-07-17T04:24:48.718261Z","iopub.status.idle":"2022-07-17T04:24:48.773226Z","shell.execute_reply.started":"2022-07-17T04:24:48.718221Z","shell.execute_reply":"2022-07-17T04:24:48.772044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:24:48.776870Z","iopub.execute_input":"2022-07-17T04:24:48.778216Z","iopub.status.idle":"2022-07-17T04:24:51.996649Z","shell.execute_reply.started":"2022-07-17T04:24:48.778153Z","shell.execute_reply":"2022-07-17T04:24:51.995084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **See**","metadata":{}},{"cell_type":"code","source":"from keras.preprocessing.image import load_img\n\n# Gallery using Matplotlib \nfig, ax = plt.subplots(15,5,figsize = (12,12), dpi = 100)\naxes = ax.ravel()\n\nfor idx,ax  in enumerate(axes):\n    path=train_1.iloc[idx,1:].values.reshape(28,28)\n    ax.imshow(path)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:24:52.002367Z","iopub.execute_input":"2022-07-17T04:24:52.002850Z","iopub.status.idle":"2022-07-17T04:24:58.831829Z","shell.execute_reply.started":"2022-07-17T04:24:52.002812Z","shell.execute_reply":"2022-07-17T04:24:58.830685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Data Pre-processing**\n\n### **Though the Image Data provided has already been Flattened, seems like we can Restore the Image Data using .Reshape()**","metadata":{}},{"cell_type":"markdown","source":"## **Train : Seggregate X & y**","metadata":{}},{"cell_type":"code","source":"y=train['label']\nX=train.drop(['label'],axis=1)\nX\n# y","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:24:58.833466Z","iopub.execute_input":"2022-07-17T04:24:58.834107Z","iopub.status.idle":"2022-07-17T04:24:59.099038Z","shell.execute_reply.started":"2022-07-17T04:24:58.834062Z","shell.execute_reply":"2022-07-17T04:24:59.097501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Train-Test Split**","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25)\n# X_train\n# X_test\n# y_train\ny_test","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:24:59.101098Z","iopub.execute_input":"2022-07-17T04:24:59.101924Z","iopub.status.idle":"2022-07-17T04:24:59.963276Z","shell.execute_reply.started":"2022-07-17T04:24:59.101871Z","shell.execute_reply":"2022-07-17T04:24:59.960925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Train : Standardizing**\n\nStandardization (or z-score normalization) scales the values between 0 to 1, while taking into account standard deviation. If the standard deviation of features is different, their range also would differ from each other. This reduces the effect of the outliers in the features.","metadata":{}},{"cell_type":"code","source":"# RobustScaler\nfrom sklearn.preprocessing import RobustScaler\n\nrs = RobustScaler()\n\nX_train= rs.fit_transform(X_train)\nX_train","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:24:59.965168Z","iopub.execute_input":"2022-07-17T04:24:59.966190Z","iopub.status.idle":"2022-07-17T04:25:01.561276Z","shell.execute_reply.started":"2022-07-17T04:24:59.966131Z","shell.execute_reply":"2022-07-17T04:25:01.559718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Test : Drop ID column**","metadata":{}},{"cell_type":"code","source":"id=test_1.id\ntest=test_1.drop(['id'],axis=1)\ntest","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:25:01.563219Z","iopub.execute_input":"2022-07-17T04:25:01.564120Z","iopub.status.idle":"2022-07-17T04:25:01.606079Z","shell.execute_reply.started":"2022-07-17T04:25:01.564053Z","shell.execute_reply":"2022-07-17T04:25:01.604826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Test : Standardizing**\n\nStandardization (or z-score normalization) scales the values between 0 to 1, while taking into account standard deviation. If the standard deviation of features is different, their range also would differ from each other. This reduces the effect of the outliers in the features.","metadata":{}},{"cell_type":"code","source":"test= rs.fit_transform(test)\ntest","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:25:01.607776Z","iopub.execute_input":"2022-07-17T04:25:01.609114Z","iopub.status.idle":"2022-07-17T04:25:01.845969Z","shell.execute_reply.started":"2022-07-17T04:25:01.609058Z","shell.execute_reply":"2022-07-17T04:25:01.844622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Pure ML : Ensemble Technique**\n\n### **Here, the Image has been Flattened already. The Image Pixel / Features are already available. A simple Logistic Regression can solve it. Any Deep Learning methods would be an OVERKILL.**","metadata":{}},{"cell_type":"markdown","source":"## **Voting Classifiers**\n\nA very simple way to create an even better classifier is to aggregate the predictions of each classifier and predict the class that gets the most votes. This majority-vote classifier is called a hard voting classifier\n\nEnsemble methods work best when the predictors are as independent from one another as possible. One way to get diverse classifiers is to train them using very different algorithms. This increases the chance that they will make very different types of errors, improving the ensemble’s accuracy.","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import VotingClassifier\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.naive_bayes import GaussianNB\n\nvc=VotingClassifier(voting='hard',estimators=[\n    ('log',LogisticRegression(penalty='elasticnet',solver='saga',l1_ratio=0.5)),\n    ('knn',KNeighborsClassifier(n_neighbors=5)),\n    ('nb',GaussianNB())\n])","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:25:01.848896Z","iopub.execute_input":"2022-07-17T04:25:01.849993Z","iopub.status.idle":"2022-07-17T04:25:02.110428Z","shell.execute_reply.started":"2022-07-17T04:25:01.849936Z","shell.execute_reply":"2022-07-17T04:25:02.109328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Train**","metadata":{}},{"cell_type":"code","source":"# Train the Model\nvc.fit(X_train,y_train)\n\n# Get the Model Parameters\nvc.get_params(deep=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:25:02.111929Z","iopub.execute_input":"2022-07-17T04:25:02.113267Z","iopub.status.idle":"2022-07-17T04:36:15.235784Z","shell.execute_reply.started":"2022-07-17T04:25:02.113213Z","shell.execute_reply":"2022-07-17T04:36:15.234199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Plot Decision Boundary**\n\n**We cannot plot the Decision Boundary of this problem as it has 284 features. It will be a hyperplane in 284D space.**\n\nSo we will plot another problem in 2D Spacw with 2 Features.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport matplotlib.pyplot as plt\nfrom sklearn import datasets\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\n\niris = datasets.load_iris()\nX = iris.data[:, [0, 2]]\ny = iris.target\n\nclf1 = DecisionTreeClassifier(max_depth=4)\nclf2 = KNeighborsClassifier(n_neighbors=7)\n\nclf1.fit(X, y)\nclf2.fit(X, y)\n\nx_min, x_max = X[:, 0].min() - 1, X[:, 0].max() + 1\ny_min, y_max = X[:, 1].min() - 1, X[:, 1].max() + 1\nxx, yy = np.meshgrid(np.arange(x_min, x_max, 0.1),\n                 np.arange(y_min, y_max, 0.1))\n\nf, axarr = plt.subplots(1,2, sharex='col', sharey='row', figsize=(15,13))\n\nfor idx, clf, tt in zip([0, 1],[clf1, clf2],\n                    ['Decision Tree (depth=4)', 'KNN (k=7)']):\n\n    Z = clf.predict(np.c_[xx.ravel(), yy.ravel()])\n    Z = Z.reshape(xx.shape)\n\n    axarr[idx].contourf(xx, yy, Z, alpha=0.4, cmap=\"brg\")\n    axarr[idx].scatter(X[:, 0], X[:, 1], c=y, cmap=\"brg\",\n                                  s=20, edgecolor='w')\n    axarr[idx].set_title(tt)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:50:38.021562Z","iopub.execute_input":"2022-07-17T04:50:38.022054Z","iopub.status.idle":"2022-07-17T04:50:38.443822Z","shell.execute_reply.started":"2022-07-17T04:50:38.022015Z","shell.execute_reply":"2022-07-17T04:50:38.442414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Predict**","metadata":{}},{"cell_type":"code","source":"pred=vc.predict(X_test)\npred","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:36:15.237760Z","iopub.execute_input":"2022-07-17T04:36:15.238219Z","iopub.status.idle":"2022-07-17T04:37:06.291541Z","shell.execute_reply.started":"2022-07-17T04:36:15.238179Z","shell.execute_reply":"2022-07-17T04:37:06.289766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **F1-Score**","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import f1_score\nf1_score(y_test, pred,average=None)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:37:06.293435Z","iopub.execute_input":"2022-07-17T04:37:06.294491Z","iopub.status.idle":"2022-07-17T04:37:06.317682Z","shell.execute_reply.started":"2022-07-17T04:37:06.294432Z","shell.execute_reply":"2022-07-17T04:37:06.316056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Confusion Matrix**","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix\ncm = confusion_matrix(y_test,pred)\n\nimport plotly.express as px\nfig = px.imshow(cm,text_auto=True,color_continuous_scale='RdBu_r')\nfig.show()                ","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:37:06.319701Z","iopub.execute_input":"2022-07-17T04:37:06.320099Z","iopub.status.idle":"2022-07-17T04:37:09.010588Z","shell.execute_reply.started":"2022-07-17T04:37:06.320064Z","shell.execute_reply":"2022-07-17T04:37:09.009265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Suggestions:-\n* Kaggle - https://www.kaggle.com/pythonkumar\n* GitHub - https://github.com/KumarPython​\n* Twitter - https://twitter.com/KumarPython\n* LinkedIn - https://www.linkedin.com/in/kumarpython/","metadata":{}},{"cell_type":"markdown","source":"# **Submission**","metadata":{}},{"cell_type":"code","source":"sub=vc.predict(test)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:37:09.012606Z","iopub.execute_input":"2022-07-17T04:37:09.013407Z","iopub.status.idle":"2022-07-17T04:37:22.367372Z","shell.execute_reply.started":"2022-07-17T04:37:09.013338Z","shell.execute_reply":"2022-07-17T04:37:22.365870Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" submission=pd.DataFrame({'id': id,\n                         'label' : sub\n                        })\nsubmission\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T04:37:22.369180Z","iopub.execute_input":"2022-07-17T04:37:22.369774Z","iopub.status.idle":"2022-07-17T04:37:22.391703Z","shell.execute_reply.started":"2022-07-17T04:37:22.369723Z","shell.execute_reply":"2022-07-17T04:37:22.389426Z"},"trusted":true},"execution_count":null,"outputs":[]}]}