{"cells":[{"metadata":{},"cell_type":"markdown","source":"# <span style=\"color:blue\">Bengali</span>\n![](http://www.ukindia.com/zip/zben02.gif)\nBengali also known by its endonym Bangla, is an Indo-Aryan language primarily spoken by the Bengalis in South Asia. It is brahmic language which is assumed to be evoloved from Sanskrit\n\n### <span style=\"color:red\">Dont forget to upvote if you like!!</span>\n## Alphabets\nThe Bengali script can be divided into vowels and vowel diacritics/marks, consonants and consonant conjuncts, diacritical and other symbols, digits, and punctuation marks. Vowels & consonants are used as alphabet and also as diacritical marks. Bengali contains 28 letters\n### Vowels\nThe Bengali script has a total of 9 vowel graphemes, each of which is called swôrôbôrnô \"vowel letter\"\n\n| Vowels  | Vowels phoneme  | \n|---|---|\n| **অ**  |ô|\n|  **আ ** | a  |\n|  **ই** | i  |\n| **ঈ** | ī/ee |\n| **উ** | u  |\n| **ঊ** | ū/oo |\n|**ঋ** |  ṛ/ri |\n| **ৠ** | ṝ/rri |\n| **ঌ** | ḷ/li |\n| **ৡ** | ḹ/lli |\n\n#### Complex Vowels\nBengali also has 4 complex vowels\n\n| Complex Vowels  | Vowels phoneme  | \n|---|---|\n| **এ ** | e |\n| **ঐ** | oi |\n| **ও** | o |\n| **ঔ** | ou |\n\n### Consonants\nConsonant letters are called bænjônbôrnô \"consonant letter\" in Bengali. The names of the letters are typically just the consonant sound plus the inherent vowel. \n![](https://i.ytimg.com/vi/2FtVGjJC68I/maxresdefault.jpg)\n\n\n\nSince they work on combination there are 19 x 9 letters to be identified."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"code","source":"\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport os\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport random\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"Train_data = pd.read_csv('/kaggle/input/bengaliai-cv19/train.csv')\nTest_data = pd.read_csv('/kaggle/input/bengaliai-cv19/test.csv')\nclass_map = pd.read_csv('/kaggle/input/bengaliai-cv19/class_map.csv')\nsample_submission = pd.read_csv('/kaggle/input/bengaliai-cv19/sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Train_data\nContains,\n1. Image_id -> Id of training Image\n2. grapheme_root -> character number (vowel + consonant)\n3. vowel_diacritic -> root vowel id number\n4. consonant_diacritic -> root consonant id number"},{"metadata":{"trusted":true},"cell_type":"code","source":"Train_data.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Class_map\n\nProvides us the number of unique elements in each class and their label to be reffered to Train_data"},{"metadata":{"trusted":true},"cell_type":"code","source":"class_map.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# count the number of classes for each target\nn_classes = class_map.groupby(by=['component_type']).count()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"n_classes.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.set(style=\"darkgrid\")\nk = ['vowel_diacritic','grapheme_root','consonant_diacritic']\nsns.countplot(data = class_map ,x = 'component_type',order = k)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"HEIGHT = 137\nWIDTH = 236","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"b_train_data = pd.read_parquet(f'/kaggle/input/bengaliai-cv19/train_image_data_0.parquet')\n\nimageid = b_train_data.iloc[:, 0]\nimage = b_train_data.iloc[:, 1:].values.reshape(-1, HEIGHT, WIDTH)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"f, ax = plt.subplots(5, 5, figsize=(16, 8))\nax = ax.flatten()\n\nfor i in range(25):\n    ax[i].imshow(image[i], cmap='Greys')\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def listToString(s):  \n    \n    # initialize an empty string \n    str1 = \"\"  \n    \n    # traverse in the string   \n    for ele in s:  \n        str1 += ele   \n    \n    # return string   \n    return str1  \n        ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from wordcloud import WordCloud\n\n\ns = listToString(Train_data['grapheme'] )\n\n# Read the whole text.\ntext = s\n\n# Generate a word cloud image\nwordcloud = WordCloud().generate(text)\n\n# Display the generated image:\n# the matplotlib way:\nimport matplotlib.pyplot as plt\n\n\n# take relative word frequencies into account, lower max_font_size\nwordcloud = WordCloud(background_color=\"white\",max_words=len(s),max_font_size=70, relative_scaling=.5, font_path = \"/kaggle/input/bengalifont/Nikosh.ttf\").generate(text)\nplt.figure(figsize=(30, 30))\nplt.imshow(wordcloud)\nplt.axis(\"off\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Histogramic Distribution of Grapheme root"},{"metadata":{"trusted":true},"cell_type":"code","source":"\n#sns.set(style=\"darkgrid\")\nplt.figure(figsize=(20, 10))\nsns.countplot(data = Train_data ,x = 'grapheme_root',color=\"c\")\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"type(image[1])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Augmentations\n## Flipping"},{"metadata":{"trusted":true},"cell_type":"code","source":"rand_num = random.randint(0,len(image))\nrand_num","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.imshow(image[rand_num])\nplt.title(\"input_img\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"flipped_img = np.fliplr(image[rand_num])\nplt.imshow(flipped_img)\nplt.title(\"flipped_img\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Gaussian Noice"},{"metadata":{"trusted":true},"cell_type":"code","source":"import skimage\nimg = image[rand_num]\nimg = skimage.util.random_noise(img, mode='gaussian', seed=None, clip=True)\nplt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Rotate"},{"metadata":{"trusted":true},"cell_type":"code","source":"from skimage.transform import rotate\nimg = image[rand_num]\nimg = rotate(img, 15)\nplt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from skimage.transform import rotate\nimg = image[rand_num]\nimg = rotate(img, -15)\nplt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Change Intensity (Make letters Clear if their intensity is low)"},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nfrom skimage import exposure\nimg = image[rand_num]\nv_min, v_max = np.percentile(img, (0.2, 99.8))\nimg = exposure.rescale_intensity(img, in_range=(v_min, v_max))\n\nplt.imshow(img)\nplt.show()\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Classical Research Papers and Repos on bengali letter classifications\n\n1. https://www.researchgate.net/publication/322112791_Handwritten_Bangla_Character_Recognition_Using_The_State-of-Art_Deep_Convolutional_Neural_Networks\n2. https://www.researchgate.net/publication/328214545_BARD_Bangla_Article_Classification_Using_a_New_Comprehensive_Dataset\n3. https://arxiv.org/html/1902.11133 (Bengali Handwritten Character Classification using Transfer Learning on Deep Convolutional Neural Network)\n4. https://github.com/srdg/bangla-dl\n5. https://github.com/dibyatanoy/Bengali-Handwritten-Character-Recognition-Using-Convolutional-Neural-Networks\n6. https://github.com/tanvirfahim15/BARD-Bangla-Article-Classifier"},{"metadata":{},"cell_type":"markdown","source":"## To Start with the Chalange go for\n## Bengali.AI : Design ResNet Layer by Layer\nhttps://www.kaggle.com/chekoduadarsh/bengali-ai-tutorial-design-resnet-layer-by-layer"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}