{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# NCA Emoji Challenge Getting Started - Wrangling Emojis","metadata":{}},{"cell_type":"code","source":"# ref \n# https://github.com/googlefonts/emoji-metadata/blob/main/emoji_14_0_ordering.json,under the Apache license, version 2.0\n# png128 - competition folder of png size 128 emoji images (Source - https://github.com/googlefonts/noto-emoji under the Apache license, version 2.0) \n# https://emojipedia.org","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport os\nimport json\nfrom IPython.display import Image\nfrom IPython.core.display import HTML\nimport re\nfrom re import finditer\nfrom collections import defaultdict","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-06T08:02:48.612113Z","iopub.execute_input":"2022-08-06T08:02:48.613364Z","iopub.status.idle":"2022-08-06T08:02:48.619791Z","shell.execute_reply.started":"2022-08-06T08:02:48.613318Z","shell.execute_reply":"2022-08-06T08:02:48.618698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### In the Data section for the NCA Emoji Challenge you will find an emojis_reference.csv\n\nUsing the emoji_14_0_ordering.json which has group for a high level grouping of emojis, base entry for each emoji as well as alternates (if any e.g. for tone modifiers) and shortcodes about the emoji, these were combined with the entries in the competition folder png128 and every emoji by codepoint from emojipedia. The result emojis_reference should contain everything you wanted to know about an emoji but were afraid to ask! The filename (fname) is provided and an indicator whether it exists in the png128 folder or not.  ","metadata":{}},{"cell_type":"code","source":"df_emjr = pd.read_csv('../input/nca-emoji-challenge/emojis_reference.csv') # load the emojis reference \ndf_emjr.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T08:02:53.612445Z","iopub.execute_input":"2022-08-06T08:02:53.612864Z","iopub.status.idle":"2022-08-06T08:02:53.670520Z","shell.execute_reply.started":"2022-08-06T08:02:53.612832Z","shell.execute_reply":"2022-08-06T08:02:53.669147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_emjr.group.unique() # view unique groups  ","metadata":{"execution":{"iopub.status.busy":"2022-08-06T08:32:21.122326Z","iopub.execute_input":"2022-08-06T08:32:21.122767Z","iopub.status.idle":"2022-08-06T08:32:21.131903Z","shell.execute_reply.started":"2022-08-06T08:32:21.122731Z","shell.execute_reply":"2022-08-06T08:32:21.130665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is a NaN group in the emojis_reference unique groups above. This has 48 entries representing combination sections such as tone modifier, hair modifier, regional indicator symbols A-Z, asterisk, digit 0-9, combining enclosed keycap.\nThese will not be needed - combination emojis will have their own entries and images. ","metadata":{}},{"cell_type":"code","source":"(df_emjr[df_emjr.group_no==9999]).head()  # group 9999 ","metadata":{"execution":{"iopub.status.busy":"2022-08-06T08:21:51.071754Z","iopub.execute_input":"2022-08-06T08:21:51.072241Z","iopub.status.idle":"2022-08-06T08:21:51.104766Z","shell.execute_reply.started":"2022-08-06T08:21:51.072192Z","shell.execute_reply":"2022-08-06T08:21:51.103239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## HTML Summaries for each group\n\nThe rest of the notebook will create HTML for each group for perhaps a nicer user experience and finding out what emojis exist and how searching could be done. Maybe to make wrangling emojis easier.","metadata":{}},{"cell_type":"code","source":"# htmlsum will hold the html for the summary of group selected with links to read more info, \nhtmlsum = defaultdict(lambda: \"\")\nURL = 'https://emojipedia.org/emoji/'  # for more information link to emojipedia for emoji ","metadata":{"execution":{"iopub.status.busy":"2022-08-06T08:37:33.862120Z","iopub.execute_input":"2022-08-06T08:37:33.862596Z","iopub.status.idle":"2022-08-06T08:37:33.868670Z","shell.execute_reply.started":"2022-08-06T08:37:33.862561Z","shell.execute_reply":"2022-08-06T08:37:33.867443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will create a dataframe excluding group 9999 and drop columns like group_no that are not needed. \nThe in_png128 column is a boolean that we will make a string and will add a column for read_more for the link.","metadata":{}},{"cell_type":"code","source":"df = df_emjr[df_emjr.group_no!=9999].copy().reset_index(drop=True)  # exclude 9999 group NaN \ndf.drop('group_no', axis=1, inplace=True) # drop col group_no \ndf['in_png128'] = df['in_png128'].apply(lambda x : str(x)) # use string for True/False, htmlstr does not like boolean\ndf['read_more'] = df.emoji.apply(lambda x: x + ' more_details')  # add read_more col to use for link to emojipedia for emoji \ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T08:31:50.262141Z","iopub.execute_input":"2022-08-06T08:31:50.262599Z","iopub.status.idle":"2022-08-06T08:31:50.297959Z","shell.execute_reply.started":"2022-08-06T08:31:50.262565Z","shell.execute_reply":"2022-08-06T08:31:50.296836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"groups = df.group.unique() # view unique groups - will create a summary html for each  \ngroups","metadata":{"execution":{"iopub.status.busy":"2022-08-06T08:37:42.689147Z","iopub.execute_input":"2022-08-06T08:37:42.689603Z","iopub.status.idle":"2022-08-06T08:37:42.699020Z","shell.execute_reply.started":"2022-08-06T08:37:42.689566Z","shell.execute_reply":"2022-08-06T08:37:42.697768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def htmlsum_create(df):\n    # for group summary\n    \n    cols = df.columns.values\n    sumgrp = df.group.unique()\n    for group in sumgrp:\n        \n        dftmp = df[df.group==group].copy().reset_index(drop=True)\n        \n        idx = dftmp.index\n        htmlstr = \"<div class='summary_output' >\"        \n        htmlstr += \"<div style='display:table; overflow-y: scroll; height: 536px;  cellpadding:2px; cellspacing:2px; border:1px; border-style:solid;'  >\"\n        htmlstr += \"<div style='display:thead; height: 536px;  overflow-y: scroll;'  >\"\n        htmlstr += \"<div style='display:table-row;   position:sticky; top: 0;background-color: yellow;'>\"\n        for i in range(len(cols)):  # table column headings - using position:sticky above with a top location will keep it in place when scrolling\n            if cols[i] != 'doi':  #  used if a column had the doi or link so not to display but use later in a href - left in as example code\n                htmlstr += \"<div style='display:table-cell; font-weight:bold; padding-left:2px; padding-right:2px; cellspacing:2px; border:1px; border-style:solid; '>\" + cols[i] + \"</div>\"\n              \n        htmlstr += \" </div>\"\n        for ix in idx:\n            htmlstr += \"<div style='display:table-row;'>\"\n            for i in range(len(cols)):\n                if cols[i] != 'doi':\n                    if cols[i] == 'read_more':  # putting link with read_more col - will be URL + emoji\n                         htmlstr += \"<div style='display:table-cell; padding-left:2px; padding-right:2px; cellspacing:2px; border:1px; border-style:solid;'>\" +  \"<span style='color:#0099cc;'> [\" + \"<a href=\"+ URL + dftmp[\"emoji\"][ix] + \">\"  + dftmp[\"read_more\"][ix] + \"</a>\" + \"] </span>\" + \"</div>\"\n                    else:    \n                        htmlstr += \"<div style='display:table-cell; padding-left:2px; padding-right:2px; cellspacing:2px; border:1px; border-style:solid;'>\" +  dftmp[cols[i]][ix] + \"</div>\"\n            htmlstr += \"</div>\"                  \n   \n        htmlstr += \"</div>\"\n        htmlstr += \"</div>\"\n        htmlsum[group] = htmlstr \n    ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-06T08:39:03.103431Z","iopub.execute_input":"2022-08-06T08:39:03.103865Z","iopub.status.idle":"2022-08-06T08:39:03.117031Z","shell.execute_reply.started":"2022-08-06T08:39:03.103833Z","shell.execute_reply":"2022-08-06T08:39:03.115719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Create the html group summary for the groups in df using the hidden function above - just pass the df.","metadata":{}},{"cell_type":"code","source":"htmlsum_create(df) ","metadata":{"execution":{"iopub.status.busy":"2022-08-06T08:39:07.880435Z","iopub.execute_input":"2022-08-06T08:39:07.881574Z","iopub.status.idle":"2022-08-06T08:39:08.168020Z","shell.execute_reply.started":"2022-08-06T08:39:07.881512Z","shell.execute_reply":"2022-08-06T08:39:08.166292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Use the group name to display the content for a group. \n\nYou could also use array entry for groups if you prefer, the unique entries set earlier. e.g.,  display(HTML(htmlsum[groups[2]]))\n\nThe last column on the right for read_more has the link to emojipedia for the emoji if you want to get more details.","metadata":{}},{"cell_type":"code","source":"groups[2]","metadata":{"execution":{"iopub.status.busy":"2022-08-06T08:56:49.961966Z","iopub.execute_input":"2022-08-06T08:56:49.962454Z","iopub.status.idle":"2022-08-06T08:56:49.969634Z","shell.execute_reply.started":"2022-08-06T08:56:49.962410Z","shell.execute_reply":"2022-08-06T08:56:49.968733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(HTML(htmlsum[\"Animals and nature\"]))  ","metadata":{"execution":{"iopub.status.busy":"2022-08-06T08:42:57.031119Z","iopub.execute_input":"2022-08-06T08:42:57.031607Z","iopub.status.idle":"2022-08-06T08:42:57.050393Z","shell.execute_reply.started":"2022-08-06T08:42:57.031571Z","shell.execute_reply":"2022-08-06T08:42:57.049529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### A bit about Flags","metadata":{}},{"cell_type":"markdown","source":"The group Flags was left in just for completeness and if any interest in viewing flags.  They are not used in the competition.  Their images are handled differently and some complications with format, colors, size, etc. that are mentioned in github issues.  A few flags exist but most are not country or region flags. So if you want an emoji for a flag, they are included and will show here with display(HTML(htmlsum[\"Flags\"])).    ","metadata":{}},{"cell_type":"markdown","source":"### Hopefully this helps with emojis for the competition!!  ","metadata":{}}]}