{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Top"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport re\nfrom tqdm.notebook import tqdm\n\ntqdm().pandas()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv(\"../input/bms-molecular-translation/train_labels.csv\")\nprint(train.shape)\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Extract Elements from Inchi"},{"metadata":{},"cell_type":"markdown","source":"The elements in the inchi can be extracted from the Chemical composition which is the second entry when the InChI is split by the `/` character."},{"metadata":{"trusted":true},"cell_type":"code","source":"train.InChI = train.InChI.progress_apply(lambda x: re.findall(r'([A-Z][a-z]?)',x.split('/')[1]) )\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We will now gather all the unique elements."},{"metadata":{"trusted":true},"cell_type":"code","source":"all_el = set()\ntrain.InChI.progress_apply(lambda x: all_el.update(x))\nprint(len(all_el), all_el)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Next, we will be counting all lines with the presence of specific elements."},{"metadata":{"trusted":true},"cell_type":"code","source":"elements = sorted(all_el)\ncounts = []\nfor e in elements:\n    counts.append((np.sum(train.InChI.progress_apply(lambda x: e in x)), e))\ncounts.sort(reverse=True)\nfor n, e in counts:\n    print(f'Element {e} count: {n} (approx {n*100/train.shape[0]:.4}%)')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We will now group each entry based on the presence of specific elements."},{"metadata":{"trusted":true},"cell_type":"code","source":"groupings = {}\nfor e in elements:\n    groupings[e] = train.InChI.progress_apply(\n        lambda x: e in x\n    )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"group_sets = {}\nfor e in elements:\n    group_sets[e] = set(train.image_id.loc[groupings[e]])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"After grouping them, we will now gather pairwise information from each pair of element."},{"metadata":{"trusted":true},"cell_type":"code","source":"#Intersections of each elements:\npair_groupings = []\nfor e1 in elements:\n    for e2 in elements:\n        if e1 == e2:\n            continue\n        n = len(group_sets[e1].intersection(group_sets[e2]))\n        pair_groupings.append((n, e1, e2))\n        \npair_groupings.sort(reverse=True)\nfor n, e1, e2 in pair_groupings[0::2]:\n    print(f'Lines with both {e1} and {e2}:\\t{n}\\t{100*n/train.shape[0]:.4}%')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Using this information, we can select part of the dataset which contains certain elements only. It can be useful when we want to minimize the train data size which spans over 2.4 Million entries.\n\nI hope this notebook helps someone out there!"},{"metadata":{},"cell_type":"markdown","source":"# End"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}