{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":67356,"databundleVersionId":8006601,"sourceType":"competition"}],"dockerImageVersionId":30673,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Published on April 04, 2023. By Marília Prata. mpwolke","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-08T20:41:42.470088Z","iopub.execute_input":"2024-04-08T20:41:42.471343Z","iopub.status.idle":"2024-04-08T20:41:45.431551Z","shell.execute_reply.started":"2024-04-08T20:41:42.471286Z","shell.execute_reply":"2024-04-08T20:41:45.429900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Competiton Citation:\n\n@misc{leash-BELKA,\n\n    author = {Andrew Blevins, Ian K Quigley, Brayden J Halverson, Nate Wilkinson, Rebecca S Levin, Agastya Pulapaka, Walter Reade, Addison Howard},\n    \n    title = {Leash Bio - Predict New Medicines with BELKA},\n    publisher = {Kaggle},\n    year = {2024},\n    url = {https://kaggle.com/competitions/leash-BELKA}\n}","metadata":{}},{"cell_type":"markdown","source":"#Dataset Big Encoded Library for Chemical Assessment (BELKA)\n\nLeash Bio - Predict New Medicines with BELKA\n\nPredict small molecule-protein interactions using the Big Encoded Library for Chemical Assessment (BELKA)\n\n\"In this competition, you’ll develop machine learning (ML) models to predict the binding affinity of small molecules to specific protein targets – a critical step in drug development for the pharmaceutical industry that would pave the way for more accurate drug discovery. You’ll help predict which drug-like small molecules (chemicals) will bind to three possible protein targets.\"\n\nhttps://www.kaggle.com/competitions/leash-BELKA/overview","metadata":{}},{"cell_type":"markdown","source":"PROTEIN STRUCTURE PREDICTION\n\n\"The most difficult problem in protein structure prediction is the modeling of proteins which have no solved structures that can be used as template, commonly referred \"ab initio\" or \"free modeling (FM)\" modeling.\"\n\n![](https://zhanglab.dcmb.med.umich.edu/image/Protein_design.gif)https://zhanggroup.org/research/","metadata":{}},{"cell_type":"markdown","source":"#Why is protein function prediction challenging?\n\nArticle: \"A (not so) Quick Introduction to Protein Function Prediction\n\nAuthor: Predrag Radivojac (Indiana University) -  September 27, 2013\n\n\"There are a number of biological and computational challenges that make protein function prediction difficult:\"\n\n\"Biologically, protein function is determined in the context of the organism and is rarely, if ever, fully determined by any single experiment or publication. In addition, some experiments cannot be performed in some organisms for a variety of biological, budgetary, or ethical reasons. Some experiments are performed in vitro and may not faithfully reflect the protein’s activity in vivo.\nFor example, if a protein is modified post-translationally, it might not carry out its function\nin vitro, because the modifying enzyme that activates the protein is missing. The available function data may also contain errors caused by misinterpretation of experiments, curation errors, experimental biases, and so on. In short, the data provided in biological databases is (seriously) incomplete, biased, and noisy. In fact, because experimental annotations are incomplete, it is fair to question the reliability of the ground truth on which we evaluate predictors.\"\n\n\"To make things worse, current biological databases are gene-centric. That is, they list functions of all protein forms corresponding to their parent gene as one entry. At this time it is difficult to disentangle which protein form has what function.\"\n\n\"Computationally, one can view protein function prediction as a multi-label classification problem or as an instance of structured-output learning in which the goal is to output a consistent subgraph (Tˆ) of a graph (O).\"\n\n\"Another challenge comes from the data integration perspective, where diverse biological data (sequence, structure, interactions, etc.) need to be dealt with.\"\n\n\"Finally, performance evaluation involves developing good similarity functions between pairs of consistent subgraphs in the ontology. It is not entirely clear how this should be done given that there are underlying relationships between terms in the ontology and differences in resolution with respect to which function is described in different branches of the ontology. The ontologies are large; there are about 40,000 terms in MFO, BPO, and CCO combined.\"\n\n\"Yet a typical protein is annotated by only a small number of terms. Even worse, the number of GO\nterms a protein is experimentally annotated with appears to have a scale-freelike distribution. Understanding which proteins are annotated with many terms and which are annotated with only a few is very important.\"\n\nhttps://www.biofunctionprediction.org/cafa-targets/Introduction_to_protein_prediction.pdf","metadata":{}},{"cell_type":"code","source":"submission = pd.read_csv('/kaggle/input/leash-BELKA/sample_submission.csv')\nsubmission.tail()","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:41:58.653396Z","iopub.execute_input":"2024-04-08T20:41:58.653772Z","iopub.status.idle":"2024-04-08T20:41:59.490789Z","shell.execute_reply.started":"2024-04-08T20:41:58.653741Z","shell.execute_reply":"2024-04-08T20:41:59.489503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#SMILES (Simplified Molecular Input Line Entry System)\n\nWhat is SMILES?\n\n\"SMILES (Simplified Molecular Input Line Entry System) is a chemical notation that allows a user to represent a chemical structure in a way that can be used by the computer. SMILES is an easily learned and flexible notation. The SMILES notation requires that you learn a handful of rules. You do not need to worry about ambiguous representations because the software will automatically reorder your entry into a unique SMILES string when necessary.\"\n\n\"SMILES was developed through funding from the U.S. Environmental Protection Agency, Mid-Continent Ecology Division-Duluth, (MED-Duluth) Duluth, MN to the Medicinal Chemistry Project at Pomona College, Claremont, CA and the Computer Sciences Corporation, Duluth, MN. Several publications discuss SMILES in more detail, including Anderson et al. 1987, Weininger 1988, Weininger et al. 1989, and Hunter et al., 1987.\"\n\n\"SMILES has five basic syntax rules which must be observed. If basic rules of chemistry are not followed in SMILES entry, the system will warn the user and ask that the structure be edited or reentered. For example, if the user places too many bonds on an atom, a SMILES warning will appear that the structure is impossible. The rules are described below and some examples are provided.\"\n\nhttps://archive.epa.gov/med/med_archive_03/web/html/smiles.html","metadata":{}},{"cell_type":"markdown","source":"![](https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcR8-7zZqV4svtho3zSFythbMva_2BnjsRxsuA&s)https://cmpe.boun.edu.tr/~hakime.ozturk/smilesvec.html","metadata":{}},{"cell_type":"markdown","source":"#SMILES\n\nbuildingblock1_smiles - The structure, in SMILES, of the first building block\n\nbuildingblock2_smiles - The structure, in SMILES, of the second building block\n\nbuildingblock3_smiles - The structure, in SMILES, of the third building block\n\nmolecule_smiles - The structure of the fully assembled molecule, in SMILES. This includes the three building blocks and the triazine core. Note we use a [Dy] as the stand-in for the DNA linker.\n\nhttps://www.kaggle.com/competitions/leash-BELKA/data","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/leash-BELKA/train.csv', nrows=1000)\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:42:07.053607Z","iopub.execute_input":"2024-04-08T20:42:07.054089Z","iopub.status.idle":"2024-04-08T20:42:07.083851Z","shell.execute_reply.started":"2024-04-08T20:42:07.054050Z","shell.execute_reply":"2024-04-08T20:42:07.082458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#I updated just to include the 1st Draw by Chemdatafarm (input 8)","metadata":{}},{"cell_type":"code","source":"#This is the cheminformatics package that will do most of the heavy lifting for us\n!pip install rdkit","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:43:32.840045Z","iopub.execute_input":"2024-04-08T20:43:32.840987Z","iopub.status.idle":"2024-04-08T20:43:51.888135Z","shell.execute_reply.started":"2024-04-08T20:43:32.840935Z","shell.execute_reply":"2024-04-08T20:43:51.886933Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Chemdatafarmer  https://www.kaggle.com/code/chemdatafarmer/additional-seh-data/notebook\n#By Meer Atif https://www.kaggle.com/code/meeratif/smiles-open-problems\n\nfrom rdkit import Chem\nfrom rdkit.Chem import Draw, AllChem\nfrom rdkit import RDLogger","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:44:01.176128Z","iopub.execute_input":"2024-04-08T20:44:01.176636Z","iopub.status.idle":"2024-04-08T20:44:01.472576Z","shell.execute_reply.started":"2024-04-08T20:44:01.176592Z","shell.execute_reply":"2024-04-08T20:44:01.471034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Visualize some of the original data","metadata":{}},{"cell_type":"code","source":"#https://www.rdkit.org/docs/source/rdkit.Chem.Draw.html\n\n#By Chemdatafarmer  https://www.kaggle.com/code/chemdatafarmer/additional-seh-data/notebook\n\n###Visualize some of the original data\nDraw.MolsToGridImage([Chem.MolFromSmiles(x) for x in train['molecule_smiles'][0:4]], molsPerRow=4, subImgSize=(400,300))","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:44:23.683326Z","iopub.execute_input":"2024-04-08T20:44:23.684659Z","iopub.status.idle":"2024-04-08T20:44:23.737860Z","shell.execute_reply.started":"2024-04-08T20:44:23.684608Z","shell.execute_reply":"2024-04-08T20:44:23.736493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Trying another visualization with more rows. And Changing the size of the images","metadata":{"execution":{"iopub.status.busy":"2024-04-08T21:03:54.774297Z","iopub.execute_input":"2024-04-08T21:03:54.774763Z","iopub.status.idle":"2024-04-08T21:03:54.812426Z","shell.execute_reply.started":"2024-04-08T21:03:54.774733Z","shell.execute_reply":"2024-04-08T21:03:54.810660Z"}}},{"cell_type":"code","source":"#https://greglandrum.github.io/rdkit-blog/posts/2023-10-25-molsmatrixtogridimage.html\n#https://www.rdkit.org/docs/source/rdkit.Chem.Draw.html\n#By Chemdatafarmer  https://www.kaggle.com/code/chemdatafarmer/additional-seh-data/notebook\n\n###Visualize some of the original data\nDraw.MolsToGridImage([Chem.MolFromSmiles(x) for x in train['buildingblock3_smiles'][0:12]], molsPerRow=3, subImgSize=(200,200))","metadata":{"execution":{"iopub.status.busy":"2024-04-08T21:13:42.354882Z","iopub.execute_input":"2024-04-08T21:13:42.355353Z","iopub.status.idle":"2024-04-08T21:13:42.402020Z","shell.execute_reply.started":"2024-04-08T21:13:42.355318Z","shell.execute_reply":"2024-04-08T21:13:42.400488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#DuckDB and SQL  to open the parquet files\n\nWHAT IS DUCKDB?\n\n\"DuckDB is an in-process SQL OLAP database, which means it runs within the same process as the application using it. This unique feature allows DuckDB to offer the advantages of a database without the complexities of managing one. But, as with any software concept, the best way to learn is to dive in and get your hands dirty.\"\n\n![](https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcTCYPPtcI1sYekiyBCZfCcKY2XE5MD3jVzpDQ&s)https://motherduck.com/blog/duckdb-tutorial-for-beginners/","metadata":{}},{"cell_type":"markdown","source":"#I wasn't able to open this huge data. Only with Andrew's D. Blevins snippets  ","metadata":{}},{"cell_type":"code","source":"!pip install duckdb","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-08T20:47:52.342387Z","iopub.execute_input":"2024-04-08T20:47:52.343012Z","iopub.status.idle":"2024-04-08T20:48:08.744144Z","shell.execute_reply.started":"2024-04-08T20:47:52.342952Z","shell.execute_reply":"2024-04-08T20:48:08.742244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Andrew D. Blevins https://www.kaggle.com/code/andrewdblevins/leash-tutorial-ecfps-and-random-forest\n\nimport duckdb\nimport pandas as pd\n\ntrain_path = '/kaggle/input/leash-BELKA/train.parquet'\ntest_path = '/kaggle/input/leash-BELKA/test.parquet'\n\ncon = duckdb.connect()\n\ndf = con.query(f\"\"\"(SELECT *\n                        FROM parquet_scan('{train_path}')\n                        WHERE binds = 0\n                        ORDER BY random()\n                        LIMIT 30000)\n                        UNION ALL\n                        (SELECT *\n                        FROM parquet_scan('{train_path}')\n                        WHERE binds = 1\n                        ORDER BY random()\n                        LIMIT 30000)\"\"\").df()\n\ncon.close()","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:48:15.828670Z","iopub.execute_input":"2024-04-08T20:48:15.829231Z","iopub.status.idle":"2024-04-08T20:49:09.650403Z","shell.execute_reply.started":"2024-04-08T20:48:15.829169Z","shell.execute_reply":"2024-04-08T20:49:09.649232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Voilà Le Parquet file opened with DuckDB and SQL","metadata":{}},{"cell_type":"code","source":"df.tail()","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:49:17.813205Z","iopub.execute_input":"2024-04-08T20:49:17.813638Z","iopub.status.idle":"2024-04-08T20:49:17.830156Z","shell.execute_reply.started":"2024-04-08T20:49:17.813602Z","shell.execute_reply":"2024-04-08T20:49:17.828994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Let's check some Smiles to apply later","metadata":{}},{"cell_type":"code","source":"#HSA sm_name - buildingblock1_smiles\n\ntrain.iloc[1,1]","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:49:37.701398Z","iopub.execute_input":"2024-04-08T20:49:37.702510Z","iopub.status.idle":"2024-04-08T20:49:37.710388Z","shell.execute_reply.started":"2024-04-08T20:49:37.702461Z","shell.execute_reply":"2024-04-08T20:49:37.708924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#HSA protein_name - buildingblock2_smiles\n\ntrain.iloc[1,2]","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:49:41.933586Z","iopub.execute_input":"2024-04-08T20:49:41.933971Z","iopub.status.idle":"2024-04-08T20:49:41.941749Z","shell.execute_reply.started":"2024-04-08T20:49:41.933935Z","shell.execute_reply":"2024-04-08T20:49:41.940359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#HSA protein_name - buildingblock3_smiles\n\ntrain.iloc[1,3]","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:49:46.231694Z","iopub.execute_input":"2024-04-08T20:49:46.232079Z","iopub.status.idle":"2024-04-08T20:49:46.240247Z","shell.execute_reply.started":"2024-04-08T20:49:46.232050Z","shell.execute_reply":"2024-04-08T20:49:46.239026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#HSA protein_name - molecule_smiles\n\ntrain.iloc[1,4]","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:49:50.528838Z","iopub.execute_input":"2024-04-08T20:49:50.529458Z","iopub.status.idle":"2024-04-08T20:49:50.539944Z","shell.execute_reply.started":"2024-04-08T20:49:50.529411Z","shell.execute_reply":"2024-04-08T20:49:50.538112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[\"protein_name\"].value_counts()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-08T20:49:55.097615Z","iopub.execute_input":"2024-04-08T20:49:55.098074Z","iopub.status.idle":"2024-04-08T20:49:55.115009Z","shell.execute_reply.started":"2024-04-08T20:49:55.098039Z","shell.execute_reply":"2024-04-08T20:49:55.113442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Rob Mulla https://www.kaggle.com/code/robikscube/sign-language-recognition-eda-twitch-stream\n\nfig, ax = plt.subplots(figsize=(4, 4))\ntrain[\"protein_name\"].value_counts().head(7).sort_values(ascending=True).plot(\n    kind=\"barh\", color='g', ax=ax, title=\"Protein Names\"\n)\nax.set_xlabel(\"Number of Training Examples\")\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-04-08T20:49:59.994026Z","iopub.execute_input":"2024-04-08T20:49:59.994504Z","iopub.status.idle":"2024-04-08T20:50:00.280054Z","shell.execute_reply.started":"2024-04-08T20:49:59.994468Z","shell.execute_reply":"2024-04-08T20:50:00.278467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install pysmiles","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-08T20:50:34.265265Z","iopub.execute_input":"2024-04-08T20:50:34.265734Z","iopub.status.idle":"2024-04-08T20:50:49.755968Z","shell.execute_reply.started":"2024-04-08T20:50:34.265702Z","shell.execute_reply":"2024-04-08T20:50:49.754385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#https://mattermodeling.stackexchange.com/questions/6460/rdkit-and-pysmiles-results-differ-on-some-smiles-strings\n#Answered by RapelPy, Jul 31, 2021\n\nfrom rdkit import Chem\nfrom rdkit.Chem.Draw import IPythonConsole\nfrom rdkit.Chem import Draw\nimport pysmiles\n\ns1 = 'C#CCOc1ccc(CN)cc1.Cl'  # aromatic\ns2 = 'Br.Br.NCC1CCCN1c1cccnn1'  # kekulized\n\n\ndef show_implicit_h(smiles):\n    m = Chem.MolFromSmiles(smiles)\n    for atom in m.GetAtoms():\n        atom.SetProp('atomLabel', str(atom.GetIdx()))\n    m = Chem.AddHs(m)\n    return Draw.MolToImage(m, size=(300, 300))\n\n\nshow_implicit_h(s1)","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:50:55.839275Z","iopub.execute_input":"2024-04-08T20:50:55.839815Z","iopub.status.idle":"2024-04-08T20:50:56.960336Z","shell.execute_reply.started":"2024-04-08T20:50:55.839776Z","shell.execute_reply":"2024-04-08T20:50:56.958864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#https://mattermodeling.stackexchange.com/questions/6460/rdkit-and-pysmiles-results-differ-on-some-smiles-strings\n\nshow_implicit_h(s2)","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:51:10.629270Z","iopub.execute_input":"2024-04-08T20:51:10.629717Z","iopub.status.idle":"2024-04-08T20:51:10.660564Z","shell.execute_reply.started":"2024-04-08T20:51:10.629686Z","shell.execute_reply":"2024-04-08T20:51:10.659564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By ButtonWood https://mattermodeling.stackexchange.com/questions/6460/rdkit-and-pysmiles-results-differ-on-some-smiles-strings\n\nfrom rdkit import Chem\nfrom rdkit.Chem.Draw import IPythonConsole\nfrom rdkit.Chem import Draw\nfrom rdkit.Chem import rdDepictor\nfrom rdkit.Chem.Draw import rdMolDraw2D\n\nfrom IPython.display import SVG\nIPythonConsole.ipython_useSVG=True  \n\n\ndef mol_with_atom_index(mol):\n    for atom in mol.GetAtoms():\n        atom.SetAtomMapNum(atom.GetIdx())\n    return mol\n\n\nmol = Chem.MolFromSmiles(\"C#CCOc1ccc(CNc2nc(NCC3CCCN3c3cccnn3)nc(N[C@@H](CC#C)CC(=O)N[Dy])n2)cc1\")\nmol = mol_with_atom_index(mol)\nmc = Chem.Mol(mol.ToBinary())\n\n\ndrawer = rdMolDraw2D.MolDraw2DSVG(450, 200) \ndrawer.DrawMolecule(mc)\ndrawer.FinishDrawing()\n\nsvg = drawer.GetDrawingText()\ndisplay(SVG(svg.replace('svg:','')))","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:51:16.367117Z","iopub.execute_input":"2024-04-08T20:51:16.367512Z","iopub.status.idle":"2024-04-08T20:51:16.419769Z","shell.execute_reply.started":"2024-04-08T20:51:16.367483Z","shell.execute_reply":"2024-04-08T20:51:16.418881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By ButtonWood https://mattermodeling.stackexchange.com/questions/6460/rdkit-and-pysmiles-results-differ-on-some-smiles-strings\n\nfrom pysmiles import read_smiles\n\nsmiles = \"C#CCOc1ccc(CNc2nc(NCC3CCCN3c3cccnn3)nc(N[C@@H](CC#C)CC(=O)N[Dy])n2)cc1\"\nmolecule = read_smiles(smiles)\n\npysmiles_list = zip(molecule.nodes(data=\"SMILES\"), molecule.nodes(data=\"hcount\"))#Don't change hcount\nfor element in pysmiles_list:\n    print(element)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-08T20:51:23.211240Z","iopub.execute_input":"2024-04-08T20:51:23.212289Z","iopub.status.idle":"2024-04-08T20:51:23.225500Z","shell.execute_reply.started":"2024-04-08T20:51:23.212246Z","shell.execute_reply":"2024-04-08T20:51:23.223900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#BRD4 protein_name - molecule_smiles\n\ntrain.iloc[3,4]","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:51:35.635147Z","iopub.execute_input":"2024-04-08T20:51:35.635624Z","iopub.status.idle":"2024-04-08T20:51:35.645863Z","shell.execute_reply.started":"2024-04-08T20:51:35.635589Z","shell.execute_reply":"2024-04-08T20:51:35.644375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Meer Atif https://www.kaggle.com/code/meeratif/smiles-open-problems\n\n# Your SMILES string\nsmiles = \"C#CCOc1ccc(CNc2nc(NCc3cccc(Br)n3)nc(N[C@@H](CC#C)CC(=O)N[Dy])n2)cc1\"\n\n# Convert the SMILES string to an RDKit molecule object\nmol = Chem.MolFromSmiles(smiles)\n\n# Check if the conversion was successful\nif mol is not None:\n    # Generate a 2D depiction of the molecule\n    img = Draw.MolToImage(mol)    \n    # Save the image to a file\n#     img.save(\"molecule.png\")\nelse:\n    print(\"Invalid SMILES string\")","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:51:42.015157Z","iopub.execute_input":"2024-04-08T20:51:42.015649Z","iopub.status.idle":"2024-04-08T20:51:42.036019Z","shell.execute_reply.started":"2024-04-08T20:51:42.015614Z","shell.execute_reply":"2024-04-08T20:51:42.034354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Meer Atif https://www.kaggle.com/code/meeratif/smiles-open-problems\n\nimg","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:51:47.336227Z","iopub.execute_input":"2024-04-08T20:51:47.336677Z","iopub.status.idle":"2024-04-08T20:51:47.354891Z","shell.execute_reply.started":"2024-04-08T20:51:47.336645Z","shell.execute_reply":"2024-04-08T20:51:47.353218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Updated https://networkx.org/documentation/stable/reference/readwrite/matrix_market.html\n\n#https://stackoverflow.com/questions/57062757/how-to-generate-a-graph-from-a-smiles-molecule-representation\n#By Davide Fiocco answered Jul 16, 2019\n\nfrom pysmiles import read_smiles\nimport networkx as nx\n\nimport scipy as sp\nimport io  # Use BytesIO as a stand-in for a Python file object\n    \nsmiles = 'C#CCOc1ccc(CNc2nc(NCc3cccc(Br)n3)nc(N[C@@H](CC#C)CC(=O)N[Dy])n2)cc1'\nmol = read_smiles(smiles)\n    \n# atom vector (C only)\nprint(mol.nodes(data='SMILES'))#Don't change data\n# adjacency matrix\nprint(nx.to_scipy_sparse_array(mol))#Original 'networkx' has no attribute 'to_numpy_matrix'      ","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-08T20:51:53.213099Z","iopub.execute_input":"2024-04-08T20:51:53.213568Z","iopub.status.idle":"2024-04-08T20:51:53.234711Z","shell.execute_reply.started":"2024-04-08T20:51:53.213534Z","shell.execute_reply":"2024-04-08T20:51:53.233577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Davide Fiocco https://stackoverflow.com/questions/57062757/how-to-generate-a-graph-from-a-smiles-molecule-representation\n\nimport matplotlib.pyplot as plt\nelements = nx.get_node_attributes(mol, name = \"SMILES\")\nnx.draw(mol, with_labels=True, labels = elements, pos=nx.spring_layout(mol))\nplt.gca().set_aspect('equal')","metadata":{"execution":{"iopub.status.busy":"2024-04-08T20:52:04.810540Z","iopub.execute_input":"2024-04-08T20:52:04.810949Z","iopub.status.idle":"2024-04-08T20:52:05.071960Z","shell.execute_reply.started":"2024-04-08T20:52:04.810919Z","shell.execute_reply":"2024-04-08T20:52:05.070729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Davide Fiocco https://stackoverflow.com/questions/57062757/how-to-generate-a-graph-from-a-smiles-molecule-representation\n\nsmiles1 = 'CC(O)Cn1cnc2c(N)ncnc21'\nmol = read_smiles(smiles1)\n    \n# atom vector (C only)\nprint(mol.nodes(data='SMILES'))#Don't change data\n# adjacency matrix\nprint(nx.to_scipy_sparse_array(mol))","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-08T20:52:20.957926Z","iopub.execute_input":"2024-04-08T20:52:20.958337Z","iopub.status.idle":"2024-04-08T20:52:20.967604Z","shell.execute_reply.started":"2024-04-08T20:52:20.958306Z","shell.execute_reply":"2024-04-08T20:52:20.966042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Davide Fiocco https://stackoverflow.com/questions/57062757/how-to-generate-a-graph-from-a-smiles-molecule-representation\n\nelements = nx.get_node_attributes(mol, name = \"SMILES\")\nnx.draw(mol, with_labels=True, labels = elements, pos=nx.spring_layout(mol))\nplt.gca().set_aspect('equal')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-04-08T20:52:33.902790Z","iopub.execute_input":"2024-04-08T20:52:33.903442Z","iopub.status.idle":"2024-04-08T20:52:34.130277Z","shell.execute_reply.started":"2024-04-08T20:52:33.903397Z","shell.execute_reply":"2024-04-08T20:52:34.128884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#I still don't know if it's an aromatic or a kekulized SMILE.\n\n![](https://image.slidesharecdn.com/kekulization-170820105806/85/we-need-to-talk-about-kekulization-aromaticity-and-smiles-10-320.jpg?cb=1666361053)https://pt.slideshare.net/slideshow/we-need-to-talk-about-kekulization-aromaticity-and-smiles/78992741","metadata":{}},{"cell_type":"markdown","source":"#Acknowledgements:\n\nAndrew D. Blevins https://www.kaggle.com/code/andrewdblevins/leash-tutorial-ecfps-and-random-forest\n\nChemdatafarmer  https://www.kaggle.com/code/chemdatafarmer/additional-seh-data/notebook\n\nRob Mulla https://www.kaggle.com/code/robikscube/sign-language-recognition-eda-twitch-stream\n\nMeer Atif https://www.kaggle.com/code/meeratif/smiles-open-problems\n\nDavide Fiocco https://stackoverflow.com/questions/57062757/how-to-generate-a-graph-from-a-smiles-molecule-representation\n\nhttps://www.rdkit.org/docs/GettingStartedInPython.html\n\nhttps://github.com/rdkit/rdkit/discussions/5769\n\nmpwolke https://www.kaggle.com/code/mpwolke/pysmiles-rdkit-scperturbations/notebook","metadata":{}}]}