{
  "id": 491394,
  "title": "📊📈🔍🧬 EDA & buildingblock SMILES",
  "url": "/competitions/leash-BELKA/discussion/491394",
  "author_name": "",
  "post_date": "2024-04-05T17:50:35.196593500Z",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<table>\n<thead>\n<tr>\n<th>binds</th>\n<th>counts</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>293,656,924</td>\n</tr>\n<tr>\n<td>1</td>\n<td>1,589,906</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<table>\n<thead>\n<tr>\n<th>buildingblock</th>\n<th># unique count</th>\n<th># unique count for binds = 1</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>buildingblock1</td>\n<td>270</td>\n<td>270</td>\n</tr>\n<tr>\n<td>buildingblock2</td>\n<td>692</td>\n<td>692</td>\n</tr>\n<tr>\n<td>buildingblock3</td>\n<td>871</td>\n<td>870 (missing - NCc1ccccn1)</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h1>is it all permutations ?</h1>\n<ul>\n<li>270 x 692 x 871 = 162737640<br>\ntotal binds = 0 combinations = 98415610 =&gt; 60.5% of permutations are in binds = 0<br>\n<strong>Is it possible to use it as augmentation to generate data instead take all training data ??</strong></li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>protein_name</th>\n<th>binds</th>\n<th>counts</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>HSA</td>\n<td>0</td>\n<td>98,007,200</td>\n</tr>\n<tr>\n<td>BRD4</td>\n<td>0</td>\n<td>97,958,646</td>\n</tr>\n<tr>\n<td>sEH</td>\n<td>0</td>\n<td>97,691,078</td>\n</tr>\n<tr>\n<td>sEH</td>\n<td>1</td>\n<td>724,532</td>\n</tr>\n<tr>\n<td>BRD4</td>\n<td>1</td>\n<td>456,964</td>\n</tr>\n<tr>\n<td>HSA</td>\n<td>1</td>\n<td>408,410</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fe4f73ab8ba366bce2f707d2010ad8589%2FBlank%2010%20Grids%20Collage.png?generation=1712338876014473&amp;alt=media\"></p>\n<ul>\n<li>from <a href=\"https://www.kaggle.com/code/seshurajup/eda-smiles?scriptVersionId=170520559\" target=\"_blank\">Notebook - EDA SMILES</a></li>\n</ul>\n<hr>\n<table>\n<thead>\n<tr>\n<th>Molecule Type</th>\n<th>Count</th>\n<th>Percentage</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Total</td>\n<td>98,415,610</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>Binds = 1</td>\n<td>1,509,779</td>\n<td>1.53%</td>\n</tr>\n<tr>\n<td>Common</td>\n<td>1,509,717</td>\n<td>1.53%</td>\n</tr>\n<tr>\n<td>Binds = 0 but not in Binds = 1</td>\n<td>96,905,831</td>\n<td>98.47%</td>\n</tr>\n<tr>\n<td>Binds = 1 but not in Binds = 0</td>\n<td>62</td>\n<td>0.00%</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><strong>Created mini dataset with 1:2 ratio of binds=1 to binds=0 ratio</strong><br>\n<a href=\"https://www.kaggle.com/datasets/seshurajup/leash-belka-mini\" target=\"_blank\">Dataset - leash-BELKA-mini</a></p>\n<p>So, we can use 270 x 692 x 871 with randomly select any 96,905,831 molecule_smiles - as binds=0 for training code faster<br>\n-- Do we need <strong>293,656,924</strong> ? if permutations approach works !!</p>",
  "messages": [
    {
      "id": "2737328",
      "postDate": "04/05/2024 17:50:35",
      "content": "<table>\n<thead>\n<tr>\n<th>binds</th>\n<th>counts</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>293,656,924</td>\n</tr>\n<tr>\n<td>1</td>\n<td>1,589,906</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<table>\n<thead>\n<tr>\n<th>buildingblock</th>\n<th># unique count</th>\n<th># unique count for binds = 1</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>buildingblock1</td>\n<td>270</td>\n<td>270</td>\n</tr>\n<tr>\n<td>buildingblock2</td>\n<td>692</td>\n<td>692</td>\n</tr>\n<tr>\n<td>buildingblock3</td>\n<td>871</td>\n<td>870 (missing - NCc1ccccn1)</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h1>is it all permutations ?</h1>\n<ul>\n<li>270 x 692 x 871 = 162737640<br>\ntotal binds = 0 combinations = 98415610 =&gt; 60.5% of permutations are in binds = 0<br>\n<strong>Is it possible to use it as augmentation to generate data instead take all training data ??</strong></li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>protein_name</th>\n<th>binds</th>\n<th>counts</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>HSA</td>\n<td>0</td>\n<td>98,007,200</td>\n</tr>\n<tr>\n<td>BRD4</td>\n<td>0</td>\n<td>97,958,646</td>\n</tr>\n<tr>\n<td>sEH</td>\n<td>0</td>\n<td>97,691,078</td>\n</tr>\n<tr>\n<td>sEH</td>\n<td>1</td>\n<td>724,532</td>\n</tr>\n<tr>\n<td>BRD4</td>\n<td>1</td>\n<td>456,964</td>\n</tr>\n<tr>\n<td>HSA</td>\n<td>1</td>\n<td>408,410</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fe4f73ab8ba366bce2f707d2010ad8589%2FBlank%2010%20Grids%20Collage.png?generation=1712338876014473&amp;alt=media\"></p>\n<ul>\n<li>from <a href=\"https://www.kaggle.com/code/seshurajup/eda-smiles?scriptVersionId=170520559\" target=\"_blank\">Notebook - EDA SMILES</a></li>\n</ul>\n<hr>\n<table>\n<thead>\n<tr>\n<th>Molecule Type</th>\n<th>Count</th>\n<th>Percentage</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Total</td>\n<td>98,415,610</td>\n<td>100.00%</td>\n</tr>\n<tr>\n<td>Binds = 1</td>\n<td>1,509,779</td>\n<td>1.53%</td>\n</tr>\n<tr>\n<td>Common</td>\n<td>1,509,717</td>\n<td>1.53%</td>\n</tr>\n<tr>\n<td>Binds = 0 but not in Binds = 1</td>\n<td>96,905,831</td>\n<td>98.47%</td>\n</tr>\n<tr>\n<td>Binds = 1 but not in Binds = 0</td>\n<td>62</td>\n<td>0.00%</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p><strong>Created mini dataset with 1:2 ratio of binds=1 to binds=0 ratio</strong><br>\n<a href=\"https://www.kaggle.com/datasets/seshurajup/leash-belka-mini\" target=\"_blank\">Dataset - leash-BELKA-mini</a></p>\n<p>So, we can use 270 x 692 x 871 with randomly select any 96,905,831 molecule_smiles - as binds=0 for training code faster<br>\n-- Do we need <strong>293,656,924</strong> ? if permutations approach works !!</p>",
      "rawMarkdown": "| binds | counts |\n|-------|-------------|\n| 0     | 293,656,924   |\n| 1     | 1,589,906     |\n\n---\n\n| buildingblock  |  # unique count | # unique count for binds = 1 |\n| --- | --- | --- |\n| buildingblock1 | 270 | 270 | \n| buildingblock2 | 692 | 692 |\n| buildingblock3 |  871 | 870 (missing - NCc1ccccn1)\n\n---\n\n# is it all permutations ? \n- 270 x 692 x 871 = 162737640\ntotal binds = 0 combinations = 98415610 => 60.5% of permutations are in binds = 0\n**Is it possible to use it as augmentation to generate data instead take all training data ??**\n\n| protein_name | binds | counts |\n|--------------|-------|--------------------|\n| HSA          | 0     | 98,007,200           |\n| BRD4         | 0     | 97,958,646           |\n| sEH          | 0     | 97,691,078           |\n| sEH          | 1     | 724,532             |\n| BRD4         | 1     | 456,964             |\n| HSA          | 1     | 408,410             |\n\n---\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fe4f73ab8ba366bce2f707d2010ad8589%2FBlank%2010%20Grids%20Collage.png?generation=1712338876014473&alt=media)\n\n- from [Notebook - EDA SMILES](https://www.kaggle.com/code/seshurajup/eda-smiles?scriptVersionId=170520559)\n\n---\n\n\n| Molecule Type                          | Count     | Percentage |\n|---------------------------------------|-----------|------------|\n| Total                                 | 98,415,610  | 100.00%    |\n| Binds = 1                             | 1,509,779   | 1.53%      |\n| Common                                | 1,509,717   | 1.53%      |\n| Binds = 0 but not in Binds = 1       | 96,905,831  | 98.47%     |\n| Binds = 1 but not in Binds = 0       | 62        | 0.00%      |\n\n---\n\n**Created mini dataset with 1:2 ratio of binds=1 to binds=0 ratio**\n[Dataset - leash-BELKA-mini](https://www.kaggle.com/datasets/seshurajup/leash-belka-mini)\n\nSo, we can use 270 x 692 x 871 with randomly select any 96,905,831 molecule_smiles - as binds=0 for training code faster\n-- Do we need **293,656,924** ? if permutations approach works !!",
      "votes": null
    },
    {
      "id": "2739583",
      "postDate": "04/07/2024 06:30:39",
      "content": "<p>Hi. can we do down sampling here?</p>",
      "rawMarkdown": "Hi. can we do down sampling here?",
      "votes": null
    },
    {
      "id": "2739594",
      "postDate": "04/07/2024 06:38:12",
      "content": "<p>Yes <a href=\"https://www.kaggle.com/faustrahul\" target=\"_blank\">@faustrahul</a>, same way i did the mini dataset with 1:2 ratio of binds=1/0  targets</p>",
      "rawMarkdown": "Yes @faustrahul, same way i did the mini dataset with 1:2 ratio of binds=1/0  targets",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2739583,
      "author_name": "faustrahul",
      "author_url": "",
      "post_date": "04/07/2024 06:30:39",
      "content": "<p>Hi. can we do down sampling here?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2739594,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "04/07/2024 06:38:12",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/faustrahul\" target=\"_blank\">@faustrahul</a>, same way i did the mini dataset with 1:2 ratio of binds=1/0  targets</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2737328": "| binds | counts |\n|-------|-------------|\n| 0     | 293,656,924   |\n| 1     | 1,589,906     |\n\n---\n\n| buildingblock  |  # unique count | # unique count for binds = 1 |\n| --- | --- | --- |\n| buildingblock1 | 270 | 270 | \n| buildingblock2 | 692 | 692 |\n| buildingblock3 |  871 | 870 (missing - NCc1ccccn1)\n\n---\n\n# is it all permutations ? \n- 270 x 692 x 871 = 162737640\ntotal binds = 0 combinations = 98415610 => 60.5% of permutations are in binds = 0\n**Is it possible to use it as augmentation to generate data instead take all training data ??**\n\n| protein_name | binds | counts |\n|--------------|-------|--------------------|\n| HSA          | 0     | 98,007,200           |\n| BRD4         | 0     | 97,958,646           |\n| sEH          | 0     | 97,691,078           |\n| sEH          | 1     | 724,532             |\n| BRD4         | 1     | 456,964             |\n| HSA          | 1     | 408,410             |\n\n---\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fe4f73ab8ba366bce2f707d2010ad8589%2FBlank%2010%20Grids%20Collage.png?generation=1712338876014473&alt=media)\n\n- from [Notebook - EDA SMILES](https://www.kaggle.com/code/seshurajup/eda-smiles?scriptVersionId=170520559)\n\n---\n\n\n| Molecule Type                          | Count     | Percentage |\n|---------------------------------------|-----------|------------|\n| Total                                 | 98,415,610  | 100.00%    |\n| Binds = 1                             | 1,509,779   | 1.53%      |\n| Common                                | 1,509,717   | 1.53%      |\n| Binds = 0 but not in Binds = 1       | 96,905,831  | 98.47%     |\n| Binds = 1 but not in Binds = 0       | 62        | 0.00%      |\n\n---\n\n**Created mini dataset with 1:2 ratio of binds=1 to binds=0 ratio**\n[Dataset - leash-BELKA-mini](https://www.kaggle.com/datasets/seshurajup/leash-belka-mini)\n\nSo, we can use 270 x 692 x 871 with randomly select any 96,905,831 molecule_smiles - as binds=0 for training code faster\n-- Do we need **293,656,924** ? if permutations approach works !!",
    "2739583": "Hi. can we do down sampling here?",
    "2739594": "Yes @faustrahul, same way i did the mini dataset with 1:2 ratio of binds=1/0  targets"
  },
  "source": "meta"
}