{
  "id": 498858,
  "title": "[lb0.602] low-resource fast 2d GNN example is here!!!",
  "url": "/competitions/leash-BELKA/discussion/498858",
  "author_name": "hengck23",
  "post_date": "2024-04-29T20:32:41.319000",
  "votes": 40,
  "comment_count": 48,
  "views": 0,
  "content": "<p>example code: <br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb6-02-graph-nn-example\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb6-02-graph-nn-example</a></p>\n<p>converted smiles string to graph object:<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F38c7a842993b6c5b5928bc426a088457%2FSelection_109.png?generation=1715558729915059&amp;alt=media\"></p>\n<h1>model :</h1>\n<ul>\n<li>pytorch geometric framework</li>\n<li>simple MPNNModel, 4 layers, hidden dim=96</li>\n<li>see: <a href=\"https://github.com/chaitjo/geometric-gnn-dojo/blob/main/geometric_gnn_101.ipynb\" target=\"_blank\">https://github.com/chaitjo/geometric-gnn-dojo/blob/main/geometric_gnn_101.ipynb</a></li>\n<li>model file size: 7.7mb  </li>\n</ul>\n<h1>graph creation:</h1>\n<ul>\n<li>atom and bond one-hot features from rdkit</li>\n<li>see :  from <a href=\"https://github.com/LiZhang30/GPCNDTA/blob/main/utils/DrugGraph.py\" target=\"_blank\">https://github.com/LiZhang30/GPCNDTA/blob/main/utils/DrugGraph.py</a></li>\n<li>node dim = 70, edge dim=6</li>\n<li>smiles to graph conversion: 1m takes 1 min (multiprocessing pool with 64 workers)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8b2f652ce3f46490ae63acb7335f9e6e%2FSelection_105.png?generation=1715489514039434&amp;alt=media\"></p>\n<h1>other tricks:</h1>\n<ul>\n<li>np.packbits to reduce ram for faster cpu-to-gpu transfer</li>\n<li>avoid using num workers&gt;0 in pgy DataLoader.<br>\nsee: <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/500877\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/discussion/500877</a></li>\n</ul>\n<h1>hyperprameters:</h1>\n<ul>\n<li>train = 50/70m samples, valid=5m (random split)</li>\n<li>batch size = 5000</li>\n<li>training speed: 1m samples takes 12 min, 1 gpu</li>\n<li>lr 1e-3/1e-4</li>\n</ul>\n<p>pyg batch collate function is not efficient enough to give 100% gpu utilization rate for large batch size of 5000.<br>\neven worse, we cannot set num workers&gt;0.<br>\ni estimate that we can probably go up to 1m samples at 5 min with better code.</p>\n<hr>\n<p>myth buster</p>\n<ul>\n<li>GNN is not correct and doesn't work ? not true. if we are memorizing sharing block predictions (i.e. train and test distribution are similiar), it needs not to be correct.</li>\n<li>GNN is slow: not true we do not need many layers.</li>\n<li>graph conversion is slow and takes much memory: not true</li>\n</ul>",
  "messages": [
    {
      "id": 2783615,
      "postDate": "2024-04-29T20:32:41.320Z",
      "content": "<p>example code: <br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb6-02-graph-nn-example\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb6-02-graph-nn-example</a></p>\n<p>converted smiles string to graph object:<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F38c7a842993b6c5b5928bc426a088457%2FSelection_109.png?generation=1715558729915059&amp;alt=media\"></p>\n<h1>model :</h1>\n<ul>\n<li>pytorch geometric framework</li>\n<li>simple MPNNModel, 4 layers, hidden dim=96</li>\n<li>see: <a href=\"https://github.com/chaitjo/geometric-gnn-dojo/blob/main/geometric_gnn_101.ipynb\" target=\"_blank\">https://github.com/chaitjo/geometric-gnn-dojo/blob/main/geometric_gnn_101.ipynb</a></li>\n<li>model file size: 7.7mb  </li>\n</ul>\n<h1>graph creation:</h1>\n<ul>\n<li>atom and bond one-hot features from rdkit</li>\n<li>see :  from <a href=\"https://github.com/LiZhang30/GPCNDTA/blob/main/utils/DrugGraph.py\" target=\"_blank\">https://github.com/LiZhang30/GPCNDTA/blob/main/utils/DrugGraph.py</a></li>\n<li>node dim = 70, edge dim=6</li>\n<li>smiles to graph conversion: 1m takes 1 min (multiprocessing pool with 64 workers)</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8b2f652ce3f46490ae63acb7335f9e6e%2FSelection_105.png?generation=1715489514039434&amp;alt=media\"></p>\n<h1>other tricks:</h1>\n<ul>\n<li>np.packbits to reduce ram for faster cpu-to-gpu transfer</li>\n<li>avoid using num workers&gt;0 in pgy DataLoader.<br>\nsee: <a href=\"https://www.kaggle.com/competitions/leash-BELKA/discussion/500877\" target=\"_blank\">https://www.kaggle.com/competitions/leash-BELKA/discussion/500877</a></li>\n</ul>\n<h1>hyperprameters:</h1>\n<ul>\n<li>train = 50/70m samples, valid=5m (random split)</li>\n<li>batch size = 5000</li>\n<li>training speed: 1m samples takes 12 min, 1 gpu</li>\n<li>lr 1e-3/1e-4</li>\n</ul>\n<p>pyg batch collate function is not efficient enough to give 100% gpu utilization rate for large batch size of 5000.<br>\neven worse, we cannot set num workers&gt;0.<br>\ni estimate that we can probably go up to 1m samples at 5 min with better code.</p>\n<hr>\n<p>myth buster</p>\n<ul>\n<li>GNN is not correct and doesn't work ? not true. if we are memorizing sharing block predictions (i.e. train and test distribution are similiar), it needs not to be correct.</li>\n<li>GNN is slow: not true we do not need many layers.</li>\n<li>graph conversion is slow and takes much memory: not true</li>\n</ul>",
      "rawMarkdown": "example code: \nhttps://www.kaggle.com/code/hengck23/lb6-02-graph-nn-example\n\nconverted smiles string to graph object:\nhttps://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F38c7a842993b6c5b5928bc426a088457%2FSelection_109.png?generation=1715558729915059&alt=media)\n\n#model :\n- pytorch geometric framework\n- simple MPNNModel, 4 layers, hidden dim=96\n- see: https://github.com/chaitjo/geometric-gnn-dojo/blob/main/geometric_gnn_101.ipynb\n- model file size: 7.7mb  \n\n#graph creation:\n- atom and bond one-hot features from rdkit\n- see :  from https://github.com/LiZhang30/GPCNDTA/blob/main/utils/DrugGraph.py\n- node dim = 70, edge dim=6\n- smiles to graph conversion: 1m takes 1 min (multiprocessing pool with 64 workers)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8b2f652ce3f46490ae63acb7335f9e6e%2FSelection_105.png?generation=1715489514039434&alt=media)\n\n#other tricks:\n- np.packbits to reduce ram for faster cpu-to-gpu transfer\n- avoid using num workers>0 in pgy DataLoader.\n  see: https://www.kaggle.com/competitions/leash-BELKA/discussion/500877\n \n\n#hyperprameters:\n- train = 50/70m samples, valid=5m (random split)\n- batch size = 5000\n- training speed: 1m samples takes 12 min, 1 gpu\n- lr 1e-3/1e-4\n\npyg batch collate function is not efficient enough to give 100% gpu utilization rate for large batch size of 5000.\neven worse, we cannot set num workers>0.\ni estimate that we can probably go up to 1m samples at 5 min with better code.\n\n---\n\nmyth buster\n- GNN is not correct and doesn't work ? not true. if we are memorizing sharing block predictions (i.e. train and test distribution are similiar), it needs not to be correct.\n- GNN is slow: not true we do not need many layers.\n- graph conversion is slow and takes much memory: not true",
      "votes": 40
    },
    {
      "id": 2821861,
      "postDate": "2024-05-18T09:45:54.390Z",
      "content": "<p>i end up wring my own collate function. Hence we need not convert to pyg graph list at all, saving 90 min conversion time and 70gb. removing pyg data loader improve speed by 2.5x.</p>\n<p>Now I can train with full 98m SMILES and my graph-NN is same speed as CNN1d.</p>\n<pre><code>def (graph, index=None, device='cpu'):\n    if index is None:\n        index = np.((graph)).()\n    batch = (\n        x=[],\n        edge_index=[],\n        edge_attr=[],\n        batch=[],\n        idx=index\n    )\n    offset = \n    for b, i in (index):\n        N, edge, node_feature, edge_feature = graph[i]\n        batch.x.(node_feature)\n        batch.edge_attr.(edge_feature)\n        batch.edge_index.(edge.(int) + offset)\n        batch.batch += N * [b]\n        offset += N\n    batch.x = torch.(np.(batch.x)).(device)\n    batch.edge_attr = torch.(np.(batch.edge_attr)).(device)\n    batch.edge_index = torch.(np.(batch.edge_index).T).(device)\n    batch.batch = torch.(batch.batch).(device)\n    return batch\n\n\n.... more code here ....\n\nwhile epoch&lt;cfg.num_epoch:\n    shuffled_idx = train_idx.()\n    np.random.(shuffled_idx)\n    for t, index in (np.(,(shuffled_idx),cfg.train_batch_size)):\n        index = shuffled_idx[index:index+cfg.train_batch_size]\n        if (index)!=cfg.train_batch_size: continue #drop last\n\n        B = (index)\n        batch = (\n            graph = (train_graph,index,device=),\n            bind = torch.(train_bind[index]).().(),\n        )\n\n        net.()\n        net.output_type = [, ]\n\n        with torch.cuda.amp.(enabled=cfg.is_amp):\n            output = (batch)  #(net,batch) #\n            bce_loss = output[]\n</code></pre>\n<p>gpu utilization is now 70% and a queue can fix that.</p>",
      "rawMarkdown": "i end up wring my own collate function. Hence we need not convert to pyg graph list at all, saving 90 min conversion time and 70gb. removing pyg data loader improve speed by 2.5x.\n\nNow I can train with full 98m SMILES and my graph-NN is same speed as CNN1d.\n```\ndef my_collate(graph, index=None, device='cpu'):\n\tif index is None:\n\t\tindex = np.arange(len(graph)).tolist()\n\tbatch = dotdict(\n\t\tx=[],\n\t\tedge_index=[],\n\t\tedge_attr=[],\n\t\tbatch=[],\n\t\tidx=index\n\t)\n\toffset = 0\n\tfor b, i in enumerate(index):\n\t\tN, edge, node_feature, edge_feature = graph[i]\n\t\tbatch.x.append(node_feature)\n\t\tbatch.edge_attr.append(edge_feature)\n\t\tbatch.edge_index.append(edge.astype(int) + offset)\n\t\tbatch.batch += N * [b]\n\t\toffset += N\n\tbatch.x = torch.from_numpy(np.concatenate(batch.x)).to(device)\n\tbatch.edge_attr = torch.from_numpy(np.concatenate(batch.edge_attr)).to(device)\n\tbatch.edge_index = torch.from_numpy(np.concatenate(batch.edge_index).T).to(device)\n\tbatch.batch = torch.LongTensor(batch.batch).to(device)\n\treturn batch\n\n\n.... more code here ....\n\nwhile epoch<cfg.num_epoch:\n\tshuffled_idx = train_idx.copy()\n\tnp.random.shuffle(shuffled_idx)\n\tfor t, index in enumerate(np.arange(0,len(shuffled_idx),cfg.train_batch_size)):\n\t\tindex = shuffled_idx[index:index+cfg.train_batch_size]\n\t\tif len(index)!=cfg.train_batch_size: continue #drop last\n\n\t\tB = len(index)\n\t\tbatch = dotdict(\n\t\t\tgraph = my_collate(train_graph,index,device='cuda'),\n\t\t\tbind = torch.from_numpy(train_bind[index]).float().cuda(),\n\t\t)\n\n\t\tnet.train()\n\t\tnet.output_type = ['loss', 'infer']\n\n\t\twith torch.cuda.amp.autocast(enabled=cfg.is_amp):\n\t\t\toutput = net(batch)  #data_parallel(net,batch) #\n\t\t\tbce_loss = output['bce_loss']\n```\n\ngpu utilization is now 70% and a queue can fix that.",
      "votes": 5,
      "replies": [
        {
          "id": 2821876,
          "postDate": "2024-05-18T09:54:06.130Z",
          "content": "<p>queue for 100% gpu utilization</p>\n<pre><code>def (queue):\n    shuffled_idx = train_idx.()\n    np.random.(shuffled_idx)\n    for t, index in (np.(, (shuffled_idx), cfg.train_batch_size)):\n        index = shuffled_idx[index:index + cfg.train_batch_size]\n        if (index)!=cfg.train_batch_size: continue #drop last\n\n        B = (index)\n        batch = (\n            graph=(train_graph, index, device=),\n            bind=torch.(train_bind[index]).(),\n        )\n        queue.(batch)\n    queue.(None)\n\n\n\nwhile epoch&lt;cfg.num_epoch:\n    queue = multiprocessing.().(maxsize=)\n    train_loader_process = (target=train_loader_func, args=(queue,))\n    train_loader_process.()\n\n    num_train_batch = ((train_idx) // cfg.train_batch_size)\n    for t in (num_train_batch+):\n        batch = queue.()\n        if batch is None:\n            break\n        #cpu to dpu outside () ....\n        batch.bind = batch.bind.()\n        batch.graph.x = batch.graph.x.()\n        batch.graph.edge_attr = batch.graph.edge_attr.()\n        batch.graph.edge_index = batch.graph.edge_index.()\n        batch.graph.batch = batch.graph.batch.()\n</code></pre>",
          "rawMarkdown": "queue for 100% gpu utilization\n\n```\ndef train_loader_func(queue):\n\tshuffled_idx = train_idx.copy()\n\tnp.random.shuffle(shuffled_idx)\n\tfor t, index in enumerate(np.arange(0, len(shuffled_idx), cfg.train_batch_size)):\n\t\tindex = shuffled_idx[index:index + cfg.train_batch_size]\n\t\tif len(index)!=cfg.train_batch_size: continue #drop last\n\n\t\tB = len(index)\n\t\tbatch = dotdict(\n\t\t\tgraph=my_collate(train_graph, index, device='cpu'),\n\t\t\tbind=torch.from_numpy(train_bind[index]).float(),\n\t\t)\n\t\tqueue.put(batch)\n\tqueue.put(None)\n\n\n\nwhile epoch<cfg.num_epoch:\n\tqueue = multiprocessing.Manager().Queue(maxsize=8)\n\ttrain_loader_process = Process(target=train_loader_func, args=(queue,))\n\ttrain_loader_process.start()\n\n\tnum_train_batch = int(len(train_idx) // cfg.train_batch_size)\n\tfor t in range(num_train_batch+1):\n\t\tbatch = queue.get()\n\t\tif batch is None:\n\t\t\tbreak\n\t\t#cpu to dpu outside queue() ....\n\t\tbatch.bind = batch.bind.cuda()\n\t\tbatch.graph.x = batch.graph.x.cuda()\n\t\tbatch.graph.edge_attr = batch.graph.edge_attr.cuda()\n\t\tbatch.graph.edge_index = batch.graph.edge_index.cuda()\n\t\tbatch.graph.batch = batch.graph.batch.cuda()\n\n```",
          "votes": 2,
          "replies": [
            {
              "id": 2857690,
              "postDate": "2024-06-06T05:02:18.200Z",
              "content": "<p>Thank you for your code! my code got stuck at \"batch = queue.get()\" when the batch size is larger than 10. What batch_size do you use? Thanks!</p>",
              "rawMarkdown": "Thank you for your code! my code got stuck at \"batch = queue.get()\" when the batch size is larger than 10. What batch_size do you use? Thanks!"
            },
            {
              "id": 2857717,
              "postDate": "2024-06-06T05:24:48.060Z",
              "content": "<p>my batch is is 2000 to 5000.<br>\nrun your code without queue first. it should get utilization more than 80%. else my simple queue don't help</p>",
              "rawMarkdown": "my batch is is 2000 to 5000.\nrun your code without queue first. it should get utilization more than 80%. else my simple queue don't help",
              "votes": 1
            },
            {
              "id": 2857729,
              "postDate": "2024-06-06T05:29:25.613Z",
              "content": "<p>there is \"train_loader_process.join()\" at the end of the code with i didn't show.<br>\n(read more about python multiprocessing Process class)</p>",
              "rawMarkdown": "there is \"train_loader_process.join()\" at the end of the code with i didn't show.\n(read more about python multiprocessing Process class)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2809601,
      "postDate": "2024-05-12T20:57:03.150Z",
      "content": "<p>start download now! converted smiles string to graph object:<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4024364f683c7fa253bccb5e2a5530c1%2FSelection_106.png?generation=1715547404927195&amp;alt=media\"></p>",
      "rawMarkdown": "start download now! converted smiles string to graph object:\nhttps://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4024364f683c7fa253bccb5e2a5530c1%2FSelection_106.png?generation=1715547404927195&alt=media)",
      "votes": 5
    },
    {
      "id": 2808134,
      "postDate": "2024-05-12T04:57:34.180Z",
      "content": "<p>Another issue 2:<br>\nany kaggle has good method to save converted graph to file on disk?<br>\nif there is good solution, i will upload my converted graph to public dataset.</p>\n<hr>\n<p>i need a solution that include file compression. it also needs to be fast to compress and decompress the list of graph objects (e.g. parallel processing). I tried cPickle alone but it is not good enough.</p>",
      "rawMarkdown": "Another issue 2:\nany kaggle has good method to save converted graph to file on disk?\nif there is good solution, i will upload my converted graph to public dataset.\n\n---\n\ni need a solution that include file compression. it also needs to be fast to compress and decompress the list of graph objects (e.g. parallel processing). I tried cPickle alone but it is not good enough.",
      "votes": 3,
      "replies": [
        {
          "id": 2808425,
          "postDate": "2024-05-12T07:46:07.397Z",
          "content": "<p>Possibly useless, but this is how I handle it</p>\n<pre><code> () -&gt; [Data, ]:\n    mol = Chem.MolFromSmiles(smiles)\n    me = Chem.MolFromSmiles()\n    dy = Chem.MolFromSmiles()\n    mol = AllChem.ReplaceSubstructs(mol, dy, me)[]\n      mol:\n          \n    mol = Chem.AddHs(mol)\n    atom_features = get_atom_features(mol)\n    bond_features = get_bond_features(mol)\n    \n    \n    \n    \n    features = torch.tensor(atom_features, dtype=torch.)\n\n    \n    edge_indices = []\n    edge_attrs = []\n     bond  mol.GetBonds():\n        start, end = bond.GetBeginAtomIdx(), bond.GetEndAtomIdx()\n        edge_indices.append((start, end))\n        edge_indices.append((end, start))  \n         _  ():  \n            edge_attrs.append(bond_features[bond.GetIdx()])\n\n    edge_index = torch.tensor(edge_indices, dtype=torch.long).t().contiguous()\n    edge_attr = torch.tensor(edge_attrs, dtype=torch.)\n    data = Data(x = features, edge_index = edge_index, edge_attr = edge_attr)\n     pickle.dumps(data)\n\n () -&gt; dd.DataFrame:\n    column_order = [\n        , , ,\n        , , \n    ]\n     column  [, , , ]:\n        df[] = df[column].apply(smiles_to_graph)\n        df = df.drop(columns=[column])\n        df = df.rename(columns={: column})\n    df = df.reindex(columns=column_order)\n     df.columns.tolist() != column_order:\n        ()\n         ValueError(, column_order, , df.columns.tolist())\n\n     df\n</code></pre>\n<p><br>\nand then I save it with joblib dump. It's pretty fast, admittedly not as fast as it can be but it's tolerable for me</p>",
          "rawMarkdown": "Possibly useless, but this is how I handle it\n```python\ndef smiles_to_graph(smiles: str) -> Union[Data, None]:\n    mol = Chem.MolFromSmiles(smiles)\n    me = Chem.MolFromSmiles('C')\n    dy = Chem.MolFromSmiles('[Dy]')\n    mol = AllChem.ReplaceSubstructs(mol, dy, me)[0]\n    if not mol:\n        return None ##To be sure RDKit can parse it\n    mol = Chem.AddHs(mol)\n    atom_features = get_atom_features(mol)\n    bond_features = get_bond_features(mol)\n    #z, pos = get_visnet_features(mol) ##Takes too long\n    #z = torch.tensor(z, dtype=torch.long)\n    #pos = torch.tensor(pos, dtype=torch.float)\n    ## Create atom feature matrix\n    features = torch.tensor(atom_features, dtype=torch.float)\n\n    ## Edge index and edge feature matrix construction\n    edge_indices = []\n    edge_attrs = []\n    for bond in mol.GetBonds():\n        start, end = bond.GetBeginAtomIdx(), bond.GetEndAtomIdx()\n        edge_indices.append((start, end))\n        edge_indices.append((end, start))  ## Since graph is undirected\n        for _ in range(2):  ## Add the same bond features for both directions\n            edge_attrs.append(bond_features[bond.GetIdx()])\n\n    edge_index = torch.tensor(edge_indices, dtype=torch.long).t().contiguous()\n    edge_attr = torch.tensor(edge_attrs, dtype=torch.float)\n    data = Data(x = features, edge_index = edge_index, edge_attr = edge_attr)#, z = z, pos = pos)\n    return pickle.dumps(data)\n\ndef process_and_replace_smiles_columns(df: dd.DataFrame) -> dd.DataFrame:\n    column_order = [\n        'buildingblock1_smiles', 'buildingblock2_smiles', 'buildingblock3_smiles',\n        'molecule_smiles', 'protein_name', 'binds'\n    ]\n    for column in ['buildingblock1_smiles', 'buildingblock2_smiles', 'buildingblock3_smiles', 'molecule_smiles']:\n        df[f'{column}_graph'] = df[column].apply(smiles_to_graph)\n        df = df.drop(columns=[column])\n        df = df.rename(columns={f'{column}_graph': column})\n    df = df.reindex(columns=column_order)\n    if df.columns.tolist() != column_order:\n        print(\"Uh oh! You are a silly goose!\")\n        raise ValueError(\"Column order is incorrect, expected:\", column_order, \"but got:\", df.columns.tolist())\n    \n    return df\n``` \nand then I save it with joblib dump. It's pretty fast, admittedly not as fast as it can be but it's tolerable for me",
          "replies": [
            {
              "id": 2808905,
              "postDate": "2024-05-12T12:52:18.827Z",
              "content": "<p>so far my best solution</p>\n<pre><code>://github.com/mxmlnkn/indexed_bzip2\n indexed_bzip2  ibz2\n load_compressed_ibz2_pickle(file):\n    with ibz2.open(file, parallelization=os.cpu_count())  f:\n        \n    return \n</code></pre>\n<p>now for 30m smiles:</p>\n<ul>\n<li>conversion 33 min</li>\n<li>save using save_compressed_bz2_pickle(), 50 min</li>\n<li>load using load_compressed_ibz2_pickle(),** 3 min !!!!**</li>\n</ul>\n<p>actually, ibz2 can also be used for parallel saving … but i haven't figure out the code yet</p>",
              "rawMarkdown": "so far my best solution\n\n```\nhttps://github.com/mxmlnkn/indexed_bzip2\nimport indexed_bzip2 as ibz2\ndef load_compressed_ibz2_pickle(file):\n\twith ibz2.open(file, parallelization=os.cpu_count()) as f:\n\t\tdata = cPickle.load(f)\n\treturn data\n\n```\n\nnow for 30m smiles:\n- conversion 33 min\n- save using save_compressed_bz2_pickle(), 50 min\n- load using load_compressed_ibz2_pickle(),** 3 min !!!!**\n\nactually, ibz2 can also be used for parallel saving ... but i haven't figure out the code yet",
              "votes": 4
            },
            {
              "id": 2825077,
              "postDate": "2024-05-20T06:31:09.403Z",
              "content": "<p>Can you share save code? Thanks so much!</p>",
              "rawMarkdown": "Can you share save code? Thanks so much!"
            }
          ]
        }
      ]
    },
    {
      "id": 2783616,
      "postDate": "2024-04-29T20:34:10.320Z",
      "content": "<p>useful functions:<br>\npickle and zip</p>\n<pre><code> _pickle   cPickle\n bz2\n save_compressed_pickle(file, \n    with bz2.(file , 'w')  f:\n        cPickle.dump(\n\n load_decompress_pickle(file):\n    \n    \n    return \n</code></pre>",
      "rawMarkdown": "useful functions:\npickle and zip\n\n```\nimport _pickle as  cPickle\nimport bz2\ndef save_compressed_pickle(file, data):\n\twith bz2.BZ2File(file , 'w') as f:\n\t\tcPickle.dump(data, f)\n\ndef load_decompress_pickle(file):\n\tdata = bz2.BZ2File(file, 'rb')\n\tdata = cPickle.load(data)\n\treturn data\n\n\n```",
      "votes": 3,
      "replies": [
        {
          "id": 2784571,
          "postDate": "2024-04-30T11:07:30.423Z",
          "content": "<p>find an interesing paper<br>\nPARAMETER-FREE MOLECULAR CLASSIFICATION AND REGRESSION WITH GZIP<br>\n<a href=\"https://github.com/daenuprobst/molzip/tree/main\" target=\"_blank\">https://github.com/daenuprobst/molzip/tree/main</a></p>\n<p>trained on compressed text … save memory and improve results?</p>",
          "rawMarkdown": "find an interesing paper\nPARAMETER-FREE MOLECULAR CLASSIFICATION AND REGRESSION WITH GZIP\nhttps://github.com/daenuprobst/molzip/tree/main\n\ntrained on compressed text ... save memory and improve results?\n"
        }
      ]
    },
    {
      "id": 2808148,
      "postDate": "2024-05-12T05:03:38.430Z",
      "content": "<p>Another issue:</p>\n<pre><code>def smile_to_pygraph(smiles):\n        N, edge, node_feature, edge_feature =smile_to_graph(smiles)\n        graph = Data(\n            #=i,\n            =torch.from_numpy(edge.T).int(),\n            =torch.from_numpy(node_feature).byte(),\n            =torch.from_numpy(edge_feature).byte(),\n        )\n        return graph\n\n\n\n\n\n    with Pool(=64) as pool:\n        train_graph = list(tqdm(pool.imap(smile_to_graph, train_smiles), =num_train))\n\n\n    with Pool(=64) as pool:\n        train_graph = list(tqdm(pool.imap(smile_to_pygraph, train_smiles), =num_train))\n\nERROR:\n  File , line 164,  recvfds\n    raise RuntimeError( %\nRuntimeError: received 0 items of ancdata\n</code></pre>\n<p>Any experienced kagglers know why?<br>\nIt seems that either pytorch or pyg Data object mess up with multiprocessing resources.</p>",
      "rawMarkdown": "Another issue:\n\n```\ndef smile_to_pygraph(smiles):\n\t\tN, edge, node_feature, edge_feature =smile_to_graph(smiles)\n\t\tgraph = Data(\n\t\t\t#idx=i,\n\t\t\tedge_index=torch.from_numpy(edge.T).int(),\n\t\t\tx=torch.from_numpy(node_feature).byte(),\n\t\t\tedge_attr=torch.from_numpy(edge_feature).byte(),\n\t\t)\n\t\treturn graph\n\n######################################3\n....\n\n#this works:\n\twith Pool(processes=64) as pool:\n\t\ttrain_graph = list(tqdm(pool.imap(smile_to_graph, train_smiles), total=num_train))\n\n#but this doesn't (out of resources?) ???\n\twith Pool(processes=64) as pool:\n\t\ttrain_graph = list(tqdm(pool.imap(smile_to_pygraph, train_smiles), total=num_train))\n\nERROR:\n  File \"/home/user/app/anaconda3.10/lib/python3.10/multiprocessing/reduction.py\", line 164, in recvfds\n    raise RuntimeError('received %d items of ancdata' %\nRuntimeError: received 0 items of ancdata\n\n\n```\n\nAny experienced kagglers know why?\nIt seems that either pytorch or pyg Data object mess up with multiprocessing resources.",
      "votes": 1,
      "replies": [
        {
          "id": 2809747,
          "postDate": "2024-05-13T01:14:17.573Z",
          "content": "<p>hello, I found a fix from <a href=\"https://stackoverflow.com/a/76327337\" target=\"_blank\">stackoverflow</a> and it works</p>\n<pre><code>torch()\nwith (processes=) as pool:\n        train_graph = ((pool(smile_to_pygraph, train_smiles), total=num_train))\n</code></pre>",
          "rawMarkdown": "hello, I found a fix from [stackoverflow](https://stackoverflow.com/a/76327337) and it works\n```\ntorch.multiprocessing.set_sharing_strategy('file_system')\nwith Pool(processes=64) as pool:\n        train_graph = list(tqdm(pool.imap(smile_to_pygraph, train_smiles), total=num_train))\n```",
          "replies": [
            {
              "id": 2809775,
              "postDate": "2024-05-13T01:43:45.737Z",
              "content": "<p>unfortunately, this doesn't work. It cannot scale beyond 1m for me. (maybe i have a bug). i still get some memory allocation from torch.</p>\n<pre><code>    storage = torch.UntypedStorage._shared_filename_cpu(manager, handle, size)\nRuntimeError: unable  mmap     &lt;/torch_4878_92960782_980&gt;: Cannot allocate memory ()\n</code></pre>\n<p>what is your size of num_train?</p>",
              "rawMarkdown": "unfortunately, this doesn't work. It cannot scale beyond 1m for me. (maybe i have a bug). i still get some memory allocation from torch.\n```\n\n    storage = torch.UntypedStorage._new_shared_filename_cpu(manager, handle, size)\nRuntimeError: unable to mmap 148 bytes from file </torch_4878_92960782_980>: Cannot allocate memory (12)\n```\nwhat is your size of num\\_train?"
            },
            {
              "id": 2809789,
              "postDate": "2024-05-13T01:58:30.543Z",
              "content": "<p>I just tried 10k.</p>",
              "rawMarkdown": "I just tried 10k."
            },
            {
              "id": 2809791,
              "postDate": "2024-05-13T02:02:47.417Z",
              "content": "<p>thanks, i will check my code again</p>",
              "rawMarkdown": "thanks, i will check my code again"
            }
          ]
        }
      ]
    },
    {
      "id": 2810984,
      "postDate": "2024-05-13T14:50:53.103Z",
      "content": "<p>not sure if different GNN will affect results, but if you want to try, you can google for github repo<br>\n<a href=\"https://github.com/waqarahmadm019/AquaPred\" target=\"_blank\">https://github.com/waqarahmadm019/AquaPred</a><br>\n(GIN, GAT,GCN, attentionFP)</p>",
      "rawMarkdown": "not sure if different GNN will affect results, but if you want to try, you can google for github repo\nhttps://github.com/waqarahmadm019/AquaPred\n(GIN, GAT,GCN, attentionFP)",
      "votes": 2
    },
    {
      "id": 2806762,
      "postDate": "2024-05-11T09:03:06.220Z",
      "content": "<p>compare GNN and conv1d results (early experiments)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0b179520cf897935ff1058c07da9cffa%2FSelection_096.png?generation=1715417931640552&amp;alt=media\"></p>",
      "rawMarkdown": "compare GNN and conv1d results (early experiments)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0b179520cf897935ff1058c07da9cffa%2FSelection_096.png?generation=1715417931640552&alt=media)\n",
      "votes": 2,
      "replies": [
        {
          "id": 2806955,
          "postDate": "2024-05-11T12:32:37.030Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2859423,
          "postDate": "2024-06-07T03:09:53.973Z",
          "content": "<p>for the 3m GNN model, do you only use 70+ node feature and 6 edge features? Thanks</p>",
          "rawMarkdown": "for the 3m GNN model, do you only use 70+ node feature and 6 edge features? Thanks",
          "replies": [
            {
              "id": 2859582,
              "postDate": "2024-06-07T05:48:31.603Z",
              "content": "<p>yes. it is the same as the given notebook link</p>",
              "rawMarkdown": "yes. it is the same as the given notebook link",
              "votes": 1
            },
            {
              "id": 2860786,
              "postDate": "2024-06-07T19:04:22.187Z",
              "content": "<p>Thanks a lot! I am working on the 3M GNN model and trying to replicate your results but got higher loss (average 0.3). Are you using BCE loss like the node book?</p>",
              "rawMarkdown": "Thanks a lot! I am working on the 3M GNN model and trying to replicate your results but got higher loss (average 0.3). Are you using BCE loss like the node book?"
            },
            {
              "id": 2860819,
              "postDate": "2024-06-07T19:31:00.040Z",
              "content": "<p>only bce. <br>\ntry a conv1d model first. that set be the baseline. (use this to check data, etc)<br>\nother models like CNN, transformer, etc should perform similarly</p>",
              "rawMarkdown": "only bce. \ntry a conv1d model first. that set be the baseline. (use this to check data, etc)\nother models like CNN, transformer, etc should perform similarly"
            },
            {
              "id": 2860844,
              "postDate": "2024-06-07T20:08:08.380Z",
              "content": "<p>I am using 1.5 m samples with bind and 1.5 m samples with no bind. And my valid set is 100000 of these 3m samples. </p>",
              "rawMarkdown": "I am using 1.5 m samples with bind and 1.5 m samples with no bind. And my valid set is 100000 of these 3m samples. "
            },
            {
              "id": 2860926,
              "postDate": "2024-06-07T21:49:36.057Z",
              "content": "<p>Finally I am able to replicate your log loss! My 3m data is consisted of 1.5 m positive samples and 1.5 m negative samples. Now I juse randomly select 3m samples and it works</p>",
              "rawMarkdown": "Finally I am able to replicate your log loss! My 3m data is consisted of 1.5 m positive samples and 1.5 m negative samples. Now I juse randomly select 3m samples and it works"
            },
            {
              "id": 2861038,
              "postDate": "2024-06-08T01:12:19.957Z",
              "content": "<p>the best sampling strategy (random or fixed ration sampling) has to be experimentally determined and verified by submission</p>",
              "rawMarkdown": "the best sampling strategy (random or fixed ration sampling) has to be experimentally determined and verified by submission"
            },
            {
              "id": 2861062,
              "postDate": "2024-06-08T02:34:27.873Z",
              "content": "<p>also, loss itself may not be important. It is the ranking of the samples that is important</p>",
              "rawMarkdown": "also, loss itself may not be important. It is the ranking of the samples that is important"
            }
          ]
        },
        {
          "id": 2865909,
          "postDate": "2024-06-11T04:23:59.740Z",
          "content": "<p>I can achieve the same micro and BRD4, HSA, sEH AP after the first 1000 iteration only if I use <strong>validation data as training data</strong> and my AP is very closed to your numbers (0.32,0.17 0.68 for 3 proteins)… Not sure what happen here and very frustrated :(</p>\n<p>For the early experiment, do you use 3m training data set and 5m validation data set?</p>",
          "rawMarkdown": "I can achieve the same micro and BRD4, HSA, sEH AP after the first 1000 iteration only if I use **validation data as training data** and my AP is very closed to your numbers (0.32,0.17 0.68 for 3 proteins)... Not sure what happen here and very frustrated :(\n\nFor the early experiment, do you use 3m training data set and 5m validation data set?",
          "replies": [
            {
              "id": 2866074,
              "postDate": "2024-06-11T06:11:50.633Z",
              "content": "<p>use conv1d model (or xgboost+ecfp) as a baseline to test your framework and data.<br>\nif you use validation as training set you are measuring train error in which all your protein bind prediction should be above 0.80 (e.g. if you set your lr to very low e.g. 1e-5, you should be able to overfit your train data and very high ap) </p>\n<p>you can only cross-check your code, results, etc using different models and dataset.<br>\ninstead of debugging on one model and data, try different model and data (and different code). starts with the simplest and least time to train.</p>",
              "rawMarkdown": "use conv1d model (or xgboost+ecfp) as a baseline to test your framework and data.\nif you use validation as training set you are measuring train error in which all your protein bind prediction should be above 0.80 (e.g. if you set your lr to very low e.g. 1e-5, you should be able to overfit your train data and very high ap) \n\nyou can only cross-check your code, results, etc using different models and dataset.\ninstead of debugging on one model and data, try different model and data (and different code). starts with the simplest and least time to train.",
              "votes": 1
            },
            {
              "id": 2867474,
              "postDate": "2024-06-11T21:03:22.623Z",
              "content": "<p>Thank you so so much! I am able to achieve same loss and average precision now! My problem is I did not realize train_reduce.parquet is ordered! Due to memory limit, I evenly split it into 10 pkl file but this will cost problem to the train_idx as similar samples will be group together. Now I am good to go. Thank you!</p>",
              "rawMarkdown": "Thank you so so much! I am able to achieve same loss and average precision now! My problem is I did not realize train_reduce.parquet is ordered! Due to memory limit, I evenly split it into 10 pkl file but this will cost problem to the train_idx as similar samples will be group together. Now I am good to go. Thank you!"
            }
          ]
        }
      ]
    },
    {
      "id": 2869040,
      "postDate": "2024-06-12T19:40:01.840Z",
      "content": "<p>Update: I add ecfp to the GNN model and it reach 0.6 ap in the first 1000 iteration and 0.7 after the first 5 epoch</p>",
      "rawMarkdown": "Update: I add ecfp to the GNN model and it reach 0.6 ap in the first 1000 iteration and 0.7 after the first 5 epoch"
    },
    {
      "id": 2859294,
      "postDate": "2024-06-07T00:03:22.200Z",
      "content": "<p>train = 50/70m samples, valid=5m (random split)<br>\ndoes it mean you use 50m for training and 70m for fine tuning the model hyperparameters (layers, embedding ..etc)? <br>\nThanks!</p>",
      "rawMarkdown": "train = 50/70m samples, valid=5m (random split)\ndoes it mean you use 50m for training and 70m for fine tuning the model hyperparameters (layers, embedding ..etc)? \nThanks!",
      "replies": [
        {
          "id": 2859380,
          "postDate": "2024-06-07T02:27:21.143Z",
          "content": "<p>yes. but these are for speeding up my early experiments.</p>\n<p>in submission, using all samples for all is the best.</p>",
          "rawMarkdown": "yes. but these are for speeding up my early experiments.\n\nin submission, using all samples for all is the best."
        }
      ]
    },
    {
      "id": 2833290,
      "postDate": "2024-05-24T06:02:35.220Z",
      "content": "<p>graph embedding space is well known for non-smooth.<br>\nthen i have an idea</p>\n<pre><code>instead of:\n model() probability\n\nwe can:\n model()  nearest  that binds\n\nthan score  distance(input  predicted nearest binding )\n</code></pre>",
      "rawMarkdown": "graph embedding space is well known for non-smooth.\nthen i have an idea\n\n```\ninstead of:\n model(x)= probability\n\nwe can:\n model(x) = nearest x that binds\n\nthan score = distance(input x, predicted nearest binding x)\n\n```",
      "replies": [
        {
          "id": 2834136,
          "postDate": "2024-05-24T15:09:03.423Z",
          "content": "<p>Nearest in what sense? How to generate labels?</p>\n<p>Is the label a KNN style distance from features_i to features_j, where j is the closest in feature space that binds?</p>",
          "rawMarkdown": "Nearest in what sense? How to generate labels?\n\nIs the label a KNN style distance from features_i to features_j, where j is the closest in feature space that binds?",
          "replies": [
            {
              "id": 2834284,
              "postDate": "2024-05-24T16:45:01.330Z",
              "content": "<p>an example is actually given here: <a href=\"https://www.kaggle.com/code/taroiwai/01-leashbio-eda-p-values\" target=\"_blank\">https://www.kaggle.com/code/taroiwai/01-leashbio-eda-p-values</a><br>\nKNN is another possibility</p>",
              "rawMarkdown": "an example is actually given here: https://www.kaggle.com/code/taroiwai/01-leashbio-eda-p-values\nKNN is another possibility",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2809614,
      "postDate": "2024-05-12T21:12:26.423Z",
      "content": "<p>This is awesome, thanks for sharing! I'll need to study this to see what I missed in my implementation. I think we're only dealing with a subset of elements ['B', 'Br', 'C', 'Cl', 'Dy', 'F', 'I', 'N', 'O', 'S', 'Si'] in our dataset, so you could reduce the node dim substantially and avoid a bunch of all-zeros dimensions from element 1-hot encoding.</p>",
      "rawMarkdown": "This is awesome, thanks for sharing! I'll need to study this to see what I missed in my implementation. I think we're only dealing with a subset of elements ['B', 'Br', 'C', 'Cl', 'Dy', 'F', 'I', 'N', 'O', 'S', 'Si'] in our dataset, so you could reduce the node dim substantially and avoid a bunch of all-zeros dimensions from element 1-hot encoding.",
      "replies": [
        {
          "id": 2811840,
          "postDate": "2024-05-14T00:08:34.680Z",
          "content": "<p>i forget to add my split:</p>\n<pre><code> ():\n\n    \n           = \n    all_valid = \n    all_train = \n    rng = np.random.RandomState()\n    index = np.arange()\n    rng.shuffle(index)\n    train_index = index[:all_train]\n    valid_index = index[all_train:]\n    \n    \n\n    \n     num_train   :\n        train_index=train_index[:num_train]\n     num_valid   :\n        valid_index=valid_index[:num_valid]\n\n    \n    \n    (, (train_index), train_index[[,,-]]) \n    (, (valid_index), valid_index[[,,-]]) \n     train_index, valid_index\n</code></pre>\n<p>i release model code, input graph and split. you should be able to repeat the experiment. if your implementation is correct you should get about the same log (loss curve) within the first 10m to 20m train samples seen.</p>",
          "rawMarkdown": "i forget to add my split:\n```\ndef make_leashbio_fold(num_train=None,num_valid=None):\n\n\t#make 5% train-valid split\n\tall       = 98_415_610\n\tall_valid = 5_000_000\n\tall_train = 93_415_610\n\trng = np.random.RandomState(123)\n\tindex = np.arange(all)\n\trng.shuffle(index)\n\ttrain_index = index[:all_train]\n\tvalid_index = index[all_train:]\n\t#train_index = np.sort(train_index)\n\t#valid_index = np.sort(valid_index)\n\n\t#subsample according to input arguments\n\tif num_train is not None:\n\t\ttrain_index=train_index[:num_train]\n\tif num_valid is not None:\n\t\tvalid_index=valid_index[:num_valid]\n\n\t#check no overlap\n\t#print('make_fold() overlap:',set(train_index).intersection(set(valid_index))) #set()\n\tprint('train_index', len(train_index), train_index[[0,1,-1]]) # print some index for debug\n\tprint('valid_index', len(valid_index), valid_index[[0,1,-1]]) #\n\treturn train_index, valid_index\n```\n\ni release model code, input graph and split. you should be able to repeat the experiment. if your implementation is correct you should get about the same log (loss curve) within the first 10m to 20m train samples seen.",
          "votes": 2,
          "replies": [
            {
              "id": 2862901,
              "postDate": "2024-06-09T05:38:24.513Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2864191,
              "postDate": "2024-06-10T02:23:45.033Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2864338,
              "postDate": "2024-06-10T05:30:12.547Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2830891,
          "postDate": "2024-05-23T12:16:09.307Z",
          "content": "<p>Thanks, here is a simple code to check the unique elements:</p>\n<pre><code>all_atoms = ()\n mol  tqdm(test[]):\n    atoms = unique_atoms(mol)\n    all_atoms = all_atoms.union(atoms)\n</code></pre>\n<p>and indeed they are ['B', 'Br', 'C', 'Cl', 'Dy', 'F', 'I', 'N', 'O', 'S', 'Si'] <br>\nWe can also remove the unnecessary bits after F_unpackbits (e.g., instead of using all 8x9 = 72 bits we can use the 69 bit that hold the node features). </p>",
          "rawMarkdown": "Thanks, here is a simple code to check the unique elements:\n```python\nall_atoms = set()\nfor mol in tqdm(test['molecule_smiles']):\n    atoms = unique_atoms(mol)\n    all_atoms = all_atoms.union(atoms)\n``` and indeed they are ['B', 'Br', 'C', 'Cl', 'Dy', 'F', 'I', 'N', 'O', 'S', 'Si'] \nWe can also remove the unnecessary bits after F_unpackbits (e.g., instead of using all 8x9 = 72 bits we can use the 69 bit that hold the node features). ",
          "votes": 1,
          "replies": [
            {
              "id": 2860848,
              "postDate": "2024-06-07T20:11:21.630Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2808421,
      "postDate": "2024-05-12T07:43:32.933Z",
      "content": "<p>I have just a question (if you're willing to share of course), how have you handled the 3d position extraction? I am asking because RDKit takes a laughable amount of time for that</p>",
      "rawMarkdown": "I have just a question (if you're willing to share of course), how have you handled the 3d position extraction? I am asking because RDKit takes a laughable amount of time for that",
      "replies": [
        {
          "id": 2809778,
          "postDate": "2024-05-13T01:48:28.820Z",
          "content": "<p>you can google for fast conformer generator (e.g. with deep learning),.<br>\none example is:<br>\n<a href=\"https://arxiv.org/pdf/2106.07802\" target=\"_blank\">https://arxiv.org/pdf/2106.07802</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F961a74e9262cb84949e5c44e0a7b9b07%2FSelection_110.png?generation=1715564906183916&amp;alt=media\"></p>",
          "rawMarkdown": "you can google for fast conformer generator (e.g. with deep learning),.\none example is:\nhttps://arxiv.org/pdf/2106.07802\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F961a74e9262cb84949e5c44e0a7b9b07%2FSelection_110.png?generation=1715564906183916&alt=media)",
          "votes": 2,
          "replies": [
            {
              "id": 2810146,
              "postDate": "2024-05-13T06:31:36.627Z",
              "content": "<p>Thanks, I will look into it!</p>",
              "rawMarkdown": "Thanks, I will look into it!"
            }
          ]
        }
      ]
    },
    {
      "id": 2879221,
      "postDate": "2024-06-19T12:41:21.833Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2821861,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-18T09:45:54.390000",
      "content": "<p>i end up wring my own collate function. Hence we need not convert to pyg graph list at all, saving 90 min conversion time and 70gb. removing pyg data loader improve speed by 2.5x.</p>\n<p>Now I can train with full 98m SMILES and my graph-NN is same speed as CNN1d.</p>\n<pre><code>def (graph, index=None, device='cpu'):\n    if index is None:\n        index = np.((graph)).()\n    batch = (\n        x=[],\n        edge_index=[],\n        edge_attr=[],\n        batch=[],\n        idx=index\n    )\n    offset = \n    for b, i in (index):\n        N, edge, node_feature, edge_feature = graph[i]\n        batch.x.(node_feature)\n        batch.edge_attr.(edge_feature)\n        batch.edge_index.(edge.(int) + offset)\n        batch.batch += N * [b]\n        offset += N\n    batch.x = torch.(np.(batch.x)).(device)\n    batch.edge_attr = torch.(np.(batch.edge_attr)).(device)\n    batch.edge_index = torch.(np.(batch.edge_index).T).(device)\n    batch.batch = torch.(batch.batch).(device)\n    return batch\n\n\n.... more code here ....\n\nwhile epoch&lt;cfg.num_epoch:\n    shuffled_idx = train_idx.()\n    np.random.(shuffled_idx)\n    for t, index in (np.(,(shuffled_idx),cfg.train_batch_size)):\n        index = shuffled_idx[index:index+cfg.train_batch_size]\n        if (index)!=cfg.train_batch_size: continue #drop last\n\n        B = (index)\n        batch = (\n            graph = (train_graph,index,device=),\n            bind = torch.(train_bind[index]).().(),\n        )\n\n        net.()\n        net.output_type = [, ]\n\n        with torch.cuda.amp.(enabled=cfg.is_amp):\n            output = (batch)  #(net,batch) #\n            bce_loss = output[]\n</code></pre>\n<p>gpu utilization is now 70% and a queue can fix that.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2821876,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-05-18T09:54:06.130000",
          "content": "<p>queue for 100% gpu utilization</p>\n<pre><code>def (queue):\n    shuffled_idx = train_idx.()\n    np.random.(shuffled_idx)\n    for t, index in (np.(, (shuffled_idx), cfg.train_batch_size)):\n        index = shuffled_idx[index:index + cfg.train_batch_size]\n        if (index)!=cfg.train_batch_size: continue #drop last\n\n        B = (index)\n        batch = (\n            graph=(train_graph, index, device=),\n            bind=torch.(train_bind[index]).(),\n        )\n        queue.(batch)\n    queue.(None)\n\n\n\nwhile epoch&lt;cfg.num_epoch:\n    queue = multiprocessing.().(maxsize=)\n    train_loader_process = (target=train_loader_func, args=(queue,))\n    train_loader_process.()\n\n    num_train_batch = ((train_idx) // cfg.train_batch_size)\n    for t in (num_train_batch+):\n        batch = queue.()\n        if batch is None:\n            break\n        #cpu to dpu outside () ....\n        batch.bind = batch.bind.()\n        batch.graph.x = batch.graph.x.()\n        batch.graph.edge_attr = batch.graph.edge_attr.()\n        batch.graph.edge_index = batch.graph.edge_index.()\n        batch.graph.batch = batch.graph.batch.()\n</code></pre>",
          "votes": 2,
          "replies": [
            {
              "id": 2857690,
              "author_name": "JoO",
              "author_url": "",
              "post_date": "2024-06-06T05:02:18.200000",
              "content": "<p>Thank you for your code! my code got stuck at \"batch = queue.get()\" when the batch size is larger than 10. What batch_size do you use? Thanks!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2857717,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-06-06T05:24:48.060000",
              "content": "<p>my batch is is 2000 to 5000.<br>\nrun your code without queue first. it should get utilization more than 80%. else my simple queue don't help</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2857729,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-06-06T05:29:25.613000",
              "content": "<p>there is \"train_loader_process.join()\" at the end of the code with i didn't show.<br>\n(read more about python multiprocessing Process class)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2809601,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-12T20:57:03.150000",
      "content": "<p>start download now! converted smiles string to graph object:<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4024364f683c7fa253bccb5e2a5530c1%2FSelection_106.png?generation=1715547404927195&amp;alt=media\"></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2808134,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-12T04:57:34.180000",
      "content": "<p>Another issue 2:<br>\nany kaggle has good method to save converted graph to file on disk?<br>\nif there is good solution, i will upload my converted graph to public dataset.</p>\n<hr>\n<p>i need a solution that include file compression. it also needs to be fast to compress and decompress the list of graph objects (e.g. parallel processing). I tried cPickle alone but it is not good enough.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2808425,
          "author_name": "Giorgio Micaletto",
          "author_url": "",
          "post_date": "2024-05-12T07:46:07.397000",
          "content": "<p>Possibly useless, but this is how I handle it</p>\n<pre><code> () -&gt; [Data, ]:\n    mol = Chem.MolFromSmiles(smiles)\n    me = Chem.MolFromSmiles()\n    dy = Chem.MolFromSmiles()\n    mol = AllChem.ReplaceSubstructs(mol, dy, me)[]\n      mol:\n          \n    mol = Chem.AddHs(mol)\n    atom_features = get_atom_features(mol)\n    bond_features = get_bond_features(mol)\n    \n    \n    \n    \n    features = torch.tensor(atom_features, dtype=torch.)\n\n    \n    edge_indices = []\n    edge_attrs = []\n     bond  mol.GetBonds():\n        start, end = bond.GetBeginAtomIdx(), bond.GetEndAtomIdx()\n        edge_indices.append((start, end))\n        edge_indices.append((end, start))  \n         _  ():  \n            edge_attrs.append(bond_features[bond.GetIdx()])\n\n    edge_index = torch.tensor(edge_indices, dtype=torch.long).t().contiguous()\n    edge_attr = torch.tensor(edge_attrs, dtype=torch.)\n    data = Data(x = features, edge_index = edge_index, edge_attr = edge_attr)\n     pickle.dumps(data)\n\n () -&gt; dd.DataFrame:\n    column_order = [\n        , , ,\n        , , \n    ]\n     column  [, , , ]:\n        df[] = df[column].apply(smiles_to_graph)\n        df = df.drop(columns=[column])\n        df = df.rename(columns={: column})\n    df = df.reindex(columns=column_order)\n     df.columns.tolist() != column_order:\n        ()\n         ValueError(, column_order, , df.columns.tolist())\n\n     df\n</code></pre>\n<p><br>\nand then I save it with joblib dump. It's pretty fast, admittedly not as fast as it can be but it's tolerable for me</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2808905,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-12T12:52:18.827000",
              "content": "<p>so far my best solution</p>\n<pre><code>://github.com/mxmlnkn/indexed_bzip2\n indexed_bzip2  ibz2\n load_compressed_ibz2_pickle(file):\n    with ibz2.open(file, parallelization=os.cpu_count())  f:\n        \n    return \n</code></pre>\n<p>now for 30m smiles:</p>\n<ul>\n<li>conversion 33 min</li>\n<li>save using save_compressed_bz2_pickle(), 50 min</li>\n<li>load using load_compressed_ibz2_pickle(),** 3 min !!!!**</li>\n</ul>\n<p>actually, ibz2 can also be used for parallel saving … but i haven't figure out the code yet</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2825077,
              "author_name": "go",
              "author_url": "",
              "post_date": "2024-05-20T06:31:09.403000",
              "content": "<p>Can you share save code? Thanks so much!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2783616,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-29T20:34:10.320000",
      "content": "<p>useful functions:<br>\npickle and zip</p>\n<pre><code> _pickle   cPickle\n bz2\n save_compressed_pickle(file, \n    with bz2.(file , 'w')  f:\n        cPickle.dump(\n\n load_decompress_pickle(file):\n    \n    \n    return \n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 2784571,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-04-30T11:07:30.423000",
          "content": "<p>find an interesing paper<br>\nPARAMETER-FREE MOLECULAR CLASSIFICATION AND REGRESSION WITH GZIP<br>\n<a href=\"https://github.com/daenuprobst/molzip/tree/main\" target=\"_blank\">https://github.com/daenuprobst/molzip/tree/main</a></p>\n<p>trained on compressed text … save memory and improve results?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2808148,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-12T05:03:38.430000",
      "content": "<p>Another issue:</p>\n<pre><code>def smile_to_pygraph(smiles):\n        N, edge, node_feature, edge_feature =smile_to_graph(smiles)\n        graph = Data(\n            #=i,\n            =torch.from_numpy(edge.T).int(),\n            =torch.from_numpy(node_feature).byte(),\n            =torch.from_numpy(edge_feature).byte(),\n        )\n        return graph\n\n\n\n\n\n    with Pool(=64) as pool:\n        train_graph = list(tqdm(pool.imap(smile_to_graph, train_smiles), =num_train))\n\n\n    with Pool(=64) as pool:\n        train_graph = list(tqdm(pool.imap(smile_to_pygraph, train_smiles), =num_train))\n\nERROR:\n  File , line 164,  recvfds\n    raise RuntimeError( %\nRuntimeError: received 0 items of ancdata\n</code></pre>\n<p>Any experienced kagglers know why?<br>\nIt seems that either pytorch or pyg Data object mess up with multiprocessing resources.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2809747,
          "author_name": "xc",
          "author_url": "",
          "post_date": "2024-05-13T01:14:17.573000",
          "content": "<p>hello, I found a fix from <a href=\"https://stackoverflow.com/a/76327337\" target=\"_blank\">stackoverflow</a> and it works</p>\n<pre><code>torch()\nwith (processes=) as pool:\n        train_graph = ((pool(smile_to_pygraph, train_smiles), total=num_train))\n</code></pre>",
          "votes": 0,
          "replies": [
            {
              "id": 2809775,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-13T01:43:45.737000",
              "content": "<p>unfortunately, this doesn't work. It cannot scale beyond 1m for me. (maybe i have a bug). i still get some memory allocation from torch.</p>\n<pre><code>    storage = torch.UntypedStorage._shared_filename_cpu(manager, handle, size)\nRuntimeError: unable  mmap     &lt;/torch_4878_92960782_980&gt;: Cannot allocate memory ()\n</code></pre>\n<p>what is your size of num_train?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2809789,
              "author_name": "xc",
              "author_url": "",
              "post_date": "2024-05-13T01:58:30.543000",
              "content": "<p>I just tried 10k.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2809791,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-13T02:02:47.417000",
              "content": "<p>thanks, i will check my code again</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2810984,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-13T14:50:53.103000",
      "content": "<p>not sure if different GNN will affect results, but if you want to try, you can google for github repo<br>\n<a href=\"https://github.com/waqarahmadm019/AquaPred\" target=\"_blank\">https://github.com/waqarahmadm019/AquaPred</a><br>\n(GIN, GAT,GCN, attentionFP)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2806762,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-11T09:03:06.220000",
      "content": "<p>compare GNN and conv1d results (early experiments)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0b179520cf897935ff1058c07da9cffa%2FSelection_096.png?generation=1715417931640552&amp;alt=media\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2806955,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-05-11T12:32:37.030000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2859423,
          "author_name": "JoO",
          "author_url": "",
          "post_date": "2024-06-07T03:09:53.973000",
          "content": "<p>for the 3m GNN model, do you only use 70+ node feature and 6 edge features? Thanks</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2859582,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-06-07T05:48:31.603000",
              "content": "<p>yes. it is the same as the given notebook link</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2860786,
              "author_name": "JoO",
              "author_url": "",
              "post_date": "2024-06-07T19:04:22.187000",
              "content": "<p>Thanks a lot! I am working on the 3M GNN model and trying to replicate your results but got higher loss (average 0.3). Are you using BCE loss like the node book?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2860819,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-06-07T19:31:00.040000",
              "content": "<p>only bce. <br>\ntry a conv1d model first. that set be the baseline. (use this to check data, etc)<br>\nother models like CNN, transformer, etc should perform similarly</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2860844,
              "author_name": "JoO",
              "author_url": "",
              "post_date": "2024-06-07T20:08:08.380000",
              "content": "<p>I am using 1.5 m samples with bind and 1.5 m samples with no bind. And my valid set is 100000 of these 3m samples. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2860926,
              "author_name": "JoO",
              "author_url": "",
              "post_date": "2024-06-07T21:49:36.057000",
              "content": "<p>Finally I am able to replicate your log loss! My 3m data is consisted of 1.5 m positive samples and 1.5 m negative samples. Now I juse randomly select 3m samples and it works</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2861038,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-06-08T01:12:19.957000",
              "content": "<p>the best sampling strategy (random or fixed ration sampling) has to be experimentally determined and verified by submission</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2861062,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-06-08T02:34:27.873000",
              "content": "<p>also, loss itself may not be important. It is the ranking of the samples that is important</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2865909,
          "author_name": "JoO",
          "author_url": "",
          "post_date": "2024-06-11T04:23:59.740000",
          "content": "<p>I can achieve the same micro and BRD4, HSA, sEH AP after the first 1000 iteration only if I use <strong>validation data as training data</strong> and my AP is very closed to your numbers (0.32,0.17 0.68 for 3 proteins)… Not sure what happen here and very frustrated :(</p>\n<p>For the early experiment, do you use 3m training data set and 5m validation data set?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2866074,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-06-11T06:11:50.633000",
              "content": "<p>use conv1d model (or xgboost+ecfp) as a baseline to test your framework and data.<br>\nif you use validation as training set you are measuring train error in which all your protein bind prediction should be above 0.80 (e.g. if you set your lr to very low e.g. 1e-5, you should be able to overfit your train data and very high ap) </p>\n<p>you can only cross-check your code, results, etc using different models and dataset.<br>\ninstead of debugging on one model and data, try different model and data (and different code). starts with the simplest and least time to train.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2867474,
              "author_name": "JoO",
              "author_url": "",
              "post_date": "2024-06-11T21:03:22.623000",
              "content": "<p>Thank you so so much! I am able to achieve same loss and average precision now! My problem is I did not realize train_reduce.parquet is ordered! Due to memory limit, I evenly split it into 10 pkl file but this will cost problem to the train_idx as similar samples will be group together. Now I am good to go. Thank you!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2869040,
      "author_name": "JoO",
      "author_url": "",
      "post_date": "2024-06-12T19:40:01.840000",
      "content": "<p>Update: I add ecfp to the GNN model and it reach 0.6 ap in the first 1000 iteration and 0.7 after the first 5 epoch</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2859294,
      "author_name": "JoO",
      "author_url": "",
      "post_date": "2024-06-07T00:03:22.200000",
      "content": "<p>train = 50/70m samples, valid=5m (random split)<br>\ndoes it mean you use 50m for training and 70m for fine tuning the model hyperparameters (layers, embedding ..etc)? <br>\nThanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2859380,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-06-07T02:27:21.143000",
          "content": "<p>yes. but these are for speeding up my early experiments.</p>\n<p>in submission, using all samples for all is the best.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2833290,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-05-24T06:02:35.220000",
      "content": "<p>graph embedding space is well known for non-smooth.<br>\nthen i have an idea</p>\n<pre><code>instead of:\n model() probability\n\nwe can:\n model()  nearest  that binds\n\nthan score  distance(input  predicted nearest binding )\n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 2834136,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2024-05-24T15:09:03.423000",
          "content": "<p>Nearest in what sense? How to generate labels?</p>\n<p>Is the label a KNN style distance from features_i to features_j, where j is the closest in feature space that binds?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2834284,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-05-24T16:45:01.330000",
              "content": "<p>an example is actually given here: <a href=\"https://www.kaggle.com/code/taroiwai/01-leashbio-eda-p-values\" target=\"_blank\">https://www.kaggle.com/code/taroiwai/01-leashbio-eda-p-values</a><br>\nKNN is another possibility</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2809614,
      "author_name": "Darek Kłeczek",
      "author_url": "",
      "post_date": "2024-05-12T21:12:26.423000",
      "content": "<p>This is awesome, thanks for sharing! I'll need to study this to see what I missed in my implementation. I think we're only dealing with a subset of elements ['B', 'Br', 'C', 'Cl', 'Dy', 'F', 'I', 'N', 'O', 'S', 'Si'] in our dataset, so you could reduce the node dim substantially and avoid a bunch of all-zeros dimensions from element 1-hot encoding.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2811840,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-05-14T00:08:34.680000",
          "content": "<p>i forget to add my split:</p>\n<pre><code> ():\n\n    \n           = \n    all_valid = \n    all_train = \n    rng = np.random.RandomState()\n    index = np.arange()\n    rng.shuffle(index)\n    train_index = index[:all_train]\n    valid_index = index[all_train:]\n    \n    \n\n    \n     num_train   :\n        train_index=train_index[:num_train]\n     num_valid   :\n        valid_index=valid_index[:num_valid]\n\n    \n    \n    (, (train_index), train_index[[,,-]]) \n    (, (valid_index), valid_index[[,,-]]) \n     train_index, valid_index\n</code></pre>\n<p>i release model code, input graph and split. you should be able to repeat the experiment. if your implementation is correct you should get about the same log (loss curve) within the first 10m to 20m train samples seen.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2862901,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-09T05:38:24.513000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2864191,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-10T02:23:45.033000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2864338,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-10T05:30:12.547000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2830891,
          "author_name": "Alhasan Abdellatif",
          "author_url": "",
          "post_date": "2024-05-23T12:16:09.307000",
          "content": "<p>Thanks, here is a simple code to check the unique elements:</p>\n<pre><code>all_atoms = ()\n mol  tqdm(test[]):\n    atoms = unique_atoms(mol)\n    all_atoms = all_atoms.union(atoms)\n</code></pre>\n<p>and indeed they are ['B', 'Br', 'C', 'Cl', 'Dy', 'F', 'I', 'N', 'O', 'S', 'Si'] <br>\nWe can also remove the unnecessary bits after F_unpackbits (e.g., instead of using all 8x9 = 72 bits we can use the 69 bit that hold the node features). </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2860848,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-07T20:11:21.630000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2808421,
      "author_name": "Giorgio Micaletto",
      "author_url": "",
      "post_date": "2024-05-12T07:43:32.933000",
      "content": "<p>I have just a question (if you're willing to share of course), how have you handled the 3d position extraction? I am asking because RDKit takes a laughable amount of time for that</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2809778,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-05-13T01:48:28.820000",
          "content": "<p>you can google for fast conformer generator (e.g. with deep learning),.<br>\none example is:<br>\n<a href=\"https://arxiv.org/pdf/2106.07802\" target=\"_blank\">https://arxiv.org/pdf/2106.07802</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F961a74e9262cb84949e5c44e0a7b9b07%2FSelection_110.png?generation=1715564906183916&amp;alt=media\"></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2810146,
              "author_name": "Giorgio Micaletto",
              "author_url": "",
              "post_date": "2024-05-13T06:31:36.627000",
              "content": "<p>Thanks, I will look into it!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2879221,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-19T12:41:21.833000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2783615": "example code: \nhttps://www.kaggle.com/code/hengck23/lb6-02-graph-nn-example\n\nconverted smiles string to graph object:\nhttps://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F38c7a842993b6c5b5928bc426a088457%2FSelection_109.png?generation=1715558729915059&alt=media)\n\n#model :\n- pytorch geometric framework\n- simple MPNNModel, 4 layers, hidden dim=96\n- see: https://github.com/chaitjo/geometric-gnn-dojo/blob/main/geometric_gnn_101.ipynb\n- model file size: 7.7mb  \n\n#graph creation:\n- atom and bond one-hot features from rdkit\n- see :  from https://github.com/LiZhang30/GPCNDTA/blob/main/utils/DrugGraph.py\n- node dim = 70, edge dim=6\n- smiles to graph conversion: 1m takes 1 min (multiprocessing pool with 64 workers)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8b2f652ce3f46490ae63acb7335f9e6e%2FSelection_105.png?generation=1715489514039434&alt=media)\n\n#other tricks:\n- np.packbits to reduce ram for faster cpu-to-gpu transfer\n- avoid using num workers>0 in pgy DataLoader.\n  see: https://www.kaggle.com/competitions/leash-BELKA/discussion/500877\n \n\n#hyperprameters:\n- train = 50/70m samples, valid=5m (random split)\n- batch size = 5000\n- training speed: 1m samples takes 12 min, 1 gpu\n- lr 1e-3/1e-4\n\npyg batch collate function is not efficient enough to give 100% gpu utilization rate for large batch size of 5000.\neven worse, we cannot set num workers>0.\ni estimate that we can probably go up to 1m samples at 5 min with better code.\n\n---\n\nmyth buster\n- GNN is not correct and doesn't work ? not true. if we are memorizing sharing block predictions (i.e. train and test distribution are similiar), it needs not to be correct.\n- GNN is slow: not true we do not need many layers.\n- graph conversion is slow and takes much memory: not true",
    "2821861": "i end up wring my own collate function. Hence we need not convert to pyg graph list at all, saving 90 min conversion time and 70gb. removing pyg data loader improve speed by 2.5x.\n\nNow I can train with full 98m SMILES and my graph-NN is same speed as CNN1d.\n```\ndef my_collate(graph, index=None, device='cpu'):\n\tif index is None:\n\t\tindex = np.arange(len(graph)).tolist()\n\tbatch = dotdict(\n\t\tx=[],\n\t\tedge_index=[],\n\t\tedge_attr=[],\n\t\tbatch=[],\n\t\tidx=index\n\t)\n\toffset = 0\n\tfor b, i in enumerate(index):\n\t\tN, edge, node_feature, edge_feature = graph[i]\n\t\tbatch.x.append(node_feature)\n\t\tbatch.edge_attr.append(edge_feature)\n\t\tbatch.edge_index.append(edge.astype(int) + offset)\n\t\tbatch.batch += N * [b]\n\t\toffset += N\n\tbatch.x = torch.from_numpy(np.concatenate(batch.x)).to(device)\n\tbatch.edge_attr = torch.from_numpy(np.concatenate(batch.edge_attr)).to(device)\n\tbatch.edge_index = torch.from_numpy(np.concatenate(batch.edge_index).T).to(device)\n\tbatch.batch = torch.LongTensor(batch.batch).to(device)\n\treturn batch\n\n\n.... more code here ....\n\nwhile epoch<cfg.num_epoch:\n\tshuffled_idx = train_idx.copy()\n\tnp.random.shuffle(shuffled_idx)\n\tfor t, index in enumerate(np.arange(0,len(shuffled_idx),cfg.train_batch_size)):\n\t\tindex = shuffled_idx[index:index+cfg.train_batch_size]\n\t\tif len(index)!=cfg.train_batch_size: continue #drop last\n\n\t\tB = len(index)\n\t\tbatch = dotdict(\n\t\t\tgraph = my_collate(train_graph,index,device='cuda'),\n\t\t\tbind = torch.from_numpy(train_bind[index]).float().cuda(),\n\t\t)\n\n\t\tnet.train()\n\t\tnet.output_type = ['loss', 'infer']\n\n\t\twith torch.cuda.amp.autocast(enabled=cfg.is_amp):\n\t\t\toutput = net(batch)  #data_parallel(net,batch) #\n\t\t\tbce_loss = output['bce_loss']\n```\n\ngpu utilization is now 70% and a queue can fix that.",
    "2809601": "start download now! converted smiles string to graph object:\nhttps://www.kaggle.com/datasets/hengck23/leash-bio-processed-dataset\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4024364f683c7fa253bccb5e2a5530c1%2FSelection_106.png?generation=1715547404927195&alt=media)",
    "2808134": "Another issue 2:\nany kaggle has good method to save converted graph to file on disk?\nif there is good solution, i will upload my converted graph to public dataset.\n\n---\n\ni need a solution that include file compression. it also needs to be fast to compress and decompress the list of graph objects (e.g. parallel processing). I tried cPickle alone but it is not good enough.",
    "2783616": "useful functions:\npickle and zip\n\n```\nimport _pickle as  cPickle\nimport bz2\ndef save_compressed_pickle(file, data):\n\twith bz2.BZ2File(file , 'w') as f:\n\t\tcPickle.dump(data, f)\n\ndef load_decompress_pickle(file):\n\tdata = bz2.BZ2File(file, 'rb')\n\tdata = cPickle.load(data)\n\treturn data\n\n\n```",
    "2808148": "Another issue:\n\n```\ndef smile_to_pygraph(smiles):\n\t\tN, edge, node_feature, edge_feature =smile_to_graph(smiles)\n\t\tgraph = Data(\n\t\t\t#idx=i,\n\t\t\tedge_index=torch.from_numpy(edge.T).int(),\n\t\t\tx=torch.from_numpy(node_feature).byte(),\n\t\t\tedge_attr=torch.from_numpy(edge_feature).byte(),\n\t\t)\n\t\treturn graph\n\n######################################3\n....\n\n#this works:\n\twith Pool(processes=64) as pool:\n\t\ttrain_graph = list(tqdm(pool.imap(smile_to_graph, train_smiles), total=num_train))\n\n#but this doesn't (out of resources?) ???\n\twith Pool(processes=64) as pool:\n\t\ttrain_graph = list(tqdm(pool.imap(smile_to_pygraph, train_smiles), total=num_train))\n\nERROR:\n  File \"/home/user/app/anaconda3.10/lib/python3.10/multiprocessing/reduction.py\", line 164, in recvfds\n    raise RuntimeError('received %d items of ancdata' %\nRuntimeError: received 0 items of ancdata\n\n\n```\n\nAny experienced kagglers know why?\nIt seems that either pytorch or pyg Data object mess up with multiprocessing resources.",
    "2810984": "not sure if different GNN will affect results, but if you want to try, you can google for github repo\nhttps://github.com/waqarahmadm019/AquaPred\n(GIN, GAT,GCN, attentionFP)",
    "2806762": "compare GNN and conv1d results (early experiments)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0b179520cf897935ff1058c07da9cffa%2FSelection_096.png?generation=1715417931640552&alt=media)\n",
    "2869040": "Update: I add ecfp to the GNN model and it reach 0.6 ap in the first 1000 iteration and 0.7 after the first 5 epoch",
    "2859294": "train = 50/70m samples, valid=5m (random split)\ndoes it mean you use 50m for training and 70m for fine tuning the model hyperparameters (layers, embedding ..etc)? \nThanks!",
    "2833290": "graph embedding space is well known for non-smooth.\nthen i have an idea\n\n```\ninstead of:\n model(x)= probability\n\nwe can:\n model(x) = nearest x that binds\n\nthan score = distance(input x, predicted nearest binding x)\n\n```",
    "2809614": "This is awesome, thanks for sharing! I'll need to study this to see what I missed in my implementation. I think we're only dealing with a subset of elements ['B', 'Br', 'C', 'Cl', 'Dy', 'F', 'I', 'N', 'O', 'S', 'Si'] in our dataset, so you could reduce the node dim substantially and avoid a bunch of all-zeros dimensions from element 1-hot encoding.",
    "2808421": "I have just a question (if you're willing to share of course), how have you handled the 3d position extraction? I am asking because RDKit takes a laughable amount of time for that",
    "2879221": ""
  }
}