{
  "id": 508270,
  "title": "how to identify collision in ECFP hashing?",
  "url": "/competitions/leash-BELKA/discussion/508270",
  "author_name": "",
  "post_date": "2024-05-28T23:50:32.529878500Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>i am exploring the use of ECFP as tokenizer in transformer model<br>\nhow do i know which bit of the e.g. 2048 bit of ECFP has collisions or not?</p>",
  "messages": [
    {
      "id": "2842184",
      "postDate": "05/28/2024 23:50:32",
      "content": "<p>i am exploring the use of ECFP as tokenizer in transformer model<br>\nhow do i know which bit of the e.g. 2048 bit of ECFP has collisions or not?</p>",
      "rawMarkdown": "i am exploring the use of ECFP as tokenizer in transformer model\nhow do i know which bit of the e.g. 2048 bit of ECFP has collisions or not?",
      "votes": null
    },
    {
      "id": "2842194",
      "postDate": "05/29/2024 00:32:30",
      "content": "<p>These posts might be helpful</p>\n<p><a href=\"https://rdkit.blogspot.com/2016/02/morgan-fingerprint-bit-statistics.html\" target=\"_blank\">https://rdkit.blogspot.com/2016/02/morgan-fingerprint-bit-statistics.html</a><br>\n<a href=\"https://rdkit.blogspot.com/2014/02/colliding-bits.html\" target=\"_blank\">https://rdkit.blogspot.com/2014/02/colliding-bits.html</a><br>\n<a href=\"https://rdkit.blogspot.com/2014/03/colliding-bits-ii.html\" target=\"_blank\">https://rdkit.blogspot.com/2014/03/colliding-bits-ii.html</a><br>\n<a href=\"https://rdkit.blogspot.com/2016/02/colliding-bits-iii.html\" target=\"_blank\">https://rdkit.blogspot.com/2016/02/colliding-bits-iii.html</a></p>\n<p>That said, I don't think you can determine which molecular features will result in collisions, at least not without exhaustive search.</p>\n<p>Morgan fingerprints calculate a set of properties for every atom in a molecule. You would have to dig around in rdkit to find what specific properties they use, but it would be stuff like atomic number, degree, number of hs, etc. The properties of a given atom and its n-hop neighbors (radius param for Morgan fingerprint) are hashed together to create a value. These hash values are unbounded and tend to have a very low rate of collision (I've never seen one but I haven't looked very hard).</p>\n<p>The unbounded hash values are then \"folded\" to a fixed length. This is the final Morgan fingerprint, which has a much higher rate of hash collisions.</p>\n<p>You can use the Sparse fingerprint representation to inspect the pre-folding values.</p>\n<p>You can find some examples of this in my notebook: <br>\n<a href=\"https://www.kaggle.com/code/towardsentropy/fingerprint-tips-and-tricks/notebook\" target=\"_blank\">https://www.kaggle.com/code/towardsentropy/fingerprint-tips-and-tricks/notebook</a></p>",
      "rawMarkdown": "These posts might be helpful\n\nhttps://rdkit.blogspot.com/2016/02/morgan-fingerprint-bit-statistics.html\nhttps://rdkit.blogspot.com/2014/02/colliding-bits.html\nhttps://rdkit.blogspot.com/2014/03/colliding-bits-ii.html\nhttps://rdkit.blogspot.com/2016/02/colliding-bits-iii.html\n\nThat said, I don't think you can determine which molecular features will result in collisions, at least not without exhaustive search.\n\nMorgan fingerprints calculate a set of properties for every atom in a molecule. You would have to dig around in rdkit to find what specific properties they use, but it would be stuff like atomic number, degree, number of hs, etc. The properties of a given atom and its n-hop neighbors (radius param for Morgan fingerprint) are hashed together to create a value. These hash values are unbounded and tend to have a very low rate of collision (I've never seen one but I haven't looked very hard).\n\nThe unbounded hash values are then \"folded\" to a fixed length. This is the final Morgan fingerprint, which has a much higher rate of hash collisions.\n\nYou can use the Sparse fingerprint representation to inspect the pre-folding values.\n\nYou can find some examples of this in my notebook: \nhttps://www.kaggle.com/code/towardsentropy/fingerprint-tips-and-tricks/notebook",
      "votes": null
    },
    {
      "id": "2842570",
      "postDate": "05/29/2024 06:50:35",
      "content": "<p>thanks these are useful</p>\n<p>meanwhile, i find an alternative, using vector quantisation vae</p>\n<p>Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules<br>\n<a href=\"https://openreview.net/forum?id=jevY-DtiZTR\" target=\"_blank\">https://openreview.net/forum?id=jevY-DtiZTR</a><br>\n<a href=\"https://github.com/Rich-XGK/GTMGC/tree/master\" target=\"_blank\">https://github.com/Rich-XGK/GTMGC/tree/master</a></p>",
      "rawMarkdown": "thanks these are useful\n\nmeanwhile, i find an alternative, using vector quantisation vae\n\nMole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules\nhttps://openreview.net/forum?id=jevY-DtiZTR\nhttps://github.com/Rich-XGK/GTMGC/tree/master",
      "votes": null
    },
    {
      "id": "2855629",
      "postDate": "06/04/2024 23:38:45",
      "content": "<p>Sort &amp; Slice: A Simple and Superior Alternative to Hash-Based Folding for Extended-Connectivity Fingerprints<br>\n<a href=\"https://arxiv.org/abs/2403.17954\" target=\"_blank\">https://arxiv.org/abs/2403.17954</a><br>\n<a href=\"https://github.com/MarkusFerdinandDablander/ECFP-substructure-pooling-Sort-and-Slice\" target=\"_blank\">https://github.com/MarkusFerdinandDablander/ECFP-substructure-pooling-Sort-and-Slice</a></p>\n<p>\"We go on to describe Sort &amp; Slice, an easy-to-implement and bit-collision-free alternative to hash-based folding\"</p>",
      "rawMarkdown": "Sort & Slice: A Simple and Superior Alternative to Hash-Based Folding for Extended-Connectivity Fingerprints\nhttps://arxiv.org/abs/2403.17954\nhttps://github.com/MarkusFerdinandDablander/ECFP-substructure-pooling-Sort-and-Slice\n\n\"We go on to describe Sort & Slice, an easy-to-implement and bit-collision-free alternative to hash-based folding\"",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2842194,
      "author_name": "towardsentropy",
      "author_url": "",
      "post_date": "05/29/2024 00:32:30",
      "content": "<p>These posts might be helpful</p>\n<p><a href=\"https://rdkit.blogspot.com/2016/02/morgan-fingerprint-bit-statistics.html\" target=\"_blank\">https://rdkit.blogspot.com/2016/02/morgan-fingerprint-bit-statistics.html</a><br>\n<a href=\"https://rdkit.blogspot.com/2014/02/colliding-bits.html\" target=\"_blank\">https://rdkit.blogspot.com/2014/02/colliding-bits.html</a><br>\n<a href=\"https://rdkit.blogspot.com/2014/03/colliding-bits-ii.html\" target=\"_blank\">https://rdkit.blogspot.com/2014/03/colliding-bits-ii.html</a><br>\n<a href=\"https://rdkit.blogspot.com/2016/02/colliding-bits-iii.html\" target=\"_blank\">https://rdkit.blogspot.com/2016/02/colliding-bits-iii.html</a></p>\n<p>That said, I don't think you can determine which molecular features will result in collisions, at least not without exhaustive search.</p>\n<p>Morgan fingerprints calculate a set of properties for every atom in a molecule. You would have to dig around in rdkit to find what specific properties they use, but it would be stuff like atomic number, degree, number of hs, etc. The properties of a given atom and its n-hop neighbors (radius param for Morgan fingerprint) are hashed together to create a value. These hash values are unbounded and tend to have a very low rate of collision (I've never seen one but I haven't looked very hard).</p>\n<p>The unbounded hash values are then \"folded\" to a fixed length. This is the final Morgan fingerprint, which has a much higher rate of hash collisions.</p>\n<p>You can use the Sparse fingerprint representation to inspect the pre-folding values.</p>\n<p>You can find some examples of this in my notebook: <br>\n<a href=\"https://www.kaggle.com/code/towardsentropy/fingerprint-tips-and-tricks/notebook\" target=\"_blank\">https://www.kaggle.com/code/towardsentropy/fingerprint-tips-and-tricks/notebook</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2842570,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/29/2024 06:50:35",
          "content": "<p>thanks these are useful</p>\n<p>meanwhile, i find an alternative, using vector quantisation vae</p>\n<p>Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules<br>\n<a href=\"https://openreview.net/forum?id=jevY-DtiZTR\" target=\"_blank\">https://openreview.net/forum?id=jevY-DtiZTR</a><br>\n<a href=\"https://github.com/Rich-XGK/GTMGC/tree/master\" target=\"_blank\">https://github.com/Rich-XGK/GTMGC/tree/master</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2855629,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/04/2024 23:38:45",
      "content": "<p>Sort &amp; Slice: A Simple and Superior Alternative to Hash-Based Folding for Extended-Connectivity Fingerprints<br>\n<a href=\"https://arxiv.org/abs/2403.17954\" target=\"_blank\">https://arxiv.org/abs/2403.17954</a><br>\n<a href=\"https://github.com/MarkusFerdinandDablander/ECFP-substructure-pooling-Sort-and-Slice\" target=\"_blank\">https://github.com/MarkusFerdinandDablander/ECFP-substructure-pooling-Sort-and-Slice</a></p>\n<p>\"We go on to describe Sort &amp; Slice, an easy-to-implement and bit-collision-free alternative to hash-based folding\"</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2842184": "i am exploring the use of ECFP as tokenizer in transformer model\nhow do i know which bit of the e.g. 2048 bit of ECFP has collisions or not?",
    "2842194": "These posts might be helpful\n\nhttps://rdkit.blogspot.com/2016/02/morgan-fingerprint-bit-statistics.html\nhttps://rdkit.blogspot.com/2014/02/colliding-bits.html\nhttps://rdkit.blogspot.com/2014/03/colliding-bits-ii.html\nhttps://rdkit.blogspot.com/2016/02/colliding-bits-iii.html\n\nThat said, I don't think you can determine which molecular features will result in collisions, at least not without exhaustive search.\n\nMorgan fingerprints calculate a set of properties for every atom in a molecule. You would have to dig around in rdkit to find what specific properties they use, but it would be stuff like atomic number, degree, number of hs, etc. The properties of a given atom and its n-hop neighbors (radius param for Morgan fingerprint) are hashed together to create a value. These hash values are unbounded and tend to have a very low rate of collision (I've never seen one but I haven't looked very hard).\n\nThe unbounded hash values are then \"folded\" to a fixed length. This is the final Morgan fingerprint, which has a much higher rate of hash collisions.\n\nYou can use the Sparse fingerprint representation to inspect the pre-folding values.\n\nYou can find some examples of this in my notebook: \nhttps://www.kaggle.com/code/towardsentropy/fingerprint-tips-and-tricks/notebook",
    "2842570": "thanks these are useful\n\nmeanwhile, i find an alternative, using vector quantisation vae\n\nMole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules\nhttps://openreview.net/forum?id=jevY-DtiZTR\nhttps://github.com/Rich-XGK/GTMGC/tree/master",
    "2855629": "Sort & Slice: A Simple and Superior Alternative to Hash-Based Folding for Extended-Connectivity Fingerprints\nhttps://arxiv.org/abs/2403.17954\nhttps://github.com/MarkusFerdinandDablander/ECFP-substructure-pooling-Sort-and-Slice\n\n\"We go on to describe Sort & Slice, an easy-to-implement and bit-collision-free alternative to hash-based folding\""
  },
  "source": "meta"
}