{
  "id": 499813,
  "title": "xgboost  & ECFP . Filter 50% low variance ecfp columns get same score.",
  "url": "/competitions/leash-BELKA/discussion/499813",
  "author_name": "",
  "post_date": "2024-05-03T06:10:36.054200200Z",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello, I have made some experiment selecting some ECFP columns and removing more that 50% of the columns I got the same score.</p>\n<h1>Get Variance for different threshold,</h1>\n<p>from sklearn.feature_selection import VarianceThreshold<br>\nfor i in range(6):<br>\n    threshold=0.1/10**i<br>\n    var_thresh = VarianceThreshold(threshold=threshold)<br>\n    var_thresh.fit(train_ecfp[:100000].A)<br>\n    var_thresh_index=var_thresh.get_support()<br>\n    print(threshold,sum(var_thresh_index))</p>\n<p>0.1 100<br>\n0.01 940 --&gt; training with just 940 columns.<br>\n0.001 1822<br>\n0.0001 1822 <br>\n1e-05 1822<br>\n1e-06 1822</p>\n<h1>Apply mask to ecfp sparse matrix.</h1>\n<p>var_thresh = VarianceThreshold(threshold=0.01)<br>\nvar_thresh.fit(train_ecfp[:100000].A)<br>\nvar_thresh_index_1=var_thresh.get_support()<br>\ntrain_ecfp_1=train_ecfp[:,var_thresh_index_1]<br>\nprint('train_ecfp Shape :',train_ecfp_1.shape)</p>\n<p>train_ecfp Shape : (9509779, 940)</p>\n<p>Notebook: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ricopue/leashbio-xgb-ecfp-10m-sample-rows</a></p>",
  "messages": [
    {
      "id": "2790322",
      "postDate": "05/03/2024 06:10:36",
      "content": "<p>Hello, I have made some experiment selecting some ECFP columns and removing more that 50% of the columns I got the same score.</p>\n<h1>Get Variance for different threshold,</h1>\n<p>from sklearn.feature_selection import VarianceThreshold<br>\nfor i in range(6):<br>\n    threshold=0.1/10**i<br>\n    var_thresh = VarianceThreshold(threshold=threshold)<br>\n    var_thresh.fit(train_ecfp[:100000].A)<br>\n    var_thresh_index=var_thresh.get_support()<br>\n    print(threshold,sum(var_thresh_index))</p>\n<p>0.1 100<br>\n0.01 940 --&gt; training with just 940 columns.<br>\n0.001 1822<br>\n0.0001 1822 <br>\n1e-05 1822<br>\n1e-06 1822</p>\n<h1>Apply mask to ecfp sparse matrix.</h1>\n<p>var_thresh = VarianceThreshold(threshold=0.01)<br>\nvar_thresh.fit(train_ecfp[:100000].A)<br>\nvar_thresh_index_1=var_thresh.get_support()<br>\ntrain_ecfp_1=train_ecfp[:,var_thresh_index_1]<br>\nprint('train_ecfp Shape :',train_ecfp_1.shape)</p>\n<p>train_ecfp Shape : (9509779, 940)</p>\n<p>Notebook: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ricopue/leashbio-xgb-ecfp-10m-sample-rows</a></p>",
      "rawMarkdown": "Hello, I have made some experiment selecting some ECFP columns and removing more that 50% of the columns I got the same score.\n\n#Get Variance for different threshold,\nfrom sklearn.feature_selection import VarianceThreshold\nfor i in range(6):\n    threshold=0.1/10**i\n    var_thresh = VarianceThreshold(threshold=threshold)\n    var_thresh.fit(train_ecfp[:100000].A)\n    var_thresh_index=var_thresh.get_support()\n    print(threshold,sum(var_thresh_index))\n\n0.1 100\n0.01 940 --> training with just 940 columns.\n0.001 1822\n0.0001 1822 \n1e-05 1822\n1e-06 1822\n\n#Apply mask to ecfp sparse matrix.\nvar_thresh = VarianceThreshold(threshold=0.01)\nvar_thresh.fit(train_ecfp[:100000].A)\nvar_thresh_index_1=var_thresh.get_support()\ntrain_ecfp_1=train_ecfp[:,var_thresh_index_1]\nprint('train_ecfp Shape :',train_ecfp_1.shape)\n\ntrain_ecfp Shape : (9509779, 940)\n\nNotebook: [https://www.kaggle.com/code/ricopue/leashbio-xgb-ecfp-10m-sample-rows](url)",
      "votes": null
    },
    {
      "id": "2792829",
      "postDate": "05/04/2024 12:14:07",
      "content": "<p>That's a nice finding, thanks for sharing!</p>\n<p>And also it checks out with the intuition:<br>\nECFP is a <strong>hashed fingerprint</strong> - it's constructed by obtaining subgraphs from the molecular graph, hashing them and performing a modulo operation on the hashed values to get a place for 1 in the resulting vector (e.g. <strong>mod 2048</strong> -&gt; resulting bit vector is of size <strong>2048</strong>). The thing is, that those subgraphs are often obtained through getting wider and wider neighborhood for each atom. On one hand it is very flexible, and you don't need to predetermine such chemical structures, but on the other it generates many substructures, which occur very frequently (and so yield not much information = low variance). That may be especially a case in this competition, where the molecules are created from predefined ~2k building blocks.</p>\n<p>That strategy probably wouldn't work well with descriptor fingerprints such as MACCS Keys fingerprint - there, the substructures are developed by chemists, so they should be very informative.</p>\n<p>Second thought, that I've got just now - because of that, it might be beneficial to use higher size of resulting vectors in fingerprints. Lower size means higher chance of <strong>bit collisions</strong> (hashes of 2 or more substructures modulo vector size result in the same position). Here, it is important to keep in mind memory limitations due to the size of the data, but if we can save and load some dataframe, then the vector size is no problem. Variance threshold will eliminate unnecessary columns and shrink the vector, but at the same time we have low chance of missing influential features.</p>",
      "rawMarkdown": "That's a nice finding, thanks for sharing!\n\nAnd also it checks out with the intuition:\nECFP is a **hashed fingerprint** - it's constructed by obtaining subgraphs from the molecular graph, hashing them and performing a modulo operation on the hashed values to get a place for 1 in the resulting vector (e.g. **mod 2048** -> resulting bit vector is of size **2048**). The thing is, that those subgraphs are often obtained through getting wider and wider neighborhood for each atom. On one hand it is very flexible, and you don't need to predetermine such chemical structures, but on the other it generates many substructures, which occur very frequently (and so yield not much information = low variance). That may be especially a case in this competition, where the molecules are created from predefined ~2k building blocks.\n\nThat strategy probably wouldn't work well with descriptor fingerprints such as MACCS Keys fingerprint - there, the substructures are developed by chemists, so they should be very informative.\n\nSecond thought, that I've got just now - because of that, it might be beneficial to use higher size of resulting vectors in fingerprints. Lower size means higher chance of **bit collisions** (hashes of 2 or more substructures modulo vector size result in the same position). Here, it is important to keep in mind memory limitations due to the size of the data, but if we can save and load some dataframe, then the vector size is no problem. Variance threshold will eliminate unnecessary columns and shrink the vector, but at the same time we have low chance of missing influential features.",
      "votes": null
    },
    {
      "id": "2792858",
      "postDate": "05/04/2024 12:40:03",
      "content": "<p>Thanks for your thought, very interesting info. My strategy in the competition is work if there are two different problems. One for tarm/test share smiles and for not share smiles. The good point for ECFP  fingerprint  is that they can be stored and managed as spare matrix. </p>",
      "rawMarkdown": "Thanks for your thought, very interesting info. My strategy in the competition is work if there are two different problems. One for tarm/test share smiles and for not share smiles. The good point for ECFP  fingerprint  is that they can be stored and managed as spare matrix.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2792829,
      "author_name": "michaszafarczyk",
      "author_url": "",
      "post_date": "05/04/2024 12:14:07",
      "content": "<p>That's a nice finding, thanks for sharing!</p>\n<p>And also it checks out with the intuition:<br>\nECFP is a <strong>hashed fingerprint</strong> - it's constructed by obtaining subgraphs from the molecular graph, hashing them and performing a modulo operation on the hashed values to get a place for 1 in the resulting vector (e.g. <strong>mod 2048</strong> -&gt; resulting bit vector is of size <strong>2048</strong>). The thing is, that those subgraphs are often obtained through getting wider and wider neighborhood for each atom. On one hand it is very flexible, and you don't need to predetermine such chemical structures, but on the other it generates many substructures, which occur very frequently (and so yield not much information = low variance). That may be especially a case in this competition, where the molecules are created from predefined ~2k building blocks.</p>\n<p>That strategy probably wouldn't work well with descriptor fingerprints such as MACCS Keys fingerprint - there, the substructures are developed by chemists, so they should be very informative.</p>\n<p>Second thought, that I've got just now - because of that, it might be beneficial to use higher size of resulting vectors in fingerprints. Lower size means higher chance of <strong>bit collisions</strong> (hashes of 2 or more substructures modulo vector size result in the same position). Here, it is important to keep in mind memory limitations due to the size of the data, but if we can save and load some dataframe, then the vector size is no problem. Variance threshold will eliminate unnecessary columns and shrink the vector, but at the same time we have low chance of missing influential features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2792858,
          "author_name": "ricopue",
          "author_url": "",
          "post_date": "05/04/2024 12:40:03",
          "content": "<p>Thanks for your thought, very interesting info. My strategy in the competition is work if there are two different problems. One for tarm/test share smiles and for not share smiles. The good point for ECFP  fingerprint  is that they can be stored and managed as spare matrix. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2790322": "Hello, I have made some experiment selecting some ECFP columns and removing more that 50% of the columns I got the same score.\n\n#Get Variance for different threshold,\nfrom sklearn.feature_selection import VarianceThreshold\nfor i in range(6):\n    threshold=0.1/10**i\n    var_thresh = VarianceThreshold(threshold=threshold)\n    var_thresh.fit(train_ecfp[:100000].A)\n    var_thresh_index=var_thresh.get_support()\n    print(threshold,sum(var_thresh_index))\n\n0.1 100\n0.01 940 --> training with just 940 columns.\n0.001 1822\n0.0001 1822 \n1e-05 1822\n1e-06 1822\n\n#Apply mask to ecfp sparse matrix.\nvar_thresh = VarianceThreshold(threshold=0.01)\nvar_thresh.fit(train_ecfp[:100000].A)\nvar_thresh_index_1=var_thresh.get_support()\ntrain_ecfp_1=train_ecfp[:,var_thresh_index_1]\nprint('train_ecfp Shape :',train_ecfp_1.shape)\n\ntrain_ecfp Shape : (9509779, 940)\n\nNotebook: [https://www.kaggle.com/code/ricopue/leashbio-xgb-ecfp-10m-sample-rows](url)",
    "2792829": "That's a nice finding, thanks for sharing!\n\nAnd also it checks out with the intuition:\nECFP is a **hashed fingerprint** - it's constructed by obtaining subgraphs from the molecular graph, hashing them and performing a modulo operation on the hashed values to get a place for 1 in the resulting vector (e.g. **mod 2048** -> resulting bit vector is of size **2048**). The thing is, that those subgraphs are often obtained through getting wider and wider neighborhood for each atom. On one hand it is very flexible, and you don't need to predetermine such chemical structures, but on the other it generates many substructures, which occur very frequently (and so yield not much information = low variance). That may be especially a case in this competition, where the molecules are created from predefined ~2k building blocks.\n\nThat strategy probably wouldn't work well with descriptor fingerprints such as MACCS Keys fingerprint - there, the substructures are developed by chemists, so they should be very informative.\n\nSecond thought, that I've got just now - because of that, it might be beneficial to use higher size of resulting vectors in fingerprints. Lower size means higher chance of **bit collisions** (hashes of 2 or more substructures modulo vector size result in the same position). Here, it is important to keep in mind memory limitations due to the size of the data, but if we can save and load some dataframe, then the vector size is no problem. Variance threshold will eliminate unnecessary columns and shrink the vector, but at the same time we have low chance of missing influential features.",
    "2792858": "Thanks for your thought, very interesting info. My strategy in the competition is work if there are two different problems. One for tarm/test share smiles and for not share smiles. The good point for ECFP  fingerprint  is that they can be stored and managed as spare matrix."
  },
  "source": "meta"
}