{
  "id": 503327,
  "title": "[place holder] experimental results after metric update ...",
  "url": "/competitions/leash-BELKA/discussion/503327",
  "author_name": "",
  "post_date": "2024-05-17T00:19:07.873932100Z",
  "votes": 26,
  "comment_count": 2,
  "views": 0,
  "content": "<p>i will need to redo many of the experiments.<br>\nthe update has made the previous results obsolete.</p>\n<p>this will update as i go along …. please wait.</p>\n<p>the change in metric and update confirm the following ….</p>\n<p>(1) external data help:<br>\nsEH submission only:  external : 0.032, kaggle only: 0.026</p>\n<p>(2) models can overfit on training on share. (i note that some models improve on share, while nonshare get significantly worse). sometimes weak model generalized  better (share =0.6, nonshare = 0.1) compared to strong ones (share=0.8, nonshare=0.01) this is because you don't have nonshare in training set and high negative class means learning is classifying unknown as negative.</p>\n<p>(3) the ranking of xgboost, cnn1d, chermberta, molformer, gnn has changed! This is a surprised for me. This methods that they perform differently for different new subsets.</p>\n<p>(4) because of (2)(3), you can get better LB if you submit different models (e.g. train with share data but different level of regularization) for share and nonshare.</p>\n<p>but the kaggle objective is to get generalized model (i.e. one model to take care of all). I need to take of how to achieve this later (e.g. distillation/fusion to one model, mixture of experts, etc)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4c4e9dab30849fcc8e874e0fceea3a58%2FSelection_129.png?generation=1715910346727473&amp;alt=media\"><br>\ni forget to snapshot the results and LB before the metric change to do comparison and analysis.<br>\nThis is lesson learned.</p>",
  "messages": [
    {
      "id": "2817493",
      "postDate": "05/17/2024 00:19:07",
      "content": "<p>i will need to redo many of the experiments.<br>\nthe update has made the previous results obsolete.</p>\n<p>this will update as i go along …. please wait.</p>\n<p>the change in metric and update confirm the following ….</p>\n<p>(1) external data help:<br>\nsEH submission only:  external : 0.032, kaggle only: 0.026</p>\n<p>(2) models can overfit on training on share. (i note that some models improve on share, while nonshare get significantly worse). sometimes weak model generalized  better (share =0.6, nonshare = 0.1) compared to strong ones (share=0.8, nonshare=0.01) this is because you don't have nonshare in training set and high negative class means learning is classifying unknown as negative.</p>\n<p>(3) the ranking of xgboost, cnn1d, chermberta, molformer, gnn has changed! This is a surprised for me. This methods that they perform differently for different new subsets.</p>\n<p>(4) because of (2)(3), you can get better LB if you submit different models (e.g. train with share data but different level of regularization) for share and nonshare.</p>\n<p>but the kaggle objective is to get generalized model (i.e. one model to take care of all). I need to take of how to achieve this later (e.g. distillation/fusion to one model, mixture of experts, etc)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4c4e9dab30849fcc8e874e0fceea3a58%2FSelection_129.png?generation=1715910346727473&amp;alt=media\"><br>\ni forget to snapshot the results and LB before the metric change to do comparison and analysis.<br>\nThis is lesson learned.</p>",
      "rawMarkdown": "i will need to redo many of the experiments.\nthe update has made the previous results obsolete.\n\nthis will update as i go along .... please wait.\n\nthe change in metric and update confirm the following ....\n\n(1) external data help:\nsEH submission only:  external : 0.032, kaggle only: 0.026\n\n(2) models can overfit on training on share. (i note that some models improve on share, while nonshare get significantly worse). sometimes weak model generalized  better (share =0.6, nonshare = 0.1) compared to strong ones (share=0.8, nonshare=0.01) this is because you don't have nonshare in training set and high negative class means learning is classifying unknown as negative.\n\n(3) the ranking of xgboost, cnn1d, chermberta, molformer, gnn has changed! This is a surprised for me. This methods that they perform differently for different new subsets.\n\n(4) because of (2)(3), you can get better LB if you submit different models (e.g. train with share data but different level of regularization) for share and nonshare.\n\nbut the kaggle objective is to get generalized model (i.e. one model to take care of all). I need to take of how to achieve this later (e.g. distillation/fusion to one model, mixture of experts, etc)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4c4e9dab30849fcc8e874e0fceea3a58%2FSelection_129.png?generation=1715910346727473&alt=media)\ni forget to snapshot the results and LB before the metric change to do comparison and analysis.\nThis is lesson learned.",
      "votes": null
    },
    {
      "id": "2857405",
      "postDate": "06/05/2024 21:26:25",
      "content": "<p>we are back to where we start!</p>\n<p>consider the \"nonshare molecule\". let's split by building blocks bb1.<br>\nNow if the nonshare bb1 is similar to the train share bb1, we would have the same predict distribution and good results.<br>\nif they are different, nonshare bb has very low scores … but the ranking \"may be preserved\" /</p>\n<p>confused? check the code below:</p>\n<pre><code>  test cases as \np=npz   your model\nt=npz\nnonshare_idx=npz  row  the reduced train csv file\nBRD4_non = (t, )\nHSA_non = (t, )\nsEH_non = (t, )\navg_non = (BRD4_non + HSA_non + sEH_non) / \n\n\n\n\n\n\n results:\navg_non  : \nBRD4_non : \nHSA_non  : \nsEH_non  : \n</code></pre>\n<pre><code> each block\n\nreduce_df = pd()\nnonshare_df = reduce_df\nnonshare_df=\nnonshare_df=\nnonshare_df=\n\ngb = nonshare_df()\nBRD4_bb = \nHSA_bb  = \nsEH_bb  = \n bb,df  gb:\n    (df)\n     = df]\n    t = df]\n    BRD4 = (t, )\n    HSA = (t, )\n    sEH = (t, )\n    BRD4_bb(BRD4)\n    HSA_bb(HSA)\n    sEH_bb(sEH)\n    ((df), bb,BRD4,HSA,sEH)\nzz=\n\n\n)\n)\n)\n)\n\n\n results:\navg: \n\n\n\n</code></pre>\n<p>put it simply:</p>\n<ul>\n<li>our task is to detect all cats in a given set of images.</li>\n<li>if we ASSUME that each image contains at least some cat, ….</li>\n</ul>",
      "rawMarkdown": "we are back to where we start!\n\nconsider the \"nonshare molecule\". let's split by building blocks bb1.\nNow if the nonshare bb1 is similar to the train share bb1, we would have the same predict distribution and good results.\nif they are different, nonshare bb has very low scores ... but the ranking \"may be preserved\" /\n\n\nconfused? check the code below:\n\n```\n#consider all test cases as \"one\"\np=npz['prob']  #from your model\nt=npz['truth']\nnonshare_idx=npz['index'] #which row in the reduced train csv file\nBRD4_non = average_precision_score(t[:, 0], p[:, 0])\nHSA_non = average_precision_score(t[:, 1], p[:, 1])\nsEH_non = average_precision_score(t[:, 2], p[:, 2])\navg_non = (BRD4_non + HSA_non + sEH_non) / 3\n\nprint(f'avg_non  : {avg_non}')\nprint(f'BRD4_non : {BRD4_non}')\nprint(f'HSA_non  : {HSA_non}')\nprint(f'sEH_non  : {sEH_non}')\n\n#printout results:\navg_non  : 0.006616452608343465\nBRD4_non : 0.013142123275586589\nHSA_non  : 0.003945994730309735\nsEH_non  : 0.002761239819134072\n\n\n```\n\n```\n#consider each block\n\nreduce_df = pd.read_parquet('/.../train.reduced.parquet')\nnonshare_df = reduce_df.loc[nonshare_idx]\nnonshare_df.loc[:,'predict_BRD4']=p[:,0]\nnonshare_df.loc[:,'predict_HSA' ]=p[:,1]\nnonshare_df.loc[:,'predict_sEH' ]=p[:,2]\n\ngb = nonshare_df.groupby('buildingblock1_id')\nBRD4_bb = []\nHSA_bb  = []\nsEH_bb  = []\nfor bb,df in gb:\n\t#print(df)\n\tp = df[['predict_BRD4', 'predict_HSA', 'predict_sEH']].values\n\tt = df[['binds_BRD4', 'binds_HSA', 'binds_sEH']].values\n\tBRD4 = average_precision_score(t[:, 0], p[:, 0])\n\tHSA = average_precision_score(t[:, 1], p[:, 1])\n\tsEH = average_precision_score(t[:, 2], p[:, 2])\n\tBRD4_bb.append(BRD4)\n\tHSA_bb.append(HSA)\n\tsEH_bb.append(sEH)\n\tprint(len(df), bb,BRD4,HSA,sEH)\nzz=0\n\n\nprint(np.mean(BRD4_bb+HSA_bb+sEH_bb))\nprint(np.mean(BRD4_bb))\nprint(np.mean(HSA_bb))\nprint(np.mean(sEH_bb))\n\n\n#printout results:\navg: 0.009247734349738766\n0.020241097506987816\n0.004497381974609815\n0.0030047235676186664\n\n\n```\n\n\nput it simply:\n- our task is to detect all cats in a given set of images.\n- if we ASSUME that each image contains at least some cat, ....",
      "votes": null
    },
    {
      "id": "2872056",
      "postDate": "06/14/2024 14:47:32",
      "content": "<p>a note about nonshare triazine molecules in public/private test set:</p>\n<pre><code>      \n                    \n\n      \n                    \n  \n  \n\n  \n  \n  \n  \n\n\n\n      \n\n             \n       \n  \n  \n\n  \n  \n  \n  \n</code></pre>",
      "rawMarkdown": "a note about nonshare triazine molecules in public/private test set:\n````\nwe perform knn on the building blocks:\n1) for bb1 of nonshare triazine test molecules, the knn from train has a has tanimoto similarity (ecfp6) of greater 0.80\n\n1449 1.00000 C#CCCC[C@H](NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O   #query test\n0100 0.84483 CCCCC(NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O               #top 3 knn train\n0215 0.82456 C#CC[C@H](NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O\n0704 0.82456 C#CC[C@@H](NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O\n \n1696 1.00000 CC(C)C(NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O\n1130 0.87379 CCC(C)C(NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O\n1278 0.85417 CC(NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O\n0483 0.84112 CC(OC(C)(C)C)C(NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O\n```\n\n```\n2)for bb2, this is only about 0.50\n\n1811 1.00000 Cl.Cl.N=C(N)c1cccc(CN)c1          #query test\n0632 0.58333 NCc1cccc(C(=O)N2CCCC2)c1  #top 3 knn train\n0208 0.50000 NCc1cccc(C(F)(F)F)c1\n0582 0.47059 C#CCOc1cccc(CN)c1.Cl\n\n1859 1.00000 Cl.Cl.NCc1ccn2ccnc2c1\n0530 0.46377 Cl.NCc1ccnc(C(N)=O)c1\n0888 0.44118 Nc1ccc2nccn2c1\n0526 0.42105 Cl.Cl.NCc1cccc(-n2ccnn2)c1\n\n````",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2857405,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/05/2024 21:26:25",
      "content": "<p>we are back to where we start!</p>\n<p>consider the \"nonshare molecule\". let's split by building blocks bb1.<br>\nNow if the nonshare bb1 is similar to the train share bb1, we would have the same predict distribution and good results.<br>\nif they are different, nonshare bb has very low scores … but the ranking \"may be preserved\" /</p>\n<p>confused? check the code below:</p>\n<pre><code>  test cases as \np=npz   your model\nt=npz\nnonshare_idx=npz  row  the reduced train csv file\nBRD4_non = (t, )\nHSA_non = (t, )\nsEH_non = (t, )\navg_non = (BRD4_non + HSA_non + sEH_non) / \n\n\n\n\n\n\n results:\navg_non  : \nBRD4_non : \nHSA_non  : \nsEH_non  : \n</code></pre>\n<pre><code> each block\n\nreduce_df = pd()\nnonshare_df = reduce_df\nnonshare_df=\nnonshare_df=\nnonshare_df=\n\ngb = nonshare_df()\nBRD4_bb = \nHSA_bb  = \nsEH_bb  = \n bb,df  gb:\n    (df)\n     = df]\n    t = df]\n    BRD4 = (t, )\n    HSA = (t, )\n    sEH = (t, )\n    BRD4_bb(BRD4)\n    HSA_bb(HSA)\n    sEH_bb(sEH)\n    ((df), bb,BRD4,HSA,sEH)\nzz=\n\n\n)\n)\n)\n)\n\n\n results:\navg: \n\n\n\n</code></pre>\n<p>put it simply:</p>\n<ul>\n<li>our task is to detect all cats in a given set of images.</li>\n<li>if we ASSUME that each image contains at least some cat, ….</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2872056,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/14/2024 14:47:32",
      "content": "<p>a note about nonshare triazine molecules in public/private test set:</p>\n<pre><code>      \n                    \n\n      \n                    \n  \n  \n\n  \n  \n  \n  \n\n\n\n      \n\n             \n       \n  \n  \n\n  \n  \n  \n  \n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2817493": "i will need to redo many of the experiments.\nthe update has made the previous results obsolete.\n\nthis will update as i go along .... please wait.\n\nthe change in metric and update confirm the following ....\n\n(1) external data help:\nsEH submission only:  external : 0.032, kaggle only: 0.026\n\n(2) models can overfit on training on share. (i note that some models improve on share, while nonshare get significantly worse). sometimes weak model generalized  better (share =0.6, nonshare = 0.1) compared to strong ones (share=0.8, nonshare=0.01) this is because you don't have nonshare in training set and high negative class means learning is classifying unknown as negative.\n\n(3) the ranking of xgboost, cnn1d, chermberta, molformer, gnn has changed! This is a surprised for me. This methods that they perform differently for different new subsets.\n\n(4) because of (2)(3), you can get better LB if you submit different models (e.g. train with share data but different level of regularization) for share and nonshare.\n\nbut the kaggle objective is to get generalized model (i.e. one model to take care of all). I need to take of how to achieve this later (e.g. distillation/fusion to one model, mixture of experts, etc)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4c4e9dab30849fcc8e874e0fceea3a58%2FSelection_129.png?generation=1715910346727473&alt=media)\ni forget to snapshot the results and LB before the metric change to do comparison and analysis.\nThis is lesson learned.",
    "2857405": "we are back to where we start!\n\nconsider the \"nonshare molecule\". let's split by building blocks bb1.\nNow if the nonshare bb1 is similar to the train share bb1, we would have the same predict distribution and good results.\nif they are different, nonshare bb has very low scores ... but the ranking \"may be preserved\" /\n\n\nconfused? check the code below:\n\n```\n#consider all test cases as \"one\"\np=npz['prob']  #from your model\nt=npz['truth']\nnonshare_idx=npz['index'] #which row in the reduced train csv file\nBRD4_non = average_precision_score(t[:, 0], p[:, 0])\nHSA_non = average_precision_score(t[:, 1], p[:, 1])\nsEH_non = average_precision_score(t[:, 2], p[:, 2])\navg_non = (BRD4_non + HSA_non + sEH_non) / 3\n\nprint(f'avg_non  : {avg_non}')\nprint(f'BRD4_non : {BRD4_non}')\nprint(f'HSA_non  : {HSA_non}')\nprint(f'sEH_non  : {sEH_non}')\n\n#printout results:\navg_non  : 0.006616452608343465\nBRD4_non : 0.013142123275586589\nHSA_non  : 0.003945994730309735\nsEH_non  : 0.002761239819134072\n\n\n```\n\n```\n#consider each block\n\nreduce_df = pd.read_parquet('/.../train.reduced.parquet')\nnonshare_df = reduce_df.loc[nonshare_idx]\nnonshare_df.loc[:,'predict_BRD4']=p[:,0]\nnonshare_df.loc[:,'predict_HSA' ]=p[:,1]\nnonshare_df.loc[:,'predict_sEH' ]=p[:,2]\n\ngb = nonshare_df.groupby('buildingblock1_id')\nBRD4_bb = []\nHSA_bb  = []\nsEH_bb  = []\nfor bb,df in gb:\n\t#print(df)\n\tp = df[['predict_BRD4', 'predict_HSA', 'predict_sEH']].values\n\tt = df[['binds_BRD4', 'binds_HSA', 'binds_sEH']].values\n\tBRD4 = average_precision_score(t[:, 0], p[:, 0])\n\tHSA = average_precision_score(t[:, 1], p[:, 1])\n\tsEH = average_precision_score(t[:, 2], p[:, 2])\n\tBRD4_bb.append(BRD4)\n\tHSA_bb.append(HSA)\n\tsEH_bb.append(sEH)\n\tprint(len(df), bb,BRD4,HSA,sEH)\nzz=0\n\n\nprint(np.mean(BRD4_bb+HSA_bb+sEH_bb))\nprint(np.mean(BRD4_bb))\nprint(np.mean(HSA_bb))\nprint(np.mean(sEH_bb))\n\n\n#printout results:\navg: 0.009247734349738766\n0.020241097506987816\n0.004497381974609815\n0.0030047235676186664\n\n\n```\n\n\nput it simply:\n- our task is to detect all cats in a given set of images.\n- if we ASSUME that each image contains at least some cat, ....",
    "2872056": "a note about nonshare triazine molecules in public/private test set:\n````\nwe perform knn on the building blocks:\n1) for bb1 of nonshare triazine test molecules, the knn from train has a has tanimoto similarity (ecfp6) of greater 0.80\n\n1449 1.00000 C#CCCC[C@H](NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O   #query test\n0100 0.84483 CCCCC(NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O               #top 3 knn train\n0215 0.82456 C#CC[C@H](NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O\n0704 0.82456 C#CC[C@@H](NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O\n \n1696 1.00000 CC(C)C(NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O\n1130 0.87379 CCC(C)C(NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O\n1278 0.85417 CC(NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O\n0483 0.84112 CC(OC(C)(C)C)C(NC(=O)OCC1c2ccccc2-c2ccccc21)C(=O)O\n```\n\n```\n2)for bb2, this is only about 0.50\n\n1811 1.00000 Cl.Cl.N=C(N)c1cccc(CN)c1          #query test\n0632 0.58333 NCc1cccc(C(=O)N2CCCC2)c1  #top 3 knn train\n0208 0.50000 NCc1cccc(C(F)(F)F)c1\n0582 0.47059 C#CCOc1cccc(CN)c1.Cl\n\n1859 1.00000 Cl.Cl.NCc1ccn2ccnc2c1\n0530 0.46377 Cl.NCc1ccnc(C(N)=O)c1\n0888 0.44118 Nc1ccc2nccn2c1\n0526 0.42105 Cl.Cl.NCc1cccc(-n2ccnn2)c1\n\n````"
  },
  "source": "meta"
}