{
  "id": 494876,
  "title": "Creating CV not containing common BBs between training data and validation data",
  "url": "/competitions/leash-BELKA/discussion/494876",
  "author_name": "",
  "post_date": "2024-04-18T18:47:44.708524500Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In this competition, test data does not contain the building blocks(BBs) which training data has. </p>\n<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).</p>\n</blockquote>\n<p>Making <strong>cv not containing common BBs between training data and validation data</strong> is important to validate our models correctly.</p>\n<p>I made the following cv:</p>\n<ul>\n<li>There is no common BBs between training data and validation data</li>\n<li>The ratio of our predicting target \"binds\" is almost same between training data and validation data, among each folds.</li>\n</ul>\n<p>Following is the detail of my cv. <code>set()</code> indicate there is no intersection.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7853876%2F83d892e8cd06ff1380dd783752c31254%2Fcv_detail.png?generation=1713464599721993&amp;alt=media\"></p>\n<p>My cv is available on kaggle dataset:<br>\n<a href=\"https://www.kaggle.com/datasets/yusaku5739/belka-cv-no-common-bbs\" target=\"_blank\">BELKA-cv-no-common-bbs</a></p>\n<p>The notebook creating my cv is comming soon…</p>\n<p>In order to get traning and validation data for each fold, please use following function.</p>\n<pre><code>def get_train_valid_df(fold, protein_name, dataset_path, df_fold_path):\n    \n    train = pd.read_parquet(os.path.join(dataset_path, ), columns=[, , ,  ], filters=[[(, , protein_name)]])\n\n    df_path = os.path.join(df_fold_path, f)\n    df = pd.read_csv(df_path)\n\n    df_t = df.query()\n    df_v = df.query()\n\n    t_smiles = df_t[].to_list()\n    v_smiles = df_v[].to_list()\n    df_train = train.query()\n    df_train = df_train.query()\n    df_valid = train.query()\n    df_valid = df_valid.query()\n\n     df_train, df_valid\n\n\ndf_train, df_valid = get_train_valid_df(, , , )\n</code></pre>",
  "messages": [
    {
      "id": "2759566",
      "postDate": "04/18/2024 18:47:44",
      "content": "<p>In this competition, test data does not contain the building blocks(BBs) which training data has. </p>\n<blockquote>\n  <p>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).</p>\n</blockquote>\n<p>Making <strong>cv not containing common BBs between training data and validation data</strong> is important to validate our models correctly.</p>\n<p>I made the following cv:</p>\n<ul>\n<li>There is no common BBs between training data and validation data</li>\n<li>The ratio of our predicting target \"binds\" is almost same between training data and validation data, among each folds.</li>\n</ul>\n<p>Following is the detail of my cv. <code>set()</code> indicate there is no intersection.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7853876%2F83d892e8cd06ff1380dd783752c31254%2Fcv_detail.png?generation=1713464599721993&amp;alt=media\"></p>\n<p>My cv is available on kaggle dataset:<br>\n<a href=\"https://www.kaggle.com/datasets/yusaku5739/belka-cv-no-common-bbs\" target=\"_blank\">BELKA-cv-no-common-bbs</a></p>\n<p>The notebook creating my cv is comming soon…</p>\n<p>In order to get traning and validation data for each fold, please use following function.</p>\n<pre><code>def get_train_valid_df(fold, protein_name, dataset_path, df_fold_path):\n    \n    train = pd.read_parquet(os.path.join(dataset_path, ), columns=[, , ,  ], filters=[[(, , protein_name)]])\n\n    df_path = os.path.join(df_fold_path, f)\n    df = pd.read_csv(df_path)\n\n    df_t = df.query()\n    df_v = df.query()\n\n    t_smiles = df_t[].to_list()\n    v_smiles = df_v[].to_list()\n    df_train = train.query()\n    df_train = df_train.query()\n    df_valid = train.query()\n    df_valid = df_valid.query()\n\n     df_train, df_valid\n\n\ndf_train, df_valid = get_train_valid_df(, , , )\n</code></pre>",
      "rawMarkdown": "In this competition, test data does not contain the building blocks(BBs) which training data has. \n\n>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).\n\nMaking **cv not containing common BBs between training data and validation data** is important to validate our models correctly.\n\nI made the following cv:\n- There is no common BBs between training data and validation data\n- The ratio of our predicting target \"binds\" is almost same between training data and validation data, among each folds.\n\nFollowing is the detail of my cv. `set()` indicate there is no intersection.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7853876%2F83d892e8cd06ff1380dd783752c31254%2Fcv_detail.png?generation=1713464599721993&alt=media)\n\nMy cv is available on kaggle dataset:\n[BELKA-cv-no-common-bbs](https://www.kaggle.com/datasets/yusaku5739/belka-cv-no-common-bbs)\n\nThe notebook creating my cv is comming soon...\n\nIn order to get traning and validation data for each fold, please use following function.\n```\ndef get_train_valid_df(fold, protein_name, dataset_path, df_fold_path):\n    \"\"\"\n    Get train data and valid data for each fold.\n    prams\n    -----------\n    fold: fold you want to get.\n    protein_name: protein name you want to get.\n    dataset_path: directory path to competition dataset. PLEASE USE DEFAULT DATASET. DO NOT USE Shrunken dataset.\n    df_fold_path: directory path to {protein_name}_fold.csv\n    \"\"\"\n    train = pd.read_parquet(os.path.join(dataset_path, \"train_raw.parquet\"), columns=[\"buildingblock1_smiles\", \"buildingblock2_smiles\", \"buildingblock3_smiles\",  \"binds\"], filters=[[(\"protein_name\", \"==\", protein_name)]])\n    \n    df_path = os.path.join(df_fold_path, f\"{protein_name}_fold.csv\")\n    df = pd.read_csv(df_path)\n    \n    df_t = df.query(\"fold!=@fold\")\n    df_v = df.query(\"fold==@fold\")\n\n    t_smiles = df_t[\"smiles\"].to_list()\n    v_smiles = df_v[\"smiles\"].to_list()\n    df_train = train.query(\"buildingblock1_smiles in @t_smiles | buildingblock3_smiles in @t_smiles | buildingblock2_smiles in @t_smiles\")\n    df_train = df_train.query(\"not buildingblock1_smiles in @v_smiles & not buildingblock3_smiles in @v_smiles & not buildingblock2_smiles in @v_smiles\")\n    df_valid = train.query(\"buildingblock1_smiles in @v_smiles | buildingblock3_smiles in @v_smiles | buildingblock2_smiles in @v_smiles\")\n    df_valid = df_valid.query(\"not buildingblock1_smiles in @t_smiles & not buildingblock3_smiles in @t_smiles & not buildingblock2_smiles in @t_smiles\")\n\n    return df_train, df_valid\n\n# get fold0 training data and validation data \ndf_train, df_valid = get_train_valid_df(0, \"sEH\", \"data\", \"data\")\n```",
      "votes": null
    },
    {
      "id": "2760134",
      "postDate": "04/19/2024 06:02:31",
      "content": "<p>Why is it called validation data if we have labels?</p>",
      "rawMarkdown": "Why is it called validation data if we have labels?",
      "votes": null
    },
    {
      "id": "2783575",
      "postDate": "04/29/2024 20:04:28",
      "content": "<p>If you had no labels for the validation set you wouldn't be able to assess your model's performance and generalization ability on unseen data, i.e. validate your model. You make this assessment using a metric function that takes predictions on some data that was not in the training set (the validation set) along with the known labels to calculate a score. The metric used to evaluate your predictions in this competition is <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.average_precision_score.html\" target=\"_blank\">Average Precision</a>, but you can use all kinds of different metrics for the evaluation of your predictions privately if you deem them helpful.</p>",
      "rawMarkdown": "If you had no labels for the validation set you wouldn't be able to assess your model's performance and generalization ability on unseen data, i.e. validate your model. You make this assessment using a metric function that takes predictions on some data that was not in the training set (the validation set) along with the known labels to calculate a score. The metric used to evaluate your predictions in this competition is [Average Precision](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.average_precision_score.html), but you can use all kinds of different metrics for the evaluation of your predictions privately if you deem them helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2760134,
      "author_name": "gyulamaloveczky4",
      "author_url": "",
      "post_date": "04/19/2024 06:02:31",
      "content": "<p>Why is it called validation data if we have labels?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2783575,
          "author_name": "frenio",
          "author_url": "",
          "post_date": "04/29/2024 20:04:28",
          "content": "<p>If you had no labels for the validation set you wouldn't be able to assess your model's performance and generalization ability on unseen data, i.e. validate your model. You make this assessment using a metric function that takes predictions on some data that was not in the training set (the validation set) along with the known labels to calculate a score. The metric used to evaluate your predictions in this competition is <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.average_precision_score.html\" target=\"_blank\">Average Precision</a>, but you can use all kinds of different metrics for the evaluation of your predictions privately if you deem them helpful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2759566": "In this competition, test data does not contain the building blocks(BBs) which training data has. \n\n>All data were generated in-house at Leash Biosciences. We are providing roughly 98M training examples per protein, 200K validation examples per protein, and 360K test molecules per protein. To test generalizability, the test set contains building blocks that are not in the training set. These datasets are very imbalanced: roughly 0.5% of examples are classified as binders; we used 3 rounds of selection in triplicate to identify binders experimentally. Following the competition, Leash will make all the data available for future use (3 targets * 3 rounds of selection * 3 replicates * 133M molecules, or 3.6B measurements).\n\nMaking **cv not containing common BBs between training data and validation data** is important to validate our models correctly.\n\nI made the following cv:\n- There is no common BBs between training data and validation data\n- The ratio of our predicting target \"binds\" is almost same between training data and validation data, among each folds.\n\nFollowing is the detail of my cv. `set()` indicate there is no intersection.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7853876%2F83d892e8cd06ff1380dd783752c31254%2Fcv_detail.png?generation=1713464599721993&alt=media)\n\nMy cv is available on kaggle dataset:\n[BELKA-cv-no-common-bbs](https://www.kaggle.com/datasets/yusaku5739/belka-cv-no-common-bbs)\n\nThe notebook creating my cv is comming soon...\n\nIn order to get traning and validation data for each fold, please use following function.\n```\ndef get_train_valid_df(fold, protein_name, dataset_path, df_fold_path):\n    \"\"\"\n    Get train data and valid data for each fold.\n    prams\n    -----------\n    fold: fold you want to get.\n    protein_name: protein name you want to get.\n    dataset_path: directory path to competition dataset. PLEASE USE DEFAULT DATASET. DO NOT USE Shrunken dataset.\n    df_fold_path: directory path to {protein_name}_fold.csv\n    \"\"\"\n    train = pd.read_parquet(os.path.join(dataset_path, \"train_raw.parquet\"), columns=[\"buildingblock1_smiles\", \"buildingblock2_smiles\", \"buildingblock3_smiles\",  \"binds\"], filters=[[(\"protein_name\", \"==\", protein_name)]])\n    \n    df_path = os.path.join(df_fold_path, f\"{protein_name}_fold.csv\")\n    df = pd.read_csv(df_path)\n    \n    df_t = df.query(\"fold!=@fold\")\n    df_v = df.query(\"fold==@fold\")\n\n    t_smiles = df_t[\"smiles\"].to_list()\n    v_smiles = df_v[\"smiles\"].to_list()\n    df_train = train.query(\"buildingblock1_smiles in @t_smiles | buildingblock3_smiles in @t_smiles | buildingblock2_smiles in @t_smiles\")\n    df_train = df_train.query(\"not buildingblock1_smiles in @v_smiles & not buildingblock3_smiles in @v_smiles & not buildingblock2_smiles in @v_smiles\")\n    df_valid = train.query(\"buildingblock1_smiles in @v_smiles | buildingblock3_smiles in @v_smiles | buildingblock2_smiles in @v_smiles\")\n    df_valid = df_valid.query(\"not buildingblock1_smiles in @t_smiles & not buildingblock3_smiles in @t_smiles & not buildingblock2_smiles in @t_smiles\")\n\n    return df_train, df_valid\n\n# get fold0 training data and validation data \ndf_train, df_valid = get_train_valid_df(0, \"sEH\", \"data\", \"data\")\n```",
    "2760134": "Why is it called validation data if we have labels?",
    "2783575": "If you had no labels for the validation set you wouldn't be able to assess your model's performance and generalization ability on unseen data, i.e. validate your model. You make this assessment using a metric function that takes predictions on some data that was not in the training set (the validation set) along with the known labels to calculate a score. The metric used to evaluate your predictions in this competition is [Average Precision](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.average_precision_score.html), but you can use all kinds of different metrics for the evaluation of your predictions privately if you deem them helpful."
  },
  "source": "meta"
}