{
  "id": 656491,
  "title": "\"Example Baseline Predictor using Code Template\" give ValueError when to be tried with k = 2 to generate 2-mers",
  "url": "/competitions/adaptive-immune-profiling-challenge-2025/discussion/656491",
  "author_name": "",
  "post_date": "2025-12-09T12:51:57.329624800Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I tried k = 1, k = 3 and k = 4 cases successfully but case k = 2 gives following error. Do you have an idea or fix for it. Thanks.</p>\n<p>Fitting model on examples in <code>/kaggle/input/adaptive-immune-profiling-challenge-2025/train_datasets/train_datasets/train_dataset_1</code>…</p>\n<h2>Encoding 2-mers: 100%|██████████| 400/400 [00:58&lt;00:00,  6.90it/s]</h2>\n<p>ValueError                                Traceback (most recent call last)\n/tmp/ipykernel_38/3335476791.py in ()\n      6 \n      7 for train_dir, test_dirs in train_test_dataset_pairs:\n----&gt; 8     main(train_dir=train_dir, test_dirs=test_dirs, out_dir=results_dir, n_jobs=4, device=\"cpu\")\n      9 \n     10 concatenate_output_files(results_dir)</p>\n<p>/tmp/ipykernel_38/300531143.py in main(train_dir, test_dirs, out_dir, n_jobs, device)\n     48     predictor = ImmuneStatePredictor(n_jobs=n_jobs,\n     49                                      device=device)  # instantiate with any other parameters as defined by you in the class\n---&gt; 50     _train_predictor(predictor, train_dir)\n     51     predictions = _generate_predictions(predictor, test_dirs)\n     52     _save_predictions(predictions, out_dir, train_dir)</p>\n<p>/tmp/ipykernel_38/300531143.py in _train_predictor(predictor, train_dir)\n      5     \"\"\"Trains the predictor on the training data.\"\"\"\n      6     print(f\"Fitting model on examples in <code>{train_dir}</code>…\")\n----&gt; 7     predictor.fit(train_dir)\n      8 \n      9 </p>\n<p>/tmp/ipykernel_38/643916291.py in fit(self, train_dir_path)\n     65         #    Example: self.model = SomeClassifier().fit(X_train, y_train)\n     66 \n---&gt; 67         X_train, y_train, train_ids = prepare_data(X_train_df, y_train_df,\n     68                                                    id_col='ID', label_col='label_positive')\n     69 </p>\n<p>/tmp/ipykernel_38/3268145801.py in prepare_data(X_df, labels_df, id_col, label_col)\n    196 \n    197     if len(common_ids) == 0:\n--&gt; 198         raise ValueError(\"No common IDs found between feature matrix and labels\")\n    199 \n    200     X = X_df.loc[common_ids]</p>\n<p>ValueError: No common IDs found between feature matrix and labels</p>",
  "messages": [
    {
      "id": "3368713",
      "postDate": "12/09/2025 12:51:57",
      "content": "<p>I tried k = 1, k = 3 and k = 4 cases successfully but case k = 2 gives following error. Do you have an idea or fix for it. Thanks.</p>\n<p>Fitting model on examples in <code>/kaggle/input/adaptive-immune-profiling-challenge-2025/train_datasets/train_datasets/train_dataset_1</code>…</p>\n<h2>Encoding 2-mers: 100%|██████████| 400/400 [00:58&lt;00:00,  6.90it/s]</h2>\n<p>ValueError                                Traceback (most recent call last)\n/tmp/ipykernel_38/3335476791.py in ()\n      6 \n      7 for train_dir, test_dirs in train_test_dataset_pairs:\n----&gt; 8     main(train_dir=train_dir, test_dirs=test_dirs, out_dir=results_dir, n_jobs=4, device=\"cpu\")\n      9 \n     10 concatenate_output_files(results_dir)</p>\n<p>/tmp/ipykernel_38/300531143.py in main(train_dir, test_dirs, out_dir, n_jobs, device)\n     48     predictor = ImmuneStatePredictor(n_jobs=n_jobs,\n     49                                      device=device)  # instantiate with any other parameters as defined by you in the class\n---&gt; 50     _train_predictor(predictor, train_dir)\n     51     predictions = _generate_predictions(predictor, test_dirs)\n     52     _save_predictions(predictions, out_dir, train_dir)</p>\n<p>/tmp/ipykernel_38/300531143.py in _train_predictor(predictor, train_dir)\n      5     \"\"\"Trains the predictor on the training data.\"\"\"\n      6     print(f\"Fitting model on examples in <code>{train_dir}</code>…\")\n----&gt; 7     predictor.fit(train_dir)\n      8 \n      9 </p>\n<p>/tmp/ipykernel_38/643916291.py in fit(self, train_dir_path)\n     65         #    Example: self.model = SomeClassifier().fit(X_train, y_train)\n     66 \n---&gt; 67         X_train, y_train, train_ids = prepare_data(X_train_df, y_train_df,\n     68                                                    id_col='ID', label_col='label_positive')\n     69 </p>\n<p>/tmp/ipykernel_38/3268145801.py in prepare_data(X_df, labels_df, id_col, label_col)\n    196 \n    197     if len(common_ids) == 0:\n--&gt; 198         raise ValueError(\"No common IDs found between feature matrix and labels\")\n    199 \n    200     X = X_df.loc[common_ids]</p>\n<p>ValueError: No common IDs found between feature matrix and labels</p>",
      "rawMarkdown": "I tried k = 1, k = 3 and k = 4 cases successfully but case k = 2 gives following error. Do you have an idea or fix for it. Thanks.\n\nFitting model on examples in ` /kaggle/input/adaptive-immune-profiling-challenge-2025/train_datasets/train_datasets/train_dataset_1 `...\nEncoding 2-mers: 100%|██████████| 400/400 [00:58<00:00,  6.90it/s]\n---------------------------------------------------------------------------\nValueError                                Traceback (most recent call last)\n/tmp/ipykernel_38/3335476791.py in <cell line: 0>()\n      6 \n      7 for train_dir, test_dirs in train_test_dataset_pairs:\n----> 8     main(train_dir=train_dir, test_dirs=test_dirs, out_dir=results_dir, n_jobs=4, device=\"cpu\")\n      9 \n     10 concatenate_output_files(results_dir)\n\n/tmp/ipykernel_38/300531143.py in main(train_dir, test_dirs, out_dir, n_jobs, device)\n     48     predictor = ImmuneStatePredictor(n_jobs=n_jobs,\n     49                                      device=device)  # instantiate with any other parameters as defined by you in the class\n---> 50     _train_predictor(predictor, train_dir)\n     51     predictions = _generate_predictions(predictor, test_dirs)\n     52     _save_predictions(predictions, out_dir, train_dir)\n\n/tmp/ipykernel_38/300531143.py in _train_predictor(predictor, train_dir)\n      5     \"\"\"Trains the predictor on the training data.\"\"\"\n      6     print(f\"Fitting model on examples in ` {train_dir} `...\")\n----> 7     predictor.fit(train_dir)\n      8 \n      9 \n\n/tmp/ipykernel_38/643916291.py in fit(self, train_dir_path)\n     65         #    Example: self.model = SomeClassifier().fit(X_train, y_train)\n     66 \n---> 67         X_train, y_train, train_ids = prepare_data(X_train_df, y_train_df,\n     68                                                    id_col='ID', label_col='label_positive')\n     69 \n\n/tmp/ipykernel_38/3268145801.py in prepare_data(X_df, labels_df, id_col, label_col)\n    196 \n    197     if len(common_ids) == 0:\n--> 198         raise ValueError(\"No common IDs found between feature matrix and labels\")\n    199 \n    200     X = X_df.loc[common_ids]\n\nValueError: No common IDs found between feature matrix and labels",
      "votes": null
    },
    {
      "id": "3368871",
      "postDate": "12/09/2025 14:59:46",
      "content": "<p>Ah, I see why that happens in the 2-mer case 😅. It is a silly oversight/bug in this line </p>\n<p><code>features_df = pd.DataFrame(repertoire_features).fillna(0).set_index('ID')</code> in <code>load_and_encode_kmers</code>. Since 'ID' can also be one of the possible 2-mer, its count replaces the repertoire ID because of the bad naming choice. Since this example notebook is only intended as an illustration to get everyone to some speed, I will not prioritise fixing this now -- but feel free to fix it yourself in your code if you choose to encode as 2-mers. Good find 🎉</p>",
      "rawMarkdown": "Ah, I see why that happens in the 2-mer case 😅. It is a silly oversight/bug in this line \n\n`features_df = pd.DataFrame(repertoire_features).fillna(0).set_index('ID')` in `load_and_encode_kmers`. Since 'ID' can also be one of the possible 2-mer, its count replaces the repertoire ID because of the bad naming choice. Since this example notebook is only intended as an illustration to get everyone to some speed, I will not prioritise fixing this now -- but feel free to fix it yourself in your code if you choose to encode as 2-mers. Good find 🎉",
      "votes": null
    },
    {
      "id": "3369594",
      "postDate": "12/10/2025 07:07:28",
      "content": "<p>Thank you. </p>",
      "rawMarkdown": "Thank you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3368871,
      "author_name": "ckanduri",
      "author_url": "",
      "post_date": "12/09/2025 14:59:46",
      "content": "<p>Ah, I see why that happens in the 2-mer case 😅. It is a silly oversight/bug in this line </p>\n<p><code>features_df = pd.DataFrame(repertoire_features).fillna(0).set_index('ID')</code> in <code>load_and_encode_kmers</code>. Since 'ID' can also be one of the possible 2-mer, its count replaces the repertoire ID because of the bad naming choice. Since this example notebook is only intended as an illustration to get everyone to some speed, I will not prioritise fixing this now -- but feel free to fix it yourself in your code if you choose to encode as 2-mers. Good find 🎉</p>",
      "votes": null,
      "replies": [
        {
          "id": 3369594,
          "author_name": "kemalsayar",
          "author_url": "",
          "post_date": "12/10/2025 07:07:28",
          "content": "<p>Thank you. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3368713": "I tried k = 1, k = 3 and k = 4 cases successfully but case k = 2 gives following error. Do you have an idea or fix for it. Thanks.\n\nFitting model on examples in ` /kaggle/input/adaptive-immune-profiling-challenge-2025/train_datasets/train_datasets/train_dataset_1 `...\nEncoding 2-mers: 100%|██████████| 400/400 [00:58<00:00,  6.90it/s]\n---------------------------------------------------------------------------\nValueError                                Traceback (most recent call last)\n/tmp/ipykernel_38/3335476791.py in <cell line: 0>()\n      6 \n      7 for train_dir, test_dirs in train_test_dataset_pairs:\n----> 8     main(train_dir=train_dir, test_dirs=test_dirs, out_dir=results_dir, n_jobs=4, device=\"cpu\")\n      9 \n     10 concatenate_output_files(results_dir)\n\n/tmp/ipykernel_38/300531143.py in main(train_dir, test_dirs, out_dir, n_jobs, device)\n     48     predictor = ImmuneStatePredictor(n_jobs=n_jobs,\n     49                                      device=device)  # instantiate with any other parameters as defined by you in the class\n---> 50     _train_predictor(predictor, train_dir)\n     51     predictions = _generate_predictions(predictor, test_dirs)\n     52     _save_predictions(predictions, out_dir, train_dir)\n\n/tmp/ipykernel_38/300531143.py in _train_predictor(predictor, train_dir)\n      5     \"\"\"Trains the predictor on the training data.\"\"\"\n      6     print(f\"Fitting model on examples in ` {train_dir} `...\")\n----> 7     predictor.fit(train_dir)\n      8 \n      9 \n\n/tmp/ipykernel_38/643916291.py in fit(self, train_dir_path)\n     65         #    Example: self.model = SomeClassifier().fit(X_train, y_train)\n     66 \n---> 67         X_train, y_train, train_ids = prepare_data(X_train_df, y_train_df,\n     68                                                    id_col='ID', label_col='label_positive')\n     69 \n\n/tmp/ipykernel_38/3268145801.py in prepare_data(X_df, labels_df, id_col, label_col)\n    196 \n    197     if len(common_ids) == 0:\n--> 198         raise ValueError(\"No common IDs found between feature matrix and labels\")\n    199 \n    200     X = X_df.loc[common_ids]\n\nValueError: No common IDs found between feature matrix and labels",
    "3368871": "Ah, I see why that happens in the 2-mer case 😅. It is a silly oversight/bug in this line \n\n`features_df = pd.DataFrame(repertoire_features).fillna(0).set_index('ID')` in `load_and_encode_kmers`. Since 'ID' can also be one of the possible 2-mer, its count replaces the repertoire ID because of the bad naming choice. Since this example notebook is only intended as an illustration to get everyone to some speed, I will not prioritise fixing this now -- but feel free to fix it yourself in your code if you choose to encode as 2-mers. Good find 🎉",
    "3369594": "Thank you."
  },
  "source": "meta"
}