{
  "id": 366961,
  "title": "1st Place Solution Summary",
  "url": "/competitions/open-problems-multimodal/discussion/366961",
  "author_name": "Shuji Suzuki",
  "post_date": "2022-11-18T13:14:22.760000",
  "votes": 104,
  "comment_count": 21,
  "views": 0,
  "content": "<p>First of all, thank you to the organizers and kaggle management and to everyone who participated with me.<br>\nSince I needed to gain experience in analyzing single cell data, this competition was an excellent experience for me.</p>\n<p>I would like to introduce the overview of my solution.</p>\n<h1>Multiome</h1>\n<h2>Model Overview</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F05d87da294f71eca450a304785333ac7%2Fmult-model-overview.png?generation=1668772959634455&amp;alt=media\" alt=\"\"></p>\n<h2>Input Preprocessing</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fdedb439f946e0682644451a531d4df33%2Fmulti-input-preprocessing.png?generation=1668773118728830&amp;alt=media\" alt=\"\"></p>\n<h2>Target Preprocessing</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F76231a7e2c23f3f588e763bf8035dcd0%2Fmulti-target-preprocessing.png?generation=1669418912414786&amp;alt=media\" alt=\"\"></p>\n<p>tSVD-based imputation method: </p>\n<ol>\n<li>Perform dimensionality reduction on the data with tSVD</li>\n<li>And then, Transform the data back to the original space</li>\n<li>Copy the value of the 0 part of the original data from the transformed values.</li>\n</ol>\n<h2>Model</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F8861945be57ca34a6bfa2b23e5627100%2Fmulti-model.png?generation=1668773599707773&amp;alt=media\" alt=\"\"></p>\n<h2>Output Postprocessing and Loss</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Ff59241b0358a15294bdc38c25722f37d%2Fmulti-postprocessing_1.png?generation=1669418940293447&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fbff80fd72fd2a9d0ab94704c641c024b%2Fmulti-postprocessing_2.png?generation=1669419002523851&amp;alt=media\" alt=\"\"></p>\n<p>In the inference phase, the model outputs the average of the five predicted target data.</p>\n<h1>CITEseq</h1>\n<h2>Model Overview</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fa9563a70e5639db2b4b72b7645ecd910%2Fcite-model-overview.png?generation=1668774675355569&amp;alt=media\" alt=\"\"></p>\n<h2>Input Preprocessing</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Ff5ddef7a059aa3a94c3afc6f3f52b440%2Fcite-input-preprocessing.png?generation=1669419233099917&amp;alt=media\" alt=\"\"></p>\n<p>In selecting important genes in CITEseq, the correlation coefficient is calculated for each batch and select only genes with high correlation in many batches.<br>\nGenes were selected from those related to the target proteins and pathway.<br>\nI use <a href=\"https://reactome.org/\" target=\"_blank\">Reactome</a> as pathway database.</p>\n<h2>Target Preprocessing</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fc538366bbc83a1de46005c5108d6b20d%2Fcite-target-preprocessing.png?generation=1668774972017361&amp;alt=media\" alt=\"\"></p>\n<h2>Model</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F3a30e371014051289d99d5cbe57009e9%2Fcite-model.png?generation=1668775048364648&amp;alt=media\" alt=\"\"></p>\n<h2>Output Postprocessing and Loss</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F677385c92a3be6a0e0a9702b3005edc8%2Fcite-postprocessing.png?generation=1668775162024292&amp;alt=media\" alt=\"\"></p>\n<p>In the inference phase, the model outputs the average of the five predicted target data.</p>\n<h1>Local evaluation</h1>\n<p>I used two evaluation schemes.</p>\n<ol>\n<li>Evaluation with cross validation:<ul>\n<li>5-fold cross validation grouped by donor and day</li></ul></li>\n<li>Evaluation for hyperparameter optimization with Optuna:<ul>\n<li>Training data set is divided into training and validation data sets. ( Training data set: 80%, validation data set: 20%. )</li></ul></li>\n</ol>\n<h1>Ensemble</h1>\n<p>I used the weighted average of predictions of the following models.</p>\n<ol>\n<li>Models trained with changing the seed </li>\n<li>Models fine-tuned on only some batches<ul>\n<li>Batch combination pattern examples: males only, female only, Day 4, 7 only, etc.</li>\n<li>Use a model trained on the full training data set as a pre-training model </li></ul></li>\n</ol>\n<h1>Code</h1>\n<p><a href=\"https://github.com/shu65/open-problems-multimodal\" target=\"_blank\">https://github.com/shu65/open-problems-multimodal</a></p>\n<h1>Update</h1>\n<p>2022/11/20 add the repository url of my solution<br>\n2022/11/26 fix some figures</p>",
  "messages": [
    {
      "id": 2034844,
      "postDate": "2022-11-18T13:14:22.760Z",
      "content": "<p>First of all, thank you to the organizers and kaggle management and to everyone who participated with me.<br>\nSince I needed to gain experience in analyzing single cell data, this competition was an excellent experience for me.</p>\n<p>I would like to introduce the overview of my solution.</p>\n<h1>Multiome</h1>\n<h2>Model Overview</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F05d87da294f71eca450a304785333ac7%2Fmult-model-overview.png?generation=1668772959634455&amp;alt=media\" alt=\"\"></p>\n<h2>Input Preprocessing</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fdedb439f946e0682644451a531d4df33%2Fmulti-input-preprocessing.png?generation=1668773118728830&amp;alt=media\" alt=\"\"></p>\n<h2>Target Preprocessing</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F76231a7e2c23f3f588e763bf8035dcd0%2Fmulti-target-preprocessing.png?generation=1669418912414786&amp;alt=media\" alt=\"\"></p>\n<p>tSVD-based imputation method: </p>\n<ol>\n<li>Perform dimensionality reduction on the data with tSVD</li>\n<li>And then, Transform the data back to the original space</li>\n<li>Copy the value of the 0 part of the original data from the transformed values.</li>\n</ol>\n<h2>Model</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F8861945be57ca34a6bfa2b23e5627100%2Fmulti-model.png?generation=1668773599707773&amp;alt=media\" alt=\"\"></p>\n<h2>Output Postprocessing and Loss</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Ff59241b0358a15294bdc38c25722f37d%2Fmulti-postprocessing_1.png?generation=1669418940293447&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fbff80fd72fd2a9d0ab94704c641c024b%2Fmulti-postprocessing_2.png?generation=1669419002523851&amp;alt=media\" alt=\"\"></p>\n<p>In the inference phase, the model outputs the average of the five predicted target data.</p>\n<h1>CITEseq</h1>\n<h2>Model Overview</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fa9563a70e5639db2b4b72b7645ecd910%2Fcite-model-overview.png?generation=1668774675355569&amp;alt=media\" alt=\"\"></p>\n<h2>Input Preprocessing</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Ff5ddef7a059aa3a94c3afc6f3f52b440%2Fcite-input-preprocessing.png?generation=1669419233099917&amp;alt=media\" alt=\"\"></p>\n<p>In selecting important genes in CITEseq, the correlation coefficient is calculated for each batch and select only genes with high correlation in many batches.<br>\nGenes were selected from those related to the target proteins and pathway.<br>\nI use <a href=\"https://reactome.org/\" target=\"_blank\">Reactome</a> as pathway database.</p>\n<h2>Target Preprocessing</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fc538366bbc83a1de46005c5108d6b20d%2Fcite-target-preprocessing.png?generation=1668774972017361&amp;alt=media\" alt=\"\"></p>\n<h2>Model</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F3a30e371014051289d99d5cbe57009e9%2Fcite-model.png?generation=1668775048364648&amp;alt=media\" alt=\"\"></p>\n<h2>Output Postprocessing and Loss</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F677385c92a3be6a0e0a9702b3005edc8%2Fcite-postprocessing.png?generation=1668775162024292&amp;alt=media\" alt=\"\"></p>\n<p>In the inference phase, the model outputs the average of the five predicted target data.</p>\n<h1>Local evaluation</h1>\n<p>I used two evaluation schemes.</p>\n<ol>\n<li>Evaluation with cross validation:<ul>\n<li>5-fold cross validation grouped by donor and day</li></ul></li>\n<li>Evaluation for hyperparameter optimization with Optuna:<ul>\n<li>Training data set is divided into training and validation data sets. ( Training data set: 80%, validation data set: 20%. )</li></ul></li>\n</ol>\n<h1>Ensemble</h1>\n<p>I used the weighted average of predictions of the following models.</p>\n<ol>\n<li>Models trained with changing the seed </li>\n<li>Models fine-tuned on only some batches<ul>\n<li>Batch combination pattern examples: males only, female only, Day 4, 7 only, etc.</li>\n<li>Use a model trained on the full training data set as a pre-training model </li></ul></li>\n</ol>\n<h1>Code</h1>\n<p><a href=\"https://github.com/shu65/open-problems-multimodal\" target=\"_blank\">https://github.com/shu65/open-problems-multimodal</a></p>\n<h1>Update</h1>\n<p>2022/11/20 add the repository url of my solution<br>\n2022/11/26 fix some figures</p>",
      "rawMarkdown": "First of all, thank you to the organizers and kaggle management and to everyone who participated with me.\nSince I needed to gain experience in analyzing single cell data, this competition was an excellent experience for me.\n\nI would like to introduce the overview of my solution.\n\n# Multiome\n## Model Overview\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F05d87da294f71eca450a304785333ac7%2Fmult-model-overview.png?generation=1668772959634455&alt=media)\n\n## Input Preprocessing\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fdedb439f946e0682644451a531d4df33%2Fmulti-input-preprocessing.png?generation=1668773118728830&alt=media)\n\n## Target Preprocessing\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F76231a7e2c23f3f588e763bf8035dcd0%2Fmulti-target-preprocessing.png?generation=1669418912414786&alt=media)\n\ntSVD-based imputation method: \n1. Perform dimensionality reduction on the data with tSVD\n2. And then, Transform the data back to the original space\n3. Copy the value of the 0 part of the original data from the transformed values.\n\n## Model\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F8861945be57ca34a6bfa2b23e5627100%2Fmulti-model.png?generation=1668773599707773&alt=media)\n\n## Output Postprocessing and Loss\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Ff59241b0358a15294bdc38c25722f37d%2Fmulti-postprocessing_1.png?generation=1669418940293447&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fbff80fd72fd2a9d0ab94704c641c024b%2Fmulti-postprocessing_2.png?generation=1669419002523851&alt=media)\n\nIn the inference phase, the model outputs the average of the five predicted target data.\n\n# CITEseq\n## Model Overview\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fa9563a70e5639db2b4b72b7645ecd910%2Fcite-model-overview.png?generation=1668774675355569&alt=media)\n\n## Input Preprocessing\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Ff5ddef7a059aa3a94c3afc6f3f52b440%2Fcite-input-preprocessing.png?generation=1669419233099917&alt=media)\n\nIn selecting important genes in CITEseq, the correlation coefficient is calculated for each batch and select only genes with high correlation in many batches.\nGenes were selected from those related to the target proteins and pathway.\nI use [Reactome](https://reactome.org/) as pathway database.\n\n## Target Preprocessing\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fc538366bbc83a1de46005c5108d6b20d%2Fcite-target-preprocessing.png?generation=1668774972017361&alt=media)\n\n## Model\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F3a30e371014051289d99d5cbe57009e9%2Fcite-model.png?generation=1668775048364648&alt=media)\n\n\n## Output Postprocessing and Loss\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F677385c92a3be6a0e0a9702b3005edc8%2Fcite-postprocessing.png?generation=1668775162024292&alt=media)\n\nIn the inference phase, the model outputs the average of the five predicted target data.\n\n# Local evaluation\nI used two evaluation schemes.\n\n1. Evaluation with cross validation:\n  * 5-fold cross validation grouped by donor and day\n2. Evaluation for hyperparameter optimization with Optuna:\n  * Training data set is divided into training and validation data sets. ( Training data set: 80%, validation data set: 20%. )\n\n# Ensemble\nI used the weighted average of predictions of the following models.\n1. Models trained with changing the seed \n2. Models fine-tuned on only some batches\n  * Batch combination pattern examples: males only, female only, Day 4, 7 only, etc.\n  * Use a model trained on the full training data set as a pre-training model \n\n# Code\nhttps://github.com/shu65/open-problems-multimodal\n\n\n# Update\n2022/11/20 add the repository url of my solution\n2022/11/26 fix some figures",
      "votes": 104
    },
    {
      "id": 2041633,
      "postDate": "2022-11-24T05:15:32.717Z",
      "content": "<p>Congratulations and thanks for posting your great solution Shuji Suzuki!</p>",
      "rawMarkdown": "Congratulations and thanks for posting your great solution Shuji Suzuki!",
      "votes": 1
    },
    {
      "id": 2046907,
      "postDate": "2022-11-28T12:53:04.240Z",
      "content": "<p>Big old 🙏🙏🙏 <br>\nMay i ask some questions emmm<br>\nThe part of multi's model is fantastic, why  not  apply to Cite problems,  by residual target this way, it is not work in the cite problems?</p>",
      "rawMarkdown": "Big old 🙏🙏🙏 \nMay i ask some questions emmm\nThe part of multi's model is fantastic, why  not  apply to Cite problems,  by residual target this way, it is not work in the cite problems?",
      "votes": 2
    },
    {
      "id": 2038496,
      "postDate": "2022-11-21T12:36:55.080Z",
      "content": "<p>Congratulations for thq 1st place and thank you for sharing your solution! <br>\nAs I see, you ve added metadata to both models. What increase (cv/ lb) did you get from adding metadata to models? </p>",
      "rawMarkdown": "Congratulations for thq 1st place and thank you for sharing your solution! \nAs I see, you ve added metadata to both models. What increase (cv/ lb) did you get from adding metadata to models? ",
      "votes": 2
    },
    {
      "id": 2036163,
      "postDate": "2022-11-19T14:09:14.427Z",
      "content": "<p>Congratulations for the 1st position and great solution!!</p>\n<p>I tried combining multiple forms of inputs/targets (initial inputs, SVD and binarized) and weight averaging their losses, but dropped it because the added complexity wasn't helping. I missed the ingenuity of your preprocessing chain and the averaging of the outputs for prediction (I took only the prediction of the original targets, and used the other outputs solely to regularize training).</p>\n<p>I have seen big PB jumps to gold medal positions - <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> comes to mind, but I don't recall ever seeing such an impressive finish in a winner. What a great job! I'm curious about how surprised you where with taking the top position after ending the competition in a position far from the gold medals in the LB.</p>",
      "rawMarkdown": "Congratulations for the 1st position and great solution!!\n\nI tried combining multiple forms of inputs/targets (initial inputs, SVD and binarized) and weight averaging their losses, but dropped it because the added complexity wasn't helping. I missed the ingenuity of your preprocessing chain and the averaging of the outputs for prediction (I took only the prediction of the original targets, and used the other outputs solely to regularize training).\n\nI have seen big PB jumps to gold medal positions - @cpmpml comes to mind, but I don't recall ever seeing such an impressive finish in a winner. What a great job! I'm curious about how surprised you where with taking the top position after ending the competition in a position far from the gold medals in the LB.",
      "votes": 2,
      "replies": [
        {
          "id": 2036235,
          "postDate": "2022-11-19T15:44:46.487Z",
          "content": "<p>You jumped quite a bit too.</p>",
          "rawMarkdown": "You jumped quite a bit too.",
          "votes": 3
        },
        {
          "id": 2036927,
          "postDate": "2022-11-20T09:44:46.323Z",
          "content": "<p>Thank you! </p>\n<p>I couldn't believe it at first either and showed the screenshot of LB to my colleagues at work to make sure I was in first place and not wrong.</p>",
          "rawMarkdown": "Thank you! \n\nI couldn't believe it at first either and showed the screenshot of LB to my colleagues at work to make sure I was in first place and not wrong.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2035038,
      "postDate": "2022-11-18T14:52:28.227Z",
      "content": "<p>One of the best posts I have seen</p>",
      "rawMarkdown": "One of the best posts I have seen"
    },
    {
      "id": 2035261,
      "postDate": "2022-11-18T18:15:49.657Z",
      "content": "<p>Thank you for this detailed post.<br>\nI see you made a lot of work preprocessing the data. But how did you come to this sort of preprocessing? Where does the idea to divide all the inputs by non-zero medians come from?</p>",
      "rawMarkdown": "Thank you for this detailed post.\nI see you made a lot of work preprocessing the data. But how did you come to this sort of preprocessing? Where does the idea to divide all the inputs by non-zero medians come from?",
      "votes": 1,
      "replies": [
        {
          "id": 2036930,
          "postDate": "2022-11-20T09:55:49.183Z",
          "content": "<p>I noticed that the library-size normalization + log1p data is not being returned well after dimensionality compression with tSVD. The correlation coefficient between the original data and the converted and reverted data is only about 0.70.</p>\n<p>For this reason, I tried to find a better preprocessing method and found it.</p>",
          "rawMarkdown": "I noticed that the library-size normalization + log1p data is not being returned well after dimensionality compression with tSVD. The correlation coefficient between the original data and the converted and reverted data is only about 0.70.\n\nFor this reason, I tried to find a better preprocessing method and found it.",
          "votes": 7
        },
        {
          "id": 2042059,
          "postDate": "2022-11-24T12:08:58.813Z",
          "content": "<p>Dividing by non-zero medians is something I've never seen in single-cell literature, but hey, it WORKS!<br>\nHow do you get to this approach? Could you elaborate on your thought process? Thanks and congratulations!</p>",
          "rawMarkdown": "Dividing by non-zero medians is something I've never seen in single-cell literature, but hey, it WORKS!\nHow do you get to this approach? Could you elaborate on your thought process? Thanks and congratulations!",
          "votes": 1
        },
        {
          "id": 2058378,
          "postDate": "2022-12-07T21:29:54.460Z",
          "content": "<p><a href=\"https://www.kaggle.com/shujisuzuki65\" target=\"_blank\">@shujisuzuki65</a> am I right in understanding you looked at many different pre-processing approaches and optimized for one that had high correlation after low-rank-approximation with tSVD? This led to selecting normalizing by median non-zero values?</p>",
          "rawMarkdown": "@shujisuzuki65 am I right in understanding you looked at many different pre-processing approaches and optimized for one that had high correlation after low-rank-approximation with tSVD? This led to selecting normalizing by median non-zero values?",
          "votes": 1
        }
      ]
    },
    {
      "id": 2128595,
      "postDate": "2023-02-03T21:02:48.433Z",
      "content": "<p>Hello Shuji. I tried to run your code with the minor modification that I only use 996 data points for both the train and test set to speed up the calculations. I get this output/error when training the cite model and don't know what to do. The error persists even when changing batch-size to 1000, for some reason it always looks for 8 batches and can't find 6 of them: <br>\ngit_hexsha Failed to get git<br>\nload input values<br>\ncompleted loading input values. elapsed time: 4.2<br>\nload targets values<br>\ncompleted loading targets values. elapsed time: 0.2<br>\nload input values<br>\ncompleted loading input values. elapsed time: 1.2<br>\nuse test_inputs_values. total size: 1992<br>\ntrain sample size: 996<br>\ndump params<br>\nkfold type  group None<br>\nskip pre_post_process fit<br>\nmodel input shape X:(664, 217) Y:(664, 128)<br>\ndataset size 664</p>\n<hr>\n<p>KeyError                                  Traceback (most recent call last)</p>\n<p>/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/script/train_model.py in <br>\n    570 <br>\n    571 if <strong>name</strong> == \"<strong>main</strong>\":<br>\n--&gt; 572     main()</p>\n<p>6 frames</p>\n<p>/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/script/train_model.py in main()<br>\n    475 <br>\n    476     cv = CrossVaridation()<br>\n--&gt; 477     result_df, k_fold_models, k_fold_pre_post_processes = cv.compute_score(<br>\n    478         x=train_inputs,<br>\n    479         y=train_target,</p>\n<p>/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/script/train_model.py in compute_score(self, x, y, metadata, x_test, metadata_test, params, build_model, build_pre_post_process, dump, dump_dir, n_splits, n_bagging, bagging_ratio, use_batch_group)<br>\n    145                 model = build_model(params=params[\"model\"])<br>\n    146                 print(f\"model input shape X:{preprocessed_x_train.shape} Y:{preprocessed_y_train.shape}\")<br>\n--&gt; 147                 model.fit(<br>\n    148                     x=x_train_bagging,<br>\n    149                     y=y_train_bagging,</p>\n<p>/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/ss_opm/model/encoder_decoder/encoder_decoder.py in fit(self, x, preprocessed_x, y, preprocessed_y, metadata, pre_post_process)<br>\n    221         self.model.to(device=self.params[\"device\"])<br>\n    222 <br>\n--&gt; 223         dummy_batch = next(iter(data_loader))<br>\n    224         dummy_batch = self._batch_to_device(dummy_batch)<br>\n    225         self._train_step_forward(dummy_batch, 1.0)</p>\n<p>/usr/local/lib/python3.8/dist-packages/torch/utils/data/dataloader.py in <strong>next</strong>(self)<br>\n    626                 # TODO(<a href=\"https://github.com/pytorch/pytorch/issues/76750\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/76750</a>)<br>\n    627                 self._reset()  # type: ignore[call-arg]<br>\n--&gt; 628             data = self._next_data()<br>\n    629             self._num_yielded += 1<br>\n    630             if self._dataset_kind == _DatasetKind.Iterable and \\</p>\n<p>/usr/local/lib/python3.8/dist-packages/torch/utils/data/dataloader.py in _next_data(self)<br>\n   1331             else:<br>\n   1332                 del self._task_info[idx]<br>\n-&gt; 1333                 return self._process_data(data)<br>\n   1334 <br>\n   1335     def _try_put_index(self):</p>\n<p>/usr/local/lib/python3.8/dist-packages/torch/utils/data/dataloader.py in _process_data(self, data)<br>\n   1357         self._try_put_index()<br>\n   1358         if isinstance(data, ExceptionWrapper):<br>\n-&gt; 1359             data.reraise()<br>\n   1360         return data<br>\n   1361 </p>\n<p>/usr/local/lib/python3.8/dist-packages/torch/_utils.py in reraise(self)<br>\n    541             # instantiate since we don't know how to<br>\n    542             raise RuntimeError(msg) from None<br>\n--&gt; 543         raise exception<br>\n    544 <br>\n    545 </p>\n<p>KeyError: Caught KeyError in DataLoader worker process 0.<br>\nOriginal Traceback (most recent call last):<br>\n  File \"/usr/local/lib/python3.8/dist-packages/torch/utils/data/_utils/worker.py\", line 302, in _worker_loop<br>\n    data = fetcher.fetch(index)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/torch/utils/data/_utils/fetch.py\", line 58, in fetch<br>\n    data = [self.dataset[idx] for idx in possibly_batched_index]<br>\n  File \"/usr/local/lib/python3.8/dist-packages/torch/utils/data/_utils/fetch.py\", line 58, in <br>\n    data = [self.dataset[idx] for idx in possibly_batched_index]<br>\n  File \"/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/ss_opm/model/torch_dataset/citeseq_dataset.py\", line 79, in <strong>getitem</strong><br>\n    info = torch.as_tensor(self.metadata.iloc[index, :][self.metadata_keys].values.astype(float), dtype=torch.float32)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/series.py\", line 1007, in <strong>getitem</strong><br>\n    return self._get_with(key)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/series.py\", line 1047, in _get_with<br>\n    return self.loc[key]<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1073, in <strong>getitem</strong><br>\n    return self._getitem_axis(maybe_callable, axis=axis)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1301, in _getitem_axis<br>\n    return self._getitem_iterable(key, axis=axis)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1239, in _getitem_iterable<br>\n    keyarr, indexer = self._get_listlike_indexer(key, axis)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1432, in _get_listlike_indexer<br>\n    keyarr, indexer = ax._get_indexer_strict(key, axis_name)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexes/base.py\", line 6070, in _get_indexer_strict<br>\n    self._raise_if_missing(keyarr, indexer, axis_name)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexes/base.py\", line 6133, in _raise_if_missing<br>\n    raise KeyError(f\"{not_found} not in index\")<br>\nKeyError: \"['batch_sv2', 'batch_sv3', 'batch_sv4', 'batch_sv5', 'batch_sv6', 'batch_sv7'] not in index\"</p>",
      "rawMarkdown": "Hello Shuji. I tried to run your code with the minor modification that I only use 996 data points for both the train and test set to speed up the calculations. I get this output/error when training the cite model and don't know what to do. The error persists even when changing batch-size to 1000, for some reason it always looks for 8 batches and can't find 6 of them: \ngit_hexsha Failed to get git\nload input values\ncompleted loading input values. elapsed time: 4.2\nload targets values\ncompleted loading targets values. elapsed time: 0.2\nload input values\ncompleted loading input values. elapsed time: 1.2\nuse test_inputs_values. total size: 1992\ntrain sample size: 996\ndump params\nkfold type <class 'sklearn.model_selection._split.KFold'> group None\nskip pre_post_process fit\nmodel input shape X:(664, 217) Y:(664, 128)\ndataset size 664\n\n---------------------------------------------------------------------------\n\nKeyError                                  Traceback (most recent call last)\n\n/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/script/train_model.py in <module>\n    570 \n    571 if __name__ == \"__main__\":\n--> 572     main()\n\n6 frames\n\n/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/script/train_model.py in main()\n    475 \n    476     cv = CrossVaridation()\n--> 477     result_df, k_fold_models, k_fold_pre_post_processes = cv.compute_score(\n    478         x=train_inputs,\n    479         y=train_target,\n\n/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/script/train_model.py in compute_score(self, x, y, metadata, x_test, metadata_test, params, build_model, build_pre_post_process, dump, dump_dir, n_splits, n_bagging, bagging_ratio, use_batch_group)\n    145                 model = build_model(params=params[\"model\"])\n    146                 print(f\"model input shape X:{preprocessed_x_train.shape} Y:{preprocessed_y_train.shape}\")\n--> 147                 model.fit(\n    148                     x=x_train_bagging,\n    149                     y=y_train_bagging,\n\n/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/ss_opm/model/encoder_decoder/encoder_decoder.py in fit(self, x, preprocessed_x, y, preprocessed_y, metadata, pre_post_process)\n    221         self.model.to(device=self.params[\"device\"])\n    222 \n--> 223         dummy_batch = next(iter(data_loader))\n    224         dummy_batch = self._batch_to_device(dummy_batch)\n    225         self._train_step_forward(dummy_batch, 1.0)\n\n/usr/local/lib/python3.8/dist-packages/torch/utils/data/dataloader.py in __next__(self)\n    626                 # TODO(https://github.com/pytorch/pytorch/issues/76750)\n    627                 self._reset()  # type: ignore[call-arg]\n--> 628             data = self._next_data()\n    629             self._num_yielded += 1\n    630             if self._dataset_kind == _DatasetKind.Iterable and \\\n\n/usr/local/lib/python3.8/dist-packages/torch/utils/data/dataloader.py in _next_data(self)\n   1331             else:\n   1332                 del self._task_info[idx]\n-> 1333                 return self._process_data(data)\n   1334 \n   1335     def _try_put_index(self):\n\n/usr/local/lib/python3.8/dist-packages/torch/utils/data/dataloader.py in _process_data(self, data)\n   1357         self._try_put_index()\n   1358         if isinstance(data, ExceptionWrapper):\n-> 1359             data.reraise()\n   1360         return data\n   1361 \n\n/usr/local/lib/python3.8/dist-packages/torch/_utils.py in reraise(self)\n    541             # instantiate since we don't know how to\n    542             raise RuntimeError(msg) from None\n--> 543         raise exception\n    544 \n    545 \n\nKeyError: Caught KeyError in DataLoader worker process 0.\nOriginal Traceback (most recent call last):\n  File \"/usr/local/lib/python3.8/dist-packages/torch/utils/data/_utils/worker.py\", line 302, in _worker_loop\n    data = fetcher.fetch(index)\n  File \"/usr/local/lib/python3.8/dist-packages/torch/utils/data/_utils/fetch.py\", line 58, in fetch\n    data = [self.dataset[idx] for idx in possibly_batched_index]\n  File \"/usr/local/lib/python3.8/dist-packages/torch/utils/data/_utils/fetch.py\", line 58, in <listcomp>\n    data = [self.dataset[idx] for idx in possibly_batched_index]\n  File \"/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/ss_opm/model/torch_dataset/citeseq_dataset.py\", line 79, in __getitem__\n    info = torch.as_tensor(self.metadata.iloc[index, :][self.metadata_keys].values.astype(float), dtype=torch.float32)\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/series.py\", line 1007, in __getitem__\n    return self._get_with(key)\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/series.py\", line 1047, in _get_with\n    return self.loc[key]\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1073, in __getitem__\n    return self._getitem_axis(maybe_callable, axis=axis)\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1301, in _getitem_axis\n    return self._getitem_iterable(key, axis=axis)\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1239, in _getitem_iterable\n    keyarr, indexer = self._get_listlike_indexer(key, axis)\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1432, in _get_listlike_indexer\n    keyarr, indexer = ax._get_indexer_strict(key, axis_name)\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexes/base.py\", line 6070, in _get_indexer_strict\n    self._raise_if_missing(keyarr, indexer, axis_name)\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexes/base.py\", line 6133, in _raise_if_missing\n    raise KeyError(f\"{not_found} not in index\")\nKeyError: \"['batch_sv2', 'batch_sv3', 'batch_sv4', 'batch_sv5', 'batch_sv6', 'batch_sv7'] not in index\"\n"
    },
    {
      "id": 2038774,
      "postDate": "2022-11-21T16:29:22.983Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/shujisuzuki65\" target=\"_blank\">@shujisuzuki65</a>. Thanks for sharing all your approach details. Well done👍</p>",
      "rawMarkdown": "Congrats @shujisuzuki65. Thanks for sharing all your approach details. Well done👍"
    },
    {
      "id": 2036736,
      "postDate": "2022-11-20T06:22:54.790Z",
      "content": "<p>Congratulations to get the 1st palce, and thanks for sharing.</p>",
      "rawMarkdown": "Congratulations to get the 1st palce, and thanks for sharing."
    },
    {
      "id": 2036585,
      "postDate": "2022-11-19T23:29:48.880Z",
      "content": "<p>Big Congrat! And thanks a lot for sharing your idear</p>",
      "rawMarkdown": "Big Congrat! And thanks a lot for sharing your idear"
    },
    {
      "id": 2035505,
      "postDate": "2022-11-18T23:53:56.370Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/shujisuzuki65\" target=\"_blank\">@shujisuzuki65</a> ! Amazing work, great job!</p>",
      "rawMarkdown": "Congratulations @shujisuzuki65 ! Amazing work, great job!"
    },
    {
      "id": 2035404,
      "postDate": "2022-11-18T20:30:49.857Z",
      "content": "<p>Thanks a lot for the detailed approach note <a href=\"https://www.kaggle.com/shujisuzuki65\" target=\"_blank\">@shujisuzuki65</a>! Hearty congratulations and best regards!</p>\n<p>I wish to know of the approaches that did not work for you too. I eagerly await the adjutant code too. </p>",
      "rawMarkdown": "Thanks a lot for the detailed approach note @shujisuzuki65! Hearty congratulations and best regards!\n\nI wish to know of the approaches that did not work for you too. I eagerly await the adjutant code too. "
    },
    {
      "id": 2035124,
      "postDate": "2022-11-18T16:25:49.437Z",
      "content": "<p>Congratulations!<br>\nI see the combination of MSE/MAE in \"compressed\" space and correlation loss in original space as a approach to model regularization<br>\nDid you select weights of these two losses ? If so, how ?<br>\nWhy you used MAE in citeseq part, but not MSE as in multiome part ?</p>",
      "rawMarkdown": "Congratulations!\nI see the combination of MSE/MAE in \"compressed\" space and correlation loss in original space as a approach to model regularization\nDid you select weights of these two losses ? If so, how ?\nWhy you used MAE in citeseq part, but not MSE as in multiome part ?\n",
      "replies": [
        {
          "id": 2036933,
          "postDate": "2022-11-20T10:02:29.223Z",
          "content": "<p>The MAE/MSE weights were set at 1.0 at the beginning of training and the weights were gradually reduced as training progressed.</p>\n<p>The detail of weight schedule is here:<br>\n<a href=\"https://github.com/shu65/open-problems-multimodal/blob/3d57dd3837b17079fed5678043e681749ba32324/ss_opm/model/encoder_decoder/cite_encoder_decoder_module.py#L81\" target=\"_blank\">https://github.com/shu65/open-problems-multimodal/blob/3d57dd3837b17079fed5678043e681749ba32324/ss_opm/model/encoder_decoder/cite_encoder_decoder_module.py#L81</a></p>\n<p>I tried MAE as in multiome part. But, the score of the model with MAE is lower than that with MSE. </p>",
          "rawMarkdown": "The MAE/MSE weights were set at 1.0 at the beginning of training and the weights were gradually reduced as training progressed.\n\nThe detail of weight schedule is here:\nhttps://github.com/shu65/open-problems-multimodal/blob/3d57dd3837b17079fed5678043e681749ba32324/ss_opm/model/encoder_decoder/cite_encoder_decoder_module.py#L81\n\nI tried MAE as in multiome part. But, the score of the model with MAE is lower than that with MSE. "
        }
      ]
    },
    {
      "id": 2615313,
      "postDate": "2024-01-23T04:10:33.873Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2351086,
      "postDate": "2023-07-19T19:51:12.627Z",
      "content": "<p>Very cool stuff</p>",
      "rawMarkdown": "Very cool stuff",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2041633,
      "author_name": "BarryZhou",
      "author_url": "",
      "post_date": "2022-11-24T05:15:32.717000",
      "content": "<p>Congratulations and thanks for posting your great solution Shuji Suzuki!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2046907,
      "author_name": "Liilili",
      "author_url": "",
      "post_date": "2022-11-28T12:53:04.240000",
      "content": "<p>Big old 🙏🙏🙏 <br>\nMay i ask some questions emmm<br>\nThe part of multi's model is fantastic, why  not  apply to Cite problems,  by residual target this way, it is not work in the cite problems?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2038496,
      "author_name": "andreylalaley",
      "author_url": "",
      "post_date": "2022-11-21T12:36:55.080000",
      "content": "<p>Congratulations for thq 1st place and thank you for sharing your solution! <br>\nAs I see, you ve added metadata to both models. What increase (cv/ lb) did you get from adding metadata to models? </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2036163,
      "author_name": "vialactea",
      "author_url": "",
      "post_date": "2022-11-19T14:09:14.427000",
      "content": "<p>Congratulations for the 1st position and great solution!!</p>\n<p>I tried combining multiple forms of inputs/targets (initial inputs, SVD and binarized) and weight averaging their losses, but dropped it because the added complexity wasn't helping. I missed the ingenuity of your preprocessing chain and the averaging of the outputs for prediction (I took only the prediction of the original targets, and used the other outputs solely to regularize training).</p>\n<p>I have seen big PB jumps to gold medal positions - <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> comes to mind, but I don't recall ever seeing such an impressive finish in a winner. What a great job! I'm curious about how surprised you where with taking the top position after ending the competition in a position far from the gold medals in the LB.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2036235,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2022-11-19T15:44:46.487000",
          "content": "<p>You jumped quite a bit too.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2036927,
          "author_name": "Shuji Suzuki",
          "author_url": "",
          "post_date": "2022-11-20T09:44:46.323000",
          "content": "<p>Thank you! </p>\n<p>I couldn't believe it at first either and showed the screenshot of LB to my colleagues at work to make sure I was in first place and not wrong.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2035038,
      "author_name": "Syed Siddique Mridul",
      "author_url": "",
      "post_date": "2022-11-18T14:52:28.227000",
      "content": "<p>One of the best posts I have seen</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2035261,
      "author_name": "Artem Fedorov",
      "author_url": "",
      "post_date": "2022-11-18T18:15:49.657000",
      "content": "<p>Thank you for this detailed post.<br>\nI see you made a lot of work preprocessing the data. But how did you come to this sort of preprocessing? Where does the idea to divide all the inputs by non-zero medians come from?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2036930,
          "author_name": "Shuji Suzuki",
          "author_url": "",
          "post_date": "2022-11-20T09:55:49.183000",
          "content": "<p>I noticed that the library-size normalization + log1p data is not being returned well after dimensionality compression with tSVD. The correlation coefficient between the original data and the converted and reverted data is only about 0.70.</p>\n<p>For this reason, I tried to find a better preprocessing method and found it.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 2042059,
          "author_name": "Mikhael Manurung",
          "author_url": "",
          "post_date": "2022-11-24T12:08:58.813000",
          "content": "<p>Dividing by non-zero medians is something I've never seen in single-cell literature, but hey, it WORKS!<br>\nHow do you get to this approach? Could you elaborate on your thought process? Thanks and congratulations!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2058378,
          "author_name": "Daniel Burkhardt",
          "author_url": "",
          "post_date": "2022-12-07T21:29:54.460000",
          "content": "<p><a href=\"https://www.kaggle.com/shujisuzuki65\" target=\"_blank\">@shujisuzuki65</a> am I right in understanding you looked at many different pre-processing approaches and optimized for one that had high correlation after low-rank-approximation with tSVD? This led to selecting normalizing by median non-zero values?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2128595,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2023-02-03T21:02:48.433000",
      "content": "<p>Hello Shuji. I tried to run your code with the minor modification that I only use 996 data points for both the train and test set to speed up the calculations. I get this output/error when training the cite model and don't know what to do. The error persists even when changing batch-size to 1000, for some reason it always looks for 8 batches and can't find 6 of them: <br>\ngit_hexsha Failed to get git<br>\nload input values<br>\ncompleted loading input values. elapsed time: 4.2<br>\nload targets values<br>\ncompleted loading targets values. elapsed time: 0.2<br>\nload input values<br>\ncompleted loading input values. elapsed time: 1.2<br>\nuse test_inputs_values. total size: 1992<br>\ntrain sample size: 996<br>\ndump params<br>\nkfold type  group None<br>\nskip pre_post_process fit<br>\nmodel input shape X:(664, 217) Y:(664, 128)<br>\ndataset size 664</p>\n<hr>\n<p>KeyError                                  Traceback (most recent call last)</p>\n<p>/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/script/train_model.py in <br>\n    570 <br>\n    571 if <strong>name</strong> == \"<strong>main</strong>\":<br>\n--&gt; 572     main()</p>\n<p>6 frames</p>\n<p>/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/script/train_model.py in main()<br>\n    475 <br>\n    476     cv = CrossVaridation()<br>\n--&gt; 477     result_df, k_fold_models, k_fold_pre_post_processes = cv.compute_score(<br>\n    478         x=train_inputs,<br>\n    479         y=train_target,</p>\n<p>/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/script/train_model.py in compute_score(self, x, y, metadata, x_test, metadata_test, params, build_model, build_pre_post_process, dump, dump_dir, n_splits, n_bagging, bagging_ratio, use_batch_group)<br>\n    145                 model = build_model(params=params[\"model\"])<br>\n    146                 print(f\"model input shape X:{preprocessed_x_train.shape} Y:{preprocessed_y_train.shape}\")<br>\n--&gt; 147                 model.fit(<br>\n    148                     x=x_train_bagging,<br>\n    149                     y=y_train_bagging,</p>\n<p>/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/ss_opm/model/encoder_decoder/encoder_decoder.py in fit(self, x, preprocessed_x, y, preprocessed_y, metadata, pre_post_process)<br>\n    221         self.model.to(device=self.params[\"device\"])<br>\n    222 <br>\n--&gt; 223         dummy_batch = next(iter(data_loader))<br>\n    224         dummy_batch = self._batch_to_device(dummy_batch)<br>\n    225         self._train_step_forward(dummy_batch, 1.0)</p>\n<p>/usr/local/lib/python3.8/dist-packages/torch/utils/data/dataloader.py in <strong>next</strong>(self)<br>\n    626                 # TODO(<a href=\"https://github.com/pytorch/pytorch/issues/76750\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/76750</a>)<br>\n    627                 self._reset()  # type: ignore[call-arg]<br>\n--&gt; 628             data = self._next_data()<br>\n    629             self._num_yielded += 1<br>\n    630             if self._dataset_kind == _DatasetKind.Iterable and \\</p>\n<p>/usr/local/lib/python3.8/dist-packages/torch/utils/data/dataloader.py in _next_data(self)<br>\n   1331             else:<br>\n   1332                 del self._task_info[idx]<br>\n-&gt; 1333                 return self._process_data(data)<br>\n   1334 <br>\n   1335     def _try_put_index(self):</p>\n<p>/usr/local/lib/python3.8/dist-packages/torch/utils/data/dataloader.py in _process_data(self, data)<br>\n   1357         self._try_put_index()<br>\n   1358         if isinstance(data, ExceptionWrapper):<br>\n-&gt; 1359             data.reraise()<br>\n   1360         return data<br>\n   1361 </p>\n<p>/usr/local/lib/python3.8/dist-packages/torch/_utils.py in reraise(self)<br>\n    541             # instantiate since we don't know how to<br>\n    542             raise RuntimeError(msg) from None<br>\n--&gt; 543         raise exception<br>\n    544 <br>\n    545 </p>\n<p>KeyError: Caught KeyError in DataLoader worker process 0.<br>\nOriginal Traceback (most recent call last):<br>\n  File \"/usr/local/lib/python3.8/dist-packages/torch/utils/data/_utils/worker.py\", line 302, in _worker_loop<br>\n    data = fetcher.fetch(index)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/torch/utils/data/_utils/fetch.py\", line 58, in fetch<br>\n    data = [self.dataset[idx] for idx in possibly_batched_index]<br>\n  File \"/usr/local/lib/python3.8/dist-packages/torch/utils/data/_utils/fetch.py\", line 58, in <br>\n    data = [self.dataset[idx] for idx in possibly_batched_index]<br>\n  File \"/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/ss_opm/model/torch_dataset/citeseq_dataset.py\", line 79, in <strong>getitem</strong><br>\n    info = torch.as_tensor(self.metadata.iloc[index, :][self.metadata_keys].values.astype(float), dtype=torch.float32)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/series.py\", line 1007, in <strong>getitem</strong><br>\n    return self._get_with(key)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/series.py\", line 1047, in _get_with<br>\n    return self.loc[key]<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1073, in <strong>getitem</strong><br>\n    return self._getitem_axis(maybe_callable, axis=axis)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1301, in _getitem_axis<br>\n    return self._getitem_iterable(key, axis=axis)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1239, in _getitem_iterable<br>\n    keyarr, indexer = self._get_listlike_indexer(key, axis)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1432, in _get_listlike_indexer<br>\n    keyarr, indexer = ax._get_indexer_strict(key, axis_name)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexes/base.py\", line 6070, in _get_indexer_strict<br>\n    self._raise_if_missing(keyarr, indexer, axis_name)<br>\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexes/base.py\", line 6133, in _raise_if_missing<br>\n    raise KeyError(f\"{not_found} not in index\")<br>\nKeyError: \"['batch_sv2', 'batch_sv3', 'batch_sv4', 'batch_sv5', 'batch_sv6', 'batch_sv7'] not in index\"</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2038774,
      "author_name": "Oscar Aguilar",
      "author_url": "",
      "post_date": "2022-11-21T16:29:22.983000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/shujisuzuki65\" target=\"_blank\">@shujisuzuki65</a>. Thanks for sharing all your approach details. Well done👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2036736,
      "author_name": "zhez017",
      "author_url": "",
      "post_date": "2022-11-20T06:22:54.790000",
      "content": "<p>Congratulations to get the 1st palce, and thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2036585,
      "author_name": "Kexin Wang",
      "author_url": "",
      "post_date": "2022-11-19T23:29:48.880000",
      "content": "<p>Big Congrat! And thanks a lot for sharing your idear</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2035505,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-11-18T23:53:56.370000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/shujisuzuki65\" target=\"_blank\">@shujisuzuki65</a> ! Amazing work, great job!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2035404,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-11-18T20:30:49.857000",
      "content": "<p>Thanks a lot for the detailed approach note <a href=\"https://www.kaggle.com/shujisuzuki65\" target=\"_blank\">@shujisuzuki65</a>! Hearty congratulations and best regards!</p>\n<p>I wish to know of the approaches that did not work for you too. I eagerly await the adjutant code too. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2035124,
      "author_name": "Arthur Mulikhov",
      "author_url": "",
      "post_date": "2022-11-18T16:25:49.437000",
      "content": "<p>Congratulations!<br>\nI see the combination of MSE/MAE in \"compressed\" space and correlation loss in original space as a approach to model regularization<br>\nDid you select weights of these two losses ? If so, how ?<br>\nWhy you used MAE in citeseq part, but not MSE as in multiome part ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2036933,
          "author_name": "Shuji Suzuki",
          "author_url": "",
          "post_date": "2022-11-20T10:02:29.223000",
          "content": "<p>The MAE/MSE weights were set at 1.0 at the beginning of training and the weights were gradually reduced as training progressed.</p>\n<p>The detail of weight schedule is here:<br>\n<a href=\"https://github.com/shu65/open-problems-multimodal/blob/3d57dd3837b17079fed5678043e681749ba32324/ss_opm/model/encoder_decoder/cite_encoder_decoder_module.py#L81\" target=\"_blank\">https://github.com/shu65/open-problems-multimodal/blob/3d57dd3837b17079fed5678043e681749ba32324/ss_opm/model/encoder_decoder/cite_encoder_decoder_module.py#L81</a></p>\n<p>I tried MAE as in multiome part. But, the score of the model with MAE is lower than that with MSE. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2615313,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-23T04:10:33.873000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2351086,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-07-19T19:51:12.627000",
      "content": "<p>Very cool stuff</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2034844": "First of all, thank you to the organizers and kaggle management and to everyone who participated with me.\nSince I needed to gain experience in analyzing single cell data, this competition was an excellent experience for me.\n\nI would like to introduce the overview of my solution.\n\n# Multiome\n## Model Overview\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F05d87da294f71eca450a304785333ac7%2Fmult-model-overview.png?generation=1668772959634455&alt=media)\n\n## Input Preprocessing\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fdedb439f946e0682644451a531d4df33%2Fmulti-input-preprocessing.png?generation=1668773118728830&alt=media)\n\n## Target Preprocessing\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F76231a7e2c23f3f588e763bf8035dcd0%2Fmulti-target-preprocessing.png?generation=1669418912414786&alt=media)\n\ntSVD-based imputation method: \n1. Perform dimensionality reduction on the data with tSVD\n2. And then, Transform the data back to the original space\n3. Copy the value of the 0 part of the original data from the transformed values.\n\n## Model\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F8861945be57ca34a6bfa2b23e5627100%2Fmulti-model.png?generation=1668773599707773&alt=media)\n\n## Output Postprocessing and Loss\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Ff59241b0358a15294bdc38c25722f37d%2Fmulti-postprocessing_1.png?generation=1669418940293447&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fbff80fd72fd2a9d0ab94704c641c024b%2Fmulti-postprocessing_2.png?generation=1669419002523851&alt=media)\n\nIn the inference phase, the model outputs the average of the five predicted target data.\n\n# CITEseq\n## Model Overview\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fa9563a70e5639db2b4b72b7645ecd910%2Fcite-model-overview.png?generation=1668774675355569&alt=media)\n\n## Input Preprocessing\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Ff5ddef7a059aa3a94c3afc6f3f52b440%2Fcite-input-preprocessing.png?generation=1669419233099917&alt=media)\n\nIn selecting important genes in CITEseq, the correlation coefficient is calculated for each batch and select only genes with high correlation in many batches.\nGenes were selected from those related to the target proteins and pathway.\nI use [Reactome](https://reactome.org/) as pathway database.\n\n## Target Preprocessing\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2Fc538366bbc83a1de46005c5108d6b20d%2Fcite-target-preprocessing.png?generation=1668774972017361&alt=media)\n\n## Model\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F3a30e371014051289d99d5cbe57009e9%2Fcite-model.png?generation=1668775048364648&alt=media)\n\n\n## Output Postprocessing and Loss\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1053132%2F677385c92a3be6a0e0a9702b3005edc8%2Fcite-postprocessing.png?generation=1668775162024292&alt=media)\n\nIn the inference phase, the model outputs the average of the five predicted target data.\n\n# Local evaluation\nI used two evaluation schemes.\n\n1. Evaluation with cross validation:\n  * 5-fold cross validation grouped by donor and day\n2. Evaluation for hyperparameter optimization with Optuna:\n  * Training data set is divided into training and validation data sets. ( Training data set: 80%, validation data set: 20%. )\n\n# Ensemble\nI used the weighted average of predictions of the following models.\n1. Models trained with changing the seed \n2. Models fine-tuned on only some batches\n  * Batch combination pattern examples: males only, female only, Day 4, 7 only, etc.\n  * Use a model trained on the full training data set as a pre-training model \n\n# Code\nhttps://github.com/shu65/open-problems-multimodal\n\n\n# Update\n2022/11/20 add the repository url of my solution\n2022/11/26 fix some figures",
    "2041633": "Congratulations and thanks for posting your great solution Shuji Suzuki!",
    "2046907": "Big old 🙏🙏🙏 \nMay i ask some questions emmm\nThe part of multi's model is fantastic, why  not  apply to Cite problems,  by residual target this way, it is not work in the cite problems?",
    "2038496": "Congratulations for thq 1st place and thank you for sharing your solution! \nAs I see, you ve added metadata to both models. What increase (cv/ lb) did you get from adding metadata to models? ",
    "2036163": "Congratulations for the 1st position and great solution!!\n\nI tried combining multiple forms of inputs/targets (initial inputs, SVD and binarized) and weight averaging their losses, but dropped it because the added complexity wasn't helping. I missed the ingenuity of your preprocessing chain and the averaging of the outputs for prediction (I took only the prediction of the original targets, and used the other outputs solely to regularize training).\n\nI have seen big PB jumps to gold medal positions - @cpmpml comes to mind, but I don't recall ever seeing such an impressive finish in a winner. What a great job! I'm curious about how surprised you where with taking the top position after ending the competition in a position far from the gold medals in the LB.",
    "2035038": "One of the best posts I have seen",
    "2035261": "Thank you for this detailed post.\nI see you made a lot of work preprocessing the data. But how did you come to this sort of preprocessing? Where does the idea to divide all the inputs by non-zero medians come from?",
    "2128595": "Hello Shuji. I tried to run your code with the minor modification that I only use 996 data points for both the train and test set to speed up the calculations. I get this output/error when training the cite model and don't know what to do. The error persists even when changing batch-size to 1000, for some reason it always looks for 8 batches and can't find 6 of them: \ngit_hexsha Failed to get git\nload input values\ncompleted loading input values. elapsed time: 4.2\nload targets values\ncompleted loading targets values. elapsed time: 0.2\nload input values\ncompleted loading input values. elapsed time: 1.2\nuse test_inputs_values. total size: 1992\ntrain sample size: 996\ndump params\nkfold type <class 'sklearn.model_selection._split.KFold'> group None\nskip pre_post_process fit\nmodel input shape X:(664, 217) Y:(664, 128)\ndataset size 664\n\n---------------------------------------------------------------------------\n\nKeyError                                  Traceback (most recent call last)\n\n/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/script/train_model.py in <module>\n    570 \n    571 if __name__ == \"__main__\":\n--> 572     main()\n\n6 frames\n\n/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/script/train_model.py in main()\n    475 \n    476     cv = CrossVaridation()\n--> 477     result_df, k_fold_models, k_fold_pre_post_processes = cv.compute_score(\n    478         x=train_inputs,\n    479         y=train_target,\n\n/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/script/train_model.py in compute_score(self, x, y, metadata, x_test, metadata_test, params, build_model, build_pre_post_process, dump, dump_dir, n_splits, n_bagging, bagging_ratio, use_batch_group)\n    145                 model = build_model(params=params[\"model\"])\n    146                 print(f\"model input shape X:{preprocessed_x_train.shape} Y:{preprocessed_y_train.shape}\")\n--> 147                 model.fit(\n    148                     x=x_train_bagging,\n    149                     y=y_train_bagging,\n\n/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/ss_opm/model/encoder_decoder/encoder_decoder.py in fit(self, x, preprocessed_x, y, preprocessed_y, metadata, pre_post_process)\n    221         self.model.to(device=self.params[\"device\"])\n    222 \n--> 223         dummy_batch = next(iter(data_loader))\n    224         dummy_batch = self._batch_to_device(dummy_batch)\n    225         self._train_step_forward(dummy_batch, 1.0)\n\n/usr/local/lib/python3.8/dist-packages/torch/utils/data/dataloader.py in __next__(self)\n    626                 # TODO(https://github.com/pytorch/pytorch/issues/76750)\n    627                 self._reset()  # type: ignore[call-arg]\n--> 628             data = self._next_data()\n    629             self._num_yielded += 1\n    630             if self._dataset_kind == _DatasetKind.Iterable and \\\n\n/usr/local/lib/python3.8/dist-packages/torch/utils/data/dataloader.py in _next_data(self)\n   1331             else:\n   1332                 del self._task_info[idx]\n-> 1333                 return self._process_data(data)\n   1334 \n   1335     def _try_put_index(self):\n\n/usr/local/lib/python3.8/dist-packages/torch/utils/data/dataloader.py in _process_data(self, data)\n   1357         self._try_put_index()\n   1358         if isinstance(data, ExceptionWrapper):\n-> 1359             data.reraise()\n   1360         return data\n   1361 \n\n/usr/local/lib/python3.8/dist-packages/torch/_utils.py in reraise(self)\n    541             # instantiate since we don't know how to\n    542             raise RuntimeError(msg) from None\n--> 543         raise exception\n    544 \n    545 \n\nKeyError: Caught KeyError in DataLoader worker process 0.\nOriginal Traceback (most recent call last):\n  File \"/usr/local/lib/python3.8/dist-packages/torch/utils/data/_utils/worker.py\", line 302, in _worker_loop\n    data = fetcher.fetch(index)\n  File \"/usr/local/lib/python3.8/dist-packages/torch/utils/data/_utils/fetch.py\", line 58, in fetch\n    data = [self.dataset[idx] for idx in possibly_batched_index]\n  File \"/usr/local/lib/python3.8/dist-packages/torch/utils/data/_utils/fetch.py\", line 58, in <listcomp>\n    data = [self.dataset[idx] for idx in possibly_batched_index]\n  File \"/content/drive/MyDrive/bio/suzuki/open-problems-multimodal-main/ss_opm/model/torch_dataset/citeseq_dataset.py\", line 79, in __getitem__\n    info = torch.as_tensor(self.metadata.iloc[index, :][self.metadata_keys].values.astype(float), dtype=torch.float32)\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/series.py\", line 1007, in __getitem__\n    return self._get_with(key)\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/series.py\", line 1047, in _get_with\n    return self.loc[key]\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1073, in __getitem__\n    return self._getitem_axis(maybe_callable, axis=axis)\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1301, in _getitem_axis\n    return self._getitem_iterable(key, axis=axis)\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1239, in _getitem_iterable\n    keyarr, indexer = self._get_listlike_indexer(key, axis)\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexing.py\", line 1432, in _get_listlike_indexer\n    keyarr, indexer = ax._get_indexer_strict(key, axis_name)\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexes/base.py\", line 6070, in _get_indexer_strict\n    self._raise_if_missing(keyarr, indexer, axis_name)\n  File \"/usr/local/lib/python3.8/dist-packages/pandas/core/indexes/base.py\", line 6133, in _raise_if_missing\n    raise KeyError(f\"{not_found} not in index\")\nKeyError: \"['batch_sv2', 'batch_sv3', 'batch_sv4', 'batch_sv5', 'batch_sv6', 'batch_sv7'] not in index\"\n",
    "2038774": "Congrats @shujisuzuki65. Thanks for sharing all your approach details. Well done👍",
    "2036736": "Congratulations to get the 1st palce, and thanks for sharing.",
    "2036585": "Big Congrat! And thanks a lot for sharing your idear",
    "2035505": "Congratulations @shujisuzuki65 ! Amazing work, great job!",
    "2035404": "Thanks a lot for the detailed approach note @shujisuzuki65! Hearty congratulations and best regards!\n\nI wish to know of the approaches that did not work for you too. I eagerly await the adjutant code too. ",
    "2035124": "Congratulations!\nI see the combination of MSE/MAE in \"compressed\" space and correlation loss in original space as a approach to model regularization\nDid you select weights of these two losses ? If so, how ?\nWhy you used MAE in citeseq part, but not MSE as in multiome part ?\n",
    "2615313": "",
    "2351086": "Very cool stuff"
  }
}