{
  "id": 366372,
  "title": "Private 41st Solution summary",
  "url": "/competitions/open-problems-multimodal/writeups/ponkots-private-41st-solution-summary",
  "author_name": "",
  "post_date": "2022-11-23T22:45:02.853Z",
  "votes": 28,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thank you for all kagglers and organizers for this competition.<br>\nThis note is a record of my work on the competition.<br>\nThis was a competition where the Private test data was from a future date that did not exist in the training data, and the so-called domain generalization performance was being tested.<br>\nOn the other hand, there was an element of variation by date, and it was expected that it would be undesirable to completely ignore the date feature.<br>\nFirst, we conducted adversarial training (a task to classify training and test data) and found that Citeseq was capable of 99% classification, and we were concerned that training with this feature set would result in overtraining on the training data.<br>\nHowever, when we reduced the number of features to reduce the accuracy of Adversarial training, the score of Public LB also dropped significantly.<br>\nTherefore, we decided to devise some kind of biological features and to improve generalization performance through model variation.</p>\n<h2>✨ Result</h2>\n<ul>\n<li>Private: 0.769</li>\n<li>Public: 0.813</li>\n</ul>\n<h2>🖼️ Solution</h2>\n<h3>🌱 Preprocess</h3>\n<ul>\n<li>Citeseq<ul>\n<li>The input data was reduced to 100 dimensions by PCA.</li>\n<li>On the other hand, the data of important columns were preserved.</li>\n<li><a href=\"https://bering-ivis.readthedocs.io/en/latest/unsupervised.html\" target=\"_blank\">Ivis unsupervised learning</a> was used to generate 100 dimensional features.</li>\n<li>In addition, we added the sum of mitochondrial RNA cells to the features.</li>\n<li>Cell type in Metadata was added to the features.</li></ul></li>\n<li>Multiome<ul>\n<li>For each group with the same column name prefix, PCA reduced the number of dimensions to approximately 100 each.</li>\n<li>Ivis unsupervised learning was used to generate 100 dimensional features.</li></ul></li>\n</ul>\n<h3>🤸 Pre Training</h3>\n<ul>\n<li>Adversarial training (a task to classify training data and test data) is performed and the misjudged training data is used as good validation data.</li>\n<li>Prediction of Cell type for Multiome is performed and added to the features.</li>\n</ul>\n<h3>🏃 Training</h3>\n<ul>\n<li>StratifiedKFold with good validation data as positive labels.</li>\n<li>Pearson correlation coefficient was used for the Loss function. XGBoost was implemented as described below.</li>\n<li>TabNet also performed pre-training. (In this competition, pre-training was more accurate.)</li>\n</ul>\n<h3>🎨 Base Models</h3>\n<ul>\n<li>Citeseq<ul>\n<li>TabNet</li>\n<li>Simple MLP</li>\n<li>ResNet</li>\n<li>1D CNN</li>\n<li>XGBoost</li></ul></li>\n<li>Multiome<ul>\n<li>1D CNN<br>\nCiteseq scored well with an ensemble of various models.<br>\nOn the other hand, Multiome had a strong 1D CNN and did not score well with ensembles of other models, so only the 1D CNN was used.</li></ul></li>\n</ul>\n<h3>🚀 Postprocess</h3>\n<ul>\n<li>Since the evaluation metric is the Pearson correlation coefficient, each inference result (including OOF results) was normalized before ensemble.</li>\n<li>Optuna was used to optimize the ensemble weights. Good validation data was used as the evaluation metric.</li>\n<li>Ensemble with Public Notebook x2 and teammate submissions.</li>\n</ul>\n<h2>💡 Tips</h2>\n<h3>Pearson Loss for XGBoost</h3>\n<p>XGBoost does not provide a Pearson Loss Function, so I implemented it as follows.<br>\nHowever, this implementation is slow in learning, and I have the impression that I would like to improve it a little more.</p>\n<pre><code>from functools import partial\nfrom typing import Any, Callable\nimport numpy as np\nimport torch\nimport torch.nn.functional as F\nimport xgboost as xgb\ndef pearson_cc_loss(inputs, targets):\n  try:\n      assert inputs.shape == targets.shape\n  except AssertionError:\n      inputs = inputs.view(targets.shape)\n   pcc = F.cosine_similarity(inputs, targets)\n  return 1.0 - pcc\n# https://towardsdatascience.com/jax-vs-pytorch-automatic-differentiation-for-xgboost-10222e1404ec\ndef torch_autodiff_grad_hess(\n  loss_function: Callable[[torch.Tensor, torch.Tensor], torch.Tensor], y_true: np.ndarray, y_pred: np.ndarray\n):\n  \"\"\"\n  Perform automatic differentiation to get the\n  Gradient and the Hessian of `loss_function`.\n  \"\"\"\n  y_true = torch.tensor(y_true, dtype=torch.float, requires_grad=False)\n  y_pred = torch.tensor(y_pred, dtype=torch.float, requires_grad=True)\n  loss_function_sum = lambda y_pred: loss_function(y_true, y_pred).sum()\n   loss_function_sum(y_pred).backward()\n  grad = y_pred.grad.reshape(-1)\n   # hess_matrix = torch.autograd.functional.hessian(loss_function_sum, y_pred, vectorize=True)\n  # hess = torch.diagonal(hess_matrix)\n  hess = np.ones(grad.shape)\n   return grad, hess\ncustom_objective = partial(torch_autodiff_grad_hess, pearson_cc_loss)\nxgb_params = dict(\n  n_estimators=10000,\n  early_stopping_rounds=20,\n  # learning_rate=0.05,\n  objective=custom_objective,  # \"binary:logistic\", \"reg:squarederror\",\n  eval_metric=pearson_cc_xgb_score,  # \"logloss\", \"rmse\",\n  random_state=440,\n  tree_method=\"gpu_hist\",\n)  # type: dict[str, Any]\nclf = xgb.XGBRegressor(**xgb_params)\n</code></pre>\n<h2>🏷️ Links</h2>\n<ul>\n<li><a href=\"https://github.com/IMOKURI/kaggle-multimodal-single-cell-integration\" target=\"_blank\">My Solution</a></li>\n</ul>",
  "messages": [
    {
      "id": "2031222",
      "postDate": "11/16/2022 00:15:40",
      "content": "<p>Thank you for all kagglers and organizers for this competition.<br>\nThis note is a record of my work on the competition.<br>\nThis was a competition where the Private test data was from a future date that did not exist in the training data, and the so-called domain generalization performance was being tested.<br>\nOn the other hand, there was an element of variation by date, and it was expected that it would be undesirable to completely ignore the date feature.<br>\nFirst, we conducted adversarial training (a task to classify training and test data) and found that Citeseq was capable of 99% classification, and we were concerned that training with this feature set would result in overtraining on the training data.<br>\nHowever, when we reduced the number of features to reduce the accuracy of Adversarial training, the score of Public LB also dropped significantly.<br>\nTherefore, we decided to devise some kind of biological features and to improve generalization performance through model variation.</p>\n<h2>✨ Result</h2>\n<ul>\n<li>Private: 0.769</li>\n<li>Public: 0.813</li>\n</ul>\n<h2>🖼️ Solution</h2>\n<h3>🌱 Preprocess</h3>\n<ul>\n<li>Citeseq<ul>\n<li>The input data was reduced to 100 dimensions by PCA.</li>\n<li>On the other hand, the data of important columns were preserved.</li>\n<li><a href=\"https://bering-ivis.readthedocs.io/en/latest/unsupervised.html\" target=\"_blank\">Ivis unsupervised learning</a> was used to generate 100 dimensional features.</li>\n<li>In addition, we added the sum of mitochondrial RNA cells to the features.</li>\n<li>Cell type in Metadata was added to the features.</li></ul></li>\n<li>Multiome<ul>\n<li>For each group with the same column name prefix, PCA reduced the number of dimensions to approximately 100 each.</li>\n<li>Ivis unsupervised learning was used to generate 100 dimensional features.</li></ul></li>\n</ul>\n<h3>🤸 Pre Training</h3>\n<ul>\n<li>Adversarial training (a task to classify training data and test data) is performed and the misjudged training data is used as good validation data.</li>\n<li>Prediction of Cell type for Multiome is performed and added to the features.</li>\n</ul>\n<h3>🏃 Training</h3>\n<ul>\n<li>StratifiedKFold with good validation data as positive labels.</li>\n<li>Pearson correlation coefficient was used for the Loss function. XGBoost was implemented as described below.</li>\n<li>TabNet also performed pre-training. (In this competition, pre-training was more accurate.)</li>\n</ul>\n<h3>🎨 Base Models</h3>\n<ul>\n<li>Citeseq<ul>\n<li>TabNet</li>\n<li>Simple MLP</li>\n<li>ResNet</li>\n<li>1D CNN</li>\n<li>XGBoost</li></ul></li>\n<li>Multiome<ul>\n<li>1D CNN<br>\nCiteseq scored well with an ensemble of various models.<br>\nOn the other hand, Multiome had a strong 1D CNN and did not score well with ensembles of other models, so only the 1D CNN was used.</li></ul></li>\n</ul>\n<h3>🚀 Postprocess</h3>\n<ul>\n<li>Since the evaluation metric is the Pearson correlation coefficient, each inference result (including OOF results) was normalized before ensemble.</li>\n<li>Optuna was used to optimize the ensemble weights. Good validation data was used as the evaluation metric.</li>\n<li>Ensemble with Public Notebook x2 and teammate submissions.</li>\n</ul>\n<h2>💡 Tips</h2>\n<h3>Pearson Loss for XGBoost</h3>\n<p>XGBoost does not provide a Pearson Loss Function, so I implemented it as follows.<br>\nHowever, this implementation is slow in learning, and I have the impression that I would like to improve it a little more.</p>\n<pre><code>from functools import partial\nfrom typing import Any, Callable\nimport numpy as np\nimport torch\nimport torch.nn.functional as F\nimport xgboost as xgb\ndef pearson_cc_loss(inputs, targets):\n  try:\n      assert inputs.shape == targets.shape\n  except AssertionError:\n      inputs = inputs.view(targets.shape)\n   pcc = F.cosine_similarity(inputs, targets)\n  return 1.0 - pcc\n# https://towardsdatascience.com/jax-vs-pytorch-automatic-differentiation-for-xgboost-10222e1404ec\ndef torch_autodiff_grad_hess(\n  loss_function: Callable[[torch.Tensor, torch.Tensor], torch.Tensor], y_true: np.ndarray, y_pred: np.ndarray\n):\n  \"\"\"\n  Perform automatic differentiation to get the\n  Gradient and the Hessian of `loss_function`.\n  \"\"\"\n  y_true = torch.tensor(y_true, dtype=torch.float, requires_grad=False)\n  y_pred = torch.tensor(y_pred, dtype=torch.float, requires_grad=True)\n  loss_function_sum = lambda y_pred: loss_function(y_true, y_pred).sum()\n   loss_function_sum(y_pred).backward()\n  grad = y_pred.grad.reshape(-1)\n   # hess_matrix = torch.autograd.functional.hessian(loss_function_sum, y_pred, vectorize=True)\n  # hess = torch.diagonal(hess_matrix)\n  hess = np.ones(grad.shape)\n   return grad, hess\ncustom_objective = partial(torch_autodiff_grad_hess, pearson_cc_loss)\nxgb_params = dict(\n  n_estimators=10000,\n  early_stopping_rounds=20,\n  # learning_rate=0.05,\n  objective=custom_objective,  # \"binary:logistic\", \"reg:squarederror\",\n  eval_metric=pearson_cc_xgb_score,  # \"logloss\", \"rmse\",\n  random_state=440,\n  tree_method=\"gpu_hist\",\n)  # type: dict[str, Any]\nclf = xgb.XGBRegressor(**xgb_params)\n</code></pre>\n<h2>🏷️ Links</h2>\n<ul>\n<li><a href=\"https://github.com/IMOKURI/kaggle-multimodal-single-cell-integration\" target=\"_blank\">My Solution</a></li>\n</ul>",
      "rawMarkdown": "Thank you for all kagglers and organizers for this competition.\n\nThis note is a record of my work on the competition.\n\nThis was a competition where the Private test data was from a future date that did not exist in the training data, and the so-called domain generalization performance was being tested.\nOn the other hand, there was an element of variation by date, and it was expected that it would be undesirable to completely ignore the date feature.\n\nFirst, we conducted adversarial training (a task to classify training and test data) and found that Citeseq was capable of 99% classification, and we were concerned that training with this feature set would result in overtraining on the training data.\nHowever, when we reduced the number of features to reduce the accuracy of Adversarial training, the score of Public LB also dropped significantly.\n\nTherefore, we decided to devise some kind of biological features and to improve generalization performance through model variation.\n\n\n## ✨ Result\n\n- Private: 0.769\n- Public: 0.813\n\n\n\n## 🖼️ Solution\n\n\n### 🌱 Preprocess\n\n- Citeseq\n    - The input data was reduced to 100 dimensions by PCA.\n    - On the other hand, the data of important columns were preserved.\n    - [Ivis unsupervised learning](https://bering-ivis.readthedocs.io/en/latest/unsupervised.html) was used to generate 100 dimensional features.\n    - In addition, we added the sum of mitochondrial RNA cells to the features.\n    - Cell type in Metadata was added to the features.\n\n- Multiome\n    - For each group with the same column name prefix, PCA reduced the number of dimensions to approximately 100 each.\n    - Ivis unsupervised learning was used to generate 100 dimensional features.\n\n### 🤸 Pre Training\n\n- Adversarial training (a task to classify training data and test data) is performed and the misjudged training data is used as good validation data.\n- Prediction of Cell type for Multiome is performed and added to the features.\n\n\n### 🏃 Training\n\n- StratifiedKFold with good validation data as positive labels.\n- Pearson correlation coefficient was used for the Loss function. XGBoost was implemented as described below.\n- TabNet also performed pre-training. (In this competition, pre-training was more accurate.)\n\n### 🎨 Base Models\n\n- Citeseq\n    - TabNet\n    - Simple MLP\n    - ResNet\n    - 1D CNN\n    - XGBoost\n- Multiome\n    - 1D CNN\n\n\nCiteseq scored well with an ensemble of various models.\nOn the other hand, Multiome had a strong 1D CNN and did not score well with ensembles of other models, so only the 1D CNN was used.\n\n### 🚀 Postprocess\n\n- Since the evaluation metric is the Pearson correlation coefficient, each inference result (including OOF results) was normalized before ensemble.\n- Optuna was used to optimize the ensemble weights. Good validation data was used as the evaluation metric.\n- Ensemble with Public Notebook x2 and teammate submissions.\n\n\n## 💡 Tips\n\n\n### Pearson Loss for XGBoost\n\nXGBoost does not provide a Pearson Loss Function, so I implemented it as follows.\nHowever, this implementation is slow in learning, and I have the impression that I would like to improve it a little more.\n\n```\nfrom functools import partial\nfrom typing import Any, Callable\n\nimport numpy as np\nimport torch\nimport torch.nn.functional as F\nimport xgboost as xgb\n\n\ndef pearson_cc_loss(inputs, targets):\n    try:\n        assert inputs.shape == targets.shape\n    except AssertionError:\n        inputs = inputs.view(targets.shape)\n\n    pcc = F.cosine_similarity(inputs, targets)\n    return 1.0 - pcc\n\n\n# https://towardsdatascience.com/jax-vs-pytorch-automatic-differentiation-for-xgboost-10222e1404ec\ndef torch_autodiff_grad_hess(\n    loss_function: Callable[[torch.Tensor, torch.Tensor], torch.Tensor], y_true: np.ndarray, y_pred: np.ndarray\n):\n    \"\"\"\n    Perform automatic differentiation to get the\n    Gradient and the Hessian of `loss_function`.\n    \"\"\"\n    y_true = torch.tensor(y_true, dtype=torch.float, requires_grad=False)\n    y_pred = torch.tensor(y_pred, dtype=torch.float, requires_grad=True)\n    loss_function_sum = lambda y_pred: loss_function(y_true, y_pred).sum()\n\n    loss_function_sum(y_pred).backward()\n    grad = y_pred.grad.reshape(-1)\n\n    # hess_matrix = torch.autograd.functional.hessian(loss_function_sum, y_pred, vectorize=True)\n    # hess = torch.diagonal(hess_matrix)\n    hess = np.ones(grad.shape)\n\n    return grad, hess\n\n\ncustom_objective = partial(torch_autodiff_grad_hess, pearson_cc_loss)\n\n\nxgb_params = dict(\n    n_estimators=10000,\n    early_stopping_rounds=20,\n    # learning_rate=0.05,\n    objective=custom_objective,  # \"binary:logistic\", \"reg:squarederror\",\n    eval_metric=pearson_cc_xgb_score,  # \"logloss\", \"rmse\",\n    random_state=440,\n    tree_method=\"gpu_hist\",\n)  # type: dict[str, Any]\n\nclf = xgb.XGBRegressor(**xgb_params)\n```\n\n\n## 🏷️ Links\n\n- [My Solution](https://github.com/IMOKURI/kaggle-multimodal-single-cell-integration)",
      "votes": null
    },
    {
      "id": "2031455",
      "postDate": "11/16/2022 05:04:57",
      "content": "<p><a href=\"https://www.kaggle.com/imokuri\" target=\"_blank\">@imokuri</a> Thank You for sharing you Solution! Best wishes!</p>",
      "rawMarkdown": "imokuri Thank You for sharing you Solution! Best wishes!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2031455,
      "author_name": "akmalmir",
      "author_url": "",
      "post_date": "11/16/2022 05:04:57",
      "content": "<p><a href=\"https://www.kaggle.com/imokuri\" target=\"_blank\">@imokuri</a> Thank You for sharing you Solution! Best wishes!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2031222": "Thank you for all kagglers and organizers for this competition.\n\nThis note is a record of my work on the competition.\n\nThis was a competition where the Private test data was from a future date that did not exist in the training data, and the so-called domain generalization performance was being tested.\nOn the other hand, there was an element of variation by date, and it was expected that it would be undesirable to completely ignore the date feature.\n\nFirst, we conducted adversarial training (a task to classify training and test data) and found that Citeseq was capable of 99% classification, and we were concerned that training with this feature set would result in overtraining on the training data.\nHowever, when we reduced the number of features to reduce the accuracy of Adversarial training, the score of Public LB also dropped significantly.\n\nTherefore, we decided to devise some kind of biological features and to improve generalization performance through model variation.\n\n\n## ✨ Result\n\n- Private: 0.769\n- Public: 0.813\n\n\n\n## 🖼️ Solution\n\n\n### 🌱 Preprocess\n\n- Citeseq\n    - The input data was reduced to 100 dimensions by PCA.\n    - On the other hand, the data of important columns were preserved.\n    - [Ivis unsupervised learning](https://bering-ivis.readthedocs.io/en/latest/unsupervised.html) was used to generate 100 dimensional features.\n    - In addition, we added the sum of mitochondrial RNA cells to the features.\n    - Cell type in Metadata was added to the features.\n\n- Multiome\n    - For each group with the same column name prefix, PCA reduced the number of dimensions to approximately 100 each.\n    - Ivis unsupervised learning was used to generate 100 dimensional features.\n\n### 🤸 Pre Training\n\n- Adversarial training (a task to classify training data and test data) is performed and the misjudged training data is used as good validation data.\n- Prediction of Cell type for Multiome is performed and added to the features.\n\n\n### 🏃 Training\n\n- StratifiedKFold with good validation data as positive labels.\n- Pearson correlation coefficient was used for the Loss function. XGBoost was implemented as described below.\n- TabNet also performed pre-training. (In this competition, pre-training was more accurate.)\n\n### 🎨 Base Models\n\n- Citeseq\n    - TabNet\n    - Simple MLP\n    - ResNet\n    - 1D CNN\n    - XGBoost\n- Multiome\n    - 1D CNN\n\n\nCiteseq scored well with an ensemble of various models.\nOn the other hand, Multiome had a strong 1D CNN and did not score well with ensembles of other models, so only the 1D CNN was used.\n\n### 🚀 Postprocess\n\n- Since the evaluation metric is the Pearson correlation coefficient, each inference result (including OOF results) was normalized before ensemble.\n- Optuna was used to optimize the ensemble weights. Good validation data was used as the evaluation metric.\n- Ensemble with Public Notebook x2 and teammate submissions.\n\n\n## 💡 Tips\n\n\n### Pearson Loss for XGBoost\n\nXGBoost does not provide a Pearson Loss Function, so I implemented it as follows.\nHowever, this implementation is slow in learning, and I have the impression that I would like to improve it a little more.\n\n```\nfrom functools import partial\nfrom typing import Any, Callable\n\nimport numpy as np\nimport torch\nimport torch.nn.functional as F\nimport xgboost as xgb\n\n\ndef pearson_cc_loss(inputs, targets):\n    try:\n        assert inputs.shape == targets.shape\n    except AssertionError:\n        inputs = inputs.view(targets.shape)\n\n    pcc = F.cosine_similarity(inputs, targets)\n    return 1.0 - pcc\n\n\n# https://towardsdatascience.com/jax-vs-pytorch-automatic-differentiation-for-xgboost-10222e1404ec\ndef torch_autodiff_grad_hess(\n    loss_function: Callable[[torch.Tensor, torch.Tensor], torch.Tensor], y_true: np.ndarray, y_pred: np.ndarray\n):\n    \"\"\"\n    Perform automatic differentiation to get the\n    Gradient and the Hessian of `loss_function`.\n    \"\"\"\n    y_true = torch.tensor(y_true, dtype=torch.float, requires_grad=False)\n    y_pred = torch.tensor(y_pred, dtype=torch.float, requires_grad=True)\n    loss_function_sum = lambda y_pred: loss_function(y_true, y_pred).sum()\n\n    loss_function_sum(y_pred).backward()\n    grad = y_pred.grad.reshape(-1)\n\n    # hess_matrix = torch.autograd.functional.hessian(loss_function_sum, y_pred, vectorize=True)\n    # hess = torch.diagonal(hess_matrix)\n    hess = np.ones(grad.shape)\n\n    return grad, hess\n\n\ncustom_objective = partial(torch_autodiff_grad_hess, pearson_cc_loss)\n\n\nxgb_params = dict(\n    n_estimators=10000,\n    early_stopping_rounds=20,\n    # learning_rate=0.05,\n    objective=custom_objective,  # \"binary:logistic\", \"reg:squarederror\",\n    eval_metric=pearson_cc_xgb_score,  # \"logloss\", \"rmse\",\n    random_state=440,\n    tree_method=\"gpu_hist\",\n)  # type: dict[str, Any]\n\nclf = xgb.XGBRegressor(**xgb_params)\n```\n\n\n## 🏷️ Links\n\n- [My Solution](https://github.com/IMOKURI/kaggle-multimodal-single-cell-integration)",
    "2031455": "imokuri Thank You for sharing you Solution! Best wishes!"
  },
  "source": "meta"
}