{
  "id": 366503,
  "title": "Up 580 positions on lb and a bronze medal for a simple catboost solution",
  "url": "/competitions/open-problems-multimodal/writeups/artem-fedorov-up-580-positions-on-lb-and-a-bronze-",
  "author_name": "",
  "post_date": "2022-11-16T12:55:50.228640300Z",
  "votes": 16,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I guess most people here understand that significant moves on lb cause emotions. Right now I feel excited with the results.</p>\n<p>My plan was to select and engineer features using catboost models and then create NN models, but a week before competition end I understood I am out of time and out of kaggle accelerator quota to test NN models anyway, so I decided to make the improvements I can and upload whatever result I am able to achieve, even if it is lower then the bulk of competition results. <br>\nThis morning the private lb showed my models aren't that bad.</p>\n<p>Now a few points about the data, my models and how I worked on it.</p>\n<ul>\n<li>spent a lost of time trying to make some usefull features from raw data. All those efforts turned out to be fruitless, as the results were better without them.</li>\n<li>I could not understand why so many notebooks published here on kaggle have no features made from metadata, I mean donors and days. I made one-hot vectors for donors and initially also made one-hot vectors for days. But I read what <a href=\"https://www.kaggle.com/AmbrosM\" target=\"_blank\">@AmbrosM</a> wrote about data being time series and tested one-hot vectors for days versus just an int value for days (tested on Multiome data only and then changed the days features to int both for cite and multiome). I guess that was an important reason why my private lb results turned out to be than much better than the public lb results.</li>\n<li>put a lot of effort into selecting best features for CITE model. By 10 batches (2000 features in a batch) I tried all non-constant CITE inputs and selected about 600 features total from all batches with highest feature_importance_</li>\n<li>also tried to select some meaningfull features for multiome among those we were given as inputs, but here I had to pre-select correlating features. Out of about 1500 pre-selected features I have found three that were as important as 30's SVD components, while all the other features were less important than any SVD components. ['svd_x_chr1:630875-631689', 'svd_x_chr1:633700-634539', 'svd_x_chr17:22520955-22521852'] </li>\n<li>had an idea to fit a second level linear model for all multiome targets individually, but finally didn't even try</li>\n<li>divided all the cite targets into groups and used catboost models with different parameters. Increased learning rates and iterations for the best models and decreased these parameters for the worst ones. Also used stronger parameters for catboost models calculating first components of target SVD's in multiome subtask.</li>\n<li>in a number of models published on kaggle I saw that people calculate SVD on train inputs only, and then just use transform for test, instead of combing train and test inputs and running fit_transform on both. I guess those people saw their positions dropped on lb after the competition end.</li>\n</ul>",
  "messages": [
    {
      "id": "2032137",
      "postDate": "11/16/2022 12:55:50",
      "content": "<p>I guess most people here understand that significant moves on lb cause emotions. Right now I feel excited with the results.</p>\n<p>My plan was to select and engineer features using catboost models and then create NN models, but a week before competition end I understood I am out of time and out of kaggle accelerator quota to test NN models anyway, so I decided to make the improvements I can and upload whatever result I am able to achieve, even if it is lower then the bulk of competition results. <br>\nThis morning the private lb showed my models aren't that bad.</p>\n<p>Now a few points about the data, my models and how I worked on it.</p>\n<ul>\n<li>spent a lost of time trying to make some usefull features from raw data. All those efforts turned out to be fruitless, as the results were better without them.</li>\n<li>I could not understand why so many notebooks published here on kaggle have no features made from metadata, I mean donors and days. I made one-hot vectors for donors and initially also made one-hot vectors for days. But I read what <a href=\"https://www.kaggle.com/AmbrosM\" target=\"_blank\">@AmbrosM</a> wrote about data being time series and tested one-hot vectors for days versus just an int value for days (tested on Multiome data only and then changed the days features to int both for cite and multiome). I guess that was an important reason why my private lb results turned out to be than much better than the public lb results.</li>\n<li>put a lot of effort into selecting best features for CITE model. By 10 batches (2000 features in a batch) I tried all non-constant CITE inputs and selected about 600 features total from all batches with highest feature_importance_</li>\n<li>also tried to select some meaningfull features for multiome among those we were given as inputs, but here I had to pre-select correlating features. Out of about 1500 pre-selected features I have found three that were as important as 30's SVD components, while all the other features were less important than any SVD components. ['svd_x_chr1:630875-631689', 'svd_x_chr1:633700-634539', 'svd_x_chr17:22520955-22521852'] </li>\n<li>had an idea to fit a second level linear model for all multiome targets individually, but finally didn't even try</li>\n<li>divided all the cite targets into groups and used catboost models with different parameters. Increased learning rates and iterations for the best models and decreased these parameters for the worst ones. Also used stronger parameters for catboost models calculating first components of target SVD's in multiome subtask.</li>\n<li>in a number of models published on kaggle I saw that people calculate SVD on train inputs only, and then just use transform for test, instead of combing train and test inputs and running fit_transform on both. I guess those people saw their positions dropped on lb after the competition end.</li>\n</ul>",
      "rawMarkdown": "I guess most people here understand that significant moves on lb cause emotions. Right now I feel excited with the results.\n\nMy plan was to select and engineer features using catboost models and then create NN models, but a week before competition end I understood I am out of time and out of kaggle accelerator quota to test NN models anyway, so I decided to make the improvements I can and upload whatever result I am able to achieve, even if it is lower then the bulk of competition results. \nThis morning the private lb showed my models aren't that bad.\n\nNow a few points about the data, my models and how I worked on it.\n- spent a lost of time trying to make some usefull features from raw data. All those efforts turned out to be fruitless, as the results were better without them.\n- I could not understand why so many notebooks published here on kaggle have no features made from metadata, I mean donors and days. I made one-hot vectors for donors and initially also made one-hot vectors for days. But I read what @AmbrosM wrote about data being time series and tested one-hot vectors for days versus just an int value for days (tested on Multiome data only and then changed the days features to int both for cite and multiome). I guess that was an important reason why my private lb results turned out to be than much better than the public lb results.\n- put a lot of effort into selecting best features for CITE model. By 10 batches (2000 features in a batch) I tried all non-constant CITE inputs and selected about 600 features total from all batches with highest feature_importance_\n- also tried to select some meaningfull features for multiome among those we were given as inputs, but here I had to pre-select correlating features. Out of about 1500 pre-selected features I have found three that were as important as 30's SVD components, while all the other features were less important than any SVD components. ['svd_x_chr1:630875-631689', 'svd_x_chr1:633700-634539', 'svd_x_chr17:22520955-22521852'] \n- had an idea to fit a second level linear model for all multiome targets individually, but finally didn't even try\n- divided all the cite targets into groups and used catboost models with different parameters. Increased learning rates and iterations for the best models and decreased these parameters for the worst ones. Also used stronger parameters for catboost models calculating first components of target SVD's in multiome subtask.\n- in a number of models published on kaggle I saw that people calculate SVD on train inputs only, and then just use transform for test, instead of combing train and test inputs and running fit_transform on both. I guess those people saw their positions dropped on lb after the competition end.",
      "votes": null
    },
    {
      "id": "2032876",
      "postDate": "11/16/2022 21:02:08",
      "content": "<p>This is great! This is testimony to the efficacy of simple approaches! Congratulations <a href=\"https://www.kaggle.com/artemfedorov\" target=\"_blank\">@artemfedorov</a>! All the best…</p>",
      "rawMarkdown": "This is great! This is testimony to the efficacy of simple approaches! Congratulations @artemfedorov! All the best...",
      "votes": null
    },
    {
      "id": "2077333",
      "postDate": "12/27/2022 13:34:53",
      "content": "<p>Congratulations on the medal and thank you for the post, <a href=\"https://www.kaggle.com/artemfedorov\" target=\"_blank\">@artemfedorov</a> ! Can you also clarify if you used other metadata features (e.g. gender or cell type), and what CV strategy did you use? </p>\n<p>I am a part of a team, which analyses the results of the competition to find something useful for bioinformatics. Your answers would help a lot :)</p>",
      "rawMarkdown": "Congratulations on the medal and thank you for the post, @artemfedorov ! Can you also clarify if you used other metadata features (e.g. gender or cell type), and what CV strategy did you use? \n\nI am a part of a team, which analyses the results of the competition to find something useful for bioinformatics. Your answers would help a lot :)",
      "votes": null
    },
    {
      "id": "2200417",
      "postDate": "03/28/2023 14:36:23",
      "content": "<p>Didn't notice your comment. Kaggle doesn't notify about replies, so I've missed your question. I understand the question is most likely no longer relevant, but I'll answer anyway.</p>\n<ol>\n<li>As for cross-validation I started with keeping out either one day or one donor, but very soon noticed that keeping out a female donor made no sense as results were so much worse, than in case if I kept out a male donor. So, most of time I worked with a 5-fold cross-validation, keeping out one of days or one of male donors. Close to the end of competition I used a 3-fold cross-validation for CITE, keeping one day out, and 2-fold cross-validation for multiome, keeping out either first or last day of train data.<br>\nThe primary reason was to reduce the accelaration quota usage, and I also wanted to focus on getting better score, and this meant predicting the last day of test data better.</li>\n<li>I didn't use cell type as metadata since this information was not available for the test dataset. Final submission used information about all donors present in train as one-hot vector (so, one of those features actually was a sex feature as only one donor was female) and information about day as an integer. Day and sex metadata features were often among the most important features for most of 140 CITE targets. But there were some targets for which either day or sex were not important.</li>\n</ol>",
      "rawMarkdown": "Didn't notice your comment. Kaggle doesn't notify about replies, so I've missed your question. I understand the question is most likely no longer relevant, but I'll answer anyway.\n\n1. As for cross-validation I started with keeping out either one day or one donor, but very soon noticed that keeping out a female donor made no sense as results were so much worse, than in case if I kept out a male donor. So, most of time I worked with a 5-fold cross-validation, keeping out one of days or one of male donors. Close to the end of competition I used a 3-fold cross-validation for CITE, keeping one day out, and 2-fold cross-validation for multiome, keeping out either first or last day of train data.\nThe primary reason was to reduce the accelaration quota usage, and I also wanted to focus on getting better score, and this meant predicting the last day of test data better.\n2. I didn't use cell type as metadata since this information was not available for the test dataset. Final submission used information about all donors present in train as one-hot vector (so, one of those features actually was a sex feature as only one donor was female) and information about day as an integer. Day and sex metadata features were often among the most important features for most of 140 CITE targets. But there were some targets for which either day or sex were not important.",
      "votes": null
    },
    {
      "id": "2200429",
      "postDate": "03/28/2023 14:47:17",
      "content": "<p>Thank you, this is super relevant!</p>",
      "rawMarkdown": "Thank you, this is super relevant!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2032876,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "11/16/2022 21:02:08",
      "content": "<p>This is great! This is testimony to the efficacy of simple approaches! Congratulations <a href=\"https://www.kaggle.com/artemfedorov\" target=\"_blank\">@artemfedorov</a>! All the best…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2077333,
      "author_name": "shitovvladimir",
      "author_url": "",
      "post_date": "12/27/2022 13:34:53",
      "content": "<p>Congratulations on the medal and thank you for the post, <a href=\"https://www.kaggle.com/artemfedorov\" target=\"_blank\">@artemfedorov</a> ! Can you also clarify if you used other metadata features (e.g. gender or cell type), and what CV strategy did you use? </p>\n<p>I am a part of a team, which analyses the results of the competition to find something useful for bioinformatics. Your answers would help a lot :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2200417,
          "author_name": "artemfedorov",
          "author_url": "",
          "post_date": "03/28/2023 14:36:23",
          "content": "<p>Didn't notice your comment. Kaggle doesn't notify about replies, so I've missed your question. I understand the question is most likely no longer relevant, but I'll answer anyway.</p>\n<ol>\n<li>As for cross-validation I started with keeping out either one day or one donor, but very soon noticed that keeping out a female donor made no sense as results were so much worse, than in case if I kept out a male donor. So, most of time I worked with a 5-fold cross-validation, keeping out one of days or one of male donors. Close to the end of competition I used a 3-fold cross-validation for CITE, keeping one day out, and 2-fold cross-validation for multiome, keeping out either first or last day of train data.<br>\nThe primary reason was to reduce the accelaration quota usage, and I also wanted to focus on getting better score, and this meant predicting the last day of test data better.</li>\n<li>I didn't use cell type as metadata since this information was not available for the test dataset. Final submission used information about all donors present in train as one-hot vector (so, one of those features actually was a sex feature as only one donor was female) and information about day as an integer. Day and sex metadata features were often among the most important features for most of 140 CITE targets. But there were some targets for which either day or sex were not important.</li>\n</ol>",
          "votes": null,
          "replies": [
            {
              "id": 2200429,
              "author_name": "shitovvladimir",
              "author_url": "",
              "post_date": "03/28/2023 14:47:17",
              "content": "<p>Thank you, this is super relevant!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2032137": "I guess most people here understand that significant moves on lb cause emotions. Right now I feel excited with the results.\n\nMy plan was to select and engineer features using catboost models and then create NN models, but a week before competition end I understood I am out of time and out of kaggle accelerator quota to test NN models anyway, so I decided to make the improvements I can and upload whatever result I am able to achieve, even if it is lower then the bulk of competition results. \nThis morning the private lb showed my models aren't that bad.\n\nNow a few points about the data, my models and how I worked on it.\n- spent a lost of time trying to make some usefull features from raw data. All those efforts turned out to be fruitless, as the results were better without them.\n- I could not understand why so many notebooks published here on kaggle have no features made from metadata, I mean donors and days. I made one-hot vectors for donors and initially also made one-hot vectors for days. But I read what @AmbrosM wrote about data being time series and tested one-hot vectors for days versus just an int value for days (tested on Multiome data only and then changed the days features to int both for cite and multiome). I guess that was an important reason why my private lb results turned out to be than much better than the public lb results.\n- put a lot of effort into selecting best features for CITE model. By 10 batches (2000 features in a batch) I tried all non-constant CITE inputs and selected about 600 features total from all batches with highest feature_importance_\n- also tried to select some meaningfull features for multiome among those we were given as inputs, but here I had to pre-select correlating features. Out of about 1500 pre-selected features I have found three that were as important as 30's SVD components, while all the other features were less important than any SVD components. ['svd_x_chr1:630875-631689', 'svd_x_chr1:633700-634539', 'svd_x_chr17:22520955-22521852'] \n- had an idea to fit a second level linear model for all multiome targets individually, but finally didn't even try\n- divided all the cite targets into groups and used catboost models with different parameters. Increased learning rates and iterations for the best models and decreased these parameters for the worst ones. Also used stronger parameters for catboost models calculating first components of target SVD's in multiome subtask.\n- in a number of models published on kaggle I saw that people calculate SVD on train inputs only, and then just use transform for test, instead of combing train and test inputs and running fit_transform on both. I guess those people saw their positions dropped on lb after the competition end.",
    "2032876": "This is great! This is testimony to the efficacy of simple approaches! Congratulations @artemfedorov! All the best...",
    "2077333": "Congratulations on the medal and thank you for the post, @artemfedorov ! Can you also clarify if you used other metadata features (e.g. gender or cell type), and what CV strategy did you use? \n\nI am a part of a team, which analyses the results of the competition to find something useful for bioinformatics. Your answers would help a lot :)",
    "2200417": "Didn't notice your comment. Kaggle doesn't notify about replies, so I've missed your question. I understand the question is most likely no longer relevant, but I'll answer anyway.\n\n1. As for cross-validation I started with keeping out either one day or one donor, but very soon noticed that keeping out a female donor made no sense as results were so much worse, than in case if I kept out a male donor. So, most of time I worked with a 5-fold cross-validation, keeping out one of days or one of male donors. Close to the end of competition I used a 3-fold cross-validation for CITE, keeping one day out, and 2-fold cross-validation for multiome, keeping out either first or last day of train data.\nThe primary reason was to reduce the accelaration quota usage, and I also wanted to focus on getting better score, and this meant predicting the last day of test data better.\n2. I didn't use cell type as metadata since this information was not available for the test dataset. Final submission used information about all donors present in train as one-hot vector (so, one of those features actually was a sex feature as only one donor was female) and information about day as an integer. Day and sex metadata features were often among the most important features for most of 140 CITE targets. But there were some targets for which either day or sex were not important.",
    "2200429": "Thank you, this is super relevant!"
  },
  "source": "meta"
}