{
  "id": 366395,
  "title": "Private 0.773 lessons learned",
  "url": "/competitions/open-problems-multimodal/discussion/366395",
  "author_name": "",
  "post_date": "2022-11-16T02:42:27.959850600Z",
  "votes": 76,
  "comment_count": 23,
  "views": 0,
  "content": "<p><em>The most important result of every endeavor is what we learn from it…</em></p>\n<h1>Time series and extrapolation</h1>\n<p>The competition's data is a time series and the private test data is an extrapolation problem (because the unseen day is in the future): Most people know that decision trees can't extrapolate, but k-nearest-neighbors can't extrapolate either, and rbf kernel ridge regression is even worse (it predicts 0 if you extrapolate far enough).</p>\n<p>Having a good cross-validation strategy is crucial: Use a GroupKFold on days for validation (with particular emphasis on the last day) but train the submission models with a ShuffleSplit.</p>\n<p>Because the model changes every day from day 2 to day 10, a good strategy for the Multiome public leaderboard consists of training on only the data of the day we want to predict.</p>\n<p>Unfortunately we looked too much at the public lb when selecting our final two submissions and got shaken down:</p>\n<p><img src=\"https://i.imgur.com/f8WqyO3.png\" alt=\"0.773\"></p>\n<h1>The log1p transformation for gene expressions is wrong</h1>\n<p>The gene expressions in this competition were presented library-size normalized and log1p-transformed. The log1p-transformation is bad: If the raw counts are 0, 1, 2, 3, 4, 5, 6, 7, normalization transforms them into 0, 100, 200, 300, 400, 500, 600, 700, and log1p transforms them into 0, 4.6, 5.3, 5.7, 6.0, 6.2, 6.4, 6.6. The transformed 0 and 1 are far away from each other although a difference of 1 in the raw counts is pure randomness. Clipping the CITEseq inputs with <code>clip(3, None)</code> before the SVD improves the predictions.</p>\n<p><img src=\"https://i.imgur.com/iOpVRmj.png\" alt=\"log1p diagram\"></p>\n<h1>The metric</h1>\n<p>Understanding the metric is important in every competition, but this competition had the peculiarity that the correlation metric can be optimized directly only by neural networks. For all other machine learning methods, the best we can do is optimizing for mean squared error. The points to know here are:</p>\n<ol>\n<li>If we optimize MSE as a surrogate for the correlation metric, we need to standardize the targets: <code>Y = Y / Y.std(axis=1).reshape(-1, 1)</code> Neglecting this standardization corresponds to optimizing for random sample weights. The same standardization is necessary before ensembling.</li>\n<li>It can be shown that if we apply this standardization correctly, the optimum model for mse loss is equivalent to the optimum model for correlation loss.</li>\n<li>MSE loss is invariant under PCA and SVD transformations, which multiply all vectors with an orthogonal matrix (so that the distances are conserved). Only this invariance legitimates the transformation and dimensionality reduction of the targets.</li>\n</ol>\n<h1>Noise and dimensionality reduction</h1>\n<p>Understanding the signal-to-noise ratio of the data and its dimensionality is important. More than 90 % of the ATACseq variance and more than 80 % of the gene expression variance seem to be noise (i.e. useless as model input and unpredictable as model output).</p>\n<p>As all three datasets live in almost linear subspaces, the dimensionality of the data can be reduced with a PCA or SVD. As a side effect, the dimensionality reduction improves the signal-to-noise ratio by concentrating the signal in the first few components (and moving the noise to the other components).</p>\n<p>The data seems to live in manifolds of the following dimensions:</p>\n<ul>\n<li>ATACseq: between 100 and 140 of 200000</li>\n<li>Gene expressions: between 70 and 130 of 20000</li>\n<li>Proteins: about 90 of 140</li>\n</ul>\n<p>The dimensionality of the input for both problems (CITEseq and Multiome) has to be reduced for all models - models with 20000 or even 200000 features are not feasible. For the prediction targets, the situation is different:</p>\n<ul>\n<li>The CITEseq proteins were best predicted when the 140 dimensions were reduced to 90, but predicting all 140 proteins directly is possible and adds diversity to the ensemble.</li>\n<li>Because of space and time constraints, most machine learning algorithms cannot predict the 20000 gene expressions in Multiome at once and need a dimensionality reduction for the output. The notable exception here is the k-nearest-neighbors algorithm which predicts 20000 targets in reasonable time. The knn predictions are somewhat noisy and profit from a postprocessing which constrains them into a 75-dimensional manifold.</li>\n</ul>\n<h1>Feature selection</h1>\n<p>In CITEseq, many of the 20000 features are pure noise. If we select a few hundred important features and give them to a machine learning algorithm, the result is better than if we only use an SVD of all 20000 features.</p>\n<p>For Multiome, I had no success in determining the most important features, but if you drop the whole Y chromosome, the model's score remains the same. This suggests that the Y chromosome has no important function in blood cells.</p>\n<h1>Using RAM efficiently</h1>\n<p>With 100 GBytes of data, probably every participant has learnt how to use the RAM efficiently:</p>\n<ul>\n<li>It is good to understand when numpy copies arrays and when it returns views.</li>\n<li>It is good to understand the memory complexity of algorithms. ExtraTrees consumes memory proportional to n_estimators and to the number of leaves per tree, kernel ridge regression with the rbf kernel computes a distance matrix which is quadratic in n_samples, etc.</li>\n<li>In the end I even coded a ridge regression for the Multiome ensemble which works without loading all the data into memory at the same time.</li>\n</ul>\n<h1>Teamwork</h1>\n<p>I'm happy that I could be part of a good team with <a href=\"https://www.kaggle.com/mtinti\" target=\"_blank\">@mtinti</a>, <a href=\"https://www.kaggle.com/callmeb\" target=\"_blank\">@callmeb</a> and <a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a>: We had a fruitful exchange of ideas (communicating through Slack), all team members contributed their best models to a very diverse final ensemble, and the virtual presence of teammembers motivates everybody during phases of stagnating scores.</p>",
  "messages": [
    {
      "id": "2031346",
      "postDate": "11/16/2022 02:42:27",
      "content": "<p><em>The most important result of every endeavor is what we learn from it…</em></p>\n<h1>Time series and extrapolation</h1>\n<p>The competition's data is a time series and the private test data is an extrapolation problem (because the unseen day is in the future): Most people know that decision trees can't extrapolate, but k-nearest-neighbors can't extrapolate either, and rbf kernel ridge regression is even worse (it predicts 0 if you extrapolate far enough).</p>\n<p>Having a good cross-validation strategy is crucial: Use a GroupKFold on days for validation (with particular emphasis on the last day) but train the submission models with a ShuffleSplit.</p>\n<p>Because the model changes every day from day 2 to day 10, a good strategy for the Multiome public leaderboard consists of training on only the data of the day we want to predict.</p>\n<p>Unfortunately we looked too much at the public lb when selecting our final two submissions and got shaken down:</p>\n<p><img src=\"https://i.imgur.com/f8WqyO3.png\" alt=\"0.773\"></p>\n<h1>The log1p transformation for gene expressions is wrong</h1>\n<p>The gene expressions in this competition were presented library-size normalized and log1p-transformed. The log1p-transformation is bad: If the raw counts are 0, 1, 2, 3, 4, 5, 6, 7, normalization transforms them into 0, 100, 200, 300, 400, 500, 600, 700, and log1p transforms them into 0, 4.6, 5.3, 5.7, 6.0, 6.2, 6.4, 6.6. The transformed 0 and 1 are far away from each other although a difference of 1 in the raw counts is pure randomness. Clipping the CITEseq inputs with <code>clip(3, None)</code> before the SVD improves the predictions.</p>\n<p><img src=\"https://i.imgur.com/iOpVRmj.png\" alt=\"log1p diagram\"></p>\n<h1>The metric</h1>\n<p>Understanding the metric is important in every competition, but this competition had the peculiarity that the correlation metric can be optimized directly only by neural networks. For all other machine learning methods, the best we can do is optimizing for mean squared error. The points to know here are:</p>\n<ol>\n<li>If we optimize MSE as a surrogate for the correlation metric, we need to standardize the targets: <code>Y = Y / Y.std(axis=1).reshape(-1, 1)</code> Neglecting this standardization corresponds to optimizing for random sample weights. The same standardization is necessary before ensembling.</li>\n<li>It can be shown that if we apply this standardization correctly, the optimum model for mse loss is equivalent to the optimum model for correlation loss.</li>\n<li>MSE loss is invariant under PCA and SVD transformations, which multiply all vectors with an orthogonal matrix (so that the distances are conserved). Only this invariance legitimates the transformation and dimensionality reduction of the targets.</li>\n</ol>\n<h1>Noise and dimensionality reduction</h1>\n<p>Understanding the signal-to-noise ratio of the data and its dimensionality is important. More than 90 % of the ATACseq variance and more than 80 % of the gene expression variance seem to be noise (i.e. useless as model input and unpredictable as model output).</p>\n<p>As all three datasets live in almost linear subspaces, the dimensionality of the data can be reduced with a PCA or SVD. As a side effect, the dimensionality reduction improves the signal-to-noise ratio by concentrating the signal in the first few components (and moving the noise to the other components).</p>\n<p>The data seems to live in manifolds of the following dimensions:</p>\n<ul>\n<li>ATACseq: between 100 and 140 of 200000</li>\n<li>Gene expressions: between 70 and 130 of 20000</li>\n<li>Proteins: about 90 of 140</li>\n</ul>\n<p>The dimensionality of the input for both problems (CITEseq and Multiome) has to be reduced for all models - models with 20000 or even 200000 features are not feasible. For the prediction targets, the situation is different:</p>\n<ul>\n<li>The CITEseq proteins were best predicted when the 140 dimensions were reduced to 90, but predicting all 140 proteins directly is possible and adds diversity to the ensemble.</li>\n<li>Because of space and time constraints, most machine learning algorithms cannot predict the 20000 gene expressions in Multiome at once and need a dimensionality reduction for the output. The notable exception here is the k-nearest-neighbors algorithm which predicts 20000 targets in reasonable time. The knn predictions are somewhat noisy and profit from a postprocessing which constrains them into a 75-dimensional manifold.</li>\n</ul>\n<h1>Feature selection</h1>\n<p>In CITEseq, many of the 20000 features are pure noise. If we select a few hundred important features and give them to a machine learning algorithm, the result is better than if we only use an SVD of all 20000 features.</p>\n<p>For Multiome, I had no success in determining the most important features, but if you drop the whole Y chromosome, the model's score remains the same. This suggests that the Y chromosome has no important function in blood cells.</p>\n<h1>Using RAM efficiently</h1>\n<p>With 100 GBytes of data, probably every participant has learnt how to use the RAM efficiently:</p>\n<ul>\n<li>It is good to understand when numpy copies arrays and when it returns views.</li>\n<li>It is good to understand the memory complexity of algorithms. ExtraTrees consumes memory proportional to n_estimators and to the number of leaves per tree, kernel ridge regression with the rbf kernel computes a distance matrix which is quadratic in n_samples, etc.</li>\n<li>In the end I even coded a ridge regression for the Multiome ensemble which works without loading all the data into memory at the same time.</li>\n</ul>\n<h1>Teamwork</h1>\n<p>I'm happy that I could be part of a good team with <a href=\"https://www.kaggle.com/mtinti\" target=\"_blank\">@mtinti</a>, <a href=\"https://www.kaggle.com/callmeb\" target=\"_blank\">@callmeb</a> and <a href=\"https://www.kaggle.com/khahuras\" target=\"_blank\">@khahuras</a>: We had a fruitful exchange of ideas (communicating through Slack), all team members contributed their best models to a very diverse final ensemble, and the virtual presence of teammembers motivates everybody during phases of stagnating scores.</p>",
      "rawMarkdown": "*The most important result of every endeavor is what we learn from it...*\n\n# Time series and extrapolation\n\nThe competition's data is a time series and the private test data is an extrapolation problem (because the unseen day is in the future): Most people know that decision trees can't extrapolate, but k-nearest-neighbors can't extrapolate either, and rbf kernel ridge regression is even worse (it predicts 0 if you extrapolate far enough).\n\nHaving a good cross-validation strategy is crucial: Use a GroupKFold on days for validation (with particular emphasis on the last day) but train the submission models with a ShuffleSplit.\n\nBecause the model changes every day from day 2 to day 10, a good strategy for the Multiome public leaderboard consists of training on only the data of the day we want to predict.\n \nUnfortunately we looked too much at the public lb when selecting our final two submissions and got shaken down:\n\n![0.773](https://i.imgur.com/f8WqyO3.png)\n\n# The log1p transformation for gene expressions is wrong\n\nThe gene expressions in this competition were presented library-size normalized and log1p-transformed. The log1p-transformation is bad: If the raw counts are 0, 1, 2, 3, 4, 5, 6, 7, normalization transforms them into 0, 100, 200, 300, 400, 500, 600, 700, and log1p transforms them into 0, 4.6, 5.3, 5.7, 6.0, 6.2, 6.4, 6.6. The transformed 0 and 1 are far away from each other although a difference of 1 in the raw counts is pure randomness. Clipping the CITEseq inputs with `clip(3, None)` before the SVD improves the predictions.\n\n![log1p diagram](https://i.imgur.com/iOpVRmj.png)\n\n# The metric\n\nUnderstanding the metric is important in every competition, but this competition had the peculiarity that the correlation metric can be optimized directly only by neural networks. For all other machine learning methods, the best we can do is optimizing for mean squared error. The points to know here are:\n1. If we optimize MSE as a surrogate for the correlation metric, we need to standardize the targets: `Y = Y / Y.std(axis=1).reshape(-1, 1)` Neglecting this standardization corresponds to optimizing for random sample weights. The same standardization is necessary before ensembling.\n2. It can be shown that if we apply this standardization correctly, the optimum model for mse loss is equivalent to the optimum model for correlation loss.\n3. MSE loss is invariant under PCA and SVD transformations, which multiply all vectors with an orthogonal matrix (so that the distances are conserved). Only this invariance legitimates the transformation and dimensionality reduction of the targets.\n\n# Noise and dimensionality reduction\n\nUnderstanding the signal-to-noise ratio of the data and its dimensionality is important. More than 90 % of the ATACseq variance and more than 80 % of the gene expression variance seem to be noise (i.e. useless as model input and unpredictable as model output).\n\nAs all three datasets live in almost linear subspaces, the dimensionality of the data can be reduced with a PCA or SVD. As a side effect, the dimensionality reduction improves the signal-to-noise ratio by concentrating the signal in the first few components (and moving the noise to the other components).\n\nThe data seems to live in manifolds of the following dimensions:\n- ATACseq: between 100 and 140 of 200000\n- Gene expressions: between 70 and 130 of 20000\n- Proteins: about 90 of 140\n\nThe dimensionality of the input for both problems (CITEseq and Multiome) has to be reduced for all models - models with 20000 or even 200000 features are not feasible. For the prediction targets, the situation is different:\n- The CITEseq proteins were best predicted when the 140 dimensions were reduced to 90, but predicting all 140 proteins directly is possible and adds diversity to the ensemble.\n- Because of space and time constraints, most machine learning algorithms cannot predict the 20000 gene expressions in Multiome at once and need a dimensionality reduction for the output. The notable exception here is the k-nearest-neighbors algorithm which predicts 20000 targets in reasonable time. The knn predictions are somewhat noisy and profit from a postprocessing which constrains them into a 75-dimensional manifold.\n\n# Feature selection\n\nIn CITEseq, many of the 20000 features are pure noise. If we select a few hundred important features and give them to a machine learning algorithm, the result is better than if we only use an SVD of all 20000 features.\n\nFor Multiome, I had no success in determining the most important features, but if you drop the whole Y chromosome, the model's score remains the same. This suggests that the Y chromosome has no important function in blood cells.\n\n# Using RAM efficiently\n\nWith 100 GBytes of data, probably every participant has learnt how to use the RAM efficiently:\n- It is good to understand when numpy copies arrays and when it returns views.\n- It is good to understand the memory complexity of algorithms. ExtraTrees consumes memory proportional to n_estimators and to the number of leaves per tree, kernel ridge regression with the rbf kernel computes a distance matrix which is quadratic in n_samples, etc.\n- In the end I even coded a ridge regression for the Multiome ensemble which works without loading all the data into memory at the same time.\n  \n# Teamwork\n\nI'm happy that I could be part of a good team with @mtinti, @callmeb and @khahuras: We had a fruitful exchange of ideas (communicating through Slack), all team members contributed their best models to a very diverse final ensemble, and the virtual presence of teammembers motivates everybody during phases of stagnating scores.",
      "votes": null
    },
    {
      "id": "2031354",
      "postDate": "11/16/2022 02:58:46",
      "content": "<p>Nice summary! What are the main differences between your 0.773 and the 0.771 models?</p>",
      "rawMarkdown": "Nice summary! What are the main differences between your 0.773 and the 0.771 models?",
      "votes": null
    },
    {
      "id": "2031454",
      "postDate": "11/16/2022 05:03:54",
      "content": "<p>Very Very thank you, through this competition, we  have learned a lot from you. Respect!!  We are looking forward to your sharing! Respect again!! </p>",
      "rawMarkdown": "Very Very thank you, through this competition, we  have learned a lot from you. Respect!!  We are looking forward to your sharing! Respect again!!",
      "votes": null
    },
    {
      "id": "2031471",
      "postDate": "11/16/2022 05:17:58",
      "content": "<p><a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a> In the 0.773 ensemble we gave much higher ensemble weights to the models with high cv score for Multiome day 7, assuming that day 7 cv is the best predictor for day 10.</p>",
      "rawMarkdown": "kingychiu In the 0.773 ensemble we gave much higher ensemble weights to the models with high cv score for Multiome day 7, assuming that day 7 cv is the best predictor for day 10.",
      "votes": null
    },
    {
      "id": "2031495",
      "postDate": "11/16/2022 05:47:31",
      "content": "<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> - you were certainly an MVP in this competition and shared many insights and important notebooks. Whatever the results you were a gold medal person and think we all learned a lot here.  Felt there were so many permutations of what could be done here and it helped to have someone (and their team) keep a focus on it all and the momentum going.   And importantly to forget the public LB and work for the private LB.</p>\n<p>It would have been interesting to have separate LB for CITEseq and Multiome just to see esp for Multiome.  Spent a lot of time aligning genes and gene locations with target genes and gene ids, and geneactivity.  wrt Multiome it seemed to help to use a reduced set and combining SVDs, using a small number of components for each SVD. Even pursued chromosomes individually and found chr13 was the best, chr16 the worst, and targets had chrM (13 cols) not in the train/test data per se but scored 0.967 on its own just using what was not in reduced.   For chrY there were only 22 targets so perhaps why it did not matter so much especially if these were not included in scores.  The Evaluation Ids csv for scoring was confusing, being a subset for scoring maybe was key to figuring out what was relevant to predict or not, what mattered or not for models.  </p>\n<p>Sklearn MLPRegressor worked surprisingly well and was useful for testing ideas, fast and cpu only. If only there were more time and were on your team!!   Just kidding.  Many many thanks and well done.  </p>",
      "rawMarkdown": "ambrosm - you were certainly an MVP in this competition and shared many insights and important notebooks. Whatever the results you were a gold medal person and think we all learned a lot here.  Felt there were so many permutations of what could be done here and it helped to have someone (and their team) keep a focus on it all and the momentum going.   And importantly to forget the public LB and work for the private LB.\n\nIt would have been interesting to have separate LB for CITEseq and Multiome just to see esp for Multiome.  Spent a lot of time aligning genes and gene locations with target genes and gene ids, and geneactivity.  wrt Multiome it seemed to help to use a reduced set and combining SVDs, using a small number of components for each SVD. Even pursued chromosomes individually and found chr13 was the best, chr16 the worst, and targets had chrM (13 cols) not in the train/test data per se but scored 0.967 on its own just using what was not in reduced.   For chrY there were only 22 targets so perhaps why it did not matter so much especially if these were not included in scores.  The Evaluation Ids csv for scoring was confusing, being a subset for scoring maybe was key to figuring out what was relevant to predict or not, what mattered or not for models.  \n\nSklearn MLPRegressor worked surprisingly well and was useful for testing ideas, fast and cpu only. If only there were more time and were on your team!!   Just kidding.  Many many thanks and well done.",
      "votes": null
    },
    {
      "id": "2031504",
      "postDate": "11/16/2022 05:57:23",
      "content": "<p>Thanks. I was doing this as well but got distracted by another idea… life is full of regret.</p>",
      "rawMarkdown": "Thanks. I was doing this as well but got distracted by another idea... life is full of regret.",
      "votes": null
    },
    {
      "id": "2031621",
      "postDate": "11/16/2022 07:23:40",
      "content": "<p>Thank you very much for your sharing. I have learned a lot from you. Best wishes~~</p>",
      "rawMarkdown": "Thank you very much for your sharing. I have learned a lot from you. Best wishes~~",
      "votes": null
    },
    {
      "id": "2031669",
      "postDate": "11/16/2022 07:55:17",
      "content": "<p>Respect! Thank you for shareing</p>",
      "rawMarkdown": "Respect! Thank you for shareing",
      "votes": null
    },
    {
      "id": "2032331",
      "postDate": "11/16/2022 15:06:21",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>!<br>\nThank you for your insightful comments, both this one and the comments you posted during the competition.<br>\nWouldn't get a bronze without your posts.</p>\n<p>I have checked the Y-chromosome values for both multiome and cite inputs and found the following pattern.<br>\nIn CITE inputs, for example, there are total 30 RNA's that are products of Y chromosome. So, I have checked their results for women donor.<br>\n25 of them were always 0 for all the woman's cells, as they should, while 5 other had values. In all five cases the average values for a female donor were about the same as for male donors, so I considered that noise mostly comes not in a form of random low counts but in a form of misinterpreting, when some RNA is counted as a different RNA. But that conclusion didn't give me reasonable ideas about how to distinguish between noisy and good features. So, I have just moved on.</p>\n<p>I ran the ensembl_rest to map the multiome targets to multiome inputs. And couldn't find a single target that had the same dimentions as a single input. Sometimes one multiome target was a product of about 70 inputs. Sometimes in inputs we only have data for one small piece of DNA that encodes that RNA. In about 30% of cases not a single piece of DNA, that encodes the RNA was present in the data. Features mean too little because they correspond to parts of DNA that are too small. And I guess, that is the reason why the data is so sparse.<br>\nHad an idea to sum the counts for multiome inputs that correspond to neighbouring parts of DNA to create more meaningful features but couldn't find time for that.</p>\n<p>I made SVD features from raw data and tried to add them to my regular model. The SVD_raw_0 was a feature with the highest feature_importance (it correlated very well with another feature, that calculated just a number of counts). But it decreased the model's results (had mixed impact in cross-validation and decreased lb result), so in the end I've just removed all the features made from raw data.</p>",
      "rawMarkdown": "Hi @ambrosm!\nThank you for your insightful comments, both this one and the comments you posted during the competition.\nWouldn't get a bronze without your posts.\n\nI have checked the Y-chromosome values for both multiome and cite inputs and found the following pattern.\nIn CITE inputs, for example, there are total 30 RNA's that are products of Y chromosome. So, I have checked their results for women donor.\n25 of them were always 0 for all the woman's cells, as they should, while 5 other had values. In all five cases the average values for a female donor were about the same as for male donors, so I considered that noise mostly comes not in a form of random low counts but in a form of misinterpreting, when some RNA is counted as a different RNA. But that conclusion didn't give me reasonable ideas about how to distinguish between noisy and good features. So, I have just moved on.\n\nI ran the ensembl_rest to map the multiome targets to multiome inputs. And couldn't find a single target that had the same dimentions as a single input. Sometimes one multiome target was a product of about 70 inputs. Sometimes in inputs we only have data for one small piece of DNA that encodes that RNA. In about 30% of cases not a single piece of DNA, that encodes the RNA was present in the data. Features mean too little because they correspond to parts of DNA that are too small. And I guess, that is the reason why the data is so sparse.\nHad an idea to sum the counts for multiome inputs that correspond to neighbouring parts of DNA to create more meaningful features but couldn't find time for that.\n\nI made SVD features from raw data and tried to add them to my regular model. The SVD_raw_0 was a feature with the highest feature_importance (it correlated very well with another feature, that calculated just a number of counts). But it decreased the model's results (had mixed impact in cross-validation and decreased lb result), so in the end I've just removed all the features made from raw data.",
      "votes": null
    },
    {
      "id": "2032601",
      "postDate": "11/16/2022 17:43:21",
      "content": "<p>Congrats and thanks for sharing lessons learned.</p>\n<p>I'm still laughing with your crying emoji when your team decide to select yours final 2 submissions.  </p>",
      "rawMarkdown": "Congrats and thanks for sharing lessons learned.\n\nI'm still laughing with your crying emoji when your team decide to select yours final 2 submissions.",
      "votes": null
    },
    {
      "id": "2032660",
      "postDate": "11/16/2022 18:29:39",
      "content": "<p>Your work is highly appreciable! Thanks alot.</p>",
      "rawMarkdown": "Your work is highly appreciable! Thanks alot.",
      "votes": null
    },
    {
      "id": "2032714",
      "postDate": "11/16/2022 19:38:29",
      "content": "<p>Thanks for sharing your valuable work </p>",
      "rawMarkdown": "Thanks for sharing your valuable work",
      "votes": null
    },
    {
      "id": "2032842",
      "postDate": "11/16/2022 20:48:00",
      "content": "<p>Thanks for your informative sharing. Cross-validation strategy is really important in this competition, but I did not really understand the problem of time series previously. All my models were trained based on grouped-by-donor strategy , and the performance on private dataset taught me a profound lesson. 😂 </p>",
      "rawMarkdown": "Thanks for your informative sharing. Cross-validation strategy is really important in this competition, but I did not really understand the problem of time series previously. All my models were trained based on grouped-by-donor strategy , and the performance on private dataset taught me a profound lesson. 😂",
      "votes": null
    },
    {
      "id": "2032945",
      "postDate": "11/16/2022 22:52:12",
      "content": "<p>Thanks a lot for your write-up! I was curious to hear more about how you concluded that the datasets were in an almost linear subspace and found their dimensionalities.</p>",
      "rawMarkdown": "Thanks a lot for your write-up! I was curious to hear more about how you concluded that the datasets were in an almost linear subspace and found their dimensionalities.",
      "votes": null
    },
    {
      "id": "2033206",
      "postDate": "11/17/2022 05:02:18",
      "content": "<p><a href=\"https://www.kaggle.com/natalienrd\" target=\"_blank\">@natalienrd</a> If models based on 100 PCA components give better results than models based on all the data, this suggests that all deviations from the 100-dimensional subspace are noise. And if your best model predicts 90 PCA components of the target and inverse-transforms them to 140 proteins, the proteins can be only 90-dimensional.</p>",
      "rawMarkdown": "natalienrd If models based on 100 PCA components give better results than models based on all the data, this suggests that all deviations from the 100-dimensional subspace are noise. And if your best model predicts 90 PCA components of the target and inverse-transforms them to 140 proteins, the proteins can be only 90-dimensional.",
      "votes": null
    },
    {
      "id": "2040670",
      "postDate": "11/23/2022 09:13:23",
      "content": "<p>Congrats! Thanks for sharing!</p>\n<p>I'm very curious about this part:</p>\n<blockquote>\n  <h3>The log1p transformation for gene expressions is wrong</h3>\n  <p>The gene expressions in this competition were presented library-size normalized and log1p-transformed. The log1p-transformation is bad: If the raw counts are 0, 1, 2, 3, 4, 5, 6, 7, normalization transforms them into 0, 100, 200, 300, 400, 500, 600, 700, and log1p transforms them into 0, 4.6, 5.3, 5.7, 6.0, 6.2, 6.4, 6.6. <strong>The transformed 0 and 1 are far away</strong> from each other although a difference of 1 in the raw counts is pure <strong>randomness</strong>. <strong>Clipping</strong> the CITEseq inputs with clip(3, None) before the SVD improves the predictions. </p>\n</blockquote>\n<p>I think log1p is the most common procedure for both bulk and single cell RNA-seq, and I'm guessing this is because we usually focus on <strong>highly expressed genes</strong> and their fold changes, and we usually <strong>ignoring</strong> expression levels <strong>around zero</strong>. </p>\n<p>But I also saw a <a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-019-1874-1\" target=\"_blank\">paper</a> that said:</p>\n<blockquote>\n  <p>In this two-step process (referred to as “log-normalization” for brevity), UMI counts are first scaled by the total sequencing depth (“size factors”) followed by pseudocount addition and log-transformation. While this approach mitigated the relationship between sequencing depth and gene expression, we found that genes with different overall abundances exhibited distinct patterns after log-normalization, and <strong>only low/medium-abundance genes in the bottom three tiers were effectively normalized</strong> (Fig. 1d). In principle, this confounding relationship could be driven by the presence of multiple cell types in human PBMC. However, when we analyzed a 10X Chromium dataset that used human brain RNA as a control (“Chromium control dataset” [5]), we observed identical patterns, and in particular, <strong>ineffectual normalization of high-abundance genes</strong> (Additional file 1: Figure S1 and S2).</p>\n</blockquote>\n<p>I'm quite confused because these three \"views\" are in conflict.</p>\n<p>What do you think about the \"right way\" to normalized gene expressions?</p>",
      "rawMarkdown": "Congrats! Thanks for sharing!\n\nI'm very curious about this part:\n\n> ### The log1p transformation for gene expressions is wrong\nThe gene expressions in this competition were presented library-size normalized and log1p-transformed. The log1p-transformation is bad: If the raw counts are 0, 1, 2, 3, 4, 5, 6, 7, normalization transforms them into 0, 100, 200, 300, 400, 500, 600, 700, and log1p transforms them into 0, 4.6, 5.3, 5.7, 6.0, 6.2, 6.4, 6.6. **The transformed 0 and 1 are far away** from each other although a difference of 1 in the raw counts is pure **randomness**. **Clipping** the CITEseq inputs with clip(3, None) before the SVD improves the predictions. \n\nI think log1p is the most common procedure for both bulk and single cell RNA-seq, and I'm guessing this is because we usually focus on **highly expressed genes** and their fold changes, and we usually **ignoring** expression levels **around zero**. \n\nBut I also saw a [paper](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-019-1874-1) that said:\n\n> In this two-step process (referred to as “log-normalization” for brevity), UMI counts are first scaled by the total sequencing depth (“size factors”) followed by pseudocount addition and log-transformation. While this approach mitigated the relationship between sequencing depth and gene expression, we found that genes with different overall abundances exhibited distinct patterns after log-normalization, and **only low/medium-abundance genes in the bottom three tiers were effectively normalized** (Fig. 1d). In principle, this confounding relationship could be driven by the presence of multiple cell types in human PBMC. However, when we analyzed a 10X Chromium dataset that used human brain RNA as a control (“Chromium control dataset” [5]), we observed identical patterns, and in particular, **ineffectual normalization of high-abundance genes** (Additional file 1: Figure S1 and S2).\n\nI'm quite confused because these three \"views\" are in conflict.\n\nWhat do you think about the \"right way\" to normalized gene expressions?",
      "votes": null
    },
    {
      "id": "2041652",
      "postDate": "11/24/2022 05:29:01",
      "content": "<p>Congratulations and thanks for your every sharing AmbrosM. There are some regrets in this competition but I hope and believe there will be more good luck with you in the future. </p>",
      "rawMarkdown": "Congratulations and thanks for your every sharing AmbrosM. There are some regrets in this competition but I hope and believe there will be more good luck with you in the future.",
      "votes": null
    },
    {
      "id": "2045729",
      "postDate": "11/27/2022 16:08:46",
      "content": "<p>From my perspective, it depends on the tasks. Normalization and log-transformed still make sense to me because, in bioinformatics research, we usually do not pay much attention to predict expression level data, but we perform data mining, for example, identify gene-peak relation. Moreover, we use highly variable genes or differential genes. However, sequencing depth is a confounder, like you said, will affect tasks like co-expression network inference. Btw, where did you find these papers? Could you please share it with us? Thanks a lot!!!</p>",
      "rawMarkdown": "From my perspective, it depends on the tasks. Normalization and log-transformed still make sense to me because, in bioinformatics research, we usually do not pay much attention to predict expression level data, but we perform data mining, for example, identify gene-peak relation. Moreover, we use highly variable genes or differential genes. However, sequencing depth is a confounder, like you said, will affect tasks like co-expression network inference. Btw, where did you find these papers? Could you please share it with us? Thanks a lot!!!",
      "votes": null
    },
    {
      "id": "2046204",
      "postDate": "11/28/2022 02:29:39",
      "content": "<p>Yes, but I don't quite understand why it's bad in the prediction task. Maybe it indicates some bias of this task?</p>\n<p>BTW, I find the paper on <a href=\"https://satijalab.org/seurat/articles/get_started.html#additional-new-methods\" target=\"_blank\">Seurat</a> website.</p>",
      "rawMarkdown": "Yes, but I don't quite understand why it's bad in the prediction task. Maybe it indicates some bias of this task?\n\nBTW, I find the paper on [Seurat](https://satijalab.org/seurat/articles/get_started.html#additional-new-methods) website.",
      "votes": null
    },
    {
      "id": "2046323",
      "postDate": "11/28/2022 05:45:50",
      "content": "<p><a href=\"https://www.kaggle.com/awater1223\" target=\"_blank\">@awater1223</a> and <a href=\"https://www.kaggle.com/llttyy\" target=\"_blank\">@llttyy</a> I think my proposition applies to the case when normalization is followed by a singular value decomposition, which doesn't deal well with outliers. <a href=\"https://www.kaggle.com/code/ambrosm/msci-citeseq-preprocessing-variants\" target=\"_blank\">This notebook</a> compares the effect of different normalization methods on the prediction result. Results are best when either normalizing to 10000 and taking log1p, or when replacing log1p by a square root as suggested by <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a> and <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a>.</p>",
      "rawMarkdown": "awater1223 and @llttyy I think my proposition applies to the case when normalization is followed by a singular value decomposition, which doesn't deal well with outliers. [This notebook](https://www.kaggle.com/code/ambrosm/msci-citeseq-preprocessing-variants) compares the effect of different normalization methods on the prediction result. Results are best when either normalizing to 10000 and taking log1p, or when replacing log1p by a square root as suggested by @baosenguo and @senkin13.",
      "votes": null
    },
    {
      "id": "2047793",
      "postDate": "11/29/2022 03:44:00",
      "content": "<p>Thanks for the explanation! The results in the notebook are very clear.</p>",
      "rawMarkdown": "Thanks for the explanation! The results in the notebook are very clear.",
      "votes": null
    },
    {
      "id": "2077200",
      "postDate": "12/27/2022 10:38:00",
      "content": "<p>Thank you for sharing this and amazing notebooks, <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> ! It would be great to know which CV strategy your team used for the final submissions. The one described in the post or something else?</p>",
      "rawMarkdown": "Thank you for sharing this and amazing notebooks, @ambrosm ! It would be great to know which CV strategy your team used for the final submissions. The one described in the post or something else?",
      "votes": null
    },
    {
      "id": "2082684",
      "postDate": "01/01/2023 19:38:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/shitovvladimir\" target=\"_blank\">@shitovvladimir</a>, our final submissions were an ensemble of models which were optimized with different CV strategies. As you can imagine, getting the weights right is difficult without a uniform metric over all ensemble members…</p>",
      "rawMarkdown": "Hi @shitovvladimir, our final submissions were an ensemble of models which were optimized with different CV strategies. As you can imagine, getting the weights right is difficult without a uniform metric over all ensemble members...",
      "votes": null
    },
    {
      "id": "2084512",
      "postDate": "01/03/2023 15:07:02",
      "content": "<p>Thanks for your sharing!</p>",
      "rawMarkdown": "Thanks for your sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2031354,
      "author_name": "kingychiu",
      "author_url": "",
      "post_date": "11/16/2022 02:58:46",
      "content": "<p>Nice summary! What are the main differences between your 0.773 and the 0.771 models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2031471,
          "author_name": "ambrosm",
          "author_url": "",
          "post_date": "11/16/2022 05:17:58",
          "content": "<p><a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a> In the 0.773 ensemble we gave much higher ensemble weights to the models with high cv score for Multiome day 7, assuming that day 7 cv is the best predictor for day 10.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2031504,
          "author_name": "kingychiu",
          "author_url": "",
          "post_date": "11/16/2022 05:57:23",
          "content": "<p>Thanks. I was doing this as well but got distracted by another idea… life is full of regret.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2031454,
      "author_name": "liilili",
      "author_url": "",
      "post_date": "11/16/2022 05:03:54",
      "content": "<p>Very Very thank you, through this competition, we  have learned a lot from you. Respect!!  We are looking forward to your sharing! Respect again!! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2031495,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "11/16/2022 05:47:31",
      "content": "<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> - you were certainly an MVP in this competition and shared many insights and important notebooks. Whatever the results you were a gold medal person and think we all learned a lot here.  Felt there were so many permutations of what could be done here and it helped to have someone (and their team) keep a focus on it all and the momentum going.   And importantly to forget the public LB and work for the private LB.</p>\n<p>It would have been interesting to have separate LB for CITEseq and Multiome just to see esp for Multiome.  Spent a lot of time aligning genes and gene locations with target genes and gene ids, and geneactivity.  wrt Multiome it seemed to help to use a reduced set and combining SVDs, using a small number of components for each SVD. Even pursued chromosomes individually and found chr13 was the best, chr16 the worst, and targets had chrM (13 cols) not in the train/test data per se but scored 0.967 on its own just using what was not in reduced.   For chrY there were only 22 targets so perhaps why it did not matter so much especially if these were not included in scores.  The Evaluation Ids csv for scoring was confusing, being a subset for scoring maybe was key to figuring out what was relevant to predict or not, what mattered or not for models.  </p>\n<p>Sklearn MLPRegressor worked surprisingly well and was useful for testing ideas, fast and cpu only. If only there were more time and were on your team!!   Just kidding.  Many many thanks and well done.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2031621,
      "author_name": "accessibletracy",
      "author_url": "",
      "post_date": "11/16/2022 07:23:40",
      "content": "<p>Thank you very much for your sharing. I have learned a lot from you. Best wishes~~</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2031669,
      "author_name": "zhaohuili123",
      "author_url": "",
      "post_date": "11/16/2022 07:55:17",
      "content": "<p>Respect! Thank you for shareing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2032331,
      "author_name": "artemfedorov",
      "author_url": "",
      "post_date": "11/16/2022 15:06:21",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>!<br>\nThank you for your insightful comments, both this one and the comments you posted during the competition.<br>\nWouldn't get a bronze without your posts.</p>\n<p>I have checked the Y-chromosome values for both multiome and cite inputs and found the following pattern.<br>\nIn CITE inputs, for example, there are total 30 RNA's that are products of Y chromosome. So, I have checked their results for women donor.<br>\n25 of them were always 0 for all the woman's cells, as they should, while 5 other had values. In all five cases the average values for a female donor were about the same as for male donors, so I considered that noise mostly comes not in a form of random low counts but in a form of misinterpreting, when some RNA is counted as a different RNA. But that conclusion didn't give me reasonable ideas about how to distinguish between noisy and good features. So, I have just moved on.</p>\n<p>I ran the ensembl_rest to map the multiome targets to multiome inputs. And couldn't find a single target that had the same dimentions as a single input. Sometimes one multiome target was a product of about 70 inputs. Sometimes in inputs we only have data for one small piece of DNA that encodes that RNA. In about 30% of cases not a single piece of DNA, that encodes the RNA was present in the data. Features mean too little because they correspond to parts of DNA that are too small. And I guess, that is the reason why the data is so sparse.<br>\nHad an idea to sum the counts for multiome inputs that correspond to neighbouring parts of DNA to create more meaningful features but couldn't find time for that.</p>\n<p>I made SVD features from raw data and tried to add them to my regular model. The SVD_raw_0 was a feature with the highest feature_importance (it correlated very well with another feature, that calculated just a number of counts). But it decreased the model's results (had mixed impact in cross-validation and decreased lb result), so in the end I've just removed all the features made from raw data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2032601,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "11/16/2022 17:43:21",
      "content": "<p>Congrats and thanks for sharing lessons learned.</p>\n<p>I'm still laughing with your crying emoji when your team decide to select yours final 2 submissions.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2032660,
      "author_name": "faizanhussaine",
      "author_url": "",
      "post_date": "11/16/2022 18:29:39",
      "content": "<p>Your work is highly appreciable! Thanks alot.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2032714,
      "author_name": "shreyamishra0307",
      "author_url": "",
      "post_date": "11/16/2022 19:38:29",
      "content": "<p>Thanks for sharing your valuable work </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2032842,
      "author_name": "gufanmingmie",
      "author_url": "",
      "post_date": "11/16/2022 20:48:00",
      "content": "<p>Thanks for your informative sharing. Cross-validation strategy is really important in this competition, but I did not really understand the problem of time series previously. All my models were trained based on grouped-by-donor strategy , and the performance on private dataset taught me a profound lesson. 😂 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2032945,
      "author_name": "natalienrd",
      "author_url": "",
      "post_date": "11/16/2022 22:52:12",
      "content": "<p>Thanks a lot for your write-up! I was curious to hear more about how you concluded that the datasets were in an almost linear subspace and found their dimensionalities.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2033206,
          "author_name": "ambrosm",
          "author_url": "",
          "post_date": "11/17/2022 05:02:18",
          "content": "<p><a href=\"https://www.kaggle.com/natalienrd\" target=\"_blank\">@natalienrd</a> If models based on 100 PCA components give better results than models based on all the data, this suggests that all deviations from the 100-dimensional subspace are noise. And if your best model predicts 90 PCA components of the target and inverse-transforms them to 140 proteins, the proteins can be only 90-dimensional.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2040670,
      "author_name": "awater1223",
      "author_url": "",
      "post_date": "11/23/2022 09:13:23",
      "content": "<p>Congrats! Thanks for sharing!</p>\n<p>I'm very curious about this part:</p>\n<blockquote>\n  <h3>The log1p transformation for gene expressions is wrong</h3>\n  <p>The gene expressions in this competition were presented library-size normalized and log1p-transformed. The log1p-transformation is bad: If the raw counts are 0, 1, 2, 3, 4, 5, 6, 7, normalization transforms them into 0, 100, 200, 300, 400, 500, 600, 700, and log1p transforms them into 0, 4.6, 5.3, 5.7, 6.0, 6.2, 6.4, 6.6. <strong>The transformed 0 and 1 are far away</strong> from each other although a difference of 1 in the raw counts is pure <strong>randomness</strong>. <strong>Clipping</strong> the CITEseq inputs with clip(3, None) before the SVD improves the predictions. </p>\n</blockquote>\n<p>I think log1p is the most common procedure for both bulk and single cell RNA-seq, and I'm guessing this is because we usually focus on <strong>highly expressed genes</strong> and their fold changes, and we usually <strong>ignoring</strong> expression levels <strong>around zero</strong>. </p>\n<p>But I also saw a <a href=\"https://genomebiology.biomedcentral.com/articles/10.1186/s13059-019-1874-1\" target=\"_blank\">paper</a> that said:</p>\n<blockquote>\n  <p>In this two-step process (referred to as “log-normalization” for brevity), UMI counts are first scaled by the total sequencing depth (“size factors”) followed by pseudocount addition and log-transformation. While this approach mitigated the relationship between sequencing depth and gene expression, we found that genes with different overall abundances exhibited distinct patterns after log-normalization, and <strong>only low/medium-abundance genes in the bottom three tiers were effectively normalized</strong> (Fig. 1d). In principle, this confounding relationship could be driven by the presence of multiple cell types in human PBMC. However, when we analyzed a 10X Chromium dataset that used human brain RNA as a control (“Chromium control dataset” [5]), we observed identical patterns, and in particular, <strong>ineffectual normalization of high-abundance genes</strong> (Additional file 1: Figure S1 and S2).</p>\n</blockquote>\n<p>I'm quite confused because these three \"views\" are in conflict.</p>\n<p>What do you think about the \"right way\" to normalized gene expressions?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2045729,
          "author_name": "llttyy",
          "author_url": "",
          "post_date": "11/27/2022 16:08:46",
          "content": "<p>From my perspective, it depends on the tasks. Normalization and log-transformed still make sense to me because, in bioinformatics research, we usually do not pay much attention to predict expression level data, but we perform data mining, for example, identify gene-peak relation. Moreover, we use highly variable genes or differential genes. However, sequencing depth is a confounder, like you said, will affect tasks like co-expression network inference. Btw, where did you find these papers? Could you please share it with us? Thanks a lot!!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2046204,
          "author_name": "awater1223",
          "author_url": "",
          "post_date": "11/28/2022 02:29:39",
          "content": "<p>Yes, but I don't quite understand why it's bad in the prediction task. Maybe it indicates some bias of this task?</p>\n<p>BTW, I find the paper on <a href=\"https://satijalab.org/seurat/articles/get_started.html#additional-new-methods\" target=\"_blank\">Seurat</a> website.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2046323,
          "author_name": "ambrosm",
          "author_url": "",
          "post_date": "11/28/2022 05:45:50",
          "content": "<p><a href=\"https://www.kaggle.com/awater1223\" target=\"_blank\">@awater1223</a> and <a href=\"https://www.kaggle.com/llttyy\" target=\"_blank\">@llttyy</a> I think my proposition applies to the case when normalization is followed by a singular value decomposition, which doesn't deal well with outliers. <a href=\"https://www.kaggle.com/code/ambrosm/msci-citeseq-preprocessing-variants\" target=\"_blank\">This notebook</a> compares the effect of different normalization methods on the prediction result. Results are best when either normalizing to 10000 and taking log1p, or when replacing log1p by a square root as suggested by <a href=\"https://www.kaggle.com/baosenguo\" target=\"_blank\">@baosenguo</a> and <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2047793,
          "author_name": "awater1223",
          "author_url": "",
          "post_date": "11/29/2022 03:44:00",
          "content": "<p>Thanks for the explanation! The results in the notebook are very clear.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2041652,
      "author_name": "songqizhou",
      "author_url": "",
      "post_date": "11/24/2022 05:29:01",
      "content": "<p>Congratulations and thanks for your every sharing AmbrosM. There are some regrets in this competition but I hope and believe there will be more good luck with you in the future. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2077200,
      "author_name": "shitovvladimir",
      "author_url": "",
      "post_date": "12/27/2022 10:38:00",
      "content": "<p>Thank you for sharing this and amazing notebooks, <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> ! It would be great to know which CV strategy your team used for the final submissions. The one described in the post or something else?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2082684,
          "author_name": "ambrosm",
          "author_url": "",
          "post_date": "01/01/2023 19:38:48",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/shitovvladimir\" target=\"_blank\">@shitovvladimir</a>, our final submissions were an ensemble of models which were optimized with different CV strategies. As you can imagine, getting the weights right is difficult without a uniform metric over all ensemble members…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2084512,
      "author_name": "addoils",
      "author_url": "",
      "post_date": "01/03/2023 15:07:02",
      "content": "<p>Thanks for your sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2031346": "*The most important result of every endeavor is what we learn from it...*\n\n# Time series and extrapolation\n\nThe competition's data is a time series and the private test data is an extrapolation problem (because the unseen day is in the future): Most people know that decision trees can't extrapolate, but k-nearest-neighbors can't extrapolate either, and rbf kernel ridge regression is even worse (it predicts 0 if you extrapolate far enough).\n\nHaving a good cross-validation strategy is crucial: Use a GroupKFold on days for validation (with particular emphasis on the last day) but train the submission models with a ShuffleSplit.\n\nBecause the model changes every day from day 2 to day 10, a good strategy for the Multiome public leaderboard consists of training on only the data of the day we want to predict.\n \nUnfortunately we looked too much at the public lb when selecting our final two submissions and got shaken down:\n\n![0.773](https://i.imgur.com/f8WqyO3.png)\n\n# The log1p transformation for gene expressions is wrong\n\nThe gene expressions in this competition were presented library-size normalized and log1p-transformed. The log1p-transformation is bad: If the raw counts are 0, 1, 2, 3, 4, 5, 6, 7, normalization transforms them into 0, 100, 200, 300, 400, 500, 600, 700, and log1p transforms them into 0, 4.6, 5.3, 5.7, 6.0, 6.2, 6.4, 6.6. The transformed 0 and 1 are far away from each other although a difference of 1 in the raw counts is pure randomness. Clipping the CITEseq inputs with `clip(3, None)` before the SVD improves the predictions.\n\n![log1p diagram](https://i.imgur.com/iOpVRmj.png)\n\n# The metric\n\nUnderstanding the metric is important in every competition, but this competition had the peculiarity that the correlation metric can be optimized directly only by neural networks. For all other machine learning methods, the best we can do is optimizing for mean squared error. The points to know here are:\n1. If we optimize MSE as a surrogate for the correlation metric, we need to standardize the targets: `Y = Y / Y.std(axis=1).reshape(-1, 1)` Neglecting this standardization corresponds to optimizing for random sample weights. The same standardization is necessary before ensembling.\n2. It can be shown that if we apply this standardization correctly, the optimum model for mse loss is equivalent to the optimum model for correlation loss.\n3. MSE loss is invariant under PCA and SVD transformations, which multiply all vectors with an orthogonal matrix (so that the distances are conserved). Only this invariance legitimates the transformation and dimensionality reduction of the targets.\n\n# Noise and dimensionality reduction\n\nUnderstanding the signal-to-noise ratio of the data and its dimensionality is important. More than 90 % of the ATACseq variance and more than 80 % of the gene expression variance seem to be noise (i.e. useless as model input and unpredictable as model output).\n\nAs all three datasets live in almost linear subspaces, the dimensionality of the data can be reduced with a PCA or SVD. As a side effect, the dimensionality reduction improves the signal-to-noise ratio by concentrating the signal in the first few components (and moving the noise to the other components).\n\nThe data seems to live in manifolds of the following dimensions:\n- ATACseq: between 100 and 140 of 200000\n- Gene expressions: between 70 and 130 of 20000\n- Proteins: about 90 of 140\n\nThe dimensionality of the input for both problems (CITEseq and Multiome) has to be reduced for all models - models with 20000 or even 200000 features are not feasible. For the prediction targets, the situation is different:\n- The CITEseq proteins were best predicted when the 140 dimensions were reduced to 90, but predicting all 140 proteins directly is possible and adds diversity to the ensemble.\n- Because of space and time constraints, most machine learning algorithms cannot predict the 20000 gene expressions in Multiome at once and need a dimensionality reduction for the output. The notable exception here is the k-nearest-neighbors algorithm which predicts 20000 targets in reasonable time. The knn predictions are somewhat noisy and profit from a postprocessing which constrains them into a 75-dimensional manifold.\n\n# Feature selection\n\nIn CITEseq, many of the 20000 features are pure noise. If we select a few hundred important features and give them to a machine learning algorithm, the result is better than if we only use an SVD of all 20000 features.\n\nFor Multiome, I had no success in determining the most important features, but if you drop the whole Y chromosome, the model's score remains the same. This suggests that the Y chromosome has no important function in blood cells.\n\n# Using RAM efficiently\n\nWith 100 GBytes of data, probably every participant has learnt how to use the RAM efficiently:\n- It is good to understand when numpy copies arrays and when it returns views.\n- It is good to understand the memory complexity of algorithms. ExtraTrees consumes memory proportional to n_estimators and to the number of leaves per tree, kernel ridge regression with the rbf kernel computes a distance matrix which is quadratic in n_samples, etc.\n- In the end I even coded a ridge regression for the Multiome ensemble which works without loading all the data into memory at the same time.\n  \n# Teamwork\n\nI'm happy that I could be part of a good team with @mtinti, @callmeb and @khahuras: We had a fruitful exchange of ideas (communicating through Slack), all team members contributed their best models to a very diverse final ensemble, and the virtual presence of teammembers motivates everybody during phases of stagnating scores.",
    "2031354": "Nice summary! What are the main differences between your 0.773 and the 0.771 models?",
    "2031454": "Very Very thank you, through this competition, we  have learned a lot from you. Respect!!  We are looking forward to your sharing! Respect again!!",
    "2031471": "kingychiu In the 0.773 ensemble we gave much higher ensemble weights to the models with high cv score for Multiome day 7, assuming that day 7 cv is the best predictor for day 10.",
    "2031495": "ambrosm - you were certainly an MVP in this competition and shared many insights and important notebooks. Whatever the results you were a gold medal person and think we all learned a lot here.  Felt there were so many permutations of what could be done here and it helped to have someone (and their team) keep a focus on it all and the momentum going.   And importantly to forget the public LB and work for the private LB.\n\nIt would have been interesting to have separate LB for CITEseq and Multiome just to see esp for Multiome.  Spent a lot of time aligning genes and gene locations with target genes and gene ids, and geneactivity.  wrt Multiome it seemed to help to use a reduced set and combining SVDs, using a small number of components for each SVD. Even pursued chromosomes individually and found chr13 was the best, chr16 the worst, and targets had chrM (13 cols) not in the train/test data per se but scored 0.967 on its own just using what was not in reduced.   For chrY there were only 22 targets so perhaps why it did not matter so much especially if these were not included in scores.  The Evaluation Ids csv for scoring was confusing, being a subset for scoring maybe was key to figuring out what was relevant to predict or not, what mattered or not for models.  \n\nSklearn MLPRegressor worked surprisingly well and was useful for testing ideas, fast and cpu only. If only there were more time and were on your team!!   Just kidding.  Many many thanks and well done.",
    "2031504": "Thanks. I was doing this as well but got distracted by another idea... life is full of regret.",
    "2031621": "Thank you very much for your sharing. I have learned a lot from you. Best wishes~~",
    "2031669": "Respect! Thank you for shareing",
    "2032331": "Hi @ambrosm!\nThank you for your insightful comments, both this one and the comments you posted during the competition.\nWouldn't get a bronze without your posts.\n\nI have checked the Y-chromosome values for both multiome and cite inputs and found the following pattern.\nIn CITE inputs, for example, there are total 30 RNA's that are products of Y chromosome. So, I have checked their results for women donor.\n25 of them were always 0 for all the woman's cells, as they should, while 5 other had values. In all five cases the average values for a female donor were about the same as for male donors, so I considered that noise mostly comes not in a form of random low counts but in a form of misinterpreting, when some RNA is counted as a different RNA. But that conclusion didn't give me reasonable ideas about how to distinguish between noisy and good features. So, I have just moved on.\n\nI ran the ensembl_rest to map the multiome targets to multiome inputs. And couldn't find a single target that had the same dimentions as a single input. Sometimes one multiome target was a product of about 70 inputs. Sometimes in inputs we only have data for one small piece of DNA that encodes that RNA. In about 30% of cases not a single piece of DNA, that encodes the RNA was present in the data. Features mean too little because they correspond to parts of DNA that are too small. And I guess, that is the reason why the data is so sparse.\nHad an idea to sum the counts for multiome inputs that correspond to neighbouring parts of DNA to create more meaningful features but couldn't find time for that.\n\nI made SVD features from raw data and tried to add them to my regular model. The SVD_raw_0 was a feature with the highest feature_importance (it correlated very well with another feature, that calculated just a number of counts). But it decreased the model's results (had mixed impact in cross-validation and decreased lb result), so in the end I've just removed all the features made from raw data.",
    "2032601": "Congrats and thanks for sharing lessons learned.\n\nI'm still laughing with your crying emoji when your team decide to select yours final 2 submissions.",
    "2032660": "Your work is highly appreciable! Thanks alot.",
    "2032714": "Thanks for sharing your valuable work",
    "2032842": "Thanks for your informative sharing. Cross-validation strategy is really important in this competition, but I did not really understand the problem of time series previously. All my models were trained based on grouped-by-donor strategy , and the performance on private dataset taught me a profound lesson. 😂",
    "2032945": "Thanks a lot for your write-up! I was curious to hear more about how you concluded that the datasets were in an almost linear subspace and found their dimensionalities.",
    "2033206": "natalienrd If models based on 100 PCA components give better results than models based on all the data, this suggests that all deviations from the 100-dimensional subspace are noise. And if your best model predicts 90 PCA components of the target and inverse-transforms them to 140 proteins, the proteins can be only 90-dimensional.",
    "2040670": "Congrats! Thanks for sharing!\n\nI'm very curious about this part:\n\n> ### The log1p transformation for gene expressions is wrong\nThe gene expressions in this competition were presented library-size normalized and log1p-transformed. The log1p-transformation is bad: If the raw counts are 0, 1, 2, 3, 4, 5, 6, 7, normalization transforms them into 0, 100, 200, 300, 400, 500, 600, 700, and log1p transforms them into 0, 4.6, 5.3, 5.7, 6.0, 6.2, 6.4, 6.6. **The transformed 0 and 1 are far away** from each other although a difference of 1 in the raw counts is pure **randomness**. **Clipping** the CITEseq inputs with clip(3, None) before the SVD improves the predictions. \n\nI think log1p is the most common procedure for both bulk and single cell RNA-seq, and I'm guessing this is because we usually focus on **highly expressed genes** and their fold changes, and we usually **ignoring** expression levels **around zero**. \n\nBut I also saw a [paper](https://genomebiology.biomedcentral.com/articles/10.1186/s13059-019-1874-1) that said:\n\n> In this two-step process (referred to as “log-normalization” for brevity), UMI counts are first scaled by the total sequencing depth (“size factors”) followed by pseudocount addition and log-transformation. While this approach mitigated the relationship between sequencing depth and gene expression, we found that genes with different overall abundances exhibited distinct patterns after log-normalization, and **only low/medium-abundance genes in the bottom three tiers were effectively normalized** (Fig. 1d). In principle, this confounding relationship could be driven by the presence of multiple cell types in human PBMC. However, when we analyzed a 10X Chromium dataset that used human brain RNA as a control (“Chromium control dataset” [5]), we observed identical patterns, and in particular, **ineffectual normalization of high-abundance genes** (Additional file 1: Figure S1 and S2).\n\nI'm quite confused because these three \"views\" are in conflict.\n\nWhat do you think about the \"right way\" to normalized gene expressions?",
    "2041652": "Congratulations and thanks for your every sharing AmbrosM. There are some regrets in this competition but I hope and believe there will be more good luck with you in the future.",
    "2045729": "From my perspective, it depends on the tasks. Normalization and log-transformed still make sense to me because, in bioinformatics research, we usually do not pay much attention to predict expression level data, but we perform data mining, for example, identify gene-peak relation. Moreover, we use highly variable genes or differential genes. However, sequencing depth is a confounder, like you said, will affect tasks like co-expression network inference. Btw, where did you find these papers? Could you please share it with us? Thanks a lot!!!",
    "2046204": "Yes, but I don't quite understand why it's bad in the prediction task. Maybe it indicates some bias of this task?\n\nBTW, I find the paper on [Seurat](https://satijalab.org/seurat/articles/get_started.html#additional-new-methods) website.",
    "2046323": "awater1223 and @llttyy I think my proposition applies to the case when normalization is followed by a singular value decomposition, which doesn't deal well with outliers. [This notebook](https://www.kaggle.com/code/ambrosm/msci-citeseq-preprocessing-variants) compares the effect of different normalization methods on the prediction result. Results are best when either normalizing to 10000 and taking log1p, or when replacing log1p by a square root as suggested by @baosenguo and @senkin13.",
    "2047793": "Thanks for the explanation! The results in the notebook are very clear.",
    "2077200": "Thank you for sharing this and amazing notebooks, @ambrosm ! It would be great to know which CV strategy your team used for the final submissions. The one described in the post or something else?",
    "2082684": "Hi @shitovvladimir, our final submissions were an ensemble of models which were optimized with different CV strategies. As you can imagine, getting the weights right is difficult without a uniform metric over all ensemble members...",
    "2084512": "Thanks for your sharing!"
  },
  "source": "meta"
}