{
  "id": 459001,
  "title": "... & I'd like to thank Kaggle, the challenge host, and everyone who made their notebooks public.",
  "url": "/competitions/open-problems-single-cell-perturbations/writeups/mehran-somayyeh-i-d-like-to-thank-kaggle-the-chall",
  "author_name": "",
  "post_date": "2023-12-02T21:49:46.901186400Z",
  "votes": 7,
  "comment_count": 10,
  "views": 0,
  "content": "<p>The prediction in this competition included more than eighteen thousand columns, which may have less history in machine learning competitions, while the number of features was also very limited. These themes made every detail seem very important. Even regardless of the result of the match, it was definitely a good experience for us.</p>\n<p>In this challenge, there are only two features, namely \"cell_type\" and \"sm_name\". That's why we used \"Feature Augmentation\". We added two new columns (two new features) separately for each prediction column as follows:</p>\n<p>If we separate the cells based on 'cell_type' and assume that the drugs will usually have similar responses on each of these divisions, we can hope that by finding the average effects, we have obtained a new feature. For example, for y0 and the new feature of zero column, the correlation coefficient is 0.24. Of course, this amount is repeated for other columns as well.</p>\n<p>Also, if we separate the cells based on 'sm_name', we get a new feature by finding the average effects. In this case, for y0 and the new feature of column zero, the correlation coefficient is 0.62, and this value is almost repeated for other columns.</p>\n<p>In addition, we added other features by using \"SMILES\", which were used for all prediction columns at the same time. At first, we added about five hundred binary features using \"fragments of SMILES\" and then we added about two thousand new binary features using \"morgan fingerprint from SMILES\".</p>\n<p>We have publicly released two notebooks that cover the above topicss:</p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/1-op2-eda-linearsvr-regressorchain\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/1-op2-eda-linearsvr-regressorchain</a></p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles</a></p>\n<p>Finally, the results of LinearSVR and neural network and NLP and PYBOOST were combined and we used \"Separately Ensembling for Each Column\". That is, Ensembling was done with different coefficients (based on the correlation value).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Ffc0fcd5fe67f70542c4e5409a76b81ad%2Fphoto_2023-12-03_01-15-24.jpg?generation=1701553548981440&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2546812",
      "postDate": "12/02/2023 21:49:46",
      "content": "<p>The prediction in this competition included more than eighteen thousand columns, which may have less history in machine learning competitions, while the number of features was also very limited. These themes made every detail seem very important. Even regardless of the result of the match, it was definitely a good experience for us.</p>\n<p>In this challenge, there are only two features, namely \"cell_type\" and \"sm_name\". That's why we used \"Feature Augmentation\". We added two new columns (two new features) separately for each prediction column as follows:</p>\n<p>If we separate the cells based on 'cell_type' and assume that the drugs will usually have similar responses on each of these divisions, we can hope that by finding the average effects, we have obtained a new feature. For example, for y0 and the new feature of zero column, the correlation coefficient is 0.24. Of course, this amount is repeated for other columns as well.</p>\n<p>Also, if we separate the cells based on 'sm_name', we get a new feature by finding the average effects. In this case, for y0 and the new feature of column zero, the correlation coefficient is 0.62, and this value is almost repeated for other columns.</p>\n<p>In addition, we added other features by using \"SMILES\", which were used for all prediction columns at the same time. At first, we added about five hundred binary features using \"fragments of SMILES\" and then we added about two thousand new binary features using \"morgan fingerprint from SMILES\".</p>\n<p>We have publicly released two notebooks that cover the above topicss:</p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/1-op2-eda-linearsvr-regressorchain\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/1-op2-eda-linearsvr-regressorchain</a></p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles</a></p>\n<p>Finally, the results of LinearSVR and neural network and NLP and PYBOOST were combined and we used \"Separately Ensembling for Each Column\". That is, Ensembling was done with different coefficients (based on the correlation value).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Ffc0fcd5fe67f70542c4e5409a76b81ad%2Fphoto_2023-12-03_01-15-24.jpg?generation=1701553548981440&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The prediction in this competition included more than eighteen thousand columns, which may have less history in machine learning competitions, while the number of features was also very limited. These themes made every detail seem very important. Even regardless of the result of the match, it was definitely a good experience for us.\n\nIn this challenge, there are only two features, namely \"cell_type\" and \"sm_name\". That's why we used \"Feature Augmentation\". We added two new columns (two new features) separately for each prediction column as follows:\n\nIf we separate the cells based on 'cell_type' and assume that the drugs will usually have similar responses on each of these divisions, we can hope that by finding the average effects, we have obtained a new feature. For example, for y0 and the new feature of zero column, the correlation coefficient is 0.24. Of course, this amount is repeated for other columns as well.\n\nAlso, if we separate the cells based on 'sm_name', we get a new feature by finding the average effects. In this case, for y0 and the new feature of column zero, the correlation coefficient is 0.62, and this value is almost repeated for other columns.\n\n\nIn addition, we added other features by using \"SMILES\", which were used for all prediction columns at the same time. At first, we added about five hundred binary features using \"fragments of SMILES\" and then we added about two thousand new binary features using \"morgan fingerprint from SMILES\".\n\nWe have publicly released two notebooks that cover the above topicss:\n\nhttps://www.kaggle.com/code/mehrankazeminia/1-op2-eda-linearsvr-regressorchain\n\nhttps://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles\n\nFinally, the results of LinearSVR and neural network and NLP and PYBOOST were combined and we used \"Separately Ensembling for Each Column\". That is, Ensembling was done with different coefficients (based on the correlation value).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Ffc0fcd5fe67f70542c4e5409a76b81ad%2Fphoto_2023-12-03_01-15-24.jpg?generation=1701553548981440&alt=media)",
      "votes": null
    },
    {
      "id": "2547264",
      "postDate": "12/03/2023 10:59:23",
      "content": "<p>Thanks a lot for sharing ! <br>\nAnd thanks a lot for sharing your approaches during the challenge !<br>\nI guess quite a lot of participants used your work/ideas (including me).</p>\n<p>PS</p>\n<p>Would be you be so kind to give more details on blending with different coefficients for columns (targets).<br>\nHow exactly you chose the blend weights and what was the rationale ?<br>\nWe also thought in that direction but not quite successfully.<br>\n(The standard approach to get blend weight by blending OOF predictions and selected best local score - seems not really working here - because CV-LB correspondence is not good , especially for different models - like boosting vs NN). <br>\nAlso would be great to have more comments on the scatterplot - what the axis and dots ?</p>",
      "rawMarkdown": "Thanks a lot for sharing ! \nAnd thanks a lot for sharing your approaches during the challenge !\nI guess quite a lot of participants used your work/ideas (including me).\n\nPS\n\nWould be you be so kind to give more details on blending with different coefficients for columns (targets).\nHow exactly you chose the blend weights and what was the rationale ?\nWe also thought in that direction but not quite successfully.\n(The standard approach to get blend weight by blending OOF predictions and selected best local score - seems not really working here - because CV-LB correspondence is not good , especially for different models - like boosting vs NN). \nAlso would be great to have more comments on the scatterplot - what the axis and dots ?",
      "votes": null
    },
    {
      "id": "2547289",
      "postDate": "12/03/2023 11:41:40",
      "content": "<p>Hi; thank you<br>\nWe usually follow two rules:</p>\n<p>1- When two columns have a negative correlation:<br>\nInstead of their Ensemble, choose one of them alone. (Probably the column that is used as a support and maybe has a lower general score will be selected)</p>\n<p>2- When two columns have a correlation close to one:<br>\nInstead of combining them, choose one of them. (Probably the column that is used as the main one and maybe has a higher general score will be selected)</p>\n<p>Please note that:</p>\n<p>In many cases, only the final results of the calculations are accessible, that is, there is no access to the validation data. In these cases, Ensembling is possible by trial and error method, and actually Ensembling should be done in the dark.</p>\n<p>But when the number of answer columns is more than one, the darkness increases. Because it is not known that after finding the right coefficient for ensembling the first columns, the same coefficient is optimal for ensembling the next columns.</p>\n<p>In this challenge, the predict contains more than eighteen thousand columns. So if a coefficient is chosen for the ensembling of the first column, it cannot be sure that the same coefficient is optimal for more than eighteen thousand ensembling.</p>\n<p>Can the benefits of ensemble be ignored? Certainly not. So what to do?</p>\n<p>This is a complex problem and there is no unique answer for all challenges. For example, in this challenge, due to the large number of columns, we decided to make the most of the correlation value of the columns.</p>\n<p>We previously used Comparative Method and Snap to Grid in the Indoor Location &amp; Navigation challenge:</p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/1-3-indoor-navigation-cost-minimization-floor/notebook\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/1-3-indoor-navigation-cost-minimization-floor/notebook</a></p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/2-3-indoor-navigation-comparative-method\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/2-3-indoor-navigation-comparative-method</a></p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/3-3-g6-snap-to-grid-fix-the-timestamps\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/3-3-g6-snap-to-grid-fix-the-timestamps</a></p>\n<p>In the Tabular Playground Series challenge - Jul 2021, we used the Smart Ensembling method:<br>\n<a href=\"https://www.kaggle.com/code/mehrankazeminia/2-tps-jul-21-smart-ensembling\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/2-tps-jul-21-smart-ensembling</a></p>\n<p>In the Tabular Playground Series - Jul 2022 challenge, we used the Clustering-Ensembling method:<br>\n<a href=\"https://www.kaggle.com/code/mehrankazeminia/3-3-tps22jul-clustering-ensembling\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/3-3-tps22jul-clustering-ensembling</a></p>",
      "rawMarkdown": "Hi; thank you\nWe usually follow two rules:\n\n1- When two columns have a negative correlation:\nInstead of their Ensemble, choose one of them alone. (Probably the column that is used as a support and maybe has a lower general score will be selected)\n\n2- When two columns have a correlation close to one:\nInstead of combining them, choose one of them. (Probably the column that is used as the main one and maybe has a higher general score will be selected)\n\nPlease note that:\n\nIn many cases, only the final results of the calculations are accessible, that is, there is no access to the validation data. In these cases, Ensembling is possible by trial and error method, and actually Ensembling should be done in the dark.\n\nBut when the number of answer columns is more than one, the darkness increases. Because it is not known that after finding the right coefficient for ensembling the first columns, the same coefficient is optimal for ensembling the next columns.\n\nIn this challenge, the predict contains more than eighteen thousand columns. So if a coefficient is chosen for the ensembling of the first column, it cannot be sure that the same coefficient is optimal for more than eighteen thousand ensembling.\n\nCan the benefits of ensemble be ignored? Certainly not. So what to do?\n\nThis is a complex problem and there is no unique answer for all challenges. For example, in this challenge, due to the large number of columns, we decided to make the most of the correlation value of the columns.\n\nWe previously used Comparative Method and Snap to Grid in the Indoor Location & Navigation challenge:\n\nhttps://www.kaggle.com/code/mehrankazeminia/1-3-indoor-navigation-cost-minimization-floor/notebook\n\nhttps://www.kaggle.com/code/mehrankazeminia/2-3-indoor-navigation-comparative-method\n\nhttps://www.kaggle.com/code/mehrankazeminia/3-3-g6-snap-to-grid-fix-the-timestamps\n\nIn the Tabular Playground Series challenge - Jul 2021, we used the Smart Ensembling method:\nhttps://www.kaggle.com/code/mehrankazeminia/2-tps-jul-21-smart-ensembling\n\nIn the Tabular Playground Series - Jul 2022 challenge, we used the Clustering-Ensembling method:\nhttps://www.kaggle.com/code/mehrankazeminia/3-3-tps22jul-clustering-ensembling",
      "votes": null
    },
    {
      "id": "2547400",
      "postDate": "12/03/2023 13:25:31",
      "content": "<p>Thanks for your detailed answer ! </p>\n<p>Some more questions: </p>\n<p>I wonder - have you compared that with the blend just making an average ?</p>\n<p>We can do the same - not only \"column-wise\" (=target-wsie), but also \"row-wise\" (=sample-wise)  - have you thought in that direction, what do you think is it promising ? </p>\n<p>Have you tried to take into account the confidence of the models in their predicts  ?   I mean we can run the model with similar params, but a little modified params - we will get new predictions - and looking on variance of these predictions - we can get certain kind of measurements for the confidence . We can do it for each column (target), or for each row (sample), or even for each element.  We tried some of that, but not successfully, may be we are missing something.</p>\n<p>PS</p>\n<p>Here is some standard ideas from statistics - if we need to ensemble two variables - we should do it taking variance into account - bigger variance - less confidence - lower weight:<br>\nthe formula:</p>\n<p>`weight_1  = variance_2 / (variance_1 + variance_2 ) </p>\n<p>weight_2  = variance_1 / (variance_1 + variance_2 ) <br>\n`<br>\nAbove is for uncorrelated variable, and for correlated: </p>\n<p>`weight_1  = (variance_2 - Cov) / (variance_1 + variance_2 - 2*Cov) </p>\n<p>weight_2  = (variance_1 - Cov) / (variance_1 + variance_2 - 2*Cov ) <br>\n`</p>\n<p>note: weight_1 + weight_2 = 1 - automatically  - for both cases - without correlation or with correlation</p>\n<p>The problem to apply to our case - we do not have reliable estimates of variances and covariance, we can stril try to make some estimates and try something like that. We tried - but not successfully, I wonder someone have ideas on that ?</p>\n<p>PSPS</p>\n<p>Here are the formulas above not from the textbook, but from the chatGPT :)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F8abccf74d1df0ddc54baabedeed4acad%2Fphoto_2023-11-16_11-40-00.jpg?generation=1701609888771636&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F0562bf499683d60fe8da07c5ccf48fda%2Fphoto_2023-11-16_11-40-23.jpg?generation=1701609924198022&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks for your detailed answer ! \n\nSome more questions: \n\nI wonder - have you compared that with the blend just making an average ?\n\nWe can do the same - not only \"column-wise\" (=target-wsie), but also \"row-wise\" (=sample-wise)  - have you thought in that direction, what do you think is it promising ? \n\nHave you tried to take into account the confidence of the models in their predicts  ?   I mean we can run the model with similar params, but a little modified params - we will get new predictions - and looking on variance of these predictions - we can get certain kind of measurements for the confidence . We can do it for each column (target), or for each row (sample), or even for each element.  We tried some of that, but not successfully, may be we are missing something.\n\nPS\n\nHere is some standard ideas from statistics - if we need to ensemble two variables - we should do it taking variance into account - bigger variance - less confidence - lower weight:\nthe formula:\n\n`weight_1  = variance_2 / (variance_1 + variance_2 ) \n\nweight_2  = variance_1 / (variance_1 + variance_2 ) \n`\nAbove is for uncorrelated variable, and for correlated: \n\n`weight_1  = (variance_2 - Cov) / (variance_1 + variance_2 - 2*Cov) \n\nweight_2  = (variance_1 - Cov) / (variance_1 + variance_2 - 2*Cov ) \n`\n\nnote: weight_1 + weight_2 = 1 - automatically  - for both cases - without correlation or with correlation\n\nThe problem to apply to our case - we do not have reliable estimates of variances and covariance, we can stril try to make some estimates and try something like that. We tried - but not successfully, I wonder someone have ideas on that ?\n\nPSPS\n\nHere are the formulas above not from the textbook, but from the chatGPT :)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F8abccf74d1df0ddc54baabedeed4acad%2Fphoto_2023-11-16_11-40-00.jpg?generation=1701609888771636&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F0562bf499683d60fe8da07c5ccf48fda%2Fphoto_2023-11-16_11-40-23.jpg?generation=1701609924198022&alt=media)",
      "votes": null
    },
    {
      "id": "2547446",
      "postDate": "12/03/2023 14:01:25",
      "content": "<p>Despite the public score, private score, mistake in selection or mistake in the categories of samples, etc., it is clear that looking at the rows can be misleading. But looking independently at each prediction column, for their Ensembling, there is no risk. Even if the prediction columns are not independent of each other. We tried this issue thousands and thousands of times and even at the end of this competition, when the private scores were determined, the issue became clear to us once again.</p>\n<p>We are looking for a coefficient for the linear combination of two lists (a pair of lists). But to combine two pairs of lists, we need to look for two coefficients. For eighteen thousand pairs of lists, we should not expect a coefficient to be the best coefficient and the optimal coefficient.</p>\n<p>Let me give you an example: you may remember that for the initial Ensembling, we even used a general score of 0.720, because it would have made the result better. Why? Because it worked well for only a few columns. If in this match, instead of eighteen thousand columns, we were to predict only one column, this would never have happened and we should not have used the general score of 0.720.</p>",
      "rawMarkdown": "Despite the public score, private score, mistake in selection or mistake in the categories of samples, etc., it is clear that looking at the rows can be misleading. But looking independently at each prediction column, for their Ensembling, there is no risk. Even if the prediction columns are not independent of each other. We tried this issue thousands and thousands of times and even at the end of this competition, when the private scores were determined, the issue became clear to us once again.\n\nWe are looking for a coefficient for the linear combination of two lists (a pair of lists). But to combine two pairs of lists, we need to look for two coefficients. For eighteen thousand pairs of lists, we should not expect a coefficient to be the best coefficient and the optimal coefficient.\n\nLet me give you an example: you may remember that for the initial Ensembling, we even used a general score of 0.720, because it would have made the result better. Why? Because it worked well for only a few columns. If in this match, instead of eighteen thousand columns, we were to predict only one column, this would never have happened and we should not have used the general score of 0.720.",
      "votes": null
    },
    {
      "id": "2547641",
      "postDate": "12/03/2023 17:26:30",
      "content": "<p>Thanks for sharing  !<br>\nThe case of of 0.720 Jax-autoencoder model - just exactly my question:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/452515\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/452515</a></p>\n<p>Among the ideas proposed in comments - it was proposed that it has bigger values more often, than other models - and that is why - it it can help in blend.</p>\n<p>The idea that is due  to ONLY SELECTED COLUMNS - is new for me - would you be so kind to provide more details ? How did you come up with that ? What are these columns ? </p>\n<p>PS<br>\nI tried to blend that 0.720 to  several  blended Pybosts models - it did not work - the LB score - went down.  So I guess it works mainly for some weaker solutions, or may be for some NN which have does not catch big values, but catch small values better.</p>",
      "rawMarkdown": "Thanks for sharing  !\nThe case of of 0.720 Jax-autoencoder model - just exactly my question:\nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/452515\n\nAmong the ideas proposed in comments - it was proposed that it has bigger values more often, than other models - and that is why - it it can help in blend.\n\nThe idea that is due  to ONLY SELECTED COLUMNS - is new for me - would you be so kind to provide more details ? How did you come up with that ? What are these columns ? \n\nPS\nI tried to blend that 0.720 to  several  blended Pybosts models - it did not work - the LB score - went down.  So I guess it works mainly for some weaker solutions, or may be for some NN which have does not catch big values, but catch small values better.",
      "votes": null
    },
    {
      "id": "2547719",
      "postDate": "12/03/2023 18:43:17",
      "content": "<p>I think the best general Py-boost notebook is from <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>. which obtained these results:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F7f3096cde929117e24e7f9a73d484bde%2Fjj101.png?generation=1701628945608551&amp;alt=media\" alt=\"\"></p>\n<p>If you are careful when Ensembling the results of this notebook with 0.720 and do not consider the columns that have a negative correlation and only combine the rest of the columns with a factor of 0.05, the results will improve. I just did this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fbcc464865525ec68ddb05de07b057ece%2Fjj102.png?generation=1701628983539562&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I think the best general Py-boost notebook is from @ambrosm. which obtained these results:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F7f3096cde929117e24e7f9a73d484bde%2Fjj101.png?generation=1701628945608551&alt=media)\n\nIf you are careful when Ensembling the results of this notebook with 0.720 and do not consider the columns that have a negative correlation and only combine the rest of the columns with a factor of 0.05, the results will improve. I just did this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fbcc464865525ec68ddb05de07b057ece%2Fjj102.png?generation=1701628983539562&alt=media)",
      "votes": null
    },
    {
      "id": "2547725",
      "postDate": "12/03/2023 18:55:08",
      "content": "<p>We named this type of results \"Results of impure golden\". If you are interested in this topic, take a look at the notebooks I mentioned above. Also, the three notebooks below are exactly about the same topic. I hope that we will publish all these topics in one coherent article soon.</p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/1-tps22nov-pseudo-genetic-algorithm\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/1-tps22nov-pseudo-genetic-algorithm</a></p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/2-tps22nov-results-of-impure-golden-eda\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/2-tps22nov-results-of-impure-golden-eda</a></p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/3-tps22nov-golden-results-knn-lgbm\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/3-tps22nov-golden-results-knn-lgbm</a></p>",
      "rawMarkdown": "We named this type of results \"Results of impure golden\". If you are interested in this topic, take a look at the notebooks I mentioned above. Also, the three notebooks below are exactly about the same topic. I hope that we will publish all these topics in one coherent article soon.\n\nhttps://www.kaggle.com/code/mehrankazeminia/1-tps22nov-pseudo-genetic-algorithm\n\nhttps://www.kaggle.com/code/mehrankazeminia/2-tps22nov-results-of-impure-golden-eda\n\nhttps://www.kaggle.com/code/mehrankazeminia/3-tps22nov-golden-results-knn-lgbm",
      "votes": null
    },
    {
      "id": "2547740",
      "postDate": "12/03/2023 19:29:19",
      "content": "<p>Wow ! Cool ! <br>\nThanks a lot  for sharing ! <br>\nPS<br>\nWhat will happen if we combine pyboost with JAX , without negatively correlation dropping out ? <br>\nIn my experiments I used several pyboost blended , and 0.1 for JAX - so cannot compare directly</p>",
      "rawMarkdown": "Wow ! Cool ! \nThanks a lot  for sharing ! \nPS\nWhat will happen if we combine pyboost with JAX , without negatively correlation dropping out ? \nIn my experiments I used several pyboost blended , and 0.1 for JAX - so cannot compare directly",
      "votes": null
    },
    {
      "id": "2547778",
      "postDate": "12/03/2023 20:27:46",
      "content": "<p>If the columns with negative correlation are ignored and the combination factor is considered equal to 0.10: (best)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fef7cd49e7d2de235c50a95e56878716d%2Fjj103.png?generation=1701635101435479&amp;alt=media\" alt=\"\"></p>\n<p>If the columns with negative correlation are ignored and the combination factor is considered equal to 0.15:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fba47781ff043bc86bfcb253ecc46c706%2Fjj104.png?generation=1701635185955055&amp;alt=media\" alt=\"\"></p>\n<p>But if all columns are considered with a combination factor equal to 0.10:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fd883d0f0b8965c4ca8f9731868383330%2Fjj105.png?generation=1701635231696726&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "If the columns with negative correlation are ignored and the combination factor is considered equal to 0.10: (best)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fef7cd49e7d2de235c50a95e56878716d%2Fjj103.png?generation=1701635101435479&alt=media)\n\nIf the columns with negative correlation are ignored and the combination factor is considered equal to 0.15:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fba47781ff043bc86bfcb253ecc46c706%2Fjj104.png?generation=1701635185955055&alt=media)\n\nBut if all columns are considered with a combination factor equal to 0.10:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fd883d0f0b8965c4ca8f9731868383330%2Fjj105.png?generation=1701635231696726&alt=media)",
      "votes": null
    },
    {
      "id": "2547794",
      "postDate": "12/03/2023 21:04:10",
      "content": "<p>Wow ! That is really striking  ! <br>\nPS<br>\nComes to my mind that it might be interesting to take tsvd and blend tsvd components in the similar way, however  - no idea should it work or not…  </p>",
      "rawMarkdown": "Wow ! That is really striking  ! \nPS\nComes to my mind that it might be interesting to take tsvd and blend tsvd components in the similar way, however  - no idea should it work or not...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2547264,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "12/03/2023 10:59:23",
      "content": "<p>Thanks a lot for sharing ! <br>\nAnd thanks a lot for sharing your approaches during the challenge !<br>\nI guess quite a lot of participants used your work/ideas (including me).</p>\n<p>PS</p>\n<p>Would be you be so kind to give more details on blending with different coefficients for columns (targets).<br>\nHow exactly you chose the blend weights and what was the rationale ?<br>\nWe also thought in that direction but not quite successfully.<br>\n(The standard approach to get blend weight by blending OOF predictions and selected best local score - seems not really working here - because CV-LB correspondence is not good , especially for different models - like boosting vs NN). <br>\nAlso would be great to have more comments on the scatterplot - what the axis and dots ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2547289,
          "author_name": "mehrankazeminia",
          "author_url": "",
          "post_date": "12/03/2023 11:41:40",
          "content": "<p>Hi; thank you<br>\nWe usually follow two rules:</p>\n<p>1- When two columns have a negative correlation:<br>\nInstead of their Ensemble, choose one of them alone. (Probably the column that is used as a support and maybe has a lower general score will be selected)</p>\n<p>2- When two columns have a correlation close to one:<br>\nInstead of combining them, choose one of them. (Probably the column that is used as the main one and maybe has a higher general score will be selected)</p>\n<p>Please note that:</p>\n<p>In many cases, only the final results of the calculations are accessible, that is, there is no access to the validation data. In these cases, Ensembling is possible by trial and error method, and actually Ensembling should be done in the dark.</p>\n<p>But when the number of answer columns is more than one, the darkness increases. Because it is not known that after finding the right coefficient for ensembling the first columns, the same coefficient is optimal for ensembling the next columns.</p>\n<p>In this challenge, the predict contains more than eighteen thousand columns. So if a coefficient is chosen for the ensembling of the first column, it cannot be sure that the same coefficient is optimal for more than eighteen thousand ensembling.</p>\n<p>Can the benefits of ensemble be ignored? Certainly not. So what to do?</p>\n<p>This is a complex problem and there is no unique answer for all challenges. For example, in this challenge, due to the large number of columns, we decided to make the most of the correlation value of the columns.</p>\n<p>We previously used Comparative Method and Snap to Grid in the Indoor Location &amp; Navigation challenge:</p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/1-3-indoor-navigation-cost-minimization-floor/notebook\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/1-3-indoor-navigation-cost-minimization-floor/notebook</a></p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/2-3-indoor-navigation-comparative-method\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/2-3-indoor-navigation-comparative-method</a></p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/3-3-g6-snap-to-grid-fix-the-timestamps\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/3-3-g6-snap-to-grid-fix-the-timestamps</a></p>\n<p>In the Tabular Playground Series challenge - Jul 2021, we used the Smart Ensembling method:<br>\n<a href=\"https://www.kaggle.com/code/mehrankazeminia/2-tps-jul-21-smart-ensembling\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/2-tps-jul-21-smart-ensembling</a></p>\n<p>In the Tabular Playground Series - Jul 2022 challenge, we used the Clustering-Ensembling method:<br>\n<a href=\"https://www.kaggle.com/code/mehrankazeminia/3-3-tps22jul-clustering-ensembling\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/3-3-tps22jul-clustering-ensembling</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2547400,
              "author_name": "alexandervc",
              "author_url": "",
              "post_date": "12/03/2023 13:25:31",
              "content": "<p>Thanks for your detailed answer ! </p>\n<p>Some more questions: </p>\n<p>I wonder - have you compared that with the blend just making an average ?</p>\n<p>We can do the same - not only \"column-wise\" (=target-wsie), but also \"row-wise\" (=sample-wise)  - have you thought in that direction, what do you think is it promising ? </p>\n<p>Have you tried to take into account the confidence of the models in their predicts  ?   I mean we can run the model with similar params, but a little modified params - we will get new predictions - and looking on variance of these predictions - we can get certain kind of measurements for the confidence . We can do it for each column (target), or for each row (sample), or even for each element.  We tried some of that, but not successfully, may be we are missing something.</p>\n<p>PS</p>\n<p>Here is some standard ideas from statistics - if we need to ensemble two variables - we should do it taking variance into account - bigger variance - less confidence - lower weight:<br>\nthe formula:</p>\n<p>`weight_1  = variance_2 / (variance_1 + variance_2 ) </p>\n<p>weight_2  = variance_1 / (variance_1 + variance_2 ) <br>\n`<br>\nAbove is for uncorrelated variable, and for correlated: </p>\n<p>`weight_1  = (variance_2 - Cov) / (variance_1 + variance_2 - 2*Cov) </p>\n<p>weight_2  = (variance_1 - Cov) / (variance_1 + variance_2 - 2*Cov ) <br>\n`</p>\n<p>note: weight_1 + weight_2 = 1 - automatically  - for both cases - without correlation or with correlation</p>\n<p>The problem to apply to our case - we do not have reliable estimates of variances and covariance, we can stril try to make some estimates and try something like that. We tried - but not successfully, I wonder someone have ideas on that ?</p>\n<p>PSPS</p>\n<p>Here are the formulas above not from the textbook, but from the chatGPT :)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F8abccf74d1df0ddc54baabedeed4acad%2Fphoto_2023-11-16_11-40-00.jpg?generation=1701609888771636&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F0562bf499683d60fe8da07c5ccf48fda%2Fphoto_2023-11-16_11-40-23.jpg?generation=1701609924198022&amp;alt=media\" alt=\"\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2547446,
                  "author_name": "mehrankazeminia",
                  "author_url": "",
                  "post_date": "12/03/2023 14:01:25",
                  "content": "<p>Despite the public score, private score, mistake in selection or mistake in the categories of samples, etc., it is clear that looking at the rows can be misleading. But looking independently at each prediction column, for their Ensembling, there is no risk. Even if the prediction columns are not independent of each other. We tried this issue thousands and thousands of times and even at the end of this competition, when the private scores were determined, the issue became clear to us once again.</p>\n<p>We are looking for a coefficient for the linear combination of two lists (a pair of lists). But to combine two pairs of lists, we need to look for two coefficients. For eighteen thousand pairs of lists, we should not expect a coefficient to be the best coefficient and the optimal coefficient.</p>\n<p>Let me give you an example: you may remember that for the initial Ensembling, we even used a general score of 0.720, because it would have made the result better. Why? Because it worked well for only a few columns. If in this match, instead of eighteen thousand columns, we were to predict only one column, this would never have happened and we should not have used the general score of 0.720.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2547641,
                      "author_name": "alexandervc",
                      "author_url": "",
                      "post_date": "12/03/2023 17:26:30",
                      "content": "<p>Thanks for sharing  !<br>\nThe case of of 0.720 Jax-autoencoder model - just exactly my question:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/452515\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/452515</a></p>\n<p>Among the ideas proposed in comments - it was proposed that it has bigger values more often, than other models - and that is why - it it can help in blend.</p>\n<p>The idea that is due  to ONLY SELECTED COLUMNS - is new for me - would you be so kind to provide more details ? How did you come up with that ? What are these columns ? </p>\n<p>PS<br>\nI tried to blend that 0.720 to  several  blended Pybosts models - it did not work - the LB score - went down.  So I guess it works mainly for some weaker solutions, or may be for some NN which have does not catch big values, but catch small values better.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2547719,
                          "author_name": "mehrankazeminia",
                          "author_url": "",
                          "post_date": "12/03/2023 18:43:17",
                          "content": "<p>I think the best general Py-boost notebook is from <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>. which obtained these results:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F7f3096cde929117e24e7f9a73d484bde%2Fjj101.png?generation=1701628945608551&amp;alt=media\" alt=\"\"></p>\n<p>If you are careful when Ensembling the results of this notebook with 0.720 and do not consider the columns that have a negative correlation and only combine the rest of the columns with a factor of 0.05, the results will improve. I just did this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fbcc464865525ec68ddb05de07b057ece%2Fjj102.png?generation=1701628983539562&amp;alt=media\" alt=\"\"></p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2547725,
                              "author_name": "mehrankazeminia",
                              "author_url": "",
                              "post_date": "12/03/2023 18:55:08",
                              "content": "<p>We named this type of results \"Results of impure golden\". If you are interested in this topic, take a look at the notebooks I mentioned above. Also, the three notebooks below are exactly about the same topic. I hope that we will publish all these topics in one coherent article soon.</p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/1-tps22nov-pseudo-genetic-algorithm\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/1-tps22nov-pseudo-genetic-algorithm</a></p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/2-tps22nov-results-of-impure-golden-eda\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/2-tps22nov-results-of-impure-golden-eda</a></p>\n<p><a href=\"https://www.kaggle.com/code/mehrankazeminia/3-tps22nov-golden-results-knn-lgbm\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/3-tps22nov-golden-results-knn-lgbm</a></p>",
                              "votes": null,
                              "replies": []
                            },
                            {
                              "id": 2547740,
                              "author_name": "alexandervc",
                              "author_url": "",
                              "post_date": "12/03/2023 19:29:19",
                              "content": "<p>Wow ! Cool ! <br>\nThanks a lot  for sharing ! <br>\nPS<br>\nWhat will happen if we combine pyboost with JAX , without negatively correlation dropping out ? <br>\nIn my experiments I used several pyboost blended , and 0.1 for JAX - so cannot compare directly</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2547778,
                                  "author_name": "mehrankazeminia",
                                  "author_url": "",
                                  "post_date": "12/03/2023 20:27:46",
                                  "content": "<p>If the columns with negative correlation are ignored and the combination factor is considered equal to 0.10: (best)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fef7cd49e7d2de235c50a95e56878716d%2Fjj103.png?generation=1701635101435479&amp;alt=media\" alt=\"\"></p>\n<p>If the columns with negative correlation are ignored and the combination factor is considered equal to 0.15:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fba47781ff043bc86bfcb253ecc46c706%2Fjj104.png?generation=1701635185955055&amp;alt=media\" alt=\"\"></p>\n<p>But if all columns are considered with a combination factor equal to 0.10:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fd883d0f0b8965c4ca8f9731868383330%2Fjj105.png?generation=1701635231696726&amp;alt=media\" alt=\"\"></p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2547794,
                                      "author_name": "alexandervc",
                                      "author_url": "",
                                      "post_date": "12/03/2023 21:04:10",
                                      "content": "<p>Wow ! That is really striking  ! <br>\nPS<br>\nComes to my mind that it might be interesting to take tsvd and blend tsvd components in the similar way, however  - no idea should it work or not…  </p>",
                                      "votes": null,
                                      "replies": []
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2546812": "The prediction in this competition included more than eighteen thousand columns, which may have less history in machine learning competitions, while the number of features was also very limited. These themes made every detail seem very important. Even regardless of the result of the match, it was definitely a good experience for us.\n\nIn this challenge, there are only two features, namely \"cell_type\" and \"sm_name\". That's why we used \"Feature Augmentation\". We added two new columns (two new features) separately for each prediction column as follows:\n\nIf we separate the cells based on 'cell_type' and assume that the drugs will usually have similar responses on each of these divisions, we can hope that by finding the average effects, we have obtained a new feature. For example, for y0 and the new feature of zero column, the correlation coefficient is 0.24. Of course, this amount is repeated for other columns as well.\n\nAlso, if we separate the cells based on 'sm_name', we get a new feature by finding the average effects. In this case, for y0 and the new feature of column zero, the correlation coefficient is 0.62, and this value is almost repeated for other columns.\n\n\nIn addition, we added other features by using \"SMILES\", which were used for all prediction columns at the same time. At first, we added about five hundred binary features using \"fragments of SMILES\" and then we added about two thousand new binary features using \"morgan fingerprint from SMILES\".\n\nWe have publicly released two notebooks that cover the above topicss:\n\nhttps://www.kaggle.com/code/mehrankazeminia/1-op2-eda-linearsvr-regressorchain\n\nhttps://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles\n\nFinally, the results of LinearSVR and neural network and NLP and PYBOOST were combined and we used \"Separately Ensembling for Each Column\". That is, Ensembling was done with different coefficients (based on the correlation value).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Ffc0fcd5fe67f70542c4e5409a76b81ad%2Fphoto_2023-12-03_01-15-24.jpg?generation=1701553548981440&alt=media)",
    "2547264": "Thanks a lot for sharing ! \nAnd thanks a lot for sharing your approaches during the challenge !\nI guess quite a lot of participants used your work/ideas (including me).\n\nPS\n\nWould be you be so kind to give more details on blending with different coefficients for columns (targets).\nHow exactly you chose the blend weights and what was the rationale ?\nWe also thought in that direction but not quite successfully.\n(The standard approach to get blend weight by blending OOF predictions and selected best local score - seems not really working here - because CV-LB correspondence is not good , especially for different models - like boosting vs NN). \nAlso would be great to have more comments on the scatterplot - what the axis and dots ?",
    "2547289": "Hi; thank you\nWe usually follow two rules:\n\n1- When two columns have a negative correlation:\nInstead of their Ensemble, choose one of them alone. (Probably the column that is used as a support and maybe has a lower general score will be selected)\n\n2- When two columns have a correlation close to one:\nInstead of combining them, choose one of them. (Probably the column that is used as the main one and maybe has a higher general score will be selected)\n\nPlease note that:\n\nIn many cases, only the final results of the calculations are accessible, that is, there is no access to the validation data. In these cases, Ensembling is possible by trial and error method, and actually Ensembling should be done in the dark.\n\nBut when the number of answer columns is more than one, the darkness increases. Because it is not known that after finding the right coefficient for ensembling the first columns, the same coefficient is optimal for ensembling the next columns.\n\nIn this challenge, the predict contains more than eighteen thousand columns. So if a coefficient is chosen for the ensembling of the first column, it cannot be sure that the same coefficient is optimal for more than eighteen thousand ensembling.\n\nCan the benefits of ensemble be ignored? Certainly not. So what to do?\n\nThis is a complex problem and there is no unique answer for all challenges. For example, in this challenge, due to the large number of columns, we decided to make the most of the correlation value of the columns.\n\nWe previously used Comparative Method and Snap to Grid in the Indoor Location & Navigation challenge:\n\nhttps://www.kaggle.com/code/mehrankazeminia/1-3-indoor-navigation-cost-minimization-floor/notebook\n\nhttps://www.kaggle.com/code/mehrankazeminia/2-3-indoor-navigation-comparative-method\n\nhttps://www.kaggle.com/code/mehrankazeminia/3-3-g6-snap-to-grid-fix-the-timestamps\n\nIn the Tabular Playground Series challenge - Jul 2021, we used the Smart Ensembling method:\nhttps://www.kaggle.com/code/mehrankazeminia/2-tps-jul-21-smart-ensembling\n\nIn the Tabular Playground Series - Jul 2022 challenge, we used the Clustering-Ensembling method:\nhttps://www.kaggle.com/code/mehrankazeminia/3-3-tps22jul-clustering-ensembling",
    "2547400": "Thanks for your detailed answer ! \n\nSome more questions: \n\nI wonder - have you compared that with the blend just making an average ?\n\nWe can do the same - not only \"column-wise\" (=target-wsie), but also \"row-wise\" (=sample-wise)  - have you thought in that direction, what do you think is it promising ? \n\nHave you tried to take into account the confidence of the models in their predicts  ?   I mean we can run the model with similar params, but a little modified params - we will get new predictions - and looking on variance of these predictions - we can get certain kind of measurements for the confidence . We can do it for each column (target), or for each row (sample), or even for each element.  We tried some of that, but not successfully, may be we are missing something.\n\nPS\n\nHere is some standard ideas from statistics - if we need to ensemble two variables - we should do it taking variance into account - bigger variance - less confidence - lower weight:\nthe formula:\n\n`weight_1  = variance_2 / (variance_1 + variance_2 ) \n\nweight_2  = variance_1 / (variance_1 + variance_2 ) \n`\nAbove is for uncorrelated variable, and for correlated: \n\n`weight_1  = (variance_2 - Cov) / (variance_1 + variance_2 - 2*Cov) \n\nweight_2  = (variance_1 - Cov) / (variance_1 + variance_2 - 2*Cov ) \n`\n\nnote: weight_1 + weight_2 = 1 - automatically  - for both cases - without correlation or with correlation\n\nThe problem to apply to our case - we do not have reliable estimates of variances and covariance, we can stril try to make some estimates and try something like that. We tried - but not successfully, I wonder someone have ideas on that ?\n\nPSPS\n\nHere are the formulas above not from the textbook, but from the chatGPT :)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F8abccf74d1df0ddc54baabedeed4acad%2Fphoto_2023-11-16_11-40-00.jpg?generation=1701609888771636&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F0562bf499683d60fe8da07c5ccf48fda%2Fphoto_2023-11-16_11-40-23.jpg?generation=1701609924198022&alt=media)",
    "2547446": "Despite the public score, private score, mistake in selection or mistake in the categories of samples, etc., it is clear that looking at the rows can be misleading. But looking independently at each prediction column, for their Ensembling, there is no risk. Even if the prediction columns are not independent of each other. We tried this issue thousands and thousands of times and even at the end of this competition, when the private scores were determined, the issue became clear to us once again.\n\nWe are looking for a coefficient for the linear combination of two lists (a pair of lists). But to combine two pairs of lists, we need to look for two coefficients. For eighteen thousand pairs of lists, we should not expect a coefficient to be the best coefficient and the optimal coefficient.\n\nLet me give you an example: you may remember that for the initial Ensembling, we even used a general score of 0.720, because it would have made the result better. Why? Because it worked well for only a few columns. If in this match, instead of eighteen thousand columns, we were to predict only one column, this would never have happened and we should not have used the general score of 0.720.",
    "2547641": "Thanks for sharing  !\nThe case of of 0.720 Jax-autoencoder model - just exactly my question:\nhttps://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/452515\n\nAmong the ideas proposed in comments - it was proposed that it has bigger values more often, than other models - and that is why - it it can help in blend.\n\nThe idea that is due  to ONLY SELECTED COLUMNS - is new for me - would you be so kind to provide more details ? How did you come up with that ? What are these columns ? \n\nPS\nI tried to blend that 0.720 to  several  blended Pybosts models - it did not work - the LB score - went down.  So I guess it works mainly for some weaker solutions, or may be for some NN which have does not catch big values, but catch small values better.",
    "2547719": "I think the best general Py-boost notebook is from @ambrosm. which obtained these results:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2F7f3096cde929117e24e7f9a73d484bde%2Fjj101.png?generation=1701628945608551&alt=media)\n\nIf you are careful when Ensembling the results of this notebook with 0.720 and do not consider the columns that have a negative correlation and only combine the rest of the columns with a factor of 0.05, the results will improve. I just did this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fbcc464865525ec68ddb05de07b057ece%2Fjj102.png?generation=1701628983539562&alt=media)",
    "2547725": "We named this type of results \"Results of impure golden\". If you are interested in this topic, take a look at the notebooks I mentioned above. Also, the three notebooks below are exactly about the same topic. I hope that we will publish all these topics in one coherent article soon.\n\nhttps://www.kaggle.com/code/mehrankazeminia/1-tps22nov-pseudo-genetic-algorithm\n\nhttps://www.kaggle.com/code/mehrankazeminia/2-tps22nov-results-of-impure-golden-eda\n\nhttps://www.kaggle.com/code/mehrankazeminia/3-tps22nov-golden-results-knn-lgbm",
    "2547740": "Wow ! Cool ! \nThanks a lot  for sharing ! \nPS\nWhat will happen if we combine pyboost with JAX , without negatively correlation dropping out ? \nIn my experiments I used several pyboost blended , and 0.1 for JAX - so cannot compare directly",
    "2547778": "If the columns with negative correlation are ignored and the combination factor is considered equal to 0.10: (best)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fef7cd49e7d2de235c50a95e56878716d%2Fjj103.png?generation=1701635101435479&alt=media)\n\nIf the columns with negative correlation are ignored and the combination factor is considered equal to 0.15:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fba47781ff043bc86bfcb253ecc46c706%2Fjj104.png?generation=1701635185955055&alt=media)\n\nBut if all columns are considered with a combination factor equal to 0.10:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4288268%2Fd883d0f0b8965c4ca8f9731868383330%2Fjj105.png?generation=1701635231696726&alt=media)",
    "2547794": "Wow ! That is really striking  ! \nPS\nComes to my mind that it might be interesting to take tsvd and blend tsvd components in the similar way, however  - no idea should it work or not..."
  },
  "source": "meta"
}