{
  "id": 351725,
  "title": "Correlations between all inputs and targets for BOTH Multiome and CITEseq",
  "url": "/competitions/open-problems-multimodal/discussion/351725",
  "author_name": "",
  "post_date": "2022-09-11T15:46:01.485010200Z",
  "votes": 30,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>So I finally decided to make some good use of the access to 128GB RAM machines from Saturn Cloud that we have been given. (see <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999\" target=\"_blank\">this discussion</a> )</p>\n<p>I decided to compute the correlation coefficients between all inputs and targets, and to try and analyze them a bit.</p>\n<p>The pre-computed correlations are in this dataset:<br>\n<a href=\"https://www.kaggle.com/datasets/fabiencrom/msci-correlations\" target=\"_blank\">https://www.kaggle.com/datasets/fabiencrom/msci-correlations</a></p>\n<p>Here is the notebook I used to generate it:<br>\n<a href=\"https://www.kaggle.com/fabiencrom/msci-generating-all-correlations-inputs-targets\" target=\"_blank\">https://www.kaggle.com/fabiencrom/msci-generating-all-correlations-inputs-targets</a></p>\n<p>Note that you can use this notebook to compute the CITEseq correlations on Kaggle; but for the Multiome data you would need around 40GB RAM (hence the usefulness of the Saturn Cloud machine…). Also note that I removed small correlations from the Multiome correlations to be able to store it as a sparse matrix. This way, the data uses less than 1GB instead of more than 10GB (there are &gt;5billion correlations for the Multiome case!)</p>\n<p>I made an initial analysis of these correlations in two notebooks:<br>\n<a href=\"https://www.kaggle.com/fabiencrom/msci-correlations-eda-multiome\" target=\"_blank\">https://www.kaggle.com/fabiencrom/msci-correlations-eda-multiome</a><br>\n<a href=\"https://www.kaggle.com/fabiencrom/msci-correlations-eda-citeseq\" target=\"_blank\">https://www.kaggle.com/fabiencrom/msci-correlations-eda-citeseq</a></p>\n<p>My main findings so far:</p>\n<p>EDIT: Following discussion with other kagglers below (thanks to them), I realized my interpretations of the correlation numbers are wrong. The main point is that <strong>it seems that when an input has a value equal to zero, we should consider it as a missing value and not a real zero</strong>. As I did not know about that, I computed the correlations on all data for all inputs. And thus the correlation obtained are a mix of the real inputs/targets correlations and some spurious correlations due to the choices made of measuring the input or not for a given cell. I should probably redo this study with this information in mind.</p>\n<p>CITEseq:</p>\n<ul>\n<li>Some inputs have a strong correlation with all targets. They should probably not be discarded if doing dimension reduction.</li>\n<li></li>\n<li></li>\n<li>On average, a target has a correlation &gt; 0.3 with about 9 inputs.</li>\n</ul>\n<p>Multiome:</p>\n<ul>\n<li>correlations are much weaker than with CITEseq</li>\n<li>560 targets in the training set are always zero! (So i guess they should be set to zero in submissions)</li>\n<li>10 inputs in the training set are always zero (I think I saw this fact mentioned here before, but could not find a link)</li>\n<li>many hints that there are both subgroups of targets and subgroups of inputs varying together.</li>\n<li>On average, a target has a correlation &gt; 0.1 with about 3 inputs</li>\n</ul>\n<p>Apart from these findings, I expect that models could use the correlation informations for regularization (e.g. enforcing some sparsity in the matrix of a linear regression by setting to zero coefficients associated with input/target pairs with low correlations)</p>\n<p>If you have in-domain knowledge, I would be very interested by your feedback on these results.</p>",
  "messages": [
    {
      "id": "1934793",
      "postDate": "09/11/2022 15:46:01",
      "content": "<p>Hi,</p>\n<p>So I finally decided to make some good use of the access to 128GB RAM machines from Saturn Cloud that we have been given. (see <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999\" target=\"_blank\">this discussion</a> )</p>\n<p>I decided to compute the correlation coefficients between all inputs and targets, and to try and analyze them a bit.</p>\n<p>The pre-computed correlations are in this dataset:<br>\n<a href=\"https://www.kaggle.com/datasets/fabiencrom/msci-correlations\" target=\"_blank\">https://www.kaggle.com/datasets/fabiencrom/msci-correlations</a></p>\n<p>Here is the notebook I used to generate it:<br>\n<a href=\"https://www.kaggle.com/fabiencrom/msci-generating-all-correlations-inputs-targets\" target=\"_blank\">https://www.kaggle.com/fabiencrom/msci-generating-all-correlations-inputs-targets</a></p>\n<p>Note that you can use this notebook to compute the CITEseq correlations on Kaggle; but for the Multiome data you would need around 40GB RAM (hence the usefulness of the Saturn Cloud machine…). Also note that I removed small correlations from the Multiome correlations to be able to store it as a sparse matrix. This way, the data uses less than 1GB instead of more than 10GB (there are &gt;5billion correlations for the Multiome case!)</p>\n<p>I made an initial analysis of these correlations in two notebooks:<br>\n<a href=\"https://www.kaggle.com/fabiencrom/msci-correlations-eda-multiome\" target=\"_blank\">https://www.kaggle.com/fabiencrom/msci-correlations-eda-multiome</a><br>\n<a href=\"https://www.kaggle.com/fabiencrom/msci-correlations-eda-citeseq\" target=\"_blank\">https://www.kaggle.com/fabiencrom/msci-correlations-eda-citeseq</a></p>\n<p>My main findings so far:</p>\n<p>EDIT: Following discussion with other kagglers below (thanks to them), I realized my interpretations of the correlation numbers are wrong. The main point is that <strong>it seems that when an input has a value equal to zero, we should consider it as a missing value and not a real zero</strong>. As I did not know about that, I computed the correlations on all data for all inputs. And thus the correlation obtained are a mix of the real inputs/targets correlations and some spurious correlations due to the choices made of measuring the input or not for a given cell. I should probably redo this study with this information in mind.</p>\n<p>CITEseq:</p>\n<ul>\n<li>Some inputs have a strong correlation with all targets. They should probably not be discarded if doing dimension reduction.</li>\n<li></li>\n<li></li>\n<li>On average, a target has a correlation &gt; 0.3 with about 9 inputs.</li>\n</ul>\n<p>Multiome:</p>\n<ul>\n<li>correlations are much weaker than with CITEseq</li>\n<li>560 targets in the training set are always zero! (So i guess they should be set to zero in submissions)</li>\n<li>10 inputs in the training set are always zero (I think I saw this fact mentioned here before, but could not find a link)</li>\n<li>many hints that there are both subgroups of targets and subgroups of inputs varying together.</li>\n<li>On average, a target has a correlation &gt; 0.1 with about 3 inputs</li>\n</ul>\n<p>Apart from these findings, I expect that models could use the correlation informations for regularization (e.g. enforcing some sparsity in the matrix of a linear regression by setting to zero coefficients associated with input/target pairs with low correlations)</p>\n<p>If you have in-domain knowledge, I would be very interested by your feedback on these results.</p>",
      "rawMarkdown": "Hi,\n\nSo I finally decided to make some good use of the access to 128GB RAM machines from Saturn Cloud that we have been given. (see [this discussion](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999) )\n\nI decided to compute the correlation coefficients between all inputs and targets, and to try and analyze them a bit.\n\nThe pre-computed correlations are in this dataset:\nhttps://www.kaggle.com/datasets/fabiencrom/msci-correlations\n\nHere is the notebook I used to generate it:\nhttps://www.kaggle.com/fabiencrom/msci-generating-all-correlations-inputs-targets\n\nNote that you can use this notebook to compute the CITEseq correlations on Kaggle; but for the Multiome data you would need around 40GB RAM (hence the usefulness of the Saturn Cloud machine...). Also note that I removed small correlations from the Multiome correlations to be able to store it as a sparse matrix. This way, the data uses less than 1GB instead of more than 10GB (there are >5billion correlations for the Multiome case!)\n\nI made an initial analysis of these correlations in two notebooks:\nhttps://www.kaggle.com/fabiencrom/msci-correlations-eda-multiome\nhttps://www.kaggle.com/fabiencrom/msci-correlations-eda-citeseq\n\nMy main findings so far:\n\nEDIT: Following discussion with other kagglers below (thanks to them), I realized my interpretations of the correlation numbers are wrong. The main point is that **it seems that when an input has a value equal to zero, we should consider it as a missing value and not a real zero**. As I did not know about that, I computed the correlations on all data for all inputs. And thus the correlation obtained are a mix of the real inputs/targets correlations and some spurious correlations due to the choices made of measuring the input or not for a given cell. I should probably redo this study with this information in mind.\n\nCITEseq:\n- Some inputs have a strong correlation with all targets. They should probably not be discarded if doing dimension reduction.\n- ~~There are some \"Global inhibitors\" (resp. \"Global Enhancers\") that are strongly negatively (resp. positively) correlated with all targets. e.g. `ENSG00000129824_RPS4Y1` is a global inhibitor and `ENSG00000229807_XIST` is a global enhancer. (Not sure \"enhancer\" is the correct term in biology; please inform me)~~\n- ~~Contrary to what was suggested in [this discussion](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349242), genes and proteins related by names do not always have a strong correlation. For example, `ENSG00000114013_CD86` do have a correlation of 0.235 with protein `CD86`. But `ENSG00000120217_CD274` only has a correlation of 0.026 with `CD274` and rank only as the 798th most correlated input with `CD274`~~\n- On average, a target has a correlation > 0.3 with about 9 inputs.\n \nMultiome:\n- correlations are much weaker than with CITEseq\n- 560 targets in the training set are always zero! (So i guess they should be set to zero in submissions)\n- 10 inputs in the training set are always zero (I think I saw this fact mentioned here before, but could not find a link)\n- many hints that there are both subgroups of targets and subgroups of inputs varying together.\n- On average, a target has a correlation > 0.1 with about 3 inputs\n\nApart from these findings, I expect that models could use the correlation informations for regularization (e.g. enforcing some sparsity in the matrix of a linear regression by setting to zero coefficients associated with input/target pairs with low correlations)\n\nIf you have in-domain knowledge, I would be very interested by your feedback on these results.",
      "votes": null
    },
    {
      "id": "1934959",
      "postDate": "09/11/2022 17:31:34",
      "content": "<p>Nice ! Thanks for sharing ! <br>\n<a href=\"https://en.wikipedia.org/wiki/Enhancer_(genetics\" target=\"_blank\">https://en.wikipedia.org/wiki/Enhancer_(genetics</a>)<br>\nIt is already in use, with the other meaning.  (It is related to Multiome task)</p>\n<p>Remark: \"Contrary to what was suggested… \" - that is indeed surprising (at least for me). I'll try to ask around. </p>\n<p>ENSG00000229807_XIST  - that is \"XIST\" - it is quite famous <a href=\"https://en.wikipedia.org/wiki/XIST\" target=\"_blank\">https://en.wikipedia.org/wiki/XIST</a><br>\nstrange to see such correlation.   I'll try to ask around. <br>\nRPS4Y1 - ribosomal protein - again, I'll try to ask. </p>\n<p>\"On average, a target has a correlation &gt; 0.3 with about 9 inputs. \"<br>\nIs it possible to make a table \"Target\", TopCorrelated1, TopCorrelated2,TopCorrelated3 ………<br>\nThan we might either use GSEA, or look by eye if something appears </p>\n<p>PS<br>\nCan you please look: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900</a><br>\nIt is interesting to understand and probably biologically interpret the feature importance . <br>\nPSPS<br>\nHere is also some analysis of correlations:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo?scriptVersionId=103869738&amp;cellId=16\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo?scriptVersionId=103869738&amp;cellId=16</a></p>\n<p>PSPSPS<br>\nJoin our discussions : <a href=\"https://t.me/sberlogacompete\" target=\"_blank\">https://t.me/sberlogacompete</a></p>",
      "rawMarkdown": "Nice ! Thanks for sharing ! \nhttps://en.wikipedia.org/wiki/Enhancer_(genetics)\nIt is already in use, with the other meaning.  (It is related to Multiome task)\n\nRemark: \"Contrary to what was suggested... \" - that is indeed surprising (at least for me). I'll try to ask around. \n\nENSG00000229807_XIST  - that is \"XIST\" - it is quite famous https://en.wikipedia.org/wiki/XIST\nstrange to see such correlation.   I'll try to ask around. \nRPS4Y1 - ribosomal protein - again, I'll try to ask. \n\n\"On average, a target has a correlation > 0.3 with about 9 inputs. \"\nIs it possible to make a table \"Target\", TopCorrelated1, TopCorrelated2,TopCorrelated3 .........\nThan we might either use GSEA, or look by eye if something appears \n\n\n\nPS\nCan you please look: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900\nIt is interesting to understand and probably biologically interpret the feature importance . \nPSPS\nHere is also some analysis of correlations:\nhttps://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo?scriptVersionId=103869738&cellId=16\n\nPSPSPS\nJoin our discussions : https://t.me/sberlogacompete",
      "votes": null
    },
    {
      "id": "1934993",
      "postDate": "09/11/2022 18:08:22",
      "content": "<blockquote>\n  <p>It is already in use, with the other meaning. (It is related to Multiome task)</p>\n</blockquote>\n<p>You mean my use of it is incorrect? What would be the proper term for a gene that prevent expression of a protein?</p>\n<blockquote>\n  <p>RPS4Y1 - ribosomal protein - again, I'll try to ask.</p>\n</blockquote>\n<p>Thank you :-)</p>\n<blockquote>\n  <p>Is it possible to make a table \"Target\", TopCorrelated1, TopCorrelated2,TopCorrelated3 ………</p>\n</blockquote>\n<p>At the bottom of <a href=\"https://www.kaggle.com/fabiencrom/msci-correlations-eda-citeseq\" target=\"_blank\">this notebook</a> (same as linked above), you have a long output that, for each target protein gives the top 5 most correlated genes (easy to change to top 10 or more in the code if you want) + \"associated genes\". Isn't that what you are suggesting? (or I am misunderstanding)</p>\n<blockquote>\n  <p>Here is also some analysis of correlations:</p>\n</blockquote>\n<p>Oh thank you, I had not checked this before. But the correlations you are displaying are between targets and not between inputs and targets, correct? (I guess I will understand if I take the time to read the code, but just to know)</p>",
      "rawMarkdown": "> It is already in use, with the other meaning. (It is related to Multiome task)\n\nYou mean my use of it is incorrect? What would be the proper term for a gene that prevent expression of a protein?\n\n> RPS4Y1 - ribosomal protein - again, I'll try to ask.\n\nThank you :-)\n\n>Is it possible to make a table \"Target\", TopCorrelated1, TopCorrelated2,TopCorrelated3 ………\n\nAt the bottom of [this notebook](https://www.kaggle.com/fabiencrom/msci-correlations-eda-citeseq) (same as linked above), you have a long output that, for each target protein gives the top 5 most correlated genes (easy to change to top 10 or more in the code if you want) + \"associated genes\". Isn't that what you are suggesting? (or I am misunderstanding)\n\n>Here is also some analysis of correlations:\n\nOh thank you, I had not checked this before. But the correlations you are displaying are between targets and not between inputs and targets, correct? (I guess I will understand if I take the time to read the code, but just to know)",
      "votes": null
    },
    {
      "id": "1935004",
      "postDate": "09/11/2022 18:26:28",
      "content": "<p>-- You mean my use of it is incorrect?<br>\nSorry, yes. <br>\n-- What would be the proper term for a gene that prevent expression of a protein?<br>\nNot sure I know that. </p>\n<p>-- At the bottom of this notebook (same as linked above) ……<br>\nNice ! First look is the following :<br>\n1)  quite often GATA1 appears - \"GATA-binding factor 1 or GATA-1 (also termed Erythroid transcription factor) \"  <a href=\"https://en.wikipedia.org/wiki/GATA1\" target=\"_blank\">https://en.wikipedia.org/wiki/GATA1</a> <br>\nThat is gene strongly transcribed for Ery cells. <br>\nThere are might be some CD marking Ery - than it would be and an explantation.<br>\n2)We also see lots of XIST, RPS - I would suggest to exclude them. <br>\n3) There are CD** genes - that probably is natural, it might be related to my analysis of targets below:</p>\n<p>-- But the correlations you are displaying are between targets <br>\nYes. <br>\nTop correlated groups:<br>\n['CD71' 'CD115' 'CD88']<br>\n0.6 correlation_threshold <br>\n['CD155' 'CD112' 'CD47' 'HLA-A-B-C' 'CD45RA' 'CD31' 'CD11a' 'CD13' 'CD29'<br>\n 'CD81' 'CD18' 'CD45' 'CD49d' 'CD162']<br>\nI will try to think on bio interpretation some time later.</p>",
      "rawMarkdown": "You mean my use of it is incorrect?\nSorry, yes. \n-- What would be the proper term for a gene that prevent expression of a protein?\nNot sure I know that. \n\n-- At the bottom of this notebook (same as linked above) ......\nNice ! First look is the following :\n1)  quite often GATA1 appears - \"GATA-binding factor 1 or GATA-1 (also termed Erythroid transcription factor) \"  https://en.wikipedia.org/wiki/GATA1 \nThat is gene strongly transcribed for Ery cells. \nThere are might be some CD marking Ery - than it would be and an explantation.\n2)We also see lots of XIST, RPS - I would suggest to exclude them. \n3) There are CD** genes - that probably is natural, it might be related to my analysis of targets below:\n\n-- But the correlations you are displaying are between targets \nYes. \nTop correlated groups:\n['CD71' 'CD115' 'CD88']\n0.6 correlation_threshold \n['CD155' 'CD112' 'CD47' 'HLA-A-B-C' 'CD45RA' 'CD31' 'CD11a' 'CD13' 'CD29'\n 'CD81' 'CD18' 'CD45' 'CD49d' 'CD162']\nI will try to think on bio interpretation some time later.",
      "votes": null
    },
    {
      "id": "1935018",
      "postDate": "09/11/2022 18:34:46",
      "content": "<p>It is possible to calculate correlations with the cell types:<br>\nsomething like ordering cell types as follows:</p>\n<p>1 HSC <br>\n2 NeuP, MoP<br>\n3 MastP<br>\n4 Eryp , MkP, BP</p>\n<p>e.g. assigning cell types these numbers .</p>\n<p>Or <br>\nbetter taking the first component UMAP1 of the following UMAP:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo?scriptVersionId=103869738&amp;cellId=22\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo?scriptVersionId=103869738&amp;cellId=22</a></p>\n<p>That would give what is related to the cell type,<br>\nand that would probably help the interpretations. </p>",
      "rawMarkdown": "It is possible to calculate correlations with the cell types:\nsomething like ordering cell types as follows:\n\n1 HSC \n2 NeuP, MoP\n3 MastP\n4 Eryp , MkP, BP\n\ne.g. assigning cell types these numbers .\n\nOr \nbetter taking the first component UMAP1 of the following UMAP:\nhttps://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo?scriptVersionId=103869738&cellId=22\n\nThat would give what is related to the cell type,\nand that would probably help the interpretations.",
      "votes": null
    },
    {
      "id": "1935253",
      "postDate": "09/12/2022 01:20:30",
      "content": "<p>This is really interesting and at the same time, incredibly confusing!  😃</p>\n<p>First off, very nice work. The summary you presented above and your notebooks are exceptional.  Regarding some of your findings…</p>\n<blockquote>\n  <p>There are some \"Global inhibitors\" (resp. \"Global Enhancers\") that are strongly negatively (resp. positively) correlated with all targets. e.g. ENSG00000129824_RPS4Y1 is a global inhibitor and ENSG00000229807_XIST is a global enhancer. (Not sure \"enhancer\" is the correct term in biology; please inform me)</p>\n</blockquote>\n<p>Regarding \"enhancers\", as <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> mentioned, this is indeed a term in cellular biology.  Typically it refers to a region of DNA that is associated with increasing (or enhancing) transcription of a gene.  In your case, I would call what you've found \"activators\", but I'm sure there is some other term that might be more appropriate.</p>\n<p>For your Global Inhibitor and Global Enhancer/Activator, these are puzzling.  </p>\n<p>The inhibitor, <a href=\"https://www.genecards.org/cgi-bin/carddisp.pl?gene=RPS4Y1\" target=\"_blank\">RPS4Y1</a> is a protein associated with a subunit of ribosomes.  Briefly, ribosomes are protein complexes that facilitate decoding mRNA in to proteins, a process know and translation.  It seems odd to me that it would be negatively related to protein levels as it is directly involved in protein production from mRNA.  Interestingly, RSP4Y1 is coded on the Y  chromosome, so only samples from males should show any levels.  There was another notebook that suggested there were 3 male and 1 female donor.</p>\n<p>For the enhancer/activator, <a href=\"https://www.genecards.org/cgi-bin/carddisp.pl?gene=XIST\" target=\"_blank\">XIST</a>, this is not a protein coding gene but rather the RNA produced from this gene is involved in X-chromosome inactivation.  Females have 2 X chromosomes but if both are active, it is toxic to the cell.  One X chromosome is rendered inactive by genes from the XIC region, and this is one of those genes.  Males have one X chromosome and one Y chromosome and so should not have any levels of XIST as their X chromosome is not inactivated.</p>\n<p>I'm not doubting your findings, but in both cases, the direction of correlation is opposite what I might expect.  Biology never ceases to offer up surprises, though.  Is there a possibility that the X- and Y-chromosome linkages and the mixture of male and female donors a possible explanation?</p>\n<blockquote>\n  <p>Contrary to what was suggested in this discussion, genes and proteins related by names do not always have a strong correlation. </p>\n</blockquote>\n<p>This doesn't surprise me.  I mentioned in another discussion <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350863#1933699\" target=\"_blank\">here</a> that correlation between mRNA levels and protein levels are not as consistent as we might hope.  Your work does a great job of illustrating the spectrum of correlations between those two.</p>",
      "rawMarkdown": "This is really interesting and at the same time, incredibly confusing!  😃\n\nFirst off, very nice work. The summary you presented above and your notebooks are exceptional.  Regarding some of your findings...\n\n> There are some \"Global inhibitors\" (resp. \"Global Enhancers\") that are strongly negatively (resp. positively) correlated with all targets. e.g. ENSG00000129824_RPS4Y1 is a global inhibitor and ENSG00000229807_XIST is a global enhancer. (Not sure \"enhancer\" is the correct term in biology; please inform me)\n\nRegarding \"enhancers\", as @alexandervc mentioned, this is indeed a term in cellular biology.  Typically it refers to a region of DNA that is associated with increasing (or enhancing) transcription of a gene.  In your case, I would call what you've found \"activators\", but I'm sure there is some other term that might be more appropriate.\n\nFor your Global Inhibitor and Global Enhancer/Activator, these are puzzling.  \n\nThe inhibitor, [RPS4Y1](https://www.genecards.org/cgi-bin/carddisp.pl?gene=RPS4Y1) is a protein associated with a subunit of ribosomes.  Briefly, ribosomes are protein complexes that facilitate decoding mRNA in to proteins, a process know and translation.  It seems odd to me that it would be negatively related to protein levels as it is directly involved in protein production from mRNA.  Interestingly, RSP4Y1 is coded on the Y  chromosome, so only samples from males should show any levels.  There was another notebook that suggested there were 3 male and 1 female donor.\n\nFor the enhancer/activator, [XIST](https://www.genecards.org/cgi-bin/carddisp.pl?gene=XIST), this is not a protein coding gene but rather the RNA produced from this gene is involved in X-chromosome inactivation.  Females have 2 X chromosomes but if both are active, it is toxic to the cell.  One X chromosome is rendered inactive by genes from the XIC region, and this is one of those genes.  Males have one X chromosome and one Y chromosome and so should not have any levels of XIST as their X chromosome is not inactivated.\n\nI'm not doubting your findings, but in both cases, the direction of correlation is opposite what I might expect.  Biology never ceases to offer up surprises, though.  Is there a possibility that the X- and Y-chromosome linkages and the mixture of male and female donors a possible explanation?\n\n> Contrary to what was suggested in this discussion, genes and proteins related by names do not always have a strong correlation. \n\nThis doesn't surprise me.  I mentioned in another discussion [here](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350863#1933699) that correlation between mRNA levels and protein levels are not as consistent as we might hope.  Your work does a great job of illustrating the spectrum of correlations between those two.",
      "votes": null
    },
    {
      "id": "1935287",
      "postDate": "09/12/2022 02:21:17",
      "content": "<p>Thank you very much for your comment. It is very interesting to me to have this kind of feedback.</p>\n<blockquote>\n  <p>I'm not doubting your findings, but in both cases, the direction of correlation is opposite what I might expect.</p>\n</blockquote>\n<p>Actually, given that both you and <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> seem to be surprised by my CITEseq results, I am starting to be afraid there is a bug in my computation (I do not consider this possibility likely, but well, stupid mistakes do happen). Whithin the next few days I will probably re-compute the correlations with a different method for CITEseq just to see if the results are consistent.</p>\n<p>Otherwise I am thinking the paradoxes could also come from the transformation/normalization on the data done by the organizers. I still did not look into the precise transforms they applied, but they could potentially introduce some spurious correlations. </p>",
      "rawMarkdown": "Thank you very much for your comment. It is very interesting to me to have this kind of feedback.\n\n> I'm not doubting your findings, but in both cases, the direction of correlation is opposite what I might expect.\n\nActually, given that both you and @alexandervc seem to be surprised by my CITEseq results, I am starting to be afraid there is a bug in my computation (I do not consider this possibility likely, but well, stupid mistakes do happen). Whithin the next few days I will probably re-compute the correlations with a different method for CITEseq just to see if the results are consistent.\n\nOtherwise I am thinking the paradoxes could also come from the transformation/normalization on the data done by the organizers. I still did not look into the precise transforms they applied, but they could potentially introduce some spurious correlations.",
      "votes": null
    },
    {
      "id": "1936183",
      "postDate": "09/12/2022 14:53:12",
      "content": "<p>Molecules that inhibit the translation of mRNA into proteins are called translational inhibitors. These so-called post-transcriptional regulatory molecules (can be RNA or protein) can inhibit ribosome formation along an mRNA or prevent the ribosome from progressing along an mRNA. There are even small molecules like cycloheximide that can act as translation inhibitors.</p>",
      "rawMarkdown": "Molecules that inhibit the translation of mRNA into proteins are called translational inhibitors. These so-called post-transcriptional regulatory molecules (can be RNA or protein) can inhibit ribosome formation along an mRNA or prevent the ribosome from progressing along an mRNA. There are even small molecules like cycloheximide that can act as translation inhibitors.",
      "votes": null
    },
    {
      "id": "1936629",
      "postDate": "09/13/2022 00:08:42",
      "content": "<p>Thank you for sharing.</p>\n<p>I found that input/output distribution is divided into 2 groups for CITESeq cases, and 3 groups for multiome cases. (Naturally distributed group, vertically aligned group, horizontally aligned group) I hope these weird distributions do not have a bad influence on correlation coefficients.</p>\n<p>CITESeq example (first protein, and first gene expression)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11130257%2Fa79d27377f09f34e8ae9f6eb0ff78e48%2Fcite.png?generation=1663026846092261&amp;alt=media\" alt=\"\"></p>\n<p>Multiome examples (randomly selected 2x2 )<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11130257%2F2901f306542dd7b4ddadb355ba1cf6c5%2Fmulti.png?generation=1663026923741978&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thank you for sharing.\n\nI found that input/output distribution is divided into 2 groups for CITESeq cases, and 3 groups for multiome cases. (Naturally distributed group, vertically aligned group, horizontally aligned group) I hope these weird distributions do not have a bad influence on correlation coefficients.\n\nCITESeq example (first protein, and first gene expression)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11130257%2Fa79d27377f09f34e8ae9f6eb0ff78e48%2Fcite.png?generation=1663026846092261&alt=media)\n\nMultiome examples (randomly selected 2x2 )\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11130257%2F2901f306542dd7b4ddadb355ba1cf6c5%2Fmulti.png?generation=1663026923741978&alt=media)",
      "votes": null
    },
    {
      "id": "1936644",
      "postDate": "09/13/2022 00:26:15",
      "content": "<p>I'm wondering if this is what is driving the global inhibitor and global activator effects seen by <a href=\"https://www.kaggle.com/fabiencrom\" target=\"_blank\">@fabiencrom</a>.  My guess is that the horizontal lines you're seeing on gene expression are related to sex-linked genes - the female donor will have no expression from the Y-chromosome linked genes.  Males will still have X-linked genes expressed, but not from the X-inactivation center (XIC), as I described above.  Of course, looking up the two genes you illustrate above, one is on Chromsome 3 and one is on Chromsome 6 - neither would have a sex-linked effect.</p>\n<p>The horizontal and vertical sets of points, however, could certainly have an impact on correlations.</p>",
      "rawMarkdown": "I'm wondering if this is what is driving the global inhibitor and global activator effects seen by @fabiencrom.  My guess is that the horizontal lines you're seeing on gene expression are related to sex-linked genes - the female donor will have no expression from the Y-chromosome linked genes.  Males will still have X-linked genes expressed, but not from the X-inactivation center (XIC), as I described above.  Of course, looking up the two genes you illustrate above, one is on Chromsome 3 and one is on Chromsome 6 - neither would have a sex-linked effect.\n\nThe horizontal and vertical sets of points, however, could certainly have an impact on correlations.",
      "votes": null
    },
    {
      "id": "1936673",
      "postDate": "09/13/2022 01:16:44",
      "content": "<p>Thank you for your comment. <br>\nAlthough I am not a biology researcher, my guess is just data is missing.<br>\nFor CITESeq case there are 22050 inputs (gene expressions), and for Multiome cases, there are 228942 inputs and 23418 targets. I guess the technology can measure a small part of them. If the gene expression is not caught for the cell, the value is filled with 0 (my guess).  So there are gaps between naturally distributed data and vertically or horizontally aligned data. </p>",
      "rawMarkdown": "Thank you for your comment. \nAlthough I am not a biology researcher, my guess is just data is missing.\nFor CITESeq case there are 22050 inputs (gene expressions), and for Multiome cases, there are 228942 inputs and 23418 targets. I guess the technology can measure a small part of them. If the gene expression is not caught for the cell, the value is filled with 0 (my guess).  So there are gaps between naturally distributed data and vertically or horizontally aligned data.",
      "votes": null
    },
    {
      "id": "1936762",
      "postDate": "09/13/2022 03:48:36",
      "content": "<p>Yes, you are probably right.  For some genes, I think the sex-linked differences will also be present.  See <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348761\" target=\"_blank\">this notebook</a> for example, </p>",
      "rawMarkdown": "Yes, you are probably right.  For some genes, I think the sex-linked differences will also be present.  See [this notebook](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348761) for example,",
      "votes": null
    },
    {
      "id": "1937255",
      "postDate": "09/13/2022 11:26:46",
      "content": "<p>FYI, I did recomputed the correlations in a different way (this time using pandas <code>corr</code> function and the original h5 files). There are some slight differences that can probably be attributed to different method and precision of calculation, but they are not significant. So the correlations numbers are correct. But I am probably quite wrong in how I interpret them. The comment from <a href=\"https://www.kaggle.com/konomuabe\" target=\"_blank\">@konomuabe</a> above explain at least part of the problem: the joint distribution of target and inputs values appear to not be unimodal due to a kind of different behaviour when input is zero.</p>",
      "rawMarkdown": "FYI, I did recomputed the correlations in a different way (this time using pandas `corr` function and the original h5 files). There are some slight differences that can probably be attributed to different method and precision of calculation, but they are not significant. So the correlations numbers are correct. But I am probably quite wrong in how I interpret them. The comment from @konomuabe above explain at least part of the problem: the joint distribution of target and inputs values appear to not be unimodal due to a kind of different behaviour when input is zero.",
      "votes": null
    },
    {
      "id": "1937290",
      "postDate": "09/13/2022 11:52:05",
      "content": "<p>Oh, good idea to actually check the joint distribution with scatter plots!</p>\n<p>So actually the joint distributions are a bit strange. They are not really unimodal and very different from a multivariate normal distribution.</p>\n<p>I agrree with you, <a href=\"https://www.kaggle.com/konomuabe\" target=\"_blank\">@konomuabe</a>, it seems that zero values for inputs should be considered a \"missing value\" and not an actual zero. </p>\n<p>I checked a bit, and I can confirm this seem to be the source of e.g. why I thought <code>ENSG00000229807_XIST</code> is an enhancer/activator. For example, taking the pair <code>ENSG00000229807_XIST</code>/ <code>CD155</code>:</p>\n<p><code>ENSG00000229807_XIST</code> and <code>CD155</code> seem to have a rather strong positive correlation of 0.43175.<br>\nBut, if we only consider cells for which <code>ENSG00000229807_XIST &gt; 0</code>, the correlation becomes slightly negative: -0.08724</p>\n<p>On the other hand, if we check the mean value of <code>CD155</code> for when <code>ENSG00000229807_XIST==0</code> and when <code>ENSG00000229807_XIST&gt;0</code>, we find:</p>\n<p><code>ENSG00000229807_XIST==0</code>  --&gt; mean value of <code>CD155</code> is 4.881356<br>\n<code>ENSG00000229807_XIST&gt;0</code>  --&gt; mean value of <code>CD155</code> is 7.372275</p>\n<p>This explains why we get a positive correlation if we compute the correlation on the data  <code>ENSG00000229807_XIST==0</code>  + <code>ENSG00000229807_XIST&gt;0</code>.</p>\n<p>So the correct interpretation is not \"<code>ENSG00000229807_XIST</code> is an activator for <code>CD155</code>\", but rather:</p>\n<ol>\n<li>For the cells for which <code>ENSG00000229807_XIST</code> is measured, it has a negative correlation with <code>CD155</code></li>\n<li>For some reason, the cells for which <code>ENSG00000229807_XIST</code> was measured tend to have a larger value of <code>CD155</code></li>\n</ol>\n<p>Conclusion 1 becomes now more in line with <a href=\"https://www.kaggle.com/kirkdco\" target=\"_blank\">@kirkdco</a>'s intuition.<br>\nThere is probably a technical reason for 2. But not sure what it is at this point.</p>\n<p>So in practice, I should probably redo all the computations with special casing the cases where an input is equal to zero.</p>\n<p>This is rather important to know if one want to fit a simple Ridge regression model (DNNs and GBMs are probably clever enough to understand that data behave differently when input is zero).</p>",
      "rawMarkdown": "Oh, good idea to actually check the joint distribution with scatter plots!\n\nSo actually the joint distributions are a bit strange. They are not really unimodal and very different from a multivariate normal distribution.\n\nI agrree with you, @konomuabe, it seems that zero values for inputs should be considered a \"missing value\" and not an actual zero. \n\nI checked a bit, and I can confirm this seem to be the source of e.g. why I thought `ENSG00000229807_XIST` is an enhancer/activator. For example, taking the pair `ENSG00000229807_XIST`/ `CD155`:\n\n`ENSG00000229807_XIST` and `CD155` seem to have a rather strong positive correlation of 0.43175.\nBut, if we only consider cells for which `ENSG00000229807_XIST > 0`, the correlation becomes slightly negative: -0.08724\n\nOn the other hand, if we check the mean value of `CD155` for when `ENSG00000229807_XIST==0` and when `ENSG00000229807_XIST>0`, we find:\n\n `ENSG00000229807_XIST==0`  --> mean value of `CD155` is 4.881356\n`ENSG00000229807_XIST>0`  --> mean value of `CD155` is 7.372275\n\nThis explains why we get a positive correlation if we compute the correlation on the data  `ENSG00000229807_XIST==0`  + `ENSG00000229807_XIST>0`.\n\nSo the correct interpretation is not \"`ENSG00000229807_XIST` is an activator for `CD155`\", but rather:\n1. For the cells for which `ENSG00000229807_XIST` is measured, it has a negative correlation with `CD155`\n2. For some reason, the cells for which `ENSG00000229807_XIST` was measured tend to have a larger value of `CD155`\n\nConclusion 1 becomes now more in line with @kirkdco's intuition.\nThere is probably a technical reason for 2. But not sure what it is at this point.\n\nSo in practice, I should probably redo all the computations with special casing the cases where an input is equal to zero.\n\nThis is rather important to know if one want to fit a simple Ridge regression model (DNNs and GBMs are probably clever enough to understand that data behave differently when input is zero).",
      "votes": null
    },
    {
      "id": "1938264",
      "postDate": "09/14/2022 04:40:54",
      "content": "<p>Thanks.<br>\nI posted a related <a href=\"https://www.kaggle.com/konomuabe/eda-multiome-targets\" target=\"_blank\">notebook EDA:Multiome targets</a> about which gene expressions are chosen to non-zero(measured.)<br>\nOne more topic, with my another <a href=\"https://www.kaggle.com/code/konomuabe/average-of-gene-id-and-day\" target=\"_blank\">notebook Average of gene_id and day</a>, which makes prediction from just average, replacing average to the average of non-zero lowered the submission 0.13. This conflicts with my guess.</p>",
      "rawMarkdown": "Thanks.\nI posted a related [notebook EDA:Multiome targets](https://www.kaggle.com/konomuabe/eda-multiome-targets) about which gene expressions are chosen to non-zero(measured.)\nOne more topic, with my another [notebook Average of gene_id and day](https://www.kaggle.com/code/konomuabe/average-of-gene-id-and-day), which makes prediction from just average, replacing average to the average of non-zero lowered the submission 0.13. This conflicts with my guess.",
      "votes": null
    },
    {
      "id": "1938440",
      "postDate": "09/14/2022 06:45:33",
      "content": "<p>Let me also put here the comment as under the notebook:</p>\n<p>Very nice ! <br>\nJust remark - such plots are typical for all single cell RNA sequencing data.<br>\nThat is called \"dropout\" - technical zeros appearence.<br>\nThere are many methods to mitigate that problem.<br>\nSee very nice gif here:<br>\n<a href=\"https://github.com/KrishnaswamyLab/MAGIC\" target=\"_blank\">https://github.com/KrishnaswamyLab/MAGIC</a><br>\nMore generally see discussion:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856</a></p>\n<p>As an example, see:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/breastcancergse161529-cell-cycle-01\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/breastcancergse161529-cell-cycle-01</a><br>\nyou can also see \"technical zeros\" on the plots</p>",
      "rawMarkdown": "Let me also put here the comment as under the notebook:\n\nVery nice ! \nJust remark - such plots are typical for all single cell RNA sequencing data.\nThat is called \"dropout\" - technical zeros appearence.\nThere are many methods to mitigate that problem.\nSee very nice gif here:\nhttps://github.com/KrishnaswamyLab/MAGIC\nMore generally see discussion:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856\n\nAs an example, see:\nhttps://www.kaggle.com/code/alexandervc/breastcancergse161529-cell-cycle-01\nyou can also see \"technical zeros\" on the plots",
      "votes": null
    },
    {
      "id": "1939117",
      "postDate": "09/14/2022 14:36:22",
      "content": "<p>Thank you very much. <br>\nI am not a biology researcher, so advice from a researcher like you helps.  </p>",
      "rawMarkdown": "Thank you very much. \nI am not a biology researcher, so advice from a researcher like you helps.",
      "votes": null
    },
    {
      "id": "1939141",
      "postDate": "09/14/2022 14:50:35",
      "content": "<p>Good job <a href=\"https://www.kaggle.com/fabiencrom\" target=\"_blank\">@fabiencrom</a>, you focus on the correlation between features and targets, but even if the metric of this competition is Correlation, the method of selecting the features depends on your type of modeling, for example, if you choose Linear models, the best features which have a high correlation are better for your models, but that's not the same when we use No-Linear models such as Gradient decision tress models.</p>",
      "rawMarkdown": "Good job @fabiencrom, you focus on the correlation between features and targets, but even if the metric of this competition is Correlation, the method of selecting the features depends on your type of modeling, for example, if you choose Linear models, the best features which have a high correlation are better for your models, but that's not the same when we use No-Linear models such as Gradient decision tress models.",
      "votes": null
    },
    {
      "id": "1939159",
      "postDate": "09/14/2022 15:00:43",
      "content": "<p>There is something hidden inside the Correlation metric, is the stationarity correlation, we can test stationarity correlation when we choose many sample sizes from the data train and calculate the Correlation for each size, you will find a nonstationarity correlation between the features and targets, that's why Linear models and No-linear models have not the same way of selecting the features.</p>",
      "rawMarkdown": "There is something hidden inside the Correlation metric, is the stationarity correlation, we can test stationarity correlation when we choose many sample sizes from the data train and calculate the Correlation for each size, you will find a nonstationarity correlation between the features and targets, that's why Linear models and No-linear models have not the same way of selecting the features.",
      "votes": null
    },
    {
      "id": "1979596",
      "postDate": "10/09/2022 15:10:28",
      "content": "<p>Great work ! <br>\nHere is some analysis of biologically motivated correlations for CITEseq<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-bio-motivated-feature-selection-citeseq\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-bio-motivated-feature-selection-citeseq</a><br>\nWould be nice to compare </p>",
      "rawMarkdown": "Great work ! \nHere is some analysis of biologically motivated correlations for CITEseq\nhttps://www.kaggle.com/code/alexandervc/mmscel-bio-motivated-feature-selection-citeseq\nWould be nice to compare",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1934959,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/11/2022 17:31:34",
      "content": "<p>Nice ! Thanks for sharing ! <br>\n<a href=\"https://en.wikipedia.org/wiki/Enhancer_(genetics\" target=\"_blank\">https://en.wikipedia.org/wiki/Enhancer_(genetics</a>)<br>\nIt is already in use, with the other meaning.  (It is related to Multiome task)</p>\n<p>Remark: \"Contrary to what was suggested… \" - that is indeed surprising (at least for me). I'll try to ask around. </p>\n<p>ENSG00000229807_XIST  - that is \"XIST\" - it is quite famous <a href=\"https://en.wikipedia.org/wiki/XIST\" target=\"_blank\">https://en.wikipedia.org/wiki/XIST</a><br>\nstrange to see such correlation.   I'll try to ask around. <br>\nRPS4Y1 - ribosomal protein - again, I'll try to ask. </p>\n<p>\"On average, a target has a correlation &gt; 0.3 with about 9 inputs. \"<br>\nIs it possible to make a table \"Target\", TopCorrelated1, TopCorrelated2,TopCorrelated3 ………<br>\nThan we might either use GSEA, or look by eye if something appears </p>\n<p>PS<br>\nCan you please look: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900</a><br>\nIt is interesting to understand and probably biologically interpret the feature importance . <br>\nPSPS<br>\nHere is also some analysis of correlations:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo?scriptVersionId=103869738&amp;cellId=16\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo?scriptVersionId=103869738&amp;cellId=16</a></p>\n<p>PSPSPS<br>\nJoin our discussions : <a href=\"https://t.me/sberlogacompete\" target=\"_blank\">https://t.me/sberlogacompete</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1934993,
          "author_name": "fabiencrom",
          "author_url": "",
          "post_date": "09/11/2022 18:08:22",
          "content": "<blockquote>\n  <p>It is already in use, with the other meaning. (It is related to Multiome task)</p>\n</blockquote>\n<p>You mean my use of it is incorrect? What would be the proper term for a gene that prevent expression of a protein?</p>\n<blockquote>\n  <p>RPS4Y1 - ribosomal protein - again, I'll try to ask.</p>\n</blockquote>\n<p>Thank you :-)</p>\n<blockquote>\n  <p>Is it possible to make a table \"Target\", TopCorrelated1, TopCorrelated2,TopCorrelated3 ………</p>\n</blockquote>\n<p>At the bottom of <a href=\"https://www.kaggle.com/fabiencrom/msci-correlations-eda-citeseq\" target=\"_blank\">this notebook</a> (same as linked above), you have a long output that, for each target protein gives the top 5 most correlated genes (easy to change to top 10 or more in the code if you want) + \"associated genes\". Isn't that what you are suggesting? (or I am misunderstanding)</p>\n<blockquote>\n  <p>Here is also some analysis of correlations:</p>\n</blockquote>\n<p>Oh thank you, I had not checked this before. But the correlations you are displaying are between targets and not between inputs and targets, correct? (I guess I will understand if I take the time to read the code, but just to know)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1935004,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "09/11/2022 18:26:28",
          "content": "<p>-- You mean my use of it is incorrect?<br>\nSorry, yes. <br>\n-- What would be the proper term for a gene that prevent expression of a protein?<br>\nNot sure I know that. </p>\n<p>-- At the bottom of this notebook (same as linked above) ……<br>\nNice ! First look is the following :<br>\n1)  quite often GATA1 appears - \"GATA-binding factor 1 or GATA-1 (also termed Erythroid transcription factor) \"  <a href=\"https://en.wikipedia.org/wiki/GATA1\" target=\"_blank\">https://en.wikipedia.org/wiki/GATA1</a> <br>\nThat is gene strongly transcribed for Ery cells. <br>\nThere are might be some CD marking Ery - than it would be and an explantation.<br>\n2)We also see lots of XIST, RPS - I would suggest to exclude them. <br>\n3) There are CD** genes - that probably is natural, it might be related to my analysis of targets below:</p>\n<p>-- But the correlations you are displaying are between targets <br>\nYes. <br>\nTop correlated groups:<br>\n['CD71' 'CD115' 'CD88']<br>\n0.6 correlation_threshold <br>\n['CD155' 'CD112' 'CD47' 'HLA-A-B-C' 'CD45RA' 'CD31' 'CD11a' 'CD13' 'CD29'<br>\n 'CD81' 'CD18' 'CD45' 'CD49d' 'CD162']<br>\nI will try to think on bio interpretation some time later.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1935018,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "09/11/2022 18:34:46",
          "content": "<p>It is possible to calculate correlations with the cell types:<br>\nsomething like ordering cell types as follows:</p>\n<p>1 HSC <br>\n2 NeuP, MoP<br>\n3 MastP<br>\n4 Eryp , MkP, BP</p>\n<p>e.g. assigning cell types these numbers .</p>\n<p>Or <br>\nbetter taking the first component UMAP1 of the following UMAP:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo?scriptVersionId=103869738&amp;cellId=22\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo?scriptVersionId=103869738&amp;cellId=22</a></p>\n<p>That would give what is related to the cell type,<br>\nand that would probably help the interpretations. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1936183,
          "author_name": "danielburkhardt",
          "author_url": "",
          "post_date": "09/12/2022 14:53:12",
          "content": "<p>Molecules that inhibit the translation of mRNA into proteins are called translational inhibitors. These so-called post-transcriptional regulatory molecules (can be RNA or protein) can inhibit ribosome formation along an mRNA or prevent the ribosome from progressing along an mRNA. There are even small molecules like cycloheximide that can act as translation inhibitors.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1935253,
      "author_name": "kirkdco",
      "author_url": "",
      "post_date": "09/12/2022 01:20:30",
      "content": "<p>This is really interesting and at the same time, incredibly confusing!  😃</p>\n<p>First off, very nice work. The summary you presented above and your notebooks are exceptional.  Regarding some of your findings…</p>\n<blockquote>\n  <p>There are some \"Global inhibitors\" (resp. \"Global Enhancers\") that are strongly negatively (resp. positively) correlated with all targets. e.g. ENSG00000129824_RPS4Y1 is a global inhibitor and ENSG00000229807_XIST is a global enhancer. (Not sure \"enhancer\" is the correct term in biology; please inform me)</p>\n</blockquote>\n<p>Regarding \"enhancers\", as <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> mentioned, this is indeed a term in cellular biology.  Typically it refers to a region of DNA that is associated with increasing (or enhancing) transcription of a gene.  In your case, I would call what you've found \"activators\", but I'm sure there is some other term that might be more appropriate.</p>\n<p>For your Global Inhibitor and Global Enhancer/Activator, these are puzzling.  </p>\n<p>The inhibitor, <a href=\"https://www.genecards.org/cgi-bin/carddisp.pl?gene=RPS4Y1\" target=\"_blank\">RPS4Y1</a> is a protein associated with a subunit of ribosomes.  Briefly, ribosomes are protein complexes that facilitate decoding mRNA in to proteins, a process know and translation.  It seems odd to me that it would be negatively related to protein levels as it is directly involved in protein production from mRNA.  Interestingly, RSP4Y1 is coded on the Y  chromosome, so only samples from males should show any levels.  There was another notebook that suggested there were 3 male and 1 female donor.</p>\n<p>For the enhancer/activator, <a href=\"https://www.genecards.org/cgi-bin/carddisp.pl?gene=XIST\" target=\"_blank\">XIST</a>, this is not a protein coding gene but rather the RNA produced from this gene is involved in X-chromosome inactivation.  Females have 2 X chromosomes but if both are active, it is toxic to the cell.  One X chromosome is rendered inactive by genes from the XIC region, and this is one of those genes.  Males have one X chromosome and one Y chromosome and so should not have any levels of XIST as their X chromosome is not inactivated.</p>\n<p>I'm not doubting your findings, but in both cases, the direction of correlation is opposite what I might expect.  Biology never ceases to offer up surprises, though.  Is there a possibility that the X- and Y-chromosome linkages and the mixture of male and female donors a possible explanation?</p>\n<blockquote>\n  <p>Contrary to what was suggested in this discussion, genes and proteins related by names do not always have a strong correlation. </p>\n</blockquote>\n<p>This doesn't surprise me.  I mentioned in another discussion <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350863#1933699\" target=\"_blank\">here</a> that correlation between mRNA levels and protein levels are not as consistent as we might hope.  Your work does a great job of illustrating the spectrum of correlations between those two.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1935287,
          "author_name": "fabiencrom",
          "author_url": "",
          "post_date": "09/12/2022 02:21:17",
          "content": "<p>Thank you very much for your comment. It is very interesting to me to have this kind of feedback.</p>\n<blockquote>\n  <p>I'm not doubting your findings, but in both cases, the direction of correlation is opposite what I might expect.</p>\n</blockquote>\n<p>Actually, given that both you and <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> seem to be surprised by my CITEseq results, I am starting to be afraid there is a bug in my computation (I do not consider this possibility likely, but well, stupid mistakes do happen). Whithin the next few days I will probably re-compute the correlations with a different method for CITEseq just to see if the results are consistent.</p>\n<p>Otherwise I am thinking the paradoxes could also come from the transformation/normalization on the data done by the organizers. I still did not look into the precise transforms they applied, but they could potentially introduce some spurious correlations. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1937255,
          "author_name": "fabiencrom",
          "author_url": "",
          "post_date": "09/13/2022 11:26:46",
          "content": "<p>FYI, I did recomputed the correlations in a different way (this time using pandas <code>corr</code> function and the original h5 files). There are some slight differences that can probably be attributed to different method and precision of calculation, but they are not significant. So the correlations numbers are correct. But I am probably quite wrong in how I interpret them. The comment from <a href=\"https://www.kaggle.com/konomuabe\" target=\"_blank\">@konomuabe</a> above explain at least part of the problem: the joint distribution of target and inputs values appear to not be unimodal due to a kind of different behaviour when input is zero.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1936629,
      "author_name": "konomuabe",
      "author_url": "",
      "post_date": "09/13/2022 00:08:42",
      "content": "<p>Thank you for sharing.</p>\n<p>I found that input/output distribution is divided into 2 groups for CITESeq cases, and 3 groups for multiome cases. (Naturally distributed group, vertically aligned group, horizontally aligned group) I hope these weird distributions do not have a bad influence on correlation coefficients.</p>\n<p>CITESeq example (first protein, and first gene expression)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11130257%2Fa79d27377f09f34e8ae9f6eb0ff78e48%2Fcite.png?generation=1663026846092261&amp;alt=media\" alt=\"\"></p>\n<p>Multiome examples (randomly selected 2x2 )<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11130257%2F2901f306542dd7b4ddadb355ba1cf6c5%2Fmulti.png?generation=1663026923741978&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1936644,
          "author_name": "kirkdco",
          "author_url": "",
          "post_date": "09/13/2022 00:26:15",
          "content": "<p>I'm wondering if this is what is driving the global inhibitor and global activator effects seen by <a href=\"https://www.kaggle.com/fabiencrom\" target=\"_blank\">@fabiencrom</a>.  My guess is that the horizontal lines you're seeing on gene expression are related to sex-linked genes - the female donor will have no expression from the Y-chromosome linked genes.  Males will still have X-linked genes expressed, but not from the X-inactivation center (XIC), as I described above.  Of course, looking up the two genes you illustrate above, one is on Chromsome 3 and one is on Chromsome 6 - neither would have a sex-linked effect.</p>\n<p>The horizontal and vertical sets of points, however, could certainly have an impact on correlations.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1936673,
          "author_name": "konomuabe",
          "author_url": "",
          "post_date": "09/13/2022 01:16:44",
          "content": "<p>Thank you for your comment. <br>\nAlthough I am not a biology researcher, my guess is just data is missing.<br>\nFor CITESeq case there are 22050 inputs (gene expressions), and for Multiome cases, there are 228942 inputs and 23418 targets. I guess the technology can measure a small part of them. If the gene expression is not caught for the cell, the value is filled with 0 (my guess).  So there are gaps between naturally distributed data and vertically or horizontally aligned data. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1936762,
          "author_name": "kirkdco",
          "author_url": "",
          "post_date": "09/13/2022 03:48:36",
          "content": "<p>Yes, you are probably right.  For some genes, I think the sex-linked differences will also be present.  See <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348761\" target=\"_blank\">this notebook</a> for example, </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1937290,
          "author_name": "fabiencrom",
          "author_url": "",
          "post_date": "09/13/2022 11:52:05",
          "content": "<p>Oh, good idea to actually check the joint distribution with scatter plots!</p>\n<p>So actually the joint distributions are a bit strange. They are not really unimodal and very different from a multivariate normal distribution.</p>\n<p>I agrree with you, <a href=\"https://www.kaggle.com/konomuabe\" target=\"_blank\">@konomuabe</a>, it seems that zero values for inputs should be considered a \"missing value\" and not an actual zero. </p>\n<p>I checked a bit, and I can confirm this seem to be the source of e.g. why I thought <code>ENSG00000229807_XIST</code> is an enhancer/activator. For example, taking the pair <code>ENSG00000229807_XIST</code>/ <code>CD155</code>:</p>\n<p><code>ENSG00000229807_XIST</code> and <code>CD155</code> seem to have a rather strong positive correlation of 0.43175.<br>\nBut, if we only consider cells for which <code>ENSG00000229807_XIST &gt; 0</code>, the correlation becomes slightly negative: -0.08724</p>\n<p>On the other hand, if we check the mean value of <code>CD155</code> for when <code>ENSG00000229807_XIST==0</code> and when <code>ENSG00000229807_XIST&gt;0</code>, we find:</p>\n<p><code>ENSG00000229807_XIST==0</code>  --&gt; mean value of <code>CD155</code> is 4.881356<br>\n<code>ENSG00000229807_XIST&gt;0</code>  --&gt; mean value of <code>CD155</code> is 7.372275</p>\n<p>This explains why we get a positive correlation if we compute the correlation on the data  <code>ENSG00000229807_XIST==0</code>  + <code>ENSG00000229807_XIST&gt;0</code>.</p>\n<p>So the correct interpretation is not \"<code>ENSG00000229807_XIST</code> is an activator for <code>CD155</code>\", but rather:</p>\n<ol>\n<li>For the cells for which <code>ENSG00000229807_XIST</code> is measured, it has a negative correlation with <code>CD155</code></li>\n<li>For some reason, the cells for which <code>ENSG00000229807_XIST</code> was measured tend to have a larger value of <code>CD155</code></li>\n</ol>\n<p>Conclusion 1 becomes now more in line with <a href=\"https://www.kaggle.com/kirkdco\" target=\"_blank\">@kirkdco</a>'s intuition.<br>\nThere is probably a technical reason for 2. But not sure what it is at this point.</p>\n<p>So in practice, I should probably redo all the computations with special casing the cases where an input is equal to zero.</p>\n<p>This is rather important to know if one want to fit a simple Ridge regression model (DNNs and GBMs are probably clever enough to understand that data behave differently when input is zero).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1938264,
          "author_name": "konomuabe",
          "author_url": "",
          "post_date": "09/14/2022 04:40:54",
          "content": "<p>Thanks.<br>\nI posted a related <a href=\"https://www.kaggle.com/konomuabe/eda-multiome-targets\" target=\"_blank\">notebook EDA:Multiome targets</a> about which gene expressions are chosen to non-zero(measured.)<br>\nOne more topic, with my another <a href=\"https://www.kaggle.com/code/konomuabe/average-of-gene-id-and-day\" target=\"_blank\">notebook Average of gene_id and day</a>, which makes prediction from just average, replacing average to the average of non-zero lowered the submission 0.13. This conflicts with my guess.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1938440,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "09/14/2022 06:45:33",
          "content": "<p>Let me also put here the comment as under the notebook:</p>\n<p>Very nice ! <br>\nJust remark - such plots are typical for all single cell RNA sequencing data.<br>\nThat is called \"dropout\" - technical zeros appearence.<br>\nThere are many methods to mitigate that problem.<br>\nSee very nice gif here:<br>\n<a href=\"https://github.com/KrishnaswamyLab/MAGIC\" target=\"_blank\">https://github.com/KrishnaswamyLab/MAGIC</a><br>\nMore generally see discussion:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856</a></p>\n<p>As an example, see:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/breastcancergse161529-cell-cycle-01\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/breastcancergse161529-cell-cycle-01</a><br>\nyou can also see \"technical zeros\" on the plots</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1939117,
          "author_name": "konomuabe",
          "author_url": "",
          "post_date": "09/14/2022 14:36:22",
          "content": "<p>Thank you very much. <br>\nI am not a biology researcher, so advice from a researcher like you helps.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1939141,
      "author_name": "youneseloiarm",
      "author_url": "",
      "post_date": "09/14/2022 14:50:35",
      "content": "<p>Good job <a href=\"https://www.kaggle.com/fabiencrom\" target=\"_blank\">@fabiencrom</a>, you focus on the correlation between features and targets, but even if the metric of this competition is Correlation, the method of selecting the features depends on your type of modeling, for example, if you choose Linear models, the best features which have a high correlation are better for your models, but that's not the same when we use No-Linear models such as Gradient decision tress models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1939159,
          "author_name": "youneseloiarm",
          "author_url": "",
          "post_date": "09/14/2022 15:00:43",
          "content": "<p>There is something hidden inside the Correlation metric, is the stationarity correlation, we can test stationarity correlation when we choose many sample sizes from the data train and calculate the Correlation for each size, you will find a nonstationarity correlation between the features and targets, that's why Linear models and No-linear models have not the same way of selecting the features.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1979596,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "10/09/2022 15:10:28",
      "content": "<p>Great work ! <br>\nHere is some analysis of biologically motivated correlations for CITEseq<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/mmscel-bio-motivated-feature-selection-citeseq\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/mmscel-bio-motivated-feature-selection-citeseq</a><br>\nWould be nice to compare </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1934793": "Hi,\n\nSo I finally decided to make some good use of the access to 128GB RAM machines from Saturn Cloud that we have been given. (see [this discussion](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999) )\n\nI decided to compute the correlation coefficients between all inputs and targets, and to try and analyze them a bit.\n\nThe pre-computed correlations are in this dataset:\nhttps://www.kaggle.com/datasets/fabiencrom/msci-correlations\n\nHere is the notebook I used to generate it:\nhttps://www.kaggle.com/fabiencrom/msci-generating-all-correlations-inputs-targets\n\nNote that you can use this notebook to compute the CITEseq correlations on Kaggle; but for the Multiome data you would need around 40GB RAM (hence the usefulness of the Saturn Cloud machine...). Also note that I removed small correlations from the Multiome correlations to be able to store it as a sparse matrix. This way, the data uses less than 1GB instead of more than 10GB (there are >5billion correlations for the Multiome case!)\n\nI made an initial analysis of these correlations in two notebooks:\nhttps://www.kaggle.com/fabiencrom/msci-correlations-eda-multiome\nhttps://www.kaggle.com/fabiencrom/msci-correlations-eda-citeseq\n\nMy main findings so far:\n\nEDIT: Following discussion with other kagglers below (thanks to them), I realized my interpretations of the correlation numbers are wrong. The main point is that **it seems that when an input has a value equal to zero, we should consider it as a missing value and not a real zero**. As I did not know about that, I computed the correlations on all data for all inputs. And thus the correlation obtained are a mix of the real inputs/targets correlations and some spurious correlations due to the choices made of measuring the input or not for a given cell. I should probably redo this study with this information in mind.\n\nCITEseq:\n- Some inputs have a strong correlation with all targets. They should probably not be discarded if doing dimension reduction.\n- ~~There are some \"Global inhibitors\" (resp. \"Global Enhancers\") that are strongly negatively (resp. positively) correlated with all targets. e.g. `ENSG00000129824_RPS4Y1` is a global inhibitor and `ENSG00000229807_XIST` is a global enhancer. (Not sure \"enhancer\" is the correct term in biology; please inform me)~~\n- ~~Contrary to what was suggested in [this discussion](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349242), genes and proteins related by names do not always have a strong correlation. For example, `ENSG00000114013_CD86` do have a correlation of 0.235 with protein `CD86`. But `ENSG00000120217_CD274` only has a correlation of 0.026 with `CD274` and rank only as the 798th most correlated input with `CD274`~~\n- On average, a target has a correlation > 0.3 with about 9 inputs.\n \nMultiome:\n- correlations are much weaker than with CITEseq\n- 560 targets in the training set are always zero! (So i guess they should be set to zero in submissions)\n- 10 inputs in the training set are always zero (I think I saw this fact mentioned here before, but could not find a link)\n- many hints that there are both subgroups of targets and subgroups of inputs varying together.\n- On average, a target has a correlation > 0.1 with about 3 inputs\n\nApart from these findings, I expect that models could use the correlation informations for regularization (e.g. enforcing some sparsity in the matrix of a linear regression by setting to zero coefficients associated with input/target pairs with low correlations)\n\nIf you have in-domain knowledge, I would be very interested by your feedback on these results.",
    "1934959": "Nice ! Thanks for sharing ! \nhttps://en.wikipedia.org/wiki/Enhancer_(genetics)\nIt is already in use, with the other meaning.  (It is related to Multiome task)\n\nRemark: \"Contrary to what was suggested... \" - that is indeed surprising (at least for me). I'll try to ask around. \n\nENSG00000229807_XIST  - that is \"XIST\" - it is quite famous https://en.wikipedia.org/wiki/XIST\nstrange to see such correlation.   I'll try to ask around. \nRPS4Y1 - ribosomal protein - again, I'll try to ask. \n\n\"On average, a target has a correlation > 0.3 with about 9 inputs. \"\nIs it possible to make a table \"Target\", TopCorrelated1, TopCorrelated2,TopCorrelated3 .........\nThan we might either use GSEA, or look by eye if something appears \n\n\n\nPS\nCan you please look: https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900\nIt is interesting to understand and probably biologically interpret the feature importance . \nPSPS\nHere is also some analysis of correlations:\nhttps://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo?scriptVersionId=103869738&cellId=16\n\nPSPSPS\nJoin our discussions : https://t.me/sberlogacompete",
    "1934993": "> It is already in use, with the other meaning. (It is related to Multiome task)\n\nYou mean my use of it is incorrect? What would be the proper term for a gene that prevent expression of a protein?\n\n> RPS4Y1 - ribosomal protein - again, I'll try to ask.\n\nThank you :-)\n\n>Is it possible to make a table \"Target\", TopCorrelated1, TopCorrelated2,TopCorrelated3 ………\n\nAt the bottom of [this notebook](https://www.kaggle.com/fabiencrom/msci-correlations-eda-citeseq) (same as linked above), you have a long output that, for each target protein gives the top 5 most correlated genes (easy to change to top 10 or more in the code if you want) + \"associated genes\". Isn't that what you are suggesting? (or I am misunderstanding)\n\n>Here is also some analysis of correlations:\n\nOh thank you, I had not checked this before. But the correlations you are displaying are between targets and not between inputs and targets, correct? (I guess I will understand if I take the time to read the code, but just to know)",
    "1935004": "You mean my use of it is incorrect?\nSorry, yes. \n-- What would be the proper term for a gene that prevent expression of a protein?\nNot sure I know that. \n\n-- At the bottom of this notebook (same as linked above) ......\nNice ! First look is the following :\n1)  quite often GATA1 appears - \"GATA-binding factor 1 or GATA-1 (also termed Erythroid transcription factor) \"  https://en.wikipedia.org/wiki/GATA1 \nThat is gene strongly transcribed for Ery cells. \nThere are might be some CD marking Ery - than it would be and an explantation.\n2)We also see lots of XIST, RPS - I would suggest to exclude them. \n3) There are CD** genes - that probably is natural, it might be related to my analysis of targets below:\n\n-- But the correlations you are displaying are between targets \nYes. \nTop correlated groups:\n['CD71' 'CD115' 'CD88']\n0.6 correlation_threshold \n['CD155' 'CD112' 'CD47' 'HLA-A-B-C' 'CD45RA' 'CD31' 'CD11a' 'CD13' 'CD29'\n 'CD81' 'CD18' 'CD45' 'CD49d' 'CD162']\nI will try to think on bio interpretation some time later.",
    "1935018": "It is possible to calculate correlations with the cell types:\nsomething like ordering cell types as follows:\n\n1 HSC \n2 NeuP, MoP\n3 MastP\n4 Eryp , MkP, BP\n\ne.g. assigning cell types these numbers .\n\nOr \nbetter taking the first component UMAP1 of the following UMAP:\nhttps://www.kaggle.com/code/alexandervc/mmscel-eda-bioinfo?scriptVersionId=103869738&cellId=22\n\nThat would give what is related to the cell type,\nand that would probably help the interpretations.",
    "1935253": "This is really interesting and at the same time, incredibly confusing!  😃\n\nFirst off, very nice work. The summary you presented above and your notebooks are exceptional.  Regarding some of your findings...\n\n> There are some \"Global inhibitors\" (resp. \"Global Enhancers\") that are strongly negatively (resp. positively) correlated with all targets. e.g. ENSG00000129824_RPS4Y1 is a global inhibitor and ENSG00000229807_XIST is a global enhancer. (Not sure \"enhancer\" is the correct term in biology; please inform me)\n\nRegarding \"enhancers\", as @alexandervc mentioned, this is indeed a term in cellular biology.  Typically it refers to a region of DNA that is associated with increasing (or enhancing) transcription of a gene.  In your case, I would call what you've found \"activators\", but I'm sure there is some other term that might be more appropriate.\n\nFor your Global Inhibitor and Global Enhancer/Activator, these are puzzling.  \n\nThe inhibitor, [RPS4Y1](https://www.genecards.org/cgi-bin/carddisp.pl?gene=RPS4Y1) is a protein associated with a subunit of ribosomes.  Briefly, ribosomes are protein complexes that facilitate decoding mRNA in to proteins, a process know and translation.  It seems odd to me that it would be negatively related to protein levels as it is directly involved in protein production from mRNA.  Interestingly, RSP4Y1 is coded on the Y  chromosome, so only samples from males should show any levels.  There was another notebook that suggested there were 3 male and 1 female donor.\n\nFor the enhancer/activator, [XIST](https://www.genecards.org/cgi-bin/carddisp.pl?gene=XIST), this is not a protein coding gene but rather the RNA produced from this gene is involved in X-chromosome inactivation.  Females have 2 X chromosomes but if both are active, it is toxic to the cell.  One X chromosome is rendered inactive by genes from the XIC region, and this is one of those genes.  Males have one X chromosome and one Y chromosome and so should not have any levels of XIST as their X chromosome is not inactivated.\n\nI'm not doubting your findings, but in both cases, the direction of correlation is opposite what I might expect.  Biology never ceases to offer up surprises, though.  Is there a possibility that the X- and Y-chromosome linkages and the mixture of male and female donors a possible explanation?\n\n> Contrary to what was suggested in this discussion, genes and proteins related by names do not always have a strong correlation. \n\nThis doesn't surprise me.  I mentioned in another discussion [here](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350863#1933699) that correlation between mRNA levels and protein levels are not as consistent as we might hope.  Your work does a great job of illustrating the spectrum of correlations between those two.",
    "1935287": "Thank you very much for your comment. It is very interesting to me to have this kind of feedback.\n\n> I'm not doubting your findings, but in both cases, the direction of correlation is opposite what I might expect.\n\nActually, given that both you and @alexandervc seem to be surprised by my CITEseq results, I am starting to be afraid there is a bug in my computation (I do not consider this possibility likely, but well, stupid mistakes do happen). Whithin the next few days I will probably re-compute the correlations with a different method for CITEseq just to see if the results are consistent.\n\nOtherwise I am thinking the paradoxes could also come from the transformation/normalization on the data done by the organizers. I still did not look into the precise transforms they applied, but they could potentially introduce some spurious correlations.",
    "1936183": "Molecules that inhibit the translation of mRNA into proteins are called translational inhibitors. These so-called post-transcriptional regulatory molecules (can be RNA or protein) can inhibit ribosome formation along an mRNA or prevent the ribosome from progressing along an mRNA. There are even small molecules like cycloheximide that can act as translation inhibitors.",
    "1936629": "Thank you for sharing.\n\nI found that input/output distribution is divided into 2 groups for CITESeq cases, and 3 groups for multiome cases. (Naturally distributed group, vertically aligned group, horizontally aligned group) I hope these weird distributions do not have a bad influence on correlation coefficients.\n\nCITESeq example (first protein, and first gene expression)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11130257%2Fa79d27377f09f34e8ae9f6eb0ff78e48%2Fcite.png?generation=1663026846092261&alt=media)\n\nMultiome examples (randomly selected 2x2 )\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11130257%2F2901f306542dd7b4ddadb355ba1cf6c5%2Fmulti.png?generation=1663026923741978&alt=media)",
    "1936644": "I'm wondering if this is what is driving the global inhibitor and global activator effects seen by @fabiencrom.  My guess is that the horizontal lines you're seeing on gene expression are related to sex-linked genes - the female donor will have no expression from the Y-chromosome linked genes.  Males will still have X-linked genes expressed, but not from the X-inactivation center (XIC), as I described above.  Of course, looking up the two genes you illustrate above, one is on Chromsome 3 and one is on Chromsome 6 - neither would have a sex-linked effect.\n\nThe horizontal and vertical sets of points, however, could certainly have an impact on correlations.",
    "1936673": "Thank you for your comment. \nAlthough I am not a biology researcher, my guess is just data is missing.\nFor CITESeq case there are 22050 inputs (gene expressions), and for Multiome cases, there are 228942 inputs and 23418 targets. I guess the technology can measure a small part of them. If the gene expression is not caught for the cell, the value is filled with 0 (my guess).  So there are gaps between naturally distributed data and vertically or horizontally aligned data.",
    "1936762": "Yes, you are probably right.  For some genes, I think the sex-linked differences will also be present.  See [this notebook](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348761) for example,",
    "1937255": "FYI, I did recomputed the correlations in a different way (this time using pandas `corr` function and the original h5 files). There are some slight differences that can probably be attributed to different method and precision of calculation, but they are not significant. So the correlations numbers are correct. But I am probably quite wrong in how I interpret them. The comment from @konomuabe above explain at least part of the problem: the joint distribution of target and inputs values appear to not be unimodal due to a kind of different behaviour when input is zero.",
    "1937290": "Oh, good idea to actually check the joint distribution with scatter plots!\n\nSo actually the joint distributions are a bit strange. They are not really unimodal and very different from a multivariate normal distribution.\n\nI agrree with you, @konomuabe, it seems that zero values for inputs should be considered a \"missing value\" and not an actual zero. \n\nI checked a bit, and I can confirm this seem to be the source of e.g. why I thought `ENSG00000229807_XIST` is an enhancer/activator. For example, taking the pair `ENSG00000229807_XIST`/ `CD155`:\n\n`ENSG00000229807_XIST` and `CD155` seem to have a rather strong positive correlation of 0.43175.\nBut, if we only consider cells for which `ENSG00000229807_XIST > 0`, the correlation becomes slightly negative: -0.08724\n\nOn the other hand, if we check the mean value of `CD155` for when `ENSG00000229807_XIST==0` and when `ENSG00000229807_XIST>0`, we find:\n\n `ENSG00000229807_XIST==0`  --> mean value of `CD155` is 4.881356\n`ENSG00000229807_XIST>0`  --> mean value of `CD155` is 7.372275\n\nThis explains why we get a positive correlation if we compute the correlation on the data  `ENSG00000229807_XIST==0`  + `ENSG00000229807_XIST>0`.\n\nSo the correct interpretation is not \"`ENSG00000229807_XIST` is an activator for `CD155`\", but rather:\n1. For the cells for which `ENSG00000229807_XIST` is measured, it has a negative correlation with `CD155`\n2. For some reason, the cells for which `ENSG00000229807_XIST` was measured tend to have a larger value of `CD155`\n\nConclusion 1 becomes now more in line with @kirkdco's intuition.\nThere is probably a technical reason for 2. But not sure what it is at this point.\n\nSo in practice, I should probably redo all the computations with special casing the cases where an input is equal to zero.\n\nThis is rather important to know if one want to fit a simple Ridge regression model (DNNs and GBMs are probably clever enough to understand that data behave differently when input is zero).",
    "1938264": "Thanks.\nI posted a related [notebook EDA:Multiome targets](https://www.kaggle.com/konomuabe/eda-multiome-targets) about which gene expressions are chosen to non-zero(measured.)\nOne more topic, with my another [notebook Average of gene_id and day](https://www.kaggle.com/code/konomuabe/average-of-gene-id-and-day), which makes prediction from just average, replacing average to the average of non-zero lowered the submission 0.13. This conflicts with my guess.",
    "1938440": "Let me also put here the comment as under the notebook:\n\nVery nice ! \nJust remark - such plots are typical for all single cell RNA sequencing data.\nThat is called \"dropout\" - technical zeros appearence.\nThere are many methods to mitigate that problem.\nSee very nice gif here:\nhttps://github.com/KrishnaswamyLab/MAGIC\nMore generally see discussion:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/350856\n\nAs an example, see:\nhttps://www.kaggle.com/code/alexandervc/breastcancergse161529-cell-cycle-01\nyou can also see \"technical zeros\" on the plots",
    "1939117": "Thank you very much. \nI am not a biology researcher, so advice from a researcher like you helps.",
    "1939141": "Good job @fabiencrom, you focus on the correlation between features and targets, but even if the metric of this competition is Correlation, the method of selecting the features depends on your type of modeling, for example, if you choose Linear models, the best features which have a high correlation are better for your models, but that's not the same when we use No-Linear models such as Gradient decision tress models.",
    "1939159": "There is something hidden inside the Correlation metric, is the stationarity correlation, we can test stationarity correlation when we choose many sample sizes from the data train and calculate the Correlation for each size, you will find a nonstationarity correlation between the features and targets, that's why Linear models and No-linear models have not the same way of selecting the features.",
    "1979596": "Great work ! \nHere is some analysis of biologically motivated correlations for CITEseq\nhttps://www.kaggle.com/code/alexandervc/mmscel-bio-motivated-feature-selection-citeseq\nWould be nice to compare"
  },
  "source": "meta"
}