{
  "id": 349242,
  "title": "Exploiting the column names",
  "url": "/competitions/open-problems-multimodal/discussion/349242",
  "author_name": "",
  "post_date": "2022-08-31T17:56:15.103027200Z",
  "votes": 54,
  "comment_count": 9,
  "views": 0,
  "content": "<p>In this competition, the features and targets are not anonymous: Genes have names, and proteins have names. </p>\n<p>The CITEseq task has genes as input and proteins as output. Genes encode proteins, and it is more or less known which genes encode which proteins. The column names provide this information: The input dataframe has the genes as column names, and the target dataframe has the proteins as column names. According to the naming convention, the gene names contain the protein name as suffix after a '_'.</p>\n<p>If we match the input column names with the target column names, we find 151 genes which encode a target protein (see the table below). It doesn't matter that some proteins are encoded by more than one gene (e.g., rows 146 and 147 of the table). We may assume that these 151 features will have a high feature importance in our models.</p>\n<pre><code>matching_names = []\nfor protein in cite_protein_names:\n    matching_names += [(gene, protein) for gene in cite_gene_names if protein in gene]\npd.DataFrame(matching_names, columns=['Gene', 'Protein'])\n</code></pre>\n<table>\n  <thead>\n    <tr>\n      <th></th>\n      <th>Gene</th>\n      <th>Protein</th>\n    </tr>\n  </thead>\n  <tbody>\n    <tr>\n      <th>0</th>\n      <td>ENSG00000114013_CD86</td>\n      <td>CD86</td>\n    </tr>\n    <tr>\n      <th>1</th>\n      <td>ENSG00000120217_CD274</td>\n      <td>CD274</td>\n    </tr>\n    <tr>\n      <th>2</th>\n      <td>ENSG00000196776_CD47</td>\n      <td>CD47</td>\n    </tr>\n    <tr>\n      <th>3</th>\n      <td>ENSG00000117091_CD48</td>\n      <td>CD48</td>\n    </tr>\n    <tr>\n      <th>4</th>\n      <td>ENSG00000101017_CD40</td>\n      <td>CD40</td>\n    </tr>\n    <tr>\n      <th>...</th>\n      <td>...</td>\n      <td>...</td>\n    </tr>\n    <tr>\n      <th>146</th>\n      <td>ENSG00000102181_CD99L2</td>\n      <td>CD9</td>\n    </tr>\n    <tr>\n      <th>147</th>\n      <td>ENSG00000223773_CD99P1</td>\n      <td>CD9</td>\n    </tr>\n    <tr>\n      <th>148</th>\n      <td>ENSG00000204592_HLA-E</td>\n      <td>HLA-E</td>\n    </tr>\n    <tr>\n      <th>149</th>\n      <td>ENSG00000085117_CD82</td>\n      <td>CD82</td>\n    </tr>\n    <tr>\n      <th>150</th>\n      <td>ENSG00000134256_CD101</td>\n      <td>CD101</td>\n    </tr>\n  </tbody>\n</table>\n<p><strong>Insight:</strong> If we apply dimensionality reduction (PCA, SVD, whatever) to the 22050 features, we should make sure that we don't reduce away the 151 features which encode the proteins we want to predict.</p>\n<p>For an implementation of this idea, see the <a href=\"https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart\" target=\"_blank\">MSCI CITEseq Quickstart</a> notebook.</p>",
  "messages": [
    {
      "id": "1921317",
      "postDate": "08/31/2022 17:56:15",
      "content": "<p>In this competition, the features and targets are not anonymous: Genes have names, and proteins have names. </p>\n<p>The CITEseq task has genes as input and proteins as output. Genes encode proteins, and it is more or less known which genes encode which proteins. The column names provide this information: The input dataframe has the genes as column names, and the target dataframe has the proteins as column names. According to the naming convention, the gene names contain the protein name as suffix after a '_'.</p>\n<p>If we match the input column names with the target column names, we find 151 genes which encode a target protein (see the table below). It doesn't matter that some proteins are encoded by more than one gene (e.g., rows 146 and 147 of the table). We may assume that these 151 features will have a high feature importance in our models.</p>\n<pre><code>matching_names = []\nfor protein in cite_protein_names:\n    matching_names += [(gene, protein) for gene in cite_gene_names if protein in gene]\npd.DataFrame(matching_names, columns=['Gene', 'Protein'])\n</code></pre>\n<table>\n  <thead>\n    <tr>\n      <th></th>\n      <th>Gene</th>\n      <th>Protein</th>\n    </tr>\n  </thead>\n  <tbody>\n    <tr>\n      <th>0</th>\n      <td>ENSG00000114013_CD86</td>\n      <td>CD86</td>\n    </tr>\n    <tr>\n      <th>1</th>\n      <td>ENSG00000120217_CD274</td>\n      <td>CD274</td>\n    </tr>\n    <tr>\n      <th>2</th>\n      <td>ENSG00000196776_CD47</td>\n      <td>CD47</td>\n    </tr>\n    <tr>\n      <th>3</th>\n      <td>ENSG00000117091_CD48</td>\n      <td>CD48</td>\n    </tr>\n    <tr>\n      <th>4</th>\n      <td>ENSG00000101017_CD40</td>\n      <td>CD40</td>\n    </tr>\n    <tr>\n      <th>...</th>\n      <td>...</td>\n      <td>...</td>\n    </tr>\n    <tr>\n      <th>146</th>\n      <td>ENSG00000102181_CD99L2</td>\n      <td>CD9</td>\n    </tr>\n    <tr>\n      <th>147</th>\n      <td>ENSG00000223773_CD99P1</td>\n      <td>CD9</td>\n    </tr>\n    <tr>\n      <th>148</th>\n      <td>ENSG00000204592_HLA-E</td>\n      <td>HLA-E</td>\n    </tr>\n    <tr>\n      <th>149</th>\n      <td>ENSG00000085117_CD82</td>\n      <td>CD82</td>\n    </tr>\n    <tr>\n      <th>150</th>\n      <td>ENSG00000134256_CD101</td>\n      <td>CD101</td>\n    </tr>\n  </tbody>\n</table>\n<p><strong>Insight:</strong> If we apply dimensionality reduction (PCA, SVD, whatever) to the 22050 features, we should make sure that we don't reduce away the 151 features which encode the proteins we want to predict.</p>\n<p>For an implementation of this idea, see the <a href=\"https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart\" target=\"_blank\">MSCI CITEseq Quickstart</a> notebook.</p>",
      "rawMarkdown": "In this competition, the features and targets are not anonymous: Genes have names, and proteins have names. \n\nThe CITEseq task has genes as input and proteins as output. Genes encode proteins, and it is more or less known which genes encode which proteins. The column names provide this information: The input dataframe has the genes as column names, and the target dataframe has the proteins as column names. According to the naming convention, the gene names contain the protein name as suffix after a '_'.\n\nIf we match the input column names with the target column names, we find 151 genes which encode a target protein (see the table below). It doesn't matter that some proteins are encoded by more than one gene (e.g., rows 146 and 147 of the table). We may assume that these 151 features will have a high feature importance in our models.\n\n```\nmatching_names = []\nfor protein in cite_protein_names:\n    matching_names += [(gene, protein) for gene in cite_gene_names if protein in gene]\npd.DataFrame(matching_names, columns=['Gene', 'Protein'])\n``` \n\n<table class=\"dataframe\" border=\"1\">\n  <thead>\n    <tr style=\"text-align: right;\">\n      <th></th>\n      <th>Gene</th>\n      <th>Protein</th>\n    </tr>\n  </thead>\n  <tbody>\n    <tr>\n      <th>0</th>\n      <td>ENSG00000114013_CD86</td>\n      <td>CD86</td>\n    </tr>\n    <tr>\n      <th>1</th>\n      <td>ENSG00000120217_CD274</td>\n      <td>CD274</td>\n    </tr>\n    <tr>\n      <th>2</th>\n      <td>ENSG00000196776_CD47</td>\n      <td>CD47</td>\n    </tr>\n    <tr>\n      <th>3</th>\n      <td>ENSG00000117091_CD48</td>\n      <td>CD48</td>\n    </tr>\n    <tr>\n      <th>4</th>\n      <td>ENSG00000101017_CD40</td>\n      <td>CD40</td>\n    </tr>\n    <tr>\n      <th>...</th>\n      <td>...</td>\n      <td>...</td>\n    </tr>\n    <tr>\n      <th>146</th>\n      <td>ENSG00000102181_CD99L2</td>\n      <td>CD9</td>\n    </tr>\n    <tr>\n      <th>147</th>\n      <td>ENSG00000223773_CD99P1</td>\n      <td>CD9</td>\n    </tr>\n    <tr>\n      <th>148</th>\n      <td>ENSG00000204592_HLA-E</td>\n      <td>HLA-E</td>\n    </tr>\n    <tr>\n      <th>149</th>\n      <td>ENSG00000085117_CD82</td>\n      <td>CD82</td>\n    </tr>\n    <tr>\n      <th>150</th>\n      <td>ENSG00000134256_CD101</td>\n      <td>CD101</td>\n    </tr>\n  </tbody>\n</table>\n\n**Insight:** If we apply dimensionality reduction (PCA, SVD, whatever) to the 22050 features, we should make sure that we don't reduce away the 151 features which encode the proteins we want to predict.\n\nFor an implementation of this idea, see the [MSCI CITEseq Quickstart](https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart) notebook.",
      "votes": null
    },
    {
      "id": "1922958",
      "postDate": "09/01/2022 21:27:10",
      "content": "<p>Very Interesting, maybe training single models with the corresponding 151 genes for each target protein will provide a better solution compared to multi_inputs -&gt; PCA -&gt; multi_outputs</p>\n<p>Does something similar happens in Multiome? Did you check?</p>",
      "rawMarkdown": "Very Interesting, maybe training single models with the corresponding 151 genes for each target protein will provide a better solution compared to multi_inputs -> PCA -> multi_outputs\n\nDoes something similar happens in Multiome? Did you check?",
      "votes": null
    },
    {
      "id": "1923176",
      "postDate": "09/02/2022 03:46:49",
      "content": "<p>It will but it would be more difficult. I think in the current multiome data we have only represent peaks, but we need to transfer peaks to tf (transcript factors) which control the expression of genes and then perform the prediction process. People who are interested in this step can have a try.</p>",
      "rawMarkdown": "It will but it would be more difficult. I think in the current multiome data we have only represent peaks, but we need to transfer peaks to tf (transcript factors) which control the expression of genes and then perform the prediction process. People who are interested in this step can have a try.",
      "votes": null
    },
    {
      "id": "1923329",
      "postDate": "09/02/2022 06:20:12",
      "content": "<p>Nice !</p>\n<p>The first question biologists ask:</p>\n<p>What are the correlations between gene transcript and corresponding(!)  gene protein level ?</p>\n<p>The other interesting thing  - analyze feature importance here. I.e what are the other gene transcipt affect protein levels ? Then we can discuss with biologists what can be the bio reason</p>",
      "rawMarkdown": "Nice !\n\nThe first question biologists ask:\n\nWhat are the correlations between gene transcript and corresponding(!)  gene protein level ?\n\nThe other interesting thing  - analyze feature importance here. I.e what are the other gene transcipt affect protein levels ? Then we can discuss with biologists what can be the bio reason",
      "votes": null
    },
    {
      "id": "1935194",
      "postDate": "09/11/2022 23:32:06",
      "content": "<p>Actually, I have computed the pearson correlation coefficients between input and output, and to my own surprise, the inputs and targets that match by name are not necessarily very well correlated (sometime they are, as for CD86, but sometime not, as for CD274). See <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/351725\" target=\"_blank\">this discussion</a>. This seems unintuitive, but that is what the data seems to say (modulo a stupid bug in my code). It could also be that the dependence exist but is not very linear (and thus not detected by pearson coefficients).</p>",
      "rawMarkdown": "Actually, I have computed the pearson correlation coefficients between input and output, and to my own surprise, the inputs and targets that match by name are not necessarily very well correlated (sometime they are, as for CD86, but sometime not, as for CD274). See [this discussion](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/351725). This seems unintuitive, but that is what the data seems to say (modulo a stupid bug in my code). It could also be that the dependence exist but is not very linear (and thus not detected by pearson coefficients).",
      "votes": null
    },
    {
      "id": "1935480",
      "postDate": "09/12/2022 06:35:52",
      "content": "<p><a href=\"https://www.kaggle.com/fabiencrom\" target=\"_blank\">@fabiencrom</a> <br>\nProbably I was not clear at some point at the previous posts.<br>\nThe fact that  RNA and corresponding Protein may be poorly correlated is quite in line with what is known in biology.<br>\nThere is so-called \"posttranscriptional\" regulation and other ideas - in general it is not well understood in modern biology<br>\nwhy that happens, but it was measured many times that something like that may happen.<br>\nSo you should not be very surprised that it might happen, but the interesting question is to try to dig a little further and later probably we can look on biological side - what my explain that. <br>\nSome questions are here: <br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900</a></p>",
      "rawMarkdown": "fabiencrom \nProbably I was not clear at some point at the previous posts.\nThe fact that  RNA and corresponding Protein may be poorly correlated is quite in line with what is known in biology.\nThere is so-called \"posttranscriptional\" regulation and other ideas - in general it is not well understood in modern biology\nwhy that happens, but it was measured many times that something like that may happen.\nSo you should not be very surprised that it might happen, but the interesting question is to try to dig a little further and later probably we can look on biological side - what my explain that. \nSome questions are here: \nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900",
      "votes": null
    },
    {
      "id": "1944580",
      "postDate": "09/18/2022 12:04:16",
      "content": "<p>Very nice. <br>\nBut some comments,<br>\nit seems the correspondence is not precisely that.<br>\ne.g. CD9 - <a href=\"https://en.wikipedia.org/wiki/CD9\" target=\"_blank\">https://en.wikipedia.org/wiki/CD9</a><br>\nWikipedia provides Ensembl Id:    <br>\nENSG00000010278<br>\nnot the ones mentioned in 146,7 positions. </p>",
      "rawMarkdown": "Very nice. \nBut some comments,\nit seems the correspondence is not precisely that.\ne.g. CD9 - https://en.wikipedia.org/wiki/CD9\nWikipedia provides Ensembl Id:\t\nENSG00000010278\nnot the ones mentioned in 146,7 positions.",
      "votes": null
    },
    {
      "id": "1944914",
      "postDate": "09/18/2022 17:04:00",
      "content": "<p>It seems that my matching rule was too simple. Of course, ENSG00000102181_CD99L2 should match with CD99L2 and not with CD9!</p>",
      "rawMarkdown": "It seems that my matching rule was too simple. Of course, ENSG00000102181_CD99L2 should match with CD99L2 and not with CD9!",
      "votes": null
    },
    {
      "id": "1945011",
      "postDate": "09/18/2022 18:34:26",
      "content": "<p>Thanks ! Is it the only potential problem ? (Sorry have not looked yet your  matching) </p>",
      "rawMarkdown": "Thanks ! Is it the only potential problem ? (Sorry have not looked yet your  matching)",
      "votes": null
    },
    {
      "id": "1959272",
      "postDate": "09/28/2022 04:34:10",
      "content": "<p>For correspondence CD -&gt; genes see file provided by orgs and discussion:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/354713\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/354713</a></p>",
      "rawMarkdown": "For correspondence CD -> genes see file provided by orgs and discussion:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/354713",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1922958,
      "author_name": "ragnar123",
      "author_url": "",
      "post_date": "09/01/2022 21:27:10",
      "content": "<p>Very Interesting, maybe training single models with the corresponding 151 genes for each target protein will provide a better solution compared to multi_inputs -&gt; PCA -&gt; multi_outputs</p>\n<p>Does something similar happens in Multiome? Did you check?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1923176,
          "author_name": "llttyy",
          "author_url": "",
          "post_date": "09/02/2022 03:46:49",
          "content": "<p>It will but it would be more difficult. I think in the current multiome data we have only represent peaks, but we need to transfer peaks to tf (transcript factors) which control the expression of genes and then perform the prediction process. People who are interested in this step can have a try.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1923329,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/02/2022 06:20:12",
      "content": "<p>Nice !</p>\n<p>The first question biologists ask:</p>\n<p>What are the correlations between gene transcript and corresponding(!)  gene protein level ?</p>\n<p>The other interesting thing  - analyze feature importance here. I.e what are the other gene transcipt affect protein levels ? Then we can discuss with biologists what can be the bio reason</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1935194,
      "author_name": "fabiencrom",
      "author_url": "",
      "post_date": "09/11/2022 23:32:06",
      "content": "<p>Actually, I have computed the pearson correlation coefficients between input and output, and to my own surprise, the inputs and targets that match by name are not necessarily very well correlated (sometime they are, as for CD86, but sometime not, as for CD274). See <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/351725\" target=\"_blank\">this discussion</a>. This seems unintuitive, but that is what the data seems to say (modulo a stupid bug in my code). It could also be that the dependence exist but is not very linear (and thus not detected by pearson coefficients).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1935480,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "09/12/2022 06:35:52",
          "content": "<p><a href=\"https://www.kaggle.com/fabiencrom\" target=\"_blank\">@fabiencrom</a> <br>\nProbably I was not clear at some point at the previous posts.<br>\nThe fact that  RNA and corresponding Protein may be poorly correlated is quite in line with what is known in biology.<br>\nThere is so-called \"posttranscriptional\" regulation and other ideas - in general it is not well understood in modern biology<br>\nwhy that happens, but it was measured many times that something like that may happen.<br>\nSo you should not be very surprised that it might happen, but the interesting question is to try to dig a little further and later probably we can look on biological side - what my explain that. <br>\nSome questions are here: <br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1944580,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/18/2022 12:04:16",
      "content": "<p>Very nice. <br>\nBut some comments,<br>\nit seems the correspondence is not precisely that.<br>\ne.g. CD9 - <a href=\"https://en.wikipedia.org/wiki/CD9\" target=\"_blank\">https://en.wikipedia.org/wiki/CD9</a><br>\nWikipedia provides Ensembl Id:    <br>\nENSG00000010278<br>\nnot the ones mentioned in 146,7 positions. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1944914,
          "author_name": "ambrosm",
          "author_url": "",
          "post_date": "09/18/2022 17:04:00",
          "content": "<p>It seems that my matching rule was too simple. Of course, ENSG00000102181_CD99L2 should match with CD99L2 and not with CD9!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1945011,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "09/18/2022 18:34:26",
          "content": "<p>Thanks ! Is it the only potential problem ? (Sorry have not looked yet your  matching) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1959272,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/28/2022 04:34:10",
      "content": "<p>For correspondence CD -&gt; genes see file provided by orgs and discussion:<br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/354713\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/354713</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1921317": "In this competition, the features and targets are not anonymous: Genes have names, and proteins have names. \n\nThe CITEseq task has genes as input and proteins as output. Genes encode proteins, and it is more or less known which genes encode which proteins. The column names provide this information: The input dataframe has the genes as column names, and the target dataframe has the proteins as column names. According to the naming convention, the gene names contain the protein name as suffix after a '_'.\n\nIf we match the input column names with the target column names, we find 151 genes which encode a target protein (see the table below). It doesn't matter that some proteins are encoded by more than one gene (e.g., rows 146 and 147 of the table). We may assume that these 151 features will have a high feature importance in our models.\n\n```\nmatching_names = []\nfor protein in cite_protein_names:\n    matching_names += [(gene, protein) for gene in cite_gene_names if protein in gene]\npd.DataFrame(matching_names, columns=['Gene', 'Protein'])\n``` \n\n<table class=\"dataframe\" border=\"1\">\n  <thead>\n    <tr style=\"text-align: right;\">\n      <th></th>\n      <th>Gene</th>\n      <th>Protein</th>\n    </tr>\n  </thead>\n  <tbody>\n    <tr>\n      <th>0</th>\n      <td>ENSG00000114013_CD86</td>\n      <td>CD86</td>\n    </tr>\n    <tr>\n      <th>1</th>\n      <td>ENSG00000120217_CD274</td>\n      <td>CD274</td>\n    </tr>\n    <tr>\n      <th>2</th>\n      <td>ENSG00000196776_CD47</td>\n      <td>CD47</td>\n    </tr>\n    <tr>\n      <th>3</th>\n      <td>ENSG00000117091_CD48</td>\n      <td>CD48</td>\n    </tr>\n    <tr>\n      <th>4</th>\n      <td>ENSG00000101017_CD40</td>\n      <td>CD40</td>\n    </tr>\n    <tr>\n      <th>...</th>\n      <td>...</td>\n      <td>...</td>\n    </tr>\n    <tr>\n      <th>146</th>\n      <td>ENSG00000102181_CD99L2</td>\n      <td>CD9</td>\n    </tr>\n    <tr>\n      <th>147</th>\n      <td>ENSG00000223773_CD99P1</td>\n      <td>CD9</td>\n    </tr>\n    <tr>\n      <th>148</th>\n      <td>ENSG00000204592_HLA-E</td>\n      <td>HLA-E</td>\n    </tr>\n    <tr>\n      <th>149</th>\n      <td>ENSG00000085117_CD82</td>\n      <td>CD82</td>\n    </tr>\n    <tr>\n      <th>150</th>\n      <td>ENSG00000134256_CD101</td>\n      <td>CD101</td>\n    </tr>\n  </tbody>\n</table>\n\n**Insight:** If we apply dimensionality reduction (PCA, SVD, whatever) to the 22050 features, we should make sure that we don't reduce away the 151 features which encode the proteins we want to predict.\n\nFor an implementation of this idea, see the [MSCI CITEseq Quickstart](https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart) notebook.",
    "1922958": "Very Interesting, maybe training single models with the corresponding 151 genes for each target protein will provide a better solution compared to multi_inputs -> PCA -> multi_outputs\n\nDoes something similar happens in Multiome? Did you check?",
    "1923176": "It will but it would be more difficult. I think in the current multiome data we have only represent peaks, but we need to transfer peaks to tf (transcript factors) which control the expression of genes and then perform the prediction process. People who are interested in this step can have a try.",
    "1923329": "Nice !\n\nThe first question biologists ask:\n\nWhat are the correlations between gene transcript and corresponding(!)  gene protein level ?\n\nThe other interesting thing  - analyze feature importance here. I.e what are the other gene transcipt affect protein levels ? Then we can discuss with biologists what can be the bio reason",
    "1935194": "Actually, I have computed the pearson correlation coefficients between input and output, and to my own surprise, the inputs and targets that match by name are not necessarily very well correlated (sometime they are, as for CD86, but sometime not, as for CD274). See [this discussion](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/351725). This seems unintuitive, but that is what the data seems to say (modulo a stupid bug in my code). It could also be that the dependence exist but is not very linear (and thus not detected by pearson coefficients).",
    "1935480": "fabiencrom \nProbably I was not clear at some point at the previous posts.\nThe fact that  RNA and corresponding Protein may be poorly correlated is quite in line with what is known in biology.\nThere is so-called \"posttranscriptional\" regulation and other ideas - in general it is not well understood in modern biology\nwhy that happens, but it was measured many times that something like that may happen.\nSo you should not be very surprised that it might happen, but the interesting question is to try to dig a little further and later probably we can look on biological side - what my explain that. \nSome questions are here: \nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/350900",
    "1944580": "Very nice. \nBut some comments,\nit seems the correspondence is not precisely that.\ne.g. CD9 - https://en.wikipedia.org/wiki/CD9\nWikipedia provides Ensembl Id:\t\nENSG00000010278\nnot the ones mentioned in 146,7 positions.",
    "1944914": "It seems that my matching rule was too simple. Of course, ENSG00000102181_CD99L2 should match with CD99L2 and not with CD9!",
    "1945011": "Thanks ! Is it the only potential problem ? (Sorry have not looked yet your  matching)",
    "1959272": "For correspondence CD -> genes see file provided by orgs and discussion:\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/354713"
  },
  "source": "meta"
}