{
  "id": 346761,
  "title": "any more gene name/ensemble-id dependencies?",
  "url": "/competitions/open-problems-multimodal/discussion/346761",
  "author_name": "",
  "post_date": "2022-08-21T07:52:12.204630Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>The data description states that columns correspond to genes given by <code>{gene_name}_{gene_ensemble-ids}</code>, but there are 22050 columns in the Cite inputs which is the same number as a number of unique <code>gene_name</code>. So is there an expectation that we can create something like a sparse matrix of <code>gene_name</code> x <code>gene_ensemble-ids</code>.</p>\n<p>Similar question to multi inputs, if the column name has any meaning or t could be decomposed to some elements? but so far if I slit it by <code>:</code> I have 228942 columns and get first element 37 and second element 228941.</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/jirkaborovec/mmscel-inst-eda-stat-predictions/notebook\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/mmscel-inst-eda-stat-predictions/notebook</a></p>\n</blockquote>",
  "messages": [
    {
      "id": "1907936",
      "postDate": "08/21/2022 07:52:12",
      "content": "<p>The data description states that columns correspond to genes given by <code>{gene_name}_{gene_ensemble-ids}</code>, but there are 22050 columns in the Cite inputs which is the same number as a number of unique <code>gene_name</code>. So is there an expectation that we can create something like a sparse matrix of <code>gene_name</code> x <code>gene_ensemble-ids</code>.</p>\n<p>Similar question to multi inputs, if the column name has any meaning or t could be decomposed to some elements? but so far if I slit it by <code>:</code> I have 228942 columns and get first element 37 and second element 228941.</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/jirkaborovec/mmscel-inst-eda-stat-predictions/notebook\" target=\"_blank\">https://www.kaggle.com/code/jirkaborovec/mmscel-inst-eda-stat-predictions/notebook</a></p>\n</blockquote>",
      "rawMarkdown": "The data description states that columns correspond to genes given by `{gene_name}_{gene_ensemble-ids}`, but there are 22050 columns in the Cite inputs which is the same number as a number of unique `gene_name`. So is there an expectation that we can create something like a sparse matrix of `gene_name` x `gene_ensemble-ids`.\n\nSimilar question to multi inputs, if the column name has any meaning or t could be decomposed to some elements? but so far if I slit it by `:` I have 228942 columns and get first element 37 and second element 228941.\n\n>https://www.kaggle.com/code/jirkaborovec/mmscel-inst-eda-stat-predictions/notebook",
      "votes": null
    },
    {
      "id": "1909264",
      "postDate": "08/22/2022 13:23:33",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a>! Re CITEseq inputs: <code>gene_name</code> and <code>gene_ensemble-ids</code> are simply two names for the same thing: a gene. So I am not sure why you should create a <code>gene_name</code>x<code>gene_ensemble-ids</code> matrix. The ATACseq columns correspond to gene fragments where the <code>:</code> separates the chromosome specifier with the genome location. Does that answer your question?</p>",
      "rawMarkdown": "Hi @jirkaborovec! Re CITEseq inputs: `gene_name` and `gene_ensemble-ids` are simply two names for the same thing: a gene. So I am not sure why you should create a `gene_name`x`gene_ensemble-ids` matrix. The ATACseq columns correspond to gene fragments where the `:` separates the chromosome specifier with the genome location. Does that answer your question?",
      "votes": null
    },
    {
      "id": "1909277",
      "postDate": "08/22/2022 13:37:48",
      "content": "<p>Just to provide some more detail here, gene names (aka gene symbols) are human readable common names that may change over time. Ensemble gene IDs are stable and do not change from release to release of genome annotations. More info about gene naming and gene IDs here:</p>\n<ul>\n<li><a href=\"https://useast.ensembl.org/info/genome/genebuild/gene_names.html\" target=\"_blank\">Gene naming</a></li>\n<li><a href=\"https://useast.ensembl.org/Help/Faq?id=488\" target=\"_blank\">FAQ - I have an Ensembl ID, what can I tell about it from the ID?</a></li>\n</ul>\n<p>In the multi input, the feature names are <a href=\"https://www.idtdna.com/pages/support/faqs/how-are-genomic-coordinates-defined\" target=\"_blank\">genomic coordinates</a> on reference genome GRCh38 provided in the 10x References - 2020-A (July 7, 2020).</p>",
      "rawMarkdown": "Just to provide some more detail here, gene names (aka gene symbols) are human readable common names that may change over time. Ensemble gene IDs are stable and do not change from release to release of genome annotations. More info about gene naming and gene IDs here:\n* [Gene naming](https://useast.ensembl.org/info/genome/genebuild/gene_names.html)\n* [FAQ - I have an Ensembl ID, what can I tell about it from the ID?](https://useast.ensembl.org/Help/Faq?id=488)\n\nIn the multi input, the feature names are [genomic coordinates](https://www.idtdna.com/pages/support/faqs/how-are-genomic-coordinates-defined) on reference genome GRCh38 provided in the 10x References - 2020-A (July 7, 2020).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1909264,
      "author_name": "peterholderrieth",
      "author_url": "",
      "post_date": "08/22/2022 13:23:33",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jirkaborovec\" target=\"_blank\">@jirkaborovec</a>! Re CITEseq inputs: <code>gene_name</code> and <code>gene_ensemble-ids</code> are simply two names for the same thing: a gene. So I am not sure why you should create a <code>gene_name</code>x<code>gene_ensemble-ids</code> matrix. The ATACseq columns correspond to gene fragments where the <code>:</code> separates the chromosome specifier with the genome location. Does that answer your question?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1909277,
          "author_name": "danielburkhardt",
          "author_url": "",
          "post_date": "08/22/2022 13:37:48",
          "content": "<p>Just to provide some more detail here, gene names (aka gene symbols) are human readable common names that may change over time. Ensemble gene IDs are stable and do not change from release to release of genome annotations. More info about gene naming and gene IDs here:</p>\n<ul>\n<li><a href=\"https://useast.ensembl.org/info/genome/genebuild/gene_names.html\" target=\"_blank\">Gene naming</a></li>\n<li><a href=\"https://useast.ensembl.org/Help/Faq?id=488\" target=\"_blank\">FAQ - I have an Ensembl ID, what can I tell about it from the ID?</a></li>\n</ul>\n<p>In the multi input, the feature names are <a href=\"https://www.idtdna.com/pages/support/faqs/how-are-genomic-coordinates-defined\" target=\"_blank\">genomic coordinates</a> on reference genome GRCh38 provided in the 10x References - 2020-A (July 7, 2020).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1907936": "The data description states that columns correspond to genes given by `{gene_name}_{gene_ensemble-ids}`, but there are 22050 columns in the Cite inputs which is the same number as a number of unique `gene_name`. So is there an expectation that we can create something like a sparse matrix of `gene_name` x `gene_ensemble-ids`.\n\nSimilar question to multi inputs, if the column name has any meaning or t could be decomposed to some elements? but so far if I slit it by `:` I have 228942 columns and get first element 37 and second element 228941.\n\n>https://www.kaggle.com/code/jirkaborovec/mmscel-inst-eda-stat-predictions/notebook",
    "1909264": "Hi @jirkaborovec! Re CITEseq inputs: `gene_name` and `gene_ensemble-ids` are simply two names for the same thing: a gene. So I am not sure why you should create a `gene_name`x`gene_ensemble-ids` matrix. The ATACseq columns correspond to gene fragments where the `:` separates the chromosome specifier with the genome location. Does that answer your question?",
    "1909277": "Just to provide some more detail here, gene names (aka gene symbols) are human readable common names that may change over time. Ensemble gene IDs are stable and do not change from release to release of genome annotations. More info about gene naming and gene IDs here:\n* [Gene naming](https://useast.ensembl.org/info/genome/genebuild/gene_names.html)\n* [FAQ - I have an Ensembl ID, what can I tell about it from the ID?](https://useast.ensembl.org/Help/Faq?id=488)\n\nIn the multi input, the feature names are [genomic coordinates](https://www.idtdna.com/pages/support/faqs/how-are-genomic-coordinates-defined) on reference genome GRCh38 provided in the 10x References - 2020-A (July 7, 2020)."
  },
  "source": "meta"
}